On July 30, 2026, the CEO of Amazon stands on an earnings call and says something that would have sounded absurd five years ago. Andy Jassy tells analysts that Amazon will spend roughly $220 billion this year — more than the GDP of most countries — building data centers, and that it still will not have enough capacity to meet all the demand AWS has in 2026.
Then he shares a detail that stopped me cold. Two large customers asked to buy all of AWS's Graviton computing capacity — every server running Amazon's own processor chips — for the entire year. All of it. Amazon said no, because saying yes would have left nothing for everyone else.
Two days earlier, Microsoft's CFO told her own analysts, for the fourth consecutive quarter, that customer demand for Azure exceeds what Microsoft can supply — and will keep exceeding it "for a number of quarters." And the day before that, the operator of America's largest power grid proposed a plan to deliberately switch off large data centers when the grid runs short.
For fifteen years, the cloud sold us one promise above all others: as much computing as you want, whenever you want it. That promise quietly expired this year.
So what do you do when the most reliable assumption in your IT strategy — that capacity is always there when you ask for it — stops being true?
What actually happened this summer#
Three announcements landed within four days of each other in late July. Each one alone is a footnote. Together they describe a new reality.
The Amazon admission — July 30
Start with the numbers, because they're the credibility. AWS grew 36.7% year over year in the second quarter — its fastest growth in eighteen quarters. Amazon raised its capital spending plan to about $220 billion for 2026. And the company is on pace to double its power capacity by the end of 2027 compared to 2025.
Here's the part that matters for you: even at that pace, Jassy said demand outruns supply. Much of Amazon's 2027 capacity — and part of its 2028 capacity — is already spoken for. Reserved. Sold before it's built.
Think of it like a restaurant that keeps adding tables and still has a line out the door — except the line is now booking tables two years in advance, and two of the customers in that line offered to rent the entire dining room for a year. When a company whose whole business is selling computing says "we cannot build it fast enough," that is not a marketing problem. That is a supply problem.
The Graviton detail deserves one more beat, because of what it says about where the market is. Graviton isn't exotic hardware for AI research labs — it's Amazon's general-purpose workhorse chip, the kind of server that runs ordinary websites, databases, and business applications. When buyers try to corner the market on ordinary computing, the scarcity has moved past the specialty aisle and into the staples. That's the moment the story stops being about someone else's AI project and starts being about your Tuesday.
The Microsoft repeat — late July
Microsoft's story is nearly identical, and it's the repetition that should get your attention. Back in the fall of 2025, Microsoft told investors Azure would be capacity-constrained "through at least June 2026." June 2026 came and went. On the fiscal fourth-quarter call in late July, CFO Amy Hood said demand still exceeds available supply — and projected it would stay that way for a number of quarters more.
To Microsoft's credit, they've gotten faster at bringing capacity online, and they say efficiency gains across their server fleet are helping close the gap. But a deadline that slips by a year isn't a deadline. It's a condition. Azure is guiding to roughly 45% cloud growth while rationing supply — which tells you the constraint isn't hurting their business. It's hurting yours, if you assumed the tap was infinite.
The grid says the quiet part out loud — July 27–28
The third announcement came not from a cloud provider but from a power company — and it's the one I'd wager most IT leaders missed.
PJM Interconnection runs the largest electrical grid in the United States, serving about 67 million people across 13 states. During the heat wave this summer, PJM was operating within about 2,000 megawatts of its all-time demand record — the grid equivalent of redlining the engine. On July 27, PJM's board proposed an emergency backstop plan, and the detail that matters is this: data centers drawing 50 megawatts or more may face temporary, compensated power cutoffs — starting June 2027 — to prevent rolling blackouts for everyone else. The U.S. Department of Energy had already granted emergency authority to curtail large power users during the hottest days.
An analogy helps here. The power grid is a highway system, and data centers went — in about three years — from being a few delivery vans to being convoys of freight trucks that never pull over. The highway isn't wide enough, and widening it is slow: the giant transformers that grid expansion depends on now take two to three years to arrive after you order one — up to four for the largest units. A 2025 industry survey (reported by Schneider Electric) found 92% of data-center operators say the power grid, not chips, is their top constraint.
That's the punchline of 2026: the bottleneck moved from silicon to electricity. You can't overnight-ship a substation.
And the money is chasing the problem at a scale that's hard to picture. The five largest data-center builders are on track to invest something on the order of $700 billion in U.S. facilities in 2026 alone, and U.S. data-center capacity is projected to roughly quadruple by 2030. But money doesn't conjure transformers, and it doesn't speed up interconnection queues. A dollar spent today on a large grid component buys delivery in 2028 or 2029. That lag — between when demand shows up and when infrastructure can physically answer it — is the window your company is living in right now, whether you noticed or not.
The pattern underneath
Put the three stories together and one sentence falls out: cloud capacity is now allocated, not assumed. The providers are growing enormously fast and still choosing who gets what. That choice used to be invisible because supply outran demand. Now it's a queue — and the question is where you stand in it.
Quota was never capacity — the fine print that just started mattering#
Here's the structural problem, and it's hiding in a distinction most companies have never had a reason to learn.
When your team sets up a cloud account, you get a quota — a numerical limit on how many servers of a given type you're allowed to request. It feels like a guarantee. It is not one. Microsoft's own documentation now says it plainly: a quota "isn't a capacity reservation or guarantee." A quota is a fishing license. It permits you to fish. It does not promise there are fish in the lake.
For fifteen years the lake was so overstocked the distinction never mattered. In 2026, it matters. Real customers with approved quotas are hitting allocation failures — the deployment simply fails because the region is out of physical machines. Microsoft has limited new sign-ups in some regions, including Northern Virginia and parts of Texas — the beating heart of American cloud infrastructure — while letting existing workloads keep growing. Early this year it temporarily stopped new deployments of certain specialized server types in its UK South region. And as of July 2026, a formal policy took effect restricting new customer accounts from deploying older-generation server series at all — the cheaper workhorses many businesses standardized on.
Notice who gets squeezed in each case: the new deployment, the new subscription, the new region. Existing workloads are protected. Growth is what gets rationed.
To be fair to the providers, the squeeze is uneven, and honesty requires saying so. Most workloads in most regions still deploy without a hiccup. On the specialized-hardware side there are even gluts: rental prices for one previous-generation AI chip have reportedly fallen by three quarters as supply caught up. "The cloud is sold out" is shorthand — the precise version is selectively sold out, in the places and hardware generations everyone wants, with the shortages moving unpredictably. But that precise version is almost worse for planning purposes. A uniform shortage you can plan around. A moving one finds whoever didn't check.
And before you reach for the standard answer — "we'll just go multi-cloud" — sit with this: Amazon, Microsoft, and Google are all telling some version of the same story, they are all buying chips from the same handful of suppliers, and they are all plugging into the same strained grid, often in the same few counties of Virginia. When every backup option shares the bottleneck, a second vendor isn't redundancy. It's a second place in the same line.
A quota is permission to ask. Capacity is the answer. For fifteen years the answer was always yes — that's the assumption that just expired.
How This Impacts Your Organization#
The principle here doesn't change with company size: everyone's cloud strategy quietly assumed supply was infinite, and it isn't. What changes with size is which version of the squeeze reaches you first, what options you actually have, and where your leverage lives.
Large Enterprises (1,000+ employees)
If you're at enterprise scale, you already live partway in this world — you have reserved instances, committed-spend agreements, and probably a named account team at your provider. Your capability isn't the problem. Your assumptions are.
The real risk at your size is organizational: capacity commitments are scattered across dozens of teams and business units, each with its own accounts, quotas, and region choices — and nobody holds the aggregate picture. One division reserves capacity it doesn't use while another silently depends on on-demand burst that may not be there during a crunch. Meanwhile, your disaster-recovery plan almost certainly assumes you can spin up replacement infrastructure in an alternate region on demand. In a rationed market, that assumption needs to be tested, not believed. And if you operate your own large facilities in PJM territory at 50 megawatts or more, you are now on a named curtailment list — backup generation and demand-response contracts just became board-agenda items, not facilities details.
Your leverage is real: scale, multi-year commitments, and the ability to get capacity assurances in writing as negotiated contract terms — something smaller companies can't demand. The organizational move: name one owner for capacity posture across the enterprise, the way you long ago named owners for security and spend. Give them the aggregate reservation map, the DR capacity test results, and a seat in the next provider negotiation.
Mid-Size Organizations (100–999 employees)
Mid-size companies feel this fastest, and I want to be direct about why: you're big enough to need real capacity — meaningful compute for real workloads, genuine disaster-recovery obligations, customers with uptime expectations — but not big enough to have a hyperscaler fighting to keep you. When two whales offer to buy an entire instance family for a year, you are not in that conversation. You compete for what's left.
The overcorrection to avoid is the same error run in reverse: panic-buying three years of reserved capacity you may not need, or bolting on a second cloud provider "for safety" and doubling your operational surface without doubling your resilience. Both moves burn your scarcest resources — money and attention — to purchase a feeling rather than an outcome.
Your win is focus and timing. Three concrete moves, none of which require a platform team: First, convert your two or three genuinely critical workloads from on-demand assumptions to actual capacity reservations — reservations are no longer just a discount mechanism; they're an availability mechanism. Second, run one honest disaster-recovery test that attempts a real restore in your failover region and watches for allocation failures — quota math on a spreadsheet is not a test. Third, pick your second-choice region now, deliberately, weighing latency and data residency — because choosing it during an incident, from whatever happens to have capacity, is how bad architecture gets born.
Small & Growing Organizations (Under 100 employees)
If you're under a hundred people, here's the honest counsel: do not let this post scare you into an infrastructure project. Your workloads are small enough to fit in the cracks of a constrained market, and most of this squeeze will never reach you as drama. It will reach you as friction — and friction is manageable if you see it coming.
The friction looks like this: the older, cheaper server series you standardized on is now restricted for new accounts. The region closest to your customers isn't accepting new sign-ups the month you happen to need it. A price increase arrives because scarcity always rolls downhill into pricing eventually. None of these is a crisis. Each one, met unprepared, costs you a week you didn't have — a migration you didn't plan, a launch delayed while you re-platform onto hardware you can actually get.
Your discipline is lightweight, and most of it is free. Know which region and which one or two instance types your business actually runs on — write it down; it fits on an index card. Know your second-choice region before you need it. If your product genuinely cannot tolerate downtime, one small reservation for your core workload is cheap insurance. And favor boring, current-generation, widely available instance types over exotic or deprecated ones — in a rationed market, popular hardware gets restocked first.
What you should not do at your size: multi-cloud architectures, capacity hedging, or committing years of spend to save single-digit percentages. That's someone else's medicine. Reassurance where reassurance is honest: small really is flexible here — you can move regions and resize in an afternoon what takes an enterprise a quarter.
What to do Monday morning#
- Put one question on this week's IT agenda: "Where do we assume capacity will just be there?" — This is a conversation, not a project, and it costs nothing. Walk through your scaling plans, your disaster-recovery runbook, and your growth forecast, and circle every step that silently assumes on-demand infrastructure will be available. You can't fix an assumption you haven't found. Most teams find three or four within the hour.
- Write down your actual footprint: regions, instance families, quotas. — Free, and shockingly rare. One page: which regions you run in, which server types you depend on, what your approved quotas are, and — the column nobody has — which of those quotas are backed by reservations versus pure on-demand hope. This page is the input to every other decision on this list.
- Test one real restore in your failover region. — Not a tabletop exercise. Actually attempt to deploy your recovery environment where your plan says it will run, and watch what the provider does. If you get an allocation failure on a Tuesday afternoon in ordinary conditions, you have learned something priceless about what happens during a regional event when everyone else is failing over too.
- Convert your most critical workload from on-demand to reserved capacity. — Reservations used to be a finance conversation about discounts. In a constrained market they're an availability conversation. Start with the single workload whose absence would stop your business, price the reservation, and treat the cost as insurance, not overhead.
- Choose your second-choice region deliberately, this month. — Weigh latency to your customers, data-residency requirements, and price — then document the choice and the reasoning. The alternative is choosing it at 2 a.m. mid-incident based on whatever region happens to have machines, which is how one-night architecture decisions become ten-year regrets.
- Ask your provider the capacity question at your next renewal. — Whatever your size, add one line to the negotiation: "What capacity assurances come with this commitment, and can we get them in writing?" Enterprises can push for contractual language. Smaller companies may get less — but the answer itself tells you how your provider ranks you, and that's worth knowing before you deepen the dependency.
If you want a place to start#
Capacity risk is really a vendor-dependency question wearing new clothes, and it rewards a structured look. If you want a starting point for mapping which suppliers — cloud providers included — your business genuinely depends on, a NIST CSF readiness toolkit like the one we built at DLegendDigital covers supplier and dependency risk in its govern-and-identify sections. And if you'd rather have a second set of eyes on your region strategy and reservation posture, that's the kind of short, focused review our PBF consulting work does. This post is independent of both — every action above stands on its own.
The answer to the question#
So what do you do when the most reliable assumption in your IT strategy stops being true? You stop treating capacity as weather — something that's just there — and start treating it as a supplier relationship: mapped, tested, and negotiated like every other dependency your business runs on.
Here's what I expect over the next twelve months: capacity language starts showing up in cloud contracts the way security language did a decade ago — first as a differentiator, then as table stakes. Reservations become the default for anything critical. And the gap between companies that tested their failover and companies that believed their quota will get expensive.
The cloud isn't broken. It's just finite now — publicly, officially finite. The companies that adjust first will barely notice. The ones that don't will find out during their worst week.
I'd rather you be in the first group.
— Charles Redding
About the author
Charles Redding
Founder of DLegendDigital. 35+ years of enterprise technology leadership across audit, risk management, cybersecurity, and AI. Former CIO, VP of Technology, and Director at organizations ranging from high-growth startups to $4.3B global enterprises.



