AWS GPU Prices Jump 20% as Multi-Year Deals Replace Spot Market
EC2 Capacity Blocks for Nvidia GPUs now cost $14.04 per accelerator-hour, as the spot market vanishes into multi-year reservations that extend through 2028.
tomshardware.com
On July 1, 2026, Amazon Web Services raised on-demand pricing for its GPU-backed EC2 Capacity Blocks by roughly 20 percent across the board. The new hourly rate for a single Nvidia accelerator inside a P6-B300 instance landed at $14.04. A P6-B200 now bills at $12.355 per accelerator-hour. Even the previous-generation P5, which runs Nvidia's H200 silicon, ticked up to $5.191 in US regions, Yahoo Finance reported on June 26. These are list prices. The actual availability of those instances at those list prices, on demand, without a prior reservation, is a separate question the public pricing page does not answer.
The price hike itself is not surprising. What is notable is that AWS published it at all. For most of 2025, the effective price of frontier GPU compute was discoverable only through a broker, a sales rep, or a three-year committed-use discount negotiated under non-disclosure. The fact that AWS has reset its public list prices upward signals that even the sticker price is now being pulled toward the market-clearing rate, a rate that the neocloud ecosystem has been setting in plain view for at least twelve months.
Chris Campbell, senior director of AI solutions and GPU-as-a-service at World Wide Technology, offered a window into what those clearing rates actually look like in a CRN report published June 17. Campbell sketched the economics of a representative deal: a one-year reservation for 1,000 Nvidia B300 accelerators, priced at $5 per chip per hour. That single contract clocks in at roughly $45 million for twelve months. Even a channel partner taking a five percent margin walks away with a couple million dollars. The math Campbell offered was illustrative, not a specific client invoice, but the $5-per-B300-per-hour figure aligns with what several neocloud providers have been quoting to enterprise buyers since early 2026.
The neocloud market, a category that barely registered in infrastructure spending reports three years ago, reached $9 billion in quarterly revenue in the fourth quarter of 2025, according to data from Synergy Research Group cited in the same CRN report. That represented a 223 percent year-over-year increase. For the full year 2025, Synergy pegged the neocloud market above $25 billion. The firm's forecast calls for the market to hit $400 billion by 2031, a compound annual growth rate of 58 percent sustained for half a decade. Those are not cloud-growth numbers. Those are the growth numbers of a parallel compute economy assembling itself in real time.
The core dynamic driving those figures is straightforward. Demand for the most advanced GPU instances far exceeds supply, and the hyperscalers cannot build data centers fast enough to close the gap. Laurelle Roseman, vice president of global partnerships at Nebius, told CRN that for every GPU cluster the company brings online, "we have four to five customers that are lined up to take it." Nebius is investing $2.3 billion to build out capacity in the UK alone, with deployments targeting 65 megawatts by 2027, and it broke ground in May on its first gigawatt-scale campus in Missouri. Even at that pace, Roseman said, "the demand is outpacing the supply just in general in the industry."
The spot market for frontier GPUs, the kind where a startup could spin up a hundred H100s for a weekend fine-tuning run with a credit card, has effectively ceased to exist. Allocation today happens through reservation contracts that run twelve to thirty-six months, often with non-refundable commitments attached. The neoclouds are filling a structural gap: enterprises that cannot get GPU instances from AWS, Google Cloud, or Azure at the scale they need, or cannot wait the six to nine months it takes for on-premise Nvidia deployments to be racked and cabled, are routing their budgets through CoreWeave, Crusoe, Lambda, Nebius, and Vultr instead.
The supply-side constraints are not limited to silicon. The memory pipeline that feeds every GPU is under strain that industry executives are now describing in historic terms. SK Hynix CEO Kwak Noh-jung said on July 10 that the global memory industry is heading for its worst-ever supply shortage in 2027, Reuters reported, with demand projected to outstrip supply beyond 2030. SK Hynix supplies the majority of high-bandwidth memory used in Nvidia's Blackwell and Vera Rubin platforms. A memory bottleneck in 2027 would directly constrain the number of GPU accelerators that can be assembled and shipped, which in turn pushes reservation prices higher for the silicon that does reach the market.
The memory crunch extends well beyond HBM. Tom's Hardware reported in October 2025 that AI data centers were absorbing flash memory and storage supply at a rate that was pushing prices across the entire DRAM stack higher, including legacy DDR4 and DDR3 modules that have no direct connection to AI workloads. When AI infrastructure consumes enough of the global memory supply to raise prices for commodity DRAM, the pricing signal travels backward through the supply chain and shows up in the per-chip cost of every accelerator that leaves the fab.
From the hyperscaler perspective, what I've heard from our partners is that they don't make margin and they don't get access. So those are two key things that kill their business., Laurelle Roseman, VP of global partnerships, Nebius
What Roseman described is the channel economics of a constrained market. Hyperscalers allocate their most sought-after GPU instances to their largest committed customers first. A solution provider attempting to resell AWS GPU capacity to a mid-market enterprise client faces two problems: the margin on a straightforward resale is negligible, and the instances may not be provisionable at any price during periods of peak demand. Neoclouds, by contrast, are building partner programs specifically designed to let channel partners wrap managed services, fine-tuning pipelines, and Kubernetes orchestration around GPU reservations and capture double-digit margins on the total contract.
The discrete GPU market, which encompasses the add-in boards that Nvidia, AMD, and Intel ship through their partner networks, tells a parallel story of supply holding but at high equilibrium prices. Jon Peddie Research reported in June that Q1 2026 discrete GPU shipments remained relatively flat quarter over quarter, with Nvidia maintaining roughly 90 percent market share. Flat shipments during a memory crisis and an AI infrastructure buildout imply that every unit that can be produced is being sold, at prices that reflect scarcity rather than manufacturing cost. The consumer GPU market has become a residual claimant on silicon that is overwhelmingly steered toward data center SKUs.
AMD's position in this market has strengthened, not because it has matched Nvidia's software ecosystem, but because hyperscalers are actively diversifying their silicon supply. Seeking Alpha reported in late June that AMD has secured multi-year, multi-gigawatt deals from Meta and OpenAI, agreements that function less like traditional chip purchase orders and more like capacity reservations at the foundry level. These deals do not create a spot market for AMD's Instinct accelerators. They are bilateral commitments that lock up supply before it reaches any public price list.
Additional supply-chain pressure comes from an unexpected direction. Forbes reported in April that helium gas, which is used as a coolant in advanced semiconductor manufacturing, is in increasingly short supply globally. Helium is not a headline input like HBM or silicon substrates, but it is non-substitutable in several process steps for the most advanced nodes. A helium shortfall constrains wafer output at TSMC, Samsung, and Intel alike, which constrains the total number of GPU dies that can be produced, which feeds directly into the reservation-only allocation regime that now governs the market.
Jeremy Duke, founder and chief analyst at Synergy Research Group, described the structural shift to CRN in terms that go beyond a simple supply-demand imbalance. "What we are observing is not merely the emergence of a new class of cloud provider, but a deeper structural realignment in the architecture of computation itself," Duke said. Traditional cloud infrastructure was built around generalized elasticity: spin up, spin down, pay by the second. AI workloads impose constraints that are far more rigid. Training a frontier model requires weeks of uninterrupted access to thousands of tightly coupled accelerators connected by high-bandwidth interconnects. You cannot preempt that workload midway through a training run and resume it later on a different cluster without incurring costs that often exceed the compute savings.
This rigidity is what killed the GPU spot market. Spot instances work when workloads are stateless, fungible, and tolerant of interruption. A batch inference job running at batch size 32 on eight GPUs can survive a two-minute preemption warning. A distributed training run spanning 4,096 H200s with a synchronized gradient step every few hundred milliseconds cannot. The economics of training have therefore bifurcated the market: inference workloads can still find spot capacity on last-generation hardware, particularly A100 and H100 instances that are being displaced by Blackwell deployments at the hyperscalers. But any workload that requires Blackwell-class accelerators with HBM3e memory is, by default, a reservation workload.
The per-token price that end users see when they call an inference API does not yet reflect the full upstream cost of the reservation regime under which the serving infrastructure was procured. A frontier model provider that reserved 2,000 B300 accelerators on a three-year commitment at $5 per chip-hour is amortizing roughly $263 million over the life of the contract. The inference revenue those chips generate depends on utilization rates, batch sizes, and the token-per-second throughput achievable on the specific model architecture. None of those figures are public. The gap between the $14.04 per-accelerator-hour list price that AWS published on July 1 and the effective per-token cost that reaches a customer invoice is where the margin in the inference stack either accumulates or evaporates.
What to watch: SK Hynix's 2027 supply warning is the nearest-term checkpoint. If HBM3e and HBM4 output falls short of the volumes that Nvidia's Vera Rubin ramp requires, the reservation-only market that now governs Blackwell allocation will harden further. The $5 per B300 per hour that Campbell cited as a going rate in mid-2026 may look cheap by the second half of 2027, not because silicon got more expensive to manufacture, but because the memory modules that make those chips useful for AI workloads simply did not ship in sufficient quantity. The spot market, such as it was, is not coming back.