
TLDR
Amazon's custom silicon business, spanning Trainium, Inferentia, Graviton and Nitro chips, has crossed a $20 billion annual revenue run rate while growing at triple-digit percentages year on year. Major AI firms including OpenAI, Anthropic, Meta and Uber have all committed to Amazon silicon, signalling a broad shift away from Nvidia-only infrastructure. Competition from Amazon's own chips is already pressing down on Nvidia's pricing power, with analysis finding an effective $8.69 per H100 equivalent ceiling in US-Virginia spot markets. For Australian businesses running AI workloads, cheaper inference costs and rising local data centre demand are the practical consequences.
KEY TAKEAWAYS
The $20bn number and what sits behind it
Amazon's custom silicon business cleared a $20 billion annual revenue run rate in the first quarter of 2026, chief executive Andy Jassy confirmed in the company's quarterly earnings release.verifiedVerified Source: ir.aboutamazon.com[1] The figure covers four chip lines: Graviton, the company's general-purpose CPU launched in 2018; Trainium, its purpose-built AI training accelerator; Inferentia, focused on inference workloads; and Nitro, the custom network interface card underpinning EC2 instances.
According to his 2025 letter to shareholders, Amazon's annual revenue run rate for its chips business (inclusive of Graviton, Trainium, and Nitro, the EC2 network interface card) is now over $20 billion, and growing triple digit percentages year-over-year.verifiedVerified Source: aboutamazon.com[3] AWS cloud overall grew 28 per cent in the same period, the fastest growth in 15 quarters by Jassy's account, with custom silicon increasingly central to that result.[1]
Why the biggest AI labs are buying non-Nvidia silicon
The customer list underpinning the $20 billion figure is not a collection of start-ups testing alternative hardware. OpenAI has committed to consuming approximately 2 gigawatts of Trainium capacity, Anthropic has secured up to 5 GW, Meta is deploying tens of millions of AWS Graviton cores, and Uber is running AI workloads on Graviton4 and Trainium3.[1] These are production commitments at scale, not pilots.
The economics explain the migration. AWS's second-generation Trainium2 delivered a 30 to 40 per cent price-performance advantage over comparable GPUs and was nearly fully subscribed, according to Jassy's shareholder letter.[3] When a chip running your training job costs 30 per cent less per unit of output and is actually available, the procurement decision simplifies quickly.
Inferentia reinforces the same logic at the inference end of the stack. AWS's second-generation Inferentia chip, which powers EC2 Inf1 instances, delivers up to 2.3 times higher throughput and up to 70 per cent lower cost per inference than comparable EC2 instances.[2] For generative AI applications processing millions of requests daily, that cost differential compounds rapidly.
What this does to Nvidia's pricing power
Nvidia retains a dominant share of AI accelerator revenue, and its H100 and H200 GPUs remain the default choice for many training workloads. Amazon's data changes the ceiling on what Nvidia can charge, even for customers who never touch a Trainium chip.
Analysis by Signwl found that AWS spot pricing for Trainium and Inferentia instances set an effective flat ceiling of $8.69 per H100 equivalent for on-demand GPU pricing in the US-Virginia region, constraining Nvidia's pricing power in that market.verifiedVerified Source: signwl.com[4] Signwl's research, which Bushletter could not independently corroborate through a separate source, attributes the ceiling effect to the credible threat of substitution: buyers can walk away, and sellers know it.
The implication is structural rather than cyclical. So long as Amazon continues scaling Trainium and Inferentia supply, the competitive pressure on Nvidia's on-demand pricing does not ease simply because AI demand grows. Demand growth increases the volume of compute purchased; it does not, on its own, restore a monopoly premium.
A pattern, not a one-off
Amazon's chip revenue milestone sits inside a broader industry realignment. OpenAI has developed its own inference chip, known internally as Jalapeño. Anthropic has explored bespoke silicon designs. Chinese AI firm DeepSeek is reported to be building its own inference chips to cut dependence on both Nvidia and Huawei hardware.
The pattern is consistent: every major AI lab at sufficient scale is treating silicon as a strategic variable rather than a commodity input. Each custom chip programme reduces the addressable market available to Nvidia at the top end of the customer pyramid, even as the total market expands. Amazon's $20 billion run rate is the most concrete measure yet of how far that substitution has already progressed.
Amazon invested in the full stack, co-designing chip, server, networking and software layers together for Trainium, rather than retrofitting a general-purpose GPU to AI workloads. That architecture decision, taken years before the current AI infrastructure boom, is now generating revenue at a scale Jassy described in April 2026 as growing triple digits year on year.[1]
What it means for Australian businesses and local data centres
The practical consequence for Australian companies running AI workloads flows directly from the inference cost numbers. Lower per-query costs reduce the barrier to deploying generative AI applications at production scale. A business that previously found API costs prohibitive at high request volumes faces a different calculation when Inferentia-backed instances cut inference costs by up to 70 per cent.[2]
AWS operates data centre regions in Sydney and Melbourne. As Australian enterprise demand for AI inference capacity grows, driven partly by falling costs, the utilisation case for expanding local infrastructure strengthens. Rising data centre demand feeds construction activity, power procurement and local employment, making the silicon story relevant well beyond the hyperscaler balance sheet.
The Signwl ceiling effect on Nvidia pricing, if it holds and extends to other regions, would also reduce the cost of on-demand GPU access for Australian developers who build on Nvidia hardware rather than Amazon's own chips, since regional pricing typically tracks US benchmark rates over time.[4] Amazon's Trainium3 is the next generation of the accelerator, with Uber among the first customers confirmed for that platform as of the April 2026 quarterly earnings release.
SOURCES & CITATIONS
FREQUENTLY ASKED QUESTIONS
What chips does Amazon's custom silicon business include?
How does Amazon's Inferentia chip compare to standard GPU instances?
Which major AI companies have committed to using Amazon's custom chips?
Does cheaper Amazon silicon affect Nvidia GPU pricing?

Nathan Cross writes about big technology companies and the economics of a generation coming up behind them. He is interested in scale, and in who pays for it.



