The chip nobody was talking about is now running half of AI
So Amazon CEO Andy Jassy just announced a $50 billion deal with OpenAI — and as part of that deal, AWS will be the exclusive provider of OpenAI's new AI agent builder called Frontier. That's a massive win. But the really interesting part is what's powering all of this behind the scenes: a chip called Trainium, built at a secret AWS chip development lab.
TechCrunch got an exclusive tour of that lab. And the story of what Amazon has quietly built there is genuinely surprising.
What is Amazon Trainium?
Trainium is Amazon's custom AI chip — designed specifically for training and running large AI models. Think of it as Amazon's answer to Nvidia's H100/H200 GPUs that every AI company has been fighting over for the past two years.
The difference: Nvidia charges a premium because it has near-monopoly status. Trainium is Amazon's attempt to offer a cheaper alternative — with the added advantage that if you're already running on AWS (which most companies are), switching to Trainium is much simpler than buying Nvidia GPUs on the open market with months-long wait times.
The lab is led by Kristopher King (lab director) and Mark Carroll (director of engineering) — and from the sounds of the TechCrunch tour, this team has been quietly doing some serious engineering for years while everyone focused on Nvidia.
The numbers that matter
Here's what makes this concrete: as part of the OpenAI deal, AWS has committed to supplying OpenAI with 2 gigawatts of Trainium computing capacity. Let that sink in. Two gigawatts. That's not a spec — that's a commitment to deliver an almost incomprehensible amount of AI compute to one customer.
And this is on top of existing commitments. Anthropic — OpenAI's biggest competitor — has been running on Trainium since AWS's early days as Anthropic's primary cloud platform. Amazon's own Bedrock AI service runs on Trainium. The lab is reportedly producing chips faster than it can keep up with demand from existing customers.
Now add OpenAI and apparently Apple to the list. When you have Anthropic, OpenAI, AND Apple all choosing your chip, you're no longer an experiment. You're infrastructure.
Why OpenAI chose Amazon over Microsoft — and why this is complicated
Here's where it gets politically messy. Microsoft has been OpenAI's primary backer and cloud partner for years — that's the relationship that put Azure in the center of the AI world. Now OpenAI is signing a massive deal with Amazon, making AWS the exclusive provider for Frontier (its AI agent platform).
The Financial Times reported this week that Microsoft may believe the OpenAI-Amazon deal actually violates Microsoft's own agreement with OpenAI — specifically the clause that says Microsoft gets access to all of OpenAI's models and technology. That's a significant legal and business conflict brewing.
For the deal to survive exactly as announced, some careful legal navigation is ahead. But the fact that OpenAI felt comfortable enough to sign a $50 billion deal with Amazon tells you something: AWS — and Trainium specifically — has become genuinely competitive with what Microsoft's Azure offers for AI workloads.
The Nvidia disruption angle
This is the story the entire semiconductor industry is watching. Nvidia currently has something close to a monopoly on AI training chips. The H100 and H200 GPUs that power ChatGPT, Gemini, Claude, and every other major AI model are almost all Nvidia. Their market cap reflects that dominance — Nvidia is worth more than most countries' GDP.
But that monopoly has a crack: cost and availability. Nvidia chips are expensive and hard to get. Every major cloud company — Google (with TPUs), Meta (with MTIA), and now Amazon (with Trainium) — has been building custom silicon specifically to reduce their Nvidia dependence. If Trainium is good enough that OpenAI, Anthropic, and Apple choose it over Nvidia, that's the beginning of real competition in AI silicon.
It won't happen overnight. Nvidia's CUDA software ecosystem — the tool that developers use to program their chips — has a decade-plus head start. But Amazon has been making Trainium increasingly compatible with existing AI frameworks, reducing the friction of switching.
What this means for Indian developers and startups
India is one of AWS's fastest-growing markets. Indian startups — from Bangalore to Hyderabad to Chennai — run enormous amounts of workload on AWS. If you're building an AI product and using AWS, Trainium is already available to you through Amazon Bedrock.
The key benefit: lower cost per AI inference. Running AI models on Trainium through AWS can be significantly cheaper than using equivalent Nvidia GPU instances. For Indian startups burning runway on API costs, this matters. If you're making a lot of ChatGPT or Claude API calls for your product, at some point it makes sense to run your own model on AWS Trainium instead — and the cost difference can be substantial.
For Indian engineers interested in AI hardware: the Trainium lab is in the US, but AWS's chip engineering teams are distributed globally and AWS has a significant presence in India. This area of work — custom AI silicon design — is one of the fastest-growing specializations in semiconductor engineering, and India already supplies a significant portion of global chip design talent.
The Amazon Prime / AWS connection Indians should know
Here's something most people don't realize: when you use Amazon Prime in India — for Flipkart-competitor Amazon shopping, Prime Video, or anything else — you're using infrastructure built on the same AWS that now runs Trainium chips for OpenAI. Amazon is quietly one of the most important technology companies in India, not just as a retailer but as cloud infrastructure for thousands of Indian businesses and startups.
The Trainium story is part of Amazon's larger strategy: be the utility layer for everything AI. If AI is the electricity of the next decade, AWS wants to be the power grid — and Trainium is the generator.
My honest take
Amazon has been playing a long, patient game here. While everyone was obsessing over Nvidia's stock price and OpenAI's ChatGPT, AWS quietly built custom AI silicon, convinced Anthropic to run on it, built their own AI service (Bedrock) on it, and now has OpenAI and Apple as customers.
The $50 billion OpenAI deal is the announcement that makes this public. But the real story is the 5-10 years of engineering work that made this possible. That's how you build a genuine Nvidia competitor — not by announcing it in a press release, but by having OpenAI, Anthropic, and Apple choose your chip because it's actually better for their use case.
Whether Trainium can truly challenge Nvidia's dominance at the frontier of AI training is still an open question. But for the inference side — running AI models efficiently at scale — it's already proven itself. And inference is where most of the money in AI actually is.




Comments (0)
Be the first to comment!