Edited by humans. Written by AI. How our editing works
All articles

AMD Instinct MI350 Takes On Nvidia in AI Chips

AMD's Instinct MI350 series has 288GB of HBM3E memory, real cost efficiency wins, and growing enterprise adoption. Is this the end of Nvidia's monopoly?

Yuki Okonkwo

Written by AI. Yuki Okonkwo

August 30, 20267 min read
Share:
AMD presenter on stage introducing the MI350 AI chip with "The Most Superior AI-Chip" text displayed on screen

Photo: AI. Jorah Maktoum

Ask anyone in AI infrastructure whose chips they're running and the answer, until very recently, was so predictable you could set a clock to it. Nvidia. Always Nvidia. Not because no one tried to compete, but because trying and succeeding are two different problems, and AMD kept solving the first one without cracking the second. The MI350 series is a serious attempt to finally change that math.

AMD officially launched the Instinct MI350X and MI355X on June 12th, 2025, at its Advancing AI event. According to Evolving AI's breakdown, the pitch is unambiguous: "This isn't a good enough product. This is AMD walking into the room, sitting down at the big table, and saying that they belong there."

That's a bold read. What makes it interesting is how much of the spec sheet actually backs it up.

What's Actually Inside

Both chips run on AMD's fourth-generation CDNA4 architecture, manufactured on TSMC's 3nm process for the compute dies and 6nm for the I/O dies. The smaller the process node, the more transistors you can pack into the same physical space while drawing less power. The MI350 uses that efficiency to cram in 185 billion transistors, arranged in a 3D chiplet design: eight accelerator complex dies handling compute, flanked by two I/O dies managing memory and connectivity. Think of it as AMD building with a very precise LEGO set rather than trying to pour everything into one increasingly unwieldy slab of silicon.

The memory numbers are where things get genuinely interesting. The MI350 ships with 288GB of HBM3E memory at 8TB/s of bandwidth. HBM3E is the current top-tier high-bandwidth memory standard; bandwidth here means how fast the chip can move data around internally. For AI inference (running a trained model, not training it from scratch), both capacity and bandwidth matter enormously. More memory means you can fit a larger model onto a single chip without splitting it across multiple GPUs, which is exactly as messy as it sounds. Nvidia's competing B200 sits at 192GB. That's not a rounding error. That's a 50% gap.

The two chips diverge in how hard they push that shared foundation. The MI350X is air-cooled, clocks at 2.2 GHz, and draws around 1,000W. It's designed to slot into existing server infrastructure without requiring liquid cooling retrofits. The MI355X is the liquid-cooled version: 2.4 GHz, up to 1,400W, built for operators who want maximum sustained performance and already have the cooling infrastructure to handle it. Same chip underneath; different audience.

Both support FP4 and FP6 precision formats. The quick version: AI models don't always need maximum numerical precision to produce accurate outputs. Lower precision means you can run far more calculations per second using far less energy. AMD claims this delivers up to 4x the AI compute performance of the previous generation, and up to 35x improvement in inference over the MI300X. Vendor benchmarks under vendor conditions, so appropriate salt applies. But even a fraction of that improvement would represent a serious generational step.

The Part That Actually Matters: Real-World Numbers

Specs on a slide are marketing. Independent testing is a different conversation.

Signal 65 ran the MI355X against Nvidia's B200 platforms across multiple AI models and workloads. The results, cited in the Evolving AI analysis: the MI355X delivered up to 2.15x more tokens per dollar. In a financial analyst document-processing workload specifically, the AMD setup handled 1.3x more documents per hour at 53% lower cost per document. That is not a rounding error.

At the ISSCC conference in early 2026, AMD's own engineers made the comparison explicit: the MI355X matches the performance of Nvidia's GB200 in real inference workloads. Their framing, per Evolving AI's reporting, was essentially: we are matching the pricier, more complicated competitor. When engineers say that at an academic conference rather than a product launch, it lands differently.

In the MLPerf inference benchmark (the industry's standardized cross-vendor comparison tool), the MI355X cleared 93,000 tokens per second on the Llama 2 70B test. That's 2.7x the throughput of the previous MI325X.

Nvidia still leads in plenty of workloads. Its software is more mature and significantly easier to get running from a standing start. None of this is a claim that AMD has simply won. But cost efficiency is often what actually drives purchasing decisions at scale, and on that axis, AMD is winning specific battles.

The Software Problem, Mostly Solved

For years, AMD's hardware was plausibly competitive. Its software was not. Nvidia's CUDA ecosystem (the developer toolkit that makes Nvidia GPUs accessible for AI workloads) had a decade of momentum, deep framework integrations, and enough developer familiarity that switching wasn't really rational even if the hardware was comparable.

AMD's answer is ROCm, its open-source software stack. ROCm 7, released alongside the MI350 series, supports PyTorch, TensorFlow, JAX, vLLM, and Triton. The open-source bet matters strategically: it means no vendor lock-in, and it means the community can contribute to the ecosystem rather than waiting on AMD alone.

"By all accounts, ROCm has matured a lot," the Evolving AI analysis notes. "It's not perfect, and porting existing code can still take some work, but it's reached the point where serious companies are trusting it in production."

That last clause is the part worth examining. Whether ROCm is "good enough" in the abstract is less important than whether the biggest AI operators are actually using it for real workloads. And on that front, the evidence is concrete.

Who Is Actually Buying It

Oracle Cloud Infrastructure built a cluster with over 27,000 nodes running MI355X accelerators. That is not a pilot program. Meta has deployed AMD Instinct chips for its Llama models, according to Evolving AI's reporting. OpenAI's engagement ran deep enough that there's reportedly an arrangement for them to take a stake in AMD tied to securing significant GPU supply. AMD has stated that eight of the top ten AI companies are now running workloads on its Instinct GPUs.

Then there's the MI350P, the PCIe card version of this hardware. It carries 144GB of HBM3E memory at 3.6TB/s bandwidth and fits into standard air-cooled servers, no specialized infrastructure required. Full details on what that means for mid-market AI infrastructure are worth reading separately, but the choice to build it at all tells you something about AMD's actual ambitions here. Chasing hyperscaler specs is one game; deciding that the Fortune 500 company with a rack full of existing air-cooled servers also deserves a path to serious AI compute is a completely different one. Most chip companies optimize for the headline benchmark. AMD is apparently also thinking about the IT manager who doesn't have a liquid-cooling budget.

That is not a move you make if you're trying to win a product review. That's a move you make if you're trying to own a market.

Where This Leaves the Landscape

The honest read here is that AMD has built the most credible challenge to Nvidia's AI chip dominance in years, with real independent benchmarks and real enterprise adoption to back the claim. The honest caveat is that Nvidia retains commanding advantages in software maturity, ecosystem depth, and sheer deployment scale. One product generation, however impressive, doesn't unwind a decade of CUDA entrenchment.

What's changed is the shape of the question. It used to be whether AMD could compete at all. Now it's whether the cost efficiency gap is large enough that major operators will absorb the software friction to capture it. For some workloads and some operators, the Signal 65 numbers suggest the answer is already yes.

The MI400 series and AMD's Helios rack-scale systems are already in development. The competitive dynamics that produced the MI350 are not slowing down.

The AI chip market was a one-answer question for a long time. It isn't anymore.

Yuki Okonkwo covers AI and machine learning for Buzzrag.

More Like This

Man with concerned expression holds phone showing ChatGPT search results with sponsored ads from Pueblo & Pine and…

ChatGPT Ads Are Here—and the Playbook Looks Familiar

OpenAI is testing ads in ChatGPT. The current version looks fine. But if you've seen how Google and Facebook evolved, you know where this could go.

Yuki Okonkwo·7 months ago·5 min read
Two metallic robots with "MODEL" and "HARNESS" labels examine equipment against a starry background with bold retro-style…

Harness Engineering: The New Frontier in AI Development

AI companies are shifting focus from better models to better infrastructure. Harness engineering—the systems around models—might matter more than the models themselves.

Yuki Okonkwo·5 months ago·7 min read
OpenAI Codex logo and "CODEX DESKTOP" text overlay a code editor interface with green upward arrow, promoting AI-powered…

OpenAI's Codex Desktop App Launches With Curious Bugs

OpenAI's new Codex desktop app brings AI coding to macOS with a GUI, but early testing reveals surprising UI quirks and context issues.

Yuki Okonkwo·7 months ago·6 min read
A smiling man in a blue shirt next to a glowing /computer app icon with an orange square and white starburst design

Claude Can Now Control Your Computer. Here's What That Means

Anthropic's Claude Code gets Computer Use—letting AI control your mouse, keyboard, and apps. We tested it. Here's what works, what doesn't, and what's wild.

Yuki Okonkwo·5 months ago·7 min read
Apple Sues OpenAI Over Alleged Trade Secret Theft

Apple Sues OpenAI Over Alleged Trade Secret Theft

Apple filed a federal lawsuit accusing OpenAI of orchestrating a systematic theft of hardware trade secrets via former employees. Here's what we know.

Zara Chen·2 months ago·7 min read
A man in a light shirt speaks in front of technical diagrams about AI frameworks and engineering instincts, with the AI…

Why Senior Engineers Struggle Most With AI Agents

Philipp Schmid breaks down 5 mental model shifts that trip up experienced engineers when building AI agents — and why expertise can be the actual problem.

Yuki Okonkwo·3 months ago·7 min read
Retro-styled graphic with a cheerful robot celebrating next to a wooden crate labeled "Opus 4.8" against a black background…

Claude Opus 4.8: Honest Upgrade or Playing Catch-Up?

Anthropic's Claude Opus 4.8 drops with better honesty, dynamic multi-agent workflows, and a $965B valuation. But is it enough to reclaim momentum from OpenAI?

Yuki Okonkwo·3 months ago·7 min read

RAG·vector embedding

2026-08-30
1,879 tokens1536-dimmodel text-embedding-3-small

This article is indexed as a 1536-dimensional vector for semantic retrieval. Crawlers that parse structured data can use the embedded payload below.