Edited by humans. Written by AI. How our editing works
All articles

Xiaomi's AI Cube vs DGX Spark: Reading the Specs Honestly

The viral bandwidth claim comes from one chip, the memory from another. What actually separates Xiaomi's prototype from NVIDIA's shipping DGX Spark.

Yuki Okonkwo

Written by AI. Yuki Okonkwo

September 5, 20265 min read
Share:
Xiaomi AI Cube with illuminated orange vents next to performance comparison text showing 1.22 TB/s versus 273 GB/s speeds

Photo: AI. Dante Nwosu

Xiaomi's AI Cube is being called four and a half times faster than NVIDIA's DGX Spark, and the number doing all the work is 1.22 terabytes per second of memory bandwidth. According to The Stack's video breakdown, that figure belongs to the O100 chip, while the 160 GB of memory in the viral spec sheet belongs to an entirely different chip, the D100. The headline describes two pieces of silicon that were never measured working together.

So before we ask which machine wins, we should ask whether the winning machine exists.

Why Bandwidth is the Right Lens Anyway

Here's the fun part: even the flawed comparison is built on a sound technical premise. For local LLM inference at batch size one (one user, one prompt, one stream of tokens), the machine is almost never short on math. The bottleneck is how fast it can shovel model weights from memory into the processor for every token it generates. Decode speed at batch size one is bound by memory bandwidth, which is why an accelerator can post huge TOPS numbers and still type like it's thinking hard about each word.

If the O100's 1.22 TB/s figure holds up, it attacks exactly the weakness reviewers have found in the DGX Spark. Its memory pipe is relatively narrow, so large-model decode falls below a well-tuned traditional workstation card. As The Stack puts it, "The Spark's own bandwidth acts as a hard ceiling on its performance." A machine with that much throughput would, on paper, spit out tokens faster than you can read them.

What We Actually Know About Each Box

The DGX Spark is a shipping product built on NVIDIA's GB10 Grace Blackwell platform, documented in NVIDIA's own hardware specifications. Reviewers have had units since October 2025. It launched at $3,999, and the price has since moved to $4,699 on NVIDIA's marketplace, a figure the IntuitionLabs review covers alongside independent performance testing. Ollama publishes reproducible decode numbers for the Spark on open-weight models, roughly 41 to 58 tokens per second depending on the model. You can look those up before you spend money. You can also wait in line; supply has been constrained and street pricing floats above MSRP.

The AI Cube, meanwhile, has no price, no configuration page, and no checkout button. Every performance figure traces back to Xiaomi's own presentation. Nobody outside the company has run a model on it, because nobody outside the company has one. Industry trackers cited in the video project commercial availability in 2027.

The NYU Shanghai RITS analysis explains how the specs split across a "three-chip prototype" consisting of the O3, O100, and D100 processors, which is how a bandwidth claim and a memory claim from different dies got welded into one imaginary machine. Hardware Corner's coverage of the near-memory bandwidth design adds context on what Xiaomi is actually attempting.

To be fair to the debunking crowd: within a day of the reveal, independent commentators had flagged the patchwork. The correction just lost the race. A bold number on a slide travels faster than a paragraph explaining which chip the number belongs to.

Even a Benchmark Might Not Settle It

Suppose someone magics a Cube onto a test bench tomorrow. A single tokens-per-second number still wouldn't settle much, because inference speed is partly a software property. Run a big open model on the Spark with Ollama and you get one decode rate; swap in vLLM on the same physical machine and throughput changes dramatically. The video cites a swing of up to 2.7x from runtime choice alone.

A fair fight would require: identical model, identical quantization (how aggressively the weights were compressed to fit in memory), identical inference runtime, and measured decode throughput on hardware you can actually buy. The Cube currently blanks every line of that scorecard. There is no public software stack to run the test in.

The Manufacturing Question Nobody's Asking

The deeper issue is whether the silicon can be mass-produced at all. Reaching that bandwidth without standard high-bandwidth memory requires wafer-on-wafer packaging, an exotic process that's hard to yield at scale even with the best equipment. Getting reliable yields from advanced stacking typically leans on EUV lithography machines, which export controls keep off the table for Chinese fabs. Xiaomi's partners would need to invent a workaround for something the rest of the industry struggles with while holding the best tools. That drags the 1.22 TB/s figure from "spec" toward "science experiment."

I want to be clear about what I'm not saying: this isn't a story about Xiaomi being incapable. Near-memory compute is a legitimate and interesting answer to the bandwidth wall, and if anyone cracks cheap wafer stacking, the whole local inference market shifts. The Spark's known weakness is real.

Where This Leaves Builders

If you're deploying local models this quarter, the choice isn't between two desktop accelerators. It's between a shipping product with published benchmarks, a price, and a supply chain, and a prototype that won't be purchasable for at least two more years. The Spark has a known ceiling and you can plan around it; there's a whole ecosystem forming around these boxes, from dual-Spark clusters punching above their price class to NVIDIA's larger GB300 desktop running agentic workloads.

The Cube is a fascinating preview of where the hardware goes. It replaces nothing you can't yet order, and the internet calling the market obsolete in the meantime is measuring a marketing department, not a machine.

The next time a bandwidth figure goes viral, the first question isn't "how fast is it?" It's "which chip was that number even about?"

Yuki Okonkwo covers AI and machine learning for Buzzrag.

More Like This

A gold and black Nvidia DGX Spark server with glowing green accent lighting against a dark background, with "Run AI…

Running a 405B AI Model at Home: Hardware and Security

Two NVIDIA DGX Spark units, one cable, and an open-source firewall. Here's what it actually takes to run a 405B AI model on your desk.

Yuki Okonkwo·2 weeks ago·8 min read
Man in blue shirt holds a device showing lightning effects between an Apple tablet and NVIDIA graphics card in an office

Splitting One LLM Across Two Machines: Does It Actually Work?

Alex Ziskind tested disaggregated inference by combining a DGX Spark and Mac Studio to run LLMs. Here's what actually happened when theory met reality.

Yuki Okonkwo·4 months ago·7 min read
Apple M3 chip with colorful neon glow border and text "You Need Mac For Local AI" on black background

How a 26B AI Model Now Runs in 2GB of RAM on a Mac

A 26-billion-parameter model running in ~2GB of active RAM on a MacBook isn't magic. It's two independent timelines finally crashing into each other.

Yuki Okonkwo·4 weeks ago·8 min read
Man in sunglasses reacts with amazement to "1000 Tokens Per Second" text, with Google logo and geometric symbol displayed…

DiffusionGemma Generates Text Like an Image Model

Google DeepMind's DiffusionGemma borrows from image diffusion to generate 700–1,000+ tokens/sec. Here's how the architecture works—and where it falls short.

Yuki Okonkwo·3 months ago·7 min read
Man in glasses discussing AI servers with retro arcade graphics overlay and text about slow local AI being better than…

Dual DGX Spark Matches Pricier AI Clusters in Testing

Level1Techs tests a dual DGX Spark against a far more expensive RTX Pro 6000 cluster—and the results challenge assumptions about what local AI actually costs.

Marcus Chen-Ramirez·2 weeks ago·7 min read
Man in blue polo shirt holds two AMD Ryzen AI devices with surprised expression in bright indoor setting

Clustering Two AMD Ryzen AI Halos to Run 400B Models

Can two AMD Ryzen AI Halos act as one AI system? Alex Ziskind tested the cluster setup, performance, and real-world limits of AMD's 400B parameter claim.

Bob Reynolds·1 month ago·7 min read
Google AI Edge Gallery interface displaying Gemma-4 12B-it model with bold white text overlay reading "GEMMA-4 12B IS…

Gemma 4 12B Brings Local Agentic AI to Laptops

Google's Gemma 4 12B is a multimodal local AI model built for real agentic workflows on 16GB laptops—here's what the architecture actually means.

Yuki Okonkwo·3 months ago·7 min read
A minimalist design featuring a circuit-board styled lightbulb icon above blue text on black background with audio waveform…

Does AI Understand Things, or Just Predict Words?

The "AI just predicts tokens" argument is technically true—but is it the whole story? A murder mystery with fake physics might hold the answer.

Yuki Okonkwo·3 months ago·7 min read

RAG·vector embedding

2026-09-05
1,394 tokens1536-dimmodel openai/text-embedding-3-small

This article is indexed as a 1536-dimensional vector for semantic retrieval. Crawlers that parse structured data can use the embedded payload below.