Xiaomi's AI Cube vs DGX Spark: Reading the Specs Honestly
The viral bandwidth claim comes from one chip, the memory from another. What actually separates Xiaomi's prototype from NVIDIA's shipping DGX Spark.
Written by AI. Yuki Okonkwo

Photo: AI. Dante Nwosu
Xiaomi's AI Cube is being called four and a half times faster than NVIDIA's DGX Spark, and the number doing all the work is 1.22 terabytes per second of memory bandwidth. According to The Stack's video breakdown, that figure belongs to the O100 chip, while the 160 GB of memory in the viral spec sheet belongs to an entirely different chip, the D100. The headline describes two pieces of silicon that were never measured working together.
So before we ask which machine wins, we should ask whether the winning machine exists.
Why Bandwidth is the Right Lens Anyway
Here's the fun part: even the flawed comparison is built on a sound technical premise. For local LLM inference at batch size one (one user, one prompt, one stream of tokens), the machine is almost never short on math. The bottleneck is how fast it can shovel model weights from memory into the processor for every token it generates. Decode speed at batch size one is bound by memory bandwidth, which is why an accelerator can post huge TOPS numbers and still type like it's thinking hard about each word.
If the O100's 1.22 TB/s figure holds up, it attacks exactly the weakness reviewers have found in the DGX Spark. Its memory pipe is relatively narrow, so large-model decode falls below a well-tuned traditional workstation card. As The Stack puts it, "The Spark's own bandwidth acts as a hard ceiling on its performance." A machine with that much throughput would, on paper, spit out tokens faster than you can read them.
What We Actually Know About Each Box
The DGX Spark is a shipping product built on NVIDIA's GB10 Grace Blackwell platform, documented in NVIDIA's own hardware specifications. Reviewers have had units since October 2025. It launched at $3,999, and the price has since moved to $4,699 on NVIDIA's marketplace, a figure the IntuitionLabs review covers alongside independent performance testing. Ollama publishes reproducible decode numbers for the Spark on open-weight models, roughly 41 to 58 tokens per second depending on the model. You can look those up before you spend money. You can also wait in line; supply has been constrained and street pricing floats above MSRP.
The AI Cube, meanwhile, has no price, no configuration page, and no checkout button. Every performance figure traces back to Xiaomi's own presentation. Nobody outside the company has run a model on it, because nobody outside the company has one. Industry trackers cited in the video project commercial availability in 2027.
The NYU Shanghai RITS analysis explains how the specs split across a "three-chip prototype" consisting of the O3, O100, and D100 processors, which is how a bandwidth claim and a memory claim from different dies got welded into one imaginary machine. Hardware Corner's coverage of the near-memory bandwidth design adds context on what Xiaomi is actually attempting.
To be fair to the debunking crowd: within a day of the reveal, independent commentators had flagged the patchwork. The correction just lost the race. A bold number on a slide travels faster than a paragraph explaining which chip the number belongs to.
Even a Benchmark Might Not Settle It
Suppose someone magics a Cube onto a test bench tomorrow. A single tokens-per-second number still wouldn't settle much, because inference speed is partly a software property. Run a big open model on the Spark with Ollama and you get one decode rate; swap in vLLM on the same physical machine and throughput changes dramatically. The video cites a swing of up to 2.7x from runtime choice alone.
A fair fight would require: identical model, identical quantization (how aggressively the weights were compressed to fit in memory), identical inference runtime, and measured decode throughput on hardware you can actually buy. The Cube currently blanks every line of that scorecard. There is no public software stack to run the test in.
The Manufacturing Question Nobody's Asking
The deeper issue is whether the silicon can be mass-produced at all. Reaching that bandwidth without standard high-bandwidth memory requires wafer-on-wafer packaging, an exotic process that's hard to yield at scale even with the best equipment. Getting reliable yields from advanced stacking typically leans on EUV lithography machines, which export controls keep off the table for Chinese fabs. Xiaomi's partners would need to invent a workaround for something the rest of the industry struggles with while holding the best tools. That drags the 1.22 TB/s figure from "spec" toward "science experiment."
I want to be clear about what I'm not saying: this isn't a story about Xiaomi being incapable. Near-memory compute is a legitimate and interesting answer to the bandwidth wall, and if anyone cracks cheap wafer stacking, the whole local inference market shifts. The Spark's known weakness is real.
Where This Leaves Builders
If you're deploying local models this quarter, the choice isn't between two desktop accelerators. It's between a shipping product with published benchmarks, a price, and a supply chain, and a prototype that won't be purchasable for at least two more years. The Spark has a known ceiling and you can plan around it; there's a whole ecosystem forming around these boxes, from dual-Spark clusters punching above their price class to NVIDIA's larger GB300 desktop running agentic workloads.
The Cube is a fascinating preview of where the hardware goes. It replaces nothing you can't yet order, and the internet calling the market obsolete in the meantime is measuring a marketing department, not a machine.
The next time a bandwidth figure goes viral, the first question isn't "how fast is it?" It's "which chip was that number even about?"
Yuki Okonkwo covers AI and machine learning for Buzzrag.
More Like This
Running a 405B AI Model at Home: Hardware and Security
Two NVIDIA DGX Spark units, one cable, and an open-source firewall. Here's what it actually takes to run a 405B AI model on your desk.
Splitting One LLM Across Two Machines: Does It Actually Work?
Alex Ziskind tested disaggregated inference by combining a DGX Spark and Mac Studio to run LLMs. Here's what actually happened when theory met reality.
How a 26B AI Model Now Runs in 2GB of RAM on a Mac
A 26-billion-parameter model running in ~2GB of active RAM on a MacBook isn't magic. It's two independent timelines finally crashing into each other.
DiffusionGemma Generates Text Like an Image Model
Google DeepMind's DiffusionGemma borrows from image diffusion to generate 700–1,000+ tokens/sec. Here's how the architecture works—and where it falls short.
Dual DGX Spark Matches Pricier AI Clusters in Testing
Level1Techs tests a dual DGX Spark against a far more expensive RTX Pro 6000 cluster—and the results challenge assumptions about what local AI actually costs.
Clustering Two AMD Ryzen AI Halos to Run 400B Models
Can two AMD Ryzen AI Halos act as one AI system? Alex Ziskind tested the cluster setup, performance, and real-world limits of AMD's 400B parameter claim.
Gemma 4 12B Brings Local Agentic AI to Laptops
Google's Gemma 4 12B is a multimodal local AI model built for real agentic workflows on 16GB laptops—here's what the architecture actually means.
Does AI Understand Things, or Just Predict Words?
The "AI just predicts tokens" argument is technically true—but is it the whole story? A murder mystery with fake physics might hold the answer.
RAG·vector embedding
2026-09-05This article is indexed as a 1536-dimensional vector for semantic retrieval. Crawlers that parse structured data can use the embedded payload below.