on-device AI
12 stories tagged on-device AI.
MiniCPM5-2B: A 2.5B Model With 131K Context for On-Device AI
OpenBMB's MiniCPM5-2B averages 53.9 across 34 benchmarks and targets on-device use. What the model-card numbers show, and what they leave untested.
Same Model, 30% to 95%: Why the Harness Beats the Weights
Same Model, 30% to 95%: Why the Harness Beats the Weights
YC Paper Club argues the scaffolding around an LLM, not the weights, drives results. Prime Agent, OpenJarvis, and QM show how far the wrapper can go.
Superwhisper's S1-mini Cleans Up ASR Transcripts On-Device
Superwhisper's S1-mini Cleans Up ASR Transcripts On-Device
Superwhisper's S1-mini is a 462 MB open-weights model that strips fillers and fixes self-corrections in speech-to-text output—entirely on your device.
Apple's Neural Engine: Specialized Chips vs. Data Centers
Apple's Neural Engine: Specialized Chips vs. Data Centers
Apple's Neural Engine isn't an AI brain—it's a multiplication machine. Here's why that distinction matters for how businesses think about AI compute costs.
Gemini Nano Gets Faster on Pixel Without Retraining
Gemini Nano Gets Faster on Pixel Without Retraining
Google's frozen Multi-Token Prediction retrofits speed gains onto existing Gemini Nano models—no retraining needed. Here's what that means for on-device AI.
When Small AI Models Beat Frontier Ones on Your Tasks
When Small AI Models Beat Frontier Ones on Your Tasks
RL Nabors walks through a real eval framework for replacing frontier model calls with local SLMs—and the results are more nuanced than the pitch suggests.
Text Diffusion AI: Speed, Privacy, and Ambient Risk
Text Diffusion AI: Speed, Privacy, and Ambient Risk
Google DeepMind's text diffusion model generates AI responses differently—and faster. Here's what that architectural shift means for privacy and everyday users.
Google's Gemma 4 Makes Powerful AI Run on Your Phone
Google's Gemma 4 Makes Powerful AI Run on Your Phone
Gemma 4 brings multimodal AI models to phones and laptops with clever architecture tricks that make 5B parameters perform like much larger models.
Google Just Made Running LLMs on Your Phone Actually Simple
Google Just Made Running LLMs on Your Phone Actually Simple
Google's AI Edge Gallery lets anyone run large language models locally on their phone—no developer account, no cloud, no data sharing. Here's what that means.
Google's Gemma 4: Running Frontier AI on Your Phone
Google's Gemma 4: Running Frontier AI on Your Phone
Google's Gemma 4 brings frontier-level AI to consumer devices. Free, open-source, and offline-capable—but does it deliver on the promise?
Quinn 3.5 Runs AI Models On Your Phone Without Internet
Quinn 3.5 Runs AI Models On Your Phone Without Internet
The Qwen 3.5 AI model runs entirely on your iPhone with zero internet connection. We tested how well local AI works when privacy actually matters.
Google's AI Edge: Revolution or Just Hype?
Google's AI Edge: Revolution or Just Hype?
Google AI Edge lets AI models run on phones sans cloud, sparking debates on privacy and performance.