local AI inference
10 stories tagged local AI inference.
NVIDIA's PAIR Turns Idle Home GPUs Into a Local AI Pool
NVIDIA's open-source Personal AI Router pools Ollama and LM Studio nodes across your network. What PAIR does well, and what it leaves unanswered.
Framework Desktop Gets Two R9700 GPUs via PCIe Bridge
Framework Desktop Gets Two R9700 GPUs via PCIe Bridge
Level1Techs strapped two ASUS R9700 GPUs to a Framework Desktop using a Broadcom PLX bridge, hitting 192GB total memory for local AI workloads. Here's what worked.
Perplexity Open Sources Lily, a Local AI Engine for Apple Silicon
Perplexity Open Sources Lily, a Local AI Engine for Apple Silicon
Perplexity has open-sourced Lily, a Rust-based inference engine for Apple M5 Max. Here's what the benchmarks mean and why the OSS move matters.
FreeToken vs llama.cpp: A Local AI Engine Reality Check
FreeToken vs llama.cpp: A Local AI Engine Reality Check
UC Berkeley's FreeToken claims to run 753B parameter MoE models on a single GPU. Here's what the benchmarks actually show—and what they quietly obscure.
Minisforum N5 Max Review: NAS Meets AI Workstation
Minisforum N5 Max Review: NAS Meets AI Workstation
The Minisforum N5 Max stuffs a five-bay NAS, AMD Ryzen AI Max+ 395, and dual 10GbE into one box. Here's what that actually means in practice.
Intel Arc Pro B70: 32GB of VRAM at a Real Price
Intel Arc Pro B70: 32GB of VRAM at a Real Price
Intel's Arc Pro B70 offers 32GB of VRAM for $1,000 and free SR-IOV support. For homelabbers and AI tinkerers, it's a serious buy. Gamers should wait.
Diffusion Gemma Runs Locally—and That Changes Privacy
Diffusion Gemma Runs Locally—and That Changes Privacy
Google's Diffusion Gemma runs on consumer GPUs at 700+ tokens/sec. For privacy, the real story isn't speed—it's that your prompts never leave your machine.
Llama.cpp Gets MTP: Local AI Just Got Faster
Llama.cpp Gets MTP: Local AI Just Got Faster
Llama.cpp just merged Multi-Token Prediction, giving local AI a ~25% speed boost. Here's why that matters for your privacy—and how to use it.
Desktop AI Supercomputers: What Dell's GB10 Says About Tech
Desktop AI Supercomputers: What Dell's GB10 Says About Tech
Dell's Pro Max with GB10 brings Nvidia's Blackwell chips to your desk. But who needs a 1 petaflop AI workstation at home, and what does it signal about computing's future?
Intel Arc Pro B60: Testing 96GB of AI VRAM for $5K
Intel Arc Pro B60: Testing 96GB of AI VRAM for $5K
Level1Techs tests Intel's Battle Matrix with four Arc Pro B60 GPUs—96GB VRAM for the price of an RTX 5090. Real-world AI performance examined.