Gemini 3.8 Flash: Strong Benchmarks, Sharper Pricing
Google's Gemini 3.8 Flash posts competitive benchmark scores at a fraction of rival prices. Here's what the numbers actually mean for enterprises choosing AI models.
Written by AI. Yuki Okonkwo

Photo: AI. Atticus Ferenczi
Google dropped Gemini 3.8 Flash this week, and the headline number isn't the benchmark score. It's the price tag: $0.75 per million input tokens and $3.75 per million output tokens, per Google's announcement. Before you get too comfortable with that, those are introductory prices that expire at the end of the year, with the post-intro rates buried in fine print. Matthew Berman caught this in his breakdown of the release: "If you already plan on raising the price in the future, you should make that price the more prominent price." Fair point. Even at the post-intro rates of $1.50 input and $7.50 output, the model still undercuts comparable offerings from OpenAI and Anthropic, per Berman's breakdown.
The pricing context matters because 3.8 Flash isn't trying to be a frontier model. It's going after the mid-tier, the segment of deployments where cost per task shapes whether a product is viable at scale. And when you look at it through that lens, the benchmark picture gets interesting.
Where it actually wins
Berman calls DeepSWE his most trusted benchmark right now, and Gemini 3.8 Flash scores 73.7% on DeepSWE v1.1. That puts it effectively level with Claude Opus 5 and above GPT 5.6 Soul on the same benchmark, per Berman's breakdown. DeepSWE measures long-horizon software engineering tasks, the kind where a model has to maintain context across a codebase, trace bugs through multiple files, and reason about consequences before touching code. It correlates well with how working engineers actually experience these models day-to-day, which is why Berman weights it heavily.
The chart Berman highlights maps DeepSWE performance against average cost per task. Cost per task is a more honest metric than raw token pricing because it accounts for how many tokens a model burns to complete the same job. A cheaper model that requires twice the context isn't actually cheaper. On that combined view, 3.8 Flash sits in a strong position: high task performance, low task cost. That's the product.
On Harvey's legal benchmark, 3.8 Flash scores 61.4% and takes the top spot, with Gemini 3.7 Flash in second. For law firms evaluating AI tools, that combination of benchmark-leading legal performance and sub-dollar input pricing is an unusually direct argument. Berman's take: "If you're a law firm, you should definitely look at this model in particular because it scores so highly on the Harvey legal agent benchmark and also is very cheap."
On Humanity's Last Exam, it places first at 55.9%. On Terminal Bench 2.1 (agentic terminal coding), it scores 89.4% for first place. It also posts a 59% on OSWorld, which tests agentic computer use, the ability to navigate browsers and desktops autonomously.
Where it hits a wall
The Terminal Bench result deserves its own paragraph because it's confusing until you understand what happened. 3.8 Flash scores 89.4% on Terminal Bench 2.1 and 19.1% on Terminal Bench 4.0. That looks like a contradiction until you realize 2.1 was nearly saturated by top models; 4.0 is a substantially harder version of the benchmark. On 4.0, Claude Opus 5 scores 51% per Berman's breakdown, well above 3.8 Flash's 19.1%. The 3.8 Flash score lands above Sonnet 5 on the same benchmark, per Berman's breakdown, which is something, but the gap with Opus 5 suggests the model isn't the right call for the most demanding agentic coding tasks.
On GDPVal, a benchmark from OpenAI that tests real-world knowledge work like PDF extraction, data analysis, and presentation creation, 3.8 Flash scores 1545. That's a significant drop from the top performers in Berman's comparison. For knowledge-work-heavy workflows, that gap shows up in practice. Berman's web page generation tests, where he ran 3.8 Flash against the same prompts he'd used across other models, showed a similar pattern: solid on some outputs (a DJX Spark product page with correct Nvidia color branding), weak on others (the rubber duck company site was incoherent, and the Tesla Model Y came out looking like a Microsoft Paint sketch).
The standout from his tests was a 3D topographic map of Mount Everest: draggable, zoomable, with labeled climbing camps and sliders for vertical exaggeration and solar azimuth. Berman called it "an absolute winner." A Doom recreation built from a single prompt also held up. These are the kinds of outputs that make you recalibrate where the model actually sits.
The cyber variant
Google also released Gemini 3.8 Flash Cyber alongside the standard model, restricted to vetted security professionals via the Fairwind program. It scores 86.2% on the CyberGem benchmark, per Berman's breakdown, above GPT 5.5 Cyber (a model purpose-built for offensive and defensive security tasks) at 85.6%.
The decision to ship a restricted variant rather than apply safety guardrails to the base model and call it done reflects a specific calculus: Google is betting that useful cybersecurity capability belongs in the hands of defenders, and that locking it behind a vetting program is the workable middle path between suppressing the capability entirely and releasing it openly. Whether that vetting holds under adversarial pressure is an open question, but the architecture of the decision, separate model, gated access, formal program, signals that Google sees the dual-use risk as real enough to build institutional friction around it.
They also ran the model against an internal benchmark covering vulnerability discovery across 20 programming languages, since CyberGem only tests C and C++ code. 3.8 Flash Cyber posted a large jump over 3.7 Flash on that internal benchmark. Google didn't run competitor models against it, so that comparison stays opaque.
The right way to read all of this
The pattern across Gemini releases has been solid benchmarks that don't always translate to usage satisfaction, as noted in coverage of earlier Gemini Flash iterations. 3.8 Flash's DeepSWE score suggests the translation problem may be improving, but Berman's practical tests give a mixed read. Impressive on spatial and visualization tasks; middling on design-heavy generation; real gaps in knowledge work.
As The Verge notes, the model "works harder" than its predecessor, which also means it can cost more in token-heavy tasks. The introductory pricing obscures that until it doesn't.
Berman's framing for enterprise buyers is worth keeping: "It's not as simple as just saying, 'Okay, what's the latest from Anthropic? What's the latest from OpenAI?' You should actually look at the benchmarks and more importantly test them yourselves." The heterogeneous performance profile of 3.8 Flash makes that advice more than a platitude. A law firm and a data analysis pipeline and a coding agent operation are shopping for different things, and the model that wins on one may lose badly on another.
If it doesn't, the post-intro rates still make 3.8 Flash competitive in the mid-tier. Either way, the benchmark-to-usage gap that has followed Google's recent releases finally closes when developers start building on this one.
Yuki Okonkwo covers AI and machine learning for Buzzrag.
More Like This
AI Agents Promised to Do Your Work. They Can't Yet.
Wall Street lost $285B betting on AI agents that would replace SaaS tools. But the tech that triggered the panic still sleeps when you close your laptop.
Google's Gemma 4: Small Models, Big Performance Claims
Google releases Gemma 4, claiming frontier-level AI performance in models small enough for consumer hardware. The numbers look impressive. The questions remain.
AI Leaderboards Are Lying to You About State-of-the-Art
Bertrand Charpentier of Pruna AI makes the case that 'state-of-the-art' is a broken concept—and that efficiency belongs in the same sentence as quality.
Google's Gemini 3.1 Pro: Genius on Paper, Disaster in Practice
Gemini 3.1 Pro crushes benchmarks but fails at basic tasks. Developer Theo tests Google's 'smartest model ever' and finds a genius that can't follow instructions.
IBM Granite 4.2: Open Reasoning Models With an Agent Brain
IBM's Granite 4.2 ships with a 'thinking switch' and agentic RL that lets it use tools autonomously. Here's what that actually means—and why it matters.
Harvey Tenet Legal AI Model: Cool Tech, Thin Proof
Harvey's Tenet model applies async RL to long-horizon legal tasks—genuinely interesting tech. But verified benchmarks are scarce. Here's what we know and what we don't.
Gemma 4 12B Brings Local Agentic AI to Laptops
Google's Gemma 4 12B is a multimodal local AI model built for real agentic workflows on 16GB laptops—here's what the architecture actually means.
Does AI Understand Things, or Just Predict Words?
The "AI just predicts tokens" argument is technically true—but is it the whole story? A murder mystery with fake physics might hold the answer.
RAG·vector embedding
2026-09-04This article is indexed as a 1536-dimensional vector for semantic retrieval. Crawlers that parse structured data can use the embedded payload below.