Cognition's Devin Deploys More Like Consulting Than Code
Cognition's Jia Wu argues AI deployment is closer to consulting than software. Here's what that means for engineering teams and how they measure it.
Written by AI. Bob Reynolds

Photo: AI. Hayden Cross
The job title says engineer. The actual work is closer to a management consultant who happens to write code — or, increasingly, direct an agent that writes code for them. That's the model Cognition is building around its AI agent Devin, and Jia Wu, Cognition's deployed engineering lead, laid out its logic at a recent AI engineering conference in enough detail to make the implications clear.
The argument goes like this: an AI agent pointed at a codebase without strategic direction is just burning compute. The expensive part was never writing the code — Wu puts that at roughly 20% of the real problem. The hard parts are testing it, reviewing it, deploying it, and keeping it working at scale. Pointing Devin at those problems intelligently requires someone who understands the customer's business deeply enough to know which problems are actually worth solving. That's not an engineering job in any traditional sense. It's a diagnostic job.
"Those four or five hours of calls actually allow us to understand very deeply what strategic initiatives are the highest leverage for the business," Wu said, describing a workday split between customer meetings and hands-on technical work. The ratio is roughly equal. If you came up through software expecting to spend most of your time in a terminal, this role would feel strange.
The Venn Diagram Problem
Wu uses a two-circle diagram to explain the core challenge. One circle is Cognition's product capabilities. The other is the customer's problem space. The forward-deployed team's job is to maximize the overlap. That framing sounds obvious until you follow it to where it leads.
Most AI tool vendors are content to widen their circle and hope customers find the overlap themselves. Cognition's bet is that the overlap doesn't find itself — that someone has to sit inside the customer's environment, learn the shape of their specific problems, and manually map them to what the agent can actually do. That's why Wu describes deployment challenges as recurring "in similar shapes" across companies. The deployed engineers are, in effect, building a pattern library: here's what enterprise COBOL migration looks like, here's what ETL modernization looks like, here's where the agent gets stuck. Each deployment is supposed to make the next one cheaper.
The feedback loop is the product, or at least half of it. "The feedback is actually like half of the loop that makes the next deployment better than the previous deployment," Wu said. Which means the deployed engineers aren't just delivering value to customers — they're generating the training signal that sharpens Devin's capabilities over time. The customer relationship and the R&D function are the same function.
What "Outcomes" Actually Means
The piece of Wu's talk that deserves the most scrutiny is the measurement framework. Most vendors in this space track token usage — how much the agent is running. Wu's argument is that this is the wrong unit entirely. You can run an agent constantly and ship nothing useful. The right unit is delta: what did the customer's delivery look like before Devin was fully activated, and what does it look like after?
Cognition claims an 82% reduction in delivery timelines across the work they targeted. That figure is measured before Devin lands and again once it's fully embedded — not against some theoretical baseline, but against what the same team was actually producing before. Wu is explicit about the methodology: every sprint metric, every ticket completion rate, everything measured before the deployment serves as the control.
The number is Cognition's. It comes from Wu's conference talk, not from independent audit. That matters. Vendors presenting their own case studies at engineering conferences occupy a specific epistemological position — they have every incentive to select the best deployments and present the best readings of those deployments. The measurement methodology Wu describes is more rigorous than "look at our token counts," but rigor in design doesn't guarantee rigor in execution.
The named case studies are suggestive if you take them at face value. Wu says a Nubank ETL migration that originally had 50 engineers attached was delivered by Devin autonomously in a fraction of the expected timeline. A separate Latin American bank, Wu claims, completed a tax identification system migration — COBOL and JCL, the languages that enterprises can't easily staff anymore — at significantly reduced effort. Built, a fintech, reportedly saw PR output multiply by an order of magnitude on a weekly basis.
These are the kinds of numbers that either represent a genuine step-change in what enterprise software delivery looks like, or they represent cherry-picked conditions and generous accounting. What they can't tell you, presented this way, is which.
The Job That Doesn't Have a Name Yet
The more durable part of Wu's talk is the hiring discussion, because it reveals something true about where enterprise AI deployment actually lives right now. Forward-deployed engineering at Cognition is "so flexible in terms of how you actually define it," Wu acknowledged. Sales engineer? Solutions architect? Technical account manager? None of those fully apply.
What Cognition is hiring for is a T-shaped person with an unusual profile: technically deep enough to understand what the agent can and can't do, business-literate enough to identify which customer problems have the highest leverage, and communicative enough to carry field intelligence back into the product roadmap in a way that product teams can act on. Wu is explicit that you can teach business acumen, but you can't easily teach technical depth on the job. The spike has to be there.
The more interesting observation is what this implies about the direction of travel. Wu said it plainly: "If the cost of software engineering is going to zero, you actually need to know how to design a product that makes sense." The engineering skill that survives is not typing faster. It's judgment — about what to build, in what order, for which customer, pointed at which problem. That's a different kind of expertise than the industry has historically rewarded.
There's also an honest accounting problem running through the whole talk. Wu said the question of how to measure ROI on AI deployment doesn't have a clean answer — that whoever cracks it will own a significant position in the market. That's not a solved problem at Cognition either. The 82% figure is their best current answer, and the methodology is more serious than most competitors offer, but the underlying difficulty of measuring knowledge work output hasn't gone away. It's been deferred.
What Forward Deployment Is Really Selling
Strip away the product specifics and what Cognition is selling is a services wrapper around a software product. The agent is the technology; the deployed engineers are the delivery mechanism and the quality control and the feedback channel all at once. That's not a criticism — it may be the only honest way to deploy AI into complex enterprise environments right now, where the gap between what a capable model can do in isolation and what it can do inside a real organization's legacy infrastructure is enormous.
The model has precedent. Palantir built a dominant enterprise position doing essentially this with data infrastructure. The technology was real; the forward-deployed teams were what made it actually work inside large organizations that couldn't operationalize it themselves. The debate about Palantir was never whether the tech worked in controlled conditions. It was about whether the white-glove deployment model scaled, and at what margin.
Wu doesn't claim Cognition has fully solved that scaling question. What he does claim is that each deployment makes the next one structurally easier, because the challenges repeat and the pattern library grows. If that's true, the model compounds. If the human-intensive deployment layer never shrinks, the economics look different.
The honest version of the case Wu made in that room is this: Cognition has built a real agent, developed a deployment methodology more rigorous than most, and is generating field data that feeds back into product development in real time. The numbers they're reporting are vendor numbers, so hold them loosely. The job they're describing — part engineer, part consultant, part product manager — is real, and it probably represents where enterprise AI deployment lives for the next several years, whether Cognition captures most of that market or not.
The question nobody at the conference asked, and that the numbers can't answer, is whether "we'll make it cheaper next time" actually happens at the pace the model requires.
Bob Reynolds is Senior Technology Correspondent at BuzzRAG.
AI Moves Fast. We Keep You Current.
Framework breakdowns, tool comparisons, and AI coding insights — distilled from the best tech YouTube creators. Free, weekly.
More Like This
When Your AI Agent Should Actually Be a Workflow
Most AI 'agents' should be workflows instead. A technical workshop reveals why autonomy isn't always better—and how to choose the right architecture.
A Factory Built Its Own AI Brain Without Data Scientists
Rushabh Doshi built a 36-agent AI system to run Machinecraft's sales ops—no data team, no custom model training. Here's what he actually built, and what it can't tell you.
Automattic's 30-Day AI Experiment Changed How Designers Work
Automattic paused its product roadmap for 30 days and let teams build freely with AI tools. What a designer shipped — and what it means for how software gets made.
Garry Tan's Blueprint for the AI-Native Company
Y Combinator's Garry Tan argues that how you organize AI matters more than which AI you use. Here's what that means in practice.
Ramp's Leo Mehr on Scoping and AI in Engineering
Ramp's Leo Mehr argues that disciplined scoping and AI automation are both essential for enterprise engineering teams—and why neither works without the other.
AI Coding Tools and the Developer Skills Trade-Off
New research shows AI coding tools may improve productivity while quietly eroding developer skills. What does the data actually say—and who should be worried?
When Walmart Sells Last-Gen GPUs Cheaper Than Amazon
A PC build experiment reveals an uncomfortable truth about 2026 hardware markets: sometimes the discount bin beats the cutting edge.
Intel's $199 Chip Outperforms AMD's $500 Flagship
Intel's Core Ultra 250K at $199 matches or beats AMD's $500+ 9950X in real-world creative workloads. The benchmarks tell an unexpected story.
RAG·vector embedding
2026-07-29This article is indexed as a 1536-dimensional vector for semantic retrieval. Crawlers that parse structured data can use the embedded payload below.