Edited by humans. Written by AI. How our editing works
All articles

Vertical AI's Real Moat Is Judgment, Not Code

Ayush Bhardwaj built AI for a hedge fund, then a pharma startup—and hit the same wall both times. His diagnosis is more interesting than most AI talks.

Marcus Chen-Ramirez

Written by AI. Marcus Chen-Ramirez

August 20, 20267 min read
Share:
Man in glasses wearing light blue shirt speaking at AI Engineer World's Fair event, with statistics about AI adoption and…

Photo: AI. Hayden Cross

Ayush Bhardwaj had a moment of clarity at a hedge fund that most engineers working in specialized AI will recognize, even if they haven't named it yet. He could build the agent. He could watch it run. What he couldn't do was tell whether its output was any good—because he wasn't a trader. He didn't have the years of instinct that let a finance professional look at a trade thesis and know, in roughly the same way an engineer knows when generated code is lazy, that something was off.

He moved to a pharma tech startup expecting the problem to change. It didn't. Different industry, identical wall.

Bhardwaj presented this experience at the AI Engineer conference in a talk that's worth taking seriously—not because it's full of novel technical claims, but because it names something the AI industry keeps dancing around: the gap between building a vertical AI product and knowing whether it works is not an engineering problem. It's an epistemological one.

The setup is deceptively familiar. Pick a narrow task—not "find me market opportunities" but something granular, like ranking US IT stocks by capital expenditure and AI investment ratios. Secure proprietary data, because everyone already has arXiv and JP Morgan sell-side reports, and the model that only knows what's publicly available is competing with ChatGPT on ChatGPT's home turf. Write prompts that encode expert reasoning. Add observability. Bhardwaj is brisk about all of this: "The mythical 10x engineers can do this stuff in minutes. All of this fits one screen." The point isn't that these steps are trivial—it's that they're commodity. The interesting problem starts after step four.

Step five is iteration, and this is where the talk gets genuinely interesting. Bhardwaj's first instinct, when he couldn't evaluate his hedge fund agent's outputs, was to use a model as the judge. He describes this, to audible laughter from the room, as "a really, really stupid mistake." An LLM evaluating a trade thesis is, in his words, "jargoning its way out. It does not understand what alpha means." The critique lands because it's structurally precise: reinforcement learning through verifiable rewards works beautifully for math and code because those domains have answer keys. You know if the code compiles. You know if the proof checks. There's no equivalent ground truth for whether a drug candidate is worth pursuing, or whether a trade thesis has actual edge. The model can't verify itself in these fields—and errors don't stay local, they compound.

The data problem is even more structural than the evaluation problem. Bhardwaj walks through why the training data that would teach a model to reason in finance or pharma is, essentially, locked away by design. Institutional managers holding over $100 million in US equities are legally required to disclose their long positions quarterly—but those disclosures cost the funds real money, because competitors reverse-engineer the strategy the moment the filing goes public. The incentive to be vague, delayed, or minimally compliant is baked in. Pharma's version of this is darker: clinical trial disclosure is legally mandated, yet according to reporting from Final Call News, roughly 30% of firms never comply. The FDA, per a formal announcement on the agency's website, publicly called out more than 2,000 sponsors in 2026 for sitting on unfavorable results. Those failed trial results are precisely the data that would make a pharma AI useful—the signal that tells you what drug pathways are dead ends. And pharma companies guard them the way, as Bhardwaj puts it, you guard a chicken that lays golden eggs. OpenAI and Anthropic don't have this data either. Nobody does, except the organizations with strong financial reasons to keep it.

So the frontier models are trained on what's public, and what's public is the curated, disclosure-compliant, strategically sanitized surface of two industries that treat their real operational knowledge as competitive moat. The model is pattern-matching on the marketing materials, not the actual work.

Bhardwaj's answer to all of this is blunt: hire the user. For his pharma team—a group of young engineers who, by his own telling, needed convincing that they should hire a domain expert at all—this meant bringing in a senior scientist. The change wasn't cosmetic. When they pitched to major pharma companies afterward, the tools resonated. "They started liking our tools because it kind of spoke their language versus the normal jargonish LLM language." The domain expert's role isn't decorative; they curate data sources, refine prompts, and—critically—judge outputs. They bring the thing the engineers don't have: a trained mental model of what good looks like in this specific field.

The learning loop Bhardwaj describes builds out from there. Error analysis is the cheapest starting point—no weight updates required, just reading logs and understanding where the model misfires, then correcting it. From there you can climb toward reinforcement learning from human feedback, which he treats as the current gold standard for actually developing edge. The catch: every time a better base model drops, you're fine-tuning again. There's no finish line.

One element of Bhardwaj's framing that deserves more friction than it usually gets is his pivot on the human-in-the-loop question. The consensus position in AI product development is that HITL is table stakes—you add human review, you maintain oversight. Bhardwaj flips it: in finance and pharma, the AI is in the loop, not the human. The expert still makes the call. A trader generates five candidate trade theses with AI help and then decides which one to run. A scientist gets a ranked list of drug candidates from the model and then applies their judgment about which to pursue. This isn't autonomous AI decision-making—it's expert-time compression. "You just reduce the time of the expert by a lot."

That's a more honest description of what specialized AI actually does in high-stakes domains than most marketing materials will admit. Whether it scales is a different question.

He pulls in Yann LeCun's argument that current models are "text statistics, not real-world models"—they correlate, they don't cause. It's a useful frame, though LeCun is also a partisan in the ongoing war over whether large language models represent genuine intelligence or very sophisticated autocomplete, and treating his position as settled science would be too generous. The honest version is: nobody knows exactly where the ceiling is. Bhardwaj's point is narrower than LeCun's, and more defensible—that for finance and pharma, where causal reasoning about novel situations is the whole job, pattern-matching on historical text is insufficient. That seems right, and it doesn't require you to take a position on AGI timelines.

The talk closes on something that sounds like a conference platitude but isn't: the model and the infrastructure are commodities. Anyone can pay for them. "What is moat, and no one will come and sell it to you—you won't have to curate it on your own—is the domain expertise. You need your data. You need other people's data. That is just not out there on the internet."

Here's what strikes me about that claim. The industry has been selling the idea that AI moat lives in the model—in the training run, the parameter count, the benchmark score. Bhardwaj is arguing that in vertical applications, the model is table stakes. The moat is the judgment that tells you whether the model is doing the right thing, and the data that nobody else can access because it's locked inside organizations that have spent decades treating it as competitive advantage. That's not a technology story. That's an organizational and epistemic story, wearing technology's clothes.

The engineers who figure out how to acquire that judgment—by hiring for it, partnering for it, or building institutional relationships to access the data—will have something that can't be replicated by scaling up a foundation model. The ones who don't will keep shipping agents that look finished, then wonder why nobody buys them.


Marcus Chen-Ramirez is a senior technology correspondent for Buzzrag, covering AI, software development, and the intersection of technology and society.

More Like This

RAG·vector embedding

2026-08-20
1,803 tokens1536-dimmodel text-embedding-3-small

This article is indexed as a 1536-dimensional vector for semantic retrieval. Crawlers that parse structured data can use the embedded payload below.