Skip to main content
Artificial Intelligence

Medical AI Just Lost to a General Model

Incremental training of AI models in healthcare may not really move the needle.

Key points

  • A new study shows everyday AI models outperform specialized LLMs in medicine.
  • Incremental training only adds a fraction more to what frontier models already know.
  • Specialized AI may still matter for rare edge cases, but these are shrinking as general AI improves.
Image by Pexels from Pixabay.
Source: Image by Pexels from Pixabay.

There's a category of digital health that has attracted serious money and credibility. It's the medically enhanced AI model. The pitch is rather intuitive: Take a frontier model, add curated medical information, and you've built something physicians can trust in a way they can't trust a general-purpose chatbot. In fact, OpenEvidence raised hundreds of millions of dollars on that premise. UpToDate built its own AI layer on the same logic. The assumption here is that more medical knowledge should produce better medical intelligence. A study recently published in Nature Medicine suggests otherwise.

The Math Someone Should Do

First off, it's interesting to look at some of the basic math. The total corpus of biomedical literature collectively represents hundreds of billions of words. Frontier AI models train on trillions. The specialized model isn't adding medicine to an empty vessel; it's adding hundreds of billions of words to a system that has already absorbed trillions, in the areas of medical information, biology, chemistry, statistics, and pharmacology. So I start to wonder if this is more like a drop in the information bucket.

My back-of-the-envelope calculation puts the incremental knowledge these specialized tools add at somewhere around one-tenth of one percent of what a standard model already knows. The specialized layer may contribute something at the margins. What this study suggests—something counterintuitive—is that it's no longer contributing enough to matter.

Paying a Premium?

Researchers at NYU Langone compared OpenEvidence and UpToDate Expert AI against three frontier models that included GPT-5.2, Gemini 3.1 Pro, and Claude Opus 4.6. The models were challenged across medical licensing examinations, clinician-alignment benchmarks, and 100 real physician queries based on actual clinical practice. The results were reviewed blindly by practicing clinicians.

The frontier models won across all three categories. But there was more: The specialized clinical tools performed no better than Google Search AI Overview. That's the browser feature most users don't know exists, let alone pay for. Here's my takeaway: A purpose-built clinical AI, marketed to physicians and priced accordingly, is performing at parity with a browser default.

We've Seen This Before

Medicine isn't the first field to make this bet, and the precedent isn't very encouraging. In 2023, Bloomberg invested heavily in a specialized financial model called BloombergGPT and was trained on billions of tokens of proprietary market data. The rationale was nearly identical to the clinical AI argument and suggested that finance was too specialized and consequential for general models to master. Despite access to an extraordinary volume of proprietary information, BloombergGPT performed comparably to general-purpose models on financial tasks.

Where Value Goes Next

The real question isn't if medical expertise matters; it does. The question is where value resides when general intelligence becomes broadly capable of handling most of what expert models were supposed to dominate. If frontier models continue to match or exceed specialized clinical AI, the competitive advantage can shift elsewhere. The new utility and competitive differentiation may shift to other areas including proprietary clinical data, workflow integration, institutional trust, governance, regulatory expertise, and the hard-won ability to deploy inside real healthcare environments. In the final analysis, the model itself becomes infrastructure as the value moves up the stack, toward the things that fine-tuning a frontier model just can't do.

A Small, but Important, Exception

It's important to note that the study's authors were honest about what their findings don't cover. Highly specialized tasks may still benefit from domain-specific approaches, and a single, even obscure clinical fact can be decisive in the right case. Those edge cases are real. They're also becoming a smaller part of the story. Healthcare AI built its identity around the belief that clinical complexity demanded clinical specialization. What the evidence now suggests is that the specialized layer matters less than assumed because the foundation beneath it has become extraordinarily capable. The moat was real. It just wasn't permanent.

advertisement
More from John Nosta
More from Psychology Today