Four Doctors of the AI Age: OpenEvidence Launches Osler, Sackett, Snow — and Keeps Darwin Locked Up
OpenEvidence ships a family of medical AI models named for medicine's greatest minds, claims the first perfect MedQA score with Darwin, and holds its most powerful model back over bioweapon-grade dual-use risk.
On September 3, 2026, medical AI company OpenEvidence did something unusual in the current race to ship ever-larger frontier models: it launched a family of models and deliberately kept the most capable one behind a locked door.
The company introduced four models named after giants in the history of medicine — Osler, Sackett, Snow, and Darwin. Three of them (Osler, Sackett, and Snow) are now rolling out to every OpenEvidence user, free and unlimited for verified clinicians. The fourth, Darwin, sits in research preview with access by application only. The reason, according to OpenEvidence, is dual-use risk: a model that can reason at the frontier of virology, immunology, and human genetics could, in the wrong hands, accelerate tightly governed research — bioweapons-relevant work and germline editing outside mainstream scientific oversight.
It is a notable moment for the AI industry, not just for medicine. At a time when most labs compete on openness and availability, OpenEvidence has effectively said: our best model is too good to give everyone.
The model family: speed, depth, and names with meaning
The three production models are differentiated not by intelligence but by how long they think and how deep they search:
- Osler (~5 seconds per answer) is the successor to the model that previously powered OpenEvidence answers, and becomes the platform’s default.
- Sackett (~30 seconds per answer) is positioned as the deeper search model, built for questions that turn on the weight of the evidence.
- Snow (~5 minutes per answer) succeeds the company’s Deep Consult feature and runs a full investigation of the medical literature before producing a report.
The names are a deliberate history lesson. William Osler moved medical teaching from the lecture hall to the bedside. David Sackett is remembered as the father of evidence-based medicine. And John Snow traced the 1854 Broad Street cholera outbreak to a single water pump, helping found modern epidemiology. Each name encodes what the model does: answer fast at the point of care (Osler), weigh evidence rigorously (Sackett), or run a deep epidemiological-style investigation (Snow).
The family is available on openevidence.com and the company’s iOS and Android apps, with a model selector letting clinicians pick the right tool for the moment. OpenEvidence says every model is held to the same standard of clinical accuracy; what differs is compute time and search depth.
Darwin: a perfect MedQA score, and a gated release
The headline numbers belong to Darwin. OpenEvidence describes it as the most advanced medical AI model in the world and says it is the first AI model to achieve a perfect score on MedQA, which the company calls the leading fully independent benchmark of medical AI.
On the companion benchmarks, Darwin is reported as state of the art on MedXpertQA at 72.8%, HealthBench Professional at 82.7%, and NOHARM at 87.2% — ahead of the next-best models, which OpenEvidence names as Anthropic’s Claude Fable 5 and Google’s Gemini 3.7.
The evaluation methodology is unusually transparent. For MedQA, OpenEvidence started from physician re-annotations of the benchmark’s 1,273-question test split, applied exclusion criteria for missing information, ambiguity, and label errors, and ran a second review pass in August 2026 in which three OpenEvidence physicians examined every question any evaluated model answered incorrectly. That produced a final set of 660 questions — which Darwin answered without error. The company is releasing its annotations and Darwin’s complete responses for all four benchmarks.
The baselines were run as claude-fable-5 with adaptive thinking, gpt-5.6-sol at default reasoning effort, and gemini-3.7-flash at default reasoning effort, using each provider’s API defaults with no tools or customized system prompts. Darwin’s MedXpertQA score sits 7.7 points above the strongest baseline; its HealthBench Professional score is 12.1 points above the next-best model; and its NOHARM severity-weighted F1 reached 0.872 against 0.740 for the closest baseline.
One honest wrinkle: Darwin’s unweighted precision on NOHARM is lower than several baselines. The company explains that the metric counts every recommendation absent from the rubric as a false positive, and that Darwin is deliberately tuned to give physicians the full set of relevant options rather than a minimal safe subset. Its severity-weighted precision is reported at 0.9.
The dual-use decision
The most consequential part of the announcement may be what was not shipped. Darwin’s access list covers institutional partners such as the National Organization for Rare Disorders (which brings Darwin its hardest cases), research collaborators, and accredited AI researchers at academic institutions benchmarking the accuracy and safety of AI in clinical medicine. Everyone else must apply.
The stated rationale is that frontier-level reasoning in virology, immunology, and human genetics is exactly the kind of capability that needs governance before distribution. OpenEvidence says Darwin’s capabilities will flow down into Osler, Sackett, and Snow as its safeguards are validated with these partners — a staged-release model that echoes the “capability-gated” deployment frameworks frontier labs have discussed in safety plans but rarely operationalized this explicitly for a domain-specific model.
The context matters. Frontier-general models like GPT-5.6, Claude Fable 5, and Gemini 3.7 are available through ordinary API access, and their biological-reasoning capabilities are the subject of ongoing biosecurity scrutiny from researchers and regulators. OpenEvidence has taken the opposite default for its domain-specific frontier model: closed until proven safe, rather than open until proven dangerous.
Why this matters for medicine
OpenEvidence has built a real foothold among US clinicians. The platform’s citation-driven answers are widely used at the point of care, and the company has been deepening its push into specialty practice — oncology decision support in particular, with precision knowledge bases for cancer care.
The new model family is a bet that clinicians don’t want one model, they want the right model for each moment: Osler for the quick drug-interaction check between patients, Sackett for the diagnostic puzzle where the literature has to be weighed, Snow for the referral report that needs a full literature investigation. Specialty-specific versions are next, starting with oncology, radiology, and clinical genetics.
For the broader AI industry, the launch is a data point in an unresolved debate: can a lab be both frontier-class and deliberately constrained? OpenEvidence is claiming it can — publishing full benchmark responses and annotations for scrutiny while gating the model itself. Whether Darwin’s safeguards can be validated fast enough to satisfy clinical partners, and whether “application only” holds as competitive pressure builds, will be worth watching.
The namesakes would have appreciated the sequencing: Osler taught at the bedside, Sackett demanded evidence, Snow mapped the outbreak. Darwin, fittingly, is still evolving — behind a door the company says it will open carefully.
Sources
- [1] https://www.businesswire.com/news/home/20260903878517/en/Introducing-the-OpenEvidence-Model-Family
- [2] https://www.unite.ai/openevidence-launches-medical-ai-model-family-with-darwin-preview/
- [3] https://www.statnews.com/2026/09/03/openevidence-launches-new-ai-models-clinicians-health-tech/
- [4] https://www.openevidence.com/user-guide/models