Loading use case index…
Loading use case index…
AI use case
Baichuan AI released Baichuan-M3, a new-generation medical-enhanced LLM that beats OpenAI's GPT-5.2 on HealthBench (+28 over M2 on HealthBench-Hard at 44.4), SCAN-bench (74.9 on…
Core facts from this catalog record. Primary narrative lives in the hero above; full raw fields follow in the next section.
Every column from the source row, in stable order. URLs open in a new tab.
Title
Baichuan-M3 Outperforms GPT-5.2 on HealthBench and SCAN-bench, Sets New SOTA for Medical LLMs
Content
Baichuan AI released Baichuan-M3, a new-generation medical-enhanced large language model that surpasses OpenAI's GPT-5.2 on clinical-reasoning benchmarks including HealthBench, HealthBench-Hard, and SCAN-bench, establishing a new state of the art in medical-enhanced language models. "Baichuan-M3 is trained to explicitly model the clinical decision-making process, aiming to improve usability and reliability in real-world medical practice" — Baichuan AI blog, Baichuan-M3 release. The model's design addresses a structural shift in healthcare AI: applications have moved from standardized benchmarks to real-world clinical decision-making, where models face incomplete information, open-ended questions, and high-stakes consequences. Baichuan-M3 is explicitly trained to proactively acquire critical clinical information, construct coherent medical reasoning pathways, and systematically constrain hallucination-prone behaviors — rather than producing "plausible-sounding answers, fluent doctor-like questioning, or high-frequency but vague recommendations" such as advising a patient to "seek medical attention as soon as possible." The release builds on Baichuan-M2 (August 2025), which had already achieved state-of-the-art on HealthBench among open-source models. In the months that followed, OpenAI released GPT-5.2 along with the ChatGPT Health product, signaling an industry-wide transition of AI toward real clinical deployment. Baichuan-M3's three-stage SCAN framework — Clinical Inquiry, Laboratory Testing, and Diagnosis — transforms single-turn question answering into transparent, step-by-step decision trajectories that better reflect real-world medical reasoning. On HealthBench-Hard, Baichuan-M3 reached a score of 44.4 — a 28-point improvement over Baichuan-M2 — and beat GPT-5.2 to rank first on the HealthBench Total leaderboard. On SCAN-bench, the model attained 74.9 in Clinical Inquiry (outperforming GPT-5.2-High by 12.4 points and exceeding the human baseline of 53.5), 72.1 in laboratory test recommendation, and 74.4 in final diagnosis — the only model to rank first across all three stations. The blog also reports lower hallucination rates than GPT-5.2 in a tool-free setting, validated by a framework that decomposes long-form responses into atomic medical claims and checks each against authoritative medical evidence. A demo case illustrates the clinical reasoning: a 10-year-old boy with two weeks of recurrent fever initially appeared "common cold/pneumonia-like," but Baichuan-M3's SCAN-driven history taking identified cross-system signals — lower-limb joint swelling, gait limitation, dysuria, oral ulcers — and ultimately diagnosed post-infectious reactive arthritis, stepping off the routine respiratory-infection pathway.
Continue exploring AI deployments in the catalog.
Back to use casesCity
Beijing
Company/Organization
Baichuan AI
Continent
Asia
Country
China
Category
Internet Software & Services
Type
Experiment
Id
ec1c1ff2-5718-42d1-8e9a-f64d4ec97758
Created At
2026-07-01T21:18:13.399448+00:00