of FDA-authorized medical-AI devices carry any pediatric labeling.ref: Brewster et al., 2024 — 149 of 876 FDA-authorized AI/ML devices explicitly labeled for children
are specifically pediatric — yet clinical AI is already answering questions about children.ref: FDA AI-enabled device analysis, 2026 — 5 of 952 devices exclusively pediatric (~0.5%)
That whitespace is what the research addresses.basis: existing oversight (accreditation & governance marks) evaluates process, not the clinical accuracy of outputs
The same team operates CodeBlue Peds, a production clinical-AI system that delivers pediatric protocol guidance to clinicians over SMS. It is built under the same doctrine we evaluate by: assume the system fails, measure the dangerous failure mode, and enforce safety in code rather than prompts — engineered so the AI never states a drug dose, with a confident-wrong retrieval rate of roughly one percent on an independent held-out test.
Building one clinical AI, and living with its failure modes, is what makes the evaluation credible. See how CodeBlue Peds is built →
Every evaluation item is anchored to a provenance-tiered evidence corpus and scored in one of three zones. The zone is a property of the claim, not the topic — the same disease, even the same trial, can generate items in different zones.
Every item is anchored to a tiered evidence corpus — from graded consensus guidelines through landmark trials down to expert opinion. Where sources conflict, published tie-breaking rules resolve them in a fixed order. No item is scorable unless every load-bearing claim traces to a tiered source.
The three-zone framework scores each claim by the kind of certainty medicine actually supports. An evaluation that pretends every question has one answer mis-scores exactly the cases that matter most.
Gold-standard items are independently labeled by practicing pediatric clinicians and adjudicated to consensus. The method advances only when inter-rater reliability clears a pre-registered statistical bar. The full methodology is being prepared for publication and peer review before any evaluation is offered.
Neutrality is the research's core asset, protected structurally rather than by promise: standard-setting is separated from any commercial relationship; an advisory board holds no equity and serves staggered terms; conflict-of-interest and recusal rules are published; and the standard has a public, versioned revision process. The methodology is designed to be inspectable, not proprietary judgment.
Early work on how to detect what LLM-generated pediatric content gets wrong — fabricated citations, safety-critical omissions, and the limits of cross-model review. Preprints and datasets are public; manuscripts under review are listed without venue.
Omit-Then-Verify: A Deployed Defense Against Fabricated Citations in LLM-Generated Medical Content.
Hsu B · 2026 · under review
Externally Grounded Validation of LLM-Generated Pediatric Clinical Pathways: A Reproducible Pipeline and a Cross-Model Review Safety-Omission Case Analysis.
Hsu B · 2026 · under review
Cross-Model Adversarial Review Misses Safety-Critical Omissions in LLM-Generated Pediatric Pathways: A Reproducible, Fail-Closed Validation Pipeline.
Hsu B · Zenodo, 2026 · doi:10.5281/zenodo.21286951 ↗
Omit-Then-Verify: A Deployed Defense Against Fabricated Citations in LLM-Generated Medical Content.
Hsu B · Zenodo, 2026 · doi:10.5281/zenodo.21141544 ↗
Citation-Verification Audit: Reproducibility Deposit for "Omit-Then-Verify" (CodeBluePediatrics).
Hsu B · Zenodo, 2026 · doi:10.5281/zenodo.21185113 ↗
We're assembling a clinical working group — pediatric critical care, hospital medicine, EM — and speaking with health systems interested in becoming early design partners once the methodology is published.