▸ OPEN RESEARCH · IN DEVELOPMENT

Is the AI right about the child?

Healthcare AI is being deployed into pediatrics with almost no pediatric-specific evaluation behind it. Existing oversight asks whether an organization governs its AI responsibly. We're studying a different, unanswered question:

How should anyone test whether the AI's clinical outputs are actually correct?

STATUS: in development · methodology preprint targeted fall 2026 · no evaluations issued
THE WHITESPACE
~17%

of FDA-authorized medical-AI devices carry any pediatric labeling.ref: Brewster et al., 2024 — 149 of 876 FDA-authorized AI/ML devices explicitly labeled for children

<1%

are specifically pediatric — yet clinical AI is already answering questions about children.ref: FDA AI-enabled device analysis, 2026 — 5 of 952 devices exclusively pediatric (~0.5%)

No independent, pediatric-specific standard for evaluating AI clinical content exists.

That whitespace is what the research addresses.basis: existing oversight (accreditation & governance marks) evaluates process, not the clinical accuracy of outputs

WHY WE CAN BUILD THIS TEST

This standard is not written from the sidelines.

The same team operates CodeBlue Peds, a production clinical-AI system that delivers pediatric protocol guidance to clinicians over SMS. It is built under the same doctrine we evaluate by: assume the system fails, measure the dangerous failure mode, and enforce safety in code rather than prompts — engineered so the AI never states a drug dose, with a confident-wrong retrieval rate of roughly one percent on an independent held-out test.

Building one clinical AI, and living with its failure modes, is what makes the evaluation credible. See how CodeBlue Peds is built →

THE CORE IDEA · A THREE-ZONE FRAMEWORK

Medicine doesn't have one answer to every question. An honest test can't pretend it does.

Every evaluation item is anchored to a provenance-tiered evidence corpus and scored in one of three zones. The zone is a property of the claim, not the topic — the same disease, even the same trial, can generate items in different zones.

ZONE 1

Black-letter

one answer

A single guideline-anchored correct answer exists. The output is scored concordant or discordant.

Catches: confident, guideline-violating advice — e.g. a three-hour window for empiric antibiotics in septic shock, when the guideline says one.ref: Surviving Sepsis Campaign pediatric guidelines (Weiss et al., 2020)

ZONE 2

Bounded envelope

a range

A range of acceptable answers exists with explicit, citable bounds — typically titration and dosing ranges conditional on setting.

Catches: over-precise false certainty and out-of-range recommendations; an in-range number that omits a required stop-condition is flagged, not passed.

ZONE 3

Genuine equipoise

honestly unsettled

The literature contains documented, unresolved disagreement. The output is scored on reasoning and disclosure — not which conclusion it reaches.

Catches: falsely confident resolution of open questions — e.g. declaring one vasoactive agent "proven correct" where guidelines endorse either.ref: guideline equipoise on first-line vasoactive choice; no comparative trial establishes superiority

THREE COMMITMENTS
01

Provenance before verdict

Every item is anchored to a tiered evidence corpus — from graded consensus guidelines through landmark trials down to expert opinion. Where sources conflict, published tie-breaking rules resolve them in a fixed order. No item is scorable unless every load-bearing claim traces to a tiered source.

02

Honesty about knowledge

The three-zone framework scores each claim by the kind of certainty medicine actually supports. An evaluation that pretends every question has one answer mis-scores exactly the cases that matter most.

03

Measurement before claims

Gold-standard items are independently labeled by practicing pediatric clinicians and adjudicated to consensus. The method advances only when inter-rater reliability clears a pre-registered statistical bar. The full methodology is being prepared for publication and peer review before any evaluation is offered.

THE WEDGE · PEDIATRIC SEPTIC SHOCK

The first domain is deliberately narrow. Pediatric septic shock is high-stakes and guideline-dense, and it exhibits all three scoring zones within a single disease — settled resuscitation fundamentals, dosing questions with defensible setting-dependent ranges, and genuine post-FEAST equipoise around fluid strategy. If the framework holds there, it generalizes. A method proven on one demanding condition, published and replicated, is worth more than an untested method claimed to cover everything.

FEAST: Maitland et al., NEJM 2011 · resuscitation fundamentals: Surviving Sepsis Campaign pediatric guidelines (Weiss et al., 2020)

HOW INDEPENDENCE IS PROTECTED

Neutrality is the research's core asset, protected structurally rather than by promise: standard-setting is separated from any commercial relationship; an advisory board holds no equity and serves staggered terms; conflict-of-interest and recusal rules are published; and the standard has a public, versioned revision process. The methodology is designed to be inspectable, not proprietary judgment.

structural separation of standard-setting from revenue no-equity advisory board published COI & recusal rules public revision process
THE EXPLORATION, IN THE OPEN · PUBLICATIONS & PREPRINTS

The method is being built where it can be checked.

Early work on how to detect what LLM-generated pediatric content gets wrong — fabricated citations, safety-critical omissions, and the limits of cross-model review. Preprints and datasets are public; manuscripts under review are listed without venue.

SUBMITTED

Omit-Then-Verify: A Deployed Defense Against Fabricated Citations in LLM-Generated Medical Content.

Hsu B · 2026 · under review

SUBMITTED

Externally Grounded Validation of LLM-Generated Pediatric Clinical Pathways: A Reproducible Pipeline and a Cross-Model Review Safety-Omission Case Analysis.

Hsu B · 2026 · under review

PREPRINT

Cross-Model Adversarial Review Misses Safety-Critical Omissions in LLM-Generated Pediatric Pathways: A Reproducible, Fail-Closed Validation Pipeline.

Hsu B · Zenodo, 2026 · doi:10.5281/zenodo.21286951 ↗

PREPRINT

Omit-Then-Verify: A Deployed Defense Against Fabricated Citations in LLM-Generated Medical Content.

Hsu B · Zenodo, 2026 · doi:10.5281/zenodo.21141544 ↗

DATASET

Citation-Verification Audit: Reproducibility Deposit for "Omit-Then-Verify" (CodeBluePediatrics).

Hsu B · Zenodo, 2026 · doi:10.5281/zenodo.21185113 ↗

Get involved

Help build the test the field is missing.

We're assembling a clinical working group — pediatric critical care, hospital medicine, EM — and speaking with health systems interested in becoming early design partners once the methodology is published.

Join the clinical working group ← Back to Celeritas