The annual CPOE Evaluation Tool is the only empirical evidence most hospitals have that their ordering system catches what would harm a patient: one scored attempt, two hard-timed three-hour windows, a 120-day lockout. Posognos runs the same class of test continuously, against your configuration, with no clinical time, no PHI and no production system.
The lineage. The benchmark is authored by the team whose research built the CPOE evaluation instrument, with David Classen as co-senior author of the methods paper, and its scenarios are reviewed by clinical leaders from Penn Medicine, Duke and Intermountain (the team).
Selling AI into hospitals? Your model may already be on the record: see how 40 models score →
The case for continuous testing is not a new budget line. It is that the existing one is enormous, manual, and concentrated in exactly the staff you can least afford to take off the floor, and that it buys you a single annual data point.
Person-hours a year to report 162 externally required quality metrics, $5.04M in personnel plus $603K in vendor fees.
JAMA 2023 · activity year 2018 [8]And $708,691 a year on quality-reporting administration alone, over half of it clinical staff.
AHA Regulatory Overload, 2017 [9]Across 12 people, mostly department heads, to prepare one Leapfrog Hospital Survey. A project manager was supplied on top of that, and the survey itself is free.
Erdmann et al., 2024 [10]Of that total, consumed by the medication reconciliation assessment, the largest single line item recorded.
Same study [10]Left column is the annual CPOE Evaluation Tool as Leapfrog specifies it for the 2026 cycle. Right column is what continuous scenario testing changes. Nothing here replaces the scored attempt. It de-risks that attempt and fills the 364 days around it.
Mechanics per the Leapfrog CPOE Evaluation Tool instructions for the 2026 cycle [1]. Leapfrog’s own documents differ between versions on the exact test-patient count, so it is given only approximately.
The honest framing. Continuous scenario testing does not eliminate the annual survey, the abstraction, or the attestation. What it removes is the part that consumes protected prescriber time to produce one annual data point, and it replaces the blind window between attempts with a measurement you can act on while there is still time to fix something.
This is the finding that makes annual testing insufficient, and it comes from the same research lineage as the benchmark itself. A medication knowledge base, tested as its vendor designed it, catches 94% of the harmful test orders. The identical knowledge base, as actually implemented across 358 hospitals, caught 61.1% in 2017, 65.7% in 2018 and 68.8% in 2019.
Medication knowledge base performance, as designed and as deployed
Share of harmful test orders correctly flagged. The regression found that EHR vendor choice explained 0.1% of the variance between hospitals. The gap is local configuration: thresholds turned down, alert categories switched off, order sets that bypass checking.
What this means operationally. Roughly a third of the safety performance a hospital believes it has bought is lost between the vendor’s bench and the ward. It is not lost to a bad product; it is lost to configuration, which drifts every time an order set is edited, a threshold is relaxed after a complaint, or a new formulary item lands without checking rules.
Configuration drifts continuously. Measurement happens annually. That gap is the entire argument for continuous testing, and it exists whether or not any AI is involved.
The override rate is a specificity problem wearing a behaviour costume. A system that alerts on everything trains the people using it to dismiss everything, which is why measuring detection alone, without measuring restraint, produces safety theatre. It is also why the AI now entering medication workflows has to be measured on both.
Run the scenario battery against your configuration and find the categories you fail before the attempt that counts. One scored try per cycle and a 120-day lockout mean that the first time you learn your score should not be the time it goes on the record.
Re-run on every material change to order sets, formulary or alert thresholds. A suppression added after a complaint in March should not be discovered in next year’s survey.
Independent scores on the models behind the tools on your shortlist, on identical scenarios and identical ground truth, plus evaluation of a specific vendor configuration before it goes live.
Two boundaries, stated up front. Scenario testing measures what a system does with a constructed order. It does not measure outcomes, and no patient has been treated on the basis of this benchmark. And a score against your configuration is not a Leapfrog score. Leapfrog administers its own instrument, and nothing here substitutes for it or predicts it.
Not a dashboard login. A versioned document your quality committee can file, your CIO can attach to an attestation, and your counsel can read three years from now and still understand. Below is the summary page as it would look for one hospital build.
Two findings requiring action. Drug-laboratory checking missed 5 of 34 scenarios, all involving renal dose adjustment where the most recent creatinine was more than 72 hours old: the rule appears to require a same-encounter value. Drug monitoring missed 3 of 14, concentrated in narrow-therapeutic-index agents where no level had been ordered.
Change since last pass (2026-07-31). Specificity fell 4.1 points after the 4 August order-set revision. Two previously-firing duplicate-therapy checks are now suppressed on the oncology admission set. Nothing else moved.
Illustrative example. Figures are constructed to show the report format, not measured from any real hospital: the only measured results published anywhere are those for the 40 frontier models on the AI vendors page.
None of them come with an instrument. Each one creates an obligation to demonstrate something about medication safety or AI performance, and each leaves the measuring to you.
Four questions on the AI applications used in medication management became required in subsection 2B, unscored and not publicly reported. Leapfrog’s stated purpose is to research “AI vendor influence on a hospital’s CPOE Test score.”
2026 Summary of Changes [11]The Protect Patient Health Information objective requires a “Yes” attestation to annual self-assessment against all eight SAFER guides. Unscored, but a “No” fails the program. The CPOE guide, practice 3.1, states that “AI technology used for medication ordering is tested on an annual basis.”
Medicare Promoting Interoperability [12]Voluntary certification requiring a demonstrated process for monitoring and validating deployed AI safety performance, and stating plainly that it “does not validate or certify individual AI products or tools.”
Responsible Use of AI in Healthcare [13]Leapfrog’s CPOE Expert Panel has concluded non-interruptive alerts “have become less effective,” and has committed to updated expectations in the 2027 CPOE Tool. The only warning will be the autumn 2026 webinars and the public-comment document. Inventory every passive alert in your ordering path now.
Announced, not yet published [14]At OMB since 28 July 2026. It proposes to delete the Decision Support Intervention source attributes (the 31 model-card fields) entirely. ASTP’s rationale: no evidence “a single doctor, nurse, or administrator has accessed… a single source attribute.” Your vendor’s obligation to hand you standardised AI transparency data is about to end.
Pending OMB review [15]Hospital Harm: Opioid-Related Adverse Events (CMS819) joins the mandatory eCQM set, taking it to six of nine. Payment-linked and publicly reported. Naloxone administration and opioid-related respiratory depression have to be measure-ready by 31 December 2026.
Finalised in FY2025 IPPS [16]Colorado HB 1139, Georgia SB 444, Utah SB 319, Rhode Island H 7349, Oregon SB 1546 and Arizona board rules all take effect, mostly utilisation review and AI-use consent. Federal preemption has not materialised; the DOJ task force has so far intervened only in Colorado’s general AI act.
Effective dates set [17]The pattern worth showing your board. Transparency obligations on vendors are being removed. HTI-5 deletes the model cards. Validation obligations on you are being added: RUAIH, SAFER, the state laws, Leapfrog’s AI questions. The evidence you used to be handed is becoming evidence you have to generate.
The practical move before your next EHR or CDS renewal: put the transparency requirements that HTI-5 is deleting into your procurement contract language in Q4 2026, and ask any vendor selling into Europe for their EU AI Act Annex IV technical documentation. That package is the same content, and it survives.
Name the model, the configuration or the hospital build. You get the price, the calendar and the scenario battery in writing before anything runs. Fees are published in advance and never contingent on a result.