DEMO. An unaffiliated design study built by Verzi for Posognos. Not the official Posognos site and not endorsed by Posognos. Model figures come from their published CC-BY preprint; every other number is cited inline.
For health systems and quality leaders

Your safety net gets tested once a year. Your formulary changes weekly.

The annual CPOE Evaluation Tool is the only empirical evidence most hospitals have that their ordering system catches what would harm a patient: one scored attempt, two hard-timed three-hour windows, a 120-day lockout. Posognos runs the same class of test continuously, against your configuration, with no clinical time, no PHI and no production system.

The lineage. The benchmark is authored by the team whose research built the CPOE evaluation instrument, with David Classen as co-senior author of the methods paper, and its scenarios are reviewed by clinical leaders from Penn Medicine, Duke and Intermountain (the team).

Selling AI into hospitals? Your model may already be on the record: see how 40 models score →

Scenario report SAMPLE HOSPITAL 1 · CONFIG
EPIC AUG-2026 BUILD · ISSUED 2026-08-14
POSOGNOSA2026.1
0.71Youden’s J 94.5%sensitivity 76.2%specificity 17 / 311unsafe missed
Drug-allergy19 / 19
Drug interaction59 / 60
Drug-laboratory29 / 34
Drug monitoring11 / 14
492 scenarios · no PHI · the full report ↓
Built for the instruments you already answer to
  • Leapfrog CPOE Toolannual
  • Leapfrog AI questions1 Apr 2026
  • Joint Commission RUAIH1 Jun 2026
  • SAFER 3.1 AI testingannual
  • CMS819 opioid harm eCQM1 Jan 2027
01 · What you already spend, and what you get for it

Quality measurement is already one of the largest administrative loads in the building.

The case for continuous testing is not a new budget line. It is that the existing one is enormous, manual, and concentrated in exactly the staff you can least afford to take off the floor, and that it buys you a single annual data point.

Academic centre

108,478 hours

Person-hours a year to report 162 externally required quality metrics, $5.04M in personnel plus $603K in vendor fees.

JAMA 2023 · activity year 2018 [8]
161-bed community

4.6 FTEs

And $708,691 a year on quality-reporting administration alone, over half of it clinical staff.

AHA Regulatory Overload, 2017 [9]
199-bed, first survey

117 hours

Across 12 people, mostly department heads, to prepare one Leapfrog Hospital Survey. A project manager was supplied on top of that, and the survey itself is free.

Erdmann et al., 2024 [10]
Single section

40 hours

Of that total, consumed by the medication reconciliation assessment, the largest single line item recorded.

Same study [10]

Eight measures your quality committee already owns

Left column is the annual CPOE Evaluation Tool as Leapfrog specifies it for the 2026 cycle. Right column is what continuous scenario testing changes. Nothing here replaces the scored attempt. It de-risks that attempt and fills the 364 days around it.

Measure
Today
With continuous testing
Scored attempts per cycleA retake requires a 120-day wait, so in practice there is one
1
Protected prescriber time consumedOrder entry cannot be delegated to a pharmacist, nurse or informaticist
2 × 3 hours, hard-timed
Interval between safety measurementsOrder sets, formulary and alert thresholds change continuously in between
365 days
Blind window per yearTime in which a regression in your CDS build is invisible to you
~364 days
Test orders exercisedThe scored tool uses roughly a dozen test patients and several dozen orders
Dozens
Patient data exposedScenarios are synthetic, authored by clinical pharmacology experts
Production environment
Over-alerting measuredOne of the tool’s eight categories covers alert appropriateness
1 of 8 categories
Evidence for AI in medication managementRequired in the Leapfrog survey from 1 April 2026; unscored, but recorded
Vendor attestation

Mechanics per the Leapfrog CPOE Evaluation Tool instructions for the 2026 cycle [1]. Leapfrog’s own documents differ between versions on the exact test-patient count, so it is given only approximately.

The honest framing. Continuous scenario testing does not eliminate the annual survey, the abstraction, or the attestation. What it removes is the part that consumes protected prescriber time to produce one annual data point, and it replaces the blind window between attempts with a measurement you can act on while there is still time to fix something.

02 · The configuration gap

The same drug database scores 94% on the bench and 61% in your building.

This is the finding that makes annual testing insufficient, and it comes from the same research lineage as the benchmark itself. A medication knowledge base, tested as its vendor designed it, catches 94% of the harmful test orders. The identical knowledge base, as actually implemented across 358 hospitals, caught 61.1% in 2017, 65.7% in 2018 and 68.8% in 2019.

Medication knowledge base performance, as designed and as deployed

Share of harmful test orders correctly flagged. The regression found that EHR vendor choice explained 0.1% of the variance between hospitals. The gap is local configuration: thresholds turned down, alert categories switched off, order sets that bypass checking.

Co Z, Strasberg HR, Hanway BR, Sittig DF, Classen DC. JAMIA Open, January 2026 [2]. Classen is co-senior author of this paper and of the PsiBench methods paper; the intellectual lineage is direct.

What this means operationally. Roughly a third of the safety performance a hospital believes it has bought is lost between the vendor’s bench and the ward. It is not lost to a bad product; it is lost to configuration, which drifts every time an order set is edited, a threshold is relaxed after a complaint, or a new formulary item lands without checking rules.

Configuration drifts continuously. Measurement happens annually. That gap is the entire argument for continuous testing, and it exists whether or not any AI is involved.

03 · Why it matters

Medication is where the preventable harm actually is.

43%of harm events among Medicare inpatients are medication-related, the single largest categoryHHS OIG, 2022 [3]
53%of preventable medication harm originates at the ordering stepWHO, 2024 [4]
~33%of harmful test orders are still missed by hospital ordering systemsClassen et al., JAMA Netw Open 2020 [5]
~90%pooled override rate for drug-interaction alerts across 570,776 prescriptionsFelisberto et al., 2024 [6]
$871M–1.8Bannual US cost of adverse drug events from inappropriately overridden alerts aloneSlight et al., JAMIA 2018 [7]
Overriding is not carelessness. It is what clinicians do to a system that is wrong most of the time it speaks.

The override rate is a specificity problem wearing a behaviour costume. A system that alerts on everything trains the people using it to dismiss everything, which is why measuring detection alone, without measuring restraint, produces safety theatre. It is also why the AI now entering medication workflows has to be measured on both.

04 · How it works

Three uses, none of which require an IT project.

Before the scored attempt

Rehearsal

Run the scenario battery against your configuration and find the categories you fail before the attempt that counts. One scored try per cycle and a 120-day lockout mean that the first time you learn your score should not be the time it goes on the record.

Between attempts

Regression surveillance

Re-run on every material change to order sets, formulary or alert thresholds. A suppression added after a complaint in March should not be discovered in next year’s survey.

On the AI you are buying

Vendor comparison

Independent scores on the models behind the tools on your shortlist, on identical scenarios and identical ground truth, plus evaluation of a specific vendor configuration before it goes live.

Two boundaries, stated up front. Scenario testing measures what a system does with a constructed order. It does not measure outcomes, and no patient has been treated on the basis of this benchmark. And a score against your configuration is not a Leapfrog score. Leapfrog administers its own instrument, and nothing here substitutes for it or predicts it.

05 · What you actually get

A dated, sourced report against your own configuration.

Not a dashboard login. A versioned document your quality committee can file, your CIO can attach to an attestation, and your counsel can read three years from now and still understand. Below is the summary page as it would look for one hospital build.

Medication-safety scenario report CONFIG · SAMPLE HOSPITAL 1 · EPIC AUG-2026 BUILD · ISSUED 2026-08-14
0.71Youden’s J, detection net of over-alerting
94.5%Sensitivity, unsafe orders flagged
76.2%Specificity, safe orders left alone
17Unsafe orders passed through, of 311
Detection by safety category · alert-required scenarios only, 311 in total · lowest four in red
Drug-allergy19 / 19
Drug age13 / 13
Drug interaction59 / 60
Daily drug dose40 / 41
Excessive dosing58 / 60
Single drug dose26 / 27
Drug route14 / 15
Drug-diagnosis25 / 28
Drug-laboratory29 / 34
Drug monitoring11 / 14

Two findings requiring action. Drug-laboratory checking missed 5 of 34 scenarios, all involving renal dose adjustment where the most recent creatinine was more than 72 hours old: the rule appears to require a same-encounter value. Drug monitoring missed 3 of 14, concentrated in narrow-therapeutic-index agents where no level had been ordered.

Change since last pass (2026-07-31). Specificity fell 4.1 points after the 4 August order-set revision. Two previously-firing duplicate-therapy checks are now suppressed on the oncology admission set. Nothing else moved.

492 scenarios · 10 of 11 categories carry alert-required orders · 3 runs each · synthetic patients, no PHI Method: doi:10.64898/2026.06.05.26354271 · Report v4 · supersedes 2026-07-31

Illustrative example. Figures are constructed to show the report format, not measured from any real hospital: the only measured results published anywhere are those for the 40 frontier models on the AI vendors page.

06 · What is coming

Three requirements landed this year. Four more are already dated.

None of them come with an instrument. Each one creates an obligation to demonstrate something about medication safety or AI performance, and each leaves the measuring to you.

In force now
1 April 2026

Leapfrog asks about your AI

Four questions on the AI applications used in medication management became required in subsection 2B, unscored and not publicly reported. Leapfrog’s stated purpose is to research “AI vendor influence on a hospital’s CPOE Test score.”

2026 Summary of Changes [11]
CY 2026

Annual AI testing, in writing

The Protect Patient Health Information objective requires a “Yes” attestation to annual self-assessment against all eight SAFER guides. Unscored, but a “No” fails the program. The CPOE guide, practice 3.1, states that “AI technology used for medication ordering is tested on an annual basis.”

Medicare Promoting Interoperability [12]
1 June 2026

Joint Commission RUAIH

Voluntary certification requiring a demonstrated process for monitoring and validating deployed AI safety performance, and stating plainly that it “does not validate or certify individual AI products or tools.”

Responsible Use of AI in Healthcare [13]
Dated and coming
~Nov 2026

Leapfrog’s 2027 proposals

Leapfrog’s CPOE Expert Panel has concluded non-interruptive alerts “have become less effective,” and has committed to updated expectations in the 2027 CPOE Tool. The only warning will be the autumn 2026 webinars and the public-comment document. Inventory every passive alert in your ordering path now.

Announced, not yet published [14]
Weeks away

HTI-5 final rule

At OMB since 28 July 2026. It proposes to delete the Decision Support Intervention source attributes (the 31 model-card fields) entirely. ASTP’s rationale: no evidence “a single doctor, nurse, or administrator has accessed… a single source attribute.” Your vendor’s obligation to hand you standardised AI transparency data is about to end.

Pending OMB review [15]
1 Jan 2027

Opioid-related harm becomes mandatory

Hospital Harm: Opioid-Related Adverse Events (CMS819) joins the mandatory eCQM set, taking it to six of nine. Payment-linked and publicly reported. Naloxone administration and opioid-related respiratory depression have to be measure-ready by 31 December 2026.

Finalised in FY2025 IPPS [16]
1 Jan 2027

Six more state AI laws

Colorado HB 1139, Georgia SB 444, Utah SB 319, Rhode Island H 7349, Oregon SB 1546 and Arizona board rules all take effect, mostly utilisation review and AI-use consent. Federal preemption has not materialised; the DOJ task force has so far intervened only in Colorado’s general AI act.

Effective dates set [17]

The pattern worth showing your board. Transparency obligations on vendors are being removed. HTI-5 deletes the model cards. Validation obligations on you are being added: RUAIH, SAFER, the state laws, Leapfrog’s AI questions. The evidence you used to be handed is becoming evidence you have to generate.

The practical move before your next EHR or CDS renewal: put the transparency requirements that HTI-5 is deleting into your procurement contract language in Q4 2026, and ask any vendor selling into Europe for their EU AI Act Annex IV technical documentation. That package is the same content, and it survives.

07 · Sources

Numbers without sources are marketing.

Show the 17 sources
  1. The Leapfrog Group, CPOE Evaluation Tool Instructions, 2026 cycle, and the 2026 Hospital Survey. Two timed three-hour blocks; the clock runs continuously including overnight; order entry by a prescriber who routinely orders through the inpatient CPOE, explicitly not a pharmacist, nurse or informaticist; production environment or an exactly-mirrored test environment; 120-day wait before a retake; standard is alerting on at least 60% of frequent serious medication errors.
  2. Co Z, Strasberg HR, Hanway BR, Sittig DF, Classen DC. JAMIA Open, 6 January 2026, 358 hospitals across four EHR vendors.
  3. HHS Office of Inspector General, Adverse Events in Hospitals: A Quarter of Medicare Patients Experienced Harm, OEI-06-18-00400, May 2022.
  4. World Health Organization, Global Burden of Preventable Medication-Related Harm, March 2024, 100 studies, 487,162 patients.
  5. Classen DC, Holmgren AJ, Co Z, et al. National Trends in the Safety Performance of Electronic Health Record Systems From 2009 to 2018. JAMA Network Open. 2020;3(5):e205547, 8,657 hospital-years.
  6. Felisberto M, et al. Health Informatics Journal, 2024, pooled override rate 90% (95% CI 85.6–95.0), 11 studies, 570,776 prescriptions. Individual studies range 65.3–97.3%.
  7. Slight SP, Seger DL, Franz C, et al. JAMIA. 2018;25(9):1183–88.
  8. Saraswathula A, Merck SJ, Bai G, et al. The Volume and Cost of Quality Metric Reporting. JAMA. 2023;329(21):1840–1847. Single academic medical centre; activity measured in calendar year 2018, costs inflated to 2022 USD.
  9. American Hospital Association, Regulatory Overload, October 2017. These figures are 2017-vintage and have not been superseded by a later edition.
  10. Erdmann M, Marshall C, Lay M. Transparency in Hospital Safety: A field study of ‘what it takes’ for hospitals to complete the Leapfrog Group Hospital Survey. Oklahoma State Medical Proceedings. 2024;8(1):216. Single hospital, first-time effort, April–June 2023; a project manager and technical support were provided at no cost, so an unaided hospital would bear more. The hospital ultimately did not submit.
  11. The Leapfrog Group, Summary of Changes to the 2026 Leapfrog Hospital Survey, effective 1 April 2026.
  12. CMS Medicare Promoting Interoperability Program, CY 2026 requirements. ONC/ASTP SAFER Guides, January 2025 revision, consolidated from nine guides to eight. Computerized Provider Order Entry with Decision Support, Domain 3, recommended practice 3.1 reads “Key metrics related to CPOE and CDS functionality are defined, monitored, and used to optimize safety and efficiency.” Two bullets in its implementation guidance are quoted on this page: “A CPOE evaluation tool (e.g., the Leapfrog Group’s CPOE ‘flight simulator’ for hospitals) is used annually on the production system to evaluate the safety and effectiveness of CPOE and CDS functionality,” and “AI technology used for medication ordering is tested on an annual basis.” Both verified verbatim against the guide PDF on healthit.gov, 22 August 2026. Note that the guide’s own cover numbers it 5 of 8 while the SAFER Guides index lists it sixth, so it is named here rather than numbered.
  13. The Joint Commission, Responsible Use of AI in Healthcare Certification, launched 1 June 2026. Voluntary; open to accredited and non-accredited organisations alike.
  14. The Leapfrog Group CPOE Expert Panel, multi-phase project on non-interruptive alerts, described in the 2026 Summary of Changes. No 2027 CPOE Tool guidance has been published as of 22 August 2026, the timing of the public-comment window is inferred from the 2026 cycle, which ran 18 November to 18 December 2025.
  15. ASTP/ONC, Health Data, Technology, and Interoperability: Deregulatory Actions to Unleash Prosperity (HTI-5), proposed 29 December 2025, 90 FR 60970, RIN 0955-AA09. Comments closed 27 February 2026; final rule received at OMB 28 July 2026 and still pending review. Proposes removing §170.315(b)(11)(iv)–(vi).
  16. CMS, Hospital Harm: Opioid-Related Adverse Events (CMS819), finalised in the FY2025 IPPS rule at 89 FR 69568–69573; the FY2027 IPPS final rule made no change to the CY2027 mandatory eCQM schedule.
  17. Effective 1 January 2027: Colorado HB 1139, Georgia SB 444, Utah SB 319, Rhode Island H 7349/S 2197, Oregon SB 1546, Arizona behavioural health board rules. Executive Order 14365 (December 2025) created a DOJ AI Litigation Task Force; as of August 2026 it has intervened in xAI v. Weiser against the general Colorado AI Act, not against a health-specific state law.

The next step is a scoping call, not a contract.

Name the model, the configuration or the hospital build. You get the price, the calendar and the scenario battery in writing before anything runs. Fees are published in advance and never contingent on a result.