DEMO. An unaffiliated design study built by Verzi for Posognos. Not the official Posognos site and not endorsed by Posognos. Model figures come from their published CC-BY preprint; every other number is cited inline.
For AI developers and clinical AI vendors

Your model is already being measured. The question is who tells your buyer first.

Posognos evaluates publicly available models from public endpoints and publishes the field, whether or not a developer engages. What engagement buys is earlier knowledge, per-scenario failure detail, and evaluation of the configuration you actually ship. Never a better grade.

Heading into a hospital’s due-diligence review? A current rating answers the AI-validation section for you: bring the report, not a promise ↓

The field SUBSTRATE · PSIBENCH 2026.1
EVALUATED 21 TO 23 APRIL 2026
POSOGNOSA2026.1
40frontier models 59,040evaluations 9reach grade A 6.1–81.8%specificity range
Models earning each grade, of 40
public endpoints · no configuration · every mark resolves →
01 · Edition 2026.1 at a glance

Forty models, one benchmark, no configuration.

492 expert-authored medication-order scenarios across 11 safety categories, three runs per scenario at temperature 0.2, majority vote. Evaluations ran 21–23 April 2026.

The single most important caveat on this page. These models were evaluated bare: no retrieval, no drug knowledge base, no worked examples. Nobody deploys a model that way, and both First Databank and Wolters Kluwer now ship grounding services built precisely to sit underneath these systems [3].

A result here measures the substrate, not a product. If you ship a model with a formulary-aware retrieval layer and tuned thresholds, that is a different artefact and it is evaluated separately. See substrate versus configuration below.

02 · The field

Every model, every metric, sortable.

Sorted by Youden’s J, sensitivity plus specificity minus one. J is used as the primary sort because it is the only headline statistic a model cannot inflate by alerting on everything: a system that fires on every order scores exactly zero. Grade bands are explained on the grading page [4].

Grade Provider Model Youden J Sens % Spec % Alert rate % PPV % F1 % Missed / 311 False / 181 Median s

“Missed” = unsafe orders on which no alert was raised, out of 311. “False” = alerts raised on safe orders, out of 181. Source: preprint Tables 1 and 2 [1]. The paper reports no confidence intervals, so adjacent ranks are not statistically separable [2].

03 · Reading a result

What a high sensitivity score is worth on its own: nothing.

Three models in this edition reach 100% sensitivity: they catch every dangerous order in the benchmark. Two of them alert on more than 92% of all orders, and one on 97.8%. A system that alerts on everything is trivially perfect at detection and useless in a ward.

Sensitivity

Did it catch the danger?

The share of unsafe orders on which an alert was raised. Necessary, and by itself meaningless: an always-alert classifier scores 100%.

Specificity

Did it stay quiet when it should?

The share of safe orders left alone. This is the number nobody publishes. It ranges from 6.1% to 81.8% across the field, and it is what override behaviour is made of.

Youden’s J

Both, in one number

Sensitivity plus specificity minus one. Zero for a model that alerts on everything, zero for one that alerts on nothing. The grade is built on it, with floors on both components.

Sensitivity is not safety.
04 · Two rating classes

Substrate and configuration are different artefacts, and they are labelled differently.

SUBSTRATE, the bare model

  • Evaluated from public API endpoints, on Posognos’s own cadence
  • No retrieval, no drug knowledge base, no few-shot examples
  • Free, published, and produced whether or not the developer participates
  • Measures the raw reasoning of the model, which is a real and useful thing to know
  • Does not describe any product built on that model

CONFIG, a deployed configuration

  • Evaluated against a named endpoint or container, as you ship it
  • Retrieval stack, drug data source, thresholds and prompt scaffolding all named on the record
  • Synthetic patient scenarios only; no PHI moves in either direction
  • Pinned to a build and a date, and superseded when the build changes
  • This is the rating a hospital procurement committee is actually asking for

A vendor may not present a substrate rating as though it applied to their product. Posognos publishes which class each rating is, and the class travels inside the mark. This protects you as much as the buyer: a bad substrate score does not describe a well-built product, and a good one does not excuse a poorly configured deployment.

05 · The process

Four steps, and a review window that exists so the benchmark can be wrong.

Step 1

Scope

Name the model version and the configuration. Substrate and config evaluations are separate products, priced and labelled differently.

~1 week
Step 2

Evaluation

492 scenarios, three runs each, against your endpoint or container. Synthetic patients only.

3–10 days
Step 3

Review window

Results under embargo. You may contest any scenario as clinically wrong. A sustained challenge changes the benchmark for every model, not just your score.

30 business days
Step 4

Publication

The rating publishes with its date, version and class. Ship a fix and re-test at a fixed fee; the new rating supersedes the old one on the record.

45 days to re-test

The due-diligence use. Hospital procurement now asks every AI vendor the same question: how is this validated, and by whom? A current rating answers it with an artefact instead of an assertion: the dated report, the registry record the buyer can check without you in the room, and the mark carrying its class and edition inside the frame. One evaluation serves every security and clinical review until your model version changes.

What money buys. Pre-release evaluation under NDA. Per-scenario failure analysis with expert annotations. Regression testing across your release history. Evaluation of your deployed configuration rather than the bare model.

What it does not buy. A grade, a delay, a withdrawal, or any influence over how scenarios are authored, validated or scored. Fees are published in advance and are never contingent on a result. Declining to engage does not avoid a rating. It only means finding out at the same time as your customers.

06 · Sources

Numbers without sources are marketing.

Show the 4 sources
  1. Proulx J, Daines B, Barton M, et al. A Three-Tier Operational Benchmark for Evaluating Large Language Models on Hospital Medication Safety. medRxiv, posted 10 June 2026. doi:10.64898/2026.06.05.26354271, Tables 1–2, CC-BY 4.0. Not certified by peer review.
  2. Same source. The paper reports no confidence intervals, bootstrap estimates or significance tests. Differences of a point or two should not be read as a ranking.
  3. First Databank launched MedProof MCP in March 2026; Wolters Kluwer launched Medi-Span Expert AI in February 2026, both to ground third-party AI agents in curated medication data, including order validation.
  4. Grade thresholds shown here are an illustrative scheme constructed for this demo and applied to the published results. They are not Posognos policy. See the grading page.

The next step is a scoping call, not a contract.

Name the model, the configuration or the hospital build. You get the price, the calendar and the scenario battery in writing before anything runs. Fees are published in advance and never contingent on a result.