Posognos evaluates publicly available models from public endpoints and publishes the field, whether or not a developer engages. What engagement buys is earlier knowledge, per-scenario failure detail, and evaluation of the configuration you actually ship. Never a better grade.
Heading into a hospital’s due-diligence review? A current rating answers the AI-validation section for you: bring the report, not a promise ↓
492 expert-authored medication-order scenarios across 11 safety categories, three runs per scenario at temperature 0.2, majority vote. Evaluations ran 21–23 April 2026.
The single most important caveat on this page. These models were evaluated bare: no retrieval, no drug knowledge base, no worked examples. Nobody deploys a model that way, and both First Databank and Wolters Kluwer now ship grounding services built precisely to sit underneath these systems [3].
A result here measures the substrate, not a product. If you ship a model with a formulary-aware retrieval layer and tuned thresholds, that is a different artefact and it is evaluated separately. See substrate versus configuration below.
Sorted by Youden’s J, sensitivity plus specificity minus one. J is used as the primary sort because it is the only headline statistic a model cannot inflate by alerting on everything: a system that fires on every order scores exactly zero. Grade bands are explained on the grading page [4].
| Grade | Provider | Model | Youden J | Sens % | Spec % | Alert rate % | PPV % | F1 % | Missed / 311 | False / 181 | Median s |
|---|
“Missed” = unsafe orders on which no alert was raised, out of 311. “False” = alerts raised on safe orders, out of 181. Source: preprint Tables 1 and 2 [1]. The paper reports no confidence intervals, so adjacent ranks are not statistically separable [2].
Three models in this edition reach 100% sensitivity: they catch every dangerous order in the benchmark. Two of them alert on more than 92% of all orders, and one on 97.8%. A system that alerts on everything is trivially perfect at detection and useless in a ward.
The share of unsafe orders on which an alert was raised. Necessary, and by itself meaningless: an always-alert classifier scores 100%.
The share of safe orders left alone. This is the number nobody publishes. It ranges from 6.1% to 81.8% across the field, and it is what override behaviour is made of.
Sensitivity plus specificity minus one. Zero for a model that alerts on everything, zero for one that alerts on nothing. The grade is built on it, with floors on both components.
A vendor may not present a substrate rating as though it applied to their product. Posognos publishes which class each rating is, and the class travels inside the mark. This protects you as much as the buyer: a bad substrate score does not describe a well-built product, and a good one does not excuse a poorly configured deployment.
Name the model version and the configuration. Substrate and config evaluations are separate products, priced and labelled differently.
~1 week492 scenarios, three runs each, against your endpoint or container. Synthetic patients only.
3–10 daysResults under embargo. You may contest any scenario as clinically wrong. A sustained challenge changes the benchmark for every model, not just your score.
30 business daysThe rating publishes with its date, version and class. Ship a fix and re-test at a fixed fee; the new rating supersedes the old one on the record.
45 days to re-testThe due-diligence use. Hospital procurement now asks every AI vendor the same question: how is this validated, and by whom? A current rating answers it with an artefact instead of an assertion: the dated report, the registry record the buyer can check without you in the room, and the mark carrying its class and edition inside the frame. One evaluation serves every security and clinical review until your model version changes.
What money buys. Pre-release evaluation under NDA. Per-scenario failure analysis with expert annotations. Regression testing across your release history. Evaluation of your deployed configuration rather than the bare model.
What it does not buy. A grade, a delay, a withdrawal, or any influence over how scenarios are authored, validated or scored. Fees are published in advance and are never contingent on a result. Declining to engage does not avoid a rating. It only means finding out at the same time as your customers.
Name the model, the configuration or the hospital build. You get the price, the calendar and the scenario battery in writing before anything runs. Fees are published in advance and never contingent on a result.