A Posognos rating reports how a model performed on 492 expert-authored medication-order scenarios [1]: how much danger it caught, how many safe orders it interrupted, and whether it flagged the right clinical reason. It is a measurement of performance on a published benchmark. It is not a safety certification, and nothing on this page should be read as one.
Ratings are dated, version-pinned, independently verifiable, and withdrawn when the model changes.
One rule of red, and nothing else. The red bar is the only brand element. It survives a 34-pixel favicon, a black-and-white print and a screenshot, and it never fights the palette of whoever hosts the mark. The edition, class and issue date sit inside the frame, not in a caption beside it. A rating without them is not a rating.
A rectangular plaque, not a circular seal. Circles read as seals of state; plaques read as findings, and this is a finding. The mark sheds elements in a fixed order as it shrinks, but the frame, the red rule and the grade never go. At 34 pixels those three are the whole mark, and they are still recognisably yours.
Grades are assigned on Youden’s J, sensitivity plus specificity minus one, with floors on both components. J is used because it is the only headline statistic a model cannot inflate by alerting on everything: a system that fires on every order scores exactly zero. Counts are the 40 models in the current edition.
Why floors, and not just a score. Three models in this edition reach 100% sensitivity: they catch every dangerous order in the benchmark. Two of them alert on more than 92% of all orders, and one on 97.8%. Without a specificity floor, perfect detection and reflexive alarm are indistinguishable, and the grade would reward exactly the behaviour that produced a 90% override rate in rule-based systems.
Each scenario presents a structured patient vignette (demographics, active diagnoses, current medications, laboratory values and allergies) plus one medication order. The model decides whether the order proceeds or an alert fires, and names the safety categories involved.
50 fatal orders against 48 deception scenarios, safe orders deliberately built to look dangerous at a glance. An always-alert classifier scores a Youden’s J of exactly zero here. This is the tier caution cannot game.
261 serious unsafe orders against 133 safe ones, including 41 contexts where rule-based systems are known to over-alert, counted here as orders a good model should let through. This approximates the alert-fatigue dynamics of a live ward.
For every correctly raised alert: was it raised for the right clinical reason? A model can flag the right order and attribute it to the wrong category, which, under the January 2026 FDA guidance conditioning enforcement discretion on a clinician’s ability to review the basis of a recommendation, is a regulatory question and not only a quality one [4].
What is not published yet. Per-model results exist for the full 492-scenario benchmark. The tier-level breakdown (the rankings that shift between discrimination, operational and attribution) is held for a second paper that has not been posted. Those columns are absent here because they are absent from the record, and this note stays until the clinical paper is public.
Every certification that has collapsed did so by letting its mark mean more than its method. These boundaries live in the licence, not only in the copy.
Modelled on how the marks that survived are actually governed [3]: UL, ENERGY STAR, the Joint Commission’s Gold Seal, URAC and Euro NCAP. Every one of these exists because a seal without governance becomes a logo, and a logo is worth nothing to the company that earned it.
Files are released to the rated party through the ratings portal at fixed sizes. There is no public asset kit and no editable master.
Ink for the frame and type, Posognos red for the single rule. Nothing else. No gradients, no outline versions, no third colour, and no version that matches its background. A one-ink black variant is supplied for fax, engraving and single-colour print, and it is equally valid.
Scale it; never redraw, restretch, recompose or partially use it. Clear space equals the cap height of the wordmark on all sides.
Below that, use the grade letter alone with the edition number, and link to the report.
The mark is not the claim. The dated report is the claim; the mark is a pointer to it.
They sit inside the plaque. Removing them voids the licence.
The mark may never stand alone as though it were the vendor’s own credential.
It may not be incorporated into a company name, product name, logo, domain or app icon.
One model version, one configuration, one indication set, one date. A substrate rating does not extend to a product built on it.
“Safe”, “approved”, “cleared”, “validated”, “certified safe”, “meets the highest standards”, “exceeds”. Permitted: “Rated [grade] by Posognos on PsiBench [edition], [date].”
On expiry, on withdrawal, or when the rated version is superseded in the vendor’s own product.
For misuse of the mark, misrepresentation of scope, or evidence that the evaluated version was not the version shipped.
The plaque on a vendor’s slide is a pointer. This is what it points at, and it is the only thing that constitutes the claim.
Name the model, the configuration or the hospital build. You get the price, the calendar and the scenario battery in writing before anything runs. Fees are published in advance and never contingent on a result.