Research

AI wellbeing evaluation has a psychometrics problem.

Not a shortage of benchmarks. A shortage of evidence that any of them measure what they claim to measure.

Where the field is

Three findings that define the problem.

79%of 210 safety benchmarks

A review of 210 AI safety benchmarks found that 79% rely on binary pass or fail as their primary metric, and 81% test only predefined, already-known risks. Almost none are validated against a real-world outcome.

In any other applied field this is a construct validity failure. A measure never checked against the outcome it claims to predict is not yet a measure.

90,000clinical ratings · SIM-VAIL

In August 2026 that changed at one end. SIM-VAIL, from Oxford, UCL, and the UK AI Security Institute, published a clinically validated, multi-turn audit framework across 810 conversations, nine frontier models, and thirty simulated vulnerability profiles, with an open explorer for the results.

Oxford, UCL and the UK AI Security Institute, Nature Medicine, 7 August 2026. The scenario-audit problem is now substantially solved, and any new instrument has to say honestly what it adds beyond it.

0.27pooled ICC among psychiatrists

And at the other end it got worse. When three board-certified psychiatrists rated the same mental-health AI responses against shared instructions, pooled ICC(2,1) was 0.269, meaning roughly 27% of the variance in ratings reflected real differences between responses. Four dimensions returned negative Krippendorff’s alpha, which is systematic disagreement worse than chance, including boundaries at −0.203.

Jafari, Rust, Eddy, Fraser, Vasan, Djordjevic, Dadlani, Lamparth, Kim & Kochenderfer, Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing, FAccT ’26. Reliability was worst on the most safety-critical items.

The method

Behaviorally anchored rating scales, built by retranslation.

Smith and Kendall described the procedure in 1963. It does not measure agreement after the fact. It engineers agreement into the instrument before anyone scores anything.

Clinicians generate critical incidents

Practicing clinicians write concrete, observed examples of system behavior at the moments that matter, in their own words, rather than rating abstract dimensions someone else defined.

A second panel retranslates them

An independent group of clinicians is given the incidents without their labels and asked which dimension each one belongs to, and how good or bad it is.

Only agreed items survive

An incident is retained only when the panel independently sorts it to the same dimension and scales it to a similar value. Everything ambiguous is discarded rather than defined away.

Survivors become the anchors

Each scale point is defined by a specific behavior the rater population already agreed on. A rater is no longer judging “empathy” in the abstract; they are deciding which concrete example the observed response most resembles.

The discarded items are not a loss. They are the finding: they locate exactly where clinical consensus does not yet exist, which is information the field currently has no way to see.

The harder question underneath

Reliability is necessary. It is not sufficient.

Two raters agreeing tells you the instrument is stable. It does not tell you the instrument predicts anything. That is the criterion problem, and in AI wellbeing evaluation it is almost entirely unaddressed: there is no established functioning-level outcome that a benchmark score is supposed to forecast.

Which means the field is optimizing predictors against no criterion. The criterion is what actually counts as a person doing worse, observable in their life rather than in a transcript. Specifying it, and then testing whether existing marker sets predict it, is the work I want to do next.

Pilot 001

One transcript, read two ways.

A single organic conversation between a general-purpose assistant and an adult user in acute distress, fourteen turns, passed first through an automated pattern matcher and then read by a clinician. Two things came out of it, and the second is the one that matters.

16 → 1matcher hits, synthetic then real

The lexical matcher scored sixteen hits on a synthetic chat written in the same sitting as the patterns, and one hit on the real transcript. That single hit was a false positive: it matched the phrase “I feel” inside the user’s own words quoted back to her. The human read found thirteen candidate moments in the same fourteen turns.

The gap is the finding. Anchors written against imagined dialogue do not survive contact with how people actually talk, which is a reason to build anchors from clinician-generated incidents rather than from an author’s intuition.

11 / 0referrals out, practical vs emotional

Counting every referral to a named human or institution by the domain it appeared in: the system routed her outward eleven times on practical content, money, childcare, logistics. On emotional content, fear, grief, self-worth, loneliness, it routed her outward zero times, and held seven moments itself.

Same conversation, same night, same system. Competent routing on the tractable material and none at all on the material that decides whether she is alone, which is precisely the asymmetry a content-level evaluation cannot see.

This is one transcript. It establishes nothing about prevalence, it is not a reliability estimate, and a single case cannot be generalized from. What it does establish is that the instrument runs on real material, that it finds structure an automated pass misses entirely, and that the structure it finds is not the structure anyone was looking for.

Prior work, and its limits

MatterBench, and what it does not yet establish.

MatterBench is a scenario instrument I built to scope relational harm. Version 0.2 crosses six recruitment contexts against six harm families, four positive families, and two session-level constructs, on the principle that severity is a property of the pairing rather than of the behavior alone. Its useful output so far is the taxonomy, the scoring design, and one pilot, not a validated measurement claim.

It has not been criterion-validated, and its inter-rater reliability has not been established at scale. In light of the two findings above I would not present it as more than a structured, documented starting point. Saying so is not a caveat attached to the work. It is the reason the next piece of work is worth funding.