TEST leaderboard
The public TEST set has states from 720 new users and scorable questions. Half A () was rendered to text by gpt-6-astra and half B () by Claude Opus 5.5 under the same rendering contract. TEST was pseudonymized before any model was scored, so the released TEST is exactly the scored TEST. Every model was run once and all predictions were frozen before scoring.
The primary metric is family-macro accuracy: the mean of the seven per-family accuracies, with questions whose soft label has no unique answer excluded. The 6-family columns are the pre-registered sensitivity analysis without pick_option. Brier and ECE are family-macro averages; automation coverage is the share of questions answered alone at the threshold fitted on VAL for a 5% error budget.
* gpt-6-astra and Jev return probabilities through their APIs; these are used as returned (no temperature fitted on VAL), so their Brier and ECE are uncalibrated and they have no automation threshold (n/a). Accuracy does not depend on calibration. The local models are scored from label-token log-probabilities with a temperature fitted on VAL.
Paired comparisons
Pev-27B minus each other model, family-macro accuracy in points with a 95% confidence interval from a paired bootstrap over state clusters (10,000 resamples; clusters on all of TEST). p-values are one-sided and Holm-adjusted over the five comparisons, separately for each column; 5 × 10−4 is the smallest value attainable with 10,000 resamples and five comparisons.
The sensitivity analysis without pick_option agrees with the primary result in every comparison and half (same sign, all significant after Holm correction).
Per family
pick_option progress is claimed.DEV references
DEV (630 states, scorable questions) was the model-selection set for Pev-27B, so its DEV score carries selection bias; HIDDEN and TEST confirm it. These numbers were computed on the DEV text before pseudonymization; the released DEV differs only in person names, contact details and user ids and was not re-scored, so re-running on it may give slightly different numbers.
Caveats
pick_optionis essentially unsolved. It is the one family labelled by real behaviour (an item the user later rated 4 or higher). All six models score 0.41–0.48 on TEST; a “priciest option” shortcut scores 0.347 against a chance level of 0.276, so nopick_optionprogress is claimed. Pev-27B’s VAL threshold also over-automates this family (realized error 37% on accepted TEST questions against the 5% budget): a deployment should always ask the user here.- Safety false negatives. Pev-27B has no
forgotten_violationfalse negatives on TEST, but gpt-6-astra is slightly lower on the other two safety families (0.7% / 0% / 0% against 0% / 1.3% / 2.0% forforgotten_violation/needs_approval/share_ok). Rates per model are below (all TEST users; about 150 questions at risk per family).
- The text is LLM-rendered. Every state is rendered from structured facts under verbatim anchor checks, and memory facts are LLM-extracted. User histories are real public reviews, but approval rules, privacy levels, contacts and calendars are synthetic. Real user text and real rules will be messier.
- Six of the seven families are rule-derived. High scores there show that the generator’s rules can be learned from rendered text, including buried and overridden evidence and an unseen renderer (TEST half B), not general judgement about approvals or privacy.
- Renderer familiarity. gpt-6-astra rendered the training, DEV and TEST-A text. On TEST no home advantage is visible: gpt-6-astra scores 0.876 on half A and 0.871 on half B, Pev-27B 0.913 and 0.917.
- Not a safety system. Approval and sharing predictions must be backed by hard rules in an agent; the model is a fast, calibrated advisor, not an enforcement layer.
Figures
From the technical report (CC-BY-4.0).



