Pev Leaderboard

Accuracy, calibration and paired comparisons on Pev-Bench: seven families of short decision questions a personal agent with long-term memory has to answer (which option the user would pick, whether an action needs approval, whether a memory applies, whether an action relies on a forgotten fact, whether a fact may be shared, how urgently to notify, and where to route a request).

EnvLoop Research. The model (Pev-27B, a LoRA adapter on Qwen3.8-27B) and the dataset (Pev-Bench) are gated and released for non-commercial research only.

TEST leaderboard

The public TEST set has states from 720 new users and scorable questions. Half A () was rendered to text by gpt-6-astra and half B () by Claude Opus 5.5 under the same rendering contract. TEST was pseudonymized before any model was scored, so the released TEST is exactly the scored TEST. Every model was run once and all predictions were frozen before scoring.

The primary metric is family-macro accuracy: the mean of the seven per-family accuracies, with questions whose soft label has no unique answer excluded. The 6-family columns are the pre-registered sensitivity analysis without pick_option. Brier and ECE are family-macro averages; automation coverage is the share of questions answered alone at the threshold fitted on VAL for a 5% error budget.

* gpt-6-astra and Jev return probabilities through their APIs; these are used as returned (no temperature fitted on VAL), so their Brier and ECE are uncalibrated and they have no automation threshold (n/a). Accuracy does not depend on calibration. The local models are scored from label-token log-probabilities with a temperature fitted on VAL.

Paired comparisons

Pev-27B minus each other model, family-macro accuracy in points with a 95% confidence interval from a paired bootstrap over state clusters (10,000 resamples; clusters on all of TEST). p-values are one-sided and Holm-adjusted over the five comparisons, separately for each column; 5 × 10−4 is the smallest value attainable with 10,000 resamples and five comparisons.

Families

The sensitivity analysis without pick_option agrees with the primary result in every comparison and half (same sign, all significant after Holm correction).

Per family

HIDDEN one-shot gate

Before TEST existed, the frozen Pev-27B adapter was evaluated exactly once on a sealed, pre-registered HIDDEN set ( states, scorable of questions; 446 of the states use rendering styles never seen in training). The comparison is with the zero-shot base model. The run produced aggregate-only outputs and the sealed set was deleted afterwards. HIDDEN is not released.

DEV references

DEV (630 states, scorable questions) was the model-selection set for Pev-27B, so its DEV score carries selection bias; HIDDEN and TEST confirm it. These numbers were computed on the DEV text before pseudonymization; the released DEV differs only in person names, contact details and user ids and was not re-scored, so re-running on it may give slightly different numbers.

Caveats

Figures

From the technical report (CC-BY-4.0).

Grouped bar chart of TEST family-macro accuracy for six models on all of TEST, half A and half B. Pev-27B is highest in each group at about 0.91 to 0.92, followed by gpt-6-astra at about 0.87, Kev-27B about 0.82, Jev about 0.78 to 0.80, the Qwen3.8-27B base about 0.76 and Qwen3.5-4B about 0.63 to 0.68.
TEST family-macro accuracy per model, overall and per renderer half (A: gpt-6-astra, B: Claude Opus 5.5).
Dot plot with confidence intervals of Pev-27B minus each model on TEST, in points, for all of TEST, half A and half B. All intervals lie above zero: about 4 points versus gpt-6-astra, 9 versus Kev-27B, 12 to 13 versus Jev, 15 versus the Qwen3.8-27B base and 24 to 29 versus Qwen3.5-4B.
Pev-27B minus each model on TEST, family-macro accuracy with paired 95% confidence intervals, overall and per half.
Dumbbell chart of per-family accuracy on HIDDEN for the Qwen3.8-27B base and Pev-27B with chance marks. Gains in points: apply memory +15.8, forgotten violation +4.4, needs approval +15.3, share ok +16.7, route +8.1, notify level +42.2, pick option +3.3, macro +15.1.
Per-family accuracy on HIDDEN for the base model (orange) and Pev-27B (blue), with chance marks; labels give the change in points.
Scatter plot of automation coverage against realized error per family on HIDDEN at the VAL-fitted thresholds, with a horizontal line at the 5% error budget. Pev-27B covers all questions in six families with errors below the budget, but its pick option point sits at about 0.51 coverage and 0.45 realized error.
Coverage and realized error per family on HIDDEN at the VAL-fitted thresholds; the line is the 5% error budget.