held-out games (composite rules) · evals of frontier & open-source models (our trained checkpoints are in RL & SFT experiments below)
training games (primitive rules only) · oracle-guided trace collections
| run | RHAE mean±std | games | lvls | avg_lvls | |
|---|---|---|---|---|---|
| 2026-07-01_sft_fablegen_slurm | 41.3 ± 27.0 (k2) | 91/155 | 608/775 | 3.92/5 | compare |
| 2026-07-20_sft_opusv5gen_slurm | 37.9 ± 23.3 (k2) | 83/150 | 605/750 | 4.03/5 | compare |
our trained checkpoints (SFT/RL) evaluated with SAME harness — numbers on the held-out games are directly comparable with frontier/open-source rows above
play · record · rank — you are scored with the SAME official RHAE (k=2) as the models above. Runs played on the v2 (teach-first) level banks carry a v2 badge and are scored against the v2 baselines; only v1-bank runs feed the empirical human-baseline steps behind the models' 'human' metric
| run | RHAE mean±std | games | lvls | avg_lvls | |
|---|---|---|---|---|---|
| human_Guanghan_20260921_185537 | 90.3 ± — (k2) | 11/11 | 220/220 | 20.00/20 | compare |
| human_Guanghan_20260724_171042 | 83.9 ± — (k2) | 4/4 | 19/20 | 4.75/5 | compare |
| human_Guanghan_20260804_154531 | 81.9 ± — (k2) | 6/6 | 120/120 | 20.00/20 | compare |
| human_Guanghan_20260729_045730 | 75.1 ± — (k2) | 5/5 | 25/25 | 5.00/5 | compare |
| human_zomglings_20260806_070913 | 12.5 ± — (k2) | 1/1 | 3/5 | 3.00/5 | compare |
| human_huu_20260921_184949 | 1.9 ± — (k2) | 1/2 | 7/10 | 3.50/5 | compare |
| human_Deniz_20260921_184949 | 0.9 ± — (k2) | 0/1 | 2/5 | 2.00/5 | compare |
| human_allie_20260921_184949 | 0.6 ± — (k2) | 0/1 | 3/5 | 3.00/5 | compare |
RHAE (Relative Human Action Efficiency): per level,
scoreL = min(115, (baselineL / actionsL)² × 100) if the level was completed, else 0;
a game's score is the completion-capped, level-weighted mean of scoreL (official bench.scoring; cap = the column's level count: RHAE-L5 = first 5 levels, the runs-page standard and ranking key; RHAE-L20 = first 20 levels, deep execution).
The cell shows the per-rollout mean over the bundle's games, mean ± std over seeds/rerolls, at the reporting convention k=2 (baselineL = 2 × BFS-optimal steps) — directly comparable to the published frontier tables.
The human metric re-scores the same runs with baselineL = average human steps from the HiScore runs (per level; levels no human has completed keep the k2 baseline), at both caps — RHAE-L5 over the first 5 levels (the runs-page standard, which the perf rank uses) and RHAE-L20 over the first 20.