ARC-Witness: DLC Games of ARC-AGI-3 [internal]

arxiv | github | blog

Click any run to watch its recorded games replayed in the browser — the board, the model's reasoning & plan & action, step by step.
compare puts two rollouts side by side on one timeline.

games
harnessharness changelogRHAE by·rank by

Model Evals

held-out games (composite rules) · evals of frontier & open-source models (our trained checkpoints are in RL & SFT experiments below)

runRHAE mean±stdgameslvlsavg_lvls
v5_fable5_20260717_01300640.9 ± 10.6 (k2)10/2072/1003.60/5compare
v5_glm52_20260716_15522313.8 ± 4.2 (k2)5/2066/1003.30/5compare
v5_gpt56sol-nr_20260719_17375425.0 ± 3.9 (k2)11/2080/1004.00/5compare
v5_gpt56sol_20260719_07243628.6 ± 7.3 (k2)8/2073/1003.65/5compare
v5_inkling-nr_20260719_1707153.6 ± 2.1 (k2)0/2034/1001.70/5compare
v5_inkling_20260719_0950123.3 ± 1.9 (k2)1/2030/1001.50/5compare
v5_kimik3_20260718_08295541.0 ± 4.3 (k2)12/2085/1004.25/5compare
v5_opus48_20260721_21431436.2 ± 7.4 (k2)11/2081/1004.05/5compare
v5_opus5_20260724_17503642.6 ± 3.3 (k2)10/2071/1003.55/5compare
v5_qwen35b_20260721_2143147.1 ± 2.5 (k2)3/2042/1002.10/5compare
v5_qwen37max_20260716_19091318.6 ± 6.0 (k2)5/2074/1003.70/5compare
v5_qwen9b_20260716_1350436.7 ± 1.7 (k2)5/2035/1001.75/5compare
v6.1_kimik3-max_20260722_06055540.2 ± 3.9 (k2)13/2091/1004.55/5compare
v6.1_opus48_20260722_05104437.7 ± 8.0 (k2)13/2089/1004.45/5compare
v6.2_qwen9b_20260722_1946542.5 ± 1.4 (k2)1/2026/1001.30/5compare
v6_inkling_20260722_0121282.9 ± 1.4 (k2)1/2030/1001.50/5compare
v6_kimik3-std_20260722_01212818.6 ± 2.7 (k2)9/2074/1003.70/5compare
v7.1_opus48_20260722_22531656.1 ± 8.7 (k2)15/2091/1004.55/5compare
v7.1_qwen9b_20260722_2253163.4 ± 3.0 (k2)3/2043/1002.15/5compare
v7.3_opus48_20260723_03291049.1 ± 5.5 (k2)10/2087/1004.35/5compare
v7.3_qwen9b_20260723_0329107.4 ± 4.0 (k2)4/2042/1002.10/5compare
v7_opus48_20260722_22531657.8 ± 9.4 (k2)14/2091/1004.55/5compare
v7_qwen9b_20260722_2253166.9 ± 3.2 (k2)3/2037/1001.85/5compare
v8_opus48_20260723_13353860.9 ± 8.7 (k2)11/1984/954.42/5compare
v8_opus5_20260724_17503658.5 ± 1.5 (k2)8/2075/1003.75/5compare
v8_qwen9b_20260723_13353813.7 ± 5.1 (k2)4/2040/1002.00/5compare

SFT Rollout Collections

training games (primitive rules only) · oracle-guided trace collections

runRHAE mean±stdgameslvlsavg_lvls
sft_fable_2026070141.3 ± 27.0 (k2)91/155608/7753.92/5compare
sft_opusv5_2026072036.4 ± 23.6 (k2)83/137556/6854.06/5compare

RL & SFT Experiments

our trained checkpoints (SFT/RL) evaluated with SAME harness — numbers on the held-out games are directly comparable with frontier/open-source rows above

runRHAE mean±stdgameslvlsavg_lvls
v6.2_qwen359b-sft-v6eff_202607226.8 ± 2.5 (k2)3/2045/1002.25/5wandb · compare
v6.2_qwen359b-sft-v8full_202607228.0 ± 1.6 (k2)1/2043/1002.15/5wandb · compare
v6.2_qwen359b-sft-v8resetaware_202607227.7 ± 2.1 (k2)5/2045/1002.25/5wandb · compare
v7.3_qwen359b-base-rl-s5_202607238.9 ± 6.5 (k2)5/2038/1001.90/5wandb · compare
v7.3_qwen359b-base_202607234.1 ± 3.8 (k2)4/2032/1001.60/5N/A · compare
v7.3_qwen359b-sft-v8resetaware-rl-s15_2026072313.4 ± 2.4 (k2)3/2047/1002.35/5wandb · compare
v7.3_qwen359b-sft-v8resetaware_2026072310.6 ± 2.7 (k2)1/2039/1001.95/5wandb · compare

HiScore — Human Leaderboard

play · record · rank — you are scored with the SAME official RHAE (k=2) as the models above, and every run you finish sharpens the empirical human-baseline steps behind the models' 'human' metric

runRHAE mean±stdgameslvlsavg_lvls
human_Guanghan_20260724_16210821.4 ± — (k2)1/48/202.00/5compare

RHAE (Relative Human Action Efficiency): per level, scoreL = min(115, (baselineL / actionsL)² × 100) if the level was completed, else 0; a game's score is the completion-capped, level-weighted mean of scoreL (official bench.scoring, cap = 5 levels). The cell shows the per-rollout mean over the bundle's games, mean ± std over seeds/rerolls, at the reporting convention k=2 (baselineL = 2 × BFS-optimal steps) — directly comparable to the published frontier tables.

The human metric re-scores the same runs with baselineL = average human steps from the HiScore runs (per level; levels no human has completed keep the k2 baseline).