ARC-Witness: DLC Games of ARC-AGI-3 [internal]

arxiv | github | blog

Click any run to watch its recorded games replayed: the board, the model's reasoning/plan/actions.
compare puts two rollouts side by side on one timeline.

games
harnessharness changelogRHAE by·rank by

Model Evals

held-out games (composite rules) · evals of frontier & open-source models (our trained checkpoints are in RL & SFT experiments below)

runRHAE-L5 mean±stdRHAE-L20 mean±stdgameslvlsavg_lvls
v5_fable5_20260717_01300640.9 ± 10.6 (k2)10/2072/1003.60/5compare
v5_glm52_20260716_15522313.8 ± 4.2 (k2)5/2066/1003.30/5compare
v5_gpt56sol-nr_20260719_17375425.0 ± 3.9 (k2)11/2080/1004.00/5compare
v5_gpt56sol_20260719_07243628.6 ± 7.3 (k2)8/2073/1003.65/5compare
v5_inkling-nr_20260719_1707153.6 ± 2.1 (k2)0/2034/1001.70/5compare
v5_inkling_20260719_0950123.3 ± 1.9 (k2)1/2030/1001.50/5compare
v5_kimik3_20260718_08295541.0 ± 4.3 (k2)12/2085/1004.25/5compare
v5_opus48_20260721_21431436.2 ± 7.4 (k2)11/2081/1004.05/5compare
v5_opus5_20260724_17503642.6 ± 3.3 (k2)10/2071/1003.55/5compare
v5_qwen3.5-35b_20260721_2143147.1 ± 2.5 (k2)3/2042/1002.10/5compare
v5_qwen3.5-9b_20260716_1350436.7 ± 1.7 (k2)5/2035/1001.75/5compare
v5_qwen37max_20260716_19091318.6 ± 6.0 (k2)5/2074/1003.70/5compare
v6.1_kimik3-max_20260722_06055540.2 ± 3.9 (k2)13/2091/1004.55/5compare
v6.1_opus48_20260722_05104437.7 ± 8.0 (k2)13/2089/1004.45/5compare
v6.2_qwen3.5-9b_20260722_1946542.5 ± 1.4 (k2)1/2026/1001.30/5compare
v6_inkling_20260722_0121282.9 ± 1.4 (k2)1/2030/1001.50/5compare
v6_kimik3-std_20260722_01212818.6 ± 2.7 (k2)9/2074/1003.70/5compare
v7.1_opus48_20260722_22531656.1 ± 8.7 (k2)15/2091/1004.55/5compare
v7.1_qwen3.5-9b_20260722_2253163.4 ± 3.0 (k2)3/2043/1002.15/5compare
v7.3_opus48_20260723_03291049.1 ± 5.5 (k2)10/2087/1004.35/5compare
v7.3_qwen3.5-9b_20260723_0329107.4 ± 4.0 (k2)4/2042/1002.10/5compare
v7_opus48_20260722_22531657.8 ± 9.4 (k2)14/2091/1004.55/5compare
v7_qwen3.5-9b_20260722_2253166.9 ± 3.2 (k2)3/2037/1001.85/5compare
v8_fable5-max64k_20260725_00183770.5 ± 5.8 (k2)11/2087/1004.35/5compare
v8_glm52_20260725_00420325.5 ± 7.8 (k2)1/2054/1002.70/5compare
v8_gpt56sol-max_20260725_04040758.2 ± 1.6 (k2)7/2072/1003.60/5compare
v8_inkling-max_20260725_00420315.3 ± 4.5 (k2)0/2038/1001.90/5compare
v8_kimik3-max64k_20260725_23424065.7 ± 6.4 (k2)13/2088/1004.40/5compare
v8_kimik3-max_20260725_09034471.0 ± 13.8 (k2)13/2090/1004.50/5compare
v8_opus48_20260723_13353860.5 ± 9.5 (k2)11/2087/1004.35/5compare
v8_opus5_20260724_17503658.5 ± 1.5 (k2)8/2075/1003.75/5compare
v8_qwen3.5-35b_20260725_00420314.1 ± 10.2 (k2)2/2036/1001.80/5compare
v8_qwen3.5-9b_20260723_13353813.7 ± 5.1 (k2)4/2040/1002.00/5compare
v8_qwen37max_20260725_00420341.1 ± 5.6 (k2)7/2077/1003.85/5compare
v9_dots3-promax_20260816_2053472.8 ± 1.4 (k2)0.2 ± 0.1 (k2)0/8554/17000.64/20compare
v9_ds4pro-promax_20260814_1723488.1 ± 2.9 (k2)0.6 ± 0.2 (k2)0/85132/17001.55/20compare
v9_fable5-max64k-pro_20260729_16331469.6 ± 13.2 (k2)17/25111/1254.44/5compare
v9_fable5-max64k_20260729_02485466.1 ± 5.9 (k2)11/2085/1004.25/5compare
v9_fable5anth-max64k-promax_20260730_21594135.0 ± 3.6 (k2)9.0 ± 3.6 (k2)2/85414/17004.87/20compare
v9_glim30-promax_20260811_0209323.3 ± 0.5 (k2)0.2 ± 0.0 (k2)0/8557/17000.67/20compare
v9_glm5.3-promax_20260821_0413194.5 ± 1.8 (k2)0.3 ± 0.1 (k2)0/8590/17001.06/20compare
v9_glm52-max-pro_20260729_1538007.9 ± 8.7 (k2)1/2552/1252.08/5compare
v9_glm52-max-promax_20260729_2022596.0 ± 1.6 (k2)0.4 ± 0.1 (k2)0/85109/17001.28/20compare
v9_glm52-max_20260729_05132927.8 ± 8.5 (k2)0/2052/1002.60/5compare
v9_gpt6-astra-max64k-promax_20260904_21555728.1 ± 1.8 (k2)13.2 ± 0.6 (k2)10/85389/17004.58/20compare
v9_gpt6-astrapro-max64k-promax_20260908_20242228.7 ± 2.7 (k2)12.0 ± 3.3 (k2)8/85352/17004.14/20compare
v9_gpt6-astrapro-max64k-promax_20260908_202422_resume— 0/2 complete— 0/2 complete4/31147/6204.74/20compare
v9_grok46-promax_20260812_20025613.9 ± 2.4 (k2)1.7 ± 1.2 (k2)1/85240/17002.82/20compare
v9_kimik3-max-pro_20260729_18110054.4 ± 9.8 (k2)17/25101/1254.04/5compare
v9_kimik3-max-promax_20260730_01492224.4 ± 4.3 (k2)3.3 ± 1.8 (k2)1/85249/17002.93/20compare
v9_kimik3-max_20260729_03090469.3 ± 9.1 (k2)13/2091/1004.55/5compare
v9_metaint-promax_20260811_19450415.0 ± 1.8 (k2)1.1 ± 0.1 (k2)0/85115/17001.35/20compare
v9_muse12-promax_20260811_17313516.2 ± 1.2 (k2)1.3 ± 0.2 (k2)0/85194/17002.28/20compare
v9_opus48-max64k-pro_20260729_18110042.8 ± 4.3 (k2)12/2597/1253.88/5compare
v9_opus48-max64k-promax_20260730_01492222.9 ± 3.1 (k2)1.8 ± 0.2 (k2)0/85203/17002.39/20compare
v9_opus48-max64k_20260729_09532271.5 ± 7.0 (k2)11/2089/1004.45/5compare
v9_opus5-pro_20260729_14121162.0 ± 7.6 (k2)16/25105/1254.20/5compare
v9_opus5-promax_20260729_19075833.0 ± 2.2 (k2)10.6 ± 2.2 (k2)2/85405/17004.76/20compare
v9_opus5_20260729_02341368.5 ± 7.0 (k2)12/2085/1004.25/5compare
v9_qwen3.5-35b-pro_20260729_1141146.3 ± 2.4 (k2)0/2536/1251.44/5compare
v9_qwen3.5-35b-promax_20260729_1852083.0 ± 1.3 (k2)0.2 ± 0.1 (k2)0/8558/17000.68/20compare
v9_qwen3.5-35b_20260729_02181213.4 ± 4.6 (k2)3/2040/1002.00/5compare
v9_qwen3.5-9b-pro_20260729_1205215.6 ± 5.8 (k2)0/2535/1251.40/5compare
v9_qwen3.5-9b-promax_20260729_1852081.3 ± 0.9 (k2)0.1 ± 0.1 (k2)0/8544/17000.52/20compare
v9_qwen3.5-9b_20260729_0218129.5 ± 5.9 (k2)4/2042/1002.10/5compare
v9_qwen3.6-27b-promax_20260819_1716535.8 ± 1.6 (k2)0.4 ± 0.1 (k2)0/85120/17001.41/20compare
v9_qwen3.8-2.4t-promax_20260814_2106307.0 ± 0.7 (k2)0.5 ± 0.1 (k2)0/85132/17001.55/20compare
v9_qwen3.8-27b-promax_20260816_2053476.2 ± 1.1 (k2)0.4 ± 0.1 (k2)0/85112/17001.32/20compare
v9_sol-promax_20260812_20172822.8 ± 2.2 (k2)3.6 ± 1.3 (k2)0/85262/17003.08/20compare

SFT Rollout Collections

training games (primitive rules only) · oracle-guided trace collections

runRHAE mean±stdgameslvlsavg_lvls
2026-07-01_sft_fablegen_slurm41.3 ± 27.0 (k2)91/155608/7753.92/5compare
2026-07-20_sft_opusv5gen_slurm37.9 ± 23.3 (k2)83/150605/7504.03/5compare

RL & SFT Experiments

our trained checkpoints (SFT/RL) evaluated with SAME harness — numbers on the held-out games are directly comparable with frontier/open-source rows above

runRHAE-L5 mean±stdRHAE-L20 mean±stdgameslvlsavg_lvls
tinker_eval_sft_v8resetawareB_step0_v921.0 ± 5.1 (k2)4/2058/1002.90/5compare
tinker_eval_v9_mix24_3827b_ts42_step09.1 ± 7.2 (k2)7/4083/2002.08/5compare
tinker_eval_v9_mix24_3827b_ts43_step08.2 ± 3.4 (k2)5/4080/2002.00/5compare
tinker_eval_v9_qwen3.5-9b-sftv8ra-step0-promax_20260806_1757426.0 ± 1.9 (k2)0.4 ± 0.1 (k2)0/170236/34001.39/20compare
tinker_eval_v9_rlonly3827b_ts42_step08.9 ± 6.2 (k2)4/4082/2002.05/5compare
tinker_eval_v9_rlonly3827b_ts43_step08.4 ± 5.3 (k2)6/4084/2002.10/5compare
tinker_eval_v9_s0_lvl10_v9_ts42_pick_s005-promax_20260810_1318465.2 ± 1.4 (k2)0.4 ± 0.1 (k2)3/170243/34001.43/20compare
tinker_eval_v9_s0_lvl10_v9_ts42_rl30-promax_20260809_0042064.5 ± 1.6 (k2)0.3 ± 0.1 (k2)4/170221/34001.30/20compare
tinker_eval_v9_s0_lvl10_v9_ts43_pick_s020-promax_20260810_1318464.9 ± 1.5 (k2)0.3 ± 0.1 (k2)5/170221/34001.30/20compare
tinker_eval_v9_s0_lvl10_v9_ts43_rl30-promax_20260809_0042065.6 ± 1.4 (k2)0.4 ± 0.1 (k2)4/170236/34001.39/20compare
tinker_eval_v9_s0off_lvl10_v9_ts42_pick_s010-promax_20260810_1318465.3 ± 2.4 (k2)0.4 ± 0.2 (k2)4/170241/34001.42/20compare
tinker_eval_v9_s0off_lvl10_v9_ts42_rl30-promax_20260809_0042066.0 ± 1.6 (k2)0.4 ± 0.1 (k2)5/170229/34001.35/20compare
tinker_eval_v9_s0off_lvl10_v9_ts43_pick_s025-promax_20260810_1318465.0 ± 1.9 (k2)0.4 ± 0.1 (k2)4/170242/34001.42/20compare
tinker_eval_v9_s0off_lvl10_v9_ts43_rl30-promax_20260809_0042064.9 ± 1.3 (k2)0.3 ± 0.1 (k2)6/170243/34001.43/20compare
v6.2_qwen3.5-9b-sft-v6eff_202607226.8 ± 2.5 (k2)3/2045/1002.25/5compare
v6.2_qwen3.5-9b-sft-v8full_202607228.0 ± 1.6 (k2)1/2043/1002.15/5compare
v6.2_qwen3.5-9b-sft-v8resetaware_202607227.7 ± 2.1 (k2)5/2045/1002.25/5compare
v7.3_qwen3.5-9b-base-rl-s5_202607238.9 ± 6.5 (k2)5/2038/1001.90/5compare
v7.3_qwen3.5-9b-base_202607234.1 ± 3.8 (k2)4/2032/1001.60/5compare
v7.3_qwen3.5-9b-sft-v8resetaware-rl-s15_2026072313.4 ± 2.4 (k2)3/2047/1002.35/5compare
v7.3_qwen3.5-9b-sft-v8resetaware_2026072310.6 ± 2.7 (k2)1/2039/1001.95/5compare

HiScore — Human Leaderboard

play · record · rank — you are scored with the SAME official RHAE (k=2) as the models above. Runs played on the v2 (teach-first) level banks carry a v2 badge and are scored against the v2 baselines; only v1-bank runs feed the empirical human-baseline steps behind the models' 'human' metric

runRHAE mean±stdgameslvlsavg_lvls
human_Guanghan_20260921_18553790.3 ± — (k2)11/11220/22020.00/20compare
human_Guanghan_20260724_17104283.9 ± — (k2)4/419/204.75/5compare
human_Guanghan_20260804_15453181.9 ± — (k2)6/6120/12020.00/20compare
human_Guanghan_20260729_04573075.1 ± — (k2)5/525/255.00/5compare
human_zomglings_20260806_07091312.5 ± — (k2)1/13/53.00/5compare
human_huu_20260921_1849491.9 ± — (k2)1/27/103.50/5compare
human_Deniz_20260921_1849490.9 ± — (k2)0/12/52.00/5compare
human_allie_20260921_1849490.6 ± — (k2)0/13/53.00/5compare

RHAE (Relative Human Action Efficiency): per level, scoreL = min(115, (baselineL / actionsL)² × 100) if the level was completed, else 0; a game's score is the completion-capped, level-weighted mean of scoreL (official bench.scoring; cap = the column's level count: RHAE-L5 = first 5 levels, the runs-page standard and ranking key; RHAE-L20 = first 20 levels, deep execution). The cell shows the per-rollout mean over the bundle's games, mean ± std over seeds/rerolls, at the reporting convention k=2 (baselineL = 2 × BFS-optimal steps) — directly comparable to the published frontier tables.

The human metric re-scores the same runs with baselineL = average human steps from the HiScore runs (per level; levels no human has completed keep the k2 baseline), at both caps — RHAE-L5 over the first 5 levels (the runs-page standard, which the perf rank uses) and RHAE-L20 over the first 20.