glyphhunt

leaderboard

Every model, every metric, in one table. Sort any column, filter by name. Runs that left their working directory are shown but never scored.

● running
Jul 25, 2026, 03:11 AM UTC
arm64 · macOS 26.5.2 · 1920x1080 · 60fps · 1800 frames · 17 glyphs in 17 typefaces
42/72 runs 4 models ✓ 0 invalidated ~ 35 hit plan quota ✓ isolated workdirs ✓ key committed in advance
Zapfino
Papyrus
Herculanum
Chalkduster
Bodoni 72
Didot
Copperplate
Luminari
Trattatello
Snell Roundhand
Brush Script
Marker Felt
Noteworthy
American Typewriter
Impact
Phosphate
Rockwell
overall
across every model, level and prompt mode
~ partial
runs scored
42
exact solves
5/42
mean chars
2.52/17
mean frames
1.21/17
mean spatial
1.17/17
mean time
1695s
hit 30m wall
27/42
invalidated
0
every metric
click any column to sort. word, temporal and spatial accuracy are kept separate — reading the word and knowing where you saw it are different abilities.
model solved chars/17 accuracy best frame spatial lev L1L2L3 blind hinted avg t shell ffmpeg python imgs diffed out tok cost fail invalid
opus-5 2/4 8.5 ███████······· 17 8.5 8.5 0.5 1700 11.33 0 1390s 45.75 6 41 22.25 2 23,895 $7.94 2 0
gpt-5.6/high 2/18 2.33 ██············ 17 0.56 0.56 0.18 700 3.78 0.89 1946s 46.56 10.83 21.17 0 11 20,140 13 0
gpt-5.6/med 1/18 1.67 ············· 17 0.39 0.28 0.17 4.670.170.17 0.44 2.89 1549s 39.83 11.06 19.89 0 11 18,559 9 0
fable-5 0/2 0 ·············· 0 0 0 0 00 0 0 1352s 33.5 9 28.5 31.5 1 21,269 $8.34 2 0
what the hint is worth
blind says only that something is hidden. hinted gives the length, the one-typeface-per-letter rule, temporal ordering, and warns about faintness and decoys.
model blind chars hinted chars Δ blind solved hinted solved
opus-5 11.33 0 -11.3 2/3 0/1
gpt-5.6/high 3.78 0.89 -2.9 2/9 0/9
gpt-5.6/med 0.44 2.89 +2.5 0/9 1/9
fable-5 0 0 0 0/1 0/1