unsigned.gg // benchmarks nobody asked for

The Frontier Leaderboard

the models we can't run at home·higher is better, allegedly·snapshot 2026-07

🐡 What is "Fugu"? Turns out it's real: Sakana AI's Fugu Ultra, launched June 2026. The catch — it's a router over a pool of models, not a single LLM, so a big score can be great routing as much as raw ability. And every number in this chart is Sakana-reported from a deck that quietly omits Claude Mythos 5 (80.3) and Fable 5 (80.0) — the two models that actually top SWE-bench Pro. Provenance: the vendor. Scroll down for the independently-tracked board with the missing rows put back.
How to read the provenance badges
independent measured by a neutral third party (or by us, on our own 5090)
corroborated vendor placed the number, but it matches an independent tracker
self-reported ⚠ COI reported by the model's own maker; conflict of interest, not reproduced
disputed independent sources contradict the label or the number

Who won what

outright wins per model across every row · ties count for everyone who tied (looking at you, GPQA)

The whole table

gold = row winner · underline = runner-up · everyone else gets participation

loading numbers off the napkin…

Same benchmark, independently tracked independent

SWE-bench Pro, from a neutral tracker — with the rows the vendor chart left out

modelSWE-bench Proprovenance
loading the honest board…

Which models are actually interesting

not "who scored highest" — who's worth a second look, and why

On the bench (literally)

candidates loitering near the hardware, waiting for the 5090 to form an opinion