Journal — Topic
LLM benchmark
Entries on LLM benchmark
· 19 min read
Grok 4.6 vs Qwen3.8-Max on our 12-brief frontend benchmark, blind-reviewed: 72.1 vs 71.3 — inside noise — six briefs each, all 24 builds linked live.
grok 4.6qwen3.8-maxllm benchmarkai coding
· 16 min read
Meta's Muse Spark 1.2, blind-reviewed on our 12-brief frontend benchmark: 70.3, 7th of 15. Human review took both podiums away — less beautiful, more honest.
muse spark 1.2llm benchmarkmeta aiai coding
· 10 min read
The best LLM for frontend work in 2026, from a 144-build benchmark: Claude Opus 5 for breadth, Kimi K3 for clean code, Fable 5 for design, Luna for volume.
llm benchmarkclaude opus 5kimi k3frontend development
· 16 min read
We made 12 frontier LLMs each build 12 real UIs, gated in a real browser, judged blind: Claude Opus 5 and Kimi K3 tied on top, Fable 5 won every open brief.
llm benchmarkai codingclaude opus 5kimi k3
· 11 min read
Self-preference bias in LLM judges, measured: we made 10 frontier models grade each other's work across 839 blind calls — including our own benchmark winner.
llm as a judgellm benchmarkai evaluationclaude opus 5