Journal — Topic
AI coding
Entries on AI coding
· 19 min read
Grok 4.6 vs Qwen3.8-Max: 24 Builds, One Blind Reviewer
Grok 4.6 vs Qwen3.8-Max on our 12-brief frontend benchmark, blind-reviewed: 72.1 vs 71.3 — inside noise — six briefs each, all 24 builds linked live.
grok 4.6qwen3.8-maxllm benchmarkai coding
· 16 min read
Muse Spark 1.2 Benchmark: Meta's Model vs the Field
Meta's Muse Spark 1.2, blind-reviewed on our 12-brief frontend benchmark: 70.3, 7th of 15. Human review took both podiums away — less beautiful, more honest.
muse spark 1.2llm benchmarkmeta aiai coding
· 16 min read
LLM Frontend Benchmark: 12 Models Built 12 Real UIs
We made 12 frontier LLMs each build 12 real UIs, gated in a real browser, judged blind: Claude Opus 5 and Kimi K3 tied on top, Fable 5 won every open brief.
llm benchmarkai codingclaude opus 5kimi k3