The Journal
What we ship, what breaks, and what actually works. Written by the people doing the work — no ghostwriters, no filler.
More entries
· 16 min read
Meta's Muse Spark 1.2, blind-reviewed on our 12-brief frontend benchmark: 70.3, 7th of 15. Human review took both podiums away — less beautiful, more honest.
muse spark 1.2llm benchmarkmeta aiai coding
· 10 min read
The best LLM for frontend work in 2026, from a 144-build benchmark: Claude Opus 5 for breadth, Kimi K3 for clean code, Fable 5 for design, Luna for volume.
llm benchmarkclaude opus 5kimi k3frontend development
· 11 min read
Claude Opus 5 (July 24, 2026) vs Fable 5: verified pricing, benchmarks, ZDR rules, and a routing guide for agentic workloads — plus where Opus 4.8 fits.
claude opus 5claude fable 5ai agentsmodel routing
· 16 min read
We made 12 frontier LLMs each build 12 real UIs, gated in a real browser, judged blind: Claude Opus 5 and Kimi K3 tied on top, Fable 5 won every open brief.
llm benchmarkai codingclaude opus 5kimi k3
· 11 min read
Self-preference bias in LLM judges, measured: we made 10 frontier models grade each other's work across 839 blind calls — including our own benchmark winner.
llm as a judgellm benchmarkai evaluationclaude opus 5
· 10 min read
Design AI agent approval workflows without vendor lock-in: risk tiers (auto-approve, notify, block-and-ask) plus n8n and MCP implementation sketches.
AI agentsapproval workflowshuman in the loopn8n
· 8 min read
How to hire an AI enablement consultant: what they do, engagement models, real fixed pricing from $5,000, and the questions that expose pretenders.
ai enablementai operatorshiring consultantsai adoption
· 9 min read
What an AI operator actually does: producing real work through AI tools. Responsibilities, skills, org placement, and a copy-paste job description template.
AI operatorAI enablementhiringAI automation
· 9 min read
AI project rescue, defined: a staged recovery framework for stalled chatbots, RAG, and automation pilots, plus how to decide rescue vs rebuild vs retire.
AI project rescuefailed AI projectsAI automationRAG
· 9 min read
What Claude Fable 5 changes for business output: verified benchmarks, the June suspension timeline, real rollout checks, and how one operator runs it.
Claude Fable 5AI enablementAI operatorClaude Code
· 11 min read
A developer's guide to cleaning up AI-generated code: nine recurring smells, before/after fixes, and the safety-net workflow that makes refactoring safe.
AI-generated codecode cleanuprefactoringcode smells
· 9 min read
A senior engineer's triage playbook for broken vibe-coded apps: stabilize, diagnose, then decide refactor vs rewrite, with the risk order that matters.
vibe codingAI-generated codecode cleanuptechnical debt