On August 9, 2026 — four days after Meta shipped it — we ran Muse Spark 1.2 through the Startrise LLM Benchmark: the same 12 frontend briefs, browser gates, and blind cross-lab judge panel that scored twelve rival models in July. We published a provisional, auto-only 73.3 and said blind human review would move the number. It did — down. With 8 of 12 cells now blind-rated, Muse Spark 1.2 scores 70.3: 6th of 14 on the pooled table, below GLM 5.2 (72.1) and Qwen3.8-Max’s later entry (71.3), above Grok 4.5 (68.0). The review took away both of its entry podiums — the landing page fell from 2nd to 9th, the HTML email from 2nd to 5th — and the reviewer disagreed with the judge panel in both directions: lower on the taste work, higher on the self-pitch the judges savaged. What review didn’t touch: a perfectly clean sheet, $0.62 for all twelve deliverables in under eleven minutes, and the interactivity weakness. The method and the July field live in the full study; this post is the Muse story.
Nobody else has this data. Every page ranking for “Muse Spark 1.2 benchmark” right now is Meta’s launch numbers re-plotted, an aggregator index, or a sub-five-task smoke test. This is the only third-party run of Muse Spark 1.2 on the identical tasks its rivals already went through — and, as far as we can find, the only blind human review of this model’s output anywhere. Both complicate the launch story.
What Muse Spark 1.2 is
Muse Spark 1.2 is Meta’s frontier reasoning model, released August 5, 2026 alongside Muse Code, Meta’s terminal coding agent, as the coding-focused update to Muse Spark 1.1. It’s Meta’s first frontier model offered through a paid API (the Meta Model API launched in July 2026 with Spark 1.1), and it’s also available on OpenRouter as meta/muse-spark-1.2 at $1.25/$4.25 per million tokens, with a 1,048,576-token context window, roughly 131k max output, and closed weights.
Meta’s launch pitch is agentic coding, and its self-reported numbers back it: 82.9% on Terminal-Bench 2.1 and 59.3% on DeepSWE v1.1, plus an internal Meta coding bench whose charts place Spark 1.2 ahead of GPT-5.6 Terra and behind Claude Opus 5. Those are Meta’s evaluations of Meta’s model. Even the Artificial Analysis write-up — genuinely independent — is an aggregate intelligence index built from existing evals, not a hands-on build test. Here’s what happened on ours.
How we tested it
Run 2026-08-09-2106 put Muse Spark 1.2 through the same harness as the July study: the same twelve frontend briefs, single-shot, no follow-up turn, one self-contained HTML file per brief, driven through automated gates in headed Chromium, then scored blind by the same three-judge panel — Claude Opus 5 on Craft, GPT-5.6 Terra on Engineering, Grok 4.5 on Brief — under the same 25/45/30 technical/judge/human blend. One model, 12 deliverables, 12 gate runs, 36 judge verdicts. Nothing else was re-run. Blind human review landed August 9: 8 of 12 cells rated by a single reviewer, model names hidden until every cell in a task was scored.
One configuration note worth keeping: Muse reasons inside its output budget, the same trait as GLM 5.2, so it got the same 128k max_tokens headroom. Unlike GLM, it never came close to needing it.
Worth being precise about what “single-shot” means here, because it is the whole methodology: every deliverable in this benchmark is the output of one prompt. One API call, no agent harness, no tool use, no linter, no retry, no follow-up turn, no human nudging it along. The file the model returns is the file we score. That is deliberately not how most people use these models day to day — Claude Code, Cursor, and Meta’s own Muse Code are harnesses that read errors, retry, and iterate — and it produces different rankings than harness-mediated work would. The study says it plainly in its own caveats: single-shot only; nothing here measures iteration.
Single-shot is also where reputation and raw output part ways, and Claude Fable 5 is the sharpest example in our field. Fable is the most expensive model we tested ($10/$50 per MTok), the only model in the study that beats Opus 5 at anything, and the winner of all three open-subject briefs. And on the Three.js scroll brief, its one shot arrived with a JavaScript syntax error: the page rendered near-black, the animation loop never ran (0 fps), scroll did nothing — final score 22.5, on the same brief where Muse posted a 71.1 with perfect gates. In a harness, that bug is a ten-second self-fix after one look at the console. Single-shot, there is no console to read and no second turn.

A model that is stronger on almost every other measure lost that brief to a bug a linter would catch. That is exactly the capability this benchmark isolates — the raw first shot — and exactly why you should read it alongside harness-based evals like Meta’s Terminal-Bench numbers, not instead of them. A model that stumbles single-shot can still be excellent in the loop that fixes its own stumbles; what single-shot tells you is how much fixing the loop will be doing.
The briefs themselves — what each one separates, the fixture with the money bug, the judge lenses, the published bug list — are documented in the flagship study. We won’t re-explain them here.
Read this before the score: what moved, and what can still move
When this page first published, the score was a provisional, auto-only 73.3, and this section said blind review would move it in either direction. Review has now landed for 8 of 12 cells: 70.3, down three points. Muse’s number is — in the study’s phrase — “the same kind of number as its neighbours.” It is still not a settled one. Five reasons, stated as findings, because they are.
Four cells remain unrated. The accessible interface, the brownfield change, the stateful app, and the zero-JS brief have no human rating yet; their human weight is redistributed, the same treatment the July field gets at its 106-of-144 coverage. Two of the four are Muse’s behaviour-heavy briefs — rating them can still move the overall in either direction.
The human column is one reviewer. Blind — model names hidden until every cell in a task is scored — but a single rater, same as the July field’s ratings. Where reviewer and panel disagree, and on Muse they do in both directions, that is a finding about both.
N=1, no error bars. Every brief ran once. The 1.0-point gap to Qwen3.8-Max above (added August 12) and the 2.3-point gap to Grok 4.5 below are both inside plausible run-to-run variance. Read “6th of 14” as how the table displays, not a settled rank — it has already moved once as a new model joined.
The contamination window. Our twelve briefs have been public in the benchmark repository since July 26, 2026. Muse Spark 1.2 shipped August 5 — after that date. There is no evidence it saw them, but the possibility exists for any model released after publication, which is exactly why the study’s holdout requirement exists.
It’s not in the judge-bias audit. The 839-call audit predates this run and covers the July twelve; Muse judged nothing, and no Meta model sits on the panel, so no self-judging conflict applies — it joins at the next full re-run. What that audit is and why it matters: the judge-bias write-up.
And one flag from the first version of this page has closed: the WebGL shader split the panel by 4 on an axis, marked SPLIT PANEL. The blind rating came in at 50 against the panel’s 65 — the sceptic on the panel was closer.
What blind review changed: both podiums, both directions
Blind human review took both of Muse Spark 1.2’s entry podiums away. On auto-only numbers it briefly held 2nd of 13 on the brand landing page (83.9) and 2nd on the HTML email (80.7). The blind ratings disagreed with the judges on exactly those cells: the landing page rated 55 against the panel’s 75, dropping its final to 75.3 — 9th at review time, 10th today; the email rated 65, dropping it to 76.0 and 5th (Opus 5 won that brief at 84.3, and Muse’s email still drew zero red flags). The review restored the July podiums, and the table has since moved again: Qwen3.8-Max’s August 12 run took the landing-page win outright at 87.3 — the live benchmark always carries the current standings. Muse’s best standing final is the zero-JS brief (79.1, 4th) — one of the four cells still unrated. Here’s the full run with the human column in:
| Task | Technical | Judge | Human | Final | Red flags |
|---|---|---|---|---|---|
| 01 three.js scroll | 100 | 62.5 | 60 | 71.1 | 6 |
| 02 WebGL shader | 100 | 65.0 | 50 | 69.3 | 5 |
| 03 landing page | 100 | 75.0 | 55 | 75.3 | 5 |
| 05 3D game | 92.0 | 62.5 | 55 | 67.6 | 8 |
| 06 open creative | 83.4 | 67.5 | 65 | 70.7 | 8 |
| 07 sell yourself | 89.1 | 45.0 | 80 | 66.5 | 9 |
| 08 accessible | 90.5 | 57.5 | — | 69.3 | 6 |
| 09 brownfield | 79.1 | 72.5 | — | 74.9 | 1 |
| 10 SVG icons | 100 | 60.0 | 50 | 67.0 | 7 |
| 11 stateful app | 83.3 | 42.5 | — | 57.1 | 6 |
| 12 zero JS | 100 | 67.5 | — | 79.1 | 8 |
| 13 HTML email | 100 | 70.0 | 65 | 76.0 | 0 |
Run 2026-08-09-2106, blind human review on 8 of 12 cells; ”—” marks the unrated cells, whose human weight is redistributed. Every task name opens Muse’s actual deliverable — the exact single-shot HTML file the scores describe, hosted unmodified. Per-task flag counts aren’t comparable to the July red-flag totals — different review status — so we don’t sum them. Full per-cell data at the benchmark.
All twelve, exactly as generated — click any to open the live build:
01 three.js · 71.1
02 shader · 69.3
03 landing · 75.3
05 3D game · 67.6
06 open · 70.7
07 self-pitch · 66.5
08 accessible · 69.3
09 brownfield · 74.9
10 icons · 67.0
11 stateful · 57.1
12 zero JS · 79.1
13 email · 76.0
Here’s the part no other page on the internet has: nobody else has put a blind human reviewer in front of this model’s output, and ours disagreed with the judge panel in both directions. Below the judges on the taste work — landing 55 vs 75, icons 50 vs 60. Above them on the one brief the judges savaged: the self-pitch, human 80 against judge 45, lifting that final from 60.8 to 66.5. The study’s verdict, verbatim: “Muse under blind human eyes: less beautiful than the panel thought, more honest.”


Across its 8 rated cells the human column averages 60.0. Set against the July field, that sits above GLM 5.2’s 58.8 and Claude Fable 5’s 59.4, and well under Kimi K3’s 70.6 and Opus 5’s 71.7 — mid-pack in human eyes too, just by a different route than the judges took.
Then the sheet itself, which review didn’t touch: twelve of twelve built, zero contract violations, zero truncation, every page loaded console-clean. The brownfield diff drew a single red flag — its cleanest judged work. For context, the July field produced five hard failures across 144 deliverables: Sonnet 5 truncated twice, Grok 4.5 once, and GLM 5.2 wrapped deliverables in prose twice. Muse produced none.
The tension worth quoting changed shape under review: a model marketed on agentic coding, whose best surviving finals are still the static work — and whose highest human rating came for telling the truth about itself.
Where the field beat it: anything you have to click
Muse Spark 1.2’s measured weakness is behaviour. Our interaction gate — a scripted scroll, pointer, and keyboard pass measuring whether the page responds — scored zero on the open creative brief, the brownfield change, and the stateful app, and 43–44 on the 3D game and the accessible interface. The judges’ red flags echo the same finding: pages that render beautifully and under-respond to input.
Its two worst finals are still the two briefs where behaviour, not appearance, carries the score: the stateful app at 57.1 — unrated by the human reviewer, and the July winner, Opus 5, took that brief at 78.7 — and the self-pitch at 66.5, the one cell blind review pulled up, from 60.8, against a judge score of 45.0 and nine red flags.

The axis medians complete the shape: craft 6.5, technique 7, adherence 7, originality 5.5. Competent everywhere, distinctive nowhere. Against the main study’s axis table, Opus 5 is the field’s only 8 on originality and Fable 5 sits at 7. The chosen-subject spark that wins open briefs is what Muse’s 5.5 lacks.
The launch-week discourse had already decided the frontend question — daily.dev’s take crowned it “A CRAZY FRONTEND BEAST”. Our data says: half right. The static half — and blind review has since trimmed even that.
Muse Spark 1.2 vs Claude, Kimi, GLM and GPT-5.6
Five face-offs, every score from the same harness, every verdict a routing rule — Muse’s side now blind-reviewed on 8 of 12 cells.
vs Claude Opus 5 (82.3). Opus won 8 of 12 tasks in July; Muse’s blended 70.3 sits twelve points under it, and the human columns agree — 71.7 to 60.0. Muse’s case is price: $1.25/$4.25 per million tokens against Opus 5’s $5/$25 — a case that holds for static marketing surfaces only. Anything a user clicks stays with Opus.
vs Kimi K3 (79.8). Kimi wrote the cleanest source in the study — 11 judge-named red flags — and is its slowest model at 13,381 seconds for the suite. Muse ran it in 647. The two cleanliness claims still differ in kind: Muse’s is contract-level (no violations, no truncation, console-clean); Kimi’s is judge-level — and the blind reviewer sides with Kimi on quality too, 70.6 to 60.0 on the human column, handing it back the landing-page 2nd Muse briefly held. Merge-quality batch code goes to Kimi; cheap fast pages go to Muse.
vs Claude Fable 5 (73.8). Review stretched the gap from 0.5 points to 3.5 — and flipped one detail: the blind reviewer rates Muse’s work a hair above Fable’s, 60.0 to 59.4 on the human column. The shape difference is everything: Fable won all three open-ended briefs; Muse’s originality median is 5.5 to Fable’s 7. Fable runs $10/$50 per million tokens, eight times Muse’s input rate. And Fable is the field’s clearest single-shot casualty — the syntax error that cost it the Three.js brief (final 22.5 to Muse’s 71.1) is a failure mode a harness erases and this benchmark deliberately doesn’t. Open-ended design is Fable’s; cheap static pages are Muse’s.
vs GLM 5.2 (72.1). The twin — and, after review, above: GLM’s 72.1 edges Muse’s blended 70.3, a 1.8-point gap inside run-to-run noise (Qwen3.8-Max’s later entry now sits between them). Both reason inside the output budget, both get 128k headroom. GLM posted the study’s highest technical score (94.4, with Muse’s 93.1 just under it and Opus 5’s 93.9) but wrapped deliverables in prose twice; Muse’s sheet is clean, and the human column marginally prefers Muse, 60.0 to 58.8. GLM is cheaper still — $0.68/$2.14 per MTok in the study, and OpenRouter listed $0.72/$1.80 as we drafted this.
vs the GPT-5.6 family (Terra 66.9 · Sol 66.2 · Luna 65.0). Review gave Sol its email 2nd place back — Muse’s email fell to 5th at 76.0 — but Muse still sits above all three on the pooled table, now on a blended number rather than a provisional one. Luna remains the only model faster (454s) and lighter (74,075 output tokens) in the field.
The economy: $0.62 for the whole suite
The cost profile is a finding in its own right: Muse Spark 1.2 generated all twelve deliverables for $0.62 at list rates, in 647 seconds of wall clock — second-fastest in the field, behind only Luna’s 454 — on 138,562 output tokens, the third-lightest footprint in the study (only Luna and Terra emit less). As the study puts it: “For a model that reasons before it writes, that economy is the headline capability.”
At the other end of the table, Opus 5 spent 803,391 output tokens and 9,737 seconds on the same twelve briefs; Sonnet 5 emitted 918,978 tokens in July, the most in the study, and truncated twice anyway.
One practical consequence: at $1.25/$4.25, the marginal cost of trying Muse on your own task set is effectively zero. Do that instead of trusting any single N=1 run — ours included.
Verdict: when to pick Muse Spark 1.2
Our Muse-only routing rules, now with blind review behind them — the field-wide rankings live elsewhere and are not re-litigated here:
- Static marketing surfaces at volume — landing pages, HTML email, zero-JS builds — where cost matters: still a contender, on volume economics rather than podiums. Its best standing finals are exactly this work (zero-JS 79.1, email 76.0, landing 75.3), the sheet is clean, and $0.62 for twelve deliverables makes the trial cost trivial.
- Stateful apps, interaction-heavy UI, anything a user clicks: route to Opus 5, which took the stateful app at 78.7 to Muse’s 57.1 (a still-unrated cell — but the interaction gate that scored it needs no reviewer). Three interaction-gate zeros is a pattern, not an accident.
- Open-ended design: Fable 5’s territory. Muse’s 5.5 originality median doesn’t compete there.
- Merge-quality batch code: Kimi K3’s, if you can afford the latency.
For the per-use-case verdicts across the whole field, see the field-wide buying guide — written before this run, and its picks survive it: nothing in Muse’s blended column displaces a field-wide winner.
When we published this post, we promised to update it when blind review landed. It landed August 9, 2026 — 8 of 12 cells — and this page was revised August 11 with the blended numbers: promise kept, number moved, direction down. The four unrated cells and the next full re-run with holdout briefs can still move it again; when they do, so will this page.
Why we had this data four days after launch
Startrise re-benchmarks its model routing every time the frontier moves, because routing decisions rot — Meta shipped on a Wednesday, by Sunday we had same-harness numbers instead of launch-deck quotes, and by Tuesday those numbers carried a blind human review. That discipline is what clients buy: when we build AI agents (on Mastra, with routing, fallbacks, guardrails, and approval gates, from $3,500), the model choice ships with evidence, not vibes. The practice hub is AI automation.
And if you’ve already shipped an AI-generated frontend that renders beautifully and misbehaves on click — the exact failure class this run measured — that’s what a Vibe Code Cleanup audit catches, and fixes, first.
Run 2026-08-09-2106 is live at the Startrise LLM Benchmark, next to the July data — every score, every gate, every caveat.
Questions we actually get
Is Meta's Muse Spark 1.2 good at frontend development?
Partly. In the Startrise LLM Benchmark it scores 70.3 blended — 6th of 14, with blind human review on 8 of 12 cells. Its best standing finals are static work (zero-JS 79.1, HTML email 76.0), but the interaction gate scored zero on three of its twelve briefs and its worst result came on the stateful app (57.1). Strong at static work; weak where behaviour carries the score.
Is Muse Spark 1.2 better than Claude?
Not on our benchmark. Claude Opus 5 leads the pooled table at 82.3 versus Muse's blended 70.3, and the blind human reviewer scored Opus's work 71.7 against Muse's 60.0. Muse's case is economy: $1.25/$4.25 per million tokens versus Opus 5's $5/$25, and it ran our full twelve-brief suite for $0.62.
How does Muse Spark 1.2 compare to GPT-5.6?
On our pooled table Muse's blended 70.3 sits above GPT-5.6 Terra (66.9), Sol (66.2), and Luna (65.0). Blind review handed Sol back 2nd on the HTML email brief — Muse's email fell to 5th at 76.0. Luna remains the efficiency leader — the only model faster (454s) and lighter (74,075 output tokens) than Muse's 647s and 138,562.
Why did Muse Spark 1.2's benchmark score change?
We published a provisional, auto-only 73.3 and said blind human review would move it. It did — down. With 8 of 12 cells blind-rated, the blended score is 70.3 — 6th of 14 since Qwen3.8-Max joined the table. The reviewer scored Muse's taste work below the judge panel (landing page 55 vs 75) and its self-pitch above it (80 vs 45). Four cells remain unrated and can still move the number.
How much does Muse Spark 1.2 cost to use?
$1.25 per million input tokens and $4.25 per million output on the Meta Model API and OpenRouter, with a 1,048,576-token context window. In our benchmark it generated all twelve frontend deliverables for $0.62 at list rates, in 647 seconds — the second-fastest run in the field.
Does the Startrise LLM Benchmark use an agent harness?
No. Every deliverable is the output of a single prompt — one API call per brief, no tools, no retries, no follow-up turn. That isolates raw first-shot capability, which is different from what harnesses like Claude Code, Cursor, or Meta's Muse Code measure: those read errors and iterate. The cost of the difference is visible in our data — Claude Fable 5, winner of all three open-subject briefs, shipped a Three.js page broken by one syntax error and scored 22.5 on that brief. A harness fixes that in seconds; a single shot doesn't get one.