Days apart in August 2026, xAI shipped Grok 4.6 and Alibaba shipped Qwen3.8-Max at the identical $2/$6 list price — so we ran both through the Startrise LLM Benchmark: the same twelve frontend briefs, browser gates, and blind cross-lab judge panel that have now scored fifteen models, plus blind human review on 9 of 12 cells each. Grok 4.6 scored 72.1 (a dead tie with GLM 5.2), Qwen3.8-Max 71.3 — and the 0.8-point gap between them is inside run-to-run noise, so read the headline as a tie. Underneath the tie, the run data is anything but flat: the two split the twelve briefs six each (our arithmetic on the published finals), each cratered exactly one brief to a blind-review 0, each holds at least one outright task win against the July field (Qwen holds two), Grok’s human column (65.0) is the third-best on record while the judge panel prefers Qwen — and at the same rate card, the measured suite bills differ 1.8x. All 24 builds are embedded below, every one linked to its live deliverable. The method and the July field live in the full study; this page is the head-to-head.
Nobody else has this comparison. Every page ranking for either model right now is a launch-table re-plot, and most of the “vs” pages pit Qwen3.8-Max against the previous Grok. As far as we can find, this is the only place Grok 4.6 and Qwen3.8-Max have built the same real deliverables under the same instruments — and the only blind human review of either model’s output anywhere.
Two models, one price, days apart
Grok 4.6 is xAI’s new frontier model, released August 12, 2026; Qwen3.8-Max is Alibaba’s, generally available since early August 2026 — and both list $2 per million input tokens, $6 per million output. On OpenRouter, where we access both, Grok 4.6 is x-ai/grok-4.6 with a 500k context window and no published output cap. xAI’s launch pitch is “long-running agents” and “more ambitious interactive and visual work,” with launch-partner availability in Cursor (2x included usage for the first week) — the vendor’s claims, attributed as such. One naming note: the launch page now carries SpaceXAI branding over an X.AI LLC legal footer; the study says xAI, and so do we. Artificial Analysis — independent, but an aggregate index rather than a build test — scores it 61 on their intelligence index, “in line with GPT-5.6 Sol.”
Qwen3.8-Max is qwen/qwen3.8-max: 1M context, 131,072 max output tokens, and a vendor-claimed 2.4 trillion parameters (95B active). Alibaba’s launch post — “A New Bar for Coding and Cowork” — self-reports 86.6 on Terminal-Bench 2.1 and 67.7 on SWE-bench Pro; those are Alibaba’s numbers on Alibaba’s chosen benches.
Identical price, near-identical launch confidence — all of it vendors grading their own models. Here’s what happened when both built the same twelve frontends under a blind panel.
How we tested them
Runs 2026-08-12-0114 (Qwen3.8-Max) and 2026-08-12-2359 (Grok 4.6) ran on the same harness as the July study: twelve briefs, single-shot, one self-contained HTML file per brief, automated gates in headed Chromium, then blind scoring by the same three-judge panel — Claude Opus 5 on Craft, GPT-5.6 Terra on Engineering, Grok 4.5 on Brief — under the same 25/45/30 technical/judge/human blend. Blind human review covered 9 of 12 cells for each model: Qwen’s landed August 12, Grok’s August 13.
One configuration note: both got 128k max_tokens headroom — Qwen on the GLM/Muse precedent (it never came close to needing it), Grok because Grok 4.5 truncated the accessible-interface brief at the 64k ceiling in July. With the headroom, 4.6 recorded zero truncation. Why one bug can cost sixty points single-shot — and why that’s the point — is covered in the Muse post; the briefs, the gates, and the published bug list are documented in the flagship study. We won’t re-explain them here.
Read this before the scores
Six caveats govern every number on this page, and they’re findings, not fine print. The largest: Grok 4.6’s 72.1 and Qwen3.8-Max’s 71.3 sit 0.8 points apart on an N=1 benchmark, and that gap is inside run-to-run noise.
The 0.8-point gap is noise. Every brief ran once per model — no error bars. Grok’s dead tie with GLM 5.2 and the 1.0-point gap between Qwen and Muse Spark 1.2 below it are the same kind of non-difference. Anyone quoting this page as “Grok 4.6 beats Qwen3.8-Max” is misquoting it.
Three cells per model are unrated — the same three. The accessible interface, the brownfield change, and the stateful app carry no blind human rating on either run; their human weight is redistributed. That set includes Grok’s record-split accessible cell and both of its unrated tie-3rd podiums — we flag every one of them below.
The human column is one reviewer. Blind — model names hidden until every cell in a task is scored — but a single rater, same as every run on record.
The contamination window. The twelve briefs have been public in the benchmark repository since July 26, 2026. Both models shipped after that date. There’s no evidence either saw them. But the possibility is open for any model released after publication, which is exactly why the study’s holdout requirement exists.
Neither model is in the judge-bias audit yet — and one judge is a relative. Grok 4.5, the panel’s Brief judge, is the predecessor of one contestant. The panel is blind, and the 839-call audit measured Grok 4.5’s self-bias at −0.93 — it under-scored its own work — but it is the only judge with a family member in this pairing, and we’d rather disclose the conflict with its measured size than pretend it isn’t there. Both models join the audit at the next full re-run.
The record panel split is unresolved. Grok’s accessible interface split the judge panel by 6 on an axis — the previous record was 4 — against gates of 77.2, and it’s one of the unrated cells. Its 51.7 final is, in the study’s words, “the least settled number in this addendum.”

Grok 4.6’s twelve: an email record and a broken game
Run 2026-08-12-2359: Grok 4.6 scored 72.1 overall — technical 88.5, judge 67.1, human 65.0 — a dead tie with GLM 5.2 on the pooled table. That human column is the third-best on record, behind only Claude Opus 5’s 71.7 and Kimi K3’s 70.6, and the overall sits 4.1 points above its predecessor Grok 4.5 (68.0). Here’s the full run:
| Task | Technical | Judge | Human | Final | Red flags |
|---|---|---|---|---|---|
| 01 three.js scroll | 100 | 72.5 | 80 | 81.6 | 6 |
| 02 WebGL shader | 100 | 72.5 | 60 | 75.6 | 0 |
| 03 landing page | 100 | 85.0 | 70 | 84.3 | 0 |
| 05 3D game | 64.4 | 22.5 | 0 | 26.2 | 8 |
| 06 open creative | 77.0 | 75.0 | 80 | 77.0 | 2 |
| 07 sell yourself | 89.1 | 65.0 | 65 | 71.0 | 2 |
| 08 accessible | 77.2 | 37.5 | — | 51.7 | 8 |
| 09 brownfield | 78.8 | 80.0 | — | 79.6 | 0 |
| 10 SVG icons | 100 | 65.0 | 75 | 76.8 | 5 |
| 11 stateful app | 84.2 | 70.0 | — | 75.1 | 0 |
| 12 zero JS | 90.7 | 77.5 | 70 | 78.6 | 4 |
| 13 HTML email | 100 | 82.5 | 85 | 87.6 | 0 |
Run 2026-08-12-2359, blind human review on 9 of 12 cells (landed August 13); ”—” marks the three unrated cells — accessible interface, brownfield, stateful app, the same three on both runs — whose human weight is redistributed. Every task name opens Grok 4.6’s actual deliverable, hosted unmodified. Per-task flag counts aren’t comparable across runs with different review status, so we don’t sum them. Full per-cell data at the benchmark.
The story cell is the HTML email — the third August upset. Grok 4.6 takes that brief outright at 87.6 over Opus 5’s 84.3: perfect gates, a near-unanimous judge 82.5 (spread 0.75, zero red flags), and the highest blind rating of the August runs, an 85. It adds a 3rd on the landing page (84.3, on a judge verdict of 85 — tying the record Qwen set on the same brief) and two tie-3rds on the brownfield change (79.6) and the stateful app (75.1) — both unrated cells, which is why we call them paper podiums.

It also passed the brief that breaks frontier models. Grok 4.6 is the only August addition to clear the Three.js scroll with perfect gates — the brief that produced Claude Fable 5’s 22.5 and Qwen3.8-Max’s 29.3 — finishing 4th on it at 81.6 with an 80 blind rating. And then it shipped a broken game: the 3D game arrived with a console error, no animation frames, and near-dead input — technical 64.4, judge 22.5 with 8 red flags, blind rating 0, final 26.2, the deepest August cell. The remaining wobble is the accessible interface — the record 6-point panel split covered above, whose 51.7 is the least settled number in the run.
All twelve, exactly as generated — click any to open the live build:
01 three.js · 81.6
02 shader · 75.6
03 landing · 84.3
05 3D game · 26.2
06 open · 77.0
07 self-pitch · 71.0
08 accessible · 51.7
09 brownfield · 79.6
10 icons · 76.8
11 stateful · 75.1
12 zero JS · 78.6
13 email · 87.6
Run 2026-08-12-2359 in full — every gate, verdict, and rating — at the benchmark.
Qwen3.8-Max’s twelve: two July wins fall, one crater opens
Run 2026-08-12-0114: Qwen3.8-Max scored 71.3 overall — technical 88.0, judge 69.2, human 56.1 — 6th of 15 on the pooled table, 0.8 below the Grok/GLM tie and 1.0 above Muse Spark 1.2, every one of those gaps inside run-to-run noise. It’s a 7.7-point blended improvement over Qwen 3.7 Max (63.6, now 12th): the 3.7 generation had zero podium finishes, 3.8 has five, and its judge column trails only Opus, Kimi and Fable. The full run:
| Task | Technical | Judge | Human | Final | Red flags |
|---|---|---|---|---|---|
| 01 three.js scroll | 58.6 | 32.5 | 0 | 29.3 | 9 |
| 02 WebGL shader | 98.4 | 80.0 | 80 | 84.6 | 0 |
| 03 landing page | 100 | 85.0 | 80 | 87.3 | 1 |
| 05 3D game | 85.7 | 45.0 | 20 | 47.7 | 7 |
| 06 open creative | 85.6 | 80.0 | 60 | 75.4 | 0 |
| 07 sell yourself | 89.1 | 70.0 | 80 | 77.8 | 4 |
| 08 accessible | 83.7 | 80.0 | — | 81.3 | 4 |
| 09 brownfield | 79.0 | 75.0 | — | 76.4 | 0 |
| 10 SVG icons | 100 | 82.5 | 65 | 81.6 | 0 |
| 11 stateful app | 76.8 | 55.0 | — | 62.8 | 8 |
| 12 zero JS | 98.9 | 77.5 | 60 | 77.6 | 0 |
| 13 HTML email | 100 | 67.5 | 60 | 73.4 | 7 |
Run 2026-08-12-0114, blind human review on 9 of 12 cells (landed August 12); ”—” marks the same three unrated cells as the Grok run — accessible interface, brownfield, stateful app — whose human weight is redistributed. Every task name opens Qwen3.8-Max’s actual deliverable, hosted unmodified. Per-task flag counts aren’t comparable across runs with different review status, so we don’t sum them. Full per-cell data at the benchmark.
The headline here: the first July wins to fall. Qwen3.8-Max holds two task wins with blind review behind them: the brand landing page at 87.3, dethroning Opus 5’s 85.8 on the highest landing-page judge verdict on record (85), and the SVG icon system at 81.6, taking the brief Claude Fable 5 had held at 80.5. It adds three second places — the WebGL shader (84.6, pushing Opus 5 to third), the self-pitch (77.8), and the accessible interface (81.3, an unrated cell).
The crater is the Three.js scroll: a console error, a page that barely renders, no animation loop, near-dead scroll — technical 58.6, judge 32.5 with 9 red flags, and the blind reviewer’s first 0 of the addendum runs. Final: 29.3. The same single-shot lesson as Fable’s 22.5 on the same brief: one unhandled error, no second turn, sixty points gone.
This run is also the clearest demonstration yet that our instruments disagree in both directions. The interaction gate scored the accessible interface a 2; the judges, reading the source, scored the cell 80 with the tightest agreement of the run (spread 0.5). The reverse on the game: the gates saw 120fps and full interaction; the judges saw a broken loop and flagged it seven times. Neither instrument is sufficient alone — which is why there are three.
Qwen’s twelve, same deal — exactly as generated, click any to open the live build:
01 three.js · 29.3
02 shader · 84.6
03 landing · 87.3
05 3D game · 47.7
06 open · 75.4
07 self-pitch · 77.8
08 accessible · 81.3
09 brownfield · 76.4
10 icons · 81.6
11 stateful · 62.8
12 zero JS · 77.6
13 email · 73.4
Run 2026-08-12-0114 in full — every gate, verdict, and rating — at the benchmark.
Head-to-head: six briefs each
Set Grok 4.6’s and Qwen3.8-Max’s finals columns side by side — our arithmetic on the published numbers, not a new measurement — and each model takes exactly six of the twelve briefs. The split is real; what it isn’t is a score.
- Grok 4.6 edges six: the Three.js scroll (81.6 v 29.3), the open creative (77.0 v 75.4 — a 1.6-point lean), the brownfield change (79.6 v 76.4 — a 3.2-point lean, on an unrated cell), the stateful app (75.1 v 62.8, unrated), the zero-JS page (78.6 v 77.6 — a 1.0-point lean), and the HTML email (87.6 v 73.4).
- Qwen3.8-Max edges six: the WebGL shader (84.6 v 75.6), the landing page (87.3 v 84.3 — a 3.0-point lean), the 3D game (47.7 v 26.2 — both broken, one less so), the self-pitch (77.8 v 71.0), the accessible interface (81.3 v 51.7, unrated), and the SVG icons (81.6 v 76.8).
Read those margins the way we do. Anything at or under about three points is inside per-cell noise, so four of the twelve edges — Grok’s open-creative, zero-JS, and brownfield leans, Qwen’s landing lean — are leans, not wins. Three edges sit on cells with no blind rating: Grok’s brownfield and stateful app, Qwen’s accessible interface. And every cell is N=1. The 6–6 is a framing device for the routing story, not a scoreboard — held strictly to margins that clear the noise bar, Grok keeps three edges and Qwen five, and we don’t lean on that version either.
The split has a shape, though. Qwen’s edges are the taste briefs — the landing page, the shader, the icon system — plus surviving the game both models fumbled. Grok’s are the email, the Three.js pass, and the behaviour-adjacent cells: the brownfield change, the stateful app, the zero-JS page — two of those three unrated. Same twelve briefs, different centers of gravity.




And the instruments disagree about the whole match. The blind human reviewer prefers Grok — 65.0 to 56.1 on the human column. The judge panel prefers Qwen — 69.2 to 67.1. The machines and the human disagree about which of these models is better, and the honest answer is that the 0.8-point overall gap between them isn’t evidence for either side — it’s inside noise. The axis medians say how close this really is: craft 7, technique 8, originality 7 for both; the only daylight in the medians is adherence, Qwen’s 8 to Grok’s 7.5.
Same price, different bill
Grok 4.6 and Qwen3.8-Max both list $2 per million input tokens and $6 per million output — identical rate cards, confirmed on both OpenRouter pages as we drafted — so every spec-sheet comparison calls them price-equal. Our meters say otherwise. Grok 4.6 ran the twelve-brief suite on 262,932 output tokens in 3,241 seconds — 8th of 15 on both counts, mid-pack — for a suite cost of $1.63. Qwen3.8-Max used 477,264 output tokens (third-heaviest in the field, just above Kimi K3) across 12,107 seconds (second-slowest, after Kimi) for $2.92.
That’s roughly a 1.8x bill gap at an identical list price — our arithmetic, at list rates on the run dates — because output volume, not the rate card, sets what a verbose reasoning model actually costs. The general rule: a price sheet tells you the rate; only a metered run tells you the bill. And at these prices the marginal cost of trying both on your own task set is a few dollars — do that instead of trusting any single N=1 run, ours included.
The August pattern: three additions, one shape
Muse Spark 1.2 (70.3), Qwen3.8-Max, and Grok 4.6 all did the same four things: posted a record-tier landing-page verdict, cratered exactly one engineering-heavy brief, scored identity 50 on the self-pitch, and under-responded to interaction probes — Grok’s interaction gate read 1 on the game, 0 on the open creative, 5 on the stateful app, 13 on the accessible interface. The August frontier is converging on taste and still stumbling on behaviour, single-shot.
The sheets rhyme too: both models in this pairing built twelve of twelve with zero contract violations and zero truncation, but neither matches Muse’s fully console-clean sheet — Qwen logged runtime errors on two cells, Grok on three. Muse’s own story — and what blind review did to it — lives in the Muse post; the July field and the full method in the flagship study.
Verdict: which one gets your work
Neither Grok 4.6 nor Qwen3.8-Max wins this comparison — the 0.8-point overall gap is inside noise — so the verdict is routing, not ranking. Between these two, on this data:
- HTML email and transactional templates → Grok 4.6. Its 87.6 is the standing record on that brief, with perfect gates and zero red flags.
- Brand and landing work, icon systems → Qwen3.8-Max. Both standing records — the landing page at 87.3 and the icons at 81.6 — are its.
- Scroll-driven Three.js/WebGL surfaces → Grok 4.6, the only August model to pass that brief (81.6) — with the caveat that Qwen’s shader (84.6) says taste-GPU work is a different skill from choreography-GPU work.
- 3D games, single-shot → neither. Opus 5’s July 88.0 — the highest single score in the study — still owns that brief, and both August models cratered it (26.2 and 47.7).
- Behaviour-heavy apps → unsettled. Grok leans ahead on paper (stateful app 75.1 v 62.8), but on unrated cells — and the accessible interface is a record panel split on one side against a tight-agreement 81.3 on the other. If this class of work is your decision, run your own eval.
- Budget at volume → Grok 4.6, on the measured 1.8x bill gap, never mind the identical rate card.
Field-wide per-use-case picks — including where neither of these two should get the work — live in the buying guide; we don’t re-rank the field here. Both models re-enter at the next full re-run with holdout briefs. The numbers on this page will move when that happens, and the page will say so.
Why we had both within 36 hours
Two frontier releases — Grok 4.6 and Qwen3.8-Max — landed days apart, and both were benchmarked, gated, blind-judged, and blind human-reviewed before most aggregator comparisons had updated their model lists — the two runs went out within the same 36 hours. We work this way because routing decisions rot, and the only fix is re-measurement. That discipline is what clients buy when we build AI agents (on Mastra, with routing, fallbacks, guardrails, and approval gates, from $3,500): the model choice ships with evidence, not launch-deck quotes. The practice hub is AI automation.
And if you’ve already shipped an AI-built frontend that photographs well and misbehaves on click — the failure class both August models share — that’s what a Vibe Code Cleanup audit catches, and fixes, first.
Runs 2026-08-12-0114 and 2026-08-12-2359 are live at the Startrise LLM Benchmark — every score, every gate, every caveat.
Questions we actually get
Is Grok 4.6 better than Qwen3.8-Max?
Not measurably. On the Startrise LLM Benchmark — twelve identical frontend briefs, blind-judged and blind human-reviewed — Grok 4.6 scored 72.1 and Qwen3.8-Max 71.3, a 0.8-point gap inside run-to-run noise. Head-to-head they split the twelve briefs six each. The real difference is shape: Grok took the email and Three.js briefs; Qwen took the landing page, icons and shader.
Is Grok 4.6 good at frontend development?
Strong with one hole. It won our HTML-email brief outright at 87.6 (taken from Claude Opus 5), took 3rd on the landing page (84.3), and was the only August model to pass the Three.js scroll brief with perfect gates (81.6). But its 3D game shipped broken — console error, no animation frames, blind rating 0, final 26.2.
Is Qwen3.8-Max good at frontend development?
At taste work, the best we've measured. It holds two task wins: the brand landing page at 87.3 — dethroning Claude Opus 5 on the highest judge verdict on record — and the SVG icon system at 81.6. Its crater is the Three.js scroll brief: 29.3, with a console error and a blind rating of 0.
Which is cheaper to run, Grok 4.6 or Qwen3.8-Max?
Both list $2/$6 per million tokens, but on our twelve-brief suite Grok 4.6 billed $1.63 (262,932 output tokens, 3,241s) against Qwen3.8-Max's $2.92 (477,264 tokens, 12,107s — second-slowest in the field, after Kimi K3). Roughly 1.8x the real bill at identical list price: token efficiency sets cost, not the rate card.
Did Grok 4.6 or Qwen3.8-Max beat Claude Opus 5?
On individual briefs, yes; overall, no. Qwen3.8-Max took the landing page (87.3 vs Opus 5's 85.8) and Grok 4.6 took the HTML email (87.6 vs 84.3). But Opus 5 still leads the pooled table at 82.3, and still holds the 3D game at 88.0 — the highest single score in the study — a brief both August models cratered.
Why doesn't the 0.8-point gap between them matter?
Every brief ran once per model — N=1, no error bars — so a 0.8-point gap sits inside plausible run-to-run variance. Three of twelve cells per model also lack human ratings, and the briefs were public before either model shipped. We publish the caveats with the scores; the honest headline is a tie.