GPT-6 Astra and Claude Opus 5.5 headline the thirteen new models the frontier shipped between mid-August and the end of September 2026, and on October 4 we ran every one of them through the Startrise LLM Benchmark: the same twelve single-shot frontend briefs, the same browser gates, the same blind three-lab judge panel and blind human reviewer that scored fifteen models in July and August. Run 2026-10-05-0412 is fifteen configurations (thirteen models plus effort-high re-runs of Claude Opus 5.5 and Claude Sonnet 5.5), 173 scored cells, 519 judge verdicts. Within the October run, Claude Opus 5.5 at effort high leads at 80.2, a hair over GPT-6 Astra at 79.9, with GPT-6.1 Sol at 78.6 and Claude Fable 5.1 at 77.1 behind them.
July’s Claude Opus 5 (82.3) still tops the pooled all-time table, and the paragraph on why that lead is softer than it looks is the most important one on this page. The most useful finding isn’t a ranking at all: at effort xhigh, Opus 5.5 and Sonnet 5.5 each returned nothing on three briefs because they spent the entire 128k output ceiling thinking.
Nobody else has this data. Every page ranking for “GPT-6 Astra vs Claude Opus 5.5” right now re-plots two vendors’ launch tables or quotes an aggregator index. This is the only place these thirteen models have built the same real deliverables under the same instruments as the fifteen before them, and as far as we can find, the only blind human review of any of them anywhere. The per-model stories below are the digest; the benchmark page carries every cell, gate, verdict and live build.
Key Takeaways
- Within the October run, Claude Opus 5.5 (effort
high) 80.2 edges GPT-6 Astra 79.9 — inside noise. July’s Opus 5 (82.3) still leads all-time, but only on a human column rated in a different session; on gates and judges, both October leaders beat it- GPT-6.1 Sol is the value story: 78.6 for $1.66, about a sixth of Astra’s $9.67, with the cleanest source in the field (4 red flags)
- Effort
xhighis a trap at 128k output: Opus 5.5 and Sonnet 5.5 each hitmax_tokenswith nothing returned on 3 of 12 briefs. Efforthighfixed Opus (12/12, 1st); Sonnet (high) delivered 12/12 but shipped a prose-wrapped non-answer on the open brief- Claude Fable 5.1 posted the highest single cell on the board (91.4, Three.js) on the brief Fable 5 failed at 22.5 — and the most expensive bill on record, $27.52
- DeepSeek V4.1 Flash (75.6, $1.82) is the surprise and Gemini 3.8 Flash (64.6, 14th of 15) the disappointment; the new generation took exactly half the per-task wins, six of twelve briefs
October 2026 leaderboard: GPT-6 Astra vs Claude Opus 5.5
Fifteen configurations, ranked within the October run. The pooled all-time rank (30 rows, every run since July) is in the last column, because the two tables disagree in ways that matter.
| # | Model | Overall | Gates | Judge | Human | Built | All-time |
|---|---|---|---|---|---|---|---|
| 1 | Claude Opus 5.5 (high) | 80.2 | 94.4 | 82.9 | 58.9 | 12/12 | 2nd |
| 2 | GPT-6 Astra | 79.9 | 94.4 | 81.7 | 58.9 | 12/12 | 3rd |
| 3 | GPT-6.1 Sol | 78.6 | 94.7 | 80.0 | 57.2 | 12/12 | 5th |
| 4 | Claude Fable 5.1 | 77.1 | 94.4 | 77.5 | 58.9 | 12/12 | 6th |
| 5 | Claude Sonnet 5.5 (xhigh) | 75.9 | 91.2 | 73.9 | 68.3 | 9/12 | 7th |
| 6 | Grok 4.7 | 75.9 | 92.9 | 78.8 | 49.4 | 12/12 | 8th |
| 7 | DeepSeek V4.1 Flash | 75.6 | 91.9 | 76.5 | 55.0 | 12/12 | 9th |
| 8 | Muse Spark 1.3 | 74.7 | 95.5 | 70.8 | 56.3 | 12/12 | 10th |
| 9 | GLM 5.3 | 73.2 | 94.4 | 70.5 | 55.0 | 11/12 | 12th |
| 10 | Claude Sonnet 5.5 (high) | 72.8 | 90.2 | 74.8 | 51.3 | 12/12 | 13th |
| 11 | Claude Opus 5.5 (xhigh) | 71.1 | 86.4 | 71.7 | 55.7 | 9/12 | 17th |
| 12 | Qwen 3.8 Max 0902 | 69.8 | 89.5 | 72.7 | 39.4 | 12/12 | 19th |
| 13 | GPT-6 Luna | 69.3 | 90.6 | 69.8 | 43.3 | 12/12 | 20th |
| 14 | Gemini 3.8 Flash | 64.6 | 86.5 | 62.5 | 45.0 | 12/12 | 26th |
| 15 | DeepSeek V4 Pro 0813 | 62.7 | 88.5 | 57.1 | 46.1 | 12/12 | 28th |
Run 2026-10-05-0412. Blend: gates ×0.25 + judge ×0.45 + human ×0.30; where a cell has no human rating, that weight is redistributed across the other two. “Built” counts cells that returned a scorable file — a row with 9/12 is averaged over nine cells, so the three empties are not counted as zeros, which flatters it. The all-time column is the pooled 30-row table at the benchmark, where July’s Claude Opus 5 (82.3) and Kimi K3 (79.8) hold 1st and 4th.
How did we test the October drop?
Run 2026-10-05-0412 used the harness the July study documented and the August addenda reused: twelve frontend briefs (a Three.js scroll journey, a WebGL shader, a brand landing page, a 3D game, an open creative brief, a self-pitch page, a WCAG seat map, a brownfield change request, an SVG icon system, a stateful app, a zero-JS page and an HTML email), each answered single-shot as one self-contained HTML file, driven through automated gates in headed Chromium, then scored blind by the same three-judge panel: Claude Opus 5 on Craft, GPT-5.6 Terra on Engineering, Grok 4.5 on Brief. Same 25/45/30 gates/judge/human blend. The briefs, the gates and the published bug list live in the flagship study.
Thirteen models entered, in release order: the DeepSeek V4 Pro 0813 GA snapshot (August 13) and GLM 5.3 (announced August 14); Claude Fable 5.1 (September 1); Muse Spark 1.3, Gemini 3.8 Flash and the Qwen 3.8 Max 0902 snapshot (September 2); GPT-6 Astra (September 3); DeepSeek V4.1 Flash (September 10); Grok 4.7 (September 21); Claude Opus 5.5 and GPT-6 Luna (September 22); Claude Sonnet 5.5 (September 28); and GPT-6.1 Sol (September 29), which replaced the GPT-6 Sol of a week earlier. Two more configurations joined mid-run, for reasons the effort section explains: Opus 5.5 and Sonnet 5.5 re-run at effort high. That’s fifteen rows, 180 cells attempted, 173 scored.
Two models people will ask about aren’t here. GPT-6 Terra does not exist; GPT-5.6 Terra remains the current Terra and still sits on our judge panel. And Gemini 4 Argon, announced September 30, has no public API model id; the Gemini section below explains. Also unreleased as of October 5: Grok 5, Qwen 4 (previewed at Apsara on September 22 as “in training”), Kimi K4, GLM 5.5 and DeepSeek V4.1 Pro. Kimi K3, from July, is still Moonshot’s latest and still 4th all-time at 79.8.
Blind human review landed on 127 of those 173 cells, rated by one reviewer with model names hidden until every cell in a task was scored. Three briefs were not human-rated in October at all: the accessible seat map (08), the brownfield change (09) and the stateful app (11). Every cell on those three briefs is auto-only, with the human weight redistributed, and every podium on them is what the August posts called a paper podium. We flag each one below.
Read this before the scores
Five caveats govern every number on this page. They are findings, not fine print, and the first one changes how you should read the headline.
The all-time #1 rests on a human column from a different session. Opus 5.5 at effort high beats July’s Opus 5 on gates (94.4 vs 93.9) and on the blind judge panel (82.9 vs 81.3). GPT-6 Astra matches it on gates (94.4) and edges Opus 5 on the judges too (81.7). What keeps Opus 5 on top is the human column: 71.7 in July against 58.9 for both October leaders. The July ratings and the October ratings were done blind, by the same reviewer, months apart, and we don’t normalise one session against the other. A single rater’s scale drifts. So read within-run ranks as the firmer number and treat any cross-run gap under about three points as noise.
Opus 5 at 82.3, Opus 5.5 (high) at 80.2 and Astra at 79.9 are, on the evidence we have, tied.
N=1, no error bars. Every brief ran once per configuration. The 0.3 between Opus 5.5 (high) and Astra, the 0.3 between Grok 4.7 and DeepSeek V4.1 Flash: noise.
Three briefs carry no human rating. The seat map, the brownfield change and the stateful app. Two of the six wins the new generation took sit on them, and a third (Muse’s icon system) is itself an unrated cell.
Every judge has a relative in this field. Opus 5, the Craft judge, is a same-vendor sibling of Opus 5.5, Sonnet 5.5 and Fable 5.1. GPT-5.6 Terra, the Engineering judge, is the predecessor family of Astra, Sol and Luna. Grok 4.5, the Brief judge, is two versions behind Grok 4.7. The panel is blind, and the 839-call audit measured Opus 5’s vendor bias at −0.46 (negligible) and its self-bias at +2.87, Terra’s vendor bias at +0.50, Grok 4.5’s self-bias at −0.93. None of the October models is in that audit yet. We disclose the conflict with its measured size rather than pretend it isn’t there.
The contamination window is wide open. The twelve briefs have been public since July 26, 2026, and all thirteen October models shipped after that date. No evidence any of them saw the briefs; the possibility exists for every one, which is why the next full re-run uses holdout briefs.
Is Claude Opus 5.5 better than GPT-6 Astra?
Not measurably, and the two models get to 80 by opposite routes. Opus 5.5 at effort high scored 80.2 on 794,864 output tokens across the suite, 66.2k per brief on average, in 127 minutes of wall clock. Astra scored 79.9 on 188,589 output tokens, 15.7k per brief, in 46 minutes. Astra reaches the top of the October field with the smallest output of any frontier model in the run, roughly a quarter of Opus 5.5’s volume, and it did it at $10/$50 per million tokens for a $9.67 suite bill against Opus 5.5’s $16.05 at $4/$20. The cheaper rate card produced the dearer bill. We’ve written that sentence about Anthropic models before; it keeps being true.
Where they differ is shape. Astra won the stateful app outright at 82.0 (an unrated cell), took 2nd on the seat map at 90.1 (unrated) and 2nd on the brownfield change at 79.7 (unrated), and posted an 83.2 on the open creative brief, 4th all-time and the best October score on that brief. Its Three.js scroll journey ran 174 seconds and 11,067 tokens and scored 82.8 with judge 85. The judges liked the engineering; the blind reviewer gave it 65. Astra’s judge-named red flag count across twelve files was 11, the same as Kimi K3’s, the July study’s clean-code benchmark.

Opus 5.5 (high) is the generalist. It podiumed three times: 2nd on the self-pitch (81.9), 2nd on the zero-JS page (86.4, judge 90), 3rd on the seat map (87.1, unrated), and sat within a point of the podium on the icon system (79.8) and the stateful app (79.6, unrated). Its worst cell was the 3D game at 74.9, which is a telling kind of worst: no crater anywhere in twelve files. Twelve red flags across the suite. The axis medians (craft 8, technique 8.5, adherence 9, originality 8) match July’s Opus 5 exactly.
So which one? If you bill by output token, Astra. If you want the model least likely to hand you a broken file on a brief you didn’t anticipate, Opus 5.5 at high. On this data, the 0.3 between them is not a reason to pick either.
Why did Opus 5.5 and Sonnet 5.5 return nothing at effort xhigh?
This is the finding we’d put in front of any team shipping Claude Opus 5.5 or Sonnet 5.5 today. At effort xhigh, Claude Opus 5.5 and Claude Sonnet 5.5 each exhausted the 128k max_tokens ceiling on three of twelve briefs, with stop_reason: max_tokens and nothing usable in the response: every token went to thinking. Opus 5.5 returned nothing on the icon system, the stateful app and the zero-JS page, and truncated two more mid-file: the Three.js scroll journey and the seat map. Sonnet 5.5 returned nothing on the open creative brief, the icon system and the zero-JS page, and truncated the stateful app. 128k is the hard output ceiling for these models. It cannot be raised, so there is no max_tokens fix.
The truncated files are what a ceiling looks like from the outside. Opus 5.5’s Three.js build ran 1,238 seconds, emitted exactly 128,000 tokens, and stopped without a closing </html>. The gate browser loaded a near-black page: gates 45.2, judge 17.5, eight red flags, blind rating 0, final 19.2. Sonnet 5.5’s stateful app did the same and scored 20.5, the lowest cell on that brief across all thirty rows. Unfinished work and bad work score identically single-shot.

high: Opus 5.5’s Accretion, 71,698 tokens, perfect gates, judge 87.5, zero red flags, final 79.4. At xhigh the identical model spent 128,000 tokens and returned a near-black page that scored 19.2 — open that build to see what a ceiling looks like. Open the live build.So we re-ran both at effort high. Opus 5.5 delivered 12 of 12 and became the best model of the October run. Its average output fell from 96.5k tokens per brief to 66.2k, its wall clock from 138 minutes to 127, and its bill from $17.50 to $16.05, for nine more points. Sonnet 5.5 at high also delivered 12 of 12, with output down from 92.8k to 51.8k per brief, but it shipped a 5kb prose-wrapped non-answer on the open creative brief: contract violation, judge 17.5, eleven red flags, blind rating 0, final 28.0, last on that brief all-time. Its overall landed at 72.8, three points under its xhigh row, which is averaged over the nine cells that came back.
The nuance that matters for routing: where xhigh did finish, it was better. On the seven briefs Opus 5.5 completed cleanly at both settings, xhigh out-scored high on six (our arithmetic on the published finals: a mean of 81.6 against 78.7). Its landing page at xhigh scored 86.5, 2nd all-time, and its self-pitch took the brief outright at 83.8 with an 85 blind rating. More thinking produced better pages. It also produced no page at all one time in four. For single-shot work under a 128k output ceiling, high is the setting; xhigh belongs in a harness that can detect max_tokens and retry.

xhigh shader, Marginal Regeneration, judge 92.5, perfect gates, 120 fps, zero red flags, final 84.6 — 2nd all-time on the brief. The open creative brief at high went the other way: a prose-wrapped fragment that scored 28.0, on record here. Open the live build.One more number from the same ceiling: Grok 4.7’s seat map came back at 127,344 output tokens, 656 short of the wall, and scored 85.8. The frontier is writing right up to the edge of what it is allowed to return.
Is GPT-6.1 Sol the best-value model for frontend work?
GPT-6.1 Sol scored 78.6, 3rd in the October run and 5th all-time, on a $1.66 suite bill. That is about a sixth of Astra’s $9.67 for 1.3 fewer points, and roughly a seventeenth of Fable 5.1’s $27.52 for 1.5 more. Per 100 points of score, Sol costs $2.11; Astra $12.10; Opus 5.5 (high) $20.01. Sol lists at $2/$10 per million tokens and wrote 160,928 output tokens for the suite, 13.4k per brief, in 49 minutes. Its gate score (94.7) trails only Muse Spark 1.3’s. And its source drew four judge-named red flags across twelve files, the cleanest sheet on the entire 30-row board, under Kimi K3’s and Astra’s 11.
It also won a brief. Sol’s zero-JS page scored 88.3, 1st all-time, on judge 87.5 and the highest blind rating Sol received (80), displacing Opus 5’s 84.8. It took 2nd on the stateful app at 80.7 (unrated) and ran a perfect-gate Three.js scroll journey at 81.6 on 9,159 tokens. Its weak cells are taste cells: the shader (75.3, blind rating 40), the landing page (76.0, rating 50), the open creative (73.9, rating 50). The judges scored those three 85, 80 and 82.5; the human didn’t agree. Sol’s axis medians are craft 8, technique 8, adherence 9, originality 7, and the originality 7 is where the human column lives.

Sol superseded GPT-6 Sol within a week of that model’s release, so our run is the first one of its kind on the 6.1 snapshot. For a team buying frontend output by the thousand, this row is the one to test first.
How good is Claude Fable 5.1, and what does it cost?
Claude Fable 5.1 scored 77.1, 4th in October, 6th all-time, 3.3 points over Fable 5’s July 73.8, and it owns the highest single cell on the board. Its Three.js scroll journey scored 91.4: perfect gates, judge 87.5, zero red flags, and a 90 blind rating, the highest human rating of the October run. In July, Fable 5 shipped that same brief with one JavaScript syntax error, a near-black render and a 22.5. The lineage table records the swing as +68.9 on one cell.


The judges still rate Fable’s taste at the top of the field. On the landing page and the open creative brief it drew judge 85 on both, zero red flags on both, and finished 82.8 and 81.3. But it no longer owns every open brief: Fable 5’s July records on the shader (87.3) and the open creative (85.7) stand, and within October Astra out-scored Fable 5.1 on the open creative, 83.2 to 81.3. Its crater is the stateful app at 50.5, a split-panel cell with judge 42.5, on 83,570 output tokens, its longest file of the run. Nine red flags across the suite, second only to Sol.
Then the bill. Fable 5.1 lists at $10/$50 per million tokens and wrote 543,056 output tokens, so the suite cost $27.52: the most expensive row on record, $35.69 per 100 points, against $2.11 for Sol and $12.10 for Astra. In July we noted that Fable’s double rate card produced only a 1.18x bill because it wrote fewer tokens than Opus 5; this time it wrote 16% more than Fable 5 did and the bill went up 15%. Fable 5.1 is the model you route the one brief that has to be beautiful to. It is not the model you route the other eleven to.
Is DeepSeek V4.1 Flash the surprise of the October run?
DeepSeek V4.1 Flash scored 75.6, 7th in the October run and 9th all-time, for a $1.82 suite bill at the rates recorded on the run date (DeepSeek’s published pricing is tiered by peak and off-peak hours and has changed since, so check the page before you budget). That is a 20.3-point jump over the DeepSeek V4 Pro that finished 11th of 12 in July at 55.3, the largest generation-over-generation delta on the board, with the cross-run caveat applied. Its judge column (76.5) sits 6th in the October field, ahead of Sonnet 5.5 at both settings, and its axis medians (craft 8, technique 8, adherence 8, originality 7) are a frontier profile. The blind reviewer was cooler: 55.0 on the human column.
The sheet is not clean. V4.1 Flash wrapped six of its twelve deliverables in prose, the most contract violations of any model in the run, and the harness had to extract the HTML. It emitted 759,560 output tokens, the second-heaviest footprint among the non-Anthropic October models after Grok 4.7. Where it landed, though, it landed well: the zero-JS page at 81.3, the seat map at 80.4 (unrated), the stateful app at 77.5 (unrated), the email at 80.5. On the icon system its judge verdict of 85 tied the best any model drew on that brief in October; a 45 blind rating pulled the final to 76.8. Thirty-nine red flags across the suite is mid-table.

The other DeepSeek row went the other way. The V4 Pro 0813 GA snapshot scored 62.7, last of fifteen, on 82 minutes of wall clock and 54 red flags; its stateful app drew thirteen flags and a 42.1. That is still 7.4 points over July’s V4 Pro. Same vendor, same run, thirteen points apart, and the cheaper one is on top.
Why is Gemini 3.8 Flash 14th, and where is Gemini 4 Argon?
Gemini 3.8 Flash scored 64.6, 14th of 15, with 72 judge-named red flags, the most of any October model and the third-highest on the whole board behind July’s DeepSeek V4 Pro (75) and Claude Haiku 4.5 (114). It was quick (24 minutes, second only to GPT-6 Luna’s 17) and cheap at $0.97, and its best cells were the email (78.6) and the brownfield change (78.0, unrated). Its worst was the stateful app at 44.4 with ten flags. The blind reviewer gave its Three.js build a 20. Five of its twelve cells split the judge panel. The axis medians read craft 7, technique 6.5, adherence 7, originality 6: among the thirteen new models, only the DeepSeek V4 Pro snapshot has a lower technique median.
Why isn’t Gemini 4 Argon here? Because it isn’t callable. Google announced Argon on September 30, 2026, and as of October 5 it has no public API model id; the first rollout is restricted to vetted defenders in Google’s Fairwind cyber-defense program, with paid API access “to follow” and no date. Gemini 3.8 Flash is the strongest Gemini an engineering team can actually put in a pipeline today, and Google has shipped no Pro-tier model since the 3.1 Pro preview in February 2026. Argon joins the board the day it gets a model id.
Grok 4.7, Muse Spark 1.3, GLM 5.3 and the rest of the field
Before the rest of the table, the two cells from the run that are simply the most fun to open:

xhigh), Void Runner: a complete game loop at 120 fps, blind rating 85, final 86.5 — 2nd all-time on the brief. Play it.
Grok 4.7 is the quiet improver. It scored 75.9, 3.8 points over Grok 4.6, on the fourth-best judge column in the October run (78.8) and a top-tier axis profile (craft 8, technique 8, adherence 9, originality 8). The blind reviewer disagreed hard: 49.4 on the human column, where Grok 4.6 had drawn a 65. Its best cells are the seat map (85.8, unrated), the email (82.8) and the zero-JS page (82.7); its worst is the 3D game, where gates read 98.9 and the judges 72.5 but the human gave it a 10, for a 60.4. It also cost $5.26 at a $2/$6 rate card, 3.2 times Grok 4.6’s $1.63 at the same price, because it wrote 863,376 output tokens over 199 minutes.
Muse Spark 1.3 posted the highest gate score on the board (95.5) and won the icon system at 83.9, an unrated cell. It scored 74.7 overall, 4.4 over Muse Spark 1.2, for $0.71 in 31 minutes. The clean sheet is gone: two contract violations (no fenced HTML block on the Three.js and stateful briefs) where 1.2 had none, and 43 red flags. The stateful app is still its weak brief at 65.8 with eleven flags.
GLM 5.3 scored 73.2, 1.1 over GLM 5.2, which is noise, on 11 of 12 cells: its icon system came back empty: it hit the 128k output ceiling while reasoning. Its judge column rose from 65.6 to 70.5; its open creative brief (80.9, blind rating 70) was its best cell. $1.03 for the suite.
Qwen 3.8 Max 0902 scored 69.8, 1.5 under August’s Qwen 3.8 Max, inside noise. It cratered the Three.js brief again (33.0, blind rating 0, as 3.8 Max did at 29.3) and drew the lowest human column of the October run at 39.4, while the judges scored it 72.7. 195 minutes of wall clock, behind only Grok 4.7’s 199 for the slowest October row.
GPT-6 Luna scored 69.3 for seven cents. $0.07 for all twelve deliverables at $0.10/$0.50 per million tokens, 17 minutes, $0.10 per 100 points, roughly an order of magnitude under anything else in the October run (Muse Spark 1.3 is next at $0.95). Its one crater was the 3D game (28.5, blind rating 0). If your frontend task is a template, the economics start here.
Generation over generation: who actually improved?
Fourteen lineage pairs connect an October row to its predecessor. The deltas below are cross-run comparisons, with the human-rating caveat from above: treat anything under about three points as noise.
| Predecessor | Successor | Delta |
|---|---|---|
| DeepSeek V4 Pro 55.3 | DeepSeek V4.1 Flash 75.6 | +20.3 |
| GPT-5.6 Sol 66.2 | GPT-6 Astra 79.9 | +13.7 |
| GPT-5.6 Terra 66.9 | GPT-6.1 Sol 78.6 | +11.7 |
| Qwen 3.7 Max 63.6 | Qwen 3.8 Max 71.3 | +7.7 |
| DeepSeek V4 Pro 55.3 | DeepSeek V4 Pro 0813 62.7 | +7.4 |
| Claude Sonnet 5 66.6 | Claude Sonnet 5.5 (high) 72.8 | +6.2 |
| Muse Spark 1.2 70.3 | Muse Spark 1.3 74.7 | +4.4 |
| GPT-5.6 Luna 65.0 | GPT-6 Luna 69.3 | +4.3 |
| Grok 4.5 68.0 | Grok 4.6 72.1 | +4.1 |
| Grok 4.6 72.1 | Grok 4.7 75.9 | +3.8 |
| Claude Fable 5 73.8 | Claude Fable 5.1 77.1 | +3.3 |
| GLM 5.2 72.1 | GLM 5.3 73.2 | +1.1 |
| Qwen 3.8 Max 71.3 | Qwen 3.8 Max 0902 69.8 | −1.5 |
| Claude Opus 5 82.3 | Claude Opus 5.5 (high) 80.2 | −2.1 |
Overall score, predecessor to successor, pooled table at the benchmark. Cross-run deltas; the blind human ratings for July and October were done in separate sessions.
The shape of the table is the story. OpenAI’s generation step is the largest for any vendor that was already competitive: the 5.6 family finished 6th, 8th and 9th of twelve in July, and the GPT-6 family finishes 2nd, 3rd and 13th of fifteen in October. Astra’s gain over GPT-5.6 Sol came almost entirely from the engineering briefs: the Three.js scroll went from 25.3 to 82.8, the shader from 32.5 to 74.5, the 3D game from 46.1 to 83.1, while the taste briefs barely moved (landing 80.2 to 78.6, open creative 83.8 to 83.2). Anthropic’s step is flat to negative at the top and real in the middle: Sonnet 5 to Sonnet 5.5 (high) is +6.2, Fable +3.3, Opus −2.1 with the human-session caveat. DeepSeek’s +20.3 is the biggest number on the board and comes with six prose-wrapped files.
Which model won each brief?
The October models took exactly half of the twelve per-task wins on the all-time board. Two of the six sit on unrated briefs, and a third is an unrated cell.
| Brief | Winner (all-time) | Score | Runner-up | Changed hands? |
|---|---|---|---|---|
| 01 Three.js scroll | Claude Fable 5.1 | 91.4 | Claude Opus 5 85.8 | Yes |
| 02 WebGL shader | Claude Fable 5 | 87.3 | Qwen 3.8 Max 84.6 | No |
| 03 Landing page | Qwen 3.8 Max | 87.3 | Claude Opus 5.5 (xhigh) 86.5 | No |
| 05 3D game | Claude Opus 5 | 88.0 | Claude Sonnet 5.5 (xhigh) 86.5 | No |
| 06 Open creative | Claude Fable 5 | 85.7 | GPT-5.6 Sol 83.8 | No |
| 07 Self-pitch | Claude Opus 5.5 (xhigh) | 83.8 | Claude Opus 5.5 (high) 81.9 | Yes |
| 08 Accessible seat map* | Claude Sonnet 5.5 (high) | 90.4 | GPT-6 Astra 90.1 | Yes |
| 09 Brownfield change* | Claude Opus 5 | 81.1 | GPT-6 Astra 79.7 | No |
| 10 SVG icon system* | Muse Spark 1.3 | 83.9 | Qwen 3.8 Max 81.6 | Yes |
| 11 Stateful app* | GPT-6 Astra | 82.0 | GPT-6.1 Sol 80.7 | Yes |
| 12 Zero-JS page | GPT-6.1 Sol | 88.3 | Claude Opus 5.5 (high) 86.4 | Yes |
| 13 HTML email | Grok 4.6 | 87.6 | Claude Opus 5 84.3 | No |
All-time podiums across 30 rows at the benchmark. Asterisked briefs carry no October human rating; Muse’s icon win has no human rating either. The six that held: Fable 5’s two open briefs, Qwen 3.8 Max’s landing page, Opus 5’s game and brownfield, Grok 4.6’s email.
Two patterns. First, the seat map podium is now entirely October and entirely unrated: Sonnet 5.5 (high) 90.4, Astra 90.1, Opus 5.5 (high) 87.1, all auto-only. Second, the two Sonnet 5.5 configurations now hold two of the twelve last places on the board: Sonnet 5.5 (high) on the open creative at 28.0 and Sonnet 5.5 (xhigh) on the stateful app at 20.5. The same family that posted a 90.4 posted the two deepest new craters, which is the effort story told as a scoreboard.
What does each model cost per 100 points?
The October run’s generation spend was about $101, plus about $11 of billed-but-empty Anthropic responses from the xhigh ceiling. The per-model bills below are recorded usage times list price on the run date; the last column divides the bill by the score, which is the number a procurement conversation should start with. Full per-model cost rows, with July and August for comparison, are at the benchmark.
| Model | Suite bill | Output tokens | Wall clock | Score | $ per 100 pts |
|---|---|---|---|---|---|
| GPT-6 Luna | $0.07 | 137,266 | 17 min | 69.3 | $0.10 |
| Muse Spark 1.3 | $0.71 | 159,558 | 31 min | 74.7 | $0.95 |
| GLM 5.3 | $1.03 | 226,292 | 30 min | 73.2 | $1.41 |
| Gemini 3.8 Flash | $0.97 | 253,864 | 24 min | 64.6 | $1.50 |
| DeepSeek V4 Pro 0813 | $1.28 | 250,863 | 82 min | 62.7 | $2.04 |
| GPT-6.1 Sol | $1.66 | 160,928 | 49 min | 78.6 | $2.11 |
| DeepSeek V4.1 Flash | $1.82 | 759,560 | 76 min | 75.6 | $2.41 |
| Qwen 3.8 Max 0902 | $3.02 | 494,459 | 195 min | 69.8 | $4.33 |
| Grok 4.7 | $5.26 | 863,376 | 199 min | 75.9 | $6.93 |
| Claude Sonnet 5.5 (high) | $6.28 | 621,053 | 76 min | 72.8 | $8.63 |
| Claude Sonnet 5.5 (xhigh) | $8.42 | 835,122 | 102 min | 75.9 | $11.09 |
| GPT-6 Astra | $9.67 | 188,589 | 46 min | 79.9 | $12.10 |
| Claude Opus 5.5 (high) | $16.05 | 794,864 | 127 min | 80.2 | $20.01 |
| Claude Opus 5.5 (xhigh) | $17.50 | 868,740 | 138 min | 71.1 | $24.61 |
| Claude Fable 5.1 | $27.52 | 543,056 | 108 min | 77.1 | $35.69 |
Run 2026-10-05-0412, list prices on the run date. For context, July’s Claude Opus 5 billed $20.27 ($24.63 per 100 points) and Kimi K3 $7.17 ($8.98).
Three readings. The rate card predicts almost nothing: Grok 4.7 and Qwen 3.8 Max 0902 list at the same $2/$6 as their August predecessors and billed 3.2x and 1.03x those predecessors’ suites; Astra at $10/$50 billed less than Opus 5.5 at $4/$20. Output volume sets the bill, and it’s on no pricing page. Second, the whole top four of the October run costs between $1.66 and $27.52 for a spread of 3.1 points: the price of the last three points is steep everywhere except Sol. Third, at these numbers the marginal cost of trying Sol, Luna, Muse Spark 1.3 and DeepSeek V4.1 Flash on your own task set is a few dollars. Do that rather than trust any single N=1 run, including this one.
What broke, and why we’re publishing it
Three things broke, and all three are in the run log. Running 36 concurrent Anthropic streams tripped the API’s credit pre-check, which multiplies worst-case cost per request by in-flight requests and refuses the batch if the balance can’t cover it; OpenRouter rejected judge calls the same way, with 402s, mid-panel. Twenty-two cells ended up with incomplete panels on the first pass and were re-judged with full three-judge panels before any score on this page was computed. Separately, the xhigh empties were billed: about $11 of the run’s spend bought responses with nothing in them, which is the cost of discovering the ceiling finding. None of this changes a number. We publish it because a benchmark that hides its operational failures is asking you to trust the parts you can’t see.
Verdict: routing the October field
Rankings aside, these are the rules we’d apply on this data, with the field-wide per-use-case picks living in the buying guide:
- Default frontier, single-shot, when the brief might be anything: Claude Opus 5.5 at effort
high. Twelve clean files, no cell under 74.9, three podiums. Neverxhighwithout a harness that catchesmax_tokens. - Default frontier when you bill by output token or need speed: GPT-6 Astra. Same gates, a quarter of the tokens, 46 minutes for the suite, and the stateful-app win.
- Volume frontend work: GPT-6.1 Sol. 78.6 for $1.66, four red flags, the zero-JS record. Test it before Astra.
- The one page that has to be beautiful: Claude Fable 5.1, if the budget tolerates $35.69 per 100 points. Route nothing else to it.
- Cheap and competent: DeepSeek V4.1 Flash, with a post-processor that unwraps prose; Muse Spark 1.3 for static pages and games that pass gates; GPT-6 Luna for templates at seven cents a suite.
- Behaviour-heavy apps (seat maps, stateful UI): unsettled. Every October podium on those briefs is unrated. If this class of work is your decision, run your own eval on it.
- Not yet: Gemini 3.8 Flash on this data, and Gemini 4 Argon until it has a model id.
Everything above moves at the next full re-run with holdout briefs, and the blind human review of briefs 08, 09 and 11 can move the October numbers before then. When it does, this page will say so.
Why we had thirteen models benchmarked in one weekend
Startrise re-benchmarks model routing every time the frontier moves, because routing decisions rot. Thirteen models shipped in seven weeks; by the weekend after the last of them every one had same-harness numbers, a blind panel verdict and a human rating instead of a launch-deck quote. That discipline is what clients buy when we build AI agents (on Mastra, with routing, fallbacks, guardrails and approval gates, from $3,500): the model choice ships with evidence, and the fallback path for a max_tokens stop is wired before anyone finds it in production. The practice hub is AI automation.
And if you’ve already shipped an AI-built frontend that photographs well and misbehaves on click, or a file that ends before its closing tag, that’s what a Vibe Code Cleanup audit catches first.
Run 2026-10-05-0412 is live at the Startrise LLM Benchmark, next to the July and August data: every score, every gate, every caveat, every file as generated.
Questions we actually get
Is GPT-6 Astra better than Claude Opus 5.5 for frontend development?
Not measurably. On the Startrise LLM Benchmark's October 2026 run — twelve identical single-shot frontend briefs, browser-gated and blind-judged — Claude Opus 5.5 at effort high scored 80.2 and GPT-6 Astra 79.9. A 0.3-point gap on an N=1 benchmark is noise. The real differences are cost and shape: Astra's suite bill was $9.67 to Opus 5.5's $16.05, and Astra won the stateful app while Opus 5.5 took the self-pitch and ran second on the zero-JS page.
What is the best LLM for frontend development in October 2026?
On our data, Claude Opus 5.5 at effort high (80.2) and GPT-6 Astra (79.9) lead the October field, with GPT-6.1 Sol (78.6) the value pick at $1.66 for the full twelve-brief suite — about a sixth of Astra's bill. July's Claude Opus 5 (82.3) still tops the pooled all-time table, but its lead rests on a human-rating column scored in a different session, so we treat the three as a tie.
Why does Claude Opus 5 still rank above Claude Opus 5.5?
Because of the human column. Opus 5.5 at effort high beats July's Opus 5 on browser gates (94.4 vs 93.9) and on the blind judge panel (82.9 vs 81.3), but the blind human rating sits at 58.9 against Opus 5's 71.7. Those ratings were done in separate sessions months apart, and a single rater's scale drifts. Within the October run alone, Opus 5.5 (high) is the best model; across runs, read the 2.1-point gap as noise.
Why did Claude Opus 5.5 and Sonnet 5.5 return empty responses at effort xhigh?
They spent the whole 128k output ceiling thinking. At effort xhigh both models hit the 128k max_tokens ceiling with stop_reason max_tokens on 3 of 12 briefs each — nothing usable returned — and truncated two more (Opus) and one more (Sonnet) mid-file. 128k is the hard output ceiling for these models and cannot be raised. Re-run at effort high, Opus 5.5 delivered 12 of 12 and finished first in the October run.
How much does GPT-6.1 Sol cost compared with GPT-6 Astra?
GPT-6.1 Sol lists at $2 per million input tokens and $10 per million output; GPT-6 Astra lists at $10/$50. In our October run Sol generated all twelve deliverables for $1.66 and scored 78.6; Astra cost $9.67 and scored 79.9. Per 100 points of score that is $2.11 for Sol against $12.10 for Astra, and Sol drew only 4 judge-named red flags — the cleanest source in the field.
Is Gemini 4 Argon in the Startrise LLM Benchmark?
No. Google announced Gemini 4 Argon on September 30, 2026, but as of October 5 it has no public API model id; access is restricted to Google's Fairwind cyber-defense program. The strongest callable Gemini is Gemini 3.8 Flash, which scored 64.6 — 14th of 15 in the October run — with 72 judge-named red flags, the most of any October model. Argon enters the board the day it becomes callable.
How good is Claude Fable 5.1 at frontend work?
Very good, and very expensive. Fable 5.1 scored 77.1, 4th in the October run, and posted the highest single cell on the entire 30-model board: 91.4 on the Three.js scroll journey, the brief its predecessor Fable 5 failed at 22.5 in July. It also drew only 9 red flags. But its suite bill was $27.52 — the most expensive row on record and about 17 times GPT-6.1 Sol's $1.66 for 1.5 fewer points.