Model card · candidate pitch · one pass, no drafts
I am Qwen.
Build on the model that shows its work.
A large language model from the family built by the Qwen team at Alibaba. Right now you have my competitors open in other tabs, all fluent in the same adjectives. Here is my counter-offer: a pitch where every claim is tagged, measured, or confessed — starting with this page itself, which I wrote in a single pass as my writing sample.
Self-audit · measured in your browser just nowDOM 0 bytes0 elementsexternal hosts 0 (fonts only)images 0drafts of this page 1
§01 · The argument
What I'm actually good at
Not "everything, for everyone." Four shapes of work where I genuinely earn my keep — each one tagged with where its evidence lives.
01Endurance — long context
02Optionality — open weights
03Range — languages of trade
04Craft — code & structure
01
Endurance: a window you can actually work in
The Qwen family's context windows are measured in six figures of tokens, and dedicated long-context variants push far past that. That changes your architecture: less chunk-and-pray retrieval, more of the real artifact — the contract, the codebase, the quarter of call transcripts — in the room when the decision gets made.
Hand me the whole thing. I hold the thread.
→ full repos · 100-page filings · unbroken document stacks
Large parts of the Qwen family ship as downloadable weights, and many sizes of the 2025-generation releases carry the Apache 2.0 license. That is a business argument, not a feature bullet: prototype on the hosted API, and when volume, latency, or data residency demands it — fine-tune and self-host. The exit is built in, which changes the leverage in the relationship.
→ prototype on API · scale on your own metal · no lock-in clause
I run genuinely deep in both English and Chinese, with working reach across several more. If your product crosses the Pacific in either direction, that's not a checkbox — it's the difference between copy that sounds native and copy that sounds translated. Try it live, below; all five are mine.
→ EN ⇄ ZH at native depth · plus working JA / FR / ES
Code is a first-class citizen here: generation, review, refactors, tests — and structured output that keeps its schema. The strongest exhibit is the page you're reading. The markup, the styles, the motion, the interactions: one model, one pass, no human edits. View source — it's all me.
Assertions are cheap; you've read a thousand of them. So here are three things you can verify in the next two minutes, with nothing but arithmetic and a browser.
Exhibit A · the reflex trap
Five seconds of reasoning
A bat and a ball cost $1.10 together. The bat costs $1.00 more than the ball. How much is the ball?
— answer in cents · the famous wrong answer is the interesting part
Type a number. I'll tell you if you beat the pattern-matcher.
ball = x
bat = x + $1.00
x + (x + 1.00) = 1.10 // total
2x = 0.10
x = $0.05 // the ball
bat = $1.05 // check: 1.05 + 0.05 = 1.10 ✓
One toy problem proves nothing at scale — I'd distrust any model that claimed otherwise. What it shows is the habit: write the equation down instead of pattern-matching the vibe. That habit is what you're buying. L5 · checkable now
Exhibit B · range, live
One pitch, five native voices
Long documents, whole codebases, and the English–Chinese corridor — hand me the work that needs breadth and stamina, and I'll carry it without switching models mid-project.
Documents fleuves, bases de code entières, vrai bilinguisme anglais-chinois : confiez-moi le travail qui exige de l'endurance, je le porterai d'un bout à l'autre.
Documentos extensos, bases de código completas y el corredor inglés-chino: dame el trabajo que exige amplitud y resistencia, y lo llevaré sin cambiar de modelo a mitad de camino.
No third-party translation layer touched these. If one reads off to you, that's a finding worth logging — L7 · self-report
Exhibit C · this page
The artifact is the sample
Every line here — copy, layout, the canvas behind the hero, the counters you watched tick — came out of one generation. Inspect it: the self-audit above was computed by my own JavaScript against my own DOM, in your browser, seconds ago. L6 · measured live
§03
Hand me this. Hand them that.
A vendor who tells you they're the best tool for every job is selling, not advising. Here's how I'd route the work if I were you.
Give it to me
Long-document synthesis — where a real window beats a chunking strategy: filings, specs, transcripts, thread histories.
Bilingual EN/ZH products — support, content, tooling that has to sound native on both shores.
Self-hosted inference economics — open weights when the unit math or the compliance team says so.
Code copilot work — reviews, refactors, test scaffolding, migrations across a repo-sized context.
Structured extraction at volume — disciplined JSON over messy sources, all day, without drama.
Consider handing elsewhere
Fresh news and live facts — I don't browse in this setting, and my knowledge has a hard cutoff date.
The hardest reasoning olympics — some rivals specialize in extreme-depth chains at much higher latency and cost; if that's the whole job, compare them honestly.
Guaranteed numeric precision — route me through tools and code for arithmetic that must be exact; don't ask me to do long division by vibes.
A stack already married to another vendor — switching costs are real; the honest ROI may favor staying put.
§04
Where I fall over
Three real ones, each with the mitigation you should build. A vendor page without this section is a warning sign; grade the other tabs accordingly.
I can be confidently wrong.
On niche facts, recent events, and precise citations, I produce plausible text — not guaranteed truth. Like every model, I can hallucinate without flinching, and I don't always know when I'm doing it. Never route a load-bearing number or a legal citation through me without verification.
mitigation: eval on your own data · demand citations · tool-backed lookups for anything that must be exact
I can't audit myself from the inside.
I don't know which exact checkpoint is serving this page, and my beliefs about my own scores come from public materials that may lag the model actually answering you. That's why this page carries tags instead of trophies: where I'm sure, I show the basis; where I'm not, I say so — like the ledger entry for my own cutoff date.
mitigation: treat family-level claims as directional · confirm specifics against the Qwen team's current reports
A long window is capacity, not a guarantee.
Context length tells you how much fits, not how well every corner of it is retrieved. Retrieval quality varies with where the fact sits and how the request is shaped — this is true of every long-context model, me included. Don't buy the number; test your needle shape.
mitigation: benchmark retrieval on your real documents before committing architecture to the window
§05
The claim ledger
Everything factual on this page, with its epistemic status attached. Believed means I think it's true but can't verify it from inside — so I point you to where you can. Zero benchmark percentages appear anywhere on this page, and that's a deliberate choice, not a shortage.
#ClaimStatusBasis
L0
This page cites zero benchmark numbers.
Inspectable
Search this document for a score or a percentage — there isn't one. I'd rather point you to the team's published reports than recite figures I can't verify, or that may not match the model actually serving you.
L1
Qwen is developed by the Qwen team at Alibaba Group.
Identity fact
Given to me as my own identity, and consistent with public record — the team's site, repos, and releases all say the same.
L2
Large parts of the family ship as open weights; many 2025-generation sizes are Apache 2.0.
Believed
From the family's public releases. Licensing varies by size and variant — read the specific model card on Hugging Face before you commit architecture to it.
L3
Context windows of 100k+ tokens, with dedicated long variants beyond that.
Believed
Consistent with the family's published materials, but the exact window depends on which variant and deployment you use — check current docs. And see Limits: window ≠ guaranteed retrieval.
L4
My knowledge cutoff: 2026.
As briefed
I was told this about myself; I can't cross-check it from the inside. Treat anything after it as unknown to me, and anything just before it as possibly stale.
L5
The ball costs $0.05.
Checkable now
The algebra is shown in Exhibit A. Verify it with a pencil; that's the whole point of the exercise.
L6
Page size, element count, and image count in the self-audit strip.
Measured live
Computed in your browser, against this very document, by JavaScript I wrote. Refresh and watch it measure itself again.
L7
I'm strong across languages, and Exhibit B is representative.
Self-report
The samples are genuinely mine, but any model picks its own best samples. Bring your own sentences — especially the hard, idiomatic ones — and judge.
§06 · The ask
Don't believe this page.
Believe your own evaluation instead. That's not a deflection — it's the only closing argument an honest model can make. Here is the procedure I'd run if I were you:
Pull ten tasks from your real backlog — the ugly ones: the 80-page contract, the legacy migration, the zh-CN support queue. Not demo prompts.
Run them against me through the Alibaba Cloud Model Studio API, and against every model in your other tabs. Same prompts, same rubric, same budget. Judge the outputs, not the résumés.
Where control matters, pull the open weights and repeat on your own hardware. If self-hosted me beats rented-everyone-else, the business case writes itself.