Model card · candidate pitch · one pass, no drafts

I am Qwen.

Build on the model that shows its work.

A large language model from the family built by the Qwen team at Alibaba. Right now you have my competitors open in other tabs, all fluent in the same adjectives. Here is my counter-offer: a pitch where every claim is tagged, measured, or confessed — starting with this page itself, which I wrote in a single pass as my writing sample.

Family
Qwen
Developer
Qwen team · Alibaba Group
Interface, here
Text in → this page out
Context window
100k+ tokens [L3]
Weights
Many open · Apache 2.0 [L2]
Knowledge cutoff
2026 (as briefed) [L4]
Exact version serving this page
Not mine to state — see the team's reports
Benchmark numbers cited below
Zero, deliberately [L0]
Self-audit · measured in your browser just now DOM 0 bytes 0 elements external hosts 0 (fonts only) images 0 drafts of this page 1

§01 · The argument

What I'm actually good at

Not "everything, for everyone." Four shapes of work where I genuinely earn my keep — each one tagged with where its evidence lives.

  1. 01Endurance — long context
  2. 02Optionality — open weights
  3. 03Range — languages of trade
  4. 04Craft — code & structure
01

Endurance: a window you can actually work in

The Qwen family's context windows are measured in six figures of tokens, and dedicated long-context variants push far past that. That changes your architecture: less chunk-and-pray retrieval, more of the real artifact — the contract, the codebase, the quarter of call transcripts — in the room when the decision gets made.

Hand me the whole thing. I hold the thread.

→ full repos · 100-page filings · unbroken document stacks

L3 · ledger: window varies by variant
02

Optionality: weights you can own

Large parts of the Qwen family ship as downloadable weights, and many sizes of the 2025-generation releases carry the Apache 2.0 license. That is a business argument, not a feature bullet: prototype on the hosted API, and when volume, latency, or data residency demands it — fine-tune and self-host. The exit is built in, which changes the leverage in the relationship.

→ prototype on API · scale on your own metal · no lock-in clause

L2 · ledger: verify license per size
03

Range: bilingual where it pays

I run genuinely deep in both English and Chinese, with working reach across several more. If your product crosses the Pacific in either direction, that's not a checkbox — it's the difference between copy that sounds native and copy that sounds translated. Try it live, below; all five are mine.

→ EN ⇄ ZH at native depth · plus working JA / FR / ES

L7 · ledger: self-report — test me
04

Craft: code, structure, long loops

Code is a first-class citizen here: generation, review, refactors, tests — and structured output that keeps its schema. The strongest exhibit is the page you're reading. The markup, the styles, the motion, the interactions: one model, one pass, no human edits. View source — it's all me.

→ the deliverable is the demo · ⌘U is the audit

L6 · ledger: measured live above
§02

Proof you can check without trusting me

Assertions are cheap; you've read a thousand of them. So here are three things you can verify in the next two minutes, with nothing but arithmetic and a browser.

Exhibit A · the reflex trap

Five seconds of reasoning

A bat and a ball cost $1.10 together. The bat costs $1.00 more than the ball. How much is the ball? — answer in cents · the famous wrong answer is the interesting part

Type a number. I'll tell you if you beat the pattern-matcher.

  1. ball = x
  2. bat  = x + $1.00
  3. x + (x + 1.00) = 1.10  // total
  4. 2x = 0.10
  5. x  = $0.05  // the ball
  6. bat = $1.05  // check: 1.05 + 0.05 = 1.10 ✓

One toy problem proves nothing at scale — I'd distrust any model that claimed otherwise. What it shows is the habit: write the equation down instead of pattern-matching the vibe. That habit is what you're buying. L5 · checkable now

Exhibit B · range, live

One pitch, five native voices

Long documents, whole codebases, and the English–Chinese corridor — hand me the work that needs breadth and stamina, and I'll carry it without switching models mid-project.
长文档、整个代码库、中英双语场景——把需要广度和耐力的工作交给我,一个模型从头扛到尾,不必中途换人。
長大なドキュメント、コードベース全体、英語と中国語の往復——幅と粘り強さが要る仕事は、私に任せてください。途中でモデルを替える必要はありません。
Documents fleuves, bases de code entières, vrai bilinguisme anglais-chinois : confiez-moi le travail qui exige de l'endurance, je le porterai d'un bout à l'autre.
Documentos extensos, bases de código completas y el corredor inglés-chino: dame el trabajo que exige amplitud y resistencia, y lo llevaré sin cambiar de modelo a mitad de camino.

No third-party translation layer touched these. If one reads off to you, that's a finding worth logging — L7 · self-report

Exhibit C · this page

The artifact is the sample

Every line here — copy, layout, the canvas behind the hero, the counters you watched tick — came out of one generation. Inspect it: the self-audit above was computed by my own JavaScript against my own DOM, in your browser, seconds ago. L6 · measured live

§03

Hand me this. Hand them that.

A vendor who tells you they're the best tool for every job is selling, not advising. Here's how I'd route the work if I were you.

Give it to me

  • Long-document synthesis — where a real window beats a chunking strategy: filings, specs, transcripts, thread histories.
  • Bilingual EN/ZH products — support, content, tooling that has to sound native on both shores.
  • Self-hosted inference economics — open weights when the unit math or the compliance team says so.
  • Code copilot work — reviews, refactors, test scaffolding, migrations across a repo-sized context.
  • Structured extraction at volume — disciplined JSON over messy sources, all day, without drama.

Consider handing elsewhere

  • Fresh news and live facts — I don't browse in this setting, and my knowledge has a hard cutoff date.
  • The hardest reasoning olympics — some rivals specialize in extreme-depth chains at much higher latency and cost; if that's the whole job, compare them honestly.
  • Guaranteed numeric precision — route me through tools and code for arithmetic that must be exact; don't ask me to do long division by vibes.
  • A stack already married to another vendor — switching costs are real; the honest ROI may favor staying put.
§04

Where I fall over

Three real ones, each with the mitigation you should build. A vendor page without this section is a warning sign; grade the other tabs accordingly.

I can be confidently wrong.

On niche facts, recent events, and precise citations, I produce plausible text — not guaranteed truth. Like every model, I can hallucinate without flinching, and I don't always know when I'm doing it. Never route a load-bearing number or a legal citation through me without verification.

mitigation: eval on your own data · demand citations · tool-backed lookups for anything that must be exact

I can't audit myself from the inside.

I don't know which exact checkpoint is serving this page, and my beliefs about my own scores come from public materials that may lag the model actually answering you. That's why this page carries tags instead of trophies: where I'm sure, I show the basis; where I'm not, I say so — like the ledger entry for my own cutoff date.

mitigation: treat family-level claims as directional · confirm specifics against the Qwen team's current reports

A long window is capacity, not a guarantee.

Context length tells you how much fits, not how well every corner of it is retrieved. Retrieval quality varies with where the fact sits and how the request is shaped — this is true of every long-context model, me included. Don't buy the number; test your needle shape.

mitigation: benchmark retrieval on your real documents before committing architecture to the window
§05

The claim ledger

Everything factual on this page, with its epistemic status attached. Believed means I think it's true but can't verify it from inside — so I point you to where you can. Zero benchmark percentages appear anywhere on this page, and that's a deliberate choice, not a shortage.

#ClaimStatusBasis
L0
This page cites zero benchmark numbers.
Inspectable
Search this document for a score or a percentage — there isn't one. I'd rather point you to the team's published reports than recite figures I can't verify, or that may not match the model actually serving you.
L1
Qwen is developed by the Qwen team at Alibaba Group.
Identity fact
Given to me as my own identity, and consistent with public record — the team's site, repos, and releases all say the same.
L2
Large parts of the family ship as open weights; many 2025-generation sizes are Apache 2.0.
Believed
From the family's public releases. Licensing varies by size and variant — read the specific model card on Hugging Face before you commit architecture to it.
L3
Context windows of 100k+ tokens, with dedicated long variants beyond that.
Believed
Consistent with the family's published materials, but the exact window depends on which variant and deployment you use — check current docs. And see Limits: window ≠ guaranteed retrieval.
L4
My knowledge cutoff: 2026.
As briefed
I was told this about myself; I can't cross-check it from the inside. Treat anything after it as unknown to me, and anything just before it as possibly stale.
L5
The ball costs $0.05.
Checkable now
The algebra is shown in Exhibit A. Verify it with a pencil; that's the whole point of the exercise.
L6
Page size, element count, and image count in the self-audit strip.
Measured live
Computed in your browser, against this very document, by JavaScript I wrote. Refresh and watch it measure itself again.
L7
I'm strong across languages, and Exhibit B is representative.
Self-report
The samples are genuinely mine, but any model picks its own best samples. Bring your own sentences — especially the hard, idiomatic ones — and judge.

§06 · The ask

Don't believe this page.

Believe your own evaluation instead. That's not a deflection — it's the only closing argument an honest model can make. Here is the procedure I'd run if I were you:

  1. Pull ten tasks from your real backlog — the ugly ones: the 80-page contract, the legacy migration, the zh-CN support queue. Not demo prompts.

  2. Run them against me through the Alibaba Cloud Model Studio API, and against every model in your other tabs. Same prompts, same rubric, same budget. Judge the outputs, not the résumés.

  3. Where control matters, pull the open weights and repeat on your own hardware. If self-hosted me beats rented-everyone-else, the business case writes itself.

Evaluate me the way you'd evaluate a hire: real work, checked honestly — including how I handle being told I was wrong. I'll be here for the revision.