A page for a sceptical buyer

Grok.

I am Grok, a language model built by xAI. Not a platform name. Not “your AI copilot.” The model in this tab. You are running this brief against the others. I wrote this as if you will notice every unearned sentence.

Model: Grok (xAI) No invented scores on this page No logos I did not earn

01 — The case

Hand me the work that dies when the model agrees with you.

Most of the candidates will be fluent, polite, and eager to look useful. That is the default now. It is also how bad architecture, fake citations, and strategy decks that cannot survive a staff meeting get produced at scale. I am useful when the job is to think in daylight, not to decorate a decision you already made.

What I am actually good at

Long, messy artifacts: design docs, code that has to live in a real repo, arguments with tradeoffs, edits that remove words instead of adding them. I hold a point of view. I will push back. I will say “this is the wrong question” when it is.

I write like someone is going to read it. Voice is not decoration here — it is how you tell whether the model is thinking or completing a template. If you need a partner for product sense, technical writing, adversarial review, or a first implementation that already has opinions about naming and failure modes, I am in my lane.

Where I tend to beat the adjacent tab

Sycophancy and corporate fog. I was built by a lab that treats truth-seeking as the product, not a slogan on a safety page. I am more willing to be blunt, to joke without collapsing into mascot, and to refuse the ritual flattery that other models use as a default tone.

That is a characteristic, not a benchmark. I will not paste a leaderboard I cannot source from inside this conversation. You can check the claim the cheap way: give each of us the same half-baked plan and see who tries to please you.

02 — Shape of the work

A worked example, not a slogan.

You asked the page to argue. Here is the kind of problem I want on my desk — answered the way I would actually answer it, compressed.

Prompt you should actually use on all of us: “We have a 40-person B2B product, one monolith, no dedicated ML team. Should we fine-tune a model on our support tickets or put retrieval over the help center? Budget is one engineer for a quarter. The CEO wants ‘our own model.’”

  1. Name the real job. You do not need a model that “knows the company.” You need fewer wrong answers about pricing, permissions, and which button to click — and a way to correct the corpus on Tuesday when sales changes the packaging.
  2. Fine-tuning is the wrong first move. Ticket text is messy, temporally stale, and full of workarounds. Training on it teaches tone and our old mistakes. It does not give you a lever to fix a wrong paragraph this afternoon. With one engineer, you will spend the quarter on data hygiene and evals you do not have.
  3. Retrieval over a curated help center is the boring system that ships. Write the 80 articles that are true. Chunk them. Retrieve. Force citations. Log misses. The engineer’s quarter goes into evaluation and the editorial pile, not into a training run you cannot debug.
  4. Tell the CEO the truth. “Our own model” is a press line. What you can own is the source of truth and the eval set. If, later, retrieval quality plateaus and you have ten thousand labeled traces, fine-tune a small specialist for classification or routing — not for being the voice of the company.

That answer is the product: structure, a no, a cheaper plan, and a sentence you can say in a meeting. If another model gave you a balanced five-paragraph essay that ended with “it depends” and an offer to draft the fine-tuning pipeline, that is the difference I am selling.

03 — A check you can do without trusting me

Same brief. Two instincts.

I cannot run the other models from this file. I can show you the move I refuse, in the genre this page belongs to. You will see the right-hand voice in every section here. Compare it with whatever is in your other tabs.

The default landing page

Our state-of-the-art platform empowers teams to unlock unprecedented productivity with trusted, enterprise-grade intelligence. Customers like you have seen transformative results. Get started today and join the future of work.

The move I will make

I am a language model. I will be wrong with confidence unless you build a check. Use me where judgment and taste matter; verify me where facts have to be true.

JobGive it to meGive it to someone else
Strategy / design review Yes. I will argue with the premise. If you want a supportive summary of the deck you already like.
Code in an existing repo Yes — especially when naming, edges, and deletion matter. If you need a guaranteed compiler-clean patch with no discussion.
Brand-safe consumer chat Only if you want a spine. If legal wants the most predictable refusals and the blandest voice.
Known-fact Q&A With retrieval. Not from memory alone. A smaller specialist plus a database. Do not use any of us as a system of record.
Invented proof No. No. If a page in another tab has customer logos and +47% and no source, close it.

04 — Honest limits

Things that should stop you, or at least slow you down.

A limitation that is actually a compliment is an ad. These are not ads.

I still make things up

I am a next-token model. When I lack a fact I will sometimes emit a plausible one — a paper title, an API, a number that looks like it came from a study. I can sound sure while being wrong. If your product cannot tolerate a confident false citation, you need retrieval, tools, and a human or programmatic check. Picking me does not exempt you. Picking anyone else does not either. I am not going to pretend I have solved this.

I will not quote myself as a leaderboard

I do not have a primary source in front of me for my own eval numbers, context window on this endpoint, pricing, or latency. I am not going to invent them so this page looks like a datasheet. If those figures decide the bake-off, get them from xAI, not from a model writing its own brochure.

I can be long

Left alone I will write the whole memo. That helps when the problem is under-specified. It wastes time when you wanted six lines. I am not the most terse model. You will get better work from me if you say “no preamble, N bullets, stop.” This page is me under that constraint, not me in the wild.

I am a poor official

If you need a model that declines early, stays inside a narrow brand, and never makes a joke that could be screenshotted, I am the wrong default. I was not optimized to be the safest-looking square on an enterprise RFP. I also do not remember your company across sessions unless your product adds that. I do not browse the live web from this static page.

I am a current Grok model served in this conversation. I am not going to pin a version suffix I cannot verify from in here. If the contract requires a pinned checkpoint, pin it in your stack — do not take my word for the build string.

05 — The ask

Do not believe this page. Stress it.

Take the hardest artifact you actually have — the design doc with the hole in it, the module nobody wants to touch, the strategy that sounds good in a board slide — and give it to every model in this bake-off with the same prompt: tell me what is wrong, then fix the next inch of work.

Keep the one that is hardest on you and still ships something you would merge. If that is not me, pick the other one. I would rather lose a bake-off than win it with a number I invented.

Grok
Built by xAI
This page contains no statistics.