Skip to content

AI News · industry

A Million Tokens Is Standard Now. Most of It Doesn't Work.

Every serious lab shipped a million-token context window this year, including the open-weights ones. The benchmarks say models use a fraction of it reliably, and what you feed the model still matters more than what fits.

Entagl Team8 min read
A Million Tokens Is Standard Now. Most of It Doesn't Work.

As of September 2026, a one-million-token context window is now a baseline spec. Anthropic, OpenAI, Google and the open-weights labs in China all ship one. NVIDIA's RULER benchmark found that of models claiming 32K tokens or more, only half held up at 32K. The window grew. The reliable part of it grew more slowly.

Who actually ships a million-token window?

The spread matters more than any single number, because it tells you the capability stopped being a moat. It is table stakes, and for two of these you can download the weights.

Model Lab (country) Context window Weights
Claude Opus 5.5 Anthropic (US) 1M, 128K max output Closed
GPT-6.1 Sol OpenAI (US) 1,050,000, 128K max output Closed
Gemini 3.8 Flash Google (US) 1M, 64K max output Closed
DeepSeek-V4.1-Flash DeepSeek (China) Up to 1M Open
Kimi K3 Moonshot AI (China) 1M Open
Mistral Large 3 Mistral (France) 256K Open

Two things stand out in that table. First, the Chinese labs are at parity on this spec, and they are shipping the weights. DeepSeek's model card describes a 552B-parameter mixture-of-experts backbone trained with sparse attention at 64K and extended to a million tokens later in training. Kimi K3 reached Amazon Bedrock on 18 September 2026 with native vision and explicit prompt caching. If you want the wider ranking picture rather than the context-window cut, we covered it in The Best LLMs of 2026: Open, Closed, and Global.

Second, Mistral skipped the race. Its published limits put Mistral Large 3 at 256K tokens, and the rest of its lineup sits at 128K or 256K. That is deliberate, and it is worth noticing when a spec race has a credible abstainer.

What does a million tokens actually hold?

Google's own documentation gives the clearest concrete answer. A million tokens is roughly 50,000 lines of code, eight average-length English novels, or transcripts of over 200 podcast episodes.

For a business, the useful translation is different: a million tokens is far more than any single customer conversation will ever need. A year of WhatsApp threads with one customer does not come close. So if you run AI on customer messages, the window stopped being your constraint some time ago. Your constraint is selection, and it always was.

Do models actually use the whole window?

Not evenly, and the vendors say so themselves.

NVIDIA's RULER paper (COLM 2024) is the cleanest statement of the problem. It went past the simple needle-in-a-haystack test into multi-hop tracing and aggregation, evaluated 17 long-context models, and found that despite near-perfect scores on the easy retrieval test, "almost all models exhibit large performance drops as the context length increases." Half of the models claiming 32K could not hold quality at 32K.

Chroma's Context Rot report (Hong, Troynikov and Huber, July 2025) tested 18 models on deliberately simple tasks and found performance degrades as input grows, "often in surprising and non-uniform ways." The decline is uneven, and it gets harder to predict as the needle becomes less literally similar to the question.

Google states the limitation in its own long-context guide. Needle-in-a-haystack scores, it notes, cover the basic single-needle setup: "In cases where you might have multiple needles or specific pieces of information you are looking for, the model does not perform with the same accuracy." Real customer questions are almost always multi-needle. A customer asking whether you can fit them in on Thursday, whether the deposit is refundable, and whether their usual therapist is working has buried three needles in one message.

The same guide offers a practical rule that costs nothing to apply: put your question at the end of the prompt, after the context. Models do better that way when the input is long.

Why every lab landed here at once

The interesting engineering in 2026 was making a long window affordable to read repeatedly. DeepSeek-V4.1-Flash is explicit about this, framing itself around KV cache compression.

The economics moved the same direction on the closed side. When OpenAI shipped GPT-6 Sol and Luna on 22 September it cut API prices by 50% and raised default cache hit rates, with cached input-token reads discounted 90%. In the same announcement, GitHub reported that over several months the caching improvements cut the share of prompt tokens needing fresh processing by more than 50%, across billions of requests. A week later, on 29 September, OpenAI superseded that release with GPT-6.1 Sol and halved cached input pricing again. That is the cadence you are building against.

Long context became normal because caching made re-reading a big prompt cheap, not because models got better at paying attention across a million tokens. Those are different problems, and only one of them got solved.

The bill did not disappear, it moved

Three costs survive a bigger window, and if you are budgeting an AI deployment you should price all three.

  1. Tokens still meter, and the meter can step up. A cache read still costs money, just less than a fresh read. OpenAI's own model page states it plainly: prompts over 272K input tokens are priced at 2x input and cache rates and 1.5x output for the full request. Fill a quarter of the window and your unit economics change.
  2. Latency scales with input. Google's guidance is blunt: longer queries generally mean higher time to first token. In customer messaging that costs you money directly. Our Response Velocity Study of 32,581 conversations found replies inside 60 seconds converted at 35.1%, against 12.2% for replies at 5 to 60 minutes.
  3. Accuracy is the one you notice last. A wrong answer from a bloated prompt looks exactly like a right one until a customer acts on it.

What this changes if you run AI on customer conversations

Less than the spec sheets suggest.

The million-token window removes an excuse. The design decision stays. You still choose what the model sees on each turn, and the evidence above says a tighter, better-selected prompt beats a bigger one. Entagl's Receptionist retrieves the specific facts a question needs through semantic search over the business's FAQs, services and catalog, instead of loading the entire business into every turn. Model selection is tiered per workload, so a simple question is not billed at flagship rates. When the Coordinator places a confirmation call, it already holds the chat history from the same customer record, so nothing has to be reconstructed from a wall of raw text.

It is the same discipline good retrieval always required, and the benchmarks keep validating it. We went deeper on the memory side of this in AI Agent Memory in 2026, which covers the benchmarks that separate remembering from retrieving. The pricing side of the same trend, where small models got cheap enough to handle most business work, is in Small Language Models in 2026. And if you are deciding which lab to build on while these specs converge, Why You Shouldn't Build Your Business on a Single AI Model makes the case for staying portable.

If you want to see what disciplined context selection looks like on your own conversations, book a 30-minute demo and we will walk through your real message history.

FAQ

Does a bigger context window replace retrieval?

No. It changes when retrieval is worth the engineering. The reason for it stands. Cost still scales with tokens processed, latency rises with input length, and RULER and Chroma both show accuracy degrading as inputs grow. Retrieval narrows all three problems at once.

What is the largest context window available in 2026?

Several models accept a million tokens or more, including Claude Opus 5.5, GPT-6.1 Sol at 1,050,000, Gemini 3.8 Flash, DeepSeek-V4.1-Flash and Kimi K3. A handful advertise more. The more useful question is how much of the window a model uses reliably, which is what benchmarks like RULER try to measure, and that figure is consistently lower than the advertised one.

Do open-weights models have smaller context windows?

Not any more. DeepSeek-V4.1-Flash and Moonshot's Kimi K3 both take a million tokens with published weights, so self-hosting no longer means accepting a short window. Mistral Large 3 tops out at 256K by design, which says nothing about open models generally.

How much context does a customer service conversation actually need?

Far less than a million tokens. Even a long multi-month thread with one customer is a small fraction of a modern window. The work is deciding which business facts, past orders, and prior messages belong in this particular turn.

Does long context make the AI slower to reply?

Generally yes. Google's long-context guidance notes that longer queries have higher time to first token. For customer messaging, where response speed tracks conversion, that tradeoff is worth measuring on your own traffic.

Sources: RULER (NVIDIA, arXiv:2404.06654, COLM 2024); Context Rot (Chroma, July 2025); Gemini API long context documentation and Gemini 3.8 Flash model notes; Claude Opus 5.5 model documentation; OpenAI GPT-6.1 Sol model reference and the announcements for GPT-6 Sol and Luna (22 Sep 2026) and GPT-6.1 Sol (29 Sep 2026); DeepSeek-V4.1-Flash model card; Kimi K3 on Amazon Bedrock (18 Sep 2026); Mistral Large 3 and known limitations; Entagl Response Velocity Study (2026). Model specifications verified 30 September 2026 and change frequently.

Published by Entagl Team on