Listen to this lesson
Unlock audio and more
Audio streaming, downloadable PDFs and certificates come with Plus and Pro.
What it means
Everything a model uses to produce an answer has to fit in its context window: your question, the conversation so far, any documents you pasted, the system instructions, and the answer being generated. It is measured in tokens rather than words. When a conversation exceeds the window, something has to be dropped or summarized, which is why long chats sometimes seem to forget earlier details.
Context windows have grown enormously, and vendors compete on the number. But capacity is not the same as attention — models reliably use information at the beginning and end of a long context better than material buried in the middle, an effect often called "lost in the middle". A large window makes something possible; it does not guarantee the model will use all of it well.
Why it matters
Context is the main lever most people have over output quality, and the main constraint. It sets how much of a document you can analyze in one pass, how long an agent can run before losing the thread, and how much a request costs — since you pay per token, a large context is a real expense on every call.
In practice
If quality degrades on a long document, the fix is usually not a bigger model but less context: extract the relevant sections and pass those. Putting the most important material at the start or end of the prompt measurably helps.
Where this shows up
Tools and models in our catalog.
Claude Opus 5.5Anthropic's leading model (September 22, 2026) and the first of the Claude 5.5 family — Anthropic says it performs at Claude Fable 5.1's level on most work while costing 40 percent less to run than Opus 5, at $4 input and $20 output per million tokens with cache reads 60 percent cheaper at 20 cents. Ships with Fable 5.1-level biology and cybersecurity safeguards.
Gemini 3.1 ProGoogle DeepMind's Pro-tier model (Feb 2026) and still the newest Pro-tier Gemini you can use, since Gemini 4 Argon is limited to vetted cyber defenders. 94.3% GPQA Diamond and 77.1% ARC-AGI-2 at launch. 1M token context, native multimodal input (text, image, video, audio), Deep Think reasoning mode. Available via Vertex AI Model Garden and Google AI Studio.