Reference

LLM context window comparison

A model's context window is the maximum number of tokens — roughly ¾ of a word each — it can consider at once, covering your prompt, any documents, and its own reply. Bigger windows let a model read whole codebases, books or hours of transcript in one pass.

Advertised limits are not the whole story: accuracy often drops toward the end of very long contexts, and some providers cap consumer apps below the API maximum. Figures below are the largest officially documented window per model.

#ModelContext window
1Llama 410M tokens
2Gemini 1.5 Pro2M tokens
3Gemini 2.5 Pro1M tokens
4GPT-5400K tokens
5Grok 4256K tokens
6Claude Sonnet 4.5200K tokens
7Claude Opus 4200K tokens
8OpenAI o3200K tokens
9OpenAI o1200K tokens
10Claude 3.5 Sonnet200K tokens
11Claude 3 Opus200K tokens
12DeepSeek R1128K tokens
13DeepSeek V3128K tokens
14GPT-4o128K tokens
15Claude 2100K tokens
16Mixtral 8x7B32K tokens
17Gemini 1.032K tokens
18Claude 19K tokens
19Llama 38.2K tokens
20Grok-18.2K tokens
21Mistral 7B8.2K tokens
22GPT-48.2K tokens
23Llama 24.1K tokens
24GPT-3.54.1K tokens
25GPT-32.0K tokens
26GPT-21.0K tokens

Related: What is a context window? · Best open source LLMs