Reference
LLM context window comparison
A model's context window is the maximum number of tokens — roughly ¾ of a word each — it can consider at once, covering your prompt, any documents, and its own reply. Bigger windows let a model read whole codebases, books or hours of transcript in one pass.
Advertised limits are not the whole story: accuracy often drops toward the end of very long contexts, and some providers cap consumer apps below the API maximum. Figures below are the largest officially documented window per model.
| # | Model | Context window |
|---|---|---|
| 1 | Llama 4 | 10M tokens |
| 2 | Gemini 1.5 Pro | 2M tokens |
| 3 | Gemini 2.5 Pro | 1M tokens |
| 4 | GPT-5 | 400K tokens |
| 5 | Grok 4 | 256K tokens |
| 6 | Claude Sonnet 4.5 | 200K tokens |
| 7 | Claude Opus 4 | 200K tokens |
| 8 | OpenAI o3 | 200K tokens |
| 9 | OpenAI o1 | 200K tokens |
| 10 | Claude 3.5 Sonnet | 200K tokens |
| 11 | Claude 3 Opus | 200K tokens |
| 12 | DeepSeek R1 | 128K tokens |
| 13 | DeepSeek V3 | 128K tokens |
| 14 | GPT-4o | 128K tokens |
| 15 | Claude 2 | 100K tokens |
| 16 | Mixtral 8x7B | 32K tokens |
| 17 | Gemini 1.0 | 32K tokens |
| 18 | Claude 1 | 9K tokens |
| 19 | Llama 3 | 8.2K tokens |
| 20 | Grok-1 | 8.2K tokens |
| 21 | Mistral 7B | 8.2K tokens |
| 22 | GPT-4 | 8.2K tokens |
| 23 | Llama 2 | 4.1K tokens |
| 24 | GPT-3.5 | 4.1K tokens |
| 25 | GPT-3 | 2.0K tokens |
| 26 | GPT-2 | 1.0K tokens |
Related: What is a context window? · Best open source LLMs