Long Context
67 items tagged with "long-context"
Models65
Ling 3.0 Tiny
A smaller Ling 3.0-series hosted foundation model added to OpenRouter with a 262,144-token context window. The OpenRouter listing identifies it as a free-access model variant.
Muse Spark 1.2
A Meta hosted foundation model added to OpenRouter with a 1,048,576-token context window. It is a newer Muse Spark release than the existing 1.1 model.
Qwen3.8 Max
Alibaba’s Qwen3.8 Max is a hosted Qwen model added to OpenRouter with a 1,000,000-token context window. It is positioned for long-context general-purpose language tasks, reasoning, and generation.
Inkling Small
Inkling Small is a hosted 524K-context model from Thinking Machines, added on OpenRouter as a smaller variant of the Inkling model family. It is suited for long-context text and reasoning workloads where a lighter model is preferred.
Qwen3.7 Flash
Qwen3.7 Flash is a fast hosted Qwen model variant added to OpenRouter with a 1,000,000-token context window. It is suited for long-context text, reasoning, and general assistant workloads where lower latency is important.
Ling 3.0 Flash
Ling 3.0 Flash is a hosted general-purpose language model from InclusionAI, newly listed on OpenRouter with a 262K-token context window. The Flash variant is positioned for fast, long-context chat and reasoning workloads.
Laguna S 2.1
Laguna S 2.1 is a long-context foundation model from Poolside, newly listed on OpenRouter and also present in the Ollama library. The OpenRouter listing reports a 1,048,576-token context window.
Gemini 3.6 Flash
Gemini 3.6 Flash is a Google Gemini model newly added to OpenRouter with a 1,048,576-token context window. It is positioned as a Flash-family hosted model for long-context text generation and reasoning workloads.
Gemini 3.5 Flash-Lite
Gemini 3.5 Flash-Lite is a lightweight Google Gemini hosted model newly added to OpenRouter with a 1,048,576-token context window. It is suited to cost- and latency-sensitive long-context text generation workloads.
LongCat 2.0
LongCat 2.0 is a Meituan long-context foundation model newly added to OpenRouter. The listing reports a 1,048,756-token context window for hosted text-generation workloads.
Inkling
Inkling is a hosted model newly listed on OpenRouter with a 1,048,576-token context window. The discovery data verifies its availability and long-context capability, but does not provide further architectural or benchmark details.
Kimi K3
Kimi K3 is a Moonshot AI long-context hosted language model added to OpenRouter with a 1,048,576-token context window. It is positioned for large-context text, reasoning, and coding workloads.
Muse Spark 1.1
Muse Spark 1.1 is a Meta model added to OpenRouter with a 1,048,576-token context window. The discovery data verifies it as a newly listed long-context hosted model.
KAT Coder Air v2.5
KAT Coder Air v2.5 is a code-focused model from the KwaiPilot namespace added to OpenRouter with a 256,000-token context window. It appears to be the lighter Air variant of the KAT Coder v2.5 family.
KAT Coder Pro v2.5
KAT Coder Pro v2.5 is a code-focused model from the KwaiPilot namespace added to OpenRouter with a 256,000-token context window. It appears to be the higher-capability Pro variant of the KAT Coder v2.5 family.
GPT-5.6 Luna Pro
OpenRouter-listed GPT-5.6 hosted model variant with a 1,050,000-token context window. It is positioned as a long-context frontier text model for demanding reasoning and productivity workloads.
GPT-5.6 Luna
OpenRouter-listed GPT-5.6 hosted model variant with a 1,050,000-token context window. It is intended for broad long-context text, coding, and reasoning use cases.
GPT-5.6 Terra Pro
OpenRouter-listed GPT-5.6 hosted model variant with a 1,050,000-token context window. It is a proprietary long-context model for advanced text generation, reasoning, and coding workflows.
GPT-5.6 Terra
OpenRouter-listed GPT-5.6 hosted model variant with a 1,050,000-token context window. It targets long-context general-purpose assistance, reasoning, and code-generation tasks.
GPT-5.6 Sol Pro
OpenRouter-listed GPT-5.6 hosted model variant with a 1,050,000-token context window. It is a proprietary long-context model for high-end reasoning, coding, and productivity workloads.
GPT-5.6
OpenAI frontier general-purpose model announced as delivering more intelligence per token, stronger performance per dollar, and scalable capability for demanding work.
Grok 4.5
xAI hosted Grok model newly added on OpenRouter with a 500k-token context window for long-context general AI tasks.
Aion 3.0
Aion Labs hosted general-purpose model newly added on OpenRouter with a 131k-token context window.
Aion 3.0 Mini
Smaller Aion Labs hosted model newly added on OpenRouter with a 131k-token context window.
Tencent HY3
Tencent long-context hosted model newly added on OpenRouter with a 262k-token context window.
Laguna XS 2.1
Laguna XS 2.1 is a Poolside model added to OpenRouter with a 262K-token context window and also available in the Ollama library. It is suited to long-context coding and text workflows.
Fugu Ultra
Fugu Ultra is a Sakana model listed on OpenRouter with a 1M-token context window. It is positioned for long-context general-purpose reasoning and text generation.
Gemini 3 Pro Image
Google Gemini Pro-tier image-focused model listed on OpenRouter with a 65,536-token context window. It targets higher-capability multimodal and image-generation use cases than Flash-tier variants.
GLM-5.2
GLM-5.2 is a Z.ai / Zhipu AI foundation model listed on OpenRouter with a 1,048,576-token context window and available in the Ollama library. It targets long-context reasoning, generation, and coding workloads.
Kimi K2.7 Code
Kimi K2.7 Code is a Moonshot AI coding-focused model listed on OpenRouter with a 262K-token context window and available in the Ollama library. It is aimed at software engineering and long-context code understanding tasks.
Claude Fable 5
Claude Fable 5 is a newly listed Anthropic model on OpenRouter with a 1,000,000-token context window. The discovery data verifies it as a new Anthropic model added during the target date range.
NVIDIA Nemotron 3 Ultra 550B A55B
NVIDIA Nemotron 3 Ultra 550B A55B is a large open-weight Nemotron-family model listed on OpenRouter with a 1M-token context window and available in the Ollama library. It is intended for long-context reasoning and general text-generation workloads.
Qwen3.7 Plus
A Qwen-family large language model added on OpenRouter with a 1,000,000-token context window for long-context general-purpose AI workloads.
Claude Opus 4.8
Anthropic's Claude Opus 4.8 is a proprietary frontier model listed on OpenRouter and announced as available on AWS. The provided data highlights its use for agentic systems and production inference workloads, with a 1,000,000-token context window.
Claude Opus 4.8 Fast
Claude Opus 4.8 Fast is an Anthropic model variant added to OpenRouter with a 1,000,000-token context window. It is positioned as the fast variant of Claude Opus 4.8 for lower-latency agentic and production workloads.
Gemini 3.5
Google’s Gemini 3.5 is a new frontier model series focused on combining strong general intelligence with agentic action/tool use, announced at Google I/O 2026.
Gemini 3.5 Flash
A fast, efficient Gemini 3.5-series model variant listed on OpenRouter, intended for low-latency agentic and general assistant workloads with a very large context window.
Claude Opus 4.7 Fast
A latency-optimised variant of Claude Opus 4.7 with a one-million-token context window, designed for real-time agentic workflows, dependency auditing, and large codebase analysis where Opus-class reasoning is required at lower response times.
gpt-chat-latest
A ChatGPT-aligned OpenAI model alias newly added to OpenRouter with a 400k token context window, intended for general conversational and assistant-style use.
Mistral Medium 3.5
A Mistral AI foundation model newly listed on OpenRouter with a 262k token context window, positioned as a balanced medium-tier model for general purpose generation and reasoning tasks.
Grok 4.3
A new Grok-series flagship model variant listed on OpenRouter with a 1M-token context window, aimed at high-context general reasoning and assistant use.
Qwen3.6 Max (Preview)
A preview flagship Qwen3.6 foundation model variant aimed at strong general-purpose reasoning and instruction following with a large context window.
Qwen3.6 Flash
A speed-optimized Qwen3.6 foundation model for low-latency chat and agent workloads while retaining a very large context window.
DeepSeek V4 Pro
DeepSeek’s V4 Pro foundation model listing with a 1M-token context window, intended for long-context reasoning and agentic workloads.
DeepSeek V4 Flash
DeepSeek’s V4 Flash foundation model listing with a 1M-token context window, optimized for lower-latency long-context tasks.
Claude Opus 4.7
A new Claude Opus-series frontier model version listed on OpenRouter with a 1M-token context window, intended for high-end reasoning and long-context workloads.
Claude Opus 4.6 Fast
A faster variant of Claude Opus 4.6 exposed via OpenRouter, aimed at high-throughput production workloads while retaining the Opus-class capability profile.
Gemma 4 26B A4B IT
An instruction-tuned Gemma 4 model listed on OpenRouter, positioned as a large open model for general-purpose chat and instruction following with a long context window.
Qwen3.6-Plus
A long-context Qwen model variant listed on OpenRouter, intended for general-purpose instruction following and long-document workloads.
Gemma 4 31B IT
An instruction-tuned Gemma 4 family model offered via OpenRouter with a very large context window, aimed at general-purpose assistant and agentic workflows.
Qwen3.6 Plus Preview
Preview release of Alibaba's Qwen 3.6 Plus model as listed on OpenRouter, offering a very large context window for general-purpose text tasks.
Mistral Small 2603
A new Mistral Small series release listed on OpenRouter with a 262k context window, positioned as a general-purpose foundation model for long-context workloads.
Grok 4.20 (Beta)
A Grok 4.20 beta model offering a very large (2M token) context window for long-context general-purpose chat and reasoning workloads.
Grok 4.20 Multi-Agent (Beta)
A Grok 4.20 beta variant positioned for multi-agent workflows, with a 2M token context window for coordinating longer multi-step tasks.
NVIDIA Nemotron 3 Super (120B, A12B)
An open model from NVIDIA designed for scalable agentic AI, described as a 120B-parameter model with 12B active parameters and optimized throughput.
Qwen3.5-9B
A 9B-parameter Qwen3.5 foundation model with a large (262k token) context window, positioned for general chat and reasoning with long-context inputs.
Gemini 3.1 Flash-Lite
Google’s fastest and most cost-efficient Gemini 3 series model, built for intelligence at scale.
Gemini 3.1 Pro Preview (Custom Tools)
A Gemini 3.1 Pro preview variant listed on OpenRouter that is explicitly labeled for custom tools, suggesting enhanced tool-use integration with a very large context window.
Qwen3.5 Flash 02-23
A Qwen3.5 Flash model snapshot (02-23) newly listed on OpenRouter with a 1M-token context window, positioned for fast, long-context inference.
Qwen3.5 122B A10B
A large Qwen3.5 Mixture-of-Experts-style model variant newly added on OpenRouter, offering a large 262k-token context window.
Qwen3.5 35B A3B
A Qwen3.5 model variant newly listed on OpenRouter with a 262k-token context window, intended as a mid-sized foundation option in the Qwen3.5 family.
Qwen3.5 27B
A Qwen3.5 27B foundation model newly added on OpenRouter, providing a 262k-token context window for general assistant workloads.
Gemini 3.1 Pro Preview
Preview release of Google's Gemini 3.1 Pro model with a very large context window, aimed at advanced general-purpose reasoning and long-context workloads.
Qwen3.5-Plus-02-15
Alibaba Qwen 3.5 'Plus' model variant as listed on OpenRouter, featuring a 1M-token context window for long-context general-purpose generation and analysis.
Qwen3.5-397B-A17B
Large-scale Qwen 3.5 model (397B with A17B MoE-style routing indicated by the name) added on OpenRouter, intended for high-end reasoning and generation with a 262K context window.
Benchmarks2
RULER (Long-Context Benchmark)
A synthetic long-context benchmark with configurable tasks measuring a model's effective context length beyond simple retrieval.
MuSR (Multistep Soft Reasoning)
A benchmark of long natural-language narratives requiring multistep commonsense and logical reasoning, such as murder mysteries and object-placement puzzles.