Long Context
103 items tagged with "long-context"
Models101
GLM-5.3-FlashX
GLM-5.3-FlashX is a Zhipu AI / Z.ai GLM-family hosted language model added to OpenRouter with a 1,048,576-token context window. It is suited for long-context text generation and reasoning workflows.
DeepSeek Pro Latest
A hosted DeepSeek Pro-class model endpoint added to OpenRouter with a 1,048,576-token context window. It is intended for long-context text generation and general-purpose assistant workloads.
DeepSeek Flash Latest
A hosted DeepSeek Flash-class model endpoint added to OpenRouter with a 1,048,576-token context window. It targets long-context text generation in a faster Flash-tier model line.
Schematron v2 Turbo
A 128k-context hosted language model in the Schematron v2 family, added to OpenRouter on 2026-09-12. The available data verifies its model ID and context length but does not provide benchmark, modality, or pricing details.
Schematron v2 Small
A 128k-context hosted language model in the Schematron v2 family, added to OpenRouter on 2026-09-12. The available data verifies its model ID and context length but does not provide benchmark, modality, or pricing details.
Fugu Ultra v2
A Sakana AI hosted foundation model with a 1,000,000-token context window, added to OpenRouter on 2026-09-11. It appears to be a new v2 update in the Fugu Ultra line, with available data verifying context length but not pricing or benchmark details.
Fugu Max
A Sakana AI hosted foundation model with a 1,000,000-token context window, added to OpenRouter on 2026-09-11. The OpenRouter data verifies it as a new Fugu-family model with very long context support, but does not provide pricing or benchmark details.
DeepSeek V4.1 Flash
DeepSeek V4.1 Flash is a long-context foundation model from DeepSeek, newly listed on OpenRouter and also present in the Ollama library. It is positioned as a fast general-purpose model with a 1,048,576-token context window.
Ling 3.0 Flash VL
Ling 3.0 Flash VL is a vision-language variant of InclusionAI's Ling 3.0 Flash family, newly added to OpenRouter. It supports multimodal use cases with a 262,144-token context window.
Mercury 2.5
Hosted language model added to OpenRouter with a 260k-token context window.
NEX N2.5 Pro
Free hosted NEX AGI language model added to OpenRouter with a 262,144-token context window.
NEX N2.5 Mini
Free hosted compact NEX AGI language model added to OpenRouter with a 262,144-token context window.
GPT-6 Astra Pro
GPT-6 Astra Pro is a higher-tier GPT-6 Astra model variant listed on OpenRouter with a 1,050,000-token context window. It appears intended for demanding hosted frontier-model workloads requiring long-context reasoning and advanced agentic capabilities.
Muse Spark 1.3
Muse Spark 1.3 is a Meta hosted foundation model newly added to OpenRouter with a 1,048,576-token context window. It appears to be the next base release in the Muse Spark series, suited for long-context general-purpose generation.
Gemini 3.8 Flash
Gemini 3.8 Flash is a Google Gemini Flash-series model newly added to OpenRouter with a 1,048,576-token context window. It is positioned as a fast, large-context general-purpose model.
Claude Fable 5.1
Claude Fable 5.1 is a new Anthropic Claude model available on Amazon Bedrock, Claude Platform on AWS, and OpenRouter. The release highlights model improvements and Enterprise Frontier Safeguards for controlled cloud deployments.
Mercury 2.5 Preview
Mercury 2.5 Preview is a hosted preview model added to OpenRouter under the inception namespace. It provides a large 260k-token context window for long-context text tasks.
Granite 4.2 8B
Granite 4.2 8B is an IBM Granite open-weight foundation model added to OpenRouter with a 131k-token context window. It is also present in the Ollama library as granite4.2 for local deployment.
HY4 Preview
Tencent HY4 Preview is a newly added hosted foundation model on OpenRouter with a 1,048,576-token context window. It appears to target general long-context chat and text generation workloads.
Ling 3.0 Flash Fin
A newly added hosted Ling 3.0 Flash financial-domain model on OpenRouter with a 262,144-token context window. It appears intended for long-context finance-oriented text generation and analysis tasks.
Qwen3.8 Flash
Qwen3.8 Flash is an Alibaba Qwen model newly added on OpenRouter with a 1,000,000-token context window. It is positioned as a fast, long-context general-purpose model.
GLM-5.3-Flash
GLM-5.3-Flash is a Zhipu AI / Z.ai GLM model newly added on OpenRouter with a 1,310,720-token context window. It is also present in the Ollama library, indicating local open-weight availability.
DeepSeek V4 Flash Vision Exp
Experimental vision-capable variant of DeepSeek V4 Flash available on OpenRouter, adding multimodal image understanding to the Flash model line with a 1M-token context window.
GLM-5.3
GLM-5.3 is a Zhipu AI / Z.ai foundation model newly added on OpenRouter. The provided data verifies a very large 1,048,576-token context window, making it suitable for long-context text tasks.
Qwen3.8-27B
Qwen3.8-27B is a Qwen 3.8 family language model from Alibaba, listed with a 262k-token context window. It appears to be a smaller open-weight member of the Qwen3.8 lineup suitable for general-purpose long-context language tasks.
Dots 3 Note Preview
OpenRouter-listed free preview model from Dots Studio with a 512k-token context window. The discovery data identifies it as a recent preview model suitable for long-context text workflows.
Gemini 3.7 Flash
Google's Gemini 3.7 Flash is a newly listed hosted model with a 1,048,576-token context window. It appears positioned as a fast long-context Gemini model for general-purpose AI workloads.
Seed 2.1 Turbo
Seed 2.1 Turbo is a newly listed ByteDance Seed hosted model on OpenRouter with a 262,144-token context window. The Turbo designation indicates a speed-oriented variant for general-purpose long-context workloads.
Qwen3.8-2.4T-A95B
Qwen3.8-2.4T-A95B is a newly listed Alibaba Qwen model with a 1,010,000-token context window. The name indicates a large mixture-style Qwen model variant intended for high-capacity long-context reasoning and generation.
Seed 2.0 Code
Seed 2.0 Code is a newly listed ByteDance Seed code-focused model with a 262,144-token context window. It is positioned for software engineering and code-generation workflows.
Grok 4.6
Grok 4.6 is a newly listed xAI hosted model with a 500,000-token context window. It appears as a new Grok-series general-purpose reasoning model.
LFM 2.5 2.6B
Liquid AI's LFM 2.5 2.6B is a compact long-context foundation model listed on OpenRouter with a 128k-token context window.
NVIDIA Nemotron 3.5 Lightning
NVIDIA's Nemotron 3.5 Lightning expands the Nemotron 3 family as an efficient open model for long-running agentic AI workloads.
Sakana Namazu
Sakana Namazu is a Sakana AI model newly listed on OpenRouter with a 262k-token context window.
Solar Pro4
Upstage Solar Pro4 is a newly listed long-context hosted foundation model on OpenRouter, supporting a 524k-token context window.
Muse Glimmer 30B
Meta's Muse Glimmer 30B is a 30B-parameter open-weight model newly listed on OpenRouter with a 131k-token context window and also present in the Ollama library.
Ling 3.0 Tiny
A smaller Ling 3.0-series hosted foundation model added to OpenRouter with a 262,144-token context window. The OpenRouter listing identifies it as a free-access model variant.
Muse Spark 1.2
A Meta hosted foundation model added to OpenRouter with a 1,048,576-token context window. It is a newer Muse Spark release than the existing 1.1 model.
Qwen3.8 Max
Alibaba’s Qwen3.8 Max is a hosted Qwen model added to OpenRouter with a 1,000,000-token context window. It is positioned for long-context general-purpose language tasks, reasoning, and generation.
Inkling Small
Inkling Small is a hosted 524K-context model from Thinking Machines, added on OpenRouter as a smaller variant of the Inkling model family. It is suited for long-context text and reasoning workloads where a lighter model is preferred.
Qwen3.7 Flash
Qwen3.7 Flash is a fast hosted Qwen model variant added to OpenRouter with a 1,000,000-token context window. It is suited for long-context text, reasoning, and general assistant workloads where lower latency is important.
Ling 3.0 Flash
Ling 3.0 Flash is a hosted general-purpose language model from InclusionAI, newly listed on OpenRouter with a 262K-token context window. The Flash variant is positioned for fast, long-context chat and reasoning workloads.
Laguna S 2.1
Laguna S 2.1 is a long-context foundation model from Poolside, newly listed on OpenRouter and also present in the Ollama library. The OpenRouter listing reports a 1,048,576-token context window.
Gemini 3.6 Flash
Gemini 3.6 Flash is a Google Gemini model newly added to OpenRouter with a 1,048,576-token context window. It is positioned as a Flash-family hosted model for long-context text generation and reasoning workloads.
Gemini 3.5 Flash-Lite
Gemini 3.5 Flash-Lite is a lightweight Google Gemini hosted model newly added to OpenRouter with a 1,048,576-token context window. It is suited to cost- and latency-sensitive long-context text generation workloads.
LongCat 2.0
LongCat 2.0 is a Meituan long-context foundation model newly added to OpenRouter. The listing reports a 1,048,756-token context window for hosted text-generation workloads.
Inkling
Inkling is a hosted model newly listed on OpenRouter with a 1,048,576-token context window. The discovery data verifies its availability and long-context capability, but does not provide further architectural or benchmark details.
Kimi K3
Kimi K3 is a Moonshot AI long-context hosted language model added to OpenRouter with a 1,048,576-token context window. It is positioned for large-context text, reasoning, and coding workloads.
Muse Spark 1.1
Muse Spark 1.1 is a Meta model added to OpenRouter with a 1,048,576-token context window. The discovery data verifies it as a newly listed long-context hosted model.
KAT Coder Air v2.5
KAT Coder Air v2.5 is a code-focused model from the KwaiPilot namespace added to OpenRouter with a 256,000-token context window. It appears to be the lighter Air variant of the KAT Coder v2.5 family.
KAT Coder Pro v2.5
KAT Coder Pro v2.5 is a code-focused model from the KwaiPilot namespace added to OpenRouter with a 256,000-token context window. It appears to be the higher-capability Pro variant of the KAT Coder v2.5 family.
GPT-5.6 Luna Pro
OpenRouter-listed GPT-5.6 hosted model variant with a 1,050,000-token context window. It is positioned as a long-context frontier text model for demanding reasoning and productivity workloads.
GPT-5.6 Luna
OpenRouter-listed GPT-5.6 hosted model variant with a 1,050,000-token context window. It is intended for broad long-context text, coding, and reasoning use cases.
GPT-5.6 Terra Pro
OpenRouter-listed GPT-5.6 hosted model variant with a 1,050,000-token context window. It is a proprietary long-context model for advanced text generation, reasoning, and coding workflows.
GPT-5.6 Terra
OpenRouter-listed GPT-5.6 hosted model variant with a 1,050,000-token context window. It targets long-context general-purpose assistance, reasoning, and code-generation tasks.
GPT-5.6 Sol Pro
OpenRouter-listed GPT-5.6 hosted model variant with a 1,050,000-token context window. It is a proprietary long-context model for high-end reasoning, coding, and productivity workloads.
GPT-5.6
OpenAI frontier general-purpose model announced as delivering more intelligence per token, stronger performance per dollar, and scalable capability for demanding work.
Grok 4.5
xAI hosted Grok model newly added on OpenRouter with a 500k-token context window for long-context general AI tasks.
Aion 3.0
Aion Labs hosted general-purpose model newly added on OpenRouter with a 131k-token context window.
Aion 3.0 Mini
Smaller Aion Labs hosted model newly added on OpenRouter with a 131k-token context window.
Tencent HY3
Tencent long-context hosted model newly added on OpenRouter with a 262k-token context window.
Laguna XS 2.1
Laguna XS 2.1 is a Poolside model added to OpenRouter with a 262K-token context window and also available in the Ollama library. It is suited to long-context coding and text workflows.
Fugu Ultra
Fugu Ultra is a Sakana model listed on OpenRouter with a 1M-token context window. It is positioned for long-context general-purpose reasoning and text generation.
Gemini 3 Pro Image
Google Gemini Pro-tier image-focused model listed on OpenRouter with a 65,536-token context window. It targets higher-capability multimodal and image-generation use cases than Flash-tier variants.
GLM-5.2
GLM-5.2 is a Z.ai / Zhipu AI foundation model listed on OpenRouter with a 1,048,576-token context window and available in the Ollama library. It targets long-context reasoning, generation, and coding workloads.
Kimi K2.7 Code
Kimi K2.7 Code is a Moonshot AI coding-focused model listed on OpenRouter with a 262K-token context window and available in the Ollama library. It is aimed at software engineering and long-context code understanding tasks.
Claude Fable 5
Claude Fable 5 is a newly listed Anthropic model on OpenRouter with a 1,000,000-token context window. The discovery data verifies it as a new Anthropic model added during the target date range.
NVIDIA Nemotron 3 Ultra 550B A55B
NVIDIA Nemotron 3 Ultra 550B A55B is a large open-weight Nemotron-family model listed on OpenRouter with a 1M-token context window and available in the Ollama library. It is intended for long-context reasoning and general text-generation workloads.
Qwen3.7 Plus
A Qwen-family large language model added on OpenRouter with a 1,000,000-token context window for long-context general-purpose AI workloads.
Claude Opus 4.8
Anthropic's Claude Opus 4.8 is a proprietary frontier model listed on OpenRouter and announced as available on AWS. The provided data highlights its use for agentic systems and production inference workloads, with a 1,000,000-token context window.
Claude Opus 4.8 Fast
Claude Opus 4.8 Fast is an Anthropic model variant added to OpenRouter with a 1,000,000-token context window. It is positioned as the fast variant of Claude Opus 4.8 for lower-latency agentic and production workloads.
Gemini 3.5
Google’s Gemini 3.5 is a new frontier model series focused on combining strong general intelligence with agentic action/tool use, announced at Google I/O 2026.
Gemini 3.5 Flash
A fast, efficient Gemini 3.5-series model variant listed on OpenRouter, intended for low-latency agentic and general assistant workloads with a very large context window.
Claude Opus 4.7 Fast
A latency-optimised variant of Claude Opus 4.7 with a one-million-token context window, designed for real-time agentic workflows, dependency auditing, and large codebase analysis where Opus-class reasoning is required at lower response times.
gpt-chat-latest
A ChatGPT-aligned OpenAI model alias newly added to OpenRouter with a 400k token context window, intended for general conversational and assistant-style use.
Mistral Medium 3.5
A Mistral AI foundation model newly listed on OpenRouter with a 262k token context window, positioned as a balanced medium-tier model for general purpose generation and reasoning tasks.
Grok 4.3
A new Grok-series flagship model variant listed on OpenRouter with a 1M-token context window, aimed at high-context general reasoning and assistant use.
Qwen3.6 Max (Preview)
A preview flagship Qwen3.6 foundation model variant aimed at strong general-purpose reasoning and instruction following with a large context window.
Qwen3.6 Flash
A speed-optimized Qwen3.6 foundation model for low-latency chat and agent workloads while retaining a very large context window.
DeepSeek V4 Pro
DeepSeek’s V4 Pro foundation model listing with a 1M-token context window, intended for long-context reasoning and agentic workloads.
DeepSeek V4 Flash
DeepSeek’s V4 Flash foundation model listing with a 1M-token context window, optimized for lower-latency long-context tasks.
Claude Opus 4.7
A new Claude Opus-series frontier model version listed on OpenRouter with a 1M-token context window, intended for high-end reasoning and long-context workloads.
Claude Opus 4.6 Fast
A faster variant of Claude Opus 4.6 exposed via OpenRouter, aimed at high-throughput production workloads while retaining the Opus-class capability profile.
Gemma 4 26B A4B IT
An instruction-tuned Gemma 4 model listed on OpenRouter, positioned as a large open model for general-purpose chat and instruction following with a long context window.
Qwen3.6-Plus
A long-context Qwen model variant listed on OpenRouter, intended for general-purpose instruction following and long-document workloads.
Gemma 4 31B IT
An instruction-tuned Gemma 4 family model offered via OpenRouter with a very large context window, aimed at general-purpose assistant and agentic workflows.
Qwen3.6 Plus Preview
Preview release of Alibaba's Qwen 3.6 Plus model as listed on OpenRouter, offering a very large context window for general-purpose text tasks.
Mistral Small 2603
A new Mistral Small series release listed on OpenRouter with a 262k context window, positioned as a general-purpose foundation model for long-context workloads.
Grok 4.20 (Beta)
A Grok 4.20 beta model offering a very large (2M token) context window for long-context general-purpose chat and reasoning workloads.
Grok 4.20 Multi-Agent (Beta)
A Grok 4.20 beta variant positioned for multi-agent workflows, with a 2M token context window for coordinating longer multi-step tasks.
NVIDIA Nemotron 3 Super (120B, A12B)
An open model from NVIDIA designed for scalable agentic AI, described as a 120B-parameter model with 12B active parameters and optimized throughput.
Qwen3.5-9B
A 9B-parameter Qwen3.5 foundation model with a large (262k token) context window, positioned for general chat and reasoning with long-context inputs.
Gemini 3.1 Flash-Lite
Google’s fastest and most cost-efficient Gemini 3 series model, built for intelligence at scale.
Gemini 3.1 Pro Preview (Custom Tools)
A Gemini 3.1 Pro preview variant listed on OpenRouter that is explicitly labeled for custom tools, suggesting enhanced tool-use integration with a very large context window.
Qwen3.5 Flash 02-23
A Qwen3.5 Flash model snapshot (02-23) newly listed on OpenRouter with a 1M-token context window, positioned for fast, long-context inference.
Qwen3.5 122B A10B
A large Qwen3.5 Mixture-of-Experts-style model variant newly added on OpenRouter, offering a large 262k-token context window.
Qwen3.5 35B A3B
A Qwen3.5 model variant newly listed on OpenRouter with a 262k-token context window, intended as a mid-sized foundation option in the Qwen3.5 family.
Qwen3.5 27B
A Qwen3.5 27B foundation model newly added on OpenRouter, providing a 262k-token context window for general assistant workloads.
Gemini 3.1 Pro Preview
Preview release of Google's Gemini 3.1 Pro model with a very large context window, aimed at advanced general-purpose reasoning and long-context workloads.
Qwen3.5-Plus-02-15
Alibaba Qwen 3.5 'Plus' model variant as listed on OpenRouter, featuring a 1M-token context window for long-context general-purpose generation and analysis.
Qwen3.5-397B-A17B
Large-scale Qwen 3.5 model (397B with A17B MoE-style routing indicated by the name) added on OpenRouter, intended for high-end reasoning and generation with a 262K context window.
Benchmarks2
RULER (Long-Context Benchmark)
A synthetic long-context benchmark with configurable tasks measuring a model's effective context length beyond simple retrieval.
MuSR (Multistep Soft Reasoning)
A benchmark of long natural-language narratives requiring multistep commonsense and logical reasoning, such as murder mysteries and object-placement puzzles.