Low Latency
13 items tagged with "low-latency"
Models11
Qwen3.7 Flash
Qwen3.7 Flash is a fast hosted Qwen model variant added to OpenRouter with a 1,000,000-token context window. It is suited for long-context text, reasoning, and general assistant workloads where lower latency is important.
Gemini 3.5 Flash-Lite
Gemini 3.5 Flash-Lite is a lightweight Google Gemini hosted model newly added to OpenRouter with a 1,048,576-token context window. It is suited to cost- and latency-sensitive long-context text generation workloads.
Gemini 3.5 Flash
A fast, efficient Gemini 3.5-series model variant listed on OpenRouter, intended for low-latency agentic and general assistant workloads with a very large context window.
Claude Opus 4.7 Fast
A latency-optimised variant of Claude Opus 4.7 with a one-million-token context window, designed for real-time agentic workflows, dependency auditing, and large codebase analysis where Opus-class reasoning is required at lower response times.
GPT-5.5 Instant
An updated default ChatGPT model focused on smarter, more accurate responses with reduced hallucinations and improved personalization controls.
Qwen3.6 Flash
A speed-optimized Qwen3.6 foundation model for low-latency chat and agent workloads while retaining a very large context window.
DeepSeek V4 Flash
DeepSeek’s V4 Flash foundation model listing with a 1M-token context window, optimized for lower-latency long-context tasks.
Claude Opus 4.6 Fast
A faster variant of Claude Opus 4.6 exposed via OpenRouter, aimed at high-throughput production workloads while retaining the Opus-class capability profile.
Gemini 3.1 Flash Live
A low-latency, live audio-capable Gemini Flash model designed for more natural, reliable real-time voice interactions across Google products.
GPT-5.3 Instant
Conversation-focused GPT-5.3 variant announced by OpenAI for smoother, more useful everyday chat interactions.
Gemini 3.1 Flash-Lite
Google’s fastest and most cost-efficient Gemini 3 series model, built for intelligence at scale.