Reasoning
52 items tagged with "reasoning"
Models32
Qwen3.8 Max
Alibaba’s Qwen3.8 Max is a hosted Qwen model added to OpenRouter with a 1,000,000-token context window. It is positioned for long-context general-purpose language tasks, reasoning, and generation.
Qwen3.7 Flash
Qwen3.7 Flash is a fast hosted Qwen model variant added to OpenRouter with a 1,000,000-token context window. It is suited for long-context text, reasoning, and general assistant workloads where lower latency is important.
Ling 3.0 Flash
Ling 3.0 Flash is a hosted general-purpose language model from InclusionAI, newly listed on OpenRouter with a 262K-token context window. The Flash variant is positioned for fast, long-context chat and reasoning workloads.
Gemini 3.6 Flash
Gemini 3.6 Flash is a Google Gemini model newly added to OpenRouter with a 1,048,576-token context window. It is positioned as a Flash-family hosted model for long-context text generation and reasoning workloads.
GPT-5.6 Luna
OpenRouter-listed GPT-5.6 hosted model variant with a 1,050,000-token context window. It is intended for broad long-context text, coding, and reasoning use cases.
GPT-5.6 Terra Pro
OpenRouter-listed GPT-5.6 hosted model variant with a 1,050,000-token context window. It is a proprietary long-context model for advanced text generation, reasoning, and coding workflows.
GPT-5.6 Terra
OpenRouter-listed GPT-5.6 hosted model variant with a 1,050,000-token context window. It targets long-context general-purpose assistance, reasoning, and code-generation tasks.
Grok 4.5
xAI hosted Grok model newly added on OpenRouter with a 500k-token context window for long-context general AI tasks.
Gemini 3.5
Google’s Gemini 3.5 is a new frontier model series focused on combining strong general intelligence with agentic action/tool use, announced at Google I/O 2026.
Claude Opus 4.7 Fast
A latency-optimised variant of Claude Opus 4.7 with a one-million-token context window, designed for real-time agentic workflows, dependency auditing, and large codebase analysis where Opus-class reasoning is required at lower response times.
GPT-5.5 Instant
An updated default ChatGPT model focused on smarter, more accurate responses with reduced hallucinations and improved personalization controls.
Mistral Medium 3.5
A Mistral AI foundation model newly listed on OpenRouter with a 262k token context window, positioned as a balanced medium-tier model for general purpose generation and reasoning tasks.
Grok 4.3
A new Grok-series flagship model variant listed on OpenRouter with a 1M-token context window, aimed at high-context general reasoning and assistant use.
Qwen3.6 Max (Preview)
A preview flagship Qwen3.6 foundation model variant aimed at strong general-purpose reasoning and instruction following with a large context window.
DeepSeek V4 Pro
DeepSeek’s V4 Pro foundation model listing with a 1M-token context window, intended for long-context reasoning and agentic workloads.
GPT-5.5
OpenAI’s flagship GPT-5.5 model, positioned as faster and more capable for complex tasks like coding, research, and data analysis across tools.
Claude Opus 4.7
A new Claude Opus-series frontier model version listed on OpenRouter with a 1M-token context window, intended for high-end reasoning and long-context workloads.
Qwen3.6 Plus Preview
Preview release of Alibaba's Qwen 3.6 Plus model as listed on OpenRouter, offering a very large context window for general-purpose text tasks.
Mistral Small 2603
A new Mistral Small series release listed on OpenRouter with a 262k context window, positioned as a general-purpose foundation model for long-context workloads.
Grok 4.20 (Beta)
A Grok 4.20 beta model offering a very large (2M token) context window for long-context general-purpose chat and reasoning workloads.
NVIDIA Nemotron 3 Super (120B, A12B)
An open model from NVIDIA designed for scalable agentic AI, described as a 120B-parameter model with 12B active parameters and optimized throughput.
GPT-5.4
OpenAI frontier foundation model positioned as more capable and efficient for professional work, with state-of-the-art coding, computer use, and tool search, plus a 1M-token context window.
GPT-5.4 Pro
Higher-tier GPT-5.4 offering listed by OpenRouter, providing a 1M-token context window for advanced professional and agentic workloads.
Gemini 3.1 Pro Preview (Custom Tools)
A Gemini 3.1 Pro preview variant listed on OpenRouter that is explicitly labeled for custom tools, suggesting enhanced tool-use integration with a very large context window.
Qwen3.5 Flash 02-23
A Qwen3.5 Flash model snapshot (02-23) newly listed on OpenRouter with a 1M-token context window, positioned for fast, long-context inference.
Qwen3.5 122B A10B
A large Qwen3.5 Mixture-of-Experts-style model variant newly added on OpenRouter, offering a large 262k-token context window.
Qwen3.5 35B A3B
A Qwen3.5 model variant newly listed on OpenRouter with a 262k-token context window, intended as a mid-sized foundation option in the Qwen3.5 family.
Qwen3.5 27B
A Qwen3.5 27B foundation model newly added on OpenRouter, providing a 262k-token context window for general assistant workloads.
GPT-5.3 Codex
A new Codex-branded GPT-5.3 model intended for code-centric use cases, listed as newly added on OpenRouter with a large context window.
Gemini 3.1 Pro Preview
Preview release of Google's Gemini 3.1 Pro model with a very large context window, aimed at advanced general-purpose reasoning and long-context workloads.
Qwen3.5-Plus-02-15
Alibaba Qwen 3.5 'Plus' model variant as listed on OpenRouter, featuring a 1M-token context window for long-context general-purpose generation and analysis.
Qwen3.5-397B-A17B
Large-scale Qwen 3.5 model (397B with A17B MoE-style routing indicated by the name) added on OpenRouter, intended for high-end reasoning and generation with a 262K context window.
Benchmarks19
MMLU-Pro
A harder, reasoning-focused successor to MMLU with ten answer options and tougher questions designed to separate frontier models that saturated the original.
GSM8K (Grade School Math 8K)
A benchmark of ~8,500 grade-school math word problems that test multi-step arithmetic reasoning with a single numeric answer.
MATH (Competition Mathematics)
A benchmark of 12,500 competition-style math problems across algebra, geometry, number theory, and calculus, graded on exact final-answer match.
BIG-bench (Beyond the Imitation Game)
A collaborative suite of 200+ diverse tasks probing reasoning, knowledge, and emergent abilities beyond conventional language benchmarks.
BBH (BIG-bench Hard)
A 23-task subset of BIG-bench focused on challenging multi-step reasoning where chain-of-thought prompting yields large gains.
HellaSwag
A commonsense sentence-completion benchmark where models pick the most plausible continuation among adversarially generated distractors.
ARC (AI2 Reasoning Challenge)
A grade-school science question benchmark split into Easy and Challenge sets, the latter built from questions retrieval methods answer incorrectly.
GPQA (Graduate-Level Google-Proof Q&A)
A benchmark of expert-written, graduate-level science questions designed to be extremely hard even with web access, testing deep domain reasoning.
MMMU (Massive Multi-discipline Multimodal Understanding)
A multimodal benchmark of college-level questions requiring joint reasoning over text and images such as diagrams, charts, and figures.
DROP (Discrete Reasoning Over Paragraphs)
A reading-comprehension benchmark requiring discrete operations like addition, counting, sorting, and comparison over passage content.
WinoGrande
A large-scale commonsense benchmark of pronoun-resolution sentence pairs designed to require world knowledge rather than lexical cues.
AGIEval
A benchmark built from human standardized exams such as college entrance, law, and civil-service tests to measure human-centric reasoning.
AIME (Competition Math Benchmark)
An olympiad-level math benchmark using American Invitational Mathematics Examination problems with integer answers, a key frontier reasoning test.
MGSM (Multilingual Grade School Math)
A multilingual extension of grade-school math word problems that tests whether LLMs can reason through arithmetic in many languages, not just English.
FRAMES (Factuality, Retrieval, And reasoning MEasurement Set)
A benchmark for retrieval-augmented generation that tests end-to-end factuality, multi-document retrieval, and multi-hop reasoning on questions needing several sources.
MuSR (Multistep Soft Reasoning)
A benchmark of long natural-language narratives requiring multistep commonsense and logical reasoning, such as murder mysteries and object-placement puzzles.
LiveBench
A contamination-resistant benchmark that continuously refreshes questions from recent sources and grades automatically against objective ground truth across many task categories.
CRUXEval (Code Reasoning, Understanding, and Execution)
A benchmark that tests whether models can reason about code execution by predicting function inputs from outputs and outputs from inputs.
ChartQA
A benchmark for answering questions about charts and plots that require visual data extraction plus arithmetic and logical reasoning over the extracted values.