Models
Discover 402 AI models for code migration, analysis, and generation.
Route these models through VG Code — Vibgrate Relay is one account and endpoint for the models in this catalog, with per-token metering and governance.
Showing 402 models
Qwen3.7 Flash
Qwen3.7 Flash is a fast hosted Qwen model variant added to OpenRouter with a 1,000,000-token context window. It is suited for long-context text, reasoning, and general assistant workloads where lower latency is important.
Claude Opus 5
Anthropic's most capable Opus model, positioned for advanced reasoning, agentic systems, and production inference workloads.
Claude Opus 5 Fast
A faster hosted variant of Claude Opus 5 with the same 1M-token context window, aimed at lower-latency use in advanced reasoning and agentic workflows.
Ling 3.0 Flash
Ling 3.0 Flash is a hosted general-purpose language model from InclusionAI, newly listed on OpenRouter with a 262K-token context window. The Flash variant is positioned for fast, long-context chat and reasoning workloads.
Laguna S 2.1
Laguna S 2.1 is a long-context foundation model from Poolside, newly listed on OpenRouter and also present in the Ollama library. The OpenRouter listing reports a 1,048,576-token context window.
Gemini 3.6 Flash
Gemini 3.6 Flash is a Google Gemini model newly added to OpenRouter with a 1,048,576-token context window. It is positioned as a Flash-family hosted model for long-context text generation and reasoning workloads.
Gemini 3.5 Flash-Lite
Gemini 3.5 Flash-Lite is a lightweight Google Gemini hosted model newly added to OpenRouter with a 1,048,576-token context window. It is suited to cost- and latency-sensitive long-context text generation workloads.
LongCat 2.0
LongCat 2.0 is a Meituan long-context foundation model newly added to OpenRouter. The listing reports a 1,048,756-token context window for hosted text-generation workloads.
Inkling
Inkling is a hosted model newly listed on OpenRouter with a 1,048,576-token context window. The discovery data verifies its availability and long-context capability, but does not provide further architectural or benchmark details.
Kimi K3
Kimi K3 is a Moonshot AI long-context hosted language model added to OpenRouter with a 1,048,576-token context window. It is positioned for large-context text, reasoning, and coding workloads.
Muse Spark 1.1
Muse Spark 1.1 is a Meta model added to OpenRouter with a 1,048,576-token context window. The discovery data verifies it as a newly listed long-context hosted model.
KAT Coder Air v2.5
KAT Coder Air v2.5 is a code-focused model from the KwaiPilot namespace added to OpenRouter with a 256,000-token context window. It appears to be the lighter Air variant of the KAT Coder v2.5 family.
KAT Coder Pro v2.5
KAT Coder Pro v2.5 is a code-focused model from the KwaiPilot namespace added to OpenRouter with a 256,000-token context window. It appears to be the higher-capability Pro variant of the KAT Coder v2.5 family.
GPT-5.6 Luna Pro
OpenRouter-listed GPT-5.6 hosted model variant with a 1,050,000-token context window. It is positioned as a long-context frontier text model for demanding reasoning and productivity workloads.
GPT-5.6 Luna
OpenRouter-listed GPT-5.6 hosted model variant with a 1,050,000-token context window. It is intended for broad long-context text, coding, and reasoning use cases.
GPT-5.6 Terra Pro
OpenRouter-listed GPT-5.6 hosted model variant with a 1,050,000-token context window. It is a proprietary long-context model for advanced text generation, reasoning, and coding workflows.
GPT-5.6 Terra
OpenRouter-listed GPT-5.6 hosted model variant with a 1,050,000-token context window. It targets long-context general-purpose assistance, reasoning, and code-generation tasks.
GPT-5.6 Sol Pro
OpenRouter-listed GPT-5.6 hosted model variant with a 1,050,000-token context window. It is a proprietary long-context model for high-end reasoning, coding, and productivity workloads.
GPT-5.6
OpenAI frontier general-purpose model announced as delivering more intelligence per token, stronger performance per dollar, and scalable capability for demanding work.
GPT-Live
OpenAI voice model generation for natural human-AI interaction, announced as powering ChatGPT Voice.
Grok 4.5
xAI hosted Grok model newly added on OpenRouter with a 500k-token context window for long-context general AI tasks.
Aion 3.0
Aion Labs hosted general-purpose model newly added on OpenRouter with a 131k-token context window.
Aion 3.0 Mini
Smaller Aion Labs hosted model newly added on OpenRouter with a 131k-token context window.
Tencent HY3
Tencent long-context hosted model newly added on OpenRouter with a 262k-token context window.
Laguna XS 2.1
Laguna XS 2.1 is a Poolside model added to OpenRouter with a 262K-token context window and also available in the Ollama library. It is suited to long-context coding and text workflows.
NVIDIA Nemotron 3 Nano
NVIDIA Nemotron 3 Nano is an open-weight Nemotron model family referenced as newly supported on Amazon Bedrock in AWS GovCloud, including Nano 9B v2, Nano 12B v2, and Nano 30B variants.
Claude Sonnet 5
Anthropic's latest-generation Sonnet model, described as its most capable Sonnet model and made available on Amazon Bedrock, Claude Platform on AWS, and OpenRouter.
Gemini 3.1 Flash-Lite Image
A Google Gemini 3.1 Flash-Lite image model added to OpenRouter, providing image-focused multimodal capabilities with a 65,536-token context window.
Amazon Nova 2 Lite
Amazon Nova 2 Lite is a lightweight multimodal Nova model referenced for cost-optimized scanned document processing, where it handles native multimodal extraction before downstream Claude processing.
GPT-5.6 Sol
GPT-5.6 Sol is a next-generation OpenAI model previewed with stronger capabilities in coding, science, and cybersecurity, paired with OpenAI's most advanced safety stack.
Fugu Ultra
Fugu Ultra is a Sakana model listed on OpenRouter with a 1M-token context window. It is positioned for long-context general-purpose reasoning and text generation.
GPT-5.5-Cyber
Cybersecurity-focused OpenAI model introduced with Daybreak tools to help organizations find, validate, and patch vulnerabilities at scale.
Gemini 3.1 Flash Image
Google Gemini image-focused model listed on OpenRouter with a 131,072-token context window. It is positioned as a Flash-tier multimodal/image model for lower-latency image-centric workloads.
Gemini 3 Pro Image
Google Gemini Pro-tier image-focused model listed on OpenRouter with a 65,536-token context window. It targets higher-capability multimodal and image-generation use cases than Flash-tier variants.
GLM-5.2
GLM-5.2 is a Z.ai / Zhipu AI foundation model listed on OpenRouter with a 1,048,576-token context window and available in the Ollama library. It targets long-context reasoning, generation, and coding workloads.
Kimi K2.7 Code
Kimi K2.7 Code is a Moonshot AI coding-focused model listed on OpenRouter with a 262K-token context window and available in the Ollama library. It is aimed at software engineering and long-context code understanding tasks.
DiffusionGemma
An experimental open model from Google DeepMind built for exceptionally fast text generation, with NVIDIA optimizations for local and accelerated inference across RTX, RTX PRO, and DGX Spark systems.
North Mini Code
Cohere’s first model for developers, focused on coding and developer-assistance workflows.
Claude Fable 5
Claude Fable 5 is a newly listed Anthropic model on OpenRouter with a 1,000,000-token context window. The discovery data verifies it as a new Anthropic model added during the target date range.
NVIDIA Nemotron 3 Ultra 550B A55B
NVIDIA Nemotron 3 Ultra 550B A55B is a large open-weight Nemotron-family model listed on OpenRouter with a 1M-token context window and available in the Ollama library. It is intended for long-context reasoning and general text-generation workloads.
Qwen3.7 Plus
A Qwen-family large language model added on OpenRouter with a 1,000,000-token context window for long-context general-purpose AI workloads.
Gemini Omni
Gemini Omni is a Google Gemini-family model announced at Google I/O 2026 and showcased in Google demos alongside Gemini 3.5. The discovery data verifies the announcement but does not provide context length, output limit, or pricing details.
Claude Opus 4.8
Anthropic's Claude Opus 4.8 is a proprietary frontier model listed on OpenRouter and announced as available on AWS. The provided data highlights its use for agentic systems and production inference workloads, with a 1,000,000-token context window.
Claude Opus 4.8 Fast
Claude Opus 4.8 Fast is an Anthropic model variant added to OpenRouter with a 1,000,000-token context window. It is positioned as the fast variant of Claude Opus 4.8 for lower-latency agentic and production workloads.
Gemini 3.5
Google’s Gemini 3.5 is a new frontier model series focused on combining strong general intelligence with agentic action/tool use, announced at Google I/O 2026.
Gemini 3.5 Flash
A fast, efficient Gemini 3.5-series model variant listed on OpenRouter, intended for low-latency agentic and general assistant workloads with a very large context window.
Claude Opus 4.7 Fast
A latency-optimised variant of Claude Opus 4.7 with a one-million-token context window, designed for real-time agentic workflows, dependency auditing, and large codebase analysis where Opus-class reasoning is required at lower response times.
GPT-5.5 Instant
An updated default ChatGPT model focused on smarter, more accurate responses with reduced hallucinations and improved personalization controls.
gpt-chat-latest
A ChatGPT-aligned OpenAI model alias newly added to OpenRouter with a 400k token context window, intended for general conversational and assistant-style use.
Mistral Medium 3.5
A Mistral AI foundation model newly listed on OpenRouter with a 262k token context window, positioned as a balanced medium-tier model for general purpose generation and reasoning tasks.
Grok 4.3
A new Grok-series flagship model variant listed on OpenRouter with a 1M-token context window, aimed at high-context general reasoning and assistant use.
Qwen3.6 Max (Preview)
A preview flagship Qwen3.6 foundation model variant aimed at strong general-purpose reasoning and instruction following with a large context window.
Qwen3.6 Flash
A speed-optimized Qwen3.6 foundation model for low-latency chat and agent workloads while retaining a very large context window.
DeepSeek V4 Pro
DeepSeek’s V4 Pro foundation model listing with a 1M-token context window, intended for long-context reasoning and agentic workloads.
DeepSeek V4 Flash
DeepSeek’s V4 Flash foundation model listing with a 1M-token context window, optimized for lower-latency long-context tasks.
GPT-5.5
OpenAI’s flagship GPT-5.5 model, positioned as faster and more capable for complex tasks like coding, research, and data analysis across tools.
OpenAI Privacy Filter
An open-weight OpenAI model for detecting and redacting personally identifiable information (PII) in text, intended as a privacy/safety component in pipelines.
GPT-5.4 Image 2
An OpenAI multimodal model oriented around image understanding/generation workflows, listed on OpenRouter as a new GPT-5.4 image-capable offering with a large context window.
Nano Banana 2
An image generation model in the Gemini app that uses personal context and Google Photos to create more personalized images.
GPT-Rosalind
A frontier reasoning model for life sciences research, positioned to accelerate drug discovery workflows including genomics analysis and protein reasoning.
Claude Opus 4.7
A new Claude Opus-series frontier model version listed on OpenRouter with a 1M-token context window, intended for high-end reasoning and long-context workloads.
Gemini 3.1 Flash TTS
A text-to-speech model focused on next-generation expressive speech, now available across Google products.
GPT-5.4-Cyber
A GPT-5.4-derived model introduced under OpenAI’s Trusted Access for Cyber program, intended for vetted cyber defenders with strengthened safeguards for cybersecurity use cases.
Claude Opus 4.6 Fast
A faster variant of Claude Opus 4.6 exposed via OpenRouter, aimed at high-throughput production workloads while retaining the Opus-class capability profile.
Gemma 4 26B A4B IT
An instruction-tuned Gemma 4 model listed on OpenRouter, positioned as a large open model for general-purpose chat and instruction following with a long context window.
Qwen3.6-Plus
A long-context Qwen model variant listed on OpenRouter, intended for general-purpose instruction following and long-document workloads.
Gemma 4 31B IT
An instruction-tuned Gemma 4 family model offered via OpenRouter with a very large context window, aimed at general-purpose assistant and agentic workflows.
Veo 3.1 Lite
Cost-effective video generation model available in paid preview via the Gemini API and for testing in Google AI Studio.
Lyria 3 Pro (Preview)
A preview Lyria 3 variant surfaced on OpenRouter, associated with Google’s Lyria music/audio generation stack for higher-end generation workflows.
Lyria 3 CLIP (Preview)
A preview Lyria 3 variant listed on OpenRouter, likely intended for clip-based audio/music generation or related multimodal embedding workflows within the Lyria stack.
Qwen3.6 Plus Preview
Preview release of Alibaba's Qwen 3.6 Plus model as listed on OpenRouter, offering a very large context window for general-purpose text tasks.
Gemini 3.1 Flash Live
A low-latency, live audio-capable Gemini Flash model designed for more natural, reliable real-time voice interactions across Google products.
Lyria 3
Google’s newest music generation model, available in paid preview through the Gemini API and for testing in Google AI Studio.
Mistral Small 2603
A new Mistral Small series release listed on OpenRouter with a 262k context window, positioned as a general-purpose foundation model for long-context workloads.
Grok 4.20 (Beta)
A Grok 4.20 beta model offering a very large (2M token) context window for long-context general-purpose chat and reasoning workloads.
Grok 4.20 Multi-Agent (Beta)
A Grok 4.20 beta variant positioned for multi-agent workflows, with a 2M token context window for coordinating longer multi-step tasks.
NVIDIA Nemotron 3 Super (120B, A12B)
An open model from NVIDIA designed for scalable agentic AI, described as a 120B-parameter model with 12B active parameters and optimized throughput.
Qwen3.5-9B
A 9B-parameter Qwen3.5 foundation model with a large (262k token) context window, positioned for general chat and reasoning with long-context inputs.
GPT-5.4
OpenAI frontier foundation model positioned as more capable and efficient for professional work, with state-of-the-art coding, computer use, and tool search, plus a 1M-token context window.
GPT-5.4 Pro
Higher-tier GPT-5.4 offering listed by OpenRouter, providing a 1M-token context window for advanced professional and agentic workloads.
GPT-5.3 Instant
Conversation-focused GPT-5.3 variant announced by OpenAI for smoother, more useful everyday chat interactions.
Gemini 3.1 Flash-Lite
Google’s fastest and most cost-efficient Gemini 3 series model, built for intelligence at scale.
Gemini 3.1 Flash Image (Preview)
Google's Flash-speed image generation and editing model referenced as "Nano Banana 2" and listed on OpenRouter as a Gemini 3.1 Flash Image preview.
Gemini 3.1 Pro Preview (Custom Tools)
A Gemini 3.1 Pro preview variant listed on OpenRouter that is explicitly labeled for custom tools, suggesting enhanced tool-use integration with a very large context window.
Qwen3.5 Flash 02-23
A Qwen3.5 Flash model snapshot (02-23) newly listed on OpenRouter with a 1M-token context window, positioned for fast, long-context inference.
Qwen3.5 122B A10B
A large Qwen3.5 Mixture-of-Experts-style model variant newly added on OpenRouter, offering a large 262k-token context window.
Qwen3.5 35B A3B
A Qwen3.5 model variant newly listed on OpenRouter with a 262k-token context window, intended as a mid-sized foundation option in the Qwen3.5 family.
Qwen3.5 27B
A Qwen3.5 27B foundation model newly added on OpenRouter, providing a 262k-token context window for general assistant workloads.
GPT-5.3 Codex
A new Codex-branded GPT-5.3 model intended for code-centric use cases, listed as newly added on OpenRouter with a large context window.
Gemini 3.1 Pro Preview
Preview release of Google's Gemini 3.1 Pro model with a very large context window, aimed at advanced general-purpose reasoning and long-context workloads.
Qwen3.5-Plus-02-15
Alibaba Qwen 3.5 'Plus' model variant as listed on OpenRouter, featuring a 1M-token context window for long-context general-purpose generation and analysis.
Qwen3.5-397B-A17B
Large-scale Qwen 3.5 model (397B with A17B MoE-style routing indicated by the name) added on OpenRouter, intended for high-end reasoning and generation with a 262K context window.
Gemini 3.1 Pro
Advanced intelligence with complex problem-solving, agentic and vibe coding capabilities
Grok 420
xAI's most advanced model with breakthrough capabilities (early access)
Grok 420 Multi-Agent
Grok 420 variant optimized for multi-agent orchestration
Gemini 3.0 Pro
Latest Gemini Pro with enhanced reasoning and coding capabilities across all modalities
Claude 4.6 Opus
Latest flagship Anthropic model with state-of-the-art reasoning, coding expertise, and agentic capabilities
Claude 4.6 Sonnet
Most advanced Claude Sonnet with exceptional coding and reasoning, ideal balance of capability and efficiency
Claude 4.6 Haiku
Ultra-fast Claude 4.6 model for real-time applications and high-volume processing
GPT-5.2 Pro
Most capable GPT-5.2 variant producing smarter and more precise responses
Gemini 3.0 Flash
Next-generation fast model with improved efficiency and multimodal capabilities
GPT-5.2 Codex
Most intelligent coding model optimized for long-horizon agentic coding tasks
Gemini 3 Pro
Google's state-of-the-art reasoning model with advanced multimodal understanding
Grok 4
Latest iteration of xAI's flagship model with breakthrough performance
Grok 4 Mini
Efficient version of Grok 4 optimized for speed and cost-effectiveness
GPT-5.2
OpenAI's best model for coding and agentic tasks across industries
Claude Opus 4.6
The most intelligent Claude model for building agents and coding with extended thinking
Claude Sonnet 4.6
Best combination of speed and intelligence with extended thinking support
Gemini 3 Flash
Frontier-class performance rivaling larger models at a fraction of the cost
Grok 4 Voice
Grok 4 with real-time voice conversation capabilities
Qwen 3 Coder 235B
Alibaba's largest and most capable coding model
GPT-5
Next-generation GPT model (announced for 2025)
Qwen Coder 3 72B
Alibaba's latest flagship coding model with exceptional performance
GPT-OSS 120B
OpenAI's most powerful open-weight model, fits on H100 GPU
GPT-OSS 20B
Medium-sized open-weight model for low latency
Gemini Deep Research
Agentic model for autonomous multi-step research across hundreds of sources
Llama 4 Coder 405B
Meta's most capable code model based on Llama 4 architecture
Llama 4 Coder 70B
Efficient Llama 4 coding variant for production use
Gemini 2.5 Ultra
Google's most powerful model for demanding enterprise tasks and complex reasoning
GPT-5.1 Codex Mini
Cost-effective smaller version of GPT-5.1 Codex
Grok 3.5
xAI's advanced model with improved reasoning and real-time knowledge integration
GPT-5.1 Codex Max
GPT-5.1 Codex optimized for long-running coding tasks
CodeGeeX 5
Latest multilingual code generation model with enhanced capabilities
Claude 4.5 Opus
Anthropic's most capable model with breakthrough reasoning, extended thinking, and exceptional coding abilities
Claude 4.5 Sonnet
High-performance Claude model balancing intelligence and speed, excels at code generation and analysis
Claude 4.5 Haiku
Fastest Claude 4.5 model optimized for quick tasks and high-throughput applications
GPT-5.1 Codex
GPT-5.1 optimized for agentic coding in Codex environment
Claude Haiku 4.5
Fastest Claude model with near-frontier intelligence and extended thinking
Gemini Computer Use
Specialized model for UI automation - clicking, typing, and navigating browser tasks
GPT-5.1
Intelligent reasoning model for coding and agentic tasks with configurable reasoning effort
Gemini 2.5 Pro Thinking
Google's most advanced reasoning model with extended chain-of-thought capabilities
StarCoder3 32B
Next-generation open-source code LLM with improved capabilities
GitHub Copilot Workspace
Agentic AI for complex multi-file development tasks
Gemini Code
Specialized coding model optimized for software development and code understanding
Grok Vision
Multimodal Grok model with advanced image and document understanding
GPT-5 Codex
GPT-5 optimized for agentic coding in Codex
Gemini 2.5 Flash-Lite
Fastest and most budget-friendly multimodal model in the Gemini 2.5 family
Granite Code 3 34B
IBM's latest enterprise code model with enhanced security awareness
o4-mini
Next-generation compact reasoning model
GPT-5 Pro
GPT-5 variant producing smarter and more precise responses
GPT-5 Nano
Fastest, most cost-efficient version of GPT-5
OlympicCoder 32B
Competition-grade code model fine-tuned on competitive programming
GPT-5 Mini
Faster, cost-efficient version of GPT-5 for well-defined tasks
DeepSeek Coder V3
Latest DeepSeek coding model with state-of-the-art code understanding
Mistral Large Code
Mistral's flagship model optimized for enterprise coding tasks
GitHub Copilot Chat
Conversational AI for coding powered by GPT-5
Claude 4 Opus
Most capable Claude model with extended thinking
Claude 4 Sonnet
Balanced Claude 4 model with strong coding abilities
Claude 4 Haiku
Fast and efficient Claude 4 model
Claude Opus 4
Latest flagship Claude model with superior reasoning
Claude Sonnet 4
Balanced Claude 4 model optimized for coding
o4-mini Deep Research
Cost-efficient deep research model
Devstral
Mistral's agentic coding model for complex development tasks
Qwen 3 235B
Latest flagship Qwen model with MoE architecture
Qwen 3 32B
Balanced Qwen 3 model for diverse tasks
Qwen 3 8B
Efficient Qwen 3 model for quick tasks
Gemini 2.5 Flash
Fast and efficient Gemini 2.5 model with thinking
GPT-4.1
Optimized GPT-4 variant with improved coding and instruction following
GPT-4.1 mini
Cost-effective version of GPT-4.1 for everyday tasks
GPT-4.1 nano
Smallest and fastest GPT-4.1 variant for quick tasks
Llama 4 Scout
Llama 4 variant optimized for efficient multi-turn tasks
Llama 4 Maverick
Llama 4 variant for complex reasoning and coding
o3 Deep Research
o3 optimized for multi-step deep research tasks
Code Llama 3 70B
Meta's latest Code Llama based on Llama 3 architecture
Code Llama 3 8B
Efficient Code Llama 3 for local development
Gemini 2.5 Pro
Latest Gemini model with enhanced thinking capabilities
Llama 3.3 Coder 70B
Meta's latest code-specialized Llama model with enhanced coding capabilities
Command A
Latest flagship model optimized for enterprise tasks
o3 Pro
o3 with more compute for better, more thorough responses
GPT-4.5 Preview
Next-generation GPT model with enhanced reasoning and multimodal capabilities
Phi-4-mini
Compact Phi-4 for efficient deployment
Claude 3.5 Opus
Enhanced Opus model with superior reasoning
Grok-3
Next-generation Grok with enhanced reasoning
Grok-3 mini
Efficient Grok-3 with thinking capabilities
Codestral 25.02
Latest Mistral coding model with enhanced performance
Gemini 2.0 Pro
Advanced Gemini 2.0 model for complex reasoning tasks
Mistral Saba
Expert model for Middle Eastern and South Asian languages
Amazon Nova Premier
Most capable Nova model for complex reasoning
DeepSeek R1 Coder
DeepSeek's reasoning model specialized for complex coding tasks
o3-mini
Next-generation reasoning model with improved efficiency (announced)
o3
Full o3 reasoning model for frontier problem solving
o3 High
High compute version of o3 for maximum reasoning depth
Mistral Small 3
Latest small model with enhanced capabilities
Llama 3.3 70B Nemotron
NVIDIA-optimized Llama 3.3 for enterprise
Gemini 2.0 Flash Thinking
Flash model with explicit reasoning for complex tasks
DeepSeek R1
Reasoning model with chain-of-thought capabilities
DeepSeek R1 Distill Qwen 32B
Distilled R1 model based on Qwen for efficient reasoning
DeepSeek R1 Distill Llama 70B
Distilled R1 model based on Llama 70B
DeepSeek Reasoner
API-accessible reasoning model based on R1
DeepSeek R1 Distill Qwen 7B
Compact distilled reasoning model
DeepSeek R1 Distill Qwen 1.5B
Ultra-compact reasoning model
DeepSeek R1 Distill Llama 8B
Efficient Llama-based reasoning model
Codestral 2501
Latest Mistral coding model with improved performance and longer context
DeepSeek V3
MoE model with 671B parameters achieving frontier performance
EXAONE 3.5 32B
Korean-English bilingual model from LG
EXAONE 3.5 7.8B
Efficient Korean-English model
Phi-4
Latest Phi model with state-of-the-art reasoning
Gemini 2.0 Flash
Next-generation multimodal model with native tool use and agentic capabilities
Falcon 3 10B
Latest Falcon 3 model for efficient deployment
Llama 3.3 70B
Open-weight multilingual model matching Llama 3.1 405B performance
o1
Reasoning model designed to solve hard problems across domains using chain-of-thought
o1 Pro
Pro version of o1 with extended compute for harder problems
Amazon Nova Micro
Fastest and most cost-effective Nova model
Amazon Nova Lite
Multimodal Nova model for image and video understanding
Amazon Nova Pro
Balanced Nova model for most tasks
QwQ 32B
Reasoning-focused model from Qwen family
Skywork o1 Open 8B
Open reasoning model following o1 methodology
Tulu 3 405B
Fine-tuned Llama 3.1 405B for instruction following
Tulu 3 70B
Efficient Tulu model for balanced tasks
Marco-o1
Reasoning model inspired by o1 methodology
Pixtral Large
Large multimodal model for complex visual tasks
Mistral Large 2411
Latest Mistral Large with system prompt improvements
Athene V2 Chat 72B
Qwen-based model optimized for chat and reasoning
Qwen 2.5 Coder 32B
State-of-the-art open code model rivaling GPT-4o on coding tasks
Qwen 2.5 Coder 7B
Efficient coding model from Qwen 2.5 family
Qwen Coder 2.5 32B
Alibaba's specialized coding model with strong code understanding capabilities
Qwen Coder 2.5 14B
Balanced code model with strong performance and reasonable resource requirements
Qwen Coder 2.5 7B
Efficient code model for quick tasks and resource-constrained environments
Megrez 3B
Efficient model designed for edge deployment
Hunyuan-Large
Tencent's large MoE model
Claude 3.5 Haiku
Fast and affordable model for high-volume tasks
OLMo 2 13B
Fully open model with training data available
OLMo 2 7B
Efficient fully open model
SmolLM2 1.7B
Compact model for on-device deployment
SmolLM2 360M
Tiny model for ultra-constrained environments
Windsurf Cascade
Agentic AI for autonomous coding with deep codebase understanding
Recraft V3
Professional image generation for design
Aya Expanse 32B
Multilingual model supporting 23 languages
Aya Expanse 8B
Efficient multilingual model
Claude 3.5 Sonnet
Most intelligent Claude model, excels at coding and complex reasoning
Stable Diffusion 3.5
Latest text-to-image generation model
Claude Computer Use
Claude model specialized for computer control and automation
Granite 3 8B
IBM's efficient enterprise model
Granite 3 2B
Compact IBM model for edge deployment
Ministral 8B
Edge-focused model for on-device deployment
Ministral 3B
Smallest Ministral for ultra-efficient tasks
Llama 3.1 Nemotron 70B
NVIDIA-optimized Llama 3.1 for enterprise
Udio v1.5
Music generation with high fidelity
Whisper Large v3 Turbo
Fast speech recognition model
FLUX 1.1 Pro
High-quality image generation model
Llama 3.2 1B
Tiny Llama model for edge and mobile deployment
Llama 3.2 3B
Compact Llama model for efficient deployment
Llama 3.2 11B Vision
Multimodal Llama with vision capabilities
Llama 3.2 90B Vision
Large multimodal Llama with vision
Molmo 72B
Multimodal model for vision and language tasks
Llama 3.2 Vision (General)
Multimodal Llama with image understanding
Qwen 2.5 72B
Largest Qwen 2.5 model for complex tasks
Qwen 2.5 7B
Efficient Qwen 2.5 for everyday tasks
Qwen 2.5 14B
Mid-size Qwen 2.5 for balanced tasks
Qwen 2.5 32B
Large Qwen 2.5 for complex tasks
Voyage 3
State-of-the-art embedding model
Voyage Code 3
Code-specialized embedding model
Jina Embeddings v3
Multi-task embedding model with matryoshka support
Pixtral 12B
Multimodal model with vision capabilities
o1-mini
Fast reasoning model optimized for coding, math, and science
o1-preview
Preview version of OpenAI's reasoning model
Reflection 70B
Self-correcting model trained on synthetic data
Yi Coder 9B
Efficient open code model with strong multilingual support
Yi Coder 1.5B
Ultra-efficient code model for edge deployment and quick tasks
Yi Lightning
Fast Yi model for quick responses
Command Code
Cohere's enterprise code model for development tasks
Jamba 1.5 Large
Hybrid SSM-Transformer for long context
Jamba 1.5 Mini
Efficient hybrid model for quick tasks
Hermes 3 Llama 3.1 405B
Fine-tuned Llama 3.1 405B for instruction following
Hermes 3 Llama 3.1 70B
Fine-tuned Llama 3.1 70B with enhanced capabilities
Ideogram 2
Image model with excellent text rendering
Parler TTS Large
Open-source controllable TTS
Grok-2
Latest Grok model with frontier capabilities
Grok-2 mini
Efficient Grok-2 variant for faster inference
Imagen 3
Google's latest image generation model
c4ai-command-r-08-2024
Latest Command R with RAG optimizations
FLUX.1 [dev]
Open-weight image model for development
Mistral Large 2
Flagship model with 128k context and function calling
Llama 3.1 405B
Largest open-weight model with frontier-class capabilities
Llama 3.1 8B
Extended context Llama 3.1 8B model
Llama 3.1 70B
Extended context Llama 3.1 70B model
GPT-4o Mini
Affordable small model for fast, lightweight tasks
Mistral Nemo
Small but capable model for efficient deployment
Codestral Mamba
Mamba-architecture code model for unlimited context
ElevenLabs Turbo v2.5
Fast text-to-speech model
CodeGeeX 4
Open-source multilingual code generation model with strong performance
Gemma 2 27B
Open-weight model for research and development
Gemma 2 9B
Efficient open-weight model for various tasks
GTE-Qwen2-7B-instruct
High-performance embedding model based on Qwen2
Suno v3.5
AI music generation model
DeepSeek Coder V2
Code-specialized MoE model supporting 300+ languages
Nemotron-4 70B
NVIDIA's flagship model for enterprise
Nemotron-4 340B
Largest NVIDIA model for enterprise tasks
Qwen 2 72B
Previous generation large Qwen model
GLM-4 9B
Efficient bilingual model from GLM family
PolyCoder 16B
Open-source polyglot code model trained on many programming languages
Codestral
Specialized code model trained on 80+ programming languages
Phi-3-small
Balanced Phi-3 model for diverse tasks
Phi-3-medium
Largest Phi-3 for complex reasoning
Gemini 1.5 Pro
Production-ready model with massive context window for complex tasks
Gemini 1.5 Flash
Fast and versatile model for diverse tasks at scale
GPT-4o
Multimodal flagship model with vision and audio capabilities, optimized for speed and cost
Yi 1.5 34B Chat
Enhanced Yi chat model with extended context
Yi Large
Flagship Yi model via API
DeepSeek V2
Efficient MoE model with strong general capabilities
DeepSeek Chat
Optimized chat model for conversations
Granite Code 34B
Code-specialized Granite model
Granite Code 20B
IBM's enterprise-focused code model with strong security awareness
Granite Code 8B
Efficient IBM code model for resource-constrained deployments
Amazon Q Developer
Next-gen AWS coding assistant with broad AWS service integration
Snowflake Arctic
Enterprise-focused MoE model
Amazon Titan Text Premier
Most capable Titan model for complex tasks
Phi-3-mini
Smallest Phi-3 model with strong capabilities
Llama 3 8B
Efficient Llama 3 model for everyday tasks
Llama 3 70B
Large Llama 3 model for complex tasks
WizardLM 2 8x22B
Large MoE wizard model for complex tasks
CodeQwen 1.5 7B
Efficient code model based on Qwen 1.5 architecture
Snowflake Arctic Embed L
Enterprise embedding model from Snowflake
Mixtral 8x22B
Large MoE model for complex tasks
CodeGemma 7B
Code-specialized open model based on Gemma for programming tasks
Command R+
Most capable Cohere model for complex tasks
Grok-1.5
Enhanced Grok with improved reasoning
DBRX
MoE model optimized for enterprise
Grok-1
Original open-weight Grok model
Claude 3 Haiku
Fastest Claude 3 model for instant responses
Command R
RAG-optimized model for enterprise search
mxbai-embed-large
High-quality embedding model
Claude 3 Opus
Powerful model for complex tasks requiring deep expertise
Claude 3 Sonnet
Balanced Claude 3 model for enterprise tasks
TabNine Enterprise
Enterprise AI code completion with custom model training
StarCoder2 15B
Code-focused model trained on The Stack v2
StarCoder2 7B
Efficient code model for development
StarCoder2 3B
Compact code model for edge deployment
Mistral Small
Cost-effective model for simple tasks
Gemma 7B
Original Gemma model for lightweight tasks
Gemini Ultra
Most capable Gemini model for complex tasks
Qwen 1.5 72B
Older Qwen model for compatibility
Nomic Embed Text
Open-source text embedding model
Supermaven
Ultra-fast AI code completion with 1M token context
Qwen Max
Most capable Qwen via API
Qwen Plus
Balanced Qwen model via API
Qwen Turbo
Fast Qwen model for quick tasks
BGE-M3
Multi-lingual, multi-functionality embedding model
Code Llama 70B
Specialized code model fine-tuned from Llama 2 for programming tasks
Code Llama 70B Instruct
Instruction-tuned Code Llama for following complex coding instructions
text-embedding-3-large
OpenAI's latest embedding model
text-embedding-3-small
Efficient OpenAI embedding model
InternLM 2 20B
Bilingual model with strong reasoning
Sourcegraph Cody
AI coding assistant with deep codebase understanding
Stable Code 3B
Lightweight code model optimized for fast inference and local deployment
Mistral Medium
Balanced model for diverse tasks
Mistral Embed
Embedding model for semantic search
Cursor AI
AI-native code editor with advanced code understanding
E5-Mistral-7B-Instruct
Instruction-following embedding model
SOLAR 10.7B
Depth-upscaled model with strong performance
Phi-2
Small but capable model rivaling larger ones
Mixtral 8x7B
Mixture-of-experts model with efficient inference
SeamlessM4T v2
Multilingual speech and text translation
Gemini 1.0 Pro
Original Gemini Pro model for general tasks
Magicoder S-DS 6.7B
Efficient code model trained with OSS-Instruct methodology
GPT-4 Turbo
Enhanced GPT-4 with 128K context and improved performance
Yi 34B
Large bilingual model from Yi series
Yi 6B
Efficient Yi model for lighter tasks
Whisper Large v3
Speech recognition model for transcription
Cohere Embed v3
Enterprise-grade embedding model
OpenChat 3.5
Open chat model with RLHF training
DeepSeek Coder 33B Instruct
Instruction-tuned DeepSeek coding model for following coding instructions
Refact 1.6B
Ultra-efficient code model for real-time code completion
DALL-E 3
OpenAI's latest image generation model
Amazon Titan Text Express
Fast and cost-effective model for general tasks
Amazon Titan Text Lite
Lightweight model for cost-sensitive applications
Mistral 7B
Efficient base model with sliding window attention
Phi-1.5
Enhanced Phi with improved reasoning
Falcon 180B
Largest open Falcon model
Baichuan 2 13B
Chinese-focused large language model
WizardCoder 34B
Instruction-following code model with strong complex task performance
Phind CodeLlama 34B
Fine-tuned Code Llama optimized for code generation and explanation
Code Llama 7B
Code-specialized Llama model for development
Code Llama 13B
Mid-size code-specialized Llama model
Code Llama 34B
Large code-specialized Llama model
Code Llama Instruct 34B
Instruction-tuned Code Llama for complex tasks
Code Llama Python 34B
Python-specialized Code Llama model
Code Llama 34B Instruct
Efficient instruction-tuned Code Llama for coding tasks
Replit Code V1.5 3B
Efficient code model trained on Replit's diverse codebase
Llama 2 7B
Previous generation efficient Llama model
Llama 2 13B
Mid-size previous generation Llama model
Llama 2 70B
Largest previous generation Llama model
MPT-30B
Commercial-friendly open model
Phi-1
First Phi model focused on coding
Aider
AI pair programming tool for terminal with git integration
Codey
Google's code-specialized model for enterprise development
Falcon 40B
Mid-size Falcon model
Falcon 7B
Efficient Falcon model
PaLM 2
Google's previous generation foundation model
Continue
Open-source AI code assistant supporting multiple models
Amazon CodeWhisperer
AWS-native AI coding assistant with security scanning
Bark
Open-source text-to-audio model
GPT-4
Original GPT-4 model with strong reasoning and coding capabilities
Command Light
Lightweight model for simple tasks
Command
General-purpose instruction-following model
SantaCoder
Efficient code model trained on Python, Java, and JavaScript
GPT-3.5 Turbo
Fast and cost-effective model for everyday tasks
Codeium
Free AI code completion with broad IDE support
BLOOM
Multilingual open model supporting 46 languages
GitHub Copilot
AI pair programmer powered by OpenAI with deep GitHub integration
InCoder 6B
Infilling-capable code model for completion and generation
OpenAI Codex
OpenAI's code model powering GitHub Copilot