Every model, unlimited possibilities
Frontier and open models, pre-warmed and ready. Route to any of them from a single Claws endpoint, swap providers without rewriting a line. Here’s what’s powering agents this week.

gpt-astra-latestThis model always redirects to the latest model in the GPT Astra family.

gpt-6-astraGPT-6 Astra is OpenAI's flagship model for demanding end-to-end work.
Claude Fable 5.1 improves on Claude Fable 5 across the board, with the biggest gains in agentic coding, long-running agentic workflows, a…

claude-fable-latestThis model always redirects to the latest model in the Claude Fable family.

fugu-ultra-v2Fugu Ultra v2 is the higher-performance model in Sakana AI's Fugu family.
Qwen3.8 Max 0902 is an updated snapshot of Qwen3.8 Max from Alibaba's Qwen team.
Muse Spark 1.3 is a multimodal reasoning model from Meta for long-running agentic, multi-agent, and coding workflows.
Grok 4.6 is SpaceXAI's smartest model with frontier performance on coding, knowledge work, and STEM.
Gemini 3.8 Flash is Google's most intelligent Flash model with significant gains from 3.7 Flash across software engineering, agentic task…
GLM-5.3 is a large-scale reasoning model from Z.ai, built for complex software engineering and long-horizon agent tasks.

kimi-k3:batchKimi K3 is a 2.8T parameter open-weight multimodal reasoning model from Moonshot AI.
Tencent: Hy4 preview is a mixture-of-experts model from Tencent, with 49B active parameters out of 770B total.
DeepSeek V4.1 Flash is a sparse mixture-of-experts model from DeepSeek, and the first built on the company's Causal Encoder-Decoder (CED)…

schematron-v2-smallSchematron V2 Small is a 3B-parameter HTML-to-JSON extraction model from Inference.net.

ling-3.0-flash-vlLing 3.0 Flash VL builds on Ling 3.0 Flash (124B total / 5.5B active MoE from InclusionAI), further strengthening its language capabiliti…

mercury-2.5Mercury 2.5 is the fastest reasoning LLM, and the latest diffusion LLM (dLLM) from Inception.

nex-n2.5-mini:freeNex-N2.5 is an agentic model built to turn goals into working, verified outcomes.

glm-latestThis model always redirects to the latest GLM model from Z.ai.

aion-3.0Aion-3.0 is a multi-model roleplaying and storytelling system from AionLabs, built on the GLM family of models.

granite-4.2-8bGranite 4.2 8B is a dense reasoning model from IBM.

seed-2-1-turboSeed 2.1 Turbo is a multimodal model from ByteDance Seed for coding and long-horizon agent workflows.

grok-latestThis model always redirects to the latest Grok model from xAI.

dots-3-note-preview:freeDots3-Note Preview is an open-weight mixture-of-experts model from Dots Studio, with 16B active parameters out of 280B total.

nemotron-3.5-lightningNVIDIA Nemotron 3.5 Lightning is an open mixture-of-experts model from NVIDIA, with 3B active parameters out of 30B total.

solar-pro4Solar Pro 4 is Upstage's cost-efficient large language model, featuring a 524K context window.

lfm-2.5-2.6b:freeLFM2.5-2.6B is a compact reasoning model from Liquid AI.

inkling-small:batchInkling Small is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 12B active parameters out of 276B to…

deepseek-v4-flash-latestThis model always redirects to the latest model in the DeepSeek V4 Flash family.

longcat-2.0LongCat 2.0 is a sparse mixture-of-experts language model from Meituan, with 48B active parameters out of 1.6T total.

kat-coder-pro-v2.5KAT-Coder-Pro V2.5 is a flagship-level Agentic Coding model that can directly hand over an entire issue or an entire business workflow to…

laguna-s-2.1Laguna S 2.1 is the latest coding agent model from [Poolside](<https://poolside.ai/>).
North Mini Code is Cohere's first agentic coding model and the debut of its North family.

kimi-latestThis model always redirects to the latest model in the Kimi family.

gemini-pro-latestThis model always redirects to the latest model in the Gemini Pro family.

minimax-m3MiniMax-M3 is a multimodal foundation model from MiniMax.

step-3.7-flashStep 3.7 Flash is StepFun's latest high-efficiency multimodal Mixture-of-Experts model.
Mistral Medium 3.5 is a dense 128B instruction-following model from Mistral AI.

perceptron-mk1Perceptron Mk1 (Mark One) is Perceptron's highest-quality vision-language model for video and embodied reasoning.** It accepts image and…
MiMo-V2.5-Pro is Xiaomi’s flagship model, delivering strong performance in general agentic capabilities, complex software engineering, an…

trinity-large-thinkingTrinity Large Thinking is a powerful open source reasoning model from the team at Arcee AI.

reka-edgeReka Edge is an extremely efficient 7B multimodal vision-language model that accepts image/video+text inputs and generates text outputs.

palmyra-x5Palmyra X5 is Writer's most advanced model, purpose-built for building and scaling AI agents across the enterprise.

freeThe simplest way to get free inference.

sonar-pro-searchExclusively available on the OpenRouter API, Sonar Pro's new Pro Search mode is Perplexity's most advanced agentic search system.

relace-searchThe relace-search model uses 4-12 `view_file` and `grep` tools in parallel to explore a codebase and return relevant files to the user re…

nova-premier-v1Amazon Nova Premier is the most capable of Amazon’s multimodal models for complex reasoning tasks and for use as the best teacher for dis…

magnum-v4-72bThis is a series of models designed to replicate the prose quality of the Claude 3 models, specifically Sonnet(https://openrouter.ai/anth…

cydonia-24b-v4.1Uncensored and creative writing model based on Mistral Small 3.2 24B with good recall, prompt adherence, and intelligence.

hermes-4-405bHermes 4 is a large-scale reasoning model built on Meta-Llama-3.1-405B and released by Nous Research.

morph-v3-largeMorph's high-accuracy apply model for complex code edits.

l3.1-euryale-70bEuryale L3.1 70B v2.2 is a model focused on creative roleplay from [Sao10k](https://ko-fi.com/sao10k).
WizardLM-2 8x22B is Microsoft AI's most advanced Wizard model.

ernie-4.5-vl-424b-a47bERNIE-4.5-VL-424B-A47B is a multimodal Mixture-of-Experts (MoE) model from Baidu’s ERNIE 4.5 series, featuring 424B total parameters with…
Meta's latest class of model (Llama 3.1) launched with a variety of sizes & flavors.

weaverAn attempt to recreate Claude-style verbosity, but don't expect the same level of coherence or memory.

remm-slerp-l2-13bA recreation trial of the original MythoMax-L2-B13 but with updated models.

dolphin-mistral-24b-venice-editionVenice Uncensored Dolphin Mistral 24B Venice Edition is a fine-tuned variant of Mistral-Small-24B-Instruct-2501, developed by dphn.ai in…

ui-tars-1.5-7bUI-TARS-1.5 is a multimodal vision-language agent optimized for GUI-based environments, including desktop interfaces, web browsers, mobil…

mythomax-l2-13bOne of the highest performing and most popular fine-tunes of Llama 2 13B, with rich descriptions and roleplay.

gpt-sol-latestThis model always redirects to the latest model in the GPT Sol family.

gpt-6-astra-proGPT-6 Astra Pro is the same underlying model as [GPT-6 Astra](https://openrouter.ai/openai/gpt-6-astra), served with `reasoning.mode` set…
Claude Fable 5 is a Mythos-class model from Anthropic, built for autonomous knowledge work and coding.

claude-opus-latestThis model always redirects to the latest model in the Claude Opus family.

fugu-maxFugu Max is the cost-performance model in Sakana AI's Fugu family.
Qwen3.8 2.4T A95B is an open-weight sparse mixture-of-experts model from Qwen and the open-weight variant of [Qwen3.8 Max](/qwen/qwen3.8-…
Muse Spark 1.3 Contributor is the cost-efficient contributor tier of Meta’s multimodal reasoning model for experimentation, learning, and…
Grok 4.5 is a model from SpaceXAI with frontier performance on coding, knowledge work, and STEM.
Gemini 3.8 Flash is Google's most intelligent Flash model with significant gains from 3.7 Flash across software engineering, agentic task…
GLM-5.3 is a large-scale reasoning model from Z.ai, built for complex software engineering and long-horizon agent tasks.

kimi-k3Kimi K3 is a 2.8T parameter open-weight multimodal reasoning model from Moonshot AI.
Hy-MT2-30B-A3B is Tencent's flagship translation model in the Hy-MT2 family.
DeepSeek V4 Flash Vision Exp is an experimental vision-enabled version of [DeepSeek V4 Flash 0731](https://openrouter.ai/deepseek/deepsee…

schematron-v2-turboSchematron V2 Turbo is a 3B-parameter HTML-to-JSON extraction model from Inference.net.

ling-3.0-flash-vl:freeLing 3.0 Flash VL builds on Ling 3.0 Flash (124B total / 5.5B active MoE from InclusionAI), further strengthening its language capabiliti…

mercury-2Mercury 2 is an extremely fast reasoning LLM, and the first reasoning diffusion LLM (dLLM).

nex-n2.5-pro:freeNex-N2.5 is an agentic model built to turn goals into working, verified outcomes.

glm-flash-latestThis model always redirects to the latest model in the GLM Flash family.

aion-3.0-miniAion-3.0 Mini is a multi-model roleplaying and storytelling system from AionLabs, built on the DeepSeek family of models.

granite-4.0-h-microGranite-4.0-H-Micro is a 3B parameter from the Granite 4 family of models.

seed-2.0-codeSeed 2.0 Code is a model from ByteDance Seed optimized for agentic coding.

nemotron-3.5-lightning:freeNVIDIA Nemotron 3.5 Lightning is an open mixture-of-experts model from NVIDIA, with 3B active parameters out of 30B total.

solar-pro-3Solar Pro 3 is Upstage's powerful Mixture-of-Experts (MoE) language model.

inkling-smallInkling Small is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 12B active parameters out of 276B to…

kat-coder-pro-v2KAT-Coder-Pro V2 is the latest high-performance model in KwaiKAT’s KAT-Coder series, designed for complex enterprise-grade software engin…

laguna-s-2.1:freeLaguna S 2.1 is the latest coding agent model from [Poolside](<https://poolside.ai/>).
Command A is an open-weights 111B parameter model with a 256k context window focused on delivering great performance across agentic, mult…

gemini-flash-latestThis model always redirects to the latest model in the Gemini Flash family.

minimax-m3:batchMiniMax-M3 is a multimodal foundation model from MiniMax.

step-3.5-flashStep 3.5 Flash is StepFun's most capable open-source foundation model.
Mistral Medium 3.5 is a dense 128B instruction-following model from Mistral AI.
MiMo-V2.5 is a native omnimodal model by Xiaomi.

reka-flash-3Reka Flash 3 is a general-purpose, instruction-tuned large language model with 21 billion parameters, developed by Reka.

auto-betaAuto Router (Beta) is a task-aware router from OpenRouter.

sonar-proNote: Sonar Pro pricing includes Perplexity search pricing.

relace-apply-3Relace Apply 3 is a specialized code-patching LLM that merges AI-suggested edits straight into your source files.

nova-2-lite-v1Nova 2 Lite is a fast, cost-effective reasoning model for everyday workloads that can process text, images, and videos to generate text.

skyfall-36b-v2Skyfall 36B v2 is an enhanced iteration of Mistral Small 2501, specifically fine-tuned for improved creativity, nuanced writing, role-pla…

hermes-3-llama-3.1-405bHermes 3 is a generalist language model with many improvements over Hermes 2, including advanced agentic capabilities, much better rolepl…

morph-v3-fastMorph's fastest apply model for code edits.

l3.3-euryale-70bEuryale L3.3 70B is a model focused on creative roleplay from [Sao10k](https://ko-fi.com/sao10k).
[Microsoft Research](/microsoft) Phi-4 is designed to perform well in complex reasoning tasks and can operate efficiently in situations w…
Llama 4 Maverick 17B Instruct (128E) is a high-capacity multimodal language model from Meta, built on a mixture-of-experts (MoE) architec…

gpt-terra-latestThis model always redirects to the latest model in the GPT Terra family.

gpt-5.5-proGPT-5.5 Pro is OpenAI’s high-capability model optimized for deep reasoning and accuracy on complex, high-stakes workloads.
Claude Fable 5.1 improves on Claude Fable 5 across the board, with the biggest gains in agentic coding, long-running agentic workflows, a…

claude-sonnet-latestThis model always redirects to the latest model in the Claude Sonnet family.

fugu-ultraFugu Ultra is the higher-performance model in Sakana AI's Fugu family.
Qwen3.8 2.4T A95B is an open-weight sparse mixture-of-experts model from Qwen and the open-weight variant of [Qwen3.8 Max](/qwen/qwen3.8-…
Muse Spark 1.2 is a reasoning model from Meta, designed for complex agentic tasks.
Grok Build 0.1 is SpaceXAI’s fast coding model trained specifically for agentic software engineering workflows.
Gemini 3.7 Flash is a multimodal model from Google for fast agentic workflows, coding, and complex multi-step reasoning.
GLM-5.3-Flash is a native multimodal model from Z.ai.

kimi-k2.7-codeMoonshotAI: Kimi K2.7 Code is a coding-focused model in Moonshot AI's Kimi K2 family, built to complete end-to-end programming tasks reli…
Hy-MT2-1.8B is a compact 1.8B-parameter translation model from Tencent.
DeepSeek V4 Pro 0813 is a large-scale mixture-of-experts model from DeepSeek.

ling-3.0-flash-sante:freeLing 3.0 Flash Sante is a health and medicine-focused mixture-of-experts model from InclusionAI, built on Ling 3.0 Flash with 5.1B active…

aion-2.0Aion-2.0 is a variant of DeepSeek V3.2 optimized for immersive roleplaying and storytelling.

seed-2.0-liteSeed-2.0-Lite is a versatile, cost‑efficient enterprise workhorse that delivers strong multimodal and agent capabilities while offering n…

nemotron-3-ultra-550b-a55bNVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (…

inklingInkling is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 41B active parameters out of 975B total.

laguna-xs-2.1Laguna XS 2.1 is the latest coding agent model in the 33B-A3B category from [Poolside](https://poolside.ai/) and a step forward from thei…
command-r-plus-08-2024 is an update of the [Command R+](/models/cohere/command-r-plus) with roughly 50% higher throughput and 25% lower l…

minimax-m2.7MiniMax-M2.7 is a next-generation large language model designed for autonomous, real-world productivity and continuous improvement.
Mistral Small 4 is the next major release in the Mistral Small family, unifying the capabilities of several flagship Mistral models into…

fusionFusion turns your prompt into a small multi-model deliberation.

sonar-reasoning-proNote: Sonar Pro pricing includes Perplexity search pricing.

nova-pro-v1Amazon Nova Pro 1.0 is a capable multimodal model from Amazon focused on providing a combination of accuracy, speed, and cost for a wide…

unslopnemo-12bUnslopNemo v4.1 is the latest addition from the creator of Rocinante, designed for adventure writing and role-play scenarios.

hermes-3-llama-3.1-70bHermes 3 is a generalist language model with many improvements over [Hermes 2](/models/nousresearch/nous-hermes-2-mistral-7b-dpo), includ…

l3-lunaris-8bLunaris 8B is a versatile generalist and roleplaying model based on Llama 3.
Llama Guard 4 is a Llama 4 Scout-derived multimodal pretrained model, fine-tuned for content safety classification.

gpt-luna-latestThis model always redirects to the latest model in the GPT Luna family.

gpt-5.5-pro:batchGPT-5.5 Pro is OpenAI’s high-capability model optimized for deep reasoning and accuracy on complex, high-stakes workloads.
Claude Opus 5 is Anthropic’s flagship model for demanding reasoning, coding, and long-horizon agentic work.

claude-haiku-latestThis model always redirects to the latest model in the Claude Haiku family.

sakana-namazuSakana Namazu is a Japanese-specialized reasoning model from Sakana AI, based on Kimi K2.6 with additional training for Japanese language…
Qwen3.8 Flash is a multimodal reasoning model from Alibaba.
Muse Spark 1.2 contributor tier is a reasoning model from Meta designed for developers who want to start building at an even lower cost.
Grok 4.3 is a reasoning model from SpaceXAI.
Gemini 3.7 Flash is a multimodal model from Google for fast agentic workflows, coding, and complex multi-step reasoning.
GLM-5.3-Flash is a native multimodal model from Z.ai.

kimi-k2.6Kimi K2.6 is Moonshot AI's next-generation multimodal model, designed for long-horizon coding, coding-driven UI/UX generation, and multi-…
Hy-MT2-7B is a 7B-parameter translation model from Tencent.
DeepSeek V4 Pro 0813 is a large-scale mixture-of-experts model from DeepSeek.

ling-3.0-flash-finLing 3.0 Flash Fin is a finance-focused mixture-of-experts model from InclusionAI, built on Ling 3.0 Flash with 5.1B active parameters ou…

aion-rp-llama-3.1-8bAion-RP-Llama-3.1-8B ranks the highest in the character evaluation portion of the RPBench-Auto benchmark, a roleplaying-specific variant…

seed-2.0-miniSeed-2.0-mini targets latency-sensitive, high-concurrency, and cost-sensitive scenarios, emphasizing fast response and flexible inference…

nemotron-3.5-content-safetyNVIDIA Nemotron 3.5 Content Safety is a compact 4B-parameter multimodal guardrail model from NVIDIA, fine-tuned from Google Gemma-3-4B.

inkling:batchInkling is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 41B active parameters out of 975B total.

laguna-xs-2.1:freeLaguna XS 2.1 is the latest coding agent model in the 33B-A3B category from [Poolside](https://poolside.ai/) and a step forward from thei…
command-r-08-2024 is an update of the [Command R](/models/cohere/command-r) with improved performance for multilingual retrieval-augmente…

minimax-m2.5MiniMax-M2.5 is a SOTA large language model designed for real-world productivity.
Mistral Small 4 is the next major release in the Mistral Small family, unifying the capabilities of several flagship Mistral models into…

pareto-codeThe Pareto Router maintains a tiered shortlist of strong coding models, ranked by [Artificial Analysis](https://artificialanalysis.ai/) c…

sonar-deep-researchSonar Deep Research is a research-focused model designed for multi-step retrieval, synthesis, and reasoning across complex topics.

nova-lite-v1Amazon Nova Lite 1.0 is a very low-cost multimodal model from Amazon that focused on fast processing of image, video, and text inputs to…
Llama 4 Scout 17B Instruct (16E) is a mixture-of-experts (MoE) language model developed by Meta, activating 17 billion parameters out of…

gpt-mini-latestThis model always redirects to the latest model in the GPT Mini family.

gpt-5.4-proGPT-5.4 Pro is OpenAI's most advanced model, building on GPT-5.4's unified architecture with enhanced reasoning capabilities for complex,…
Claude Fable 5 is a Mythos-class model from Anthropic, built for autonomous knowledge work and coding.
Qwen3.8 27B is an open-weight dense vision-language model from Qwen.
Muse Glimmer 30B is a dense, open-weight multimodal model from Meta Superintelligence Labs, distilled from Muse Spark and optimized for a…
Grok 4.3 is a reasoning model from SpaceXAI.
Gemini 3.6 Flash is a high-efficiency model from Google for coding, agentic workflows, and web and app development.
GLM 5.2 is a large-scale reasoning model from Z.ai.

kimi-k2.5Kimi K2.5 is Moonshot AI's native multimodal model, delivering state-of-the-art visual coding capability and a self-directed agent swarm…
Hy3 is a 295B-parameter Mixture-of-Experts model from Tencent (21B active, 192 experts with top-8 routing) built for reasoning, agentic w…
DeepSeek V4 Flash Vision Exp is an experimental vision-enabled version of [DeepSeek V4 Flash 0731](https://openrouter.ai/deepseek/deepsee…

ling-3.0-flash-fin:freeLing 3.0 Flash Fin is a finance-focused mixture-of-experts model from InclusionAI, built on Ling 3.0 Flash with 5.1B active parameters ou…

seed-1.6Seed 1.6 is a general-purpose model released by the ByteDance Seed team.

nemotron-3.5-content-safety:freeNVIDIA Nemotron 3.5 Content Safety is a compact 4B-parameter multimodal guardrail model from NVIDIA, fine-tuned from Google Gemma-3-4B.

inkling-small:freeInkling Small is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 12B active parameters out of 276B to…
Command R7B (12-2024) is a small, fast update of the Command R+ model, delivered in December 2024.

minimax-m2-herMiniMax M2-her is a dialogue-first large language model built for immersive roleplay, character-driven chat, and expressive multi-turn co…
Devstral 2 is a state-of-the-art open-source model by Mistral AI specializing in agentic coding.

bodybuilderTransform your natural language requests into structured OpenRouter API request objects.

sonarSonar is lightweight, affordable, fast, and simple to use — now featuring citations and the ability to customize sources.

nova-micro-v1Amazon Nova Micro 1.0 is a text-only model that delivers the lowest latency responses in the Amazon Nova family of models at a very low c…
The Meta Llama 3.3 multilingual large language model (LLM) is a pretrained and instruction tuned generative model in 70B (text in/text out).

gpt-5.4-pro:batchGPT-5.4 Pro is OpenAI's most advanced model, building on GPT-5.4's unified architecture with enhanced reasoning capabilities for complex,…
Claude Opus 5 is Anthropic’s flagship model for demanding reasoning, coding, and long-horizon agentic work.
Qwen3.7 Flash is a vision-language reasoning model from Alibaba.
Muse Glimmer 30B is a dense, open-weight multimodal model from Meta Superintelligence Labs, distilled from Muse Spark and optimized for a…
Grok 4.20 Multi-Agent is a variant of SpaceXAI’s Grok 4.20 designed for collaborative, agent-based workflows.
Gemini 3.6 Flash is a high-efficiency model from Google for coding, agentic workflows, and web and app development.
GLM 5.2 is a large-scale reasoning model from Z.ai.

kimi-k2-thinkingKimi K2 Thinking is Moonshot AI’s most advanced open reasoning model to date, extending the K2 series into agentic, long-horizon reasoning.
Hy3 preview is a high-efficiency Mixture-of-Experts model from Tencent designed for agentic workflows and production use.
DeepSeek V4 Flash 0731 is a sparse mixture-of-experts model from DeepSeek, with 13B active parameters out of 284B total.

ling-3.0-flash*Ling-3.0-flash* is a *124B-parameter Mixture-of-Experts (MoE) model*, with approximately *5.1B parameters activated per token*.

seed-1.6-flashSeed 1.6 Flash is an ultra-fast multimodal deep thinking model by ByteDance Seed, supporting both text and visual understanding.

nemotron-3-ultra-550b-a55b:freeNVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (…

inkling:freeInkling is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 41B active parameters out of 975B total.

minimax-m2.1MiniMax-M2.1 is a lightweight, state-of-the-art large language model optimized for coding, agentic workflows, and modern application deve…
Mistral Large 3 2512 is Mistral’s most capable model to date, featuring a sparse mixture-of-experts architecture with 41B active paramete…

autoThe Auto Router automatically selects the best model for your prompt, powered by the wisdom of the market.
Llama 3.2 3B is a 3-billion-parameter multilingual large language model, optimized for advanced natural language processing tasks like di…

gpt-6-astra:batchGPT-6 Astra is OpenAI's flagship model for demanding end-to-end work.
Claude Opus 4.8 is Anthropic's most capable generally available model in the Opus family.
Qwen3.7-Max is the flagship model in Alibaba's Qwen3.7 series.
Muse Spark 1.1 is a multimodal reasoning model from Meta, built for agentic tasks.
Grok 4.20 is a reasoning model from SpaceXAI with industry-leading speed and agentic tool calling capabilities.
Nano Banana Pro is Google’s most advanced image-generation and editing model, built on Gemini 3 Pro.
GLM-5.1 delivers a major leap in coding capability, with particularly significant gains in handling long-horizon tasks.

kimi-k2-0905Kimi K2 0905 is the September update of [Kimi K2 0711](moonshotai/kimi-k2).
Hunyuan-A13B is a 13B active parameter Mixture-of-Experts (MoE) language model developed by Tencent, with a total parameter count of 80B…
DeepSeek V4 Flash 0731 is a sparse mixture-of-experts model from DeepSeek, with 13B active parameters out of 284B total.

nemotron-3-nano-omni-30b-a3b-reasoning:freeNVIDIA Nemotron™ 3 Nano Omni is a 30B-A3B open multimodal model designed to function as a perception and context sub-agent in enterprise…

minimax-m2MiniMax-M2 is a compact, high-efficiency large language model optimized for end-to-end coding and agentic workflows.
Mistral Large 3 2512 is Mistral’s most capable model to date, featuring a sparse mixture-of-experts architecture with 41B active paramete…
Meta's latest class of model (Llama 3.1) launched with a variety of sizes & flavors.

gpt-6-astra-pro:batchGPT-6 Astra Pro is the same underlying model as [GPT-6 Astra](https://openrouter.ai/openai/gpt-6-astra), served with `reasoning.mode` set…
Sonnet 5 is Anthropic's most capable Sonnet-class model, with frontier performance across coding, agents, and professional work.
Qwen3.7-Plus is a cost-effective model in Alibaba's Qwen3.7 series.
Gemini 3.5 Flash Lite is a high-efficiency model from Google with upgraded agentic capabilities.
GLM-5V-Turbo is Z.ai’s first native multimodal agent foundation model, built for vision-based coding and agent-driven tasks.

kimi-k2Kimi K2 Instruct is a large-scale Mixture-of-Experts (MoE) language model developed by Moonshot AI, featuring 1 trillion total parameters…
DeepSeek V4 Pro is a large-scale Mixture-of-Experts model from DeepSeek with 1.6T total parameters and 49B activated parameters, supporti…

nemotron-3-super-120b-a12bNVIDIA Nemotron 3 Super is a 120B-parameter open hybrid MoE model, activating just 12B parameters for maximum compute efficiency and accu…

minimax-m1MiniMax-M1 is a large-scale, open-weight reasoning model designed for extended context and high-efficiency inference.
The largest model in the Ministral 3 family, Ministral 3 14B offers frontier capabilities and performance comparable to its larger Mistra…
Llama 3.2 1B is a 1-billion-parameter language model focused on efficiently performing natural language tasks, such as summarization, dia…

gpt-5.2-proGPT-5.2 Pro is OpenAI’s most advanced model, offering major improvements in agentic coding and long context performance over GPT-5 Pro.
Opus 4.7 is the next generation of Anthropic's Opus family, built for long-running, asynchronous agents.
Qwen3.6-Max-Preview is a proprietary frontier model from Alibaba Cloud built on a sparse mixture-of-experts architecture with approximate…
Gemini 3.5 Flash Lite is a high-efficiency model from Google with upgraded agentic capabilities.
GLM-5 Turbo is a new model from Z.ai designed for fast inference and strong performance in agent-driven environments such as OpenClaw sce…
DeepSeek V4 Flash is an efficiency-optimized Mixture-of-Experts model from DeepSeek with 284B total parameters and 13B activated paramete…

nemotron-3-super-120b-a12b:freeNVIDIA Nemotron 3 Super is a 120B-parameter open hybrid MoE model, activating just 12B parameters for maximum compute efficiency and accu…

minimax-01MiniMax-01 is a combines MiniMax-Text-01 for text generation and MiniMax-VL-01 for image understanding.
A balanced model in the Ministral 3 family, Ministral 3 8B is a powerful, efficient tiny language model with vision capabilities.

gpt-5.4-image-2[GPT-5.4](https://openrouter.ai/openai/gpt-5.4) Image 2 combines OpenAI's GPT-5.4 model with state-of-the-art image generation capabiliti…
Sonnet 5 is Anthropic's most capable Sonnet-class model, with frontier performance across coding, agents, and professional work.
Qwen3.5 Plus (April 2026) is a large-scale multimodal language model from Alibaba.
Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image) is Google's fastest, most cost-efficient Gemini image model, built for high-velocity dev…
GLM-5 is Z.ai’s flagship open-source foundation model engineered for complex systems design and long-horizon agent workflows.
DeepSeek-V3.2 is a large language model designed to harmonize high computational efficiency with strong reasoning and agentic tool-use pe…

nemotron-3-nano-30b-a3bNVIDIA Nemotron 3 Nano 30B A3B is a small language MoE model with highest compute efficiency and accuracy for developers to build special…
The smallest model in the Ministral 3 family, Ministral 3 3B is a powerful, efficient tiny language model with vision capabilities.

gpt-5.6-terra-proGPT-5.6 Terra Pro is the same underlying model as [GPT-5.6 Terra](https://openrouter.ai/openai/gpt-5.6-terra), served with `reasoning.mod…
Claude Opus 4.8 is Anthropic's most capable generally available model in the Opus family.
Qwen3.6 27B is a dense 27-billion-parameter language model from the Qwen Team at Alibaba, released in April 2026.
Gemini 3.1 Flash Image, a.k.a.
As a 30B-class SOTA model, GLM-4.7-Flash offers a new option that balances performance and efficiency.
DeepSeek-V3.2-Exp is an experimental large language model released by DeepSeek as an intermediate step between V3.1 and future architectu…
A balanced model in the Ministral 3 family, Ministral 3 8B is a powerful, efficient tiny language model with vision capabilities.

gpt-5.6-terraGPT-5.6 Terra is a balanced model in OpenAI's GPT-5.6 series, positioned between the flagship Sol tier and the cost-efficient Luna tier.
Claude Opus 4.1 is an updated version of Anthropic’s flagship model, offering improved performance in coding, reasoning, and agentic tasks.
Qwen3.6 Flash is a fast, efficient language model from Alibaba's Qwen 3.6 series.
Gemini 3.5 Flash is Google's high-efficiency multimodal model, bringing near-Pro level coding and reasoning at Flash-tier cost and speed.
GLM-4.7 is Z.ai’s latest flagship model, featuring upgrades in two key areas: enhanced programming capabilities and more stable multi-ste…
DeepSeek R1 Distill Llama 70B is a distilled large language model based on [Llama-3.3-70B-Instruct](/meta-llama/llama-3.3-70b-instruct),…
Voxtral Small is an enhancement of Mistral Small 3, incorporating state-of-the-art audio input capabilities while retaining best-in-class…

gpt-5.6-sol-proGPT-5.6 Sol Pro is the same underlying model as [GPT-5.6 Sol](https://openrouter.ai/openai/gpt-5.6-sol), served with `reasoning.mode` set…
Claude Opus 4 is benchmarked as the world’s best coding model, at time of release, bringing sustained performance on complex, long-runnin…
Qwen3.6-35B-A3B is an open-weight multimodal model from Alibaba Cloud with 35 billion total parameters and 3 billion active parameters pe…
Gemini 3.5 Flash is Google's high-efficiency multimodal model, bringing near-Pro level coding and reasoning at Flash-tier cost and speed.
GLM-4.6V is a large multimodal model designed for high-fidelity visual understanding and long-context reasoning across images, documents,…
DeepSeek-V3.1 Terminus is an update to [DeepSeek V3.1](/deepseek/deepseek-chat-v3.1) that maintains the model's original capabilities whi…
This is Mistral AI's flagship model, Mistral Large 2 (version mistral-large-2407).

gpt-5.6-solGPT-5.6 Sol is the flagship model in OpenAI's GPT-5.6 series.
Opus 4.7 is the next generation of Anthropic's Opus family, built for long-running, asynchronous agents.
Qwen 3.6 Plus builds on a hybrid architecture that combines efficient linear attention with sparse mixture-of-experts routing, enabling s…
Gemini 3.1 Flash Lite is Google’s GA high-efficiency multimodal model optimized for low-latency, high-volume workloads.
Compared with GLM-4.5, this generation brings several key improvements: Longer context window: The context window has been expanded from…
DeepSeek R1 is here: Performance on par with [OpenAI o1](/openai/o1), but open-sourced and with fully open reasoning tokens.
Mistral's official instruct fine-tuned version of [Mixtral 8x22B](/models/mistralai/mixtral-8x22b).

gpt-chat-latestGPT Chat Latest points to OpenAI's stable API alias `chat-latest` that always resolves to the latest Instant chat model used in ChatGPT.
Opus 4.6 is Anthropic’s strongest model for coding and long-running professional tasks.
Qwen3.5-9B is a multimodal foundation model from the Qwen3.5 family, designed to deliver strong reasoning, coding, and visual understandi…
Gemini 3.1 Flash Lite is Google’s GA high-efficiency multimodal model optimized for low-latency, high-volume workloads.
GLM-4.5V is a vision-language foundation model for multimodal agent applications.
May 28th update to the [original DeepSeek R1](/deepseek/deepseek-r1) Performance on par with [OpenAI o1](/openai/o1), but open-sourced an…
This is Mistral AI's flagship model, Mistral Large 2 (version `mistral-large-2407`).

gpt-5-proGPT-5 Pro is OpenAI’s most advanced model, offering major improvements in reasoning, code quality, and user experience.
Sonnet 4.6 is Anthropic's most capable Sonnet-class model yet, with frontier performance across coding, agents, and professional work.
Qwen3.5-9B is a multimodal foundation model from the Qwen3.5 family, designed to deliver strong reasoning, coding, and visual understandi…
Gemma 4 31B Instruct is Google DeepMind's 30.7B dense multimodal model supporting text and image input with text output.
GLM-4.5 is our latest flagship foundation model, purpose-built for agent-based applications.
DeepSeek-V3 is the latest model from the DeepSeek team, building upon the instruction following and coding abilities of the previous vers…
Mistral Medium 3.1 is an updated version of Mistral Medium 3, which is a high-performance enterprise-grade language model designed to del…

gpt-5.5GPT-5.5 is OpenAI’s frontier model designed for complex professional workloads, building on GPT-5.4 with stronger reasoning, higher relia…
Opus 4.6 is Anthropic’s strongest model for coding and long-running professional tasks.
The Qwen3.5 Series 35B-A3B is a native vision-language model designed with a hybrid architecture that integrates linear attention mechani…
Gemini 3.1 Pro Preview Custom Tools is a variant of Gemini 3.1 Pro that improves tool selection behavior by preventing overuse of a gener…
GLM-4.5-Air is the lightweight variant of our latest flagship model family, also purpose-built for agent-centric applications.
DeepSeek-V3.1 is a large hybrid reasoning model (671B parameters, 37B active) that supports both thinking and non-thinking modes via prom…
Mistral Medium 3 is a high-performance enterprise-grade language model designed to deliver frontier-level capabilities at significantly r…

gpt-5.6-terra-pro:batchGPT-5.6 Terra Pro is the same underlying model as [GPT-5.6 Terra](https://openrouter.ai/openai/gpt-5.6-terra), served with `reasoning.mod…
Sonnet 4.6 is Anthropic's most capable Sonnet-class model yet, with frontier performance across coding, agents, and professional work.
The Qwen3.5 122B-A10B native vision-language model is built on a hybrid architecture that integrates a linear attention mechanism with a…
Gemma 4 31B Instruct is Google DeepMind's 30.7B dense multimodal model supporting text and image input with text output.
DeepSeek V3, a 685B-parameter, mixture-of-experts model, is the latest iteration of the flagship chat model family from the DeepSeek team.
Mistral Small 3.1 24B Instruct is an upgraded variant of Mistral Small 3 (2501), featuring 24 billion parameters with advanced multimodal…

gpt-5.6-terra:batchGPT-5.6 Terra is a balanced model in OpenAI's GPT-5.6 series, positioned between the flagship Sol tier and the cost-efficient Luna tier.
Claude Opus 4.5 is Anthropic’s frontier reasoning model optimized for complex software engineering, agentic workflows, and long-horizon c…
The Qwen3.5 27B native vision-language Dense model incorporates a linear attention mechanism, delivering fast response times while balanc…
Gemma 4 26B A4B IT is an instruction-tuned Mixture-of-Experts (MoE) model from Google DeepMind.
Mistral's cutting-edge language model for coding released end of July 2025.

gpt-5.6-sol-pro:batchGPT-5.6 Sol Pro is the same underlying model as [GPT-5.6 Sol](https://openrouter.ai/openai/gpt-5.6-sol), served with `reasoning.mode` set…
Claude Opus 4.1 is an updated version of Anthropic’s flagship model, offering improved performance in coding, reasoning, and agentic tasks.
The Qwen3.5 series 397B-A17B native vision-language model is built on a hybrid architecture that integrates a linear attention mechanism…
Gemma 4 26B A4B IT is an instruction-tuned Mixture-of-Experts (MoE) model from Google DeepMind.
Mistral Medium 3.1 is an updated version of Mistral Medium 3, which is a high-performance enterprise-grade language model designed to del…

gpt-5.6-sol:batchGPT-5.6 Sol is the flagship model in OpenAI's GPT-5.6 series.
Claude Opus 4.5 is Anthropic’s frontier reasoning model optimized for complex software engineering, agentic workflows, and long-horizon c…
The Qwen3.5 native vision-language Flash models are built on a hybrid architecture that integrates a linear attention mechanism with a sp…
Gemma 4 31B Instruct is Google DeepMind's 30.7B dense multimodal model supporting text and image input with text output.
Mistral Saba is a 24B-parameter language model specifically designed for the Middle East and South Asia, delivering accurate and contextu…

gpt-5.6-luna-proGPT-5.6 Luna Pro is the same underlying model as [GPT-5.6 Luna](https://openrouter.ai/openai/gpt-5.6-luna), served with `reasoning.mode`…
Claude Sonnet 4.5 is Anthropic’s most advanced Sonnet model to date, optimized for real-world agents and coding workflows.
Qwen3-Max-Thinking is the flagship reasoning model in the Qwen3 series, designed for high-stakes cognitive tasks that require deep, multi…
Gemini 3.1 Pro Preview is Google’s frontier reasoning model, delivering enhanced software engineering performance, improved agentic relia…
Mistral's cutting-edge language model for coding released end of July 2025.

gpt-5.6-lunaGPT-5.6 Luna is a fast, cost-efficient model in OpenAI's GPT-5.6 series.
Claude Sonnet 4 significantly enhances the capabilities of its predecessor, Sonnet 3.7, excelling in both coding and reasoning tasks with…
The Qwen3.5 native vision-language series Plus models are built on a hybrid architecture that integrates linear attention mechanisms with…
Full-length songs are priced at $0.08 per song.
Mistral-Small-3.2-24B-Instruct-2506 is an updated 24B parameter model from Mistral optimized for instruction following, repetition reduct…

gpt-5.6-luna-pro:batchGPT-5.6 Luna Pro is the same underlying model as [GPT-5.6 Luna](https://openrouter.ai/openai/gpt-5.6-luna), served with `reasoning.mode`…
Claude Haiku 4.5 is Anthropic’s fastest and most efficient model, delivering near-frontier intelligence at a fraction of the cost and lat…
Qwen3-Coder-Next is an open-weight causal language model optimized for coding agents and local development workflows.
30 second duration clips are priced at $0.04 per clip.
Mistral Small 3 is a 24B-parameter language model optimized for low-latency performance across common AI tasks.

gpt-5.6-luna:batchGPT-5.6 Luna is a fast, cost-efficient model in OpenAI's GPT-5.6 series.
Claude Sonnet 4.5 is Anthropic’s most advanced Sonnet model to date, optimized for real-world agents and coding workflows.
Qwen3-VL-32B-Instruct is a large-scale multimodal vision-language model designed for high-precision understanding and reasoning across te…
Gemini 3.1 Pro Preview is Google’s frontier reasoning model, delivering enhanced software engineering performance, improved agentic relia…
A 12B parameter model with a 128k token context length built by Mistral in collaboration with NVIDIA.

o3-proThe o-series of models are trained with reinforcement learning to think before they answer and perform complex reasoning.
Claude Haiku 4.5 is Anthropic’s fastest and most efficient model, delivering near-frontier intelligence at a fraction of the cost and lat…
Qwen3-VL-8B-Thinking is the reasoning-optimized variant of the Qwen3-VL-8B multimodal model, designed for advanced visual and textual rea…
Gemini 3.1 Flash Image Preview, a.k.a.

o1-proThe o1 series of models are trained with reinforcement learning to think before they answer and perform complex reasoning.
Claude 3 Haiku is Anthropic's fastest and most compact model for near-instant responsiveness.
Qwen3-VL-8B-Instruct is a multimodal vision-language model from the Qwen3-VL series, built for high-fidelity understanding and reasoning…
Gemini 3.1 Flash Lite Preview is Google's high-efficiency model optimized for high-volume use cases.

o1The latest and strongest model family from OpenAI, o1 is designed to spend more time thinking before responding.
Qwen3-VL-30B-A3B-Thinking is a multimodal model that unifies strong text generation with visual understanding for images and videos.
Nano Banana Pro is Google’s most advanced image-generation and editing model, built on Gemini 3 Pro.

gpt-4OpenAI's flagship model, GPT-4 is a large-scale multimodal language model capable of solving difficult problems with greater accuracy tha…
Qwen3-VL-30B-A3B-Instruct is a multimodal model that unifies strong text generation with visual understanding for images and videos.
Gemini 3 Flash Preview is a high speed, high value thinking model designed for agentic workflows, multi turn chat, and coding assistance.

gpt-5.2-pro:batchGPT-5.2 Pro is OpenAI’s most advanced model, offering major improvements in agentic coding and long context performance over GPT-5 Pro.
Qwen3-Max is an updated release built on the Qwen3 series, offering major improvements in reasoning, instruction following, multilingual…
Gemini 3 Flash Preview is a high speed, high value thinking model designed for agentic workflows, multi turn chat, and coding assistance.

gpt-5.5:batchGPT-5.5 is OpenAI’s frontier model designed for complex professional workloads, building on GPT-5.4 with stronger reasoning, higher relia…
Qwen3 Coder Plus is Alibaba's proprietary version of the Open Source Qwen3 Coder 480B A35B.
Gemini 2.5 Flash Image, a.k.a.

gpt-5-image[GPT-5](https://openrouter.ai/openai/gpt-5) Image combines OpenAI's GPT-5 model with state-of-the-art image generation capabilities.
Qwen3-VL-235B-A22B Thinking is a multimodal model that unifies strong text generation with visual understanding across images and video.
Gemini 2.5 Pro is Google’s state-of-the-art AI model designed for advanced reasoning, coding, mathematics, and scientific tasks.

gpt-5.4GPT-5.4 is OpenAI’s latest frontier model, unifying the Codex and GPT lines into a single system.
Qwen2.5-VL is proficient in recognizing common objects such as flowers, birds, fish, and insects.
Gemini 2.5 Pro is Google’s state-of-the-art AI model designed for advanced reasoning, coding, mathematics, and scientific tasks.

gpt-5.4-miniGPT-5.4 mini brings the core capabilities of GPT-5.4 to a faster, more efficient model optimized for high-throughput workloads.
Qwen3-VL-235B-A22B Instruct is an open-weight multimodal model that unifies strong text generation with visual understanding across image…
Gemini 2.5 Pro is Google’s state-of-the-art AI model designed for advanced reasoning, coding, mathematics, and scientific tasks.

gpt-4-turboThe latest GPT-4 Turbo model with vision capabilities.
Qwen2.5-Coder is the latest series of Code-Specific Qwen large language models (formerly known as CodeQwen).
Gemma 2 27B by Google is an open model built from the same research and technology used to create the [Gemini models](/models?q=gemini).

gpt-4-turbo-previewThe preview GPT-4 model with improved instruction following, JSON mode, reproducible outputs, parallel function calling, and more.
Qwen3-235B-A22B is a 235B parameter mixture-of-experts (MoE) model developed by Qwen, activating 22B parameters per forward pass.
Gemini 2.5 Pro is Google’s state-of-the-art AI model designed for advanced reasoning, coding, mathematics, and scientific tasks.

gpt-5.3-codexGPT-5.3-Codex is OpenAI’s most advanced agentic coding model, combining the frontier software engineering performance of GPT-5.2-Codex wi…
Qwen3 Coder Flash is Alibaba's fast and cost efficient version of their proprietary Qwen3 Coder Plus.
Gemini 2.5 Flash is Google's state-of-the-art workhorse model, specifically designed for advanced reasoning, coding, mathematics, and sci…

gpt-5.4:batchGPT-5.4 is OpenAI’s latest frontier model, unifying the Codex and GPT lines into a single system.
Qwen2.5 72B is the latest series of Qwen large language models.
Gemini 2.5 Flash is Google's state-of-the-art workhorse model, specifically designed for advanced reasoning, coding, mathematics, and sci…

gpt-5.4-mini:batchGPT-5.4 mini brings the core capabilities of GPT-5.4 to a faster, more efficient model optimized for high-throughput workloads.
Qwen3-Coder-480B-A35B-Instruct is a Mixture-of-Experts (MoE) code generation model developed by the Qwen team.
Gemini 2.5 Flash-Lite is a lightweight reasoning model in the Gemini 2.5 family, optimized for ultra-low latency and cost efficiency.

gpt-5.4-nanoGPT-5.4 nano is the most lightweight and cost-efficient variant of the GPT-5.4 family, optimized for speed-critical and high-volume tasks.
Qwen Plus 0728, based on the Qwen3 foundation model, is a 1 million context hybrid reasoning model with a balanced performance, speed, an…
Gemma 3 introduces multimodality, supporting vision-language input and text outputs.

gpt-5.4-nano:batchGPT-5.4 nano is the most lightweight and cost-efficient variant of the GPT-5.4 family, optimized for speed-critical and high-volume tasks.
Qwen-Plus, based on the Qwen2.5 foundation model, is a 131K context model with a balanced performance, speed, and cost combination.
Gemini 2.5 Flash-Lite is a lightweight reasoning model in the Gemini 2.5 family, optimized for ultra-low latency and cost efficiency.

gpt-audioThe gpt-audio model is OpenAI's first generally available audio model.
Qwen3-235B-A22B-Thinking-2507 is a high-performance, open-weight Mixture-of-Experts (MoE) language model optimized for complex reasoning…
Gemma 3 introduces multimodality, supporting vision-language input and text outputs.

gpt-5-pro:batchGPT-5 Pro is OpenAI’s most advanced model, offering major improvements in reasoning, code quality, and user experience.
Qwen3-14B is a dense 14.8B parameter causal language model from the Qwen3 series, designed for both complex reasoning and efficient dialo…
Gemma 3 introduces multimodality, supporting vision-language input and text outputs.

gpt-5.2-codexGPT-5.2-Codex is an upgraded version of GPT-5.1-Codex optimized for software engineering and coding workflows.
Qwen3-30B-A3B-Thinking-2507 is a 30B parameter Mixture-of-Experts reasoning model optimized for complex tasks requiring extended multi-st…

gpt-audio-miniA cost-efficient version of GPT Audio.
Qwen3-Next-80B-A3B-Thinking is a reasoning-first chat model in the Qwen3-Next line that outputs structured “thinking” traces by default.

gpt-5.2-chatGPT-5.2 Chat (AKA Instant) is the fast, lightweight member of the 5.2 family, optimized for low-latency chat while retaining strong gener…
Qwen3, the latest generation in the Qwen large language model series, features both dense and mixture-of-experts (MoE) architectures to e…

gpt-5.2GPT-5.2 is the latest frontier-grade model in the GPT-5 series, offering stronger agentic and long context perfomance compared to GPT-5.1.
Qwen3-8B is a dense 8.2B parameter causal language model from the Qwen3 series, designed for both reasoning-heavy tasks and efficient dia…

gpt-5.1-codex-maxGPT-5.1-Codex-Max is OpenAI’s latest agentic coding model, designed for long-running, high-context software development tasks.
Qwen2.5 7B is the latest series of Qwen large language models.

gpt-5.2:batchGPT-5.2 is the latest frontier-grade model in the GPT-5 series, offering stronger agentic and long context perfomance compared to GPT-5.1.
Qwen3-Next-80B-A3B-Instruct is an instruction-tuned chat model in the Qwen3-Next series optimized for fast, stable responses without “thi…

gpt-4o-2024-05-13GPT-4o ("o" for "omni") is OpenAI's latest AI model, supporting both text and image inputs with text outputs.
Qwen3-30B-A3B-Instruct-2507 is a 30.5B-parameter mixture-of-experts language model from Qwen, with 3.3B active parameters per inference.

gpt-4-turbo:batchThe latest GPT-4 Turbo model with vision capabilities.
Qwen3-235B-A22B-Instruct-2507 is a multilingual, instruction-tuned mixture-of-experts language model based on the Qwen3-235B architecture…

gpt-5.1GPT-5.1 is the latest frontier-grade model in the GPT-5 series, offering stronger general-purpose reasoning, improved instruction adheren…
Qwen3-32B is a dense 32.8B parameter causal language model from the Qwen3 series, optimized for both complex reasoning and efficient dial…

gpt-5.1-codexGPT-5.1-Codex is a specialized version of GPT-5.1 optimized for software engineering and coding workflows.
Qwen3-Coder-30B-A3B-Instruct is a 30.5B parameter Mixture-of-Experts (MoE) model with 128 experts (8 active per forward pass), designed f…

gpt-5-image-miniGPT-5 Image Mini combines OpenAI's advanced language capabilities, powered by [GPT-5 Mini](https://openrouter.ai/openai/gpt-5-mini), with…

gpt-5.1:batchGPT-5.1 is the latest frontier-grade model in the GPT-5 series, offering stronger general-purpose reasoning, improved instruction adheren…

gpt-5.1-codex-miniGPT-5.1-Codex-Mini is a smaller and faster version of GPT-5.1-Codex

gpt-3.5-turbo-16kThis model offers four times the context length of gpt-3.5-turbo, allowing it to support approximately 20 pages of text in a single reque…

gpt-4o-2024-11-20The 2024-11-20 version of GPT-4o offers a leveled-up creative writing ability with more natural, engaging, and tailored writing to improv…

gpt-4o-2024-08-06The 2024-08-06 version of GPT-4o offers improved performance in structured outputs, with the ability to supply a JSON schema in the respo…

gpt-4oGPT-4o ("o" for "omni") is OpenAI's latest AI model, supporting both text and image inputs with text outputs.

gpt-oss-safeguard-20bgpt-oss-safeguard-20b is a safety reasoning model from OpenAI built upon gpt-oss-20b.

o3o3 is a well-rounded and powerful model across domains.

gpt-4.1GPT-4.1 is a flagship large language model optimized for advanced instruction following, real-world software engineering, and long-contex…

gpt-3.5-turbo-instructThis model is a variant of GPT-3.5 Turbo tuned for instructional prompts and omitting chat-related optimizations.

gpt-5GPT-5 is OpenAI’s most advanced model, offering major improvements in reasoning, code quality, and user experience.

gpt-4o:batchGPT-4o ("o" for "omni") is OpenAI's latest AI model, supporting both text and image inputs with text outputs.

o4-mini-highOpenAI o4-mini-high is the same model as [o4-mini](/openai/o4-mini) with reasoning_effort set to high.

o4-miniOpenAI o4-mini is a compact reasoning model in the o-series, optimized for fast, cost-efficient performance while retaining strong multim…

o3-mini-highOpenAI o3-mini-high is the same model as [o3-mini](/openai/o3-mini) with reasoning_effort set to high.

o3-miniOpenAI o3-mini is a cost-efficient language model optimized for STEM reasoning tasks, particularly excelling in science, mathematics, and…

o3:batcho3 is a well-rounded and powerful model across domains.

gpt-4.1:batchGPT-4.1 is a flagship large language model optimized for advanced instruction following, real-world software engineering, and long-contex…

gpt-3.5-turbo-0613GPT-3.5 Turbo is OpenAI's fastest model.

gpt-5:batchGPT-5 is OpenAI’s most advanced model, offering major improvements in reasoning, code quality, and user experience.

o4-mini:batchOpenAI o4-mini is a compact reasoning model in the o-series, optimized for fast, cost-efficient performance while retaining strong multim…

o3-mini:batchOpenAI o3-mini is a cost-efficient language model optimized for STEM reasoning tasks, particularly excelling in science, mathematics, and…

gpt-3.5-turboGPT-3.5 Turbo is OpenAI's fastest model.

gpt-4.1-miniGPT-4.1 Mini is a mid-sized model delivering performance competitive with GPT-4o at substantially lower latency and cost.

gpt-5-miniGPT-5 Mini is a compact version of GPT-5, designed to handle lighter-weight reasoning tasks.

gpt-3.5-turbo:batchGPT-3.5 Turbo is OpenAI's fastest model.

gpt-4.1-mini:batchGPT-4.1 Mini is a mid-sized model delivering performance competitive with GPT-4o at substantially lower latency and cost.

gpt-oss-120b:batchgpt-oss-120b is an open-weight, 117B-parameter Mixture-of-Experts (MoE) language model from OpenAI designed for high-reasoning, agentic,…

gpt-4o-miniGPT-4o mini is OpenAI's newest model after [GPT-4 Omni](/models/openai/gpt-4o), supporting both text and image inputs with text outputs.

gpt-4o-mini-2024-07-18GPT-4o mini is OpenAI's newest model after [GPT-4 Omni](/models/openai/gpt-4o), supporting both text and image inputs with text outputs.

gpt-5-mini:batchGPT-5 Mini is a compact version of GPT-5, designed to handle lighter-weight reasoning tasks.

gpt-4.1-nanoFor tasks that demand low latency, GPT‑4.1 nano is the fastest and cheapest model in the GPT-4.1 series.

gpt-4o-mini:batchGPT-4o mini is OpenAI's newest model after [GPT-4 Omni](/models/openai/gpt-4o), supporting both text and image inputs with text outputs.

gpt-5-nanoGPT-5-Nano is the smallest and fastest variant in the GPT-5 system, optimized for developer tools, rapid interactions, and ultra-low late…

gpt-oss-20b:batchgpt-oss-20b is an open-weight 21B parameter model released by OpenAI under the Apache 2.0 license.

gpt-4.1-nano:batchFor tasks that demand low latency, GPT‑4.1 nano is the fastest and cheapest model in the GPT-4.1 series.

gpt-oss-120bgpt-oss-120b is an open-weight, 117B-parameter Mixture-of-Experts (MoE) language model from OpenAI designed for high-reasoning, agentic,…

gpt-oss-20bgpt-oss-20b is an open-weight 21B parameter model released by OpenAI under the Apache 2.0 license.

gpt-5-nano:batchGPT-5-Nano is the smallest and fastest variant in the GPT-5 system, optimized for developer tools, rapid interactions, and ultra-low late…
Pick a model and ship today.
Spin up any agent framework free, route to any of these models, and only pay for the tokens you use.