{"data":{"models":[{"id":"claude-fable-5-1","provider":"anthropic","family":"claude-fable","display_name":"Claude Fable 5.1","description":"Anthropic's current flagship ('Latest'); successor to Claude Fable 5. Released Sept 1, 2026, GA across Claude API, AWS, Google Cloud, and Microsoft Foundry. Same base input/output pricing as Fable 5, but cache reads cut to 0.025x the input price (vs 0.1x on every other Claude model) - up to ~45% cheaper on long, cache-heavy agentic sessions, ~25% cheaper for typical workloads. Adaptive thinking always on, default effort 'high', 1M context, 128k max output. Strongest gains vs Fable 5 in long-running agentic coding, document/spreadsheet/slide work, multistep research, dense-document vision, and computer use. Breaking changes vs Fable 5: forced tool_choice ('any'/'tool') now returns a 400 error; thinking blocks are one-way compatible (Fable 5.1 can read earlier models' thinking blocks, but no earlier model can read its own); editing/reordering earlier turns invalidates later thinking blocks (enforced for accounts created on/after Aug 31, 2026). New: beta per-message effort, beta turn-scoped system messages, beta 'display: updates' progress narration, and Anthropic's statistical text watermark on all generated text. Temperature/top_p/top_k not supported. Training data cutoff June 2026.","type":"chat","input_modalities":["text","image"],"output_modalities":["text"],"status":"ga","released":"2026-09-01","limits":{"context_window":1000000,"max_output":128000},"pricing":{"currency":"USD","input_per_mtok":10,"cached_input_per_mtok":0.25,"cache_write_5m_per_mtok":12.5,"cache_write_1h_per_mtok":20,"output_per_mtok":50,"batch_input_per_mtok":5,"batch_output_per_mtok":25,"notes":"Same base input/output/cache-write pricing as Claude Fable 5. Cache reads (hits and refreshes) are priced at 0.025x base input ($0.25/MTok) instead of the standard 0.1x multiplier used by every other Claude model."},"capabilities":{"tools":true,"parallel_tools":true,"structured_output":true,"streaming":true,"batching":true,"prompt_caching":true,"prefill":false,"system_prompt":true,"stop_sequences":true,"reasoning":true,"reasoning_modes":["low","medium","high","xhigh","max"],"reasoning_default":"high","vision":true,"document_input":true,"computer_use":true,"web_search":true,"code_execution":true},"performance":{"benchmarks":{"terminal_bench_science":52.6,"terminal_bench":55.8,"cursor_bench":73.4,"osworld":41.7}},"api":{"endpoint":"https://api.anthropic.com/v1/messages","protocol":"anthropic","openai_compatible":false,"auth":{"type":"header","header_name":"x-api-key","env_var":"ANTHROPIC_API_KEY"},"sdks":{"anthropic":{"package":"@anthropic-ai/sdk","model_id":"claude-fable-5-1"},"vercel-ai":{"package":"@ai-sdk/anthropic","model_id":"claude-fable-5-1","factory":"anthropic('claude-fable-5-1')"}},"supported_params":["max_tokens","stop_sequences","system","tools","tool_choice","stream","metadata","thinking","service_tier"],"unsupported_params":["temperature","top_p","top_k"],"notes":"Adaptive thinking always on (no off mode); thinking:{type:\"enabled\"/\"disabled\"} both return 400. tool_choice of type \"any\" or \"tool\" (forced tool use) returns a 400 error - only \"auto\" (default) and \"none\" are supported; use strict tool use or structured outputs instead. Assistant-turn prefill returns a 400 error. Temperature, top_p, and top_k return a 400 error when set to non-default values. Thinking blocks are bound to the model that produced them: earlier Claude models cannot read this model's thinking blocks, and editing/reordering/removing an earlier turn invalidates later thinking blocks (checked for accounts created on/after 2026-08-31). Full 1M context at standard pricing. 30-day data retention; not available under zero data retention unless expressly authorized."},"aggregator_ids":{"openrouter":"anthropic/claude-fable-5.1","aws-bedrock":"anthropic.claude-fable-5-1","gcp-vertex":"claude-fable-5-1"},"sources":{"spec":"https://platform.claude.com/docs/en/models/fable-5-1/overview","pricing":"https://platform.claude.com/docs/en/about-claude/pricing","release_notes":"https://www.anthropic.com/claude-fable-and-mythos-5-1","docs":"https://platform.claude.com/docs/en/models/fable-5-1/whats-new-fable-5-1"},"last_verified":"2026-09-28"},{"id":"claude-mythos-5-1","provider":"anthropic","family":"claude-mythos","display_name":"Claude Mythos 5.1","description":"Same underlying model as Claude Fable 5.1 with lighter safeguards; successor to Claude Mythos 5. Released Sept 1, 2026, offered by invitation only through Project Glasswing's Cyber Verification Program (CVP, defensive security professionals) and Life Sciences Verification Program (LSVP, vetted researchers, currently US organizations only). Identical specs and pricing to Claude Fable 5.1, including the 0.025x cache-read discount. Contact your Anthropic, AWS, or Google Cloud account team for access. Training data cutoff June 2026. As of Sept 17, 2026, Anthropic's new Life Sciences Verification Program (beta) Standard Use tier also grants approved life-science organizations more permissive biology safeguards on Claude Opus 5 and Claude Sonnet 5 (not just Mythos 5.1) - a policy/access change, not a new model id.","type":"chat","input_modalities":["text","image"],"output_modalities":["text"],"status":"preview","released":"2026-09-01","limits":{"context_window":1000000,"max_output":128000},"pricing":{"currency":"USD","input_per_mtok":10,"cached_input_per_mtok":0.25,"cache_write_5m_per_mtok":12.5,"cache_write_1h_per_mtok":20,"output_per_mtok":50,"batch_input_per_mtok":5,"batch_output_per_mtok":25,"notes":"Invitation-only access via Project Glasswing (CVP/LSVP). Same pricing as Claude Fable 5.1, including the 0.025x cache-read multiplier."},"capabilities":{"tools":true,"parallel_tools":true,"structured_output":true,"streaming":true,"batching":true,"prompt_caching":true,"prefill":false,"system_prompt":true,"stop_sequences":true,"reasoning":true,"reasoning_modes":["low","medium","high","xhigh","max"],"reasoning_default":"high","vision":true,"document_input":true,"computer_use":true,"web_search":true,"code_execution":true},"performance":{"benchmarks":{"terminal_bench":60.9}},"api":{"endpoint":"https://api.anthropic.com/v1/messages","protocol":"anthropic","openai_compatible":false,"auth":{"type":"header","header_name":"x-api-key","env_var":"ANTHROPIC_API_KEY"},"sdks":{"anthropic":{"package":"@anthropic-ai/sdk","model_id":"claude-mythos-5-1"}},"unsupported_params":["temperature","top_p","top_k"],"notes":"Limited availability through Project Glasswing (invitation only). See https://anthropic.com/glasswing for access. Shares Claude Fable 5.1's API behavior (adaptive thinking always on, forced tool_choice unsupported, thinking-block binding), except the earlier-turn-edit invalidation check does not run on this model."},"aggregator_ids":{"aws-bedrock":"anthropic.claude-mythos-5-1","gcp-vertex":"claude-mythos-5-1"},"sources":{"spec":"https://platform.claude.com/docs/en/models/mythos-5-1/overview","pricing":"https://platform.claude.com/docs/en/about-claude/pricing","release_notes":"https://www.anthropic.com/claude-fable-and-mythos-5-1","docs":"https://www.anthropic.com/news/life-sciences-verification-program"},"last_verified":"2026-09-28"},{"id":"claude-fable-5","provider":"anthropic","family":"claude-fable","display_name":"Claude Fable 5","description":"Legacy model as of Sept 1, 2026 - superseded by Claude Fable 5.1 (still fully available, no deprecation notice). Released June 9, 2026. Was suspended worldwide June 12 - June 30, 2026 under a US export-control directive; the directive was lifted and Fable 5 was restored globally starting July 1, 2026 (Claude API, Claude Platform on AWS, Amazon Bedrock, Google Cloud Vertex AI, Microsoft Foundry). Adaptive thinking always on, 1M context, 128k max output, new tokenizer (up to 35% more tokens vs. pre-4.7 models), web search, code execution, computer use. Temperature/top_p/top_k not supported. Training data cutoff January 2026.","type":"chat","input_modalities":["text","image"],"output_modalities":["text"],"status":"ga","released":"2026-06-09","limits":{"context_window":1000000,"max_output":128000},"pricing":{"currency":"USD","input_per_mtok":10,"cached_input_per_mtok":1,"cache_write_5m_per_mtok":12.5,"cache_write_1h_per_mtok":20,"output_per_mtok":50,"batch_input_per_mtok":5,"batch_output_per_mtok":25},"capabilities":{"tools":true,"parallel_tools":true,"structured_output":true,"streaming":true,"batching":true,"prompt_caching":true,"prefill":false,"system_prompt":true,"stop_sequences":true,"reasoning":true,"reasoning_modes":["low","medium","high","xhigh","max"],"reasoning_default":"high","vision":true,"document_input":true,"computer_use":true,"web_search":true,"code_execution":true},"performance":{"intelligence_index":90},"api":{"endpoint":"https://api.anthropic.com/v1/messages","protocol":"anthropic","openai_compatible":false,"auth":{"type":"header","header_name":"x-api-key","env_var":"ANTHROPIC_API_KEY"},"sdks":{"anthropic":{"package":"@anthropic-ai/sdk","model_id":"claude-fable-5"},"vercel-ai":{"package":"@ai-sdk/anthropic","model_id":"claude-fable-5","factory":"anthropic('claude-fable-5')"}},"supported_params":["max_tokens","stop_sequences","system","tools","tool_choice","stream","metadata","thinking","service_tier"],"unsupported_params":["temperature","top_p","top_k"],"notes":"Adaptive thinking always on (no off mode). Temperature, top_p, and top_k return a 400 error when set to non-default values. New tokenizer: same text uses ~30% more tokens vs. pre-4.7 models. Full 1M context at standard pricing."},"aggregator_ids":{"openrouter":"anthropic/claude-fable-5","aws-bedrock":"anthropic.claude-fable-5","gcp-vertex":"claude-fable-5"},"sources":{"spec":"https://platform.claude.com/docs/en/models/fable-5/overview","pricing":"https://platform.claude.com/docs/en/about-claude/pricing","release_notes":"https://www.anthropic.com/news/redeploying-fable-5"},"last_verified":"2026-09-28"},{"id":"claude-opus-5-5","provider":"anthropic","family":"claude-5","display_name":"Claude Opus 5.5","description":"New flagship as of Sept 22, 2026 - first model in the 5.5 family, succeeding Claude Opus 5 as Anthropic's recommended default for most workloads. Performs at roughly Claude Fable 5.1's level on agentic coding, knowledge work, and computer use at a much lower price ($4/$20 vs Fable 5.1's $10/$50), while costing ~40% less than Opus 5 on typical workloads and producing output >30% faster. 1M context, 128k max output (300k on Batch API with output-300k-2026-03-24 beta header). Adaptive thinking always on (cannot be disabled - breaking change vs Opus 5, which allowed thinking:{type:\"disabled\"} at effort high or below); default effort 'medium' (vs Opus 5's 'high'). Breaking changes vs Opus 5: thinking cannot be disabled (400 error), forced tool_choice ('any'/'tool') returns a 400 error, thinking blocks are tied to the producing model and conversation (enforced for accounts created on/after Aug 31, 2026), and on the Claude API/Google Cloud the earlier computer_20251124 tool is rejected in favor of computer_toolset_20260801 (Bedrock keeps computer_20251124 support). Non-breaking behavior change: text between tool calls now returns in (often empty, at default display settings) thinking blocks rather than text blocks. Adds a biology safety classifier alongside the cybersecurity one, plus a reasoning_extraction refusal category. Sharper reading of dense charts/diagrams/screenshots. Temperature/top_p/top_k and assistant-turn prefill not supported. Fast mode (research preview, Claude API only): $8/$40 per MTok. Training data cutoff June 2026. Not yet listed under the Life Sciences Verification Program as of Sept 28, 2026 (LSVP Standard Use currently names Mythos 5.1, Opus 5, and Sonnet 5 only).","type":"chat","input_modalities":["text","image"],"output_modalities":["text"],"status":"ga","released":"2026-09-22","limits":{"context_window":1000000,"max_output":128000},"pricing":{"currency":"USD","input_per_mtok":4,"cached_input_per_mtok":0.2,"cache_write_5m_per_mtok":5,"cache_write_1h_per_mtok":8,"output_per_mtok":20,"batch_input_per_mtok":2,"batch_output_per_mtok":10,"notes":"Cache reads (hits/refreshes) priced at 0.05x base input ($0.20/MTok) - lower than the standard 0.1x multiplier but higher than Fable 5.1/Mythos 5.1's 0.025x. Fast mode (research preview, Claude API only): $8 input / $40 output per MTok; not available on Amazon Bedrock, Claude Platform on AWS, Google Cloud, or Microsoft Foundry, and not available with the Batch API. Prompt-cache minimum is 512 tokens."},"capabilities":{"tools":true,"parallel_tools":true,"structured_output":true,"streaming":true,"batching":true,"prompt_caching":true,"prefill":false,"system_prompt":true,"stop_sequences":true,"reasoning":true,"reasoning_modes":["low","medium","high","xhigh","max"],"reasoning_default":"medium","vision":true,"document_input":true,"computer_use":true,"web_search":true,"code_execution":true},"performance":{"benchmarks":{"terminal_bench":66.4,"cursor_bench":57.8,"osworld":81.8}},"api":{"endpoint":"https://api.anthropic.com/v1/messages","protocol":"anthropic","openai_compatible":false,"auth":{"type":"header","header_name":"x-api-key","env_var":"ANTHROPIC_API_KEY"},"sdks":{"anthropic":{"package":"@anthropic-ai/sdk","model_id":"claude-opus-5-5"},"vercel-ai":{"package":"@ai-sdk/anthropic","model_id":"claude-opus-5-5","factory":"anthropic('claude-opus-5-5')"}},"supported_params":["max_tokens","stop_sequences","system","tools","tool_choice","stream","metadata","thinking","service_tier"],"unsupported_params":["temperature","top_p","top_k"],"notes":"Adaptive thinking always on; thinking:{type:\"disabled\"} or a manual budget returns a 400 error - use effort instead. tool_choice of type \"any\" or \"tool\" (forced tool use) returns a 400 error - only \"auto\" (default) and \"none\" are supported. Thinking blocks are bound to the producing model and to the conversation prefix (system/tools/earlier messages); mismatches error by default for accounts created on/after 2026-08-31 unless the thinking-binding-controls-2026-08-01 beta header is used. On the Claude API and Google Cloud, the computer_20251124 tool is rejected - use the computer_toolset_20260801 toolset instead (Amazon Bedrock still accepts computer_20251124). Assistant-turn prefill returns a 400 error. Temperature, top_p, and top_k return a 400 error when set to non-default values. Full 1M context at standard pricing."},"aggregator_ids":{"openrouter":"anthropic/claude-opus-5.5","aws-bedrock":"anthropic.claude-opus-5-5","gcp-vertex":"claude-opus-5-5"},"manual_tags":["best_for_coding","smartest"],"sources":{"spec":"https://platform.claude.com/docs/en/models/opus-5-5/overview","pricing":"https://platform.claude.com/docs/en/about-claude/pricing","release_notes":"https://www.anthropic.com/claude-opus-5-5","docs":"https://platform.claude.com/docs/en/models/opus-5-5/whats-new-opus-5-5"},"last_verified":"2026-09-28"},{"id":"claude-opus-5","provider":"anthropic","family":"claude-5","display_name":"Claude Opus 5","description":"Legacy model as of Sept 22, 2026 - superseded by Claude Opus 5.5 as Anthropic's recommended default (still fully available on the Claude API, AWS, Google Cloud, and Microsoft Foundry; no deprecation notice, retirement floor unchanged at 'not sooner than July 24, 2027'). For complex agentic coding and enterprise work. Step-change over Claude Opus 4.8, with the largest gains in deep reasoning, agentic/long-horizon tasks, and test-time compute scaling. 1M context window (default and maximum, no smaller variant), 128k max output. Adaptive thinking is on by default (breaking change vs. Opus 4.8, where omitting thinking meant no thinking); thinking:{type:\"disabled\"} is accepted only at effort high or below (400 at xhigh/max). Same price as Opus 4.8. Temperature/top_p/top_k and assistant-turn prefill are not supported. Web fetch tool is not available on this model (an exception to Opus feature parity). Training data cutoff May 2026. As of Sept 17, 2026, approved life-science orgs under Anthropic's Life Sciences Verification Program (beta) Standard Use tier get more permissive biology safeguards on this model.","type":"chat","input_modalities":["text","image"],"output_modalities":["text"],"status":"ga","released":"2026-07-24","limits":{"context_window":1000000,"max_output":128000},"pricing":{"currency":"USD","input_per_mtok":5,"cached_input_per_mtok":0.5,"cache_write_5m_per_mtok":6.25,"cache_write_1h_per_mtok":10,"output_per_mtok":25,"batch_input_per_mtok":2.5,"batch_output_per_mtok":12.5,"notes":"Fast mode (research preview, Claude API only, incl. Claude Managed Agents): $10 input / $50 output per MTok; not available on Amazon Bedrock, Google Cloud, or Microsoft Foundry, and not available with the Batch API or Priority Tier. Prompt-cache minimum is 512 tokens (down from 1024 on Opus 4.8)."},"capabilities":{"tools":true,"parallel_tools":true,"structured_output":true,"streaming":true,"batching":true,"prompt_caching":true,"prefill":false,"system_prompt":true,"stop_sequences":true,"reasoning":true,"reasoning_modes":["low","medium","high","xhigh","max"],"reasoning_default":"high","vision":true,"document_input":true,"computer_use":true,"web_search":true,"code_execution":true},"api":{"endpoint":"https://api.anthropic.com/v1/messages","protocol":"anthropic","openai_compatible":false,"auth":{"type":"header","header_name":"x-api-key","env_var":"ANTHROPIC_API_KEY"},"sdks":{"anthropic":{"package":"@anthropic-ai/sdk","model_id":"claude-opus-5"},"vercel-ai":{"package":"@ai-sdk/anthropic","model_id":"claude-opus-5","factory":"anthropic('claude-opus-5')"}},"supported_params":["max_tokens","stop_sequences","system","tools","tool_choice","stream","metadata","thinking","service_tier"],"unsupported_params":["temperature","top_p","top_k"],"notes":"Adaptive thinking on by default (unlike Opus 4.8, where omitting thinking ran without it); thinking:{type:\"disabled\"} is accepted only at effort high or below, returning 400 at xhigh/max. Temperature, top_p, and top_k return a 400 error when set to non-default values. Assistant-turn prefill returns a 400 error — use output_config.format or system prompt instructions instead. Web fetch tool is not available on this model. Full 1M context (default and maximum) at standard pricing."},"aggregator_ids":{"openrouter":"anthropic/claude-opus-5","aws-bedrock":"anthropic.claude-opus-5","gcp-vertex":"claude-opus-5"},"manual_tags":["best_for_coding"],"sources":{"spec":"https://platform.claude.com/docs/en/about-claude/models/whats-new-opus-5","pricing":"https://platform.claude.com/docs/en/about-claude/pricing","release_notes":"https://www.anthropic.com/news/claude-opus-5","docs":"https://www.anthropic.com/news/life-sciences-verification-program"},"last_verified":"2026-09-28"},{"id":"claude-sonnet-5","provider":"anthropic","family":"claude-5","display_name":"Claude Sonnet 5","description":"Default model for Free and Pro plans; available on Max, Team, and Enterprise. Announced June 30, 2026. Introductory pricing of $2/$10 per MTok was made permanent on Aug 10, 2026 — the previously planned Sep 1, 2026 increase to $3/$15 will not occur, so Sonnet 5 is now permanently cheaper than Sonnet 4.6. Adaptive thinking on by default (no manual extended-thinking budget_tokens — returns a 400 error). 1M context, 128k max output (300k on Batch API with output-300k-2026-03-24 beta header). Training data cutoff January 2026. First Sonnet-tier model with real-time cybersecurity safeguards (refusals return HTTP 200, stop_reason: \"refusal\"). Temperature/top_p/top_k not supported.","type":"chat","input_modalities":["text","image"],"output_modalities":["text"],"status":"ga","released":"2026-06-30","limits":{"context_window":1000000,"max_output":128000},"pricing":{"currency":"USD","input_per_mtok":2,"cached_input_per_mtok":0.2,"cache_write_5m_per_mtok":2.5,"cache_write_1h_per_mtok":4,"output_per_mtok":10,"batch_input_per_mtok":1,"batch_output_per_mtok":5,"notes":"Introductory $2/$10 per-MTok pricing was made permanent on Aug 10, 2026 — the previously announced increase to $3/$15 on Sep 1, 2026 will not occur. $2/$10 is now the standard (non-introductory) rate; no future increase is scheduled."},"capabilities":{"tools":true,"parallel_tools":true,"structured_output":true,"streaming":true,"batching":true,"prompt_caching":true,"prefill":false,"system_prompt":true,"stop_sequences":true,"reasoning":true,"reasoning_modes":["low","medium","high","xhigh","max"],"reasoning_default":"high","vision":true,"document_input":true,"computer_use":true,"web_search":true,"code_execution":true},"api":{"endpoint":"https://api.anthropic.com/v1/messages","protocol":"anthropic","openai_compatible":false,"auth":{"type":"header","header_name":"x-api-key","env_var":"ANTHROPIC_API_KEY"},"sdks":{"anthropic":{"package":"@anthropic-ai/sdk","model_id":"claude-sonnet-5"},"vercel-ai":{"package":"@ai-sdk/anthropic","model_id":"claude-sonnet-5","factory":"anthropic('claude-sonnet-5')"}},"supported_params":["max_tokens","stop_sequences","system","tools","tool_choice","stream","metadata","thinking","service_tier"],"unsupported_params":["temperature","top_p","top_k"],"notes":"Adaptive thinking on by default; manual extended-thinking budget_tokens removed (returns 400). Effort parameter defaults to high on the API and Claude Code. Not available with Priority Tier. Supports zero data retention for ZDR orgs."},"aggregator_ids":{"openrouter":"anthropic/claude-sonnet-5","aws-bedrock":"anthropic.claude-sonnet-5","gcp-vertex":"claude-sonnet-5"},"manual_tags":["best_for_coding"],"sources":{"spec":"https://platform.claude.com/docs/en/about-claude/models/whats-new-sonnet-5","pricing":"https://platform.claude.com/docs/en/about-claude/pricing","release_notes":"https://www.anthropic.com/news/claude-sonnet-5","docs":"https://www.anthropic.com/news/life-sciences-verification-program"},"last_verified":"2026-09-28"},{"id":"claude-mythos-5","provider":"anthropic","family":"claude-mythos","display_name":"Claude Mythos 5","description":"Legacy model as of Sept 1, 2026 - superseded by Claude Mythos 5.1 (still fully available, no deprecation notice). Limited access via Project Glasswing (invitation only). Released June 9, 2026; suspended worldwide June 12 - June 26, 2026 under a US export-control directive. Restored June 26-30, 2026, but only to a specific set of US organizations approved by federal authorities under Project Glasswing — not broad GA. Expansion to more domestic/overseas partners is an ongoing discussion with the US government. Same token pricing as Claude Fable 5. Training data cutoff January 2026.","type":"chat","input_modalities":["text","image"],"output_modalities":["text"],"status":"preview","released":"2026-06-09","limits":{"context_window":1000000,"max_output":128000},"pricing":{"currency":"USD","input_per_mtok":10,"cached_input_per_mtok":1,"cache_write_5m_per_mtok":12.5,"cache_write_1h_per_mtok":20,"output_per_mtok":50,"batch_input_per_mtok":5,"batch_output_per_mtok":25,"notes":"Invitation-only access via Project Glasswing. Same token pricing as Claude Fable 5."},"capabilities":{"tools":true,"parallel_tools":true,"structured_output":true,"streaming":true,"batching":true,"prompt_caching":true,"prefill":false,"system_prompt":true,"stop_sequences":true,"reasoning":true,"reasoning_modes":["low","medium","high","xhigh","max"],"reasoning_default":"high","vision":true,"document_input":true,"computer_use":true,"web_search":true,"code_execution":true},"performance":{"intelligence_index":92},"api":{"endpoint":"https://api.anthropic.com/v1/messages","protocol":"anthropic","openai_compatible":false,"auth":{"type":"header","header_name":"x-api-key","env_var":"ANTHROPIC_API_KEY"},"sdks":{"anthropic":{"package":"@anthropic-ai/sdk","model_id":"claude-mythos-5"}},"unsupported_params":["temperature","top_p","top_k"],"notes":"Limited availability through Project Glasswing (invitation only). See https://anthropic.com/glasswing for access."},"aggregator_ids":{"aws-bedrock":"anthropic.claude-mythos-5","gcp-vertex":"claude-mythos-5"},"sources":{"spec":"https://platform.claude.com/docs/en/models/mythos-5/overview","pricing":"https://platform.claude.com/docs/en/about-claude/pricing","release_notes":"https://www.anthropic.com/news/redeploying-fable-5"},"last_verified":"2026-09-28"},{"id":"claude-opus-4-8","provider":"anthropic","family":"claude-4","display_name":"Claude Opus 4.8","description":"Anthropic's most capable Opus-tier model for complex reasoning and long-horizon agentic coding. 1M context, 128k max output, adaptive thinking (defaults to high effort). No extended thinking. Temperature/top_p/top_k not supported. Fast mode available at $10/$50 per MTok. Training data cutoff January 2026.","type":"chat","input_modalities":["text","image"],"output_modalities":["text"],"status":"ga","released":"2026-05-28","limits":{"context_window":1000000,"max_output":128000},"pricing":{"currency":"USD","input_per_mtok":5,"cached_input_per_mtok":0.5,"cache_write_5m_per_mtok":6.25,"cache_write_1h_per_mtok":10,"output_per_mtok":25,"batch_input_per_mtok":2.5,"batch_output_per_mtok":12.5,"notes":"Fast mode (research preview): $10 input / $50 output per MTok. Batch API not available in fast mode."},"capabilities":{"tools":true,"parallel_tools":true,"structured_output":true,"streaming":true,"batching":true,"prompt_caching":true,"prefill":false,"system_prompt":true,"stop_sequences":true,"reasoning":true,"reasoning_modes":["low","medium","high","xhigh","max"],"reasoning_default":"high","vision":true,"document_input":true,"computer_use":true,"web_search":true,"code_execution":true},"performance":{"intelligence_index":85},"api":{"endpoint":"https://api.anthropic.com/v1/messages","protocol":"anthropic","openai_compatible":false,"auth":{"type":"header","header_name":"x-api-key","env_var":"ANTHROPIC_API_KEY"},"sdks":{"anthropic":{"package":"@anthropic-ai/sdk","model_id":"claude-opus-4-8"},"vercel-ai":{"package":"@ai-sdk/anthropic","model_id":"claude-opus-4-8","factory":"anthropic('claude-opus-4-8')"}},"supported_params":["max_tokens","stop_sequences","system","tools","tool_choice","stream","metadata","thinking","service_tier"],"unsupported_params":["temperature","top_p","top_k"],"notes":"Adaptive thinking only; effort defaults to 'high'. Temperature, top_p, and top_k return a 400 error when set to non-default values. New tokenizer (same as Opus 4.7+). Full 1M context at standard pricing."},"aggregator_ids":{"openrouter":"anthropic/claude-opus-4.8","aws-bedrock":"anthropic.claude-opus-4-8","gcp-vertex":"claude-opus-4-8"},"manual_tags":["best_for_coding"],"sources":{"spec":"https://platform.claude.com/docs/en/about-claude/models/all-models","pricing":"https://platform.claude.com/docs/en/about-claude/pricing","release_notes":"https://www.anthropic.com/news/claude-opus-4-8"},"last_verified":"2026-09-28"},{"id":"claude-opus-4-7","provider":"anthropic","family":"claude-4","display_name":"Claude Opus 4.7","description":"Legacy Opus model; use Claude Opus 4.8 for new projects. 1M context, 128k max output, adaptive thinking. New tokenizer (up to 35% more tokens vs. pre-4.7 models). Temperature/top_p/top_k not supported. Training data cutoff January 2026.","type":"chat","input_modalities":["text","image"],"output_modalities":["text"],"status":"ga","released":"2026-04-16","limits":{"context_window":1000000,"max_output":128000},"pricing":{"currency":"USD","input_per_mtok":5,"cached_input_per_mtok":0.5,"cache_write_5m_per_mtok":6.25,"cache_write_1h_per_mtok":10,"output_per_mtok":25,"batch_input_per_mtok":2.5,"batch_output_per_mtok":12.5,"notes":"Fast mode (research preview) has been removed from this model as of 2026-07-27 confirmation: requests with speed=\"fast\" now return an error (unlike Opus 4.6, which silently falls back to standard speed/pricing). The model itself remains fully available at standard speed. Migrate to claude-opus-5 or claude-opus-4-8 to continue using fast mode."},"capabilities":{"tools":true,"parallel_tools":true,"structured_output":true,"streaming":true,"batching":true,"prompt_caching":true,"prefill":false,"system_prompt":true,"stop_sequences":true,"reasoning":true,"reasoning_modes":["low","medium","high","xhigh","max"],"reasoning_default":"medium","vision":true,"document_input":true,"computer_use":true,"web_search":true,"code_execution":true},"performance":{"intelligence_index":82,"benchmarks":{"swe_bench_verified":87.6,"swe_bench_pro":64.3,"cursor_bench":70}},"api":{"endpoint":"https://api.anthropic.com/v1/messages","protocol":"anthropic","openai_compatible":false,"auth":{"type":"header","header_name":"x-api-key","env_var":"ANTHROPIC_API_KEY"},"sdks":{"anthropic":{"package":"@anthropic-ai/sdk","model_id":"claude-opus-4-7"},"vercel-ai":{"package":"@ai-sdk/anthropic","model_id":"claude-opus-4-7","factory":"anthropic('claude-opus-4-7')"}},"supported_params":["max_tokens","stop_sequences","system","tools","tool_choice","stream","metadata","thinking","service_tier"],"unsupported_params":["temperature","top_p","top_k"],"notes":"Adaptive thinking only (no extended thinking switch). New tokenizer that may use up to ~35% more tokens for the same text. Full 1M context at standard pricing."},"aggregator_ids":{"openrouter":"anthropic/claude-opus-4.7","aws-bedrock":"anthropic.claude-opus-4-7","gcp-vertex":"claude-opus-4-7"},"sources":{"spec":"https://platform.claude.com/docs/en/about-claude/models/all-models","pricing":"https://platform.claude.com/docs/en/about-claude/pricing","docs":"https://platform.claude.com/docs/en/build-with-claude/fast-mode"},"last_verified":"2026-09-28"},{"id":"claude-opus-4-6","provider":"anthropic","family":"claude-4","display_name":"Claude Opus 4.6","description":"Legacy Opus model. 1M context, 128k max output, extended thinking and adaptive thinking, prompt caching. Fast Mode (research preview) was removed June 29, 2026; speed=\"fast\" requests now run at standard pricing. Training data cutoff August 2025 (reliable knowledge cutoff May 2025).","type":"chat","input_modalities":["text","image"],"output_modalities":["text"],"status":"ga","released":"2026-02-05","limits":{"context_window":1000000,"max_output":128000},"pricing":{"currency":"USD","input_per_mtok":5,"cached_input_per_mtok":0.5,"cache_write_5m_per_mtok":6.25,"cache_write_1h_per_mtok":10,"output_per_mtok":25,"batch_input_per_mtok":2.5,"batch_output_per_mtok":12.5,"notes":"Fast Mode was removed June 29, 2026 — speed=\"fast\" requests now silently run at standard speed/pricing (no error)."},"capabilities":{"tools":true,"parallel_tools":true,"structured_output":true,"streaming":true,"batching":true,"prompt_caching":true,"prefill":false,"system_prompt":true,"stop_sequences":true,"reasoning":true,"reasoning_modes":["off","low","medium","high","max"],"vision":true,"document_input":true,"computer_use":true,"web_search":true,"code_execution":true},"performance":{"intelligence_index":78,"benchmarks":{"swe_bench_verified":80.8}},"api":{"endpoint":"https://api.anthropic.com/v1/messages","protocol":"anthropic","openai_compatible":false,"auth":{"type":"header","header_name":"x-api-key","env_var":"ANTHROPIC_API_KEY"},"sdks":{"anthropic":{"package":"@anthropic-ai/sdk","model_id":"claude-opus-4-6"},"vercel-ai":{"package":"@ai-sdk/anthropic","model_id":"claude-opus-4-6","factory":"anthropic('claude-opus-4-6')"}},"supported_params":["max_tokens","temperature","top_p","top_k","stop_sequences","system","tools","tool_choice","stream","metadata","thinking"]},"aggregator_ids":{"openrouter":"anthropic/claude-opus-4.6","aws-bedrock":"anthropic.claude-opus-4-6-v1","gcp-vertex":"claude-opus-4-6"},"sources":{"spec":"https://platform.claude.com/docs/en/about-claude/models/all-models","pricing":"https://platform.claude.com/docs/en/about-claude/pricing","docs":"https://platform.claude.com/docs/en/models/opus-4-6/overview"},"last_verified":"2026-09-28"},{"id":"claude-opus-4-5-20251101","provider":"anthropic","family":"claude-4","display_name":"Claude Opus 4.5","description":"Legacy Opus model. 200k context, 64k max output, extended thinking, prompt caching. Active but superseded by Opus 4.8. Training data cutoff August 2025 (reliable knowledge cutoff May 2025).","type":"chat","input_modalities":["text","image"],"output_modalities":["text"],"status":"ga","released":"2025-11-01","limits":{"context_window":200000,"max_output":64000},"pricing":{"currency":"USD","input_per_mtok":5,"cached_input_per_mtok":0.5,"cache_write_5m_per_mtok":6.25,"cache_write_1h_per_mtok":10,"output_per_mtok":25,"batch_input_per_mtok":2.5,"batch_output_per_mtok":12.5},"capabilities":{"tools":true,"parallel_tools":true,"structured_output":true,"streaming":true,"batching":true,"prompt_caching":true,"prefill":true,"system_prompt":true,"stop_sequences":true,"reasoning":true,"reasoning_modes":["off","low","medium","high"],"vision":true,"computer_use":true},"performance":{"intelligence_index":75},"api":{"endpoint":"https://api.anthropic.com/v1/messages","protocol":"anthropic","openai_compatible":false,"auth":{"type":"header","header_name":"x-api-key","env_var":"ANTHROPIC_API_KEY"},"sdks":{"anthropic":{"package":"@anthropic-ai/sdk","model_id":"claude-opus-4-5"}}},"aggregator_ids":{"aws-bedrock":"anthropic.claude-opus-4-5-20251101-v1:0","gcp-vertex":"claude-opus-4-5@20251101"},"sources":{"spec":"https://platform.claude.com/docs/en/about-claude/models/all-models","docs":"https://platform.claude.com/docs/en/models/opus-4-5/overview"},"last_verified":"2026-09-28"},{"id":"claude-sonnet-4-6","provider":"anthropic","family":"claude-4","display_name":"Claude Sonnet 4.6","description":"Best price-to-performance Claude model. 1M context window, 128k max output, extended thinking and adaptive thinking, prompt caching. 79.6% SWE-bench Verified, 72.5% OSWorld. Training data cutoff January 2026 (reliable knowledge cutoff August 2025).","type":"chat","input_modalities":["text","image"],"output_modalities":["text"],"status":"ga","released":"2026-02-17","limits":{"context_window":1000000,"max_output":128000},"pricing":{"currency":"USD","input_per_mtok":3,"cached_input_per_mtok":0.3,"cache_write_5m_per_mtok":3.75,"cache_write_1h_per_mtok":6,"output_per_mtok":15,"batch_input_per_mtok":1.5,"batch_output_per_mtok":7.5},"capabilities":{"tools":true,"parallel_tools":true,"structured_output":true,"streaming":true,"batching":true,"prompt_caching":true,"prefill":false,"system_prompt":true,"stop_sequences":true,"reasoning":true,"reasoning_modes":["off","low","medium","high"],"vision":true,"document_input":true,"computer_use":true,"web_search":true,"code_execution":true},"performance":{"intelligence_index":73,"benchmarks":{"swe_bench_verified":79.6,"osworld":72.5}},"api":{"endpoint":"https://api.anthropic.com/v1/messages","protocol":"anthropic","openai_compatible":false,"auth":{"type":"header","header_name":"x-api-key","env_var":"ANTHROPIC_API_KEY"},"sdks":{"anthropic":{"package":"@anthropic-ai/sdk","model_id":"claude-sonnet-4-6"},"vercel-ai":{"package":"@ai-sdk/anthropic","model_id":"claude-sonnet-4-6","factory":"anthropic('claude-sonnet-4-6')"}},"supported_params":["max_tokens","temperature","top_p","top_k","stop_sequences","system","tools","tool_choice","stream","metadata","thinking"]},"aggregator_ids":{"openrouter":"anthropic/claude-sonnet-4.6","aws-bedrock":"anthropic.claude-sonnet-4-6","gcp-vertex":"claude-sonnet-4-6"},"sources":{"spec":"https://platform.claude.com/docs/en/about-claude/models/all-models","pricing":"https://platform.claude.com/docs/en/about-claude/pricing","docs":"https://platform.claude.com/docs/en/models/sonnet-4-6/overview"},"last_verified":"2026-09-28"},{"id":"claude-sonnet-4-5-20250929","provider":"anthropic","family":"claude-4","display_name":"Claude Sonnet 4.5","description":"Legacy Sonnet model. 200k context, 64k max output, extended thinking, prompt caching. Active but superseded by Sonnet 4.6. Anthropic's model-deprecations page still lists 'Current state: Active' with 'Deprecated: N/A' and a tentative retirement floor of 'not sooner than September 29, 2026' - reconfirmed as still just a floor date, not an actual deprecation notice, as of Sept 28, 2026 (one day before that floor date; no deprecated_on/sunset_on published yet).","type":"chat","input_modalities":["text","image"],"output_modalities":["text"],"status":"ga","released":"2025-09-29","limits":{"context_window":200000,"max_output":64000},"pricing":{"currency":"USD","input_per_mtok":3,"cached_input_per_mtok":0.3,"cache_write_5m_per_mtok":3.75,"cache_write_1h_per_mtok":6,"output_per_mtok":15,"batch_input_per_mtok":1.5,"batch_output_per_mtok":7.5},"capabilities":{"tools":true,"parallel_tools":true,"structured_output":true,"streaming":true,"batching":true,"prompt_caching":true,"prefill":true,"system_prompt":true,"stop_sequences":true,"reasoning":true,"reasoning_modes":["off","low","medium","high"],"vision":true,"computer_use":true},"performance":{"intelligence_index":70},"api":{"endpoint":"https://api.anthropic.com/v1/messages","protocol":"anthropic","openai_compatible":false,"auth":{"type":"header","header_name":"x-api-key","env_var":"ANTHROPIC_API_KEY"},"sdks":{"anthropic":{"package":"@anthropic-ai/sdk","model_id":"claude-sonnet-4-5"}}},"aggregator_ids":{"aws-bedrock":"anthropic.claude-sonnet-4-5-20250929-v1:0","gcp-vertex":"claude-sonnet-4-5@20250929"},"sources":{"spec":"https://platform.claude.com/docs/en/about-claude/models/all-models","docs":"https://platform.claude.com/docs/en/about-claude/model-deprecations"},"last_verified":"2026-09-28"},{"id":"claude-haiku-4-5-20251001","provider":"anthropic","family":"claude-4","display_name":"Claude Haiku 4.5","description":"Fastest Claude with near-frontier intelligence. 200k context, 64k max output, extended thinking, prompt caching, vision. Training data cutoff July 2025 (reliable knowledge cutoff February 2025).","type":"chat","input_modalities":["text","image"],"output_modalities":["text"],"status":"ga","released":"2025-10-01","limits":{"context_window":200000,"max_output":64000},"pricing":{"currency":"USD","input_per_mtok":1,"cached_input_per_mtok":0.1,"cache_write_5m_per_mtok":1.25,"cache_write_1h_per_mtok":2,"output_per_mtok":5,"batch_input_per_mtok":0.5,"batch_output_per_mtok":2.5},"capabilities":{"tools":true,"parallel_tools":true,"structured_output":true,"streaming":true,"batching":true,"prompt_caching":true,"prefill":true,"system_prompt":true,"stop_sequences":true,"reasoning":true,"reasoning_modes":["off","low","medium","high"],"vision":true},"performance":{"speed_tps":130,"ttft_ms":400,"intelligence_index":60},"api":{"endpoint":"https://api.anthropic.com/v1/messages","protocol":"anthropic","openai_compatible":false,"auth":{"type":"header","header_name":"x-api-key","env_var":"ANTHROPIC_API_KEY"},"sdks":{"anthropic":{"package":"@anthropic-ai/sdk","model_id":"claude-haiku-4-5"},"vercel-ai":{"package":"@ai-sdk/anthropic","model_id":"claude-haiku-4-5","factory":"anthropic('claude-haiku-4-5')"}}},"aggregator_ids":{"openrouter":"anthropic/claude-haiku-4.5","aws-bedrock":"anthropic.claude-haiku-4-5-20251001-v1:0","gcp-vertex":"claude-haiku-4-5@20251001"},"sources":{"spec":"https://platform.claude.com/docs/en/about-claude/models/all-models","pricing":"https://platform.claude.com/docs/en/about-claude/pricing","docs":"https://platform.claude.com/docs/en/models/haiku-4-5/overview"},"last_verified":"2026-09-28"},{"id":"deepseek-flash","provider":"deepseek","family":"deepseek-v4.1","display_name":"DeepSeek V4.1 Flash","description":"New flagship Flash model, GA 2026-09-10, replacing BOTH deepseek-v4-flash and deepseek-v4-flash-vision-exp in a single release (both officially retired the same day; their legacy ids are still accepted but temporarily route to this model and bill at its rate, per the official 2026-09-10 changelog entry and confirmed live on the Models & Pricing / Vision / Your First API Call / Change Log pages as of 2026-09-28, no further changes since the 2026-09-22 check). First model on DeepSeek's new 'Causal Encoder-Decoder' architecture: a 552B-parameter MoE with an asymmetric 8B active parameters for input / 16B active for output, per the official release notes and the linked tech report (huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash) — independently corroborated by OpenRouter's own model page and launch-volume announcement (1T tokens in 24h). Native multimodal (text+image) understanding trained in from pretraining, unlike the bolted-on vision-exp sibling it replaces. Same 1M context / 384k max output as the models it replaces. Thinking mode on by default (effort high); on/off toggle plus low/high/max effort control unchanged from the prior V4 family (confirmed via the official Thinking Mode guide, re-checked 2026-09-28). Vision handling changed materially from the retired vision-exp model: images now resize to a ~544x544-1300x1300 pixel-count band (was ~384x384-800x800) and are capped at 1024 tokens/image (was 384 tokens/image), still billed as ordinary input tokens; a new optional 'detail' parameter (low/high/original/auto) lets callers force a 512x512 downscale for cheaper/faster calls. Other vision limits are unchanged from the retired model's guide, re-verified 2026-09-28: 8192px max image dimension (4096px once a request holds 15+ images), 32 MiB per inline (base64/external-URL) image vs 64 MiB per Files-API image, 600 images/request, 64 MiB total request images (200 MiB including Files-API references), 8192-character external URL cap, 48 MiB total request body. Pricing was cut with this release (DeepSeek's own framing: 'API prices have been reduced accordingly') — see pricing.notes for the off-peak/peak schedule, unchanged in hours/multiplier since 2026-09-14 (peak 01:00-04:00 and 06:00-10:00 UTC weekdays only, off-peak = half of peak) the footnote wording (which by 2026-09-22 had grown an 'excluding Chinese public holidays' carve-out) remains unchanged as of 2026-09-28 — see pricing.notes. New pricing took effect 04:00 UTC 2026-09-10. Benchmark suite below is a full replacement set from the 2026-09-10 release notes; DeepSeek also reports an alternate HLE figure of 39.1 explicitly footnoted as 'tested only on the pure-text subset of the HLE benchmark set' (not recorded as a separate field here to avoid inventing a naming convention for a partial-subset score — the hle_no_tools value below, 36.8, is the general/primary figure DeepSeek reports). DeepSeek Harness (an agent-tooling layer) remained in developer preview as of 2026-09-22 (latest public preview build 0.1.2-rc.1, no GA date announced) — not re-verified on 2026-09-28 (no changelog signal either way), and official partners WorkBuddy/CodeBuddy and OpenCode support this model out of the box, per the release notes. No new DeepSeek model has shipped since this 2026-09-10 release as of 2026-09-28 (re-checked official Change Log at api-docs.deepseek.com/updates).","type":"chat","input_modalities":["text","image"],"output_modalities":["text"],"status":"ga","released":"2026-09-10","limits":{"context_window":1000000,"max_output":384000,"max_image_size_mb":32,"max_images_per_request":600},"pricing":{"currency":"USD","input_per_mtok":0.15,"cached_input_per_mtok":0.003,"output_per_mtok":0.6,"notes":"Off-peak (default/shown) rates apply all hours except 01:00-04:00 and 06:00-10:00 UTC Monday-Friday, EXCLUDING Chinese public holidays (weekends AND Chinese public holidays are fully off-peak) — same schedule as the retired deepseek-v4-flash and as deepseek-v4-pro. During weekday peak hours, rates are exactly double: input (cache miss) $0.30/Mtok, input (cache hit) $0.006/Mtok, output $1.20/Mtok. New pricing (a cut from the retired deepseek-v4-flash's $0.22/$0.007/$0.66 off-peak rates) took effect 04:00 UTC 2026-09-10 alongside the DeepSeek-V4.1-Flash release. Footnote wording changed again between 2026-09-14 and 2026-09-22: the pricing page now explicitly adds an 'excluding Chinese public holidays' carve-out to the peak-hour definition (previously only weekends were called out as fully off-peak) — the hours (01:00-04:00, 06:00-10:00 UTC) and 2x/0.5x multiplier are unchanged. Re-confirmed unchanged via a fresh fetch of the pricing page on 2026-09-28 (same rates, same footnote wording as 2026-09-22)."},"modalities":{"input":{"text":{"pricing_per_mtok":0.15,"pricing_cached_per_mtok":0.003},"image":{"max_size_mb":32,"max_count_per_request":600,"supported_formats":["JPEG","PNG","GIF","WebP"],"notes":"Images are auto-resized (upscaled toward ~544x544, downscaled toward ~1300x1300 total pixel count) then tokenized at up to 1024 tokens/image, billed as input tokens at the text rate; an optional 'detail' param (low/high/original/auto) can force a 512x512 downscale for cheaper calls. Images only supported in user messages. Per the official Vision guide (re-checked 2026-09-28, unchanged since 2026-09-22): max image dimension is 8192px per side, dropping to 4096px per side once a request contains 15+ images; max image size is 32 MiB inline (base64/external URL) vs 64 MiB via the Files API file_id; max total image size per request is 64 MiB without file_id images or up to 200 MiB including them; external image URLs are capped at 8192 characters; total request body is capped at 48 MiB."}},"output":{"text":{"max_tokens":384000,"pricing_per_mtok":0.6}}},"capabilities":{"tools":true,"json_mode":true,"structured_output":true,"streaming":true,"prompt_caching":true,"prefill":true,"system_prompt":true,"reasoning":true,"reasoning_modes":["off","low","high","max"],"reasoning_default":"high","vision":true},"performance":{"benchmarks":{"gpqa_diamond":90.9,"hle_no_tools":36.8,"hle_with_tools":63.9,"codeforces_rating":3471,"matharena_apex":65.6,"terminal_bench_2_1":90.6,"terminal_bench_3_0":30,"terminal_bench_4_0":31.2,"deepswe_v1_1":74.2,"programbench":20.3,"nl2repo":65.4,"cybergym":88.1,"sec_bench_pro":62.8,"exploitgym":15.3,"automation_bench_public":54.8,"agent_last_exam":31.8,"chartography":78.9,"babyvision":89.6,"zerobench_main":49}},"api":{"endpoint":"https://api.deepseek.com/chat/completions","protocol":"openai_compat","openai_compatible":true,"auth":{"type":"bearer","env_var":"DEEPSEEK_API_KEY"},"sdks":{"openai":{"package":"openai","model_id":"deepseek-flash"},"anthropic":{"package":"@anthropic-ai/sdk","model_id":"deepseek-flash"}},"supported_params":["max_tokens","temperature","top_p","stop","tools","tool_choice","response_format","stream","reasoning","reasoning_effort"],"notes":"OpenAI-compatible at https://api.deepseek.com; Anthropic-compatible endpoint at https://api.deepseek.com/anthropic; also natively supports the OpenAI Responses API format. Legacy model names deepseek-v4-flash and deepseek-v4-flash-vision-exp are still accepted but both retired 2026-09-10 — requests are served by this model and billed at its rate. temperature/presence_penalty/frequency_penalty are accepted but have no effect in thinking mode; top_p is honored in thinking mode but floored at 0.95, while in non-thinking mode it's fixed at 1.0 and ignored. Images accepted via base64 data URL, external URL (max 8192 chars, 60s download timeout), or Files API file_id; new optional 'detail' param (low/high/original/auto). DeepSeek Harness is in developer preview for agent-tool builders (still developer preview as of 2026-09-22, no GA); official partners WorkBuddy/CodeBuddy and OpenCode support this model natively. Concurrency limit per the rate-limit docs (re-checked 2026-09-28, unchanged since 2026-09-22): 2500 concurrent requests."},"aggregator_ids":{"openrouter":"deepseek/deepseek-v4.1-flash"},"manual_tags":["cheapest"],"sources":{"spec":"https://api-docs.deepseek.com/quick_start/pricing","release_notes":"https://api-docs.deepseek.com/news/news260910","docs":"https://api-docs.deepseek.com/guides/vision"},"last_verified":"2026-09-28"},{"id":"deepseek-v4-pro","provider":"deepseek","family":"deepseek-v4","display_name":"DeepSeek V4 Pro","description":"Pro V4 model with strong reasoning. 1M context, 384k max output. Full reasoning model with thinking and non-thinking modes. Reached GA on 2026-08-13 as the DeepSeek-V4-Pro-0813 build, with major agent-capability gains over V4-Pro-Preview (e.g. Terminal Bench 2.1 72.1->87.9, DeepSWE 12.8->62.7, per the official changelog). Natively supports the OpenAI Responses API format (adapted for Codex) since the 2026-08-13 GA release. Thinking mode offers three effort levels (low/high/max, default high) alongside the on/off toggle, identical to deepseek-v4-flash. Peak/off-peak pricing (see pricing.notes) took effect 16:00 UTC 2026-08-16 (Mon-Fri peak hours 01:00-04:00 and 06:00-10:00 UTC, off-peak = half of peak; footnote gained an 'excluding Chinese public holidays' carve-out by 2026-09-22 — see pricing.notes). NEAR-RETIREMENT, REVERSED (2026-09-10/2026-09-14): the 2026-09-10 DeepSeek-V4.1-Flash release announcement (news260910) originally stated \"We're phasing out V4-Pro... Starting at 04:00 UTC on Sept 14, 2026, all deepseek-v4-pro requests will route to V4.1-Flash at V4.1-Flash rates. This will continue until V4.1-Pro launches\" — but the live Change Log entry for that same date and the current Models & Pricing / Your First API Call pages read instead: \"In response to user demand, we have decided to continue providing API services for DeepSeek V4 Pro after September 14, 2026, with the billing method remaining unchanged. We will provide further notice should there be any changes.\" Treating the living docs pages (updated after the original announcement) as authoritative over the static news post: deepseek-v4-pro remains GA, unchanged pricing/limits, still the V4-Pro-0813 build — NOT retired and NOT rerouted; re-confirmed live and unchanged via a fresh fetch on 2026-09-28 (Models & Pricing page still lists model version 'DeepSeek-V4-Pro-0813', no successor build; unchanged since the 2026-09-22 check). STILL NO V4.1-PRO as of 2026-09-28: no model card, id, pricing, or release date on any official DeepSeek page (api-docs.deepseek.com Change Log's newest entry is still 2026-09-10, re-checked at /updates); the announcement's own phrasing (\"until V4.1-Pro launches\") continues to signal an eventual successor is planned but third-party trackers (e.g. orcarouter.ai) report only unofficial/uncorroborated speculation about a delay reason — not applied here. Separately, third-party press (not an official DeepSeek page) reported a closed-door DeepSeek investor meeting on 2026-09-21 where CEO Liang Wenfeng described a future ~20T-parameter model in training (with an even larger ~80T-parameter model floated beyond that) — no id, spec, pricing, or timeline, and nothing on any official api-docs.deepseek.com or platform.deepseek.com page; not added here, tracked only as a watch item. See the standing watch item in docs/ASSIGNED.md.","type":"reasoning","input_modalities":["text"],"output_modalities":["text"],"status":"ga","limits":{"context_window":1000000,"max_output":384000,"max_reasoning_tokens":32000},"pricing":{"currency":"USD","input_per_mtok":0.66,"cached_input_per_mtok":0.022,"output_per_mtok":1.98,"notes":"Off-peak (default/shown) rates apply all hours except 01:00-04:00 and 06:00-10:00 UTC Monday-Friday, EXCLUDING Chinese public holidays (weekends AND Chinese public holidays are fully off-peak). During those weekday peak hours (7 of 24h on weekdays), rates are exactly double: input (cache miss) $1.32/Mtok, input (cache hit) $0.044/Mtok, output $3.96/Mtok. New tiered pricing took effect 16:00 UTC 2026-08-16, replacing the flat $0.435/$0.003625/$0.87 rates in force before then. The pricing page footnote picked up the \"Monday through Friday\" qualifier (i.e. weekends fully off-peak) sometime between 2026-08-17 and 2026-08-24. Footnote wording changed again between 2026-09-14 and 2026-09-22: it now explicitly adds an 'excluding Chinese public holidays' carve-out to the peak-hour definition — the hours and 2x/0.5x multiplier are unchanged. Re-confirmed unchanged via a fresh fetch of the pricing page on 2026-09-28 (same rates, same footnote wording as 2026-09-22)."},"capabilities":{"tools":true,"json_mode":true,"structured_output":true,"streaming":true,"prompt_caching":true,"prefill":true,"system_prompt":true,"reasoning":true,"reasoning_modes":["off","low","high","max"],"reasoning_default":"high"},"performance":{"intelligence_index":78,"benchmarks":{"hle_no_tools":42.7,"hle_with_tools":60,"terminal_bench_2_1":87.9,"nl2repo":61.5,"cybergym":83.3,"deepswe":62.7,"toolathlon_verified":74.1,"agent_last_exam":25.7,"automation_bench_public":31.8,"dsbench_fullstack":71.1,"dsbench_hard":67.2}},"api":{"endpoint":"https://api.deepseek.com/chat/completions","protocol":"openai_compat","openai_compatible":true,"auth":{"type":"bearer","env_var":"DEEPSEEK_API_KEY"},"sdks":{"openai":{"package":"openai","model_id":"deepseek-v4-pro"},"anthropic":{"package":"@anthropic-ai/sdk","model_id":"deepseek-v4-pro"}},"supported_params":["max_tokens","temperature","top_p","stop","tools","tool_choice","response_format","stream","reasoning","reasoning_effort"],"notes":"OpenAI-compatible at https://api.deepseek.com; Anthropic-compatible endpoint at https://api.deepseek.com/anthropic. As of the 2026-08-13 GA release, also natively supports the OpenAI Responses API format (base_url https://api.deepseek.com, adapted for Codex). reasoning.effort accepts none/low/high/max on the Responses API (none disables thinking). temperature/top_p are accepted but have no effect in thinking mode. Concurrency limit per the rate-limit docs (re-checked 2026-09-28, unchanged since 2026-09-22): 500 concurrent requests. Re-confirmed live and unretired on 2026-09-28 (still model version DeepSeek-V4-Pro-0813) — see description for the reversed retirement scare and the still-unreleased V4.1-Pro."},"aggregator_ids":{"openrouter":"deepseek/deepseek-v4-pro"},"manual_tags":["best_value"],"sources":{"spec":"https://api-docs.deepseek.com/quick_start/pricing","release_notes":"https://api-docs.deepseek.com/news/news260813","docs":"https://api-docs.deepseek.com/news/news260910"},"last_verified":"2026-09-28"},{"id":"gemini-3.5-flash","provider":"google","family":"gemini-3.5","display_name":"Gemini 3.5 Flash","description":"Google's most intelligent widely-available model for sustained frontier performance on agentic and coding tasks. 1M context, multimodal (text/image/audio/video), thinking support. Stable release.","type":"chat","input_modalities":["text","image","audio","video"],"output_modalities":["text"],"status":"ga","limits":{"context_window":1048576,"max_output":65536,"max_video_seconds":7200},"pricing":{"currency":"USD","input_per_mtok":1.5,"cached_input_per_mtok":0.15,"output_per_mtok":9,"batch_input_per_mtok":0.75,"batch_output_per_mtok":4.5,"notes":"Priority tier: $2.70 input / $16.20 output per MTok. Flex/Batch at 50% of standard."},"capabilities":{"tools":true,"parallel_tools":true,"json_mode":true,"structured_output":true,"streaming":true,"batching":true,"prompt_caching":true,"system_prompt":true,"reasoning":true,"reasoning_modes":["off","low","medium","high"],"vision":true,"audio_input":true,"video_input":true,"document_input":true,"web_search":true,"code_execution":true},"performance":{"intelligence_index":86},"api":{"endpoint":"https://generativelanguage.googleapis.com/v1beta/models/gemini-3.5-flash:generateContent","protocol":"gemini","openai_compatible":true,"auth":{"type":"bearer","env_var":"GOOGLE_AI_API_KEY"},"sdks":{"google":{"package":"@google/genai","model_id":"gemini-3.5-flash"},"vercel-ai":{"package":"@ai-sdk/google","model_id":"gemini-3.5-flash","factory":"google('gemini-3.5-flash')"}},"supported_params":["max_output_tokens","temperature","top_p","top_k","stop_sequences","tools","response_mime_type","response_schema","thinking_config"]},"aggregator_ids":{"openrouter":"google/gemini-3.5-flash","gcp-vertex":"gemini-3.5-flash"},"manual_tags":["smartest","best_for_coding"],"sources":{"spec":"https://ai.google.dev/gemini-api/docs/models","pricing":"https://ai.google.dev/gemini-api/docs/pricing","release_notes":"https://ai.google.dev/gemini-api/docs/interactions/whats-new-gemini-3.5"},"last_verified":"2026-09-22"},{"id":"gemini-3.6-flash","provider":"google","family":"gemini-3.6","display_name":"Gemini 3.6 Flash","description":"Workhorse Gemini model, released July 21, 2026 alongside gemini-3.5-flash-lite. Positioned by Google as the recommended replacement for gemini-2.5-flash and gemini-2.0-flash. Reduces output token usage ~17% vs 3.5 Flash at a lower price, with built-in client-side computer-use tool. 1M context, multimodal (text/image/audio/video/PDF) input, thinking support. Superseded as the newest Flash model by gemini-3.7-flash (Aug 13, 2026), but remains GA with no shutdown date announced.","type":"chat","input_modalities":["text","image","audio","video"],"output_modalities":["text"],"status":"ga","released":"2026-07-21","limits":{"context_window":1048576,"max_output":65536},"pricing":{"currency":"USD","input_per_mtok":0.75,"cached_input_per_mtok":0.075,"output_per_mtok":3.75,"batch_input_per_mtok":0.375,"batch_output_per_mtok":1.875,"notes":"Corrected 2026-08-17: live pricing page (verified via two independent fetches) shows promotional launch pricing in effect through Dec 31, 2026 -- $0.75 input / $3.75 output / $0.075 cached (batch: $0.375/$1.875) -- which the registry had missed; it previously recorded only the post-promo standard rate ($1.50/$7.50, batch $0.75/$3.75) as if it were current. Standard rate of $1.50 input / $7.50 output / $0.15 cached (batch $0.75/$3.75/$0.075) resumes Jan 1, 2027. Priority tier: $1.35 input / $6.75 output through 12/31/26, then $2.70/$13.50 from 1/1/27. Context cache storage: $0.50/MTok/hr through 12/31/26, then $1.00/MTok/hr. Sampling params temperature/top_p/top_k are deprecated for this model; prefilling model turns (ending a request on a model turn) returns HTTP 400. Re-verified 2026-09-22 against the live pricing page: promo pricing unchanged."},"capabilities":{"tools":true,"structured_output":true,"streaming":true,"batching":true,"prompt_caching":true,"system_prompt":true,"reasoning":true,"vision":true,"audio_input":true,"video_input":true,"document_input":true,"computer_use":true,"web_search":true,"code_execution":true},"api":{"endpoint":"https://generativelanguage.googleapis.com/v1beta/models/gemini-3.6-flash:generateContent","protocol":"gemini","openai_compatible":true,"auth":{"type":"bearer","env_var":"GOOGLE_AI_API_KEY"},"sdks":{"google":{"package":"@google/genai","model_id":"gemini-3.6-flash"}},"notes":"temperature, top_p, and top_k sampling parameters are deprecated starting with this model. Prefilling is not supported."},"sources":{"spec":"https://ai.google.dev/gemini-api/docs/models/gemini-3.6-flash","pricing":"https://ai.google.dev/gemini-api/docs/pricing","release_notes":"https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-6-flash-3-5-flash-lite-3-5-flash-cyber/"},"last_verified":"2026-09-22"},{"id":"gemini-3.7-flash","provider":"google","family":"gemini-3.7","display_name":"Gemini 3.7 Flash","description":"Google's most intelligent workhorse Flash model, released August 13-14, 2026 -- just three weeks after gemini-3.6-flash -- via an in-place algorithmic/post-training update rather than a from-scratch retrain. Targets complex coding, web development, and agentic/multi-step tool-use workflows; Google reports strong gains over 3.6 Flash on coding benchmarks (DeepSWE v1.1 49.0%->65.3%, FrontierCode 1.1 Main 34.4%->43.6%). 1M context, multimodal (text/image/audio/video/PDF) input, thinking support (low/medium/high; minimal not supported), built-in computer-use tool (preview). Does not support audio/image generation or the Live API. Launched at promotional (roughly half-price) rates through end of 2026.","type":"chat","input_modalities":["text","image","audio","video"],"output_modalities":["text"],"status":"ga","released":"2026-08-13","limits":{"context_window":1048576,"max_output":65536},"pricing":{"currency":"USD","input_per_mtok":0.75,"cached_input_per_mtok":0.075,"output_per_mtok":3.75,"batch_input_per_mtok":0.375,"batch_output_per_mtok":1.875,"notes":"Promotional launch pricing through Dec 31, 2026: $0.75 input / $3.75 output / $0.075 cached (batch: $0.375/$1.875/$0.0375). Standard rate resumes Jan 1, 2027 at $1.50 input / $7.50 output / $0.15 cached (batch $0.75/$3.75/$0.075) -- prices are scheduled to double. Priority tier: $1.35 input / $6.75 output through 12/31/26, then $2.70/$13.50 from 1/1/27. Context cache storage: $0.50/MTok/hr through 12/31/26, then $1.00/MTok/hr. Free tier available (rate-limited). Re-verified 2026-09-22 against the live pricing page: promo pricing unchanged."},"capabilities":{"tools":true,"structured_output":true,"streaming":true,"batching":true,"prompt_caching":true,"system_prompt":true,"reasoning":true,"reasoning_modes":["low","medium","high"],"vision":true,"audio_input":true,"video_input":true,"document_input":true,"computer_use":true,"web_search":true,"code_execution":true},"api":{"endpoint":"https://generativelanguage.googleapis.com/v1beta/models/gemini-3.7-flash:generateContent","protocol":"gemini","openai_compatible":true,"auth":{"type":"bearer","env_var":"GOOGLE_AI_API_KEY"},"sdks":{"google":{"package":"@google/genai","model_id":"gemini-3.7-flash"}},"notes":"Supports caching, code execution, computer use (preview), file search, function calling, structured outputs, thinking, Google Maps/Search grounding, URL context, and Batch/Flex/Priority inference. Does not support audio generation, image generation, or Live API access."},"sources":{"spec":"https://ai.google.dev/gemini-api/docs/models/gemini-3.7-flash","pricing":"https://ai.google.dev/gemini-api/docs/pricing","release_notes":"https://blog.google/innovation-and-ai/models-and-research/gemini-models/introducing-gemini-3-7-flash/"},"last_verified":"2026-09-22"},{"id":"gemini-3.8-flash","provider":"google","family":"gemini-3.8","display_name":"Gemini 3.8 Flash","description":"Google's most intelligent Flash model, released September 2, 2026 -- three weeks after gemini-3.7-flash -- via an in-place algorithmic/post-training update. Engineered for long-horizon software engineering, autonomous agents, and complex enterprise/legal/finance-domain reasoning at Flash speed and cost. Google reports substantial gains over 3.7 Flash on DeepSWE v1.1, Vals Finance Agent V2, and Harvey's Legal Agent Benchmark; HLE-Verified 54.9%. Coexists with (does not deprecate) gemini-3.6-flash or gemini-3.7-flash, which Google says remain fully supported for efficiency-first workloads. Launched at the same introductory (roughly half-price) rate as 3.7 Flash through end of 2026. A gated, non-public 'Gemini 3.8 Flash Cyber' security-tuned variant also shipped the same day, accessible only via Google's Fairwind Program vetting scheme -- no public id or pricing.","type":"chat","input_modalities":["text","image","audio","video"],"output_modalities":["text"],"status":"ga","released":"2026-09-02","limits":{"context_window":1048576,"max_output":65536},"pricing":{"currency":"USD","input_per_mtok":0.75,"cached_input_per_mtok":0.075,"output_per_mtok":3.75,"batch_input_per_mtok":0.375,"batch_output_per_mtok":1.875,"notes":"Promotional launch pricing through Dec 31, 2026: $0.75 input / $3.75 output / $0.075 cached (batch: $0.375/$1.875). Standard rate resumes Jan 1, 2027 at $1.50 input / $7.50 output / $0.15 cached (batch $0.75/$3.75/$0.075) -- prices are scheduled to double, same schedule as gemini-3.7-flash. Free tier available (rate-limited). Re-verified 2026-09-22 against the live pricing page: promo pricing unchanged."},"capabilities":{"tools":true,"structured_output":true,"streaming":true,"batching":true,"prompt_caching":true,"system_prompt":true,"reasoning":true,"reasoning_modes":["low","medium","high"],"vision":true,"audio_input":true,"video_input":true,"document_input":true,"computer_use":true,"web_search":true,"code_execution":true},"api":{"endpoint":"https://generativelanguage.googleapis.com/v1beta/models/gemini-3.8-flash:generateContent","protocol":"gemini","openai_compatible":true,"auth":{"type":"bearer","env_var":"GOOGLE_AI_API_KEY"},"sdks":{"google":{"package":"@google/genai","model_id":"gemini-3.8-flash"}},"notes":"Supports caching, code execution, computer use (preview), file search, function calling, structured outputs, thinking, Google Maps/Search grounding, URL context, and Batch/Flex/Priority inference. Does not support audio generation, image generation, or Live API access."},"sources":{"spec":"https://ai.google.dev/gemini-api/docs/models/gemini-3.8-flash","pricing":"https://ai.google.dev/gemini-api/docs/pricing","release_notes":"https://blog.google/innovation-and-ai/models-and-research/gemini-models/3-8-flash-and-3-8-flash-cyber/"},"last_verified":"2026-09-22"},{"id":"gemini-3.5-flash-lite","provider":"google","family":"gemini-3.5","display_name":"Gemini 3.5 Flash-Lite","description":"Fastest, most cost-effective 3.5-class model, released July 21, 2026 alongside gemini-3.6-flash. Targets high-throughput, low-latency, high-volume workloads with configurable thinking levels; distinct from and not a direct replacement for gemini-3.1-flash-lite, which remains Google's official migration target for gemini-2.5-flash-lite. 1M context, multimodal (text/image/audio/video/PDF) input.","type":"chat","input_modalities":["text","image","audio","video"],"output_modalities":["text"],"status":"ga","released":"2026-07-21","limits":{"context_window":1048576,"max_output":65536},"pricing":{"currency":"USD","input_per_mtok":0.3,"cached_input_per_mtok":0.03,"output_per_mtok":2.5,"batch_input_per_mtok":0.15,"batch_output_per_mtok":1.25,"notes":"Added 2026-09-14: cached-input rate was missing from the registry; confirmed $0.03/MTok on the live pricing page (Batch: $0.02/MTok)."},"capabilities":{"tools":true,"structured_output":true,"streaming":true,"batching":true,"prompt_caching":true,"system_prompt":true,"reasoning":true,"vision":true,"audio_input":true,"video_input":true,"document_input":true,"computer_use":true,"web_search":true,"code_execution":true},"performance":{"speed_tps":350},"api":{"endpoint":"https://generativelanguage.googleapis.com/v1beta/models/gemini-3.5-flash-lite:generateContent","protocol":"gemini","openai_compatible":true,"auth":{"type":"bearer","env_var":"GOOGLE_AI_API_KEY"},"sdks":{"google":{"package":"@google/genai","model_id":"gemini-3.5-flash-lite"}}},"sources":{"spec":"https://ai.google.dev/gemini-api/docs/models/gemini-3.5-flash-lite","pricing":"https://ai.google.dev/gemini-api/docs/pricing","release_notes":"https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-6-flash-3-5-flash-lite-3-5-flash-cyber/"},"last_verified":"2026-09-22"},{"id":"gemini-3.5-transcribe","provider":"google","family":"gemini-3.5","display_name":"Gemini 3.5 Transcribe","description":"Speech-to-text model based on Gemini's audio understanding, reaching GA August 26, 2026. Low-latency, accurate transcription of pre-recorded audio (up to 1 hour per request; capped at 30 minutes when diarization or word-level timestamps are enabled) with utterance-based language auto-detection across 85+ languages, speaker diarization (up to 8 speakers), word-level timestamps, smart formatting, and custom vocabulary biasing (up to 1,000 terms). Companion real-time variant: gemini-3.5-transcribe-live.","type":"stt","input_modalities":["audio"],"output_modalities":["text"],"status":"ga","released":"2026-08-26","limits":{"max_audio_seconds":3600},"pricing":{"currency":"USD","input_per_mtok":2,"output_per_mtok":12,"notes":"Priced per million tokens like other Gemini models, not per character/minute flat-rate: effective blended rate is roughly $0.003/min audio input and $0.002/min text output. Free tier available."},"capabilities":{"streaming":false,"word_timestamps":true,"speaker_diarization":true},"api":{"endpoint":"https://generativelanguage.googleapis.com/v1beta/models/gemini-3.5-transcribe:generateContent","protocol":"gemini","openai_compatible":false,"auth":{"type":"bearer","env_var":"GOOGLE_AI_API_KEY"},"sdks":{"google":{"package":"@google/genai","model_id":"gemini-3.5-transcribe"}}},"sources":{"spec":"https://ai.google.dev/gemini-api/docs/models/gemini-3.5-transcribe","pricing":"https://ai.google.dev/gemini-api/docs/pricing","release_notes":"https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-5-transcribe/"},"last_verified":"2026-09-22"},{"id":"gemini-3.5-transcribe-live","provider":"google","family":"gemini-3.5","display_name":"Gemini 3.5 Transcribe Live","description":"Real-time streaming counterpart to gemini-3.5-transcribe, reaching GA August 26, 2026. Live/WebSocket transcription with the same 85+-language auto-detection and custom vocabulary biasing (up to 1,000 terms) as the file-based model, but without diarization or word-level timestamps. Max 10 minutes per live session.","type":"stt","input_modalities":["audio"],"output_modalities":["text"],"status":"ga","released":"2026-08-26","limits":{"max_audio_seconds":600},"pricing":{"currency":"USD","input_per_mtok":3.5,"output_per_mtok":21,"notes":"Priced per million tokens: effective blended rate is roughly $0.005/min audio input and $0.0315/min output. Free tier available."},"capabilities":{"streaming":true},"api":{"endpoint":"wss://generativelanguage.googleapis.com/v1beta/live","protocol":"gemini","openai_compatible":false,"auth":{"type":"bearer","env_var":"GOOGLE_AI_API_KEY"},"sdks":{"google":{"package":"@google/genai","model_id":"gemini-3.5-transcribe-live"}}},"sources":{"pricing":"https://ai.google.dev/gemini-api/docs/pricing","release_notes":"https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-5-transcribe/"},"last_verified":"2026-09-22"},{"id":"gemini-3.5-live-translate-preview","provider":"google","family":"gemini-3.5","display_name":"Gemini Live 3.5 Translate (Preview)","description":"Audio-to-audio real-time translation model based on Gemini 3 Pro. Released June 9, 2026.","type":"audio_chat","input_modalities":["audio"],"output_modalities":["audio"],"status":"preview","released":"2026-06-09","limits":{"context_window":131072,"max_output":65536},"pricing":{"currency":"USD","input_per_mtok":3.5,"output_per_mtok":21,"notes":"Audio tokens at ~25 tok/sec (≈ $0.0368/min effective at output rate)."},"capabilities":{"streaming":true,"audio_input":true,"audio_output":true},"api":{"endpoint":"https://generativelanguage.googleapis.com/v1beta/models/gemini-3.5-live-translate-preview:streamGenerateContent","protocol":"gemini","openai_compatible":false,"auth":{"type":"bearer","env_var":"GOOGLE_AI_API_KEY"}},"sources":{"spec":"https://ai.google.dev/gemini-api/docs/models/gemini-3.5-live-translate-preview","pricing":"https://ai.google.dev/gemini-api/docs/pricing","release_notes":"https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-live-3-5-translate/"},"last_verified":"2026-09-22"},{"id":"gemini-3.8-live","provider":"google","family":"gemini-3.8","display_name":"Gemini 3.8 Live","description":"New GA (Sept 15, 2026) default Live API model for low-latency, real-time voice-to-voice agent experiences without reasoning-induced delays. Supports interleaved reasoning, asynchronous function calling, full session client content updates, and built-in audio streaming. Google's models/deprecations pages now name this the recommended replacement for gemini-3.1-flash-live-preview, which is relabeled \"Legacy\"; both models share the same pricing table on the live pricing page. Added to registry 2026-09-22 (previously untracked). Companion high-reasoning variant: gemini-3.8-live-extended-thinking. Spot-checked 2026-09-28 (Sept 22 addition re-verification): pricing table (raw-HTML fetch) unchanged -- $0.75/MTok text, $3.00/MTok or ~$0.005/min audio, $1.00/MTok or ~$0.002/min image/video input; $4.50/MTok text, $12.00/MTok or ~$0.018/min audio output.","type":"audio_chat","input_modalities":["text","image","audio","video"],"output_modalities":["text","audio"],"status":"ga","released":"2026-09-15","limits":{"context_window":131072,"max_output":65536},"pricing":{"currency":"USD","input_per_mtok":0.75,"output_per_mtok":4.5,"notes":"Shares a combined pricing table with gemini-3.8-live-extended-thinking and gemini-3.1-flash-live-preview. Input: $0.75/MTok text, $3.00/MTok or ~$0.005/min audio, $1.00/MTok or ~$0.002/min image/video. Output: $4.50/MTok text, $12.00/MTok or ~$0.018/min audio."},"capabilities":{"tools":true,"structured_output":false,"streaming":true,"batching":false,"prompt_caching":false,"reasoning":true,"vision":true,"audio_input":true,"video_input":true,"audio_output":true,"web_search":true,"code_execution":false},"api":{"endpoint":"wss://generativelanguage.googleapis.com/v1beta/live","protocol":"gemini","openai_compatible":false,"auth":{"type":"bearer","env_var":"GOOGLE_AI_API_KEY"},"sdks":{"google":{"package":"@google/genai","model_id":"gemini-3.8-live"}},"notes":"thinking_level/thinking_config is not supported (interleaved reasoning is automatic, not configurable). Caching, code execution, file search, image generation, Maps grounding, structured outputs, URL context, and Batch API are not supported."},"sources":{"spec":"https://ai.google.dev/gemini-api/docs/models/gemini-3.8-live","pricing":"https://ai.google.dev/gemini-api/docs/pricing"},"last_verified":"2026-09-28"},{"id":"gemini-3.8-live-extended-thinking","provider":"google","family":"gemini-3.8","display_name":"Gemini 3.8 Live Extended Thinking","description":"New GA (Sept 15, 2026) high-reasoning audio-to-audio Live API model, recommended when higher background reasoning is required during real-time voice interactions. Processes background reasoning and asynchronous (async-only) tool calls while streaming continuous audio responses; turnComplete:true no longer means the model is idle since background reasoning/tool calls may continue after. Added to registry 2026-09-22 (previously untracked).","type":"audio_chat","input_modalities":["text","image","audio","video"],"output_modalities":["text","audio"],"status":"ga","released":"2026-09-15","limits":{"context_window":131072,"max_output":65536},"pricing":{"currency":"USD","input_per_mtok":0.75,"output_per_mtok":4.5,"notes":"Shares a combined pricing table with gemini-3.8-live and gemini-3.1-flash-live-preview. Input: $0.75/MTok text, $3.00/MTok or ~$0.005/min audio, $1.00/MTok or ~$0.002/min image/video. Output: $4.50/MTok text, $12.00/MTok or ~$0.018/min audio."},"capabilities":{"tools":true,"structured_output":false,"streaming":true,"batching":false,"prompt_caching":false,"reasoning":true,"vision":true,"audio_input":true,"video_input":true,"audio_output":true,"web_search":true,"code_execution":false},"api":{"endpoint":"wss://generativelanguage.googleapis.com/v1beta/live","protocol":"gemini","openai_compatible":false,"auth":{"type":"bearer","env_var":"GOOGLE_AI_API_KEY"},"sdks":{"google":{"package":"@google/genai","model_id":"gemini-3.8-live-extended-thinking"}},"notes":"Function calling is async-only. Caching, code execution, file search, image generation, Maps grounding, structured outputs, URL context, and Batch API are not supported."},"sources":{"spec":"https://ai.google.dev/gemini-api/docs/models/gemini-3.8-live-extended-thinking","pricing":"https://ai.google.dev/gemini-api/docs/pricing"},"last_verified":"2026-09-22"},{"id":"gemini-3.1-flash-live-preview","provider":"google","family":"gemini-3.1","display_name":"Gemini 3.1 Flash Live (Preview)","description":"Live API model for low-latency, real-time voice/video dialogue. Announced March 26, 2026 as Google's highest-quality audio-to-audio (A2A) Live model, powering upgrades to Gemini Live and Search Live. Bidirectional streaming over WebSockets. Async function calling, proactive audio, and affective dialogue are not yet supported. Updated 2026-09-22: Google's models page now labels this a \"Legacy Live API preview model\" and recommends updating to gemini-3.8-live, which shipped Sept 15, 2026 and shares this model's pricing. The deprecations table still lists it as \"No shutdown date announced\" (no formal deprecated_on/sunset_on yet), so status is left as preview, but replacement_id now points at gemini-3.8-live.","type":"audio_chat","input_modalities":["text","image","audio","video"],"output_modalities":["text","audio"],"status":"preview","released":"2026-03-26","replacement_id":"gemini-3.8-live","limits":{"context_window":131072,"max_output":65536},"pricing":{"currency":"USD","input_per_mtok":0.75,"output_per_mtok":4.5,"notes":"Audio input: $3.00/MTok (~$0.005/min). Image/video input: $1.00/MTok (~$0.002/min). Audio output: $12.00/MTok (~$0.018/min)."},"capabilities":{"tools":true,"streaming":true,"reasoning":true,"vision":true,"audio_input":true,"video_input":true,"audio_output":true,"web_search":true},"api":{"endpoint":"wss://generativelanguage.googleapis.com/v1beta/live","protocol":"gemini","openai_compatible":false,"auth":{"type":"bearer","env_var":"GOOGLE_AI_API_KEY"},"sdks":{"google":{"package":"@google/genai","model_id":"gemini-3.1-flash-live-preview"}},"notes":"Function calling is synchronous only (async function calling not yet supported). Caching, code execution, file search, image generation, Maps grounding, structured outputs, URL context, and Batch API are not supported."},"sources":{"spec":"https://ai.google.dev/gemini-api/docs/models/gemini-3.1-flash-live-preview","pricing":"https://ai.google.dev/gemini-api/docs/pricing"},"last_verified":"2026-09-22"},{"id":"gemini-3.1-flash-tts-preview","provider":"google","family":"gemini-3.1","display_name":"Gemini 3.1 Flash TTS (Preview)","description":"Native text-to-speech model launched April 15, 2026. Steerable via natural-language style/pace/accent/tone instructions and audio tags, native multi-speaker dialogue (up to 2 speakers), 70+ languages with automatic detection. Streaming is supported (unlike the older gemini-2.5 TTS preview models). Updated 2026-09-28: Google's deprecations table now names gemini-3.8-flash-tts and gemini-3.8-flash-lite-tts (both GA Sept 22, 2026) as the recommended replacements for this model. Still no shutdown date announced, so status remains preview.","type":"tts","input_modalities":["text"],"output_modalities":["audio"],"status":"preview","released":"2026-04-15","replacement_id":"gemini-3.8-flash-tts","limits":{"context_window":32768,"supported_voices":["Zephyr","Puck","Charon","Kore","Fenrir","Leda","Orus","Aoede","Callirrhoe","Autonoe","Enceladus","Iapetus","Umbriel","Algieba","Despina","Erinome","Algenib","Rasalgethi","Laomedeia","Achernar","Alnilam","Schedar","Gacrux","Pulcherrima","Achird","Zubenelgenubi","Vindemiatrix","Sadachbia","Sadaltager","Sulafat"]},"pricing":{"currency":"USD","input_per_mtok":1,"output_per_mtok":20,"batch_input_per_mtok":0.5,"batch_output_per_mtok":10,"notes":"Priced per million tokens, not per character: audio output is billed as tokens like other Gemini models, unlike some other providers' per-character TTS pricing."},"capabilities":{"streaming":true,"speed_control":true},"api":{"endpoint":"https://generativelanguage.googleapis.com/v1beta/models/gemini-3.1-flash-tts-preview:generateContent","protocol":"gemini","openai_compatible":false,"auth":{"type":"bearer","env_var":"GOOGLE_AI_API_KEY"},"sdks":{"google":{"package":"@google/genai","model_id":"gemini-3.1-flash-tts-preview"}}},"sources":{"spec":"https://ai.google.dev/gemini-api/docs/speech-generation","pricing":"https://ai.google.dev/gemini-api/docs/pricing"},"last_verified":"2026-09-28"},{"id":"gemini-3.1-pro-preview","provider":"google","family":"gemini-3.1","display_name":"Gemini 3.1 Pro","description":"Google's most advanced reasoning and coding model, released February 19, 2026. 1M context, multimodal (text/image/audio/video), thinking config. Successor to Gemini 3 Pro.","type":"chat","input_modalities":["text","image","audio","video"],"output_modalities":["text"],"status":"preview","released":"2026-02-19","limits":{"context_window":1048576,"max_output":65536,"max_video_seconds":7200},"pricing":{"currency":"USD","input_per_mtok":2,"cached_input_per_mtok":0.2,"output_per_mtok":12,"batch_input_per_mtok":1,"batch_output_per_mtok":6,"notes":"Pricing tier for prompts ≤200K. >200K: $4 input, $0.40 cached, $18 output."},"capabilities":{"tools":true,"parallel_tools":true,"json_mode":true,"structured_output":true,"streaming":true,"batching":true,"prompt_caching":true,"system_prompt":true,"reasoning":true,"reasoning_modes":["off","low","medium","high"],"vision":true,"audio_input":true,"video_input":true,"document_input":true,"web_search":true,"code_execution":true},"performance":{"intelligence_index":84},"api":{"endpoint":"https://generativelanguage.googleapis.com/v1beta/models/gemini-3.1-pro-preview:generateContent","protocol":"gemini","openai_compatible":true,"auth":{"type":"bearer","env_var":"GOOGLE_AI_API_KEY"},"sdks":{"google":{"package":"@google/genai","model_id":"gemini-3.1-pro-preview"},"vercel-ai":{"package":"@ai-sdk/google","model_id":"gemini-3.1-pro-preview","factory":"google('gemini-3.1-pro-preview')"}},"supported_params":["max_output_tokens","temperature","top_p","top_k","stop_sequences","tools","response_mime_type","response_schema","thinking_config"],"notes":"OpenAI-compatible endpoint at https://generativelanguage.googleapis.com/v1beta/openai/. As of April 1, 2026, Gemini Pro models are paid-only (no free tier)."},"aggregator_ids":{"openrouter":"google/gemini-3.1-pro-preview","gcp-vertex":"gemini-3.1-pro"},"manual_tags":["smartest","best_for_coding"],"sources":{"spec":"https://docs.cloud.google.com/vertex-ai/generative-ai/docs/models/gemini/3-1-pro","pricing":"https://ai.google.dev/gemini-api/docs/pricing","release_notes":"https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-1-pro/"},"last_verified":"2026-09-22"},{"id":"gemini-3-flash-preview","provider":"google","family":"gemini-3","display_name":"Gemini 3 Flash","description":"Fast tier in Gemini 3 family. Audio/video/image input, structured outputs, prompt caching.","type":"chat","input_modalities":["text","image","audio","video"],"output_modalities":["text"],"status":"preview","limits":{"context_window":1048576,"max_output":65536},"pricing":{"currency":"USD","input_per_mtok":0.5,"cached_input_per_mtok":0.05,"output_per_mtok":3,"batch_input_per_mtok":0.25,"batch_output_per_mtok":1.5,"notes":"Audio input: $1/MTok. Cached audio: $0.10/MTok."},"capabilities":{"tools":true,"json_mode":true,"structured_output":true,"streaming":true,"batching":true,"prompt_caching":true,"reasoning":true,"vision":true,"audio_input":true,"video_input":true},"performance":{"speed_tps":250,"intelligence_index":70},"api":{"endpoint":"https://generativelanguage.googleapis.com/v1beta/models/gemini-3-flash-preview:generateContent","protocol":"gemini","openai_compatible":true,"auth":{"type":"bearer","env_var":"GOOGLE_AI_API_KEY"},"sdks":{"google":{"package":"@google/genai","model_id":"gemini-3-flash-preview"}}},"aggregator_ids":{"openrouter":"google/gemini-3-flash-preview","gcp-vertex":"gemini-3-flash"},"sources":{"pricing":"https://ai.google.dev/gemini-api/docs/pricing"},"last_verified":"2026-09-22"},{"id":"gemini-3.1-flash-lite","provider":"google","family":"gemini-3.1","display_name":"Gemini 3.1 Flash-Lite","description":"Cheapest Gemini 3.1 model. Free tier available with reduced daily quotas. Google's official deprecations table lists a concrete shutdown date (2027-05-07) with gemini-3.5-flash-lite as the named successor — status kept as GA since no explicit deprecation banner was found on the model card itself, but this is the official migration target and should be watched for a formal status change. Re-checked 2026-09-14 (raw-HTML fetch of the deprecations table): still lists release date May 7, 2026, shutdown date May 7, 2027, replacement gemini-3.5-flash-lite — but no separate status column and no \"Deprecated models\" grouping applied to this row (unlike, e.g., gemini-omni-flash-preview, which the table explicitly buckets under \"Deprecated models\"). Still no formal deprecation banner; no change from the Aug 24/Sept 7 checks. Re-checked 2026-09-22 (raw-HTML fetch): identical row still under the stable \"Gemini 3 models\" table (not the \"Preview models\" subsection or a \"Deprecated models\" grouping) — release May 7, 2026, shutdown May 7, 2027, replacement gemini-3.5-flash-lite, no status column. No change. Re-checked 2026-09-28 (raw-HTML fetch of the deprecations table): row unchanged -- still under the stable table, same May 7 2026 / May 7 2027 dates and gemini-3.5-flash-lite replacement, still no explicit deprecation banner. No change.","type":"chat","input_modalities":["text","image","audio","video"],"output_modalities":["text"],"status":"ga","released":"2026-05-07","sunset_on":"2027-05-07","replacement_id":"gemini-3.5-flash-lite","limits":{"context_window":1048576,"max_output":65536},"pricing":{"currency":"USD","input_per_mtok":0.25,"cached_input_per_mtok":0.025,"output_per_mtok":1.5,"batch_input_per_mtok":0.125,"batch_output_per_mtok":0.75,"notes":"Audio input: $0.50/MTok. Cached audio: $0.05/MTok. Non-global routing adds 10% premium."},"capabilities":{"tools":true,"json_mode":true,"structured_output":true,"streaming":true,"batching":true,"prompt_caching":true,"vision":true,"audio_input":true,"document_input":true},"performance":{"speed_tps":300,"intelligence_index":60},"api":{"endpoint":"https://generativelanguage.googleapis.com/v1beta/models/gemini-3.1-flash-lite:generateContent","protocol":"gemini","openai_compatible":true,"auth":{"type":"bearer","env_var":"GOOGLE_AI_API_KEY"},"sdks":{"google":{"package":"@google/genai","model_id":"gemini-3.1-flash-lite"}}},"sources":{"pricing":"https://ai.google.dev/gemini-api/docs/pricing"},"last_verified":"2026-09-28"},{"id":"gemini-3-pro-image","provider":"google","family":"gemini-3","display_name":"Gemini 3 Pro Image (Nano Banana Pro)","description":"GA image generation and editing model on Gemini 3 Pro backbone. Multi-turn editing, Search grounding, infographics, 4K resolution. Replaces gemini-3-pro-image-preview.","type":"image_generation","input_modalities":["text","image"],"output_modalities":["image","text"],"status":"ga","limits":{"context_window":65536,"max_output":32768,"supported_resolutions":["1024x1024","2048x2048","4096x4096"]},"pricing":{"currency":"USD","input_per_mtok":2,"output_per_mtok":12,"batch_input_per_mtok":1,"batch_output_per_mtok":6,"image_tiers":[{"label":"1K/2K image","per_image":0.134},{"label":"4K image","per_image":0.24}],"notes":"Image output charged at $120/MTok in addition to text I/O (flat per-image rate applies across Standard and Priority tiers). Batch: $1.00/$6.00 text I/O; per-image batch rate is 50% off ($0.067 for 1K/2K, $0.12 for 4K). Priority: $3.60/$21.60 text I/O."},"capabilities":{"tools":true,"vision":true,"image_output":true,"web_search":true,"image_to_image":true,"inpainting":true},"api":{"endpoint":"https://generativelanguage.googleapis.com/v1beta/models/gemini-3-pro-image:generateContent","protocol":"gemini","openai_compatible":false,"auth":{"type":"bearer","env_var":"GOOGLE_AI_API_KEY"},"sdks":{"google":{"package":"@google/genai","model_id":"gemini-3-pro-image"}}},"aggregator_ids":{"openrouter":"google/gemini-3-pro-image"},"sources":{"pricing":"https://ai.google.dev/gemini-api/docs/pricing"},"last_verified":"2026-09-22"},{"id":"gemini-3.1-flash-image","provider":"google","family":"gemini-3.1","display_name":"Gemini 3.1 Flash Image","description":"GA image generation and editing on Gemini 3.1 Flash backbone. Fast and cost-efficient. Replaces imagen-4.0-fast-generate-001.","type":"image_generation","input_modalities":["text","image"],"output_modalities":["image","text"],"status":"ga","limits":{"context_window":65536,"max_output":32768},"pricing":{"currency":"USD","input_per_mtok":0.5,"output_per_mtok":3,"batch_input_per_mtok":0.25,"batch_output_per_mtok":1.5,"image_tiers":[{"label":"0.5K image","per_image":0.045},{"label":"1K image","per_image":0.067},{"label":"2K image","per_image":0.101},{"label":"4K image","per_image":0.151}],"notes":"Corrected 2026-07-27: previous single $0.039/image figure was stale/wrong; official pricing page lists resolution-based tiers (Standard: $0.045 @0.5K, $0.067 @1K, $0.101 @2K, $0.151 @4K). Image output charged at $60/MTok in addition to text I/O. Batch: $0.25/$1.50 text I/O; per-image batch is 50% off ($0.022/$0.034/$0.050/$0.076 for 0.5K/1K/2K/4K)."},"capabilities":{"vision":true,"image_output":true,"image_to_image":true,"inpainting":true},"api":{"endpoint":"https://generativelanguage.googleapis.com/v1beta/models/gemini-3.1-flash-image:generateContent","protocol":"gemini","openai_compatible":false,"auth":{"type":"bearer","env_var":"GOOGLE_AI_API_KEY"},"sdks":{"google":{"package":"@google/genai","model_id":"gemini-3.1-flash-image"}}},"aggregator_ids":{"openrouter":"google/gemini-3.1-flash-image"},"sources":{"pricing":"https://ai.google.dev/gemini-api/docs/pricing"},"last_verified":"2026-09-22"},{"id":"gemini-3.1-flash-lite-image","provider":"google","family":"gemini-3.1","display_name":"Gemini 3.1 Flash-Lite Image (Nano Banana 2 Lite)","description":"GA lightweight image generation and editing on the Gemini 3.1 Flash-Lite backbone. Released June 30, 2026 as the cheapest tier of the Nano Banana 2 image family.","type":"image_generation","input_modalities":["text","image","video"],"output_modalities":["image","text"],"status":"ga","released":"2026-06-30","pricing":{"currency":"USD","input_per_mtok":0.25,"output_per_mtok":1.5,"batch_input_per_mtok":0.125,"batch_output_per_mtok":0.75,"notes":"Image output charged at $30/MTok (~$0.0336/1K-resolution image), $15/MTok on Batch."},"capabilities":{"vision":true,"video_input":true,"image_output":true,"image_to_image":true,"inpainting":true},"api":{"endpoint":"https://generativelanguage.googleapis.com/v1beta/models/gemini-3.1-flash-lite-image:generateContent","protocol":"gemini","openai_compatible":false,"auth":{"type":"bearer","env_var":"GOOGLE_AI_API_KEY"},"sdks":{"google":{"package":"@google/genai","model_id":"gemini-3.1-flash-lite-image"}}},"aggregator_ids":{"openrouter":"google/gemini-3.1-flash-lite-image"},"sources":{"spec":"https://ai.google.dev/gemini-api/docs/image-generation","pricing":"https://ai.google.dev/gemini-api/docs/pricing"},"last_verified":"2026-09-22"},{"id":"gemini-omni-1.1-flash","provider":"google","family":"gemini-omni","display_name":"Gemini Omni 1.1 Flash","description":"GA (Aug 27, 2026) successor to gemini-omni-flash-preview -- a high-speed multimodal model for text/image-to-video generation and conversational editing via the Interactions API. Adds video extension (append in 3-10s increments up to a total of ~30-40s), start/end-frame interpolation, explicit output-resolution controls (360p/720p/1080p/4K upscaled), and reference-to-video generation from images and short video clips.","type":"video_generation","input_modalities":["text","image","video"],"output_modalities":["video"],"status":"ga","released":"2026-08-27","limits":{"max_video_seconds":10,"max_video_resolution":"4K"},"pricing":{"currency":"USD","input_per_mtok":1.5,"output_per_mtok":17.5,"notes":"Input tokens (text/image/video/audio) billed at a flat $1.50/MTok. Output: $9.00/MTok for text, $17.50/MTok for video (~$0.10/sec effective). No free tier."},"capabilities":{"video_input":true,"image_to_image":true},"api":{"endpoint":"https://generativelanguage.googleapis.com/v1beta/models/gemini-omni-1.1-flash:generateVideo","protocol":"gemini","openai_compatible":false,"auth":{"type":"bearer","env_var":"GOOGLE_AI_API_KEY"}},"sources":{"spec":"https://ai.google.dev/gemini-api/docs/omni","pricing":"https://ai.google.dev/gemini-api/docs/pricing","release_notes":"https://ai.google.dev/gemini-api/docs/changelog"},"last_verified":"2026-09-22"},{"id":"gemini-2.5-pro","provider":"google","family":"gemini-2.5","display_name":"Gemini 2.5 Pro","description":"Previous-generation GA flagship. 1M context, multimodal, thinking. Corrected 2026-08-10: this model had been marked deprecated (sunset Oct 16, 2026) since June 15, but a fresh structural check of Google's own Gemini Developer API deprecations table (ai.google.dev/gemini-api/docs/deprecations, last updated 2026-08-03) shows the bare gemini-2.5-pro row now reads \"No shutdown date announced\" with an empty replacement column, and the pricing page carries no deprecation warning banner for it either (unlike gemini-2.0-flash, Imagen 4, or Veo 2/3, which still do). Both checks independently agree, so status reverted to GA. Note: the separate Vertex AI / Gemini Enterprise Agent Platform retirement table still lists an October 20, 2026 retirement date with Gemini 3.5 Flash as replacement — this registry entry tracks the Gemini Developer API (generativelanguage.googleapis.com), which is the discrepancy; flagged as a watch item. Re-confirmed 2026-09-14 via both raw-HTML and AI-summarized fetches: still \"No shutdown date announced\" on the Developer API deprecations table -- no reversal. Also corrected 2026-09-14: cached_input_per_mtok was recorded as 0.13, a rounding drift from the official rate of 0.125 (>200K tier is correctly 0.25, exactly double). Re-confirmed 2026-09-22 via raw-HTML fetch: still \"No shutdown date announced\", $1.25/$2.50>200k input and $10/$15>200k output pricing unchanged, cached $0.125/$0.25>200k unchanged. Re-confirmed 2026-09-28 (raw-HTML fetch): still \"No shutdown date announced\" on the Developer API deprecations table, and Google now explicitly states in a note on that page (added on or before 2026-09-28): \"To ensure reliable performance for everyone, we are limiting access to the 2.5 models to users who have actively used them in the past. These models are not deprecated and will continue to be served until further notice through the API. For any new projects, use our latest models: 3.5 Flash-Lite or 3.8 Flash.\" This confirms GA/not-deprecated status directly rather than by absence of a banner; access restriction is a capacity measure for new callers, not a deprecation. Pricing unchanged.","type":"chat","input_modalities":["text","image","audio","video"],"output_modalities":["text"],"status":"ga","released":"2025-06-17","limits":{"context_window":1000000,"max_output":65536},"pricing":{"currency":"USD","input_per_mtok":1.25,"cached_input_per_mtok":0.125,"output_per_mtok":10,"batch_input_per_mtok":0.625,"batch_output_per_mtok":5,"notes":">200K context: $2.50/$0.25 cached/$15 output."},"capabilities":{"tools":true,"parallel_tools":true,"json_mode":true,"structured_output":true,"streaming":true,"batching":true,"prompt_caching":true,"system_prompt":true,"reasoning":true,"vision":true,"audio_input":true,"video_input":true,"document_input":true,"web_search":true,"code_execution":true},"performance":{"intelligence_index":75,"benchmarks":{"mmlu_pro":78,"gpqa":72}},"api":{"endpoint":"https://generativelanguage.googleapis.com/v1beta/models/gemini-2.5-pro:generateContent","protocol":"gemini","openai_compatible":true,"auth":{"type":"bearer","env_var":"GOOGLE_AI_API_KEY"},"sdks":{"google":{"package":"@google/genai","model_id":"gemini-2.5-pro"}}},"aggregator_ids":{"openrouter":"google/gemini-2.5-pro","gcp-vertex":"gemini-2.5-pro"},"sources":{"pricing":"https://ai.google.dev/gemini-api/docs/pricing"},"last_verified":"2026-09-28"},{"id":"gemini-2.5-computer-use-preview-10-2025","provider":"google","family":"gemini-2.5","display_name":"Gemini 2.5 Computer Use Preview","description":"Computer-use specialized variant of Gemini 2.5 Pro, optimized for building browser control agents that automate tasks. Corrected 2026-08-10: the registry had tracked this under a stale/non-callable id (gemini-2.5-pro-computer-use, never re-verified since 2026-06-09). Confirmed via a structural fetch of the official pricing page that the real, currently-callable id is gemini-2.5-computer-use-preview-10-2025; the old id string doesn't resolve. This model does not appear on Google's deprecations table (no shutdown date announced), but Google's own computer-use docs now describe it as a \"Legacy model\" and steer new integrations toward gemini-3.6-flash's built-in computer_use tool instead.","type":"chat","input_modalities":["text","image"],"output_modalities":["text"],"status":"preview","limits":{"context_window":1000000,"max_output":65536},"pricing":{"currency":"USD","input_per_mtok":1.25,"output_per_mtok":10,"notes":"Pricing tier for prompts ≤200K. >200K: $2.50 input / $15.00 output per MTok."},"capabilities":{"tools":true,"streaming":true,"reasoning":true,"vision":true,"computer_use":true},"api":{"endpoint":"https://generativelanguage.googleapis.com/v1beta/models/gemini-2.5-computer-use-preview-10-2025:generateContent","protocol":"gemini","openai_compatible":false,"auth":{"type":"bearer","env_var":"GOOGLE_AI_API_KEY"},"sdks":{"google":{"package":"@google/genai","model_id":"gemini-2.5-computer-use-preview-10-2025"}}},"sources":{"pricing":"https://ai.google.dev/gemini-api/docs/pricing","docs":"https://ai.google.dev/gemini-api/docs/computer-use"},"last_verified":"2026-09-22"},{"id":"gemini-2.5-flash","provider":"google","family":"gemini-2.5","display_name":"Gemini 2.5 Flash","description":"Previous-generation fast multimodal model. 1M context, native audio in/out, thinking, code execution. Corrected 2026-08-10: reverted from deprecated (sunset Oct 16, 2026) back to GA — same finding as gemini-2.5-pro. Google's Gemini Developer API deprecations table (last updated 2026-08-03) now shows \"No shutdown date announced\" for the bare gemini-2.5-flash row with no replacement listed, and the pricing page has no deprecation warning banner for it (unlike gemini-2.0-flash/Imagen 4/Veo 2/3, which still carry one). The Vertex AI / Gemini Enterprise Agent Platform retirement table separately still lists October 20, 2026 with Gemini 3.5 Flash-Lite/3.1 Flash-Lite as replacement — noted as a cross-platform discrepancy, not applied here since this entry tracks the Developer API. Re-confirmed 2026-09-14 via both raw-HTML and AI-summarized fetches: still \"No shutdown date announced\" on the Developer API deprecations table -- no reversal. Re-confirmed 2026-09-22 via raw-HTML fetch: still \"No shutdown date announced\", pricing unchanged ($0.30/$2.50, cached $0.03). Re-confirmed 2026-09-28 (raw-HTML fetch): still \"No shutdown date announced\"; Google's deprecations page now carries an explicit note that the 2.5 models are \"not deprecated and will continue to be served until further notice,\" with new-project access limited to prior users as a capacity measure (new projects steered to 3.5 Flash-Lite or 3.8 Flash) -- this is an access restriction, not a deprecation. Pricing unchanged.","type":"chat","input_modalities":["text","image","audio","video"],"output_modalities":["text"],"status":"ga","released":"2025-04-17","limits":{"context_window":1000000,"max_output":65536},"pricing":{"currency":"USD","input_per_mtok":0.3,"cached_input_per_mtok":0.03,"output_per_mtok":2.5,"batch_input_per_mtok":0.15,"batch_output_per_mtok":1.25,"notes":"Audio input: $1/MTok. Cached audio: $0.10/MTok."},"capabilities":{"tools":true,"json_mode":true,"structured_output":true,"streaming":true,"batching":true,"prompt_caching":true,"reasoning":true,"vision":true,"audio_input":true,"video_input":true,"web_search":true,"code_execution":true},"performance":{"speed_tps":250,"ttft_ms":350,"intelligence_index":65,"benchmarks":{"mmlu_pro":70}},"api":{"endpoint":"https://generativelanguage.googleapis.com/v1beta/models/gemini-2.5-flash:generateContent","protocol":"gemini","openai_compatible":true,"auth":{"type":"bearer","env_var":"GOOGLE_AI_API_KEY"},"sdks":{"google":{"package":"@google/genai","model_id":"gemini-2.5-flash"}}},"aggregator_ids":{"openrouter":"google/gemini-2.5-flash","gcp-vertex":"gemini-2.5-flash"},"sources":{"pricing":"https://ai.google.dev/gemini-api/docs/pricing"},"last_verified":"2026-09-28"},{"id":"gemini-2.5-flash-lite","provider":"google","family":"gemini-2.5","display_name":"Gemini 2.5 Flash-Lite","description":"Previous-generation cheapest Gemini 2.5 model. 1M context, multimodal input. Free tier available. Corrected 2026-08-10: reverted from deprecated (sunset Oct 16, 2026) back to GA — same finding as gemini-2.5-pro/gemini-2.5-flash. Google's Gemini Developer API deprecations table (last updated 2026-08-03) now shows \"No shutdown date announced\" for this bare id with no replacement listed, and the pricing page carries no deprecation warning banner for it. The Vertex AI / Gemini Enterprise Agent Platform retirement table separately still lists October 20, 2026 (replacement Gemini 3.1 Flash-Lite or Gemma 4) — noted as a cross-platform discrepancy, not applied here since this entry tracks the Developer API. Re-confirmed 2026-09-14 via both raw-HTML and AI-summarized fetches: still \"No shutdown date announced\" on the Developer API deprecations table -- no reversal. Re-confirmed 2026-09-22 via raw-HTML fetch: still \"No shutdown date announced\", pricing unchanged ($0.10/$0.40, cached $0.01; audio $0.30/MTok). Re-confirmed 2026-09-28 (raw-HTML fetch): still \"No shutdown date announced\"; same explicit \"not deprecated, will continue to be served\" note now on the deprecations page applies here too (access to 2.5 models limited to prior users as a capacity measure, new projects steered to 3.5 Flash-Lite/3.8 Flash). Pricing unchanged.","type":"chat","input_modalities":["text","image","audio"],"output_modalities":["text"],"status":"ga","limits":{"context_window":1000000,"max_output":65536},"pricing":{"currency":"USD","input_per_mtok":0.1,"cached_input_per_mtok":0.01,"output_per_mtok":0.4,"batch_input_per_mtok":0.05,"batch_output_per_mtok":0.2,"notes":"Audio input: $0.30/MTok. Cached audio: $0.03/MTok."},"capabilities":{"tools":true,"json_mode":true,"streaming":true,"batching":true,"prompt_caching":true,"vision":true,"audio_input":true},"performance":{"speed_tps":350,"intelligence_index":55},"api":{"endpoint":"https://generativelanguage.googleapis.com/v1beta/models/gemini-2.5-flash-lite:generateContent","protocol":"gemini","openai_compatible":true,"auth":{"type":"bearer","env_var":"GOOGLE_AI_API_KEY"},"sdks":{"google":{"package":"@google/genai","model_id":"gemini-2.5-flash-lite"}}},"aggregator_ids":{"openrouter":"google/gemini-2.5-flash-lite","gcp-vertex":"gemini-2.5-flash-lite"},"sources":{"pricing":"https://ai.google.dev/gemini-api/docs/pricing"},"last_verified":"2026-09-28"},{"id":"gemini-2.5-flash-native-audio-preview-12-2025","provider":"google","family":"gemini-2.5","display_name":"Gemini 2.5 Flash Live Preview","description":"Live API for low-latency voice/video conversations. Bidirectional streaming over WebSockets. Corrected 2026-08-10: the registry had tracked this under a non-existent id (gemini-2.5-flash-live, never re-verified since 2026-06-09) whose pricing figures actually matched this model exactly. The real id, confirmed via a structural fetch of the official pricing page and cross-checked against Google's deprecations table (\"No shutdown date announced\"), is gemini-2.5-flash-native-audio-preview-12-2025, released December 12, 2025. Note: two other, similarly-named ids exist and are already retired — gemini-live-2.5-flash-preview (shutdown Dec 9, 2025) and gemini-2.0-flash-live-001 (shutdown Dec 9, 2025) — both replaced by gemini-3.1-flash-live-preview, not this model. Updated 2026-09-22: Google's deprecations table now lists gemini-3.8-live (released Sept 15, 2026) as the recommended replacement instead of gemini-3.1-flash-live-preview. Still \"No shutdown date announced\", so status is unchanged.","type":"audio_chat","input_modalities":["text","audio","video","image"],"output_modalities":["text","audio"],"status":"preview","released":"2025-12-12","replacement_id":"gemini-3.8-live","limits":{"context_window":1000000,"max_output":8192},"pricing":{"currency":"USD","input_per_mtok":0.5,"output_per_mtok":2,"notes":"Audio input: $3/MTok. Audio output: $12/MTok. Video/image input: $3/MTok."},"capabilities":{"tools":true,"streaming":true,"vision":true,"audio_input":true,"video_input":true,"audio_output":true},"api":{"endpoint":"wss://generativelanguage.googleapis.com/v1beta/live","protocol":"gemini","openai_compatible":false,"auth":{"type":"bearer","env_var":"GOOGLE_AI_API_KEY"},"sdks":{"google":{"package":"@google/genai","model_id":"gemini-2.5-flash-native-audio-preview-12-2025"}}},"sources":{"spec":"https://ai.google.dev/gemini-api/docs/models/gemini-2.5-flash-native-audio-preview-12-2025","pricing":"https://ai.google.dev/gemini-api/docs/pricing"},"last_verified":"2026-09-22"},{"id":"gemini-embedding-2","provider":"google","family":"gemini-embedding","display_name":"Gemini Embedding 2","description":"Google's latest multimodal embedding model. Supports text, image, audio, and video inputs for semantic search and retrieval. Successor to Gemini Embedding 001.","type":"embedding","input_modalities":["text","image","audio","video"],"output_modalities":["embedding"],"status":"ga","released":"2026-04-22","pricing":{"currency":"USD","input_per_mtok":0.2,"notes":"Text: $0.20/M tokens. Image: $0.45/M ($0.00012/image). Audio: $6.50/M ($0.00016/sec). Video: $12.00/M ($0.00079/frame). Batch rates: 50% off."},"capabilities":{"batching":true,"vision":true,"audio_input":true,"video_input":true},"api":{"endpoint":"https://generativelanguage.googleapis.com/v1beta/models/gemini-embedding-2:embedContent","protocol":"gemini","openai_compatible":false,"auth":{"type":"bearer","env_var":"GOOGLE_AI_API_KEY"},"sdks":{"google":{"package":"@google/genai","model_id":"gemini-embedding-2"}}},"manual_tags":["smartest_embedding"],"sources":{"pricing":"https://ai.google.dev/gemini-api/docs/pricing"},"last_verified":"2026-09-22"},{"id":"lyria-3-clip-preview","provider":"google","family":"lyria","display_name":"Lyria 3 Clip (Preview)","description":"Google DeepMind music generation model. Short clip generation (≤30 seconds). Released March 25, 2026. Re-checked 2026-09-07: still live and served on the Gemini API/Vertex AI alongside the newer lyria-3.5, which Google has not stated will replace it.","type":"audio_chat","input_modalities":["text"],"output_modalities":["audio"],"status":"preview","released":"2026-03-25","pricing":{"currency":"USD","notes":"$0.04/song."},"api":{"endpoint":"https://generativelanguage.googleapis.com/v1beta/models/lyria-3-clip-preview:predict","protocol":"gemini","openai_compatible":false,"auth":{"type":"bearer","env_var":"GOOGLE_AI_API_KEY"}},"sources":{"spec":"https://deepmind.google/models/lyria/","pricing":"https://ai.google.dev/gemini-api/docs/pricing"},"last_verified":"2026-09-22"},{"id":"lyria-3-pro-preview","provider":"google","family":"lyria","display_name":"Lyria 3 Pro (Preview)","description":"Google DeepMind music generation model. Full song generation (up to 3 minutes). Released March 25, 2026. Re-checked 2026-09-07: still live and served on the Gemini API/Vertex AI alongside the newer lyria-3.5, which Google has not stated will replace it.","type":"audio_chat","input_modalities":["text"],"output_modalities":["audio"],"status":"preview","released":"2026-03-25","pricing":{"currency":"USD","notes":"$0.08/song."},"api":{"endpoint":"https://generativelanguage.googleapis.com/v1beta/models/lyria-3-pro-preview:predict","protocol":"gemini","openai_compatible":false,"auth":{"type":"bearer","env_var":"GOOGLE_AI_API_KEY"}},"sources":{"spec":"https://deepmind.google/models/lyria/","pricing":"https://ai.google.dev/gemini-api/docs/pricing"},"last_verified":"2026-09-22"},{"id":"lyria-3.5","provider":"google","family":"lyria","display_name":"Lyria 3.5","description":"Google DeepMind's newest music generation model, in preview. Rolled out inside Google Flow Music July 29, 2026, then opened more broadly through Google AI Studio, the Gemini API, and the Gemini app on September 4, 2026. Successor to Lyria 3 Pro for full-length song generation with improved musicality, lyric quality, vocal expression, and tempo/duration control (multiple verses, choruses, bridges; 44.1kHz stereo). Google has not stated whether it formally replaces lyria-3-pro-preview, which remains separately available.","type":"audio_chat","input_modalities":["text","image"],"output_modalities":["audio","text"],"status":"preview","released":"2026-09-04","limits":{"context_window":131072},"pricing":{"currency":"USD","notes":"$0.08/song. No free tier."},"api":{"endpoint":"https://generativelanguage.googleapis.com/v1beta/models/lyria-3.5:predict","protocol":"gemini","openai_compatible":false,"auth":{"type":"bearer","env_var":"GOOGLE_AI_API_KEY"}},"sources":{"spec":"https://ai.google.dev/gemini-api/docs/models/lyria-3.5","pricing":"https://ai.google.dev/gemini-api/docs/pricing","docs":"https://ai.google.dev/gemini-api/docs/music-generation"},"last_verified":"2026-09-22"},{"id":"veo-3.1-generate-preview","provider":"google","family":"veo","display_name":"Veo 3.1","description":"Latest Veo video generation model. Native audio. Fast and Standard tiers. Still carries a preview id/status in Google's own docs despite being the current recommended Veo model.","type":"video_generation","input_modalities":["text","image"],"output_modalities":["video"],"status":"preview","limits":{"max_video_seconds":8,"max_video_resolution":"1080p"},"pricing":{"currency":"USD","video_per_second":0.4,"notes":"Veo 3.1 Standard: $0.40/sec at 720p/1080p, $0.60/sec at 4k (audio included). Veo 3.1 Fast (veo-3.1-fast-generate-preview): $0.10/sec at 720p, $0.12/sec at 1080p, $0.30/sec at 4k — not separately tracked here. The Lite tier is tracked separately as veo-3.1-lite-generate-preview."},"capabilities":{"audio_output":true,"image_to_image":true},"api":{"endpoint":"https://generativelanguage.googleapis.com/v1beta/models/veo-3.1-generate-preview:generateVideo","protocol":"gemini","openai_compatible":false,"auth":{"type":"bearer","env_var":"GOOGLE_AI_API_KEY"}},"aggregator_ids":{"openrouter":"google/veo-3.1"},"sources":{"spec":"https://ai.google.dev/gemini-api/docs/veo"},"last_verified":"2026-09-22"},{"id":"veo-3.1-lite-generate-preview","provider":"google","family":"veo","display_name":"Veo 3.1 Lite (Preview)","description":"Most cost-effective Veo 3.1 tier, launched March 31, 2026 (less than 50% of the cost of Veo 3.1 Fast at the same speed). Was a pre-existing registry gap -- mentioned only in veo-3.1-generate-preview's pricing notes since it was added -- closed 2026-09-14. Generates 4s/6s/8s clips at 24fps in 720p or 1080p (no 4K), landscape (16:9) or portrait (9:16), text-to-video and image-to-video, video output includes audio. Does not support the Extension feature that some other Veo 3.1 tiers offer. Text prompt input capped at 1,024 tokens.","type":"video_generation","input_modalities":["text","image"],"output_modalities":["video"],"status":"preview","released":"2026-03-31","limits":{"max_video_seconds":8,"max_video_resolution":"1080p"},"pricing":{"currency":"USD","video_per_second":0.08,"notes":"Video-with-audio price (default): $0.05/sec at 720p, $0.08/sec at 1080p. 4K output is not supported."},"capabilities":{"audio_output":true,"image_to_image":true},"api":{"endpoint":"https://generativelanguage.googleapis.com/v1beta/models/veo-3.1-lite-generate-preview:generateVideo","protocol":"gemini","openai_compatible":false,"auth":{"type":"bearer","env_var":"GOOGLE_AI_API_KEY"}},"sources":{"spec":"https://ai.google.dev/gemini-api/docs/models/veo-3.1-lite-generate-preview","pricing":"https://ai.google.dev/gemini-api/docs/pricing","release_notes":"https://blog.google/innovation-and-ai/technology/ai/veo-3-1-lite/"},"last_verified":"2026-09-22"},{"id":"gemma-4-26b-a4b-it","provider":"google","family":"gemma","display_name":"Gemma 4 26B A4B (Instruction-Tuned)","description":"Open-weights Gemma 4 Mixture-of-Experts model served via the Gemini API. 25.2B total parameters, 3.8B active per token (8 active / 128 total experts, plus 1 shared expert). Corrected 2026-08-10: id renamed from the non-callable \"gemma-4-26b-it\" (tracked since the June 9 initial pass, never re-verified) to the real, currently-listed id gemma-4-26b-a4b-it, confirmed via a structural fetch of Google's official Gemma-on-Gemini-API docs (every code sample on that page uses this exact id). Context window and modalities also corrected: official model card lists 256K (262,144) context and text+image input (not text-only, 128K as previously recorded); audio input is only supported on the smaller E2B/E4B/12B Gemma 4 variants, not this one. Pricing corrected: the Gemini API pricing page currently lists Gemma 4 as free of charge with no paid tier available (previous $0.15/$0.60 figures were unconfirmed/unsourced).","type":"chat","input_modalities":["text","image"],"output_modalities":["text"],"status":"ga","released":"2026-07-30","limits":{"context_window":262144},"pricing":{"currency":"USD","input_per_mtok":0,"output_per_mtok":0,"notes":"Free of charge via the Gemini API (Google AI Studio); no paid tier currently available (confirmed on the official pricing page)."},"capabilities":{"tools":true,"streaming":true,"prompt_caching":true,"system_prompt":true,"reasoning":true,"vision":true},"api":{"endpoint":"https://generativelanguage.googleapis.com/v1beta/models/gemma-4-26b-a4b-it:generateContent","protocol":"gemini","openai_compatible":true,"auth":{"type":"bearer","env_var":"GOOGLE_AI_API_KEY"},"sdks":{"google":{"package":"@google/genai","model_id":"gemma-4-26b-a4b-it"}}},"sources":{"spec":"https://ai.google.dev/gemma/docs/core/model_card_4","pricing":"https://ai.google.dev/gemini-api/docs/pricing","docs":"https://ai.google.dev/gemma/docs/core/gemma_on_gemini_api"},"last_verified":"2026-09-22"},{"id":"gemma-4-31b-it","provider":"google","family":"gemma","display_name":"Gemma 4 31B (Instruction-Tuned)","description":"Open-weights Gemma 4 dense model served via the Gemini API. 30.7B total parameters (dense, not MoE), released alongside gemma-4-26b-a4b-it. Was entirely missing from the registry — added 2026-08-10 after confirming both ids via a structural fetch of Google's official Gemma-on-Gemini-API docs and model card. 140+ pre-trained languages, 35+ with native support.","type":"chat","input_modalities":["text","image"],"output_modalities":["text"],"status":"ga","released":"2026-07-30","limits":{"context_window":262144},"pricing":{"currency":"USD","input_per_mtok":0,"output_per_mtok":0,"notes":"Free of charge via the Gemini API (Google AI Studio); no paid tier currently available (confirmed on the official pricing page)."},"capabilities":{"tools":true,"streaming":true,"prompt_caching":true,"system_prompt":true,"reasoning":true,"vision":true},"api":{"endpoint":"https://generativelanguage.googleapis.com/v1beta/models/gemma-4-31b-it:generateContent","protocol":"gemini","openai_compatible":true,"auth":{"type":"bearer","env_var":"GOOGLE_AI_API_KEY"},"sdks":{"google":{"package":"@google/genai","model_id":"gemma-4-31b-it"}}},"sources":{"spec":"https://ai.google.dev/gemma/docs/core/model_card_4","pricing":"https://ai.google.dev/gemini-api/docs/pricing","docs":"https://ai.google.dev/gemma/docs/core/gemma_on_gemini_api"},"last_verified":"2026-09-22"},{"id":"gemini-robotics-er-2-preview","provider":"google","family":"gemini-robotics","display_name":"Gemini Robotics-ER 2 (Preview)","description":"Embodied-reasoning model for robotics (spatial understanding, affordance/grasp prediction, trajectory planning) built on a Gemini 3.5 Flash backbone. Launched 2026-07-30. Replaces gemini-robotics-er-1.6-preview, which shut down 2026-08-31. Corrected 2026-09-14: official pricing was published for the first time this pass (previously the registry carried an editorial estimate) and turns out to be introductory/promotional -- $1.00 input / $5.00 output per MTok through Dec 31, 2026, doubling to $2.00/$10.00 on Jan 1, 2027 -- confirmed via two independent fetches (raw HTML + AI-summarized) of the live pricing page.","type":"chat","input_modalities":["text","image","video","audio"],"output_modalities":["text"],"status":"preview","released":"2026-07-30","limits":{"context_window":131072,"max_output":65536},"pricing":{"currency":"USD","input_per_mtok":1,"output_per_mtok":5,"batch_input_per_mtok":0.5,"batch_output_per_mtok":2.5,"notes":"Introductory pricing confirmed on the official pricing page through Dec 31, 2026: $1.00 input / $5.00 output per MTok (text/image/video/audio), Standard context caching $0.10/MTok ($0.50/MTok/hr storage), Batch $0.50 input / $2.50 output. Standard rate doubles Jan 1, 2027 to $2.00 input / $10.00 output (caching $0.20/MTok, $1.00/MTok/hr storage). Re-verified 2026-09-22 against the live pricing page: unchanged."},"capabilities":{"tools":true,"structured_output":true,"streaming":true,"batching":true,"prompt_caching":true,"reasoning":true,"vision":true,"audio_input":true,"video_input":true,"computer_use":true,"web_search":true,"code_execution":true},"api":{"endpoint":"https://generativelanguage.googleapis.com/v1beta/models/gemini-robotics-er-2-preview:generateContent","protocol":"gemini","openai_compatible":true,"auth":{"type":"bearer","env_var":"GOOGLE_AI_API_KEY"},"sdks":{"google":{"package":"@google/genai","model_id":"gemini-robotics-er-2-preview"}}},"sources":{"spec":"https://ai.google.dev/gemini-api/docs/models/gemini-robotics-er-2-preview","pricing":"https://ai.google.dev/gemini-api/docs/pricing","release_notes":"https://ai.google.dev/gemini-api/docs/changelog","docs":"https://deepmind.google/models/model-cards/gemini-robotics-er-2/"},"last_verified":"2026-09-22"},{"id":"gemini-robotics-er-2-streaming-preview","provider":"google","family":"gemini-robotics","display_name":"Gemini Robotics-ER 2 Streaming (Preview)","description":"Live/streaming variant of gemini-robotics-er-2-preview for real-time robotics control loops. Launched 2026-07-30. Live API models don't support batching, prompt caching, code execution, computer use, or structured output — only the base capabilities carry over from the primary variant. Exact websocket endpoint form not explicitly confirmed on the model page; inferred by analogy with other Gemini Live API models in this registry (e.g. gemini-3.1-flash-live-preview) — re-verify before relying on it for a live integration. Corrected 2026-09-14: official pricing (published for the first time this pass) is introductory -- $1.00 input / $5.00 output per MTok through Dec 31, 2026, doubling to $2.00/$10.00 on Jan 1, 2027 -- same promo structure as the base gemini-robotics-er-2-preview, confirmed via two independent fetches of the live pricing page.","type":"audio_chat","input_modalities":["text","image","video","audio"],"output_modalities":["text"],"status":"preview","released":"2026-07-30","limits":{"context_window":131072,"max_output":65536},"pricing":{"currency":"USD","input_per_mtok":1,"output_per_mtok":5,"notes":"Introductory pricing confirmed on the official pricing page through Dec 31, 2026: $1.00 input / $5.00 output per MTok (text/image/video/audio). Standard rate doubles to $2.00/$10.00 on Jan 1, 2027. Re-verified 2026-09-22 against the live pricing page: unchanged."},"capabilities":{"tools":true,"streaming":true,"reasoning":true,"vision":true,"audio_input":true,"video_input":true},"api":{"endpoint":"https://generativelanguage.googleapis.com/v1beta/models/gemini-robotics-er-2-streaming-preview:generateContent","protocol":"gemini","openai_compatible":true,"auth":{"type":"bearer","env_var":"GOOGLE_AI_API_KEY"},"sdks":{"google":{"package":"@google/genai","model_id":"gemini-robotics-er-2-streaming-preview"}}},"sources":{"spec":"https://ai.google.dev/gemini-api/docs/models/gemini-robotics-er-2-preview","pricing":"https://ai.google.dev/gemini-api/docs/pricing","docs":"https://deepmind.google/models/model-cards/gemini-robotics-er-2/"},"last_verified":"2026-09-22"},{"id":"gemini-3.8-flash-tts","provider":"google","family":"gemini-3.8","display_name":"Gemini 3.8 Flash TTS","description":"New GA (Sept 22, 2026) flagship creative text-to-speech model, engineered for studio-grade voice fidelity, expressive acting, authentic regional accents, and long-form multi-turn stability. Detects input language automatically across 130+ languages. Supports single- and multi-speaker (up to 2 speakers) generation, streaming output, voice replication (clone a speaker's voice from reference + consent audio), and steerable style/accent/pace/tone via structured speech_metadata plus inline vocal tags. Named on Google's deprecations table as the (joint, with gemini-3.8-flash-lite-tts) recommended replacement for gemini-3.1-flash-tts-preview and the older gemini-2.5-flash/pro-preview-tts models. Added to registry 2026-09-28 (previously untracked).","type":"tts","input_modalities":["text"],"output_modalities":["audio"],"status":"ga","released":"2026-09-22","limits":{"context_window":8192,"max_output":16384,"supported_voices":["Zephyr","Puck","Charon","Kore","Fenrir","Leda","Orus","Aoede","Callirrhoe","Autonoe","Enceladus","Iapetus","Umbriel","Algieba","Despina","Erinome","Algenib","Rasalgethi","Laomedeia","Achernar","Alnilam","Schedar","Gacrux","Pulcherrima","Achird","Zubenelgenubi","Vindemiatrix","Sadachbia","Sadaltager","Sulafat"]},"pricing":{"currency":"USD","input_per_mtok":0.5,"cached_input_per_mtok":0.125,"output_per_mtok":9,"batch_input_per_mtok":0.25,"batch_output_per_mtok":4.5,"notes":"Promotional pricing through Dec 31, 2026: $0.50 input (text) / $9.00 output (audio) per MTok; standard rate doubles Jan 1, 2027 to $1.00/$18.00. Context caching: $0.125 input-caching through 12/31/26 -> $0.25 from 1/1/27, storage $0.50/MTok/hr -> $1.00/MTok/hr. Batch: $0.25/$4.50 -> $0.50/$9.00 (caching $0.0625 -> $0.125). Flex: same standard/batch rate as Batch tier ($0.25/$4.50 -> $0.50/$9.00), caching $0.025 -> $0.05. Priority: $0.90/$16.20 -> $1.80/$32.40, caching $0.225 -> $0.45. Free tier available."},"capabilities":{"streaming":true,"batching":true,"prompt_caching":true,"voice_cloning":true,"speed_control":true},"api":{"endpoint":"https://generativelanguage.googleapis.com/v1beta/models/gemini-3.8-flash-tts:generateContent","protocol":"gemini","openai_compatible":false,"auth":{"type":"bearer","env_var":"GOOGLE_AI_API_KEY"},"sdks":{"google":{"package":"@google/genai","model_id":"gemini-3.8-flash-tts"}},"notes":"Multi-speaker generation (speech_config.speakers) supports up to 2 speakers with prebuilt voices; hundreds more available via the Extended Voice Library (/v1beta/voices). Does not support code execution, file search, function calling, Maps/Search grounding, image generation, Live API, structured outputs, thinking, or URL context."},"sources":{"spec":"https://ai.google.dev/gemini-api/docs/models/gemini-3.8-flash-tts","pricing":"https://ai.google.dev/gemini-api/docs/pricing","release_notes":"https://ai.google.dev/gemini-api/docs/changelog","docs":"https://ai.google.dev/gemini-api/docs/speech-generation"},"last_verified":"2026-09-28"},{"id":"gemini-3.8-flash-lite-tts","provider":"google","family":"gemini-3.8","display_name":"Gemini 3.8 Flash-Lite TTS","description":"New GA (Sept 22, 2026) fast, cost-efficient workhorse text-to-speech model built for high-throughput production workloads, real-time voice-agent cascades, and everyday single-speaker generation across 100+ auto-detected languages. Shares the same capability surface as gemini-3.8-flash-tts (streaming, up to 2-speaker dialogue, voice replication, steerable style/accent/pace/tone) at a lower price point. Named on Google's deprecations table as the (joint, with gemini-3.8-flash-tts) recommended replacement for gemini-3.1-flash-tts-preview and the older gemini-2.5-flash/pro-preview-tts models. Added to registry 2026-09-28 (previously untracked).","type":"tts","input_modalities":["text"],"output_modalities":["audio"],"status":"ga","released":"2026-09-22","limits":{"context_window":8192,"max_output":16384,"supported_voices":["Zephyr","Puck","Charon","Kore","Fenrir","Leda","Orus","Aoede","Callirrhoe","Autonoe","Enceladus","Iapetus","Umbriel","Algieba","Despina","Erinome","Algenib","Rasalgethi","Laomedeia","Achernar","Alnilam","Schedar","Gacrux","Pulcherrima","Achird","Zubenelgenubi","Vindemiatrix","Sadachbia","Sadaltager","Sulafat"]},"pricing":{"currency":"USD","input_per_mtok":0.5,"cached_input_per_mtok":0.125,"output_per_mtok":6,"batch_input_per_mtok":0.25,"batch_output_per_mtok":3,"notes":"Promotional pricing through Dec 31, 2026: $0.50 input (text) / $6.00 output (audio) per MTok; standard rate doubles Jan 1, 2027 to $1.00/$12.00. Context caching: $0.125 input-caching through 12/31/26 -> $0.25 from 1/1/27, storage $0.50/MTok/hr -> $1.00/MTok/hr. Batch: $0.25/$3.00 -> $0.50/$6.00 (caching $0.0625 -> $0.125). Flex: same standard/batch rate as Batch tier ($0.25/$3.00 -> $0.50/$6.00), caching $0.025 -> $0.05. Priority: $0.90/$10.80 -> $1.80/$21.60, caching $0.225 -> $0.45. Free tier available."},"capabilities":{"streaming":true,"batching":true,"prompt_caching":true,"voice_cloning":true,"speed_control":true},"api":{"endpoint":"https://generativelanguage.googleapis.com/v1beta/models/gemini-3.8-flash-lite-tts:generateContent","protocol":"gemini","openai_compatible":false,"auth":{"type":"bearer","env_var":"GOOGLE_AI_API_KEY"},"sdks":{"google":{"package":"@google/genai","model_id":"gemini-3.8-flash-lite-tts"}},"notes":"Multi-speaker generation (speech_config.speakers) supports up to 2 speakers with prebuilt voices; hundreds more available via the Extended Voice Library (/v1beta/voices). Does not support code execution, file search, function calling, Maps/Search grounding, image generation, Live API, structured outputs, thinking, or URL context."},"sources":{"spec":"https://ai.google.dev/gemini-api/docs/models/gemini-3.8-flash-lite-tts","pricing":"https://ai.google.dev/gemini-api/docs/pricing","release_notes":"https://ai.google.dev/gemini-api/docs/changelog","docs":"https://ai.google.dev/gemini-api/docs/speech-generation"},"last_verified":"2026-09-28"},{"id":"gemini-2.5-flash-preview-tts","provider":"google","family":"gemini-2.5","display_name":"Gemini 2.5 Flash Preview TTS","description":"Previously untracked -- added 2026-09-22 during a routine audit prompted by discovering it on the live models page (\"Gemini 2.5 Flash TTS\", listed under both the Gemini 2.5 Flash and Audio models sections). Price-performant, low-latency, controllable text-to-speech model, price-per-token like other Gemini models rather than per-character. Updated 2026-09-28: Google's deprecations table now lists the recommended replacement as gemini-3.8-flash-tts or gemini-3.8-flash-lite-tts (both GA Sept 22, 2026), superseding the prior gemini-3.1-flash-tts-preview recommendation. Still no shutdown date announced, so status is kept GA/preview as documented rather than deprecated.","type":"tts","input_modalities":["text"],"output_modalities":["audio"],"status":"preview","replacement_id":"gemini-3.8-flash-tts","pricing":{"currency":"USD","input_per_mtok":0.5,"output_per_mtok":10,"batch_input_per_mtok":0.25,"batch_output_per_mtok":5,"notes":"Priced per million tokens, not per character: $0.50/MTok text input, $10.00/MTok audio output. Batch: $0.25/$5.00."},"capabilities":{"streaming":false},"api":{"endpoint":"https://generativelanguage.googleapis.com/v1beta/models/gemini-2.5-flash-preview-tts:generateContent","protocol":"gemini","openai_compatible":false,"auth":{"type":"bearer","env_var":"GOOGLE_AI_API_KEY"},"sdks":{"google":{"package":"@google/genai","model_id":"gemini-2.5-flash-preview-tts"}}},"sources":{"spec":"https://ai.google.dev/gemini-api/docs/models","pricing":"https://ai.google.dev/gemini-api/docs/pricing"},"last_verified":"2026-09-28"},{"id":"gemini-2.5-pro-preview-tts","provider":"google","family":"gemini-2.5","display_name":"Gemini 2.5 Pro Preview TTS","description":"Previously untracked -- added 2026-09-22 alongside gemini-2.5-flash-preview-tts during a routine audit. High-fidelity text-to-speech model optimized for quality in structured workflows like podcasts and audiobooks, with more natural, steerable outputs than the Flash TTS tier. Updated 2026-09-28: Google's deprecations table now lists the recommended replacement as gemini-3.8-flash-tts or gemini-3.8-flash-lite-tts (both GA Sept 22, 2026), superseding the prior gemini-3.1-flash-tts-preview recommendation. Still no shutdown date announced.","type":"tts","input_modalities":["text"],"output_modalities":["audio"],"status":"preview","replacement_id":"gemini-3.8-flash-tts","pricing":{"currency":"USD","input_per_mtok":1,"output_per_mtok":20,"notes":"Priced per million tokens, not per character: $1.00/MTok text input, $20.00/MTok audio output."},"capabilities":{"streaming":false},"api":{"endpoint":"https://generativelanguage.googleapis.com/v1beta/models/gemini-2.5-pro-preview-tts:generateContent","protocol":"gemini","openai_compatible":false,"auth":{"type":"bearer","env_var":"GOOGLE_AI_API_KEY"},"sdks":{"google":{"package":"@google/genai","model_id":"gemini-2.5-pro-preview-tts"}}},"sources":{"spec":"https://ai.google.dev/gemini-api/docs/models","pricing":"https://ai.google.dev/gemini-api/docs/pricing"},"last_verified":"2026-09-28"},{"id":"deep-research-pro-preview-12-2025","provider":"google","family":"deep-research","display_name":"Deep Research Pro (Preview)","description":"Resolves a watch item open since the Sept 14 pass. A powerful agentic researcher, powered by Gemini 3.1 Pro, designed for autonomous multi-step investigations that synthesize complex information into comprehensive, cited reports across hundreds of public web sources and private workspace data (Gmail/Drive). Called via the Interactions API using an \"agent\" code rather than a chat model id. Confirmed live and publicly documented (own doc page returns HTTP 200) as of 2026-09-22, but superseded on Google's main models listing by deep-research-preview-04-2026 / deep-research-max-preview-04-2026 (April 2026) -- it no longer appears in the \"Tool and agent models\" table, only reachable via its dedicated docs page. Not listed on the deprecations table at all (no shutdown date, no replacement column), so treated as still-live preview rather than formally deprecated. No fixed per-agent price: billed at standard Gemini 3.1 Pro token rates plus tool fees (Search grounding, URL context, File Search) at their normal rates; compute/sandbox is not separately billed during preview. Spot-checked 2026-09-28 (Sept 22 addition re-verification): still absent from the deprecations table (raw-HTML grep for the id returns no shutdown/replacement row), consistent with the Sept 22 finding.","type":"chat","input_modalities":["text","image","audio","video"],"output_modalities":["text"],"status":"preview","released":"2025-12-01","limits":{"context_window":1048576,"max_output":65536},"pricing":{"currency":"USD","notes":"No fixed per-agent price. Billed at standard Gemini 3.1 Pro list rates for model inference (input/output/reasoning tokens) plus normal tool fees for Search grounding, URL context, and File Search. Environment compute is not billed during the preview period."},"capabilities":{"tools":true,"streaming":true,"reasoning":true,"vision":true,"audio_input":true,"video_input":true,"document_input":true,"web_search":true},"api":{"endpoint":"https://generativelanguage.googleapis.com/v1beta/interactions","protocol":"gemini","openai_compatible":false,"auth":{"type":"bearer","env_var":"GOOGLE_AI_API_KEY"},"notes":"Invoked as an agent code via the Interactions API (not a plain generateContent model id)."},"sources":{"spec":"https://ai.google.dev/gemini-api/docs/models/deep-research-pro-preview-12-2025","pricing":"https://ai.google.dev/gemini-api/docs/pricing","docs":"https://ai.google.dev/gemini-api/docs/deep-research"},"last_verified":"2026-09-28"},{"id":"deep-research-preview-04-2026","provider":"google","family":"deep-research","display_name":"Deep Research (Preview)","description":"Resolves a watch item open since the Sept 14 pass. Current-generation agentic researcher (April 2026) that autonomously plans and executes multi-step research across hundreds of sources to produce cited, interactive reports; supports collaborative planning, visualization, MCP servers, and File Search. Optimized for speed/efficiency and streaming back to a client UI, as opposed to the Max variant. Called via the Interactions API using an \"agent\" code. Not on the deprecations table (no shutdown date announced). No fixed per-agent price: billed at standard underlying Gemini token rates plus tool fees; compute/sandbox not separately billed during preview.","type":"chat","input_modalities":["text","image","audio","video"],"output_modalities":["text","image"],"status":"preview","released":"2026-04-01","limits":{"context_window":1048576,"max_output":65536},"pricing":{"currency":"USD","notes":"No fixed per-agent price. Billed at standard underlying Gemini model list rates for inference (input/output/reasoning tokens) plus normal tool fees for Search grounding, URL context, and File Search. Environment compute is not billed during the preview period."},"capabilities":{"tools":true,"streaming":true,"reasoning":true,"vision":true,"audio_input":true,"video_input":true,"document_input":true,"web_search":true},"api":{"endpoint":"https://generativelanguage.googleapis.com/v1beta/interactions","protocol":"gemini","openai_compatible":false,"auth":{"type":"bearer","env_var":"GOOGLE_AI_API_KEY"},"notes":"Invoked as an agent code via the Interactions API (not a plain generateContent model id)."},"sources":{"spec":"https://ai.google.dev/gemini-api/docs/models/deep-research-preview-04-2026","pricing":"https://ai.google.dev/gemini-api/docs/pricing","docs":"https://ai.google.dev/gemini-api/docs/deep-research"},"last_verified":"2026-09-22"},{"id":"deep-research-max-preview-04-2026","provider":"google","family":"deep-research","display_name":"Deep Research Max (Preview)","description":"Resolves a watch item open since the Sept 14 pass. Maximum-comprehensiveness sibling of deep-research-preview-04-2026, optimized for long-running, accuracy-critical investigations synthesizing hundreds of public web sources and private workspace data into cited reports; supports collaborative planning, visualization, MCP servers, and File Search. Called via the Interactions API using an \"agent\" code. Not on the deprecations table (no shutdown date announced). No fixed per-agent price: billed at standard underlying Gemini token rates plus tool fees.","type":"chat","input_modalities":["text","image","audio","video"],"output_modalities":["text","image"],"status":"preview","released":"2026-04-01","limits":{"context_window":1048576,"max_output":65536},"pricing":{"currency":"USD","notes":"No fixed per-agent price. Billed at standard underlying Gemini model list rates for inference (input/output/reasoning tokens) plus normal tool fees for Search grounding, URL context, and File Search. Environment compute is not billed during the preview period."},"capabilities":{"tools":true,"streaming":true,"reasoning":true,"vision":true,"audio_input":true,"video_input":true,"document_input":true,"web_search":true},"api":{"endpoint":"https://generativelanguage.googleapis.com/v1beta/interactions","protocol":"gemini","openai_compatible":false,"auth":{"type":"bearer","env_var":"GOOGLE_AI_API_KEY"},"notes":"Invoked as an agent code via the Interactions API (not a plain generateContent model id)."},"sources":{"spec":"https://ai.google.dev/gemini-api/docs/models/deep-research-max-preview-04-2026","pricing":"https://ai.google.dev/gemini-api/docs/pricing","docs":"https://ai.google.dev/gemini-api/docs/deep-research"},"last_verified":"2026-09-22"},{"id":"antigravity-preview-09-2026","provider":"google","family":"antigravity","display_name":"Antigravity Agent (Preview, September 2026)","description":"New current version of the Antigravity managed agent, released Sept 17, 2026 (5 days before this pass) as the replacement for antigravity-preview-05-2026 per Google's deprecations table. General-purpose managed agent that autonomously plans, reasons, runs code, manages files, and browses the web inside a secure, isolated Linux sandbox hosted by Google; called via the Interactions API with a configurable underlying Gemini model (defaults to gemini-3.8-flash per the Antigravity agent guide's current code samples). No dedicated model-card page at /gemini-api/docs/models/antigravity-preview-09-2026 yet (404 as of 2026-09-22) -- confirmed instead via the deprecations table and the antigravity-agent guide's live code samples, which use this id throughout. No fixed per-agent price: billed at the configured underlying Gemini model's standard token rates; compute/sandbox not separately billed during preview. Context window/output limits carried over from the prior version pending an official model card (re-verify before relying on them). Per the Sept 17, 2026 changelog entry, this version also changed the agent's tool call parameters from snake_case to PascalCase and added line-range-replacement file editing. Re-checked 2026-09-28: deprecations table still shows \"No shutdown date announced\" for this id; no change.","type":"chat","input_modalities":["text","image"],"output_modalities":["text"],"status":"preview","released":"2026-09-17","limits":{"context_window":1048576,"max_output":65536},"pricing":{"currency":"USD","notes":"No fixed per-agent price. Billed at the configured underlying Gemini model's standard list rates for inference (input/output/reasoning tokens); default underlying model is gemini-3.8-flash (configurable, e.g. gemini-3.5-flash-lite). Environment compute (CPU, memory, sandbox execution) is not billed during the preview period."},"capabilities":{"tools":true,"streaming":true,"reasoning":true,"vision":true,"web_search":true,"code_execution":true},"api":{"endpoint":"https://generativelanguage.googleapis.com/v1beta/interactions","protocol":"gemini","openai_compatible":false,"auth":{"type":"bearer","env_var":"GOOGLE_AI_API_KEY"},"notes":"Invoked as an agent code via the Interactions API (not a plain generateContent model id). No standalone /models/ doc page as of 2026-09-22; confirmed via ai.google.dev/gemini-api/docs/deprecations and ai.google.dev/gemini-api/docs/antigravity-agent."},"sources":{"pricing":"https://ai.google.dev/gemini-api/docs/pricing","docs":"https://ai.google.dev/gemini-api/docs/antigravity-agent"},"last_verified":"2026-09-28"},{"id":"lyria-realtime-exp","provider":"google","family":"lyria","display_name":"Lyria RealTime (Experimental)","description":"Resolves a watch item open since the Sept 14 pass. Experimental engine for high-fidelity, low-latency streaming musical synthesis via WebSockets, released May 20, 2025. Takes weighted text prompts as input and streams raw 16-bit PCM audio (48kHz stereo) as output, with a maximum control latency of 2 seconds for steering the generation live -- best for AI-assisted songwriting and instrumental generation (no vocals). Older than lyria-3-clip-preview/lyria-3-pro-preview/lyria-3.5 but Google has not deprecated it: still listed with \"No shutdown date announced\" on the deprecations table and no recommended replacement. No entry on the live pricing page at all (no billable rate found via raw-HTML grep), consistent with its \"Experimental\" labeling -- likely free of charge or usage-based pricing not yet published; do not assume free without re-confirming from the account/billing console. Spot-checked 2026-09-28 (Sept 22 addition re-verification): deprecations row and pricing-page absence both unchanged from the Sept 22 check.","type":"audio_chat","input_modalities":["text"],"output_modalities":["audio"],"status":"preview","released":"2025-05-20","pricing":{"currency":"USD","notes":"No priced row found on the live Gemini Developer API pricing page as of 2026-09-22 (unlike lyria-3.5/lyria-3-clip-preview/lyria-3-pro-preview, which are all priced per-song). Consistent with \"Experimental\" status; exact billing treatment unconfirmed -- do not assume free of charge without re-checking."},"capabilities":{"streaming":true},"api":{"endpoint":"wss://generativelanguage.googleapis.com/v1beta/models/lyria-realtime-exp:streamGenerateContent","protocol":"gemini","openai_compatible":false,"auth":{"type":"bearer","env_var":"GOOGLE_AI_API_KEY"},"sdks":{"google":{"package":"@google/genai","model_id":"lyria-realtime-exp"}}},"sources":{"spec":"https://ai.google.dev/gemini-api/docs/models/lyria-realtime-exp","pricing":"https://ai.google.dev/gemini-api/docs/pricing","docs":"https://ai.google.dev/gemini-api/docs/realtime-music-generation"},"last_verified":"2026-09-28"},{"id":"openai/gpt-oss-120b","provider":"groq","family":"gpt-oss","display_name":"GPT-OSS 120B (Groq)","description":"OpenAI's open-weights 120B parameter MoE model served on Groq with low/medium/high reasoning modes (default: medium, per Groq's API reference). RE-VERIFIED 2026-09-22 via raw HTML of console.groq.com/docs/models and the per-model page: pricing unchanged ($0.15 input / $0.075 cached input / $0.60 output per MTok, confirmed on the per-model spec card), 500 tps, 131,072 context / 65,536 max output, 250K TPM / 1K RPM developer-plan rate limit — no discrepancies found. RE-VERIFIED 2026-09-28 via console.groq.com/docs/models live table: unchanged — $0.15/$0.60 per MTok, 500 tps, 131,072/65,536 — no discrepancies found.","type":"chat","input_modalities":["text"],"output_modalities":["text"],"status":"ga","limits":{"context_window":131072,"max_output":65536},"pricing":{"currency":"USD","input_per_mtok":0.15,"cached_input_per_mtok":0.075,"output_per_mtok":0.6},"capabilities":{"tools":true,"parallel_tools":true,"json_mode":true,"structured_output":true,"streaming":true,"prompt_caching":true,"system_prompt":true,"reasoning":true,"reasoning_modes":["low","medium","high"]},"performance":{"speed_tps":500,"intelligence_index":65},"api":{"endpoint":"https://api.groq.com/openai/v1/chat/completions","protocol":"openai_compat","openai_compatible":true,"auth":{"type":"bearer","env_var":"GROQ_API_KEY"},"sdks":{"openai":{"package":"openai","model_id":"openai/gpt-oss-120b"},"groq":{"package":"groq-sdk","model_id":"openai/gpt-oss-120b"}}},"aggregator_ids":{"openrouter":"openai/gpt-oss-120b"},"manual_tags":["best_value"],"sources":{"spec":"https://console.groq.com/docs/model/openai/gpt-oss-120b","pricing":"https://console.groq.com/docs/models"},"last_verified":"2026-09-28"},{"id":"openai/gpt-oss-20b","provider":"groq","family":"gpt-oss","display_name":"GPT-OSS 20B (Groq)","description":"OpenAI's open-weights 20B parameter model. Very fast inference with low/medium/high reasoning modes (default: medium, per Groq's API reference). RE-VERIFIED 2026-09-22 via raw HTML of console.groq.com/docs/models and the per-model page: pricing unchanged ($0.075 input / $0.0375 cached input, shown rounded as $0.037 on the per-model card / $0.30 output per MTok), 1000 tps, 131,072 context / 65,536 max output, 250K TPM / 1K RPM developer-plan rate limit — no discrepancies found. RE-VERIFIED 2026-09-28 via console.groq.com/docs/models live table: unchanged — $0.075/$0.30 per MTok, 1000 tps, 131,072/65,536 — no discrepancies found.","type":"chat","input_modalities":["text"],"output_modalities":["text"],"status":"ga","limits":{"context_window":131072,"max_output":65536},"pricing":{"currency":"USD","input_per_mtok":0.075,"cached_input_per_mtok":0.0375,"output_per_mtok":0.3},"capabilities":{"tools":true,"json_mode":true,"structured_output":true,"streaming":true,"prompt_caching":true,"reasoning":true,"reasoning_modes":["low","medium","high"]},"performance":{"speed_tps":1000,"intelligence_index":50},"api":{"endpoint":"https://api.groq.com/openai/v1/chat/completions","protocol":"openai_compat","openai_compatible":true,"auth":{"type":"bearer","env_var":"GROQ_API_KEY"},"sdks":{"openai":{"package":"openai","model_id":"openai/gpt-oss-20b"}}},"aggregator_ids":{"openrouter":"openai/gpt-oss-20b"},"manual_tags":["fastest"],"sources":{"spec":"https://console.groq.com/docs/model/openai/gpt-oss-20b","pricing":"https://console.groq.com/docs/models"},"last_verified":"2026-09-28"},{"id":"openai/gpt-oss-safeguard-20b","provider":"groq","family":"gpt-oss","display_name":"GPT-OSS Safeguard 20B (Groq)","description":"Safety-tuned variant of GPT-OSS 20B. Used for content moderation and policy classification. RE-VERIFIED 2026-09-22 via raw HTML of console.groq.com/docs/models and the per-model page: still Preview Models (no Enterprise tag), pricing unchanged ($0.075 input / $0.0375 cached input / $0.30 output per MTok), 1000 tps, 131,072 context / 65,536 max output, 150K TPM / 1K RPM developer-plan rate limit — no discrepancies found. RE-VERIFIED 2026-09-28 via console.groq.com/docs/models live table: still Preview Models, unchanged — $0.075/$0.30 per MTok, 1000 tps, 131,072/65,536 — no discrepancies found.","type":"chat","input_modalities":["text"],"output_modalities":["text"],"status":"preview","limits":{"context_window":131072,"max_output":65536},"pricing":{"currency":"USD","input_per_mtok":0.075,"cached_input_per_mtok":0.0375,"output_per_mtok":0.3},"capabilities":{"tools":true,"json_mode":true,"streaming":true,"reasoning":true,"reasoning_modes":["low","medium","high"]},"performance":{"speed_tps":1000},"api":{"endpoint":"https://api.groq.com/openai/v1/chat/completions","protocol":"openai_compat","openai_compatible":true,"auth":{"type":"bearer","env_var":"GROQ_API_KEY"},"sdks":{"openai":{"package":"openai","model_id":"openai/gpt-oss-safeguard-20b"}}},"sources":{"spec":"https://console.groq.com/docs/model/openai/gpt-oss-safeguard-20b","pricing":"https://console.groq.com/docs/models"},"last_verified":"2026-09-28"},{"id":"whisper-large-v3-turbo","provider":"groq","family":"whisper","display_name":"Whisper Large v3 Turbo (Groq)","description":"Fastest Whisper variant on Groq — 216x speed factor (corrected 2026-07-27, previously listed as 228x). $0.04/hour transcribed. RE-VERIFIED 2026-09-22 via raw HTML of console.groq.com/docs/models and the per-model page: 216x speed factor and $0.04/hour pricing both confirmed unchanged; 400K ASH / 400 RPM developer-plan rate limit, 100MB max file size. RE-VERIFIED 2026-09-28 via console.groq.com/docs/models live table: unchanged.","type":"stt","input_modalities":["audio"],"output_modalities":["text"],"status":"ga","limits":{"max_audio_size_mb":100,"supported_audio_formats":["flac","mp3","mp4","mpeg","mpga","m4a","ogg","wav","webm"]},"pricing":{"currency":"USD","audio_input_per_minute":0.000667,"notes":"Billed per audio hour: $0.04/hour."},"capabilities":{"word_timestamps":true},"performance":{"speed_tps":0},"api":{"endpoint":"https://api.groq.com/openai/v1/audio/transcriptions","protocol":"openai_compat","openai_compatible":true,"auth":{"type":"bearer","env_var":"GROQ_API_KEY"},"sdks":{"openai":{"package":"openai","model_id":"whisper-large-v3-turbo"}}},"manual_tags":["cheapest_stt"],"sources":{"spec":"https://console.groq.com/docs/speech-to-text","pricing":"https://console.groq.com/docs/models"},"last_verified":"2026-09-28"},{"id":"whisper-large-v3","provider":"groq","family":"whisper","display_name":"Whisper Large v3 (Groq)","description":"Standard Whisper Large v3 on Groq — 189x speed factor (corrected 2026-07-27, previously listed as 217x). $0.111/hour transcribed. RE-VERIFIED 2026-09-22 via raw HTML of console.groq.com/docs/models and the per-model page: 189x speed factor and $0.111/hour pricing both confirmed unchanged; 200K ASH / 300 RPM developer-plan rate limit, 100MB max file size. RE-VERIFIED 2026-09-28 via console.groq.com/docs/models live table: unchanged.","type":"stt","input_modalities":["audio"],"output_modalities":["text"],"status":"ga","limits":{"max_audio_size_mb":100,"supported_audio_formats":["flac","mp3","mp4","mpeg","mpga","m4a","ogg","wav","webm"]},"pricing":{"currency":"USD","audio_input_per_minute":0.00185,"notes":"Billed per audio hour: $0.111/hour."},"capabilities":{"word_timestamps":true},"api":{"endpoint":"https://api.groq.com/openai/v1/audio/transcriptions","protocol":"openai_compat","openai_compatible":true,"auth":{"type":"bearer","env_var":"GROQ_API_KEY"},"sdks":{"openai":{"package":"openai","model_id":"whisper-large-v3"}}},"sources":{"spec":"https://console.groq.com/docs/speech-to-text","pricing":"https://console.groq.com/docs/models"},"last_verified":"2026-09-28"},{"id":"canopylabs/orpheus-v1-english","provider":"groq","family":"orpheus","display_name":"Orpheus v1 English (Groq)","description":"Canopy Labs Orpheus TTS, English. Replaced PlayAI TTS on Groq following the Jan 30, 2026 platform-wide migration. RE-VERIFIED 2026-09-22 via raw HTML of console.groq.com/docs/models and the per-model page: still Preview Models, $22.00/1M characters confirmed, per-model page text confirms \"Keep input text under 200 characters maximum per request\" (matches max_input_chars). Note: the models table also shows generic CONTEXT WINDOW 4,000 / MAX COMPLETION TOKENS 50,000 token-based columns for this row, but these appear to be shared generic table columns not specific to TTS input-character limits — the per-model page's explicit 200-character guidance is used instead. RE-VERIFIED 2026-09-28 via console.groq.com/docs/models live table: still Preview, $22.00/1M chars unchanged.","type":"tts","input_modalities":["text"],"output_modalities":["audio"],"status":"preview","limits":{"supported_audio_formats":["wav","mp3"],"max_input_chars":200},"pricing":{"currency":"USD","per_million_characters":22},"capabilities":{"streaming":true,"speed_control":true},"api":{"endpoint":"https://api.groq.com/openai/v1/audio/speech","protocol":"openai_compat","openai_compatible":true,"auth":{"type":"bearer","env_var":"GROQ_API_KEY"},"sdks":{"openai":{"package":"openai","model_id":"canopylabs/orpheus-v1-english"}}},"sources":{"spec":"https://console.groq.com/docs/text-to-speech/orpheus","pricing":"https://console.groq.com/docs/models"},"last_verified":"2026-09-28"},{"id":"canopylabs/orpheus-arabic-saudi","provider":"groq","family":"orpheus","display_name":"Orpheus Arabic Saudi (Groq)","description":"Canopy Labs Orpheus TTS, Arabic (Saudi). Replaced PlayAI TTS Arabic Saudi on Groq following the Jan 30, 2026 platform-wide migration. RE-VERIFIED 2026-09-22 via raw HTML of console.groq.com/docs/models: still Preview Models, $40.00/1M characters confirmed unchanged, 50K TPM / 250 RPM developer-plan rate limit. RE-VERIFIED 2026-09-28 via console.groq.com/docs/models live table: unchanged.","type":"tts","input_modalities":["text"],"output_modalities":["audio"],"status":"preview","limits":{"supported_languages":["ar-SA"],"max_input_chars":200},"pricing":{"currency":"USD","per_million_characters":40},"capabilities":{"streaming":true},"api":{"endpoint":"https://api.groq.com/openai/v1/audio/speech","protocol":"openai_compat","openai_compatible":true,"auth":{"type":"bearer","env_var":"GROQ_API_KEY"},"sdks":{"openai":{"package":"openai","model_id":"canopylabs/orpheus-arabic-saudi"}}},"sources":{"spec":"https://console.groq.com/docs/text-to-speech/orpheus","pricing":"https://console.groq.com/docs/models"},"last_verified":"2026-09-28"},{"id":"meta-llama/llama-prompt-guard-2-22m","provider":"groq","family":"llama-prompt-guard","display_name":"Llama Prompt Guard 2 (22M)","description":"Specialized classifier model to detect and prevent prompt injection/jailbreak attacks. Listed under Preview Models with confirmed pricing on the canonical console.groq.com/docs/models table. RE-VERIFIED 2026-09-22 via raw HTML of console.groq.com/docs/models: $0.03/$0.03 input/output pricing, 512/512 context/max-output, 30K TPM / 100 RPM developer-plan rate limit — all unchanged. groq.com/pricing remains unreachable (still 308-redirects to the homepage). RE-VERIFIED 2026-09-28: unchanged; groq.com/pricing confirmed still 308-redirecting via curl -D -.","type":"moderation","input_modalities":["text"],"output_modalities":["text"],"status":"preview","limits":{"context_window":512,"max_output":512},"pricing":{"currency":"USD","input_per_mtok":0.03,"output_per_mtok":0.03},"api":{"endpoint":"https://api.groq.com/openai/v1/chat/completions","protocol":"openai_compat","openai_compatible":true,"auth":{"type":"bearer","env_var":"GROQ_API_KEY"},"sdks":{"openai":{"package":"openai","model_id":"meta-llama/llama-prompt-guard-2-22m"}}},"sources":{"spec":"https://console.groq.com/docs/model/meta-llama/llama-prompt-guard-2-22m","pricing":"https://console.groq.com/docs/models"},"last_verified":"2026-09-28"},{"id":"meta-llama/llama-prompt-guard-2-86m","provider":"groq","family":"llama-prompt-guard","display_name":"Llama Prompt Guard 2 (86M)","description":"Specialized classifier model to detect and prevent prompt injection/jailbreak attacks, larger/more accurate variant. Listed under Preview Models with confirmed pricing on the canonical console.groq.com/docs/models table. RE-VERIFIED 2026-09-22 via raw HTML of console.groq.com/docs/models: $0.04/$0.04 input/output pricing, 512/512 context/max-output, 30K TPM / 100 RPM developer-plan rate limit — all unchanged. groq.com/pricing remains unreachable (still 308-redirects to the homepage). RE-VERIFIED 2026-09-28: unchanged; groq.com/pricing confirmed still 308-redirecting via curl -D -.","type":"moderation","input_modalities":["text"],"output_modalities":["text"],"status":"preview","limits":{"context_window":512,"max_output":512},"pricing":{"currency":"USD","input_per_mtok":0.04,"output_per_mtok":0.04},"api":{"endpoint":"https://api.groq.com/openai/v1/chat/completions","protocol":"openai_compat","openai_compatible":true,"auth":{"type":"bearer","env_var":"GROQ_API_KEY"},"sdks":{"openai":{"package":"openai","model_id":"meta-llama/llama-prompt-guard-2-86m"}}},"sources":{"spec":"https://console.groq.com/docs/model/meta-llama/llama-prompt-guard-2-86m","pricing":"https://console.groq.com/docs/models"},"last_verified":"2026-09-28"},{"id":"minimaxai/minimax-m2.7","provider":"groq","family":"minimax","display_name":"MiniMax M2.7 (Groq)","description":"New addition to Groq's catalog, added 2026-08-24 (first seen in this registry; not present as of the 2026-08-17 pass). 229B-parameter Mixture-of-Experts model (~10B active per token) from MiniMax, built for agentic workflows and software engineering with interleaved thinking across multi-step tool-calling tasks. Listed under Preview Models on console.groq.com/docs/models. Enterprise-only on Groq: page text reads \"MiniMax M2.7 is available to Enterprise customers. Contact sales for access\" — no public per-token pricing is shown anywhere (page has no Pricing section), so pricing is intentionally left unset rather than invented. Verified via raw HTML fetch of the per-model spec card (context window, max output, speed, capabilities icons for Tool Use / JSON Object Mode / Reasoning). Re-checked 2026-09-07 via raw HTML fetch of both console.groq.com/docs/models and the per-model page: no change — still Preview/Enterprise-only, still no public pricing, specs unchanged. Re-checked again 2026-09-14 (raw HTML): unchanged — Preview Models table, Enterprise-only, Contact Sales/Contact Sales, 260 tps, 196,608/131,072 context/max-output, no pricing section on the per-model page. Note: console.groq.com/docs/changelog's most recent \"Added\" entry is still dated Apr 18, 2026 (\"MiniMax M2.5 and Qwen3-VL 32B Instruct\", an older Enterprise pair that no longer appears anywhere on the live models table) — the changelog has not been updated to reflect minimax-m2.7 or qwen3.8-27b at all, so it is stale/unreliable for spotting new releases; the models table itself remains the authoritative source and was used instead. RE-CHECKED 2026-09-22 (raw HTML of console.groq.com/docs/models and the per-model page): unchanged — still Preview Models, Enterprise-only, Contact Sales/Contact Sales, 260 tps, 196,608/131,072 context/max-output, page text still reads \"MiniMax M2.7 is available to Enterprise customers\" with no pricing section. RE-CHECKED 2026-09-28 via console.groq.com/docs/models and the per-model page: unchanged — still Preview Models, Enterprise-only, Contact Sales/Contact Sales, 260 tps, 196,608/131,072 context/max-output, no pricing section.","type":"chat","input_modalities":["text"],"output_modalities":["text"],"status":"preview","limits":{"context_window":196608,"max_output":131072},"capabilities":{"tools":true,"json_mode":true,"streaming":true,"reasoning":true},"performance":{"speed_tps":260},"api":{"endpoint":"https://api.groq.com/openai/v1/chat/completions","protocol":"openai_compat","openai_compatible":true,"auth":{"type":"bearer","env_var":"GROQ_API_KEY"},"sdks":{"openai":{"package":"openai","model_id":"minimaxai/minimax-m2.7"}},"notes":"Enterprise-tier access only on Groq; contact Groq sales for API access. No public pricing published."},"sources":{"spec":"https://console.groq.com/docs/model/minimaxai/minimax-m2.7","pricing":"https://console.groq.com/docs/models"},"last_verified":"2026-09-28"},{"id":"qwen/qwen3.8-27b","provider":"groq","family":"qwen-3.8","display_name":"Qwen 3.8 27B (Groq)","description":"New addition to Groq's catalog, first seen in this registry 2026-09-07 (not present as of the 2026-08-24 pass). Dense 27B-parameter model from Alibaba's Qwen series (64 layers, hybrid Gated DeltaNet + Gated Attention design, 5120 hidden dim) with a dual-mode system: thinking mode (reasoning_effort=\"default\"/\"low\"/\"medium\"/\"high\") for complex reasoning/math/coding, and instruct mode (reasoning_effort=\"none\") for efficient dialogue. Accepts text and image input (up to 3 images per request, each counted as 2048 input tokens for billing) for OCR/visual QA/chart understanding. Listed under Preview Models on console.groq.com/docs/models, confirmed via raw HTML of both the models table and the per-model spec page. Official benchmarks (from the per-model spec page, not fabricated): GPQA Diamond 89.2%, LiveCodeBench v6 90.3%, SWE-bench Pro 61.7%, Terminal-Bench 2.1 73.0%, IFBench 79.5%. Re-checked 2026-09-14 via raw HTML of console.groq.com/docs/models and the per-model page: unchanged — still Preview Models, $0.80/$4.00, 131,042/16,384 context/max-output, 20MB/3-image limits, ~450+ tps. Re-checked again 2026-09-22 (raw HTML of console.groq.com/docs/models and the per-model page): unchanged — still Preview Models, $0.80/$4.00, 450 tps, 131,042/16,384 context/max-output, 20MB max file size, 250K TPM/1K RPM developer-plan rate limit; benchmarks (GPQA Diamond 89.2%, LiveCodeBench v6 90.3%, SWE-bench Pro 61.7%, Terminal-Bench 2.1 73.0%, IFBench 79.5%) reconfirmed on the per-model Performance Metrics section verbatim. Re-checked again 2026-09-28 via console.groq.com/docs/models and the per-model page: unchanged — still Preview Models, $0.80/$4.00 per MTok, 450 tps, 131,072/16,384 context/max-output. Per Groq's API reference, qwen/qwen3.8-27b's reasoning_effort default is \"none\" and its \"high\" value selects the model's native \"xhigh\" mode — both already reflected in this entry's reasoning_modes. Note: console.groq.com/docs/deprecations gained an entry this pass announcing qwen/qwen3.6-27b's same-day deprecation+retirement in favor of this model (qwen3.8-27b) as its direct successor — see qwen/qwen3.6-27b's entry.","type":"chat","input_modalities":["text","image"],"output_modalities":["text"],"status":"preview","limits":{"context_window":131042,"max_output":16384,"max_image_size_mb":20,"max_images_per_request":3},"pricing":{"currency":"USD","input_per_mtok":0.8,"output_per_mtok":4},"capabilities":{"tools":true,"json_mode":true,"structured_output":true,"streaming":true,"reasoning":true,"reasoning_modes":["none","low","medium","default","high"],"vision":true},"performance":{"speed_tps":450,"benchmarks":{"gpqa_diamond":89.2,"livecodebench_v6":90.3,"swebench_pro":61.7,"terminal_bench_2_1":73,"ifbench":79.5}},"api":{"endpoint":"https://api.groq.com/openai/v1/chat/completions","protocol":"openai_compat","openai_compatible":true,"auth":{"type":"bearer","env_var":"GROQ_API_KEY"},"sdks":{"openai":{"package":"openai","model_id":"qwen/qwen3.8-27b"},"groq":{"package":"groq-sdk","model_id":"qwen/qwen3.8-27b"}}},"aggregator_ids":{"openrouter":"qwen/qwen3.8-27b"},"sources":{"spec":"https://console.groq.com/docs/model/qwen/qwen3.8-27b","pricing":"https://console.groq.com/docs/models"},"last_verified":"2026-09-28"},{"id":"mistral-medium-3-5","provider":"mistral","family":"mistral-medium","display_name":"Mistral Medium 3.5","description":"Frontier-class 128B dense multimodal model optimized for agentic and coding use cases. Released April 28, 2026. Corrected 2026-07-27: canonical API id is 'mistral-medium-3-5' (aliases 'mistral-medium-3', 'mistral-medium-latest'), not 'mistral-medium-2604' — the docs.mistral.ai model card and changelog never use a dated '-2604' suffix for this release.","type":"chat","input_modalities":["text","image"],"output_modalities":["text"],"status":"ga","released":"2026-04-28","limits":{"context_window":262144,"max_output":16384},"pricing":{"currency":"USD","input_per_mtok":1.5,"output_per_mtok":7.5,"batch_input_per_mtok":0.75,"batch_output_per_mtok":3.75},"capabilities":{"tools":true,"parallel_tools":true,"json_mode":true,"structured_output":true,"streaming":true,"batching":true,"system_prompt":true,"reasoning":true,"vision":true,"document_input":true},"performance":{"intelligence_index":70,"benchmarks":{"mmlu_pro":75,"humaneval":88}},"api":{"endpoint":"https://api.mistral.ai/v1/chat/completions","protocol":"openai_compat","openai_compatible":true,"auth":{"type":"bearer","env_var":"MISTRAL_API_KEY"},"sdks":{"mistralai":{"package":"@mistralai/mistralai","model_id":"mistral-medium-3-5"},"vercel-ai":{"package":"@ai-sdk/mistral","model_id":"mistral-medium-3-5","factory":"mistral('mistral-medium-3-5')"}}},"aggregator_ids":{"openrouter":"mistralai/mistral-medium-3-5"},"manual_tags":["best_value"],"sources":{"spec":"https://docs.mistral.ai/models/model-cards/mistral-medium-3-5-26-04","pricing":"https://mistral.ai/pricing"},"last_verified":"2026-09-28"},{"id":"mistral-large-2512","provider":"mistral","family":"mistral-large","display_name":"Mistral Large 3","description":"Open-weight, state-of-the-art general-purpose multimodal model. Released December 2025.","type":"chat","input_modalities":["text","image"],"output_modalities":["text"],"status":"ga","released":"2025-12-02","limits":{"context_window":262144,"max_output":32768},"pricing":{"currency":"USD","input_per_mtok":0.5,"output_per_mtok":1.5,"batch_input_per_mtok":0.25,"batch_output_per_mtok":0.75},"capabilities":{"tools":true,"parallel_tools":true,"json_mode":true,"structured_output":true,"streaming":true,"batching":true,"system_prompt":true,"vision":true,"document_input":true},"performance":{"intelligence_index":68,"benchmarks":{"mmlu_pro":73}},"api":{"endpoint":"https://api.mistral.ai/v1/chat/completions","protocol":"openai_compat","openai_compatible":true,"auth":{"type":"bearer","env_var":"MISTRAL_API_KEY"},"sdks":{"mistralai":{"package":"@mistralai/mistralai","model_id":"mistral-large-latest"}}},"aggregator_ids":{"openrouter":"mistralai/mistral-large-3"},"sources":{"spec":"https://docs.mistral.ai/models/model-cards/mistral-large-3-25-12","pricing":"https://mistral.ai/pricing/api"},"last_verified":"2026-09-28"},{"id":"mistral-small-2603","provider":"mistral","family":"mistral-small","display_name":"Mistral Small 4","description":"Hybrid model for instruct, reasoning and coding. Open-weights, released March 2026.","type":"chat","input_modalities":["text","image"],"output_modalities":["text"],"status":"ga","released":"2026-03-16","limits":{"context_window":262144,"max_output":8192},"pricing":{"currency":"USD","input_per_mtok":0.15,"output_per_mtok":0.6,"batch_input_per_mtok":0.075,"batch_output_per_mtok":0.3,"notes":"Corrected 2026-07-06 from $0.10/$0.30 to the current official rate of $0.15/$0.60."},"capabilities":{"tools":true,"json_mode":true,"streaming":true,"batching":true,"reasoning":true,"vision":true},"performance":{"intelligence_index":60},"api":{"endpoint":"https://api.mistral.ai/v1/chat/completions","protocol":"openai_compat","openai_compatible":true,"auth":{"type":"bearer","env_var":"MISTRAL_API_KEY"},"sdks":{"mistralai":{"package":"@mistralai/mistralai","model_id":"mistral-small-latest"}}},"aggregator_ids":{"openrouter":"mistralai/mistral-small-4"},"manual_tags":["cheapest"],"sources":{"spec":"https://docs.mistral.ai/models/model-cards/mistral-small-4-0-26-03","pricing":"https://mistral.ai/pricing/api"},"last_verified":"2026-09-28"},{"id":"zai-glm-5-2","provider":"mistral","family":"zai-glm","display_name":"Z.ai GLM 5.2","description":"Third-party open-source text model from Z.ai (GLM 5.2), hosted by Mistral on La Plateforme for long-context agentic coding and coding workflows; served without Mistral modifications. Added to the catalog 2026-08-06 per the official model card and pricing page. Available through Mistral's global endpoint and EU regional endpoint; not yet available through the US regional endpoint specifically (the global endpoint still serves worldwide, including US callers). Canonical API id 'zai-glm-5-2' confirmed via the model card's primary 'Click to copy' badge (docs.mistral.ai/models/zai-glm-5-2). Re-checked 2026-08-24 via raw HTML source inspection: the model card's status badge still reads 'Public Preview' (yellow) — kept as preview, not GA (a summarized fetch of docs.mistral.ai/models/overview this pass mis-stated it as 'GA'; the raw page source is authoritative and contradicts that). Pricing re-confirmed unchanged against the pricing page's raw data widgets: input $1.4/M, cached input $0.14/M. Re-checked 2026-09-07: same pattern recurred (docs.mistral.ai/getting-started/models AI-summarized as 'GA') but a raw-HTML fetch of the model card itself shows the badge (data-badge-type=\"yellow\") still rendering 'Public Preview' — kept as preview. Pricing unchanged. Re-checked 2026-09-14 via raw-HTML fetch of docs.mistral.ai/models/zai-glm-5-2: badge still data-badge-type=\"yellow\" 'Public Preview' with an adjacent 'Third-party' badge — unchanged. Pricing re-confirmed against the raw data-prices widgets on mistral.ai/pricing/api: input $1.4/M, cached $0.14/M, output $4.4/M — all unchanged. Re-checked 2026-09-22 via raw curl'd HTML (not AI-summarized) of docs.mistral.ai/models/zai-glm-5-2: badge is still data-badge-type=\"yellow\" 'Public Preview', id badge still 'zai-glm-5-2' — unchanged. Re-checked 2026-09-28 via curl'd raw HTML of the same URL: badge is still data-badge-type=\"yellow\" 'Public Preview', id badge still 'zai-glm-5-2' — unchanged, no GA promotion. Pricing re-confirmed against mistral.ai/pricing/api's raw data-prices widgets: input $1.4/M, output $4.4/M — unchanged (cached rate still not shown on that pricing card). Note: Mistral added a sibling model, Z.ai GLM 5.3 (zai-glm-5-3), on 2026-09-15 (tracked as a new separate entry in this registry); GLM 5.2's own model card shows no deprecation banner or note that it has been superseded by 5.3, so both remain listed independently.","type":"chat","input_modalities":["text"],"output_modalities":["text"],"status":"preview","released":"2026-08-06","limits":{"context_window":1000000,"max_output":131072},"pricing":{"currency":"USD","input_per_mtok":1.4,"cached_input_per_mtok":0.14,"output_per_mtok":4.4},"capabilities":{"tools":true,"structured_output":true,"streaming":true,"batching":true},"api":{"endpoint":"https://api.mistral.ai/v1/chat/completions","protocol":"openai_compat","openai_compatible":true,"auth":{"type":"bearer","env_var":"MISTRAL_API_KEY"},"sdks":{"mistralai":{"package":"@mistralai/mistralai","model_id":"zai-glm-5-2"}},"notes":"Third-party model (Z.ai) served on Mistral's own api.mistral.ai infrastructure without modification. No 'latest' alias observed on the model card."},"sources":{"spec":"https://docs.mistral.ai/models/zai-glm-5-2","pricing":"https://mistral.ai/pricing/api"},"last_verified":"2026-09-28"},{"id":"zai-glm-5-3","provider":"mistral","family":"zai-glm","display_name":"Z.ai GLM 5.3","description":"Third-party open-source text model from Z.ai (GLM 5.3), hosted by Mistral on La Plateforme for long-context agentic and coding workflows; served without Mistral modifications. New model, first spotted on this pass (2026-09-22); not present as of the 2026-09-14 verification. Released 2026-09-15 per the model card's 'September 15, 2026' release text and confirmed as a 'New' badge on the mistral.ai/pricing/api pricing card. Canonical API id 'zai-glm-5-3' confirmed via curl'd raw HTML of the model card's primary 'Click to copy' badge (data-badge-type=\"outline\">zai-glm-5-3) at docs.mistral.ai/models/zai-glm-5-3 — not an AI-summarized fetch. Status badge confirmed via the same raw-HTML method: data-badge-type=\"yellow\" 'Public Preview' (same yellow/Public-Preview pattern as the sibling zai-glm-5-2 entry), not GA. Context window (1M tokens) and max output (128k) confirmed via raw page text. Pricing confirmed via the raw data-driven pricing card on mistral.ai/pricing/api (text extracted from the HTML, not summarized): input $1.4/M, output $4.4/M; cached input $0.14/M corroborated by the model card and independently by a web search of the same page, but not independently re-confirmed against the raw pricing-page HTML in this pass (the pricing card's visible text only showed input/output, not a cached rate) — flagged as slightly less certain than the input/output figures, though it exactly matches GLM 5.2's cached rate and is very likely correct. Released alongside GLM 5.2, which remains listed separately with no deprecation banner (see zai-glm-5-2 entry). No 'latest' alias observed on the model card, consistent with the zai-glm-5-2 pattern. Re-checked 2026-09-28 via curl'd raw HTML of docs.mistral.ai/models/zai-glm-5-3: badge is still data-badge-type=\"yellow\" 'Public Preview', id badge still 'zai-glm-5-3' — unchanged, no GA promotion. Pricing re-confirmed via mistral.ai/pricing/api's raw data-prices widgets: input $1.4/M, output $4.4/M — unchanged; cached input rate still not independently shown on the pricing card this pass.","type":"chat","input_modalities":["text"],"output_modalities":["text"],"status":"preview","released":"2026-09-15","limits":{"context_window":1000000,"max_output":131072},"pricing":{"currency":"USD","input_per_mtok":1.4,"cached_input_per_mtok":0.14,"output_per_mtok":4.4,"notes":"Cached input rate ($0.14/M) taken from the model card and matches the zai-glm-5-2 sibling's rate exactly, but was not independently re-confirmed against the raw pricing-page HTML this pass (only input/output appeared in that card's visible text)."},"capabilities":{"tools":true,"structured_output":true,"streaming":true,"batching":true},"api":{"endpoint":"https://api.mistral.ai/v1/chat/completions","protocol":"openai_compat","openai_compatible":true,"auth":{"type":"bearer","env_var":"MISTRAL_API_KEY"},"sdks":{"mistralai":{"package":"@mistralai/mistralai","model_id":"zai-glm-5-3"}},"notes":"Third-party model (Z.ai) served on Mistral's own api.mistral.ai infrastructure without modification. No 'latest' alias observed on the model card."},"sources":{"spec":"https://docs.mistral.ai/models/zai-glm-5-3","pricing":"https://mistral.ai/pricing/api"},"last_verified":"2026-09-28"},{"id":"codestral-2508","provider":"mistral","family":"codestral","display_name":"Codestral","description":"Code completion / fill-in-the-middle model. 128K context. Strong on HumanEval. Corrected 2026-07-27: context window is 128K (131072), not 256K.","type":"code","input_modalities":["text"],"output_modalities":["text"],"status":"ga","released":"2025-07-30","limits":{"context_window":131072,"max_output":8192},"pricing":{"currency":"USD","input_per_mtok":0.3,"output_per_mtok":0.9,"batch_input_per_mtok":0.15,"batch_output_per_mtok":0.45},"capabilities":{"tools":true,"json_mode":true,"streaming":true,"batching":true},"performance":{"intelligence_index":65,"benchmarks":{"humaneval":86}},"api":{"endpoint":"https://api.mistral.ai/v1/chat/completions","protocol":"openai_compat","openai_compatible":true,"auth":{"type":"bearer","env_var":"MISTRAL_API_KEY"},"sdks":{"mistralai":{"package":"@mistralai/mistralai","model_id":"codestral-latest"}}},"aggregator_ids":{"openrouter":"mistralai/codestral-2508"},"manual_tags":["best_for_coding"],"sources":{"spec":"https://docs.mistral.ai/models/model-cards/codestral-25-08","pricing":"https://mistral.ai/pricing/api"},"last_verified":"2026-09-28"},{"id":"ministral-14b-2512","provider":"mistral","family":"ministral","display_name":"Ministral 3 14B","description":"Compact 14B Ministral with vision. Open-weight edge model. Corrected 2026-07-27: canonical API id is 'ministral-14b-2512' (no '3' segment), per the official model card.","type":"chat","input_modalities":["text","image"],"output_modalities":["text"],"status":"ga","released":"2025-12-02","limits":{"context_window":262144,"max_output":8192},"pricing":{"currency":"USD","input_per_mtok":0.2,"output_per_mtok":0.2},"capabilities":{"tools":true,"json_mode":true,"streaming":true,"vision":true},"performance":{"intelligence_index":50},"api":{"endpoint":"https://api.mistral.ai/v1/chat/completions","protocol":"openai_compat","openai_compatible":true,"auth":{"type":"bearer","env_var":"MISTRAL_API_KEY"},"sdks":{"mistralai":{"package":"@mistralai/mistralai","model_id":"ministral-14b-latest"}}},"sources":{"spec":"https://docs.mistral.ai/models/model-cards/ministral-3-14b-25-12","pricing":"https://mistral.ai/pricing/api"},"last_verified":"2026-09-28"},{"id":"ministral-8b-2512","provider":"mistral","family":"ministral","display_name":"Ministral 3 8B","description":"Compact 8B Ministral with vision. Cheaper than 14B. Corrected 2026-07-27: canonical API id is 'ministral-8b-2512' (no '3' segment), per the official model card.","type":"chat","input_modalities":["text","image"],"output_modalities":["text"],"status":"ga","released":"2025-12-02","limits":{"context_window":262144,"max_output":8192},"pricing":{"currency":"USD","input_per_mtok":0.15,"output_per_mtok":0.15},"capabilities":{"tools":true,"json_mode":true,"streaming":true,"vision":true},"performance":{"intelligence_index":45},"api":{"endpoint":"https://api.mistral.ai/v1/chat/completions","protocol":"openai_compat","openai_compatible":true,"auth":{"type":"bearer","env_var":"MISTRAL_API_KEY"},"sdks":{"mistralai":{"package":"@mistralai/mistralai","model_id":"ministral-8b-latest"}}},"sources":{"spec":"https://docs.mistral.ai/models/model-cards/ministral-3-8b-25-12","pricing":"https://mistral.ai/pricing/api"},"last_verified":"2026-09-28"},{"id":"ministral-3b-2512","provider":"mistral","family":"ministral","display_name":"Ministral 3 3B","description":"Smallest Ministral. Edge / on-device friendly. Corrected 2026-07-27: canonical API id is 'ministral-3b-2512' (single '3b' segment), and context window is 256K, not 128K, per the official model card.","type":"chat","input_modalities":["text","image"],"output_modalities":["text"],"status":"ga","released":"2025-12-02","limits":{"context_window":262144,"max_output":8192},"pricing":{"currency":"USD","input_per_mtok":0.1,"output_per_mtok":0.1},"capabilities":{"tools":true,"streaming":true,"vision":true},"performance":{"intelligence_index":38},"api":{"endpoint":"https://api.mistral.ai/v1/chat/completions","protocol":"openai_compat","openai_compatible":true,"auth":{"type":"bearer","env_var":"MISTRAL_API_KEY"},"sdks":{"mistralai":{"package":"@mistralai/mistralai","model_id":"ministral-3b-latest"}}},"sources":{"spec":"https://docs.mistral.ai/models/model-cards/ministral-3-3b-25-12","pricing":"https://mistral.ai/pricing/api"},"last_verified":"2026-09-28"},{"id":"voxtral-mini-tts-2603","provider":"mistral","family":"voxtral","display_name":"Voxtral TTS","description":"Text-to-speech with voice cloning. Released March 2026. Corrected 2026-07-27: canonical API id is 'voxtral-mini-tts-2603' (includes 'mini'), not 'voxtral-tts-2603', per the official model card.","type":"tts","input_modalities":["text"],"output_modalities":["audio"],"status":"ga","released":"2026-03-23","limits":{"supported_audio_formats":["wav","mp3","ogg"],"max_input_chars":8000},"pricing":{"currency":"USD","per_million_characters":16,"notes":"Quoted as $0.016 / 1,000 characters."},"capabilities":{"streaming":true,"voice_cloning":true,"speed_control":true},"api":{"endpoint":"https://api.mistral.ai/v1/audio/speech","protocol":"openai_compat","openai_compatible":false,"auth":{"type":"bearer","env_var":"MISTRAL_API_KEY"},"sdks":{"mistralai":{"package":"@mistralai/mistralai","model_id":"voxtral-mini-tts-latest"}}},"sources":{"spec":"https://docs.mistral.ai/models/model-cards/voxtral-tts-26-03","pricing":"https://mistral.ai/pricing/api"},"last_verified":"2026-09-28"},{"id":"voxtral-mini-2602","provider":"mistral","family":"voxtral","display_name":"Voxtral Mini Transcribe 2","description":"Multilingual speech-to-text transcription. Cost-efficient. Corrected 2026-07-27: canonical API id is 'voxtral-mini-2602', not 'voxtral-mini-transcribe-2602', per the official model card.","type":"stt","input_modalities":["audio"],"output_modalities":["text"],"status":"ga","released":"2026-02-04","limits":{"max_audio_size_mb":50,"supported_audio_formats":["mp3","wav","flac","m4a","ogg"],"supported_languages":["multilingual"]},"pricing":{"currency":"USD","audio_input_per_minute":0.003},"capabilities":{"streaming":true,"word_timestamps":true,"speaker_diarization":true},"api":{"endpoint":"https://api.mistral.ai/v1/audio/transcriptions","protocol":"openai_compat","openai_compatible":true,"auth":{"type":"bearer","env_var":"MISTRAL_API_KEY"},"sdks":{"mistralai":{"package":"@mistralai/mistralai","model_id":"voxtral-mini-2602"}}},"sources":{"spec":"https://docs.mistral.ai/models/model-cards/voxtral-mini-transcribe-26-02","pricing":"https://mistral.ai/pricing/api","release_notes":"https://mistral.ai/news/voxtral-transcribe-2/"},"last_verified":"2026-09-28"},{"id":"voxtral-mini-transcribe-realtime-2602","provider":"mistral","family":"voxtral","display_name":"Voxtral Mini Transcribe Realtime 2","description":"Multilingual speech-to-text transcription optimized for live, low-latency streaming (vs. the batch-oriented Voxtral Mini Transcribe 2). Was missing from the registry — added 2026-07-20 during a routine audit.","type":"stt","input_modalities":["audio"],"output_modalities":["text"],"status":"ga","released":"2026-02-04","limits":{"max_audio_size_mb":50,"supported_audio_formats":["mp3","wav","flac","m4a","ogg"],"supported_languages":["multilingual"]},"pricing":{"currency":"USD","audio_input_per_minute":0.006},"capabilities":{"streaming":true,"word_timestamps":true},"api":{"endpoint":"https://api.mistral.ai/v1/audio/transcriptions","protocol":"openai_compat","openai_compatible":true,"auth":{"type":"bearer","env_var":"MISTRAL_API_KEY"},"sdks":{"mistralai":{"package":"@mistralai/mistralai","model_id":"voxtral-mini-transcribe-realtime-2602"}}},"sources":{"spec":"https://docs.mistral.ai/models/model-cards/voxtral-mini-transcribe-realtime-26-02","pricing":"https://mistral.ai/pricing/api"},"last_verified":"2026-09-28"},{"id":"voxtral-small-2507","provider":"mistral","family":"voxtral","display_name":"Voxtral Small","description":"Audio-aware multimodal LLM. Accepts audio + text input.","type":"audio_chat","input_modalities":["text","audio"],"output_modalities":["text"],"status":"ga","released":"2025-07-15","limits":{"context_window":32000,"max_output":8192},"pricing":{"currency":"USD","input_per_mtok":0.1,"output_per_mtok":0.4,"audio_input_per_minute":0.004,"notes":"Corrected 2026-08-24: official model card and mistral.ai/pricing/api both show output at $0.4/M tokens (data-prices priceUsd 0.4 in the pricing widget), not the previously listed $0.3/M. Input ($0.1/M) and audio input ($0.004/min) were already correct and unchanged. Re-confirmed 2026-09-14 via the raw data-prices widgets on mistral.ai/pricing/api: input $0.1/M, output $0.4/M, audio input $0.004/min — all unchanged."},"capabilities":{"tools":true,"streaming":true,"audio_input":true},"api":{"endpoint":"https://api.mistral.ai/v1/chat/completions","protocol":"openai_compat","openai_compatible":true,"auth":{"type":"bearer","env_var":"MISTRAL_API_KEY"},"sdks":{"mistralai":{"package":"@mistralai/mistralai","model_id":"voxtral-small-2507"}}},"sources":{"spec":"https://docs.mistral.ai/models/model-cards/voxtral-small-25-07","pricing":"https://mistral.ai/pricing/api"},"last_verified":"2026-09-28"},{"id":"mistral-embed-2312","provider":"mistral","family":"mistral-embed","display_name":"Mistral Embed","description":"Original Mistral embedding model. 1024 dimensions. Corrected 2026-08-10: canonical API id is 'mistral-embed-2312' per the model card's primary 'Click to copy' badge (docs.mistral.ai/models/model-cards/mistral-embed-23-12); 'mistral-embed' is listed as a secondary alias on the same page, same pattern as every other dated-vs-'latest' pair in this file. Previously this registry used the bare alias as the canonical id.","type":"embedding","input_modalities":["text"],"output_modalities":["embedding"],"status":"ga","released":"2023-12-11","limits":{"max_input_tokens":8192,"max_dimensions":1024},"pricing":{"currency":"USD","input_per_mtok":0.1,"batch_input_per_mtok":0.05},"capabilities":{"batching":true},"performance":{"benchmarks":{"mteb":62.5}},"api":{"endpoint":"https://api.mistral.ai/v1/embeddings","protocol":"openai_compat","openai_compatible":true,"auth":{"type":"bearer","env_var":"MISTRAL_API_KEY"},"sdks":{"mistralai":{"package":"@mistralai/mistralai","model_id":"mistral-embed-2312"}}},"sources":{"spec":"https://docs.mistral.ai/models/model-cards/mistral-embed-23-12","pricing":"https://mistral.ai/pricing/api"},"last_verified":"2026-09-28"},{"id":"codestral-embed-2505","provider":"mistral","family":"codestral-embed","display_name":"Codestral Embed","description":"Code semantic embeddings. Optimized for code search and retrieval.","type":"embedding","input_modalities":["text","code"],"output_modalities":["embedding"],"status":"ga","released":"2025-05-28","limits":{"max_input_tokens":8192,"max_dimensions":1536},"pricing":{"currency":"USD","input_per_mtok":0.15,"batch_input_per_mtok":0.075},"capabilities":{"batching":true},"api":{"endpoint":"https://api.mistral.ai/v1/embeddings","protocol":"openai_compat","openai_compatible":true,"auth":{"type":"bearer","env_var":"MISTRAL_API_KEY"},"sdks":{"mistralai":{"package":"@mistralai/mistralai","model_id":"codestral-embed-2505"}}},"sources":{"spec":"https://docs.mistral.ai/models/model-cards/codestral-embed-25-05","pricing":"https://mistral.ai/pricing/api"},"last_verified":"2026-09-28"},{"id":"mistral-ocr-4-0","provider":"mistral","family":"mistral-ocr","display_name":"Mistral OCR 4","description":"Document AI service: OCR 4 with enhanced layout extraction, structured outputs, and improved accuracy. Released June 23, 2026. Replaces mistral-ocr-2512. Canonical API id is 'mistral-ocr-4-0' (corrected 2026-07-27, not 'mistral-ocr-4-2606'). Updated 2026-08-17: Mistral released OCR 4.1 (mistral-ocr-4-1, tracked separately in this registry) on 2026-07-16 and reassigned the 'mistral-ocr-4' and 'mistral-ocr-latest' aliases to point to it — this model's own model-card page no longer lists those as aliases of mistral-ocr-4-0. Still 'GA' status (not deprecated) per docs.mistral.ai/models/overview as of this pass, so kept as a distinct still-supported entry, but the SDK model_id below was changed from the now-reassigned 'mistral-ocr-latest' to the explicit dated id to avoid silently resolving to the wrong model. Re-checked 2026-08-24 via raw HTML source inspection of the model card: status badge still reads 'GA', no deprecation banner rendered — unchanged. Re-checked 2026-09-07, same method: still 'GA', no deprecation banner, unaffected by OCR 4.1's promotion to GA the same week. Re-checked 2026-09-14, same raw-HTML method against docs.mistral.ai/models/model-cards/ocr-4-0: badge is still 'GA' (data-badge-type=\"outline\"), id badge still reads 'mistral-ocr-4-0', no deprecation banner — unchanged. Re-checked 2026-09-28 via curl'd raw HTML of the same model-card URL: badge is still 'GA' (data-badge-type=\"outline\"), id badge still 'mistral-ocr-4-0', no instantiated deprecation/retirement banner found. New watch item this pass: mistral-ocr-4-0 has disappeared from mistral.ai/pricing/api's rendered pricing rows (only an 'OCR 4.1' pricing card remains; no separate 'OCR 4.0' row) — the same pattern previously seen with OCR 3 and with the already-deprecated Devstral Small 2 being quietly dropped from the sell page. Unlike OCR 3, this is a still-GA model with no deprecation signal on its own model card, so this is flagged as a watch item rather than a status change; no new official price was found on the pricing page to re-confirm this pass, so the existing $4/$5-per-1,000-pages figures are left as last confirmed 2026-07-27, not re-verified.","type":"vision","input_modalities":["image"],"output_modalities":["text"],"status":"ga","released":"2026-06-23","pricing":{"currency":"USD","notes":"$4 per 1,000 pages (standard OCR). $5 per 1,000 annotated pages (Document AI/structured extraction), confirmed 2026-07-27 via the official model card. Batch discount (50% off, per Mistral's general batch policy) previously noted as ~$2/1,000 pages but not independently reconfirmed this pass."},"capabilities":{"structured_output":true,"batching":true,"vision":true,"document_input":true},"api":{"endpoint":"https://api.mistral.ai/v1/ocr","protocol":"openai_compat","openai_compatible":false,"auth":{"type":"bearer","env_var":"MISTRAL_API_KEY"},"sdks":{"mistralai":{"package":"@mistralai/mistralai","model_id":"mistral-ocr-4-0"}}},"sources":{"spec":"https://docs.mistral.ai/models/model-cards/ocr-4-0","pricing":"https://mistral.ai/pricing/api"},"last_verified":"2026-09-28"},{"id":"mistral-ocr-4-1","provider":"mistral","family":"mistral-ocr","display_name":"Mistral OCR 4.1","description":"Document AI service: OCR 4.1, the newest and current 'mistral-ocr-latest' / 'mistral-ocr-4' model — adds additional confidence-score granularity options to the OCR API on top of OCR 4.0's layout extraction and structured outputs. Released 2026-07-16 per docs.mistral.ai/resources/changelogs. Added to this registry 2026-08-17 (new since the 2026-08-10 pass). Canonical API id 'mistral-ocr-4-1' confirmed via the model card's primary 'Click to copy' badge; aliases 'mistral-ocr-4' and 'mistral-ocr-latest' confirmed via the same page's alias list. Status badge on the model card read 'Public Preview' through 2026-08-24. **Flipped to GA 2026-09-07**: the official changelog (docs.mistral.ai/resources/changelogs) carries an August 31, 2026 entry reading \"OCR 4.1 (mistral-ocr-4-1) is now Generally Available\", and a raw-HTML fetch of the model card the same pass shows the status badge now rendering 'GA' (not 'Public Preview') — two independent official confirmations. Re-checked 2026-09-14: same raw-HTML badge scan of docs.mistral.ai/models/model-cards/ocr-4-1 shows data-badge-type=\"outline\" 'GA' still, id badge still 'mistral-ocr-4-1' — unchanged. Changelog's latest entry is still the same 2026-08-31 GA entry; no entries have posted since. Pricing re-confirmed unchanged via the raw data-prices widgets ($4/$5 per 1,000 pages).","type":"vision","input_modalities":["image"],"output_modalities":["text"],"status":"ga","released":"2026-07-16","pricing":{"currency":"USD","notes":"$4 per 1,000 pages (standard OCR). $5 per 1,000 annotated pages (Document AI/structured extraction). Same rate as OCR 4.0, confirmed via both the model card and mistral.ai/pricing/api on 2026-08-17."},"capabilities":{"structured_output":true,"batching":true,"vision":true,"document_input":true},"api":{"endpoint":"https://api.mistral.ai/v1/ocr","protocol":"openai_compat","openai_compatible":false,"auth":{"type":"bearer","env_var":"MISTRAL_API_KEY"},"sdks":{"mistralai":{"package":"@mistralai/mistralai","model_id":"mistral-ocr-latest"}}},"sources":{"spec":"https://docs.mistral.ai/models/model-cards/ocr-4-1","pricing":"https://mistral.ai/pricing/api","release_notes":"https://docs.mistral.ai/resources/changelogs"},"last_verified":"2026-09-28"},{"id":"mistral-ocr-2512","provider":"mistral","family":"mistral-ocr","display_name":"Mistral OCR 3","description":"Document AI service: OCR, layout extraction, structured outputs. OCR 4 is the newer recommended model, but as of 2026-07-06 Mistral's official governance page lists OCR 3 as Active with no retirement date — it remains fully supported for existing integrations. (Corrected 2026-07-06: previously marked deprecated without an official sunset date; that was an error.) Re-checked 2026-09-14: legal.mistral.ai/ai-governance/models/mistral-ocr-3 still lists status 'Active', retirement date 'N/A' — unchanged, still not deprecated. It also still appears in the non-deprecated model grid on docs.mistral.ai/getting-started/models alongside OCR 4.0/4.1. New watch item this pass: mistral.ai/pricing/api's raw HTML no longer contains any 'OCR 3' pricing row at all (only OCR 4.1's $4/$5-per-1000-pages row remains) — same pattern previously seen with Devstral Small 2 (deprecated model quietly dropped from the sell page while still technically live). Since OCR 3 is still Active (not deprecated) per governance, this is flagged as a watch item rather than a status change; no new official price was found, so the existing $2/$3-per-1000-pages figures are left as last confirmed (2026-07-27), not re-verified this pass. Re-checked 2026-09-22 via curl'd raw HTML (not AI-summarized) of legal.mistral.ai/ai-governance/models/mistral-ocr-3: structured page data reads \"retirementDate\":[0,null],\"status\":[0,\"active\"] — still Active, no retirement date. Still absent from mistral.ai/pricing/api's rendered pricing rows and from the docs.mistral.ai/getting-started/models catalog grid. Status unchanged; watch item remains open. Re-checked 2026-09-28 via curl'd raw HTML of legal.mistral.ai/ai-governance/models/mistral-ocr-3: structured page data still reads \"retirementDate\":[0,null],\"status\":[0,\"active\"] — unchanged, still Active, no retirement date. Still absent from mistral.ai/pricing/api's rendered rows (re-confirmed this pass) and from the docs.mistral.ai/getting-started/models catalog listing. Status unchanged; watch item remains open. Also noted this pass: OCR 4.0 (mistral-ocr-4-0, tracked separately) has now also disappeared from the same pricing page's rendered rows — see that entry.","type":"vision","input_modalities":["image"],"output_modalities":["text"],"status":"ga","released":"2025-12-18","pricing":{"currency":"USD","notes":"Corrected 2026-07-27: official model card shows $2 per 1,000 pages (standard OCR) and $3 per 1,000 annotated pages (Document AI/structured extraction), not the previously listed $1 per 1,000 pages."},"capabilities":{"structured_output":true,"vision":true,"document_input":true},"api":{"endpoint":"https://api.mistral.ai/v1/ocr","protocol":"openai_compat","openai_compatible":false,"auth":{"type":"bearer","env_var":"MISTRAL_API_KEY"},"sdks":{"mistralai":{"package":"@mistralai/mistralai","model_id":"mistral-ocr-2512"}}},"sources":{"spec":"https://legal.mistral.ai/ai-governance/models/mistral-ocr-3","pricing":"https://docs.mistral.ai/models/model-cards/ocr-3-25-12"},"last_verified":"2026-09-28"},{"id":"labs-leanstral-1-5","provider":"mistral","family":"leanstral","display_name":"Leanstral 1.5","description":"Labs (experimental) model for Lean 4 formal theorem proving. 119B total / 6.5B active parameters (MoE), Apache 2.0 weights. Released June 30, 2026, superseding the earlier Leanstral (labs-leanstral-2603). Free during the Labs preview. Corrected 2026-08-10: canonical API id is 'labs-leanstral-1-5' (with the 'labs-' prefix), per the model card's primary 'Click to copy' badge (docs.mistral.ai/models/model-cards/leanstral-1-5) — the prior registry entry dropped the 'labs-' prefix ('leanstral-1.5'), incorrectly implying the prefix had been dropped for this release; it has not. Re-verified 2026-08-17: model card badge still 'labs-leanstral-1-5', status 'Public Preview', pricing still free — unchanged. Noting a pricing-page-only inconsistency (not actioned): the 'Leanstral' card on mistral.ai/pricing/api links to the leanstral-1-5 model-card page but its copy-to-clipboard id still reads the older 'labs-leanstral-2603' — appears to be a stale label on that specific pricing widget, since the model card itself (the more authoritative source per this registry's convention) consistently shows 'labs-leanstral-1-5'. Re-checked 2026-08-24: model card badge still 'labs-leanstral-1-5', status still 'Public Preview', pricing still free, no deprecation/retirement banner rendered on the page (checked via raw HTML) — unchanged. A third-party aggregator's August claim that Leanstral 1.5 'will be retired on September 30, 2026' was flagged then as unconfirmed since the model-card page itself showed no retirement banner. **Sunset date CONFIRMED official 2026-09-07**: a raw-HTML fetch of the official changelog (docs.mistral.ai/resources/changelogs, June 30, 2026 entry) reads \"We released Leanstral 1.5 (labs-leanstral-1-5) ... This model will be retired on September 30, 2026\" — the date was on Mistral's own changelog all along, just not surfaced on the model-card page itself (which still shows no banner and 'Public Preview' as of this pass, consistent with the retirement being 3+ weeks out, not yet in effect). No replacement_id announced. Independently corroborated via a web search of the same changelog page. sunset_on set accordingly; status left 'preview' (not yet 'deprecated') since no deprecation banner/date has been posted ahead of the retirement itself. Re-checked 2026-09-14 via raw-HTML fetch of docs.mistral.ai/models/model-cards/leanstral-1-5: badge is still data-badge-type=\"yellow\" 'Public Preview' (id badge still 'labs-leanstral-1-5'), no deprecation/retirement banner rendered, pricing still free — unchanged, now only 16 days from the confirmed 2026-09-30 sunset_on with no deprecation notice posted yet. The stale 'labs-leanstral-2603' id on the mistral.ai/pricing/api card also persists unchanged. Watch closely next pass — the retirement date is now imminent. Re-checked 2026-09-22 via curl'd raw HTML (not AI-summarized, to avoid the AI-summary Preview/GA mis-statement bug seen elsewhere in this registry) of docs.mistral.ai/models/model-cards/leanstral-1-5: badge is still data-badge-type=\"yellow\" 'Public Preview' (id badge still 'labs-leanstral-1-5'); searched the full page source for any 'deprecat'/'retir'/'sunset' occurrence outside the shared i18n string-template bundle (which always contains the generic 'This model is retired...' template regardless of this model's actual state) and found none instantiated — no banner, no structured status/sunset JSON field with a real value. Pricing still free. **sunset_on (2026-09-30) is now only 8 days away with still no deprecation banner posted** — status intentionally left 'preview' per registry convention (banner/date determines the status flip, not our own countdown), but this is now the single most urgent open watch item in this file. No replacement_id has been announced anywhere (changelog, model card, or news) as of this pass. Re-checked 2026-09-28 via curl'd raw HTML (not AI-summarized) of docs.mistral.ai/models/leanstral-1-5: badge is still data-badge-type=\"outline\" reading the id 'labs-leanstral-1-5' with an adjacent data-badge-type=\"yellow\" 'Public Preview' status badge; scanned the full page source for any instantiated deprecation/retirement banner referencing this model by name and found none — only the generic i18n string-template bundle (present on every model card regardless of status) contains the 'This model is deprecated/retired...' template text, not filled in with this model's name. No structured sunset/status JSON field with a real value was found either. Cross-checked the official changelog (docs.mistral.ai/resources/changelogs) raw HTML again: the June 30, 2026 entry still reads 'This model will be retired on September 30, 2026' (future tense; no new changelog entry has posted marking the retirement as having occurred, and the changelog's most recent entry overall is still 2026-08-31). Also checked mistral.ai/pricing/api's raw HTML: the Leanstral card still shows 'Free' pricing, and in this snapshot the card no longer displays any copyable model-id text at all (the previously-noted stale 'labs-leanstral-2603' label is no longer present). **sunset_on (2026-09-30) is now only 2 DAYS AWAY with still no deprecation/retirement banner posted anywhere** — status intentionally left 'preview' per registry convention (the banner/date on the model card determines the status flip, not our own countdown), but this is flagged as extremely urgent: the next check (on or after 2026-09-30) must re-verify immediately and flip to 'retired' the moment the banner or changelog confirms it actually happened, even off the normal weekly cadence if possible.","type":"reasoning","input_modalities":["text"],"output_modalities":["text"],"status":"preview","released":"2026-06-30","sunset_on":"2026-09-30","limits":{"context_window":256000},"pricing":{"currency":"USD","input_per_mtok":0,"output_per_mtok":0,"notes":"Free during the Labs experimental preview."},"capabilities":{"reasoning":true,"code_execution":true},"api":{"endpoint":"https://api.mistral.ai/v1/chat/completions","protocol":"openai_compat","openai_compatible":true,"auth":{"type":"bearer","env_var":"MISTRAL_API_KEY"},"sdks":{"mistralai":{"package":"@mistralai/mistralai","model_id":"labs-leanstral-1-5"}}},"sources":{"spec":"https://docs.mistral.ai/models/model-cards/leanstral-1-5","release_notes":"https://docs.mistral.ai/resources/changelogs","docs":"https://mistral.ai/news/leanstral-1-5/"},"last_verified":"2026-09-28"},{"id":"mistral-moderation-2603","provider":"mistral","family":"mistral-moderation","display_name":"Mistral Moderation 2","description":"Content moderation classifier with jailbreak detection.","type":"moderation","input_modalities":["text"],"output_modalities":["text"],"status":"ga","released":"2026-03-01","limits":{"context_window":131072},"pricing":{"currency":"USD","input_per_mtok":0,"notes":"Updated 2026-08-10: previously listed at $0.1/M input tokens with a known conflict against the model card's '$0' widget (flagged unresolved on 2026-07-27 and 2026-08-03). This pass, both mistral.ai/pricing/api (which now explicitly reads 'A classifier service for text content moderation. Free') and the model-card page agree it is free — resolving the prior conflict as a genuine pricing change rather than a page glitch. Set to $0 accordingly."},"api":{"endpoint":"https://api.mistral.ai/v1/moderations","protocol":"openai_compat","openai_compatible":true,"auth":{"type":"bearer","env_var":"MISTRAL_API_KEY"},"sdks":{"mistralai":{"package":"@mistralai/mistralai","model_id":"mistral-moderation-latest"}}},"sources":{"spec":"https://docs.mistral.ai/models/model-cards/mistral-moderation-26-03","pricing":"https://mistral.ai/pricing/api"},"last_verified":"2026-09-28"},{"id":"gpt-6-astra","provider":"openai","family":"gpt-6","display_name":"GPT-6 Astra","description":"OpenAI's new flagship model, GA September 3, 2026 (\"our most capable model, built for the hardest end-to-end work\"). 1.05M context (max 922K input tokens), 128k max output, reasoning effort levels low/medium/high/xhigh/max (no \"none\"/minimal tier, unlike the GPT-5.6 family). Knowledge cutoff Apr 30, 2026. First model to cross OpenAI's Preparedness Framework \"Critical\" cybersecurity capability threshold, so some cyber-related tool use is gated. Does not deprecate or replace the GPT-5.6 family (Sol/Terra/Luna), which remains GA as of this pass. Cross-verified via the official model spec page, OpenAI's own Deployment Safety Hub system card, and independent press coverage (CNBC, Fortune, 9to5Mac) before adding. Re-confirmed 2026-09-14: pricing/context/output unchanged against the live model page and pricing page. Moved the `smartest`/`best_for_coding` manual_tags pins here from gpt-5.5/gpt-5.5-pro, which had been left stale since Astra's Sept 3 launch (same class of bug as the Sept 7 Claude Fable 5.1 fix) — Astra is OpenAI's current flagship per its own launch messaging even though no official intelligence_index-comparable composite benchmark has been published for it yet. Third-party sources (OpenRouter, DataCamp, Artificial Analysis) cite GPQA Diamond 96.0% and a DeepSWE v1.1 74% score attributed to an official OpenAI blog post (openai.com/index/gpt-6-astra), but that page still 403s on direct fetch (re-tried 2026-09-22 with both default and Googlebot user agents, both blocked) and the fetchable Deployment Safety Hub system card does not itself contain a clean top-line capability table — so `performance` remains deliberately unset pending an independently re-fetchable official source (standing watch item, re-checked and still unresolved). Pricing/context/output reconfirmed unchanged against the live pricing page's raw JSON payload this pass.","type":"chat","input_modalities":["text","image"],"output_modalities":["text"],"status":"ga","released":"2026-09-03","limits":{"context_window":1050000,"max_output":128000},"pricing":{"currency":"USD","input_per_mtok":10,"cached_input_per_mtok":1,"cache_write_5m_per_mtok":12.5,"output_per_mtok":50,"batch_input_per_mtok":5,"batch_output_per_mtok":25,"notes":"Standard tier, short-context (≤272K input tokens) rates shown above. Long-context (>272K) Standard: $20.00/MTok input ($2.00 cached, $25.00 cache write) / $75.00/MTok output (batch: $10.00/$1.00/$12.50 write/$37.50 output). Fast mode (2x Standard price, up to 2x speed): short-context $20.00/MTok input ($2.00 cached, $25.00 write) / $100.00 output; long-context $40.00/$4.00/$50.00 write/$150.00 output. Flex tier pricing matches Batch tier ($5.00/$0.50/$6.25 write/$25.00 output short-context; $10.00/$1.00/$12.50 write/$37.50 output long-context). Regional data-residency processing adds a 10% surcharge (applies to models released on or after Mar 5, 2026)."},"capabilities":{"tools":true,"structured_output":true,"streaming":true,"batching":true,"prompt_caching":true,"system_prompt":true,"reasoning":true,"reasoning_modes":["low","medium","high","xhigh","max"],"vision":true,"computer_use":true,"web_search":true,"code_execution":true},"api":{"endpoint":"https://api.openai.com/v1/responses","protocol":"openai","openai_compatible":true,"auth":{"type":"bearer","env_var":"OPENAI_API_KEY"},"sdks":{"openai":{"package":"openai","model_id":"gpt-6-astra"},"vercel-ai":{"package":"@ai-sdk/openai","model_id":"gpt-6-astra","factory":"openai('gpt-6-astra')"}},"supported_params":["max_output_tokens","temperature","top_p","tools","tool_choice","response_format","stream","reasoning_effort"]},"manual_tags":["smartest","best_for_coding"],"sources":{"spec":"https://developers.openai.com/api/docs/models/gpt-6-astra","pricing":"https://developers.openai.com/api/docs/pricing","release_notes":"https://developers.openai.com/api/docs/changelog"},"last_verified":"2026-09-28"},{"id":"gpt-6-sol","provider":"openai","family":"gpt-6","display_name":"GPT-6 Sol","description":"New: OpenAI's mid-tier GPT-6 model, GA September 22, 2026 (\"built for complex coding and agentic workflows\"), released alongside gpt-6-luna beneath the Sept 3, 2026 flagship gpt-6-astra. 1.05M context (max 922K input tokens), 128k max output, reasoning effort levels none/low/medium(default)/high/xhigh/max (unlike gpt-6-astra, which drops the \"none\" tier and custom temperature/top_p/logprobs per its own launch notes). Knowledge cutoff Apr 20, 2026. Standard-tier pricing ($2/$10 per MTok) is roughly half gpt-5.6-sol's ($4/$20) and close to gpt-5.6-terra's promotional rate ($2/$12) -- but does not deprecate gpt-5.6-sol, gpt-5.6-terra, or gpt-5.6-luna, none of which carry a deprecation banner on their own spec pages or an entry on the official deprecations table as of this pass (all three remain GA per this registry's \"deprecated only on explicit official notice\" convention). Press (TechCrunch, MacRumors) reports GPT-6 Sol produces roughly half as many factual errors as GPT-5.6 Sol on OpenAI's internal evaluations; no official benchmark table has been published, so `performance` is left unset. A Sept 25, 2026 changelog entry notes a since-fixed image-encoding bug that had degraded image understanding (including computer-use workflows) on this model and gpt-6-luna between Sept 22-25. Confirmed via the official model spec page, the official changelog, and the pricing page's raw embedded JSON (Standard/Batch/Flex/Fast tiers, short- and long-context bands, cross-checked against the model page's own \">272K input tokens priced at 2x input/cache and 1.5x output\" formula).","type":"chat","input_modalities":["text","image"],"output_modalities":["text"],"status":"ga","released":"2026-09-22","limits":{"context_window":1050000,"max_output":128000},"pricing":{"currency":"USD","input_per_mtok":2,"cached_input_per_mtok":0.2,"cache_write_5m_per_mtok":2.5,"output_per_mtok":10,"batch_input_per_mtok":1,"batch_output_per_mtok":5,"notes":"Standard tier, short-context (≤272K input tokens) rates shown above. Long-context (>272K) Standard: $4.00/MTok input ($0.40 cached, $5.00 cache write) / $15.00/MTok output (batch: $2.00/$0.20/$2.50 write/$7.50 output). Fast mode (2x Standard price): short-context $4.00/MTok input ($0.40 cached, $5.00 write) / $20.00 output; long-context $8.00/MTok input ($0.80 cached, $10.00 write) / $30.00/MTok output. Flex tier pricing matches Batch tier. Regional data-residency processing adds a 10% surcharge; EU data residency is available only with Standard processing."},"capabilities":{"tools":true,"structured_output":true,"streaming":true,"batching":true,"prompt_caching":true,"system_prompt":true,"reasoning":true,"reasoning_modes":["none","low","medium","high","xhigh","max"],"reasoning_default":"medium","vision":true,"computer_use":true,"web_search":true,"code_execution":true},"api":{"endpoint":"https://api.openai.com/v1/responses","protocol":"openai","openai_compatible":true,"auth":{"type":"bearer","env_var":"OPENAI_API_KEY"},"sdks":{"openai":{"package":"openai","model_id":"gpt-6-sol"},"vercel-ai":{"package":"@ai-sdk/openai","model_id":"gpt-6-sol","factory":"openai('gpt-6-sol')"}},"supported_params":["max_output_tokens","temperature","top_p","tools","tool_choice","response_format","stream","reasoning_effort"],"notes":"Chat Completions supports function calling only with reasoning_effort set to none; use the Responses API for built-in tools and function calling at other effort levels."},"sources":{"spec":"https://developers.openai.com/api/docs/models/gpt-6-sol","pricing":"https://developers.openai.com/api/docs/pricing","release_notes":"https://developers.openai.com/api/docs/changelog"},"last_verified":"2026-09-28"},{"id":"gpt-6-luna","provider":"openai","family":"gpt-6","display_name":"GPT-6 Luna","description":"New: OpenAI's cheapest/highest-volume GPT-6 model, GA September 22, 2026 (\"our most efficient model for focused, high-volume tasks\"), released alongside gpt-6-sol beneath the Sept 3, 2026 flagship gpt-6-astra -- completing the three-tier GPT-6 lineup (Astra/Sol/Luna; there is no gpt-6-terra). 1.05M context (max 922K input tokens), 128k max output, reasoning effort levels none/low/medium(default)/high/xhigh/max. Knowledge cutoff May 18, 2026. Standard-tier pricing ($0.10/$0.50 per MTok) is roughly half gpt-5.6-luna's ($0.20/$1.20) -- but does not deprecate gpt-5.6-luna (or gpt-5.6-sol/gpt-5.6-terra), which carries no deprecation banner on its own spec page or an entry on the official deprecations table as of this pass, so kept GA per this registry's convention. No official benchmark table has been published for this model, so `performance` is left unset. A Sept 25, 2026 changelog entry notes a since-fixed image-encoding bug that had degraded image understanding (including computer-use workflows) on this model and gpt-6-sol between Sept 22-25. Confirmed via the official model spec page, the official changelog, and the pricing page's raw embedded JSON (Standard/Batch/Flex/Fast tiers, short- and long-context bands, cross-checked against the model page's own \">272K input tokens priced at 2x input/cache and 1.5x output\" formula).","type":"chat","input_modalities":["text","image"],"output_modalities":["text"],"status":"ga","released":"2026-09-22","limits":{"context_window":1050000,"max_output":128000},"pricing":{"currency":"USD","input_per_mtok":0.1,"cached_input_per_mtok":0.01,"cache_write_5m_per_mtok":0.125,"output_per_mtok":0.5,"batch_input_per_mtok":0.05,"batch_output_per_mtok":0.25,"notes":"Standard tier, short-context (≤272K input tokens) rates shown above. Long-context (>272K) Standard: $0.20/MTok input ($0.02 cached, $0.25 cache write) / $0.75/MTok output (batch: $0.10/$0.01/$0.125 write/$0.375 output). Fast mode (2x Standard price): short-context $0.20/MTok input ($0.02 cached, $0.25 write) / $1.00 output; long-context $0.40/MTok input ($0.04 cached, $0.50 write) / $1.50/MTok output. Flex tier pricing matches Batch tier. Regional data-residency processing adds a 10% surcharge; EU data residency is available only with Standard processing."},"capabilities":{"tools":true,"structured_output":true,"streaming":true,"batching":true,"prompt_caching":true,"system_prompt":true,"reasoning":true,"reasoning_modes":["none","low","medium","high","xhigh","max"],"reasoning_default":"medium","vision":true,"computer_use":true,"web_search":true,"code_execution":true},"api":{"endpoint":"https://api.openai.com/v1/responses","protocol":"openai","openai_compatible":true,"auth":{"type":"bearer","env_var":"OPENAI_API_KEY"},"sdks":{"openai":{"package":"openai","model_id":"gpt-6-luna"},"vercel-ai":{"package":"@ai-sdk/openai","model_id":"gpt-6-luna","factory":"openai('gpt-6-luna')"}},"supported_params":["max_output_tokens","temperature","top_p","tools","tool_choice","response_format","stream","reasoning_effort"],"notes":"Chat Completions supports function calling only with reasoning_effort set to none; use the Responses API for built-in tools and function calling at other effort levels."},"sources":{"spec":"https://developers.openai.com/api/docs/models/gpt-6-luna","pricing":"https://developers.openai.com/api/docs/pricing","release_notes":"https://developers.openai.com/api/docs/changelog"},"last_verified":"2026-09-28"},{"id":"gpt-5.6-sol","provider":"openai","family":"gpt-5.6","display_name":"GPT-5.6 Sol","description":"OpenAI's \"durable capability tier\" family (flagship tier), GA July 9, 2026. 1.05M context, 128k max output, reasoning effort levels (none/low/medium/high/xhigh/max). Knowledge cutoff Feb 16, 2026. Promotional price cut effective Aug 21, 2026 (20% lower input, 33% lower output), available at least through Nov 21, 2026. Re-confirmed 2026-09-07: still GA and still the deprecations table's default replacement target (unaffected by the Sept 3, 2026 gpt-6-astra launch, which sits above this tier rather than replacing it); Ultrafast mode (Cerebras-backed, previewed Aug 13) still has no published token pricing.","type":"chat","input_modalities":["text","image"],"output_modalities":["text"],"status":"ga","released":"2026-07-09","limits":{"context_window":1050000,"max_output":128000},"pricing":{"currency":"USD","input_per_mtok":4,"cached_input_per_mtok":0.4,"output_per_mtok":20,"batch_input_per_mtok":2,"batch_output_per_mtok":10,"notes":"Corrected 2026-08-24: per the official Aug 21, 2026 changelog entry and the live pricing page, Standard-tier (short-context, i.e. ≤272K input tokens) pricing dropped to $4.00/MTok input ($0.40 cached, $5.00 cache write) / $20.00/MTok output — a promotional cut (20% lower input, 33% lower output vs. the prior $5.00/$0.50/$30.00) confirmed available at least through Nov 21, 2026; batch pricing is $2.00/$0.20/$10.00 accordingly. Long-context (>272K input tokens) Standard pricing is $8.00/MTok input ($0.80 cached, $10.00 cache write) / $30.00/MTok output (batch: $4.00/$0.40/$15.00). Fast mode (2x list price, up to 2.5x speed; replaces old \"Priority Processing\" tier) is now also split by context length: short-context $8.00/MTok input ($0.80 cached, $10.00 cache write) / $40.00/MTok output; long-context $16.00/MTok input ($1.60 cached, $20.00 cache write) / $60.00/MTok output. Ultrafast mode announced 2026-08-13 (Cerebras-hardware-backed, up to 14x faster than Standard / ~750 output tokens/sec, same underlying model/intelligence/context): limited preview for select API customers only, no published token pricing as of 2026-08-24 — not reflected in the pricing fields above pending an official price."},"capabilities":{"tools":true,"parallel_tools":true,"json_mode":true,"structured_output":true,"streaming":true,"batching":true,"prompt_caching":true,"system_prompt":true,"log_probs":true,"seed":true,"reasoning":true,"reasoning_modes":["none","low","medium","high","xhigh","max"],"reasoning_default":"medium","vision":true,"computer_use":true,"web_search":true,"code_execution":true},"api":{"endpoint":"https://api.openai.com/v1/responses","protocol":"openai","openai_compatible":true,"auth":{"type":"bearer","env_var":"OPENAI_API_KEY"},"sdks":{"openai":{"package":"openai","model_id":"gpt-5.6-sol"},"vercel-ai":{"package":"@ai-sdk/openai","model_id":"gpt-5.6-sol","factory":"openai('gpt-5.6-sol')"}},"supported_params":["max_output_tokens","temperature","top_p","tools","tool_choice","response_format","stream","logprobs","reasoning_effort"]},"sources":{"spec":"https://developers.openai.com/api/docs/models/gpt-5.6-sol","pricing":"https://developers.openai.com/api/docs/pricing","release_notes":"https://developers.openai.com/api/docs/changelog"},"last_verified":"2026-09-28"},{"id":"gpt-5.6-terra","provider":"openai","family":"gpt-5.6","display_name":"GPT-5.6 Terra","description":"OpenAI's \"durable capability tier\" family (mid tier), GA July 9, 2026. 1.05M context, 128k max output, reasoning effort levels (none/low/medium/high/xhigh/max). Knowledge cutoff Feb 16, 2026.","type":"chat","input_modalities":["text","image"],"output_modalities":["text"],"status":"ga","released":"2026-07-09","limits":{"context_window":1050000,"max_output":128000},"pricing":{"currency":"USD","input_per_mtok":2,"cached_input_per_mtok":0.2,"output_per_mtok":12,"batch_input_per_mtok":1,"batch_output_per_mtok":6,"notes":"Knowledge cutoff Feb 16, 2026. Standard-tier price cut 20% effective 2026-07-30 (was $2.50/$0.25/$15.00 input/cached/output). Batch cached input: $0.10/MTok. Re-confirmed 2026-08-24 against the live pricing page, unchanged: short-context (≤272K input tokens) Standard is $2.00/$0.20 cached/$2.50 cache write/$12.00 output; long-context (>272K) Standard is $4.00/$0.40 cached/$5.00 cache write/$18.00 output. Fast mode (2x price, up to 2.5x speed, added 2026-07-30): short-context $4.00/MTok input, $0.40/MTok cached input, $5.00/MTok cache write, $24.00/MTok output; long-context (extended 2026-08-05) $8.00/$0.80/$10.00 write/$36.00 output."},"capabilities":{"tools":true,"parallel_tools":true,"json_mode":true,"structured_output":true,"streaming":true,"batching":true,"prompt_caching":true,"system_prompt":true,"log_probs":true,"seed":true,"reasoning":true,"reasoning_modes":["none","low","medium","high","xhigh","max"],"reasoning_default":"medium","vision":true,"computer_use":true,"web_search":true,"code_execution":true},"api":{"endpoint":"https://api.openai.com/v1/responses","protocol":"openai","openai_compatible":true,"auth":{"type":"bearer","env_var":"OPENAI_API_KEY"},"sdks":{"openai":{"package":"openai","model_id":"gpt-5.6-terra"},"vercel-ai":{"package":"@ai-sdk/openai","model_id":"gpt-5.6-terra","factory":"openai('gpt-5.6-terra')"}},"supported_params":["max_output_tokens","temperature","top_p","tools","tool_choice","response_format","stream","logprobs","reasoning_effort"]},"sources":{"spec":"https://developers.openai.com/api/docs/models/gpt-5.6-terra","pricing":"https://developers.openai.com/api/docs/pricing","release_notes":"https://developers.openai.com/api/docs/changelog"},"last_verified":"2026-09-28"},{"id":"gpt-5.6-luna","provider":"openai","family":"gpt-5.6","display_name":"GPT-5.6 Luna","description":"OpenAI's \"durable capability tier\" family (fastest/cheapest tier), GA July 9, 2026. 1.05M context, 128k max output, reasoning effort levels (none/low/medium/high/xhigh/max). Knowledge cutoff Feb 16, 2026.","type":"chat","input_modalities":["text","image"],"output_modalities":["text"],"status":"ga","released":"2026-07-09","limits":{"context_window":1050000,"max_output":128000},"pricing":{"currency":"USD","input_per_mtok":0.2,"cached_input_per_mtok":0.02,"output_per_mtok":1.2,"batch_input_per_mtok":0.1,"batch_output_per_mtok":0.6,"notes":"Knowledge cutoff Feb 16, 2026. Standard-tier price cut 80% effective 2026-07-30 (was $1.00/$0.10/$6.00 input/cached/output). Batch cached input: $0.01/MTok. Re-confirmed 2026-08-24 against the live pricing page, unchanged: short-context (≤272K input tokens) Standard is $0.20/$0.02 cached/$0.25 cache write/$1.20 output; long-context (>272K) Standard is $0.40/$0.04 cached/$0.50 cache write/$1.80 output. Fast mode (2x price, up to 2.5x speed, added 2026-07-30): short-context $0.40/MTok input, $0.04/MTok cached input, $0.50/MTok cache write, $2.40/MTok output; long-context (extended 2026-08-05) $0.80/$0.08/$1.00 write/$3.60 output."},"capabilities":{"tools":true,"parallel_tools":true,"json_mode":true,"structured_output":true,"streaming":true,"batching":true,"prompt_caching":true,"system_prompt":true,"log_probs":true,"seed":true,"reasoning":true,"reasoning_modes":["none","low","medium","high","xhigh","max"],"reasoning_default":"medium","vision":true,"computer_use":true,"web_search":true,"code_execution":true},"api":{"endpoint":"https://api.openai.com/v1/responses","protocol":"openai","openai_compatible":true,"auth":{"type":"bearer","env_var":"OPENAI_API_KEY"},"sdks":{"openai":{"package":"openai","model_id":"gpt-5.6-luna"},"vercel-ai":{"package":"@ai-sdk/openai","model_id":"gpt-5.6-luna","factory":"openai('gpt-5.6-luna')"}},"supported_params":["max_output_tokens","temperature","top_p","tools","tool_choice","response_format","stream","logprobs","reasoning_effort"]},"sources":{"spec":"https://developers.openai.com/api/docs/models/gpt-5.6-luna","pricing":"https://developers.openai.com/api/docs/pricing","release_notes":"https://developers.openai.com/api/docs/changelog"},"last_verified":"2026-09-28"},{"id":"gpt-5.5","provider":"openai","family":"gpt-5.5","display_name":"GPT-5.5","description":"OpenAI's frontier model, released April 24, 2026. 1.05M context, 128k max output, reasoning effort levels (none/low/medium/high/xhigh). Current snapshot: gpt-5.5-2026-04-23. Updated 2026-09-14: `smartest`/`best_for_coding` manual_tags pins moved to gpt-6-astra (GA Sept 3, 2026), OpenAI's current flagship per its own launch messaging — still GA and still a strong option, but no longer the top-of-lineup pin.","type":"chat","input_modalities":["text","image"],"output_modalities":["text"],"status":"ga","released":"2026-04-24","limits":{"context_window":1050000,"max_output":128000},"pricing":{"currency":"USD","input_per_mtok":5,"cached_input_per_mtok":0.5,"output_per_mtok":30,"batch_input_per_mtok":2.5,"batch_output_per_mtok":15,"notes":"Knowledge cutoff Dec 1, 2025. Regional data residency adds 10% surcharge."},"capabilities":{"tools":true,"parallel_tools":true,"json_mode":true,"structured_output":true,"streaming":true,"batching":true,"prompt_caching":true,"system_prompt":true,"log_probs":true,"seed":true,"reasoning":true,"reasoning_modes":["none","low","medium","high","xhigh"],"reasoning_default":"medium","vision":true,"computer_use":true,"web_search":true,"code_execution":true},"performance":{"intelligence_index":88},"api":{"endpoint":"https://api.openai.com/v1/responses","protocol":"openai","openai_compatible":true,"auth":{"type":"bearer","env_var":"OPENAI_API_KEY"},"sdks":{"openai":{"package":"openai","model_id":"gpt-5.5"},"vercel-ai":{"package":"@ai-sdk/openai","model_id":"gpt-5.5","factory":"openai('gpt-5.5')"}},"supported_params":["max_output_tokens","temperature","top_p","tools","tool_choice","response_format","stream","logprobs","reasoning_effort"]},"aggregator_ids":{"openrouter":"openai/gpt-5.5","azure":"gpt-5.5"},"sources":{"spec":"https://developers.openai.com/api/docs/models/gpt-5.5","pricing":"https://developers.openai.com/api/docs/pricing","release_notes":"https://openai.com/index/introducing-gpt-5-5/"},"last_verified":"2026-09-28"},{"id":"gpt-5.5-pro","provider":"openai","family":"gpt-5.5","display_name":"GPT-5.5 Pro","description":"Highest-quality variant of GPT-5.5 for the most complex reasoning tasks. Premium pricing with no caching discount.","type":"reasoning","input_modalities":["text","image"],"output_modalities":["text"],"status":"ga","released":"2026-04-24","limits":{"context_window":400000,"max_output":128000},"pricing":{"currency":"USD","input_per_mtok":30,"output_per_mtok":180,"batch_input_per_mtok":15,"batch_output_per_mtok":90},"capabilities":{"tools":true,"parallel_tools":true,"structured_output":true,"streaming":true,"batching":true,"system_prompt":true,"reasoning":true,"reasoning_modes":["low","medium","high","xhigh"],"vision":true},"performance":{"intelligence_index":90},"api":{"endpoint":"https://api.openai.com/v1/responses","protocol":"openai","openai_compatible":true,"auth":{"type":"bearer","env_var":"OPENAI_API_KEY"},"sdks":{"openai":{"package":"openai","model_id":"gpt-5.5-pro"}}},"aggregator_ids":{"openrouter":"openai/gpt-5.5-pro"},"sources":{"pricing":"https://developers.openai.com/api/docs/pricing"},"last_verified":"2026-09-28"},{"id":"gpt-5.4","provider":"openai","family":"gpt-5.4","display_name":"GPT-5.4","description":"Frontier model from before GPT-5.5. 1M context, 128k output. Knowledge cutoff Aug 31, 2025. Lower-cost frontier option versus gpt-5.5.","type":"chat","input_modalities":["text","image"],"output_modalities":["text"],"status":"ga","limits":{"context_window":1000000,"max_output":128000},"pricing":{"currency":"USD","input_per_mtok":2.5,"cached_input_per_mtok":0.25,"output_per_mtok":15,"batch_input_per_mtok":1.25,"batch_output_per_mtok":7.5},"capabilities":{"tools":true,"parallel_tools":true,"json_mode":true,"structured_output":true,"streaming":true,"batching":true,"prompt_caching":true,"system_prompt":true,"log_probs":true,"seed":true,"reasoning":true,"reasoning_modes":["none","low","medium","high"],"vision":true,"web_search":true,"code_execution":true},"performance":{"intelligence_index":80},"api":{"endpoint":"https://api.openai.com/v1/responses","protocol":"openai","openai_compatible":true,"auth":{"type":"bearer","env_var":"OPENAI_API_KEY"},"sdks":{"openai":{"package":"openai","model_id":"gpt-5.4"}}},"aggregator_ids":{"openrouter":"openai/gpt-5.4"},"sources":{"pricing":"https://developers.openai.com/api/docs/pricing"},"last_verified":"2026-09-28"},{"id":"gpt-5.4-mini","provider":"openai","family":"gpt-5.4","display_name":"GPT-5.4 Mini","description":"Smaller, faster variant of GPT-5.4. Designed for routed production traffic with strong cost-performance.","type":"chat","input_modalities":["text","image"],"output_modalities":["text"],"status":"ga","limits":{"context_window":400000,"max_output":128000},"pricing":{"currency":"USD","input_per_mtok":0.75,"cached_input_per_mtok":0.075,"output_per_mtok":4.5,"batch_input_per_mtok":0.375,"batch_output_per_mtok":2.25},"capabilities":{"tools":true,"parallel_tools":true,"json_mode":true,"structured_output":true,"streaming":true,"batching":true,"prompt_caching":true,"reasoning":true,"reasoning_modes":["none","low","medium","high"],"vision":true},"performance":{"speed_tps":130,"intelligence_index":70},"api":{"endpoint":"https://api.openai.com/v1/responses","protocol":"openai","openai_compatible":true,"auth":{"type":"bearer","env_var":"OPENAI_API_KEY"},"sdks":{"openai":{"package":"openai","model_id":"gpt-5.4-mini"}}},"aggregator_ids":{"openrouter":"openai/gpt-5.4-mini"},"sources":{"pricing":"https://developers.openai.com/api/docs/pricing"},"last_verified":"2026-09-28"},{"id":"gpt-5.4-nano","provider":"openai","family":"gpt-5.4","display_name":"GPT-5.4 Nano","description":"Smallest, cheapest GPT-5.4 variant. Optimized for high-throughput, latency-sensitive workloads.","type":"chat","input_modalities":["text","image"],"output_modalities":["text"],"status":"ga","limits":{"context_window":400000,"max_output":128000},"pricing":{"currency":"USD","input_per_mtok":0.2,"cached_input_per_mtok":0.02,"output_per_mtok":1.25,"batch_input_per_mtok":0.1,"batch_output_per_mtok":0.625},"capabilities":{"tools":true,"json_mode":true,"structured_output":true,"streaming":true,"batching":true,"prompt_caching":true,"vision":true},"performance":{"speed_tps":200,"intelligence_index":55},"api":{"endpoint":"https://api.openai.com/v1/responses","protocol":"openai","openai_compatible":true,"auth":{"type":"bearer","env_var":"OPENAI_API_KEY"},"sdks":{"openai":{"package":"openai","model_id":"gpt-5.4-nano"}}},"aggregator_ids":{"openrouter":"openai/gpt-5.4-nano"},"sources":{"pricing":"https://developers.openai.com/api/docs/pricing"},"last_verified":"2026-09-28"},{"id":"gpt-5.4-pro","provider":"openai","family":"gpt-5.4","display_name":"GPT-5.4 Pro","description":"Pro-tier reasoning variant of GPT-5.4 with extended thinking. Premium pricing.","type":"reasoning","input_modalities":["text","image"],"output_modalities":["text"],"status":"ga","limits":{"context_window":400000,"max_output":128000},"pricing":{"currency":"USD","input_per_mtok":30,"output_per_mtok":180,"batch_input_per_mtok":15,"batch_output_per_mtok":90},"capabilities":{"tools":true,"structured_output":true,"streaming":true,"batching":true,"system_prompt":true,"reasoning":true,"reasoning_modes":["low","medium","high","xhigh"],"vision":true},"performance":{"intelligence_index":85},"api":{"endpoint":"https://api.openai.com/v1/responses","protocol":"openai","openai_compatible":true,"auth":{"type":"bearer","env_var":"OPENAI_API_KEY"},"sdks":{"openai":{"package":"openai","model_id":"gpt-5.4-pro"}}},"aggregator_ids":{"openrouter":"openai/gpt-5.4-pro"},"sources":{"pricing":"https://developers.openai.com/api/docs/pricing"},"last_verified":"2026-09-28"},{"id":"gpt-5-pro","provider":"openai","family":"gpt-5","display_name":"GPT-5 Pro","description":"Added 2026-08-24: a previously-flagged coverage gap, now confirmed as a live, generally-available model with its own official spec page (not just a pricing-table row). Highest-quality GPT-5 variant, using extra compute for harder problems; defaults to and only supports reasoning.effort: high. Responses API only (no Chat Completions/Realtime); does not support code interpreter. Default snapshot gpt-5-pro-2025-10-06 — that dated snapshot is listed on the deprecations table shutting down Dec 11, 2026 (replacement gpt-5.6-sol with reasoning.mode: pro), but the bare `gpt-5-pro` id itself carries no deprecation notice.","type":"reasoning","input_modalities":["text","image"],"output_modalities":["text"],"status":"ga","limits":{"context_window":400000,"max_output":272000},"pricing":{"currency":"USD","input_per_mtok":15,"output_per_mtok":120,"batch_input_per_mtok":7.5,"batch_output_per_mtok":60,"notes":"No cached-input discount published for this model (pricing page shows '-' for cached input). Knowledge cutoff Sep 30, 2024."},"capabilities":{"tools":true,"structured_output":true,"streaming":true,"batching":true,"system_prompt":true,"reasoning":true,"reasoning_modes":["high"],"reasoning_default":"high","vision":true,"web_search":true,"code_execution":false},"api":{"endpoint":"https://api.openai.com/v1/responses","protocol":"openai","openai_compatible":true,"auth":{"type":"bearer","env_var":"OPENAI_API_KEY"},"sdks":{"openai":{"package":"openai","model_id":"gpt-5-pro"}}},"sources":{"spec":"https://developers.openai.com/api/docs/models/gpt-5-pro","pricing":"https://developers.openai.com/api/docs/pricing","release_notes":"https://developers.openai.com/api/docs/deprecations"},"last_verified":"2026-09-28"},{"id":"gpt-5.1","provider":"openai","family":"gpt-5.1","display_name":"GPT-5.1","description":"Added 2026-09-22: a previously-undiscovered coverage gap. Point release between gpt-5 and gpt-5.2, released Nov 13, 2025 (single snapshot gpt-5.1-2025-11-13). Its own model page still markets it as OpenAI's flagship for coding/agentic tasks, but that appears to be stale copy predating gpt-5.2/gpt-5.4+; gpt-6-astra is the actual current flagship. Not listed under its bare id on the official deprecations table (only the derived gpt-5.1-chat-latest/-codex/-codex-max snapshots are, shut down Jul 23, 2026) and carries no deprecated banner on its own page, so tracked as GA per this registry's established convention (deprecated only on explicit official notice). Confirmed via the official model spec page and the pricing page's full model table, which requires a 'show more/all models' toggle to reveal older-generation rows beyond the default collapsed view — the likely reason this and gpt-5.2/gpt-5.2-pro were missed across 7 consecutive prior research passes. Cross-checked against independent third-party pricing trackers. Some capability flags (parallel_tools, json_mode, log_probs, seed, web_search, code_execution) could not be confirmed from the fetched page and are deliberately left unset rather than assumed from the sibling gpt-5 entry.","type":"chat","input_modalities":["text","image"],"output_modalities":["text"],"status":"ga","released":"2025-11-13","limits":{"context_window":400000,"max_output":128000},"pricing":{"currency":"USD","input_per_mtok":1.25,"cached_input_per_mtok":0.125,"output_per_mtok":10,"batch_input_per_mtok":0.625,"batch_output_per_mtok":5,"notes":"Knowledge cutoff Sep 30, 2024. Fast mode (2x standard price, per the live pricing page): $2.50/MTok input ($0.25 cached), $20.00/MTok output."},"capabilities":{"tools":true,"structured_output":true,"streaming":true,"batching":true,"prompt_caching":true,"system_prompt":true,"reasoning":true,"reasoning_modes":["none","low","medium","high"],"reasoning_default":"none","vision":true},"api":{"endpoint":"https://api.openai.com/v1/responses","protocol":"openai","openai_compatible":true,"auth":{"type":"bearer","env_var":"OPENAI_API_KEY"},"sdks":{"openai":{"package":"openai","model_id":"gpt-5.1"}}},"sources":{"spec":"https://developers.openai.com/api/docs/models/gpt-5.1","pricing":"https://developers.openai.com/api/docs/pricing"},"last_verified":"2026-09-28"},{"id":"gpt-5.2","provider":"openai","family":"gpt-5.2","display_name":"GPT-5.2","description":"Added 2026-09-22: a previously-undiscovered coverage gap. Point release after gpt-5.1, released Dec 11, 2025 (single snapshot gpt-5.2-2025-12-11). Its own model page explicitly frames it as OpenAI's 'previous flagship model for complex professional work' and recommends gpt-6-astra for new use — but carries no deprecated banner and is absent from the official deprecations table under its bare id (only the derived gpt-5.2-chat-latest/-codex snapshots are, both already shut down), so tracked as GA per this registry's convention (deprecated only on explicit official notice), consistent with how gpt-5.4/gpt-5.5 stay GA despite being superseded by later tiers. Confirmed via the official model spec page and the pricing page's full model table (hidden behind a 'show all models' toggle by default). Some capability flags (parallel_tools, json_mode, log_probs, seed, web_search, code_execution) could not be confirmed from the fetched page and are deliberately left unset.","type":"chat","input_modalities":["text","image"],"output_modalities":["text"],"status":"ga","released":"2025-12-11","limits":{"context_window":400000,"max_output":128000},"pricing":{"currency":"USD","input_per_mtok":1.75,"cached_input_per_mtok":0.175,"output_per_mtok":14,"batch_input_per_mtok":0.875,"batch_output_per_mtok":7,"notes":"Knowledge cutoff Aug 31, 2025. Fast mode (2x standard price, per the live pricing page): $3.50/MTok input ($0.35 cached), $28.00/MTok output."},"capabilities":{"tools":true,"structured_output":true,"streaming":true,"batching":true,"prompt_caching":true,"system_prompt":true,"reasoning":true,"reasoning_modes":["none","low","medium","high","xhigh"],"reasoning_default":"none","vision":true},"api":{"endpoint":"https://api.openai.com/v1/responses","protocol":"openai","openai_compatible":true,"auth":{"type":"bearer","env_var":"OPENAI_API_KEY"},"sdks":{"openai":{"package":"openai","model_id":"gpt-5.2"}}},"sources":{"spec":"https://developers.openai.com/api/docs/models/gpt-5.2","pricing":"https://developers.openai.com/api/docs/pricing"},"last_verified":"2026-09-28"},{"id":"gpt-5.2-pro","provider":"openai","family":"gpt-5.2","display_name":"GPT-5.2 Pro","description":"Added 2026-09-22: a previously-undiscovered coverage gap, found alongside gpt-5.1/gpt-5.2. Higher-compute pro variant of gpt-5.2, released Dec 11, 2025 (single snapshot gpt-5.2-pro-2025-12-11). Its own model page frames it as OpenAI's 'previous pro model for complex professional work' and recommends gpt-5.5-pro for the latest pro tier, but carries no deprecated banner and is absent from the deprecations table under its bare id, so tracked as GA per this registry's convention. Responses API only (no Chat Completions/Realtime/Assistants) — 'available in the Responses API only to enable support for multi-turn model interactions... requests may take several minutes to finish, try using background mode' per its official page, same pattern as gpt-5-pro/gpt-5.4-pro/o3-pro. Structured outputs explicitly NOT supported on this model per its official Features table (unlike the non-pro gpt-5.2). Confirmed via the official model spec page and the pricing page's full model table.","type":"reasoning","input_modalities":["text","image"],"output_modalities":["text"],"status":"ga","released":"2025-12-11","limits":{"context_window":400000,"max_output":128000},"pricing":{"currency":"USD","input_per_mtok":21,"output_per_mtok":168,"batch_input_per_mtok":10.5,"batch_output_per_mtok":84,"notes":"No cached-input discount published for this model (pricing page shows '-' for cached input). Knowledge cutoff Aug 31, 2025."},"capabilities":{"tools":true,"streaming":true,"batching":true,"system_prompt":true,"reasoning":true,"reasoning_modes":["medium","high","xhigh"],"vision":true},"api":{"endpoint":"https://api.openai.com/v1/responses","protocol":"openai","openai_compatible":true,"auth":{"type":"bearer","env_var":"OPENAI_API_KEY"},"sdks":{"openai":{"package":"openai","model_id":"gpt-5.2-pro"}}},"sources":{"spec":"https://developers.openai.com/api/docs/models/gpt-5.2-pro","pricing":"https://developers.openai.com/api/docs/pricing"},"last_verified":"2026-09-28"},{"id":"gpt-5.3-codex","provider":"openai","family":"gpt-5-codex","display_name":"GPT-5.3 Codex","description":"Coding-specialized variant. Replaces gpt-5-codex (deprecating Jul 23, 2026).","type":"code","input_modalities":["text","image"],"output_modalities":["text"],"status":"ga","limits":{"context_window":400000,"max_output":128000},"pricing":{"currency":"USD","input_per_mtok":1.75,"cached_input_per_mtok":0.175,"output_per_mtok":14,"batch_input_per_mtok":0.875,"batch_output_per_mtok":7,"notes":"Fast mode (confirmed on the official pricing page 2026-08-17): $3.50/MTok input, $0.35/MTok cached input, $28.00/MTok output."},"capabilities":{"tools":true,"structured_output":true,"streaming":true,"batching":true,"prompt_caching":true,"system_prompt":true,"reasoning":true,"code_execution":true},"performance":{"intelligence_index":80},"api":{"endpoint":"https://api.openai.com/v1/responses","protocol":"openai","openai_compatible":true,"auth":{"type":"bearer","env_var":"OPENAI_API_KEY"},"sdks":{"openai":{"package":"openai","model_id":"gpt-5.3-codex"}}},"manual_tags":["best_for_coding"],"sources":{"pricing":"https://developers.openai.com/api/docs/pricing"},"last_verified":"2026-09-28"},{"id":"o3-pro","provider":"openai","family":"o-series","display_name":"OpenAI o3-pro","description":"Added 2026-08-24: a previously-flagged coverage gap, now confirmed as a live, generally-available model with its own official spec page. Uses more compute than o3 for consistently better answers on complex reasoning. Responses API only (no Chat Completions/Realtime); background mode recommended since requests may take several minutes. Default snapshot o3-pro-2025-06-10 — that dated snapshot is listed on the deprecations table shutting down Dec 11, 2026 (replacement gpt-5.6-sol with reasoning.mode: pro), but the bare `o3-pro` id itself carries no deprecation notice, distinct from bare `o3` which the registry already tracks as deprecated.","type":"reasoning","input_modalities":["text","image"],"output_modalities":["text"],"status":"ga","limits":{"context_window":200000,"max_output":100000},"pricing":{"currency":"USD","input_per_mtok":20,"output_per_mtok":80,"batch_input_per_mtok":10,"batch_output_per_mtok":40,"notes":"No cached-input discount published for this model (pricing page shows '-' for cached input). Knowledge cutoff Jun 1, 2024."},"capabilities":{"tools":true,"structured_output":true,"batching":true,"reasoning":true,"vision":true,"web_search":true},"api":{"endpoint":"https://api.openai.com/v1/responses","protocol":"openai","openai_compatible":true,"auth":{"type":"bearer","env_var":"OPENAI_API_KEY"},"sdks":{"openai":{"package":"openai","model_id":"o3-pro"}}},"sources":{"spec":"https://developers.openai.com/api/docs/models/o3-pro","pricing":"https://developers.openai.com/api/docs/pricing","release_notes":"https://developers.openai.com/api/docs/deprecations"},"last_verified":"2026-09-28"},{"id":"gpt-realtime-2.1","provider":"openai","family":"gpt-realtime","display_name":"GPT Realtime 2.1","description":"Incremental update to GPT Realtime 2 (released July 6, 2026): improved alphanumeric recognition, silence/noise handling, and interruption behavior.","type":"audio_chat","input_modalities":["text","audio","image"],"output_modalities":["text","audio"],"status":"ga","released":"2026-07-06","limits":{"context_window":128000,"max_output":32000},"pricing":{"currency":"USD","input_per_mtok":4,"cached_input_per_mtok":0.4,"output_per_mtok":24,"notes":"Audio: $32.0/MTok input ($0.4 cached), $64.0/MTok output. Image input: $5.0/MTok ($0.5 cached)."},"capabilities":{"tools":true,"streaming":true,"reasoning":true,"reasoning_modes":["minimal","low","medium","high","xhigh"],"reasoning_default":"low","vision":true,"audio_input":true,"audio_output":true},"api":{"endpoint":"https://api.openai.com/v1/realtime","protocol":"openai","openai_compatible":true,"auth":{"type":"bearer","env_var":"OPENAI_API_KEY"},"sdks":{"openai":{"package":"openai","model_id":"gpt-realtime-2.1"}}},"sources":{"spec":"https://developers.openai.com/api/docs/models/gpt-realtime-2.1","pricing":"https://developers.openai.com/api/docs/pricing","release_notes":"https://developers.openai.com/api/docs/changelog"},"last_verified":"2026-09-28"},{"id":"gpt-realtime-2.1-mini","provider":"openai","family":"gpt-realtime","display_name":"GPT Realtime 2.1 Mini","description":"Distilled/faster variant of GPT Realtime 2.1, released July 6, 2026.","type":"audio_chat","input_modalities":["text","audio","image"],"output_modalities":["text","audio"],"status":"ga","released":"2026-07-06","limits":{"context_window":128000,"max_output":32000},"pricing":{"currency":"USD","input_per_mtok":0.6,"cached_input_per_mtok":0.06,"output_per_mtok":2.4,"notes":"Audio: $10.0/MTok input ($0.3 cached), $20.0/MTok output. Image input: $0.8/MTok ($0.08 cached)."},"capabilities":{"tools":true,"streaming":true,"reasoning":true,"reasoning_modes":["minimal","low","medium","high","xhigh"],"reasoning_default":"low","vision":true,"audio_input":true,"audio_output":true},"api":{"endpoint":"https://api.openai.com/v1/realtime","protocol":"openai","openai_compatible":true,"auth":{"type":"bearer","env_var":"OPENAI_API_KEY"},"sdks":{"openai":{"package":"openai","model_id":"gpt-realtime-2.1-mini"}}},"sources":{"spec":"https://developers.openai.com/api/docs/models/gpt-realtime-2.1-mini","pricing":"https://developers.openai.com/api/docs/pricing","release_notes":"https://developers.openai.com/api/docs/changelog"},"last_verified":"2026-09-28"},{"id":"gpt-realtime-2","provider":"openai","family":"gpt-realtime","display_name":"GPT Realtime 2","description":"Real-time speech-in/speech-out model with reasoning. 128k context (up from 32k in 1.5). Reasoning effort levels: minimal, low, medium, high, xhigh.","type":"audio_chat","input_modalities":["text","audio","image"],"output_modalities":["text","audio"],"status":"ga","released":"2026-04-08","limits":{"context_window":128000,"max_output":32000},"pricing":{"currency":"USD","input_per_mtok":4,"cached_input_per_mtok":0.4,"output_per_mtok":24,"notes":"Audio: $32/MTok input ($0.40 cached), $64/MTok output. Text input: $4/MTok ($0.40 cached), output: $24/MTok."},"capabilities":{"tools":true,"streaming":true,"reasoning":true,"reasoning_modes":["minimal","low","medium","high","xhigh"],"reasoning_default":"low","vision":true,"audio_input":true,"audio_output":true},"api":{"endpoint":"https://api.openai.com/v1/realtime","protocol":"openai","openai_compatible":true,"auth":{"type":"bearer","env_var":"OPENAI_API_KEY"},"sdks":{"openai":{"package":"openai","model_id":"gpt-realtime-2"}}},"sources":{"spec":"https://developers.openai.com/api/docs/models/gpt-realtime-2","pricing":"https://developers.openai.com/api/docs/pricing"},"last_verified":"2026-09-28"},{"id":"gpt-realtime-1.5","provider":"openai","family":"gpt-realtime","display_name":"GPT Realtime 1.5","description":"Earlier-generation realtime speech model (32k context, predates the 128k-context gpt-realtime-2/2.1 line) with no reasoning support. Still GA and individually documented; it is the official recommended replacement for the retired gpt-4o-realtime-preview snapshot family (May 7, 2026 shutdown). Resolves a standing registry watch item — confirmed via the official pricing page and its own /docs/models page as a real, currently-listed, non-deprecated id, distinct from the also-real but deprecated bare `gpt-realtime`.","type":"audio_chat","input_modalities":["text","audio","image"],"output_modalities":["text","audio"],"status":"ga","limits":{"context_window":32000,"max_output":4096},"pricing":{"currency":"USD","input_per_mtok":4,"cached_input_per_mtok":0.4,"output_per_mtok":16,"notes":"Audio: $32.00/MTok input ($0.40 cached), $64.00/MTok output. Image input: $5.00/MTok ($0.50 cached). Knowledge cutoff Sep 30, 2024. Current pricing is identical to the deprecated bare `gpt-realtime`, but this model has a newer knowledge cutoff and remains GA (not deprecated). Re-confirmed 2026-08-24 against the live pricing page — unchanged."},"capabilities":{"tools":true,"streaming":true,"vision":true,"audio_input":true,"audio_output":true},"api":{"endpoint":"https://api.openai.com/v1/realtime","protocol":"openai","openai_compatible":true,"auth":{"type":"bearer","env_var":"OPENAI_API_KEY"},"sdks":{"openai":{"package":"openai","model_id":"gpt-realtime-1.5"}}},"sources":{"spec":"https://developers.openai.com/api/docs/models/gpt-realtime-1.5","pricing":"https://developers.openai.com/api/docs/pricing","release_notes":"https://developers.openai.com/api/docs/deprecations"},"last_verified":"2026-09-28"},{"id":"gpt-realtime-translate","provider":"openai","family":"gpt-realtime","display_name":"GPT Realtime Translate","description":"Speech-to-speech translation model. Per-minute pricing.","type":"audio_chat","input_modalities":["audio"],"output_modalities":["audio"],"status":"ga","pricing":{"currency":"USD","audio_input_per_minute":0.034},"capabilities":{"streaming":true,"audio_input":true,"audio_output":true},"api":{"endpoint":"https://api.openai.com/v1/realtime","protocol":"openai","openai_compatible":true,"auth":{"type":"bearer","env_var":"OPENAI_API_KEY"},"sdks":{"openai":{"package":"openai","model_id":"gpt-realtime-translate"}}},"sources":{"pricing":"https://developers.openai.com/api/docs/pricing"},"last_verified":"2026-09-28"},{"id":"gpt-realtime-whisper","provider":"openai","family":"gpt-realtime","display_name":"GPT Realtime Whisper","description":"Streaming speech-to-text model for applications that need low-latency transcript deltas from live audio. Was missing from the registry — added 2026-07-20 during a routine audit. Announced alongside gpt-realtime-2 and gpt-realtime-translate in OpenAI's \"Advancing voice intelligence\" post. Relationship to gpt-live-transcribe clarified 2026-08-17: OpenAI's official migration cookbook (developers.openai.com/cookbook, \"Migrate from Whisper to GPT-Transcribe and GPT-Live-Transcribe\") frames gpt-live-transcribe as the newer recommended default for realtime/streaming transcription, but does not formally deprecate this model — no deprecations-table entry exists for it, so status remains ga. Both stay independently documented and identically priced ($0.017/min); closing the standing watch item on this basis, no status change made.","type":"stt","input_modalities":["audio"],"output_modalities":["text"],"status":"ga","limits":{"context_window":16000,"max_output":2000},"pricing":{"currency":"USD","audio_input_per_minute":0.017},"capabilities":{"streaming":true},"api":{"endpoint":"https://api.openai.com/v1/realtime","protocol":"openai","openai_compatible":true,"auth":{"type":"bearer","env_var":"OPENAI_API_KEY"},"sdks":{"openai":{"package":"openai","model_id":"gpt-realtime-whisper"}}},"sources":{"spec":"https://developers.openai.com/api/docs/models/gpt-realtime-whisper","pricing":"https://developers.openai.com/api/docs/pricing","release_notes":"https://openai.com/index/advancing-voice-intelligence-with-new-models-in-the-api/","docs":"https://developers.openai.com/cookbook/examples/migrating_from_whisper_to_gpt_transcribe"},"last_verified":"2026-09-28"},{"id":"gpt-live-1","provider":"openai","family":"gpt-live","display_name":"GPT-Live 1","description":"New: full-duplex voice model, GA September 10, 2026, on a new dedicated Live API surface (v1/live/sessions) separate from the Realtime API. Listens and speaks simultaneously rather than turn-by-turn, and delegates reasoning/tool use to a backend Responses model (or a self-hosted backend via client delegation) instead of reasoning itself. Drops the image-input support its gpt-realtime-2.1 predecessor had. Backend model and tool usage are billed separately at that model's own rates, on top of this model's $0.05/minute voice-session price. Knowledge cutoff Jul 31, 2025. Independent press (not yet an official OpenAI benchmark page) reports a 30-point gain over gpt-realtime-2.1 on OpenAI's internal Full Duplex Bench — not recorded here pending an official source. Confirmed via the official changelog, pricing page, and model spec page, cross-checked against independent press coverage.","type":"audio_chat","input_modalities":["text","audio"],"output_modalities":["text","audio"],"status":"ga","released":"2026-09-10","pricing":{"currency":"USD","audio_input_per_minute":0.05,"notes":"Voice-session price only ($0.05/minute, billed per second, not rounded up to a whole minute). Backend Responses-model and tool usage is billed separately at that model's own token/tool rates — this is not an all-in price."},"capabilities":{"tools":true,"streaming":true,"audio_input":true,"audio_output":true},"api":{"endpoint":"https://api.openai.com/v1/live/sessions","protocol":"openai","openai_compatible":true,"auth":{"type":"bearer","env_var":"OPENAI_API_KEY"},"sdks":{"openai":{"package":"openai","model_id":"gpt-live-1"}}},"sources":{"spec":"https://developers.openai.com/api/docs/models/gpt-live-1","pricing":"https://developers.openai.com/api/docs/pricing","release_notes":"https://developers.openai.com/api/docs/changelog"},"last_verified":"2026-09-28"},{"id":"gpt-audio-1.5","provider":"openai","family":"gpt-audio","display_name":"GPT Audio 1.5","description":"Current GA flagship of OpenAI's Chat-Completions-based audio-in/audio-out family — a separate API surface from the Realtime-API gpt-realtime line (same tagline pattern: \"the best voice model for audio in, audio out\", but \"with Chat Completions\"). Recommended replacement for gpt-audio, gpt-audio-mini, gpt-4o-audio, gpt-4o-mini-audio, and the legacy gpt-4o-audio-preview/gpt-4o-mini-audio-preview snapshots. This entire gpt-audio family was previously untracked in the registry — added this pass after being found on the official deprecations table alongside the gpt-realtime family gap.","type":"audio_chat","input_modalities":["text","audio"],"output_modalities":["text","audio"],"status":"ga","limits":{"context_window":128000,"max_output":16384},"pricing":{"currency":"USD","input_per_mtok":2.5,"output_per_mtok":10,"notes":"Audio: $32.00/MTok input, $64.00/MTok output (no cached-input discount on this family — cached-input columns are '-' on the official pricing page). Knowledge cutoff Sep 30, 2024. Uses the Chat Completions API endpoint, not the Realtime API."},"capabilities":{"tools":true,"streaming":true,"audio_input":true,"audio_output":true},"api":{"endpoint":"https://api.openai.com/v1/chat/completions","protocol":"openai","openai_compatible":true,"auth":{"type":"bearer","env_var":"OPENAI_API_KEY"},"sdks":{"openai":{"package":"openai","model_id":"gpt-audio-1.5"}}},"sources":{"spec":"https://developers.openai.com/api/docs/models/gpt-audio-1.5","pricing":"https://developers.openai.com/api/docs/pricing","release_notes":"https://developers.openai.com/api/docs/deprecations"},"last_verified":"2026-09-28"},{"id":"gpt-transcribe","provider":"openai","family":"gpt-transcribe","display_name":"GPT Transcribe","description":"High-accuracy speech-to-text model for file uploads and Realtime transcription sessions. Supports free-form context and keyword hints for multi-language transcription. Launched 2026-07-28. Now the designated (co-)replacement for whisper-1, gpt-4o-transcribe, gpt-4o-mini-transcribe, and gpt-4o-transcribe-diarize, deprecated 2026-08-26 (sunset Feb 26, 2027).","type":"stt","input_modalities":["audio"],"output_modalities":["text"],"status":"ga","released":"2026-07-28","pricing":{"currency":"USD","audio_input_per_minute":0.0045},"capabilities":{"streaming":true},"api":{"endpoint":"https://api.openai.com/v1/audio/transcriptions","protocol":"openai","openai_compatible":true,"auth":{"type":"bearer","env_var":"OPENAI_API_KEY"},"sdks":{"openai":{"package":"openai","model_id":"gpt-transcribe"}}},"sources":{"spec":"https://developers.openai.com/api/docs/models/gpt-transcribe","pricing":"https://developers.openai.com/api/docs/pricing","release_notes":"https://developers.openai.com/api/docs/changelog"},"last_verified":"2026-09-28"},{"id":"gpt-live-transcribe","provider":"openai","family":"gpt-transcribe","display_name":"GPT Live Transcribe","description":"Low-latency speech-to-text model for realtime transcription sessions only (not available on the batch transcription endpoint). Tunable latency, supports keyword hints. Launched 2026-07-28 alongside gpt-transcribe. Relationship to gpt-realtime-whisper clarified 2026-08-17: OpenAI's official migration cookbook positions gpt-live-transcribe as the newer recommended default for realtime/streaming transcription (the streaming counterpart to gpt-transcribe replacing whisper-1 for file transcription), but gpt-realtime-whisper is not formally deprecated (absent from the deprecations table) and both remain identically priced ($0.017/min) and independently GA. Update 2026-09-07: the official Aug 26, 2026 deprecations entry formally names gpt-live-transcribe (or gpt-transcribe) as the replacement for the now-deprecated whisper-1/gpt-4o-transcribe/gpt-4o-mini-transcribe/gpt-4o-transcribe-diarize family — this registry records gpt-live-transcribe as the replacement_id on each of those four entries.","type":"stt","input_modalities":["audio"],"output_modalities":["text"],"status":"ga","released":"2026-07-28","pricing":{"currency":"USD","audio_input_per_minute":0.017},"capabilities":{"streaming":true},"api":{"endpoint":"https://api.openai.com/v1/realtime","protocol":"openai","openai_compatible":true,"auth":{"type":"bearer","env_var":"OPENAI_API_KEY"},"sdks":{"openai":{"package":"openai","model_id":"gpt-live-transcribe"}}},"sources":{"spec":"https://developers.openai.com/api/docs/models/gpt-live-transcribe","pricing":"https://developers.openai.com/api/docs/pricing","release_notes":"https://developers.openai.com/api/docs/changelog","docs":"https://developers.openai.com/cookbook/examples/migrating_from_whisper_to_gpt_transcribe"},"last_verified":"2026-09-28"},{"id":"gpt-image-2.5-sunburst","provider":"openai","family":"gpt-image","display_name":"GPT Image 2.5 Sunburst","description":"New: OpenAI's highest-quality image generation/editing model, GA September 8, 2026, via the Image API (v1/images/generations, v1/images/edits) and the Responses API image-generation tool. Positioned for workflows where editing precision matters most (production creative, polished product imagery) — higher image quality than gpt-image-2 per OpenAI's own model comparison, at the same token rates. Adds `xhigh` and `max` quality tiers on top of gpt-image-2's low/medium/high/auto settings. \"Token rates match GPT Image 2\" per OpenAI's own pricing note; confirmed identical on the live pricing page. Confirmed via the official changelog, model spec page, pricing page, and image-generation guide, cross-checked against independent press coverage.","type":"image_generation","input_modalities":["text","image"],"output_modalities":["image"],"status":"ga","released":"2026-09-08","pricing":{"currency":"USD","input_per_mtok":5,"cached_input_per_mtok":1.25,"notes":"Text input $5.00/MTok ($1.25 cached). Image input $8.00/MTok ($2.00 cached). Image output $30.00/MTok (no separate output pricing for text, since the model only outputs images). Same per-token rates as gpt-image-2; no separate batch-tier row published yet for this model as of 2026-09-14."},"capabilities":{"negative_prompt":false,"image_to_image":true,"inpainting":true},"api":{"endpoint":"https://api.openai.com/v1/images/generations","protocol":"openai","openai_compatible":true,"auth":{"type":"bearer","env_var":"OPENAI_API_KEY"},"sdks":{"openai":{"package":"openai","model_id":"gpt-image-2.5-sunburst"}}},"sources":{"spec":"https://developers.openai.com/api/docs/models/gpt-image-2.5-sunburst","pricing":"https://developers.openai.com/api/docs/pricing","release_notes":"https://developers.openai.com/api/docs/changelog","docs":"https://developers.openai.com/api/docs/guides/image-generation"},"last_verified":"2026-09-28"},{"id":"gpt-image-2.5-flare","provider":"openai","family":"gpt-image","display_name":"GPT Image 2.5 Flare","description":"New: OpenAI's fastest image generation/editing model for everyday use, GA September 8, 2026, via the Image API (v1/images/generations, v1/images/edits) and the Responses API image-generation tool. Positioned for fast, high-quality everyday generation (vs. Sunburst's editing-precision focus); \"very fast\" per OpenAI's own speed rating. Adds `xhigh` and `max` quality tiers on top of gpt-image-2's low/medium/high/auto settings. Identically priced to gpt-image-2.5-sunburst (\"token rates match GPT Image 2\" per OpenAI's own pricing note). Confirmed via the official changelog, model spec page, pricing page, and image-generation guide, cross-checked against independent press coverage.","type":"image_generation","input_modalities":["text","image"],"output_modalities":["image"],"status":"ga","released":"2026-09-08","pricing":{"currency":"USD","input_per_mtok":5,"cached_input_per_mtok":1.25,"notes":"Text input $5.00/MTok ($1.25 cached). Image input $8.00/MTok ($2.00 cached). Image output $30.00/MTok (no separate output pricing for text, since the model only outputs images). Same per-token rates as gpt-image-2 and gpt-image-2.5-sunburst; no separate batch-tier row published yet for this model as of 2026-09-14."},"capabilities":{"negative_prompt":false,"image_to_image":true,"inpainting":true},"api":{"endpoint":"https://api.openai.com/v1/images/generations","protocol":"openai","openai_compatible":true,"auth":{"type":"bearer","env_var":"OPENAI_API_KEY"},"sdks":{"openai":{"package":"openai","model_id":"gpt-image-2.5-flare"}}},"sources":{"spec":"https://developers.openai.com/api/docs/models/gpt-image-2.5-flare","pricing":"https://developers.openai.com/api/docs/pricing","release_notes":"https://developers.openai.com/api/docs/changelog","docs":"https://developers.openai.com/api/docs/guides/image-generation"},"last_verified":"2026-09-28"},{"id":"gpt-image-2","provider":"openai","family":"gpt-image","display_name":"GPT Image 2","description":"Latest OpenAI image generation model. Supports inpainting, image-to-image, and instruction-following at the prompt level.","type":"image_generation","input_modalities":["text","image"],"output_modalities":["image"],"status":"ga","limits":{"supported_resolutions":["1024x1024","1024x1536","1536x1024","2048x2048"]},"pricing":{"currency":"USD","input_per_mtok":5,"cached_input_per_mtok":1.25,"batch_input_per_mtok":2.5,"notes":"Corrected/enriched 2026-08-24 against the live pricing page: Text input $5.00/MTok ($1.25 cached, $2.50/$0.625 cached batch). Image input $8.00/MTok ($2.00 cached, $4.00/$1.00 cached batch). Image output $30.00/MTok ($15.00 batch). Per-image effective: ~$0.011-$0.211 depending on quality and resolution. Transparent-background output (background: \"transparent\", png/webp only) released in preview 2026-08-20 for gpt-image-2 / gpt-image-2-2026-04-21 in the Images API and Responses API image tool; no separate pricing tier."},"capabilities":{"negative_prompt":false,"image_to_image":true,"inpainting":true},"api":{"endpoint":"https://api.openai.com/v1/images/generations","protocol":"openai","openai_compatible":true,"auth":{"type":"bearer","env_var":"OPENAI_API_KEY"},"sdks":{"openai":{"package":"openai","model_id":"gpt-image-2"}}},"sources":{"pricing":"https://developers.openai.com/api/docs/pricing","release_notes":"https://developers.openai.com/api/docs/changelog"},"last_verified":"2026-09-28"},{"id":"text-embedding-3-large","provider":"openai","family":"embedding-3","display_name":"Text Embedding 3 Large","description":"High-quality embeddings, up to 3072 dimensions, supports Matryoshka truncation.","type":"embedding","input_modalities":["text"],"output_modalities":["embedding"],"status":"ga","released":"2024-01-25","limits":{"max_input_tokens":8191,"max_dimensions":3072,"matryoshka_dimensions":[256,1024,3072]},"pricing":{"currency":"USD","input_per_mtok":0.13,"batch_input_per_mtok":0.065},"capabilities":{"batching":true},"performance":{"benchmarks":{"mteb":64.6}},"api":{"endpoint":"https://api.openai.com/v1/embeddings","protocol":"openai","openai_compatible":true,"auth":{"type":"bearer","env_var":"OPENAI_API_KEY"},"sdks":{"openai":{"package":"openai","model_id":"text-embedding-3-large"}}},"aggregator_ids":{"openrouter":"openai/text-embedding-3-large"},"manual_tags":["smartest_embedding"],"sources":{"pricing":"https://developers.openai.com/api/docs/pricing"},"last_verified":"2026-09-28"},{"id":"text-embedding-3-small","provider":"openai","family":"embedding-3","display_name":"Text Embedding 3 Small","description":"Cost-efficient embeddings. 1536 dims by default, supports Matryoshka truncation.","type":"embedding","input_modalities":["text"],"output_modalities":["embedding"],"status":"ga","released":"2024-01-25","limits":{"max_input_tokens":8191,"max_dimensions":1536,"matryoshka_dimensions":[512,1536]},"pricing":{"currency":"USD","input_per_mtok":0.02,"batch_input_per_mtok":0.01},"capabilities":{"batching":true},"performance":{"benchmarks":{"mteb":62.3}},"api":{"endpoint":"https://api.openai.com/v1/embeddings","protocol":"openai","openai_compatible":true,"auth":{"type":"bearer","env_var":"OPENAI_API_KEY"},"sdks":{"openai":{"package":"openai","model_id":"text-embedding-3-small"}}},"aggregator_ids":{"openrouter":"openai/text-embedding-3-small"},"manual_tags":["cheapest_embedding"],"sources":{"pricing":"https://developers.openai.com/api/docs/pricing"},"last_verified":"2026-09-28"},{"id":"text-embedding-ada-002","provider":"openai","family":"embedding-ada","display_name":"Text Embedding Ada 002","description":"Legacy embedding model. Still available; superseded by text-embedding-3 family.","type":"embedding","input_modalities":["text"],"output_modalities":["embedding"],"status":"ga","limits":{"max_input_tokens":8191,"max_dimensions":1536},"pricing":{"currency":"USD","input_per_mtok":0.1,"batch_input_per_mtok":0.05},"capabilities":{"batching":true},"performance":{"benchmarks":{"mteb":60.99}},"api":{"endpoint":"https://api.openai.com/v1/embeddings","protocol":"openai","openai_compatible":true,"auth":{"type":"bearer","env_var":"OPENAI_API_KEY"},"sdks":{"openai":{"package":"openai","model_id":"text-embedding-ada-002"}}},"sources":{"pricing":"https://developers.openai.com/api/docs/pricing"},"last_verified":"2026-09-28"},{"id":"gpt-4o-mini-tts","provider":"openai","family":"tts","display_name":"GPT-4o Mini TTS","description":"Text-to-speech with GPT-4o-mini backbone. Corrected 2026-07-27: previously marked deprecated with sunset 2026-07-23, but the official deprecations table's shutdown row only names the dated snapshot gpt-4o-mini-tts-2025-03-20 (replaced by gpt-4o-mini-tts-2025-12-15) -- never the bare id. The bare id's own docs page is live with no shutdown banner and now defaults to the 2025-12-15 snapshot as the current, generally-available model. The prior deprecated/sunset data on the bare id appears to have been an entry error conflating it with the retired 2025-03-20 snapshot; flipped back to ga.","type":"tts","input_modalities":["text"],"output_modalities":["audio"],"status":"ga","limits":{"max_input_tokens":2000,"supported_voices":["alloy","ash","ballad","coral","echo","fable","onyx","nova","sage","shimmer","verse"]},"pricing":{"currency":"USD","input_per_mtok":0.6,"output_per_mtok":12,"notes":"Corrected 2026-07-27: model is token-based ($0.60/MTok text input, $12.00/MTok audio output) per its official model page, not per-character as previously recorded here."},"capabilities":{"streaming":true,"speed_control":true},"api":{"endpoint":"https://api.openai.com/v1/audio/speech","protocol":"openai","openai_compatible":true,"auth":{"type":"bearer","env_var":"OPENAI_API_KEY"},"sdks":{"openai":{"package":"openai","model_id":"gpt-4o-mini-tts"}}},"sources":{"spec":"https://developers.openai.com/api/docs/deprecations","pricing":"https://developers.openai.com/api/docs/models/gpt-4o-mini-tts"},"last_verified":"2026-09-28"},{"id":"tts-1-hd","provider":"openai","family":"tts-1","display_name":"TTS-1 HD","description":"Legacy high-quality TTS. Available; superseded by gpt-4o TTS family.","type":"tts","input_modalities":["text"],"output_modalities":["audio"],"status":"ga","limits":{"supported_audio_formats":["mp3","opus","aac","flac","wav","pcm"],"supported_voices":["alloy","echo","fable","onyx","nova","shimmer"],"max_input_chars":4096},"pricing":{"currency":"USD","per_million_characters":30},"capabilities":{"streaming":true,"speed_control":true},"api":{"endpoint":"https://api.openai.com/v1/audio/speech","protocol":"openai","openai_compatible":true,"auth":{"type":"bearer","env_var":"OPENAI_API_KEY"},"sdks":{"openai":{"package":"openai","model_id":"tts-1-hd"}}},"last_verified":"2026-09-28"},{"id":"tts-1","provider":"openai","family":"tts-1","display_name":"TTS-1","description":"Legacy TTS. Available; superseded by gpt-4o TTS family.","type":"tts","input_modalities":["text"],"output_modalities":["audio"],"status":"ga","limits":{"supported_audio_formats":["mp3","opus","aac","flac","wav","pcm"],"supported_voices":["alloy","echo","fable","onyx","nova","shimmer"],"max_input_chars":4096},"pricing":{"currency":"USD","per_million_characters":15},"capabilities":{"streaming":true,"speed_control":true},"api":{"endpoint":"https://api.openai.com/v1/audio/speech","protocol":"openai","openai_compatible":true,"auth":{"type":"bearer","env_var":"OPENAI_API_KEY"},"sdks":{"openai":{"package":"openai","model_id":"tts-1"}}},"manual_tags":["cheapest_tts"],"last_verified":"2026-09-28"},{"id":"omni-moderation-latest","provider":"openai","family":"moderation","display_name":"Omni Moderation","description":"Free multimodal moderation for OpenAI API users. Detects harmful content across text and images.","type":"moderation","input_modalities":["text","image"],"output_modalities":["text"],"status":"ga","pricing":{"currency":"USD","input_per_mtok":0,"notes":"Free for OpenAI API users."},"capabilities":{"vision":true},"api":{"endpoint":"https://api.openai.com/v1/moderations","protocol":"openai","openai_compatible":true,"auth":{"type":"bearer","env_var":"OPENAI_API_KEY"},"sdks":{"openai":{"package":"openai","model_id":"omni-moderation-latest"}}},"last_verified":"2026-09-28"},{"id":"gpt-4.1","provider":"openai","family":"gpt-4.1","display_name":"GPT-4.1","description":"Cost-effective production model. Still GA; recommended replacement for legacy gpt-4 / gpt-4-turbo (which retire Oct 23, 2026).","type":"chat","input_modalities":["text","image"],"output_modalities":["text"],"status":"ga","released":"2025-04-14","limits":{"context_window":1000000,"max_output":32768},"pricing":{"currency":"USD","input_per_mtok":2,"cached_input_per_mtok":0.5,"output_per_mtok":8,"batch_input_per_mtok":1,"batch_output_per_mtok":4},"capabilities":{"tools":true,"parallel_tools":true,"json_mode":true,"structured_output":true,"streaming":true,"batching":true,"prompt_caching":true,"vision":true},"performance":{"intelligence_index":60},"api":{"endpoint":"https://api.openai.com/v1/chat/completions","protocol":"openai","openai_compatible":true,"auth":{"type":"bearer","env_var":"OPENAI_API_KEY"},"sdks":{"openai":{"package":"openai","model_id":"gpt-4.1"}}},"aggregator_ids":{"openrouter":"openai/gpt-4.1"},"last_verified":"2026-09-28"},{"id":"gpt-4.1-mini","provider":"openai","family":"gpt-4.1","display_name":"GPT-4.1 Mini","description":"Smaller GPT-4.1 variant. Recommended replacement for gpt-3.5-turbo and gpt-4o-search-preview.","type":"chat","input_modalities":["text","image"],"output_modalities":["text"],"status":"ga","released":"2025-04-14","limits":{"context_window":1000000,"max_output":32768},"pricing":{"currency":"USD","input_per_mtok":0.4,"cached_input_per_mtok":0.1,"output_per_mtok":1.6,"batch_input_per_mtok":0.2,"batch_output_per_mtok":0.8},"capabilities":{"tools":true,"json_mode":true,"structured_output":true,"streaming":true,"batching":true,"prompt_caching":true,"vision":true},"api":{"endpoint":"https://api.openai.com/v1/chat/completions","protocol":"openai","openai_compatible":true,"auth":{"type":"bearer","env_var":"OPENAI_API_KEY"},"sdks":{"openai":{"package":"openai","model_id":"gpt-4.1-mini"}}},"aggregator_ids":{"openrouter":"openai/gpt-4.1-mini"},"last_verified":"2026-09-28"},{"id":"gpt-4o","provider":"openai","family":"gpt-4o","display_name":"GPT-4o","description":"Multimodal model from 2024. Retired from ChatGPT Feb 13, 2026. Still in API but specific snapshots are deprecated.","type":"chat","input_modalities":["text","image","audio"],"output_modalities":["text"],"status":"ga","released":"2024-05-13","limits":{"context_window":128000,"max_output":16384},"pricing":{"currency":"USD","input_per_mtok":2.5,"cached_input_per_mtok":1.25,"output_per_mtok":10,"batch_input_per_mtok":1.25,"batch_output_per_mtok":5},"capabilities":{"tools":true,"parallel_tools":true,"json_mode":true,"structured_output":true,"streaming":true,"batching":true,"prompt_caching":true,"vision":true,"audio_input":true},"api":{"endpoint":"https://api.openai.com/v1/chat/completions","protocol":"openai","openai_compatible":true,"auth":{"type":"bearer","env_var":"OPENAI_API_KEY"},"sdks":{"openai":{"package":"openai","model_id":"gpt-4o"}}},"aggregator_ids":{"openrouter":"openai/gpt-4o"},"last_verified":"2026-09-28"},{"id":"gpt-4o-mini","provider":"openai","family":"gpt-4o","display_name":"GPT-4o Mini","description":"Smaller, cheaper GPT-4o. Still GA in API.","type":"chat","input_modalities":["text","image"],"output_modalities":["text"],"status":"ga","released":"2024-07-18","limits":{"context_window":128000,"max_output":16384},"pricing":{"currency":"USD","input_per_mtok":0.15,"cached_input_per_mtok":0.075,"output_per_mtok":0.6,"batch_input_per_mtok":0.075,"batch_output_per_mtok":0.3},"capabilities":{"tools":true,"json_mode":true,"structured_output":true,"streaming":true,"batching":true,"prompt_caching":true,"vision":true},"api":{"endpoint":"https://api.openai.com/v1/chat/completions","protocol":"openai","openai_compatible":true,"auth":{"type":"bearer","env_var":"OPENAI_API_KEY"},"sdks":{"openai":{"package":"openai","model_id":"gpt-4o-mini"}}},"aggregator_ids":{"openrouter":"openai/gpt-4o-mini"},"last_verified":"2026-09-28"},{"id":"grok-4.7","provider":"xai","family":"grok-4.7","display_name":"Grok 4.7","description":"xAI's flagship model, GA September 21, 2026 (per docs.x.ai/developers/grok-4-7, docs.x.ai/developers/models/grok-4.7, and x.ai/news/grok-4-7; also live on docs.x.ai/developers/pricing's model table and OpenRouter's x-ai/grok-4.7 listing — cross-checked and in agreement). Shipped ~9 days after Musk's last-stated slip (\"needs a few more days\" on Sept 11, following earlier missed targets of Sept 12 and \"3-4 weeks\"). Uses a new, larger base model (2.1T parameters per press coverage, up from a reported 1.5T for Grok 4.6) with a longer RL run weighted toward long-running, hours-long tasks, plus improved self-verification and long-context handling; xAI's announcement frames it as a same-price/same-speed upgrade over Grok 4.6, not a deprecation — Grok 4.6 remains GA. 500k context, no stated max-output-token cap. Reasoning effort configurable (low/medium/high/xhigh, default high). Supports tool calling, web search, X search, and code execution (confirmed via docs.x.ai/developers/models/grok-4.7 and the Tools Overview / Web Search docs pages). Batch API not supported. Knowledge cutoff reported inconsistently across sources (May 2026 per one fetch of the model spec page, June 2026 pretraining + August 2026 supplemental per a third-party analysis) — left out of structured fields pending a single authoritative figure. A separate 'Grok 4.7 Fast' variant (2x speed at 2x price) is available only through Cursor and Grok Build, not the public API — not added as its own registry entry since it has no standalone API model id. Available via the API, Cursor, Grok Build, and third-party gateways (OpenRouter, Vercel, Cloudflare). **Spot-checked 2026-09-28, unchanged:** docs.x.ai/developers/models/grok-4.7 and the live docs.x.ai/developers/pricing table both re-confirm 500k context, $2/$0.50/$6 (<200k) and $4/$1/$12 (≥200k) per MTok, and reasoning modes low/medium/high/xhigh (default high) — no drift since the Sept 22 pass.","type":"reasoning","input_modalities":["text","image"],"output_modalities":["text"],"status":"ga","released":"2026-09-21","limits":{"context_window":500000},"pricing":{"currency":"USD","input_per_mtok":2,"cached_input_per_mtok":0.5,"output_per_mtok":6,"notes":"Long-context tier (>200k tokens): $4.00/MTok input, $1.00/MTok cached input, $12.00/MTok output — same structure as Grok 4.6. Confirmed via three independent official fetches in agreement: docs.x.ai/developers/pricing's rendered model table, docs.x.ai/developers/models/grok-4.7's own spec page, and docs.x.ai/developers/grok-4-7; also matches OpenRouter's x-ai/grok-4.7 listing (OpenRouter's own margin-adjusted rate of $1.60/$4.80 is a routing markup/discount, not the base xAI rate, and is not used here)."},"capabilities":{"tools":true,"parallel_tools":true,"json_mode":true,"structured_output":true,"streaming":true,"system_prompt":true,"reasoning":true,"reasoning_modes":["low","medium","high","xhigh"],"reasoning_default":"high","vision":true,"web_search":true,"code_execution":true},"api":{"endpoint":"https://api.x.ai/v1/chat/completions","protocol":"openai_compat","openai_compatible":true,"auth":{"type":"bearer","env_var":"XAI_API_KEY"},"sdks":{"openai":{"package":"openai","model_id":"grok-4.7"},"vercel-ai":{"package":"@ai-sdk/xai","model_id":"grok-4.7","factory":"xai('grok-4.7')"}},"supported_params":["max_tokens","temperature","top_p","stop","tools","tool_choice","response_format","stream"]},"aggregator_ids":{"openrouter":"x-ai/grok-4.7"},"sources":{"spec":"https://docs.x.ai/developers/models/grok-4.7","pricing":"https://docs.x.ai/developers/pricing","release_notes":"https://x.ai/news/grok-4-7"},"last_verified":"2026-09-28"},{"id":"grok-4.6","provider":"xai","family":"grok-4.6","display_name":"Grok 4.6","description":"xAI's model, GA August 12, 2026 (per docs.x.ai/developers/grok-4-6 and x.ai/news/grok-4-6). Builds on Grok 4.5 with extended agentic RL training for long-running agents and coding; the announcement explicitly does not deprecate or replace Grok 4.5. 500k context, text output has no stated max-token cap. Reasoning effort configurable (low/medium/high/xhigh, default high) — xhigh is new for this release. Also supports X search in addition to web search and code execution. Knowledge cutoff February 1, 2026. Available via the API, Grok Build, Cursor, GitHub Copilot (as of Aug 14), and third-party gateways (OpenRouter, Vercel, Cloudflare). Long-context tier (>200k tokens) uses higher pricing. **Update 2026-09-21/22:** superseded as xAI's flagship by Grok 4.7 (see that entry), which xAI explicitly frames as a same-price/same-speed upgrade rather than a deprecation of 4.6 — this model remains GA and live on the pricing catalog, still confirmed on docs.x.ai/developers/pricing and docs.x.ai/developers/models with unchanged pricing this pass.","type":"reasoning","input_modalities":["text","image"],"output_modalities":["text"],"status":"ga","released":"2026-08-12","limits":{"context_window":500000},"pricing":{"currency":"USD","input_per_mtok":2,"cached_input_per_mtok":0.5,"output_per_mtok":6,"notes":"Long-context tier (>200k tokens): $4.00/MTok input, $1.00/MTok cached input, $12.00/MTok output — requests whose prompt reaches the 200k-token threshold are billed at the higher rate for the entire request. Confirmed on docs.x.ai/developers/pricing and cross-checked against docs.x.ai/developers/grok-4-6 and OpenRouter's x-ai/grok-4.6 listing on 2026-08-17 (all three agree). Re-confirmed unchanged on 2026-08-24 against the live docs.x.ai/developers/pricing table. Re-confirmed unchanged again 2026-09-07 (docs.x.ai/developers/pricing table row and the models catalog page both match). Re-confirmed unchanged again 2026-09-14 against the raw pricing/billing config embedded in both docs.x.ai/developers/pricing and docs.x.ai/developers/models (identical values in each). Re-confirmed unchanged again 2026-09-22 against the live docs.x.ai/developers/pricing model table, now shown alongside the newly-added Grok 4.7 row at identical pricing."},"capabilities":{"tools":true,"parallel_tools":true,"json_mode":true,"structured_output":true,"streaming":true,"system_prompt":true,"reasoning":true,"reasoning_modes":["low","medium","high","xhigh"],"reasoning_default":"high","vision":true,"web_search":true,"code_execution":true},"api":{"endpoint":"https://api.x.ai/v1/chat/completions","protocol":"openai_compat","openai_compatible":true,"auth":{"type":"bearer","env_var":"XAI_API_KEY"},"sdks":{"openai":{"package":"openai","model_id":"grok-4.6"},"vercel-ai":{"package":"@ai-sdk/xai","model_id":"grok-4.6","factory":"xai('grok-4.6')"}},"supported_params":["max_tokens","temperature","top_p","stop","tools","tool_choice","response_format","stream"]},"aggregator_ids":{"openrouter":"x-ai/grok-4.6"},"sources":{"spec":"https://docs.x.ai/developers/models/grok-4.6","pricing":"https://docs.x.ai/developers/pricing","release_notes":"https://x.ai/news/grok-4-6"},"last_verified":"2026-09-28"},{"id":"grok-4.5","provider":"xai","family":"grok-4.5","display_name":"Grok 4.5","description":"xAI's model, GA July 8, 2026 (aliases: grok-4.5-latest, grok-build-latest). 500k context, configurable reasoning effort. Long-context tier (>200k tokens) uses higher pricing. **Update 2026-09-14:** the reasoning effort options now include `xhigh` (previously only low/medium/high were documented for this model) — confirmed directly on docs.x.ai/developers/models/grok-4.5's own spec page (\"Reasoning efforts (supported): low, medium, high, xhigh\") and cross-checked against the identical `reasoningEffortOptions` field in the raw pricing/billing config embedded in docs.x.ai/developers/pricing and docs.x.ai/developers/models (both agree). Default remains `high`. Unclear whether this is a genuine new capability or a pre-existing gap in the registry; treated as a correction either way.","type":"reasoning","input_modalities":["text","image"],"output_modalities":["text"],"status":"ga","released":"2026-07-08","limits":{"context_window":500000},"pricing":{"currency":"USD","input_per_mtok":2,"cached_input_per_mtok":0.3,"output_per_mtok":6,"notes":"Long-context tier (>200k tokens): $4.00/MTok input, $0.60/MTok cached input, $12.00/MTok output. As of 2026-08-03, docs.x.ai/developers/release-notes (July 2026) states Grok 4.5 is now available in the API console for EU users, but the model spec page still lists inference regions as us-east-1/us-west-2 only (no EU inference region) — treat as partial/console-level EU availability, not a new EU data region, pending clarification. Pricing re-confirmed exact-match on 2026-08-10 against the live pricing config embedded in docs.x.ai/developers/pricing (promptTextTokenPrice/completionTextTokenPrice/cachedPromptTokenPrice raw values). Re-confirmed unchanged again on 2026-08-17 following Grok 4.6's Aug 12 GA launch — xAI's announcement explicitly frames 4.6 as building on 4.5 rather than replacing it, and Grok 4.5 remains listed and priced identically on the live catalog. Re-confirmed unchanged again on 2026-08-24 against the live docs.x.ai/developers/pricing table. Re-confirmed unchanged again 2026-09-07. Re-confirmed unchanged again 2026-09-14 against the raw pricing config on both docs.x.ai/developers/pricing and docs.x.ai/developers/models. Re-confirmed unchanged again 2026-09-22 against the live docs.x.ai/developers/pricing model table; Grok 4.5 remains listed and priced identically alongside the newly-added Grok 4.7."},"capabilities":{"tools":true,"parallel_tools":true,"json_mode":true,"structured_output":true,"streaming":true,"system_prompt":true,"reasoning":true,"reasoning_modes":["low","medium","high","xhigh"],"reasoning_default":"high","vision":true,"web_search":true,"code_execution":true},"api":{"endpoint":"https://api.x.ai/v1/chat/completions","protocol":"openai_compat","openai_compatible":true,"auth":{"type":"bearer","env_var":"XAI_API_KEY"},"sdks":{"openai":{"package":"openai","model_id":"grok-4.5"},"vercel-ai":{"package":"@ai-sdk/xai","model_id":"grok-4.5","factory":"xai('grok-4.5')"}},"supported_params":["max_tokens","temperature","top_p","stop","tools","tool_choice","response_format","stream"]},"sources":{"spec":"https://docs.x.ai/developers/models/grok-4.5","pricing":"https://docs.x.ai/developers/pricing","release_notes":"https://docs.x.ai/developers/release-notes"},"last_verified":"2026-09-28"},{"id":"grok-4.20-0309-reasoning","provider":"xai","family":"grok-4.20","display_name":"Grok 4.20 (Reasoning)","description":"Multi-agent reasoning model released March 10, 2026. Built-in 4-agent system (Grok captain, Harper research, Benjamin logic, etc.) for complex tasks. 1M context, text/image input (docs.x.ai does not list video as a supported input modality).","type":"reasoning","input_modalities":["text","image"],"output_modalities":["text"],"status":"ga","released":"2026-03-10","limits":{"context_window":1000000,"max_output":32768},"pricing":{"currency":"USD","input_per_mtok":1.25,"cached_input_per_mtok":0.2,"output_per_mtok":2.5,"notes":"Long-context tier (≥200k tokens): $2.50/MTok input, $0.40/MTok cached input, $5.00/MTok output."},"capabilities":{"tools":true,"parallel_tools":true,"json_mode":true,"structured_output":true,"streaming":true,"system_prompt":true,"reasoning":true,"reasoning_modes":["off","on"],"vision":true,"web_search":true,"code_execution":true},"performance":{"intelligence_index":80,"benchmarks":{"mmlu_pro":80,"gpqa":75}},"api":{"endpoint":"https://api.x.ai/v1/chat/completions","protocol":"openai_compat","openai_compatible":true,"auth":{"type":"bearer","env_var":"XAI_API_KEY"},"sdks":{"openai":{"package":"openai","model_id":"grok-4.20-0309-reasoning","factory":"openai({ baseURL: 'https://api.x.ai/v1' })"},"vercel-ai":{"package":"@ai-sdk/xai","model_id":"grok-4.20-0309-reasoning","factory":"xai('grok-4.20-0309-reasoning')"}},"supported_params":["max_tokens","temperature","top_p","stop","tools","tool_choice","response_format","stream"]},"aggregator_ids":{"openrouter":"x-ai/grok-4.20"},"sources":{"spec":"https://docs.x.ai/developers/models/grok-4.20-0309-reasoning","pricing":"https://docs.x.ai/developers/pricing","release_notes":"https://x.ai/news"},"last_verified":"2026-09-28"},{"id":"grok-4.20-0309-non-reasoning","provider":"xai","family":"grok-4.20","display_name":"Grok 4.20 (Non-Reasoning)","description":"Non-reasoning variant of Grok 4.20 for latency-sensitive use cases. 1M context.","type":"chat","input_modalities":["text","image"],"output_modalities":["text"],"status":"ga","released":"2026-03-10","limits":{"context_window":1000000,"max_output":32768},"pricing":{"currency":"USD","input_per_mtok":1.25,"cached_input_per_mtok":0.2,"output_per_mtok":2.5,"notes":"Long-context tier (≥200k tokens): $2.50/MTok input, $0.40/MTok cached input, $5.00/MTok output."},"capabilities":{"tools":true,"parallel_tools":true,"json_mode":true,"structured_output":true,"streaming":true,"system_prompt":true,"vision":true,"web_search":true},"performance":{"speed_tps":100,"intelligence_index":70},"api":{"endpoint":"https://api.x.ai/v1/chat/completions","protocol":"openai_compat","openai_compatible":true,"auth":{"type":"bearer","env_var":"XAI_API_KEY"},"sdks":{"openai":{"package":"openai","model_id":"grok-4.20-0309-non-reasoning"}}},"aggregator_ids":{"openrouter":"x-ai/grok-4.20-non-reasoning"},"sources":{"spec":"https://docs.x.ai/developers/models/grok-4.20-0309-non-reasoning","pricing":"https://docs.x.ai/developers/pricing"},"last_verified":"2026-09-28"},{"id":"grok-4.20-multi-agent-0309","provider":"xai","family":"grok-4.20","display_name":"Grok 4.20 Multi-Agent","description":"Explicit multi-agent variant of Grok 4.20. Coordinates 4 specialized agents in parallel. 1M context.","type":"reasoning","input_modalities":["text","image"],"output_modalities":["text"],"status":"ga","released":"2026-03-10","limits":{"context_window":1000000,"max_output":32768},"pricing":{"currency":"USD","input_per_mtok":1.25,"cached_input_per_mtok":0.2,"output_per_mtok":2.5,"notes":"Long-context tier (≥200k tokens): $2.50/MTok input, $0.40/MTok cached input, $5.00/MTok output."},"capabilities":{"tools":true,"parallel_tools":true,"structured_output":true,"streaming":true,"system_prompt":true,"reasoning":true,"vision":true,"web_search":true,"code_execution":true},"performance":{"intelligence_index":82},"api":{"endpoint":"https://api.x.ai/v1/chat/completions","protocol":"openai_compat","openai_compatible":true,"auth":{"type":"bearer","env_var":"XAI_API_KEY"},"sdks":{"openai":{"package":"openai","model_id":"grok-4.20-multi-agent-0309"}}},"aggregator_ids":{"openrouter":"x-ai/grok-4.20-multi-agent"},"sources":{"spec":"https://docs.x.ai/developers/models","pricing":"https://docs.x.ai/developers/pricing"},"last_verified":"2026-09-28"},{"id":"grok-build-0.1","provider":"xai","family":"grok-build","display_name":"Grok Build 0.1","description":"xAI's development-focused model. 256k context, optimized for building and iterative coding tasks at lower cost.","type":"chat","input_modalities":["text"],"output_modalities":["text"],"status":"ga","limits":{"context_window":256000,"max_output":16384},"pricing":{"currency":"USD","input_per_mtok":1,"cached_input_per_mtok":0.2,"output_per_mtok":2,"notes":"Long-context tier (≥200k tokens): $2.00/MTok input, $0.40/MTok cached input, $4.00/MTok output."},"capabilities":{"tools":true,"json_mode":true,"structured_output":true,"streaming":true,"system_prompt":true},"api":{"endpoint":"https://api.x.ai/v1/chat/completions","protocol":"openai_compat","openai_compatible":true,"auth":{"type":"bearer","env_var":"XAI_API_KEY"},"sdks":{"openai":{"package":"openai","model_id":"grok-build-0.1"}}},"sources":{"spec":"https://docs.x.ai/developers/models/grok-build-0.1","pricing":"https://docs.x.ai/developers/pricing"},"last_verified":"2026-09-28"},{"id":"grok-4.3","provider":"xai","family":"grok-4.3","display_name":"Grok 4.3","description":"Chat, reasoning, and coding model. 1M context. Strong general purpose model from xAI. **Fixed 2026-09-14:** the registry previously listed this as a plain chat model with no reasoning support — docs.x.ai/developers/models/grok-4.3 actually documents full configurable reasoning effort (\"none, low, medium, high, xhigh\", default `low`), and the identical `reasoningEffortOptions` field appears in the raw pricing/billing config embedded in both docs.x.ai/developers/pricing and docs.x.ai/developers/models (two independent official fetches agree); OpenRouter's third-party listing independently corroborates reasoning support. Type changed chat -> reasoning and reasoning capabilities/effort levels added accordingly. Unclear whether this is a new capability or a pre-existing registry gap; not otherwise a functional change (pricing/context window unchanged).","type":"reasoning","input_modalities":["text","image"],"output_modalities":["text"],"status":"ga","limits":{"context_window":1000000,"max_output":16384},"pricing":{"currency":"USD","input_per_mtok":1.25,"cached_input_per_mtok":0.2,"output_per_mtok":2.5,"notes":"Long-context tier (≥200k tokens): $2.50/MTok input, $0.40/MTok cached input, $5.00/MTok output. Re-confirmed unchanged 2026-09-14 against the raw pricing config on docs.x.ai/developers/pricing and docs.x.ai/developers/models. Re-confirmed unchanged again 2026-09-22 against the live docs.x.ai/developers/pricing model table."},"capabilities":{"tools":true,"parallel_tools":true,"json_mode":true,"structured_output":true,"streaming":true,"system_prompt":true,"reasoning":true,"reasoning_modes":["none","low","medium","high","xhigh"],"reasoning_default":"low","vision":true,"web_search":true},"performance":{"speed_tps":90,"intelligence_index":72},"api":{"endpoint":"https://api.x.ai/v1/chat/completions","protocol":"openai_compat","openai_compatible":true,"auth":{"type":"bearer","env_var":"XAI_API_KEY"},"sdks":{"openai":{"package":"openai","model_id":"grok-4.3"}}},"aggregator_ids":{"openrouter":"x-ai/grok-4.3"},"sources":{"spec":"https://docs.x.ai/developers/models/grok-4.3","pricing":"https://docs.x.ai/developers/pricing"},"last_verified":"2026-09-28"},{"id":"grok-imagine-image-pro","provider":"xai","family":"grok-imagine","display_name":"Grok Imagine Image Pro","description":"Reactivated July 2026 as an alias id for grok-imagine-image-quality (the original standalone \"Pro\" tier model under this id was retired May 15, 2026; xAI has since reused this id as an alias pointing at the quality tier). Reconfirmed 2026-08-17: fetching docs.x.ai/developers/models/grok-imagine-image-pro serves the grok-imagine-image-quality page content directly, and that page's own alias list explicitly names both grok-imagine-image-quality-latest and grok-imagine-image-pro — alias relationship unchanged despite grok-imagine-image-pro no longer appearing as its own row on the pricing page table. **Update 2026-09-07:** grok-imagine-image-quality (this id's alias target) has an announced 2026-11-02 retirement (announced 2026-09-02, docs.x.ai/developers/migration/imagine-image-quality-nov-2). That migration page explicitly states requests to grok-imagine-image-pro, which already redirect to grok-imagine-image-quality, will follow it on to grok-imagine-image-2.0 (quality forced to 'low') once the Nov 2 retirement takes effect — so this id's ultimate target is changing from grok-imagine-image-quality to grok-imagine-image-2.0, still kept `ga`/status unchanged here since it keeps resolving throughout (same treatment as the earlier May 15 alias reactivation), only replacement_id updated to reflect the new end-of-chain target.","type":"image_generation","input_modalities":["text"],"output_modalities":["image"],"status":"ga","replacement_id":"grok-imagine-image-2.0","limits":{"supported_aspect_ratios":["1:1","16:9","9:16","4:3","3:4"]},"pricing":{"currency":"USD","image_tiers":[{"label":"1K","resolution":"1024x1024","per_image":0.05},{"label":"2K","resolution":"2048x2048","per_image":0.07}],"notes":"Input image: $0.01/image. Alias of grok-imagine-image-quality — see that entry for the canonical id. Confirmed via raw pricing config on docs.x.ai/developers/pricing (2026-08-10): 'grok-imagine-image-pro' is listed as an alias of the 'grok-imagine-image-quality' model entry (1K=$0.05, 2K=$0.07, input image=$0.01) — pricing unchanged."},"api":{"endpoint":"https://api.x.ai/v1/images/generations","protocol":"openai_compat","openai_compatible":true,"auth":{"type":"bearer","env_var":"XAI_API_KEY"},"sdks":{"openai":{"package":"openai","model_id":"grok-imagine-image-pro"}}},"sources":{"spec":"https://docs.x.ai/developers/models","release_notes":"https://docs.x.ai/developers/migration/imagine-image-quality-nov-2"},"last_verified":"2026-09-28"},{"id":"grok-imagine-image","provider":"xai","family":"grok-imagine","display_name":"Grok Imagine Image","description":"xAI image generation model.","type":"image_generation","input_modalities":["text","image"],"output_modalities":["image"],"status":"ga","limits":{"supported_resolutions":["1024x1024","2048x2048"]},"pricing":{"currency":"USD","image_tiers":[{"label":"1K","resolution":"1024x1024","per_image":0.02},{"label":"2K","resolution":"2048x2048","per_image":0.02}],"notes":"Input image: $0.002/image."},"capabilities":{"image_to_image":true},"api":{"endpoint":"https://api.x.ai/v1/images/generations","protocol":"openai_compat","openai_compatible":true,"auth":{"type":"bearer","env_var":"XAI_API_KEY"},"sdks":{"openai":{"package":"openai","model_id":"grok-imagine-image"}}},"sources":{"spec":"https://docs.x.ai/developers/pricing"},"last_verified":"2026-09-28"},{"id":"grok-imagine-image-2.0","provider":"xai","family":"grok-imagine","display_name":"Grok Imagine Image 2.0","description":"xAI's newest image generation/editing model, GA August 7, 2026 (per x.ai/news/grok-imagine-image-2 and docs.x.ai/developers/models/grok-imagine-image-2.0) — a pre-existing registry gap discovered this pass, live since launch. xAI's own announcement reports it ranking #2 on the Arena text-to-image and image-edit leaderboards at launch (behind OpenAI's gpt-image-2). Multi-reference editing accepts up to 5 input images in a single generation (magic-wand region editing, segmentation-based selection, background removal). **Update 2026-09-07:** a dated-in-August release note added a `quality` parameter (`low`/`medium`/`auto`) — default changed from `medium` to `auto` (auto resolves to `low` for generation, `medium` for editing), and images are billed at whichever quality they're actually served at; passing `low`/`medium` explicitly pins it. The same update raised the multi-image-editing cap from 3 to 5 reference images (already reflected above) and added `21:9` and `5:2` aspect ratios alongside the existing set. Also, this model becomes the redirect target for the retiring grok-imagine-image-quality (and its grok-imagine-image-pro alias) as of 2026-11-02 — see that entry. **Pricing fix 2026-09-14:** the flat $0.04/image rate used last pass was wrong/stale — docs.x.ai/developers/pricing's rendered table now shows a full quality x resolution breakdown for this model (1K Low $0.04, 2K Low $0.06, 1K Medium $0.06, 2K Medium $0.08, plus a $0.01/image input-image surcharge), matching the raw pricing/billing config embedded on both docs.x.ai/developers/pricing and docs.x.ai/developers/models and independently matching the dedicated docs.x.ai/developers/models/grok-imagine-image-2.0 spec page's own pricing table (three-way agreement). This resolves last pass's flagged ambiguity between the quality-based billing language and the previously-flat displayed rate — the tiers were not previously rendered on the pricing page (or were missed) and now are.","type":"image_generation","input_modalities":["text","image"],"output_modalities":["image"],"status":"ga","released":"2026-08-07","limits":{"max_images_per_request":5,"supported_aspect_ratios":["1:1","16:9","9:16","4:3","3:4","3:2","2:3","2:1","1:2","19.5:9","9:19.5","20:9","9:20","21:9","5:2"]},"pricing":{"currency":"USD","image_tiers":[{"label":"1K","quality":"low","resolution":"1024x1024","per_image":0.04},{"label":"2K","quality":"low","resolution":"2048x2048","per_image":0.06},{"label":"1K","quality":"medium","resolution":"1024x1024","per_image":0.06},{"label":"2K","quality":"medium","resolution":"2048x2048","per_image":0.08}],"notes":"Input image: $0.01/image. Confirmed 2026-09-14 via three independent official fetches all in agreement: the rendered pricing table on docs.x.ai/developers/pricing (columns: Media Input $0.01/img, then per-row Resolution x Quality -> Output price), the raw pricing/billing config embedded in both docs.x.ai/developers/pricing and docs.x.ai/developers/models (resolutionPricing array with quality-tagged entries), and the dedicated docs.x.ai/developers/models/grok-imagine-image-2.0 spec page's own pricing table (identical four rows). `auto` quality (the default) resolves to `low` for generation and `medium` for editing per the capabilities update, billed at whichever tier is actually served. Supersedes the prior pass's flat $0.04/image figure, which did not reflect the tier breakdown."},"capabilities":{"image_to_image":true},"api":{"endpoint":"https://api.x.ai/v1/images/generations","protocol":"openai_compat","openai_compatible":true,"auth":{"type":"bearer","env_var":"XAI_API_KEY"},"sdks":{"openai":{"package":"openai","model_id":"grok-imagine-image-2.0"}}},"sources":{"spec":"https://docs.x.ai/developers/models/grok-imagine-image-2.0","pricing":"https://docs.x.ai/developers/pricing","release_notes":"https://docs.x.ai/developers/release-notes","docs":"https://docs.x.ai/developers/model-capabilities/images/generation"},"last_verified":"2026-09-28"},{"id":"grok-imagine-video","provider":"xai","family":"grok-imagine","display_name":"Grok Imagine Video","description":"xAI video generation model. 720p output. Released January 28, 2026.","type":"video_generation","input_modalities":["text","image"],"output_modalities":["video"],"status":"ga","released":"2026-01-28","limits":{"max_video_resolution":"720p"},"pricing":{"currency":"USD","video_per_second":0.05,"notes":"$0.050/sec at 480p, $0.070/sec at 720p (max resolution)."},"capabilities":{"image_to_image":true},"api":{"endpoint":"https://api.x.ai/v1/videos/generations","protocol":"openai_compat","openai_compatible":true,"auth":{"type":"bearer","env_var":"XAI_API_KEY"},"sdks":{"openai":{"package":"openai","model_id":"grok-imagine-video"}}},"sources":{"spec":"https://docs.x.ai/developers/models"},"last_verified":"2026-09-28"},{"id":"grok-imagine-video-1.5","provider":"xai","family":"grok-imagine","display_name":"Grok Imagine Video 1.5","description":"xAI's enhanced video generation model. Includes native audio. GA June 16, 2026. Updated July 31, 2026 (x.ai/news/grok-imagine-video-1-5-references, same model id — no new id issued): added image and voice references for character/scene consistency, text-to-video generation without a starting image, and native 1080p output (already reflected in max_video_resolution below). Per the announcement, image references/text-to-video/1080p are live via the API; voice references are 'available on request'. No pricing change.","type":"video_generation","input_modalities":["text","image"],"output_modalities":["video"],"status":"ga","released":"2026-06-16","limits":{"max_video_resolution":"1080p"},"pricing":{"currency":"USD","video_per_second":0.08,"notes":"$0.080/sec at 480p, $0.250/sec at 1080p (max resolution, audio included)."},"capabilities":{"audio_output":true,"image_to_image":true},"api":{"endpoint":"https://api.x.ai/v1/videos/generations","protocol":"openai_compat","openai_compatible":true,"auth":{"type":"bearer","env_var":"XAI_API_KEY"},"sdks":{"openai":{"package":"openai","model_id":"grok-imagine-video-1.5"}}},"sources":{"spec":"https://docs.x.ai/developers/models"},"last_verified":"2026-09-28"},{"id":"grok-voice-think-fast-2.0","provider":"xai","family":"grok-voice","display_name":"Grok Voice Think Fast 2.0","description":"Realtime speech-to-speech model with tool calling (web search, X search, file search, MCP, custom functions), session resumption, and 5 built-in cloneable voices (Eve, Ara, Rex, Sal, Leo). Auto-detects 20+ languages. Successor to grok-voice-think-fast-1.0. The grok-voice-latest alias was scheduled to repoint from 1.0 to this model on 2026-08-05; as of 2026-08-10 the raw routing config embedded in docs.x.ai/developers/pricing lists 'grok-voice-latest' as an alias of grok-voice-think-fast-2.0 (not 1.0), confirming the switch has taken effect — note the human-readable prose on docs.x.ai/developers/model-capabilities/audio/speech-to-speech still read 'grok-voice-latest currently points to grok-voice-think-fast-1.0' at last fetch, which looks like stale/un-updated copy on that page; the machine-readable pricing config is treated as authoritative here. Release date (2026-07-29) is corroborated by secondary sources (release-notes page placement, press coverage) since the primary x.ai/news announcement page returned HTTP 403 on direct fetch — re-verify against that page if precision matters. Re-confirmed 2026-08-24: pricing/status unchanged; docs.x.ai/developers/pricing now explicitly labels the predecessor grok-voice-think-fast-1.0 row 'Deprecated' (see that entry).","type":"audio_chat","input_modalities":["text","audio"],"output_modalities":["text","audio"],"status":"ga","released":"2026-07-29","pricing":{"currency":"USD","notes":"$0.08/min audio (input/output not separately broken out by xAI), plus $0.004 per text input. Predecessor grok-voice-think-fast-1.0 is $0.05/min audio + $0.004/text input. Cached conversation history expires after 30 min inactivity when using session resumption."},"capabilities":{"tools":true,"streaming":true,"audio_input":true,"audio_output":true,"voice_cloning":true},"api":{"endpoint":"https://api.x.ai/v1/realtime","protocol":"openai_compat","openai_compatible":true,"auth":{"type":"bearer","env_var":"XAI_API_KEY"}},"sources":{"spec":"https://docs.x.ai/developers/model-capabilities/audio/speech-to-speech","pricing":"https://docs.x.ai/developers/pricing","release_notes":"https://docs.x.ai/developers/release-notes"},"last_verified":"2026-09-28"},{"id":"grok-tts","provider":"xai","family":"grok-voice","display_name":"Grok TTS","description":"Standalone text-to-speech API, built on the same stack as Grok Voice. Launched April 18, 2026 per xAI/press coverage. 5 built-in voices (Eve default, Ara, Rex, Sal, Leo) plus custom voice cloning, 20 languages (BCP-47), character-level timestamps, inline speech tags (e.g. [laugh], [sigh], <whisper>), 0.7x-1.5x speed control, and configurable output codec/sample rate (default MP3 24kHz/128kbps). Discovered during the 2026-08-10 research pass as a pre-existing registry gap — the model has been live since April 2026, this is not a change from the prior week; flagging in case the omission was intentional.","type":"tts","input_modalities":["text"],"output_modalities":["audio"],"status":"ga","released":"2026-04-18","limits":{"supported_voices":["eve","ara","rex","sal","leo"],"max_input_chars":15000},"pricing":{"currency":"USD","per_million_characters":15,"notes":"Confirmed directly on docs.x.ai/developers/pricing (2026-08-10): 'Text to Speech | $15.00 / 1M chars', cross-checked against the page's embedded pricing config (perCharacter raw value 150000, consistent with the site's 1e10 fixed-point scale used elsewhere on the page). Note: third-party April 2026 launch coverage (e.g. MarkTechPost, KuCoin) cited an introductory ~$4.20/M characters rate — price appears to have risen since launch (or the early figure was promotional); the current live docs.x.ai number is used here. Re-confirmed unchanged ($15.00/1M chars) on 2026-08-24."},"capabilities":{"streaming":true,"voice_cloning":true,"speed_control":true,"word_timestamps":true},"api":{"endpoint":"https://api.x.ai/v1/tts","protocol":"rest","openai_compatible":false,"auth":{"type":"bearer","env_var":"XAI_API_KEY"}},"sources":{"spec":"https://docs.x.ai/developers/model-capabilities/audio/text-to-speech","pricing":"https://docs.x.ai/developers/pricing","release_notes":"https://x.ai/news/grok-stt-and-tts-apis"},"last_verified":"2026-09-28"},{"id":"grok-voice-transcribe-1.0","provider":"xai","family":"grok-voice","display_name":"Grok Voice Transcribe 1.0","description":"Standalone speech-to-text API, built on the same stack as Grok Voice. Launched April 18, 2026 per xAI/press coverage (informally called \"grok-stt\" at launch, per the x.ai/news/grok-stt-and-tts-apis announcement title and endpoint path /v1/stt). 25+ language support, word-level timestamps, speaker diarization, multichannel audio, and inverse text normalization; batch (REST) and realtime streaming modes. Accepts WAV, MP3, OGG, Opus, FLAC, AAC, MP4, M4A, MKV (container) and PCM, mu-law, A-law (raw) formats, up to 500MB per request. **Renamed 2026-09-22:** discovered this pass that docs.x.ai now exclusively documents this model under a versioned id, `grok-voice-transcribe-1.0` (docs.x.ai/developers/model-capabilities/audio/speech-to-text's own code samples and \"Original model. Pin this slug to keep it.\" copy) — the plain `grok-stt` string used in the registry previously does not appear anywhere on the current docs pages, so the id here has been updated to match; whether `grok-stt` still works as a legacy alias is unconfirmed either way. **No longer the default**: as of grok-voice-transcribe-2.0's 2026-09-18 release (x.ai/news/grok-voice-transcribe-2), this 1.0 model must now be explicitly pinned by name — the `model` parameter defaults to 2.0 when omitted. xAI's 2.0 announcement states 1.0 \"deprecation planned for the coming weeks\" but gives no specific date; status kept `ga` here pending an explicit dated notice, per the same policy applied to grok-voice-think-fast-1.0.","type":"stt","input_modalities":["audio"],"output_modalities":["text"],"status":"ga","released":"2026-04-18","limits":{"max_audio_size_mb":500,"supported_audio_formats":["wav","mp3","ogg","opus","flac","aac","mp4","m4a","mkv","pcm","mu-law","a-law"]},"pricing":{"currency":"USD","audio_input_per_minute":0.001667,"notes":"Confirmed directly on docs.x.ai/developers/pricing (2026-08-10): '$0.10 / hr (REST), $0.20 / hr (Streaming)'. audio_input_per_minute above reflects the batch/REST rate ($0.10/hr = $0.001667/min); the realtime streaming rate is roughly double, $0.20/hr ($0.003333/min) — schema has no separate streaming-rate field, noted here instead. Re-confirmed unchanged ($0.10/hr REST, $0.20/hr streaming) on 2026-08-24, and again 2026-09-22 — grok-voice-transcribe-2.0 launched at the same price, so this row's rate is unchanged."},"capabilities":{"streaming":true,"audio_input":true,"word_timestamps":true,"speaker_diarization":true},"api":{"endpoint":"https://api.x.ai/v1/stt","protocol":"rest","openai_compatible":false,"auth":{"type":"bearer","env_var":"XAI_API_KEY"},"sdks":{"openai":{"package":"openai","model_id":"grok-voice-transcribe-1.0"}}},"sources":{"spec":"https://docs.x.ai/developers/model-capabilities/audio/speech-to-text","pricing":"https://docs.x.ai/developers/pricing","release_notes":"https://x.ai/news/grok-stt-and-tts-apis"},"last_verified":"2026-09-28"},{"id":"grok-voice-transcribe-2.0","provider":"xai","family":"grok-voice","display_name":"Grok Voice Transcribe 2.0","description":"Standalone speech-to-text API, second-generation model. GA September 18, 2026 (per x.ai/news/grok-voice-transcribe-2 and MarkTechPost coverage dated the same day; also live on docs.x.ai/developers/model-capabilities/audio/speech-to-text, whose code samples use `model=grok-voice-transcribe-2.0` and which is now the default when `model` is omitted, superseding grok-voice-transcribe-1.0). xAI's announcement claims 2x the accuracy of 1.0 at the same price — self-reported multilingual short-phrase word-error-rate improvement from 20.6% to 6.8%, and #1 accuracy ranking among 32 streaming models on the (third-party) Artificial Analysis leaderboard; these are vendor-reported figures, not independently verified here, so not entered under `performance.benchmarks`. Same feature set as 1.0 (batch/streaming modes, word-level timestamps, speaker diarization, multichannel, key-term biasing, text formatting, filler-word removal, Smart Turn end-of-turn detection) plus improved accuracy, especially on telephony, conversational, credentials, and short-voice-command audio. Same format/size support as 1.0 (WAV, MP3, OGG, Opus, FLAC, AAC, MP4, M4A, MKV, PCM, mu-law, A-law; up to 500MB/request). Pricing unchanged from 1.0: $0.10/hr batch (REST), $0.20/hr streaming. **Spot-checked 2026-09-28, unchanged:** raw-HTML curl of docs.x.ai/developers/pricing's embedded config confirms grok-voice-transcribe-2.0 still lists identical per-second pricing (perAudioSecond 277778 / perAudioSecondStreaming 555556 fixed-point units, matching $0.10/hr REST and $0.20/hr streaming) alongside grok-voice-transcribe-1.0 (identical figures, still labeled \"Original model. Pin this slug to keep it.\" and not marked deprecated in the machine-readable config).","type":"stt","input_modalities":["audio"],"output_modalities":["text"],"status":"ga","released":"2026-09-18","limits":{"max_audio_size_mb":500,"supported_audio_formats":["wav","mp3","ogg","opus","flac","aac","mp4","m4a","mkv","pcm","mu-law","a-law"]},"pricing":{"currency":"USD","audio_input_per_minute":0.001667,"notes":"Confirmed via docs.x.ai/developers/model-capabilities/audio/speech-to-text and x.ai/news/grok-voice-transcribe-2 (2026-09-18 launch announcement), both stating pricing is unchanged from grok-voice-transcribe-1.0: $0.10/hr batch (REST), $0.20/hr streaming. audio_input_per_minute above reflects the batch/REST rate ($0.10/hr = $0.001667/min); streaming is roughly double ($0.20/hr = $0.003333/min) — schema has no separate streaming-rate field, noted here instead."},"capabilities":{"streaming":true,"audio_input":true,"word_timestamps":true,"speaker_diarization":true},"api":{"endpoint":"https://api.x.ai/v1/stt","protocol":"rest","openai_compatible":false,"auth":{"type":"bearer","env_var":"XAI_API_KEY"},"sdks":{"openai":{"package":"openai","model_id":"grok-voice-transcribe-2.0"}}},"sources":{"spec":"https://docs.x.ai/developers/model-capabilities/audio/speech-to-text","pricing":"https://docs.x.ai/developers/pricing","release_notes":"https://x.ai/news/grok-voice-transcribe-2"},"last_verified":"2026-09-28"}],"total":149,"scope":{}},"meta":{"version":"223@2026-09-28","as_of":"2026-09-28T01:27:41.513Z","request_id":"a510fcd5023a2256"},"errors":[]}