Releases
New models. Clearer choices.
Pricing, benchmarks, and availability. Sources included.
Editorial reference
25 sourced · 22 routable · revised 2026-09-29
Live catalog signal
24 arrivals · through 29 Sept 2026
Model releases
Pricing, specs, and reported benchmarks from publisher announcements.
Newest first · revised 2026-09-29
- OpenAI
GPT-6.1 Sol
Routable on MiniRouter
Read release - Anthropic
Claude Sonnet 5.5
Routable on MiniRouter
Read release - SpaceXAI
Grok 4.7
Not in MiniRouter's catalog
Read release - StepFun
Step 5 Preview
Routable on MiniRouter
Read release - PrismML
Bonsai 2 27B
Routable on MiniRouter
Read release - DeepSeek
DeepSeek V4.1 Flash
Routable on MiniRouter
Read release - OpenAI
GPT-6 Astra
Routable on MiniRouter
Read release - Google
Gemini 3.8 Flash
Routable on MiniRouter
Read release - Alibaba
Qwen3.8-Max-0902
Routable on MiniRouter
Read release - Anthropic
Claude Fable 5.1
Routable on MiniRouter
Read release - Tencent
Tencent Hy4 preview
Not in MiniRouter's catalog
Read release - Z.ai
GLM-5.3-Flash
Routable on MiniRouter
Read release - Alibaba
Qwen3.8-Flash
Routable on MiniRouter
Read release - Alibaba
Qwen3.8-Flash-Next
Not in MiniRouter's catalog
Read release - DeepSeek
DeepSeek-V4-Flash-Vision-Exp
Routable on MiniRouter
Read release - Z.ai
GLM-5.3
Routable on MiniRouter
Read release - Google
Gemini 3.7 Flash
Routable on MiniRouter
Read release - DeepSeek
DeepSeek-V4-Pro
Routable on MiniRouter
Read release - xAI
Grok 4.6
Routable on MiniRouter
Read release - OpenAI
GPT-5.4 mini
Routable on MiniRouter
Read release - Meta
Muse Glimmer 30B
Routable on MiniRouter
Read release - MiniMax
MiniMax M3
Routable on MiniRouter
Read release - Moonshot AI
Kimi K3
Routable on MiniRouter
Read release - Alibaba
Qwen3.8-2.4T-A95B
Routable on MiniRouter
Read release - OpenAI
GPT-5.6 Sol
Routable on MiniRouter
Read release
Recent gateway arrivals
Vercel AI Gateway arrivals from the last 14 days.
24 current arrivals · latest 29 Sept 2026
- inclusionai
Ling 3.1 Flash
Not yet in MiniRouter’s catalog
inclusionai/ling-3.1-flash
Ling 3.1 Flash is InclusionAI's hybrid reasoning language model for coding, multi-step analysis, and tool-using agents. It has 560B total parameters with 25B activated parameters and supports long-context text workflows.
Gateway list rates: $0.00 input / $0.00 output per 1M · 262,144 context
- inclusionai
Ling 3.1 Flash (Free)
Not yet in MiniRouter’s catalog
inclusionai/ling-3.1-flash-free
Ling 3.1 Flash is InclusionAI's hybrid reasoning language model for coding, multi-step analysis, and tool-using agents. It has 560B total parameters with 25B activated parameters and supports long-context text workflows.
Gateway list rates: $0.00 input / $0.00 output per 1M · 262,144 context
- openai
GPT-6.1 Sol
Routable on MiniRouter
openai/gpt-6.1-sol
GPT-6.1 Sol is OpenAI's reasoning model for complex coding, computer use, and professional work. It accepts text and images, generates text, and supports a 1,050,000-token context window with up to 128,000 output tokens. Configurable reasoning effort, structured outputs, prompt caching, and Responses API tools support multi-step workflows.
Gateway list rates: $2 input / $10 output per 1M · 1,050,000 context
- openai
GPT-6.1 Sol (Fast)
Routable on MiniRouter
openai/gpt-6.1-sol-fast
GPT-6.1 Sol is OpenAI's reasoning model for complex coding, computer use, and professional work. It accepts text and images, generates text, and supports a 1,050,000-token context window with up to 128,000 output tokens. Configurable reasoning effort, structured outputs, prompt caching, and Responses API tools support multi-step workflows.
Gateway list rates: $4 input / $20 output per 1M · 1,050,000 context
- anthropic
Claude Sonnet 5.5
Routable on MiniRouter
anthropic/claude-sonnet-5.5
Claude Sonnet 5.5 is a step-change improvement over Sonnet 5 for well-scoped everyday work like building features, fixing bugs, and creating polished documents, slides, and spreadsheets. It also writes more clearly than Sonnet 5.
Gateway list rates: $2 input / $10 output per 1M · 1,000,000 context
- meituan
LongCat 2.5 Preview
Routable on MiniRouter
meituan/longcat-2.5-preview
LongCat-2.5-Preview is Meituan’s multimodal reasoning model for coding and agentic workflows, with a 1M-token context window and support for text and image inputs. It combines image understanding, visual question answering, and content summarization with code generation, code understanding, and automated programming, and supports tool calling and optional thinking for complex tasks.
Gateway list rates: $0.30 input / $1.2 output per 1M · 1,048,576 context
- stealth
Pixel Canary
Not yet in MiniRouter’s catalog
stealth/pixel-canary
Pixel Canary is an anonymous large model with strong coding capabilities and adjustable reasoning effort. It is well suited to web and mobile app development, helping developers build and refine applications across both platforms. Prompts and outputs may be retained for training by the provider.
Gateway list rates: $0.00 input / $0.00 output per 1M · 262,144 context
- alibaba
Qwen 3.8 Max Prime
Routable on MiniRouter
alibaba/qwen3.8-max-prime
Qwen 3.8 Max Prime is the high-speed edition of Qwen 3.8 Max, retaining its 2.4-trillion-parameter MoE architecture and full capabilities while delivering 1.5–2× higher output throughput. Built for coding, office automation, and long-running agent workflows, it supports autonomous multi-day development, professional knowledge work, and visual understanding with a 1M-token context window.
Gateway list rates: $4 input / $12 output per 1M · 1,000,000 context
- fireworks
Ember-1
Routable on MiniRouter
fireworks/ember-1
Ember-1 is Fireworks’ research-preview reasoning model built on Kimi K3 for coding and agentic workflows. It supports image input, function calling, and a 1M-token context window. Fireworks reports approximately 40% fewer generated tokens while maintaining comparable quality across its evaluations, helping reduce the cost of workloads with extensive reasoning.
Gateway list rates: $3 input / $15 output per 1M · 1,048,576 context
- anthropic
Claude Opus 5.5
Routable on MiniRouter
anthropic/claude-opus-5.5
Claude Opus 5.5 is a step-change improvement over Opus 5 for agentic coding, long-running agentic tasks, knowledge work, communication, and vision. It is more token-efficient per completed task and reports progress, findings, and next steps in clear language.
Gateway list rates: $4 input / $20 output per 1M · 1,000,000 context
- anthropic
Claude Opus 5.5 (Fast)
Routable on MiniRouter
anthropic/claude-opus-5.5-fast
Claude Opus 5.5 is a step-change improvement over Opus 5 for agentic coding, long-running agentic tasks, knowledge work, communication, and vision. It is more token-efficient per completed task and reports progress, findings, and next steps in clear language.
Gateway list rates: $8 input / $40 output per 1M · 1,000,000 context
- openai
GPT-6 Luna
Routable on MiniRouter
openai/gpt-6-luna
GPT-6 Luna is OpenAI's efficient reasoning model for focused, high-volume tasks. It accepts text and images, generates text, and supports a 1,050,000-token context window with up to 128,000 output tokens. Configurable reasoning effort, structured outputs, prompt caching, and Responses API tools support cost-sensitive applications, coding tasks, and automated workflows.
Gateway list rates: $0.10 input / $0.50 output per 1M · 1,050,000 context
- openai
GPT-6 Luna (Fast)
Routable on MiniRouter
openai/gpt-6-luna-fast
GPT-6 Luna is OpenAI's efficient reasoning model for focused, high-volume tasks. It accepts text and images, generates text, and supports a 1,050,000-token context window with up to 128,000 output tokens. Configurable reasoning effort, structured outputs, prompt caching, and Responses API tools support cost-sensitive applications, coding tasks, and automated workflows.
Gateway list rates: $0.20 input / $1 output per 1M · 1,050,000 context
- openai
GPT-6 Sol
Routable on MiniRouter
openai/gpt-6-sol
GPT-6 Sol is OpenAI's reasoning model for complex coding and agentic workflows. It accepts text and images, generates text, and supports a 1,050,000-token context window with up to 128,000 output tokens. Configurable reasoning effort, structured outputs, prompt caching, and Responses API tools make it suited to software development and multi-step automation.
Gateway list rates: $2 input / $10 output per 1M · 1,050,000 context
- openai
GPT-6 Sol (Fast)
Routable on MiniRouter
openai/gpt-6-sol-fast
GPT-6 Sol is OpenAI's reasoning model for complex coding and agentic workflows. It accepts text and images, generates text, and supports a 1,050,000-token context window with up to 128,000 output tokens. Configurable reasoning effort, structured outputs, prompt caching, and Responses API tools make it suited to software development and multi-step automation.
Gateway list rates: $4 input / $20 output per 1M · 1,050,000 context
- spacexai
Grok 4.7
Not yet in MiniRouter’s catalog
spacexai/grok-4.7
Grok 4.7 is SpaceXAI’s advanced AI model for coding and professional knowledge work, built to tackle complex, multi-hour tasks with improved self-verification and long-context handling. It strengthens software engineering, document creation, and presentation workflows while maintaining Grok 4.6’s speed and pricing.
Gateway list rates: $2 input / $6 output per 1M · 500,000 context
- xiaomi
MiMo V2.6 Flash
Routable on MiniRouter
xiaomi/mimo-v2.6-flash
MiMo V2.6 Flash is Xiaomi's efficient multimodal reasoning model for coding, automation, and everyday agent workflows. It accepts text, images, audio, and video within a 1M-token context window, with up to 128K tokens of output. Deep thinking, tool calling, JSON mode, and prompt caching make it suitable for applications that need multimodal understanding at a lower token cost.
Gateway list rates: $0.14 input / $0.28 output per 1M · 1,048,576 context
- xiaomi
MiMo V2.6 Pro
Routable on MiniRouter
xiaomi/mimo-v2.6-pro
MiMo V2.6 Pro is Xiaomi's flagship multimodal reasoning model for complex software engineering, long-running agent tasks, and professional workflows. It supports text, image, audio, and video inputs with a 1M-token context window and up to 128K tokens of output. Deep thinking, tool calling, JSON mode, and prompt caching support applications that combine large inputs with multi-step reasoning.
Gateway list rates: $0.435 input / $0.87 output per 1M · 1,048,576 context
- xiaomi
MiMo V2.6 Pro UltraSpeed
Routable on MiniRouter
xiaomi/mimo-v2.6-pro-ultraspeed
MiMo V2.6 Pro UltraSpeed is Xiaomi's accelerated inference offering for MiMo V2.6 Pro, designed for interactive agents and workflows where response latency matters. It combines multimodal understanding of text, images, audio, and video with a 1M-token context window and up to 128K tokens of output. It supports deep thinking, tool calling, JSON mode, and prompt caching.
Gateway list rates: $4.35 input / $8.7 output per 1M · 1,048,576 context
- stepfun
Step 5 Preview
Routable on MiniRouter
stepfun/step-5-preview
Step 5 Preview is StepFun’s flagship AI model for agentic coding, professional knowledge work, and financial analysis. Built on a sparse Mixture-of-Experts architecture with 600B total parameters and 27B active per token, it combines a 1M-token context window with vision input. From building interactive applications and debugging code to conducting research and producing analytical reports, it supports complex workflows that require sustained reasoning, tool use, and iterative refinement.
Gateway list rates: $1 input / $2.7 output per 1M · 1,000,000 context
- zai
GLM 5.3 FlashX
Routable on MiniRouter
zai/glm-5.3-flashx
GLM-5.3-FlashX is the high-speed serving option for Z.ai’s native multimodal coding model, featuring 320B total parameters, 18B activated parameters, and a 1M-token context window. Its efficient hybrid attention architecture supports visual coding, tool use, and end-to-end professional workflows across code, browsers, documents, and graphical interfaces.
Gateway list rates: $0.37 input / $1.25 output per 1M · 1,000,000 context
- alibaba
Qwen 3.8 Omni Flash
Routable on MiniRouter
alibaba/qwen3.8-omni-flash
Qwen3.8-Omni-Flash is Alibaba’s native multimodal model for understanding text, images, audio, and video and generating text. Built on Qwen3.8-Flash-Next, it supports a 1M-token context window, reasoning, and tool calling for coding, knowledge work, and multimedia analysis.
Gateway list rates: $0.15 input / $0.47 output per 1M · 1,000,000 context
- quiverai
Arrow 2
Not yet in MiniRouter’s catalog
quiverai/arrow-2
SVG generation model with agentic refinement and tool calling.
Gateway list rates: $4 input / $20 output per 1M · 131,072 context
- quiverai
Arrow 2 Telos
Not yet in MiniRouter’s catalog
quiverai/arrow-2-telos
Higher-refinement Arrow 2 variant combining Arrow's speed with frontier-model reasoning.
Gateway list rates: $6 input / $30 output per 1M · 131,072 context
Source: Vercel AI Gateway's live model catalog ↗