google/gemini-3.1-pro-preview, anthropic/claude-sonnet-4.6, google/gemini-3-pro-image-preview are Now Live
Claude Sonnet 4 will be retired by both providers: AWS Bedrock and Anthropic on June 14, 2026. We recommend upgrading to Claude Sonnet 4.6.
Claude Opus 4 will be retired by both providers: AWS Bedrock on May 31, 2026 and Anthropic on June 14, 2026. We recommend upgrading to Claude opus 4.7.
Qwen3 Max Thinking Preview will be deprecated and removed on April 24, 2026. We recommend upgrading to Qwen3-Max-Thinking
Caution: On Azure, gpt-5-chat will be deprecated and removed on May 15, 2026.
Claude 3.7 Sonnet will be retired by both providers: AWS Bedrock on April 28, 2026, and Google Cloud Vertex AI on May 11, 2026. We recommend upgrading to Claude Sonnet 4.6.
Get 5% off service fee on top-up, and exclusive gifts for referring friends
A billing anomaly was identified. Affected users will receive compensation. Read more
Caution: On Azure, gpt-5-chat will be deprecated and removed on May 15, 2026.
A billing anomaly was identified. Affected users will receive compensation. Read more
- 1
- 2
- 3
- 4
- 5
- 6
- 7
- 8
Studio
Models
PricingDevelopers Analytics About Us
Sign In
Input Modalities
Text
Image
File
Audio
Video
Output Modalities
Text
Image
File
Audio
Video
Embeddings
Context Length
4K64K1M
Maker
Alibaba
Anthropic
Baidu
Show More
Providers
OpenAI
Anthropic
Tbox
Show More
Supported Parameters
max_completion_tokens
temperature
top_p
Show More
Supported Protocol
OpenAI Chat Completions
OpenAI Responses
OpenAI Embeddings
Anthropic Messages
Google Gemini
Google Imagen
Google Video
Reasoning
No Reasoning
Toggleable Reasoning
Always-On Reasoning
Models
157 models
Newest
zenmux/auto
ZenMux's automatic routing feature selects the most cost-effective and high-performing AI models based on your query.
Input type
Output Type
Input-
Output-
Context-
Max Output-
qwen/qwen3.6-max-preview
5.25Mtokens
Qwen3.6-Max-Preview is a proprietary frontier model from Alibaba Cloud built on a sparse mixture-of-experts architecture with approximately 1 trillion total parameters. It is optimized for agentic coding, tool use, and long-context reasoning, supporting a 262K token context window. The model includes an integrated thinking mode that preserves reasoning traces across multi-turn conversations and supports structured output and function calling. Access is available exclusively through the Alibaba Cloud Model Studio and Qwen Studio APIs; no open weights are provided.
Input type
Output Type
Input1.3-2/M tokens
Output7.8-12/M tokens
Context262.14K
Max Output65.54K
Available on 1 provider
Apr 30, 2026 6:00 PM-
alibaba/happyhorse-1.0
500.00Ktokens
Happy Horse 1.0 is described as an open-source state-of-the-art AI video generator with native joint audio-video generation — meaning the Happy Horse AI video model produces video frames and the corresponding audio track (dialogue, ambient sound, Foley) together in a single forward pass, rather than generating silent video and dubbing it afterward. According to community-compiled architecture notes, the model is built around a 15-billion-parameter unified self-attention Transformer that processes text, image, video, and audio tokens within a single token sequence. It is reportedly built without dedicated cross-attention branches and without a separate audio module. Combined with DMD-2 distillation, the distilled variant is reported to generate 1080p video in roughly 38 seconds on an NVIDIA H100, using only 8 denoising steps without classifier-free guidance.
Input type
Output Type
Input-
Output-
Context-
Max Output-
Available on 1 provider
Apr 30, 2026 6:00 PM-
openai/gpt-5.5
3.42Btokens
GPT‑5.5 understands what you’re trying to do faster and can carry more of the work itself. It excels at writing and debugging code, researching online, analyzing data, creating documents and spreadsheets, operating software, and moving across tools until a task is finished. Instead of carefully managing every step, you can give GPT‑5.5 a messy, multi-part task and trust it to plan, use tools, check its work, navigate through ambiguity, and keep going.
Input type
Output Type
Input5-10/M tokens
Output30-45/M tokens
Context1.05M
Max Output128.00K
Available on 2 providers
Apr 30, 2026 6:00 PM-
openai/gpt-5.5-pro
79.26Mtokens
GPT-5.5 Pro is OpenAI’s high-capability model optimized for deep reasoning and accuracy on complex, high-stakes workloads. It features a 1M+ token context window (922K input, 128K output) with support for text and image inputs, and is designed for long-horizon problem solving, agentic coding, and precise execution across multi-step workflows.
Input type
Output Type
Input30-60/M tokens
Output180-270/M tokens
Context1.05M
Max Output128.00K
Available on 1 provider
Apr 30, 2026 6:00 PM-
deepseek/deepseek-v4-flash
1.93Btokens
DeepSeek-V4-Flash is the efficiency-oriented variant of the DeepSeek V4 series, released as a preview and open-sourced alongside the flagship V4-Pro. It is designed for developers who need the V4 generation's long-context and reasoning capability at a faster, more economical API tier. Compared with V4-Pro, V4-Flash uses smaller total parameters and active parameters, resulting in faster response times and lower API cost. It retains reasoning capability close to V4-Pro and matches V4-Pro on simple agent tasks, with a measurable gap appearing only on the most demanding agent workflows. World knowledge is slightly below V4-Pro but remains competitive within the open-source landscape.
Like V4-Pro, V4-Flash inherits the new attention mechanism built on token-dimension compression and DeepSeek Sparse Attention (DSA), supports a 1M context window as standard, and offers both thinking and non-thinking modes with a reasoning_effort parameter (high / max)._
Note on migration: the legacy model names deepseek-chat and deepseek-reasoner currently route to V4-Flash in non-thinking and thinking mode respectively, and will be retired on 2026-07-24. Existing integrations should migrate to the explicit model names deepseek-v4-flash or deepseek-v4-pro. The model is accessible via OpenAI ChatCompletions and Anthropic interfaces, with weights open-sourced on Hugging Face and ModelScope.
Input type
Output Type
Input0.14/M tokens
Output0.28/M tokens
Context1000.00K
Max Output384.00K
Available on 2 providers
Apr 30, 2026 6:00 PM-
DeepSeek: DeepSeek V4 Flash (Free)
deepseek/deepseek-v4-flash-free
3.21Btokens
Free
Rate Limit
DeepSeek latest model: DeepSeek V4 Flash.
Input type
Output Type
Input0/M tokens
Output0/M tokens
Context1000.00K
Max Output384.00K
Available on 1 provider
Apr 30, 2026 6:00 PM-
deepseek/deepseek-v4-pro
17.52Btokens
75% OFF
The deepseek-v4-pro model is currently offered through the official DeepSeek direct channel at a limited-time 75% discount, valid until 2026/05/31 15:59 UTC. DeepSeek-V4-Pro is the flagship model of the DeepSeek V4 series, released as a preview and open-sourced alongside its efficiency-tier sibling V4-Flash. It is positioned as the performance-first tier in the V4 lineup, designed to push agentic capability, world knowledge, and reasoning performance to levels competitive with leading closed-source models.
The model introduces a new attention mechanism that performs compression along the token dimension, combined with DeepSeek Sparse Attention (DSA). The design makes a 1M context window the default across all official DeepSeek services while substantially reducing compute and memory overhead compared with prior approaches.
DeepSeek-V4-Pro has been specifically adapted and optimized for mainstream agent products such as Claude Code, OpenClaw, OpenCode, and CodeBuddy. According to DeepSeek, V4-Pro reaches the top tier among open-source models on Agentic Coding benchmarks, and is currently used internally at DeepSeek as the default Agentic Coding model—with internal evaluation reports describing a usage experience above Sonnet 4.5 and delivery quality close to Opus 4.6 in non-thinking mode, while still trailing Opus 4.6 in thinking mode. On world knowledge, it leads open-source models and trails only Gemini-Pro-3.1; on math, STEM, and competitive coding, it delivers results on par with top closed-source models.
The model supports both thinking and non-thinking modes, with the thinking mode exposing a reasoning_effort parameter (high / max) for complex agent workflows. It is available via App, Web, and API through the OpenAI ChatCompletions and Anthropic interfaces, with weights open-sourced on Hugging Face and ModelScope._
Input type
Output Type
Input0.435/M tokens
Output0.87/M tokens
Context1000.00K
Max Output384.00K
Available on 2 providers
Apr 30, 2026 6:00 PM-
DeepSeek: DeepSeek V4 Pro (Free)
deepseek/deepseek-v4-pro-free
17.49Btokens
Free
Rate Limit
DeepSeek-V4-Pro is the flagship model of the DeepSeek V4 series, released as a preview and open-sourced alongside its efficiency-tier sibling V4-Flash. It is positioned as the performance-first tier in the V4 lineup, designed to push agentic capability, world knowledge, and reasoning performance to levels competitive with leading closed-source models.
Input type
Output Type
Input0/M tokens
Output0/M tokens
Context1000.00K
Max Output384.00K
Available on 1 provider
Apr 30, 2026 6:00 PM-
sapiens-ai/agnes-1.5-flash
116.32Mtokens
Agnes-1.5-Flash is a model derived from the Agnes-1.5-Pro architecture, optimized for high efficiency without compromising performance. Through advanced quantization techniques, it delivers performance comparable to significantly larger models while maintaining much lower compute requirements and latency. It retains strong capabilities in conversation and content generation, making it ideal for scalable, cost-sensitive, and real-time applications.
Input type
Output Type
Input0.07/M tokens
Output0.15/M tokens
Context256.00K
Max Output65.54K
Available on 1 provider
Apr 30, 2026 6:00 PM-
inclusionai/ling-2.6-1t
3.40Btokens
Free
Ling-2.6-1T is an instant (instruct) model from inclusionAI and the company’s trillion-parameter flagship, designed for real-world agents that require fast execution and high efficiency at scale. It uses a “fast thinking” approach to reduce costs to roughly a quarter of comparable models while maintaining top-tier performance.
The model achieves state-of-the-art results on benchmarks such as AIME26 and SWE-bench Verified, and is well suited for advanced coding, complex reasoning, and large-scale agent workflows where both capability and efficiency are critical.
Input type
Output Type
Input0/M tokens
Output0/M tokens
Context262.14K
Max Output32.77K
Available on 1 provider
Apr 30, 2026 6:00 PM-
tencent/hy3-preview
1.39Btokens
Hy3 Preview is a high-efficiency Mixture-of-Experts model from Tencent designed for agentic workflows and production use. It supports configurable reasoning levels across disabled, low, and high modes, allowing it to balance speed and depth depending on the task, while delivering strong code generation and reliable performance across multi-step, real-world workflows.
Input type
Output Type
Input0.172-0.286/M tokens
Output0.572-1.144/M tokens
Context262.14K
Max Output131.07K
Available on 1 provider
Apr 30, 2026 6:00 PM-
xiaomi/mimo-v2.5
80.61Mtokens
MiMo-V2.5 is a native omnimodal model by Xiaomi. It delivers Pro-level agentic performance at roughly half the inference cost, while surpassing MiMo-V2-Omni in multimodal perception across image and video understanding tasks. Its 1M context window supports complete documents, extended conversations, and complex task contexts in a single pass, making it ideal for integration with agent frameworks where strong reasoning, rich perception, and cost efficiency all matter.
Input type
Output Type
Input0.4-0.8/M tokens
Output2-4/M tokens
Context1.05M
Max Output131.07K
Available on 1 provider
Apr 30, 2026 6:00 PM-
xiaomi/mimo-v2.5-pro
354.27Mtokens
MiMo-V2.5-Pro is Xiaomi’s flagship model, delivering strong performance in general agentic capabilities, complex software engineering, and long-horizon tasks, with top rankings on benchmarks such as ClawEval, GDPVal, and SWE-bench Pro. It can independently and autonomously complete professional tasks that would take human experts days or weeks, involving more than a thousand tool calls. Its context length of up to 1M makes it well suited for integration with a wide range of agent frameworks.
Input type
Output Type
Input1-2/M tokens
Output3-6/M tokens
Context1.05M
Max Output131.07K
Available on 1 provider
Apr 30, 2026 6:00 PM-
baidu/ernie-image-turbo
60.00Ktokens
ERNIE-Image-Turbo is an open text-to-image generation model developed by the ERNIE-Image team at Baidu. It is the distilled release of ERNIE-Image, built on the same single-stream Diffusion Transformer (DiT) family and designed for fast generation with strong fidelity in only 8 inference steps. The model retains strong controllability in practical generation scenarios where accurate content realization matters as much as aesthetics. In particular, ERNIE-Image-Turbo remains strong on complex instruction following, text rendering, and structured image generation, making it well suited for posters, comics, multi-panel layouts, and other content creation tasks that require both visual quality and efficiency. It also supports a broad range of visual styles, including realistic photography, design-oriented imagery, and stylized aesthetic outputs.
Input type
Output Type
Input-
Output-
Context-
Max Output-
Available on 1 provider
Apr 30, 2026 6:00 PM-
inclusionai/ling-2.6-flash
3.90Mtokens
Ling-2.6-flash is a 100B-parameter text model focused on intelligence efficiency, delivering strong performance while minimizing token usage. It supports a 256K context window with up to 32K output tokens, function calling, structured output, and prompt caching. It is particularly well-suited for code completion and debugging, rapid document processing, and lightweight agent interactions.
Input type
Output Type
Input0.1/M tokens
Output0.3/M tokens
Context262.14K
Max Output32.77K
Available on 2 providers
Apr 30, 2026 6:00 PM-
openai/gpt-image-2
102.40Mtokens
GPT Image 2 — OpenAI's next-generation image generation model. It delivers rapid, high-quality image generation and editing capabilities, with support for flexible image dimensions and high-fidelity image inputs for seamless creative workflows.
Input type
Output Type
Input5/M tokens
Output-
Context-
Max Output-
Available on 2 providers
Apr 30, 2026 6:00 PM-
moonshotai/kimi-k2.6
2.09Btokens
Kimi’s most intelligent model to date, achieving open-source SoTA performance in Agent, code, visual understanding, and a range of general intelligent tasks. It is also Kimi’s most versatile model to date, featuring a native multimodal architecture that supports both visual and text input, thinking and non-thinking modes, and dialogue and Agent tasks. Context 256k
Input type
Output Type
Input0.95/M tokens
Output4/M tokens
Context262.14K
Max Output262.14K
Available on 2 providers
Apr 30, 2026 6:00 PM-
skyreels/skyreels-v4
240.00Ktokens
SkyReels V4 is a unified multimodal video foundation model that generates, edits, and inpaints video and audio simultaneously using a dual-stream architecture with shared text encoding and efficient high-resolution processing.
Input type
Output Type
Input-
Output0.14/seconds
Context1.28K
Max Output1.28K
Available on 1 provider
Apr 30, 2026 6:00 PM-
anthropic/claude-opus-4.7
134.90Btokens
Opus 4.7 is the next generation of Anthropic's Opus family, built for long-running, asynchronous agents. Building on the coding and agentic strengths of Opus 4.6, it delivers stronger performance on complex, multi-step tasks and more reliable agentic execution across extended workflows. It is especially effective for asynchronous agent pipelines where tasks unfold over time - large codebases, multi-stage debugging, and end-to-end project orchestration.
Beyond coding, Opus 4.7 brings improved knowledge work capabilities - from drafting documents and building presentations to analyzing data. It maintains coherence across very long outputs and extended sessions, making it a strong default for tasks that require persistence, judgment, and follow-through.
Input type
Output Type
Input5/M tokens
Output25/M tokens
Context1000.00K
Max Output128.00K
Available on 3 providers
Apr 30, 2026 6:00 PM-
ByteDance: Doubao-Seedance-2.0
bytedance/doubao-seedance-2.0
82.38Mtokens
Seedance 2.0 delivers ultra-realistic, highly stable audiovisual output: with outstanding motion stability and fine visual detail, it produces imagery with a live-action-quality look and an almost indistinguishable blend of real and virtual visual impact. It can handle complex scenes with ease—vividly recreating everything from subtle micro-expressions and intense physical confrontations to dynamic, high-energy song-and-dance performances. It also comes with professional camera movement, multi-shot storytelling, and text-to-video generation capabilities, enhancing narrative tension. Audio and visuals are generated natively in sync, accurately matching the visuals with rich sound effects. It supports performances in multiple languages, accents, and dialects, giving videos exceptional completeness and immersion. Seedance 2.0 is deeply optimized for three key scenarios: commercial advertising, film/TV production, and social media marketing. With industrial-grade generation quality, the hit rate for successful generations is significantly improved, lowering the barrier and cost of producing high-quality content, streamlining the workflow from idea to final cut, and delivering substantial efficiency gains for the industry.
Input type
Output Type
Input-
Output4.1-6.74/M tokens
Context12.80K
Max Output12.80K
Available on 1 provider
Apr 30, 2026 6:00 PM-
google/veo-3.1-fast-generate-001
32.00Ktokens
Veo 3.1 Fast is a speed-optimized variant of Google DeepMind's flagship video generation model. It is designed to generate high-quality video significantly faster and at a lower cost than the standard Veo 3.1 Quality model, making it ideal for rapid prototyping and high-volume content creation.
Input type
Output Type
Input-
Output0.15-0.35/seconds
Context-
Max Output-
Available on 1 provider
Apr 30, 2026 6:00 PM-
google/veo-3.1-lite-generate-001
-tokens
Veo 3.1 Lite is Google DeepMind's most cost-efficient AI video generation model, released on March 31, 2026. It is designed to provide professional-grade video capabilities at a significantly lower price point, making it ideal for developers and content teams who need to scale high-volume video production
Input type
Output Type
Input-
Output0.05-0.08/seconds
Context-
Max Output-
Available on 1 provider
Apr 30, 2026 6:00 PM-
z-ai/glm-5.1
44.98Btokens
GLM-5.1 is Zai’s new-generation flagship foundation model, designed for Agentic Engineering, capable of providing reliable productivity in complex system engineering and long-range Agent tasks. In terms of Coding and Agent capabilities, GLM-5 has achieved state-of-the-art (SOTA) performance in open source, with its usability in real programming scenarios approaching that of Claude Opus 4.5.
Input type
Output Type
Input0.8781-1.1709/M tokens
Output3.5126-4.098/M tokens
Context200.00K
Max Output128.00K
Available on 3 providers
Apr 30, 2026 6:00 PM-
qwen/qwen3.6-plus
6.84Btokens
Qwen3.6-Plus is Alibaba’s next-generation Qwen large language model released on April 2, 2026. Compared with version 3.5, Qwen 3.6 has made significant overall performance improvements and has exhibited remarkably strong agent-oriented programming capabilities.
Input type
Output Type
Input0.5-2/M tokens
Output3-6/M tokens
Context1000.00K
Max Output64.00K
Available on 1 provider
Apr 30, 2026 6:00 PM-
z-ai/glm-5v-turbo
418.86Mtokens
GLM-5V-Turbo is Zhipu’s first multimodal coding base model, designed for vision-based programming tasks. It can natively process multimodal inputs such as images, video, and text, and excels at long-horizon planning, complex programming, and action execution. It is deeply adapted to Agent workflows and can closely collaborate with Agents like Claude Code and OpenClaw to complete a full closed loop of “understanding the environment → planning actions → executing tasks.
Input type
Output Type
Input0.726-1.0165/M tokens
Output3.1946-3.7754/M tokens
Context200.00K
Max Output128.00K
Available on 1 provider
Apr 30, 2026 6:00 PM-
kuaishou/kat-coder-pro-v2
106.74Mtokens
KAT-Coder-Pro V2 is theKwaiKAT’s Latest Flagship Agentic Coding Model. Engineered for multi-scaffold generalization, the model is natively compatible with over 10 mainstream AI coding tools—such as Claude Code, Cline, Kilo, and OpenCode, offering unparalleled flexibility. Through dedicated full-pipeline optimization for OpenClaw, KAT-Coder-Pro V2 is trained from the ground up to master complex, real-world application workflows with ease. Beyond logic, KAT-Coder-Pro V2 achieves a breakthrough in frontend aesthetic generation. In Landing Page and PPT scenarios, the model delivers a paradigm shift in user experience: No Structured Spec Needed: Users no longer need to provide rigid design specifications. Simply describe what you want in plain language to receive production-grade, high-quality output that rivals structured design inputs. Mass Market Expansion: This evolution expands the model's service boundary from the 1% of power users to hundreds of millions of ordinary users, truly democratizing professional-grade creation.
Input type
Output Type
Input0.3/M tokens
Output1.2/M tokens
Context256.00K
Max Output80.00K
Available on 1 provider
Apr 30, 2026 6:00 PM-
sapiens-ai/agnes-1.5-lite
2.18Btokens
Agnes-1.5-Lite is a model derived from the Agnes-1.5-Pro architecture, optimized for high efficiency without compromising performance. Through advanced quantization techniques, it delivers performance comparable to significantly larger models while maintaining much lower compute requirements and latency. It retains strong capabilities in conversation and content generation, making it ideal for scalable, cost-sensitive, and real-time applications.
Input type
Output Type
Input0.07/M tokens
Output0.15/M tokens
Context256.00K
Max Output65.54K
Available on 1 provider
Apr 30, 2026 6:00 PM-
sapiens-ai/agnes-video-v1.2
-tokens
Agnes-Video-V1.2 is a cinematic-grade video generation model that delivers high-fidelity visuals and fully synchronized audio in a single pass. It natively generates aligned dialogue and rich environmental sound alongside film-quality imagery, ensuring coherence between speech, motion, and scene dynamics. Designed for storytelling, marketing, and immersive experiences, it transforms simple prompts into production-ready videos with cinematic realism—eliminating the need for separate audio pipelines or post-processing.
Input type
Output Type
Input-
Output0.018/seconds
Context-
Max Output-
Available on 1 provider
Apr 30, 2026 6:00 PM-
sapiens-ai/agnes-1.5-pro
4.48Btokens
Agnes-1.5-Pro is a large-scale text foundation model built with tens of billions of parameters, delivering strong capabilities in natural language understanding and generation. It demonstrates excellent performance across complex semantic modeling, multi-turn dialogue, and reasoning tasks.
Through continuous optimization, Agnes leverages parameter-efficient fine-tuning techniques and task-driven training strategies to enhance its adaptability to real-world business scenarios. In addition, it incorporates native tool calling capabilities, enabling the model not only to understand and generate language, but also to autonomously select and invoke external tools based on task requirements—bridging the gap between comprehension and execution.
Input type
Output Type
Input0.16/M tokens
Output0.8/M tokens
Context256.00K
Max Output256.00K
Available on 1 provider
Apr 30, 2026 6:00 PM-
xiaomi/mimo-v2-omni
1.18Btokens
MiMo-V2-Omni is purpose-built for complex multimodal interaction and execution scenarios in the real world. We constructed a fully omnimodal foundation from the ground up — one that natively integrates text, vision, and speech — and deeply couples perception with action through a unified architecture. This not only breaks free from the limitations of conventional models that prioritize understanding over execution, but also equips the model with native capabilities spanning multimodal perception, tool invocation, function execution, and GUI manipulation. With seamless integration into mainstream agent frameworks, MiMo-V2-Omni bridges the gap between comprehension and control, dramatically lowering the barrier to deploying full-stack multimodal agents in production.
Input type
Output Type
Input0.4/M tokens
Output2/M tokens
Context265.00K
Max Output265.00K
Available on 1 provider
Apr 30, 2026 6:00 PM-
xiaomi/mimo-v2-pro
7.82Btokens
Xiaomi MiMo-V2-Pro is purpose-built for high-intensity agent workflows in real-world scenarios. It features over 1 trillion total parameters (with 42B active parameters), employs an innovative hybrid attention architecture, and supports an ultra-long context window of 1M tokens. Building upon this powerful model foundation, we've continuously scaled compute across a broader spectrum of agent scenarios, further expanding the intelligent action space and achieving significant generalization — from Coding to Claw.
Input type
Output Type
Input1-2/M tokens
Output3-6/M tokens
Context1000.00K
Max Output256.00K
Available on 2 providers
Apr 30, 2026 6:00 PM-
minimax/minimax-m2.7
9.69Btokens
M2.7 delivers outstanding performance in real-world software engineering, including end-to-end complete project delivery, log analysis and bug triaging, code security, machine learning, and more. On the benchmark SWE-Pro, M2.7 scores 56.22%, nearly matching the level of Opus. This capability also extends to end-to-end complete project delivery scenarios (VIBE-Pro 55.6%) and deep understanding of complex engineering systems on Terminal Bench 2 (57.0%).
In the professional office domain, we have improved the model's specialized knowledge and task delivery capabilities across various fields. On GDPval-AA, its ELO score is 1495, the highest among open-source models. M2.7's ability to perform complex editing in the Office suite (Excel/PPT/Word) has significantly improved, enabling better multi-round revisions and high-fidelity editing. M2.7 is capable of interacting with complex environments. Across 40 complex skills (> 2000 tokens) cases, M2.7 still maintains a 97% skill adherence rate. In OpenClaw usage, M2.7 has shown significant improvement compared to M2.5, scoring close to the latest Sonnet 4.6 in the MMClaw evaluation.
M2.7 possesses excellent identity retention capabilities and emotional intelligence. Beyond productivity use cases, it also opens up space for innovation in interactive entertainment scenarios.
Input type
Output Type
Input0.3/M tokens
Output1.2/M tokens
Context204.80K
Max Output131.07K
Available on 2 providers
Apr 30, 2026 6:00 PM-
MiniMax: MiniMax M2.7 highspeed
minimax/minimax-m2.7-highspeed
795.48Mtokens
M2.7 highspeed: Same performance, faster, more agile
Input type
Output Type
Input0.611/M tokens
Output2.4439/M tokens
Context204.80K
Max Output131.07K
Available on 1 provider
Apr 30, 2026 6:00 PM-
openai/gpt-5.4-mini
5.98Btokens
GPT-5.4 mini brings the strengths of GPT-5.4 to a faster, more efficient model designed for high-volume workloads.
Input type
Output Type
Input0.75/M tokens
Output4.5/M tokens
Context400.00K
Max Output128.00K
Available on 2 providers
Apr 30, 2026 6:00 PM-
openai/gpt-5.4-nano
468.02Mtokens
GPT-5.4 nano is designed for tasks where speed and cost matter most like classification, data extraction, ranking, and sub-agents.
Input type
Output Type
Input0.2/M tokens
Output1.25/M tokens
Context400.00K
Max Output128.00K
Available on 2 providers
Apr 30, 2026 6:00 PM-
inclusionai/llada2.1-flash
119.66Mtokens
LLaDA 2.1 is a diffusion language model in the LLaDA series, enhanced with editing capabilities. It delivers strong task performance while significantly improving inference speed.
Input type
Output Type
Input0.28/M tokens
Output2.85/M tokens
Context32.00K
Max Output32.00K
Available on 1 provider
Apr 30, 2026 6:00 PM-
z-ai/glm-5-turbo
4.18Btokens
GLM-5 Turbo is a new model from Z.ai designed for fast inference and strong performance in agent-driven environments such as OpenClaw scenarios. It is deeply optimized for real-world agent workflows involving long execution chains, with improved complex instruction decomposition, tool use, scheduled and persistent execution, and overall stability across extended tasks.
Input type
Output Type
Input0.73-1.02/M tokens
Output3.19-3.77/M tokens
Context200.00K
Max Output128.00K
Available on 1 provider
Apr 30, 2026 6:00 PM-
x-ai/grok-4.2-fast
1.34Btokens
Grok 4.20 Beta features a powerful 4-agent, multi-agent architecture (Grok, Harper, Benjamin, Lucas) that operates in parallel to research, code, and reason through complex tasks. It offers 2M token context, reduced hallucinations, better LaTeX for STEM, improved image/video understanding, and 95% MMLU-Pro accuracy, targeting superior reasoning
Input type
Output Type
Input2-4/M tokens
Output6-12/M tokens
Context2.00M
Max Output30.00K
Available on 1 provider
Apr 30, 2026 6:00 PM-
xAI: Grok 4.2 Fast Non Reasoning
x-ai/grok-4.2-fast-non-reasoning
218.53Mtokens
Input type
Output Type
Input2-4/M tokens
Output6-12/M tokens
Context2.00M
Max Output30.00K
Available on 1 provider
Apr 30, 2026 6:00 PM-
openai/gpt-5.4
83.40Btokens
GPT‑5.4, released March 5, 2026, is OpenAI’s newest frontier model aimed at complex professional work. It’s available in the API as gpt-5.4 and in ChatGPT as GPT‑5.4 Thinking, with a higher-end GPT‑5.4 Pro tier for maximum performance. It supports adjustable deliberation through reasoning.effort, letting developers trade latency/cost for deeper reasoning. The model accepts text and image inputs and produces text output, with a 1,050,000‑token context window and up to 128,000 output tokens; its published knowledge cutoff is Aug 31, 2025. OpenAI also published a system card describing new cybersecurity-focused mitigations for the Thinking variant.
Input type
Output Type
Input2.5-5/M tokens
Output15-22.5/M tokens
Context1.05M
Max Output128.00K
Available on 2 providers
Apr 30, 2026 6:00 PM-
openai/gpt-5.4-pro
371.13Mtokens
GPT-5.4 Pro is OpenAI's most advanced model, building on GPT-5.4's unified architecture with enhanced reasoning capabilities for complex, high-stakes tasks. It features a 1M+ token context window (922K input, 128K output) with support for text and image inputs. Optimized for step-by-step reasoning, instruction following, and accuracy, GPT-5.4 Pro excels at agentic coding, long-context workflows, and multi-step problem solving.
Input type
Output Type
Input30-60/M tokens
Output180-270/M tokens
Context1.05M
Max Output128.00K
Available on 2 providers
Apr 30, 2026 6:00 PM-
Google: Gemini 3.1 Flash Lite Preview
google/gemini-3.1-flash-lite-preview
15.07Btokens
Gemini 3.1 Flash-Lite is Google's most cost-efficient Gemini model, optimized for low latency use cases for high-volume, cost-sensitive LLM traffic. It provides a significant quality increase over Gemini 2.0/2.5 Flash-Lite models, matching Gemini 2.5 Flash performance across key capability areas.
Input type
Output Type
Input0.25/M tokens
Output1.5/M tokens
Context1.05M
Max Output65.53K
Available on 2 providers
Apr 30, 2026 6:00 PM-
openai/gpt-5.3-chat
459.50Mtokens
GPT-5.3 Chat points to the GPT-5.3 Instant snapshot currently used in ChatGPT. We recommend GPT-5.2 for API usage, but feel free to use this GPT-5.3 Chat model to test our latest improvements for chat use cases.
Input type
Output Type
Input1.75/M tokens
Output14/M tokens
Context128.00K
Max Output16.38K
Available on 2 providers
Apr 30, 2026 6:00 PM-
qwen/qwen-image-2.0
678.00Ktokens
Qwen-Image-2.0 is an image generation model launched by Qwen AI.
Input type
Output Type
Input-
Output0.0289/counts
Context-
Max Output-
Available on 1 provider
Apr 30, 2026 6:00 PM-
qwen/qwen-image-2.0-pro
1.03Mtokens
Qwen-Image-2.0-Pro is an image generation model launched by Qwen AI.
Input type
Output Type
Input-
Output0.073/counts
Context-
Max Output-
Available on 1 provider
Apr 30, 2026 6:00 PM-
ByteDance: Doubao-Seedream-5.0-lite
bytedance/doubao-seedream-5.0-lite
844.00Ktokens
“Doubao-Seedream-5.0-lite is ByteDance’s latest image-generation model. For the first time, it integrates an online retrieval feature, enabling it to incorporate real-time web information and improve the timeliness of generated images. At the same time, the model’s intelligence has been further upgraded, allowing it to accurately interpret complex instructions and visual content. In addition, the model has been enhanced in terms of the breadth of world knowledge, reference consistency, and generation quality in professional scenarios, better meeting enterprise-level visual creation needs.”
Input type
Output Type
Input-
Output0.032/counts
Context-
Max Output-
Available on 1 provider
Apr 30, 2026 6:00 PM-
Google: Nano Banana 2 (Gemini 3.1 Flash Image Preview)
google/gemini-3.1-flash-image-preview
743.23Mtokens
Designed for speed and efficiency, the Gemini 3.1 Flash Image generation model is effective for quick, interactive responses and high throughput.
Preview models may change before becoming stable and have more restrictive rate limits.
Input type
Output Type
Input0.5/M tokens
Output3/M tokens
Context65.54K
Max Output32.77K
Available on 2 providers
Apr 30, 2026 6:00 PM-
qwen/qwen3.5-flash
3.18Btokens
Qwen3.5 native vision-language Flash model series, based on a hybrid architecture design, integrates linear attention mechanisms with sparse Mixture-of-Experts (MoE) models, achieving higher inference efficiency. The model's performance has made a leap forward compared to the 3 series in both pure text and multimodal capabilities; it offers fast response times, balancing inference speed and performance.
Input type
Output Type
Input0.1/M tokens
Output0.4/M tokens
Context1.02M
Max Output1.02M
Available on 1 provider
Apr 30, 2026 6:00 PM-
qwen/qwen3.6-flash
2.47Mtokens
Qwen3.6 Flash is a fast, efficient language model from Alibaba's Qwen 3.6 series. It supports text, image, and video input with a 1M token context window. Tiered pricing kicks in above 256K tokens. Prompt caching is supported, with both explicit cache read and cache creation pricing.
Input type
Output Type
Input0.25-1/M tokens
Output1.5-4/M tokens
Context1000.00K
Max Output65.54K
Available on 1 provider
Apr 30, 2026 6:00 PM-
Google: Gemini 3.1 Pro Preview
google/gemini-3.1-pro-preview
25.93Btokens
Gemini 3.1 Pro is the next generation in the Gemini series of models, a suite of highly-capable, natively multimodal, reasoning models. Gemini 3 Pro is now Google’s most advanced model for complex tasks, and can comprehend vast datasets, challenging problems from different information sources, including text, audio, images, video, and entire code repositories
Input type
Output Type
Input2-4/M tokens
Output12-18/M tokens
Context1.05M
Max Output65.53K
Available on 2 providers
Apr 30, 2026 6:00 PM-
anthropic/claude-sonnet-4.6
440.15Btokens
Sonnet 4.6 is Anthropic's most capable Sonnet-class model yet, with frontier performance across coding, agents, and professional work. It excels at iterative development, complex codebase navigation, end-to-end project management with memory, polished document creation, and confident computer use for web QA and workflow automation.
Input type
Output Type
Input3/M tokens
Output15/M tokens
Context1000.00K
Max Output64.00K
Available on 3 providers
Apr 30, 2026 6:00 PM-
qwen/qwen3.5-plus
4.62Btokens
Qwen3.5 Native Visual Language Series Plus model, based on a hybrid architecture design, integrates linear attention mechanisms with sparse mixture-of-experts models, achieving higher reasoning efficiency. In multiple task evaluations, the 3.5 series has demonstrated outstanding performance comparable to current top-tier cutting-edge models, with model performance showing leaps and bounds of progress compared to the 3 series in both pure text and multimodal aspects.
Input type
Output Type
Input0.4-0.5/M tokens
Output2.4-3/M tokens
Context1000.00K
Max Output64.00K
Available on 1 provider
Apr 30, 2026 6:00 PM-
ByteDance: Doubao-Seed-2.0-Code
bytedance/doubao-seed-2.0-code
75.53Mtokens
Doubao-Seed-2.0-Code is optimized for enterprise-level programming needs. Building on the excellent Agent and VLM capabilities of Seed 2.0, it has significantly enhanced code capabilities. Not only does it have outstanding front-end performance, but it has also been specially optimized for the multi-language coding requirements common in enterprises, making it suitable for integration with various AI programming tools.
Input type
Output Type
Input0.45-1.34/M tokens
Output2.24-6.71/M tokens
Context256.00K
Max Output32.00K
Available on 1 provider
Apr 30, 2026 6:00 PM-
ByteDance: Doubao-Seed-2.0-lite
bytedance/doubao-seed-2.0-lite
387.96Mtokens
Doubao-Seed-2.0-lite is a balanced model designed for high-frequency enterprise scenarios, balancing performance and cost, and its overall capabilities surpass those of its predecessor, Doubao-Seed-1.8. It excels in production-oriented tasks such as unstructured information processing, content creation, search and recommendation, and data analysis, supporting long contexts, multi-source information fusion, multi-step instruction execution, and high-fidelity structured output. It significantly optimizes costs while ensuring stable performance.
Input type
Output Type
Input0.09-0.25/M tokens
Output0.51-1.51/M tokens
Context256.00K
Max Output32.00K
Available on 1 provider
Apr 30, 2026 6:00 PM-
ByteDance: Doubao-Seed-2.0-mini
bytedance/doubao-seed-2.0-mini
461.63Mtokens
Doubao-Seed-2.0-mini is designed for low-latency, high-concurrency, and cost-sensitive scenarios, emphasizing rapid response and flexible inference deployment. Its model performance is comparable to Doubao-Seed-1.6. It supports 256k context, four levels of think length, and multimodal understanding, making it suitable for lightweight tasks where cost and speed are paramount.
Input type
Output Type
Input0.03-0.12/M tokens
Output0.28-1.12/M tokens
Context256.00K
Max Output32.00K
Available on 1 provider
Apr 30, 2026 6:00 PM-
ByteDance: Doubao-Seed-2.0-pro
bytedance/doubao-seed-2.0-pro
227.83Mtokens
Doubao-Seed-2.0-pro is a flagship, all-around general-purpose model designed for complex reasoning and long-chain task execution scenarios in the Agent era. It emphasizes multimodal understanding, long-context reasoning, structured generation, and tool-enhanced execution. Its capabilities in executing complex instructions and multiple constraints are outstanding, stably handling scenarios such as multi-step complex planning, complex graph and text reasoning, video content understanding, and high-difficulty analysis.
Input type
Output Type
Input0.45-1.34/M tokens
Output2.24-6.71/M tokens
Context256.00K
Max Output32.00K
Available on 1 provider
Apr 30, 2026 6:00 PM-
minimax/minimax-m2.5
6.13Btokens
MiniMax-M2.5 is a lightweight, state-of-the-art large language model optimized for coding, agentic workflows, and modern application development. With only 10 billion activated parameters, it delivers a major jump in real-world capability while maintaining exceptional latency, scalability, and cost efficiency.
Compared to its predecessor, M2.1 delivers cleaner, more concise outputs and faster perceived response times. It shows leading multilingual coding performance across major systems and application languages, achieving 49.4% on Multi-SWE-Bench and 72.5% on SWE-Bench Multilingual, and serves as a versatile agent “brain” for IDEs, coding tools, and general-purpose assistance.
To avoid degrading this model's performance, MiniMax highly recommends preserving reasoning between turns.
Input type
Output Type
Input0.3/M tokens
Output1.2/M tokens
Context204.80K
Max Output131.07K
Available on 2 providers
Apr 30, 2026 6:00 PM-
MiniMax: MiniMax M2.5 highspeed
minimax/minimax-m2.5-lightning
971.02Mtokens
M2.5 highspeed: Same performance, faster, more agile
Input type
Output Type
Input0.6/M tokens
Output2.4/M tokens
Context204.80K
Max Output131.07K
Available on 1 provider
Apr 30, 2026 6:00 PM-
z-ai/glm-5
30.88Btokens
GLM-5 is Zai’s new-generation flagship foundation model, designed for Agentic Engineering, capable of providing reliable productivity in complex system engineering and long-range Agent tasks. In terms of Coding and Agent capabilities, GLM-5 has achieved state-of-the-art (SOTA) performance in open source, with its usability in real programming scenarios approaching that of Claude Opus 4.5.
Input type
Output Type
Input0.58-0.87/M tokens
Output2.6-3.18/M tokens
Context200.00K
Max Output128.00K
Available on 3 providers
Apr 30, 2026 6:00 PM-
openai/gpt-5.3-codex
11.76Btokens
GPT‑5.3‑Codex is OpenAI’s most capable agentic coding model, designed to handle not just code generation but longer, end‑to‑end work on a computer. It combines the frontier coding strength of GPT‑5.2‑Codex with the reasoning and professional knowledge capabilities of GPT‑5.2, and is engineered to run faster (about 25% faster in Codex). It’s built for long‑running, multi‑step tasks that can include research, tool use, debugging, and complex execution, while staying interactive—so you can steer it, ask questions, and refine direction as it works without losing context. In evaluations, it sets new highs on software engineering and agentic benchmarks (including SWE‑Bench Pro and Terminal‑Bench 2.0) and shows strong results on OSWorld‑Verified and GDPval. It’s also deployed with strengthened cybersecurity safeguards and is trained to help identify software vulnerabilities.
Input type
Output Type
Input1.75/M tokens
Output14/M tokens
Context400.00K
Max Output128.00K
Available on 2 providers
Apr 30, 2026 6:00 PM-
anthropic/claude-opus-4.6
736.43Btokens
Opus 4.6 is Anthropic’s strongest model for coding and long-running professional tasks. It is built for agents that operate across entire workflows rather than single prompts, making it especially effective for large codebases, complex refactors, and multi-step debugging that unfolds over time. The model shows deeper contextual understanding, stronger problem decomposition, and greater reliability on hard engineering tasks than prior generations.
Beyond coding, Opus 4.6 excels at sustained knowledge work. It produces near-production-ready documents, plans, and analyses in a single pass, and maintains coherence across very long outputs and extended sessions. This makes it a strong default for tasks that require persistence, judgment, and follow-through, such as technical design, migration planning, and end-to-end project execution.
For users upgrading from earlier Opus versions, see our official migration guide here
Input type
Output Type
Input5/M tokens
Output25/M tokens
Context1000.00K
Max Output128.00K
Available on 3 providers
Apr 30, 2026 6:00 PM-
stepfun/step-3.5-flash
2.27Btokens
Step 3.5 Flash is StepFun's most capable open-source foundation model. Built on a sparse Mixture of Experts (MoE) architecture, it selectively activates only 11B of its 196B parameters per token. It is a reasoning model that is incredibly speed efficient even at long contexts.
Input type
Output Type
Input0.1/M tokens
Output0.3/M tokens
Context256.00K
Max Output256.00K
Available on 1 provider
Apr 30, 2026 6:00 PM-
moonshotai/kimi-k2.5
10.37Btokens
Kimi K2.5 is Kimi's most intelligent model to date, achieving open-source SoTA performance in Agent capabilities, coding, visual understanding, and a range of general intelligence tasks. At the same time, Kimi K2.5 is also Kimi's most versatile model yet. Built with a natively multimodal architecture, it seamlessly supports both visual and text inputs, thinking and non-thinking modes, as well as conversational and Agent-based tasks.
Input type
Output Type
Input0.58/M tokens
Output3.02/M tokens
Context262.14K
Max Output262.14K
Available on 2 providers
Apr 30, 2026 6:00 PM-
minimax/minimax-m2-her
2.13Mtokens
M2-her text chat model, designed for role-playing, multi-turn conversations and dialogue scenarios.
Input type
Output Type
Input0.3/M tokens
Output1.2/M tokens
Context64.00K
Max Output2.05K
Available on 1 provider
Apr 30, 2026 6:00 PM-
qwen/qwen3-max
969.48Mtokens
Qwen3-Max-Thinking is an updated release built on the Qwen3 series, offering major improvements in reasoning, instruction following, multilingual support, and long-tail knowledge coverage compared to the January 2025 version. It delivers higher accuracy in math, coding, logic, and science tasks, follows complex instructions in Chinese and English more reliably, reduces hallucinations, and produces higher-quality responses for open-ended Q&A, writing, and conversation. The model supports over 100 languages with stronger translation and commonsense reasoning, and is optimized for retrieval-augmented generation (RAG) and tool calling, though it does not include a dedicated “thinking” mode.
Input type
Output Type
Input1.2-3/M tokens
Output6-15/M tokens
Context256.00K
Max Output32.00K
Available on 1 provider
Apr 30, 2026 6:00 PM-
baidu/ernie-5.0-thinking-preview
740.68Mtokens
The new-generation Wenxin model, Wenxin 5.0, is a natively multimodal large model. It adopts a native unified multimodal modeling approach to jointly model text, images, audio, and video, providing comprehensive multimodal capabilities. Wenxin 5.0’s core abilities have been comprehensively upgraded and it performs excellently on benchmark datasets, with particularly strong results in multimodal understanding, instruction following, creative writing, factuality, agent planning, and tool use
Input type
Output Type
Input0.84-1.41/M tokens
Output3.37-5.62/M tokens
Context128.00K
Max Output64.00K
Available on 1 provider
Apr 30, 2026 6:00 PM-
z-ai/glm-4.7-flash-free
804.85Mtokens
Free
Rate Limit
As a 30B-class SOTA model, GLM-4.7-Flash offers a new option that balances performance and efficiency. It is further optimized for agentic coding use cases, strengthening coding capabilities, long-horizon task planning, and tool collaboration, and has achieved leading performance among open-source models of the same size on several current public benchmark leaderboards.
Input type
Output Type
Input0/M tokens
Output0/M tokens
Context200.00K
Max Output128.00K
Available on 2 providers
Apr 30, 2026 6:00 PM-
z-ai/glm-4.7-flashx
315.94Mtokens
As a 30B-class SOTA model, GLM-4.7-FlashX offers a new option that balances performance and efficiency. It is further optimized for agentic coding use cases, strengthening coding capabilities, long-horizon task planning, and tool collaboration, and has achieved leading performance among open-source models of the same size on several current public benchmark leaderboards.
Input type
Output Type
Input0.0728/M tokens
Output0.4367/M tokens
Context200.00K
Max Output128.00K
Available on 2 providers
Apr 30, 2026 6:00 PM-
z-ai/glm-image
138.00Ktokens
GLM-Image is an image generation model adopts a hybrid autoregressive + diffusion decoder architecture. In general image generation quality, GLM‑Image aligns with mainstream latent diffusion approaches, but it shows significant advantages in text-rendering and knowledge‑intensive generation scenarios. It performs especially well in tasks requiring precise semantic understanding and complex information expression, while maintaining strong capabilities in high‑fidelity and fine‑grained detail generation. In addition to text‑to‑image generation, GLM‑Image also supports a rich set of image‑to‑image tasks including image editing, style transfer, identity‑preserving generation, and multi‑subject consistency.
Input type
Output Type
Input-
Output0.0146/counts
Context-
Max Output-
Available on 2 providers
Apr 30, 2026 6:00 PM-
openai/gpt-5.2-codex
9.38Btokens
GPT-5.2-Codex is an upgraded version of GPT-5.2 optimized for agentic coding tasks in Codex or similar environments. GPT-5.2-Codex supports low, medium, high, and xhigh reasoning effort settings.
Input type
Output Type
Input1.75/M tokens
Output14/M tokens
Context400.00K
Max Output128.00K
Available on 2 providers
Apr 30, 2026 6:00 PM-
google/veo-3.1-generate-001
352.04Ktokens
Veo 3.1 is Google's state-of-the-art model for generating high-fidelity, 8-second 720p, 1080p or 4k videos featuring stunning realism and natively generated audio. You can access this model programmatically using the Gemini API. To learn more about the available Veo model variants, see the Model Versions section.
Input type
Output Type
Input-
Output0.4-0.6/seconds
Context-
Max Output-
Available on 1 provider
Apr 30, 2026 6:00 PM-
z-ai/glm-4.7
15.01Btokens
Pricing: The model's official pricing is (GLM-4.7: Input $0.28 ~ $0.57, Cached Input $0.057$0.11, Output $1.14$2.27). Our platform is currently running a limited-time promotion, during which you will receive a discount on the official price.
GLM-4.7 is Zhipu’s latest flagship model. Tailored for agentic coding scenarios, GLM-4.7 strengthens coding capabilities, long-horizon task planning, and tool collaboration, and delivers leading performance among open-source models on the latest leaderboards of multiple public benchmarks. Its general capabilities have also improved, with responses that are more concise and natural and writing that feels more immersive. When executing complex agent tasks and invoking tools, it follows instructions more reliably, while the visual quality of artifacts and agentic coding front ends—as well as long-horizon task completion efficiency—are further enhanced.
Input type
Output Type
Input0.2911-0.5823/M tokens
Output1.1645-2.3291/M tokens
Context200.00K
Max Output128.00K
Available on 3 providers
Apr 30, 2026 6:00 PM-
minimax/minimax-m2.1
1.04Btokens
MiniMax-M2.1 is a lightweight, state-of-the-art large language model optimized for coding, agentic workflows, and modern application development. With only 10 billion activated parameters, it delivers a major jump in real-world capability while maintaining exceptional latency, scalability, and cost efficiency.
To avoid degrading this model's performance, MiniMax highly recommends preserving reasoning between turns.
Input type
Output Type
Input0.3/M tokens
Output1.2/M tokens
Context204.80K
Max Output131.07K
Available on 1 provider
Apr 30, 2026 6:00 PM-
bytedance/doubao-seed-1.8
197.57Mtokens
An all-new model purpose-built and optimized for multimodal agent scenarios. Stronger agent capabilities, upgraded multimodal understanding, and more flexible context management
Input type
Output Type
Input0.11-0.34/M tokens
Output0.28-3.41/M tokens
Context256.00K
Max Output32.00K
Available on 1 provider
Apr 30, 2026 6:00 PM-
Google: Gemini 3 Flash Preview
google/gemini-3-flash-preview
79.40Btokens
Gemini 3 Flash Preview is a low-latency model in the Gemini 3 family, optimized for fast, high-throughput inference. It retains the core multimodal and reasoning capabilities of Gemini 3 while prioritizing responsiveness and execution efficiency. Built on the same architecture as Gemini 3 Pro, Gemini 3 Flash Preview supports native multimodal inputs—including text, images, and audio—and incorporates the improved reasoning and long-context handling introduced in the Gemini 3 generation. It is designed for real-time and scalable workloads where latency and cost efficiency are primary considerations.
Input type
Output Type
Input0.5/M tokens
Output3/M tokens
Context1.05M
Max Output65.53K
Available on 2 providers
Apr 30, 2026 6:00 PM-
xiaomi/mimo-v2-flash
8.51Btokens
MiMo-V2-Flash is an open-source foundation language model developed by Xiaomi. It is a Mixture-of-Experts model with 309B total parameters and 15B active parameters, adopting hybrid attention architecture. MiMo-V2-Flash supports a hybrid-thinking toggle and a 256K context window, and excels at reasoning, coding, and agent scenarios. On SWE-bench Verified and SWE-bench Multilingual, MiMo-V2-Flash ranks as the top #1 open-source model globally, delivering performance comparable to Claude Sonnet 4.5 while costing only about 3.5% as much.
Input type
Output Type
Input0.1/M tokens
Output0.3/M tokens
Context262.14K
Max Output262.14K
Available on 2 providers
Apr 30, 2026 6:00 PM-
openai/gpt-image-1.5
8.08Mtokens
OpenAI's GPT Image 1.5 is the latest evolution of its AI image generation, offering superior instruction following, photorealism, text rendering, and editing control, making it ideal for detailed creative and production work with faster speeds and lower costs than predecessors like DALL-E 3. It excels at complex requests, maintaining character/style consistency, rendering crisp text in visuals, and understanding nuanced prompts through built-in reasoning, integrated into ChatGPT and available via API.
Input type
Output Type
Input5/M tokens
Output10/M tokens
Context-
Max Output-
Available on 1 provider
Apr 30, 2026 6:00 PM-
ByteDance/Doubao-Seedance-1.5-pro
bytedance/doubao-seedance-1.5-pro
12.41Mtokens
Seedance 1.5 Pro is a new-generation professional-grade audio-visual co-generation video model released by the Doubao large-model team. Building on its predecessor’s multi-shot storytelling and high-definition generation capabilities, it natively supports integrated audio-and-video output, aiming to deliver an end-to-end synchronized creation experience across visuals, voice, music, and sound effects. The model also includes a built-in first-and-last-frame feature: creators only need to set the opening and ending frames of a video to precisely lock in its style, composition, and characters, which then drives the model to generate smooth, natural motion between frames. By combining audio-visual co-generation with first/last-frame control, Seedance 1.5 Pro significantly improves the efficiency, controllability, and artistic expressiveness of professional video creation.
Input type
Output Type
Input-
Output2.33/M tokens
Context-
Max Output-
Available on 1 provider
Apr 30, 2026 6:00 PM-
openai/gpt-5.2-pro
519.87Mtokens
GPT-5.2 Pro is OpenAI’s most advanced model, offering major improvements in agentic coding and long context performance over GPT-5 Pro. It is optimized for complex tasks that require step-by-step reasoning, instruction following, and accuracy in high-stakes use cases. It supports test-time routing features and advanced prompt understanding, including user-specified intent like "think hard about this." Improvements include reductions in hallucination, sycophancy, and better performance in coding, writing, and health-related tasks.
Input type
Output Type
Input21/M tokens
Output168/M tokens
Context400.00K
Max Output128.00K
Available on 1 provider
Apr 30, 2026 6:00 PM-
openai/gpt-5.2
34.28Btokens
GPT-5.2 is the newest frontier-grade model in the GPT-5 family, outperforming GPT-5.1 in agentic capabilities and long-context handling. It uses adaptive reasoning to dynamically allocate compute—responding quickly to straightforward questions while applying deeper analysis to more complex tasks. Designed for broad task coverage, GPT-5.2 delivers consistent improvements across math, coding, science, and tool-calling workloads, producing more coherent long-form responses and more reliable tool use.
Input type
Output Type
Input1.75/M tokens
Output14/M tokens
Context400.00K
Max Output128.00K
Available on 2 providers
Apr 30, 2026 6:00 PM-
openai/gpt-5.2-chat
407.44Mtokens
GPT-5.2 Chat (AKA Instant) is the fast, lightweight member of the 5.2 family, optimized for low-latency chat while retaining strong general intelligence. It uses adaptive reasoning to selectively “think” on harder queries, improving accuracy on math, coding, and multi-step tasks without slowing down typical conversations. The model is warmer and more conversational by default, with better instruction following and more stable short-form reasoning. GPT-5.2 Chat is designed for high-throughput, interactive workloads where responsiveness and consistency matter more than deep deliberation.
Input type
Output Type
Input1.75/M tokens
Output14/M tokens
Context128.00K
Max Output16.38K
Available on 2 providers
Apr 30, 2026 6:00 PM-
z-ai/glm-4.6v
105.49Mtokens
GLM-4.6V represents a significant evolution of the GLM series into the multimodal domain. It features a 128k-token training context window and sets a new state-of-the-art in visual understanding accuracy for its parameter scale. Pioneeringly, it is the first model to natively integrate tool-calling capabilities into its visual architecture, bridging the gap from visual perception to executable actions. This makes it a unified technical foundation for multimodal Agents in real-world business scenarios. Pricing: The model's official pricing is (GLM-4.6v: Input $0.3, Cached Input $0.05, Output $0.9). Our platform is currently running a limited-time promotion, during which you will receive a discount on the official price.
Input type
Output Type
Input0.1456-0.2911/M tokens
Output0.4367-0.8734/M tokens
Context200.00K
Max Output128.00K
Available on 2 providers
Apr 30, 2026 6:00 PM-
z-ai/glm-4.6v-flash-free
291.69Mtokens
Free
Rate Limit
GLM-4.6V-Flash is the free version of GLM-4.6V, representing a significant iteration in the GLM series for multimodal capabilities. It supports toggling reasoning modes, with a training-time context window expanded to 128k tokens. Achieving state-of-the-art (SOTA) visual understanding accuracy at its parameter scale, it is the first visual model to natively integrate Function Call capability into its architecture. This establishes a seamless pipeline from "visual perception" to "executable actions (Action)", offering a unified technical foundation for multimodal agents in real-world business scenarios.
Input type
Output Type
Input0/M tokens
Output0/M tokens
Context200.00K
Max Output128.00K
Available on 2 providers
Apr 30, 2026 6:00 PM-
z-ai/glm-4.6v-flash
2.08Btokens
GLM-4.6V-FlashX is the the paid version, offering higher capacity and stability, representing a significant iteration in the GLM series for multimodal capabilities. It supports toggling reasoning modes, with a training-time context window expanded to 128k tokens. Achieving state-of-the-art (SOTA) visual understanding accuracy at its parameter scale, it is the first visual model to natively integrate Function Call capability into its architecture. This establishes a seamless pipeline from "visual perception" to "executable actions (Action)", offering a unified technical foundation for multimodal agents in real-world business scenarios.
Input type
Output Type
Input0.0218-0.0437/M tokens
Output0.2184-0.4367/M tokens
Context200.00K
Max Output128.00K
Available on 2 providers
Apr 30, 2026 6:00 PM-
deepseek/deepseek-v3.2
4.07Btokens
DeepSeek-V3.2 is a reasoning-first large language model released by DeepSeek as the official successor to V3.2-Exp. It is designed with a focus on agentic capabilities and integrated reasoning for tool-use scenarios.
If the request to the deepseek-reasoner model includes the tools parameter, the request will actually be processed using the deepseek-chat model.
The model introduces a new large-scale agent training data synthesis method covering over 1,800 environments and 85,000+ complex instructions. DeepSeek-V3.2 is the first model from DeepSeek to integrate thinking directly into tool-use, supporting both thinking and non-thinking modes during tool interactions.
According to DeepSeek's benchmarks, the model delivers performance comparable to GPT-5 level while balancing inference quality and output length. It is available via App, Web, and API, making it suitable for general-purpose daily use as well as complex agentic workflows.
Input type
Output Type
Input0.293/M tokens
Output0.4395/M tokens
Context128.00K
Max Output8.00K
Available on 2 providers
Apr 30, 2026 6:00 PM-
mistralai/mistral-large-2512
74.11Mtokens
Mistral Large 3, is a state-of-the-art, open-weight, general-purpose multimodal model with a granular Mixture-of-Experts architecture. It features 41B active parameters and 675B total parameters.
Input type
Output Type
Input0.5/M tokens
Output1.5/M tokens
Context256.00K
Max Output256.00K
Available on 2 providers
Apr 30, 2026 6:00 PM-
DeepSeek: DeepSeek-V3.2 (Non-thinking Mode)
deepseek/deepseek-chat
13.82Btokens
DeepSeek-V3.2 (Non-thinking Mode) is DeepSeek's latest production model, currently served under the deepseek-chat model slug, which is automatically updated as new versions are released.
If the request to the deepseek-reasoner model includes the tools parameter, the request will actually be processed using the deepseek-chat model.
Input type
Output Type
Input0.14/M tokens
Output0.28/M tokens
Context128.00K
Max Output8.00K
Available on 1 provider
Apr 30, 2026 6:00 PM-
DeepSeek: DeepSeek-V3.2 (Thinking Mode)
deepseek/deepseek-reasoner
4.45Btokens
DeepSeek-V3.2 (Thinking Mode) is DeepSeek's latest production model, currently served under the deepseek-reasoner model slug, which is automatically updated as new versions are released.
If the request to the deepseek-reasoner model includes the tools parameter, the request will actually be processed using the deepseek-chat model.
Input type
Output Type
Input0.14/M tokens
Output0.28/M tokens
Context128.00K
Max Output64.00K
Available on 1 provider
Apr 30, 2026 6:00 PM-
anthropic/claude-opus-4.5
63.05Btokens
Claude Opus 4.5 is Anthropic's latest frontier reasoning model, purpose-built for complex software engineering, agentic workflows, and long-horizon computer use. It delivers strong multimodal capabilities, competitive performance on real-world coding and reasoning benchmarks, and improved robustness against prompt injection attacks. The model is designed to operate efficiently across varied effort levels, allowing developers to balance speed, depth, and token usage based on their specific task requirements—you can fine-tune token efficiency through the OpenRouter Verbosity parameter, which offers low, medium, and high settings. Beyond that, Opus 4.5 supports advanced tool use, extended context management, and coordinated multi-agent setups, making it ideal for autonomous research, debugging, multi-step planning, and spreadsheet or browser manipulation. Compared to previous Opus generations, it brings substantial improvements in structured reasoning, execution reliability, and alignment, while reducing token overhead and delivering more consistent performance on long-running tasks.
Input type
Output Type
Input5/M tokens
Output25/M tokens
Context200.00K
Max Output32.00K
Available on 3 providers
Apr 30, 2026 6:00 PM-
Google: Nano Banana Pro (Gemini 3 Pro Image Preview)
google/gemini-3-pro-image-preview
1.68Btokens
Nano Banana Pro is an AI image generation model that is a significant upgrade from its predecessor, built on Google's Gemini 3 Pro. It promises to move beyond simple pattern matching to a more reasoning-driven system with improved physics understanding, text rendering, and image consistency. Key features include faster processing, native 2K resolution, and the ability to edit existing images with greater control, aiming to produce more reliable and professional-grade results.
Input type
Output Type
Input2-4/M tokens
Output12-18/M tokens
Context65.54K
Max Output32.77K
Available on 2 providers
Apr 30, 2026 6:00 PM-
x-ai/grok-4.1-fast
9.26Btokens
Grok 4.1 Fast is xAI's best agentic tool calling model that shines in real-world use cases like customer support and deep research. 2M context window.
Reasoning can be enabled/disabled using the reasoning``enabled parameter in the API.
Input type
Output Type
Input0.2-0.4/M tokens
Output0.5-1/M tokens
Context2.00M
Max Output30.00K
Available on 1 provider
Apr 30, 2026 6:00 PM-
xAI: Grok 4.1 Fast Non Reasoning
x-ai/grok-4.1-fast-non-reasoning
21.19Btokens
Grok 4.1 Fast Non Reasoning is xAI's best agentic tool calling model that shines in real-world use cases like customer support and deep research. 2M context window.
Input type
Output Type
Input0.2-0.4/M tokens
Output0.5-1/M tokens
Context2.00M
Max Output30.00K
Available on 1 provider
Apr 30, 2026 6:00 PM-
openai/gpt-5.1
1.40Btokens
GPT-5.1 is the latest frontier-grade model in the GPT-5 series, offering stronger general-purpose reasoning, improved instruction adherence, and a more natural conversational style compared to GPT-5. It uses adaptive reasoning to allocate computation dynamically, responding quickly to simple queries while spending more depth on complex tasks. The model produces clearer, more grounded explanations with reduced jargon, making it easier to follow even on technical or multi-step problems.
Built for broad task coverage, GPT-5.1 delivers consistent gains across math, coding, and structured analysis workloads, with more coherent long-form answers and improved tool-use reliability. It also features refined conversational alignment, enabling warmer, more intuitive responses without compromising precision. GPT-5.1 serves as the primary full-capability successor to GPT-5
Input type
Output Type
Input1.25/M tokens
Output10/M tokens
Context400.00K
Max Output128.00K
Available on 2 providers
Apr 30, 2026 6:00 PM-
openai/gpt-5.1-chat
156.49Mtokens
GPT-5.1 Chat (AKA Instant is the fast, lightweight member of the 5.1 family, optimized for low-latency chat while retaining strong general intelligence. It uses adaptive reasoning to selectively “think” on harder queries, improving accuracy on math, coding, and multi-step tasks without slowing down typical conversations. The model is warmer and more conversational by default, with better instruction following and more stable short-form reasoning. GPT-5.1 Chat is designed for high-throughput, interactive workloads where responsiveness and consistency matter more than deep deliberation.
Input type
Output Type
Input1.25/M tokens
Output10/M tokens
Context128.00K
Max Output16.38K
Available on 2 providers
Apr 30, 2026 6:00 PM-
openai/gpt-5.1-codex
922.49Mtokens
GPT-5.1-Codex is a specialized version of GPT-5.1 optimized for software engineering and coding workflows. It is designed for both interactive development sessions and long, independent execution of complex engineering tasks. The model supports building projects from scratch, feature development, debugging, large-scale refactoring, and code review. Compared to GPT-5.1, Codex is more steerable, adheres closely to developer instructions, and produces cleaner, higher-quality code outputs.
Codex integrates into developer environments including the CLI, IDE extensions, GitHub, and cloud tasks. It adapts reasoning effort dynamically—providing fast responses for small tasks while sustaining extended multi-hour runs for large projects. The model is trained to perform structured code reviews, catching critical flaws by reasoning over dependencies and validating behavior against tests. It also supports multimodal inputs such as images or screenshots for UI development and integrates tool use for search, dependency installation, and environment setup. Codex is intended specifically for agentic coding applications.
Input type
Output Type
Input1.25/M tokens
Output10/M tokens
Context400.00K
Max Output128.00K
Available on 2 providers
Apr 30, 2026 6:00 PM-
openai/gpt-5.1-codex-mini
780.89Mtokens
GPT-5.1-Codex-Mini is a smaller and faster version of GPT-5.1-Codex
Input type
Output Type
Input0.25/M tokens
Output2/M tokens
Context400.00K
Max Output100.00K
Available on 2 providers
Apr 30, 2026 6:00 PM-
bytedance/doubao-seed-code
1.03Mtokens
Doubao-Seed-Code has been deeply optimized for Agentic Programming tasks, delivering exceptional performance across multiple authoritative benchmarks — including Terminal Bench, SWE-Bench-Verified-Openhands, and Multi-SWE-Bench-Flash-Openhands — outperforming domestic counterparts and supporting a context window of up to 256k tokens.
Input type
Output Type
Input0.17-0.39/M tokens
Output1.12-2.25/M tokens
Context256.00K
Max Output32.00K
Available on 1 provider
Apr 30, 2026 6:00 PM-
tencent/hunyuan-2.0-thinking
4.22Mtokens
Tencent HY 2.0 Think, a large language model fully developed end-to-end by Tencent, leads the industry with outstanding performance in high-quality content creation, mathematical logic reasoning, code generation and multi-turn conversations; its API supports an internet-connected AI search plugin that integrates Tencent's premium content ecosystem to provide powerful real-time, in-depth content retrieval and AI question-answering capabilities, and this release upgrades the model base from TurboS to HY 2.0 for overall capability enhancement, with significant improvements in complex instruction following, multi-turn and long-text comprehension, code generation, Agent support and reasoning abilities.
Input type
Output Type
Input0.57-0.76/M tokens
Output2.29-3.05/M tokens
Context128.00K
Max Output64.00K
Available on 1 provider
Apr 30, 2026 6:00 PM-
moonshotai/kimi-k2-thinking
501.44Mtokens
A thinking model with general agentic and reasoning capabilities, specializing in deep reasoning tasks
Input type
Output Type
Input0.6/M tokens
Output2.5/M tokens
Context262.14K
Max Output262.14K
Available on 1 provider
Apr 30, 2026 6:00 PM-
MoonshotAI: Kimi K2 Thinking Turbo
moonshotai/kimi-k2-thinking-turbo
138.11Mtokens
Context length 256k. High-speed version of kimi-k2-thinking, suitable for scenarios requiring both deep reasoning and extremely fast responses
Input type
Output Type
Input1.15/M tokens
Output8/M tokens
Context262.14K
Max Output262.14K
Available on 1 provider
Apr 30, 2026 6:00 PM-
OpenAI: Text Embedding 3 Small
openai/text-embedding-3-small
41.14Mtokens
text-embedding-3-small is OpenAI's improved, more performant version of the ada embedding model. Embeddings are a numerical representation of text that can be used to measure the relatedness between two pieces of text. Embeddings are useful for search, clustering, recommendations, anomaly detection, and classification tasks.
Input type
Output Type
Input0.02/M tokens
Output0/M tokens
Context8.19K
Max Output8.19K
Available on 1 provider
Apr 30, 2026 6:00 PM-
minimax/minimax-m2
775.27Mtokens
MiniMax-M2 is a compact, high-efficiency large language model optimized for end-to-end coding and agentic workflows. With 10 billion activated parameters (230 billion total), it delivers near-frontier intelligence across general reasoning, tool use, and multi-step task execution while maintaining low latency and deployment efficiency. The model excels in code generation, multi-file editing, compile-run-fix loops, and test-validated repair, showing strong results on SWE-Bench Verified, Multi-SWE-Bench, and Terminal-Bench. It also performs competitively in agentic evaluations such as BrowseComp and GAIA, effectively handling long-horizon planning, retrieval, and recovery from execution errors. Benchmarked by Artificial Analysis, MiniMax-M2 ranks among the top open-source models for composite intelligence, spanning mathematics, science, and instruction-following. Its small activation footprint enables fast inference, high concurrency, and improved unit economics, making it well-suited for large-scale agents, developer assistants, and reasoning-driven applications that require responsiveness and cost efficiency.
Input type
Output Type
Input0.3/M tokens
Output1.2/M tokens
Context204.80K
Max Output128.00K
Available on 1 provider
Apr 30, 2026 6:00 PM-
anthropic/claude-haiku-4.5
99.78Btokens
Claude Haiku 4.5 is Anthropic’s fastest and most efficient model, delivering near-frontier intelligence at a fraction of the cost and latency of larger Claude models. Matching Claude Sonnet 4’s performance across reasoning, coding, and computer-use tasks, Haiku 4.5 brings frontier-level capability to real-time and high-volume applications.
It introduces extended thinking to the Haiku line; enabling controllable reasoning depth, summarized or interleaved thought output, and tool-assisted workflows with full support for coding, bash, web search, and computer-use tools. Scoring >73% on SWE-bench Verified, Haiku 4.5 ranks among the world’s best coding models while maintaining exceptional responsiveness for sub-agents, parallelized execution, and scaled deployment.
Input type
Output Type
Input1/M tokens
Output5/M tokens
Context200.00K
Max Output64.00K
Available on 2 providers
Apr 30, 2026 6:00 PM-
inclusionai/ring-1t
1.38Btokens
Ring-1T is a trillion-parameter sparse mixture-of-experts (MoE) thinking model developed by inclusionAI. It adopts the Ling 2.0 architecture and is trained on the Ling-1T-base foundation model, which contains 1 trillion total parameters with 50 billion activated parameters, supporting a context window of up to 128K tokens. Building upon the preview version released at the end of September, Ring-1T has undergone continued scaling with large-scale verifiable reward reinforcement learning (RLVR) training, further unlocking the natural language reasoning capabilities of the trillion-parameter foundation model.
Input type
Output Type
Input0.56/M tokens
Output2.24/M tokens
Context128.00K
Max Output32.00K
Available on 1 provider
Apr 30, 2026 6:00 PM-
inclusionai/ling-1t
7.75Btokens
Ling-1T is a trillion-parameter sparse mixture-of-experts (MoE) model developed by inclusionAI, optimized for efficient and scalable reasoning. Featuring approximately 50 billion active parameters per token, it is pre-trained on over 20 trillion reasoning-dense tokens, supports a 128K context length, and utilizes an Evolutionary Chain-of-Thought (Evo-CoT) process to enhance its reasoning depth. The model achieves state-of-the-art performance across complex benchmarks, demonstrating strong capabilities in code generation, software development, and advanced mathematics. In addition to its core reasoning skills, Ling-1T possesses specialized abilities in front-end code generation—combining semantic understanding with visual aesthetics—and exhibits emergent agentic capabilities, such as proficient tool use with minimal instruction tuning. Its primary use cases span software engineering, professional mathematics, complex logical reasoning, and agent-based workflows that demand a balance of high performance and efficiency.
Input type
Output Type
Input0.56/M tokens
Output2.24/M tokens
Context128.00K
Max Output32.00K
Available on 1 provider
Apr 30, 2026 6:00 PM-
Google: Gemini 2.5 Flash Image (Nano Banana)
google/gemini-2.5-flash-image
233.13Mtokens
Gemini 2.5 Flash Image, a.k.a. "Nano Banana," is now generally available. It is a state of the art image generation model with contextual understanding. It is capable of image generation, edits, and multi-turn conversations.
Input type
Output Type
Input0.3/M tokens
Output2.5/M tokens
Context32.77K
Max Output8.19K
Available on 1 provider
Apr 30, 2026 6:00 PM-
openai/gpt-5-pro
32.03Mtokens
GPT-5 Pro is OpenAI’s most advanced model, offering major improvements in reasoning, code quality, and user experience. It is optimized for complex tasks that require step-by-step reasoning, instruction following, and accuracy in high-stakes use cases. It supports test-time routing features and advanced prompt understanding, including user-specified intent like "think hard about this." Improvements include reductions in hallucination, sycophancy, and better performance in coding, writing, and health-related tasks.
Input type
Output Type
Input15/M tokens
Output120/M tokens
Context400.00K
Max Output128.00K
Available on 2 providers
Apr 30, 2026 6:00 PM-
z-ai/glm-4.6
1.31Btokens
GLM-4.6 is a flagship model from Zhishen with 355B total parameters and 32B active parameters. The model's context window has been expanded from 128K to 200K, enabling it to handle longer code and agent tasks. In programming capabilities, GLM-4.6's performance is comparable to Claude Sonnet 4 on public benchmarks and real-world programming tasks. The model supports tool calling during inference and features optimized search and tool-use performance within agent frameworks. Furthermore, enhancements have been made to its writing style, readability, role-playing, and cross-lingual task processing abilities.
Pricing: The model's official pricing is (GLM-4.6: Input $0.6, Cached Input $0.11, Output $2.2). Our platform is currently running a limited-time promotion, during which you will receive a discount on the official price.
Input type
Output Type
Input0.2911-0.5823/M tokens
Output1.1645-2.3291/M tokens
Context200.00K
Max Output128.00K
Available on 3 providers
Apr 30, 2026 6:00 PM-
anthropic/claude-sonnet-4.5
117.01Btokens
Claude Sonnet 4.5 is Anthropic’s most advanced Sonnet model to date, optimized for real-world agents and coding workflows. It delivers state-of-the-art performance on coding benchmarks such as SWE-bench Verified, with improvements across system design, code security, and specification adherence. The model is designed for extended autonomous operation, maintaining task continuity across sessions and providing fact-based progress tracking.
Sonnet 4.5 also introduces stronger agentic capabilities, including improved tool orchestration, speculative parallel execution, and more efficient context and memory management. With enhanced context tracking and awareness of token usage across tool calls, it is particularly well-suited for multi-context and long-running workflows. Use cases span software engineering, cybersecurity, financial analysis, research agents, and other domains requiring sustained reasoning and tool use.
Input type
Output Type
Input3-6/M tokens
Output15-22.5/M tokens
Context1000.00K
Max Output64.00K
Available on 3 providers
Apr 30, 2026 6:00 PM-
deepseek/deepseek-v3.2-exp
214.87Mtokens
DeepSeek-V3.2-Exp is an experimental model version, serving as an intermediate step toward the next-generation architecture. Built on the foundation of V3.1-Terminus, it introduces the DeepSeek Sparse Attention (DSA) mechanism—a sparse attention mechanism designed to explore and validate the optimization of training and inference efficiency in long-context scenarios. This experimental version represents the team's continuous research on more efficient Transformer architectures, with a specific focus on improving computational efficiency when processing long text sequences. For the first time, DSA enables fine-grained sparse attention, which significantly enhances the efficiency of long-context training and inference while maintaining almost unchanged model output quality.
Input type
Output Type
Input0.216/M tokens
Output0.328/M tokens
Context163.84K
Max Output65.54K
Available on 1 provider
Apr 30, 2026 6:00 PM-
tencent/hunyuan-image3
126.00Ktokens
HunyuanImage-3.0 is a groundbreaking native multimodal model that unifies multimodal understanding and generation within an autoregressive framework. Our text-to-image and image-to-image model achieves performance comparable to or surpassing leading closed-source models.
Input type
Output Type
Input-
Output0.029/counts
Context-
Max Output-
Available on 1 provider
Apr 30, 2026 6:00 PM-
openai/gpt-5-codex
299.24Mtokens
GPT-5-Codex is a specialized version of GPT-5 optimized for software engineering and coding workflows. It is designed for both interactive development sessions and long, independent execution of complex engineering tasks. The model supports building projects from scratch, feature development, debugging, large-scale refactoring, and code review. Compared to GPT-5, Codex is more steerable, adheres closely to developer instructions, and produces cleaner, higher-quality code outputs. Reasoning effort can be adjusted with the reasoning.effort parameter. Read the docs here
Input type
Output Type
Input1.25/M tokens
Output10/M tokens
Context400.00K
Max Output128.00K
Available on 2 providers
Apr 30, 2026 6:00 PM-
qwen/qwen3-vl-plus
82.38Mtokens
The Qwen3 series VL models effectively integrates thinking and non-thinking modes, achieving world-leading performance in visual agent capabilities on public benchmark datasets such as OS World. This version features comprehensive upgrades in areas like visual coding, spatial perception, and multimodal reasoning, significantly enhancing visual perception and recognition abilities, and supporting the understanding of ultra-long videos.
Input type
Output Type
Input0.2-0.6/M tokens
Output1.6-4.8/M tokens
Context262.14K
Max Output32.77K
Available on 1 provider
Apr 30, 2026 6:00 PM-
x-ai/grok-4-fast
10.34Btokens
Grok 4 Fast is xAI's latest multimodal model with SOTA cost-efficiency and a 2M token context window. It comes in two flavors: non-reasoning and reasoning. Read more about the model on xAI's news post. Reasoning can be enabled using the reasoning``enabled parameter in the API. Learn more in our docs
Prompts and completions may be used by xAI or OpenRouter to improve future models.
Input type
Output Type
Input0.2-0.4/M tokens
Output0.5-1/M tokens
Context2.00M
Max Output30.00K
Available on 1 provider
Apr 30, 2026 6:00 PM-
xAI: Grok 4 Fast None Reasoning
x-ai/grok-4-fast-non-reasoning
3.55Btokens
Prompts and completions may be used by xAI or OpenRouter to improve future models.
Input type
Output Type
Input0.2-0.4/M tokens
Output0.5-1/M tokens
Context2.00M
Max Output30.00K
Available on 1 provider
Apr 30, 2026 6:00 PM-
inclusionai/ling-flash-2.0
228.50Mtokens
Ling-flash-2.0 is an open-source Mixture-of-Experts (MoE) language model developed under the Ling 2.0 architecture. It features 100 billion total parameters, with 6.1 billion activated during inference (4.8B non-embedding).
Trained on over 20 trillion tokens and refined with supervised fine-tuning and multi-stage reinforcement learning, the model demonstrates strong performance against dense models up to 40B parameters. It excels in complex reasoning, code generation, and frontend development.
Input type
Output Type
Input0.28/M tokens
Output2.8/M tokens
Context128.00K
Max Output32.00K
Available on 1 provider
Apr 30, 2026 6:00 PM-
inclusionai/ring-flash-2.0
218.15Mtokens
Input type
Output Type
Input0.28/M tokens
Output2.8/M tokens
Context128.00K
Max Output32.00K
Available on 1 provider
Apr 30, 2026 6:00 PM-
baidu/ernie-x1.1-preview
3.67Mtokens
The Wenxin Large Model X1.1 delivers significantly enhanced performance in question answering, tool invocation, agent capabilities, instruction following, logical reasoning, mathematical problem-solving, and coding tasks, with markedly improved factual accuracy. Its context window has been extended to 64K tokens, enabling longer inputs and dialogue histories, while maintaining response speed and improving the coherence of long-chain reasoning.
Input type
Output Type
Input0.14/M tokens
Output0.56/M tokens
Context65.54K
Max Output65.54K
Available on 1 provider
Apr 30, 2026 6:00 PM-
moonshotai/kimi-k2-0905
121.16Mtokens
Kimi K2 0905 is the September update of Kimi K2 0711 [blocked]. It is a large-scale Mixture-of-Experts (MoE) language model developed by Moonshot AI, featuring 1 trillion total parameters with 32 billion active per forward pass. It supports long-context inference up to 256k tokens, extended from the previous 128k.
This update improves agentic coding with higher accuracy and better generalization across scaffolds, and enhances frontend coding with more aesthetic and functional outputs for web, 3D, and related tasks. Kimi K2 is optimized for agentic capabilities, including advanced tool use, reasoning, and code synthesis. It excels across coding (LiveCodeBench, SWE-bench), reasoning (ZebraLogic, GPQA), and tool-use (Tau2, AceBench) benchmarks. The model is trained with a novel stack incorporating the MuonClip optimizer for stable large-scale MoE training.
Input type
Output Type
Input0.58/M tokens
Output2.33/M tokens
Context262.10K
Max Output262.10K
Available on 1 provider
Apr 30, 2026 6:00 PM-
inclusionai/ling-mini-2.0
215.92Mtokens
Ling-mini-2.0 is an open-source Mixture-of-Experts (MoE) large language model designed to balance strong task performance with high inference efficiency. It has 16B total parameters, with approximately 1.4B activated per token (about 789M non-embedding). Trained on over 20T tokens and refined via multi-stage supervised fine-tuning and reinforcement learning, it is reported to deliver strong results in complex reasoning and instruction following while keeping computational costs low. According to the upstream release, it reaches top-tier performance among sub-10B dense LLMs and in some cases matches or surpasses larger MoE models.
Input type
Output Type
Input0.07/M tokens
Output0.28/M tokens
Context128.00K
Max Output32.00K
Available on 1 provider
Apr 30, 2026 6:00 PM-
inclusionai/ring-mini-2.0
1.23Btokens
Ring-mini-2.0 is a Mixture-of-Experts (MoE) model oriented toward high-throughput inference and extensively optimized on the Ling 2.0 architecture. It uses 16B total parameters with approximately 1.4B activated per token and is reported to deliver comprehensive reasoning performance comparable to sub-10B dense LLMs. The model shows strong results on logical reasoning, code generation, and mathematical tasks, supports 128K context windows, and reports generation speeds of 300+ tokens per second.
Input type
Output Type
Input0.07/M tokens
Output0.7/M tokens
Context128.00K
Max Output32.00K
Available on 1 provider
Apr 30, 2026 6:00 PM-
x-ai/grok-code-fast-1
1.18Btokens
Grok Code Fast 1 is a speedy and economical reasoning model that excels at agentic coding. With reasoning traces visible in the response, developers can steer Grok Code for high-quality work flows.
Input type
Output Type
Input0.2/M tokens
Output1.5/M tokens
Context256.00K
Max Output10.00K
Available on 1 provider
Apr 30, 2026 6:00 PM-
deepseek/deepseek-chat-v3.1
1.49Btokens
DeepSeek-V3.1 is a large hybrid reasoning model (671B parameters, 37B active) that supports both thinking and non-thinking modes via prompt templates. It extends the DeepSeek-V3 base with a two-phase long-context training process, reaching up to 128K tokens, and uses FP8 microscaling for efficient inference.
The model improves tool use, code generation, and reasoning efficiency, achieving performance comparable to DeepSeek-R1 on difficult benchmarks while responding more quickly. It supports structured tool calling, code agents, and search agents, making it suitable for research, coding, and agentic workflows.
It succeeds the DeepSeek V3-0324 model and performs well on a variety of tasks.
Input type
Output Type
Input0.56/M tokens
Output1.68/M tokens
Context128.00K
Max Output65.54K
Available on 2 providers
Apr 30, 2026 6:00 PM-
openai/gpt-5
2.48Btokens
GPT-5 is OpenAI’s most advanced model, offering major improvements in reasoning, code quality, and user experience. It is optimized for complex tasks that require step-by-step reasoning, instruction following, and accuracy in high-stakes use cases. It supports test-time routing features and advanced prompt understanding, including user-specified intent like "think hard about this." Improvements include reductions in hallucination, sycophancy, and better performance in coding, writing, and health-related tasks.
Input type
Output Type
Input1.25/M tokens
Output10/M tokens
Context400.00K
Max Output128.00K
Available on 2 providers
Apr 30, 2026 6:00 PM-
openai/gpt-5-chat
100.41Mtokens
GPT-5 Chat is designed for advanced, natural, multimodal, and context-aware conversations for enterprise applications.
Input type
Output Type
Input1.25/M tokens
Output10/M tokens
Context128.00K
Max Output16.38K
Available on 2 providers
Apr 30, 2026 6:00 PM-
openai/gpt-5-mini
2.35Btokens
GPT-5 Mini is a compact version of GPT-5, designed to handle lighter-weight reasoning tasks. It provides the same instruction-following and safety-tuning benefits as GPT-5, but with reduced latency and cost. GPT-5 Mini is the successor to OpenAI's o4-mini model.
Input type
Output Type
Input0.25/M tokens
Output2/M tokens
Context400.00K
Max Output128.00K
Available on 2 providers
Apr 30, 2026 6:00 PM-
openai/gpt-5-nano
816.62Mtokens
GPT-5-Nano is the smallest and fastest variant in the GPT-5 system, optimized for developer tools, rapid interactions, and ultra-low latency environments. While limited in reasoning depth compared to its larger counterparts, it retains key instruction-following and safety features. It is the successor to GPT-4.1-nano and offers a lightweight option for cost-sensitive or real-time applications.
Input type
Output Type
Input0.05/M tokens
Output0.4/M tokens
Context400.00K
Max Output128.00K
Available on 2 providers
Apr 30, 2026 6:00 PM-
anthropic/claude-opus-4.1
283.17Mtokens
Claude Opus 4.1 is an updated version of Anthropic’s flagship model, offering improved performance in coding, reasoning, and agentic tasks. It achieves 74.5% on SWE-bench Verified and shows notable gains in multi-file code refactoring, debugging precision, and detail-oriented reasoning. The model supports extended thinking up to 64K tokens and is optimized for tasks involving research, data analysis, and tool-assisted reasoning.
Input type
Output Type
Input15/M tokens
Output75/M tokens
Context200.00K
Max Output32.00K
Available on 3 providers
Apr 30, 2026 6:00 PM-
stepfun/step-3
37.12Mtokens
Step-3 is a brand-new multimodal reasoning model that can process both image and text inputs and produce text responses. It is capable of deep thinking and can autonomously carry out a reasoning process. That is, before generating the final output, it completes a “thinking” phase (for example, presenting reasoning information via a reasoning field), which improves the accuracy of the final result and the depth of reasoning. When calling the model, developers do not need to preset too many system prompts (sys_prompt), as the model can automatically leverage its built-in deep-thinking capability._
Input type
Output Type
Input0.21-0.57/M tokens
Output0.57-1.42/M tokens
Context65.54K
Max Output65.54K
Available on 1 provider
Apr 30, 2026 6:00 PM-
klingai/kling-v2
740.00Ktokens
kling-v2 is an image generation model launched by Kling AI.
Input type
Output Type
Input-
Output0.014/counts
Context-
Max Output-
Available on 1 provider
Apr 30, 2026 6:00 PM-
z-ai/glm-4.5
124.68Mtokens
GLM-4.5 is our latest flagship foundation model, purpose-built for agent-based applications. It leverages a Mixture-of-Experts (MoE) architecture and supports a context length of up to 128k tokens. GLM-4.5 delivers significantly enhanced capabilities in reasoning, code generation, and agent alignment. It supports a hybrid inference mode with two options, a "thinking mode" designed for complex reasoning and tool use, and a "non-thinking mode" optimized for instant responses. Users can control the reasoning behaviour with the reasoning``enabled boolean.
Input type
Output Type
Input0.2911-0.5823/M tokens
Output1.1645-2.3291/M tokens
Context128.00K
Max Output96.00K
Available on 2 providers
Apr 30, 2026 6:00 PM-
z-ai/glm-4.5-air
271.10Mtokens
GLM-4.5-Air is the lightweight variant of our latest flagship model family, also purpose-built for agent-centric applications. Like GLM-4.5, it adopts the Mixture-of-Experts (MoE) architecture but with a more compact parameter size. GLM-4.5-Air also supports hybrid inference modes, offering a "thinking mode" for advanced reasoning and tool use, and a "non-thinking mode" for real-time interaction. Users can control the reasoning behaviour with the reasoning``enabled boolean. Learn more in our docs
Input type
Output Type
Input0.1165-0.1747/M tokens
Output0.2911-1.1645/M tokens
Context128.00K
Max Output96.00K
Available on 3 providers
Apr 30, 2026 6:00 PM-
qwen/qwen3-coder-plus
535.18Mtokens
Powered by Qwen3, this is a powerful Coding Agent that excels in tool calling and environment interaction to achieve autonomous programming. It combines outstanding coding proficiency with versatile general-purpose abilities.
Input type
Output Type
Input1-6/M tokens
Output5-60/M tokens
Context1000.00K
Max Output65.54K
Available on 1 provider
Apr 30, 2026 6:00 PM-
google/gemini-2.5-flash-lite
24.72Btokens
Gemini 2.5 Flash-Lite is a lightweight reasoning model in the Gemini 2.5 family, optimized for ultra-low latency and cost efficiency. It offers improved throughput, faster token generation, and better performance across common benchmarks compared to earlier Flash models. By default, "thinking" (i.e. multi-pass reasoning) is disabled to prioritize speed, but developers can enable it via the Reasoning API parameter to selectively trade off cost for intelligence.
Input type
Output Type
Input0.1/M tokens
Output0.4/M tokens
Context1.05M
Max Output65.53K
Available on 2 providers
Apr 30, 2026 6:00 PM-
Qwen: Qwen3 235B A22B Instruct 2507
qwen/qwen3-235b-a22b-2507
354.87Mtokens
Qwen3-235B-A22B-Instruct-2507 is a multilingual, instruction-tuned mixture-of-experts language model based on the Qwen3-235B architecture, with 22B active parameters per forward pass. It is optimized for general-purpose text generation, including instruction following, logical reasoning, math, code, and tool usage. The model supports a native 262K context length and does not implement "thinking mode" (
Compared to its base variant, this version delivers significant gains in knowledge coverage, long-context reasoning, coding benchmarks, and alignment with open-ended tasks. It is particularly strong on multilingual understanding, math reasoning (e.g., AIME, HMMT), and alignment evaluations like Arena-Hard and WritingBench.
Input type
Output Type
Input0.28/M tokens
Output1.11/M tokens
Context256.00K
Max Output128.00K
Available on 1 provider
Apr 30, 2026 6:00 PM-
Qwen: Qwen3 235B A22B Thinking 2507
qwen/qwen3-235b-a22b-thinking-2507
21.68Mtokens
Qwen3-235B-A22B-Thinking-2507 is a high-performance, open-weight Mixture-of-Experts (MoE) language model optimized for complex reasoning tasks. It activates 22B of its 235B parameters per forward pass and natively supports up to 262,144 tokens of context. This "thinking-only" variant enhances structured logical reasoning, mathematics, science, and long-form generation, showing strong benchmark performance across AIME, SuperGPQA, LiveCodeBench, and MMLU-Redux. It enforces a special reasoning mode () and is designed for high-token outputs (up to 81,920 tokens) in challenging domains.
The model is instruction-tuned and excels at step-by-step reasoning, tool use, agentic workflows, and multilingual tasks. This release represents the most capable open-source variant in the Qwen3-235B series, surpassing many closed models in structured reasoning use cases.
Input type
Output Type
Input0.28/M tokens
Output2.78/M tokens
Context256.00K
Max Output128.00K
Available on 1 provider
Apr 30, 2026 6:00 PM-
qwen/qwen3-coder
36.56Mtokens
Qwen3-Coder-480B-A35B-Instruct is a Mixture-of-Experts (MoE) code generation model developed by the Qwen team. It is optimized for agentic coding tasks such as function calling, tool use, and long-context reasoning over repositories. The model features 480 billion total parameters, with 35 billion active per forward pass (8 out of 160 experts).
Pricing for the Alibaba endpoints varies by context length. Once a request is greater than 128k input tokens, the higher pricing is used.
Input type
Output Type
Input1.25/M tokens
Output5.01/M tokens
Context256.00K
Max Output128.00K
Available on 1 provider
Apr 30, 2026 6:00 PM-
moonshotai/kimi-k2-0711
15.53Mtokens
Kimi K2 Instruct is a large-scale Mixture-of-Experts (MoE) language model developed by Moonshot AI, featuring 1 trillion total parameters with 32 billion active per forward pass. It is optimized for agentic capabilities, including advanced tool use, reasoning, and code synthesis. Kimi K2 excels across a broad range of benchmarks, particularly in coding (LiveCodeBench, SWE-bench), reasoning (ZebraLogic, GPQA), and tool-use (Tau2, AceBench) tasks. It supports long-context inference up to 128K tokens and is designed with a novel training stack that includes the MuonClip optimizer for stable large-scale MoE training.
Input type
Output Type
Input0.56/M tokens
Output2.23/M tokens
Context128.00K
Max Output32.00K
Available on 2 providers
Apr 30, 2026 6:00 PM-
x-ai/grok-4
294.85Mtokens
Grok 4 is xAI's latest reasoning model with a 256k context window. It supports parallel tool calling, structured outputs, and both image and text inputs. Note that reasoning is not exposed, reasoning cannot be disabled, and the reasoning effort cannot be specified. Pricing increases once the total tokens in a given request is greater than 128k tokens. See more details on the xAI docs
Input type
Output Type
Input3-6/M tokens
Output15-30/M tokens
Context256.00K
Max Output256.00K
Available on 1 provider
Apr 30, 2026 6:00 PM-
google/gemini-2.5-flash
4.26Btokens
Gemini 2.5 Flash is Google's state-of-the-art workhorse model, specifically designed for advanced reasoning, coding, mathematics, and scientific tasks. It includes built-in "thinking" capabilities, enabling it to provide responses with greater accuracy and nuanced context handling.
Additionally, Gemini 2.5 Flash is configurable through the "max tokens for reasoning" parameter, as described in the documentation.
Input type
Output Type
Input0.3/M tokens
Output2.5/M tokens
Context1.05M
Max Output65.53K
Available on 2 providers
Apr 30, 2026 6:00 PM-
google/gemini-2.5-pro
3.25Btokens
Gemini 2.5 Pro is Google’s state-of-the-art AI model designed for advanced reasoning, coding, mathematics, and scientific tasks. It employs “thinking” capabilities, enabling it to reason through responses with enhanced accuracy and nuanced context handling. Gemini 2.5 Pro achieves top-tier performance on multiple benchmarks, including first-place positioning on the LMArena leaderboard, reflecting superior human-preference alignment and complex problem-solving abilities.
Input type
Output Type
Input1.25-2.5/M tokens
Output10-15/M tokens
Context1.05M
Max Output65.53K
Available on 2 providers
Apr 30, 2026 6:00 PM-
deepseek/deepseek-r1-0528
908.74Mtokens
May 28th update to the original DeepSeek R1 Performance on par with OpenAI o1, but open-sourced and with fully open reasoning tokens. It's 671B parameters in size, with 37B active in an inference pass. Fully open-source model.
Input type
Output Type
Input0.56/M tokens
Output2.23/M tokens
Context64.00K
Max Output64.00K
Available on 2 providers
Apr 30, 2026 6:00 PM-
anthropic/claude-opus-4
219.16Mtokens
Claude Opus 4 is benchmarked as the world’s best coding model, at time of release, bringing sustained performance on complex, long-running tasks and agent workflows. It sets new benchmarks in software engineering, achieving leading results on SWE-bench (72.5%) and Terminal-bench (43.2%). Opus 4 supports extended, agentic workflows, handling thousands of task steps continuously for hours without degradation.
Input type
Output Type
Input15/M tokens
Output75/M tokens
Context200.00K
Max Output32.00K
Available on 3 providers
Apr 30, 2026 6:00 PM-
anthropic/claude-sonnet-4
8.49Btokens
Claude Sonnet 4 significantly enhances the capabilities of its predecessor, Sonnet 3.7, excelling in both coding and reasoning tasks with improved precision and controllability. Achieving state-of-the-art performance on SWE-bench (72.7%), Sonnet 4 balances capability and computational efficiency, making it suitable for a broad range of applications from routine coding tasks to complex software development projects. Key enhancements include improved autonomous codebase navigation, reduced error rates in agent-driven workflows, and increased reliability in following intricate instructions. Sonnet 4 is optimized for practical everyday use, providing advanced reasoning capabilities while maintaining efficiency and responsiveness in diverse internal and external scenarios.
Input type
Output Type
Input3-6/M tokens
Output15-22.5/M tokens
Context1000.00K
Max Output64.00K
Available on 3 providers
Apr 30, 2026 6:00 PM-
google/gemma-3-12b-it
491.80Mtokens
Gemma 3 introduces multimodality, supporting vision-language input and text outputs. It handles context windows up to 128k tokens, understands over 140 languages, and offers improved math, reasoning, and chat capabilities, including structured outputs and function calling. Gemma 3 12B is the second largest in the family of Gemma 3 models after Gemma 3 27B [blocked]
Input type
Output Type
Input0.024/M tokens
Output0.096/M tokens
Context128.00K
Max Output128.00K
Available on 1 provider
Apr 30, 2026 6:00 PM-
qwen/qwen3-14b
94.63Mtokens
Qwen3-14B is a dense 14.8B parameter causal language model from the Qwen3 series, designed for both complex reasoning and efficient dialogue. It supports seamless switching between a "thinking" mode for tasks like math, programming, and logical inference, and a "non-thinking" mode for general-purpose conversation. The model is fine-tuned for instruction-following, agent tool use, creative writing, and multilingual tasks across 100+ languages and dialects. It natively handles 32K token contexts and can extend to 131K tokens using YaRN-based scaling.
Input type
Output Type
Input0.14/M tokens
Output1.4/M tokens
Context32.00K
Max Output32.00K
Available on 1 provider
Apr 30, 2026 6:00 PM-
openai/o4-mini
131.78Mtokens
OpenAI o4-mini is a compact reasoning model in the o-series, optimized for fast, cost-efficient performance while retaining strong multimodal and agentic capabilities. It supports tool use and demonstrates competitive reasoning and coding performance across benchmarks like AIME (99.5% with Python) and SWE-bench, outperforming its predecessor o3-mini and even approaching o3 in some domains.
Despite its smaller size, o4-mini exhibits high accuracy in STEM tasks, visual problem solving (e.g., MathVista, MMMU), and code editing. It is especially well-suited for high-throughput scenarios where latency or cost is critical. Thanks to its efficient architecture and refined reinforcement learning training, o4-mini can chain tools, generate structured outputs, and solve multi-step tasks with minimal delay—often in under a minute.
Input type
Output Type
Input1.1/M tokens
Output4.4/M tokens
Context200.00K
Max Output100.00K
Available on 2 providers
Apr 30, 2026 6:00 PM-
openai/gpt-4.1
3.33Btokens
GPT-4.1 is a flagship large language model optimized for advanced instruction following, real-world software engineering, and long-context reasoning. It supports a 1 million token context window and outperforms GPT-4o and GPT-4.5 across coding (54.6% SWE-bench Verified), instruction compliance (87.4% IFEval), and multimodal understanding benchmarks. It is tuned for precise code diffs, agent reliability, and high recall in large document contexts, making it ideal for agents, IDE tooling, and enterprise knowledge retrieval.
Input type
Output Type
Input2/M tokens
Output8/M tokens
Context1.05M
Max Output32.77K
Available on 2 providers
Apr 30, 2026 6:00 PM-
openai/gpt-4.1-mini
3.57Btokens
GPT-4.1 Mini is a mid-sized model delivering performance competitive with GPT-4o at substantially lower latency and cost. It retains a 1 million token context window and scores 45.1% on hard instruction evals, 35.8% on MultiChallenge, and 84.1% on IFEval. Mini also shows strong coding ability (e.g., 31.6% on Aider’s polyglot diff benchmark) and vision understanding, making it suitable for interactive applications with tight performance constraints.
Input type
Output Type
Input0.4/M tokens
Output1.6/M tokens
Context1.05M
Max Output32.77K
Available on 2 providers
Apr 30, 2026 6:00 PM-
openai/gpt-4.1-nano
1.89Btokens
For tasks that demand low latency, GPT‑4.1 nano is the fastest and cheapest model in the GPT-4.1 series. It delivers exceptional performance at a small size with its 1 million token context window, and scores 80.1% on MMLU, 50.3% on GPQA, and 9.8% on Aider polyglot coding – even higher than GPT‑4o mini. It’s ideal for tasks like classification or autocompletion.
Input type
Output Type
Input0.1/M tokens
Output0.4/M tokens
Context1.05M
Max Output32.77K
Available on 2 providers
Apr 30, 2026 6:00 PM-
meta/llama-4-scout-17b-16e-instruct
34.60Mtokens
Llama 4 Scout 17B Instruct (16E) is a mixture-of-experts (MoE) language model developed by Meta, activating 17 billion parameters out of a total of 109B. It supports native multimodal input (text and image) and multilingual output (text and code) across 12 supported languages. Designed for assistant-style interaction and visual reasoning, Scout uses 16 experts per forward pass and features a context length of 10 million tokens, with a training corpus of ~40 trillion tokens. Built for high efficiency and local or commercial deployment, Llama 4 Scout incorporates early fusion for seamless modality integration. It is instruction-tuned for use in multilingual chat, captioning, and image understanding tasks. Released under the Llama 4 Community License, it was last trained on data up to August 2024 and launched publicly on April 5, 2025.
Input type
Output Type
Input0.08/M tokens
Output0.4/M tokens
Context-
Max Output-
Available on 1 provider
Apr 30, 2026 6:00 PM-
anthropic/claude-3.7-sonnet
3.70Btokens
Sunset: 2026/04/28
Claude 3.7 Sonnet is an advanced large language model with improved reasoning, coding, and problem-solving capabilities. It introduces a hybrid reasoning approach, allowing users to choose between rapid responses and extended, step-by-step processing for complex tasks. The model demonstrates notable improvements in coding, particularly in front-end development and full-stack updates, and excels in agentic workflows, where it can autonomously navigate multi-step processes.
Claude 3.7 Sonnet maintains performance parity with its predecessor in standard mode while offering an extended reasoning mode for enhanced accuracy in math, coding, and instruction-following tasks.
Input type
Output Type
Input3/M tokens
Output15/M tokens
Context200.00K
Max Output64.00K
Available on 1 provider
Apr 30, 2026 6:00 PM-
meta/llama-3.3-70b-instruct
245.77Mtokens
The Meta Llama 3.3 multilingual large language model (LLM) is a pretrained and instruction tuned generative model in 70B (text in/text out). The Llama 3.3 instruction tuned text only model is optimized for multilingual dialogue use cases and outperforms many of the available open source and closed chat models on common industry benchmarks.
Supported languages: English, German, French, Italian, Portuguese, Hindi, Spanish, and Thai.
Input type
Output Type
Input0.6/M tokens
Output1.2/M tokens
Context-
Max Output-
Available on 1 provider
Apr 30, 2026 6:00 PM-
anthropic/claude-3.5-haiku
278.03Mtokens
Claude 3.5 Haiku features offers enhanced capabilities in speed, coding accuracy, and tool use. Engineered to excel in real-time applications, it delivers quick response times that are essential for dynamic tasks such as chat interactions and immediate coding suggestions.
This makes it highly suitable for environments that demand both speed and precision, such as software development, customer service bots, and data management systems.
Input type
Output Type
Input0.8/M tokens
Output4/M tokens
Context200.00K
Max Output8.19K
Available on 2 providers
Apr 30, 2026 6:00 PM-
openai/gpt-4o-mini
1.57Btokens
GPT-4o mini is OpenAI's newest model after GPT-4 Omni, supporting both text and image inputs with text outputs.
As their most advanced small model, it is many multiples more affordable than other recent frontier models, and more than 60% cheaper than GPT-3.5 Turbo. It maintains SOTA intelligence, while being significantly more cost-effective.
GPT-4o mini achieves an 82% score on MMLU and presently ranks higher than GPT-4 on chat preferences common leaderboards.
Input type
Output Type
Input0.15/M tokens
Output0.6/M tokens
Context128.00K
Max Output16.38K
Available on 2 providers
Apr 30, 2026 6:00 PM-
openai/gpt-4o
577.81Mtokens
GPT-4o ("o" for "omni") is OpenAI's latest AI model, supporting both text and image inputs with text outputs. It maintains the intelligence level of GPT-4 Turbowhile being twice as fast and 50% more cost-effective. GPT-4o also offers improved performance in processing non-English languages and enhanced visual capabilities.
For benchmarking against other models, it was briefly called "im-also-a-good-gpt2-chatbot"
Input type
Output Type
Input2.5/M tokens
Output10/M tokens
Context128.00K
Max Output16.38K
Available on 2 providers
Apr 30, 2026 6:00 PM-
StripeM-Inner