google/gemini-3.1-pro-preview, anthropic/claude-sonnet-4.6, google/gemini-3-pro-image-preview are Now Live

Claude Sonnet 4 will be retired by both providers: AWS Bedrock and Anthropic on June 14, 2026. We recommend upgrading to Claude Sonnet 4.6.

Claude Opus 4 will be retired by both providers: AWS Bedrock on May 31, 2026 and Anthropic on June 14, 2026. We recommend upgrading to Claude opus 4.7.

Qwen3 Max Thinking Preview will be deprecated and removed on April 24, 2026. We recommend upgrading to Qwen3-Max-Thinking

Caution: On Azure, gpt-5-chat will be deprecated and removed on May 15, 2026.

Claude 3.7 Sonnet will be retired by both providers: AWS Bedrock on April 28, 2026, and Google Cloud Vertex AI on May 11, 2026. We recommend upgrading to Claude Sonnet 4.6.

Get 5% off service fee on top-up, and exclusive gifts for referring friends

A billing anomaly was identified. Affected users will receive compensation. Read more

Caution: On Azure, gpt-5-chat will be deprecated and removed on May 15, 2026.

A billing anomaly was identified. Affected users will receive compensation. Read more

  • 1
  • 2
  • 3
  • 4
  • 5
  • 6
  • 7
  • 8

Studio

Models

PricingDevelopers Analytics About Us

Sign In

Input Modalities

Text

Image

File

Audio

Video

Output Modalities

Text

Image

File

Audio

Video

Embeddings

Context Length

4K64K1M

Maker

Alibaba

Anthropic

Baidu

Show More

Providers

OpenAI

Anthropic

Tbox

Show More

Supported Parameters

max_completion_tokens

temperature

top_p

Show More

Supported Protocol

OpenAI Chat Completions

OpenAI Responses

OpenAI Embeddings

Anthropic Messages

Google Gemini

Google Imagen

Google Video

Reasoning

No Reasoning

Toggleable Reasoning

Always-On Reasoning

Models

157 models

Newest

ZenMux: Auto Router

zenmux/auto

API Request Chat

ZenMux's automatic routing feature selects the most cost-effective and high-performing AI models based on your query.

Input type

Output Type

Input-

Output-

Context-

Max Output-

Qwen: Qwen3.6 Max Preview

qwen/qwen3.6-max-preview

API Request Chat

5.25Mtokens

Qwen3.6-Max-Preview is a proprietary frontier model from Alibaba Cloud built on a sparse mixture-of-experts architecture with approximately 1 trillion total parameters. It is optimized for agentic coding, tool use, and long-context reasoning, supporting a 262K token context window. The model includes an integrated thinking mode that preserves reasoning traces across multi-turn conversations and supports structured output and function calling. Access is available exclusively through the Alibaba Cloud Model Studio and Qwen Studio APIs; no open weights are provided.

Input type

Output Type

Input1.3-2/M tokens

Output7.8-12/M tokens

Context262.14K

Max Output65.54K

Available on 1 provider

Apr 30, 2026 6:00 PM-

Alibaba: HappyHorse 1.0

alibaba/happyhorse-1.0

API Request Create

500.00Ktokens

Happy Horse 1.0 is described as an open-source state-of-the-art AI video generator with native joint audio-video generation — meaning the Happy Horse AI video model produces video frames and the corresponding audio track (dialogue, ambient sound, Foley) together in a single forward pass, rather than generating silent video and dubbing it afterward. According to community-compiled architecture notes, the model is built around a 15-billion-parameter unified self-attention Transformer that processes text, image, video, and audio tokens within a single token sequence. It is reportedly built without dedicated cross-attention branches and without a separate audio module. Combined with DMD-2 distillation, the distilled variant is reported to generate 1080p video in roughly 38 seconds on an NVIDIA H100, using only 8 denoising steps without classifier-free guidance.

Input type

Output Type

Input-

Output-

Context-

Max Output-

Available on 1 provider

Apr 30, 2026 6:00 PM-

OpenAI: GPT-5.5

openai/gpt-5.5

API Request Chat

3.42Btokens

GPT‑5.5 understands what you’re trying to do faster and can carry more of the work itself. It excels at writing and debugging code, researching online, analyzing data, creating documents and spreadsheets, operating software, and moving across tools until a task is finished. Instead of carefully managing every step, you can give GPT‑5.5 a messy, multi-part task and trust it to plan, use tools, check its work, navigate through ambiguity, and keep going.

Input type

Output Type

Input5-10/M tokens

Output30-45/M tokens

Context1.05M

Max Output128.00K

Available on 2 providers

Apr 30, 2026 6:00 PM-

OpenAI: GPT-5.5 Pro

openai/gpt-5.5-pro

API Request Chat

79.26Mtokens

GPT-5.5 Pro is OpenAI’s high-capability model optimized for deep reasoning and accuracy on complex, high-stakes workloads. It features a 1M+ token context window (922K input, 128K output) with support for text and image inputs, and is designed for long-horizon problem solving, agentic coding, and precise execution across multi-step workflows.

Input type

Output Type

Input30-60/M tokens

Output180-270/M tokens

Context1.05M

Max Output128.00K

Available on 1 provider

Apr 30, 2026 6:00 PM-

DeepSeek: DeepSeek V4 Flash

deepseek/deepseek-v4-flash

API Request Chat

1.93Btokens

DeepSeek-V4-Flash is the efficiency-oriented variant of the DeepSeek V4 series, released as a preview and open-sourced alongside the flagship V4-Pro. It is designed for developers who need the V4 generation's long-context and reasoning capability at a faster, more economical API tier. Compared with V4-Pro, V4-Flash uses smaller total parameters and active parameters, resulting in faster response times and lower API cost. It retains reasoning capability close to V4-Pro and matches V4-Pro on simple agent tasks, with a measurable gap appearing only on the most demanding agent workflows. World knowledge is slightly below V4-Pro but remains competitive within the open-source landscape.

Like V4-Pro, V4-Flash inherits the new attention mechanism built on token-dimension compression and DeepSeek Sparse Attention (DSA), supports a 1M context window as standard, and offers both thinking and non-thinking modes with a reasoning_effort parameter (high / max)._

Note on migration: the legacy model names deepseek-chat and deepseek-reasoner currently route to V4-Flash in non-thinking and thinking mode respectively, and will be retired on 2026-07-24. Existing integrations should migrate to the explicit model names deepseek-v4-flash or deepseek-v4-pro. The model is accessible via OpenAI ChatCompletions and Anthropic interfaces, with weights open-sourced on Hugging Face and ModelScope.

Input type

Output Type

Input0.14/M tokens

Output0.28/M tokens

Context1000.00K

Max Output384.00K

Available on 2 providers

Apr 30, 2026 6:00 PM-

DeepSeek: DeepSeek V4 Flash (Free)

deepseek/deepseek-v4-flash-free

API Request Chat

3.21Btokens

Free

Rate Limit

DeepSeek latest model: DeepSeek V4 Flash.

Input type

Output Type

Input0/M tokens

Output0/M tokens

Context1000.00K

Max Output384.00K

Available on 1 provider

Apr 30, 2026 6:00 PM-

DeepSeek: DeepSeek V4 Pro

deepseek/deepseek-v4-pro

API Request Chat

17.52Btokens

75% OFF

The deepseek-v4-pro model is currently offered through the official DeepSeek direct channel at a limited-time 75% discount, valid until 2026/05/31 15:59 UTC. DeepSeek-V4-Pro is the flagship model of the DeepSeek V4 series, released as a preview and open-sourced alongside its efficiency-tier sibling V4-Flash. It is positioned as the performance-first tier in the V4 lineup, designed to push agentic capability, world knowledge, and reasoning performance to levels competitive with leading closed-source models.

The model introduces a new attention mechanism that performs compression along the token dimension, combined with DeepSeek Sparse Attention (DSA). The design makes a 1M context window the default across all official DeepSeek services while substantially reducing compute and memory overhead compared with prior approaches.

DeepSeek-V4-Pro has been specifically adapted and optimized for mainstream agent products such as Claude Code, OpenClaw, OpenCode, and CodeBuddy. According to DeepSeek, V4-Pro reaches the top tier among open-source models on Agentic Coding benchmarks, and is currently used internally at DeepSeek as the default Agentic Coding model—with internal evaluation reports describing a usage experience above Sonnet 4.5 and delivery quality close to Opus 4.6 in non-thinking mode, while still trailing Opus 4.6 in thinking mode. On world knowledge, it leads open-source models and trails only Gemini-Pro-3.1; on math, STEM, and competitive coding, it delivers results on par with top closed-source models.

The model supports both thinking and non-thinking modes, with the thinking mode exposing a reasoning_effort parameter (high / max) for complex agent workflows. It is available via App, Web, and API through the OpenAI ChatCompletions and Anthropic interfaces, with weights open-sourced on Hugging Face and ModelScope._

Input type

Output Type

Input0.435/M tokens

Output0.87/M tokens

Context1000.00K

Max Output384.00K

Available on 2 providers

Apr 30, 2026 6:00 PM-

DeepSeek: DeepSeek V4 Pro (Free)

deepseek/deepseek-v4-pro-free

API Request Chat

17.49Btokens

Free

Rate Limit

DeepSeek-V4-Pro is the flagship model of the DeepSeek V4 series, released as a preview and open-sourced alongside its efficiency-tier sibling V4-Flash. It is positioned as the performance-first tier in the V4 lineup, designed to push agentic capability, world knowledge, and reasoning performance to levels competitive with leading closed-source models.

Input type

Output Type

Input0/M tokens

Output0/M tokens

Context1000.00K

Max Output384.00K

Available on 1 provider

Apr 30, 2026 6:00 PM-

Sapiens AI: Agnes-1.5-Flash

sapiens-ai/agnes-1.5-flash

API Request Chat

116.32Mtokens

Agnes-1.5-Flash is a model derived from the Agnes-1.5-Pro architecture, optimized for high efficiency without compromising performance. Through advanced quantization techniques, it delivers performance comparable to significantly larger models while maintaining much lower compute requirements and latency. It retains strong capabilities in conversation and content generation, making it ideal for scalable, cost-sensitive, and real-time applications.

Input type

Output Type

Input0.07/M tokens

Output0.15/M tokens

Context256.00K

Max Output65.54K

Available on 1 provider

Apr 30, 2026 6:00 PM-

inclusionAI: Ling-2.6-1T

inclusionai/ling-2.6-1t

API Request Chat

3.40Btokens

Free

Ling-2.6-1T is an instant (instruct) model from inclusionAI and the company’s trillion-parameter flagship, designed for real-world agents that require fast execution and high efficiency at scale. It uses a “fast thinking” approach to reduce costs to roughly a quarter of comparable models while maintaining top-tier performance.

The model achieves state-of-the-art results on benchmarks such as AIME26 and SWE-bench Verified, and is well suited for advanced coding, complex reasoning, and large-scale agent workflows where both capability and efficiency are critical.

Input type

Output Type

Input0/M tokens

Output0/M tokens

Context262.14K

Max Output32.77K

Available on 1 provider

Apr 30, 2026 6:00 PM-

Tencent: Hy3 preview

tencent/hy3-preview

API Request Chat

1.39Btokens

Hy3 Preview is a high-efficiency Mixture-of-Experts model from Tencent designed for agentic workflows and production use. It supports configurable reasoning levels across disabled, low, and high modes, allowing it to balance speed and depth depending on the task, while delivering strong code generation and reliable performance across multi-step, real-world workflows.

Input type

Output Type

Input0.172-0.286/M tokens

Output0.572-1.144/M tokens

Context262.14K

Max Output131.07K

Available on 1 provider

Apr 30, 2026 6:00 PM-

Xiaomi: MiMo-V2.5

xiaomi/mimo-v2.5

API Request Chat

80.61Mtokens

MiMo-V2.5 is a native omnimodal model by Xiaomi. It delivers Pro-level agentic performance at roughly half the inference cost, while surpassing MiMo-V2-Omni in multimodal perception across image and video understanding tasks. Its 1M context window supports complete documents, extended conversations, and complex task contexts in a single pass, making it ideal for integration with agent frameworks where strong reasoning, rich perception, and cost efficiency all matter.

Input type

Output Type

Input0.4-0.8/M tokens

Output2-4/M tokens

Context1.05M

Max Output131.07K

Available on 1 provider

Apr 30, 2026 6:00 PM-

Xiaomi: MiMo-V2.5-Pro

xiaomi/mimo-v2.5-pro

API Request Chat

354.27Mtokens

MiMo-V2.5-Pro is Xiaomi’s flagship model, delivering strong performance in general agentic capabilities, complex software engineering, and long-horizon tasks, with top rankings on benchmarks such as ClawEval, GDPVal, and SWE-bench Pro. It can independently and autonomously complete professional tasks that would take human experts days or weeks, involving more than a thousand tool calls. Its context length of up to 1M makes it well suited for integration with a wide range of agent frameworks.

Input type

Output Type

Input1-2/M tokens

Output3-6/M tokens

Context1.05M

Max Output131.07K

Available on 1 provider

Apr 30, 2026 6:00 PM-

Baidu: ERNIE-Image-Turbo

baidu/ernie-image-turbo

API Request Chat

60.00Ktokens

ERNIE-Image-Turbo is an open text-to-image generation model developed by the ERNIE-Image team at Baidu. It is the distilled release of ERNIE-Image, built on the same single-stream Diffusion Transformer (DiT) family and designed for fast generation with strong fidelity in only 8 inference steps. The model retains strong controllability in practical generation scenarios where accurate content realization matters as much as aesthetics. In particular, ERNIE-Image-Turbo remains strong on complex instruction following, text rendering, and structured image generation, making it well suited for posters, comics, multi-panel layouts, and other content creation tasks that require both visual quality and efficiency. It also supports a broad range of visual styles, including realistic photography, design-oriented imagery, and stylized aesthetic outputs.

Input type

Output Type

Input-

Output-

Context-

Max Output-

Available on 1 provider

Apr 30, 2026 6:00 PM-

inclusionAI: Ling-2.6-flash

inclusionai/ling-2.6-flash

API Request Chat

3.90Mtokens

Ling-2.6-flash is a 100B-parameter text model focused on intelligence efficiency, delivering strong performance while minimizing token usage. It supports a 256K context window with up to 32K output tokens, function calling, structured output, and prompt caching. It is particularly well-suited for code completion and debugging, rapid document processing, and lightweight agent interactions.

Input type

Output Type

Input0.1/M tokens

Output0.3/M tokens

Context262.14K

Max Output32.77K

Available on 2 providers

Apr 30, 2026 6:00 PM-

OpenAI: GPT-Image-2

openai/gpt-image-2

API Request Chat

102.40Mtokens

GPT Image 2 — OpenAI's next-generation image generation model. It delivers rapid, high-quality image generation and editing capabilities, with support for flexible image dimensions and high-fidelity image inputs for seamless creative workflows.

Input type

Output Type

Input5/M tokens

Output-

Context-

Max Output-

Available on 2 providers

Apr 30, 2026 6:00 PM-

MoonshotAI: Kimi K2.6

moonshotai/kimi-k2.6

API Request Chat

2.09Btokens

Kimi’s most intelligent model to date, achieving open-source SoTA performance in Agent, code, visual understanding, and a range of general intelligent tasks. It is also Kimi’s most versatile model to date, featuring a native multimodal architecture that supports both visual and text input, thinking and non-thinking modes, and dialogue and Agent tasks. Context 256k

Input type

Output Type

Input0.95/M tokens

Output4/M tokens

Context262.14K

Max Output262.14K

Available on 2 providers

Apr 30, 2026 6:00 PM-

SkyReels V4

skyreels/skyreels-v4

API Request Create

240.00Ktokens

SkyReels V4 is a unified multimodal video foundation model that generates, edits, and inpaints video and audio simultaneously using a dual-stream architecture with shared text encoding and efficient high-resolution processing.

Input type

Output Type

Input-

Output0.14/seconds

Context1.28K

Max Output1.28K

Available on 1 provider

Apr 30, 2026 6:00 PM-

Anthropic: Claude Opus 4.7

anthropic/claude-opus-4.7

API Request Chat

134.90Btokens

Opus 4.7 is the next generation of Anthropic's Opus family, built for long-running, asynchronous agents. Building on the coding and agentic strengths of Opus 4.6, it delivers stronger performance on complex, multi-step tasks and more reliable agentic execution across extended workflows. It is especially effective for asynchronous agent pipelines where tasks unfold over time - large codebases, multi-stage debugging, and end-to-end project orchestration.

Beyond coding, Opus 4.7 brings improved knowledge work capabilities - from drafting documents and building presentations to analyzing data. It maintains coherence across very long outputs and extended sessions, making it a strong default for tasks that require persistence, judgment, and follow-through.

Input type

Output Type

Input5/M tokens

Output25/M tokens

Context1000.00K

Max Output128.00K

Available on 3 providers

Apr 30, 2026 6:00 PM-

ByteDance: Doubao-Seedance-2.0

bytedance/doubao-seedance-2.0

API Request Create

82.38Mtokens

Seedance 2.0 delivers ultra-realistic, highly stable audiovisual output: with outstanding motion stability and fine visual detail, it produces imagery with a live-action-quality look and an almost indistinguishable blend of real and virtual visual impact. It can handle complex scenes with ease—vividly recreating everything from subtle micro-expressions and intense physical confrontations to dynamic, high-energy song-and-dance performances. It also comes with professional camera movement, multi-shot storytelling, and text-to-video generation capabilities, enhancing narrative tension. Audio and visuals are generated natively in sync, accurately matching the visuals with rich sound effects. It supports performances in multiple languages, accents, and dialects, giving videos exceptional completeness and immersion. Seedance 2.0 is deeply optimized for three key scenarios: commercial advertising, film/TV production, and social media marketing. With industrial-grade generation quality, the hit rate for successful generations is significantly improved, lowering the barrier and cost of producing high-quality content, streamlining the workflow from idea to final cut, and delivering substantial efficiency gains for the industry.

Input type

Output Type

Input-

Output4.1-6.74/M tokens

Context12.80K

Max Output12.80K

Available on 1 provider

Apr 30, 2026 6:00 PM-

Google: Veo 3.1 Fast

google/veo-3.1-fast-generate-001

API Request Create

32.00Ktokens

Veo 3.1 Fast is a speed-optimized variant of Google DeepMind's flagship video generation model. It is designed to generate high-quality video significantly faster and at a lower cost than the standard Veo 3.1 Quality model, making it ideal for rapid prototyping and high-volume content creation.

Input type

Output Type

Input-

Output0.15-0.35/seconds

Context-

Max Output-

Available on 1 provider

Apr 30, 2026 6:00 PM-

Google: Veo 3.1 Lite

google/veo-3.1-lite-generate-001

API Request Create

-tokens

Veo 3.1 Lite is Google DeepMind's most cost-efficient AI video generation model, released on March 31, 2026. It is designed to provide professional-grade video capabilities at a significantly lower price point, making it ideal for developers and content teams who need to scale high-volume video production

Input type

Output Type

Input-

Output0.05-0.08/seconds

Context-

Max Output-

Available on 1 provider

Apr 30, 2026 6:00 PM-

Z.AI: GLM 5.1

z-ai/glm-5.1

API Request Chat

44.98Btokens

GLM-5.1 is Zai’s new-generation flagship foundation model, designed for Agentic Engineering, capable of providing reliable productivity in complex system engineering and long-range Agent tasks. In terms of Coding and Agent capabilities, GLM-5 has achieved state-of-the-art (SOTA) performance in open source, with its usability in real programming scenarios approaching that of Claude Opus 4.5.

Input type

Output Type

Input0.8781-1.1709/M tokens

Output3.5126-4.098/M tokens

Context200.00K

Max Output128.00K

Available on 3 providers

Apr 30, 2026 6:00 PM-

Qwen: Qwen3.6-Plus

qwen/qwen3.6-plus

API Request Chat

6.84Btokens

Qwen3.6-Plus is Alibaba’s next-generation Qwen large language model released on April 2, 2026. Compared with version 3.5, Qwen 3.6 has made significant overall performance improvements and has exhibited remarkably strong agent-oriented programming capabilities.

Input type

Output Type

Input0.5-2/M tokens

Output3-6/M tokens

Context1000.00K

Max Output64.00K

Available on 1 provider

Apr 30, 2026 6:00 PM-

Z.AI: GLM 5V Turbo

z-ai/glm-5v-turbo

API Request Chat

418.86Mtokens

GLM-5V-Turbo is Zhipu’s first multimodal coding base model, designed for vision-based programming tasks. It can natively process multimodal inputs such as images, video, and text, and excels at long-horizon planning, complex programming, and action execution. It is deeply adapted to Agent workflows and can closely collaborate with Agents like Claude Code and OpenClaw to complete a full closed loop of “understanding the environment → planning actions → executing tasks.

Input type

Output Type

Input0.726-1.0165/M tokens

Output3.1946-3.7754/M tokens

Context200.00K

Max Output128.00K

Available on 1 provider

Apr 30, 2026 6:00 PM-

KwaiKAT: KAT-Coder-Pro-V2

kuaishou/kat-coder-pro-v2

API Request Chat

106.74Mtokens

KAT-Coder-Pro V2 is theKwaiKAT’s Latest Flagship Agentic Coding Model. Engineered for multi-scaffold generalization, the model is natively compatible with over 10 mainstream AI coding tools—such as Claude Code, Cline, Kilo, and OpenCode, offering unparalleled flexibility. Through dedicated full-pipeline optimization for OpenClaw, KAT-Coder-Pro V2 is trained from the ground up to master complex, real-world application workflows with ease. Beyond logic, KAT-Coder-Pro V2 achieves a breakthrough in frontend aesthetic generation. In Landing Page and PPT scenarios, the model delivers a paradigm shift in user experience: No Structured Spec Needed: Users no longer need to provide rigid design specifications. Simply describe what you want in plain language to receive production-grade, high-quality output that rivals structured design inputs. Mass Market Expansion: This evolution expands the model's service boundary from the 1% of power users to hundreds of millions of ordinary users, truly democratizing professional-grade creation.

Input type

Output Type

Input0.3/M tokens

Output1.2/M tokens

Context256.00K

Max Output80.00K

Available on 1 provider

Apr 30, 2026 6:00 PM-

Sapiens AI: Agnes-1.5-Lite

sapiens-ai/agnes-1.5-lite

API Request Chat

2.18Btokens

Agnes-1.5-Lite is a model derived from the Agnes-1.5-Pro architecture, optimized for high efficiency without compromising performance. Through advanced quantization techniques, it delivers performance comparable to significantly larger models while maintaining much lower compute requirements and latency. It retains strong capabilities in conversation and content generation, making it ideal for scalable, cost-sensitive, and real-time applications.

Input type

Output Type

Input0.07/M tokens

Output0.15/M tokens

Context256.00K

Max Output65.54K

Available on 1 provider

Apr 30, 2026 6:00 PM-

Sapiens AI: Agnes-Video-V1.2

sapiens-ai/agnes-video-v1.2

API Request Create

-tokens

Agnes-Video-V1.2 is a cinematic-grade video generation model that delivers high-fidelity visuals and fully synchronized audio in a single pass. It natively generates aligned dialogue and rich environmental sound alongside film-quality imagery, ensuring coherence between speech, motion, and scene dynamics. Designed for storytelling, marketing, and immersive experiences, it transforms simple prompts into production-ready videos with cinematic realism—eliminating the need for separate audio pipelines or post-processing.

Input type

Output Type

Input-

Output0.018/seconds

Context-

Max Output-

Available on 1 provider

Apr 30, 2026 6:00 PM-

Sapiens AI: Agnes-1.5-Pro

sapiens-ai/agnes-1.5-pro

API Request Chat

4.48Btokens

Agnes-1.5-Pro is a large-scale text foundation model built with tens of billions of parameters, delivering strong capabilities in natural language understanding and generation. It demonstrates excellent performance across complex semantic modeling, multi-turn dialogue, and reasoning tasks.

Through continuous optimization, Agnes leverages parameter-efficient fine-tuning techniques and task-driven training strategies to enhance its adaptability to real-world business scenarios. In addition, it incorporates native tool calling capabilities, enabling the model not only to understand and generate language, but also to autonomously select and invoke external tools based on task requirements—bridging the gap between comprehension and execution.

Input type

Output Type

Input0.16/M tokens

Output0.8/M tokens

Context256.00K

Max Output256.00K

Available on 1 provider

Apr 30, 2026 6:00 PM-

Xiaomi: MiMo-V2-Omni

xiaomi/mimo-v2-omni

API Request Chat

1.18Btokens

MiMo-V2-Omni is purpose-built for complex multimodal interaction and execution scenarios in the real world. We constructed a fully omnimodal foundation from the ground up — one that natively integrates text, vision, and speech — and deeply couples perception with action through a unified architecture. This not only breaks free from the limitations of conventional models that prioritize understanding over execution, but also equips the model with native capabilities spanning multimodal perception, tool invocation, function execution, and GUI manipulation. With seamless integration into mainstream agent frameworks, MiMo-V2-Omni bridges the gap between comprehension and control, dramatically lowering the barrier to deploying full-stack multimodal agents in production.

Input type

Output Type

Input0.4/M tokens

Output2/M tokens

Context265.00K

Max Output265.00K

Available on 1 provider

Apr 30, 2026 6:00 PM-

Xiaomi: MiMo-V2-Pro

xiaomi/mimo-v2-pro

API Request Chat

7.82Btokens

Xiaomi MiMo-V2-Pro is purpose-built for high-intensity agent workflows in real-world scenarios. It features over 1 trillion total parameters (with 42B active parameters), employs an innovative hybrid attention architecture, and supports an ultra-long context window of 1M tokens. Building upon this powerful model foundation, we've continuously scaled compute across a broader spectrum of agent scenarios, further expanding the intelligent action space and achieving significant generalization — from Coding to Claw.

Input type

Output Type

Input1-2/M tokens

Output3-6/M tokens

Context1000.00K

Max Output256.00K

Available on 2 providers

Apr 30, 2026 6:00 PM-

MiniMax: MiniMax M2.7

minimax/minimax-m2.7

API Request Chat

9.69Btokens

M2.7 delivers outstanding performance in real-world software engineering, including end-to-end complete project delivery, log analysis and bug triaging, code security, machine learning, and more. On the benchmark SWE-Pro, M2.7 scores 56.22%, nearly matching the level of Opus. This capability also extends to end-to-end complete project delivery scenarios (VIBE-Pro 55.6%) and deep understanding of complex engineering systems on Terminal Bench 2 (57.0%).

In the professional office domain, we have improved the model's specialized knowledge and task delivery capabilities across various fields. On GDPval-AA, its ELO score is 1495, the highest among open-source models. M2.7's ability to perform complex editing in the Office suite (Excel/PPT/Word) has significantly improved, enabling better multi-round revisions and high-fidelity editing. M2.7 is capable of interacting with complex environments. Across 40 complex skills (> 2000 tokens) cases, M2.7 still maintains a 97% skill adherence rate. In OpenClaw usage, M2.7 has shown significant improvement compared to M2.5, scoring close to the latest Sonnet 4.6 in the MMClaw evaluation.

M2.7 possesses excellent identity retention capabilities and emotional intelligence. Beyond productivity use cases, it also opens up space for innovation in interactive entertainment scenarios.

Input type

Output Type

Input0.3/M tokens

Output1.2/M tokens

Context204.80K

Max Output131.07K

Available on 2 providers

Apr 30, 2026 6:00 PM-

MiniMax: MiniMax M2.7 highspeed

minimax/minimax-m2.7-highspeed

API Request Chat

795.48Mtokens

M2.7 highspeed: Same performance, faster, more agile

Input type

Output Type

Input0.611/M tokens

Output2.4439/M tokens

Context204.80K

Max Output131.07K

Available on 1 provider

Apr 30, 2026 6:00 PM-

OpenAI: GPT-5.4 Mini

openai/gpt-5.4-mini

API Request Chat

5.98Btokens

GPT-5.4 mini brings the strengths of GPT-5.4 to a faster, more efficient model designed for high-volume workloads.

Input type

Output Type

Input0.75/M tokens

Output4.5/M tokens

Context400.00K

Max Output128.00K

Available on 2 providers

Apr 30, 2026 6:00 PM-

OpenAI: GPT-5.4 Nano

openai/gpt-5.4-nano

API Request Chat

468.02Mtokens

GPT-5.4 nano is designed for tasks where speed and cost matter most like classification, data extraction, ranking, and sub-agents.

Input type

Output Type

Input0.2/M tokens

Output1.25/M tokens

Context400.00K

Max Output128.00K

Available on 2 providers

Apr 30, 2026 6:00 PM-

inclusionAI: LLaDA2.1-flash

inclusionai/llada2.1-flash

API Request Chat

119.66Mtokens

LLaDA 2.1 is a diffusion language model in the LLaDA series, enhanced with editing capabilities. It delivers strong task performance while significantly improving inference speed.

Input type

Output Type

Input0.28/M tokens

Output2.85/M tokens

Context32.00K

Max Output32.00K

Available on 1 provider

Apr 30, 2026 6:00 PM-

Z.ai: GLM 5 Turbo

z-ai/glm-5-turbo

API Request Chat

4.18Btokens

GLM-5 Turbo is a new model from Z.ai designed for fast inference and strong performance in agent-driven environments such as OpenClaw scenarios. It is deeply optimized for real-world agent workflows involving long execution chains, with improved complex instruction decomposition, tool use, scheduled and persistent execution, and overall stability across extended tasks.

Input type

Output Type

Input0.73-1.02/M tokens

Output3.19-3.77/M tokens

Context200.00K

Max Output128.00K

Available on 1 provider

Apr 30, 2026 6:00 PM-

xAI: Grok 4.2 Fast

x-ai/grok-4.2-fast

API Request Chat

1.34Btokens

Grok 4.20 Beta features a powerful 4-agent, multi-agent architecture (Grok, Harper, Benjamin, Lucas) that operates in parallel to research, code, and reason through complex tasks. It offers 2M token context, reduced hallucinations, better LaTeX for STEM, improved image/video understanding, and 95% MMLU-Pro accuracy, targeting superior reasoning

Input type

Output Type

Input2-4/M tokens

Output6-12/M tokens

Context2.00M

Max Output30.00K

Available on 1 provider

Apr 30, 2026 6:00 PM-

xAI: Grok 4.2 Fast Non Reasoning

x-ai/grok-4.2-fast-non-reasoning

API Request Chat

218.53Mtokens

Input type

Output Type

Input2-4/M tokens

Output6-12/M tokens

Context2.00M

Max Output30.00K

Available on 1 provider

Apr 30, 2026 6:00 PM-

OpenAI: GPT-5.4

openai/gpt-5.4

API Request Chat

83.40Btokens

GPT‑5.4, released March 5, 2026, is OpenAI’s newest frontier model aimed at complex professional work. It’s available in the API as gpt-5.4 and in ChatGPT as GPT‑5.4 Thinking, with a higher-end GPT‑5.4 Pro tier for maximum performance. It supports adjustable deliberation through reasoning.effort, letting developers trade latency/cost for deeper reasoning. The model accepts text and image inputs and produces text output, with a 1,050,000‑token context window and up to 128,000 output tokens; its published knowledge cutoff is Aug 31, 2025. OpenAI also published a system card describing new cybersecurity-focused mitigations for the Thinking variant.

Input type

Output Type

Input2.5-5/M tokens

Output15-22.5/M tokens

Context1.05M

Max Output128.00K

Available on 2 providers

Apr 30, 2026 6:00 PM-

OpenAI: GPT-5.4 Pro

openai/gpt-5.4-pro

API Request Chat

371.13Mtokens

GPT-5.4 Pro is OpenAI's most advanced model, building on GPT-5.4's unified architecture with enhanced reasoning capabilities for complex, high-stakes tasks. It features a 1M+ token context window (922K input, 128K output) with support for text and image inputs. Optimized for step-by-step reasoning, instruction following, and accuracy, GPT-5.4 Pro excels at agentic coding, long-context workflows, and multi-step problem solving.

Input type

Output Type

Input30-60/M tokens

Output180-270/M tokens

Context1.05M

Max Output128.00K

Available on 2 providers

Apr 30, 2026 6:00 PM-

Google: Gemini 3.1 Flash Lite Preview

google/gemini-3.1-flash-lite-preview

API Request Chat

15.07Btokens

Gemini 3.1 Flash-Lite is Google's most cost-efficient Gemini model, optimized for low latency use cases for high-volume, cost-sensitive LLM traffic. It provides a significant quality increase over Gemini 2.0/2.5 Flash-Lite models, matching Gemini 2.5 Flash performance across key capability areas.

Input type

Output Type

Input0.25/M tokens

Output1.5/M tokens

Context1.05M

Max Output65.53K

Available on 2 providers

Apr 30, 2026 6:00 PM-

OpenAI: GPT-5.3 Chat

openai/gpt-5.3-chat

API Request Chat

459.50Mtokens

GPT-5.3 Chat points to the GPT-5.3 Instant snapshot currently used in ChatGPT. We recommend GPT-5.2 for API usage, but feel free to use this GPT-5.3 Chat model to test our latest improvements for chat use cases.

Input type

Output Type

Input1.75/M tokens

Output14/M tokens

Context128.00K

Max Output16.38K

Available on 2 providers

Apr 30, 2026 6:00 PM-

Qwen: Qwen-Image-2.0

qwen/qwen-image-2.0

API Request Chat

678.00Ktokens

Qwen-Image-2.0 is an image generation model launched by Qwen AI.

Input type

Output Type

Input-

Output0.0289/counts

Context-

Max Output-

Available on 1 provider

Apr 30, 2026 6:00 PM-

Qwen: Qwen-Image-2.0-Pro

qwen/qwen-image-2.0-pro

API Request Chat

1.03Mtokens

Qwen-Image-2.0-Pro is an image generation model launched by Qwen AI.

Input type

Output Type

Input-

Output0.073/counts

Context-

Max Output-

Available on 1 provider

Apr 30, 2026 6:00 PM-

ByteDance: Doubao-Seedream-5.0-lite

bytedance/doubao-seedream-5.0-lite

API Request Chat

844.00Ktokens

“Doubao-Seedream-5.0-lite is ByteDance’s latest image-generation model. For the first time, it integrates an online retrieval feature, enabling it to incorporate real-time web information and improve the timeliness of generated images. At the same time, the model’s intelligence has been further upgraded, allowing it to accurately interpret complex instructions and visual content. In addition, the model has been enhanced in terms of the breadth of world knowledge, reference consistency, and generation quality in professional scenarios, better meeting enterprise-level visual creation needs.”

Input type

Output Type

Input-

Output0.032/counts

Context-

Max Output-

Available on 1 provider

Apr 30, 2026 6:00 PM-

Google: Nano Banana 2 (Gemini 3.1 Flash Image Preview)

google/gemini-3.1-flash-image-preview

API Request Chat

743.23Mtokens

Designed for speed and efficiency, the Gemini 3.1 Flash Image generation model is effective for quick, interactive responses and high throughput.

Preview models may change before becoming stable and have more restrictive rate limits.

Input type

Output Type

Input0.5/M tokens

Output3/M tokens

Context65.54K

Max Output32.77K

Available on 2 providers

Apr 30, 2026 6:00 PM-

Qwen: Qwen3.5-Flash

qwen/qwen3.5-flash

API Request Chat

3.18Btokens

Qwen3.5 native vision-language Flash model series, based on a hybrid architecture design, integrates linear attention mechanisms with sparse Mixture-of-Experts (MoE) models, achieving higher inference efficiency. The model's performance has made a leap forward compared to the 3 series in both pure text and multimodal capabilities; it offers fast response times, balancing inference speed and performance.

Input type

Output Type

Input0.1/M tokens

Output0.4/M tokens

Context1.02M

Max Output1.02M

Available on 1 provider

Apr 30, 2026 6:00 PM-

Qwen: Qwen3.6 Flash

qwen/qwen3.6-flash

API Request Chat

2.47Mtokens

Qwen3.6 Flash is a fast, efficient language model from Alibaba's Qwen 3.6 series. It supports text, image, and video input with a 1M token context window. Tiered pricing kicks in above 256K tokens. Prompt caching is supported, with both explicit cache read and cache creation pricing.

Input type

Output Type

Input0.25-1/M tokens

Output1.5-4/M tokens

Context1000.00K

Max Output65.54K

Available on 1 provider

Apr 30, 2026 6:00 PM-

Google: Gemini 3.1 Pro Preview

google/gemini-3.1-pro-preview

API Request Chat

25.93Btokens

Gemini 3.1 Pro is the next generation in the Gemini series of models, a suite of highly-capable, natively multimodal, reasoning models. Gemini 3 Pro is now Google’s most advanced model for complex tasks, and can comprehend vast datasets, challenging problems from different information sources, including text, audio, images, video, and entire code repositories

Input type

Output Type

Input2-4/M tokens

Output12-18/M tokens

Context1.05M

Max Output65.53K

Available on 2 providers

Apr 30, 2026 6:00 PM-

Anthropic: Claude Sonnet 4.6

anthropic/claude-sonnet-4.6

API Request Chat

440.15Btokens

Sonnet 4.6 is Anthropic's most capable Sonnet-class model yet, with frontier performance across coding, agents, and professional work. It excels at iterative development, complex codebase navigation, end-to-end project management with memory, polished document creation, and confident computer use for web QA and workflow automation.

Input type

Output Type

Input3/M tokens

Output15/M tokens

Context1000.00K

Max Output64.00K

Available on 3 providers

Apr 30, 2026 6:00 PM-

Qwen: Qwen3.5-Plus

qwen/qwen3.5-plus

API Request Chat

4.62Btokens

Qwen3.5 Native Visual Language Series Plus model, based on a hybrid architecture design, integrates linear attention mechanisms with sparse mixture-of-experts models, achieving higher reasoning efficiency. In multiple task evaluations, the 3.5 series has demonstrated outstanding performance comparable to current top-tier cutting-edge models, with model performance showing leaps and bounds of progress compared to the 3 series in both pure text and multimodal aspects.

Input type

Output Type

Input0.4-0.5/M tokens

Output2.4-3/M tokens

Context1000.00K

Max Output64.00K

Available on 1 provider

Apr 30, 2026 6:00 PM-

ByteDance: Doubao-Seed-2.0-Code

bytedance/doubao-seed-2.0-code

API Request Chat

75.53Mtokens

Doubao-Seed-2.0-Code is optimized for enterprise-level programming needs. Building on the excellent Agent and VLM capabilities of Seed 2.0, it has significantly enhanced code capabilities. Not only does it have outstanding front-end performance, but it has also been specially optimized for the multi-language coding requirements common in enterprises, making it suitable for integration with various AI programming tools.

Input type

Output Type

Input0.45-1.34/M tokens

Output2.24-6.71/M tokens

Context256.00K

Max Output32.00K

Available on 1 provider

Apr 30, 2026 6:00 PM-

ByteDance: Doubao-Seed-2.0-lite

bytedance/doubao-seed-2.0-lite

API Request Chat

387.96Mtokens

Doubao-Seed-2.0-lite is a balanced model designed for high-frequency enterprise scenarios, balancing performance and cost, and its overall capabilities surpass those of its predecessor, Doubao-Seed-1.8. It excels in production-oriented tasks such as unstructured information processing, content creation, search and recommendation, and data analysis, supporting long contexts, multi-source information fusion, multi-step instruction execution, and high-fidelity structured output. It significantly optimizes costs while ensuring stable performance.

Input type

Output Type

Input0.09-0.25/M tokens

Output0.51-1.51/M tokens

Context256.00K

Max Output32.00K

Available on 1 provider

Apr 30, 2026 6:00 PM-

ByteDance: Doubao-Seed-2.0-mini

bytedance/doubao-seed-2.0-mini

API Request Chat

461.63Mtokens

Doubao-Seed-2.0-mini is designed for low-latency, high-concurrency, and cost-sensitive scenarios, emphasizing rapid response and flexible inference deployment. Its model performance is comparable to Doubao-Seed-1.6. It supports 256k context, four levels of think length, and multimodal understanding, making it suitable for lightweight tasks where cost and speed are paramount.

Input type

Output Type

Input0.03-0.12/M tokens

Output0.28-1.12/M tokens

Context256.00K

Max Output32.00K

Available on 1 provider

Apr 30, 2026 6:00 PM-

ByteDance: Doubao-Seed-2.0-pro

bytedance/doubao-seed-2.0-pro

API Request Chat

227.83Mtokens

Doubao-Seed-2.0-pro is a flagship, all-around general-purpose model designed for complex reasoning and long-chain task execution scenarios in the Agent era. It emphasizes multimodal understanding, long-context reasoning, structured generation, and tool-enhanced execution. Its capabilities in executing complex instructions and multiple constraints are outstanding, stably handling scenarios such as multi-step complex planning, complex graph and text reasoning, video content understanding, and high-difficulty analysis.

Input type

Output Type

Input0.45-1.34/M tokens

Output2.24-6.71/M tokens

Context256.00K

Max Output32.00K

Available on 1 provider

Apr 30, 2026 6:00 PM-

MiniMax: MiniMax M2.5

minimax/minimax-m2.5

API Request Chat

6.13Btokens

MiniMax-M2.5 is a lightweight, state-of-the-art large language model optimized for coding, agentic workflows, and modern application development. With only 10 billion activated parameters, it delivers a major jump in real-world capability while maintaining exceptional latency, scalability, and cost efficiency.

Compared to its predecessor, M2.1 delivers cleaner, more concise outputs and faster perceived response times. It shows leading multilingual coding performance across major systems and application languages, achieving 49.4% on Multi-SWE-Bench and 72.5% on SWE-Bench Multilingual, and serves as a versatile agent “brain” for IDEs, coding tools, and general-purpose assistance.

To avoid degrading this model's performance, MiniMax highly recommends preserving reasoning between turns.

Input type

Output Type

Input0.3/M tokens

Output1.2/M tokens

Context204.80K

Max Output131.07K

Available on 2 providers

Apr 30, 2026 6:00 PM-

MiniMax: MiniMax M2.5 highspeed

minimax/minimax-m2.5-lightning

API Request Chat

971.02Mtokens

M2.5 highspeed: Same performance, faster, more agile

Input type

Output Type

Input0.6/M tokens

Output2.4/M tokens

Context204.80K

Max Output131.07K

Available on 1 provider

Apr 30, 2026 6:00 PM-

Z.AI: GLM 5

z-ai/glm-5

API Request Chat

30.88Btokens

GLM-5 is Zai’s new-generation flagship foundation model, designed for Agentic Engineering, capable of providing reliable productivity in complex system engineering and long-range Agent tasks. In terms of Coding and Agent capabilities, GLM-5 has achieved state-of-the-art (SOTA) performance in open source, with its usability in real programming scenarios approaching that of Claude Opus 4.5.

Input type

Output Type

Input0.58-0.87/M tokens

Output2.6-3.18/M tokens

Context200.00K

Max Output128.00K

Available on 3 providers

Apr 30, 2026 6:00 PM-

OpenAI: GPT-5.3-Codex

openai/gpt-5.3-codex

API Request Chat

11.76Btokens

GPT‑5.3‑Codex is OpenAI’s most capable agentic coding model, designed to handle not just code generation but longer, end‑to‑end work on a computer. It combines the frontier coding strength of GPT‑5.2‑Codex with the reasoning and professional knowledge capabilities of GPT‑5.2, and is engineered to run faster (about 25% faster in Codex). It’s built for long‑running, multi‑step tasks that can include research, tool use, debugging, and complex execution, while staying interactive—so you can steer it, ask questions, and refine direction as it works without losing context. In evaluations, it sets new highs on software engineering and agentic benchmarks (including SWE‑Bench Pro and Terminal‑Bench 2.0) and shows strong results on OSWorld‑Verified and GDPval. It’s also deployed with strengthened cybersecurity safeguards and is trained to help identify software vulnerabilities.

Input type

Output Type

Input1.75/M tokens

Output14/M tokens

Context400.00K

Max Output128.00K

Available on 2 providers

Apr 30, 2026 6:00 PM-

Anthropic: Claude Opus 4.6

anthropic/claude-opus-4.6

API Request Chat

736.43Btokens

Opus 4.6 is Anthropic’s strongest model for coding and long-running professional tasks. It is built for agents that operate across entire workflows rather than single prompts, making it especially effective for large codebases, complex refactors, and multi-step debugging that unfolds over time. The model shows deeper contextual understanding, stronger problem decomposition, and greater reliability on hard engineering tasks than prior generations.

Beyond coding, Opus 4.6 excels at sustained knowledge work. It produces near-production-ready documents, plans, and analyses in a single pass, and maintains coherence across very long outputs and extended sessions. This makes it a strong default for tasks that require persistence, judgment, and follow-through, such as technical design, migration planning, and end-to-end project execution.

For users upgrading from earlier Opus versions, see our official migration guide here

Input type

Output Type

Input5/M tokens

Output25/M tokens

Context1000.00K

Max Output128.00K

Available on 3 providers

Apr 30, 2026 6:00 PM-

StepFun: Step 3.5 Flash

stepfun/step-3.5-flash

API Request Chat

2.27Btokens

Step 3.5 Flash is StepFun's most capable open-source foundation model. Built on a sparse Mixture of Experts (MoE) architecture, it selectively activates only 11B of its 196B parameters per token. It is a reasoning model that is incredibly speed efficient even at long contexts.

Input type

Output Type

Input0.1/M tokens

Output0.3/M tokens

Context256.00K

Max Output256.00K

Available on 1 provider

Apr 30, 2026 6:00 PM-

MoonshotAI: Kimi K2.5

moonshotai/kimi-k2.5

API Request Chat

10.37Btokens

Kimi K2.5 is Kimi's most intelligent model to date, achieving open-source SoTA performance in Agent capabilities, coding, visual understanding, and a range of general intelligence tasks. At the same time, Kimi K2.5 is also Kimi's most versatile model yet. Built with a natively multimodal architecture, it seamlessly supports both visual and text inputs, thinking and non-thinking modes, as well as conversational and Agent-based tasks.

Input type

Output Type

Input0.58/M tokens

Output3.02/M tokens

Context262.14K

Max Output262.14K

Available on 2 providers

Apr 30, 2026 6:00 PM-

MiniMax: MiniMax M2-her

minimax/minimax-m2-her

API Request Chat

2.13Mtokens

M2-her text chat model, designed for role-playing, multi-turn conversations and dialogue scenarios.

Input type

Output Type

Input0.3/M tokens

Output1.2/M tokens

Context64.00K

Max Output2.05K

Available on 1 provider

Apr 30, 2026 6:00 PM-

Qwen: Qwen3-Max-Thinking

qwen/qwen3-max

API Request Chat

969.48Mtokens

Qwen3-Max-Thinking is an updated release built on the Qwen3 series, offering major improvements in reasoning, instruction following, multilingual support, and long-tail knowledge coverage compared to the January 2025 version. It delivers higher accuracy in math, coding, logic, and science tasks, follows complex instructions in Chinese and English more reliably, reduces hallucinations, and produces higher-quality responses for open-ended Q&A, writing, and conversation. The model supports over 100 languages with stronger translation and commonsense reasoning, and is optimized for retrieval-augmented generation (RAG) and tool calling, though it does not include a dedicated “thinking” mode.

Input type

Output Type

Input1.2-3/M tokens

Output6-15/M tokens

Context256.00K

Max Output32.00K

Available on 1 provider

Apr 30, 2026 6:00 PM-

Baidu: ERNIE 5.0

baidu/ernie-5.0-thinking-preview

API Request Chat

740.68Mtokens

The new-generation Wenxin model, Wenxin 5.0, is a natively multimodal large model. It adopts a native unified multimodal modeling approach to jointly model text, images, audio, and video, providing comprehensive multimodal capabilities. Wenxin 5.0’s core abilities have been comprehensively upgraded and it performs excellently on benchmark datasets, with particularly strong results in multimodal understanding, instruction following, creative writing, factuality, agent planning, and tool use

Input type

Output Type

Input0.84-1.41/M tokens

Output3.37-5.62/M tokens

Context128.00K

Max Output64.00K

Available on 1 provider

Apr 30, 2026 6:00 PM-

Z.AI: GLM 4.7 Flash (Free)

z-ai/glm-4.7-flash-free

API Request Chat

804.85Mtokens

Free

Rate Limit

As a 30B-class SOTA model, GLM-4.7-Flash offers a new option that balances performance and efficiency. It is further optimized for agentic coding use cases, strengthening coding capabilities, long-horizon task planning, and tool collaboration, and has achieved leading performance among open-source models of the same size on several current public benchmark leaderboards.

Input type

Output Type

Input0/M tokens

Output0/M tokens

Context200.00K

Max Output128.00K

Available on 2 providers

Apr 30, 2026 6:00 PM-

Z.AI: GLM 4.7 FlashX

z-ai/glm-4.7-flashx

API Request Chat

315.94Mtokens

As a 30B-class SOTA model, GLM-4.7-FlashX offers a new option that balances performance and efficiency. It is further optimized for agentic coding use cases, strengthening coding capabilities, long-horizon task planning, and tool collaboration, and has achieved leading performance among open-source models of the same size on several current public benchmark leaderboards.

Input type

Output Type

Input0.0728/M tokens

Output0.4367/M tokens

Context200.00K

Max Output128.00K

Available on 2 providers

Apr 30, 2026 6:00 PM-

Z.AI: GLM-Image

z-ai/glm-image

API Request Chat

138.00Ktokens

GLM-Image is an image generation model adopts a hybrid autoregressive + diffusion decoder architecture. In general image generation quality, GLM‑Image aligns with mainstream latent diffusion approaches, but it shows significant advantages in text-rendering and knowledge‑intensive generation scenarios. It performs especially well in tasks requiring precise semantic understanding and complex information expression, while maintaining strong capabilities in high‑fidelity and fine‑grained detail generation. In addition to text‑to‑image generation, GLM‑Image also supports a rich set of image‑to‑image tasks including image editing, style transfer, identity‑preserving generation, and multi‑subject consistency.

Input type

Output Type

Input-

Output0.0146/counts

Context-

Max Output-

Available on 2 providers

Apr 30, 2026 6:00 PM-

OpenAI: GPT-5.2-Codex

openai/gpt-5.2-codex

API Request Chat

9.38Btokens

GPT-5.2-Codex is an upgraded version of GPT-5.2 optimized for agentic coding tasks in Codex or similar environments. GPT-5.2-Codex supports low, medium, high, and xhigh reasoning effort settings.

Input type

Output Type

Input1.75/M tokens

Output14/M tokens

Context400.00K

Max Output128.00K

Available on 2 providers

Apr 30, 2026 6:00 PM-

Google: Veo 3.1

google/veo-3.1-generate-001

API Request Create

352.04Ktokens

Veo 3.1 is Google's state-of-the-art model for generating high-fidelity, 8-second 720p, 1080p or 4k videos featuring stunning realism and natively generated audio. You can access this model programmatically using the Gemini API. To learn more about the available Veo model variants, see the Model Versions section.

Input type

Output Type

Input-

Output0.4-0.6/seconds

Context-

Max Output-

Available on 1 provider

Apr 30, 2026 6:00 PM-

Z.AI: GLM 4.7

z-ai/glm-4.7

API Request Chat

15.01Btokens

Pricing: The model's official pricing is (GLM-4.7: Input $0.28 ~ $0.57, Cached Input $0.057$0.11, Output $1.14$2.27). Our platform is currently running a limited-time promotion, during which you will receive a discount on the official price. GLM-4.7 is Zhipu’s latest flagship model. Tailored for agentic coding scenarios, GLM-4.7 strengthens coding capabilities, long-horizon task planning, and tool collaboration, and delivers leading performance among open-source models on the latest leaderboards of multiple public benchmarks. Its general capabilities have also improved, with responses that are more concise and natural and writing that feels more immersive. When executing complex agent tasks and invoking tools, it follows instructions more reliably, while the visual quality of artifacts and agentic coding front ends—as well as long-horizon task completion efficiency—are further enhanced.

Input type

Output Type

Input0.2911-0.5823/M tokens

Output1.1645-2.3291/M tokens

Context200.00K

Max Output128.00K

Available on 3 providers

Apr 30, 2026 6:00 PM-

MiniMax: MiniMax M2.1

minimax/minimax-m2.1

API Request Chat

1.04Btokens

MiniMax-M2.1 is a lightweight, state-of-the-art large language model optimized for coding, agentic workflows, and modern application development. With only 10 billion activated parameters, it delivers a major jump in real-world capability while maintaining exceptional latency, scalability, and cost efficiency.

To avoid degrading this model's performance, MiniMax highly recommends preserving reasoning between turns.

Input type

Output Type

Input0.3/M tokens

Output1.2/M tokens

Context204.80K

Max Output131.07K

Available on 1 provider

Apr 30, 2026 6:00 PM-

ByteDance: Doubao-Seed-1.8

bytedance/doubao-seed-1.8

API Request Chat

197.57Mtokens

An all-new model purpose-built and optimized for multimodal agent scenarios. Stronger agent capabilities, upgraded multimodal understanding, and more flexible context management

Input type

Output Type

Input0.11-0.34/M tokens

Output0.28-3.41/M tokens

Context256.00K

Max Output32.00K

Available on 1 provider

Apr 30, 2026 6:00 PM-

Google: Gemini 3 Flash Preview

google/gemini-3-flash-preview

API Request Chat

79.40Btokens

Gemini 3 Flash Preview is a low-latency model in the Gemini 3 family, optimized for fast, high-throughput inference. It retains the core multimodal and reasoning capabilities of Gemini 3 while prioritizing responsiveness and execution efficiency. Built on the same architecture as Gemini 3 Pro, Gemini 3 Flash Preview supports native multimodal inputs—including text, images, and audio—and incorporates the improved reasoning and long-context handling introduced in the Gemini 3 generation. It is designed for real-time and scalable workloads where latency and cost efficiency are primary considerations.

Input type

Output Type

Input0.5/M tokens

Output3/M tokens

Context1.05M

Max Output65.53K

Available on 2 providers

Apr 30, 2026 6:00 PM-

Xiaomi: MiMo-V2-Flash

xiaomi/mimo-v2-flash

API Request Chat

8.51Btokens

MiMo-V2-Flash is an open-source foundation language model developed by Xiaomi. It is a Mixture-of-Experts model with 309B total parameters and 15B active parameters, adopting hybrid attention architecture. MiMo-V2-Flash supports a hybrid-thinking toggle and a 256K context window, and excels at reasoning, coding, and agent scenarios. On SWE-bench Verified and SWE-bench Multilingual, MiMo-V2-Flash ranks as the top #1 open-source model globally, delivering performance comparable to Claude Sonnet 4.5 while costing only about 3.5% as much.

Input type

Output Type

Input0.1/M tokens

Output0.3/M tokens

Context262.14K

Max Output262.14K

Available on 2 providers

Apr 30, 2026 6:00 PM-

OpenAI: GPT-Image-1.5

openai/gpt-image-1.5

API Request Chat

8.08Mtokens

OpenAI's GPT Image 1.5 is the latest evolution of its AI image generation, offering superior instruction following, photorealism, text rendering, and editing control, making it ideal for detailed creative and production work with faster speeds and lower costs than predecessors like DALL-E 3. It excels at complex requests, maintaining character/style consistency, rendering crisp text in visuals, and understanding nuanced prompts through built-in reasoning, integrated into ChatGPT and available via API.

Input type

Output Type

Input5/M tokens

Output10/M tokens

Context-

Max Output-

Available on 1 provider

Apr 30, 2026 6:00 PM-

ByteDance/Doubao-Seedance-1.5-pro

bytedance/doubao-seedance-1.5-pro

API Request Create

12.41Mtokens

Seedance 1.5 Pro is a new-generation professional-grade audio-visual co-generation video model released by the Doubao large-model team. Building on its predecessor’s multi-shot storytelling and high-definition generation capabilities, it natively supports integrated audio-and-video output, aiming to deliver an end-to-end synchronized creation experience across visuals, voice, music, and sound effects. The model also includes a built-in first-and-last-frame feature: creators only need to set the opening and ending frames of a video to precisely lock in its style, composition, and characters, which then drives the model to generate smooth, natural motion between frames. By combining audio-visual co-generation with first/last-frame control, Seedance 1.5 Pro significantly improves the efficiency, controllability, and artistic expressiveness of professional video creation.

Input type

Output Type

Input-

Output2.33/M tokens

Context-

Max Output-

Available on 1 provider

Apr 30, 2026 6:00 PM-

OpenAI: GPT-5.2 Pro

openai/gpt-5.2-pro

API Request Chat

519.87Mtokens

GPT-5.2 Pro is OpenAI’s most advanced model, offering major improvements in agentic coding and long context performance over GPT-5 Pro. It is optimized for complex tasks that require step-by-step reasoning, instruction following, and accuracy in high-stakes use cases. It supports test-time routing features and advanced prompt understanding, including user-specified intent like "think hard about this." Improvements include reductions in hallucination, sycophancy, and better performance in coding, writing, and health-related tasks.

Input type

Output Type

Input21/M tokens

Output168/M tokens

Context400.00K

Max Output128.00K

Available on 1 provider

Apr 30, 2026 6:00 PM-

OpenAI: GPT-5.2

openai/gpt-5.2

API Request Chat

34.28Btokens

GPT-5.2 is the newest frontier-grade model in the GPT-5 family, outperforming GPT-5.1 in agentic capabilities and long-context handling. It uses adaptive reasoning to dynamically allocate compute—responding quickly to straightforward questions while applying deeper analysis to more complex tasks. Designed for broad task coverage, GPT-5.2 delivers consistent improvements across math, coding, science, and tool-calling workloads, producing more coherent long-form responses and more reliable tool use.

Input type

Output Type

Input1.75/M tokens

Output14/M tokens

Context400.00K

Max Output128.00K

Available on 2 providers

Apr 30, 2026 6:00 PM-

OpenAI: GPT-5.2 Chat

openai/gpt-5.2-chat

API Request Chat

407.44Mtokens

GPT-5.2 Chat (AKA Instant) is the fast, lightweight member of the 5.2 family, optimized for low-latency chat while retaining strong general intelligence. It uses adaptive reasoning to selectively “think” on harder queries, improving accuracy on math, coding, and multi-step tasks without slowing down typical conversations. The model is warmer and more conversational by default, with better instruction following and more stable short-form reasoning. GPT-5.2 Chat is designed for high-throughput, interactive workloads where responsiveness and consistency matter more than deep deliberation.

Input type

Output Type

Input1.75/M tokens

Output14/M tokens

Context128.00K

Max Output16.38K

Available on 2 providers

Apr 30, 2026 6:00 PM-

Z.AI: GLM 4.6V

z-ai/glm-4.6v

API Request Chat

105.49Mtokens

GLM-4.6V represents a significant evolution of the GLM series into the multimodal domain. It features a 128k-token training context window and sets a new state-of-the-art in visual understanding accuracy for its parameter scale. Pioneeringly, it is the first model to natively integrate tool-calling capabilities into its visual architecture, bridging the gap from visual perception to executable actions. This makes it a unified technical foundation for multimodal Agents in real-world business scenarios. Pricing: The model's official pricing is (GLM-4.6v: Input $0.3, Cached Input $0.05, Output $0.9). Our platform is currently running a limited-time promotion, during which you will receive a discount on the official price.

Input type

Output Type

Input0.1456-0.2911/M tokens

Output0.4367-0.8734/M tokens

Context200.00K

Max Output128.00K

Available on 2 providers

Apr 30, 2026 6:00 PM-

Z.AI: GLM 4.6V Flash (Free)

z-ai/glm-4.6v-flash-free

API Request Chat

291.69Mtokens

Free

Rate Limit

GLM-4.6V-Flash is the free version of GLM-4.6V, representing a significant iteration in the GLM series for multimodal capabilities. It supports toggling reasoning modes, with a training-time context window expanded to 128k tokens. Achieving state-of-the-art (SOTA) visual understanding accuracy at its parameter scale, it is the first visual model to natively integrate Function Call capability into its architecture. This establishes a seamless pipeline from "visual perception" to "executable actions (Action)", offering a unified technical foundation for multimodal agents in real-world business scenarios.

Input type

Output Type

Input0/M tokens

Output0/M tokens

Context200.00K

Max Output128.00K

Available on 2 providers

Apr 30, 2026 6:00 PM-

Z.AI: GLM 4.6V FlashX

z-ai/glm-4.6v-flash

API Request Chat

2.08Btokens

GLM-4.6V-FlashX is the the paid version, offering higher capacity and stability, representing a significant iteration in the GLM series for multimodal capabilities. It supports toggling reasoning modes, with a training-time context window expanded to 128k tokens. Achieving state-of-the-art (SOTA) visual understanding accuracy at its parameter scale, it is the first visual model to natively integrate Function Call capability into its architecture. This establishes a seamless pipeline from "visual perception" to "executable actions (Action)", offering a unified technical foundation for multimodal agents in real-world business scenarios.

Input type

Output Type

Input0.0218-0.0437/M tokens

Output0.2184-0.4367/M tokens

Context200.00K

Max Output128.00K

Available on 2 providers

Apr 30, 2026 6:00 PM-

DeepSeek: DeepSeek V3.2

deepseek/deepseek-v3.2

API Request Chat

4.07Btokens

DeepSeek-V3.2 is a reasoning-first large language model released by DeepSeek as the official successor to V3.2-Exp. It is designed with a focus on agentic capabilities and integrated reasoning for tool-use scenarios.

If the request to the deepseek-reasoner model includes the tools parameter, the request will actually be processed using the deepseek-chat model.

The model introduces a new large-scale agent training data synthesis method covering over 1,800 environments and 85,000+ complex instructions. DeepSeek-V3.2 is the first model from DeepSeek to integrate thinking directly into tool-use, supporting both thinking and non-thinking modes during tool interactions.

According to DeepSeek's benchmarks, the model delivers performance comparable to GPT-5 level while balancing inference quality and output length. It is available via App, Web, and API, making it suitable for general-purpose daily use as well as complex agentic workflows.

Input type

Output Type

Input0.293/M tokens

Output0.4395/M tokens

Context128.00K

Max Output8.00K

Available on 2 providers

Apr 30, 2026 6:00 PM-

Mistral: Mistral Large 3

mistralai/mistral-large-2512

API Request Chat

74.11Mtokens

Mistral Large 3, is a state-of-the-art, open-weight, general-purpose multimodal model with a granular Mixture-of-Experts architecture. It features 41B active parameters and 675B total parameters.

Input type

Output Type

Input0.5/M tokens

Output1.5/M tokens

Context256.00K

Max Output256.00K

Available on 2 providers

Apr 30, 2026 6:00 PM-

DeepSeek: DeepSeek-V3.2 (Non-thinking Mode)

deepseek/deepseek-chat

API Request Chat

13.82Btokens

DeepSeek-V3.2 (Non-thinking Mode) is DeepSeek's latest production model, currently served under the deepseek-chat model slug, which is automatically updated as new versions are released.

If the request to the deepseek-reasoner model includes the tools parameter, the request will actually be processed using the deepseek-chat model.

Input type

Output Type

Input0.14/M tokens

Output0.28/M tokens

Context128.00K

Max Output8.00K

Available on 1 provider

Apr 30, 2026 6:00 PM-

DeepSeek: DeepSeek-V3.2 (Thinking Mode)

deepseek/deepseek-reasoner

API Request Chat

4.45Btokens

DeepSeek-V3.2 (Thinking Mode) is DeepSeek's latest production model, currently served under the deepseek-reasoner model slug, which is automatically updated as new versions are released.

If the request to the deepseek-reasoner model includes the tools parameter, the request will actually be processed using the deepseek-chat model.

Input type

Output Type

Input0.14/M tokens

Output0.28/M tokens

Context128.00K

Max Output64.00K

Available on 1 provider

Apr 30, 2026 6:00 PM-

Anthropic: Claude Opus 4.5

anthropic/claude-opus-4.5

API Request Chat

63.05Btokens

Claude Opus 4.5 is Anthropic's latest frontier reasoning model, purpose-built for complex software engineering, agentic workflows, and long-horizon computer use. It delivers strong multimodal capabilities, competitive performance on real-world coding and reasoning benchmarks, and improved robustness against prompt injection attacks. The model is designed to operate efficiently across varied effort levels, allowing developers to balance speed, depth, and token usage based on their specific task requirements—you can fine-tune token efficiency through the OpenRouter Verbosity parameter, which offers low, medium, and high settings. Beyond that, Opus 4.5 supports advanced tool use, extended context management, and coordinated multi-agent setups, making it ideal for autonomous research, debugging, multi-step planning, and spreadsheet or browser manipulation. Compared to previous Opus generations, it brings substantial improvements in structured reasoning, execution reliability, and alignment, while reducing token overhead and delivering more consistent performance on long-running tasks.

Input type

Output Type

Input5/M tokens

Output25/M tokens

Context200.00K

Max Output32.00K

Available on 3 providers

Apr 30, 2026 6:00 PM-

Google: Nano Banana Pro (Gemini 3 Pro Image Preview)

google/gemini-3-pro-image-preview

API Request Chat

1.68Btokens

Nano Banana Pro is an AI image generation model that is a significant upgrade from its predecessor, built on Google's Gemini 3 Pro. It promises to move beyond simple pattern matching to a more reasoning-driven system with improved physics understanding, text rendering, and image consistency. Key features include faster processing, native 2K resolution, and the ability to edit existing images with greater control, aiming to produce more reliable and professional-grade results.

Input type

Output Type

Input2-4/M tokens

Output12-18/M tokens

Context65.54K

Max Output32.77K

Available on 2 providers

Apr 30, 2026 6:00 PM-

xAI: Grok 4.1 Fast

x-ai/grok-4.1-fast

API Request Chat

9.26Btokens

Grok 4.1 Fast is xAI's best agentic tool calling model that shines in real-world use cases like customer support and deep research. 2M context window.

Reasoning can be enabled/disabled using the reasoning``enabled parameter in the API.

Input type

Output Type

Input0.2-0.4/M tokens

Output0.5-1/M tokens

Context2.00M

Max Output30.00K

Available on 1 provider

Apr 30, 2026 6:00 PM-

xAI: Grok 4.1 Fast Non Reasoning

x-ai/grok-4.1-fast-non-reasoning

API Request Chat

21.19Btokens

Grok 4.1 Fast Non Reasoning is xAI's best agentic tool calling model that shines in real-world use cases like customer support and deep research. 2M context window.

Input type

Output Type

Input0.2-0.4/M tokens

Output0.5-1/M tokens

Context2.00M

Max Output30.00K

Available on 1 provider

Apr 30, 2026 6:00 PM-

OpenAI: GPT-5.1

openai/gpt-5.1

API Request Chat

1.40Btokens

GPT-5.1 is the latest frontier-grade model in the GPT-5 series, offering stronger general-purpose reasoning, improved instruction adherence, and a more natural conversational style compared to GPT-5. It uses adaptive reasoning to allocate computation dynamically, responding quickly to simple queries while spending more depth on complex tasks. The model produces clearer, more grounded explanations with reduced jargon, making it easier to follow even on technical or multi-step problems.

Built for broad task coverage, GPT-5.1 delivers consistent gains across math, coding, and structured analysis workloads, with more coherent long-form answers and improved tool-use reliability. It also features refined conversational alignment, enabling warmer, more intuitive responses without compromising precision. GPT-5.1 serves as the primary full-capability successor to GPT-5

Input type

Output Type

Input1.25/M tokens

Output10/M tokens

Context400.00K

Max Output128.00K

Available on 2 providers

Apr 30, 2026 6:00 PM-

OpenAI: GPT-5.1 Chat

openai/gpt-5.1-chat

API Request Chat

156.49Mtokens

GPT-5.1 Chat (AKA Instant is the fast, lightweight member of the 5.1 family, optimized for low-latency chat while retaining strong general intelligence. It uses adaptive reasoning to selectively “think” on harder queries, improving accuracy on math, coding, and multi-step tasks without slowing down typical conversations. The model is warmer and more conversational by default, with better instruction following and more stable short-form reasoning. GPT-5.1 Chat is designed for high-throughput, interactive workloads where responsiveness and consistency matter more than deep deliberation.

Input type

Output Type

Input1.25/M tokens

Output10/M tokens

Context128.00K

Max Output16.38K

Available on 2 providers

Apr 30, 2026 6:00 PM-

OpenAI: GPT-5.1-Codex

openai/gpt-5.1-codex

API Request Chat

922.49Mtokens

GPT-5.1-Codex is a specialized version of GPT-5.1 optimized for software engineering and coding workflows. It is designed for both interactive development sessions and long, independent execution of complex engineering tasks. The model supports building projects from scratch, feature development, debugging, large-scale refactoring, and code review. Compared to GPT-5.1, Codex is more steerable, adheres closely to developer instructions, and produces cleaner, higher-quality code outputs.

Codex integrates into developer environments including the CLI, IDE extensions, GitHub, and cloud tasks. It adapts reasoning effort dynamically—providing fast responses for small tasks while sustaining extended multi-hour runs for large projects. The model is trained to perform structured code reviews, catching critical flaws by reasoning over dependencies and validating behavior against tests. It also supports multimodal inputs such as images or screenshots for UI development and integrates tool use for search, dependency installation, and environment setup. Codex is intended specifically for agentic coding applications.

Input type

Output Type

Input1.25/M tokens

Output10/M tokens

Context400.00K

Max Output128.00K

Available on 2 providers

Apr 30, 2026 6:00 PM-

OpenAI: GPT-5.1-Codex-Mini

openai/gpt-5.1-codex-mini

API Request Chat

780.89Mtokens

GPT-5.1-Codex-Mini is a smaller and faster version of GPT-5.1-Codex

Input type

Output Type

Input0.25/M tokens

Output2/M tokens

Context400.00K

Max Output100.00K

Available on 2 providers

Apr 30, 2026 6:00 PM-

ByteDance: Doubao-Seed-Code

bytedance/doubao-seed-code

API Request Chat

1.03Mtokens

Doubao-Seed-Code has been deeply optimized for Agentic Programming tasks, delivering exceptional performance across multiple authoritative benchmarks — including Terminal Bench, SWE-Bench-Verified-Openhands, and Multi-SWE-Bench-Flash-Openhands — outperforming domestic counterparts and supporting a context window of up to 256k tokens.

Input type

Output Type

Input0.17-0.39/M tokens

Output1.12-2.25/M tokens

Context256.00K

Max Output32.00K

Available on 1 provider

Apr 30, 2026 6:00 PM-

Tencent: HY 2.0 Think

tencent/hunyuan-2.0-thinking

API Request Chat

4.22Mtokens

Tencent HY 2.0 Think, a large language model fully developed end-to-end by Tencent, leads the industry with outstanding performance in high-quality content creation, mathematical logic reasoning, code generation and multi-turn conversations; its API supports an internet-connected AI search plugin that integrates Tencent's premium content ecosystem to provide powerful real-time, in-depth content retrieval and AI question-answering capabilities, and this release upgrades the model base from TurboS to HY 2.0 for overall capability enhancement, with significant improvements in complex instruction following, multi-turn and long-text comprehension, code generation, Agent support and reasoning abilities.

Input type

Output Type

Input0.57-0.76/M tokens

Output2.29-3.05/M tokens

Context128.00K

Max Output64.00K

Available on 1 provider

Apr 30, 2026 6:00 PM-

MoonshotAI: Kimi K2 Thinking

moonshotai/kimi-k2-thinking

API Request Chat

501.44Mtokens

A thinking model with general agentic and reasoning capabilities, specializing in deep reasoning tasks

Input type

Output Type

Input0.6/M tokens

Output2.5/M tokens

Context262.14K

Max Output262.14K

Available on 1 provider

Apr 30, 2026 6:00 PM-

MoonshotAI: Kimi K2 Thinking Turbo

moonshotai/kimi-k2-thinking-turbo

API Request Chat

138.11Mtokens

Context length 256k. High-speed version of kimi-k2-thinking, suitable for scenarios requiring both deep reasoning and extremely fast responses

Input type

Output Type

Input1.15/M tokens

Output8/M tokens

Context262.14K

Max Output262.14K

Available on 1 provider

Apr 30, 2026 6:00 PM-

OpenAI: Text Embedding 3 Small

openai/text-embedding-3-small

API Request Chat

41.14Mtokens

text-embedding-3-small is OpenAI's improved, more performant version of the ada embedding model. Embeddings are a numerical representation of text that can be used to measure the relatedness between two pieces of text. Embeddings are useful for search, clustering, recommendations, anomaly detection, and classification tasks.

Input type

Output Type

Input0.02/M tokens

Output0/M tokens

Context8.19K

Max Output8.19K

Available on 1 provider

Apr 30, 2026 6:00 PM-

MiniMax: MiniMax M2

minimax/minimax-m2

API Request Chat

775.27Mtokens

MiniMax-M2 is a compact, high-efficiency large language model optimized for end-to-end coding and agentic workflows. With 10 billion activated parameters (230 billion total), it delivers near-frontier intelligence across general reasoning, tool use, and multi-step task execution while maintaining low latency and deployment efficiency. The model excels in code generation, multi-file editing, compile-run-fix loops, and test-validated repair, showing strong results on SWE-Bench Verified, Multi-SWE-Bench, and Terminal-Bench. It also performs competitively in agentic evaluations such as BrowseComp and GAIA, effectively handling long-horizon planning, retrieval, and recovery from execution errors. Benchmarked by Artificial Analysis, MiniMax-M2 ranks among the top open-source models for composite intelligence, spanning mathematics, science, and instruction-following. Its small activation footprint enables fast inference, high concurrency, and improved unit economics, making it well-suited for large-scale agents, developer assistants, and reasoning-driven applications that require responsiveness and cost efficiency.

Input type

Output Type

Input0.3/M tokens

Output1.2/M tokens

Context204.80K

Max Output128.00K

Available on 1 provider

Apr 30, 2026 6:00 PM-

Anthropic: Claude Haiku 4.5

anthropic/claude-haiku-4.5

API Request Chat

99.78Btokens

Claude Haiku 4.5 is Anthropic’s fastest and most efficient model, delivering near-frontier intelligence at a fraction of the cost and latency of larger Claude models. Matching Claude Sonnet 4’s performance across reasoning, coding, and computer-use tasks, Haiku 4.5 brings frontier-level capability to real-time and high-volume applications.

It introduces extended thinking to the Haiku line; enabling controllable reasoning depth, summarized or interleaved thought output, and tool-assisted workflows with full support for coding, bash, web search, and computer-use tools. Scoring >73% on SWE-bench Verified, Haiku 4.5 ranks among the world’s best coding models while maintaining exceptional responsiveness for sub-agents, parallelized execution, and scaled deployment.

Input type

Output Type

Input1/M tokens

Output5/M tokens

Context200.00K

Max Output64.00K

Available on 2 providers

Apr 30, 2026 6:00 PM-

inclusionAI: Ring-1T

inclusionai/ring-1t

API Request Chat

1.38Btokens

Ring-1T is a trillion-parameter sparse mixture-of-experts (MoE) thinking model developed by inclusionAI. It adopts the Ling 2.0 architecture and is trained on the Ling-1T-base foundation model, which contains 1 trillion total parameters with 50 billion activated parameters, supporting a context window of up to 128K tokens. Building upon the preview version released at the end of September, Ring-1T has undergone continued scaling with large-scale verifiable reward reinforcement learning (RLVR) training, further unlocking the natural language reasoning capabilities of the trillion-parameter foundation model.

Input type

Output Type

Input0.56/M tokens

Output2.24/M tokens

Context128.00K

Max Output32.00K

Available on 1 provider

Apr 30, 2026 6:00 PM-

inclusionAI: Ling-1T

inclusionai/ling-1t

API Request Chat

7.75Btokens

Ling-1T is a trillion-parameter sparse mixture-of-experts (MoE) model developed by inclusionAI, optimized for efficient and scalable reasoning. Featuring approximately 50 billion active parameters per token, it is pre-trained on over 20 trillion reasoning-dense tokens, supports a 128K context length, and utilizes an Evolutionary Chain-of-Thought (Evo-CoT) process to enhance its reasoning depth. The model achieves state-of-the-art performance across complex benchmarks, demonstrating strong capabilities in code generation, software development, and advanced mathematics. In addition to its core reasoning skills, Ling-1T possesses specialized abilities in front-end code generation—combining semantic understanding with visual aesthetics—and exhibits emergent agentic capabilities, such as proficient tool use with minimal instruction tuning. Its primary use cases span software engineering, professional mathematics, complex logical reasoning, and agent-based workflows that demand a balance of high performance and efficiency.

Input type

Output Type

Input0.56/M tokens

Output2.24/M tokens

Context128.00K

Max Output32.00K

Available on 1 provider

Apr 30, 2026 6:00 PM-

Google: Gemini 2.5 Flash Image (Nano Banana)

google/gemini-2.5-flash-image

API Request Chat

233.13Mtokens

Gemini 2.5 Flash Image, a.k.a. "Nano Banana," is now generally available. It is a state of the art image generation model with contextual understanding. It is capable of image generation, edits, and multi-turn conversations.

Input type

Output Type

Input0.3/M tokens

Output2.5/M tokens

Context32.77K

Max Output8.19K

Available on 1 provider

Apr 30, 2026 6:00 PM-

OpenAI: GPT-5 Pro

openai/gpt-5-pro

API Request Chat

32.03Mtokens

GPT-5 Pro is OpenAI’s most advanced model, offering major improvements in reasoning, code quality, and user experience. It is optimized for complex tasks that require step-by-step reasoning, instruction following, and accuracy in high-stakes use cases. It supports test-time routing features and advanced prompt understanding, including user-specified intent like "think hard about this." Improvements include reductions in hallucination, sycophancy, and better performance in coding, writing, and health-related tasks.

Input type

Output Type

Input15/M tokens

Output120/M tokens

Context400.00K

Max Output128.00K

Available on 2 providers

Apr 30, 2026 6:00 PM-

Z.AI: GLM 4.6

z-ai/glm-4.6

API Request Chat

1.31Btokens

GLM-4.6 is a flagship model from Zhishen with 355B total parameters and 32B active parameters. The model's context window has been expanded from 128K to 200K, enabling it to handle longer code and agent tasks. In programming capabilities, GLM-4.6's performance is comparable to Claude Sonnet 4 on public benchmarks and real-world programming tasks. The model supports tool calling during inference and features optimized search and tool-use performance within agent frameworks. Furthermore, enhancements have been made to its writing style, readability, role-playing, and cross-lingual task processing abilities.

Pricing: The model's official pricing is (GLM-4.6: Input $0.6, Cached Input $0.11, Output $2.2). Our platform is currently running a limited-time promotion, during which you will receive a discount on the official price.

Input type

Output Type

Input0.2911-0.5823/M tokens

Output1.1645-2.3291/M tokens

Context200.00K

Max Output128.00K

Available on 3 providers

Apr 30, 2026 6:00 PM-

Anthropic: Claude Sonnet 4.5

anthropic/claude-sonnet-4.5

API Request Chat

117.01Btokens

Claude Sonnet 4.5 is Anthropic’s most advanced Sonnet model to date, optimized for real-world agents and coding workflows. It delivers state-of-the-art performance on coding benchmarks such as SWE-bench Verified, with improvements across system design, code security, and specification adherence. The model is designed for extended autonomous operation, maintaining task continuity across sessions and providing fact-based progress tracking.

Sonnet 4.5 also introduces stronger agentic capabilities, including improved tool orchestration, speculative parallel execution, and more efficient context and memory management. With enhanced context tracking and awareness of token usage across tool calls, it is particularly well-suited for multi-context and long-running workflows. Use cases span software engineering, cybersecurity, financial analysis, research agents, and other domains requiring sustained reasoning and tool use.

Input type

Output Type

Input3-6/M tokens

Output15-22.5/M tokens

Context1000.00K

Max Output64.00K

Available on 3 providers

Apr 30, 2026 6:00 PM-

DeepSeek: DeepSeek-V3.2-Exp

deepseek/deepseek-v3.2-exp

API Request Chat

214.87Mtokens

DeepSeek-V3.2-Exp is an experimental model version, serving as an intermediate step toward the next-generation architecture. Built on the foundation of V3.1-Terminus, it introduces the DeepSeek Sparse Attention (DSA) mechanism—a sparse attention mechanism designed to explore and validate the optimization of training and inference efficiency in long-context scenarios. This experimental version represents the team's continuous research on more efficient Transformer architectures, with a specific focus on improving computational efficiency when processing long text sequences. For the first time, DSA enables fine-grained sparse attention, which significantly enhances the efficiency of long-context training and inference while maintaining almost unchanged model output quality.

Input type

Output Type

Input0.216/M tokens

Output0.328/M tokens

Context163.84K

Max Output65.54K

Available on 1 provider

Apr 30, 2026 6:00 PM-

hunyuan-image3

tencent/hunyuan-image3

API Request Chat

126.00Ktokens

HunyuanImage-3.0 is a groundbreaking native multimodal model that unifies multimodal understanding and generation within an autoregressive framework. Our text-to-image and image-to-image model achieves performance comparable to or surpassing leading closed-source models.

Input type

Output Type

Input-

Output0.029/counts

Context-

Max Output-

Available on 1 provider

Apr 30, 2026 6:00 PM-

OpenAI: GPT-5 Codex

openai/gpt-5-codex

API Request Chat

299.24Mtokens

GPT-5-Codex is a specialized version of GPT-5 optimized for software engineering and coding workflows. It is designed for both interactive development sessions and long, independent execution of complex engineering tasks. The model supports building projects from scratch, feature development, debugging, large-scale refactoring, and code review. Compared to GPT-5, Codex is more steerable, adheres closely to developer instructions, and produces cleaner, higher-quality code outputs. Reasoning effort can be adjusted with the reasoning.effort parameter. Read the docs here

Input type

Output Type

Input1.25/M tokens

Output10/M tokens

Context400.00K

Max Output128.00K

Available on 2 providers

Apr 30, 2026 6:00 PM-

Qwen: Qwen3-VL-Plus

qwen/qwen3-vl-plus

API Request Chat

82.38Mtokens

The Qwen3 series VL models effectively integrates thinking and non-thinking modes, achieving world-leading performance in visual agent capabilities on public benchmark datasets such as OS World. This version features comprehensive upgrades in areas like visual coding, spatial perception, and multimodal reasoning, significantly enhancing visual perception and recognition abilities, and supporting the understanding of ultra-long videos.

Input type

Output Type

Input0.2-0.6/M tokens

Output1.6-4.8/M tokens

Context262.14K

Max Output32.77K

Available on 1 provider

Apr 30, 2026 6:00 PM-

xAI: Grok 4 Fast

x-ai/grok-4-fast

API Request Chat

10.34Btokens

Grok 4 Fast is xAI's latest multimodal model with SOTA cost-efficiency and a 2M token context window. It comes in two flavors: non-reasoning and reasoning. Read more about the model on xAI's news post. Reasoning can be enabled using the reasoning``enabled parameter in the API. Learn more in our docs

Prompts and completions may be used by xAI or OpenRouter to improve future models.

Input type

Output Type

Input0.2-0.4/M tokens

Output0.5-1/M tokens

Context2.00M

Max Output30.00K

Available on 1 provider

Apr 30, 2026 6:00 PM-

xAI: Grok 4 Fast None Reasoning

x-ai/grok-4-fast-non-reasoning

API Request Chat

3.55Btokens

Prompts and completions may be used by xAI or OpenRouter to improve future models.

Input type

Output Type

Input0.2-0.4/M tokens

Output0.5-1/M tokens

Context2.00M

Max Output30.00K

Available on 1 provider

Apr 30, 2026 6:00 PM-

inclusionAI: Ling-flash-2.0

inclusionai/ling-flash-2.0

API Request Chat

228.50Mtokens

Ling-flash-2.0 is an open-source Mixture-of-Experts (MoE) language model developed under the Ling 2.0 architecture. It features 100 billion total parameters, with 6.1 billion activated during inference (4.8B non-embedding).

Trained on over 20 trillion tokens and refined with supervised fine-tuning and multi-stage reinforcement learning, the model demonstrates strong performance against dense models up to 40B parameters. It excels in complex reasoning, code generation, and frontend development.

Input type

Output Type

Input0.28/M tokens

Output2.8/M tokens

Context128.00K

Max Output32.00K

Available on 1 provider

Apr 30, 2026 6:00 PM-

inclusionAI: Ring-flash-2.0

inclusionai/ring-flash-2.0

API Request Chat

218.15Mtokens

Input type

Output Type

Input0.28/M tokens

Output2.8/M tokens

Context128.00K

Max Output32.00K

Available on 1 provider

Apr 30, 2026 6:00 PM-

Baidu: ERNIE-X1.1-Preview

baidu/ernie-x1.1-preview

API Request Chat

3.67Mtokens

The Wenxin Large Model X1.1 delivers significantly enhanced performance in question answering, tool invocation, agent capabilities, instruction following, logical reasoning, mathematical problem-solving, and coding tasks, with markedly improved factual accuracy. Its context window has been extended to 64K tokens, enabling longer inputs and dialogue histories, while maintaining response speed and improving the coherence of long-chain reasoning.

Input type

Output Type

Input0.14/M tokens

Output0.56/M tokens

Context65.54K

Max Output65.54K

Available on 1 provider

Apr 30, 2026 6:00 PM-

MoonshotAI: Kimi K2 0905

moonshotai/kimi-k2-0905

API Request Chat

121.16Mtokens

Kimi K2 0905 is the September update of Kimi K2 0711 [blocked]. It is a large-scale Mixture-of-Experts (MoE) language model developed by Moonshot AI, featuring 1 trillion total parameters with 32 billion active per forward pass. It supports long-context inference up to 256k tokens, extended from the previous 128k.

This update improves agentic coding with higher accuracy and better generalization across scaffolds, and enhances frontend coding with more aesthetic and functional outputs for web, 3D, and related tasks. Kimi K2 is optimized for agentic capabilities, including advanced tool use, reasoning, and code synthesis. It excels across coding (LiveCodeBench, SWE-bench), reasoning (ZebraLogic, GPQA), and tool-use (Tau2, AceBench) benchmarks. The model is trained with a novel stack incorporating the MuonClip optimizer for stable large-scale MoE training.

Input type

Output Type

Input0.58/M tokens

Output2.33/M tokens

Context262.10K

Max Output262.10K

Available on 1 provider

Apr 30, 2026 6:00 PM-

inclusionAI: Ling-mini-2.0

inclusionai/ling-mini-2.0

API Request Chat

215.92Mtokens

Ling-mini-2.0 is an open-source Mixture-of-Experts (MoE) large language model designed to balance strong task performance with high inference efficiency. It has 16B total parameters, with approximately 1.4B activated per token (about 789M non-embedding). Trained on over 20T tokens and refined via multi-stage supervised fine-tuning and reinforcement learning, it is reported to deliver strong results in complex reasoning and instruction following while keeping computational costs low. According to the upstream release, it reaches top-tier performance among sub-10B dense LLMs and in some cases matches or surpasses larger MoE models.

Input type

Output Type

Input0.07/M tokens

Output0.28/M tokens

Context128.00K

Max Output32.00K

Available on 1 provider

Apr 30, 2026 6:00 PM-

inclusionAI: Ring-mini-2.0

inclusionai/ring-mini-2.0

API Request Chat

1.23Btokens

Ring-mini-2.0 is a Mixture-of-Experts (MoE) model oriented toward high-throughput inference and extensively optimized on the Ling 2.0 architecture. It uses 16B total parameters with approximately 1.4B activated per token and is reported to deliver comprehensive reasoning performance comparable to sub-10B dense LLMs. The model shows strong results on logical reasoning, code generation, and mathematical tasks, supports 128K context windows, and reports generation speeds of 300+ tokens per second.

Input type

Output Type

Input0.07/M tokens

Output0.7/M tokens

Context128.00K

Max Output32.00K

Available on 1 provider

Apr 30, 2026 6:00 PM-

xAI: Grok Code Fast 1

x-ai/grok-code-fast-1

API Request Chat

1.18Btokens

Grok Code Fast 1 is a speedy and economical reasoning model that excels at agentic coding. With reasoning traces visible in the response, developers can steer Grok Code for high-quality work flows.

Input type

Output Type

Input0.2/M tokens

Output1.5/M tokens

Context256.00K

Max Output10.00K

Available on 1 provider

Apr 30, 2026 6:00 PM-

DeepSeek: DeepSeek V3.1

deepseek/deepseek-chat-v3.1

API Request Chat

1.49Btokens

DeepSeek-V3.1 is a large hybrid reasoning model (671B parameters, 37B active) that supports both thinking and non-thinking modes via prompt templates. It extends the DeepSeek-V3 base with a two-phase long-context training process, reaching up to 128K tokens, and uses FP8 microscaling for efficient inference.

The model improves tool use, code generation, and reasoning efficiency, achieving performance comparable to DeepSeek-R1 on difficult benchmarks while responding more quickly. It supports structured tool calling, code agents, and search agents, making it suitable for research, coding, and agentic workflows.

It succeeds the DeepSeek V3-0324 model and performs well on a variety of tasks.

Input type

Output Type

Input0.56/M tokens

Output1.68/M tokens

Context128.00K

Max Output65.54K

Available on 2 providers

Apr 30, 2026 6:00 PM-

OpenAI: GPT-5

openai/gpt-5

API Request Chat

2.48Btokens

GPT-5 is OpenAI’s most advanced model, offering major improvements in reasoning, code quality, and user experience. It is optimized for complex tasks that require step-by-step reasoning, instruction following, and accuracy in high-stakes use cases. It supports test-time routing features and advanced prompt understanding, including user-specified intent like "think hard about this." Improvements include reductions in hallucination, sycophancy, and better performance in coding, writing, and health-related tasks.

Input type

Output Type

Input1.25/M tokens

Output10/M tokens

Context400.00K

Max Output128.00K

Available on 2 providers

Apr 30, 2026 6:00 PM-

OpenAI: GPT-5 Chat

openai/gpt-5-chat

API Request Chat

100.41Mtokens

GPT-5 Chat is designed for advanced, natural, multimodal, and context-aware conversations for enterprise applications.

Input type

Output Type

Input1.25/M tokens

Output10/M tokens

Context128.00K

Max Output16.38K

Available on 2 providers

Apr 30, 2026 6:00 PM-

OpenAI: GPT-5 Mini

openai/gpt-5-mini

API Request Chat

2.35Btokens

GPT-5 Mini is a compact version of GPT-5, designed to handle lighter-weight reasoning tasks. It provides the same instruction-following and safety-tuning benefits as GPT-5, but with reduced latency and cost. GPT-5 Mini is the successor to OpenAI's o4-mini model.

Input type

Output Type

Input0.25/M tokens

Output2/M tokens

Context400.00K

Max Output128.00K

Available on 2 providers

Apr 30, 2026 6:00 PM-

OpenAI: GPT-5 Nano

openai/gpt-5-nano

API Request Chat

816.62Mtokens

GPT-5-Nano is the smallest and fastest variant in the GPT-5 system, optimized for developer tools, rapid interactions, and ultra-low latency environments. While limited in reasoning depth compared to its larger counterparts, it retains key instruction-following and safety features. It is the successor to GPT-4.1-nano and offers a lightweight option for cost-sensitive or real-time applications.

Input type

Output Type

Input0.05/M tokens

Output0.4/M tokens

Context400.00K

Max Output128.00K

Available on 2 providers

Apr 30, 2026 6:00 PM-

Anthropic: Claude Opus 4.1

anthropic/claude-opus-4.1

API Request Chat

283.17Mtokens

Claude Opus 4.1 is an updated version of Anthropic’s flagship model, offering improved performance in coding, reasoning, and agentic tasks. It achieves 74.5% on SWE-bench Verified and shows notable gains in multi-file code refactoring, debugging precision, and detail-oriented reasoning. The model supports extended thinking up to 64K tokens and is optimized for tasks involving research, data analysis, and tool-assisted reasoning.

Input type

Output Type

Input15/M tokens

Output75/M tokens

Context200.00K

Max Output32.00K

Available on 3 providers

Apr 30, 2026 6:00 PM-

StepFun: Step-3

stepfun/step-3

API Request Chat

37.12Mtokens

Step-3 is a brand-new multimodal reasoning model that can process both image and text inputs and produce text responses. It is capable of deep thinking and can autonomously carry out a reasoning process. That is, before generating the final output, it completes a “thinking” phase (for example, presenting reasoning information via a reasoning field), which improves the accuracy of the final result and the depth of reasoning. When calling the model, developers do not need to preset too many system prompts (sys_prompt), as the model can automatically leverage its built-in deep-thinking capability._

Input type

Output Type

Input0.21-0.57/M tokens

Output0.57-1.42/M tokens

Context65.54K

Max Output65.54K

Available on 1 provider

Apr 30, 2026 6:00 PM-

KlingAI: Kling-v2

klingai/kling-v2

API Request Chat

740.00Ktokens

kling-v2 is an image generation model launched by Kling AI.

Input type

Output Type

Input-

Output0.014/counts

Context-

Max Output-

Available on 1 provider

Apr 30, 2026 6:00 PM-

Z.AI: GLM 4.5

z-ai/glm-4.5

API Request Chat

124.68Mtokens

GLM-4.5 is our latest flagship foundation model, purpose-built for agent-based applications. It leverages a Mixture-of-Experts (MoE) architecture and supports a context length of up to 128k tokens. GLM-4.5 delivers significantly enhanced capabilities in reasoning, code generation, and agent alignment. It supports a hybrid inference mode with two options, a "thinking mode" designed for complex reasoning and tool use, and a "non-thinking mode" optimized for instant responses. Users can control the reasoning behaviour with the reasoning``enabled boolean.

Input type

Output Type

Input0.2911-0.5823/M tokens

Output1.1645-2.3291/M tokens

Context128.00K

Max Output96.00K

Available on 2 providers

Apr 30, 2026 6:00 PM-

Z.AI: GLM 4.5 Air

z-ai/glm-4.5-air

API Request Chat

271.10Mtokens

GLM-4.5-Air is the lightweight variant of our latest flagship model family, also purpose-built for agent-centric applications. Like GLM-4.5, it adopts the Mixture-of-Experts (MoE) architecture but with a more compact parameter size. GLM-4.5-Air also supports hybrid inference modes, offering a "thinking mode" for advanced reasoning and tool use, and a "non-thinking mode" for real-time interaction. Users can control the reasoning behaviour with the reasoning``enabled boolean. Learn more in our docs

Input type

Output Type

Input0.1165-0.1747/M tokens

Output0.2911-1.1645/M tokens

Context128.00K

Max Output96.00K

Available on 3 providers

Apr 30, 2026 6:00 PM-

Qwen: Qwen3-Coder-Plus

qwen/qwen3-coder-plus

API Request Chat

535.18Mtokens

Powered by Qwen3, this is a powerful Coding Agent that excels in tool calling and environment interaction to achieve autonomous programming. It combines outstanding coding proficiency with versatile general-purpose abilities.

Input type

Output Type

Input1-6/M tokens

Output5-60/M tokens

Context1000.00K

Max Output65.54K

Available on 1 provider

Apr 30, 2026 6:00 PM-

Google: Gemini 2.5 Flash Lite

google/gemini-2.5-flash-lite

API Request Chat

24.72Btokens

Gemini 2.5 Flash-Lite is a lightweight reasoning model in the Gemini 2.5 family, optimized for ultra-low latency and cost efficiency. It offers improved throughput, faster token generation, and better performance across common benchmarks compared to earlier Flash models. By default, "thinking" (i.e. multi-pass reasoning) is disabled to prioritize speed, but developers can enable it via the Reasoning API parameter to selectively trade off cost for intelligence.

Input type

Output Type

Input0.1/M tokens

Output0.4/M tokens

Context1.05M

Max Output65.53K

Available on 2 providers

Apr 30, 2026 6:00 PM-

Qwen: Qwen3 235B A22B Instruct 2507

qwen/qwen3-235b-a22b-2507

API Request Chat

354.87Mtokens

Qwen3-235B-A22B-Instruct-2507 is a multilingual, instruction-tuned mixture-of-experts language model based on the Qwen3-235B architecture, with 22B active parameters per forward pass. It is optimized for general-purpose text generation, including instruction following, logical reasoning, math, code, and tool usage. The model supports a native 262K context length and does not implement "thinking mode" ( blocks).

Compared to its base variant, this version delivers significant gains in knowledge coverage, long-context reasoning, coding benchmarks, and alignment with open-ended tasks. It is particularly strong on multilingual understanding, math reasoning (e.g., AIME, HMMT), and alignment evaluations like Arena-Hard and WritingBench.

Input type

Output Type

Input0.28/M tokens

Output1.11/M tokens

Context256.00K

Max Output128.00K

Available on 1 provider

Apr 30, 2026 6:00 PM-

Qwen: Qwen3 235B A22B Thinking 2507

qwen/qwen3-235b-a22b-thinking-2507

API Request Chat

21.68Mtokens

Qwen3-235B-A22B-Thinking-2507 is a high-performance, open-weight Mixture-of-Experts (MoE) language model optimized for complex reasoning tasks. It activates 22B of its 235B parameters per forward pass and natively supports up to 262,144 tokens of context. This "thinking-only" variant enhances structured logical reasoning, mathematics, science, and long-form generation, showing strong benchmark performance across AIME, SuperGPQA, LiveCodeBench, and MMLU-Redux. It enforces a special reasoning mode () and is designed for high-token outputs (up to 81,920 tokens) in challenging domains.

The model is instruction-tuned and excels at step-by-step reasoning, tool use, agentic workflows, and multilingual tasks. This release represents the most capable open-source variant in the Qwen3-235B series, surpassing many closed models in structured reasoning use cases.

Input type

Output Type

Input0.28/M tokens

Output2.78/M tokens

Context256.00K

Max Output128.00K

Available on 1 provider

Apr 30, 2026 6:00 PM-

Qwen: Qwen3-Coder

qwen/qwen3-coder

API Request Chat

36.56Mtokens

Qwen3-Coder-480B-A35B-Instruct is a Mixture-of-Experts (MoE) code generation model developed by the Qwen team. It is optimized for agentic coding tasks such as function calling, tool use, and long-context reasoning over repositories. The model features 480 billion total parameters, with 35 billion active per forward pass (8 out of 160 experts).

Pricing for the Alibaba endpoints varies by context length. Once a request is greater than 128k input tokens, the higher pricing is used.

Input type

Output Type

Input1.25/M tokens

Output5.01/M tokens

Context256.00K

Max Output128.00K

Available on 1 provider

Apr 30, 2026 6:00 PM-

MoonshotAI: Kimi K2 0711

moonshotai/kimi-k2-0711

API Request Chat

15.53Mtokens

Kimi K2 Instruct is a large-scale Mixture-of-Experts (MoE) language model developed by Moonshot AI, featuring 1 trillion total parameters with 32 billion active per forward pass. It is optimized for agentic capabilities, including advanced tool use, reasoning, and code synthesis. Kimi K2 excels across a broad range of benchmarks, particularly in coding (LiveCodeBench, SWE-bench), reasoning (ZebraLogic, GPQA), and tool-use (Tau2, AceBench) tasks. It supports long-context inference up to 128K tokens and is designed with a novel training stack that includes the MuonClip optimizer for stable large-scale MoE training.

Input type

Output Type

Input0.56/M tokens

Output2.23/M tokens

Context128.00K

Max Output32.00K

Available on 2 providers

Apr 30, 2026 6:00 PM-

xAI: Grok 4

x-ai/grok-4

API Request Chat

294.85Mtokens

Grok 4 is xAI's latest reasoning model with a 256k context window. It supports parallel tool calling, structured outputs, and both image and text inputs. Note that reasoning is not exposed, reasoning cannot be disabled, and the reasoning effort cannot be specified. Pricing increases once the total tokens in a given request is greater than 128k tokens. See more details on the xAI docs

Input type

Output Type

Input3-6/M tokens

Output15-30/M tokens

Context256.00K

Max Output256.00K

Available on 1 provider

Apr 30, 2026 6:00 PM-

Google: Gemini 2.5 Flash

google/gemini-2.5-flash

API Request Chat

4.26Btokens

Gemini 2.5 Flash is Google's state-of-the-art workhorse model, specifically designed for advanced reasoning, coding, mathematics, and scientific tasks. It includes built-in "thinking" capabilities, enabling it to provide responses with greater accuracy and nuanced context handling.

Additionally, Gemini 2.5 Flash is configurable through the "max tokens for reasoning" parameter, as described in the documentation.

Input type

Output Type

Input0.3/M tokens

Output2.5/M tokens

Context1.05M

Max Output65.53K

Available on 2 providers

Apr 30, 2026 6:00 PM-

Google: Gemini 2.5 Pro

google/gemini-2.5-pro

API Request Chat

3.25Btokens

Gemini 2.5 Pro is Google’s state-of-the-art AI model designed for advanced reasoning, coding, mathematics, and scientific tasks. It employs “thinking” capabilities, enabling it to reason through responses with enhanced accuracy and nuanced context handling. Gemini 2.5 Pro achieves top-tier performance on multiple benchmarks, including first-place positioning on the LMArena leaderboard, reflecting superior human-preference alignment and complex problem-solving abilities.

Input type

Output Type

Input1.25-2.5/M tokens

Output10-15/M tokens

Context1.05M

Max Output65.53K

Available on 2 providers

Apr 30, 2026 6:00 PM-

DeepSeek: DeepSeek R1 0528

deepseek/deepseek-r1-0528

API Request Chat

908.74Mtokens

May 28th update to the original DeepSeek R1 Performance on par with OpenAI o1, but open-sourced and with fully open reasoning tokens. It's 671B parameters in size, with 37B active in an inference pass. Fully open-source model.

Input type

Output Type

Input0.56/M tokens

Output2.23/M tokens

Context64.00K

Max Output64.00K

Available on 2 providers

Apr 30, 2026 6:00 PM-

Anthropic: Claude Opus 4

anthropic/claude-opus-4

API Request Chat

219.16Mtokens

Claude Opus 4 is benchmarked as the world’s best coding model, at time of release, bringing sustained performance on complex, long-running tasks and agent workflows. It sets new benchmarks in software engineering, achieving leading results on SWE-bench (72.5%) and Terminal-bench (43.2%). Opus 4 supports extended, agentic workflows, handling thousands of task steps continuously for hours without degradation.

Input type

Output Type

Input15/M tokens

Output75/M tokens

Context200.00K

Max Output32.00K

Available on 3 providers

Apr 30, 2026 6:00 PM-

Anthropic: Claude Sonnet 4

anthropic/claude-sonnet-4

API Request Chat

8.49Btokens

Claude Sonnet 4 significantly enhances the capabilities of its predecessor, Sonnet 3.7, excelling in both coding and reasoning tasks with improved precision and controllability. Achieving state-of-the-art performance on SWE-bench (72.7%), Sonnet 4 balances capability and computational efficiency, making it suitable for a broad range of applications from routine coding tasks to complex software development projects. Key enhancements include improved autonomous codebase navigation, reduced error rates in agent-driven workflows, and increased reliability in following intricate instructions. Sonnet 4 is optimized for practical everyday use, providing advanced reasoning capabilities while maintaining efficiency and responsiveness in diverse internal and external scenarios.

Input type

Output Type

Input3-6/M tokens

Output15-22.5/M tokens

Context1000.00K

Max Output64.00K

Available on 3 providers

Apr 30, 2026 6:00 PM-

Google: Gemma 3 12B

google/gemma-3-12b-it

API Request Chat

491.80Mtokens

Gemma 3 introduces multimodality, supporting vision-language input and text outputs. It handles context windows up to 128k tokens, understands over 140 languages, and offers improved math, reasoning, and chat capabilities, including structured outputs and function calling. Gemma 3 12B is the second largest in the family of Gemma 3 models after Gemma 3 27B [blocked]

Input type

Output Type

Input0.024/M tokens

Output0.096/M tokens

Context128.00K

Max Output128.00K

Available on 1 provider

Apr 30, 2026 6:00 PM-

Qwen: Qwen3-14B

qwen/qwen3-14b

API Request Chat

94.63Mtokens

Qwen3-14B is a dense 14.8B parameter causal language model from the Qwen3 series, designed for both complex reasoning and efficient dialogue. It supports seamless switching between a "thinking" mode for tasks like math, programming, and logical inference, and a "non-thinking" mode for general-purpose conversation. The model is fine-tuned for instruction-following, agent tool use, creative writing, and multilingual tasks across 100+ languages and dialects. It natively handles 32K token contexts and can extend to 131K tokens using YaRN-based scaling.

Input type

Output Type

Input0.14/M tokens

Output1.4/M tokens

Context32.00K

Max Output32.00K

Available on 1 provider

Apr 30, 2026 6:00 PM-

OpenAI: o4 Mini

openai/o4-mini

API Request Chat

131.78Mtokens

OpenAI o4-mini is a compact reasoning model in the o-series, optimized for fast, cost-efficient performance while retaining strong multimodal and agentic capabilities. It supports tool use and demonstrates competitive reasoning and coding performance across benchmarks like AIME (99.5% with Python) and SWE-bench, outperforming its predecessor o3-mini and even approaching o3 in some domains.

Despite its smaller size, o4-mini exhibits high accuracy in STEM tasks, visual problem solving (e.g., MathVista, MMMU), and code editing. It is especially well-suited for high-throughput scenarios where latency or cost is critical. Thanks to its efficient architecture and refined reinforcement learning training, o4-mini can chain tools, generate structured outputs, and solve multi-step tasks with minimal delay—often in under a minute.

Input type

Output Type

Input1.1/M tokens

Output4.4/M tokens

Context200.00K

Max Output100.00K

Available on 2 providers

Apr 30, 2026 6:00 PM-

OpenAI: GPT-4.1

openai/gpt-4.1

API Request Chat

3.33Btokens

GPT-4.1 is a flagship large language model optimized for advanced instruction following, real-world software engineering, and long-context reasoning. It supports a 1 million token context window and outperforms GPT-4o and GPT-4.5 across coding (54.6% SWE-bench Verified), instruction compliance (87.4% IFEval), and multimodal understanding benchmarks. It is tuned for precise code diffs, agent reliability, and high recall in large document contexts, making it ideal for agents, IDE tooling, and enterprise knowledge retrieval.

Input type

Output Type

Input2/M tokens

Output8/M tokens

Context1.05M

Max Output32.77K

Available on 2 providers

Apr 30, 2026 6:00 PM-

OpenAI: GPT-4.1 Mini

openai/gpt-4.1-mini

API Request Chat

3.57Btokens

GPT-4.1 Mini is a mid-sized model delivering performance competitive with GPT-4o at substantially lower latency and cost. It retains a 1 million token context window and scores 45.1% on hard instruction evals, 35.8% on MultiChallenge, and 84.1% on IFEval. Mini also shows strong coding ability (e.g., 31.6% on Aider’s polyglot diff benchmark) and vision understanding, making it suitable for interactive applications with tight performance constraints.

Input type

Output Type

Input0.4/M tokens

Output1.6/M tokens

Context1.05M

Max Output32.77K

Available on 2 providers

Apr 30, 2026 6:00 PM-

OpenAI: GPT-4.1 Nano

openai/gpt-4.1-nano

API Request Chat

1.89Btokens

For tasks that demand low latency, GPT‑4.1 nano is the fastest and cheapest model in the GPT-4.1 series. It delivers exceptional performance at a small size with its 1 million token context window, and scores 80.1% on MMLU, 50.3% on GPQA, and 9.8% on Aider polyglot coding – even higher than GPT‑4o mini. It’s ideal for tasks like classification or autocompletion.

Input type

Output Type

Input0.1/M tokens

Output0.4/M tokens

Context1.05M

Max Output32.77K

Available on 2 providers

Apr 30, 2026 6:00 PM-

Meta: Llama 4 Scout Instruct

meta/llama-4-scout-17b-16e-instruct

API Request Chat

34.60Mtokens

Llama 4 Scout 17B Instruct (16E) is a mixture-of-experts (MoE) language model developed by Meta, activating 17 billion parameters out of a total of 109B. It supports native multimodal input (text and image) and multilingual output (text and code) across 12 supported languages. Designed for assistant-style interaction and visual reasoning, Scout uses 16 experts per forward pass and features a context length of 10 million tokens, with a training corpus of ~40 trillion tokens. Built for high efficiency and local or commercial deployment, Llama 4 Scout incorporates early fusion for seamless modality integration. It is instruction-tuned for use in multilingual chat, captioning, and image understanding tasks. Released under the Llama 4 Community License, it was last trained on data up to August 2024 and launched publicly on April 5, 2025.

Input type

Output Type

Input0.08/M tokens

Output0.4/M tokens

Context-

Max Output-

Available on 1 provider

Apr 30, 2026 6:00 PM-

Anthropic: Claude 3.7 Sonnet

anthropic/claude-3.7-sonnet

API Request Chat

3.70Btokens

Sunset: 2026/04/28

Claude 3.7 Sonnet is an advanced large language model with improved reasoning, coding, and problem-solving capabilities. It introduces a hybrid reasoning approach, allowing users to choose between rapid responses and extended, step-by-step processing for complex tasks. The model demonstrates notable improvements in coding, particularly in front-end development and full-stack updates, and excels in agentic workflows, where it can autonomously navigate multi-step processes.

Claude 3.7 Sonnet maintains performance parity with its predecessor in standard mode while offering an extended reasoning mode for enhanced accuracy in math, coding, and instruction-following tasks.

Input type

Output Type

Input3/M tokens

Output15/M tokens

Context200.00K

Max Output64.00K

Available on 1 provider

Apr 30, 2026 6:00 PM-

Meta: Llama 3.3 70B Instruct

meta/llama-3.3-70b-instruct

API Request Chat

245.77Mtokens

The Meta Llama 3.3 multilingual large language model (LLM) is a pretrained and instruction tuned generative model in 70B (text in/text out). The Llama 3.3 instruction tuned text only model is optimized for multilingual dialogue use cases and outperforms many of the available open source and closed chat models on common industry benchmarks.

Supported languages: English, German, French, Italian, Portuguese, Hindi, Spanish, and Thai.

Input type

Output Type

Input0.6/M tokens

Output1.2/M tokens

Context-

Max Output-

Available on 1 provider

Apr 30, 2026 6:00 PM-

Anthropic: Claude 3.5 Haiku

anthropic/claude-3.5-haiku

API Request Chat

278.03Mtokens

Claude 3.5 Haiku features offers enhanced capabilities in speed, coding accuracy, and tool use. Engineered to excel in real-time applications, it delivers quick response times that are essential for dynamic tasks such as chat interactions and immediate coding suggestions.

This makes it highly suitable for environments that demand both speed and precision, such as software development, customer service bots, and data management systems.

Input type

Output Type

Input0.8/M tokens

Output4/M tokens

Context200.00K

Max Output8.19K

Available on 2 providers

Apr 30, 2026 6:00 PM-

OpenAI: GPT-4o-mini

openai/gpt-4o-mini

API Request Chat

1.57Btokens

GPT-4o mini is OpenAI's newest model after GPT-4 Omni, supporting both text and image inputs with text outputs.

As their most advanced small model, it is many multiples more affordable than other recent frontier models, and more than 60% cheaper than GPT-3.5 Turbo. It maintains SOTA intelligence, while being significantly more cost-effective.

GPT-4o mini achieves an 82% score on MMLU and presently ranks higher than GPT-4 on chat preferences common leaderboards.

Input type

Output Type

Input0.15/M tokens

Output0.6/M tokens

Context128.00K

Max Output16.38K

Available on 2 providers

Apr 30, 2026 6:00 PM-

OpenAI: GPT-4o

openai/gpt-4o

API Request Chat

577.81Mtokens

GPT-4o ("o" for "omni") is OpenAI's latest AI model, supporting both text and image inputs with text outputs. It maintains the intelligence level of GPT-4 Turbowhile being twice as fast and 50% more cost-effective. GPT-4o also offers improved performance in processing non-English languages and enhanced visual capabilities.

For benchmarking against other models, it was briefly called "im-also-a-good-gpt2-chatbot"

Input type

Output Type

Input2.5/M tokens

Output10/M tokens

Context128.00K

Max Output16.38K

Available on 2 providers

Apr 30, 2026 6:00 PM-

StripeM-Inner