ZenMux Benchmark
Model Updates
google/gemini-3.1-pro-preview, anthropic/claude-sonnet-4.6, google/gemini-3-pro-image-preview are Now Live.
Claude Sonnet 4 will be retired by both providers: AWS Bedrock and Anthropic on June 14, 2026. We recommend upgrading to Claude Sonnet 4.6.
Claude Opus 4 will be retired by both providers: AWS Bedrock on May 31, 2026 and Anthropic on June 14, 2026. We recommend upgrading to Claude Opus 4.7.
Qwen3 Max Thinking Preview will be deprecated and removed on April 24, 2026. We recommend upgrading to Qwen3-Max-Thinking.
Caution: On Azure, gpt-5-chat will be deprecated and removed on May 15, 2026.
Claude 3.7 Sonnet will be retired by both providers: AWS Bedrock on April 28, 2026, and Google Cloud Vertex AI on May 11, 2026. We recommend upgrading to Claude Sonnet 4.6.
Offers
- Get 5% off service fee on top-up, and exclusive gifts for referring friends.
Notifications
- A billing anomaly was identified. Affected users will receive compensation. Read more.
Model Cost Performance
Cost-Performance Curve
Performance Table
| Model | Provider | Update at | Cost | usage |
|---|---|---|---|---|
| openai/gpt-5 | openai | Sep 22, 2025 12:29 PM | $646.9 | 293.81 MTokens |
Evaluation Models
| Ranking | Model | Provider | Metrics | Update at | Latency | Throughput | Cost | usage | LogsFile |
|---|---|---|---|---|---|---|---|---|---|
| 1 | openai/gpt-5 | openai | openai/gpt-5 25.43 ±1.84 Calib Err: 50.2 |
Sep 22, 2025 12:29 PM | 16583.37 ms | 50.13 tokens/s | $178.65 | 18.46 MTokens | View |
| 4 | google/gemini-2.5-pro | google-vertex | google/gemini-2.5-pro 20.02 ±1.69 Calib Err: 72.96 |
Sep 22, 2025 12:29 PM | 5951.89 ms | 113.13 tokens/s | $226.46 | 23.24 MTokens | View |
About ZenMux Benchmark
ZenMux-Benchmark is a dynamic AI model evaluation leaderboard maintained by the ZenMux. We regularly conduct systematic evaluations of all AI models across every provider channel available on the ZenMux platform to ensure access to the latest and most accurate performance data.
Transparency
All test code, procedures, and results are publicly available on GitHub. Project repository: ZenMux-Benchmark.
Dataset
We use Scale AI's publicly released dataset, Humanity's Last Exam (Text Only), as our primary evaluation benchmark. For more about the dataset, see: Humanity's Last Exam (Text Only).
Feedback
ZenMux-Benchmark aims to build a dynamically updated real-time leaderboard that enables tracking of the latest performances of AI models. We welcome community feedback and suggestions.