# ZenMux Benchmark

## Model Updates

- **[google/gemini-3.1-pro-preview](/content/google/gemini-3.1-pro-preview/index.html)**, **[anthropic/claude-sonnet-4.6](/content/anthropic/claude-sonnet-4.6)**, **[google/gemini-3-pro-image-preview](/content/google/gemini-3-pro-image-preview/index.html)** are Now Live.

- **[Claude Sonnet 4](/content/anthropic/claude-sonnet-4/index.html)** will be retired by both providers: AWS Bedrock and Anthropic on June 14, 2026. We recommend upgrading to **[Claude Sonnet 4.6](/content/anthropic/claude-sonnet-4.6)**.

- **[Claude Opus 4](/content/anthropic/claude-opus-4/index.html)** will be retired by both providers: AWS Bedrock on May 31, 2026 and Anthropic on June 14, 2026. We recommend upgrading to **[Claude Opus 4.7](/content/anthropic/claude-opus-4.7)**.

- **[Qwen3 Max Thinking Preview](/content/qwen/qwen3-max-preview/index.html)** will be deprecated and removed on April 24, 2026. We recommend upgrading to **[Qwen3-Max-Thinking](/content/qwen/qwen3-max/index.html)**.

- Caution: On Azure, **[gpt-5-chat](/content/openai/gpt-5-chat/index.html)** will be deprecated and removed on May 15, 2026.

- **[Claude 3.7 Sonnet](/content/anthropic/claude-3.7-sonnet/index.html)** will be retired by both providers: AWS Bedrock on April 28, 2026, and Google Cloud Vertex AI on May 11, 2026. We recommend upgrading to **[Claude Sonnet 4.6](/content/anthropic/claude-sonnet-4.6)**.

## Offers

- Get **[5% off service fee](/content/platform/pay-as-you-go?from=topflag/index.html)** on top-up, and exclusive gifts for **[referring friends](/content/benchmark#/index.html)**.

## Notifications

- A billing anomaly was identified. Affected users will receive compensation. **[Read more](/content/blog/billing-anomaly-notice-and-compensation/index.html)**.

## Model Cost Performance

### Cost-Performance Curve

### Performance Table

| Model | Provider | Update at | Cost | usage |
| --- | --- | --- | --- | --- |
| **[openai/gpt-5](/content/openai/gpt-5/index.html)** | **[openai](/content/provider/openai/index.html)** | Sep 22, 2025 12:29 PM | $646.9 | 293.81 MTokens |

### Evaluation Models

| Ranking | Model | Provider | Metrics | Update at | Latency | Throughput | Cost | usage | LogsFile |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
| 1 | **[openai/gpt-5](/content/openai/gpt-5/index.html)** | **[openai](/content/provider/openai/index.html)** | openai/gpt-5<br>25.43<br>±1.84<br>Calib Err: 50.2 | Sep 22, 2025 12:29 PM | 16583.37 ms | 50.13 tokens/s | $178.65 | 18.46 MTokens | View |
| 4 | **[google/gemini-2.5-pro](/content/google/gemini-2.5-pro/index.html)** | **[google-vertex](/content/provider/google-vertex/index.html)** | google/gemini-2.5-pro<br>20.02<br>±1.69<br>Calib Err: 72.96 | Sep 22, 2025 12:29 PM | 5951.89 ms | 113.13 tokens/s | $226.46 | 23.24 MTokens | View |

## About ZenMux Benchmark

ZenMux-Benchmark is a dynamic AI model evaluation leaderboard maintained by the ZenMux. We regularly conduct systematic evaluations of all AI models across every provider channel available on the ZenMux platform to ensure access to the latest and most accurate performance data.

### Transparency

All test code, procedures, and results are publicly available on GitHub. Project repository: **[ZenMux-Benchmark](https://github.com/ZenMux/zenmux-benchmark)**.

### Dataset

We use Scale AI's publicly released dataset, **Humanity's Last Exam (Text Only)**, as our primary evaluation benchmark. For more about the dataset, see: **[Humanity's Last Exam (Text Only)](https://scale.com/leaderboard/humanitys_last_exam_text_only)**.

### Feedback

ZenMux-Benchmark aims to build a dynamically updated real-time leaderboard that enables tracking of the latest performances of AI models. We welcome community feedback and suggestions.
