ZenMux Benchmark

Model Updates

Offers

Notifications

  • A billing anomaly was identified. Affected users will receive compensation. Read more.

Model Cost Performance

Cost-Performance Curve

Performance Table

Model Provider Update at Cost usage
openai/gpt-5 openai Sep 22, 2025 12:29 PM $646.9 293.81 MTokens

Evaluation Models

Ranking Model Provider Metrics Update at Latency Throughput Cost usage LogsFile
1 openai/gpt-5 openai openai/gpt-5
25.43
±1.84
Calib Err: 50.2
Sep 22, 2025 12:29 PM 16583.37 ms 50.13 tokens/s $178.65 18.46 MTokens View
4 google/gemini-2.5-pro google-vertex google/gemini-2.5-pro
20.02
±1.69
Calib Err: 72.96
Sep 22, 2025 12:29 PM 5951.89 ms 113.13 tokens/s $226.46 23.24 MTokens View

About ZenMux Benchmark

ZenMux-Benchmark is a dynamic AI model evaluation leaderboard maintained by the ZenMux. We regularly conduct systematic evaluations of all AI models across every provider channel available on the ZenMux platform to ensure access to the latest and most accurate performance data.

Transparency

All test code, procedures, and results are publicly available on GitHub. Project repository: ZenMux-Benchmark.

Dataset

We use Scale AI's publicly released dataset, Humanity's Last Exam (Text Only), as our primary evaluation benchmark. For more about the dataset, see: Humanity's Last Exam (Text Only).

Feedback

ZenMux-Benchmark aims to build a dynamically updated real-time leaderboard that enables tracking of the latest performances of AI models. We welcome community feedback and suggestions.