Google: Gemini 3.1 Pro Preview

google/gemini-3.1-pro-preview

Chat

Gemini 3.1 Pro is the next generation in the Gemini series of models, a suite of highly-capable, natively multimodal, reasoning models. Gemini 3 Pro is now Google’s most advanced model for complex tasks, and can comprehend vast datasets, challenging problems from different information sources, including text, audio, images, video, and entire code repositories

by Google
Input type: Output type
Publish time: 2026-02-19

Recent activity on Gemini 3.1 Pro Preview

Tokens processed per day

Throughput

(tokens/s)

Providers Min (tokens/s) Max (tokens/s) Avg (tokens/s)
Google Vertex 7.89 83.64 42.67
SkyRouter 25.33 25.33 25.33

First Token Latency

(ms)

Providers Min (ms) Max (ms) Avg (ms)
Google Vertex 3626 10021 7078.97
SkyRouter 6437 6437 6437.00

Uptime stats for Gemini 3.1 Pro Preview

Uptime stats for Gemini 3.1 Pro Preview across all providers

Providers Uptime (Percent)
Google Vertex 100.00
SkyRouter -

Price

Tiered Pricing

Input Size Price
0 ≤ Input < 200k $2/ M tokens

Model limitation

Context: 1.05M
Max output: 65.53K
Supported Parameters: max_completion_tokens, temperature, top_p, frequency_penalty, presence_penalty, seed, logit_bias, logprobs, top_logprobs, response_format, stop, tools, tool_choice, parallel_tool_calls

Sample code and API for Gemini 3.1 Pro Preview

ZenMux normalizes requests and responses across providers for you.

OpenAI: Python-SDK

from openai import OpenAI

client = OpenAI(
  base_url="https://zenmux.ai/api/v1",
  api_key="<ZENMUX_API_KEY>",
)

# Chat Completion
completion = client.chat.completions.create(
  model="google/gemini-3.1-pro-preview",
  messages=[
    {
      "role": "user",
      "content": "What is the meaning of life?"
    }
  ]
)
print(completion.choices[0].message.content)

# Responses API
responses = client.responses.create(
  model="google/gemini-3.1-pro-preview",
  input="What is the meaning of life?"
)
print(responses)