Supported Models
Current list of AI models available through SoxAI — Claude 4.x, GPT-5.x, Gemini 2.5/3.x, GLM 5.x, MiniMax M2.x, and Gemma. Use the model ID exactly as shown.
Supported Models
All models listed below are accessible via the SoxAI gateway at https://api.soxai.io/v1. Use the model ID exactly as shown.
This page is a snapshot of the catalog. The live catalog — including any model added between doc updates — is always available at /api/public/models (JSON) or console.soxai.io/models (browseable). If you see a model mentioned in our blog or release notes that isn't here yet, check the live catalog first.
Pricing note
Prices below are list prices in USD per 1 million tokens. During the launch promo, SoxAI bills you 15% below list — the discount applies at settlement, not at quote time, and shows up as a line item on your invoice. See /vs-openrouter for the pricing comparison.
Anthropic — Claude 4.x
| Model ID | Context | Input $/1M | Output $/1M | Notes |
|---|---|---|---|---|
claude-opus-4-7 | 1M | $5.00 | $25.00 | Flagship — best Anthropic model |
claude-opus-4-6 | 1M | $5.00 | $25.00 | Previous flagship |
claude-sonnet-4-6 | 1M | $3.00 | $15.00 | Balanced — most popular |
claude-haiku-4-5 | 200k | $1.00 | $5.00 | Fast + cheap, latest Haiku |
claude-haiku-4-5-20251001 | 200k | $1.00 | $5.00 | Dated pin of Haiku 4.5 |
All Anthropic models support the Messages API (prompt caching, tool use, extended thinking). Set ANTHROPIC_BASE_URL to SoxAI and your existing Anthropic SDK code works unchanged — see the Claude Code guide for the full CLI setup.
OpenAI — GPT-5.x
| Model ID | Context | Input $/1M | Output $/1M | Notes |
|---|---|---|---|---|
gpt-5.4 | 1.05M | $2.50 | $15.00 | Flagship chat model |
gpt-5.4-pro | 1.05M | $30.00 | $180.00 | Top-tier reasoning |
gpt-5.4-mini | 400k | $0.75 | $4.50 | Cost-efficient 5.4 |
gpt-5.4-nano | 400k | $0.20 | $1.25 | Smallest, fastest |
gpt-5-mini | 400k | $0.25 | $2.00 | Legacy 5.0-era mini |
gpt-5.3-codex | 400k | $1.75 | $14.00 | Codex CLI default |
gpt-4o | 128k | $2.50 | $10.00 | GPT-4 omni (legacy) |
gpt-4o-mini | 128k | $0.15 | $0.60 | GPT-4 omni mini (legacy) |
Image generation
| Model ID | Notes |
|---|---|
gpt-image-1.5 | Latest image model |
gpt-image-1 | Previous generation |
gpt-image-1-mini | Smaller, cheaper |
Image generation is priced per image, not per token. See /docs/api-reference/images for the request shape.
Google — Gemini 2.5 / 3.x
Gemini 3 (latest)
| Model ID | Context | Input $/1M | Output $/1M | Notes |
|---|---|---|---|---|
gemini-3.1-pro-preview | 1M | $2.00 | $12.00 | Flagship Gemini |
gemini-3.1-pro-preview-customtools | 1M | $2.00 | $12.00 | With custom tool support |
gemini-3-pro-preview | 1M | $2.00 | $12.00 | Previous Pro preview |
gemini-3.1-flash-lite-preview | 1M | $0.25 | $1.50 | Flash Lite 3.1 |
gemini-3-flash-preview | 1M | $0.50 | $3.00 | Flash 3.0 preview |
gemini-3.1-flash-image-preview | 128k | $0.25 | $60.00 | Image output variant |
Gemini 2.5
| Model ID | Context | Input $/1M | Output $/1M | Notes |
|---|---|---|---|---|
gemini-2.5-pro | 1M | $1.25 | $10.00 | Stable Pro — most compatible |
gemini-2.5-flash | 1M | $0.30 | $2.50 | Stable Flash |
gemini-2.5-flash-lite | 1M | $0.10 | $0.40 | Cheapest 2.5 |
gemini-flash-latest | 1M | $0.30 | $2.50 | Auto-updating Flash alias |
gemini-flash-lite-latest | 1M | $0.10 | $0.40 | Auto-updating Flash Lite alias |
Gemini 2.0 and 1.5 (legacy, still supported)
| Model ID | Context | Input $/1M | Output $/1M |
|---|---|---|---|
gemini-2.0-flash | 1M | $0.10 | $0.40 |
gemini-2.0-flash-lite | 1M | $0.075 | $0.30 |
gemini-1.5-pro | 1M | $1.25 | $5.00 |
gemini-1.5-flash | 1M | $0.075 | $0.30 |
gemini-1.5-flash-8b | 1M | $0.0375 | $0.15 |
Several dated previews (gemini-2.5-flash-preview-*, gemini-2.5-pro-preview-*) are also pinned for reproducibility — see the live catalog for the full list.
Multimodal variants
| Model ID | Context | Input $/1M | Output $/1M | Notes |
|---|---|---|---|---|
gemini-2.5-flash-image | 32k | $0.30 | $30.00 | Image generation |
gemini-2.5-flash-preview-tts | 8k | $0.50 | $10.00 | Text-to-speech |
gemini-2.5-pro-preview-tts | 8k | $1.00 | $20.00 | TTS, higher fidelity |
gemini-live-2.5-flash | 128k | $0.50 | $2.00 | Live streaming conversation |
gemini-live-2.5-flash-preview-native-audio | 128k | $0.50 | $2.00 | Live with native audio |
gemini-embedding-001 | 2k | $0.15 | — | Embeddings |
GLM (Zhipu) — 5.x
| Model ID | Context | Input $/1M | Output $/1M | Notes |
|---|---|---|---|---|
glm-5.1 | 200k | $1.40 | $4.40 | Latest GLM |
glm-5 | 200k | $1.00 | $3.20 | Previous GLM |
MiniMax — M2.x
| Model ID | Context | Input $/1M | Output $/1M | Notes |
|---|---|---|---|---|
MiniMax-M2.7 | 200k | $0.30 | $1.20 | Latest M2 |
MiniMax-M2.7-highspeed | 200k | $0.60 | $2.40 | Priority throughput |
MiniMax-M2.5 | 200k | $0.30 | $1.20 | Stable |
MiniMax-M2.5-highspeed | 200k | $0.60 | $2.40 | Priority throughput |
MiniMax-M2.1 | 200k | $0.30 | $1.20 | Earlier release |
MiniMax-M2 | 196k | $0.30 | $1.20 | Original M2 |
Gemma (open weights, free tier)
Served from SoxAI's internal compute — no per-token charge during the early-access window.
| Model ID | Context | Notes |
|---|---|---|
gemma-4-31b-it | 256k | Largest open Gemma |
gemma-4-26b-it | 256k | Mid-tier 4.x |
gemma-3-27b-it | 128k | Stable Gemma 3 |
gemma-3-12b-it | 32k | Balanced |
gemma-3-4b-it | 32k | Fast |
gemma-3n-e4b-it | 8k | On-device class (4B effective) |
gemma-3n-e2b-it | 8k | On-device class (2B effective) |
What about DeepSeek / Mistral / Cohere?
Earlier catalog versions listed DeepSeek, Mistral, and Cohere as native providers. Those are currently not in the public catalog — if your workload depends on them, reach out to [email protected] to discuss adding them as a BYOK channel on your account.
Checking the live catalog
Use the List Models endpoint to see exactly which models your token can access right now:
curl https://api.soxai.io/v1/models \
-H "Authorization: Bearer $SOXAI_API_KEY"Or browse the public catalog without signing in:
curl https://console.soxai.io/api/public/modelsThe public endpoint returns the same list this page is generated from, but updated in real time — always the source of truth.