SoxAIDocs
Models

Supported Models

Current list of AI models available through SoxAI — Claude 4.x, GPT-5.x, Gemini 2.5/3.x, GLM 5.x, MiniMax M2.x, and Gemma. Use the model ID exactly as shown.

Supported Models

All models listed below are accessible via the SoxAI gateway at https://api.soxai.io/v1. Use the model ID exactly as shown.

This page is a snapshot of the catalog. The live catalog — including any model added between doc updates — is always available at /api/public/models (JSON) or console.soxai.io/models (browseable). If you see a model mentioned in our blog or release notes that isn't here yet, check the live catalog first.

Pricing note

Prices below are list prices in USD per 1 million tokens. During the launch promo, SoxAI bills you 15% below list — the discount applies at settlement, not at quote time, and shows up as a line item on your invoice. See /vs-openrouter for the pricing comparison.


Anthropic — Claude 4.x

Model IDContextInput $/1MOutput $/1MNotes
claude-opus-4-71M$5.00$25.00Flagship — best Anthropic model
claude-opus-4-61M$5.00$25.00Previous flagship
claude-sonnet-4-61M$3.00$15.00Balanced — most popular
claude-haiku-4-5200k$1.00$5.00Fast + cheap, latest Haiku
claude-haiku-4-5-20251001200k$1.00$5.00Dated pin of Haiku 4.5

All Anthropic models support the Messages API (prompt caching, tool use, extended thinking). Set ANTHROPIC_BASE_URL to SoxAI and your existing Anthropic SDK code works unchanged — see the Claude Code guide for the full CLI setup.


OpenAI — GPT-5.x

Model IDContextInput $/1MOutput $/1MNotes
gpt-5.41.05M$2.50$15.00Flagship chat model
gpt-5.4-pro1.05M$30.00$180.00Top-tier reasoning
gpt-5.4-mini400k$0.75$4.50Cost-efficient 5.4
gpt-5.4-nano400k$0.20$1.25Smallest, fastest
gpt-5-mini400k$0.25$2.00Legacy 5.0-era mini
gpt-5.3-codex400k$1.75$14.00Codex CLI default
gpt-4o128k$2.50$10.00GPT-4 omni (legacy)
gpt-4o-mini128k$0.15$0.60GPT-4 omni mini (legacy)

Image generation

Model IDNotes
gpt-image-1.5Latest image model
gpt-image-1Previous generation
gpt-image-1-miniSmaller, cheaper

Image generation is priced per image, not per token. See /docs/api-reference/images for the request shape.


Google — Gemini 2.5 / 3.x

Gemini 3 (latest)

Model IDContextInput $/1MOutput $/1MNotes
gemini-3.1-pro-preview1M$2.00$12.00Flagship Gemini
gemini-3.1-pro-preview-customtools1M$2.00$12.00With custom tool support
gemini-3-pro-preview1M$2.00$12.00Previous Pro preview
gemini-3.1-flash-lite-preview1M$0.25$1.50Flash Lite 3.1
gemini-3-flash-preview1M$0.50$3.00Flash 3.0 preview
gemini-3.1-flash-image-preview128k$0.25$60.00Image output variant

Gemini 2.5

Model IDContextInput $/1MOutput $/1MNotes
gemini-2.5-pro1M$1.25$10.00Stable Pro — most compatible
gemini-2.5-flash1M$0.30$2.50Stable Flash
gemini-2.5-flash-lite1M$0.10$0.40Cheapest 2.5
gemini-flash-latest1M$0.30$2.50Auto-updating Flash alias
gemini-flash-lite-latest1M$0.10$0.40Auto-updating Flash Lite alias

Gemini 2.0 and 1.5 (legacy, still supported)

Model IDContextInput $/1MOutput $/1M
gemini-2.0-flash1M$0.10$0.40
gemini-2.0-flash-lite1M$0.075$0.30
gemini-1.5-pro1M$1.25$5.00
gemini-1.5-flash1M$0.075$0.30
gemini-1.5-flash-8b1M$0.0375$0.15

Several dated previews (gemini-2.5-flash-preview-*, gemini-2.5-pro-preview-*) are also pinned for reproducibility — see the live catalog for the full list.

Multimodal variants

Model IDContextInput $/1MOutput $/1MNotes
gemini-2.5-flash-image32k$0.30$30.00Image generation
gemini-2.5-flash-preview-tts8k$0.50$10.00Text-to-speech
gemini-2.5-pro-preview-tts8k$1.00$20.00TTS, higher fidelity
gemini-live-2.5-flash128k$0.50$2.00Live streaming conversation
gemini-live-2.5-flash-preview-native-audio128k$0.50$2.00Live with native audio
gemini-embedding-0012k$0.15—Embeddings

GLM (Zhipu) — 5.x

Model IDContextInput $/1MOutput $/1MNotes
glm-5.1200k$1.40$4.40Latest GLM
glm-5200k$1.00$3.20Previous GLM

MiniMax — M2.x

Model IDContextInput $/1MOutput $/1MNotes
MiniMax-M2.7200k$0.30$1.20Latest M2
MiniMax-M2.7-highspeed200k$0.60$2.40Priority throughput
MiniMax-M2.5200k$0.30$1.20Stable
MiniMax-M2.5-highspeed200k$0.60$2.40Priority throughput
MiniMax-M2.1200k$0.30$1.20Earlier release
MiniMax-M2196k$0.30$1.20Original M2

Gemma (open weights, free tier)

Served from SoxAI's internal compute — no per-token charge during the early-access window.

Model IDContextNotes
gemma-4-31b-it256kLargest open Gemma
gemma-4-26b-it256kMid-tier 4.x
gemma-3-27b-it128kStable Gemma 3
gemma-3-12b-it32kBalanced
gemma-3-4b-it32kFast
gemma-3n-e4b-it8kOn-device class (4B effective)
gemma-3n-e2b-it8kOn-device class (2B effective)

What about DeepSeek / Mistral / Cohere?

Earlier catalog versions listed DeepSeek, Mistral, and Cohere as native providers. Those are currently not in the public catalog — if your workload depends on them, reach out to [email protected] to discuss adding them as a BYOK channel on your account.

Checking the live catalog

Use the List Models endpoint to see exactly which models your token can access right now:

curl https://api.soxai.io/v1/models \
  -H "Authorization: Bearer $SOXAI_API_KEY"

Or browse the public catalog without signing in:

curl https://console.soxai.io/api/public/models

The public endpoint returns the same list this page is generated from, but updated in real time — always the source of truth.