SoxAIDocs
Models

Provider Routing

How SoxAI routes requests across multiple upstream AI providers

Provider Routing

SoxAI intelligently routes requests across multiple AI providers using a priority-based fallback system with session stickiness.

Priority Fallback Chain

Each model in SoxAI can be backed by multiple upstream providers (channels), ordered by priority. When you make a request for gpt-5.4-mini, SoxAI might have three channels configured:

gpt-5.4-mini: priority=1 (OpenAI Direct)
      → priority=2 (Azure OpenAI East US)
      → priority=3 (Azure OpenAI West EU)

Normal operation: Priority 1 handles all requests.

On failure: If Priority 1 returns a 5xx error, times out, or returns a 429 (rate limited), SoxAI retries up to 3 times on that channel, then falls back to Priority 2, and so on.

All channels exhausted: If every channel fails, SoxAI returns a 503 No Available Channel response.

Health Scoring

Each channel has a continuous health score (0–100) based on recent performance:

Health Score = (Success Rate × 70%) + (Latency Score × 30%)
Health ScoreStatusRouting Behavior
80–100HealthyNormal priority routing
40–79DegradedDeprioritized in routing
0–39UnhealthySent only as last resort

Channels recover automatically — every 5 minutes a probe request is sent to unhealthy channels, and their score rises on success.

Session Stickiness

For multi-turn conversations, it is often desirable to route all requests in a session to the same upstream provider. This ensures context consistency, especially when providers maintain conversation state differently.

To enable session stickiness, include the X-Session-ID header:

curl https://api.soxai.io/v1/chat/completions \
  -H "Authorization: Bearer $SOXAI_API_KEY" \
  -H "X-Session-ID: conv-user-123-session-abc" \
  -H "Content-Type: application/json" \
  -d '{"model": "gpt-5.4-mini", "messages": [...]}'

On the first request with a new session ID, SoxAI selects a channel using the normal priority algorithm and records the binding in Redis (TTL: 24 hours by default).

Subsequent requests with the same session ID are routed to the bound channel, bypassing the priority algorithm.

Automatic unbinding: If the bound channel fails 3 consecutive times, SoxAI clears the binding and selects the next best channel, recording a new binding.

Choosing Session IDs

Use a stable, unique identifier for each logical conversation:

  • User ID + conversation ID: user-42-conv-789
  • Session token hash: sess-a1b2c3
  • UUID: 550e8400-e29b-41d4-a716-446655440000

Do not use session IDs that are guessable or correlate to sensitive data — they are not secret, only used for routing decisions.

Load Balancing

When multiple channels have the same priority, SoxAI distributes load across them using weighted round-robin based on their health scores. Higher-health channels receive more traffic.

Admin Configuration

Channel routing is configured by administrators in the Console under Admin → Channels. Each channel specifies:

  • Provider type (OpenAI, Anthropic, Azure, etc.)
  • API key (encrypted at rest)
  • Supported models
  • Priority level
  • Request limits (optional, for budget control per channel)