SoxAIDocs
API Reference

Error Codes

HTTP error codes and error response format for the SoxAI API

Error Codes

All errors are returned as JSON with a consistent structure:

{
  "error": {
    "code": "machine_readable_code",
    "message": "Human-readable explanation of the error.",
    "request_id": "req_01jq4abc..."
  }
}

Use the request_id when contacting support — it maps to a trace ID for full request diagnostics.

HTTP Status Codes

StatusCodeMeaningAction
400invalid_requestMalformed JSON or missing required parameterFix the request body
400model_not_foundThe specified model is not availableCheck Supported Models
400context_length_exceededInput tokens exceed the model's context windowTruncate the conversation history
400content_policy_violationRequest was blocked by content policyReview prompt content
401invalid_tokenToken is missing, malformed, or revokedCheck the Authorization header
401token_expiredToken has passed its expiration dateCreate a new token in the Console
403model_not_allowedYour token's model allowlist excludes this modelRequest access or use a different model
403ip_not_allowedRequest originated from a disallowed IPUpdate the token's IP allowlist
429quota_exceededA quota window limit has been reachedWait for the window to reset or contact your admin
429rate_limitedToo many requests per secondImplement exponential backoff
500internal_errorUnexpected server errorRetry with exponential backoff; contact support if persistent
503no_available_channelAll upstream providers for this model are unavailableRetry after a delay

Handling 429 Errors

When you receive a 429, check the error code:

  • quota_exceeded — A quota window has been exceeded. The message includes which window (1h, 5h, 24h, 7d) and the reset time.
  • rate_limited — You are sending too many requests per second. Implement backoff.
import time
from openai import OpenAI, RateLimitError

client = OpenAI(
    api_key="sox-your-token-here",
    base_url="https://api.soxai.io/v1",
)

def chat_with_retry(messages, max_retries=3):
    for attempt in range(max_retries):
        try:
            return client.chat.completions.create(
                model="gpt-5.4-mini",
                messages=messages,
            )
        except RateLimitError as e:
            if attempt == max_retries - 1:
                raise
            wait = 2 ** attempt  # exponential backoff: 1s, 2s, 4s
            time.sleep(wait)

Handling 503 Errors

A 503 means all configured upstream providers for the requested model failed their fallback chain. This is rare and usually temporary.

Retry with exponential backoff:

async function chatWithRetry(messages: Message[], maxRetries = 3) {
  for (let attempt = 0; attempt < maxRetries; attempt++) {
    try {
      return await client.chat.completions.create({ model: "gpt-5.4-mini", messages });
    } catch (err) {
      if (err instanceof OpenAI.APIError && err.status === 503) {
        if (attempt === maxRetries - 1) throw err;
        await new Promise((r) => setTimeout(r, 1000 * 2 ** attempt));
        continue;
      }
      throw err;
    }
  }
}