API Reference
Error Codes
HTTP error codes and error response format for the SoxAI API
Error Codes
All errors are returned as JSON with a consistent structure:
{
"error": {
"code": "machine_readable_code",
"message": "Human-readable explanation of the error.",
"request_id": "req_01jq4abc..."
}
}Use the request_id when contacting support — it maps to a trace ID for full request diagnostics.
HTTP Status Codes
| Status | Code | Meaning | Action |
|---|---|---|---|
400 | invalid_request | Malformed JSON or missing required parameter | Fix the request body |
400 | model_not_found | The specified model is not available | Check Supported Models |
400 | context_length_exceeded | Input tokens exceed the model's context window | Truncate the conversation history |
400 | content_policy_violation | Request was blocked by content policy | Review prompt content |
401 | invalid_token | Token is missing, malformed, or revoked | Check the Authorization header |
401 | token_expired | Token has passed its expiration date | Create a new token in the Console |
403 | model_not_allowed | Your token's model allowlist excludes this model | Request access or use a different model |
403 | ip_not_allowed | Request originated from a disallowed IP | Update the token's IP allowlist |
429 | quota_exceeded | A quota window limit has been reached | Wait for the window to reset or contact your admin |
429 | rate_limited | Too many requests per second | Implement exponential backoff |
500 | internal_error | Unexpected server error | Retry with exponential backoff; contact support if persistent |
503 | no_available_channel | All upstream providers for this model are unavailable | Retry after a delay |
Handling 429 Errors
When you receive a 429, check the error code:
quota_exceeded— A quota window has been exceeded. The message includes which window (1h,5h,24h,7d) and the reset time.rate_limited— You are sending too many requests per second. Implement backoff.
import time
from openai import OpenAI, RateLimitError
client = OpenAI(
api_key="sox-your-token-here",
base_url="https://api.soxai.io/v1",
)
def chat_with_retry(messages, max_retries=3):
for attempt in range(max_retries):
try:
return client.chat.completions.create(
model="gpt-5.4-mini",
messages=messages,
)
except RateLimitError as e:
if attempt == max_retries - 1:
raise
wait = 2 ** attempt # exponential backoff: 1s, 2s, 4s
time.sleep(wait)Handling 503 Errors
A 503 means all configured upstream providers for the requested model failed their fallback chain. This is rare and usually temporary.
Retry with exponential backoff:
async function chatWithRetry(messages: Message[], maxRetries = 3) {
for (let attempt = 0; attempt < maxRetries; attempt++) {
try {
return await client.chat.completions.create({ model: "gpt-5.4-mini", messages });
} catch (err) {
if (err instanceof OpenAI.APIError && err.status === 503) {
if (attempt === maxRetries - 1) throw err;
await new Promise((r) => setTimeout(r, 1000 * 2 ** attempt));
continue;
}
throw err;
}
}
}