SoxAIDocs
API Reference

Rate Limits

How rate limiting and quota enforcement works in SoxAI

Rate Limits

SoxAI enforces two distinct types of limits: rate limits (requests per second) and quota limits (usage over time windows).

Rate Limits

Rate limits cap the number of requests your token can make per second. These are enforced to protect gateway stability and ensure fair sharing.

PlanRequests/secondBurst
Free25
Starter1030
Pro50100
EnterpriseCustomCustom

When a rate limit is hit, you receive a 429 response with code rate_limited. Implement exponential backoff.

Quota Limits

Quota limits cap cumulative usage over rolling time windows. Your administrator configures these at the team or user level.

Time Windows

WindowDescription
1hRolling 1-hour window
5hRolling 5-hour window
24hRolling 24-hour window
7dRolling 7-day window

Quota Dimensions

Each window can independently limit:

DimensionDescription
input_tokensTotal prompt tokens consumed
output_tokensTotal completion tokens generated
requestsTotal API requests made
amount_centsTotal spend in USD cents

Quota Exceeded Response

{
  "error": {
    "code": "quota_exceeded",
    "message": "Hourly input token quota exceeded (100,000/100,000). Resets in 23 minutes.",
    "request_id": "req_01jq4abc"
  }
}

The message indicates which window and dimension was exceeded, and when it resets.

Checking Your Usage

View current usage in the Console under Dashboard → Usage. You can see usage broken down by time window, model, and team member.

Quota Priority

Quotas are applied in this priority order:

  1. User-level override — set by your team admin for your specific account
  2. Group policy — applied to all members of your group
  3. Tenant default — the platform-wide default
  4. No limit — system admin accounts have no quota by default

The first matching policy wins.

Best Practices

Pre-estimate token usage. For large batch jobs, estimate tokens before sending to avoid mid-batch quota errors. Most tokenizers are available as open-source libraries (tiktoken for OpenAI models).

Monitor usage proactively. Use the Console dashboard or set up webhook alerts (see Webhooks) to get notified before you hit quota.

Use streaming for long responses. Streaming does not change quota accounting, but it improves user experience when approaching token limits.

Contact your admin for limit increases. Quota limits are configurable. If your workload requires higher limits, ask your team administrator to update your quota policy.