Rate Limits
How rate limiting and quota enforcement works in SoxAI
Rate Limits
SoxAI enforces two distinct types of limits: rate limits (requests per second) and quota limits (usage over time windows).
Rate Limits
Rate limits cap the number of requests your token can make per second. These are enforced to protect gateway stability and ensure fair sharing.
| Plan | Requests/second | Burst |
|---|---|---|
| Free | 2 | 5 |
| Starter | 10 | 30 |
| Pro | 50 | 100 |
| Enterprise | Custom | Custom |
When a rate limit is hit, you receive a 429 response with code rate_limited. Implement exponential backoff.
Quota Limits
Quota limits cap cumulative usage over rolling time windows. Your administrator configures these at the team or user level.
Time Windows
| Window | Description |
|---|---|
1h | Rolling 1-hour window |
5h | Rolling 5-hour window |
24h | Rolling 24-hour window |
7d | Rolling 7-day window |
Quota Dimensions
Each window can independently limit:
| Dimension | Description |
|---|---|
input_tokens | Total prompt tokens consumed |
output_tokens | Total completion tokens generated |
requests | Total API requests made |
amount_cents | Total spend in USD cents |
Quota Exceeded Response
{
"error": {
"code": "quota_exceeded",
"message": "Hourly input token quota exceeded (100,000/100,000). Resets in 23 minutes.",
"request_id": "req_01jq4abc"
}
}The message indicates which window and dimension was exceeded, and when it resets.
Checking Your Usage
View current usage in the Console under Dashboard → Usage. You can see usage broken down by time window, model, and team member.
Quota Priority
Quotas are applied in this priority order:
- User-level override — set by your team admin for your specific account
- Group policy — applied to all members of your group
- Tenant default — the platform-wide default
- No limit — system admin accounts have no quota by default
The first matching policy wins.
Best Practices
Pre-estimate token usage. For large batch jobs, estimate tokens before sending to avoid mid-batch quota errors. Most tokenizers are available as open-source libraries (tiktoken for OpenAI models).
Monitor usage proactively. Use the Console dashboard or set up webhook alerts (see Webhooks) to get notified before you hit quota.
Use streaming for long responses. Streaming does not change quota accounting, but it improves user experience when approaching token limits.
Contact your admin for limit increases. Quota limits are configurable. If your workload requires higher limits, ask your team administrator to update your quota policy.