Quota Management
Configure and manage usage quotas for users and teams
Quota Management
SoxAI's quota system lets administrators enforce fine-grained usage limits across multiple time windows simultaneously.
Quota Dimensions
Each quota policy can limit usage across four dimensions:
| Dimension | What It Limits |
|---|---|
input_tokens | Total prompt tokens per window |
output_tokens | Total completion tokens per window |
requests | Total API calls per window |
amount_cents | Total spend in USD cents per window |
Any single dimension reaching its limit blocks further requests until the window resets.
Time Windows
| Window | Use Case |
|---|---|
1h | Prevent burst abuse, protect against runaway scripts |
5h | Medium-term throttling |
24h | Daily budget control |
7d | Weekly spending cap |
Windows are rolling (not fixed calendar windows). A 1h limit tracks the last 60 minutes from the current moment.
Creating a Quota Policy
Navigate to Admin → Quota Policies → Create Policy:
{
"name": "Standard Developer",
"allowed_models": [],
"limits": {
"1h": {
"input_tokens": 100000,
"output_tokens": 50000,
"requests": 200,
"amount_cents": 500
},
"24h": {
"input_tokens": 1000000,
"output_tokens": 500000,
"requests": 2000,
"amount_cents": 3000
},
"7d": {
"input_tokens": -1,
"output_tokens": -1,
"requests": 10000,
"amount_cents": 15000
}
}
}A value of -1 means no limit for that dimension/window combination.
Assigning Policies
Policies are assigned at three levels (highest priority wins):
1. User-Level Override
Assign a specific policy to one user, overriding their group policy:
Console: Team → Members → [User] → Quota Override → Select Policy
2. Group Policy
All members of a group inherit the group's policy:
Console: Team → Groups → [Group] → Edit → Quota Policy
3. Tenant Default
Applied to all users without a more specific policy:
Console: Admin → Tenants → [Tenant] → Default Quota Policy
Model Restrictions
The allowed_models field in a quota policy acts as an allowlist. Leave empty to permit all models, or specify a list:
{
"allowed_models": ["gpt-5.4-nano", "gpt-5.4-mini", "gemini-embedding-001"]
}Requests to unlisted models return 403 Model Not Allowed.
Monitoring Quota Usage
Per user: Console → Dashboard → Usage → Select time window
Per team: Console → Team → Usage
Admin view: Admin → Users → [User] → Usage
API Endpoint
curl https://api.soxai.io/api/usage/me \
-H "Authorization: Bearer $SOXAI_API_KEY"{
"windows": {
"1h": {
"input_tokens": {"used": 12500, "limit": 100000},
"output_tokens": {"used": 3200, "limit": 50000},
"requests": {"used": 47, "limit": 200},
"amount_cents": {"used": 48, "limit": 500}
}
}
}Quota Exceeded Response
When any limit is hit, requests return 429 Quota Exceeded:
{
"error": {
"code": "quota_exceeded",
"message": "1-hour input token quota exceeded (100,000/100,000). Resets in 34 minutes.",
"request_id": "req_01jq4abc"
}
}Best Practices
Start permissive, tighten over time. Begin with generous limits and observe actual usage patterns before enforcing strict quotas.
Use the amount_cents limit as a safety net. Even if token limits are generous, a spending cap prevents runaway costs from expensive models.
Create model-restricted policies for external integrations. If you embed the API key in a customer-facing app, restrict it to cheap models like gpt-5.4-nano.