SoxAIDocs
Guides

Quota Management

Configure and manage usage quotas for users and teams

Quota Management

SoxAI's quota system lets administrators enforce fine-grained usage limits across multiple time windows simultaneously.

Quota Dimensions

Each quota policy can limit usage across four dimensions:

DimensionWhat It Limits
input_tokensTotal prompt tokens per window
output_tokensTotal completion tokens per window
requestsTotal API calls per window
amount_centsTotal spend in USD cents per window

Any single dimension reaching its limit blocks further requests until the window resets.

Time Windows

WindowUse Case
1hPrevent burst abuse, protect against runaway scripts
5hMedium-term throttling
24hDaily budget control
7dWeekly spending cap

Windows are rolling (not fixed calendar windows). A 1h limit tracks the last 60 minutes from the current moment.

Creating a Quota Policy

Navigate to Admin → Quota Policies → Create Policy:

{
  "name": "Standard Developer",
  "allowed_models": [],
  "limits": {
    "1h": {
      "input_tokens": 100000,
      "output_tokens": 50000,
      "requests": 200,
      "amount_cents": 500
    },
    "24h": {
      "input_tokens": 1000000,
      "output_tokens": 500000,
      "requests": 2000,
      "amount_cents": 3000
    },
    "7d": {
      "input_tokens": -1,
      "output_tokens": -1,
      "requests": 10000,
      "amount_cents": 15000
    }
  }
}

A value of -1 means no limit for that dimension/window combination.

Assigning Policies

Policies are assigned at three levels (highest priority wins):

1. User-Level Override

Assign a specific policy to one user, overriding their group policy:

Console: Team → Members → [User] → Quota Override → Select Policy

2. Group Policy

All members of a group inherit the group's policy:

Console: Team → Groups → [Group] → Edit → Quota Policy

3. Tenant Default

Applied to all users without a more specific policy:

Console: Admin → Tenants → [Tenant] → Default Quota Policy

Model Restrictions

The allowed_models field in a quota policy acts as an allowlist. Leave empty to permit all models, or specify a list:

{
  "allowed_models": ["gpt-5.4-nano", "gpt-5.4-mini", "gemini-embedding-001"]
}

Requests to unlisted models return 403 Model Not Allowed.

Monitoring Quota Usage

Per user: Console → Dashboard → Usage → Select time window

Per team: Console → Team → Usage

Admin view: Admin → Users → [User] → Usage

API Endpoint

curl https://api.soxai.io/api/usage/me \
  -H "Authorization: Bearer $SOXAI_API_KEY"
{
  "windows": {
    "1h": {
      "input_tokens": {"used": 12500, "limit": 100000},
      "output_tokens": {"used": 3200, "limit": 50000},
      "requests": {"used": 47, "limit": 200},
      "amount_cents": {"used": 48, "limit": 500}
    }
  }
}

Quota Exceeded Response

When any limit is hit, requests return 429 Quota Exceeded:

{
  "error": {
    "code": "quota_exceeded",
    "message": "1-hour input token quota exceeded (100,000/100,000). Resets in 34 minutes.",
    "request_id": "req_01jq4abc"
  }
}

Best Practices

Start permissive, tighten over time. Begin with generous limits and observe actual usage patterns before enforcing strict quotas.

Use the amount_cents limit as a safety net. Even if token limits are generous, a spending cap prevents runaway costs from expensive models.

Create model-restricted policies for external integrations. If you embed the API key in a customer-facing app, restrict it to cheap models like gpt-5.4-nano.