Data Loss Prevention (DLP)
Scan every AI request for credentials, PII, and sensitive data before it reaches the upstream model.
Data Loss Prevention (DLP)
Overview
DLP runs in the relay hot path after channel selection and before the upstream request. It scans the JSON request body for sensitive content and applies a policy action — mask, block, or audit — before the request leaves SoxAI. Masking rewrites the body in place; the upstream receives a redacted version. Blocking rejects the request before it reaches any external service.
Relay Pipeline
tokenAuth → quotaCheck → channelSelect
↓
✦ applyDLP ← scans + rewrites/blocks here
↓
applyRequestTransform
↓
upstream (model provider)Builtin Detectors
SoxAI ships 15 builtin detectors. All are enabled by default and can be suppressed per policy.
| Code | Description | Post-validator | Category |
|---|---|---|---|
credit_card | Visa, Mastercard, Amex, Discover | Luhn algorithm | Financial |
cn_id_card | Chinese national ID (18-digit) | GB 11643 check digit | PII |
openai_api_key | OpenAI sk-... keys | — | Credentials |
anthropic_api_key | Anthropic sk-ant-... keys | — | Credentials |
aws_access_key | AWS AKIA... access key IDs | — | Credentials |
aws_secret_key | 40-char AWS secret keys | — | Credentials |
gcp_service_key | GCP service account JSON keys | — | Credentials |
generic_jwt | eyJ... JWT tokens (any algorithm) | — | Credentials |
pem_private_key | PEM-encoded RSA/EC/Ed25519 private keys | — | Credentials |
us_ssn | US Social Security Numbers (NNN-NN-NNNN) | — | PII |
email | Email addresses | — | PII |
cn_phone | Chinese mobile numbers (11-digit 1xx) | — | PII |
ipv4_internal | Private IPv4 addresses | RFC 1918 ranges | PII |
bitcoin_wallet | Bitcoin mainnet addresses | — | Financial |
mac_address | IEEE 802 MAC addresses | — | PII |
Custom Detectors
Create tenant-private detectors in Console → Security → DLP → Detectors → New Detector.
Regex Detectors
Use RE2 syntax. Supply a confidence score (0–100) that flows into policy decisions. Optionally add a post-validator to reduce false positives.
Example — match an internal project code like PROJ-12345:
Pattern: \bPROJ-\d{5}\b
Confidence: 90Dictionary Detectors
Supply a newline-separated list of terms. The engine uses Aho-Corasick for O(n) multi-term matching. Useful for codenames, internal system identifiers, or customer lists.
Example terms:
project-phoenix
operation-nightfall
internal-codename-xPolicies
A policy binds one or more detectors to a scope and an action. Create policies in Console → Security → DLP → Policies → New Policy.
Actions
| Action | What happens |
|---|---|
mask | Matched spans are replaced with [REDACTED_<detector_code>]. The rewritten body is forwarded to the upstream. The caller receives a normal response. |
block | The request is rejected before reaching the upstream. The caller receives a 403 error. |
audit_only | The request passes unchanged. A finding is recorded for review. |
Scope Bindings
A policy can target any combination of scopes. Multiple bindings use union semantics — if any scope matches, the policy applies.
| Scope | Description |
|---|---|
| Team | All requests from members of a specific team |
| User | Requests from a specific user |
| Token | Requests authenticated with a specific API token |
Minimum Confidence
Set a minimum confidence threshold (0–100) per policy. Findings below the threshold are ignored by that policy. This lets you enable a detector globally (e.g., email) but only act on high-confidence matches.
API Behavior
When a Request Is Blocked
HTTP status 403 Forbidden. Response body:
{
"error": {
"code": "dlp_blocked",
"message": "Request blocked: credit card number detected in prompt.",
"detector": "credit_card",
"request_id": "req_01abc..."
}
}message describes which detector triggered. detector is the detector code (e.g. credit_card, openai_api_key). request_id correlates with gateway logs.
When a Request Is Masked
The request completes normally (200). The upstream receives the rewritten body. The caller is not notified that masking occurred — the response reflects what the model produced given the redacted input.
When Audit-Only
The request completes normally (200). No visible change to the caller or the upstream.
Findings & Audit
Every mask and block finding (and all audit_only findings) is encrypted and stored.
- Encryption: AES-256-GCM. The matched text is never stored in plaintext.
- Viewing findings: Console → Security → DLP → Findings. Findings show detector code, action taken, timestamp, and scope metadata. The matched text is stored encrypted and is not visible until decrypted.
- Decrypting a finding: Only
system_adminaccounts can decrypt. Every decrypt triggers an immutable entry in the audit log. Navigate to a finding and click Decrypt — you will be prompted for step-up authentication.
See Also
- Prompt Guard — Detect adversarial prompt injection attacks
- Data Privacy — What data SoxAI stores and how long