Streaming
How to use Server-Sent Events (SSE) for token streaming
Streaming
SoxAI supports streaming responses via Server-Sent Events (SSE). When stream: true is set, tokens are delivered as they are generated by the upstream model rather than waiting for the full completion.
Enabling Streaming
Add "stream": true to your request body:
{
"model": "gpt-5.4-mini",
"messages": [{"role": "user", "content": "Tell me a story."}],
"stream": true
}SSE Format
Each event is delivered as:
data: <json>\n\nThe stream ends with:
data: [DONE]\n\nChunk Structure
{
"id": "chatcmpl-abc123",
"object": "chat.completion.chunk",
"created": 1743350400,
"model": "gpt-5.4-mini",
"choices": [
{
"index": 0,
"delta": {
"role": "assistant",
"content": "Once"
},
"finish_reason": null
}
]
}The final chunk before [DONE] will have finish_reason set to "stop" (or another finish reason) and an empty delta.content.
Python Example
from openai import OpenAI
client = OpenAI(
api_key="sox-your-token-here",
base_url="https://api.soxai.io/v1",
)
with client.chat.completions.stream(
model="gpt-5.4-mini",
messages=[{"role": "user", "content": "Write a poem about the sea."}],
) as stream:
for text in stream.text_stream:
print(text, end="", flush=True)Node.js Example
const stream = await client.chat.completions.create({
model: "gpt-5.4-mini",
messages: [{ role: "user", content: "Write a poem about the sea." }],
stream: true,
});
for await (const chunk of stream) {
const content = chunk.choices[0]?.delta?.content ?? "";
process.stdout.write(content);
}Raw SSE Parsing
If you are not using an SDK, parse SSE manually:
import httpx
with httpx.stream(
"POST",
"https://api.soxai.io/v1/chat/completions",
headers={"Authorization": "Bearer sox-your-token-here"},
json={
"model": "gpt-5.4-mini",
"messages": [{"role": "user", "content": "Hello!"}],
"stream": True,
},
) as response:
for line in response.iter_lines():
if line.startswith("data: "):
payload = line[6:]
if payload == "[DONE]":
break
import json
chunk = json.loads(payload)
content = chunk["choices"][0]["delta"].get("content", "")
print(content, end="", flush=True)First Token Latency (TTFT)
SoxAI adds less than 10 ms of overhead to the time-to-first-token compared to calling the upstream provider directly. This overhead includes authentication, quota checking, billing pre-consume, and channel selection.
Handling Disconnections
If the client disconnects mid-stream, SoxAI detects this and cancels the upstream request where possible. Billing is settled based on tokens generated up to the point of cancellation.