SoxAIDocs
API Reference

Streaming

How to use Server-Sent Events (SSE) for token streaming

Streaming

SoxAI supports streaming responses via Server-Sent Events (SSE). When stream: true is set, tokens are delivered as they are generated by the upstream model rather than waiting for the full completion.

Enabling Streaming

Add "stream": true to your request body:

{
  "model": "gpt-5.4-mini",
  "messages": [{"role": "user", "content": "Tell me a story."}],
  "stream": true
}

SSE Format

Each event is delivered as:

data: <json>\n\n

The stream ends with:

data: [DONE]\n\n

Chunk Structure

{
  "id": "chatcmpl-abc123",
  "object": "chat.completion.chunk",
  "created": 1743350400,
  "model": "gpt-5.4-mini",
  "choices": [
    {
      "index": 0,
      "delta": {
        "role": "assistant",
        "content": "Once"
      },
      "finish_reason": null
    }
  ]
}

The final chunk before [DONE] will have finish_reason set to "stop" (or another finish reason) and an empty delta.content.

Python Example

from openai import OpenAI

client = OpenAI(
    api_key="sox-your-token-here",
    base_url="https://api.soxai.io/v1",
)

with client.chat.completions.stream(
    model="gpt-5.4-mini",
    messages=[{"role": "user", "content": "Write a poem about the sea."}],
) as stream:
    for text in stream.text_stream:
        print(text, end="", flush=True)

Node.js Example

const stream = await client.chat.completions.create({
  model: "gpt-5.4-mini",
  messages: [{ role: "user", content: "Write a poem about the sea." }],
  stream: true,
});

for await (const chunk of stream) {
  const content = chunk.choices[0]?.delta?.content ?? "";
  process.stdout.write(content);
}

Raw SSE Parsing

If you are not using an SDK, parse SSE manually:

import httpx

with httpx.stream(
    "POST",
    "https://api.soxai.io/v1/chat/completions",
    headers={"Authorization": "Bearer sox-your-token-here"},
    json={
        "model": "gpt-5.4-mini",
        "messages": [{"role": "user", "content": "Hello!"}],
        "stream": True,
    },
) as response:
    for line in response.iter_lines():
        if line.startswith("data: "):
            payload = line[6:]
            if payload == "[DONE]":
                break
            import json
            chunk = json.loads(payload)
            content = chunk["choices"][0]["delta"].get("content", "")
            print(content, end="", flush=True)

First Token Latency (TTFT)

SoxAI adds less than 10 ms of overhead to the time-to-first-token compared to calling the upstream provider directly. This overhead includes authentication, quota checking, billing pre-consume, and channel selection.

Handling Disconnections

If the client disconnects mid-stream, SoxAI detects this and cancels the upstream request where possible. Billing is settled based on tokens generated up to the point of cancellation.