SoxAI
researchagentsengineeringlanggraph

Building SoxAI Deep Research: From LangGraph Pipelines to Real-Time Chat

An inside look at how we built SoxAI Deep Research — a multi-model deep research agent powered by LangGraph, Tavily, and our own AI gateway. Architecture, design decisions, and the road ahead.

SoxAI Team·

Building SoxAI Deep Research: From LangGraph Pipelines to Real-Time Chat

When we started SoxAI, the core mission was straightforward: give teams a single, reliable endpoint to access any AI model. But as we built and used the platform ourselves, a question kept coming up — what if SoxAI could do research for you, not just route your requests?

That question became SoxAI Deep Research: a fully integrated AI research agent that turns a single question into a structured, cited report in minutes. Built on our own gateway, it also became the ultimate stress test for everything we had built.

Why We Built This

Most AI research tools today are black boxes. You type a question, wait 2-5 minutes, and get a report. You have no idea how much it cost, which models were used, or which sources were actually read versus skimmed.

We wanted something different:

  • Cost transparency: Every research session shows an exact breakdown — how much went to LLM calls, how much to search APIs, total to the cent.
  • Model choice: We're not locked into one provider. A research session can use Gemini to plan, GPT-4o Mini to search, and Claude Opus to synthesize.
  • Real-time progress: Watch the research unfold — see each sub-question, each source, each chunk of the report as it's being written.
  • No subscription lock-in: Pay per research session starting at ~$0.20, with no monthly commitment.

And beyond the product: building a research agent on top of our own gateway was the best possible dogfood. If SoxAI couldn't handle 25 parallel LLM calls with accurate per-session cost attribution, we needed to know.

The Architecture

SoxAI Deep Research runs as a standalone Python microservice alongside the main Go backend. The separation was intentional — AI agent workloads have fundamentally different operational characteristics than API gateway traffic.

Browser (research.soxai.io)
  │
  ├─── Next.js BFF (apps/console)
  │      └─ /api/research/* → research-agent
  │
  └─── SSE stream ──→ research-agent (FastAPI + LangGraph)
                           │
                ┌──────────┼──────────────┐
                │          │              │
           SoxAI      SoxAI Gateway   External APIs
           Server     (all LLM calls)  (Tavily, Jina)
        (auth, quota)

Three independent processes, communicating only through well-defined HTTP APIs:

ServiceTechRole
apps/consoleNext.js 15UI + BFF proxy
apps/research-agentPython 3.12 + FastAPI + LangGraphAgent orchestration + SSE
apps/serverGo + GinAuth, quota, billing

All LLM traffic — every planner call, every searcher call, every synthesis token — routes through the SoxAI gateway. This gives us accurate cost attribution per session, per node, per model, without any custom billing code in the research service itself.

The LangGraph Pipeline

The research pipeline has three nodes: Planner, Searcher, and Synthesizer. We chose LangGraph for orchestration because it gives us stateful graph execution with a clean model for fan-out parallelism.

State

Everything flows through a single typed state object:

class ResearchState(TypedDict):
    session_id: str
    user_id: int
    tier: str  # "draft" | "standard" | "deep"
    question: str

    # Planner output
    sub_questions: list[dict]

    # Searcher output
    sources: list[dict]

    # Synthesizer output
    report_md: str

    # Cost tracking
    total_cost_units: int
    llm_calls: list[dict]
    external_api_calls: list[dict]

    # LangGraph message history
    messages: Annotated[list, add_messages]

Planner Node

The Planner receives the original question and decomposes it into focused sub-questions. The number of sub-questions — and the model used — depends on the research tier:

TierSub-questionsPlanner ModelCost Ceiling
Draft6Gemini Flash~$0.20
Standard12GPT-4o Mini~$0.50
Deep25Claude Sonnet~$2.00

The prompt asks the model to detect the question's language and generate sub-questions in that same language — so a Chinese research question produces Chinese sub-questions throughout the pipeline. This was a small addition with a meaningful impact on report coherence.

Searcher Fan-out

After planning, the searcher node fans out across all sub-questions in parallel using asyncio.gather with a concurrency cap:

async def searcher_node(state: ResearchState) -> dict:
    semaphore = asyncio.Semaphore(5)  # max 5 concurrent Tavily calls

    async def search_with_limit(sub_q):
        async with semaphore:
            return await search_single(sub_q, state)

    results = await asyncio.gather(
        *[search_with_limit(sq) for sq in state["sub_questions"]],
        return_exceptions=True,
    )
    return {"sources": [r for r in results if not isinstance(r, Exception)]}

Each search call:

  1. Calls Tavily Advanced search (with include_raw_content=True)
  2. Falls back to Jina Reader for full-page extraction when needed
  3. Writes sources to the database immediately (persist-before-parse pattern)
  4. Emits a source SSE event as each source is found

For a Deep tier research session, this means up to 25 Tavily queries running concurrently (subject to the semaphore limit of 5).

Synthesizer Node

The Synthesizer receives all collected sources and writes the final report as a streaming response. It:

  • Cites sources inline with [1], [2] notation
  • Produces tables and Mermaid diagrams where appropriate
  • Streams tokens back to the browser via Redis pub/sub → SSE

We use graph.astream_events() — LangGraph's unified async event stream — to capture token-level output and forward it to the SSE endpoint. The Synthesizer prompt varies by tier: Draft gets a concise summary prompt, Deep gets an instruction set that resembles a senior analyst's briefing document template.

Crash Recovery

LangGraph supports PostgreSQL-backed checkpointing via langgraph-checkpoint-postgres, which can resume a graph from the last completed node after a worker crash. We plan to wire this up in a future release — for now, the Arq queue handles retries at the job level, and the reconciliation reaper (see Billing section) issues refunds for sessions that fail mid-run.

Real-Time Progress with SSE

One of the design goals was that users can watch the research happen. We implemented this with Server-Sent Events rather than WebSockets — SSE is simpler, works through proxies, and supports automatic reconnection via Last-Event-ID.

The event flow from backend to browser:

LangGraph node emits event
  → normalized to typed SSE event
  → published to Redis channel research:session:{id}:events
  → FastAPI SSE endpoint subscribes and forwards
  → browser EventSource receives and renders

Six application event types cover the full lifecycle:

status         → "planning" | "searching" | "synthesizing"
sub_question   → one per planned sub-question (streamed as they're generated)
source         → each source as it's found during search
report_chunk   → each token of the final report (streaming)
done           → cost breakdown (LLM + external API split)
failed         → sanitized error message

The SSE transport uses sse-starlette's built-in 15-second ping to keep idle connections alive through Cloudflare Tunnel, which cuts connections idle for more than 100 seconds.

If the SSE connection drops, the frontend reconnects with Last-Event-ID and the backend replays any missed events from the session's event log.

Billing Integration

Every research session is backed by SoxAI's quota system:

  1. Pre-consume: At session start, we reserve the tier's cost ceiling from the user's balance (e.g., $0.50 for Standard). If insufficient balance, return 402 before any LLM call is made.
  2. Track: Each gateway LLM call returns X-SoxAI-Cost-Units in the response header. Each external API call (Tavily, Jina) has a fixed price per call. Both are written to research.llm_calls and research.external_api_calls immediately after the call.
  3. Settle: When the session completes, we sum actual costs and settle: actual = Σ llm_calls + Σ external_api_calls, refund ceiling - actual back to the user.

The settle logic handles all failure modes:

OutcomeWhat the user pays
SuccessActual cost (always ≤ ceiling)
User cancelled (ESC)Calls made before cancel
Timeout (>5 min)Full refund
Total failureFull refund

We added a reconciliation reaper that runs hourly, scanning for sessions with dangling pre-consumes that were never settled. These are almost always dead workers — the reaper issues refunds automatically.

The Chat Evolution

The original research UI was "one question in, one report out." Users loved the reports but immediately asked: can I ask a follow-up?

We've now redesigned the UI around a chat-first paradigm. A research session is one turn in a conversation. After the report arrives, you can:

  • Ask follow-up questions that get answered using the report as context
  • Start a new research thread on a related topic
  • Attach images to provide visual context for your question

The underlying data model evolved from sessions as the top-level entity to conversations containing turns, where a turn can be either a chat message or a research task. The research history sidebar now shows conversation titles, not session IDs.

This shift also unlocked model selection at the conversation level — you pick Claude Opus, GPT-4o, or any model in your plan's whitelist, and follow-up messages use that model directly (no LangGraph overhead for simple follow-ups).

What We Learned Building This

Persist before parse. Our original synthesizer parsed the LLM response before writing the llm_calls record. When the parser threw an exception, the cost record was lost — we had charged the user but had no record. The fix: always write the raw cost to the database immediately after the gateway call returns, before any processing.

SSE and Cloudflare need heartbeats. Cloudflare Tunnel cuts connections idle for >100 seconds. 15-second heartbeats are non-negotiable in production.

Tier configuration drift is a real problem. Early on, we had tier settings in three places: a YAML file loaded at startup, a DB row fetched per-session, and some hardcoded fallbacks. Sessions would silently use the wrong model. We fixed this by taking a config_snapshot JSONB into each session row at creation time — the session always knows exactly what configuration it was run with.

The gateway makes cost attribution trivial. Because every LLM call goes through our own gateway with X-SoxAI-Cost-Units in the response header, the research agent doesn't need any model-specific cost calculation logic. This was one of our most validating moments — the gateway abstraction paid for itself.

Performance and Cost

In production, a Standard tier research (12 sub-questions) takes 90-120 seconds end-to-end and costs users $0.15-$0.35. The actual breakdown:

PhaseTimeCost share
Planning3-5s~5%
Searching (12 parallel)40-60s~35% (Tavily)
Synthesis (streaming)45-60s~60% (LLM)

Deep tier (25 sub-questions) runs 3-5 minutes and costs $0.80-$1.50 in practice, well below the $2.00 ceiling. The ceiling exists to protect users, not to extract revenue.

Roadmap

We're continuing to invest heavily in Research. The features we're actively building:

Q2 2026

  • PDF and DOCX export — download reports in publication-ready formats
  • Project knowledge base — upload your own documents and have them included as context in research sessions
  • Admin model configuration — system admins can configure which models are available per tier without touching YAML files

Q3 2026

  • Brave Search and Firecrawl fallback — additional search provider options for different content types
  • Research sharing — generate a public link to share a research report
  • Scheduled research — run a research session on a schedule (e.g., weekly competitive analysis)
  • White-label research — enterprise tenants can brand the research interface with their own domain and identity

Q4 2026 and beyond

  • Multi-agent collaboration — multiple specialized agents (financial data, legal databases, scientific literature) working in parallel on Deep sessions
  • Citation verification — automatically check that cited content actually supports the claims made in the report
  • Custom research personas — define domain-specific instructions that shape how the agent approaches questions in your field

Getting Started

SoxAI Deep Research is available to all SoxAI users. You need a positive balance to start a session — Draft sessions cost as little as $0.20 with no subscription required.

  1. Sign up at console.soxai.io — free $5 credit included
  2. Navigate to Research in the sidebar
  3. Type your question, pick a tier, and watch the research happen

If you're a developer interested in the technical details — the LangGraph pipeline, the SSE event schema, the billing integration — reach out at [email protected]. We're happy to talk architecture.


SoxAI Deep Research uses LangGraph for agent orchestration, Tavily for web search, and runs all LLM calls through the SoxAI gateway for unified cost attribution.