How to Build AI Agents That Use Multiple LLMs (With Fallback)
A practical guide to building AI agents that can switch between GPT-4o, Claude, Gemini, and DeepSeek dynamically — with automatic failover when a provider goes down.
How to Build AI Agents That Use Multiple LLMs (With Fallback)
AI agents are only as reliable as the models behind them. If your agent depends on a single LLM provider and that provider has an outage, rate-limits you, or degrades in quality — your agent stops working.
The better approach: build agents that can dynamically switch between models based on the task, cost, and availability.
Why Multi-Model Agents?
Every LLM has strengths:
| Task | Best Model | Why |
|---|---|---|
| Complex reasoning | Claude Opus 4 | Strongest on multi-step logic |
| Fast Q&A / triage | GPT-4o Mini | Cheapest at $0.15/1M input |
| Code generation | Claude Sonnet 4 | Top coding benchmarks |
| Long document analysis | Gemini 2.5 Pro | 1M token context window |
| Cost-sensitive batch jobs | DeepSeek V3 | $0.27/1M, near GPT-4o quality |
A well-designed agent picks the right model for each step — like a team of specialists instead of one generalist.
The Naive Approach: Multiple SDKs
# Don't do this — it's fragile and hard to maintain
from openai import OpenAI
import anthropic
import google.generativeai as genai
def agent_step(task, complexity):
if complexity == "high":
# Anthropic SDK — different API format
client = anthropic.Anthropic(api_key="sk-ant-...")
response = client.messages.create(
model="claude-sonnet-4-20250514",
messages=[{"role": "user", "content": task}]
)
return response.content[0].text
elif complexity == "low":
# OpenAI SDK
client = OpenAI(api_key="sk-...")
response = client.chat.completions.create(
model="gpt-4o-mini",
messages=[{"role": "user", "content": task}]
)
return response.choices[0].message.content
else:
# Google SDK — yet another format
genai.configure(api_key="...")
model = genai.GenerativeModel("gemini-2.5-pro")
return model.generate_content(task).textProblems with this approach:
- Three SDKs — three sets of dependencies, three API formats
- Three API keys — three billing accounts to monitor
- No failover — if Claude is down, you get an error, not a fallback
- No cost tracking — you need to aggregate billing from three dashboards
The Better Approach: Unified API
Use a single OpenAI-compatible endpoint that routes to any provider:
from openai import OpenAI
# One client, one API key — routes to any provider
client = OpenAI(
api_key="your-gateway-key",
base_url="https://api.soxai.io/v1",
)
def agent_step(task: str, model: str = "gpt-4o-mini") -> str:
"""Execute one agent step with any model, same SDK."""
response = client.chat.completions.create(
model=model,
messages=[{"role": "user", "content": task}],
)
return response.choices[0].message.contentNow your agent can use any model with zero code changes:
# Agent workflow: analyze → plan → execute
analysis = agent_step(document, model="gemini-2.5-pro") # Long context
plan = agent_step(analysis, model="claude-sonnet-4-20250514") # Best reasoning
results = [
agent_step(task, model="gpt-4o-mini") # Cheap execution
for task in plan.split("\n")
]Adding Automatic Failover
What happens when Claude returns a 503? With the naive approach, your agent crashes. With a gateway, failover is handled at the infrastructure level:
Your Agent → Gateway → Claude (primary)
↓ (503 error)
→ GPT-4o (fallback)
↓ (success)
Response returned to agentYour agent code doesn't change at all — the gateway retries and falls back automatically. You configure fallback chains once:
claude-sonnet-4 → gpt-4o → gemini-2.5-pro
gpt-4o-mini → gemini-2.5-flash → deepseek-v3Real-World Agent Example: Document Processor
Here's a complete agent that processes documents using multiple models optimally:
from openai import OpenAI
client = OpenAI(
api_key="your-soxai-key",
base_url="https://api.soxai.io/v1",
)
def process_document(doc: str) -> dict:
# Step 1: Classify (cheap, fast)
doc_type = client.chat.completions.create(
model="gpt-4o-mini",
messages=[{
"role": "user",
"content": f"Classify this document in one word (invoice/contract/report/other):\n{doc[:500]}"
}],
).choices[0].message.content.strip().lower()
# Step 2: Extract (needs quality)
extraction = client.chat.completions.create(
model="claude-sonnet-4-20250514",
messages=[{
"role": "system",
"content": f"Extract structured data from this {doc_type}. Return JSON.",
}, {
"role": "user",
"content": doc,
}],
).choices[0].message.content
# Step 3: Summarize (long context if needed)
model = "gemini-2.5-pro" if len(doc) > 100_000 else "gpt-4o-mini"
summary = client.chat.completions.create(
model=model,
messages=[{
"role": "user",
"content": f"Summarize in 3 bullet points:\n{doc}"
}],
).choices[0].message.content
return {
"type": doc_type,
"data": extraction,
"summary": summary,
}Cost breakdown for 1,000 documents:
- Classification (GPT-4o Mini): ~$0.08
- Extraction (Claude Sonnet): ~$18.00
- Summarization (mixed): ~$2.50
- Total: ~$20.58 vs ~$45.00 if using Claude for everything
Cost Tracking per Agent Step
When you're running agents at scale, you need to know which steps cost the most. With unified billing, you see all providers in one dashboard:
| Agent Step | Model | Avg Cost/Call | Daily Calls | Daily Cost |
|---|---|---|---|---|
| Classify | GPT-4o Mini | $0.00008 | 10,000 | $0.80 |
| Extract | Claude Sonnet | $0.018 | 10,000 | $180.00 |
| Summarize | Mixed | $0.0025 | 10,000 | $25.00 |
Now you know: extraction is 87% of your cost. Maybe DeepSeek V3 can handle 60% of extractions at 1/10th the price? Test it without changing your agent code — just update the model name.
Getting Started
- Sign up for SoxAI — free $5 credit, no card required
- Get your API key from the dashboard
- Replace your
base_url— your existing OpenAI code works immediately - Start using any model from any provider with the same SDK
All 200+ models, one API key, automatic failover, unified billing.