Smart routing, on by the token
`helix/auto` picks the cheapest model that clears your quality bar, then drops to a faster one mid-stream if a turn is simple. You set a latency budget; we hold the line.
route: nearest-edge · budget_ms: 40 01Helix is the inference API for agents that can't wait. Sub-40 ms first token, streamed from 312 edge nodes — one endpoint, every frontier model.
We routed a million agent turns through Helix and the median provider. Same models, same prompts. The gap is the part your users feel.
Swap a base URL and you're done. Helix speaks the chat-completions dialect you already wrote against — then adds the parts the spec forgot.
`helix/auto` picks the cheapest model that clears your quality bar, then drops to a faster one mid-stream if a turn is simple. You set a latency budget; we hold the line.
route: nearest-edge · budget_ms: 40 01True token-level SSE with mid-flight cancel and resume. Kill a runaway generation without burning the rest of your budget.
SSE · cancel() · resume() 02Prompts and completions are dropped the instant a request closes. SOC 2 Type II, no training on your data — ever, on every tier.
soc2 · zdr · region-pinned 03Per-agent, per-route cost telemetry streamed to your dashboard in real time — not a CSV that lands three days into the next billing cycle. Set hard caps; we stop the bleed before it's an incident.
live cost stream · per-agent caps · alerts 04Helix runs inference on 312 nodes pressed up against the backbone. Requests never hairpin to a single region — they resolve at the metro nearest the caller.
Spin up a key, point your base URL at Helix, and watch your first-token times collapse. $5 in credits on signup — enough to run the benchmark yourself.