Respan | LLM Engineering Platform

Route, observe, and evaluate every LLM call

Route LLM traffic through one gateway. Observe calls in traces, watch spend in metrics, and run evals.

Trace

Evaluate

Monitor

Optimize

Behaviors

001 Gateway

Reach every model, and stay up.

Point the SDK you already use at one endpoint to reach 1,000+ models, with automatic fallbacks, response caching, and spend limits built in.

One API for every model

Send every request to one endpoint, reach 1,000+ models across every major provider, and switch between them by changing a single word.

99.99% with fallbacks
96.43% single provider
Rspn with fallbacks 100% Single provider 66%

One model fails, the next takes over

List the models to fall back to, and Respan moves to the next the moment one errors or rate-limits, so an outage never takes you down.

Cache hits 128,402
Cost saved $412.90
Tokens saved 8.2M
Last used
Expiry
Hit count
Cache

4s ago 28d 1,872 summarize the invoice into three bullet points
18s ago 27d 3,214 classify this support ticket by intent
41s ago 25d 2,905 extract the line items and totals as JSON
2m ago 24d 1,530 translate this reply into Spanish
5m ago 22d 1,204 draft a reply to the refund request

Repeat requests served from cache

Cache a request and its response, then serve the same call again instantly, cutting the cost and latency of every repeat to nothing.

$3,000 $0 $5,000

Budgets and limits at every scope

Set budgets and rate limits per key, per customer, or org-wide, get warned as you near them, and block requests before spend runs away.

002 Evals

Ship the version that scores best.

Score every output against the bar you set, so you know it holds up on a test set and in production, instead of just hoping it does.

Evaluator

rag_quality_v1

Every way to grade an output, in one evaluator

An LLM judge, a deterministic code check, or a human reviewer, composed into one evaluator that turns any output into a single score.

Production traffic

Sampling 2% of 12,480 requests
support_replies 240 rows

{ "ques": "refund the duplicate charge" } Refund issued to card ending 4242
{ "ques": "cancel my subscription" } Downgraded to the Free plan
{ "ques": "where's my order #4021?" } Shipped Monday, arriving Tuesday

Test cases sampled straight from production

Pull real requests from your logs by filter and sampling rate, or upload a CSV, so every case is one your users actually sent.

support_agent· d7b2e94

See which version scores better, before you ship

Run a model or prompt against the dataset, then compare their score distributions side by side, so you ship the change that measurably wins.

Faithfulness

Scoring 5% of production spans

Span Input Score Cost
support_agent upgrade me to the pro plan 0.90 $0.0004
support_agent track my package status 0.95 $0.0004
draft_reply change my delivery address 0.63 $0.0005
classify_intent is the invoice tax included 0.91 $0.0003
retrieve_docs how do I reset my password 0.88 $0.0006
draft_reply I was double charged this month 0.58 $0.0005
support_agent update my payment method 0.94 $0.0004
draft_reply why was I charged twice 0.61 $0.0006
classify_intent cancel my subscription today 0.92 $0.0003
support_agent apply my discount code 0.93 $0.0003
summarize_thread the agent never resolved my issue 0.55 $0.0007
draft_reply this is the third time asking 0.66 $0.0005
retrieve_docs where is my refund 0.86 $0.0007

The same evaluator, now scoring live traffic

Deploy an evaluator on production spans, filtered by status, customer, or thread and sampled to control cost, so regressions surface in real time.

003 Monitoring

Know the moment things change.

Track requests, errors, cost, latency, and tokens on one dashboard, and get alerted the moment any of them crosses a threshold you set.

Error rate
Requests 128,402 +4.2%
Error rate 0.38% -0.1%
Cost $412.90 +2.6%
Tokens 48.2M +3.1%

Top models by requests 24h

Every metric on one dashboard

Requests, errors, cost, latency, and tokens across all your traffic, sliced by model, key, or user, so a spike is easy to spot.

004 Tracing

See exactly what your agents did.

Every LLM call, tool run, retrieval, and agent turn becomes a span in one trace, with its input, output, latency, and cost captured.

support_agent· 583eb39

Each step of a request, in one trace

LLM calls, tool runs, retrievals, and agent turns each become a span in one trace, showing exactly where a run's time and cost go.

draft_reply.generation· a1e77c2

Latency: 0.74s Cost: $0.0119 Tokens: 2,754

Input 1,412w · 9,240c

{
  "model": "claude-opus-4-8",
  "temperature": 0.2,
  "max_tokens": 1024,
  "messages": [
    { "role": "system", "content": "You are a support agent. Answer only from the retrieved context." },
    { "role": "user", "content": "Why was I charged twice this month?" },
    { "role": "assistant", "content": "Let me check your recent invoices and payment retries." }
  ]
}

Any span, down to the last field

Model, latency, cost, tokens, and the exact input and output sit on every span, showing precisely what each step ran and returned.

The AI observability platform behind 80 trillion+ tokens.

Loved by world-class founders, engineers, and product teams.

"Imagine jumping to a log immediately after every LLM call. This is the dream for debugging."

— Daniel Wolf, Product Lead, AlphaSense

"We scaled from 5M to 500M+ monthly API calls quickly. Respan gave us the debugging layer to resolve production issues 10x faster."

— Zexia Zhang, CTO, Retell AI

"Respan legit has some of the best UX/DX I’ve ever seen in my life. I truly don’t think I’ve ever integrated a product that was as easy."

— Rahul Behal, Co-founder, Gumloop

"This one felt pretty nice."

— Fabian Hedin, CTO, Lovable

"Such a no brainer choice over LangSmith or anything else and super easy to set up."

— Andy Wang, CEO, Finta

"Respan has been key in helping us scale to trillions of tokens reliably with real-time observability."

— Deshraj Yadav, CTO, Mem0

"Great product - really love the metrics dashboard."

— Esha Dinne, CTO, Giga

Certified where it counts.

Built to meet the security and privacy standards that enterprise and healthcare teams require.