Observability

Every token, accounted for.

Full traces, per-user spend, routing decisions, tool usage and audit logs — stitched into one dashboard. Replay any agent decision, export to your SIEM.

DASHBOARD · LIVE trace #8842
2,148
req/s
41ms
p50 latency
$0.0031
avg cost/req
quality score96.2
compression−53%
routed tollama-70b · groq
Full traces

Replay any agent decision.

Every request is a trace — policy checks, compression, routing decision, tool calls, model response — stitched together with latency and cost at each step.

TRACE #8842 · support-agent 41ms · $0.0031 · quality 96.2
Step Action Detail Latency Tokens Cost
01 policy check rate.limit ✓ · cost.ceil ✓ 0.8ms
02 compress 4 transforms applied 3.2ms 4,096 → 1,925 · −53% −$0.0076
03 route complexity 0.31 → economy 2.1ms
04 Ggithub.pull repo: cohesor/api · #482 142ms 3,847 → 1,612 $0.0011
05 Mgmail.search tool call" · max 5 89ms 412 $0.0003
06 → model · llama-70b final llm call 92ms 1,925 in · 412 out $0.0015
total latency 41ms tokens saved −2,171 (53%) total cost $0.0031 quality 96.2
Dashboards

The whole fleet,
on one screen.

Traffic over time, spend by model, top users, top tools — live dashboards that answer "what's happening right now?" without a query.

  • Traffic over time. Requests per second, per agent, per model — live and historical.
  • Spend by model. Which models are costing the most, and what routing is saving you.
  • Top users & tools. Who's spending, what they're calling, and where the waste is.
  • No queries needed. The dashboard is pre-built — just open it.
DASHBOARD · OVERVIEW live
TRAFFIC · LAST 10 MIN
SPEND BY MODEL
llama-70b · groq
$86
gpt-5
$39
claude-opus-4.6
$14
TOP TOOLS
G github 842 calls
M gmail 531 calls
C calendar 318 calls
D drive 204 calls
ROUTING DECISIONS per request

See why each request went where.

Every routing decision is logged with the complexity score, the threshold, the chosen model and the savings. No black box — just a decision log you can audit.

09:41:02 economy · score 0.31 · llama-70b
09:41:01 strong · score 0.92 · claude-opus
09:41:01 economy · score 0.28 · llama-70b
09:41:00 economy · score 0.44 · gpt-5
09:40:59 strong · score 0.88 · claude-opus
09:40:59 economy · score 0.22 · llama-70b
09:40:58 economy · score 0.35 · llama-70b
AUDIT LOG streamable

Every decision, on the record.

Policy decisions, tool calls, model selections — all logged with user, timestamp, arguments and cost. Stream to your SIEM in real time, or export for compliance.

09:41:02 rate.limit · 142/s
09:41:02 cost.ceil · 142/s
09:41:01 rate.limit · 1 deny
09:41:01 cost.ceil · 142/s
09:41:00 rate.limit · 142/s
09:41:00 cost.ceil · 142/s
09:40:59 budget.cap · key blocked
09:40:59 rate.limit · 142/s
41ms
p50 trace

Full request traces — policy, compress, route, tools, model — stitched at p50 41ms.

100+
models tracked

Spend, latency and quality per model — across every provider, in one view.

SIEM
streamable

Audit logs stream to your SIEM in real time — no export step, no delay.

Get started

See every token. Replay any decision.

Full traces, live dashboards and audit logs — built in, no extra setup. Open the dashboard on day one.