Models
Per-model usage, cost, cache, and responsiveness — measured client-side, as your users feel it.
Model breakdown
View
| Model | Provider | Requests | Errors | Cost | % of spend | In tokens | Out tokens | Cache rate | p50 | p95 | Price breakdown | Custom price |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| gpt-4o | openai | 84 | — | $0.93 | 68.2% | 202,471 | 41,940 | 21.0% | 2410ms | 4007ms | ||
Embedding: 8.1k tok | ||||||||||||
| claude-sonnet-4 | anthropic | 28 | 1 | $0.33 | 24.4% | 49,582 | 12,197 | 9.0% | 2835ms | 5322ms | ||
Embedding: 2.0k tok | ||||||||||||
| gpt-4o-mini | openai | 162 | 4 | $0.08 | 6.0% | 268,732 | 69,201 | 21.0% | 895ms | 1381ms | ||
Embedding: 10.7k tok | ||||||||||||
| claude-3-5-haiku | anthropic | 5 | — | $0.02 | 1.3% | 7,998 | 2,841 | 9.0% | 608ms | 905ms | ||
Embedding: 320 tok | ||||||||||||
| gpt-4o-internal | openai | 1 | — | $0.00 | 0.0% | 12,400 | 980 | 21.0% | 1040ms | 1040ms | ||
Embedding: 496 tok | ||||||||||||
Provider breakdown
| Provider | Requests | Share | Errors | Cost | % of spend | In tokens | Out tokens | Cache rate | p50 | p95 |
|---|---|---|---|---|---|---|---|---|---|---|
| openai | 247 | 88.2% | 4 | $1.01 | 74.2% | 483,603 | 112,121 | 21.0% | 1.16s | 3.74s |
| anthropic | 33 | 11.8% | 1 | $0.35 | 25.8% | 57,580 | 15,038 | 9.0% | 2.66s | 5.32s |
Typical latency (p50)
1.23s
Slow calls (p95)
3.77s
Latency over time
All calls · wall-clock duration
Error rate & reliability
Failures by kind & status
| Kind | Status | Count |
|---|---|---|
| invalid_request | 400 | 2 |
| rate_limited | 429 | 1 |
| timeout | 504 | 1 |
| overloaded | 529 | 1 |