Ranking of agent-observability platforms on how well their read APIs give trace data back — the same three traced agent scenarios written into five OTEL-based backends, then benchmarked with executable scripts on 9 weighted criteria.
| Rank ⓘ | Platform | Score ⓘ | Δ vs prev ⓘ | Latency ⓘ | Freshness ⓘ | Query flex ⓘ | Exports ⓘ |
|---|
Where each total comes from: every segment is one criterion's weighted contribution, in the fixed rubric order (completeness → otel-fidelity). Solid = the 7 judge-scored criteria (84 pts max); hatched = the 2 measured metrics, latency + freshness (16 pts max). Hover any segment.
Judge scores 0–10 per criterion (weight in parentheses). Hover a cell for its weighted contribution to the total.
The arena above queries projects holding 3 demo traces, so its latency column is dominated by HTTP round-trip. A separate benchmark loads ~10k byte-identical synthetic spans into all five platforms and measures the engine itself.
The ranking inverts under load. Logfire — 4th on simple request latency — aggregates 10k spans per model in 149 ms (SQL GROUP BY) and exports at 19,992 rows/s, an order of magnitude ahead of everyone. Even after its v2 migration cut everything 7–11x, LangSmith still takes 30 s to hand back its own 10k rows (pages cap at 100).
Two different races. Phoenix stays the sprinter (fastest small reads, 0.1 s write-to-read, quick filtered scans) but has no server-side aggregation at all — per-model token totals require downloading everything. Speed and analytical power are different axes; the arena score alone shows neither.
Corpus: 9,996 spans, seed 43, parity-gated · dataset/
What you actually call, and what gets in the way.
POST /v2/query · SQL · pylf_v2 read token
Strength: full SQL over spans, and the widest export menu (JSON, NDJSON, CSV, Arrow).
Friction: returns Arrow binary unless you send Accept: application/json; min_timestamp is required; ~10 queries/min rate limit.
POST /btql · SQL over project_logs() · Bearer
Strength: standard SQL (contributed via PR #7 by Braintrust), three real export formats via fmt — json, ndjson, parquet — and sub-second freshness (0.4 s in every cohort).
Friction: csv returns 400; /fetch pagination can re-return rows — dedupe by id; demo traces aged out of the project between cohorts (retention), so evidence is fetched deterministically by span name.
POST /api/v2/runs/query · POST /api/v2/traces/query · X-Api-Key + X-Tenant-Id
Strength: expressive filter DSL, cursor pagination (next_cursor), and — since the v2 migration — one of the fastest read paths in the field.
Friction: projects addressed by UUID plus an X-Tenant-Id header (resolve both first); Parquet export is plan-gated. The deprecated v1 query path was the platform's old bottleneck — migrating to v2 took retrieval from 1.9 s to 123 ms and freshness to 0.1 s (Aug 13 cohort).
GET /api/public/v2/observations · HTTP Basic pk/sk
Strength: biggest incumbent jump July → August (65.63 → ~80 continuous) after moving to the observations v2 API; sub-140 ms retrieval in the Aug 12 cohort.
Friction: rows are observations — group by traceId to reconstruct traces; I/O fields need explicit ?fields= opt-in; slowest ingest freshness in both August cohorts (~6–11 s).
GET /s/{space}/v1/projects/{id}/spans · Bearer
Strength: fastest read path (87.5 ms) and write-to-read lag (0.1 s) measured in any cohort; verbatim gen_ai.* keys with W3C ids plus an OTLP-JSON export endpoint; the space serves its own OpenAPI spec.
Friction: deliberately simple filters (time, name, attribute, span_kind) — no server-side aggregation and no SQL/DSL, so analysis is client-side; traces reconstructed by grouping on context.trace_id.