Query Bench

Query-engine performance over a 9,996-span synthetic corpus — byte-identical OTLP data in all five platforms (seed 43, 12-hour window, parity-gated). This measures what the 3-trace cohorts could not: the engine behind the endpoint, not the HTTP round trip.

Aug 15, 2026 (re-run: LangSmith v2 + Braintrust SQL-over-/btql) · dataset/out_v2/query_bench.json · pipeline: dataset/

Filtered scan

Server-side filter "LLM spans where model = claude-haiku-4-5", first 100 rows · median of 3 · each platform's native filter (SQL / BTQL / DSL / params)

Server-side aggregation

"Span count + Σ input tokens per model" in one call · median of 3

Phoenix has no server-side aggregation of any kind — its rollup was computed client-side during the export pass (not comparable, shown as ✗). † LangSmith only offers fixed-shape /runs/stats (no GROUP BY): its 5.6 s bar answers a weaker question than the other three.

Full-corpus export throughput

Page through all ~10k rows in the window · one timed pass · log scale — every gridline is 10×

All numbers

PlatformScan p50Agg p50 Export wallRows/sMethod

Read these before quoting the numbers

  • Logfire timings exclude its client-side pacing. The ~10 queries/min rate limit is paid with untimed 7 s sleeps between calls; wall-clock including pacing is 14.7 s for the export. The engine is fast; the meter is small.
  • Braintrust now measures through SQL over /btql (its canonical path since PR #7): the time-window filter runs server-side and export uses cursor-paged SQL — 7x faster than the legacy /fetch route (62 → 8.5 s). Its pilot project still holds the earlier 7-day corpus, but the SQL WHERE excludes it server-side.
  • LangSmith's numbers are from the v2 query APIs. The deprecated v1 path measured 2,177 ms scans and a 329 s export in the first pass — v2 cut those 7–11x. What remains: pages cap at 100 rows (so a 10k export is ~100 round trips) and the only server-side aggregate is the fixed-shape v1 /runs/stats (5.6 s, no GROUP BY).
  • Langfuse and Logfire share projects with the live benchmark data (time-window + service-name isolation); the other three use dedicated pilot projects.
  • Single machine, single session, n=3 medians for scans/aggregations, n=1 for exports. 10k spans is the pilot scale — the 200k run is where these gaps should be re-measured.