/metrics on each pod’s HTTP port. This page lists the metrics worth watching for each component and how to read them. For scrape and export configuration, see Configure SmithDB observability.
Names are as they appear on /metrics. If your collector adds a namespace or prefix, adjust accordingly.
For Datadog and Grafana dashboards built on these metrics, with the scrape annotations each one expects, see the SmithDB observability examples in the LangSmith Helm chart repository.
This page covers metrics SmithDB emits. Kubernetes signals such as OOM kills, container restarts, CPU and memory against limits, and cache disk usage are worth alerting on but come from your infrastructure monitoring, not from SmithDB.
Baseline metrics
If the full reference is more than you need, the following metrics answer whether SmithDB is healthy. All other metrics provide additional detail.
Compaction has two entries because the failure modes are separate: jobs can fail repeatedly while the queue length stays flat, and the queue can grow while every job that runs succeeds.
Ingestion
Query
Compaction
Migration
Migration Job pods emit metrics during a historical migration. How you aggregate a metric across pods depends on its type:- Gauges: Report totals for the whole migration. The values come from TaskDB and refresh every two minutes. Every pod reports the same values, so aggregate with
max, notsum. - Counters: Track the work each pod does. Sum their rates across pods.
Migration does not retry failed tasks or jobs on its own. Any
failed task, or a failed or validation_failed job count, keeps the migration Job from reaching Complete. The Job keeps running until you resolve the failure. See Migration Job failures.
All components
LangSmith ingestion path
Emitted by LangSmith rather than by SmithDB, and labeledstore="clickhouse|smithdb", so the two stores can be compared directly during dual ingestion.
Connect these docs to your agent of choice via MCP for real-time answers.

