> ## Documentation Index
> Fetch the complete documentation index at: https://langchain-5e9cc07a-preview-docsby-1791319236-3be7a15.mintlify.site/llms.txt
> Use this file to discover all available pages before exploring further.

# SmithDB metrics reference

> Metrics to watch on each SmithDB component, what they measure, and how to read them.

SmithDB components expose Prometheus metrics at `/metrics` on each pod's HTTP port. This page lists the metrics worth watching for each component and how to read them. For scrape and export configuration, see [Configure SmithDB observability](/langsmith/self-host-smithdb-observability).

Names are as they appear on `/metrics`. If your collector adds a namespace or prefix, adjust accordingly.

For Datadog and Grafana dashboards built on these metrics, with the scrape annotations each one expects, see the [SmithDB observability examples](https://github.com/langchain-ai/helm/tree/main/charts/langsmith/examples/smithdb-observability) in the LangSmith Helm chart repository.

<Note>
  This page covers metrics SmithDB emits. Kubernetes signals such as OOM kills, container restarts, CPU and memory against limits, and cache disk usage are worth alerting on but come from your infrastructure monitoring, not from SmithDB.
</Note>

## Baseline metrics

If the full reference is more than you need, the following metrics answer whether SmithDB is healthy. All other metrics provide additional detail.

| Question | Metric | Component |
| - | - | - |
| Are queries succeeding, and fast? | `server_requests_duration` | Query |
| Are writes landing, and fast? | `ingest_batch_latency` | Ingestion |
| Is ingestion keeping up with arrivals? | `ingest_pending_segment_entry_count` | Ingestion |
| Is compaction keeping up? | `compaction_queue_job_count` | Compaction |
| Is compaction succeeding? | `compaction_job_completion` | Compaction |

Compaction has two entries because the failure modes are separate: jobs can fail repeatedly while the queue length stays flat, and the queue can grow while every job that runs succeeds.

## Ingestion

| Metric | Type | Description |
| - | - | - |
| `ingest_batch_latency` | Histogram | End-to-end latency from batch received to flush completed. |
| `ingest_flush_count` | Counter | Flushes to object storage. A flat rate while traffic arrives means writes are not landing. |
| `ingest_flush_rows` | Histogram | Rows per flush, by level. Falling rows at a steady flush rate suggests premature flushing. |
| `ingest_flush_duration` | Histogram | How long flushes take. Rising durations point at object store latency or undersized ingestion. |
| `ingest_flush_reason` | Counter | What triggered each flush, labeled `reason` (`size`, `time`, or `shutdown`) and `level` (`L0` or `L1`). L1 flushes on accumulated run count, or a timeout. |
| `ingest_pending_segment_entry_count` | Gauge | Entries not yet flushed. Sustained growth means flushing is backlogged. |
| `semaphore_permits_in_use` | Gauge | Concurrency permits held, against `semaphore_max_permits`. Sustained saturation means writes are queuing. |
| `server_requests_duration` | Histogram | Ingestion request latency. Split by `status` for success and failure counts. |

## Query

| Metric | Type | Description |
| - | - | - |
| `server_requests_duration` | Histogram | Query latency. Split by `status` for success and failure counts, where `status` is the gRPC code (`0` is success) or the HTTP code, and by `path` to find the slow endpoint. |
| `query_fanout_partitions` | Histogram | Active scan partitions per query. Rising fanout drives both latency and memory. |
| `query_fanout_limit_exceeded_total` | Counter | Queries rejected for excessive fanout. Any sustained rate means queries are failing outright. |
| `query_execution_max_mem_used_bytes` | Histogram | Peak memory for the largest operator in a query plan. Approaching the configured limit precedes termination. |

## Compaction

| Metric | Type | Description |
| - | - | - |
| `compaction_queue_job_count` | Gauge | Jobs waiting in the compaction queue. Sustained growth means compaction is falling behind. |
| `compaction_jobs_scheduled` | Counter | Work created by the pipeline. Compare against jobs executed. |
| `compaction_jobs_executed` | Counter | Work consumed. Persistently below scheduled means under-provisioned workers. |
| `compaction_job_completion` | Counter | Job outcomes, labeled `status` (`success` or `failure`) and `job_kind`. A `job_kind` dropping to zero means that job type has stopped running. |
| `compaction_queue_jobs_skipped_capacity` | Counter | Jobs skipped for exceeding a worker's remaining capacity. A sustained rate means worker capacity is too small for the jobs being produced. |
| `compaction_queue_oldest_pending_created_at_seconds` | Gauge | Unix timestamp when the oldest pending job was enqueued, or zero when the queue is empty. Compute age as `time() - metric`, and exclude the zero case. |
| `compaction_job_latency_seconds` | Histogram | Job creation to completion. Includes queue wait, so it rises when workers saturate as well as when jobs are slow. |
| `compaction_worker_capacity_used` | Gauge | In-flight capacity cost by `job_kind`, against `compaction_worker_capacity_limit`. Sustained use near the limit explains skipped jobs and a growing queue. |
| `compaction_worker_running_tasks` | Gauge | Tasks running by `job_kind`. Zero while the queue is non-empty means workers are stalled, not busy. |

## Migration

Migration Job pods emit metrics during a [historical migration](/langsmith/self-host-smithdb-migrate). How you aggregate a metric across pods depends on its type:

* **Gauges**: Report totals for the whole migration. The values come from TaskDB and refresh every two minutes. Every pod reports the same values, so aggregate with `max`, not `sum`.
* **Counters**: Track the work each pod does. Sum their rates across pods.

Migration Job pods emit the following metrics:

| Metric | Type | Description |
| - | - | - |
| `migration_tasks` | Gauge | Migration tasks, labeled `kind` (`run` or `feedback`) and `status` (`pending`, `running`, `completed`, or `failed`). Progress is the `completed` count against the total across statuses. |
| `migration_jobs` | Gauge | Migration jobs, labeled `kind` and `status`. Run jobs finish as `promoted` and feedback jobs as `validated`. `failed` and `validation_failed` are failures, and every other status means the job is still in progress. |
| `migration_run_tasks_completed_total` | Counter | Run tasks completed by the pod. The rate is migration throughput in tasks. |
| `migration_run_task_planned_rows_migrated_total` | Counter | Rows in the run tasks the pod completed, as counted in ClickHouse when the tasks were planned. The rate is migration throughput in rows. |

Migration does not retry failed tasks or jobs on its own. Any `failed` task, or a `failed` or `validation_failed` job count, keeps the migration Job from reaching `Complete`. The Job keeps running until you resolve the failure. See [Migration Job failures](/langsmith/self-host-smithdb-troubleshooting#migration-job-failures).

## All components

| Metric | Type | Description |
| - | - | - |
| `object_store_op_duration` | Histogram | Object store operation latency. |
| `sys_jemalloc_resident_bytes` | Gauge | Resident process memory. Track against the pod memory limit. |

## LangSmith ingestion path

Emitted by LangSmith rather than by SmithDB, and labeled `store="clickhouse|smithdb"`, so the two stores can be compared directly during [dual ingestion](/langsmith/self-host-smithdb-install#step-4-enable-dual-ingestion).

| Metric | Type | Description |
| - | - | - |
| `langsmith_ingestion_e2e_latency_seconds` | Histogram | API receipt to store write ack, per run. The user-facing number; compare `store="smithdb"` against `store="clickhouse"`. |
| `langsmith_ingestion_api_to_worker_latency_seconds` | Histogram | Queue wait before the worker starts. Rising here is a queue problem, not a SmithDB problem. |
| `langsmith_ingestion_worker_to_store_latency_seconds` | Histogram | Store write time alone. Isolates SmithDB from queue delay. |
| `langsmith_asynq_ingestion_queue_pending` | Gauge | Tasks waiting in the LangSmith ingestion queue before a worker picks them up. Sustained growth means the queue is not keeping up with arrivals, upstream of SmithDB. |

***

<div className="source-links">
  <Callout icon="terminal-2">
    [Connect these docs](/use-these-docs) to your agent of choice via MCP for real-time answers.
  </Callout>

  <Callout icon="edit">
    [Edit this page on GitHub](https://github.com/langchain-ai/docs/edit/main/src/langsmith/self-host-smithdb-metrics.mdx) or [file an issue](https://github.com/langchain-ai/docs/issues/new/choose).
  </Callout>
</div>
