> ## Documentation Index
> Fetch the complete documentation index at: https://docs.goldsky.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Advanced metrics dashboard

> Reference for every panel on the per-pipeline Advanced metrics dashboard: what it measures and when to use it

## Overview

The **Advanced metrics** dashboard is a Grafana dashboard that shows detailed metrics for a single Turbo pipeline. It goes deeper than the project-wide [Health overview](/turbo-pipelines/health-dashboard): it breaks metrics down per component (each source, transform, and sink), and it adds checkpoint internals and container resource usage.

To open it:

1. Sign in to the [Goldsky dashboard](https://app.goldsky.com/dashboard/pipelines) and open a pipeline from the **Pipelines** page.
2. Click **Advanced metrics**. Grafana opens in a new tab with the dashboard filtered to that pipeline.

You can also open it by clicking a pipeline name in the **Pipeline Status** panel of the [Health overview](/turbo-pipelines/health-dashboard#pipeline-status).

<Tip>
  Use the Health overview to find the pipeline that has a problem. Then use this dashboard to find the component that causes it.
</Tip>

## Filters

The filters at the top of the dashboard apply to all panels.

| Filter | What it does |
| - | - |
| **Service Instance ID** | The pipeline to show. Set automatically when you open the dashboard from a pipeline. |
| **Component ID** | Limit panels to one or more components, by the name you gave them in the pipeline YAML. |
| **Node Type** | Limit panels to `source`, `transform`, or `sink` components. |
| **Operator Type** | Limit panels to one implementation, for example `sql`, `postgres`, or `webhook`. |
| **Rate Time Window** | The time window that rate and increase calculations use (1m to 1h). A longer window gives smoother lines. A shorter window shows short spikes. |
| **Kafka Partition** | Limit the Kafka panels to one or more partitions. |

<Note>
  Some rows apply only to some pipelines. For example, the **Solana Source** row has data only for Solana sources. A panel that shows **No data** usually means that the pipeline does not use that feature. It does not mean that there is a problem.
</Note>

## Pipeline overview

General throughput, latency, and lag for each component.

| Panel | What it shows | Use it to |
| - | - | - |
| **Component Summary - Total Records** | A table of each component with its node type, operator type, and total input and output records since the pipeline started. | Confirm that records reach every component. A transform whose output is much lower than its input is filtering records. |
| **Cumulative Records Processed** | Total input and output records for each component over time. | Find the time when a component stopped receiving or sending records (the line goes flat). |
| **Overall Pipeline Throughput (records/sec)** | Output records per second for each component. | Compare throughput between components and find the slowest one. |
| **Avg Batch Size by Component** | Average number of records in each batch that a component outputs. | Check how full batches are. Many small batches cause more overhead than fewer large batches. |
| **Record Batch Processing Latency (p50, p95)** | Time that each component takes to process one batch, at the 50th and 95th percentile. | Find a slow transform or sink. A high p95 with a normal p50 means that some batches are much slower than others. |
| **Record Batch Average Processing Time** | Average time that each component takes to process one batch. | See the general trend of processing time, with less noise than the percentile panel. |
| **Block Lag (Max)** | How far behind the chain tip the pipeline is, in seconds, over time. | Find out when the pipeline started to fall behind, and whether it is catching up. |
| **Block Lag Gauge** | The current block lag, in seconds. | See the current lag at a glance. |

Block lag is available only for pipelines whose output includes a block number or timestamp column. Block lag is an end-to-end metric: a slow sink can cause it to grow. For more information, see [Block lag](/turbo-pipelines/health-dashboard#block-lag) on the Health overview page.

## Checkpoints

A checkpoint saves the position of the pipeline after all sinks confirm that they wrote a set of records. These panels show how long that takes and whether it succeeds. For more information about checkpoints, see [Delivery guarantees](/turbo-pipelines/delivery-guarantees).

| Panel | What it shows | Use it to |
| - | - | - |
| **Checkpoint Epochs** | Number of checkpoints that succeeded and failed in each time window. | Detect failed checkpoints. Any failed checkpoint needs investigation. |
| **Checkpoint Epochs In Flight** | Number of checkpoints in progress. | Check that checkpoints complete. A value that stays above zero for a long time means that a checkpoint is stuck, usually because a sink did not confirm its write. |
| **Checkpoint Epoch Duration (p50, p95)** | Time from the start of a checkpoint to its completion. | Measure how long records stay in the pipeline before the pipeline confirms delivery. |
| **Checkpoint Sink Flush (p50, p95)** | Time that each sink takes to write its data and confirm a checkpoint. | Find the sink that makes checkpoints slow. |
| **Checkpoint Messages** | Number of checkpoint markers sent, confirmations (acks) received, and finalizations sent. | Check that every marker gets a confirmation. If acks received stays lower than markers sent, a sink is not confirming. |
| **Checkpoint Marker Arrival (p50, p95)** | Time that a checkpoint marker takes to move through the pipeline and arrive at each component. | Find a component that delays other components. A slow component upstream makes markers arrive late at all components after it. |

<Tip>
  If **Checkpoint Epoch Duration** is high, compare **Checkpoint Sink Flush** and **Checkpoint Marker Arrival**:

  * High sink flush: the sink writes slowly. Check the destination database or service.
  * High marker arrival: records wait upstream of the sink. Check the latency of transforms, or the batch settings.
</Tip>

## Kafka

Metrics for components that read from Kafka (consumer) or write to Kafka (producer). EVM dataset sources in streaming mode read from Kafka. The producer panels have data only if the pipeline has a [Kafka sink](/turbo-pipelines/sinks/kafka).

| Panel | What it shows | Use it to |
| - | - | - |
| **Kafka Consumer Lag (Aggregated)** | Number of messages that the source has not read yet, summed across all partitions. | Check whether the source keeps up with new data. Lag that increases steadily means that the pipeline falls behind. |
| **Kafka Consumer Lag by Partition** | Number of unread messages for each partition. | Find a single partition that falls behind while the others keep up. |
| **Kafka Consumer Message Decode Latency (p50, p95)** | Time to decode messages read from Kafka. | Find out if decoding is the bottleneck for a source. |
| **Kafka Producer Message Encode Latency (p50, p95)** | Time to encode messages before the sink sends them to Kafka. | Find out if encoding is the bottleneck for a Kafka sink. |
| **Kafka Producer Batch Send Latency (p50, p95)** | Time to send a batch of messages to the Kafka broker. | Find a slow broker or network problems between the pipeline and your Kafka cluster. |
| **Kafka Consumer Row Kind Count Rate** | Number of records read from Kafka, by row kind (insert, update, or delete). | See the mix of operations that the source receives. |

High Kafka lag during a backfill is normal. For more information, see [Pipeline status](/turbo-pipelines/health-dashboard#pipeline-status) on the Health overview page.

## Solana source

These panels have data only for pipelines with a [Solana source](/turbo-pipelines/sources/solana).

| Panel | What it shows | Use it to |
| - | - | - |
| **Block Rate** | Number of blocks per second that the source processes. | Check the speed of the source. |
| **Buffer Size** | Number of blocks in the source's internal buffer. | Detect backpressure. A buffer that grows means that the rest of the pipeline cannot process blocks as fast as the source fetches them. |
| **Block Slot** | The next Solana slot that the source will process. | Track progress. A flat line means that the source has stopped. |
| **Fetch Duration** | Time to fetch data from the Solana data source, at the 95th percentile. | Find out if the slow part is the fetch itself. |

## Handler/Webhook HTTP metrics

These panels have data only for pipelines with an [HTTP handler transform](/turbo-pipelines/transforms/http-handler) or a [webhook sink](/turbo-pipelines/sinks/webhook).

| Panel | What it shows | Use it to |
| - | - | - |
| **Requests per minute by status** | Number of HTTP requests per minute, by status class: `2xx`, `4xx`, `5xx`, or, when no response arrives, `timeout` or `network_error`. | Find errors that your endpoint returns, and requests that do not reach it. |
| **Requests per minute by outcome** | Number of HTTP requests per minute, by outcome: `success`, `retriable_error` (the pipeline tries the request again), or `non_retriable_error`. | Find out if errors are temporary or permanent. A steady rate of `non_retriable_error` usually means a configuration problem, such as a wrong URL or bad credentials. |
| **Request Latency** | Response time of your endpoint at the 95th and 99th percentile. | Find out if a slow endpoint limits the pipeline. |

## Kubernetes

Resource usage of the containers that run the pipeline.

| Panel | What it shows | Use it to |
| - | - | - |
| **Pod Status by Phase** | The phase of each pipeline pod over time (for example `Running` or `Pending`). | Detect rescheduling. A new pod name, or a phase other than `Running`, means that the pipeline moved to a new pod. |
| **CPU Usage vs Provisioned** | CPU use, compared to the CPU request and limit. | Find out if the pipeline needs more CPU. Usage at the limit for a long time means that you can increase [`resource_size`](/turbo-pipelines/pipeline-config). |
| **Memory Usage vs Provisioned** | Memory use, compared to the memory request and limit. | Find out if the pipeline is near its memory limit. A pod that reaches its memory limit restarts. |
| **Network Bytes (RX/TX)** | Bytes received and sent per second. | Compare network traffic with throughput, and find unexpected traffic. |
| **Network connections** | The highest number of connections opened, closed, failed, and refused in each time window. | Find connection problems with a source or sink. Failed or refused connections usually mean that a destination is not reachable. |
| **Network Dropped Packets (RX/TX)** | Packets received and sent that were dropped. | Find network problems. This value is usually zero. |

## Next steps

* Use the [Health overview](/turbo-pipelines/health-dashboard) to monitor all pipelines in your project.
* Set up [custom alerts](/turbo-pipelines/custom-alerts) on these metrics.
* Send the same metrics to your own monitoring system with the [Prometheus integration](/turbo-pipelines/prometheus-integration).


## Related topics

- [Health dashboard](/turbo-pipelines/health-dashboard.md)
- [Custom alerts](/turbo-pipelines/custom-alerts.md)
- [Kafka](/turbo-pipelines/sinks/kafka.md)
- [Prometheus integration](/turbo-pipelines/prometheus-integration.md)
- [Monitoring](/edge-rpc/platform/monitoring.md)
