# Introduction

groundcover is a full stack, cloud-native observability platform, developed to break all industry paradigms - from making instrumentation a thing of the past, to decoupling cost from data volumes

The [groundcover ](https://www.groundcover.com)platform consolidates all your traces, metrics, logs, and Kubernetes events into a single pane of glass, allowing you to identify issues faster than ever before and conduct granular investigations for quick remediation and long-term prevention.

Our[ pricing](https://www.groundcover.com/pricing) is not impacted by the volume of data generated by the environments you monitor, so you can dare to start monitoring environments that had been blind spots until now - such as your Dev and Staging clusters. This, in turn, provides you visibility into all your environments, making it much more likely to identify issues in the early stages of development, rather than in your live product.

groundcover introduces game-changing concepts to observability:

<table data-view="cards"><thead><tr><th></th><th data-hidden data-card-target data-type="content-ref"></th><th data-hidden data-card-cover data-type="files"></th></tr></thead><tbody><tr><td><a href="#ebpf-sensor"><strong>eBPF sensor</strong></a></td><td><a href="#ebpf-sensor">#ebpf-sensor</a></td><td></td></tr><tr><td><a href="#bring-your-own-cloud-byoc-architecture"><strong>BYOC architecture</strong></a></td><td><a href="https://docs.groundcover.com/#bring-your-own-cloud-byoc-architecture">https://docs.groundcover.com/#bring-your-own-cloud-byoc-architecture</a></td><td></td></tr><tr><td><a href="#disruptive-pricing-model"><strong>Disruptive pricing</strong></a></td><td><a href="#disruptive-pricing-model">#disruptive-pricing-model</a></td><td></td></tr></tbody></table>

## eBPF sensor <a href="#ebpf-sensor" id="ebpf-sensor"></a>

[eBPF](https://www.groundcover.com/ebpf) (extended Berkeley Packet Filter) is a groundbreaking technology that has significantly impacted the Linux kernel, offering a new way to safely and efficiently extend its capabilities.

By powering our sensor with eBPF, groundcover unlocks unprecedented granularity on your cloud environment, while also practically eliminating the need for human involvement in the installation and deployment process. Our unique sensor collects data directly from the Linux kernel with near-zero impact on CPU and memory.

**Advantages of our eBPF sensor:**

* **Zero instrumentation:** groundcover's eBPF sensor gathers granular observability data without the need for integrating an SDK or changing your applications' code in any way. This enables all your logs, metrics, traces, and other observability data to flow automatically into the platform. In minutes, you gain full visibility into application and infrastructure health, performance, resource usage, and more.
* **Minimal resources footprint:** groundcover’s sensor in installed on a dedicated node in each monitored cluster, operating separately from the applications it is monitoring. Without interference with the application's primary functions, the groundcover platform operates with near-zero impact on your resources, maintaining the applications' performance and avoiding unexpected overhead on the infrastructure.
* **A new level of insight granularity:** With direct access to the Linux kernel, our eBPF sensor enables the collection of data straight from the source. This guarantees that the data is clean, unaltered, and precise. It also offers access to unique insight on your application and infrastructure, such as the ability to view the full traces of payloads, or analyzing network performance over time.

<figure><img src="/files/8dQvYd39hTwGuZBH0oOv" alt=""><figcaption></figcaption></figure>

## **Bring-your-own-cloud (BYOC) architecture**

The one-of-a-kind architecture on which groundcover was built eliminates all requirements to stream your logs, metrics, traces, and other monitoring data outside of your environment and into a third-party's cloud. By leveraging integrations with best-of-breed technologies, including ClickHouse and Victoria Metrics, all your observability is stored data locally, with the ability of being fully managed by groundcover.

**Advantages of our BYOC architecture:**

* By separating the data plane from the control plane, you get the advantages of a SaaS solution, without its security and privacy challenges.
* With multiple deployment models available, you also get to choose the level of security and privacy your organization needs, up to the highest standards (FedRamp-level).
* Automated deployment, maintenance & resource optimization with our [BYOC](/architecture/byoc) deployment option.

*This concept is unique to groundcover, and takes a while to grasp. Read about our BYOC architecture more in detail in* [*this dedicated section*](/architecture/overview)*.*

{% hint style="info" %}
Learn about groundcover [BYOC](/architecture/byoc) (currently available only on a [paid plan](https://www.groundcover.com/pricing)), which enables you to deploy groundcover's control plane inside your own environment and delegate the entire setup and management of the groundcover platform.
{% endhint %}

## **Disruptive pricing model**

Enabled by our unique BYOC architecture, groundcover's vision is to revolutionize the industry by offering a pricing model that is unheard of anywhere else. Our fully transparent pricing model is based only on the number of nodes being monitored, and the costs of hosting the groundcover backend in your environment. Volume of logs, metrics, traces, and all other observability data don’t affect your cost. This results in savings of 60-90% compared to SaaS platforms.

In addition, all our subscription tiers never limit your access to features and capabilities.

**Advantages of our nodes-based pricing model:**

* Cost is predictable and transparent, becoming an enabler of growth and expansion.
* The ability to deploy groundcover in data-intensive environments enables the monitoring of Dev and Staging clusters, which promotes early identification of issues.
* No cardinality or retention limits

Read our latest customer stories to learn how organization of varying sizes all reduce their observability costs dramatically by migrating to groundcover:

<table data-card-size="large" data-view="cards"><thead><tr><th></th><th></th><th data-hidden data-card-cover data-type="files"></th><th data-hidden data-card-target data-type="content-ref"></th></tr></thead><tbody><tr><td><strong>Nobl9 expands monitoring to cover production e2e, including testing and staging environments</strong></td><td>Replacing Datadog with groundcover cut Nobl9’s observability costs in half while improving log coverage, providing deeper granularity on traces with eBPF, and enabling operational growth and scalability.</td><td><a href="/files/HhgzZ6aa2RkuqwMf231i">/files/HhgzZ6aa2RkuqwMf231i</a></td><td><a href="https://www.groundcover.com/customer-stories/nobl9">https://www.groundcover.com/customer-stories/nobl9</a></td></tr><tr><td><strong>Tracr eliminates blind spots with native-K8s observability and eBPF tracing</strong></td><td>Tracr migrates from a fragmented observability stack to groundcover, gaining deep Kubernetes visibility, automated eBPF tracing, and a cost-effective monitoring solution. This transition streamlined troubleshooting, expanded observability across teams, and enhanced the reliability of their blockchain infrastructure.</td><td><a href="/files/l14pO93bD8XMmwhXxW09">/files/l14pO93bD8XMmwhXxW09</a></td><td><a href="https://www.groundcover.com/customer-stories/tracr">https://www.groundcover.com/customer-stories/tracr</a></td></tr></tbody></table>

#### Stream processing

groundcover applies a stream processing approach to collect and control the continuous flow of data to gain immediate insights, detect anomalies, and respond to changing conditions. Unlike batch processing, where data is collected over a period and then analyzed, stream processing analyzes the data as it flows through the system.

Our platform uses a distributed stream processing engine that enables it to ingest huge amounts of data (such as logs, traces and Kubernetes events) in real time. It also processes all that data and instantly generates complex insights (such as metrics and context) based on it.

As a result, the volume of raw data stored dramatically decreases which, in turn, further reduces the overall cost of observability.

## Capabilities

### Log Management

Designed for high scalability and rapid query performance, enabling quick and efficient log analysis from all your environments. Each log is enriched with actionable context and correlated with relevant metrics and traces, providing a comprehensive view for fast troubleshooting.

[**Learn more** →](/capabilities/log-management)

### Infrastructure Monitoring

The groundcover platform provides cloud-native infrastructure monitoring, enabling automatic collection and real-time monitoring of infrastructure health and efficiency.

[**Learn more →**](/capabilities/infrastructure-monitoring)

### Application Performance Monitoring (APM)

Gain end-to-end observability into your applications performance, identify and resolve issues instantly, all with zero code changes.

[**Learn more** →](/capabilities/application-performance-monitoring-apm)

### Real User Monitoring (RUM)

Real User Monitoring (RUM) extends groundcover’s observability platform to the client side, providing visibility into actual user interactions and front-end performance. It tracks key aspects of your web application as experienced by real users, then correlates them with backend metrics, logs, and traces for a full-stack view of your system.

[**Learn more →**](/capabilities/real-user-monitoring-rum)

<figure><img src="/files/t0ixyBNEQum439aSAMoD" alt=""><figcaption></figcaption></figure>


# FAQ

## How much does groundcover cost?

groundcover's unique pricing model is the first to decouple data volumes from cost of owning and operating the solution. Check out our [pricing plans](https://www.groundcover.com/pricing) for details.

Overall, the cost of owning and operating groundcover is based on two factors:

* The number of nodes (hosts) you are running in the environment you are monitoring
* The costs of hosting groundcover's backend in your environment

Check out our [TCO calculator](https://www.groundcover.com/calculator) to simulate your total cost of ownership for groundcover.

## Can I use groundcover across multiple clusters?

Definitely. As you deploy groundcover each cluster is automatically assigned the unique name it holds inside your cloud environment. You can browse and select all your clusters at one place with our UI experience.

## What K8s flavors are supported?

groundcover has been tested and validated on the most common K8s distributions. See full list in the [Requirements](/getting-started/requirements) section.

## What protocols are supported?

groundcover supports the most common protocols in most K8s production environments out-of-the-box. See full list [here](/integrations/overview#supported-protocols-out-of-the-box).

## What types of data does groundcover collect?

groundcover's kernel-level eBPF sensor automatically collects your logs, application metrics (such as latency, throughput, error rate and much more), infrastructure metrics (such as deployment updates, container crashes etc.), traces, and Kubernetes events. You can control which data is left out of the automatic collection using [data obfuscation](/customization/customize-usage/sensitive-data-obfuscation).

## Where is my data being stored?

groundcover stores all the data it collects inside your environment, using the state-of-the-art storage services of ClickHouse and Victoria Metrics, with the option to offload data to object storage such as S3 for long-term retention. See our [Architecture](/architecture/overview) section for more details.

## Is my data secure?

groundcover stores the data it collects in-cluster, inside your environment without ever leaving the cluster to be stored anywhere else.

Our SaaS UI experience stores only information related to the account, user access and general K8s metadata used for governance (like the number of nodes per cluster, the name given to the cluster etc.).

All the information served to the UI experience is encrypted all the way to the in-cluster data sources. groundcover has no access to your collected data, which is accessible only to an authenticated user from your organization.\
\
groundcover does collect telemetry information (*opt-out is of course possible*) which includes metrics about the performance of the deployment (e.g. resource consumption metrics) and logs reported from the groundcover components running in the cluster.

All telemetry information is anonymized, and contains no data related to your environment.

Regardless, groundcover is **SOC2 and ISO 27001 compliant** and follows best practices.

## How can I invite my team to my workspace?

If you used your business email to create your groundcover account, you can invite your team to your workspace by clicking on the purple "Invite" button on the upper menu. This will open a pop-up where you can enter the emails of the people you want to invite. You also have an option to copy and share your private link.

**Note:** The Admin of the account (i.e. the person that created it) can also invite users outside of your email domain. Non-admin users can only invite users that share the same email domain.\
\
If you used a private email, you can only share the link to your workspace by clicking the "Share" button on the top bar.

Read more about invites in our [quick start guide](/getting-started/5-quick-steps-to-get-you-started).

## Is groundcover open source?

groundcover's [CLI tool](https://github.com/groundcover-com/cli) is currently Open Source along side more projects like [Murre](https://github.com/groundcover-com/murre) and [Caretta](https://github.com/groundcover-com/caretta). We're working on releasing more parts of our solution to Open Source very soon. Stay tuned in our [GitHub](https://github.com/groundcover-com) page!

## What operating system (OS) do I need to use groundcover?

groundcover’s sensor uses [eBPF](https://www.groundcover.com/ebpf), which means it can only deployed on a Kubernetes cluster that is running on a Linux system.

Installing using the CLI command is currently only supported on Linux and Mac.

You can install using the Helm command from any operating system.

Once installed, accessing the groundcover platform is possible from any web browser, on any operating system.

## Is Agent Mode available for on-premises deployments?

Yes. On AWS, GCP, or Azure, groundcover can provision the native cloud LLM service: AWS Bedrock, Google Cloud Vertex AI, or Microsoft Foundry. You can also use your own Anthropic API key or a custom Anthropic-compatible proxy on any supported backend. Native cloud services require network access to the selected service, which can use private connectivity. The direct Anthropic API requires outbound internet access. A custom proxy can remain inside your network, allowing Agent Mode in an air-gapped environment when your groundcover deployment can access the proxy and the proxy can access its model backend without internet egress. See [Requirements & Compatibility](/use-groundcover/agent-mode/requirements) for details.

## Can Agent Mode use LiteLLM, Bifrost, or another Anthropic-compatible proxy?

Yes. Configure **Custom Anthropic Domain** for the backend. The proxy must support streaming Anthropic Messages API requests and API key authentication.

Enter the proxy's full base URL, including any path it requires. For Bifrost, use a URL ending in `/anthropic`. The `/v1/messages` path is added automatically. Use HTTPS for external endpoints. Use HTTP only for internal endpoints protected by authenticated encryption, such as a TLS service mesh. Provide the URL without query parameters or fragments. See [Configure a custom Anthropic domain](/use-groundcover/agent-mode/configuring-settings#configure-a-custom-anthropic-domain) for detailed steps.

## How do I enable or disable Agent Mode?

Admins can toggle Agent Mode on or off per backend in **Settings > AI & Agents > Management**. See [Configuring Settings](/use-groundcover/agent-mode/configuring-settings) for instructions.

## Does Agent Mode have usage limits?

Admins can set monthly spending limits at the organization and per-user level. See [Cost Management](/use-groundcover/agent-mode/cost-management) for details.


# Log Management

Stream, store, and query your logs at any scale, for a fixed cost.

## Overview

Our Log Management solution is built for high scale and fast query performance so you can analyze logs quickly and effectively from all your cloud environments.

**Gain context** - Each log data is enriched with actionable context and correlated with relevant metrics and traces in one single view so you can find what you’re looking for and troubleshoot, faster.

**Centralize to maximize -** The groundcover platform can act as a limitless, centralized log management hub. Your [subscription costs](https://www.groundcover.com/pricing) are completely unaffected by the amount of logs you choose to store or query. It's entirely up to you to decide.

## Collection

### Seamless log collection

groundcover ensures a seamless log collection experience with our [proprietary eBPF sensor](/getting-started/requirements/kernel-requirements-for-ebpf-sensor), which automatically collects and aggregates all logs in all formats - including JSON, plain text, NGINX logs, and more. All this without any configuration needed.

This sensor is deployed as a DaemonSet, running a single pod on each node within your Kubernetes cluster. This configuration enables the groundcover platform to automatically collect logs from all of your pods, across all namespaces in your cluster. This means that once you've installed groundcover, no further action is needed on your part for log collection. The logs collected by each sensor instance are then channeled to the `OTel Collector`.

### OTel Collector: A vendor-agnostic way to receive, process and export telemetry data.

Acting as the central processing hub, the `OTel Collector` is a vendor-agnostic tool that receives logs from various `sensor` pods. It processes, enriches, and forwards the data into groundcover's `ClickHouse database`, where all log data from your cluster is [securely stored](/architecture/security-considerations).

### Logs Attributes

Logs `Attributes` enable advanced filtering capabilities and is currently supported for the formats:

* JSON
* Common Log Format (CLF) - like those from NGNIX and Kong
* logfmt

groundcover automatically detects the format of these logs, extracting key:value pairs from the original log records as `Attributes`.

Each attribute can be added to your filters and search queries.

Example: filtering a log in a supported format with a field of a request path "/status" will look as follows: `@request.path:"/status"`. Syntax can be found [here](#search-and-filter).

## Configuration

groundcover offers the flexibility to craft tailored collection filtering rules, you can choose to set up filters and collect only the logs that are essential for your analysis, avoiding unnecessary data noise. For guidance on configuring your filters, explore our [Customize Logs Collection](/customization/customize-usage/custom-logs-collection) section.

You also have the option to [define the retention period](/customization/customize-usage/custom-data-retention) for your logs in the ClickHouse database. By default, logs are retained for 3 days. To adjust this period to your preferences, visit our [Customize Retention](/customization/customize-usage/custom-data-retention) section for instructions.

## Log Explorer

Once logs are collected and ingested, they are available within the groundcover platform in the Log Explorer, which is designed for quick searches and seamless exploration of your logs data. Using the Log Explorer you can troubleshoot and explore your logs with advanced search capabilities and filters, all within a clear and fast interface.

<figure><img src="/files/JEDCt1Rb7gvuebCREpWL" alt=""><figcaption></figcaption></figure>

### Search and filter

The Log Explorer integrates dynamic filters and a versatile search functionality that enables you to quickly and easily identify the right data. You can filter out logs by selecting one or multiple criteria, including log-level, workload, namespace and more, and can limit your search to a specific time range.

[Learn more about how to use our search syntaxes](/use-groundcover/search-and-filter)

### Log Pipelines

groundcover natively supports setting up log pipelines using OTTL transforms. This allows full flexibility in the processing and manipulation of collected logs — including parsing additional patterns with regex, renaming attributes, and much more.

[Learn more about how to configure log pipelines](https://docs.groundcover.com/use-groundcover/data-pipelines/log-pipelines)


# Infrastructure Monitoring

Get complete visibility into your cloud infrastructure performance at any scale, easily access all your metrics in one place and optimize infrastructure efficiency.

## Overview

The groundcover platform offers infrastructure monitoring capabilities that were built for cloud-native environments. It enables you to track the health and efficiency of your infrastructure instantly, with an effortless deployment process.

**Troubleshoot efficiently -** acting as a centralized hub for all your infrastructure, application and customer metrics allows you to query, correlate and troubleshoot your cloud environments using real time data and insight on your entire stack.

**Store it all, without a sweat -** store any metrics volume without worrying about cardinality or retention limits. [Your subscription costs](https://www.groundcover.com/pricing) remain unaffected by the granularity of metrics you store or query.

## Collection

groundcover's proprietary eBPF sensor leverages all its innovative powers to collect comprehensive data across your cloud environments without the burden of performance overhead. This data is sourced from various Kubernetes components, including kube-system workloads, cluster information via the Kubernetes API, and the applications' interactions with the Kubernetes infrastructure. This level of detailed collection at the kernel level enables the ability to provide actionable insights into the health of your Kubernetes clusters, which are indispensable for troubleshooting existing issues and taking proactive steps to future-proof your cloud environments.

## Configuration

You also have the option to [define the retention period](/customization/customize-usage/custom-data-retention) for your metrics in the VictoriaMetrics database. By default, logs are retained for 7 days, but you can adjust this period to your preferences.

## Enrichment

Beyond collecting data, groundcover's methodology involves a strategic layer of data enrichment that seeks to correlate Kubernetes metrics with application performance indicators. This correlation is crucial for creating a transparent image of the Kubernetes ecosystem. It enables a deep understanding of how Kubernetes interacts with applications, identifying [potential points of failure](broken://pages/Vrq2aaqMQwh101oRKyeZ) across the interconnected environment. By monitoring Kubernetes not as an isolated platform but as an integral part of the application infrastructure, groundcover ensures that the monitoring strategy aligns with your dynamic and complex cloud operations.

<figure><img src="/files/XSpdYDOY4w9XcUoTPXiW" alt=""><figcaption></figcaption></figure>

## Infrastructure Metrics

Monitoring a cluster involves tracking resources that are critical to the performance and stability of the entire system. Monitoring these essential metrics is crucial for maintaining a healthy Kubernetes cluster:

* **CPU consumption**: It's essential to track the CPU resources being utilized against the total capacity to prevent workloads from failing due to insufficient CPU availability.
* **Memory utilization**: Keeping an eye on the remaining memory resources ensures that your cluster doesn't encounter disruptions due to memory shortages.
* **Disk space allocation**: For Kubernetes clusters running stateful applications or requiring persistent storage for data, such as etcd databases, tracking the available disk space is crucial to avert potential storage deficiencies.
* **Network usage:** Visualize traffic rates and connections being established and closed on a service-to-service level of granularity, and easily pinpoint cross availability zone communication to investigate misconfigurations and surging costs.

### Container CPU and Memory

<mark style="color:green;background-color:yellow;">`Available Labels`</mark>

**`type`**

**`clusterId`** **`region`** **`namespace`** **`node_name`** **`workload_name`**

**`pod_name`** **`container_name`** **`container_image`**

<mark style="color:green;background-color:yellow;">`Available Metrics`</mark>

<table><thead><tr><th width="397.3333333333333">Name</th><th>Description</th><th>Type<select><option value="c476a927c24d4e67b8115feb2e039751" label="Counter" color="blue"></option><option value="655b3932f5bd49b0aacc43978ad78425" label="Gauge" color="blue"></option><option value="c0106f6abc0244adbf70f29447aa3875" label="Summary" color="blue"></option></select></th></tr></thead><tbody><tr><td>groundcover_container_cpu_usage_rate_millis</td><td>CPU usage in mCPU</td><td><span data-option="655b3932f5bd49b0aacc43978ad78425">Gauge</span></td></tr><tr><td>groundcover_container_cpu_request_m_cpu</td><td>K8s container CPU request (mCPU)</td><td><span data-option="655b3932f5bd49b0aacc43978ad78425">Gauge</span></td></tr><tr><td>groundcover_container_cpu_limit_m_cpu</td><td>K8s container CPU limit (mCPU)</td><td><span data-option="655b3932f5bd49b0aacc43978ad78425">Gauge</span></td></tr><tr><td>groundcover_container_memory_working_set_bytes</td><td>current memory working set (B)</td><td><span data-option="655b3932f5bd49b0aacc43978ad78425">Gauge</span></td></tr><tr><td>groundcover_container_memory_rss_bytes</td><td>current memory RSS (B)</td><td><span data-option="655b3932f5bd49b0aacc43978ad78425">Gauge</span></td></tr><tr><td>groundcover_container_memory_request_bytes</td><td>K8s container memory request (B)</td><td><span data-option="655b3932f5bd49b0aacc43978ad78425">Gauge</span></td></tr><tr><td>groundcover_container_memory_limit_bytes</td><td>K8s container memory limit (B)</td><td><span data-option="655b3932f5bd49b0aacc43978ad78425">Gauge</span></td></tr><tr><td>groundcover_container_cpu_delay_seconds</td><td>K8s container CPU delay accounting in seconds</td><td><span data-option="c476a927c24d4e67b8115feb2e039751">Counter</span></td></tr><tr><td>groundcover_container_disk_delay_seconds</td><td>K8s container disk delay accounting in seconds</td><td><span data-option="c476a927c24d4e67b8115feb2e039751">Counter</span></td></tr><tr><td>groundcover_container_cpu_throttled_seconds_total</td><td>K8s container total CPU throttling in seconds</td><td><span data-option="c476a927c24d4e67b8115feb2e039751">Counter</span></td></tr></tbody></table>

### Node CPU, Memory and Disk

<mark style="color:green;background-color:yellow;">`Available Labels`</mark>

**`type`** **`clusterId`** **`region`** **`node_name`**

<mark style="color:green;background-color:yellow;">`Available Metrics`</mark>

<table><thead><tr><th width="397.3333333333333">Name</th><th>Description</th><th>Type<select><option value="c476a927c24d4e67b8115feb2e039751" label="Counter" color="blue"></option><option value="655b3932f5bd49b0aacc43978ad78425" label="Gauge" color="blue"></option><option value="c0106f6abc0244adbf70f29447aa3875" label="Summary" color="blue"></option></select></th></tr></thead><tbody><tr><td>groundcover_node_allocatable_cpum_cpu</td><td>amount of allocatable CPU in the current node (mCPU)</td><td><span data-option="655b3932f5bd49b0aacc43978ad78425">Gauge</span></td></tr><tr><td>groundcover_node_allocatable_mem_bytes</td><td>amount of allocatable memory in the current node (B)</td><td><span data-option="655b3932f5bd49b0aacc43978ad78425">Gauge</span></td></tr><tr><td>groundcover_node_mem_used_percent</td><td>percent of used memory in current node (0-100)</td><td><span data-option="655b3932f5bd49b0aacc43978ad78425">Gauge</span></td></tr><tr><td>groundcover_node_used_disk_space</td><td>current used disk space in current node (B)</td><td><span data-option="655b3932f5bd49b0aacc43978ad78425">Gauge</span></td></tr><tr><td>groundcover_node_free_disk_space</td><td>amount of free disk space in current node (B)</td><td><span data-option="655b3932f5bd49b0aacc43978ad78425">Gauge</span></td></tr><tr><td>groundcover_node_total_disk_space</td><td>amount of total disk space in current node (B)</td><td><span data-option="655b3932f5bd49b0aacc43978ad78425">Gauge</span></td></tr><tr><td>groundcover_node_used_percent_disk_space</td><td>percent of used disk space in current node (0-100)</td><td><span data-option="655b3932f5bd49b0aacc43978ad78425">Gauge</span></td></tr></tbody></table>

### PVC Usage

<mark style="color:green;background-color:yellow;">`Available Labels`</mark>

**`type`** **`clusterId`** **`region`** **`name`** **`namespace`**

<mark style="color:green;background-color:yellow;">`Available Metrics`</mark>

<table><thead><tr><th width="397.3333333333333">Name</th><th>Description</th><th>Type<select><option value="c476a927c24d4e67b8115feb2e039751" label="Counter" color="blue"></option><option value="655b3932f5bd49b0aacc43978ad78425" label="Gauge" color="blue"></option><option value="c0106f6abc0244adbf70f29447aa3875" label="Summary" color="blue"></option></select></th></tr></thead><tbody><tr><td>groundcover_pvc_usage_bytes</td><td>PVC used bytes (B)</td><td><span data-option="655b3932f5bd49b0aacc43978ad78425">Gauge</span></td></tr><tr><td>groundcover_pvc_capacity_bytes</td><td>PVC capacity bytes (B)</td><td><span data-option="655b3932f5bd49b0aacc43978ad78425">Gauge</span></td></tr><tr><td>groundcover_pvc_available_bytes</td><td>PVC available bytes (B)</td><td><span data-option="655b3932f5bd49b0aacc43978ad78425">Gauge</span></td></tr><tr><td>groundcover_pvc_usage_percent</td><td>percent of used pvc storage (0-100)</td><td><span data-option="655b3932f5bd49b0aacc43978ad78425">Gauge</span></td></tr></tbody></table>

### Network Usage

<mark style="color:green;background-color:yellow;">`Available Labels`</mark>

**`clusterId workload_name`** **`namespace`** **`container_name`** **`remote_service_name`** **`remote_namespace`** **`remote_is_external`** **`availability_zone`** **`region`** **`remote_availability_zone`** **`remote_region`** **`is_cross_az`** **`protocol`** **`role`** **`server_port`** **`encryption`** **`transport_protocol`** **`is_loopback`**

Note&#x73;**:**

* `is_loopback` and `remote_is_external` are special labels that indicate the remote service is either the same service as the recording side (loopback) or resides in an external network, e.g managed service outside of the cluster (external).
  * In both cases the `remote_service_name` and the `remote_namespace` labels will be empty
* `is_cross_az` means the traffic was sent and/or received between two different availability zones. This is a helpful flag to quickly identify this special kind of communication.
  * The actual zones are detailed in the `availability_zone` and `remote_availability_zone` labels

<mark style="color:green;background-color:yellow;">`Available Metrics`</mark>

<table><thead><tr><th width="417.3333333333333">Name</th><th>Description</th><th>Type<select><option value="c476a927c24d4e67b8115feb2e039751" label="Counter" color="blue"></option><option value="655b3932f5bd49b0aacc43978ad78425" label="Gauge" color="blue"></option><option value="c0106f6abc0244adbf70f29447aa3875" label="Summary" color="blue"></option></select></th></tr></thead><tbody><tr><td>groundcover_network_rx_bytes_total</td><td>Bytes received by the workload (B)</td><td><span data-option="c476a927c24d4e67b8115feb2e039751">Counter</span></td></tr><tr><td>groundcover_network_tx_bytes_total</td><td>Bytes sent by the workload (B)</td><td><span data-option="c476a927c24d4e67b8115feb2e039751">Counter</span></td></tr><tr><td>groundcover_network_connections_opened_total</td><td>Connections opened by the workload</td><td><span data-option="c476a927c24d4e67b8115feb2e039751">Counter</span></td></tr><tr><td>groundcover_network_connections_closed_total</td><td>Connections closed by the workload</td><td><span data-option="c476a927c24d4e67b8115feb2e039751">Counter</span></td></tr><tr><td>groundcover_network_connections_opened_failed_total</td><td>Connections attempts failed per workload (including refused connections)</td><td><span data-option="c476a927c24d4e67b8115feb2e039751">Counter</span></td></tr><tr><td>groundcover_network_connections_opened_refused_total</td><td>Connections attempts refused per workload</td><td><span data-option="c476a927c24d4e67b8115feb2e039751">Counter</span></td></tr></tbody></table>


# Application Performance Monitoring (APM)

Gain end-to-end observability into your applications performance, identify and resolve issues instantly - all with zero code changes.

### Overview

The groundcover platform collects data all across your stack using the power of eBPF instrumentation. Our [proprietary eBPF sensor](/getting-started/requirements/kernel-requirements-for-ebpf-sensor) is installed in seconds and provides 100% coverage into application metrics and traces with zero code changes or configurations.

**Resolve faster -** By seamlessly correlating traces with application metrics, logs, and infrastructure events, groundcover’s APM enables you to detect and resolve root issues faster.

**Improve user experience -** Optimize your application performance and resource utilization faster than ever before, avoid downtimes and make poor end-user experience a thing of the past.

### Collection

Our revolutionary eBPF sensor, [Flora](https://www.groundcover.com/blog/ebpf-observability-agent), is deployed as a DaemonSet in your Kubernetes cluster. This approach allows us to inspect every packet that each service is sending or receiving, achieving 100% coverage. No sampling rates or relying on statistical luck - all requests and responses are observed.

This approach would not be feasible without a resource-efficient eBPF-powered sensor. eBPF not only extends the ability to pinpoint issues - it does so with much less overhead than any other method. eBPF can be used to analyze traffic originating from every programming language and SDK - even for encrypted connections!

{% hint style="info" %}
Click [here](/capabilities/application-performance-monitoring-apm/supported-technologies) for a full list of supported technologies
{% endhint %}

### Reconstruction

After being collected by our eBPF code, the traffic is then classified according to its protocol - which is identified directly from the underlying traffic, or the library from which it originated. Connections are reconstructed, and we can generate transactions - HTTP requests and responses, SQL queries and responses etc.

### Enrichment

In order to give as much context as possible each transaction is enriched with as much metadata as possible. Some examples might include the pods that took part in this transaction (both client and server), the nodes on which these pods are scheduled, and the state of container at the time of the request.

It is important to emphasize the impressive granularity level with which this process takes place - every single transaction observed is fully enriched. This allows us to perform more advanced aggregations.

### Aggregation

After being enriched by as much context as possible, the transactions as grouped together into meaningful aggregations. These could be defined by the workloads involved, the protocols detected and the resources that were accessed in the operations. These aggregations will mostly come into play when displaying [golden signals](/capabilities/application-performance-monitoring-apm/application-metrics#golden-signals).

### Exporting

After collecting the data, contextualizing it and putting it together in meaningful aggregations - we can now create [metrics](/capabilities/application-performance-monitoring-apm/application-metrics) and [traces](/capabilities/application-performance-monitoring-apm/traces) to provide meaningful insights into the services' behaviors.

### Metrics

Learn how groundcover's application metrics work:

{% content-ref url="/pages/889fayVViR0TfQtx9mUG" %}
[Application Metrics](/capabilities/application-performance-monitoring-apm/application-metrics)
{% endcontent-ref %}

### Traces

Learn how groundcover's application traces work:

{% content-ref url="/pages/AWaEze0cqLjCPV0W8Jy4" %}
[Traces](/capabilities/application-performance-monitoring-apm/traces)
{% endcontent-ref %}


# Application Metrics

## Our metrics philosophy

The groundcover platform generates 100% of its metrics from the actual data. There are no sample rates or complex interpolations to make up for partial coverage. Our measurements represent the real, complete flow of data in your environment.

[Stream processing](/#stream-data-processing) allows us to construct the majority of the metrics on the very node where the raw transactions are recorded. This means the raw data is turned into numbers the moment it becomes possible - removing the need for storing or sending it elsewhere.

Metrics are stored in groundcover's `victoria-metrics` deployment, ensuring top-notch performance on every scale.

### Golden signals

In the world of excessive data, it's important to have a rule of thumb for knowing where to start looking. For application metrics, we rely on our [golden signals](https://www.groundcover.com/blog/monitor-the-four-golden-signals).

The following metrics are generated for each resource being [aggregated](/capabilities/application-performance-monitoring-apm#aggregation):

* Requests per second (RPS)
* Errors rate
* Latencies (p50 and p95)

The golden signals are then displayed in two important ways: **Workload** and **Resource** aggregations.

{% hint style="info" %}
See [below](#golden-signals-metrics) for the full list of generated workload and resource golden metrics.
{% endhint %}

**Resource** aggregations are highly granularity metrics, providing insights into individual APIs.

<figure><img src="/files/6rYhux7vwrx7SQLDs0Bm" alt=""><figcaption></figcaption></figure>

**Workload** aggregations are designed to show an overview of each service, enabling a higher level inspection. These are constructed using all of the resources recorded for each service.

<figure><img src="/files/NQUGx37u1IkZRr5pgVAS" alt=""><figcaption></figcaption></figure>

### Controlling retention

groundcover allows full control over the retention of your metrics. Learn more [here](/customization/customize-usage/custom-data-retention).

### List of available metrics

Below you will find the full list of our APM metrics, as well as the labels we export for each. These labels are designed with high granularity in mind for maximal insight depth. All of the metrics listed are available out of the box after installing groundcover, without any further setup.

{% hint style="info" %}
We fully support the ingestion of [custom metrics](/integrations/data-sources/prometheus) to further expand the visibility into your environment.

We also allow for building [custom dashboards](/use-groundcover/dashboards-and-alerts), enabling full freedom in deciding how to display your metrics - building on groundcover's metrics below plus every custom metric ingested.
{% endhint %}

### Our labels

<table><thead><tr><th width="256">Label name</th><th width="266">Description</th><th>Relevant types</th></tr></thead><tbody><tr><td>clusterId</td><td>Name identifier of the K8s cluster</td><td></td></tr><tr><td>region</td><td>Cloud provider region name</td><td></td></tr><tr><td>namespace</td><td>K8s namespace</td><td></td></tr><tr><td>workload_name</td><td>K8s workload (or service) name</td><td></td></tr><tr><td>pod_name</td><td>K8s pod name</td><td></td></tr><tr><td>container_name</td><td>K8s container name</td><td></td></tr><tr><td>container_image</td><td>K8s container image name</td><td></td></tr><tr><td>remote_namespace</td><td>Remote K8s namespace (other side of the communication)</td><td></td></tr><tr><td>remote_service_name</td><td>Remote K8s service name (other side of the communication)</td><td></td></tr><tr><td>remote_container_name</td><td>Remote K8s container name (other side of the communication)</td><td></td></tr><tr><td>type</td><td>The protocol in use (HTTP, gRPC, Kafka, DNS etc.)</td><td></td></tr><tr><td>role</td><td>Role in the communication (client or server)</td><td></td></tr><tr><td>clustered_path</td><td>HTTP / gRPC aggregated resource path (e.g. /metrics/*)</td><td>http, grpc</td></tr><tr><td>method</td><td>HTTP / gRPC method (e.g GET)</td><td>http, grpc</td></tr><tr><td>response_status_code</td><td>Return status code of a HTTP / gPRC request (e.g. 200 in HTTP)</td><td>http, grpc</td></tr><tr><td>dialect</td><td>SQL dialect (MySQL or PostgreSQL)</td><td>mysql, postgresql</td></tr><tr><td>response_status</td><td>Return status code of a SQL query (e.g 42P01 for undefined table)</td><td>mysql, postgresql</td></tr><tr><td>client_type</td><td>Kafka client type (Fetcher / Producer)</td><td>kafka</td></tr><tr><td>topic</td><td>Kafka topic name</td><td>kafka</td></tr><tr><td>partition</td><td>Kafka partition identifier</td><td>kafka</td></tr><tr><td>error_code</td><td>Kafka return status code</td><td>kafka</td></tr><tr><td>query_type</td><td>type of DNS query (e.g. AAAA)</td><td>dns</td></tr><tr><td>response_return_code</td><td>Return status code of a DNS resolution request (e.g. Name Error)</td><td>dns</td></tr><tr><td>method_name, method_class_name</td><td>Method code for the operation</td><td>amqp</td></tr><tr><td>response_method_name, response_method_class_name</td><td>Method code for the operation's response</td><td>amqp</td></tr><tr><td>exit_code</td><td>K8s container termination exit code</td><td>container_state, container_crash</td></tr><tr><td>state</td><td>K8s container current state (Running, Waiting or Terminated)</td><td>container_state</td></tr><tr><td>state_reason</td><td>K8s container state transition reason (e.g CrashLoopBackOff or OOMKilled)</td><td>container_state</td></tr><tr><td>crash_reason</td><td>K8s container crash reason (e.g Error, OOMKilled)</td><td>container_crash</td></tr><tr><td>pvc_name</td><td>K8s PVC name</td><td>storage</td></tr></tbody></table>

{% hint style="info" %}
Summary based metrics have an additional ***quantile*** label, representing the percentile. Available values: \[`”0.5”, “0.95”, 0.99”`].
{% endhint %}

{% hint style="warning" %}
[**groundcover**](https://www.groundcover.com/) uses a set of internal labels which are not relevant in most use-cases. Find them interesting? [Let us know over Slack!](https://www.groundcover.com/join-slack)

**`issue_id`** **`entity_id`** **`resource_id`** **`query_id`** **`aggregation_id`** **`parent_entity_id`** **`perspective_entity_id`** **`perspective_entity_is_external`** **`perspective_entity_issue_id`** **`perspective_entity_name`** **`perspective_entity_namespace`** **`perspective_entity_resource_id`**
{% endhint %}

### Golden Signals Metrics

{% hint style="info" %}
In the lists below, we describe **error** and **issue** counters. Every issue flagged by groundcover is an error; but not every error is flagged as an issue.
{% endhint %}

#### Resource metrics

<table><thead><tr><th width="357.3333333333333">Name</th><th width="258">Description</th><th>Type<select><option value="c476a927c24d4e67b8115feb2e039751" label="Counter" color="blue"></option><option value="655b3932f5bd49b0aacc43978ad78425" label="Gauge" color="blue"></option><option value="c0106f6abc0244adbf70f29447aa3875" label="Summary" color="blue"></option></select></th></tr></thead><tbody><tr><td>groundcover_resource_total_counter</td><td>total amount of resource requests</td><td><span data-option="c476a927c24d4e67b8115feb2e039751">Counter</span></td></tr><tr><td>groundcover_resource_error_counter</td><td>total amount of requests with error status codes</td><td><span data-option="c476a927c24d4e67b8115feb2e039751">Counter</span></td></tr><tr><td>groundcover_resource_issue_counter</td><td>total amount of requests which were flagged as issues</td><td><span data-option="c476a927c24d4e67b8115feb2e039751">Counter</span></td></tr><tr><td>groundcover_resource_success_counter</td><td>total amount of resource requests with OK status codes</td><td><span data-option="c476a927c24d4e67b8115feb2e039751">Counter</span></td></tr><tr><td>groundcover_resource_latency_seconds</td><td>resource latency [sec]</td><td><span data-option="c0106f6abc0244adbf70f29447aa3875">Summary</span></td></tr></tbody></table>

#### Workload metrics

<table><thead><tr><th width="358.3333333333333">Name</th><th width="258">Description</th><th>Type<select><option value="c476a927c24d4e67b8115feb2e039751" label="Counter" color="blue"></option><option value="655b3932f5bd49b0aacc43978ad78425" label="Gauge" color="blue"></option><option value="c0106f6abc0244adbf70f29447aa3875" label="Summary" color="blue"></option></select></th></tr></thead><tbody><tr><td>groundcover_workload_total_counter</td><td>total amount of requests handled by the workload</td><td><span data-option="c476a927c24d4e67b8115feb2e039751">Counter</span></td></tr><tr><td>groundcover_workload_error_counter</td><td>total amount of requests handled by the workload with error status codes</td><td><span data-option="c476a927c24d4e67b8115feb2e039751">Counter</span></td></tr><tr><td>groundcover_workload_issue_counter</td><td>total amount of requests handled by the workload which were flagged as issues</td><td><span data-option="c476a927c24d4e67b8115feb2e039751">Counter</span></td></tr><tr><td>groundcover_workload_success_counter</td><td>total amount of requests handled by the workload with OK status codes</td><td><span data-option="c476a927c24d4e67b8115feb2e039751">Counter</span></td></tr><tr><td>groundcover_workload_latency_seconds</td><td>resource latency across all of the workload APIs [sec]</td><td><span data-option="c0106f6abc0244adbf70f29447aa3875">Summary</span></td></tr></tbody></table>

### Storage usage metrics

<table><thead><tr><th width="329">Name</th><th width="285">Description</th><th>Type<select><option value="579160c3a58647ae94a346a3c517890a" label="Counter" color="blue"></option><option value="5a23e1b2296a4288be65fa1bc492c829" label="Gauge" color="blue"></option><option value="26c560ac591747cdb4f95ed12ed73b2c" label="Summary" color="blue"></option></select></th></tr></thead><tbody><tr><td>groundcover_pvc_read_bytes_total</td><td>total amount of bytes read by the workload from the PVC</td><td><span data-option="579160c3a58647ae94a346a3c517890a">Counter</span></td></tr><tr><td>groundcover_pvc_write_bytes_total</td><td>total amount of bytes written by the workload to the PVC</td><td><span data-option="579160c3a58647ae94a346a3c517890a">Counter</span></td></tr><tr><td>groundcover_pvc_reads_total</td><td>total amount of read operations done by the workload from the PVC</td><td><span data-option="579160c3a58647ae94a346a3c517890a">Counter</span></td></tr><tr><td>groundcover_pvc_writes_total</td><td>total amount of write operations done by the workload to the PVC</td><td><span data-option="579160c3a58647ae94a346a3c517890a">Counter</span></td></tr><tr><td>groundcover_pvc_read_latency</td><td>latency of read operation by the workload from the PVC, in microseconds</td><td><span data-option="26c560ac591747cdb4f95ed12ed73b2c">Summary</span></td></tr><tr><td>groundcover_pvc_write_latency</td><td>latency of write operation by the workload to the PVC, in microseconds</td><td><span data-option="26c560ac591747cdb4f95ed12ed73b2c">Summary</span></td></tr></tbody></table>

### Kafka specific metrics

<table><thead><tr><th width="302.3333333333333">Name</th><th width="258">Description</th><th>Type<select><option value="c476a927c24d4e67b8115feb2e039751" label="Counter" color="blue"></option><option value="655b3932f5bd49b0aacc43978ad78425" label="Gauge" color="blue"></option><option value="c0106f6abc0244adbf70f29447aa3875" label="Summary" color="blue"></option></select></th></tr></thead><tbody><tr><td>groundcover_client_offset</td><td>client last message offset (for producer the last offset produced, for consumer the last requested offset)</td><td><span data-option="655b3932f5bd49b0aacc43978ad78425">Gauge</span></td></tr><tr><td>groundcover_workload_client_offset</td><td>client last message offset (for producer the last offset produced, for consumer the last requested offset), aggregated by workload</td><td><span data-option="655b3932f5bd49b0aacc43978ad78425">Gauge</span></td></tr><tr><td>groundcover_calc_lagged_messages</td><td>current lag in messages</td><td><span data-option="655b3932f5bd49b0aacc43978ad78425">Gauge</span></td></tr><tr><td>groundcover_workload_calc_lagged_messages</td><td>current lag in messages, aggregated by workload</td><td><span data-option="655b3932f5bd49b0aacc43978ad78425">Gauge</span></td></tr><tr><td>groundcover_calc_lag_seconds</td><td>current lag in time [sec]</td><td><span data-option="655b3932f5bd49b0aacc43978ad78425">Gauge</span></td></tr><tr><td>groundcover_workload_calc_lag_seconds</td><td>current lag in time, aggregated by workload [sec]</td><td><span data-option="655b3932f5bd49b0aacc43978ad78425">Gauge</span></td></tr></tbody></table>


# Traces

## Our traces philosophy

Traces are a powerful [observability pillar](https://www.groundcover.com/blog/kubernetes-observability), providing granular insights into microservice interactions. Traditionally, they were hard to implement, requiring coordination of multiple teams and constant code changes, making this critical aspect very challenging to maintain.

groundcover's eBPF sensor disrupts the famous tradeoff, empowering developers to gain full visibility into their applications, effortlessly and without any code changes.

<figure><img src="/files/aTh045Tb2eT78yFQ1O7L" alt=""><figcaption></figcaption></figure>

The platform supports two kinds of traces:

### **eBPF traces**

These traces are automatically generated for every [supported](/capabilities/application-performance-monitoring-apm/supported-technologies) service in your stack. They are available out-of-the-box and within seconds of installation. These traces always include critical information such as:

* All services that took part in the interaction (both client and server)
* Accessed resource
* Full payloads, including:
  * All headers
  * All query parameters
  * All bodies - for both the request and response

### **3rd-party traces**

These can be ingested into the platform, allowing to leverage already existing instrumentation to create a single pane of glass for all of your traces.

Traces are stored in groundcover's `ClickHouse` deployment, ensuring top notch performance on every scale.

{% hint style="info" %}
For more details about ingesting 3rd party traces, see the [datasources page](/integrations/data-sources).
{% endhint %}

<figure><img src="/files/nKoOt3K1sQdsR3me1nQ6" alt=""><figcaption></figcaption></figure>

## Sampling

groundcover further disrupts the customary traces experience by reinventing the concept of sampling. This innovation differs between the different types of traces:

### **eBPF traces**

These are generated by using 100% of the data, always processing every request being made, on every scale. However, the groundcover platform utilizes **smart sampling** to only store a fraction of the traces, while still generating an accurate picture. In general, sampling is performed according to these rules:

* Requests with unusually high or low latencies, measured per resource
* Requests which returned an error response (e.g 500 status code for HTTP)
* "Normal" requests which form the baseline for each resource

Lastly, [stream processing](/#stream-data-processing) is utilized to make the sampling decisions on the node itself, without having to send or save any redundant traces.

{% hint style="info" %}
Certain aspects of our sampling algorithm are configurable - read more [here](/customization/customize-usage/controlling-the-ebpf-sampling-mechanism).
{% endhint %}

### **3rd-party** **traces**

Various mechanisms control the sampling performed over 3rd party traces. Read more here:

* [OpenTelemetry](https://opentelemetry.io/docs/concepts/signals/traces/)
* [Datadog](https://docs.groundcover.com/integrations/data-sources/datadog/shipping-from-the-datadog-agent)

{% hint style="warning" %}
When integrating 3rd-party traces, it is often wise to configure some sampling mechanism according to the specific use case.
{% endhint %}

## Additional Context

Each trace is enriched with additional information to give as much context as possible for the service which generated the trace. This includes:

* Container information - image, environment variables, pod name
* Logs generated by the service around the time of the trace
* [Golden Signals](/capabilities/application-performance-monitoring-apm/application-metrics#golden-signals) of the resource around the time of the trace
* Kubernetes events relevant to the service
* CPU and Memory utilization of the service and the node it is scheduled on

## Distributed Tracing

One of the advantages of ingesting 3rd-party traces is the ability to leverage their distributed tracing feature. groundcover natively displays the full trace for ingested traces in the `Traces` page.

<figure><img src="/files/cQLE9SvLUUxlsbuaGJ21" alt=""><figcaption></figcaption></figure>

## Trace Attributes

Trace `Attributes` enable advanced filtering and search capabilities. groundcover support attributes across all trace types. This encompasses a diverse range of protocols such as HTTP, MongoDB, PostgreSQL, and others, as well as varied sources including eBPF or manual instrumentations (for example - OpenTelemetry).

groundcover enriches your original traces and generates meaningful metadata as key-value pairs. This metadata includes critical information, such as protocol type, `http.path`, `db.statement`, and similar attributes, aligning with OTel conventions. Furthermore, groundcover seamlessly incorporates this metadata from spans received through supported manual instrumentations. For an in-depth understanding of attributes in OTel, please refer to [OTel Attributes Documentation](https://opentelemetry.io/docs/concepts/signals/traces/#attributes) (external link to OpenTelemtry website).

Each attribute can be effortlessly integrated into your filters and search queries. You can add them directly from the trace side-panel with a simple click or input them manually into the search bar.

Example: If you want to filter all HTTP traces that contain the path "/products". The query would be formatted as: `@http.path:"/products"`. For a comprehensive guide on the query syntax, see Syntax table below.

## Trace Tags

Trace `Tags` enable advanced filtering and search capabilities. groundcover support tags across all trace types. This encompasses a diverse range of protocols such as HTTP, MongoDB, PostgreSQL, and others, as well as varied sources including eBPF or manual instrumentations (for example - OpenTelemetry).

Tags are powerful metadata components, structured as key-value pairs. They offer insightful information about the resource generating the span, like: `container.image.name` ,`host.name` and more.

Tags include metadata enriched by the our sensor and additional metadata if provided by manual instrumentations (such as OpenTelemetry traces) . Utilizing these Tags enhances understanding and context of your traces, allowing for more comprehensive analysis and easier filtering by the relevant information.

Each tag can be effortlessly integrated into your filters and search queries. You can add them directly from the trace side-panel with a simple click or input them manually into the search bar.

Example: If you want to filter all traces from mysql containers - The query would be formatted as: `container.image.name:mysql`. For a comprehensive guide on the query syntax, see Syntax table below.

## Search and filter

The Trace Explorer integrates dynamic filters and a versatile search functionality, to enhance your trace data analysis. You can filter out traces using specific criteria, including trace-status, workload, namespace and more, as well as limit your search to a specific time range.

[Learn more about how to use our search syntaxes](/use-groundcover/search-and-filter)

## Traces Pipelines

groundcover natively supports setting up log pipelines using [Vector transforms.](https://vector.dev/docs/reference/configuration/transforms/) This allow for full flexibility in the processing and manipulation of traces being collected - parsing additional patterns by regex, renaming attributes, and many more.

[Learn more about how to configure traces pipelines](broken://pages/cQtjVs3eRtELfe2cHcWa)

## Controlling retention

groundcover allows full control over the retention of your traces. [Read here](/customization/customize-usage/custom-data-retention) to learn more.

## Custom Configuration

Tracing can be customized in several ways:

* [Configuring which protocols should be traced](/customization/customize-usage/disable-tracing-for-specific-protocols)
* [Configuring obfuscation for sensitive payload data](/customization/customize-usage/sensitive-data-obfuscation)
* [Configuring the sampling mechanism](/customization/customize-usage/controlling-the-ebpf-sampling-mechanism)


# Supported Technologies

groundcover will work out-of-the-box on all protocols, encryption libraries and runtimes below - generating [traces](/capabilities/application-performance-monitoring-apm/traces) and [metrics](/capabilities/application-performance-monitoring-apm/application-metrics) with zero code changes.

{% hint style="success" %}
We're growing our coverage all the time.\
Cant find what you're looking for? [let us know over Slack.](https://www.groundcover.com/join-slack)
{% endhint %}

### Supported protocols

<table><thead><tr><th>Protocol</th><th width="259.3333333333333">Status</th><th>Comments</th></tr></thead><tbody><tr><td>HTTP</td><td><code>supported</code></td><td></td></tr><tr><td>gRPC</td><td><code>supported</code></td><td></td></tr><tr><td>MySQL</td><td><code>supported</code></td><td></td></tr><tr><td>PostgreSQL</td><td><code>supported</code></td><td></td></tr><tr><td>Redis</td><td><code>supported</code></td><td></td></tr><tr><td>DNS</td><td><code>supported</code></td><td></td></tr><tr><td>Kafka</td><td><code>supported</code></td><td></td></tr><tr><td>MongoDB</td><td><code>supported</code></td><td>v3.6+</td></tr><tr><td>AMQP</td><td><code>supported</code></td><td>AMQP 0-9-1</td></tr><tr><td>GraphQL</td><td><code>supported</code></td><td></td></tr><tr><td>AWS S3</td><td><code>supported</code></td><td></td></tr><tr><td>AWS SQS</td><td><code>supported</code></td><td></td></tr></tbody></table>

### Supported encryption libraries and runtimes

groundcover seamlessly supports APM for encrypted communication - as long as it's listed below.

<table><thead><tr><th width="278">Encryption Library/Runtime</th><th width="230.33333333333331">Status</th><th>Comments</th></tr></thead><tbody><tr><td>crypto/tls (golang)</td><td><code>supported</code></td><td></td></tr><tr><td>OpenSSL (c, c++, Python)</td><td><code>supported</code></td><td></td></tr><tr><td>NodeJS</td><td><code>supported</code></td><td></td></tr><tr><td>JavaSSL</td><td><code>supported</code></td><td>Java 11+ is supported. Requires <a href="/pages/v0D5SfCVE6pboI3rdu1h">enabling the groundcover Java agent</a></td></tr></tbody></table>

{% hint style="warning" %}
Encryption is unsupported for binaries which have been compiled without debug symbols ("stripped"). Known cases:

* Crossplane
  {% endhint %}


# Real User Monitoring (RUM)

Monitor front-end applications and connect it to your backend — all inside your cloud.

{% hint style="info" %}
This capability is only available in BYOC deployments.
{% endhint %}

Capture real end-user experiences directly from their browsers and unify these insights with your backend observability data.

Follow our [instructions on how to connect RUM](/getting-started/installation-and-updating/connect-rum) guide to add RUM to your platform.

### Overview

Real User Monitoring (RUM) extends groundcover’s observability platform to the client side, providing visibility into actual user interactions and front-end performance. It tracks key aspects of your web application as experienced by real users, then correlates them with backend metrics, logs, and traces for a full-stack view of your system.

Main benefits of using RUM:

* **Understand the user experience** - capture every interaction, page load, and performance metric from the end-user perspective to pinpoint front-end issues in real time. View sessions with session recording and navigate from an event to the relevant timestamp in the replay.
* **Resolve issues faster** - seamlessly tie front-end events to backend traces and logs in one platform, enabling end-to-end troubleshooting of user journeys.
* **Monitor frontend performance** - track your app performance and get alerted before your users report it to you.
* **Privacy first** - groundcover’s Bring Your Own Cloud (BYOC) model ensures all RUM data stays in your own cloud environment. Sensitive user data never leaves your infrastructure, ensuring privacy and compliance without sacrificing insight.
* **AI-powered investigation** - let groundcover’s [Agent](/use-groundcover/agent-mode) explain a session or pinpoint the root cause of a front-end exception, turning raw RUM data into an answer in seconds.

### Collection

groundcover RUM collects a wide range of data from users’ browsers through a lightweight JavaScript SDK. Once integrated into your web application, the SDK automatically gathers and sends the following telemetry from each user session to the groundcover platform:

* **User interactions:** RUM tracks user interactions such as clicks, keydown, and navigation events. By recording which elements users interact with and when, groundcover helps you reconstruct user flows and understand the sequence of actions leading up to any issue or performance problem.
* **Session recording**: groundcover also supports the recording of your users session to visually understand the user behavior that led to specific events. Unlike the other telemetry here, session recording does not start on `init()` — it must be started explicitly via [`startReplayRecording()`](/getting-started/installation-and-updating/connect-rum#session-replay).
* **Network requests:** Every HTTP request initiated by the browser (such as API calls) is captured as a trace. Each client-side request can be linked with its corresponding server-side trace, giving you a complete picture of the request from the user’s click to the backend response.
* **Front-end logs:** Client-side log messages (e.g., `console.log` outputs, warnings, and errors) are collected and forwarded to groundcover’s log management. This ensures that browser logs are stored alongside your application’s server logs for unified analysis.
* **Exceptions:** Uncaught JavaScript exceptions and errors are automatically captured with full stack traces and contextual data (browser type, URL, etc.). These front-end errors can be used to set up groundcover monitors, letting you quickly identify and debug issues in the user’s environment.
* **Performance metrics (Core Web Vitals):** Key performance indicators like page load time and Core Web Vitals like Largest Contentful Paint, Interaction to Next Paint, and Cumulative Layout Shift are measured for each page view. groundcover RUM records these metrics to help you track real-world performance and detect slowdowns affecting users.
* **Custom events:** You can instrument your application to send custom events via the RUM SDK. This allows you to capture domain-specific actions or business events (for example, a checkout completion or a specific UI gesture) with associated metadata, providing deeper insight into user behavior beyond automatic captures.

All collected data is streamed securely to your groundcover deployment. Because groundcover runs in your environment, RUM data (including potentially sensitive details from user sessions) is stored in the observability backend within your own cloud. From there, it is aggregated and indexed just like other telemetry, ready to be searched and analyzed in the groundcover UI.

The SDK is built for minimal overhead: sensitive data is masked by default, and CPU-heavy work such as session-replay packing and payload compression is offloaded to a dedicated Web Worker so your application’s main thread stays responsive. See the [Connect RUM](/getting-started/installation-and-updating/connect-rum) guide for setup, configuration, and privacy controls.

### Full-Stack Visibility

One of the core advantages of groundcover RUM is its native integration with backend observability data. Every front-end trace, log, or event captured via RUM is contextualized alongside server-side data:

* **Trace correlation:** Client-side traces (from browser network requests) are automatically correlated with server-side traces captured via OpenTelemetry (OTel) instrumentation. This means when a user triggers an API call, you can see the complete distributed trace that spans the browser and the backend services in a single, unified view. Similarly, users can navigate from a trace to the matching user session in order to understand the user behavior that led to that specific trace.

<figure><img src="/files/sq7YnnoF4GadjdtV6hd7" alt=""><figcaption></figcaption></figure>

* **Unified logging:** Front-end log entries and error reports are ingested into the same backend as your server-side logs. In the Data Explorer, you can filter and search across logs from both client and server, using common fields (like timestamp, session ID, or trace ID) to connect events.
* **End-to-end troubleshooting:** With full-stack data in one platform, you can pivot easily between a user’s session replay, the front-end events, and the backend metrics/traces involved. This end-to-end context significantly reduces the time to isolate whether an issue originated in the frontend (browser/UI) or the backend (services/infrastructure), helping teams pinpoint problems faster across the entire stack.

By bridging the gap between the user’s browser and your cloud infrastructure, groundcover’s RUM capability ensures that no part of the user journey is invisible to your monitoring. This holistic view is critical for optimizing user experience and rapidly resolving issues that span multiple layers of your application.

### Sessions Explorer

Once RUM data is collected, it becomes available in the groundcover platform via the Sessions Explorer - a dedicated view for inspecting and troubleshooting user sessions. The Sessions Explorer allows you to navigate through user journeys and understand how your users experience your application.

<figure><img src="/files/DSIhf4o2UW1wCu4ysHdh" alt=""><figcaption></figcaption></figure>

Clicking on any session opens the **Session Drawer**, where you can inspect a full timeline of the user’s experience along with a recording if added. This view shows every key event captured during the session - including clicks, navigations, network requests, logs, custom events, and errors.

<figure><img src="/files/pFH17muGKYIETHdw00zv" alt=""><figcaption></figcaption></figure>

Each event is displayed in sequence with full context like timestamps, URLs, and stack traces. The Session View helps you understand exactly what the user did and what the system reported at each step, making it easier to trace issues and user flows.

### AI-Powered Session Analysis

Beyond manual inspection, groundcover can analyze sessions for you using [Agent Mode](/use-groundcover/agent-mode). From within a session, choose **Explain this session** to get a natural-language summary of what the user did, where they hit friction, and which events matter most.

The Agent performs semantic analysis of the session — including the recording — so you can understand a user journey without replaying it end to end, and jump straight to the moments that led to an error or performance issue.

<figure><img src="/files/m9TR1nxDKlMDcwQhM9gU" alt=""><figcaption><p>The Agent summarizing a session with "Explain this session"</p></figcaption></figure>

### Exception Exploration

Front-end exceptions captured by RUM have a dedicated exploration experience for triaging and debugging client-side errors. Group and browse exceptions, drill into individual occurrences with full stack traces and contextual data (browser, URL, session), and pivot from an exception to the session and recording where it happened.

<figure><img src="/files/fi05WddsUm5DIj7IpTBb" alt=""><figcaption><p>Browsing and grouping front-end exceptions</p></figcaption></figure>

<figure><img src="/files/YvgbQRXztXK7z1wiJZ7t" alt=""><figcaption><p>An individual exception with its full stack trace and context</p></figcaption></figure>

For faster diagnosis, groundcover adds **AI root-cause analysis** powered by [Agent Mode](/use-groundcover/agent-mode): the Agent examines the exception, its stack trace, and the surrounding session context to suggest the likely root cause and where to look next. With [source maps](/getting-started/installation-and-updating/connect-rum#source-maps) uploaded, stack traces resolve to your original file names and line numbers, making both manual and AI-assisted investigation far more precise.

<figure><img src="/files/sPxRguB1JaDiL1Jq5kes" alt=""><figcaption><p>The Agent investigating an exception and producing a root-cause analysis</p></figcaption></figure>

### Web Vitals

Web Vitals provide information about the performance of your web pages. This page presents the following important web vitals metrics:

* **Largest Contentful Paint (LCP)** - how long it took the largest pixel area content to be presented, i.e. the time it took your users to see the main content of the page.
* **First Contentful Paint (FCP)** - how long it took for the first content to be presented, i.e. the time it took for the user to see anything on the page.
* **Interaction to Next Paint (INP)** - the time from any user interaction till the next paint occurs.
* **Time to First Byte (TTFB)** - The time for the first byte of response from the moment the user navigated to the page.
* **Cumulative Layout Shift (CLS)** - The sum of layout shifts that occur when elements moved unexpectedly.

For each of the above metrics, users can see a trend line and have the abiliy to focus on a specific page or a specific population in order to detect suspicious patterns.

<figure><img src="/files/8QYbNY84ElfTXLU9PcmS" alt=""><figcaption></figcaption></figure>


# AI Observability

{% hint style="info" %}
AI Observability requires an up-to-date sensor and backend. See [Installation & Updating](/getting-started/installation-and-updating).
{% endhint %}

## Overview

AI Observability gives you full visibility into how your services use AI — models, cost, latency, prompts, agent behavior, and tool execution. All data stays in your infrastructure with [BYOC (Bring Your Own Cloud)](/architecture/byoc); groundcover never processes your AI data outside your environment.

groundcover captures AI telemetry from two sources: **eBPF auto-detection** for immediate per-call visibility with zero code changes, and **OpenTelemetry instrumentation** for the full agent-level picture.

If your sensor is up to date and your services call a supported provider, you already have data. Open [AI Observability](https://app.groundcover.com/ai-observability) in your console to see what's there.

For instrumentation setup, privacy controls, and troubleshooting, see [Using AI Observability](/use-groundcover/ai-observability).

{% hint style="info" %}
**Monitoring AI coding tools (Claude Code, Claude Cowork, Codex)?** Those integrations ship tool-usage telemetry (logs, plus metrics for Claude Code), not GenAI traces, and don't surface here. See [AI Tools Observability](/integrations/data-sources/ai-tools-observability) for setup, dashboards, and verification queries.
{% endhint %}

***

## Two Levels of Visibility

eBPF captures every AI API call automatically — model, tokens, cost, latency, and full prompt/response content, with zero code changes. The sensor auto-detects **OpenAI**, **Anthropic**, and **AWS Bedrock** traffic from anything running in your monitored environment: production services, CI/CD pipelines, development pods, staging. If a process runs on a monitored node and makes an HTTPS call to a supported provider, groundcover captures it.

{% hint style="info" %}
For providers not yet auto-detected, use [SDK instrumentation](/use-groundcover/ai-observability#sdk-instrumentation) to send GenAI traces directly. [Contact us](https://www.groundcover.com/contact) if you need a specific provider — eBPF support is extended based on customer requests.
{% endhint %}

eBPF gives you the calls. To see the full picture — which agent triggered which call, how tool results shaped the next prompt, how a multi-step workflow reasons from start to finish — add OpenTelemetry instrumentation. SDK spans give you trace trees, agent workflows, tool execution chains, and conversation threading: not just what your AI costs, but how it thinks.

When both are active, the same call may appear twice — once from eBPF and once from the SDK. Both are correct; cost is not double-counted. When groundcover detects an SDK span for the same call, eBPF cost and token data is excluded from aggregations.

See [Using AI Observability](/use-groundcover/ai-observability) for setup instructions.

***

## Cost Tracking

groundcover calculates cost per AI span at ingestion time using a maintained model pricing table that updates as providers release new models. Each span includes input tokens, output tokens, and cached tokens.

By default, every AI span is stored — no sampling, no dropping. GenAI calls are high-value and low-volume compared to typical service traffic. Every call matters for cost analysis and debugging. If you need to disable GenAI storage entirely, see [Privacy Controls](/use-groundcover/ai-observability/privacy-controls#disable-genai-data-collection).

Cost and token data is available for filtering and sorting in the span list — find your most expensive calls by model, service, or agent. See [Example Queries](/use-groundcover/ai-observability/example-queries) for cost analysis patterns.


# Synthetics

Synthetics allow you to proactively monitor the health, availability, and performance of your endpoints. By simulating requests from your infrastructure, you can verify that your critical services are reachable and returning correct data, even when no real user traffic is active.

### Overview

groundcover Synthetics execute checks from your installed groundcover backend, working on BYOC deployments only.

* **Source**: Checks run from within your backend, when having multiple groundcover backends you can select the specific backend to use. We will support region selection for running tests from specific locations.
* **Supported Protocols**: HTTP/HTTPS is fully supported from the UI. SSL/TLS, TCP, and DNS checks are supported via the [Terraform provider](https://registry.terraform.io/providers/groundcover-com/groundcover/latest/docs/resources/synthetic_test) while the UI is being rolled out. Additional protocols (gRPC, ICMP/Ping, UDP, WebSocket) are coming soon.
* **Alerting**: Creating a Synthetic test automatically creates a corresponding Monitor (See: [Monitors](/use-groundcover/monitors)). Using monitors you can get alerted on failed synthetic tests, see: [Notification Channels](/integrations/workflow-integrations) . The monitor is uneditable.
* **Trace Integration**: We generate traces for all synthetic tests, which you can see as first-class citizens in groundcover platform. You can query these traces by using `source:synthetics` in traces page.

### Creating a Synthetic Test

Navigate to Monitors > Synthetics and click `+ Create Synthetic Test` .

{% hint style="info" %}
Only Editors can create/edit/delete synthetic tests, see [Role-Based Access Control (RBAC)](/use-groundcover/role-based-access-control-rbac#default-policies)
{% endhint %}

#### Request Configuration

Define the endpoint and parameters for the test. The available settings depend on the selected synthetic type.

**Common settings (all types)**

* **Synthetic Test Name**: A descriptive name for the test.
* **Interval**: Frequency of the check (e.g., every 60s). Must be at least 15s and less than 1h. The interval must also be greater than the total worker timeout (including retries).
* **Custom Labels**: Add custom labels that will exist on traces generated by checks. You can use these labels to filter traces.

**HTTP settings**

* **Target**: Select the HTTP method and URL. Include HTTP scheme (`http://` or `https://`), for example: `https://api.groundcover.com/api/backends/list`. Supported methods: `GET`, `POST`, `PUT`, `PATCH`, `DELETE`, `HEAD`, `OPTIONS`.
  * *Tip:* Use Import from cURL to paste a command and auto-fill these fields.
* **Follow redirects**: Whether the test should follow 3xx responses. Defaults to `true`. When disabled, the test returns the 3xx response as the result set for assertions.
* **Allow insecure**: Disables SSL/TLS certificate verification. Defaults to `false`. Use this only for internal testing or self-signed certificates — it exposes you to Man-in-the-Middle attacks on production endpoints.
* **HTTP Version**: The UI exposes `HTTP/1.0`, `HTTP/1.1` and `HTTP/2.0`; default is `HTTP/1.1`.
* **Timeout**: Max duration to wait before marking an attempt as failed. Required for HTTP checks (e.g., `30s`, `60s`). **The interval must be greater than the total worker timeout** (timeout × retry count + retry intervals).
* **Payload**: Set the body type and content if your request requires data (e.g., POST/PUT). Supported body types: `json`, `text`, `raw`. Omit the body entirely to send no payload.
* **Headers**: Custom headers as key/value pairs.
* **Authentication**: Supported types: `none` (default), `basic` (requires username and password), `bearer` (requires a token).

**SSL settings**

* **Host**: Hostname to connect to (e.g., `example.com`).
* **Port**: Port to connect to (typically `443`). Must be between 1 and 65535.
* **Verify**: Verify the certificate chain. Set to `false` to accept self-signed or invalid certificates.
* **Min Version**: Minimum accepted TLS version (e.g., `1.2`, `1.3`).
* **SNI**: Server Name Indication override.
* **Timeout**: Max duration for the TLS handshake. Default is 10s.

**TCP settings**

* **Host**: Hostname or IP to connect to.
* **Port**: TCP port to connect to. Must be between 1 and 65535.
* **Send**: Optional string payload to send after connecting.
* **Expect Response**: Whether to read a response back.
* **Receive Max Bytes**: Maximum bytes to read from the response.
* **Timeout**: Max duration for the connection and read. Default is 10s.

**DNS settings**

* **Domain**: Domain name to query (e.g., `example.com`).
* **Record Type**: One of `A`, `AAAA`, `CNAME`, `TXT`, `MX`, `NS`, `SOA`, `PTR`.
* **Resolver**: Optional DNS resolver to use. Defaults to the system resolver.
* **Port**: Optional DNS port. Must be between 1 and 65535 when set.
* **DNSSEC**: Whether to enable DNSSEC validation.
* **Timeout**: Max duration to wait for the DNS response. Default is 10s.

#### Assertions (Validation Logic)

Assertions are the rules that determine the outcome of a check. You can add multiple assertions to a single test. Each assertion evaluates independently to either **pass** or **fail**. The overall check outcome depends on the `severity` of the assertions that evaluate to fail:

* If any assertion with `severity: critical` evaluates to fail, the check outcome is **failed**. This triggers the auto-generated monitor alert.
* If no critical assertions fail but any assertion with `severity: degraded` evaluates to fail, the check outcome is **degraded**. This does *not* trigger the auto-generated monitor alert.
* If all assertions pass, the check outcome is **passed**.

An assertion has the following fields:

* **`source`** — which part of the response to inspect (e.g., `statusCode`, `jsonBody`, `certificateValid`, `tcpConnection`).
* **`property`** — required for `responseHeader` (header name) and `jsonBody` (dot-notation JSON path). Not used by other sources.
* **`operator`** — how to compare the `source` value against the `target` (e.g., `eq`, `contains`, `gt`).
* **`target`** — the expected value. Required for most operators except `exists`/`notExists`.
* **`severity`** *(optional)* — `critical` (default) or `degraded`. Controls how a failed assertion affects the check outcome: `critical` failures fail the check; `degraded` failures mark it as degraded without failing it.

**Assertion Sources per Synthetic Type**

The `source` field must be compatible with the synthetic type. The table below lists every valid source, the synthetic types it applies to, and the operators the worker currently validates.

| Source                 | Applies To          | Property                                                  | Operators                                                                                     | Target                                                                                                                                                                      |
| ---------------------- | ------------------- | --------------------------------------------------------- | --------------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `statusCode`           | HTTP                | *not used*                                                | `eq`, `ne`, `gt`, `lt`                                                                        | HTTP status code as a numeric string (e.g. `"200"`, `"404"`)                                                                                                                |
| `responseHeader`       | HTTP                | Header name (case-insensitive), e.g. `Content-Type`       | `exists`, `notExists`, `eq`, `ne`, `contains`                                                 | Expected header value (required for `eq`/`ne`/`contains`)                                                                                                                   |
| `responseBody`         | HTTP                | *not used*                                                | `contains`, `eq`, `ne`, `exists`, `notExists`                                                 | Substring or full body to match against. `exists`/`notExists` check whether the body is non-empty.                                                                          |
| `jsonBody`             | HTTP                | Dot-notation JSON path (e.g. `data.user.name`). Required. | `exists`, `notExists`, `eq`, `ne`, `contains`                                                 | Expected value at the JSON path. For `eq`, values that look like JSON (`{…}`, `[…]`, `"..."`), numbers, or booleans are auto-parsed; anything else is compared as a string. |
| `responseTime`         | HTTP, SSL, TCP, DNS | *not used*                                                | `gt`, `lt`                                                                                    | Threshold in milliseconds as a string (e.g. `"5000"`)                                                                                                                       |
| `certificateValid`     | SSL                 | *not used*                                                | `eq`                                                                                          | `"true"` or `"false"`                                                                                                                                                       |
| `certificateExpiresIn` | SSL                 | *not used*                                                | `gt`, `lt`                                                                                    | Days until expiry as a string (e.g. `"30"`)                                                                                                                                 |
| `tlsVersion`           | SSL                 | *not used*                                                | `eq`, `gt`, `lt`                                                                              | TLS version string (e.g. `"1.2"`, `"1.3"`)                                                                                                                                  |
| `chainValid`           | SSL                 | *not used*                                                | `eq`                                                                                          | `"true"` or `"false"`                                                                                                                                                       |
| `cipherSuite`          | SSL                 | *not used*                                                | `eq`, `contains`                                                                              | Cipher name (e.g. `"AES256"`)                                                                                                                                               |
| `tcpConnection`        | TCP                 | *not used*                                                | `exists`, `notExists`, `eq`, `ne`                                                             | `"true"`/`"false"` for `eq`/`ne`                                                                                                                                            |
| `responseContains`     | TCP                 | *not used*                                                | `contains`                                                                                    | Substring to match in received payload (requires `send` + `expectResponse` in the TCP request)                                                                              |
| `dnsAnswer`            | DNS                 | *not used*                                                | `exists`/`notExists` (does DNS resolve at all), `eq`, `ne`, `contains` (match a record value) | Expected DNS record value (e.g. an IP for A records). Not used for `exists`/`notExists`.                                                                                    |

{% hint style="warning" %}
The `target` field is always a **JSON string**, even when the value represents a number or boolean. The API rejects native JSON numbers and booleans (e.g., `"target": 200` returns 400). Use quoted values: `"target": "200"`, `"target": "true"`. The worker coerces numeric and boolean strings automatically for sources that require them (`statusCode`, `responseTime`, `certificateExpiresIn`, `tlsVersion`, `certificateValid`, `chainValid`). For `jsonBody` with operator `eq`, string targets that look like JSON, numbers, or booleans (`"true"`, `"false"`, `"123"`, `"{...}"`) are auto-parsed before comparison.
{% endhint %}

{% hint style="info" %}
The API schema also accepts the operators `startsWith`, `endsWith`, `regex`, and `oneOf`, and the sources `udp` and `websocket`. These are reserved for future use — the worker does not currently validate them and assertions using them will fail with an "Unsupported operator" or "Unknown assertion source" result.
{% endhint %}

**Assertion Operators**

The `operator` defines the logic applied to the source/property.

| Operator    | Function                                                                                                                                             | Example                                  |
| ----------- | ---------------------------------------------------------------------------------------------------------------------------------------------------- | ---------------------------------------- |
| `eq`        | Equals. Numeric compare for `statusCode`, `responseTime`, `tlsVersion`, `certificateExpiresIn`; otherwise string equality (case-sensitive).          | `statusCode` `eq` `200`                  |
| `ne`        | Not equal.                                                                                                                                           | `statusCode` `ne` `500`                  |
| `gt`        | Numeric greater-than. Used with `statusCode`, `responseTime`, `certificateExpiresIn`, `tlsVersion`.                                                  | `responseTime` `gt` `100`                |
| `lt`        | Numeric less-than. Same sources as `gt`.                                                                                                             | `responseTime` `lt` `5000`               |
| `contains`  | Checks if a substring exists within the source value.                                                                                                | `responseBody` `contains` `"error"`      |
| `exists`    | Checks that the field/resource is present. For `dnsAnswer`, checks that DNS resolves. For `tcpConnection`, checks that the TCP connection succeeded. | `responseHeader` (`set-cookie`) `exists` |
| `notExists` | Opposite of `exists`.                                                                                                                                | `jsonBody` (`error_message`) `notExists` |

### Examples

All examples below represent the JSON payload accepted by the `POST /api/synthetics/v1/rules` endpoint. HTTP checks can also be created via the UI, which renders this structure as form fields. SSL/TLS, TCP, and DNS checks are currently supported via the [Terraform provider](https://registry.terraform.io/providers/groundcover-com/groundcover/latest/docs/resources/synthetic_test) while UI support is being rolled out. For Terraform HCL examples, see the [provider documentation](https://registry.terraform.io/providers/groundcover-com/groundcover/latest/docs/resources/synthetic_test).

#### HTTP — health check with multiple assertions

Validates that an API endpoint returns `200`, completes within 5 seconds, exposes a `Content-Type` header, and that the JSON body's `status` field equals `"ok"`.

```json
{
  "version": 1,
  "name": "Homepage Health Check",
  "interval": "1m",
  "checkConfig": {
    "kind": "http",
    "metadata": { "syntheticName": "Homepage Health Check" },
    "request": {
      "http": {
        "kind": "http",
        "url": "https://example.com/api/health",
        "method": "GET",
        "timeout": "30s",
        "followRedirects": true
      }
    },
    "executionPolicy": {
      "assertions": [
        { "source": "statusCode",     "operator": "eq",     "target": "200" },
        { "source": "responseTime",   "operator": "lt",     "target": "5000" },
        { "source": "responseHeader", "operator": "exists", "property": "Content-Type" },
        { "source": "jsonBody",       "operator": "eq",     "property": "status", "target": "ok" }
      ]
    }
  }
}
```

#### SSL — certificate validity and expiration

Checks that the certificate served by `example.com:443` is valid, its chain is valid, and it expires in more than 30 days.

```json
{
  "version": 1,
  "name": "SSL Certificate Check",
  "interval": "15m",
  "checkConfig": {
    "kind": "ssl",
    "metadata": { "syntheticName": "SSL Certificate Check" },
    "request": {
      "ssl": {
        "kind": "ssl",
        "host": "example.com",
        "port": 443,
        "verify": true,
        "minVersion": "1.2"
      }
    },
    "executionPolicy": {
      "assertions": [
        { "source": "certificateValid",     "operator": "eq", "target": "true" },
        { "source": "chainValid",           "operator": "eq", "target": "true" },
        { "source": "certificateExpiresIn", "operator": "gt", "target": "30"   }
      ]
    }
  }
}
```

#### TCP — connectivity and response

Opens a TCP connection to `db.example.com:5432`, asserts the connection succeeded, and asserts the handshake completed in under 1 second.

```json
{
  "version": 1,
  "name": "Database TCP Check",
  "interval": "30s",
  "checkConfig": {
    "kind": "tcp",
    "metadata": { "syntheticName": "Database TCP Check" },
    "request": {
      "tcp": {
        "kind": "tcp",
        "host": "db.example.com",
        "port": 5432,
        "timeout": "5s"
      }
    },
    "executionPolicy": {
      "assertions": [
        { "source": "tcpConnection", "operator": "exists", "target": "true" },
        { "source": "responseTime", "operator": "lt", "target": "1000" }
      ]
    }
  }
}
```

#### DNS — A record resolution

Queries `example.com` for its `A` record and asserts the answer contains the expected IP.

```json
{
  "version": 1,
  "name": "DNS A Record Check",
  "interval": "5m",
  "checkConfig": {
    "kind": "dns",
    "metadata": { "syntheticName": "DNS A Record Check" },
    "request": {
      "dns": {
        "kind": "dns",
        "domain": "example.com",
        "recordType": "A"
      }
    },
    "executionPolicy": {
      "assertions": [
        { "source": "dnsAnswer",    "operator": "contains", "target": "93.184.216.34" },
        { "source": "responseTime", "operator": "lt",       "target": "2000" }
      ]
    }
  }
}
```

#### Assertion severity and retries

Assertions with `severity: critical` (the default) that evaluate to fail mark the check as **failed**. Assertions with `severity: degraded` that evaluate to fail mark the check as **degraded** — a warning state that does not count as a failure. Use `retries` to re-attempt the check before settling on the final outcome, avoiding false alerts on transient failures.

In the example below, if the status code assertion (`severity: critical`) evaluates to fail the check is failed; if only the response time assertion (`severity: degraded`) evaluates to fail the check is degraded but not failed.

```json
{
  "version": 1,
  "name": "Performance Check",
  "interval": "30s",
  "checkConfig": {
    "kind": "http",
    "metadata": { "syntheticName": "Performance Check" },
    "request": {
      "http": {
        "kind": "http",
        "url": "https://api.example.com/status",
        "method": "GET",
        "timeout": "5s"
      }
    },
    "executionPolicy": {
      "retries": { "count": 3, "interval": "500ms" },
      "assertions": [
        { "source": "statusCode",   "operator": "eq", "target": "200", "severity": "critical" },
        { "source": "responseTime", "operator": "lt", "target": "300", "severity": "degraded" }
      ]
    }
  }
}
```

### Auto-Generated Monitors

When you create a Synthetic Test, groundcover eliminates the need to manually configure separate alert rules. A Monitor is automatically generated and permanently bound to your test. See: [Monitors](/use-groundcover/monitors) .

* **Managed Logic**: The monitor queries for checks with outcome `failed` (i.e., a `severity: critical` assertion evaluated to fail). Checks with outcome `degraded` (only `severity: degraded` assertions failed) do **not** trigger the monitor.
* **Lifecycle**: The monitor transitions between `Pending`, `Firing` (when the test outcome is `failed`), and `Resolved` (when the test passes or is only degraded).
* **Zero Maintenance**: You do not need to edit this monitor's query. Any changes you make to the Synthetic Test (such as changing the target URL or assertions) are automatically synced to the Monitor.

> Note: To prevent configuration drift, these auto-generated monitors are read-only. You cannot edit their query logic directly; you simply edit the Synthetic Test itself.


# Requirements

To ensure a seamless experience with groundcover, it's important to confirm that your environment meets the necessary requirements. Please review the detailed requirements for Kubernetes, our eBPF sensor, and the necessary hardware and resources to guarantee optimal performance.

## Kubernetes requirements

groundcover supports a wide range of Kubernetes versions and distributions, including popular platforms like EKS, AKS, and GKE.

[**Learn more ->**](/getting-started/requirements/kubernetes-requirements)

## Kernel requirements for eBPF sensor

Our state-of-the-art eBPF sensor leverages advanced kernel features to deliver comprehensive monitoring with minimal overhead, requiring specific Linux kernel versions, permissions, and CO:RE support.

[**Learn more ->**](/getting-started/requirements/kernel-requirements-for-ebpf-sensor)

## Hardware and resource requirements

groundcover fully supports both x86 and ARM processors, ensuring compatibility across diverse environments.

[**Learn more ->**](/getting-started/requirements/cpu-architectures)

## ClickHouse resources

groundcover operates ClickHouse to support many of its core features. This requires suitable resources given to the deployment, which groundcover takes care of according to your data usage.


# Kubernetes requirements

### Kubernetes version

groundcover supports any K8s version from **v1.21.**

{% hint style="info" %}
groundcover may work on many other K8s flavors, but we might just didn't get a chance to test it yet. Can't find yours in the list? [let us know over Slack.](https://www.groundcover.com/join-slack)
{% endhint %}

### Kubernetes distributions

<table><thead><tr><th width="209.33333333333331">K8s distribution</th><th>Status</th><th>Comments</th></tr></thead><tbody><tr><td>EKS</td><td><code>supported</code></td><td></td></tr><tr><td>AKS</td><td><code>supported</code></td><td></td></tr><tr><td>GKE</td><td><code>supported</code></td><td></td></tr><tr><td>OKE</td><td><code>supported</code></td><td></td></tr><tr><td>OpenShift</td><td><code>supported</code></td><td></td></tr><tr><td>Rancher</td><td><code>supported</code></td><td></td></tr><tr><td>Self-managed</td><td><code>supported</code></td><td></td></tr><tr><td>minikube</td><td><code>supported</code></td><td></td></tr><tr><td>kind</td><td><code>supported</code></td><td></td></tr><tr><td>Rancher Desktop</td><td><code>supported</code></td><td></td></tr><tr><td>k0s</td><td><code>supported</code></td><td></td></tr><tr><td>k3s</td><td><code>supported</code></td><td></td></tr><tr><td>k3d</td><td><code>supported</code></td><td></td></tr><tr><td>microk8s</td><td><code>supported</code></td><td></td></tr><tr><td>AWS Fargate</td><td><mark style="color:red;"><code>not supported</code></mark></td><td></td></tr><tr><td>Docker-desktop</td><td><mark style="color:red;"><code>not supported</code></mark></td><td></td></tr></tbody></table>

### Kubernetes RBAC permissions

For the installation to complete successfully, permissions to deploy the following objects are required:

* StatefulSet
* Deployment
* DaemonSet (With privileged containers for loading our [eBPF sensor](/getting-started/requirements/kernel-requirements-for-ebpf-sensor))
* ConfigMap
* Secret
* PVC

To learn more about groundcover's architecture and components visit our [**Architecture Section**](/architecture/overview)

### Outgoing traffic

groundcover's <mark style="color:purple;">`portal`</mark> pod sends HTTP requests to the cloud platform `app.groundcover.com` on port 443.

This unique [architecture](/architecture/overview) keeps the the data inside the cluster and fetches it on-demand keeping the data encrypted all the way without the need to open the cluster for incoming traffic via ingresses.


# Kernel requirements for eBPF sensor

### Intro

groundcover’s eBPF sensor uses state-of-the-art kernel features to provide full coverage at low overhead. In order to do so it requires certain kernel features which are listed below.

{% hint style="info" %}
groundcover may work on many other linux kernels, but we might just didn't get a chance to test it yet. Can't find yours in the list? [let us know over Slack.](https://www.groundcover.com/join-slack)
{% endhint %}

### Kernel Version

Version v5.3 or higher (anything since 2020).

### Linux Distributions

| Name                    | Supported Versions     |
| ----------------------- | ---------------------- |
| Debian                  | 11+                    |
| RedHat Enterprise Linux | 8.2+                   |
| Ubuntu                  | 20.10+                 |
| CentOS                  | 7.3+                   |
| Fedora                  | 31+                    |
| BottlerocketOS          | 1.10+                  |
| Amazon Linux            | All off the shelf AMIs |
| Google COS              | All off the shelf AMIs |
| Azure Linux             | All off the shelf AMIs |
| Talos                   | 1.7.3+                 |

### Permissions

Loading eBPF code requires running privileged containers. While this might seem unusual, there's nothing to worry about - eBPF is [safe by design!](https://www.groundcover.com/ebpf#ebpf-security-how-to-maximize-ebpf-safety)

### CO:RE support

Our sensor uses eBPF’s [CO:RE](https://nakryiko.com/posts/bpf-portability-and-co-re/) feature in order to support the vast variety of linux kernels and distributions detailed above. This feature requires the kernel to be compiled with BTF information (enabled using the CONFIG\_BTF\_ENABLE=Y kernel compilation flag). This is the case for most common [distributions](https://github.com/libbpf/libbpf#bpf-co-re-compile-once--run-everywhere) nowadays.

You can check if your kernel has CO:RE support by manually looking for the BTF file:

```c
$ ls -la /sys/kernel/btf/vmlinux

- r--r--r--. 1 root root 3541561 Jun 2 18:16 /sys/kernel/btf/vmlinux
```

If the file exists, congratulations! Your kernel supported CO:RE.

### What happens if my kernel is not supported?

If your system does not fit into any of the above - unfortunately, our eBPF sensor will not be able to run on your environment. However, this does not mean groundcover won’t collect any data. You will still be able to inspect your [k8s environment](/capabilities/infrastructure-monitoring), see all [collected logs](/capabilities/log-management) and use [integrations ](/integrations/overview)with outer data sources.


# CPU architectures

The following architectures are fully supported for all groundcover workloads:

* x86
* ARM


# Login and Create a Workspace

Get started with groundcover

{% hint style="success" %}
This is the first step to start with groundcover for all types of plans :rocket:
{% endhint %}

### Sign up to groundcover

The first thing you need to do to start using groundcover is [sign up](https://console.groundcover.com/) to groundcover console and create a workspace and a backend. Signing up is only possible using a computer and will not be possible using a mobile phone or tablet.

{% hint style="info" %}
It is highly recommended you use your corporate email address, as it will make it easier to use other features such as inviting your colleagues to your workspace. However, signing up using Gmail, Outlook or any other public domains is also possible.
{% endhint %}

### Setup a backend

Follow the guidelines in the console to create a groundcover BYOC backend. Read more about groundcover architecture [here](/architecture/byoc).

<figure><img src="/files/qGPJhakQZzsn2y3cIRY0" alt=""><figcaption></figcaption></figure>

After the backend was successfully provisioned, click on Go to App to see your groundcover workspace.

### Workspace Selection

When [signing in](https://app.groundcover.com) to groundcover for the first time, the platform automatically detects based on your authentication details the workspaces you joined or eligible to join. If you already joined one or more workspaces, you will be redirected to the last viewed workspace. If not and you have existing workspaces available, the workspace selection screen will be displayed, where you can choose the workspaces to join.

Available workspaces will be displayed **only** if either of the following applies:

* You have been invited to join existing workspaces and haven't joined them yet
* Someone has previously created a workspace that has auto-join enabled for the email domain that you used to sign in (applicable for corporate email domains only)

<figure><img src="/files/Wf4JI5ccy7vgjt4LephM" alt=""><figcaption></figcaption></figure>

### To join an existing workspace:

1. Click the **Join** button next to the desired workspace
2. You will be added as a user to that workspace with the user privileges that were assigned by default or those that were assigned to you specifically when the invite was sent.
3. You will automatically be redirected to that workspace.

### Workspace Auto-joining

Workspace admins can allow teammates that log in with the same email domain as them to join the Workspace they created automatically, without an admin approval. This capability is called "Auto-join". It is disabled by default, but can be switched on any time in the workspace settings.

{% hint style="warning" %}
If you logged in with a **public email domain** (Gmail, Yahoo, Proton, etc.) and are creating a new Workspace, you will not be able to switch on Auto-join for that Workspace.
{% endhint %}


# I can't find my workspace

If you signed in to groundcover but cannot find the workspace you expected to join, there are a few common reasons. Use the checklist below to understand what to verify and when to contact your workspace admin.

### 1. You were not invited to the workspace

A workspace will appear in the workspace selection screen only if you are eligible to join it.

You may be eligible if:

* You were invited to the workspace and have not joined it yet.
* The workspace has auto-join enabled for your email domain.

Learn more in [Login and Create a Workspace](https://docs.groundcover.com/getting-started/login-and-create-a-workspace#workspace-selection).

If you do not see the workspace, ask a workspace admin to invite you. Admins can invite users from the groundcover UI using the **Invite** button.

Learn more in [How can I invite my team to my workspace?](https://docs.groundcover.com/welcome/faq#how-can-i-invite-my-team-to-my-workspace).

### 2. Auto-join is not enabled for your domain

If your organization uses workspace auto-joining, users with an approved corporate email domain can join the workspace without a manual invite.

Ask a workspace admin to verify that:

* Auto-join is enabled for the workspace.
* Your email address belongs to the approved domain.
* You are signing in with the correct corporate email address.

Admins can review and manage auto-join from the workspace settings.

Learn more in [Workspace Auto-joining](https://docs.groundcover.com/getting-started/login-and-create-a-workspace#workspace-auto-joining).

{% hint style="warning" %}
Auto-join is available for corporate email domains. It cannot be enabled for public email domains such as Gmail, Yahoo, or Proton.
{% endhint %}

### 3. Your SSO profile is not mapped correctly

If your organization uses SSO, groundcover uses your identity provider profile to authenticate you and determine workspace access.

Ask your IdP or workspace admin to verify that:

* You are assigned to the groundcover application in the identity provider.
* Your SSO profile uses the expected email address.
* Your email domain matches the domain configured for the workspace.
* Your SSO attributes, groups, or claims are being sent correctly to groundcover.

Learn more about SSO support in [Security considerations](https://docs.groundcover.com/architecture/security-considerations#single-sign-on-sso-support-with-oidc-and-saml).

For Okta-based SSO, see [Okta SSO - onboarding](https://docs.groundcover.com/architecture/security-considerations/okta-sso-onboarding).

### 4. Your SSO policy or RBAC policy does not match your expected access

If you can access the workspace but do not have the expected permissions, or you cannot see the expected data, your assigned policy may not match the access you need.

For example, you may expect to have:

* Admin, Editor, or Viewer access.
* Access to specific clusters, environments, or namespaces.
* Different data permissions than the ones currently assigned to you.

Ask your admin to check both sides of the configuration:

1. In the identity provider, verify the policy, group, role, or claim assigned to your user.
2. In groundcover, go to **Settings → Policy** and verify that the matching groundcover policy grants the correct permission level and data scope.

groundcover RBAC policies define both the user’s permission level and the data scope they can access.

Learn more in [Role-Based Access Control (RBAC)](https://docs.groundcover.com/use-groundcover/role-based-access-control-rbac).

### Still can't find your workspace?

Contact your workspace admin and share the following details:

* The email address you used to sign in.
* Whether you are using SSO.
* The workspace name you expected to join.
* Whether you received an invite.
* Whether your organization expects you to join through auto-join.
* The role or data access you expected to have, if relevant.


# Installation & Updating

Multiple ways to connect your infrastructure and applications to groundcover

groundcover is designed to support data ingestion from multiple sources, giving you comprehensive observability across your entire stack. Choose the installation method that best fits your infrastructure and monitoring needs.

## Available Installation Options

### Kubernetes Clusters

Connect your Kubernetes clusters using groundcover's eBPF-based sensor for automatic instrumentation and deep observability.

* [**Connect Kubernetes clusters**](/getting-started/installation-and-updating/connect-kubernetes-cluster) - Deploy groundcover's sensor to monitor containerized workloads, infrastructure, and applications with zero code changes

### Standalone Linux Hosts

Monitor individual Linux servers, virtual machines, or cloud instances outside of Kubernetes.

* [**Connect Linux hosts**](/getting-started/installation-and-updating/connect-linux-hosts) - Install groundcover on standalone Linux hosts such as EC2 instances, bare metal servers, or VMs

### Real User Monitoring (RUM)

Gain visibility into frontend performance and user experience with client-side monitoring.

* [**Connect RUM**](/getting-started/installation-and-updating/connect-rum) - Monitor real user interactions, page loads, and frontend performance in web applications

### External Data Sources

Integrate with existing observability tools and send data from your current monitoring stack.

* [**Ship from OpenTelemetry**](/integrations/data-sources/opentelemetry) - Forward traces, metrics, and logs from existing OpenTelemetry collectors
* [**Ship from Datadog Agent**](/integrations/data-sources/datadog) - Send data from Datadog agents while maintaining your existing setup

## Getting Started

1. [**Login and Create a Workspace**](/getting-started/login-and-create-a-workspace) - Set up your groundcover account and workspace
2. **Review** [**Requirements**](/getting-started/requirements) - Ensure your environment meets the necessary prerequisites
3. **Choose your installation method** - Select the option that matches your infrastructure setup
4. **Follow the** [**5 quick steps**](/getting-started/5-quick-steps-to-get-you-started) - Get oriented with groundcover's interface and features

## Need Help?

If you're unsure which installation method is right for you, or if you have specific requirements, check our [FAQ](/welcome/faq) or reach out to our support team.


# Connect Kubernetes clusters

Get up and running in minutes in Kubernetes

Before installing groundcover in Kubernetes, please make sure your cluster meets the [requirements](/getting-started/requirements).

After ensuring your cluster meets the requirements, complete the [login and workspace setup](/getting-started/login-and-create-a-workspace), then choose your preferred installation method:

* [CLI](#installing-using-cli)
* [Helm](#installing-using-helm)
* [ArgoCD](#installing-using-argocd)

{% hint style="info" %}
Coverage policy covers all nodes excluding control plane and fargate. [See details here](/customization/customize-deployment/configuring-sensor-deployment-coverage).
{% endhint %}

## Creating helm values file

Sensor deployment requires installation values similar to these stored in a `values.yaml` file

```yaml
global:
  backend:
    enabled: false
  ingress:
    site: {BYOC_ENDPOINT}

clusterId: "your-cluster-name" # CLI will automatically detect cluster name
env: "your-environment-name" # Add it to differentiate between different environments
```

{% hint style="info" %}
**Default values:** If `clusterId` is not set, the CLI will auto-detect the cluster name. If `env` is not configured, the environment will display as **"none"** in the Cluster Picker. We recommend setting both values for better organization. [Learn more about environment labels](/use-groundcover/add-custom-environment-labels).
{% endhint %}

`{BYOC_ENDPOINT}` is your unique groundcover ingestion endpoint, you can locate it in the [ingestion keys tab](https://app.groundcover.com/settings?selectedTab=ingestion-keys).

## Installing using CLI

Use groundcover CLI to automate the installation process. The main advantages of using this installation method are:

* Auto-detection of cluster incompatibility issues
* Tolerations setup automation
* Tuning of resources according to cluster size
* Supports passing helm overrides
* Automated detection of new versions and upgrades suggestions

Read more [here](https://github.com/groundcover-com/cli).

{% hint style="success" %}
The CLI will automatically use existing ingestion keys or provision a new one if none exist
{% endhint %}

#### **Installing groundcover CLI**

```bash
sh -c "$(curl -fsSL https://groundcover.com/install.sh)"
```

**Deploying groundcover using the CLI**

```bash
groundcover deploy -f values.yaml
```

To upgrade groundcover to the latest version, simply re-run the `groundcover deploy` command with your desired overrides (such as `-f values.yaml`). The CLI will automatically detect and apply the latest available version during the deployment process.

## Installing using Helm

### Step 1 - Install groundcover CLI

```bash
sh -c "$(curl -fsSL https://groundcover.com/install.sh)"
```

### Step 2 - Generate Installation Key

For more details about ingestion keys, refer to our [ingestion key documentation](/use-groundcover/remote-access-and-apis/ingestion-keys).

```bash
groundcover auth get-ingestion-key sensor
```

### Step 3 - Add Sensor Ingestion Key to Values File

Add the recently created sensor key and your BYOC endpoint to the values.yaml file. Find your BYOC endpoint in the [ingestion keys tab](https://app.groundcover.com/settings?selectedTab=ingestion-keys).

```yaml
global:
  groundcover_token: {sensor_key}
  backend:
    enabled: false
  ingress:
    site: {BYOC_ENDPOINT}
    
clusterId: "your-cluster-name"
env: "your-environment-name"
```

{% hint style="info" %}
**Note:** Clusters without `env` configured will show **"none"** in the Cluster Picker. Set this value to organize clusters by environment (e.g., `prod`, `staging`, `dev`).
{% endhint %}

### Step 4 - Add Helm Repository

```bash
# Add groundcover Helm repository and fetch latest chart
helm repo add groundcover https://helm.groundcover.com && helm repo update groundcover
```

### Step 5 - Install groundcover

Initial installation:

```bash
helm upgrade \
    groundcover \
    groundcover/groundcover \
    -i \
    --create-namespace \
    -n groundcover \
    -f values.yaml
```

Upgrade groundcover:

```bash
helm repo update groundcover && helm upgrade \
    groundcover \
    groundcover/groundcover \
    -n groundcover \
    -f values.yaml
```

## Installing using ArgoCD

For CI/CD deployments using ArgoCD, refer to our [ArgoCD deployment guide](/customization/customize-deployment/argo-cd).

## **What can you do next?**

Check out our [5 quick steps to get you started](/getting-started/5-quick-steps-to-get-you-started)

## Uninstalling

### CLI

```bash
groundcover delete
```

### Helm

```bash
helm uninstall groundcover -n groundcover
```

```bash
# delete the namespace in order to remove the PVCs as well
kubectl delete ns groundcover
```


# Connect Linux hosts

Linux hosts sensor

## Supported Environments

We currently support running on eBPF-enabled linux machines (See more [Kernel requirements for eBPF sensor](/getting-started/requirements/kernel-requirements-for-ebpf-sensor))

{% hint style="info" %}
In case your linux machine doesn't comply with the Kernel requirements, post installation, run the following commands:\
\
`echo "FLORA_EBPFENABLED=false" | sudo tee -a /etc/opt/groundcover/env.conf > /dev/null`\
`sudo systemctl restart groundcover-sensor`
{% endhint %}

Supported architectures: `AMD64` + `ARM64`

For the following providers, we will fetch the machine metadata from the provider's API.

<table><thead><tr><th width="144" align="center">Provider</th><th width="119" align="center">Supported</th></tr></thead><tbody><tr><td align="center">AWS</td><td align="center">✅</td></tr><tr><td align="center">GCP</td><td align="center">✅</td></tr><tr><td align="center">Azure</td><td align="center">✅</td></tr><tr><td align="center">Linode</td><td align="center">✅</td></tr></tbody></table>

## Sensor capabilities

* Infrastructure Host metrics: CPU/Memory/Disk usage
* Logs
  * Natively from docker containers running on the machine
  * JournalD ([requires configuration](https://docs.groundcover.com/customization/customize-usage/custom-logs-collection#configure-journal-logs))
  * Static log files on the machine ([requires configuration](https://docs.groundcover.com/customization/customize-usage/custom-logs-collection#configure-log-file-targets))
* Traces
  * Natively from docker containers running on the machine
* APM metrics and insights from the traces

\
How to install?
---------------

Installation currently requires running a script on the machine.

The script will pull the latest sensor version and install it as a service: **groundcover-sensor (requires privileges)**

### Install/Upgrade existing sensor:

```bash
curl -fsSL https://groundcover.com/install-groundcover-sensor.sh | sudo env API_KEY='{ingestion_Key}' GC_ENV_NAME='{selected_Env}' GC_DOMAIN='{BYOC_ENDPOINT}' bash -s -- install
```

Where:

* {ingestion\_Key} - A dedicated ingestion key, you can generate or find existing ones from Settings -> Access -> Ingestion Keys
  * Ingestion Key needs to be of Type `Sensor`
* {BYOC\_Endpoint} - Your BYOC public ingestion endpoint
* {selected\_Env} - The **Environment** that will group those machines on the cluster drop down in the top right corner (We recommend setting a separate one for non k8s deployments)

### Check installed sensor status:

* Check service status: `systemctl status groundcover-sensor`
* View sensor logs: `journalctl -u groundcover-sensor`

{% hint style="info" %}
Initial data may take a few minutes to appear in the app after installation
{% endhint %}

### Remove installed sensor:

```bash
curl -fsSL https://groundcover.com/install-groundcover-sensor.sh | sudo bash -s -- uninstall
```

## Customize sensor configuration:

The sensor supports overriding its default configuration by writing to the file is located in:

`/etc/opt/groundcover/overrides.yaml`.\
\
After writing it you should restart the sensor service using:

`systemctl restart groundcover-sensor`

Example 1 - override Docker max log line size:

```bash
echo "# Local overrides to sensor configuration
k8sLogs:
  scraper:
    dockerMaxLogSize: 102400
" | sudo tee /etc/opt/groundcover/overrides.yaml && sudo systemctl restart groundcover-sensor
```

Example 2 - add static labels to your metrics:

```yaml
// add custom static labels to metrics
pipelines:
  metrics:
    additionalMetricLabels:
      label1: "label1_value"
      label2: "label2_value"
```


# Connect RUM

{% hint style="info" %}
This capability is only available to BYOC deployments. Check out our [pricing page](https://www.groundcover.com/pricing) for more information about subscription plans and the available deployment modes.
{% endhint %}

groundcover’s Real User Monitoring (RUM) SDK captures front-end **performance**, **user interactions**, **errors**, **logs**, **distributed traces**, and **session replay** from your web application — with privacy masking **on by default**.

**Start capturing RUM data** by installing the [browser SDK](https://www.npmjs.com/package/@groundcover/browser) in your web app.

This guide walks you through installing and initializing the SDK, the full configuration reference, identifying users, sending custom events and logs, capturing exceptions, session management, source maps, and session replay.

### Install the SDK

```bash
npm install @groundcover/browser
# or
yarn add @groundcover/browser
```

### Initialize the SDK

A single `init()` call installs every instrumentation (page loads, DOM interactions, network requests, errors, console logs, navigation, and performance) and starts sending data. Session replay is the one exception — it must be started explicitly with [`startReplayRecording()`](#session-replay).

```typescript
import groundcover from '@groundcover/browser';

groundcover.init({
  apiKey: 'your-ingestion-key',
  dsn: 'your-dsn',
  cluster: 'your-cluster',
  appId: 'your-app-id',
  environment: 'production',
});
```

From here you can enrich it:

```typescript
// Tie events to a user
groundcover.identifyUser({ id: 'u_123', email: 'john@acme.com', organization: 'acme' });

// Capture a handled error
groundcover.captureException(new Error('Checkout failed'), { feature: 'checkout' });

// Emit a structured log
groundcover.logger.warn('Payment retry', { provider: 'stripe', attempt: 2 });

// Emit a custom business event
groundcover.sendCustomEvent({ event: 'plan_upgraded', attributes: { plan: 'pro' } });
```

## Configuration

`init()` takes your connection and identity fields at the top level, plus an `options` object for behavioral configuration.

### Connection & identity

| Field                                         | Required | Description                                                                                                                                                                                                                                      |
| --------------------------------------------- | -------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| `apiKey`                                      | ✅        | A dedicated Ingestion Key of type `RUM` (Settings → Access → Ingestion Keys).                                                                                                                                                                    |
| `dsn`                                         | ✅        | Your public groundcover endpoint, in the format `https://example.platform.grcv.io`, where `example.platform.grcv.io` is your `ingress.site` installation value.                                                                                  |
| `cluster`                                     | ✅        | Identifier for your cluster; helps filter RUM data by cluster.                                                                                                                                                                                   |
| `appId`                                       | ✅        | Application identifier; reported as `service.name`.                                                                                                                                                                                              |
| `environment`                                 | —        | Deployment environment (e.g. `production`, `staging`) used for filtering.                                                                                                                                                                        |
| `namespace`, `releaseId`, `user`, `sessionId` | —        | Optional identity/grouping fields. `releaseId` associates uploaded [source maps](#source-maps) with a release; `user` matches [`identifyUser`](#identify-users); `sessionId` enables [shared sessions](#micro-frontend-session-synchronization). |

### Behavioral options

All behavioral configuration lives under `options`, grouped by concern — sampling (`sessionSampleRate`, `eventSampleRate`), enabled instrumentations (`enabledEvents`), `excludedUrls`, the `beforeSend` / `enrichEvent` hooks, and the `privacy`, `tracing`, `transport`, and `replay` groups. You can update it at runtime with `groundcover.updateConfig(...)`.

For the complete, always-current configuration reference — every option, type, and default — see the [`@groundcover/browser` package on npm](https://www.npmjs.com/package/@groundcover/browser). Data masking is **on by default**; see [Privacy and data masking](#privacy-and-data-masking) below.

## Privacy and data masking

Masking is **on by default** (`privacy.level: 'mask-sensitive'`). A single `level` is the master switch; finer toggles and hooks refine it.

| `level`                        | Replay inputs    | Replay text                               | DOM events       | Network / logs / errors               |
| ------------------------------ | ---------------- | ----------------------------------------- | ---------------- | ------------------------------------- |
| `mask-sensitive` **(default)** | sensitive masked | `[data-private]` + `maskSelectors` masked | sensitive masked | redacted                              |
| `mask-all`                     | masked           | masked (`*`)                              | all masked       | redacted                              |
| `allow`                        | —                | —                                         | —                | off (auth headers are still stripped) |

Under `mask-sensitive`, an input/element is masked when it is a `type="password"`, sits under a `[data-private]` ancestor or a `maskSelectors` match, or has an `id`/`name`/`class`/`aria-label`/`placeholder` matching a built-in sensitive-key pattern or your `sensitiveKeys`. Non-sensitive inputs and static page text stay visible; use `[data-private]` / `maskSelectors` to mask static content.

```typescript
options: {
  privacy: {
    level: 'mask-sensitive',
    maskSelectors: ['.pii', '#ssn'],
    sensitiveKeys: ['account_no'],
  },
}
```

**Built-in sensitive patterns** (always treated as sensitive, case-insensitive; your `sensitiveKeys` merge on top):

* **Key substrings** — matched inside a body/query key or a DOM element attribute (`id`/`name`/`class`/`aria-label`/`placeholder`): `token`, `secret`, `passwd`, `password`, `api_key`, `access_key`, `write_key`, `auth`, `bearer`, `credential`, `cvv`, `ssn`, `credit_card`, `card_number` (the `_` in the last five is optional — `apikey` / `api-key` also match).
* **Request/response headers** — always stripped regardless of `level`: `authorization`, `cookie`, `set-cookie`, plus any header name containing `token`, `key`, `secret`, `passwd`, `password`, `auth`, `bearer`, or `credential`.
* **Query / form param names** — matched as a whole key (not inside JSON bodies), for OAuth-style callbacks: `code`, `state`, `session_state`, `id_token`, `access_token`, `refresh_token`, `token`.

To turn masking off, set `privacy: { level: 'allow' }`. Note this disables masking only — auth/request headers are **still stripped** regardless of `level` (see the header list above).

## Identify Users

Link RUM data to specific users. An omitted key leaves the current identity untouched; pass `null` via `updateConfig` to clear it (e.g. on logout).

```typescript
groundcover.identifyUser({
  id: 'u_123',
  email: 'john@acme.com',
  name: 'John Doe',
  role: 'admin',
  organization: 'acme',
  properties: { plan: 'pro' },
});
```

## Send Custom Events

Instrument key user interactions or business events:

```typescript
groundcover.sendCustomEvent({
  event: 'PurchaseCompleted',
  attributes: { orderId: 1234, amount: 99.99 },
});
```

{% hint style="info" %}
Custom event payloads are **not** auto-redacted (they’re deliberately provided). Scrub sensitive fields yourself, or via `enrichEvent`.
{% endhint %}

## Capture Exceptions

Manually track caught errors with optional context:

```typescript
try {
  performAction();
} catch (error) {
  groundcover.captureException(error, { userId: '123', feature: 'checkout' });
}
```

## Send Logs

`groundcover.logger` provides one method per level — `log`, `info`, `warn`, `error`, `debug`, `trace`. The second argument is an attributes object; nested objects are flattened to dotted keys.

```typescript
groundcover.logger.info('User entered new experience', { releaseId: '1.5.3' });

groundcover.logger.warn('Checkout failed', {
  orderId: 'ord_42',
  cart: { items: 3, total: 99.99 }, // → cart.items, cart.total
});
```

The SDK also auto-captures `console.*` calls; when any argument is a plain object, its keys are promoted to structured log attributes. Reserved keys (`message`, `level`, `location`) are always set by the SDK and can’t be overridden.

## Session Management

Read or override the current session id:

```typescript
const id = groundcover.getSessionId();

groundcover.setSessionId('my-session-id'); // omit the argument to mint a fresh id
```

### Micro-frontend session synchronization

Pass a shared `sessionId` so multiple frontends report under one session:

```typescript
const sharedSessionId = 'session-12345';
groundcover.init({ /* …app… */ apiKey, dsn, cluster, appId: 'shell', sessionId: sharedSessionId });
groundcover.init({ /* …mfe… */ apiKey, dsn, cluster, appId: 'micro-frontend', sessionId: sharedSessionId });
```

## Session lifecycle

`sessionMaxDuration` sets a **target** maximum wall-clock session length (default 4 hours; must be between 1 minute and 8 hours).

It is enforced **lazily, on activity — not by a background timer**, so it is **not a hard upper bound**. Once the cap has elapsed, the *next* user/business event (click, navigation, log, custom, network, exception, …) flushes pending events under the current session id, mints a fresh id, and resumes replay recording if it had been active.

Because rotation is activity-gated, a session that goes idle keeps its id past the cap until the next qualifying event. Sessions are also bounded by a 30-minute inactivity gap, enforced the same lazy way. The flush is best-effort and the rotation always proceeds regardless of delivery success. Invalid values fall back to the default with a `console.warn`.

## Manual navigation

When `navigation` isn’t auto-tracked (e.g. a custom router), you can bracket navigation spans manually:

```typescript
groundcover.startNavigation({ to: '/checkout' });
// …route transition…
groundcover.endNavigation({ to: '/checkout' });
```

## Source Maps

Source maps map minified/bundled code back to your original source files. With them, stack traces in RUM (e.g. in session details and exceptions) show your real file names, line numbers, and function names instead of bundle names and minified positions.

Pre-condition: source maps were enabled in your CI operation.

To enable source maps in your account, upload them from your CI job via the API:

```
POST /api/rum/sourcemaps
Content-Type: multipart/form-data
```

Required form fields:

* `app_id` - your application identifier (alphanumeric, dots, hyphens, underscores), as provided in the RUM init call.
* `release_id` - the release/version being deployed (same character restrictions), as provided in the RUM init call.
* `file` - the source map file.

Required headers:

* `Authorization` - groundcover api key - [here is how to generate one](/use-groundcover/remote-access-and-apis/api-keys#creation-and-storage).
* `X-Backend-Id` - the relevant BYOC backend, displayed in the api keys page.

Here is an example of source maps uploading using curl:

```shellscript
curl -X POST "https://app.groundcover.com/api/rum/sourcemaps" \
-H "Authorization: Bearer <API_KEY>" \
-H "X-Backend-Id: $BACKEND_ID" \
-F "app_id=my-web-app" \
-F "release_id=1.2.3" \
-F "file=@./dist/main.js.map"
```

Response (201):

```json
{
"status": "uploaded",
"app_id": "my-web-app",
"release_id": "1.2.3",
"filename": "main.js.map",
"size_bytes": 524288
}
```

Files will be stored in the selected provider at the path: `sourcemaps/<app_id>/<release_id>/<filename>`.

Every new RUM trace that holds a stack trace will be automatically converted based on the source map.

## Session Replay

Session replay records your users’ behavior with [rrweb](https://github.com/rrweb-io/rrweb) so you can replay the actions that led to specific RUM events.

Replay recording does **not** start automatically — even when `replay` is in `enabledEvents`. You must start it explicitly (for example, after obtaining user consent), and can stop it before a sensitive section of your app:

```typescript
groundcover.startReplayRecording();
groundcover.stopReplayRecording();
```

Recording is stored on your BYOC server and is deleted along with the RUM session, in accordance with your retention policy.

### Masking replay content

Replay masking is driven by your [`privacy`](#privacy-and-data-masking) config — masking is **on by default**. To mask specific static content, add the `data-private` attribute or a `maskSelectors` match:

```typescript
options: {
  privacy: {
    level: 'mask-sensitive',      // default
    maskSelectors: ['.pii'],       // always masked in replay + DOM
  },
  replay: {
    blockedSelectors: ['.grammarly-extension'], // excluded from recording (noise reduction)
  },
}
```

{% hint style="info" %}
rrweb mask options are fixed at `record()` time. Changing privacy config at runtime via `updateConfig` automatically **restarts** the active recording so the new masking applies.
{% endhint %}

### Viewing sessions

In the [summary page](https://app.groundcover.com/rum/summary), you will see an indication next to sessions with a recording.

<figure><img src="/files/CenWUozMBBzJdwjpHzfx" alt=""><figcaption></figcaption></figure>

Within the drawer, open the Session Replay tab to see the recording.

<figure><img src="/files/pFH17muGKYIETHdw00zv" alt=""><figcaption></figcaption></figure>

## API reference

All methods are available on the default export and on `window.groundcover`.

| Method                                                           | Description                                                     |
| ---------------------------------------------------------------- | --------------------------------------------------------------- |
| `init(config)`                                                   | Initialize the SDK and install instrumentation.                 |
| `identifyUser(user)`                                             | Attach user identity to subsequent events.                      |
| `sendCustomEvent({ event, attributes })`                         | Emit a custom business event.                                   |
| `captureException(error, metadata?)`                             | Capture a handled error with optional context.                  |
| `logger.{log,info,warn,error,debug,trace}(message, attributes?)` | Structured logging.                                             |
| `updateConfig({ options?, user?, … })`                           | Update config at runtime (merges nested groups one level deep). |
| `startNavigation(metadata)` / `endNavigation(metadata)`          | Manual navigation spans (when `navigation` isn’t auto-tracked). |
| `getSessionId()` / `setSessionId(id?)`                           | Read / override the current session id.                         |
| `startReplayRecording()` / `stopReplayRecording()`               | Manually control session replay.                                |

## Migrating to 1.0.0

`1.0.0` restructures `options` by concern (a clean break from `0.x`) and removes the deprecated masking flags. Update your config as follows:

| `0.x`                                        | `1.0.0`                                                               |
| -------------------------------------------- | --------------------------------------------------------------------- |
| `environment` *(duplicated in `options`)*    | top-level `environment` only                                          |
| `userIdentifier`                             | `user`                                                                |
| `options.sessionReplay.blockedSelectors`     | `options.replay.blockedSelectors`                                     |
| `options.tracePropagationUrls`               | `options.tracing.propagationUrls`                                     |
| `options.tracePropagationHeaders`            | `options.tracing.propagationHeaders`                                  |
| `options.tracePropagationTraceIdHeaderName`  | `options.tracing.traceIdHeaderName`                                   |
| `options.tracePropagationSpanIdHeaderName`   | `options.tracing.spanIdHeaderName`                                    |
| `options.traceOrigin`                        | `options.tracing.origin`                                              |
| `options.batchSize` / `options.batchTimeout` | `options.transport.batchSize` / `options.transport.batchTimeout`      |
| `options.enableCompression`                  | `options.transport.compression`                                       |
| `options.enableMasking: true` *(removed)*    | set `options.privacy.level: 'mask-all'`                               |
| `options.enableMasking: false` *(removed)*   | set `options.privacy.level: 'allow'`                                  |
| `options.maskFields` *(removed)*             | set `options.privacy.maskSelectors` / `options.privacy.sensitiveKeys` |

{% hint style="warning" %}
The removed masking flags are **ignored, not auto-mapped** — passing `enableMasking` / `maskFields` logs a `console.warn` and has **no effect**. You must set the corresponding `privacy` option yourself. Masking remains **on by default** (`mask-sensitive`); if you previously relied on masking being off, set `privacy: { level: 'allow' }` explicitly.
{% endhint %}


# 5 quick steps to get you started

Once installed, we recommend following these steps to help you quickly gain the most our of groundcover's unique observability platform.

## 1. Get a complete view of your workloads

The "Home page" of the groundcover app is our Workloads page. From here, you can get a service-centric view,

[**Go to Workloads →**](https://app.groundcover.com/workloads)

<figure><img src="/files/xD4oH8NlACcS0X6rGeGS" alt=""><figcaption></figcaption></figure>

## 2. Check out full payloads of traces

A highly impactful advantage of leveraging eBPF in our proprietary sensor is that it enables visibility on the [full payloads of both request and response](/capabilities/application-performance-monitoring-apm/traces) - including headers! This allows you to very quickly understand issues and provides context.

[**Go to Traces →**](https://app.groundcover.com/traces)

<figure><img src="/files/bwu7MkiJbbqWUEcYCWLl" alt=""><figcaption></figcaption></figure>

## 3. Build a native dashboard

groundcover enables to very easily [build custom dashboards](/use-groundcover/dashboards-and-alerts/create-a-dashboard) to visualize your data using our intuitive Query Builder as a guide, or using your own queries.

[**Go to Dashboards →**](https://app.groundcover.com/dashboards)

<figure><img src="/files/cJpNkx0VBuzbuA6xQHDc" alt=""><figcaption></figcaption></figure>

## 4. Set up a Monitor

[Define custom alerts](/use-groundcover/monitors/create-a-new-monitor) using our native Monitors, which you can configure using groundcover data and custom metrics. You can also choose from our Monitors Catalog, which contains multiple pre-built Monitors that cover a few of the most common use cases and needs.

[**Go to Monitors →**](https://app.groundcover.com/monitors)

<figure><img src="/files/9TLh2JDWmv4hvSMHPGGr" alt=""><figcaption></figcaption></figure>

## 5. Invite your team

Invites lets you share your workspaces with your colleagues in just a couple of clicks. You can find the "Invite Members" option at the bottom of the left navigation bar. Type in the email addresses of the teammates you want to invite, and set their user permissions (Admin, Editor, Read Only), then click "Send Invites".

<figure><img src="/files/27HFWmR8k7GR27wL2dYJ" alt="" width="375"><figcaption></figcaption></figure>


# groundcover MCP

Supercharge your AI agents with o11y superpowers using the groundcover MCP server. Bring logs, traces, metrics, events, K8s resources, and more directly into your agent’s context — and troubleshoot side by side with your AI assistant to solve issues in minutes.

> **Status:** Work in progress. We keep adding tools and polishing the experience. Got an idea or question? Ping us on Slack!

## What is MCP?

**MCP (Model Context Protocol)** is an open standard that enables AI tools to access external data sources like APIs, observability platforms, and documentation directly within their working context. MCP servers expose these resources in a consistent way, so the AI agent can query and use them as needed to streamline workflows.

groundcover’s MCP server brings your live observability data into the picture - making your agents smarter, faster, and more accurate.

## How Can groundcover's MCP Help You and Your Agent?

By connecting your agent to groundcover’s MCP server, you enable it to:

* Query live logs, traces, and metrics for a workload, container, or issue.
* Run root cause analysis (RCA) on issues, right inside your IDE or chat interface.
* Auto-debug code with observability context built in.
* Monitor deployed code and validate fixes without switching tools.

See examples in our [Getting-started Prompts](/getting-started/groundcover-mcp/getting-started-prompts) and [Real-world Use Cases](/getting-started/groundcover-mcp/real-world-use-cases), or jump to the [MCP Tools Reference](/getting-started/groundcover-mcp/mcp-tools-reference) for the full list of tools and parameters.

## Install groundcover's MCP Server

Set up is quick and agent-friendly. We support both **OAuth** (recommended) and **API key** flows.

Head to [Configure groundcover's MCP Server](/getting-started/groundcover-mcp/configure-groundcovers-mcp-server) for setup instructions and client-specific guides.


# Configure groundcover's MCP Server

Set up your agent to talk to groundcover's MCP server. Use OAuth for a quick login, or an API key for service accounts.

The MCP endpoint is:

```
https://mcp.groundcover.com/api/mcp
```

{% hint style="info" %}
Self-hosted (on-prem / air-gapped) deployments use their own MCP URL. Replace `https://mcp.groundcover.com/api/mcp` with your environment's endpoint everywhere in this guide.
{% endhint %}

**The MCP server supports two methods:**

* [**OAuth**](#oauth-recommended) (Recommended for IDEs)
* [**API Key**](#api-key)

## OAuth (recommended)

**OAuth is the default** if your agent supports it.

Add the config below to your MCP client. The first time it connects, your browser opens and prompts you to log in with your groundcover credentials. If you have access to more than one workspace, your agent picks the target workspace using the `list_workspaces` tool — or you can [pin a specific one](#optional-headers).

{% hint style="info" %}
**Pro tip**: You can copy a ready-to-go config from the UI.\
Go to the sidebar → Click your profile picture → **"Connect to our MCP"**
{% endhint %}

<figure><img src="/files/OfgJpBR7zEDKp6VgZTNX" alt="" width="563"><figcaption></figcaption></figure>

**Cursor / generic `mcp.json`**

```json
{
  "mcpServers": {
    "groundcover": {
      "type": "http",
      "url": "https://mcp.groundcover.com/api/mcp"
    }
  }
}
```

**Claude Code**

```bash
claude mcp add --transport http groundcover https://mcp.groundcover.com/api/mcp
```

Then run `/mcp` inside Claude Code to complete the browser login.

**Codex**

```bash
codex mcp add groundcover --url https://mcp.groundcover.com/api/mcp
codex mcp login groundcover
```

**Other clients (e.g. Claude Web)**

Point the client at the remote endpoint:

```
https://mcp.groundcover.com/api/mcp
```

## API Key

If your agent doesn't support OAuth, or if you want to connect a service account, use an API key. The client authenticates with a `Bearer` token — the key is already scoped to a tenant, and its backend is selected automatically, so the token is all you need.

### Prerequisites

1. **Service‑account API key** – create one or use an existing API Key. Learn more at [groundcover API keys](https://docs.groundcover.com/use-groundcover/api-keys).

### Configuration Example

```json
{
  "mcpServers": {
    "groundcover": {
      "type": "http",
      "url": "https://mcp.groundcover.com/api/mcp",
      "headers": {
        "Authorization": "Bearer <your_token>"
      }
    }
  }
}
```

For **Claude Code**:

```bash
claude mcp add --transport http groundcover https://mcp.groundcover.com/api/mcp \
  --header "Authorization: Bearer <your_token>"
```

To target a specific backend (when the tenant has more than one) or set your time zone, add an [optional header](#optional-headers).

## Optional Headers

A connection is always scoped to one workspace (tenant + backend), and it's resolved automatically: with **OAuth** your agent selects the workspace via `list_workspaces` (or uses your only one); with an **API key** the tenant comes from the key and a single backend is auto-selected.

To override that, add any of these headers to your config's `headers` block (or as `--header` flags). They work with both OAuth and API-key connections:

| Header          | When to use                                                                                                                               |
| --------------- | ----------------------------------------------------------------------------------------------------------------------------------------- |
| `X-Tenant-UUID` | Pin an OAuth connection to a single workspace (you must be a member). Not needed with an API key — its tenant is fixed by the key.        |
| `X-Backend-Id`  | Choose the backend when the tenant has more than one.                                                                                     |
| `X-Timezone`    | Your [IANA time zone](#how-to-find-your-time-zone) (for example `America/New_York`), so relative time windows resolve to your local time. |

For example, to pin an OAuth client to a single workspace:

```json
{
  "mcpServers": {
    "groundcover": {
      "type": "http",
      "url": "https://mcp.groundcover.com/api/mcp",
      "headers": {
        "X-Tenant-UUID": "<your_tenant_uuid>"
      }
    }
  }
}
```

**Where to find your tenant UUID and backend ID**

* Go to the sidebar → Click your **profile picture** → **"Connect to our MCP"**.
* Alternatively: **Settings → Access → API Keys** tab.

## **How to find your time zone**

| OS                 | Command                           |
| ------------------ | --------------------------------- |
| macOS              | `sudo systemsetup -gettimezone`   |
| Linux              | `timedatectl \| grep "Time zone"` |
| Windows PowerShell | `Get-TimeZone`                    |

## **Client‑specific Guides**

Depending on your client, you can usually set up the MCP server through the UI - or just ask the client to add it for you. Here are quick links for common tools:

* [Instructions for Claude Desktop](https://modelcontextprotocol.io/quickstart/user#2-add-the-filesystem-mcp-server)
* [Instructions for Claude Web](https://support.anthropic.com/en/articles/11175166-about-custom-integrations-using-remote-mcp)
* [Instructions for Cursor](https://docs.cursor.com/context/model-context-protocol#configuring-mcp-servers)
* [Instructions for Windsurf](https://docs.windsurf.com/windsurf/mcp)
* [Instructions for VS Code](https://code.visualstudio.com/docs/copilot/chat/mcp-servers#_add-an-mcp-server-to-your-workspace)


# MCP Tools Reference

The groundcover MCP server exposes tools for **querying** observability data (logs, traces, events, live entities, monitor issues, metrics, monitor definitions) and for **discovering metadata** (field names and values) before writing those queries. Most query tools accept a [gcQL](/use-groundcover/querying-your-groundcover-data/groundcover-query-language/groundcover-query-language-gcql-reference) pipeline; `query_metrics` accepts PromQL.

{% hint style="info" %}
**Tenant & backend scoping**

Each tool call runs against a single workspace (tenant and backend), resolved automatically: with OAuth your agent selects it via the `list_workspaces` tool; with an API key the tenant comes from the key and a single backend is auto-selected. When the choice is ambiguous (you have multiple workspaces, or a tenant with multiple backends), target one explicitly with the optional `tenant_uuid` / `backend_id` parameters described under [`list_workspaces`](#list_workspaces).
{% endhint %}

## Tool Overview

| Tool                                                  | Category            | Signal                | Query language   | Time-bound      |
| ----------------------------------------------------- | ------------------- | --------------------- | ---------------- | --------------- |
| [`list_workspaces`](#list_workspaces)                 | Workspace discovery | —                     | n/a              | no              |
| [`query_logs`](#query_logs)                           | Signal query        | Logs                  | gcQL             | yes             |
| [`query_traces`](#query_traces)                       | Signal query        | Traces / spans        | gcQL             | yes             |
| [`query_events`](#query_events)                       | Signal query        | Kubernetes events     | gcQL             | yes             |
| [`query_entities`](#query_entities)                   | Signal query        | Live entities         | gcQL             | no (live state) |
| [`query_issues`](#query_issues)                       | Signal query        | Monitor issue firings | gcQL             | yes             |
| [`query_metrics`](#query_metrics)                     | Signal query        | Metrics               | PromQL (4 modes) | yes             |
| [`query_monitors`](#query_monitors)                   | Signal query        | Monitor definitions   | gcQL filter      | no              |
| [`search_logs_metadata`](#search_logs_metadata)       | Metadata discovery  | Logs                  | n/a              | yes             |
| [`search_traces_metadata`](#search_traces_metadata)   | Metadata discovery  | Traces                | n/a              | yes             |
| [`search_events_metadata`](#search_events_metadata)   | Metadata discovery  | Events                | n/a              | yes             |
| [`search_metrics_metadata`](#search_metrics_metadata) | Metadata discovery  | Metrics               | n/a              | yes             |

## Workspace Discovery

### list\_workspaces

List the tenants (workspaces) and their backends that the authenticated caller can query. Call this first to discover which `tenant_uuid` and `backend_id` values to pass to the query tools.

For OAuth (member) auth, it returns every tenant you belong to; for API-key (service-account) auth, it returns the single tenant the credential is scoped to.

**Parameters**

None.

**Returns**

A list of workspaces, sorted by organization name. Each entry contains:

| Field         | Type      | Description                                                                                     |
| ------------- | --------- | ----------------------------------------------------------------------------------------------- |
| `tenant_uuid` | string    | Tenant UUID — pass as `tenant_uuid` to a query tool to target this workspace.                   |
| `org_name`    | string    | Organization name.                                                                              |
| `tenant_name` | string    | Tenant name, when set.                                                                          |
| `backends`    | string\[] | Active backend IDs for the tenant — pass one as `backend_id` when the tenant has more than one. |

**Routing query tools**

Every `query_*` and `search_*` tool accepts two optional routing parameters:

| Parameter     | Type   | Required | Description                                                                             |
| ------------- | ------ | -------- | --------------------------------------------------------------------------------------- |
| `tenant_uuid` | string | no       | Tenant to route the call to. Only needed when your account has more than one workspace. |
| `backend_id`  | string | no       | Backend to route the call to. Only needed when the tenant has more than one backend.    |

A single workspace (and single backend) is selected automatically, so you only need these when the choice is ambiguous.

## Signal Query Tools (gcQL)

### Common Parameters

All gcQL signal tools share the parameters below. `query_entities` is the exception on time: it queries live state and ignores `start` / `end` / `period`.

| Parameter | Type                       | Required | Description                                                                      |
| --------- | -------------------------- | -------- | -------------------------------------------------------------------------------- |
| `query`   | string                     | yes      | gcQL query string. Must start with a filter or `*`. Always include `\| limit N`. |
| `start`   | string (RFC3339)           | no       | Window start. Defaults to 1 hour ago.                                            |
| `end`     | string (RFC3339)           | no       | Window end. Defaults to now.                                                     |
| `period`  | string (ISO 8601 duration) | no       | Relative window, for example `PT15M`, `PT1H`, `P1D`. Defaults to `PT1H`.         |

`query_metrics` and `query_monitors` use a different shape - see their dedicated sections below.

Every gcQL `query_*` tool returns the JSON response of the executed pipeline. When the row count exactly equals the effective `limit`, the server appends a follow-up text block warning that results were truncated, plus a link to the gcQL reference - the agent should refine the query rather than paginate.

### query\_logs

Run a gcQL query against logs.

**When to use it**

* Investigate errors, warnings, or arbitrary log text for a workload, namespace, or trace.
* Aggregate log volume, error rates, or latest message per group.
* Correlate logs with traces or issues via `_from` subqueries inside `join` / `union` / `in()`.

**Examples**

Latest 10 production errors, newest first:

```
filter env:production level:error | sort by (_time desc) | limit 10
```

Error rate per workload, surfacing only workloads above 5% errors, in a single query:

```
* | stats by (env, cluster, namespace, workload)
    count() as total,
    count() if (level:error) as errors
| math errors / total * 100 as error_pct
| filter error_pct > 5
| sort by (error_pct desc)
| limit 20
```

### query\_traces

Run a gcQL query against trace spans.

**When to use it**

* Identify slow services, endpoints, or operations.
* Quantify HTTP error rates (4xx / 5xx) or span-level errors.
* Pull the spans behind a specific `trace_id` to investigate a request.

**Examples**

Slow checkout calls (>500 ms):

```
* | filter service.name:checkout | filter duration_seconds>0.5 | limit 10
```

5xx count per service:

```
* | filter http_status_code>=500 | stats by (service.name) count() as errors | sort by (errors desc) | limit 20
```

p95 latency per workload, top 5:

```
* | stats by (workload) quantile(0.95, duration_seconds) as p95
| sort by (p95 desc)
| limit 5
```

### query\_events

Run a gcQL query against Kubernetes events.

**When to use it**

* Find OOMKills, crash loops, scheduling failures, or any abnormal cluster event.
* Correlate cluster activity with logs and traces during an incident.

**Examples**

Recent OOMKilled events:

```
* | filter type:Warning reason:OOMKilled | sort by (_time desc) | limit 10
```

Top warning reasons in production:

```
* | filter env:production type:Warning
| stats by (reason) count() as count
| sort by (count desc)
| limit 20
```

### query\_entities

Run a gcQL query against the **live state** of entities tracked by groundcover. This includes Kubernetes resources (Pods, Deployments, Services, Nodes, etc.) as well as non-Kubernetes entities. Reflects the current snapshot - does **not** accept time parameters.

**When to use it**

* Inspect the spec or status of a specific resource (a Deployment, a Pod, a Node).
* Count or group entities by current state (Running / Pending / Failed).
* Discover which entity kinds and fields exist in the cluster.

**Examples**

Fetch a single deployment:

```
kind:Deployment name:chat-app | limit 1
```

Running pods, top 10:

```
kind:Pod status_phase:Running | limit 10
```

Deployment readiness summary:

```
kind:Deployment | fields name, namespace, ready_replicas, replicas | limit 20
```

Node count by status:

```
kind:Node | stats by (status) count() | limit 10
```

**Drilling into the underlying Kubernetes object**

The full Kubernetes object is exposed as discoverable `raw_json.*` paths (not as a single scalar column - `| fields raw_json` returns empty). Project specific paths and filter by them directly:

```
kind:Pod namespace:flux-system | fields name, raw_json.spec.serviceAccountName, raw_json.spec.containers | limit 5
```

The exact paths depend on the kind. For Pods the container/SA paths are `raw_json.spec.containers` and `raw_json.spec.serviceAccountName`. For Deployments and other workloads that wrap a Pod template they are `raw_json.spec.template.spec.containers` and `raw_json.spec.template.spec.serviceAccountName`. Use `kind:<EntityKind> | field_names` to enumerate every available path for that kind.

{% hint style="info" %}
**Notes**

* Field discovery for entities is **kind-scoped**: `kind:Pod | field_names`, `kind:Deployment | field_names`, etc. The available fields differ between kinds. See the [field discovery cheat sheet](#field-discovery-cheat-sheet).
* `raw_json` itself has no scalar column; it surfaces only via dotted paths like `raw_json.metadata.*`, `raw_json.spec.*`, `raw_json.status.*`.
  {% endhint %}

### query\_issues

Run a gcQL query against monitor issue instances - active alerts and historical firings produced by your monitors. Use [`query_monitors`](#query_monitors) first to find a monitor's identity, then drill into its firings here.

**When to use it**

* List active alerts in an environment or namespace.
* Identify the noisiest monitors (most firings) over a window.
* Drill from a monitor's identity into the underlying signals (with `query_logs` / `query_traces` afterward).

**Examples**

Recent issues in production:

```
* | filter env:production | sort by (_time desc) | limit 10
```

CPU-related monitors, most recent firing first:

```
* | filter monitor_name:*cpu* | sort by (_time desc) | limit 10
```

Firing count per monitor:

```
* | stats by (monitor_name) count() as firings | sort by (firings desc) | limit 20
```

{% hint style="info" %}
**Notes**

* Common fields: `id`, `monitor_id`, `monitor_name`, `state`, `previous_state`, `severity`, `cluster`, `env`, `namespace`, `workload`, `summary`, `silence_ids`, `timestamp`. Use `* | field_names` to discover the full list.
* An issue is silenced when `silence_ids` is non-empty.
  {% endhint %}

***

## query\_metrics

Query metrics from groundcover. Unlike the gcQL signal tools, `query_metrics` runs **PromQL** and exposes several modes:

| Mode            | Purpose                                                                    |
| --------------- | -------------------------------------------------------------------------- |
| `get_names`     | Discover metric names, optionally filtered by substring or required names. |
| `get_labels`    | List the label keys for a specific metric.                                 |
| `query_instant` | Execute a PromQL query at a point in time.                                 |
| `query_range`   | Execute a PromQL range query over a time window.                           |

### Parameters

| Parameter                  | Type              | Required         | Description                                                       |
| -------------------------- | ----------------- | ---------------- | ----------------------------------------------------------------- |
| `mode`                     | enum              | yes              | One of `get_names`, `get_labels`, `query_instant`, `query_range`. |
| `metricName`               | string            | for `get_labels` | The metric to inspect.                                            |
| `promql`                   | string            | for query modes  | PromQL expression. Use names verified via `get_names`.            |
| `filter`                   | string            | no               | Substring filter for `get_names` / `get_labels`.                  |
| `required`                 | string\[]         | no               | Required metric names for `get_names`.                            |
| `step`                     | string            | no               | Step interval for range queries (default `1m`).                   |
| `clusters` / `envs`        | string\[]         | no               | Optional cluster / environment filters.                           |
| `start` / `end` / `period` | see common params | no               | Time window.                                                      |
| `limit` / `skip`           | integer           | no               | Pagination for discovery modes.                                   |

### Examples

Discover CPU-related groundcover metrics:

```json
{
  "mode": "get_names",
  "filter": "groundcover_node_cpu",
  "limit": 10
}
```

Get the labels for a metric:

```json
{
  "mode": "get_labels",
  "metricName": "groundcover_node_capacity_cpum_cpu",
  "period": "PT15M"
}
```

Instant query - total CPU capacity per cluster:

```json
{
  "mode": "query_instant",
  "promql": "sum(groundcover_node_capacity_cpum_cpu) by (cluster)",
  "period": "PT5M"
}
```

{% hint style="info" %}
**Notes**

* Native groundcover metrics start with the `groundcover_` prefix and include both a description and a unit in the `get_names` response.
* Always pull names with `get_names` first - the response carries the canonical name to plug into `promql`.
* Returned timestamps are in UTC.
  {% endhint %}

***

## query\_monitors

List configured monitor definitions and their current health status. Use this to discover monitors before drilling into their firings via [`query_issues`](#query_issues).

### Parameters

| Parameter        | Type    | Required | Description                                                                                                      |
| ---------------- | ------- | -------- | ---------------------------------------------------------------------------------------------------------------- |
| `query`          | string  | no       | gcQL filter string. Supported fields: `monitor_name`, `type`. Examples: `monitor_name:*api*`, `type:prometheus`. |
| `limit` / `skip` | integer | no       | Pagination.                                                                                                      |

This tool is **not time-bound** and does not accept `start` / `end` / `period`.

### Examples

List all monitors (paged):

```
(no query)
```

Find checkout-related Prometheus monitors:

```
monitor_name:*checkout* type:prometheus
```

Find all traces-driven monitors:

```
type:traces
```

{% hint style="info" %}
**Notes**

* Returns objects shaped `{ uuid, title, type }`. Note the field-name mapping when crossing tools: the **filter** field here is `monitor_name`, but the **result** field is `title`; both carry the same value. In `query_issues`, the corresponding fields are `monitor_id` (matches `uuid`) and `monitor_name` (matches `title`).
* Use `query_monitors` for definitions, `query_issues` for instances - don't try to fetch alert history here.
  {% endhint %}

***

## Metadata Discovery Tools

These tools take a list of keywords and return matching field names with sample values. Use them to discover what fields exist before writing a `query_*` call.

### Common Parameters

| Parameter  | Type                       | Required | Description                                                                         |
| ---------- | -------------------------- | -------- | ----------------------------------------------------------------------------------- |
| `keywords` | string\[]                  | yes      | Search terms matched against field names and values. At least one keyword required. |
| `limit`    | integer                    | no       | Max results to return. Defaults to 1000.                                            |
| `start`    | string (RFC3339)           | no       | Window start. Defaults to 1 hour ago.                                               |
| `end`      | string (RFC3339)           | no       | Window end. Defaults to now.                                                        |
| `period`   | string (ISO 8601 duration) | no       | Relative window. Defaults to `PT1H`.                                                |

### Field Discovery Cheat Sheet

Pick the right discovery method based on what you're querying:

| To discover...             | Use                                                                                                                          |
| -------------------------- | ---------------------------------------------------------------------------------------------------------------------------- |
| Log fields                 | `search_logs_metadata`                                                                                                       |
| Trace / span fields        | `search_traces_metadata`                                                                                                     |
| Event fields               | `search_events_metadata`                                                                                                     |
| Metric names and labels    | `search_metrics_metadata` (or `query_metrics` with `mode:get_names` / `mode:get_labels`)                                     |
| Entity fields (live state) | `kind:<EntityKind> \| field_names` inside `query_entities` (kind-scoped)                                                     |
| Issue fields               | `* \| field_names` inside `query_issues`                                                                                     |
| Monitor definitions        | `query_monitors` (filter on `monitor_name` / `type`; results return `title` and `uuid` - same values, different field names) |

### search\_logs\_metadata

Search log field names and sample values.

**When to use it**

* The agent doesn't know the exact log attribute name (`workload` vs `k8s.workload.name`, etc.).
* The user mentions a domain term ("checkout", "tenant", "request\_id") and the agent needs to find which fields surface it.

**Example call**

```json
{
  "keywords": ["checkout"],
  "period": "PT1H"
}
```

### search\_traces\_metadata

Search span attribute keys and sample values.

**When to use it**

* Discover trace attributes for a service (`service.name`, `http.path`, `http_status_code`, etc.).
* Find which span attributes carry a particular business identifier.

**Example call**

```json
{
  "keywords": ["http", "status"],
  "period": "PT15M"
}
```

### search\_events\_metadata

Search Kubernetes event fields and sample values.

**When to use it**

* Look up event reasons (`OOMKilled`, `BackOff`, `FailedScheduling`) before filtering.
* Find the correct entity / object fields for a specific resource kind.

**Example call**

```json
{
  "keywords": ["OOMKilled"],
  "period": "P1D"
}
```

### search\_metrics\_metadata

Search metric names and their label keys / values.

**When to use it**

* Quick keyword-based discovery of what metrics exist.
* For more structured discovery (filtered name lists, label keys for a single metric), prefer `query_metrics` with `mode:get_names` or `mode:get_labels`.

**Example call**

```json
{
  "keywords": ["cpu", "container"],
  "period": "PT1H"
}
```


# Getting-started Prompts

Once your MCP server is connected, you can dive right in.

Here are a few prompts to try. They work out of the box with agents like Cursor, Claude, or VS Code:

> 💡 Starting your request with “**Use groundcover**” is a helpful nudge - it pushes the agent toward MCP tools and context.

## Basic Prompts to Try

MCP supports complex, multi-step flows, but starting simple is the best way to ramp up.

### Pull Logs

**Prompt:**

{% code overflow="wrap" %}

```
Use groundcover to get 5 logs from the workload news-service from the past 15 minutes.
```

{% endcode %}

**Expected behavior:**\
The agent should call `query_logs` and show recent logs for that workload.

### Inspect a K8s Resource

**Prompt:**

{% code overflow="wrap" %}

```
Use groundcover to show the current state of the chat-app deployment.
```

{% endcode %}

**Expected behavior:**\
The agent should call `query_entities` with `kind:Deployment name:chat-app | limit 1` and return the deployment's live-state fields (replicas, status, labels, etc.) or a summary of them.

### Find Slow Workloads

**Prompt:**

{% code overflow="wrap" %}

```
Use groundcover to show the top 5 workloads by P95 latency.
```

{% endcode %}

**Expected behavior:**\
The agent should call `query_traces` with a query like `* | stats by (workload) quantile(0.95, duration_seconds) as p95 | sort by (p95 desc) | limit 5` and return the top 5 workloads.

<figure><img src="/files/WWDytCbOkhriBR8BrX46" alt=""><figcaption></figcaption></figure>

## Investigate Issues

When something breaks, your agent can help investigate and make sense of it.

### Paste an Issue Link

**Prompt:**

{% code overflow="wrap" %}

```
I got an alert for this critical groundcover issue. Can you investigate it?
https://app.groundcover.com/monitors/issues?...
```

{% endcode %}

**Expected behavior:**\
The agent should use `query_issues`, pull issue details, and kick off a relevant investigation using logs, traces, and metadata.

### Investigate Multiple Issues

**Prompt:**

{% code overflow="wrap" %}

```
I got multiple alerts in the staging-env namespace. Can you help me look into them using groundcover?
```

{% endcode %}

**Expected behavior:**\
The agent should call `query_issues` with a scoped query like `* | filter namespace:staging-env | sort by (_time desc) | limit 20` and walk through the returned issues one by one.

## Automate Coding & Debugging

groundcover’s MCP can also be your coding sidekick.\
Instead of digging through tests and logs manually, deploy your changes and let the agent take over.

### Iterate Over Test Results

**Prompt:**

{% code overflow="wrap" %}

```
Use groundcover to debug this code. For each test, print relevant logs with test_id, and dive into any error logs.
```

{% endcode %}

**Expected behavior:**\
The agent should update the code with log statements, deploy it, and use `query_logs` to trace and debug.

### Deploy & Verify

**Prompt:**

{% code overflow="wrap" %}

```
Please deploy this service and verify everything works using groundcover.
```

{% endcode %}

**Expected behavior:**\
The agent should assist with deployment, then check for issues, error logs, and traces via groundcover.


# Real-world Use Cases

These are patterns we've seen in the wild. Agents use groundcover to debug, monitor, and close the loop.

### Test → Logs → Fix

Cursor generates tests, tags each with a `test_id`, logs them, and then uses groundcover to instantly fetch related log lines.

### Investigate Issues via Cursor

Got a monitor firing? Drop the alert into Cursor. The agent runs a quick RCA, queries groundcover, and even suggests a patch based on recent logs and traces.

### **Support Workflow**

Support rep gets an error ID → uses MCP to query groundcover → jumps straight to the root cause by exploring traces and logs around the error.

### **The Autonomous Loop**

An agent picks up a ticket, writes tests, ships code to staging, monitors it with groundcover, checks logs and traces, and verifies the fix end to end.\
Yes, really. Full loop. Almost no hands.


# TCO Calculator

The [groundcover Cost Calculator](https://www.groundcover.com/calculator) is a tool designed to help you estimate the total cost of setting up and running a managed groundcover backend, deployed in your own cloud account via [console.groundcover.com](https://console.groundcover.com).

It is intended to give a realistic, data-driven estimate before you begin, so you can plan your infrastructure spend and subscription costs with confidence.

## What the calculator estimates

The calculator reflects the **Total Cost of Ownership (TCO)** for groundcover. This consists of two components:

* **Subscription cost**: based on the number of nodes monitored and the plan you select.
* **Infrastructure cost**: the cloud resources provisioned in your account to run the groundcover backend (compute, storage, networking, managed services, etc.).

Both components are combined into a single yearly estimate.

## Input parameters

### Nodes

The number of Kubernetes nodes or Virtual Machines you plan to monitor. This is the only driver of subscription cost, and also influences the size of the backend infrastructure needed to ingest and store your telemetry.

### Logs

The estimated daily volume of log data ingested, typically measured in GB/day. Higher log volumes require more storage, processing capacity, and networking, which increases infrastructure costs.

### Traces

The estimated daily volume of trace data ingested, measured in GB/day. Trace volume affects the compute, storage, and networking resources needed for the tracing backend.

### Metrics

The number of active time series (also called cardinality) across all your monitored environments. Higher metric cardinality increases the memory, storage, and networking requirements of the metrics backend.

### Retention

How long you want to retain each data type (logs, traces, metrics). Longer retention periods directly increase storage costs, since more data must be kept on disk or in object storage.

### Region

The cloud region where your BYOC backend will be deployed. Infrastructure pricing varies by region across all cloud providers, so selecting your intended region produces a more accurate estimate.

## Frequently asked questions

**Does the price include savings plans or Enterprise Discount Programs (EDPs) on the backend?**

No. The calculator shows list price. Most customers benefit from additional savings based on their own cloud discounts (e.g. AWS EDP, GCP CUD, Azure reservations). Your actual cost is likely lower once those are applied.

**Does the price include tax?**

No. The estimate reflects expected usage and resource costs only, before any applicable taxes.

**How is this estimate calculated?**

The calculator is based on historical data from hundreds of customer environments, ranging from early-stage startups to Fortune 500 companies. It takes into account groundcover's specific architecture and the resource patterns observed across those deployments to produce a realistic, grounded estimate rather than a theoretical one.

**Why am I seeing price ranges instead of a single number?**

Real-life costs vary based on actual volumes and usage patterns. Rather than showing a single point estimate that may be misleading, the calculator displays a range that more accurately reflects this variance.

**Does it work for all cloud providers (AWS, GCP, Azure)?**

Yes. Expected costs are similar across all major cloud providers. The region selector accounts for regional pricing differences, and the underlying model applies equally to AWS, GCP, and Azure deployments.


# Migrations

### Overview

When switching to a new observability platform, there is the concern of high effort of switch due to the many assets already created in the previous platform.

groundcover's migration flow discovers, translates, and installs your observability setup into groundcover.

### Getting Started

Navigate to the [migrations page](https://app.groundcover.com/migrations) to get started. Note that Admin role is required to access this page.

Select the relevant platform to migrate from and get started.

### How it works

1. **Fetch & discover:** Provide API keys. groundcover pulls your monitors, dashboards, and data sources.
2. **Automatic translation:** All the pulled assets are automatically being converted to groundcover assets within the migration page.
3. **Migrate assets:** Review and migrate missing data sources, monitors and dashboards with one click.

API keys are not stored.

### What we migrate

#### Monitors

Includes alert conditions, thresholds, and evaluation windows.

#### Dashboards

Complete dashboard migration with preserved layouts:

* All widget types groundcover supports
* Query translations
* Time ranges and filters
* Visual settings and arrangements

#### Data Sources

Based on the metrics in use, we detect which data sources are missing and help you set it up in groundcover.

#### Data & Mapping

We don't just copy configurations - we ensure the data flows:

* **Automatic metric mapping:** metric names translated to groundcover equivalents
* **Label translation:** Tags become labels with intelligent mapping
* **Query conversion:** query syntax of the source platform converted to groundcover
* **Data validation:** We verify all referenced metrics and data sources exist

### Supported providers

#### Datadog

Full migration support available now.

[Migrate from Datadog →](/getting-started/migrations/migrate-from-datadog)

#### Other providers

Additional vendors coming soon.


# Migrate from Datadog

Complete guide for migrating your Datadog setup to groundcover.

### Prerequisites

**Access required:**

* Admin role in groundcover
* Datadog API key and application key with read permissions

**Datadog application key scopes:**

* [`dashboards_read`](https://docs.datadoghq.com/account_management/rbac/permissions/#dashboards) - List and retrieve dashboards
* [`monitors_read`](https://docs.datadoghq.com/account_management/rbac/permissions/#monitors) - View monitors
* [`metrics_read`](https://docs.datadoghq.com/account_management/rbac/permissions/#metrics) - Query timeseries data
* [`integrations_read`](https://docs.datadoghq.com/account_management/rbac/permissions/#integrations) - View AWS, GCP, Azure integrations

### Create a migration project

Navigate to [**Migrations**](https://app.groundcover.com/migrations) from the main page.

<figure><img src="/files/5NjGNm6t4x6kHgz9V6Ai" alt=""><figcaption></figcaption></figure>

1. Click **Start** on the Datadog card
2. Enter a project name (e.g., "Production Migration", "US5 Migration")
3. Click **Create**

**Tip:** Use descriptive names. You can run multiple migration projects for different environments or teams.

### Fetch assets from Datadog

<figure><img src="/files/OFCx99vbOucLnD6fTi48" alt=""><figcaption></figcaption></figure>

Provide your Datadog credentials:

#### Datadog site

The domain of your Datadog console. The options are:

* `US1` - app.datadoghq.com
* `US3` - us3.datadoghq.com
* `US5` - us5.datadoghq.com
* `EU1` - app.datadoghq.eu
* `AP1` - ap1.datadoghq.com

You can find your Datadog site by looking at your console's URL.

#### API key

A regular Datadog API key. Find this under **Organization Settings →** [**API Keys**](https://docs.datadoghq.com/account_management/api-app-keys/#api-keys).

#### Application key

Create one under **Organization Settings →** [**Application Keys**](https://docs.datadoghq.com/account_management/api-app-keys/#application-keys) with the required scopes listed above.

{% hint style="info" %}
**Important:** groundcover does not store these keys. Assets are fetched, then the keys are discarded.
{% endhint %}

Click **Fetch Assets**. This typically takes 10 seconds depending on the number of assets.

### Review migration summary

<figure><img src="/files/LVyZulMOQxT34sr8NZSu" alt=""><figcaption></figcaption></figure>

Once fetched, you see:

* **Progress overview:** Total assets discovered and their support status
* **Asset cards:** Monitors, Dashboards, Data Sources, etc
* **Support breakdown:** How many assets are fully supported, partial, or unsupported

The overview shows everything we found in your Datadog account and what we'll bring over.

### Migrate data sources

Start with adding missing data sources to ensure relevant data is in place and ease the evaluation process of monitors and dashboards.

#### What we detect

Based on the full set of metrics in use by your dashboards and monitors, we map the list of data sources that should be added.

For each, we show the number of monitors and dashboards that rely on this data source.

<figure><img src="/files/lpgJqXq5O5rIbxt2d4mg" alt=""><figcaption></figcaption></figure>

Clicking on any data source will show the list of monitors and dashboards to provide more context for users.

Clicking on connect data source will show a wizard explaining how to connect a data source. Users will need to provide the connection details for the data source. Whenever data mapping is required, it will already be populated. For example, in case there is a need to connect cloudwatch, only the user AWS namespaces will be selected by default.

### Migrate monitors

<figure><img src="/files/pD5k4Zdix95YOivWWv91" alt=""><figcaption></figcaption></figure>

Once data sources are ready, migrate your monitors.

#### Monitor status indicators

* **✓ Supported:** Fully compatible. Migrate as-is.
* **⚠ Partial:** Migrates with warnings. Review before installing.
* **✗ Unsupported:** Requires manual attention.

#### Review warnings

For monitors with warnings, click **View Warnings**:

* See what adjustments were made
* Understand query translations
* Get recommendations for post-migration verification

Warnings don't block migration — they inform you of changes so you can verify behavior.

#### Migrate monitors

**Single monitor:**

1. Preview the monitor
2. Click **Migrate**
3. Monitor installs immediately

**Bulk migrate:**

1. Select multiple monitors using checkboxes
2. Click **Migrate Selected**
3. All install in parallel

Migrated monitors appear instantly in **Monitors → Monitor List**.

### Migrate dashboards

<figure><img src="/files/YUxP1swmhckiN1Aqwaq1" alt=""><figcaption></figcaption></figure>

Dashboards preserve:

* Layout and widget positions
* Query logic and filters
* Time ranges and visualization settings
* Colors and formatting

Check out the dashboard preview to confirm the migration worked and that all your assets came through successfully.

#### Migrate dashboards

Click **Migrate** to install. Dashboards appear under **Dashboards** immediately.

**Tip:** Migrate critical dashboards first. Verify queries return expected data before bulk migrating.


# Monitors

Monitors offers the ability to define custom alerts, which you can configure using groundcover data and custom metrics.

## What is a Monitor

A `Monitor` defines a set of rules and conditions that track the state of your system. When a monitor's conditions are met, it triggers an issue that is displayed on the [Issues page](/use-groundcover/monitors/issues-page) and can be used for alerting using your [integrations](/integrations/workflow-integrations) and [workflows](/use-groundcover/workflows).

Easily create a new Monitor by [using our guide](/use-groundcover/monitors/create-a-new-monitor).

<figure><img src="/files/T5dfubNX58zpl7uRafye" alt=""><figcaption><p>Example of a groundcover Monitor</p></figcaption></figure>


# Create a new Monitor

Learn how to create and configure monitors using the Wizard, Monitor Catalog, or Import options. The following guide will help you set up queries, thresholds, and alert routing for effective monitoring.

> You can either create monitors using our web application following this guide, or use our API, see: [Monitors](/use-groundcover/monitors) or use our Terraform provider, see: [groundcover Terraform Provider](/use-groundcover/groundcover-terraform-provider).

In the Monitors section (left navigation bar), navigate to the **Issues** page or the **Monitor List** page to create a new Monitor. Click on the “Create Monitor” button at the top right and select one of the following options from the dropdown:

<table data-view="cards"><thead><tr><th></th><th data-hidden data-card-cover data-type="files"></th><th data-hidden data-card-target data-type="content-ref"></th></tr></thead><tbody><tr><td><a href="#using-the-monitor-wizard"><strong>Monitor Wizard</strong></a></td><td></td><td><a href="https://docs.groundcover.com/use-groundcover/monitors/create-a-new-monitor#using-the-monitor-wizard">https://docs.groundcover.com/use-groundcover/monitors/create-a-new-monitor#using-the-monitor-wizard</a></td></tr><tr><td><a href="/pages/mcvychJhqFoZCh029eDF"><strong>Monitor Catalog</strong></a></td><td></td><td><a href="/pages/mcvychJhqFoZCh029eDF">/pages/mcvychJhqFoZCh029eDF</a></td></tr><tr><td><a href="#using-the-import-option"><strong>Import</strong></a></td><td></td><td><a href="https://docs.groundcover.com/use-groundcover/monitors/create-a-new-monitor#using-the-import-option">https://docs.groundcover.com/use-groundcover/monitors/create-a-new-monitor#using-the-import-option</a></td></tr></tbody></table>

## Using the Monitor Wizard

### Overview

The Monitor Wizard is a guided, user-friendly approach to creating and configuring monitors tailored to your observability needs. By breaking down the process into simple steps, it ensures consistency and accuracy.

### Section 1: Query

Select the data source, build the query and define thresholds for the monitor.

{% hint style="info" %}
If you're unfamiliar with query building in groundcover, refer to the [Query Builder section](/use-groundcover/querying-your-groundcover-data/explore-and-monitors-query-builder) for full details on the different components.
{% endhint %}

* **Data Source (Required):**
  * Select the type of data (Metrics, Logs, Traces, Events, APM, or Ingestion).
  * **Ingestion** monitors how much data groundcover ingests (volume in bytes or entry count, for logs or traces). Useful for alerting on sudden ingestion drops or spikes. See [Ingestion](/use-groundcover/querying-your-groundcover-data/explore-and-monitors-query-builder#ingestion) in the Query Builder for the available signals and options.
* **Query Functions:**
  * Choose how to process the data (e.g., average, count).
  * Add aggregation (group by) clauses if applicable, you MUST use aggregations if you want to add labels to your issues.
  * **Important**: The labels used for aggregation (group by) maybe also be used for notification routes and the issue summary & description.
  * **Examples:** `cluster`, `node`, `container_name`
* **Time Window (Required):**
  * Specify the period over which data is aggregated (the look-behind window).
  * **Example:** “Over the last 5 minutes.”
* **Window Aggregation (Required):**
  * Specify the aggregation function to be used on the selected time window.
  * **Example:** "`avg` over the last 5m"
* **Threshold Conditions (Required):**
  * Define when a monitor should trigger an Issue. You can use:
    * Greater Than - Trigger when the value exceeds X.
    * Lower Than - Trigger when the value falls below X.
    * Within Range - Trigger when the value is between X and Y.
    * Outside Range - Trigger when the value is not between X and Y.
  * **Important:** The units in which the threshold is being measured in must be the same as the units the query uses.
    * For metrics queries the threshold should match the unit the metric is measured in.
    * For APM queries the threshold should match the selected metric's unit (e.g. seconds for latency, % for Error Rate, a plain number for request/error counts).
    * For logs, traces, and events, it's just a number.
  * **Example:** “Trigger if disk space usage is greater than 10%.”
* **Preview Settings (Optional):**

  * Preview data using Stacked Bar or Line Chart for better clarity while building the monitor.
  * Choose the Y axis units.
  * Choose the rollup to present.
  * **Important**: These configurations only affect the preview graph, not the monitor's evaluation.

  <figure><img src="/files/e0ZBFRRCmDtpNoVyREUT" alt=""><figcaption></figcaption></figure>
* **Advanced (Optional):**

  * **Evaluation Interval:**
    * Specify how often the monitor evaluates the query
    * Example: “Evaluate every 1 minute.”
  * **Pending Period:**
    * Specify how many times the evaluation needs to pass the threshold in order to trigger an Issue. This refers to a consecutive evaluations passing the threshold.
    * Monitors that have entered the pending period (the first evaluation passed the threshold) will be in 'Pending' state, only after all consecutive evaluations passed the threshold, the monitor will be 'Firing' and an Issue will be created. If even 1 of the evaluations did not pass the threshold, the Monitor will be set right back to 'Normal'.
    * **Example**: “When Evaluation Interval of 5m, setting this to 2 (10m) ensures the condition must be evaluated 3 times before a monitor will fire.
      * Evaluation #1 at 0m
      * Evaluation #2 at 5m
      * Evaluation #3 at 10m -> If all 3 passed the threshold, an the monitor will 'Fire'
    * **Note**: This ensures that transient conditions do not trigger alerts, reducing false positives or smoothing sudden unwanted spikes.
    * **Important**: The default configuration is 0, which means the monitor will trigger an Issue immediately when an evaluation was run and the threshold was passed.
  * **Evaluation delay:**
    * Evaluates the query against a window that ends this many seconds in the past instead of "now". Useful for sources that backfill recent data (like AWS CloudWatch and GCP), which can otherwise cause false alerts.
    * Input is in seconds (0–3600). Leave it empty or `0` for no delay (the default). The monitor shifts its whole evaluation window back by the delay without changing your query, and the Preview graph reflects the same shift.
    * For new metric monitors, groundcover prefills a recommended delay from the metric name (`aws_` → 15m, `gcp` → 5m), which you can override or clear.
  * **Treat No Data As**:
    * Controls what the monitor does when its query returns **no data** (an empty result set) on an evaluation. Choose one of:
      * **Normal** — Treated as healthy. No Issue is created and no notification is sent; the monitor stays green.
      * **Firing** — Creates an Issue (and sends notifications, if a notification route matches), exactly like a threshold breach. Choose this to be alerted when data stops arriving — see [Alerting on No Data](/use-groundcover/monitors/notification-routes#alerting-on-no-data).
      * **No Data** — Sets the monitor's status to **No Data**, shown on the [Monitor List](/use-groundcover/monitors/monitor-list-page) and in the monitor's timeline and facets. No Issue is created and no notification is sent: No Data is a monitor **status only, not an Issue**.
    * **Example**: "I want to be notified if the metric has a gap for the entire look-behind window of the query, so I will set it to 'Firing'."
    * **No data vs. a zero value**: "No data" means the query returned **no rows**, not a row whose value is `0`. For example, `workload:agent | stats count() as cnt` returns a single row with `cnt = 0` when nothing matches — that is data (a zero value), so it is **not** treated as No Data. To make the query return no rows when there are no matches, drop the zero row with a filter:

      ```gcql
      workload:agent | stats count() as cnt | filter cnt > 0
      ```

  <figure><img src="/files/DnIm0Z9ZnZBfQACUoWbp" alt="Monitor advanced evaluation settings, including Treat No Data As"><figcaption></figcaption></figure>

### Section 2: Monitor Details

Set up the basic information for the monitor.

* **Monitor Name (Required):**
  * Add a title for the monitor. The title will appear in notifications and in the [Monitor List page](/use-groundcover/monitors/monitor-list-page).
  * Give the Monitor a clear, short name, that describes its function **at a high level**.
  * **Examples:**
    * `“Workload High API Error Rate”`
    * `“Node High Memory”`

{% hint style="info" %}
The title will appear in the monitors page table and be accessible in notification routes.
{% endhint %}

* **Severity (required):**
  * Use severity to categorize alerts by importance.
  * Select a severity level (S1-S4).
  * **Important**: For Destinations (OpsGenie, PagerDuty) that require specific severities like P1-P4 or Critical-Info, we translate automatically to the relevant respective severity.
* **Custom Labels (formally called 'metadata labels'):**

  * Add custom labels (key:value) that will be added to all Issues generated by this monitor
  * Example: To create a notification route for my team's issues, add "Team:Infra" and use it in the notification route's scope

  <figure><img src="/files/oOBjPavAnVXb7tXAX4lI" alt=""><figcaption></figcaption></figure>

### Section 3: Issue Details

Customize how the Monitor’s Issues will appear and what content will be sent in it's notifications. This section also includes a live preview of the way it will appear in the notification.

{% hint style="info" %}
Ensure that the labels you wish to use dynamically (e.g., `cluster`, `workload`) or statically (e.g. `team:infra`) are defined in the query and monitor details.
{% endhint %}

* **Issue Summary (required):**

  * Define a short title for issues that this Monitor will raise. It's useful to include variables that can be informative at first glance.
  * **Example:** Adding `{{ labels.statusCode }}` to the header will inject the status code to the name of the issue - this becomes especially useful when one Monitor raises multiple issues and you want to quickly understand their content without having to open each one.
    * `“HTTP API Error {{ labels.status_code }}”` -> `HTTP API Error 500`
    * `“Workload {{ labels.workload }} Pod Restart”` -> `Workload frontend Pod Restart`
    * `“{{ labels.team }} APIs High Latency”` -> `Infra APIs High Latency`
  * **Note:** Autocomplete is supported to view what is usable in the Issue and will help ensure you put in the variables correctly.

  <figure><img src="/files/nErm9QuJlvgHU2qQwbyk" alt="" width="375"><figcaption></figcaption></figure>

{% hint style="warning" %}
The new format for templating variables is `{{ variable }}` or `{{ labels.<label> }}` , but the previous format `{{ alert.labels.statusCode }}` used for keep is still supported.
{% endhint %}

* **Description:**
  * Used as the body of the message for the Issue.
  * The default templating uses simple `{{ variable }}` and `{{ labels.<label> }}` substitution. Opt in to full Jinja2 by setting `display.templateLanguage: jinja2` in the monitor YAML.
    * **Example**: Adding all the labels to be shown in the Slack message's body should be inserted into here using `{{labels.<label>}}` , you can add the severity `{{severity}}`, the monitor's name `{{monitor_name}}` and many more.
  * URLs can be rendered using <`url link`|`url title`>
    * **Example:** `<www.groundcover.com/{{labels.env}}|text>` will add the env label from the issue to the URL link and put the link inside a text called 'text'

{% hint style="info" %}
Full Jinja2 features like `{% if %}` blocks and filters require `display.templateLanguage: jinja2` in the monitor YAML. Without that field, only plain `{{ variable }}` substitution is supported.
{% endhint %}

<figure><img src="/files/aEFO2GgzaF9LiVCswUV8" alt=""><figcaption></figcaption></figure>

* **Advanced (Optional)**
  * **Display Labels (formally called 'context labels'):**

    * These Labels will be displayed and filterable in Monitors>Issues page.
    * This list gets automatically populated based on the labels used in the aggregation function in the Query.
    * **Note:** You can remove labels from this list if you do not wish to see them in the Issues page.

    <figure><img src="/files/Hkknf6340lYX4GtCsHT6" alt=""><figcaption></figcaption></figure>

### Section 4: Notifications

Set up notifications behavior for issues from this monitor.

{% hint style="info" %}
Workflows (used for Keep) and Notification Routes may work in parallel and do not affect each other.
{% endhint %}

Choose one of three notification methods:

* **Based on matching notification routes** (default)
  * The issues generated by this monitor will be evaluated by the [Notification Route's](/use-groundcover/monitors/notification-routes) scopes and rules and notifications will be sent accordingly.
  * **Note**: The **Preview** can be used in order to align expectations on which notification routes may match this monitor's future issues.
* **Directly to destinations**
  * Send notifications directly to specific Destinations for this monitor only, bypassing notification routes.
  * Click **Set rule**, choose an **issue status filter** (Firing, Resolved, or both), then in the **Send to** dropdown select the Destination.
  * For a [Slack App](/use-groundcover/connectors/slack) Destination, the dropdown opens a drilldown panel where you pick the target Slack channel. The bot can post to any public channel in the workspace without being added; private channels only appear in the picker if the bot has been invited to them.
  * Click **+** to add more destinations — additional Slack channels (same Slack App), other Slack Apps, or any other Destination.
  * **Note**: Issues from this monitor will be skipped by all matching notification routes when this method is selected.
* **Don't send notifications**
  * Suppresses all notifications for this monitor. The monitor still evaluates and creates issues, but no Destinations or notification routes are notified.

{% hint style="info" %}
The wizard labels issue states as **Firing** and **Resolved**. The underlying YAML schema uses the enum values `Alerting` and `Resolved` (see [Monitor YAML structure](/use-groundcover/monitors/monitor-yaml-structure#notificationsettings)) — `Firing` in the UI is the same state as `Alerting` in YAML.
{% endhint %}

To set up the Slack App connector itself, see [Slack](/use-groundcover/connectors/slack).

* **Routing (Workflows) (Optional)**
  * **Select Workflow:**
    * Route alerts to existing workflows **only**, this means that other workflows will not process them. Use this to send alerts for a critical application such as Slack or PagerDuty.
  * **No Routing:**
    * This means that any workflow (without filters), will process the issue.

{% hint style="warning" %}
Configure Routing for Keep Workflows only
{% endhint %}

<figure><img src="/files/TsBGk4mF62FpvHvGXLKV" alt=""><figcaption></figcaption></figure>

* **Advanced (Optional)**

  * Override Renotify Interval
    * Used to override the interval configured on the Notification Route for when a certain monitior's issue should send another notification at a different inerval.
    * **Example**: If it's set to 1m while the evaluation interval is 1m a notification will be sent with every firing evaluation. If it's set to 2d, even if the monitor evaluates every 1m, a notification will be sent once every 2 days.
    * **Note**: If the Issue stops firing and starts firing again, a new notification will be sent, this is not considered 'renotification'
    * **Important**: Minimum interval is the evaluation interval

  <figure><img src="/files/oqirgxQISGJue9GaVW6d" alt=""><figcaption></figcaption></figure>

## Using the Import option

{% hint style="warning" %}
This is an advanced feature, please use it with caution.
{% endhint %}

<figure><img src="/files/MyzdTEuJwU5RXb3OEPtt" alt=""><figcaption></figcaption></figure>

In the "Import Bulk Monitors" you can add multiple monitors using an array of Monitors that follows the [Monitor YAML structure](/use-groundcover/monitors/monitor-yaml-structure).

Example of importing multiple monitors

```
monitors:
- title: K8s Cluster High Memory Requests Monitor
  display:
    header: K8s Cluster High Memory Requests
    description: Alerts when a K8s Cluster's total Container Memory Requests exceeds 90% of the Allocatable Memory of all the Nodes for 5 minutes    
    contextHeaderLabels:
      - env
      - cluster
  severity: S1
  measurementType: state
  model:
    queries:
      - name: threshold_input_query
        expression: avg_over_time( (((sum(groundcover_node_rt_mem_requests_bytes{}) by (cluster, env)) / (sum(groundcover_node_rt_allocatable_mem_bytes{}) by (cluster, env))) * 100)[5m] )
        queryType: instant
        datasourceType: prometheus
    thresholds:
      - name: threshold_1
        inputName: threshold_input_query
        operator: gt
        values:
          - 90
  noDataState: OK
  evaluationInterval:
    interval: 1m
    pendingFor: 0s
- title: K8s PVC Pending For 5 Minutes Monitor
  display:
    header: K8s PVC Pending Over 5 Minutes
    description: This monitor triggers an alert when a PVC remains in a Pending state for more than 5 minutes.
    contextHeaderLabels:
      - cluster
      - namespace
      - persistentvolumeclaim
  severity: S2
  measurementType: state
  model:
    queries:
      - name: threshold_input_query
        expression: last_over_time(max(groundcover_kube_persistentvolumeclaim_status_phase{phase="Pending"}) by (cluster, namespace, persistentvolumeclaim)[1m])
        queryType: instant
        datasourceType: prometheus
    thresholds:
      - name: threshold_1
        inputName: threshold_input_query
        operator: gt
        values:
          - 0
  executionErrorState: OK
  noDataState: OK
  evaluationInterval:
    interval: 1m
    pendingFor: 5m
```

Click on "**Create Monitors**" to create them.

## Query Best Practices

### Performance Recommendations

To ensure reliable monitor evaluation and avoid timeouts:

* **Avoid excessively long time ranges** with free-text search queries. Maximum recommended range is 7 days for most query types.
* **Use attribute filters instead of free-text search** when possible - free-text searches across large time ranges are expensive.
* **Set appropriate dashboard refresh intervals** - avoid refreshing complex dashboards every 1-2 minutes with week-long queries.
* **Consider logs-to-metrics** for aggregation queries that would otherwise scan large log volumes.
* **Use parsing rules** to add attributes for frequently-filtered fields instead of relying on free-text search.

### MetricsQL Query Limitations

When using MetricsQL in monitors:

**Regex Operators**: The regex not-match operator `!~` may have limitations with certain patterns:

* Pattern `.+` (one or more characters) may not work as expected in some cases
* Use `.*` (zero or more characters) as an alternative: `label!~".*Error"` instead of `label!~".+Error"`

**Switching between data sources**: Changing the data source type (e.g., from Metrics to Logs) will clear the current query.

### gcQL Query Limitations in Monitors

Some gcQL operations behave differently in monitors compared to the Data Explorer:

| Operation                          | Data Explorer | Monitors                                   |
| ---------------------------------- | ------------- | ------------------------------------------ |
| `grok` parsing                     | Full support  | Supported (fixed in recent versions)       |
| `join` operations                  | Full support  | Limited - complex joins may timeout        |
| Free-text search on message fields | Works         | May have limitations in Explore filter bar |

If your query works in Data Explorer but fails in monitors, try simplifying the query or breaking it into multiple monitors.

## Troubleshooting

### Monitor Not Firing

1. **Check the preview graph** - Verify the query returns data and exceeds the threshold
2. **Review pending period** - If set, the condition must be met for multiple consecutive evaluations
3. **Check "No Data" handling** - If data is intermittent, "Treat No Data As: Firing" may cause unexpected behavior
4. **Verify aggregation labels** - Ensure "group by" labels match what you expect

### Labels Not Appearing in Notifications

1. **Labels must be in "group by"** - Only labels used in aggregation are available as `{{ labels.<name> }}`
2. **Check variable syntax** - Use `{{ labels.workload }}` not `{{ alert.labels.workload }}` (though both are supported)
3. **Preview the notification** - Use the live preview to verify variable expansion

### Query Works in Explorer but Not in Monitor

1. **Time alignment** - Explorer and monitors may use different time window handling
2. **Aggregation function order** - For metrics like `sum(sum_over_time(...))`, ensure the step and aggregation window match to avoid inflated values
3. **gcQL limitations** - Some advanced operations may not be fully supported in monitor queries

### Dashboard Shows Different Values Than Monitor

This can occur due to:

* **Step vs aggregation window mismatch** - Using `sum_over_time(...[5m])` with step=1m creates a sliding window that can inflate values. Align step with aggregation window.
* **Different aggregation functions** - Verify both use the same aggregation
* **Time range differences** - Monitors use a specific look-behind window


# Issues page

View and analyze monitor issues with detailed timelines, metadata, and context to quickly identify and resolve problems in your environment.

The Issues page provides a detailed view of active and resolved issues triggered by Monitors. This page helps users investigate, analyze, and resolve problems in their environment by visualizing issue trends and providing in-depth context through an issue drawer.

## Issues drawer

Clicking on an issue in the Issues List opens the Issue drawer, which provides an in-depth view of the Monitor and its triggered issue. You can also navigate if possible to related entities like workload, node, pod, etc.

<figure><img src="/files/JOFWutugFMTwUfXnwBwG" alt=""><figcaption></figcaption></figure>

### Details tab

Displays metadata about the issue, including:

* **Monitor Name:** Name of the Monitor that triggered the issue, including a link to it.
* **Description:** Explains what the Monitor tracks and why it was triggered.
* **Severity:** Shows the assigned severity level (e.g., S3).
* **Labels:** Lists contextual labels like `cluster`, `namespace`, and `workload`.
* **Creation Time:** Shows when the issue started firing.

### **Events tab**

Displays the Kubernetes events related to the selected issue within the timeframe selected in the Time Picker dropdown (upper right of the issue drawer).

### **Traces tab**

When creating a Monitor using a traces query, the Traces tab will display the matching traces generated within the timeframe selected in the Time Picker dropdown (upper right of the issue drawer). Click on "View in Traces" to navigate to the Traces section with all relevant filters automatically applied.

### **Logs tab**

When creating a monitor using a log query, the Logs tab will display the matching logs generated within the timeframe selected in the Time Picker dropdown (upper right of the issue drawer). Click on "View in Logs" to navigate to the Logs section with all relevant filters automatically applied.

### **Map tab**

A focused visualization of the interactions between workloads related to the selected issue.


# Monitor List page

View, filter, and manage all monitors in one place, and quickly identify issues or create new monitors.

The Monitor List is the central hub for managing and monitoring all active and configured Monitors. It provides a clear, filterable table view of your Monitors, with their current status and key details, such as creation date, severity, and live issues. Use this page to review your Monitors performance, identify issues, and take appropriate action.

## **Key Features**

### Monitors Table

* Displays the following columns:
  * **Name:** Title of the monitor.
  * **Creation Date:** When the monitor was created.
  * **Live Issues:** Number of live issues currently firing.
  * **Status:** The Monitor's current status — **Firing** (alerts active), **Pending** (within its pending period), **Normal** (no alerts), or **No Data** (the query returned no data on its latest evaluation). No Data is a status only and does not create an Issue; see [Treat No Data As](/use-groundcover/monitors/create-a-new-monitor).

<figure><img src="/files/nAR8THZr8X8ICk7OX1Gs" alt=""><figcaption><p>Monitors Table</p></figcaption></figure>

### **Create Monitor**

You can create a new Monitor by clicking on Create Monitor, then choosing between the different options: Monitor Wizard, Monitor Catalog, or Import. For further guidance, [check out our guide](/use-groundcover/monitors/create-a-new-monitor).

### **Filters Panel**

Use filters to narrow down monitors by:

* **Severity:** S1, S2, S3, or custom severity levels.
* **Status:** Firing, Pending, Normal, or No Data.
* **Silenced:** Exclude silenced monitors.

**Tip:** Toggle multiple filters to refine your view.

### Search Bar

Quickly locate monitors by typing a name, status, category, or other keywords.

### Cluster and Environment Filters

Located at the top-right corner, use these to focus on monitors for specific clusters or environments.


# Notification Routes

Notification Routes let you automatically send notifications to Destinations when monitor issues change state.

{% hint style="info" %}
This capability is only available to BYOC deployments. Check out our [pricing page](https://www.groundcover.com/pricing) for more information about subscription plans and the available deployment modes.
{% endhint %}

{% hint style="info" %}
Editor permissions are required for creating and editing Notification Routes.\
Reader permissions may list the existing Notification Routes.
{% endhint %}

### How It Works

1. **Scope**: A gcQL query that filters which monitors' issues will trigger this route
2. **Rules**: Define what happens when an issue is in a specific state (Firing or Resolved)
3. [**Destinations**](/integrations/connected-apps): Choose where to send the notification (e.g., Slack Webhook, PagerDuty, etc)

### Prerequisites

Before creating notification routes, set up your Destinations in **Integrations → Destinations**.

### Create a Notification Route

1. Go to **Monitors → Notification Routes**
2. Click **Create Notification Route**
3. Complete the wizard:

#### Step 1: Route Name

Give your route a descriptive name (e.g., `prod-critical-alerts`, `infra-team-notifications`).

#### Step 2: Scope Monitors

Define which monitors this route applies to using gcQL on:

1. The grouping labels defined in the query
2. The custom labels
3. The Monitor's metadata such as the name or severity

**Examples:**

* `env:prod` — All monitors with grouping key 'env' and possible value of 'prod'
* `env:prod AND severity:S1` — Only critical production alerts
* `team:platform` — Monitors with a custom label for the platform team
* `*:*` — To match all monitors

#### Step 3: Rules

Rules define what happens when a scoped monitor's issue changes state.

Each rule has:

* **Status**: When to trigger — `Firing` (issue is active) or `Resolved` (issue cleared)
* **Destinations**: Where to send the notification

**Example setup:**

* When **Firing or Resolved** → Send to `#prod-alerts` Slack channel
* When **Firing or Resolved** → Create and synchronize a [Linear issue](/use-groundcover/connectors/linear#automate-linear-issues-with-a-notification-route)
* When **Firing** (only) → Send to `Pagerduty` service directory

Click **Add Rule** to create multiple rules with different status/destination combinations.

**Selecting Slack App channels**

When you add a [Slack App](/integrations/connected-apps) Destination to a rule, the destination dropdown opens a drilldown so you can pick the exact channel to deliver to:

1. Click the destination dropdown on the rule.
2. Select your Slack App from the list — it is marked with an **App** tag.
3. A second panel opens listing the channels in the workspace. Use the search box to filter by name.
4. Click the target channel. The bot can post to any public channel in the workspace without being added; private channels show a lock icon and only appear in the picker if the bot has been invited to them.
5. To deliver to additional channels (on the same Slack App or a different Destination), click **+** on the rule and repeat.

{% hint style="info" %}
You can route the same monitor to multiple Slack channels by adding the Slack App once per channel in the same rule, or by creating multiple rules.
{% endhint %}

To set up the Slack App connector itself, see [Slack](/use-groundcover/connectors/slack).

#### Re-notification Interval

Configure how long to wait before sending another notification while an alert is still firing.

Options: 1m, 5m, 10m, 30m, 1h, 2h, 4h, 8h, 12h, 1d, 2d

This prevents notification fatigue from long-running alerts.

### Managing Notification Routes

The Notification Routes page shows all your routes with:

* **Name**: Route identifier
* **Scope**: The gcQL query defining which monitors are affected
* **Destinations**: Summary of destinations by type
* **Creator**: Who created the route

#### Edit, Duplicate, or Delete

Hover over any row to access the action menu:

* **Edit**: Modify the route configuration
* **Duplicate**: Create a copy as a starting point
* **Delete**: Remove the route

### Example Use Cases

#### Route Critical Production Alerts to PagerDuty and Slack

```
Name: prod-critical-pagerduty
Scope: env:prod AND severity:S1

Rules:
- When Firing → PagerDuty On-Call
- When Firing or Resolved → #critical-alerts (Slack)
```

#### Separate Routes by Team

```
Name: platform-team-alerts
Scope: team:platform

Rules:
- When Firing → #platform-alerts (Slack)
- When Resolved → #platform-alerts (Slack)
```

```
Name: backend-team-alerts
Scope: team:backend

Rules:
- When Firing → #backend-alerts (Slack)
- When Resolved → #backend-alerts (Slack)
```

#### Development Alerts — Firing Only

```
Name: dev-notifications
Scope: env:dev

Rules:
- When Firing → #dev-alerts (Slack)
```

### Alerting on No Data

Notification routes match **Issues** — so to alert on no data, the monitor first has to **create an Issue** when it returns no data. The **No Data** status on its own is not an Issue and produces no notification, so there is nothing for a route to match. To be notified, set the monitor to **Treat No Data As → Firing** (which turns a no-data evaluation into an Issue), then scope a route to those Issues:

#### Step 1: Set the monitor to treat No Data as Firing

In the monitor wizard, set **Treat No Data As → Firing** (or `noDataState: Alerting` in YAML). An evaluation that returns no data then creates an Issue, just like a threshold breach.

{% hint style="info" %}
A query returns "no data" only when it produces **no rows**, not when it returns a row whose value is `0`. To turn an empty result into No Data, filter out the zero row — e.g. `... | stats count() as cnt | filter cnt > 0`. See [No data vs. a zero value](/use-groundcover/monitors/create-a-new-monitor).
{% endhint %}

#### Step 2: Scope a notification route to the No Data reason

Create a notification route and, in **Step 2: Scope Monitors**, use the gcQL condition:

```gcql
state_change_reason:NoData
```

This matches only the Issues that were created because the monitor returned no data. Combine it with any other condition as needed, e.g. `state_change_reason:NoData AND severity:S1`. Then add a rule and destination in **Step 3** (e.g. *When Firing → #my-alerts*).

<figure><img src="/files/KpSUez3tTNW8hagbGuTN" alt="A notification route scoped with state_change_reason:NoData"><figcaption><p>A notification route that alerts on No Data, scoped with <code>state_change_reason:NoData</code></p></figcaption></figure>

### Testing Notifications

To test your notification configuration before enabling it on production monitors:

1. **Test Destinations directly**: In **Integrations → Destinations**, click the "Test" button on any configured destination to send a test notification with sample data
2. **Test from a monitor**: Open any monitor, click the "Test" button to send a test notification using that monitor's actual payload

{% hint style="info" %}
The test button in Destinations requires Admin permissions. The test button within monitors works for Editor permissions and respects your RBAC data scope.
{% endhint %}

### Re-notification Behavior

Understanding how re-notification intervals work:

* The re-notification timer starts when the first notification is sent
* If an alert **resolves and fires again** within the re-notification window, and the fingerprint is the same, the timer continues (no new notification)
* The `renotification_count` field in payloads starts at 0 for the first notification and increments with each re-notification
* When an alert resolves, the counter resets for the next firing cycle

### Permissions

| Action                          | Required Permission |
| ------------------------------- | ------------------- |
| Create/Edit Notification Routes | Editor              |
| View Notification Routes        | Reader              |
| Create/Edit Destinations        | Admin               |
| Test Destinations               | Admin               |
| Test from Monitor               | Editor              |

### How Routes Apply to Monitors

Notification Routes apply **automatically** based on label matching. You do not need to edit existing monitors to use routes:

1. Routes match monitors based on the scope query (gcQL)
2. When a monitor fires, all routes whose scope matches the monitor's labels will trigger
3. Multiple routes can match the same monitor

### Troubleshooting

#### Notifications not being sent

1. Verify the scope query matches your monitor's labels
2. Check that the Destination is properly configured (use Test button)
3. Review notification delivery traces in groundcover for delivery errors
4. Ensure the monitor has transitioned state (not just hovering at threshold)

#### Wrong time range in issue links

If clicking notification links shows "Last 5 minutes" instead of the actual alert time, this typically affects legacy Workflow-based notifications. Migrate to Notification Routes for proper time range handling.

#### Debugging notification delivery

Notification delivery is performed by the `dispatch-center` service, which emits OpenTelemetry traces for every signal it processes. To inspect delivery attempts and errors, filter traces in groundcover by the route's identifier:

* `route.id` — exact match on the notification route's UUID
* `route.name` — match by the route's display name

Useful span names include `dispatch.signal` (one per signal evaluation) and `dispatch.scan` (per scan iteration). For Slack App deliveries specifically, the `slack.channels` attribute lists the channel IDs the message was sent to.


# Silences page

Manage and create silences to suppress Monitor notifications during maintenance or specific periods. Use one-time silences for ad-hoc needs, or recurring silences for scheduled maintenance windows tha

## Overview

The Silences page lists all Silences you and your team created for your Monitors. In this section, you can create and manage your Silence rules to suppress notifications and Issues noise for a specified period of time. Silences are a great way to reduce alert fatigue, which can lead to missing important issues, and help focus on the most critical issues during specific operational scenarios such as scheduled maintenances.

groundcover supports two types of silences:

* **One-time Silences** — Suppress alerts for a fixed time window. Useful for ad-hoc maintenance or known one-off events.
* **Recurring Silences** — Automatically suppress alerts on a repeating schedule (daily or weekly). Useful for planned maintenance windows, batch jobs, or any predictable operational period.

## Create a one-time Silence

Follow these simple steps to create a new one-time Silence.

<div data-with-frame="true"><figure><img src="/files/GN56G7v8RKykai5ynTSd" alt=""><figcaption></figcaption></figure></div>

### Section 1: Schedule

Specify the timeframe for the silence rule. Note that the starting point doesn't have to be now, and can also be any time in the future.

Below the From / Until boxes, you'll see a Silence summary, showing its approximate length (rounded down to full days) and starting date.

### Section 2: Matchers

Define the criteria for Monitors or Issues to be silenced.

1. Click **Add Matcher** to specify match conditions (e.g., cluster, namespace, span\_name).
2. Each matcher has a label **name**, a **value**, and a **match type**:

| Match type      | Description                                                                                                                      |
| --------------- | -------------------------------------------------------------------------------------------------------------------------------- |
| **Equals**      | The label value exactly matches the value you enter.                                                                             |
| **Not Equals**  | The label value is anything other than the value you enter.                                                                      |
| **Contains**    | Matches any label value that contains the value you enter (`*value*`). You can also add your own `*` wildcards for more control. |
| **Not Contain** | Matches any label value that does *not* contain the value you enter.                                                             |

3. Combine multiple matchers for more granular control.

**Example:** Silence all Monitors in the "demo" namespace.

### Section 3: Comment

Add notes or context for the Silence rule. These comments help you and other users understand the purpose of the rule.

### Section 4: Affected Active Issues

Preview the issues currently affected by the Silence rule, based on any defined Matchers. This list contains only actively firing Issues.

**Tip:** Use this preview to see the list of impacted issues and adjust your Matchers before finishing to create the Silence.

## Recurring Silences

Recurring Silences let you define alert suppression rules that repeat on a schedule. Instead of manually creating a new silence before every maintenance window, you configure the recurrence pattern once and groundcover automatically creates silence instances at the right times.

### How it works

When you create a recurring silence, you define:

1. **Recurrence type** — How often the silence repeats (daily or weekly)
2. **Timeframes** — The specific time windows during which alerts should be silenced
3. **Timezone** — The timezone in which the schedule is evaluated
4. **Matchers** — Which alerts to silence (same matcher system as one-time silences)

groundcover evaluates your recurring silence rules in the background and automatically creates one-time silence instances when a scheduled window is approaching. These instances appear on the Silences page alongside manually created silences, with a **\[recurring]** prefix in their comment to indicate they were generated automatically.

### Create a Recurring Silence

To create a new Recurring Silence, open the Silences page and select the **Recurring** silence type.

#### Recurrence type

Choose how often the silence should repeat:

| Recurrence type | Description                             | Timeframe keys                                                               |
| --------------- | --------------------------------------- | ---------------------------------------------------------------------------- |
| **Daily**       | Repeats every day at the specified time | `every_day`                                                                  |
| **Weekly**      | Repeats on selected days of the week    | `monday`, `tuesday`, `wednesday`, `thursday`, `friday`, `saturday`, `sunday` |

#### Timeframes

For each recurrence type, define one or more time windows using **start time** and **end time** in `HH:MM` format (24-hour clock).

**Key behaviors:**

* **Start time must be before end time** — Each time range must have `startTime < endTime`. The one exception is `00:00` to `00:00`, which represents a full day.
* **Overnight windows** — Overnight windows are not supported as a single time range. To cover an overnight period (e.g., 10 PM to 6 AM), split it into two ranges: `22:00–00:00` and `00:00–06:00`.
* **Multiple windows per day** — You can define more than one time range for the same day. This is especially useful for overnight windows that need to be split.
* **All-day window** — Use `00:00` to `00:00` to silence for the entire day.
* **Weekly selections** — For weekly recurrence, select one or more days and define time windows for each. Different days can have different time windows.

#### Timezone

All time windows are interpreted in a specific timezone, including handling of daylight saving time (DST) transitions.

* **Via the UI** — The timezone is automatically set to the local timezone of the user creating the silence.
* **Via the API** — The timezone is provided as an explicit input using an IANA timezone name (e.g. `America/New_York`, `UTC`, `Asia/Jerusalem`). See the [Recurring Silences API](/use-groundcover/remote-access-and-apis/api-examples/recurring-silences-api) for details.

#### Matchers

Matchers work the same way as one-time silences. Define which alerts should be suppressed when the recurring silence is active.

1. Click **Add Matcher** to specify match conditions.
2. Each matcher has a label **name**, a **value**, and a **match type** (Equals, Not Equals, Contains, or Not Contain). See [Section 2: Matchers](#section-2-matchers) for how match types and wildcards work.
3. Combine multiple matchers for precise targeting.

#### Comment

Add a description to help your team understand the purpose of this recurring silence (e.g., "Weekly database maintenance window"). If left empty, a default comment is added with the created timestamp.

### Managing Recurring Silences

Recurring silences appear in the Silences page alongside one-time silences. From there you can:

* **Edit** — Modify the schedule, matchers, timezone, or comment of an existing recurring silence.
* **Delete** — Remove a recurring silence. Any future silence instances that haven't started yet are also removed.

{% hint style="info" %}
Editing a recurring silence takes effect on the next scheduled window. Silence instances that are already active are not affected by the change.
{% endhint %}

### Example use cases

#### Weekly maintenance window

Suppress alerts every Sunday night from 10 PM to Monday 6 AM for your production cluster. Since overnight windows must be split, this uses two ranges across Sunday and Monday:

```
Recurrence type: Weekly
Timeframes:
  monday: 00:00 – 06:00
  sunday: 22:00 – 00:00
Timezone: America/New_York
Matchers:
  cluster = "production"
```

#### Daily nightly silence

Suppress alerts every night from 8 PM to 8 AM using two time windows:

```
Recurrence type: Daily
Timeframes:
  every_day: 20:00 – 00:00, 00:00 – 08:00
Timezone: UTC
Matchers:
  namespace = "etl"
```

#### Weekend silence for non-critical alerts

Reduce noise from low-severity alerts during weekends:

```
Recurrence type: Weekly
Timeframes:
  saturday: 00:00 – 00:00
  sunday:   00:00 – 00:00
Timezone: Asia/Jerusalem
Matchers:
  severity = "S4"
```

{% hint style="info" %}
`00:00 – 00:00` represents a full-day window.
{% endhint %}


# Monitor Catalog page

Explore and select pre-built Monitors from the catalog to quickly set up observability for your environment. Customize and deploy Monitors in just a few clicks.

## Overview

The Monitor Catalog is a library of pre-built templates to efficiently create new Monitors. Browse and select one or more Monitors to quickly configure their environment with a single click. The Catalog groups monitors into "Packs", based on different use cases.

<figure><img src="/files/eLfCC4fwE4yXiJLA1k2Y" alt=""><figcaption><p>Monitors Catalog</p></figcaption></figure>

## Key Features

### Batch Monitor Creation

You can select as many monitors as you wish, and add them all in one click. Select a complete pack or multiple Monitors from different packs, then click "Create Monitor". All Monitors will be automatically created. You can always edit them later.

### Single Monitor Creation

You can also create a single Monitor from the Catalog. When hovering over a Monitor, a "Wizard" button will appear. Clicking on it will direct you to the [Monitor Creation Wizard](/use-groundcover/monitors/create-a-new-monitor#monitor-creation-wizard) where you can review and edit before creation.


# Monitor YAML structure

While we strongly suggest building monitors using our [Wizard](/use-groundcover/monitors/create-a-new-monitor#using-the-monitor-wizard) or [Catalog](/use-groundcover/monitors/monitor-catalog-page), groundcover also supports building and editing monitors directly in YAML. This page documents the current schema.

For the query language used inside monitor queries, see the [gcQL Reference](/use-groundcover/querying-your-groundcover-data/groundcover-query-language/groundcover-query-language-gcql-reference). For ClickHouse SQL escape hatch monitors, see [SQL Based Monitors](/use-groundcover/monitors/sql-based-monitors).

## Top-level fields

<table data-full-width="true"><thead><tr><th width="220">Field</th><th>Description</th><th width="180">Allowed values</th></tr></thead><tbody><tr><td><strong>title</strong> <em>(required)</em></td><td>Human-readable name of the monitor. Shown in the Monitor List.</td><td>string</td></tr><tr><td><strong>display</strong></td><td>Display settings controlling how issues from this monitor are rendered. See <a href="#display">Display</a>.</td><td>object</td></tr><tr><td><strong>severity</strong></td><td>Severity reported on firing issues.</td><td><code>S1</code>, <code>S2</code>, <code>S3</code>, <code>S4</code></td></tr><tr><td><strong>measurementType</strong></td><td>Type of measurement the monitor represents.</td><td><code>state</code>, <code>event</code></td></tr><tr><td><strong>model</strong> <em>(required)</em></td><td>Queries, reducers, and thresholds that define what the monitor evaluates. See <a href="#model">Model</a>.</td><td>object</td></tr><tr><td><strong>labels</strong></td><td>Static or templated labels attached to the issue. Values can reference query results via <code>{{ $values.&#x3C;threshold_name>.Labels.&#x3C;key> }}</code>.</td><td><code>map&#x3C;string,string></code></td></tr><tr><td><strong>annotations</strong></td><td>Annotations attached to the alert, often used to wire monitors into workflows.</td><td><code>map&#x3C;string,string></code></td></tr><tr><td><strong>category</strong></td><td>Free-form category used for grouping in the Monitor List.</td><td>string</td></tr><tr><td><strong>executionErrorState</strong></td><td>State the monitor enters when query execution fails. Set in the wizard as <strong>Treat Evaluation Errors As</strong>: <code>OK</code> (<strong>Normal</strong>) treats the failure as healthy; <code>Error</code> (<strong>Error</strong>) sets a distinct Error status; <code>Alerting</code> (<strong>Firing</strong>) creates an Issue. Defaults to <code>OK</code>.</td><td><code>OK</code>, <code>Error</code>, <code>Alerting</code></td></tr><tr><td><strong>noDataState</strong></td><td>State the monitor enters when the query returns no rows. Set in the wizard as <strong>Treat No Data As</strong>: <code>OK</code> (<strong>Normal</strong>) treats it as healthy; <code>NoData</code> (<strong>No Data</strong>) sets the No Data status — visible on the monitor but not an Issue; <code>Alerting</code> (<strong>Firing</strong>) creates an Issue. Defaults to <code>NoData</code>. To notify on no data, see <a href="/pages/iJQsKADR1EaNDH0bsmEp#alerting-on-no-data">Alerting on No Data</a>.</td><td><code>OK</code>, <code>NoData</code>, <code>Alerting</code></td></tr><tr><td><strong>evaluationInterval</strong></td><td>Evaluation cadence and pending window. See <a href="#evaluationinterval">EvaluationInterval</a>.</td><td>object</td></tr><tr><td><strong>notificationSettings</strong></td><td>How alerts are delivered. See <a href="#notificationsettings">NotificationSettings</a>.</td><td>object</td></tr><tr><td><strong>autoResolve</strong></td><td>When true, issues from this monitor automatically resolve once the condition no longer holds. Optional.</td><td>boolean</td></tr><tr><td><strong>isPaused</strong></td><td>When true, the monitor is defined but not evaluated.</td><td>boolean</td></tr></tbody></table>

## display

<table data-full-width="true"><thead><tr><th width="260">Field</th><th>Description</th></tr></thead><tbody><tr><td><strong>header</strong></td><td>Template for the issue header. Supports alert label substitution, e.g. <code>"gRPC API Error {{ labels.status_code }}"</code>.</td></tr><tr><td><strong>description</strong></td><td>Template for the issue description. Supports the same substitutions as <code>header</code>.</td></tr><tr><td><strong>resourceHeaderLabels</strong></td><td>List of labels identifying the <em>resource</em> the issue relates to. Rendered as the secondary header across Issues tables. Example: <code>["span_name", "role"]</code>.</td></tr><tr><td><strong>contextHeaderLabels</strong></td><td>List of labels identifying the <em>location</em> of the issue. Rendered as a subset of the issue's labels. Example: <code>["cluster", "namespace", "workload"]</code>.</td></tr><tr><td><strong>templateLanguage</strong></td><td>Template engine for <code>header</code> and <code>description</code>. Set to <code>jinja2</code> to opt into Jinja2 syntax (enablesblocks and filters). Omit for the default Go-template syntax.</td></tr></tbody></table>

## model

<table data-full-width="true"><thead><tr><th width="200">Field</th><th>Description</th></tr></thead><tbody><tr><td><strong>queries</strong> <em>(required)</em></td><td>One or more queries that produce the data the monitor evaluates. See <a href="#modelqueries">model.queries</a>.</td></tr><tr><td><strong>reducers</strong></td><td>Aggregations applied on top of queries before thresholds run. See <a href="#modelreducers">model.reducers</a>.</td></tr><tr><td><strong>thresholds</strong></td><td>Conditions evaluated against a query or reducer output. See <a href="#modelthresholds">model.thresholds</a>.</td></tr></tbody></table>

### model.queries

Each query describes one data source and one expression. The combination of `dataType` and the query body determines which engine runs the query.

<table data-full-width="true"><thead><tr><th width="220">Field</th><th>Description</th></tr></thead><tbody><tr><td><strong>name</strong> <em>(required)</em></td><td>Identifier used by reducers and thresholds to reference this query's output.</td></tr><tr><td><strong>dataType</strong></td><td>Data source for a gcQL query. One of <code>logs</code>, <code>traces</code>, <code>events</code>. <strong>Omit <code>dataType</code> for MetricsQL queries</strong> — see the <a href="#query-engine-field-compatibility">field compatibility</a> note.</td></tr><tr><td><strong>expression</strong></td><td><p>The query itself. The language depends on <code>dataType</code>:</p><ul><li><strong>gcQL</strong> for <code>logs</code>, <code>traces</code>, <code>events</code>. See the <a href="/pages/0rYEo6UxOzVzcXriFK6Y">gcQL Reference</a>.</li><li><a href="https://docs.victoriametrics.com/metricsql/"><strong>MetricsQL</strong></a> when <code>dataType</code> is omitted. MetricsQL is VictoriaMetrics' query language and is backwards-compatible with PromQL, with extra functions (<code>topk_last</code>, <code>rollup_rate</code>, etc.). It does <em>not</em> use the pipe (<code>|</code>) operator; combine operations with nested functions or arithmetic.</li></ul></td></tr><tr><td><strong>datasourceType</strong></td><td>Required for MetricsQL queries. Set to <code>prometheus</code> (the name refers to the Prometheus-compatible API served by the metrics backend).</td></tr><tr><td><strong>queryType</strong></td><td>Required for MetricsQL queries. Use <code>instant</code>.</td></tr><tr><td><strong>filters</strong></td><td>Optional standalone gcQL filter expression applied to the query. For most monitors, put filters directly inside <code>expression</code> instead.</td></tr><tr><td><strong>relativeTimerange</strong></td><td>Time window relative to evaluation time. Object with <code>from</code> and optional <code>to</code> durations (e.g. <code>from: 5m</code>).</td></tr><tr><td><strong>instantRollup</strong></td><td>Bucket size for <strong>gcQL</strong> queries (logs/traces/events), e.g. <code>1 minutes</code>, <code>5 minutes</code>. Controls the time granularity the monitor evaluates over.</td></tr><tr><td><strong>rollup</strong></td><td><strong>Required</strong> for <strong>MetricsQL</strong> queries (whenever <code>datasourceType: prometheus</code> is set). Server-side rollup. Object with <code>function</code> (<code>avg</code>, <code>max</code>, <code>min</code>, <code>sum</code>, <code>count</code>, <code>stddev</code>, <code>stdvar</code>, <code>last</code>) and <code>time</code> (duration). Not interchangeable with <code>instantRollup</code>.</td></tr><tr><td><strong>evaluationDelay</strong></td><td>Optional. Number of <strong>seconds</strong> (0–3600) to shift the evaluation window into the past, so the monitor evaluates a window ending <code>evaluationDelay</code> seconds before "now" (window length unchanged). Useful for sources that backfill recent data (AWS CloudWatch, GCP). Omit or set to <code>0</code> for no delay. See <a href="/pages/jhQvSfcnZfvR8NzVhJUL#section-1-query">Evaluation delay</a>.</td></tr></tbody></table>

#### Query-engine field compatibility

The query fields are tied to the query engine, and several of them must appear **together**. Mixing them across engines is the most common cause of a rejected monitor.

<table data-header-hidden data-full-width="true"><thead><tr><th></th><th></th><th></th></tr></thead><tbody><tr><td>Engine</td><td>Required fields</td><td>Must <em>not</em> be set</td></tr><tr><td><strong>gcQL</strong><br>(<code>dataType</code> = <code>logs</code>, <code>traces</code>, <code>events</code>)</td><td><code>dataType</code>, <code>expression</code>, <code>instantRollup</code></td><td><code>datasourceType</code>, <code>queryType</code>, <code>rollup</code></td></tr><tr><td><strong>MetricsQL</strong><br>(no <code>dataType</code>)</td><td><code>expression</code>, <code>datasourceType: prometheus</code>, <code>queryType: instant</code>, <code>rollup</code></td><td><code>dataType</code>, <code>instantRollup</code></td></tr></tbody></table>

Key rules the backend enforces:

* `datasourceType: prometheus` **requires** `rollup` (object with `function` + `time`). Omitting it is rejected with *"rollup is required for prometheus datasource type"*.
* For MetricsQL, **omit `dataType`** — `datasourceType: prometheus` together with the MetricsQL `expression` and `rollup` is all you need.
* `rollup` (MetricsQL) and `instantRollup` (gcQL) are **not** interchangeable — pick the one for your engine.

### model.reducers

Reducers aggregate a query's output into a single value (or per-group value) before thresholds run. This is how you turn a timeseries into a single number to compare against a threshold.

<table data-full-width="true"><thead><tr><th width="200">Field</th><th>Description</th></tr></thead><tbody><tr><td><strong>name</strong> <em>(required)</em></td><td>Identifier used by thresholds.</td></tr><tr><td><strong>inputName</strong></td><td>Name of the query (or another reducer) to read from. Required unless <code>type: math</code>.</td></tr><tr><td><strong>type</strong> <em>(required)</em></td><td>One of <code>last</code>, <code>min</code>, <code>max</code>, <code>mean</code>, <code>sum</code>, <code>count</code>, <code>math</code>.</td></tr><tr><td><strong>expression</strong></td><td>Required when <code>type: math</code>. An arithmetic expression over reducer outputs, e.g. <code>$errors / $total * 100</code>.</td></tr><tr><td><strong>relativeTimerange</strong></td><td>Optional time window specific to this reducer.</td></tr></tbody></table>

### model.thresholds

Thresholds are the final condition that determines whether the monitor fires.

<table data-full-width="true"><thead><tr><th width="200">Field</th><th>Description</th></tr></thead><tbody><tr><td><strong>name</strong> <em>(required)</em></td><td>Identifier for this threshold.</td></tr><tr><td><strong>inputName</strong> <em>(required)</em></td><td>Name of the query or reducer this threshold evaluates.</td></tr><tr><td><strong>operator</strong> <em>(required)</em></td><td>One of <code>gt</code>, <code>lt</code>, <code>gte</code>, <code>lte</code>, <code>eq</code>, <code>neq</code>, <code>within_range</code>, <code>outside_range</code>, <code>within_range_included</code>, <code>outside_range_included</code>.</td></tr><tr><td><strong>values</strong> <em>(required)</em></td><td>Array of numbers. One value for comparison operators; two values for <code>within_range</code> / <code>outside_range</code> and their inclusive variants.</td></tr><tr><td><strong>relativeTimerange</strong></td><td>Optional time window specific to this threshold.</td></tr><tr><td><strong>customResolveThreshold</strong></td><td>Optional hysteresis recovery condition. When set, a firing issue resolves only once this separate predicate is met, which reduces flapping for values that hover around the firing threshold. Object with <code>operator</code> and <code>values</code>. The <code>operator</code> must be the directional opposite of the threshold's <code>operator</code> (<code>gt</code>↔<code>lt</code>, <code>within_range</code>↔<code>outside_range</code>) and is only supported for <code>gt</code>, <code>lt</code>, <code>within_range</code>, and <code>outside_range</code>. The resolve <code>values</code> must not overlap the firing range, e.g. fire when <code>gt 100</code> and resolve when <code>lt 80</code>.</td></tr></tbody></table>

## evaluationInterval

<table data-full-width="true"><thead><tr><th width="200">Field</th><th>Description</th></tr></thead><tbody><tr><td><strong>interval</strong></td><td>How often the monitor is evaluated, e.g. <code>1m</code>, <code>5m</code>.</td></tr><tr><td><strong>pendingFor</strong></td><td>Duration the threshold must hold before the issue transitions from Pending to Alerting. Use <code>0s</code> to fire immediately.</td></tr></tbody></table>

## notificationSettings

<table data-full-width="true"><thead><tr><th width="240">Field</th><th>Description</th></tr></thead><tbody><tr><td><strong>method</strong></td><td>How notifications are delivered. <code>notificationRoutes</code> uses the matching routes defined in <a href="/pages/iJQsKADR1EaNDH0bsmEp">Notification Routes</a>; <code>connectedApps</code> sends directly to the apps listed in <code>connectedApps</code>, bypassing routes; <code>noNotifications</code> suppresses all notifications for this monitor.</td></tr><tr><td><strong>connectedApps</strong></td><td>List of Destination IDs. Used with <code>method: connectedApps</code>. (The YAML field name <code>connectedApps</code> is retained for backward compatibility; in the UI these are called <a href="/pages/fo2WdWlqpxCCjttUSWFa">Destinations</a>.)</td></tr><tr><td><strong>connectedAppParams</strong></td><td>Per-destination delivery options keyed by Destination ID. For Slack App connectors, use <code>channels</code> to set the target channels — see <a href="/pages/Y2tO4Ei086NA8DhVmBzS">Slack</a>. Example: <code>{ "&#x3C;slack-app-id>": { "channels": [{ "id": "C123456", "name": "#alerts" }] } }</code>. Used with <code>method: connectedApps</code>.</td></tr><tr><td><strong>renotificationInterval</strong></td><td>Duration between repeat notifications while the issue remains firing, e.g. <code>4h</code>.</td></tr><tr><td><strong>disableRenotification</strong></td><td>When true, suppresses repeat notifications.</td></tr><tr><td><strong>statusFilters</strong></td><td>List of issue statuses that trigger notifications. Allowed values: <code>Alerting</code>, <code>Resolved</code>. Used with <code>method: connectedApps</code>. Note: the monitor wizard labels these as <strong>Firing</strong> and <strong>Resolved</strong> — <code>Firing</code> in the UI corresponds to <code>Alerting</code> in YAML.</td></tr></tbody></table>

## Examples

### Traces monitor (gcQL)

Fires when gRPC traces return a non-zero status code.

```yaml
title: gRPC API Errors Monitor
display:
  header: gRPC API Error {{ labels.status_code }}
  description: |-
    This monitor detects gRPC API errors by identifying responses with a status code indicating failure.
    Cluster: {{ labels.cluster }}
    Namespace: {{ labels.namespace }}
    Workload: {{ labels.workload }}
    Span: {{ labels.span_name }}
  resourceHeaderLabels:
    - span_name
    - role
  contextHeaderLabels:
    - env
    - cluster
    - namespace
    - workload
  templateLanguage: jinja2
severity: S3
measurementType: event
model:
  queries:
    - name: threshold_input_query
      dataType: traces
      expression: >
        span_type:grpc status_code:!=0 status:error source:eBPF
        | stats by (env, cluster, namespace, workload, status_code, span_name, role) count() errors_total
      instantRollup: 1 minutes
  thresholds:
    - name: threshold_1
      inputName: threshold_input_query
      operator: gt
      values:
        - 0
executionErrorState: OK
noDataState: OK
evaluationInterval:
  interval: 1m
  pendingFor: 0s
```

### Logs monitor (gcQL)

Fires when sensor logs contain panic or fatal errors.

```yaml
title: Sensor Panic / Fatal Errors
display:
  header: Sensor Panic or Fatal Errors
  description: |-
    Detects panic or fatal errors in sensor logs.
    Cluster: {{ labels.cluster }}
    Namespace: {{ labels.namespace }}
    Pod: {{ labels.pod }}
  contextHeaderLabels:
    - pod
    - workload
    - cluster
    - env
    - namespace
severity: S2
measurementType: event
model:
  queries:
    - name: threshold_input_query
      dataType: logs
      expression: >
        container:sensor level:in(panic, fatal)
        | stats by (env, cluster, namespace, workload, pod) count() count_all_result
      instantRollup: 5 minutes
  thresholds:
    - name: threshold_1
      inputName: threshold_input_query
      operator: gt
      values:
        - 0
executionErrorState: OK
noDataState: OK
evaluationInterval:
  interval: 5m
  pendingFor: 0s
notificationSettings:
  renotificationInterval: 4h
```

### Monitor that delivers directly to Slack channels

Bypasses notification routes and sends issues from this monitor straight to two Slack channels via the [Slack](/use-groundcover/connectors/slack) Destination. Replace the `slack-app-id` placeholder with the ID of your Slack App Destination, and the channel `id` values with the Slack channel IDs (e.g., `C0123456789`) you want to deliver to.

```yaml
title: Checkout Service 5xx Spike
display:
  header: Checkout 5xx Spike in {{ labels.cluster }}
  description: |-
    HTTP 5xx responses from the checkout service are above threshold.
    Cluster: {{ labels.cluster }}
    Namespace: {{ labels.namespace }}
    Workload: {{ labels.workload }}
severity: S2
measurementType: event
model:
  queries:
    - name: threshold_input_query
      dataType: traces
      expression: >
        workload:checkout status_code:>=500
        | stats by (env, cluster, namespace, workload) count() errors_total
      instantRollup: 1 minutes
  thresholds:
    - name: threshold_1
      inputName: threshold_input_query
      operator: gt
      values:
        - 10
executionErrorState: OK
noDataState: OK
evaluationInterval:
  interval: 1m
  pendingFor: 0s
notificationSettings:
  method: connectedApps
  connectedApps:
    - <slack-app-id>
  connectedAppParams:
    <slack-app-id>:
      channels:
        - id: C0123456789
          name: "#checkout-alerts"
        - id: C0987654321
          name: "#oncall-prod"
  statusFilters:
    - Alerting
    - Resolved
  renotificationInterval: 1h
```

{% hint style="info" %}
Channel `id` is the canonical Slack channel ID (it stays stable if the channel is renamed); `name` is optional and only used for display. You can find the channel ID in Slack via **Channel name → View channel details → About** (the ID is shown at the bottom).
{% endhint %}

### Metrics monitor (MetricsQL)

Fires when a Kubernetes pod is in `CrashLoopBackOff` for more than 5 minutes. Uses a reducer to collapse the timeseries before the threshold runs.

{% hint style="info" %}
Metrics queries use [MetricsQL](https://docs.victoriametrics.com/metricsql/) (PromQL-compatible). kube-state-metrics names are prefixed with `groundcover_`, and the node identity label is `node_name` (not `node`).
{% endhint %}

```yaml
title: K8s Pod Crash Looping Monitor
display:
  header: K8s Pod Crash Looping
  description: Kubernetes pod has been in CrashLoopBackOff for more than 5 minutes.
  resourceHeaderLabels:
    - workload
  contextHeaderLabels:
    - env
    - cluster
    - namespace
severity: S2
measurementType: state
model:
  queries:
    - name: crash_looping_query
      expression: >
        avg_over_time(
          avg by (env, cluster, namespace, workload) (
            groundcover_kube_pod_container_status_waiting_reason{reason="CrashLoopBackOff"}
          )[5m]
        )
      datasourceType: prometheus
      queryType: instant
      rollup:
        function: avg
        time: 5m
  reducers:
    - name: crash_looping_mean
      inputName: crash_looping_query
      type: mean
  thresholds:
    - name: threshold_1
      inputName: crash_looping_mean
      operator: gt
      values:
        - 0
executionErrorState: OK
noDataState: OK
evaluationInterval:
  interval: 1m
  pendingFor: 5m
```

### ClickHouse SQL monitor

For advanced cases that need joins, CTEs, or comparisons across time windows, you can drop to ClickHouse SQL. See [SQL Based Monitors](/use-groundcover/monitors/sql-based-monitors) for details and examples.


# SQL Based Monitors

Sometimes there are use cases that involve complex queries and conditions for triggering a monitor. This might go beyond the built-in query logic that is provided within the groundcover logs page query language.

An example for such a use case could be the need to compare some logs to the same ones in a past period. This is not something that is regularly available for log search but can definitely be something to alert on. If the number of errors for a group of logs dramatically changes from a previous week, this could be an event to alert and investigate.

For such use cases you can harness the powerful ClickHouse SQL language to create an SQL based monitor within groundcover.

### ClickHouse within groundcover

Log and Trace telemetry data is stored within a ClickHouse database.

You can directly query this data using SQL statements and create powerful monitors.

To create and test your SQL queries use the [Grafana Explore](https://app.groundcover.com/grafana/explore) page within the groundcover app.

Select the ClickHouse\@groundcover datasource with the SQL Editor option to start crafting your SQL queries

<figure><img src="/files/7XmUtW9NpImYsZ39YAOs" alt=""><figcaption></figcaption></figure>

Start with `show tables;` to see of all the available tables to use for your queries: `logs` and `traces` would be popular choices (table names are case sensitive).

### Query best practices

While testing your queries always use LIMIT to limit your results to a small set of data.

```sql
SELECT * FROM logs LIMIT 10;
```

To apply the Grafana timeframe on your queries make sure to add the following conditions:

Logs: `WHERE $__timeFilter(timestamp)`

Traces: `WHERE $__timeFilter(start_timestamp)`

**Note:** When querying logs with SQL, it's crucial to use efficient filters to prevent timeouts and enhance performance. Implementing primary filters like `cluster`, `workload`, `namespace`, and `env` will significantly speed up queries. Always integrate these filters when writing your queries to avoid inefficient queries.

### Filtering on attributes and tags

Traces and Logs have rich context that is normally stored in dedicated columns in json format. Accessing the context for filtering and retrieving values is a popular need when querying the data.

To get to the relevant context item, either in the attributes or tags you can use the following syntax:

`WHERE string_attributes['host_name'] = 'my.host'`

`WHERE string_tags['cloud.name'] = 'aws'`

`WHERE float_attributes['hradline_count'] = 4`

`WHERE float_tags['container.size'] = 22.4`

To use the `float` context ensure that the relevant attributes or tags are indeed numeric. To do that, check the relevant log in json format to see if the referenced field is not wrapped with quotes (for example, `headline_count` in the screenshot below)

<figure><img src="/files/I4C22WWRWByoVSpEPruX" alt=""><figcaption></figcaption></figure>

### SQL Query structure for a monitor

In order to be able to use an SQL query to create a monitor you must make sure the query returns no more than a single numeric field - this is the monitored field on which the threshold is placed.

The query can also contain any number of "group by" fields that are passed to the monitor as context labels.

Here is an exmaple of an SQL query that can be used for a monitor

```sql
with engineStatusLastWeek as (
  select string_attributes['tenantID'] tenantID, , string_attributes['env'] env, max(float_attributes['engineStatus.numCylinders']) cylinders
  from logs
  where timestamp >= now() - interval 7 days
    and workload = 'engine-processing'
    and string_attributes['tenantID'] != ''
  group by tenantID, env
),
engineStatusNow as (
  select string_attributes['tenantID'] tenantID, string_attributes['env'] env, min(float_attributes['engineStatus.numCylinders']) cylinders
  from logs
  where timestamp >= now() - interval 10 minutes
    and workload = 'engine-processing'
    and string_attributes['tenantID'] != ''
  group by tenantID, env
)
select n.tenantID, n.env, n.cylinders/lw.cylinders AS threshold
from engineStatusNow n
left join engineStatusLastWeek lw using (tenantID)
where n.cylinders/lw.cylinders <= 0.5
```

In this query the threshold field is a ratio between some value measured on the last week and in the last 10 minutes.

`tenantID` and `env` are the group by labels that are passed to the monitor as context labels.

***

Here is another query example (check the percentage of errors in a set of logs):

A single numeric value is calculated and grouped by cluster, namespace and workload

```sql
SELECT cluster, namespace, workload, 
    round( 100.0 * countIf(level = 'error') / 
    nullIf(count(), 0), 2 ) AS error_ratio_pct 
FROM "groundcover"."logs" 
WHERE timestamp >= now() - interval '10 minute' AND 
namespace IN ('refurbished', 'interface') GROUP BY cluster, namespace, workload
```

### Applying the SQL query as a monitor

Applying an SQL query can only happens in YAML mode. You can use the following YAML template to add your query

```yaml
title: "[SQL] Monitor name"
display:
  header: Monitor description
severity: S2
measurementType: event
model:
  queries:
    - name: threshold_input_query
      expression: "[YOUR SQL QUERY GOES HERE]"
      datasourceType: clickhouse
      queryType: instant
  thresholds:
    - name: threshold_1
      inputName: threshold_input_query
      operator: lt
      values:
        - 0.5
annotations:
  [Workflow Name]: enabled
executionErrorState: OK
noDataState: OK
evaluationInterval:
  interval: 3m
  pendingFor: 2m
isPaused: false

```

1. Give your monitor a name and a description
2. Paste your SQL query in the `expression` field
3. Set the threshold value and the relevant operator - in this example this is "lower than" 0.5 (< 0.5)
4. Set your workflow name in the annotations section
5. Set the check interval and the pending time
6. Save the monitor

### Trigger an alert when no logs are coming from a Linux host

Use [YAML mode](/use-groundcover/monitors/monitor-yaml-structure) to add the following template to your monitor.

```yaml
title: Host not sending logs more than 5 minutes
display:
  header: Host "{{host}}" is not sending logs for more than 5 minutes
severity: S2
measurementType: event
model:
  queries:
    - name: threshold_input_query
      expression: "
      WITH
          (
          SELECT groupArray(DISTINCT host)
          FROM logs
          WHERE timestamp >= now() - INTERVAL 24 HOUR
          AND env_type = 'host'
          ) AS all_hosts
      SELECT
          host,
          coalesce(log_count, 0) AS log_count
      FROM
          (
          SELECT arrayJoin(all_hosts) AS host
          ) AS h
          LEFT JOIN
              (
              SELECT host, count(*) AS log_count
              FROM logs
              WHERE timestamp >= now() - INTERVAL 5 MINUTE
              AND env_type = 'host'
              GROUP BY host
              ) AS l
      USING (host)
      ORDER BY host
      "
      datasourceType: clickhouse
      queryType: instant
  thresholds:
    - name: threshold_1
      inputName: threshold_input_query
      operator: lt
      values:
        - 10
annotations:
  {Put Your Workflow Name Here}: enabled
executionErrorState: Error
noDataState: NoData
evaluationInterval:
  interval: 5m
  pendingFor: 0s
isPaused: false
```

In this example we are creating a list of Linux hosts that were sending logs in the last 24 hours and then checking if there were any logs collected from those hosts in the last 5 minutes.

This monitor can be used e.g. to catch when the host is down.

{% hint style="info" %}
It would be helpful if you add an indication on the monitor name that this is SQL based. For example, add an \[SQL] prefix or suffix to the monitor name as shown in the example
{% endhint %}

## SQL Mode vs Builder Mode

### Key Differences

| Feature                 | Builder Mode           | SQL Mode                 |
| ----------------------- | ---------------------- | ------------------------ |
| Query interface         | Visual query builder   | Raw SQL/PromQL           |
| Notification Routes     | UI selection available | Via YAML only            |
| Switching between modes | Can switch freely      | Cannot switch to Builder |
| Data sources            | All supported          | ClickHouse only          |

### Limitations of SQL Mode

1. **Cannot switch back to Builder mode**: Once a monitor is created in SQL mode, you cannot convert it to Builder mode
2. **Notification settings via YAML only**: To configure how a SQL-based monitor sends notifications, use the `notificationSettings` field in YAML. The wizard does not expose this section for SQL monitors. See [Monitor YAML structure](/use-groundcover/monitors/monitor-yaml-structure#notificationsettings) for the full field reference.

   To route through matching notification routes (default):

   ```yaml
   title: "[SQL] My Monitor"
   # ... other monitor config ...
   notificationSettings:
     method: notificationRoutes
   ```

   To deliver directly to specific Destinations (for example, a Slack App with channels), see the [Monitor that delivers directly to Slack channels](/use-groundcover/monitors/monitor-yaml-structure#monitor-that-delivers-directly-to-slack-channels) example.
3. **No visual threshold preview**: The query preview doesn't show threshold lines like Builder mode

### When to Use SQL Mode

Use SQL mode when you need:

* Complex queries with CTEs, JOINs, or subqueries
* Comparison between current and historical data
* Custom aggregations not available in Builder
* Direct access to ClickHouse functions


# Migrating from Issues to Monitors Issues Page

The legacy Issues page is being deprecated in favor of a fully customizable, monitor-based experience that gives you more control over what constitutes an issue in your environment.

While the new page introduces powerful capabilities, no core functionality is being removed, the key change is that the old auto-created issue rules will no longer be automatically generated. Instead, you’ll define your own monitors, or choose from a rich catalog of prebuilt ones.

All the existing issues in the legacy page can be easily added to the monitors via the Monitors Catalog's "Started Pack". See [Getting Started](#getting-started) below for more info.

### Why migrate to the new Issues experience?

The new Issues page is built on top of the Monitors engine, enabling deeper customization and automation:

1. **Define what qualifies as an issue**

   Use filters in monitor definitions to include or exclude workloads, namespaces, HTTP status codes, clusters, and more tailor it to your context.
2. **Silence issues with precision**

   Silence issues based on any label, such as `status_code`, `cluster`, or `workload`, to reduce noise and keep focus.
3. **Clean, scoped issue view**

   Only see issues relevant to your environment, based on your configured monitors and silencing rules, no clutter.
4. **Get alerted on new issues**

   Trigger alerts through your preferred integrations (Slack, PagerDuty, Webhooks, etc.) when a new issue is detected.
5. **Define custom issues using all your data**

   Build monitors using metrics, traces, logs, and events, and correlate them to uncover complex problems.
6. **Manage everything as code**

   Use Terraform to manage monitors and issues at scale, ensuring consistency and auditability.

### What’s Changing?

| Aspect                              | Legacy Issues Page | New Issues Page       |
| ----------------------------------- | ------------------ | --------------------- |
| Issue Source                        | Auto-created rules | User-defined monitors |
| Custom Filtering                    | ❌                  | ✅                     |
| Silencing by Labels                 | ❌                  | ✅                     |
| Alerts via Integrations             | ❌                  | ✅                     |
| Terraform Support                   | ❌                  | ✅                     |
| Issues Based on Traces/Logs/Metrics | Limited            | Full support          |

### Getting Started

All the built-in rules you’re used to are already available in the [Monitors Catalog](/use-groundcover/monitors/monitor-catalog-page), you can add them all with a single click.

Adding the monitors in the "Started Pack" will match all the existing issues in the legacy page.

Head to:

**Monitors → Create Monitor -> Monitor Catalog → Recommended Monitors**

<figure><img src="/files/cEczsGGYRFnXY4oqigg7" alt=""><figcaption></figcaption></figure>

{% hint style="info" %}
Only users with Editor/Admin roles can create monitors
{% endhint %}

### Learn More

* [Create a new Monitor](/use-groundcover/monitors/create-a-new-monitor)
* [Issues page](/use-groundcover/monitors/issues-page)
* [Silences page](/use-groundcover/monitors/silences-page)


# Dashboards

Learn how to build custom dashboards using groundcover

Dashboards let you build **persistent, shareable views** over your observability data. Use them for on-call playbooks, service health boards, incident investigation layouts, and any view you want to revisit without rebuilding queries in Explore.

Dashboards are built from **widgets** (charts, tables, stats, text, and grouped sections), filtered with **variables**, and scoped to a **time range**. Changes are saved per dashboard and can be shared with your team.

Easily create a new Dashboard [using our guide](/use-groundcover/dashboards-and-alerts/create-a-dashboard).

You can also install a ready-made dashboard from the [Dashboard Catalog](/use-groundcover/dashboards-and-alerts/dashboard-catalog), a library of pre-built dashboards that groundcover manages and keeps current automatically.

<figure><img src="/files/oxIfhHndU3n1cgxaHFqM" alt=""><figcaption></figcaption></figure>


# Creating Dashboards

{% hint style="info" %}
**Note**: Only users with Write or Admin permissions can create and edit dashboards.
{% endhint %}

## **Overview**

Dashboards let you build **persistent, shareable views** over your observability data. Use them for on-call runbooks, service health boards, incident investigation layouts, and any view you want to revisit without rebuilding queries in Explore.

## **How to create a new dashboard in groundcover?**

{% hint style="info" %}
Don't want to build from scratch? Install a ready-made dashboard from the [Dashboard Catalog](/use-groundcover/dashboards-and-alerts/dashboard-catalog) instead.
{% endhint %}

1. Navigate to the [**Dashboards**](https://app.groundcover.com/dashboards) pag&#x65;**.** This page shows all dashboards for the selected backend.
2. Click on the **Create Dashboard** button.
3. Provide a meaningful name for your dashboard and, optionally, a description.

<figure><img src="/files/cA6hViE2QSgGqGeuQlMJ" alt=""><figcaption></figcaption></figure>

An empty dashboard layout will appear.

Follow up the steps below to populate your dashboard with widgets.

### Create a new Widget

Widgets can be added by clicking on the **Create New Widget** button in case of a new Dashboard or Create Widget at the top right in case there is at least one widget.

### **Choose a Widget Type**

Widgets are the main building blocks of dashboards. groundcover supports the following widget types:

* **Chart Widget**: Visualize your data through various display types.
* **Textual Widget**: Add context to your dashboard, such as headers or instructions for issue investigations.
* **Section Widget**: Group related widgets together

{% hint style="info" %}
Since selecting a Textual or Section Widget is the last step for this type of widget, the rest of this guide is relevant only to Chart Widgets.
{% endhint %}

### **Select a Data Type**

**One Chart Widget is selected, the widget builder will open.**

**The first step is to select the data type to query out of the following:**

| Data type     | Use for                                    |
| ------------- | ------------------------------------------ |
| **Metrics**   | Metric time series, rates, aggregations    |
| **Logs**      | Log volume, counts, grouped log tables     |
| **Traces**    | Trace latency, errors, grouped trace views |
| **Events**    | Kubernetes and platform events             |
| **Entities**  | Infrastructure entity data                 |
| **Issues**    | Issue-oriented signals                     |
| **RUM**       | Real user monitoring data                  |
| **APM**       | Application performance views              |
| **Ingestion** | Ingestion and pipeline metrics             |

### Build your query

Build one or more queries and formulas to be visualized. Queries can be built using a visual Builder or by providing a MetricsQL/gcQL query.

{% hint style="info" %}
If you're unfamiliar with query building in groundcover, refer to the [Query Builder section](/use-groundcover/querying-your-groundcover-data/explore-and-monitors-query-builder) for full details on the different components.
{% endhint %}

<figure><img src="/files/fByYZLSiIwElx71e3QV5" alt=""><figcaption></figcaption></figure>

### **Choose a Visualization Type**

There are seven ways to visualize your query, each one with different configuration options.

Refer to the [visualization types](/use-groundcover/dashboards-and-alerts/dashboard-visualizations) page to read more about the visualization and configuration options.

## **Working with widgets**

### Sharing a widget

Every widget and section has a **Copy link** action (widget menu) that generates a URL pointing directly to that widget. Opening the link scrolls to and highlights the widget automatically — useful for sharing one specific chart in Slack or a ticket instead of the whole dashboard.

### Keyboard shortcuts

While hovering over a widget, the following shortcuts apply:

| Key | Action           |
| --- | ---------------- |
| `E` | Edit             |
| `V` | View / exit edit |
| `F` | Fullscreen       |
| `D` | Duplicate        |
| `R` | Remove           |

### Synced crosshair

Toggle **crosshair sync** from the dashboard's view menu to move the hover crosshair across all time-series widgets on the dashboard together, making it easier to compare the same point in time across multiple charts.

## **Layout**

Dashboards support two layout modes, available from the layout menu:

* **Ordered** — widgets flow automatically into rows; reordering a widget shifts the others around it.
* **Floating** — widgets can be freely positioned and resized on a grid, independent of one another.

Drag a widget's corner to resize it, or drag its header to reposition it, in either mode.

## **Variables**

Variables dynamically filter your entire dashboard or specific widgets with just one click. They consist of a key-value pair that you define once and reuse across multiple widgets.

### **Adding a Variable**

1. Click on **Add Variable** and configure the variable using the following fields.\
   ![](/files/KlKzlyuDMACVKo5QGN63)
2. The **data source** to be used:
   1. `Suggested Variables` - A predefined list of popular variables which use 'All Datasources' behind the scenes.
   2. `All Data Sources` - Will show values for the chosen key across all data sources.
   3. Specific data types - Only show values for the chosen key for data coming from the selected source.
3. Choose the **label key** to be used to fetch values from the data source selected.
4. Choose the **name** of the variable to be used in the widgets with `$` as explained below.

### **Using a Variable**

Variables can be referenced in the Filter Bar of the Widget Builder Modal using their name.

1. In this following example we selected Clusters from the predefined list, and named it 'clusters'.
2. While creating or editing a Chart Widget, add a reference to the variable using a dollar sign in the filter bar, (for example, `$clusters`).
3. The data will automatically filtered by the variable's key with the selected values. If all values are selected, the filter will be followed by an asterisk (for example, `cluster:.*)`<br>

   <figure><img src="/files/JcGdDLuj7zxBKNnmZOTR" alt=""><figcaption></figcaption></figure>
4. After configuring the Variable in the widget queries, you may select the values to filter and choose the default to be used when the dashboard loads on first time.
5. After selecting values in at least one Variable, all other relevant Variables will render an 'Associated Values' section in the dropdown list. This list renders the values of the selected variable's key which are associated with the values of the currently selected variables' keys.
   1. For example- selecting the value `production` in a variable called `cluster` which uses the key `cluster` from `All Data Sources.` When going to the `workloads` variable, the 'Associated Values' section will list the `workload` values that are in the `production` cluster.
   2. Below in the 'Additional Values' will be shown all other values.
   3. The association is done by relevant data types only, if you are getting unexpected associated results you may be advised to narrow down the data sources that the variable uses from 'All Data Sources' to a specific type or a specific metric.
   4. Limitations and tips-
      1. It's possible that there are associated values which don't appear in the list, this list is not hermetic, but anything associated is necessarily associated.
      2. Start to type the value you are searching to narrow down the list.
      3. It's possible that Additional Values will also relate to the chosen values of other variables.

<figure><img src="/files/sl2DYJ4MQ4cBiHKQymUD" alt=""><figcaption></figcaption></figure>

## **Adding a chart from elsewhere**

You don't always have to build a widget from scratch inside a dashboard. A chart built while querying data in Explore, or a chart generated inline by the AI assistant, can be saved directly into an existing dashboard via its **Add to Dashboard** action — pick the target dashboard and the chart is added as a new widget with its query intact.

## **Finding and organizing dashboards**

The Dashboards list supports:

* **Tags** — apply free-form labels to a dashboard and filter the list by tag.
* **Search** — filter by name, tag, owner, or description.
* **Archive / Restore** — archive a dashboard to hide it from the default list without permanently deleting it, and restore it later. Archived dashboards can still be permanently deleted if no longer needed.

## **Keeping dashboards in sync**

If someone else saves changes to a dashboard while you're viewing it, groundcover shows a notice that a newer version is available. Refresh to pick up their changes before making your own edits, so you don't unknowingly overwrite them.

If you do need to recover from an overwritten or unwanted change, see [Dashboard Version History](/use-groundcover/dashboards-and-alerts/dashboard-version-history).


# Dashboard Catalog

Browse groundcover's library of pre-built dashboards and install the ones you need. Installed dashboards are managed by groundcover and stay current automatically. Customize one to make it your own.

## **Overview**

The Dashboard Catalog is a library of pre-built dashboards for common technologies and use cases, including Kubernetes, hosts, AWS services, databases, messaging, CI/CD, AI & LLM, and platform usage. Browse it, preview a dashboard against your own data, and install what you need in one click.

Use the filters at the top to narrow the catalog by technology or use case, for example Kubernetes, AWS, or Databases.

<figure><img src="/files/tiOlUNraIBSFI4Rn4xSZ" alt=""><figcaption><p>The Dashboard Catalog</p></figcaption></figure>

## **Installed dashboards are managed by groundcover**

There's one thing to know before you install.

An installed dashboard is **managed by groundcover and always current**. groundcover maintains its content, so when we improve a dashboard your installed copy updates automatically. You never re-install to get changes.

A managed dashboard is **read-only** and marked **Managed by groundcover** in its header. To edit it, use **Customize Dashboard**, described [below](#customize-a-dashboard).

## **Browse and preview**

1. Open the catalog from the [**Dashboards**](https://app.groundcover.com/dashboards) page.
2. Filter by category, or use **Search** to find a dashboard by name or tag. The **Recommended** filter shows the dashboards most teams start with.
3. Click a dashboard to preview it, rendered live against your own data.

<figure><img src="/files/OaakGujxnDhrCyuQN3Kf" alt=""><figcaption><p>Previewing a dashboard against live data</p></figcaption></figure>

## **Install a dashboard**

In the preview, click **Install Dashboard**. It's added to your Dashboards list and managed by groundcover. The catalog card then shows **Installed**, and **Open Dashboard** takes you to it.

<figure><img src="/files/DU7o3k8uG9mXFRusZNtt" alt=""><figcaption><p>Install Dashboard from the preview</p></figcaption></figure>

## **Customize a dashboard**

To change a managed dashboard, open it and click **Customize Dashboard**. This creates an editable copy you own, with full control over its widgets, layout, and variables.

The copy is independent and no longer tracks the catalog, so it won't receive future updates. The original managed dashboard stays in place and stays current.

<figure><img src="/files/my2oszE5BgBYcAFuMINe" alt=""><figcaption><p>Customize Dashboard on a managed dashboard</p></figcaption></figure>

## **Dashboards for your data sources**

Dashboards relevant to a connected data source are available from that data source, so you can go from a new integration to a working dashboard without browsing the full catalog.

<figure><img src="/files/pXVWVfipDHBBXcan09CR" alt=""><figcaption><p>Related dashboards from a connected data source</p></figcaption></figure>


# Dashboard Visualizations

This page documents the **visualization types** available in dashboard widgets and the **configuration options** for each one.

For the overall dashboard workflow (creating widgets, variables, saving, and permissions), see [Creating Dashboards](/use-groundcover/dashboards-and-alerts/create-a-dashboard).

### Visualization types at a glance

| Visualization   | Best for                                           |
| --------------- | -------------------------------------------------- |
| **Time series** | Trends and rates over time                         |
| **Table**       | Grouped or row-level tabular results               |
| **Stat**        | One or more KPI numbers                            |
| **Top list**    | Ranked values as horizontal bars                   |
| **Pie**         | Share of total across a small number of categories |
| **Treemap**     | Hierarchical share of total across many categories |
| **Data**        | Raw log, trace, or event rows (list-style)         |

In the Widget Builder, switch types with the visualization tabs under your query. The preview updates immediately; configuration panels below the preview change based on the selected type.

<figure><img src="/files/Iw0xpPLIdLxZBxgrgxDs" alt=""><figcaption></figcaption></figure>

### Which visualizations each data type supports

Not every data type supports every visualization.

<table data-header-hidden><thead><tr><th width="156.90234375"></th><th width="98.63671875"></th><th width="83.734375"></th><th width="69.59765625"></th><th width="100.24609375"></th><th width="67.9140625"></th><th width="85.85546875"></th></tr></thead><tbody><tr><td>Data type</td><td>Time series</td><td>Table</td><td>Stat</td><td>Top list</td><td>Pie</td><td>Data</td></tr><tr><td><strong>Metrics</strong></td><td>✓</td><td>✓</td><td>✓</td><td>✓</td><td>✓</td><td>—</td></tr><tr><td><strong>Logs</strong></td><td>✓</td><td>✓</td><td>✓</td><td>✓</td><td>✓</td><td>✓</td></tr><tr><td><strong>Traces</strong></td><td>✓</td><td>✓</td><td>✓</td><td>✓</td><td>✓</td><td>✓</td></tr><tr><td><strong>Events</strong></td><td>✓</td><td>✓</td><td>✓</td><td>✓</td><td>✓</td><td>✓</td></tr><tr><td><strong>Entities</strong></td><td>—</td><td>✓</td><td>✓</td><td>✓</td><td>✓</td><td>—</td></tr><tr><td><strong>Issues</strong></td><td>✓</td><td>✓</td><td>✓</td><td>✓</td><td>✓</td><td>—</td></tr><tr><td><strong>RUM</strong></td><td>✓</td><td>✓</td><td>✓</td><td>✓</td><td>✓</td><td>—</td></tr><tr><td><strong>APM</strong></td><td>✓</td><td>✓</td><td>✓</td><td>✓</td><td>✓</td><td>—</td></tr><tr><td><strong>Ingestion</strong></td><td>✓</td><td>✓</td><td>✓</td><td>✓</td><td>✓</td><td>—</td></tr></tbody></table>

### Multiple queries in one widget

Some visualizations let you add **query B, C, …** or a **formula** query in the same widget.

| Visualization | Multiple queries | Formula |
| ------------- | ---------------- | ------- |
| Time series   | ✓                | ✓       |
| Table         | ✓                | ✓       |
| Stat          | ✓                | ✓       |
| Top list      | ✓                | ✓       |
| Pie           | ✓                | ✓       |
| Data          | —                | —       |

When a widget has more than one query (or a formula), results are shown in a **combined view.** For example, multiple series on one chart or merged table columns.

**Data** widgets always use a single query.

Formulas let you combine queries with arithmetic (for example `(A / B) * 100` for a ratio). See [Formulas](/use-groundcover/querying-your-groundcover-data/explore-and-monitors-query-builder#formulas) in the Query Builder guide for supported operators and examples.

### Comparing to a previous period

Metrics and APM queries support a **time offset**, which shifts a query back by a fixed duration (for example `1d`, `1w`) without changing the dashboard's time range. Add an offset to a query to overlay it against its own past — for example, plotting today's request rate next to the same metric from exactly one week ago on the same chart.

Each query in a widget can have its own offset, so you can compare more than one prior period at a time.

***

### Cross-cutting configuration options

These settings apply across several visualization types.

#### Rollup interval (bucket size)

For time-bucketed queries, the **rollup interval** controls how data is grouped in time (for example `1m`, `5m`, `1h`). It appears on the query row when the data type supports it.

The effective interval may be **adjusted automatically** to respect minimum bucket sizes for certain data types (for example ingestion queries enforce a minimum of ten minutes).

#### Data unit

The **Data Unit** panel controls how numeric values are formatted (labels, axes, legend stats, and stat tiles).

| Unit category | Examples                    |
| ------------- | --------------------------- |
| General       | Number, Percentage, Raw     |
| Time          | Nanoseconds through Minutes |
| Data size     | Bytes, Bytes/sec, KB/s      |
| Other         | mCPU, Dollars, Req/s        |

For **tables** with multiple numeric columns, you can set a **unit per column.**

Only stats column units can be edited.

#### High series count warning

Time series, table, top list, and pie widgets may return many series or groups. When the result count exceeds the default display limit (**1,000** series), a warning appears above the preview.

| Control          | Behavior                                               |
| ---------------- | ------------------------------------------------------ |
| **Show all**     | Renders more series (up to a hard cap of **5,000**)    |
| Performance note | Showing all series can slow rendering on dense queries |

**Stat** widgets do not offer this toggle; they display up to **50** stat tiles and truncate beyond that.

***

### Time series

**Use when:** You need to see how a metric or aggregated signal changes over the dashboard time range.

#### Chart display type

Under **Additional Styling → Display**:

| Option           | Appearance                                             |
| ---------------- | ------------------------------------------------------ |
| **Lines**        | Separate line per series (default for most data types) |
| **Stacked Bars** | Time buckets as stacked bars                           |
| **Area**         | Stacked area chart                                     |

#### Missing buckets

When a time bucket has no data point, you can choose how gaps are rendered:

| Option              | Effect                                    |
| ------------------- | ----------------------------------------- |
| **Nulls** (default) | Breaks the line / leaves the bucket empty |
| **Zeros**           | Treats missing buckets as zero            |

#### Color palette

Under **Additional Styling → Color**:

| Palette                                         | Description                                       |
| ----------------------------------------------- | ------------------------------------------------- |
| **Classic**                                     | Default product palette                           |
| **Smart**                                       | Semantic coloring (status-aware where applicable) |
| **Red / Green / Blue / Orange / Gray / Purple** | Fixed theme palettes                              |

Color applies across series in the chart (eight-color rotation for time series).

#### Legend

Under **Legend → Mode**:

| Mode        | Behavior                                  |
| ----------- | ----------------------------------------- |
| **Compact** | Inline legend chips below the chart       |
| **Table**   | Tabular legend with optional stat columns |
| **None**    | Legend hidden                             |

When **Table** mode is selected, choose which stat columns appear:

| Column    | Shows                               |
| --------- | ----------------------------------- |
| **Avg**   | Average over the visible time range |
| **Min**   | Minimum value                       |
| **Max**   | Maximum value                       |
| **Sum**   | Sum over the range                  |
| **Value** | Latest / instant value              |

#### Y-axis controls

| Setting                 | Description                                                            |
| ----------------------- | ---------------------------------------------------------------------- |
| **Scale**               | **Linear** (default) or **Log**                                        |
| **Min**                 | Fixed lower bound (leave empty for auto)                               |
| **Max**                 | Fixed upper bound (leave empty for auto)                               |
| **Always Include Zero** | Forces the axis to include zero when compatible with min/max and scale |

**Log scale** cannot include zero. If min is above zero or max is below zero, **Always Include Zero** is disabled.

#### Thresholds

Add reference lines or bands to mark meaningful values on the chart — for example an error rate ceiling or an SLO target.

| Setting      | Description                                                  |
| ------------ | ------------------------------------------------------------ |
| **Value**    | Where the threshold line/band is drawn on the Y-axis         |
| **Label**    | Optional text shown next to the threshold                    |
| **Severity** | **Error**, **Warning**, **OK**, or **Info** — controls color |
| **Style**    | **Solid** or **Dashed** line                                 |

Add up to **10 thresholds** per widget from the **Thresholds** tab in the widget builder.

#### Interaction on the dashboard

* **Zoom** on the time axis refines the dashboard time range (when chart zoom is enabled on the dashboard view).
* **Hover** over a series will show the value at a specific point along with indication for the step size.

### Table

**Use when:** You need sortable columns, grouped aggregates, or row-level detail in a grid.

#### Configuration in the widget builder

| Option               | Where           | Description                                          |
| -------------------- | --------------- | ---------------------------------------------------- |
| **Data unit**        | Data Unit panel | Default formatting for numeric value columns         |
| **Per-column units** | Data Unit panel | Available when multiple numeric columns are detected |
| **Rollup interval**  | Query row       | For time-bucketed SQL/metrics table queries          |

#### Combined tables

When multiple queries feed one table widget, columns from each query are **merged** into a single grid. Per-column units can be set for numeric columns from each query.

#### Exporting table data

Use **Export → CSV** in the widget menu to download the table's current rows as a CSV file.

#### Interaction on the dashboard

Hovering over a specific line will show the option to drill down further either in Explore, for aggregated drill down, or to the relevant data type page (e.g. Logs), to see the actual rows.

### Stat

**Use when:** You want a large numeric KPI - one number per series or group.

#### Configuration

| Option              | Description                        |
| ------------------- | ---------------------------------- |
| **Data unit**       | Controls value unit of measurement |
| **Rollup interval** | For time-aggregated stats          |

#### Behavior

* Each series or group renders as a **stat tile** in a responsive grid.
* Up to **50** stats are shown; additional series are truncated.
* With **multiple queries** or a formula, stats from all queries appear in the combined stat grid.

#### Conditional formatting

Color a stat tile based on its value instead of using a single fixed color. Define one or more rules, each with a comparison (`>`, `≥`, `<`, `≤`, `=`, or a range) and a color; the first matching rule wins.

| Setting        | Description                                        |
| -------------- | -------------------------------------------------- |
| **Rule**       | A comparison and value (or range) to match against |
| **Color**      | The color applied when the rule matches            |
| **Applies to** | **Text** color or the tile's **background**        |

Add up to **10 rules** per widget from the **Conditional Formatting** tab in the widget builder.

### Top list

**Use when:** You want a ranked horizontal bar chart (for example “top 10 namespaces by log volume”).

#### Configuration

| Option              | Where              | Description                                       |
| ------------------- | ------------------ | ------------------------------------------------- |
| **Data unit**       | Data Unit panel    | Value axis formatting                             |
| **Color**           | Additional Styling | Single-bar palette (one color theme for all bars) |
| **Rollup interval** | Query row          | When the query is time-bucketed                   |

#### Behavior

* Bars are ranked by the query’s sort/limit.
* Works with **multiple queries** in combined view (each query contributes bars).
* Supports the high-series warning and **Show all** toggle (same limits as time series).

### Pie

**Use when:** You want proportional breakdowns across a **small** number of categories.

#### Configuration

| Option              | Where              | Description                                       |
| ------------------- | ------------------ | ------------------------------------------------- |
| **Data unit**       | Data Unit panel    | Slice value formatting                            |
| **Color**           | Additional Styling | Multi-slice palette (same options as time series) |
| **Rollup interval** | Query row          | When applicable                                   |

#### Behavior

* Best suited to a limited number of slices (use query **limit** to keep the chart readable).
* Supports **multiple queries** in combined view.
* Supports the high-series warning and **Show all** toggle.

### Data

**Use when:** You want a **scrollable list of raw records -** similar to the Logs, Traces, or Events explorers - embedded in a dashboard tile.

#### Availability

| Requirement | Detail                                              |
| ----------- | --------------------------------------------------- |
| Data types  | **Logs**, **Traces**, or **Events** only            |
| Query mode  | **Builder** mode only (not the raw SQL/editor mode) |
| Query count | **Single query** per widget                         |

#### Configuration

| Option               | Description                                                 |
| -------------------- | ----------------------------------------------------------- |
| **Column selectors** | Choose which fields appear (builder query row)              |
| **Filters**          | Standard gcQL / filter builder for the signal type          |
| **Limit**            | Controls page size; scroll or load more for additional rows |


# Dashboard Version History

Every save to a dashboard creates a new revision. Use Version History to review past changes, compare them against the current version, and recover from an unwanted edit — without duplicating the dashboard or relying on Terraform state.

## Opening Version History

1. Open the dashboard.
2. From the **Actions** menu, select **Version History**.

A panel opens showing every revision, grouped by date, most recent first.

## Reviewing a revision

Select any revision in the list to preview a diff against the dashboard's current state — added, removed, and changed widgets are highlighted.

## Restoring a revision

From a previewed revision, choose **Restore** to make it the dashboard's current version. Restoring creates a new revision on top of the history rather than deleting anything, so you can always go back further if needed.

## Exporting a past revision

A revision doesn't have to be restored to be reused. Choose **Export** on any revision to get its JSON (or Terraform, see [Managing Dashboards with Terraform](/use-groundcover/dashboards-and-alerts/managing-dashboards-with-terraform)) without changing the dashboard's current state — useful for diffing a past layout or pulling an old widget's query.

{% hint style="info" %}
For dashboards managed by Terraform, Version History is a good first check before assuming a change needs to be re-applied from code — a revision may already contain what you're looking for.
{% endhint %}


# Managing Dashboards with Terraform

## Create & manage dashboards with Terraform

Use Terraform to **create, update, delete, and list** groundcover dashboards as code. Managing dashboards with infrastructure‑as‑code (IaC) lets you version changes, review them in pull requests, promote the same definitions across environments, and detect drift between what’s applied and what’s running in your account.

***

### Prerequisites

* A groundcover account with permissions to create/edit Dashboards
* A Terraform environment (groundcover provider >v1.1.1)
* The **groundcover Terraform provider** configured with your API credentials

> See also: [**groundcover Terraform provider**](/use-groundcover/groundcover-terraform-provider) reference for provider configuration and authentication details.

***

### 1) Creating a Dashboard via Terraform

#### 1.1) Create a dashboard directly from the UI

In order to create a dashboard using Terraform you first need to create a dashboard manually in order to export it in a Terraform format.

See [Creating dashboards](/use-groundcover/dashboards-and-alerts/create-a-dashboard) to learn more.

#### 1.2) Export the dashboard in Terraform format

You can export a Dashboard into as a Terraform resource:

1. Open the Dashboard.
2. Click **Actions → Export**.
3. Download or copy the **Terraform** tab’s content and paste it into your `.tf` file (see placeholder above).

<figure><img src="/files/wlUuTta5TZHtiihAUcBd" alt="" width="563"><figcaption></figcaption></figure>

#### 1.3) Add the dashboard resource to your Terraform configuration

{% hint style="info" %}
The example below is a placeholder, paste your generated snippet or hand‑write your own.
{% endhint %}

```hcl
resource "groundcover_dashboard" "llm_observability" {
  name             = "LLM Observability"
  description      = "Dashboard to monitor OpenAI and Anthropic usage"
  preset           = "{\"widgets\":[{\"id\":\"B\",\"type\":\"widget\",\"name\":\"Total LLM Calls\",\"queries\":[{\"id\":\"A\",\"expr\":\"span_type:openai span_type:anthropic | stats by(span_type) count() count_all_result | sort by (count_all_result desc) | limit 5\",\"dataType\":\"traces\",\"editorMode\":\"builder\"}],\"visualizationConfig\":{\"type\":\"stat\"}},{\"id\":\"D\",\"type\":\"widget\",\"name\":\"LLM Calls Rate\",\"queries\":[{\"id\":\"A\",\"expr\":\"sum(rate(groundcover_resource_total_counter{type=~\\\"openai|anthropic\\\",status_code=\\\"ok\\\"})) by (gen_ai_request_model)\",\"dataType\":\"metrics\",\"editorMode\":\"builder\"}],\"visualizationConfig\":{\"type\":\"time-series\",\"selectedChartType\":\"stackedBar\"}},{\"id\":\"E\",\"type\":\"widget\",\"name\":\"Average LLM Response Time\",\"queries\":[{\"id\":\"A\",\"expr\":\"avg(groundcover_resource_latency_seconds{type=~\\\"openai|anthropic\\\"}) by (type)\",\"dataType\":\"metrics\",\"step\":\"disabled\",\"editorMode\":\"builder\"}],\"visualizationConfig\":{\"type\":\"stat\",\"step\":\"disabled\",\"selectedUnit\":\"Seconds\"}},{\"id\":\"A\",\"type\":\"widget\",\"name\":\"Total LLM Tokens Used\",\"queries\":[{\"id\":\"A\",\"expr\":\"span_type:openai span_type:anthropic | stats by(span_type) sum(gen_ai.response.usage.total_tokens) sum_result | sort by (sum_result desc) | limit 5\",\"dataType\":\"traces\",\"editorMode\":\"builder\"}],\"visualizationConfig\":{\"type\":\"stat\",\"step\":\"disabled\"}},{\"id\":\"C\",\"type\":\"widget\",\"name\":\"AVG Input Tokens Per LLM Call \",\"queries\":[{\"id\":\"A\",\"expr\":\"span_type:openai OR span_type:anthropic | stats by(span_type) avg(gen_ai.response.usage.input_tokens) avg_result | sort by (avg_result desc) | limit 5\",\"dataType\":\"traces\",\"editorMode\":\"builder\"}],\"visualizationConfig\":{\"type\":\"stat\"}},{\"id\":\"F\",\"type\":\"widget\",\"name\":\"AVG Output Tokens Per LLM Call \",\"queries\":[{\"id\":\"A\",\"expr\":\"span_type:openai OR span_type:anthropic | stats by(span_type) avg(gen_ai.response.usage.output_tokens) avg_result | sort by (avg_result desc) | limit 5\",\"dataType\":\"traces\",\"editorMode\":\"builder\"}],\"visualizationConfig\":{\"type\":\"stat\",\"step\":\"disabled\"}},{\"id\":\"G\",\"type\":\"widget\",\"name\":\"Top Used Models\",\"queries\":[{\"id\":\"A\",\"expr\":\"span_type:openai OR span_type:anthropic | stats by(gen_ai.request.model) count() count_all_result | sort by (count_all_result desc) | limit 100\",\"dataType\":\"traces\",\"editorMode\":\"builder\"}],\"visualizationConfig\":{\"type\":\"bar\",\"step\":\"disabled\"}},{\"id\":\"H\",\"type\":\"widget\",\"name\":\"Total LLM Errors \",\"queries\":[{\"id\":\"A\",\"expr\":\"(span_type:openai OR span_type:anthropic) status:error | stats by(span_type) count() count_all_result | sort by (count_all_result desc) | limit 1\",\"dataType\":\"traces\",\"editorMode\":\"builder\"}],\"visualizationConfig\":{\"type\":\"stat\"}},{\"id\":\"I\",\"type\":\"widget\",\"name\":\"AVG TTFT Over Time by Model\",\"queries\":[{\"id\":\"A\",\"expr\":\"avg(groundcover_workload_latency_seconds{gen_ai_system=~\\\"openai|anthropic\\\",quantile=\\\"0.50\\\"}) by (gen_ai_request_model)\",\"dataType\":\"metrics\",\"editorMode\":\"builder\"}],\"visualizationConfig\":{\"type\":\"time-series\",\"selectedChartType\":\"line\",\"selectedUnit\":\"Seconds\"}},{\"id\":\"J\",\"type\":\"widget\",\"name\":\"Avg Output Tokens Per Second by Model\",\"queries\":[{\"id\":\"A\",\"expr\":\"avg(groundcover_gen_ai_response_usage_output_tokens{}) by (gen_ai_request_model)\",\"dataType\":\"metrics\",\"editorMode\":\"builder\"},{\"id\":\"B\",\"expr\":\"avg(groundcover_workload_latency_seconds{quantile=\\\"0.50\\\"}) by (gen_ai_request_model)\",\"dataType\":\"metrics\",\"editorMode\":\"builder\"},{\"id\":\"formula-A\",\"expr\":\"A / B\",\"dataType\":\"metrics-formula\",\"editorMode\":\"builder\"}],\"visualizationConfig\":{\"type\":\"time-series\",\"selectedUnit\":\"Number\"}}],\"layout\":[{\"id\":\"B\",\"x\":0,\"y\":0,\"w\":4,\"h\":6,\"minH\":4},{\"id\":\"D\",\"x\":0,\"y\":30,\"w\":24,\"h\":6,\"minH\":4},{\"id\":\"E\",\"x\":8,\"y\":0,\"w\":8,\"h\":6,\"minH\":4},{\"id\":\"A\",\"x\":16,\"y\":0,\"w\":8,\"h\":6,\"minH\":4},{\"id\":\"C\",\"x\":0,\"y\":12,\"w\":8,\"h\":6,\"minH\":4},{\"id\":\"F\",\"x\":8,\"y\":24,\"w\":8,\"h\":6,\"minH\":4},{\"id\":\"G\",\"x\":16,\"y\":24,\"w\":8,\"h\":6,\"minH\":4},{\"id\":\"H\",\"x\":4,\"y\":0,\"w\":4,\"h\":6,\"minH\":4},{\"id\":\"I\",\"x\":0,\"y\":18,\"w\":24,\"h\":6,\"minH\":4},{\"id\":\"J\",\"x\":0,\"y\":3,\"w\":24,\"h\":6,\"minH\":4}],\"duration\":\"Last 15 minutes\",\"variables\":{},\"spec\":{\"layoutType\":\"ordered\"},\"schemaVersion\":4}"
}
```

After saving this file as `main.tf` along with the provider details, type:

```
terraform plan
terraform apply
```

***

### 2) Managing existing provisioned Dashboard

#### 2.1) "Provisioned" badge for IaC‑managed Dashboards

Dashboards added via Terraform are marked as **Provisioned** in the UI so you can quickly distinguish IaC‑managed Dashboards from manually created ones, both from the Dashboard List and inside the Dashboard itself.

<figure><img src="/files/Ou0QmaUvzSYII2PVNPxO" alt="" width="207"><figcaption></figcaption></figure>

#### 2.2) Edit behavior for Provisioned Dashboards

Provisioned Dashboards are **read‑only by default** to protect the source of truth in your Terraform code.

* To make a quick change, click **Unlock dashboard**. This allows editing directly in the UI, all changes are automatically saved as always.\\

<figure><img src="/files/moWtvEVOEArkArYlYTOt" alt="" width="375"><figcaption></figcaption></figure>

* **Important:** Any changes can be **overwritten** the next time your provisioner runs `terraform apply`.
* Safer alternative: **Duplicate** the Dashboard and edit the copy, then migrate those changes back into code.
* If changes are overwritten unexpectedly, [Dashboard Version History](/use-groundcover/dashboards-and-alerts/dashboard-version-history) can recover the pre-`apply` revision without waiting on another Terraform run.

#### 2.3) Editing dashboards via Terraform

Changing the resource and reapplying Terraform willupdate the Dashboard in groundcover.

Deleting the resource from your code (and applying) will delete it from groundcover.

See more examples on our [Github repo](https://github.com/groundcover-com/terraform-provider-groundcover/tree/main/examples/resources/groundcover_dashboard).

***

### 3) Importing existing Dashboards into Terraform

Already have a Dashboard in groundcover? Bring it under Terraform management without recreating it:

```bash
# Syntax
terraform import groundcover_dashboard.<local_name> <dashboard_id>

# Example
terraform import groundcover_dashboard.service_overview dsh_1234567890
```

After importing, run `terraform plan` to view the state and align your config with what exists.

***

### Reference

* [**Creating dashboars**](/use-groundcover/dashboards-and-alerts/create-a-dashboard) – how to build widgets and layouts in the UI
* [groundcover Terraform provider documentation](/use-groundcover/groundcover-terraform-provider)
* [**groundcover Terraform provider Github repo**](https://github.com/groundcover-com/terraform-provider-groundcover) – resource schema, arguments, and examples


# Agent Mode

AI-powered assistant for investigating, exploring, and building with your groundcover data

{% hint style="info" %}
**Availability:** Agent Mode is available for groundcover Cloud, all BYOC deployments (AWS, GCP, Azure), and on-premises deployments with a supported LLM provider. Some environments require additional setup. See [Requirements & Compatibility](/use-groundcover/agent-mode/requirements).
{% endhint %}

groundcover Agent Mode is an AI-powered assistant built into the groundcover platform. Interact with it in natural language to investigate issues, explore your data, create dashboards and monitors, and more - across logs, traces, metrics, events, and entities.

## What You Can Do

### Investigate Issues

Ask the Agent to diagnose problems across your infrastructure. It queries logs, traces, metrics, and Kubernetes events, then reports what it found and what likely caused the issue.

> "Why is the checkout service throwing 500 errors?" "What changed recently in the payments namespace?" "Triage the alert on high-error-rate monitor"

### Explore Your Data

Query any signal type using plain language.

> "Show me the top 10 error-producing workloads in production" "What are the slowest endpoints in the API gateway?" "Which pods are consuming the most memory?"

### Create Dashboards & Monitors

Build production-ready dashboards and monitoring rules from a description.

> "Build a dashboard for the payments namespace" "Create a monitor that alerts when error rate exceeds 5% on the checkout service"

### Parse & Manage Logs

Generate log parsing rules and drop rules without learning OTTL syntax.

> "Parse the unstructured logs from the nginx workload" "Create a drop rule for health check logs"

### Navigate the Platform

The Agent can take you to the right page in the groundcover UI with filters already applied.

> "Take me to the traces view for the auth service filtered to errors"

## Signals

The Agent works across all groundcover signal types:

| Signal       | What It Contains                                                         | Example Use                                  |
| ------------ | ------------------------------------------------------------------------ | -------------------------------------------- |
| **Logs**     | Application log lines with level, format, and content                    | Find error patterns, parse unstructured logs |
| **Traces**   | Distributed traces with duration, status codes, and service dependencies | Diagnose latency, trace error paths          |
| **Metrics**  | Prometheus-compatible time-series data                                   | Monitor resource usage, track SLIs           |
| **Events**   | Kubernetes events and infrastructure changes (deploys, crashes, scaling) | Correlate changes with incidents             |
| **Entities** | Live Kubernetes resource state (pods, deployments, nodes, services)      | Check current health, find resource issues   |
| **Issues**   | Active and resolved monitor alerts with severity levels                  | Triage alerts, review incident history       |

## Next Steps

* [Requirements & Compatibility](/use-groundcover/agent-mode/requirements) - Supported deployments, provider options, and prerequisites
* [Getting Started](/use-groundcover/agent-mode/getting-started) - Learn how to interact with Agent Mode
* [Skills](/use-groundcover/agent-mode/skills) - Teach Agent Mode personal and organizational workflows and conventions
* [Example Prompts](/use-groundcover/agent-mode/example-prompts) - Copy-paste prompts for common scenarios
* [Configuring Settings](/use-groundcover/agent-mode/configuring-settings) - Enable Agent Mode, choose an LLM provider, set Org Instructions, and manage budgets
* [Privacy & Security](/use-groundcover/agent-mode/privacy-and-security) - Tenant isolation, session management, data retention, and access control
* [Cost Management](/use-groundcover/agent-mode/cost-management) - Set monthly spending limits for Agent Mode
* [Connectors](/use-groundcover/connectors) - Connect external tools like Cursor so Agent Mode can take actions on your behalf
* [@groundcover Agent in Slack](/use-groundcover/connectors/slack/slack-agent) - Use the Agent directly from Slack via @mentions
* [MCP Integration](/getting-started/groundcover-mcp) - Use groundcover as a tool for external AI agents (Cursor, Claude Desktop, etc.)


# Requirements & Compatibility

Deployment types and prerequisites for using Agent Mode

Agent Mode requires network access to a supported LLM endpoint and the authentication method that endpoint expects. The endpoint can be a native cloud LLM service, the Anthropic API, or an Anthropic-compatible proxy. This page covers which deployment types support Agent Mode and what setup is needed. See [Privacy & Security](/use-groundcover/agent-mode/privacy-and-security#llm-provider) for provider-specific data handling.

## Supported Deployments

| Deployment Type                 | Agent Mode | Setup Required                                                                                                  |
| ------------------------------- | ---------- | --------------------------------------------------------------------------------------------------------------- |
| **groundcover Cloud**           | Yes        | None                                                                                                            |
| **BYOC AWS**                    | Yes        | None                                                                                                            |
| **BYOC GCP**                    | Yes        | [Enable models in Vertex AI](/architecture/byoc/setup-byoc-with-gcp/enable-ai-models-for-agent-mode)            |
| **BYOC Azure**                  | Yes        | None ([what gets provisioned](/architecture/byoc/setup-byoc-with-azure/enable-ai-models-for-agent-mode))        |
| **On-premises (AWS/GCP/Azure)** | Yes        | Use your own Anthropic API key or compatible endpoint, or contact your account team for a native cloud provider |
| **On-premises (other)**         | Yes        | Use your own Anthropic API key or compatible endpoint                                                           |

{% hint style="info" %}
**No cost until you use it.** Agent Mode is enabled by default on groundcover Cloud and BYOC deployments once provisioning completes, and LLM inference costs are incurred only when users start using the Agent. If you use your own Anthropic API key, Anthropic bills the usage to your account. GCP BYOC requires a manual setup step because Google Cloud requires customers to explicitly approve third-party Anthropic models in Vertex AI Model Garden before they can be invoked. AWS Bedrock and Microsoft Foundry have no equivalent gate.
{% endhint %}

For on-premises deployments on AWS, GCP, or Azure, groundcover can provision the native LLM service: AWS Bedrock, Google Cloud Vertex AI, or Microsoft Foundry. You can also use a direct Anthropic API key or a custom Anthropic-compatible endpoint on any supported backend. All configurations require network access to the selected LLM provider or proxy.

## Custom LLM providers

You can override the native cloud provider for an individual backend from **Settings > AI & Agents > Management**.

### Anthropic API key

Agent Mode supports using your own Anthropic API key for inference on any backend where AI features are available. This is useful when you want usage billed through your Anthropic account or when a native cloud-provider LLM service is not available for your deployment.

Workspace administrators configure the key separately for each backend. Inference requests are sent directly to Anthropic after the key is configured. See [Configuring Settings](/use-groundcover/agent-mode/configuring-settings#anthropic-api-key) for the self-service setup and key rotation steps.

### Custom Anthropic domain

**Custom Anthropic Domain** sends inference requests to an Anthropic-compatible endpoint using its base URL and API key. You can use LiteLLM, Bifrost, or any other proxy or gateway that implements the Anthropic Messages API and exposes the Claude models Agent Mode uses.

Your groundcover deployment must have network access to the custom endpoint. The endpoint must support streaming Anthropic Messages API responses. See [Configuring Settings](/use-groundcover/agent-mode/configuring-settings#choose-an-llm-provider) for setup instructions and endpoint examples.

## GCP BYOC setup

GCP BYOC deployments require manually enabling the Anthropic Claude models in your project's Vertex AI Model Garden before Agent Mode will work. See [Enable AI models for Agent Mode](/architecture/byoc/setup-byoc-with-gcp/enable-ai-models-for-agent-mode) for step-by-step instructions.


# Getting Started

Learn how to interact with groundcover Agent Mode

No setup required - the Agent is available directly in the groundcover UI.

## Opening the Agent

Click the **Agent** icon in the sidebar or press **Cmd+Shift+A** (Mac) / **Ctrl+Shift+A** (Windows/Linux) to toggle the chat panel. The panel stays open as you navigate the platform so you can ask questions in context.

Press **Cmd+Shift+E** to expand the Agent to full screen for longer investigations.

## Template Cards

When you start a new conversation, the Agent shows **template cards** - pre-built prompts covering common scenarios like error investigation, performance analysis, infrastructure health, and more. Click a card, fill in the placeholders for your specific workload or namespace, and the Agent starts working.

## @Mentions

Type **@** in the chat input to reference specific entities. A picker appears with your resources grouped by category:

* **Workloads** - Deployments, StatefulSets, DaemonSets, CronJobs, Jobs
* **Pods**
* **Nodes**

Recently mentioned entities appear at the top for quick access. You can search by name or click a category header to filter.

> Investigate errors in @checkout-service Compare latency between @api-gateway and @auth-service

When you mention an entity, its full context (namespace, cluster, environment) is automatically included in your request.

## / for Skills

Type **/** in the chat input to attach a [Skill](/use-groundcover/agent-mode/skills) to your next message. A picker shows your Skills; select one and the Agent will follow its Instructions for that prompt.

The Agent also activates matching Skills automatically based on each Skill's *When to use* field - typing `/` is for when you want to force a specific Skill.

## Personal Instructions

Open **Instructions** from the Agent sidebar to set guidance the Agent should follow in every conversation you start. Use Personal Instructions for preferences that should always apply to your own Agent sessions, such as response format, naming conventions, query defaults, or systems you commonly prioritize.

Personal Instructions are different from Skills:

* **Personal Instructions** are always included in your conversations.
* **Skills** apply when you attach them with `/` or when the Agent matches their **When to use** field.
* **Organizational Skills** can be shared with the whole organization, while Personal Instructions stay scoped to your user.

To reset your Personal Instructions, clear the editor and save the empty value.

## Adding Context from the UI

While browsing the platform, you can send specific items directly to the Agent as context. Hover over any row in a table or open a detail drawer and click the **Add to context** button (star icon). The item and its full metadata are attached to your next Agent message.

This lets you point the Agent at exactly what you're looking at without copy-pasting IDs or describing it manually. For example, add a failing span to context and ask "why is this span erroring?" - the Agent already has all the details.

You can also **paste a groundcover URL** directly into the chat. The Agent extracts the full context from the URL - page, filters, time range, and entity.

## Exploring Tool Results

As the Agent works, each step it takes - every query, lookup, or action - appears as an expandable item in the conversation. Click on any step to see what the Agent actually did: the query it ran, the parameters it used, and the raw results it got back.

This gives you full transparency into the Agent's reasoning and lets you verify its work, reuse queries, or continue exploring the data yourself.

## Suggestions

After completing an investigation or answering a question, the Agent may offer **follow-up suggestions** - clickable buttons that continue the conversation in a useful direction. Click a suggestion to automatically send it as your next prompt.

## Conversations

The Agent maintains full context within a conversation. Ask follow-ups naturally:

> **You:** What's causing high latency in the payments service?
>
> **Agent:** *(investigates and reports findings)*
>
> **You:** Drill into the database calls specifically
>
> **Agent:** *(focuses on database spans)*
>
> **You:** Create a monitor for this

Previous conversations are saved in the sidebar. Start a new one with **Cmd+Shift+O**.

## Sharing & Forking

You can share any conversation with teammates. Right-click a conversation in the sidebar and select **Share** to copy a shareable link to your clipboard. Shared conversations have an expiration time.

When someone opens a shared link, they see a read-only view of the full conversation. To continue the investigation from where it left off, they can click **Continue this conversation** - this forks the conversation into a new, editable copy owned by them. The original shared conversation stays unchanged.

## Keyboard Shortcuts

| Shortcut                 | Action                     |
| ------------------------ | -------------------------- |
| **Cmd/Ctrl + Shift + A** | Toggle the Agent panel     |
| **Cmd/Ctrl + Shift + O** | New conversation           |
| **Cmd/Ctrl + Shift + E** | Expand to full screen      |
| **Cmd/Ctrl + Shift + S** | Share conversation         |
| **@**                    | Open entity mention picker |
| **Cmd/Ctrl + /**         | Show keyboard shortcuts    |


# Skills

Teach the Agent your team's workflows, runbooks, and conventions with custom Skills

Skills are reusable instructions you write once and the Agent follows whenever they apply. Use them to encode personal preferences, team runbooks, naming conventions, investigation playbooks, or any context you'd otherwise repeat in every prompt.

{% hint style="info" %}
Skills can be personal or organizational. Personal Skills are visible only to you. Organizational Skills are readable and usable by everyone in the organization, and only admins can create, edit, or delete them.
{% endhint %}

## Personal vs. Organizational Skills

Use the Skill scope to decide who should be able to use the instructions:

| Scope                    | Who can use it               | Who can manage it                           | Best for                                                                 |
| ------------------------ | ---------------------------- | ------------------------------------------- | ------------------------------------------------------------------------ |
| **Personal Skill**       | Only the user who created it | The creator, if they have Skill permissions | Individual response preferences, personal shortcuts, private workflows   |
| **Organizational Skill** | Everyone in the organization | Admins only                                 | Shared investigation runbooks, team conventions, common operating guides |

Organizational Skills appear with an **Org** badge in the Skills list and detail view. They can be selected from the Skill picker and can also auto-activate from their **When to use** field, just like personal Skills.

## Creating a Skill

Open the **Skills** page from the Agent sidebar. Click **New Skill** and fill in:

| Field                    | Required | Purpose                                                                                                                   |
| ------------------------ | -------- | ------------------------------------------------------------------------------------------------------------------------- |
| **Name**                 | Yes      | A short identifier shown in the picker (e.g. `Incident Triage`).                                                          |
| **When to use**          | Yes      | Plain-language description of when this Skill should apply. The Agent reads this to decide whether to activate the Skill. |
| **Description**          | No       | Short summary shown alongside the name in the picker.                                                                     |
| **Instructions**         | Yes      | Markdown describing how the Agent should operate when this Skill is active.                                               |
| **Organizational Skill** | No       | Admin-only toggle. Turn it on to make the Skill available to everyone in the organization.                                |

Click **Save**. Personal Skills are available immediately in your next Agent message. Organizational Skills are available to everyone in the organization after they are saved.

## Ask the Agent to draft a Skill for you

You don't have to write Skills from scratch. Once the Agent has done useful work in a conversation - an RCA, a triage walk-through, a parsing-rule investigation, a dashboard build - you can ask it to turn that work into a Skill without leaving the chat. The Agent drafts a Name, When to use, and Instructions from the current conversation, and you review and save. If you are an admin, you can choose whether the draft stays personal or becomes an Organizational Skill before saving.

The better your ask, the better the resulting Skill. Thin asks produce thin Skills:

> create a skill for investigating DB-related errors for checkout-service

A better ask spells out the activation trigger, the scope, and which parts of the conversation are the reusable pattern:

> Turn this investigation into a Skill. Call it "Postgres slow-query triage for checkout". Activate it when I ask about checkout-service latency or errors that look database-bound. The Instructions should keep the steps we just did: pull `pg_stat_statements` top queries over the incident window, correlate with the deploy timeline of `checkout-api`, then trace one failing request end-to-end. Default lookback is 1 hour, and always use P95 for latency.

Or:

> Save this conversation as a Skill for Redis authentication failures. Name it "Redis auth troubleshooting". Activate it when Redis clients in `cache-*` namespaces start failing auth after a deploy or config change. Instructions should capture exactly what we just checked - Redis config history, client library version, and TLS cert expiry - in that order. Skip the side tangent about node labels.

The Agent produces a better Skill when you tell it:

* **The activation condition** - which services, namespaces, alert shapes, or user phrasings should trigger it. Vague activation causes the Skill to either miss the prompt or auto-pull on unrelated questions.
* **Which parts of the current conversation are the reusable pattern** vs. incidental details. "Keep the three checks we did, drop the bit about node labels."
* **Conventions that should survive** - your preferred metric (P95 vs P99), default lookback, output format, namespaces to ignore, services to exclude.
* **A Name** - or let the Agent propose one. "Postgres slow-query triage" beats "DB stuff."

The draft is a starting point - edit any field before saving, or tweak the Skill from the Skills page later.

## Using a Skill

There are two ways a Skill gets activated in a conversation:

**Explicit** - type **`/`** in the Agent input. A picker shows your Skills; select one to attach it to your next message.

> `/incident-triage` investigate the alert on checkout-service

**Automatic** - the Agent reads the **When to use** field of all your Skills and activates the relevant ones based on your prompt. For example, if your Skill's When to use says "Use when I ask for incident triage or RCA follow-ups", asking the Agent to investigate an incident will pull it in automatically.

When a Skill is active, its Instructions are injected into the Agent's context alongside the base system prompt. The Agent will mention which Skills it used, and you can see them attached to the message.

## Writing Good Instructions

The Instructions field is Markdown - structure it like a runbook or spec. A few patterns that work well:

**Goals and scope**

```markdown
# Goal
Run a 5-whys root cause analysis for production incidents.

# Scope
Only for workloads in the `prod-*` namespaces.
```

**Step-by-step procedures**

```markdown
# Steps
1. Pull error logs from the last 30 minutes, grouped by pattern.
2. Correlate with deploy events in the same window.
3. Check upstream and downstream services for related errors.
4. Summarize findings with Who / What / Where / When / Why.
```

**Conventions and preferences**

```markdown
# Output format
- Lead with a one-line summary.
- Use a bulleted timeline for the chronology.
- Always include the monitor name and namespace.

# Query style
- Prefer P95 over P99 for latency charts.
- Default lookback is 1 hour unless I say otherwise.
```

### Tips

* **Be specific about when it applies** - the Agent uses *When to use* to decide whether to pull the Skill in. Vague descriptions lead to unexpected activations or missed ones.
* **Write for the Agent, not for a human reader** - imperative instructions ("Always include X", "Never do Y") are followed more reliably than prose.
* **Keep it focused** - one Skill per workflow. Two Skills with overlapping triggers are fine; one giant Skill covering every scenario is harder to maintain.
* **Iterate** - if the Agent doesn't behave as expected, refine the Instructions and try again. Changes take effect on the next message.

## A Full Example: Payments Incident Triage

Here's a complete, copy-ready Skill that codifies how one engineer wants the Agent to handle incidents on their payments service. Use it as a starting point and adapt it to your own stack.

**Name**

```
Payments Incident Triage
```

**When to use**

```
Use when I ask to triage an alert, investigate an incident, or do an RCA on any service in the `payments-*` namespaces.
```

**Description**

```
5-whys triage playbook for the payments platform.
```

**Instructions**

````markdown
# Goal
Triage a production incident on the payments platform end-to-end and produce a written summary I can paste into the incident channel.

# Scope
Only for workloads in `payments-api`, `payments-worker`, `payments-ledger`, and `payments-gateway` namespaces. If the alert is for a different service, say so and stop.

# Procedure
1. **Anchor the timeline.** Find when the alert first fired and when symptoms started in the data. Use the earlier of the two as t=0.
2. **Check the blast radius.** For the affected workload, report:
   - Error rate (last 1h, broken down by endpoint)
   - P95 latency (last 1h vs. the prior 24h baseline)
   - Pod-level restarts, OOMKills, and crashloops in the window
3. **Look for change correlation.** In a ±30 min window around t=0, list:
   - Deploys (image tag changes) for any payments-* workload
   - ConfigMap / Secret updates
   - HPA scaling events
   - Upstream incidents (check `auth-*` and `fraud-*` dependencies)
4. **Trace a bad request.** Pull one representative failing trace and walk through the span tree. Identify the first span that shows the error.
5. **Check the database.** Query latency and error rate on the Postgres calls from `payments-ledger`. Flag any query pattern that crossed 500ms P95.
6. **Summarize** in this exact format:

   ```
   **Impact:** <one line>
   **Likely cause:** <one line, cite the evidence>
   **Timeline:**
   - HH:MM — <event>
   - HH:MM — <event>
   **Suggested next step:** <one action>
   **Evidence:** <links to the groundcover views you used>
   ```

# Conventions
- Always use P95 for latency, never average.
- Default lookback is 1 hour; extend to 6 hours only if the alert is older than that.
- Ignore `payments-staging-*` namespaces.
- If you can't find deploy events, say so explicitly — don't guess.
````

### Why this Skill works

A few things make this Skill reliable in practice:

* **The&#x20;*****When to use*****&#x20;is specific.** It names the namespace prefix and the types of prompts that should trigger it. The Agent uses this field to decide whether to auto-activate the Skill, so vague wording like "use for incidents" would cause it to fire on unrelated services.
* **The scope is bounded.** Explicitly listing the four payments namespaces and telling the Agent to stop if the alert is elsewhere prevents it from applying payments-specific assumptions to, say, an auth service outage.
* **Steps are imperative and ordered.** "Anchor the timeline", "Check the blast radius", "Trace a bad request" — each step is a concrete action the Agent can execute against groundcover data. Prose like "investigate thoroughly" produces inconsistent results.
* **The output format is pinned.** Giving the Agent an exact template (Impact / Likely cause / Timeline / Next step / Evidence) means every triage summary looks the same and is safe to paste into an incident channel without reformatting.
* **Defaults and exceptions are explicit.** "Default lookback is 1 hour", "Ignore staging namespaces", "Don't guess if deploys aren't found" — each of these is a decision the Agent would otherwise make inconsistently from one run to the next.

## Managing Skills

On the Skills page you can:

* **Search** Skills by name.
* **Edit** a personal Skill - changes apply to new messages immediately.
* **Delete** a personal Skill - removes it from the picker and stops auto-activation.
* **Manage** an Organizational Skill - admins can edit or delete it from the same page.

If an admin changes an Organizational Skill back to a personal Skill, groundcover asks for confirmation. Once the Skill is personal, only its creator can see it.

Skills are versioned - when a conversation uses a Skill, the revision in use at send time is what the Agent saw.


# Example Prompts

Ready-to-use prompts for common investigation, exploration, and creation scenarios

Copy-paste these prompts to get started. Replace the `@mentions` (e.g. `@workload`, `@namespace`) with the actual entity names from your environment - the Agent autocompletes them as you type.

## Incident Investigation

**Error Investigation**

> Investigate errors in `@workload`. Check error rates, identify the most common error patterns in logs and traces, and correlate with recent events.

**Root Cause Analysis**

> Perform a root cause analysis on `@workload`: Who is affected, What is failing, Where in the call chain, When it started, and Why.

**Alert Triage**

> Triage the alert on `@monitor`. Check the current state, review recent firings, and investigate the underlying cause.

**Crashloop Diagnosis**

> Why is `@workload` crashlooping? Check for OOMKill events, crash events, and recent changes.

***

## Performance

**Slow Endpoints**

> Find the slowest endpoints in `@workload`. Show P50, P95, and P99 latency broken down by path.

**Slow Database Queries**

> Find the slowest database queries called by `@workload`. Show query patterns, durations, and which endpoints trigger them.

**Resource Usage**

> Investigate CPU and memory usage for `@workload`. Show trends, compare against requests/limits, and flag anomalies.

***

## Data Exploration

**Top Errors**

> Show me the top 10 error-producing workloads in `@namespace` over the last 24 hours.

**Log Patterns**

> What are the most common log patterns for `@workload`? Group by pattern and show hit counts.

**Change Correlation**

> What changed recently in `@workload`? Check for image updates, config changes, scaling events, and restarts.

**Dependencies**

> Show me the dependencies of `@workload` - which services does it call, and which services call it?

***

## Building

**Dashboard**

> Build a dashboard for `@namespace` covering: error rates by workload, latency P95, pod restarts, and resource utilization.

**Monitor**

> Create a monitor that fires when error rate on `@workload` exceeds 5% for more than 5 minutes, with severity S2.

**Log Parsing**

> Parse the unstructured logs from `@workload`. Extract structured fields from the raw log content.

**Drop Rules**

> Create a drop rule to filter out health check logs from `@workload`.

***

## Infrastructure

**Node Pressure**

> Are any nodes experiencing CPU, memory, or disk pressure? Show affected nodes and the workloads running on them.

**Over-Provisioned Resources**

> Find workloads in `@namespace` that are over-provisioned - using significantly less CPU or memory than their requests.

**Pod Health**

> Show me all unhealthy pods in `@namespace` - crashlooping, pending, or evicted.

***

## Smart Filtering

{% hint style="info" %}
**Replaces auto-generated filters.** The legacy auto-generated-filters feature has been retired. Use the prompts below in Agent Mode to get the same filter suggestions, with more control over scope and intent.
{% endhint %}

Use the Agent to suggest relevant filters for your investigations - just describe what you're looking at.

**Suggest Filters**

> I'm investigating errors in `@namespace`. What filters should I use to narrow down the root cause?

**Find Log Patterns**

> What are the most important log patterns I should filter for when investigating `@workload`?

**Identify Key Attributes**

> What attributes should I filter on to find performance issues in `@workload`?

**Scope by Error Type**

> Help me isolate connection-timeout errors in `@namespace`.

***

## Tips for Effective Prompts

* **Be specific about scope** - include the workload, namespace, or cluster you care about
* **State what you want to see** - "Show me error rates broken down by endpoint" beats "look at errors"
* **Mention the signal** - if you specifically want traces vs. logs, say so
* **Ask follow-ups** - the Agent keeps context, so narrow down iteratively
* **Use @mentions** - reference entities with `@name` for precise resolution; the Agent autocompletes from your live data
* **Turn repeated context into a** [**Skill**](/use-groundcover/agent-mode/skills) - if you find yourself writing the same framing ("always use P95", "this is a payments service", "start with the monitor") across prompts, save it as a Skill and the Agent will apply it automatically
* **Let the UI do the work** - if you're already filtered to the right scope, just ask the question


# Configuring Settings

How to configure Agent Mode access, LLM providers, instructions, and budgets for your workspace

Workspace administrators control Agent Mode availability, organization-wide instructions, and budget settings from the Settings page.

## Accessing AI Settings

Navigate to **Settings > AI & Agents** in the groundcover UI. The page is split into three sections:

* **Management**: enable or disable Agent Mode and choose an LLM provider per backend
* **Org Instructions**: set backend-wide guidance the Agent follows for everyone
* **Cost Control**: set monthly spending limits at the organization and per-user level

{% hint style="info" %}
**Admin role required.** Only users with the Admin role can view and modify AI settings. If you don't see the **AI & Agents** option in the Settings sidebar, contact your workspace administrator.
{% endhint %}

## Management

The Management section controls whether Agent Mode is available and which LLM provider it uses for each backend. Expand a backend to view its settings.

**When enabled:**

* The Agent panel appears in the sidebar
* Users can open the Agent with **Cmd/Ctrl + Shift + A**
* All AI-related UI elements are visible

**When disabled:**

* The Agent panel and all AI-related UI elements are hidden
* Keyboard shortcuts for the Agent are inactive

This setting is per backend. You can enable Agent Mode for some backends while keeping it disabled for others, and each enabled backend can use a different LLM provider.

### Choose an LLM provider

| Provider                    | When to use it                                                                                                                                                                                                            |
| --------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Native Cloud Provider**   | Use the LLM service included with your deployment, such as AWS Bedrock, Google Cloud Vertex AI, or Microsoft Foundry. See [Requirements & Compatibility](/use-groundcover/agent-mode/requirements#supported-deployments). |
| **Anthropic Key**           | Send requests directly to Anthropic and bill usage to your Anthropic account.                                                                                                                                             |
| **Custom Anthropic Domain** | Send requests to an Anthropic-compatible proxy or gateway, such as LiteLLM, Bifrost, or another service that implements the Anthropic Messages API.                                                                       |

## Anthropic API key

You can bring your own Anthropic API key and configure Agent Mode to send inference requests directly to Anthropic. The provider selection and key are configured separately for each backend.

To configure an Anthropic API key:

1. Go to **Settings > AI & Agents > Management**.
2. Find the backend you want to configure and make sure **AI enabled** is turned on.
3. Expand the backend settings.
4. Under **Choose your LLM Provider**, select **Anthropic Key**.
5. Enter your Anthropic API key.
6. Select **Save**.

groundcover validates the key before activating it. A valid key takes effect without restarting the backend. If validation fails, the existing provider configuration remains active.

After the key is saved, the API key field displays a masked value. To rotate the key, expand the backend settings, select **Edit** next to the API key field, enter the replacement key, and select **Save**.

To return to the LLM provider configured for your deployment, select **Native Cloud Provider** and then select **Save**. A native provider must already be configured for the backend.

See [Requirements & Compatibility](/use-groundcover/agent-mode/requirements#anthropic-api-key) for deployment support and [Privacy & Security](/use-groundcover/agent-mode/privacy-and-security#llm-provider) for data handling and key storage details.

## Configure a custom Anthropic domain

Use **Custom Anthropic Domain** to route Claude traffic through an internal gateway, centralized model proxy, or another Anthropic-compatible endpoint.

Before you begin, make sure the proxy:

* Is accessible from your groundcover deployment
* Implements the Anthropic Messages API, including streaming responses
* Supports API key authentication

To configure the endpoint:

1. Go to **Settings > AI & Agents > Management**.
2. Expand the backend you want to configure.
3. Select **Custom Anthropic Domain** under **Choose your LLM Provider**.
4. Enter the proxy's Anthropic base URL in **Endpoint URL**, including any required path.
   * The `/v1/messages` path is added automatically.
   * Use HTTPS for external endpoints. Use HTTP only for internal endpoints protected by authenticated encryption, such as a TLS service mesh.
   * Use a URL without query parameters or fragments.
5. Enter the API key accepted by the proxy.
6. Click **Save**.

For example, depending on how your proxy is exposed, the base URL could be:

```
https://litellm.example.com
https://bifrost.example.com/anthropic
https://litellm.ai-gateway.svc.cluster.local:4000
```

{% hint style="info" %}
When you save, groundcover makes a model request to validate the endpoint, API key, and model access before applying the change. If validation fails, the existing provider remains active. Successful validation confirms access only at save time. Later availability depends on the proxy and its upstream provider.
{% endhint %}

## Org Instructions

The Org Instructions section lets admins set guidance the Agent follows in Agent Mode for everyone in the selected backend. Use Org Instructions for broad defaults that should apply across the team, such as:

* Preferred terminology and service names
* Default response structure
* Systems or namespaces the Agent should prioritize
* Organization-wide investigation or escalation expectations

Org Instructions are included alongside each user's [Personal Instructions](/use-groundcover/agent-mode/getting-started#personal-instructions) and any active [Skills](/use-groundcover/agent-mode/skills). Keep Org Instructions broad and stable; use Organizational Skills for reusable runbooks or workflows that should activate only for specific prompts.

To update Org Instructions, edit the text and click **Save**. To remove them, clear the editor and save the empty value.

## Cost Control

The Cost Control section lets admins set monthly spending limits. See [Cost Management](/use-groundcover/agent-mode/cost-management) for full details on configuring organization budgets, default per-user budgets, and per-user overrides.

## Related

* [Getting Started](/use-groundcover/agent-mode/getting-started#personal-instructions) - Configure Personal Instructions for your own Agent conversations
* [Skills](/use-groundcover/agent-mode/skills) - Create personal or Organizational Skills for reusable Agent guidance
* [Cost Management](/use-groundcover/agent-mode/cost-management) - Set monthly spending limits at the organization and per-user level
* [Privacy & Security](/use-groundcover/agent-mode/privacy-and-security) - Data handling and access control details


# Privacy & Security

Data handling, LLM providers, tenant isolation, session management, and access control for groundcover Agent Mode

{% hint style="info" %}
Agent Mode requires access to a supported LLM endpoint. Some deployment types require additional setup. See [Requirements & Compatibility](/use-groundcover/agent-mode/requirements).
{% endhint %}

## LLM Provider

The groundcover Agent uses a supported LLM provider for inference. The provider depends on your deployment and selected configuration:

| Deployment                  | LLM Provider                                                                      |
| --------------------------- | --------------------------------------------------------------------------------- |
| groundcover Cloud (default) | **AWS Bedrock** - Anthropic Claude models                                         |
| BYOC AWS                    | **AWS Bedrock** - Anthropic Claude models                                         |
| BYOC GCP                    | **Google Cloud Vertex AI** - Anthropic Claude models                              |
| BYOC Azure                  | **Microsoft Foundry** - Anthropic Claude models                                   |
| On-premises (AWS)           | **AWS Bedrock** - Anthropic Claude models (provisioned by groundcover)            |
| On-premises (GCP)           | **Google Cloud Vertex AI** - Anthropic Claude models (provisioned by groundcover) |
| On-premises (Azure)         | **Microsoft Foundry** - Anthropic Claude models (provisioned by groundcover)      |
| On-premises (other)         | **Anthropic API** - Anthropic Claude models (your API key)                        |

You can override the deployment provider for an individual backend from **Settings > AI & Agents > Management**:

| Override                    | Request destination                                                                 |
| --------------------------- | ----------------------------------------------------------------------------------- |
| **Anthropic Key**           | Anthropic's API, authenticated with your Anthropic API key                          |
| **Custom Anthropic Domain** | The Anthropic-compatible endpoint you configure, such as a LiteLLM or Bifrost proxy |

Data handling for **Anthropic Key** follows your Anthropic agreement and account settings. Data handling for **Custom Anthropic Domain** follows the policies of your proxy and its upstream model provider. See [Configuring Settings](/use-groundcover/agent-mode/configuring-settings#choose-an-llm-provider) for custom endpoint setup and [Anthropic API key](/use-groundcover/agent-mode/configuring-settings#anthropic-api-key) for direct Anthropic setup and key rotation.

When you use **Native Cloud Provider**, the cloud-provider services listed above share the following data handling guarantees:

* **No model training on your data** - your prompts and telemetry data are not used to train or improve the underlying models
* **No data retention** - inputs and outputs are not stored by the LLM provider beyond the request lifecycle
* **Data stays in your cloud account** - requests are processed within your configured cloud account. For BYOC AWS and BYOC GCP, inference runs in the same region as your cluster. For BYOC Azure, Anthropic models on Foundry are only available in select regions and inference may run in a different region from your AKS cluster, but always within your Azure subscription (see [Enable AI models for Agent Mode](/architecture/byoc/setup-byoc-with-azure/enable-ai-models-for-agent-mode))

For provider-specific security documentation:

* AWS Bedrock: [AWS Bedrock Security](https://docs.aws.amazon.com/bedrock/latest/userguide/security.html)
* Google Cloud Vertex AI: [Vertex AI Data Governance](https://cloud.google.com/vertex-ai/docs/general/data-governance)
* Microsoft Foundry: [Microsoft Foundry documentation](https://learn.microsoft.com/azure/ai-foundry/)

## What Data Reaches the LLM

When you ask the Agent a question:

1. Your prompt, current session context, and UI context (page, filters, time range) are sent to the Agent service running within your groundcover deployment
2. The Agent service queries your telemetry data (logs, traces, metrics) through internal APIs
3. Relevant query results are passed to the LLM to generate analysis
4. The response streams back to your browser

Only the data needed to answer your specific question is sent to the LLM. The Agent does not send your entire dataset.

## Tenant Isolation

All Agent operations are scoped to a single tenant. The Agent service enforces tenant boundaries at every layer:

* All telemetry queries are executed against your tenant's data store - the Agent has no path to query data belonging to another tenant
* Conversation history and session state are stored with a tenant identifier and cannot be accessed across tenant boundaries
* LLM requests are constructed using only data from the requesting tenant's context

This means cross-tenant data access is not possible through the Agent, regardless of prompt content.

## Session Management

Conversations are organized into sessions. Each session maintains its own isolated message history, which provides the Agent with context for follow-up questions and multi-step investigations.

* **Session TTL** - sessions expire after **30 days** of inactivity by default
* **Session cleanup** - expired sessions and their associated message history are automatically deleted when the TTL is reached
* **Session scope** - conversation history is scoped to the individual session; the Agent does not carry context between separate conversations unless you explicitly share or fork them

Starting a new conversation (Cmd/Ctrl+Shift+O) creates a new session with no prior context.

## Conversation Storage & Data Retention

Conversation history is stored in a database within your groundcover deployment. This data is:

* Scoped to the originating user and tenant - other users and tenants cannot access your conversation history
* Retained for the duration of the session TTL (30 days of inactivity), after which it is deleted
* Stored entirely within your own infrastructure in BYOC deployments

## Access Control

**Authentication** - The Agent is only accessible to authenticated groundcover users. There is no anonymous or public access.

**Permission scope** - All queries the Agent executes run with your user-level permissions, not elevated or admin privileges. The Agent respects your existing groundcover [RBAC configuration](/use-groundcover/role-based-access-control-rbac) and can only access data your account is authorized to see.

**Instructions and Skills** - Personal Instructions and personal Skills are scoped to the user who created them. Org Instructions apply to everyone in the selected backend, and Organizational Skills are readable and usable across the organization. They do not grant additional data access: the Agent still follows the current user's RBAC permissions for every query and action.

**AI & Agents settings**: Admins can enable or disable AI features and choose an LLM provider per backend from **Settings > AI & Agents > Management**. See [Configuring Settings](/use-groundcover/agent-mode/configuring-settings) for instructions.

## Questions

Contact your groundcover account team if you have questions about data handling or want to discuss your organization's specific security requirements.


# Cost Management

Control monthly Agent Mode spending with organization, default, and per-user cost limits

Admins can set monthly cost limits (in USD) at three levels to control Agent Mode spending across the organization. Budgets reset automatically at the start of each calendar month. Budgets are configured per groundcover backend — each backend has its own independent budget settings and usage tracking.

## How Budgets Work

### Budget Hierarchy

Cost management uses a three-tier budget hierarchy. Each tier is independent — configure any combination based on your needs.

| Level                       | Scope                                               | When Exceeded         |
| --------------------------- | --------------------------------------------------- | --------------------- |
| **Organization budget**     | All users in the organization combined              | All users are blocked |
| **Default per-user budget** | Any user without an individual override             | That user is blocked  |
| **Per-user override**       | A specific user (takes precedence over the default) | That user is blocked  |

Both the organization budget and the user-level budget are checked before each agent request. A user is blocked if **either** limit is reached.

{% hint style="info" %}
If no budgets are configured, usage is unlimited. You can set an organization budget, a default per-user budget, per-user overrides, or any combination of the three.
{% endhint %}

### Monthly Reset

Budgets reset at the start of each calendar month (UTC). Usage does not carry over from previous months.

### Fail-Open Behavior

{% hint style="info" %}
If the budget system is temporarily unavailable (for example, due to a database connectivity issue), agent requests are allowed to proceed. This ensures transient infrastructure issues do not block users.
{% endhint %}

### Cost Calculation

Costs are calculated per token based on the model used for each request. Input tokens, output tokens, and cache tokens each have different rates. Cache read tokens are significantly cheaper than regular input tokens, and cache creation tokens are priced slightly above regular input.

The total cost of each agent request is the sum of input cost (regular + cache) and output cost. Costs are tracked automatically — no configuration is needed.

## Configuring Budgets

Budget configuration requires an **admin** role. Navigate to **Settings > AI & Agents** and open the **Cost Control** section.

### Month Picker

Use the month picker at the top of the section to view current or historical usage. The usage graph displays the full calendar month, from the 1st to the last day. Budget editing is only available for the current month, historical months are read-only.

### Organization Budget

The organization budget sets an aggregate monthly cost cap across all users. When total usage reaches this limit, every user in the organization is blocked until the next month.

1. In the **Cost Control** section, find **Organization monthly limit**
2. Enter the dollar amount
3. Click **Update**

To remove the organization budget, click **Remove**. Without an organization budget, there is no aggregate cap.

### Default Per-User Budget

The default per-user budget is a fallback limit applied to every user who does not have an individual override. When removed, users without overrides become unlimited.

1. In the **Default per-user monthly limit** area, enter the dollar amount
2. Click **Update**

To remove, click **Remove**.

### Per-User Overrides

Per-user overrides let you set individual budgets that take precedence over the default. Use these to give specific users higher or lower limits.

**Setting an override for an existing user:**

1. Find the user in the **Per-User Usage & Limits** table (use the search filter to narrow results)
2. Click the edit icon next to their limit
3. Enter the dollar amount and confirm

**Adding a new user override:**

1. Click **Add User**
2. Select the user from the dropdown
3. Enter the dollar amount and confirm

**Removing an override:**

1. Click the remove icon next to the user's override
2. Confirm the removal

When an override is removed, the user falls back to the default per-user budget. If no default is set, the user becomes unlimited.

### Understanding the Usage Table

The per-user table shows each user's current month usage with visual indicators:

| Indicator           | Meaning                      |
| ------------------- | ---------------------------- |
| Green progress bar  | Under 75% of limit           |
| Yellow progress bar | 75%–99% of limit             |
| Red progress bar    | At or over 100% of limit     |
| **Blocked** tag     | User has reached their limit |

## Budget Status in the Chat

All users see their budget status below the chat input field. The status updates in real time as the agent processes requests.

* **Progress bar with percentage** — shows "X% of monthly budget used" when a budget is active
* **"Limit reached" tag** — shown in red when the budget is exceeded

{% hint style="info" %}
Budget status updates in real time during agent conversations. After each agent response, the indicator reflects the latest usage.
{% endhint %}

### What Happens When Blocked

When a budget is exceeded, the agent returns an error message instead of processing the request:

* **User budget exceeded** — the user sees a message with their current cost and limit
* **Organization budget exceeded** — all users see a message indicating the organization's monthly budget has been reached

The user cannot start new agent conversations until the next monthly reset or an admin increases or removes the budget.

## API Reference

{% hint style="info" %}
Admin endpoints require admin-level permissions. All endpoints are accessed through the groundcover API gateway.
{% endhint %}

### Budget Configuration (Admin)

| Method   | Endpoint                                   | Description                    |
| -------- | ------------------------------------------ | ------------------------------ |
| `GET`    | `/api/agent/token-budgets`                 | List all configured budgets    |
| `PUT`    | `/api/agent/token-budgets/tenant`          | Set organization budget        |
| `DELETE` | `/api/agent/token-budgets/tenant`          | Remove organization budget     |
| `PUT`    | `/api/agent/token-budgets/default`         | Set default per-user budget    |
| `DELETE` | `/api/agent/token-budgets/default`         | Remove default per-user budget |
| `PUT`    | `/api/agent/token-budgets/users/{user_id}` | Set per-user override          |
| `DELETE` | `/api/agent/token-budgets/users/{user_id}` | Remove per-user override       |

**Request body** for `PUT` endpoints:

```json
{
  "monthly_cost_limit": 25.00
}
```

The `monthly_cost_limit` value is in USD and must be greater than zero.

### Usage Reporting (Admin)

| Method | Endpoint                           | Description                  |
| ------ | ---------------------------------- | ---------------------------- |
| `GET`  | `/api/agent/token-usage`           | All users' usage for a month |
| `GET`  | `/api/agent/token-usage/tenant`    | Organization total usage     |
| `GET`  | `/api/agent/token-usage/{user_id}` | Specific user's usage        |
| `GET`  | `/api/agent/token-usage/history`   | Multi-month usage history    |

All usage endpoints accept an optional `month` query parameter in `YYYY-MM` format (defaults to the current month). The history endpoint accepts `months` (default: 12).

### Budget Status (Any User)

Any authenticated user can check their own budget status:

`GET /api/agent/token-budget/status`

**Response:**

```json
{
  "allowed": true,
  "cost": 12.50,
  "limit": 25.00,
  "remaining": 12.50
}
```

| Field       | Type           | Description                                    |
| ----------- | -------------- | ---------------------------------------------- |
| `allowed`   | boolean        | Whether the user can make agent requests       |
| `cost`      | number         | Current month's accumulated cost in USD        |
| `limit`     | number or null | Effective monthly limit (null means unlimited) |
| `remaining` | number or null | Remaining budget in USD (null means unlimited) |


# Connectors

Connectors let individual users link their external tool accounts to groundcover, enabling Agent Mode to take actions on their behalf

Connectors link groundcover to external tools such as GitHub, Notion, PagerDuty, Claude Managed Agents, Linear, Slack, and Cursor. Depending on the connector, they can power organization-wide notification workflows, personal [Agent Mode](/use-groundcover/agent-mode) tools, or both.

Notion, Pylon, GitHub, PagerDuty, Rootly, incident.io, and Atlassian Rovo use their providers' MCP services behind first-class connector cards. Select the card directly; you do not need to configure an MCP server URL. Claude Managed Agents connects directly to Anthropic instead of using MCP.

{% hint style="info" %}
The PagerDuty, Rootly, and incident.io connector cards give **Agent Mode** access to those services. To send monitor notifications, configure a [PagerDuty](/integrations/connected-apps/pagerduty), [Rootly](/integrations/connected-apps/incident.io-1), or [incident.io](/integrations/connected-apps/incident.io) Destination instead.
{% endhint %}

## Connectors vs Destinations

groundcover has two types of external integrations that serve different purposes:

| Aspect             | Connectors                                                                | Destinations                                                     |
| ------------------ | ------------------------------------------------------------------------- | ---------------------------------------------------------------- |
| **Scope**          | Personal connections, sometimes backed by an organization-managed app     | Organization, admins configure shared integrations               |
| **Purpose**        | Enable Agent Mode to take actions in external tools on behalf of the user | Send notifications and alerts to external services               |
| **Authentication** | User authorizes their account with OAuth or provides personal credentials | Admin configures org-wide credentials or authorizes a shared app |
| **Used by**        | groundcover Agent Mode during conversations                               | Notification Routes, Monitors, and Workflows                     |
| **Setup location** | **Integrations → Connectors**                                             | **Integrations → Destinations**                                  |

In short: **Destinations** push alerts *out* of groundcover at the org level, while **Connectors** let Agent Mode reach *into* external tools at the user level.

Some integrations support both roles. Setting up [Linear](/use-groundcover/connectors/linear) or [Slack](/use-groundcover/connectors/slack) under **Org Connectors** creates the organization-level connection used by monitor workflows, while **My Connections** lets each user authorize the same tool for Agent Mode.

## How It Works

Connectors use a two-tier model:

1. **Organization setup**: A workspace admin enables or configures a connector for the organization. Depending on the connector, the admin may authorize with OAuth, supply an API token, or register an OAuth application and enter its Client ID and Client Secret.
2. **User connection**: Individual users authorize their account with OAuth or provide a personal token. Agent Mode uses the connection to act on the user's behalf.

{% hint style="info" %}
Admin permissions are required to enable or disable a connector for the organization.\
Any workspace member with [Agent Mode access](/use-groundcover/agent-mode/privacy-and-security) can connect their own account once the connector is enabled.
{% endhint %}

<div data-with-frame="true"><figure><img src="/files/dmcbwQUz43o38opidlPv" alt="Connector list in groundcover"><figcaption><p>Connectors are available as first-class cards under Integrations → Connectors.</p></figcaption></figure></div>

## Multiple Instances per Connector

Every connector supports **multiple instances**, so you can connect the same connector type to more than one external account or environment. Each instance is identified by a unique **Connector Name** that you choose at setup time.

Common use cases:

* **Slack** — connect more than one Slack workspace by adding a separate Slack App per workspace (e.g., `groundcover Prod`, `groundcover Staging`). When [selecting Slack channels](/use-groundcover/connectors/slack#selecting-slack-channels-in-notification-routes), each Slack App appears separately in the destination picker so you can route alerts to the right workspace.
* **Linear** — connect multiple Linear workspaces by creating a separate OAuth application and organization connector for each workspace. Each connector appears separately when creating issues or configuring Notification Routes.
* **Cursor** — connect multiple Cursor API keys (e.g., one per repository owner or environment). When you invoke Cursor through Agent Mode, you'll be asked which credential to use if more than one is connected.
* **Remote MCP Servers** — allowlist as many MCP servers as you need; each one is added independently and exposes its own set of tools.
* **Other connectors** — the same model applies to all current and future connector types.

Constraints:

* The **Connector Name** must be unique across all Destinations in the workspace — including other connector types and Destinations configured under **Integrations → Destinations**. For example, you cannot reuse `prod-alerts` as both a Slack App connector name and a PagerDuty Destination name.
* For connectors bound to an external workspace (such as Slack), each instance is bound to one workspace; you cannot reuse the same workspace across two instances.
* For personal credentials (such as Cursor), there is no upper limit on how many you can add to your account.

## Available Connectors

| Connector                                                                      | Description                                                                         |
| ------------------------------------------------------------------------------ | ----------------------------------------------------------------------------------- |
| [**Slack**](/use-groundcover/connectors/slack)                                 | Integrate with Slack for monitor notifications and agent interactions via @mentions |
| [**Linear**](/use-groundcover/connectors/linear)                               | Create and manage Linear issues from monitors, Notification Routes, and Agent Mode  |
| [**Cursor Cloud Agents**](/use-groundcover/connectors/cursor)                  | Run Cursor's cloud coding agent on your repositories directly from Agent Mode       |
| [**Claude Managed Agents**](/use-groundcover/connectors/claude-managed-agents) | Run coding tasks using a Claude agent and environment configured in Anthropic       |
| [**Notion**](/use-groundcover/connectors/notion)                               | Search workspace knowledge and publish investigation findings                       |
| [**Pylon**](/use-groundcover/connectors/pylon)                                 | Investigate customer issues and take support actions                                |
| [**GitHub**](/use-groundcover/connectors/github)                               | Investigate code and manage issues, pull requests, and development work             |
| [**PagerDuty**](/use-groundcover/connectors/pagerduty)                         | Investigate and manage incidents, services, schedules, and on-call work             |
| [**Rootly**](/use-groundcover/connectors/rootly)                               | Investigate incidents and coordinate response and follow-up work                    |
| [**incident.io**](/use-groundcover/connectors/incident-io)                     | Analyze operations and manage incidents, escalations, and follow-ups                |
| [**Atlassian Rovo**](/use-groundcover/connectors/atlassian-rovo)               | Read and update Jira issues and Confluence knowledge                                |
| [**Remote MCP Servers**](/use-groundcover/connectors/mcp)                      | Allowlist custom MCP servers with token or OAuth authentication                     |

## Managing Connectors

The Connectors page has two tabs:

* **My Connections** — where any user can link their personal account for an enabled connector
* **Org Connectors** — where admins enable connector types and configure org-level credentials (e.g., the Slack App)

### For Admins

Go to **Integrations → Connectors** and open the **Org Connectors** tab to enable or disable connector types for your organization, and to configure org-level credentials where required (e.g., the Slack App). Disabling a connector does not delete existing user credentials, but prevents them from being used until re-enabled.

### For Users

Once a connector is enabled by an admin:

1. Go to **Integrations → Connectors → My Connections**
2. Select the connector you want to set up
3. Authorize your account or provide your credentials, then configure any defaults
4. When you ask Agent Mode to use a specific tool (e.g., "use Cursor to fix the failing test"), it will use your connected credentials to perform the action

You can update your credentials, change configuration, or disconnect at any time from the same page.

## Next Steps

* [Slack](/use-groundcover/connectors/slack): Set up the Slack App connector and link your personal account
* [@groundcover Agent in Slack](/use-groundcover/connectors/slack/slack-agent): Use the groundcover Agent directly from Slack via @mentions
* [Linear](/use-groundcover/connectors/linear): Connect a Linear workspace for monitor issues, Notification Routes, and Agent Mode
* [Cursor Cloud Agents](/use-groundcover/connectors/cursor): Set up and use the Cursor connector with Agent Mode
* [Claude Managed Agents](/use-groundcover/connectors/claude-managed-agents): Connect a Claude agent, environment, vaults, and memory store
* [Notion](/use-groundcover/connectors/notion), [Pylon](/use-groundcover/connectors/pylon), [GitHub](/use-groundcover/connectors/github), [PagerDuty](/use-groundcover/connectors/pagerduty), [Rootly](/use-groundcover/connectors/rootly), [incident.io](/use-groundcover/connectors/incident-io), and [Atlassian Rovo](/use-groundcover/connectors/atlassian-rovo): Set up each first-class connector and explore Agent Mode examples
* [Remote MCP Servers](/use-groundcover/connectors/mcp): Allowlist custom MCP servers with token or OAuth authentication and manage tool restrictions


# Slack

{% hint style="info" %}
This capability is only available to BYOC deployments. Check out our [pricing page](https://www.groundcover.com/pricing) for more information about subscription plans and the available deployment modes.
{% endhint %}

The Slack connector integrates groundcover with your Slack workspace. Once configured, groundcover can send monitor notifications to specific Slack channels and enable users to interact with groundcover's agent directly from Slack using [@groundcover mentions](/use-groundcover/connectors/slack/slack-agent).

{% hint style="info" %}
Admin permissions are required to create and manage the Slack connector.\
Any workspace member can connect their personal Slack account once the connector is set up.
{% endhint %}

## Setting Up the Slack App

Setting up the Slack connector is a one-time process that involves creating a Slack App, providing its credentials, and authorizing it in your workspace.

### Step 1: Create the Slack App from Manifest

1. In groundcover, go to **Integrations → Connectors** and open the **Org Connectors** tab.<br>

   <div data-with-frame="true"><figure><img src="/files/zTyzwESP4pcT0WRR2fGX" alt=""><figcaption></figcaption></figure></div>
2. Select **Slack**. The setup form opens directly for the first Slack App
   1. If you already have one or more Slack Apps connected, click **Add Slack Connector** to add another one (see [Connecting Multiple Slack Workspaces](#connecting-multiple-slack-workspaces)).
3. Click **Create App** — this opens Slack's app creation page with a pre-filled manifest containing all the required permissions and configuration.<br>

   <div data-with-frame="true"><figure><img src="/files/PTPMhGyvFxHHV8mftvGO" alt=""><figcaption></figcaption></figure></div>
4. In Slack, **choose the workspace** to install the app in, click **Next**, review the manifest, and click **Create**.

{% hint style="info" %}
You can also reach the same setup form from **Integrations → Destinations → Slack App**, which is convenient when you're configuring Slack alongside other Destinations like PagerDuty.
{% endhint %}

{% hint style="info" %}
You can click **View Manifest** to inspect the full manifest JSON before creating the app. The manifest includes all required OAuth scopes, event subscriptions, and Socket Mode configuration. Do not modify it — groundcover relies on these exact permissions to function.

The manifest requests a broad set of permissions on purpose: some are needed for current features (notifications, agent @mentions, channel-aware AI replies) and some are reserved for features we plan to release. Requesting them upfront means you won't have to reinstall the app — or get users to re-approve it — when those features ship.
{% endhint %}

For more details on creating apps from manifests, see [Slack's manifest guide](https://docs.slack.dev/app-manifests/configuring-apps-with-app-manifests/).

### Step 2: Upload the App Icon

1. Click **Download Icon** in the groundcover setup wizard to save the groundcover app icon (PNG).

   <div data-with-frame="true"><figure><img src="/files/id6nVsmjO5osL28KltWE" alt=""><figcaption></figcaption></figure></div>
2. In your Slack App settings, go to **Settings → Basic Information → Display Information → App Icon & Preview**.

   <div data-with-frame="true"><figure><img src="/files/Twb1B9cJDcbVcadsdhJA" alt=""><figcaption></figcaption></figure></div>
3. Click **Add App Icon** and upload the downloaded image.

   <div data-with-frame="true"><figure><img src="/files/U1qqoqZWV2bCdPcPyZmv" alt=""><figcaption></figcaption></figure></div>

### Step 3: Fill in App Credentials

You need three values from your Slack App. All of them are found in your Slack App's settings page.

| Field               | Where to find it in Slack                                                |
| ------------------- | ------------------------------------------------------------------------ |
| **Client ID**       | Settings → Basic Information → App Credentials → Client ID               |
| **Client Secret**   | Settings → Basic Information → App Credentials → Client Secret           |
| **App-Level Token** | Settings → Basic Information → App-Level Tokens (see instructions below) |

<div data-with-frame="true"><figure><img src="/files/Io0etfWfU91WCPAiaT9Z" alt=""><figcaption></figcaption></figure></div>

#### Generating the App-Level Token

The App-Level Token is not created automatically — you need to generate it:

1. In your Slack App settings, go to **Basic Information → App-Level Tokens**.

   <div data-with-frame="true"><figure><img src="/files/wXKNccLwfuTNxAszX8vD" alt=""><figcaption></figcaption></figure></div>
2. Click **Generate Token and Scopes**.
3. Enter a name (e.g., `groundcover`).
4. Click **Add Scope** and select **`connections:write`**.

   <div data-with-frame="true"><figure><img src="/files/Vp0Om6VGdRnCLr4MZfhM" alt=""><figcaption></figcaption></figure></div>
5. Click **Generate**.
6. Copy the token — it starts with `xapp-`.

{% hint style="warning" %}
The App-Level Token must start with `xapp-`. groundcover validates it against Slack's API during setup. If the token is invalid or missing the `connections:write` scope, the setup will fail with an error.
{% endhint %}

Back in the groundcover setup wizard, fill in:

* **Connector Name** — A unique label to identify this connector (e.g., `groundcover Prod`).
* **Client ID** — From your Slack App credentials.
* **Client Secret** — From your Slack App credentials.
* **App-Level Token** — The `xapp-` token you just generated.

<div data-with-frame="true"><figure><img src="/files/Nl6k1nMnNdDMhZdhNBHt" alt=""><figcaption></figcaption></figure></div>

Click **Connect** to save and proceed to authorization.

### Step 4: Authorize the App (OAuth)

After saving the credentials, you are redirected to Slack to authorize the app in your workspace. This grants groundcover's bot the permissions defined in the manifest (sending messages, reading channels, etc.).

<div data-with-frame="true"><figure><img src="/files/ToqzGTrxnpFQ47cxCFRi" alt=""><figcaption></figcaption></figure></div>

Once authorization completes:

* The connector status changes to **active**.

<div data-with-frame="true"><figure><img src="/files/Hjl9tpz2NkVfmZk6NWW6" alt=""><figcaption></figcaption></figure></div>

* A Destination of type `slack-app` is automatically created and bound to your Slack workspace. It appears alongside other Destinations (like Slack Webhook and PagerDuty) in the [Notification Routes](/use-groundcover/monitors/notification-routes) destination picker and in the monitor wizard. You do not configure this Destination separately — it is managed by the connector.

<div data-with-frame="true"><figure><img src="/files/sD59p62DmUxxIWLQT3XG" alt=""><figcaption></figcaption></figure></div>

## Connecting Multiple Slack Workspaces

You can connect more than one Slack workspace by repeating the [Setting Up the Slack App](#setting-up-the-slack-app) flow for each workspace:

1. Create a separate Slack App in each workspace (each workspace will need its own Client ID, Client Secret, and App-Level Token).
2. Add each one as a separate connector in **Integrations → Connectors → Org Connectors → Slack** with a distinct **Connector Name** (e.g., `groundcover Prod`, `groundcover Staging`, `groundcover Customer-Success`).
3. After OAuth, each connector is bound to its workspace and creates its own `slack-app` Destination.

When you [select a Slack channel](#selecting-slack-channels-in-notification-routes) in a notification route or monitor, each Slack App appears separately in the picker labeled by its connector name, so you can route alerts to the correct workspace.

{% hint style="info" %}
A single Slack connector is bound to one workspace. You cannot point two connectors at the same Slack workspace.
{% endhint %}

## Connecting Your Personal Slack Account

After an admin sets up the Slack connector, individual users can link their personal Slack accounts from **Integrations → Connectors → My Connections**. This enables features like interacting with groundcover's agent via [@groundcover mentions](/use-groundcover/connectors/slack/slack-agent) in Slack.

See [Connecting Your Personal Slack Account](/use-groundcover/connectors/slack/slack-agent#connecting-your-personal-slack-account) for the full walkthrough.

## Selecting Slack Channels in Notification Routes

After the Slack connector is active, you can route monitor notifications to specific Slack channels.

1. Go to **Monitors → Notification Routes**.
2. Click **Create Notification Route** (or edit an existing one).
3. In **Step 3 (Rules)**, click the **destination dropdown**.
4. Select your Slack App — it is marked with an **App** tag. A second panel appears listing the available channels in your workspace.<br>

   <div data-with-frame="true"><figure><img src="/files/Mh3lqP1h72C048eKwer4" alt=""><figcaption></figcaption></figure></div>
5. Search for and select the target channel.<br>

   <div data-with-frame="true"><figure><img src="/files/HQlNS13wk7cFndjFjO6t" alt="" width="174"><figcaption></figcaption></figure></div>

<div data-with-frame="true"><figure><img src="/files/qijjbUMI3LULNr4dojBa" alt="" width="185"><figcaption></figcaption></figure></div>

1. Click **Add Rule** to add more rules with different channels or status filters.

Private channels are marked with a lock icon. The bot can post to any **public** channel in the workspace, even if it has not been added to it. For **private** channels, the bot must first be invited to the channel — only private channels the bot has been added to appear in the picker.

{% hint style="info" %}
You can create multiple rules to route to different channels based on issue status. For example, send **Firing** alerts to `#critical-alerts` and **Resolved** notifications to `#alerts-resolved`.
{% endhint %}

For more details on notification routes, see [Notification Routes](/use-groundcover/monitors/notification-routes).

## Selecting Slack Channels in the Monitor Wizard

When creating or editing a monitor, you can send notifications directly to Slack channels without using notification routes:

1. Open the monitor wizard (create or edit a monitor).
2. In the **Notifications** section, select **Directly to destinations**.
3. Click **Set rule**.
4. Choose an issue status filter (Firing, Resolved, or both).
5. In the **Send to** dropdown, select your Slack App.<br>

   <div data-with-frame="true"><figure><img src="/files/k9FD4TpLt1hH5edjpSZU" alt=""><figcaption></figcaption></figure></div>
6. Pick a target channel from the drilldown panel.
7. Click **+** to add more destinations (additional Slack channels or other Destinations).

{% hint style="info" %}
Choosing **Directly to destinations** means this monitor's issues will be skipped by notification routes. If you want notifications to flow through your configured routes instead, select **Based on matching notification routes**.
{% endhint %}

The selected channels are persisted in the monitor under `notificationSettings.connectedAppParams.<slack-app-id>.channels` — see [Monitor YAML structure](/use-groundcover/monitors/monitor-yaml-structure#notificationsettings) for the full schema.

## Updating Credentials

If you need to rotate the Client Secret or App-Level Token after the initial setup:

1. Go to **Integrations → Connectors** and open the **Org Connectors** tab.
2. Select **Slack** and open your Slack connector.
3. Enter the new Client Secret and/or App-Level Token. Leave a field blank to keep its current value.
4. Click **Save Changes**.

{% hint style="warning" %}
If the App-Level Token becomes invalid or is revoked in Slack, the connector will stop receiving events from Slack and move to an **invalid token** state. Update the token to restore functionality.
{% endhint %}

## Deleting the Connector

1. Go to **Integrations → Connectors** and open the **Org Connectors** tab.
2. Select **Slack** and click the delete action on the Slack connector.
3. Confirm deletion.

{% hint style="warning" %}
A Slack connector that is used by notification routes or monitors cannot be deleted. Remove it from all routes and monitors first.
{% endhint %}

## Troubleshooting

### The App-Level Token is rejected during setup

Verify that:

* The token starts with `xapp-`.
* The token was generated with the `connections:write` scope.
* The token has not been revoked in Slack.

You can regenerate it in your Slack App under **Basic Information → App-Level Tokens**.

### OAuth authorization fails

* Ensure you are signing into the correct Slack workspace.
* Verify the Client ID and Client Secret match your Slack App's credentials (check for extra spaces when pasting).

### User sign-in fails with "workspace does not match"

Each Slack connector is bound to a specific workspace. If a user tries to sign in with a different workspace, the connection will fail. Make sure users sign into the same workspace the admin authorized during initial setup.

### Notifications are not being sent to Slack

1. Verify the connector status is **active** in **Integrations → Connectors → Org Connectors**.
2. Check that the target channel is selected in your notification route or monitor.
3. Verify the notification route's scope query matches the monitor's labels.
4. Use the **Test** button in the monitor to send a test notification.
5. Filter traces by `workload:dispatch-center` in groundcover to inspect delivery attempts.

### A private channel does not appear in the channel picker

The bot can post to all public channels without being added, but it must be invited to a private channel before that channel appears in the picker. Open the channel in Slack and run `/invite @groundcover` (or use the channel settings → **Integrations** → **Add apps**), then reopen the picker.

### The channel picker is empty

* If channels were recently created, try closing and reopening the dropdown to refresh the list.
* Verify the Slack App connector status is **active** in **Integrations → Connectors → Org Connectors**.


# @groundcover Agent in Slack

Interact with groundcover's Agent directly from Slack by mentioning @groundcover in any channel

{% hint style="info" %}
This capability is only available to BYOC deployments. Check out our [pricing page](https://www.groundcover.com/pricing) for more information about subscription plans and the available deployment modes.
{% endhint %}

Once the [Slack connector](/use-groundcover/connectors/slack) is set up and you have [connected your personal Slack account](#connecting-your-personal-slack-account), you can interact with groundcover's Agent from any Slack channel where the groundcover app is present by mentioning **@groundcover**.

This gives your team a shared interface for investigating issues, querying observability data, and triaging alerts — without leaving Slack.

## How It Works

1. In any Slack channel where the groundcover app is present, type `@groundcover` followed by your question or request.
2. The Agent processes your message and replies in a thread under your message.
3. You can continue the conversation in the thread — the Agent maintains context within that thread, just like a regular [Agent Mode](/use-groundcover/agent-mode) session.

The Agent in Slack has the same capabilities as Agent Mode in the groundcover UI: investigating issues, querying logs/traces/metrics, creating monitors and dashboards, and more. See [Agent Mode](/use-groundcover/agent-mode) for the full list of capabilities.

{% hint style="info" %}
The Agent responds in a thread to keep the channel clean. All follow-up messages in that thread are part of the same conversation context.
{% endhint %}

## Identity, Permissions & Budget

Every `@groundcover` mention is tied to the **Slack user who sent the message**. The Agent runs as that user's agent session, which means:

| Aspect       | Behavior                                                                                                                                                                                                                                                                      |
| ------------ | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Identity** | The Agent session is associated with the groundcover user whose personal Slack account matches the Slack user who sent the message                                                                                                                                            |
| **RBAC**     | The Agent can only access data the mentioning user is authorized to see, based on their [RBAC configuration](/use-groundcover/role-based-access-control-rbac). If the user's role restricts them to specific namespaces or clusters, the Agent respects those same boundaries |
| **Budget**   | The request counts against the mentioning user's [Agent Mode budget](/use-groundcover/agent-mode/cost-management). If the user has reached their monthly limit, the Agent will not process the request                                                                        |

{% hint style="warning" %}
**Visibility:** While the Agent enforces the mentioning user's permissions for data access, the Agent's response is visible to **everyone in the Slack channel**. Be mindful when querying sensitive data in public or broadly-shared channels — the response will be visible to all channel members, even those who may not have the same level of access in groundcover.
{% endhint %}

### What Happens If You're Not Connected

If a Slack user mentions `@groundcover` but has not [connected their personal Slack account](#connecting-your-personal-slack-account), the Agent will respond with a message asking them to connect their account first. The Agent cannot process requests from unconnected users because it has no way to determine which groundcover user they are.

## Connecting Your Personal Slack Account

Before you can use `@groundcover` mentions, you need to link your personal Slack account to your groundcover user. This is a one-time setup.

1. In groundcover, go to **Integrations → Connectors → My Connections**.
2. Find **Slack** and locate the Slack App for the workspace you want to connect to.
3. Click **Connect**.
4. You are redirected to Slack to sign in and approve access.
5. Once approved, you are redirected back to groundcover. Your Slack identity is now linked to your groundcover account.

{% hint style="info" %}
You must sign into the same Slack workspace that the admin authorized during the [Slack connector setup](/use-groundcover/connectors/slack#step-4-authorize-the-app-oauth). Signing into a different workspace will fail.
{% endhint %}

To disconnect, click **Disconnect** on the Slack App entry in **My Connections**. You can reconnect at any time.

### Multiple Workspaces

If your organization has [multiple Slack workspaces connected](/use-groundcover/connectors/slack#connecting-multiple-slack-workspaces), you can connect your personal account to each one independently. The Agent will respond to your mentions in any workspace where your account is linked.

## Tips

* **Be specific** — the Agent works best with clear, specific questions. Instead of "what's wrong?", try "why is the checkout service throwing 500 errors in the production cluster?"
* **Use threads for follow-ups** — continue the conversation in the same thread to maintain context. Starting a new top-level mention creates a new, independent session.
* **Channel selection** — use a dedicated channel (e.g., `#groundcover-agent`) for agent interactions to keep other channels focused. The Agent works in any channel it has been added to.
* **Private channels** — the Agent can respond in private channels, but it must first be invited to the channel. Run `/invite @groundcover` to add it.


# Linear

Connect Linear to create and manage issues from monitor incidents, notification routes, and groundcover Agent Mode

The Linear connector links a Linear workspace to groundcover. After it is configured, you can:

* Create a Linear issue directly from a monitor issue
* Create and update Linear issues automatically through [Notification Routes](/use-groundcover/monitors/notification-routes)
* Connect your personal Linear account so [Agent Mode](/use-groundcover/agent-mode) can use Linear MCP tools on your behalf

{% hint style="info" %}
A groundcover workspace admin must set up the organization connector. Afterward, individual users can connect their own Linear accounts for Agent Mode.
{% endhint %}

## How the Linear connections work

Linear uses two related connections:

| Connection                 | Who configures it             | What it enables                                                                                      |
| -------------------------- | ----------------------------- | ---------------------------------------------------------------------------------------------------- |
| **Organization connector** | A groundcover workspace admin | Manual issue creation from monitor issues and automated issue management through Notification Routes |
| **Personal connection**    | Each groundcover user         | Linear MCP tools in Agent Mode, using that user's Linear permissions                                 |

The organization connector is required before users can create personal connections.

## Prerequisites

Before you begin, make sure you have:

* Admin permissions in groundcover
* Permission to create an OAuth application in the Linear workspace you want to connect
* Access to that Linear workspace during the OAuth approval flow

## Set up the organization connector

Setting up the organization connector is a one-time process for each Linear workspace.

### Step 1: Open the Linear connector setup

1. In groundcover, go to **Integrations → Connectors**.
2. Open the **Org Connectors** tab.
3. In the Linear section, enter a unique **Connector Name**. This name identifies the Linear workspace throughout groundcover.
4. Click **Create App**.

<div data-with-frame="true"><figure><img src="/files/DM6DUDbpsEmfvKU7V4i6" alt="Linear organization connector setup in groundcover"><figcaption><p>Create a Linear OAuth app and enter its Client ID in the organization connector setup.</p></figcaption></figure></div>

{% hint style="info" %}
Click **View Manifest** if you want to inspect the application manifest before continuing.
{% endhint %}

### Step 2: Create the OAuth application in Linear

Linear opens in a new tab with the application fields populated from the groundcover manifest. Review the target workspace, then click **Create**. No changes to the pre-filled settings are required.

<div data-with-frame="true"><figure><img src="/files/dZE3zbbP3zeg9pXea6Uj" alt="Pre-filled groundcover OAuth application in Linear"><figcaption><p>Linear populates the application name, description, and redirect URI from the manifest.</p></figcaption></figure></div>

The application remains private to your Linear workspace.

### Step 3: Copy the Linear Client ID

After Linear creates the application, open its settings and scroll to **OAuth credentials**. Copy the **Client ID**.

<div data-with-frame="true"><figure><img src="/files/iE3XYfBVxcYtEdvRMpRx" alt="Client ID in the Linear OAuth credentials section"><figcaption><p>Copy the Client ID from the application's OAuth credentials.</p></figcaption></figure></div>

{% hint style="warning" %}
Only copy the **Client ID** into groundcover. Do not copy or share the Client secret.
{% endhint %}

### Step 4: Connect and authorize the application

1. Return to the groundcover connector setup tab.
2. Paste the value into **Linear OAuth Client ID**.
3. Click **Connect**.
4. On the Linear OAuth approval page, select the correct workspace and approve the application.

After authorization returns you to groundcover, the Linear workspace appears under **Org Connectors** and is ready to use.

<div data-with-frame="true"><figure><img src="/files/gg6vWUKHXwjrJTqZfExG" alt="Connected Linear workspace in groundcover organization connectors"><figcaption><p>The connected Linear workspace is available to the organization.</p></figcaption></figure></div>

To connect another Linear workspace, click **Add Another** and repeat the setup with a unique Connector Name.

## Connect your personal Linear account for Agent Mode

After an admin configures the organization connector, each user can enable Linear MCP tools for Agent Mode:

1. Go to **Integrations → Connectors**.
2. Open **My Connections**.
3. Find the Linear workspace you want to use and click **Connect**.
4. Approve the connection in Linear.

<div data-with-frame="true"><figure><img src="/files/EQIDztfGnzpPrEQuAAO0" alt="Linear connector available under My Connections"><figcaption><p>Connect your Linear account to make its MCP tools available to Agent Mode.</p></figcaption></figure></div>

Agent Mode can now use Linear within the permissions of your connected Linear account. For example, you can ask it to create an issue, find issues assigned to you, or update an existing issue.

## Create a Linear issue from a monitor issue

To create and link a Linear issue manually:

1. Go to **Monitors → Monitor Issues** and open an issue.
2. Click the Linear icon in the issue actions. Its tooltip reads **Create Linear issue**.

<div data-with-frame="true"><figure><img src="/files/DiqIrmK3TfUAdWM1ZGHM" alt="Create Linear issue action on a groundcover monitor issue"><figcaption><p>Create a Linear issue directly from the monitor issue details.</p></figcaption></figure></div>

3. Review the pre-filled title and description.
4. Select a Linear team and, optionally, a project, labels, assignee, and status.
5. Create the issue.

The new Linear issue includes monitor context and links back to groundcover. After an issue is linked, groundcover shows actions for opening it or adding an update instead of creating a duplicate.

## Automate Linear issues with a Notification Route

Use a Notification Route to keep a Linear issue synchronized with a monitor issue throughout its lifecycle:

1. Go to **Monitors → Notification Routes**.
2. Create a route or edit an existing one.
3. Define the monitor scope for the route.
4. In **Step 3: Rules**, select **Firing** and **Resolved**.
5. Under **Send to**, select your Linear connector.
6. Select the required Linear **Team**.
7. **Auto-resolve Linear issue** is enabled by default. While it is enabled, select a required **Resolved status**. If you do not want groundcover to transition the Linear issue when the monitor resolves, disable auto-resolve instead.
8. Optionally, select a project, labels, or assignee.
9. Save the route.

<div data-with-frame="true"><figure><img src="/files/TWcOEUPiav17BttaAGhW" alt="Linear connector selected as a notification route destination"><figcaption><p>Send firing and resolved monitor events to the connected Linear workspace.</p></figcaption></figure></div>

When a matching monitor issue fires, groundcover creates one linked Linear issue. Re-notifications add updates to that issue instead of creating duplicates. When the monitor issue resolves, groundcover adds a resolution comment and, when auto-resolve is enabled, moves the Linear issue to the configured resolved status.

For more information about route scopes and rules, see [Notification Routes](/use-groundcover/monitors/notification-routes).

## Troubleshooting

### Create App does not open Linear

Allow pop-ups for the groundcover application, then click **Create App** again.

### You cannot create the OAuth application

Ask a Linear workspace administrator to create the application or grant you permission to manage OAuth applications.

### The connector does not appear under My Connections

Confirm that a workspace admin completed the organization connector setup and authorized the intended Linear workspace.

### Linear actions are unavailable

Confirm that the organization connector is still listed under **Org Connectors**. For Agent Mode actions, also confirm that your personal connection is active under **My Connections**.


# Cursor Cloud Agents

Connect your Cursor account to let groundcover Agent Mode run code changes on your repositories using Cursor's cloud coding agent

The Cursor connector lets [Agent Mode](/use-groundcover/agent-mode) launch [Cursor Background Agents](https://docs.cursor.com/background-agent) on your repositories. When you ask Agent Mode to make a code change, it runs Cursor's cloud coding agent remotely using your Cursor API key, creating branches, writing code, and opening pull requests on your behalf.

{% hint style="info" %}
The Cursor connector must be enabled by a workspace admin before users can connect their accounts. See [Connectors](/use-groundcover/connectors) for details.
{% endhint %}

## Prerequisites

* A [Cursor](https://www.cursor.com/) account with API access
* A Cursor API key (generated from your Cursor account settings)
* The Cursor connector enabled for your organization (by an admin)

## Setup

### Admin: Enable the Cursor Connector

Before users can connect, a workspace admin must enable the Cursor connector for the organization:

1. Go to **Integrations → Connectors** and open the **Org Connectors** tab
2. Select **Cursor Cloud Agents**
3. Toggle the connector to **Enabled**

No org-level credentials are required — enabling simply makes the connector available to users.

### Generate a Cursor API Key

1. Open Cursor and go to **Settings → API Keys** (or visit your Cursor account settings page)
2. Create a new API key
3. Copy the key. You'll need it in the next step

### Connect Your Account in groundcover

1. Go to **Integrations → Connectors → My Connections**
2. Select **Cursor Cloud Agents**
3. Paste your API key
4. Give the credential a name (e.g., `my-cursor-key`)
5. Click **Connect**

Once connected, you'll see the configuration page where you can set defaults.

## Configuration

After connecting, you can configure the following settings from **Integrations → Connectors → My Connections → Cursor Cloud Agents**:

| Setting                  | Description                                                                           | Default        |
| ------------------------ | ------------------------------------------------------------------------------------- | -------------- |
| **Default Repository**   | The repository the Agent uses when you don't specify one                              | None           |
| **Default Model**        | The AI model Cursor uses for code generation                                          | None           |
| **Create Pull Requests** | Whether Cursor automatically creates PRs when a job completes, or waits for approval  | Needs approval |
| **Create Branch**        | Whether Cursor automatically creates branches for code changes, or waits for approval | Auto           |
| **API Key**              | Your Cursor API key. Can be updated at any time                                       | -              |

{% hint style="info" %}
You can have multiple Cursor credentials connected to the same account. If you have more than one, the Agent will ask you which one to use.
{% endhint %}

## Using Cursor Through the Agent

### Starting a Code Change

Ask the Agent to make a code change in natural language:

> "Use Cursor to add a CPU limit of 500m to the checkout deployment in my-org/infra-manifests"

> "Use Cursor to scale up the payments deployment to 5 replicas in my-org/k8s-config"

The Agent will:

1. Select your Cursor connector (or ask you to choose if you have multiple)
2. Identify the target repository
3. Prepare the code change and present an **Approve** button

### Approval Flow

The Agent always asks for your permission before running Cursor. Before any code is changed, the Agent presents its plan and an **Approve Cursor Job** button in the conversation. Only after you click approve does the Agent send the instructions to Cursor for execution.

This ensures you stay in control. No code is modified without your explicit approval.

### Monitoring Progress

Once approved, the Agent displays a live status card showing:

* **Status**: The current state of the job (queued, running, completed, error, expired, cancelled)
* **Repository**: Which repo and branch Cursor is working on
* **Branch**: The branch created for the changes (once available)
* **Pull Request**: A link to the PR (once created)

You can check on a running job at any time:

> "What's the status of the Cursor job?"

### Following Up

While a Cursor job is running (or after it completes), you can send follow-up instructions:

> "Also update the HPA to match the new replica count"

> "The PR looks good but add a comment explaining why we raised the CPU limit"

Follow-ups also require approval before being sent.

### Cancelling a Job

To stop a running Cursor job:

> "Cancel the Cursor job"

## How It Works

When you use the Cursor connector through the Agent:

1. The Agent resolves your Cursor credentials from the connector you configured
2. It sends instructions to the Cursor Cloud API using your API key
3. Cursor's cloud coding agent runs remotely, cloning the repository, making changes, and pushing code
4. The Agent polls for status updates and displays progress in the conversation
5. When complete, Cursor creates a branch and (optionally) a pull request

All actions are performed using your Cursor API key and your permissions. groundcover acts as a proxy and does not store or access your source code.

## Troubleshooting

| Issue                            | Resolution                                                                                      |
| -------------------------------- | ----------------------------------------------------------------------------------------------- |
| "No active Cursor connectors"    | Ensure you've connected your Cursor account in **Integrations → Connectors → My Connections**   |
| "Connector not enabled"          | Ask a workspace admin to enable the Cursor connector for your organization                      |
| Agent can't find your repository | Verify the repository is accessible with your Cursor API key and try listing repositories first |
| Job fails with rate limit error  | Cursor enforces rate limits on API usage. Wait a few minutes and try again                      |
| Job status shows "expired"       | Cursor jobs have a time limit. Try again with a more focused prompt                             |


# Claude Managed Agents

Connect Claude Managed Agents to let groundcover Agent Mode run coding tasks with your configured Claude agent and environment

The Claude Managed Agents connector lets [Agent Mode](/use-groundcover/agent-mode) delegate coding tasks to a Claude agent configured in the Anthropic Console. Unlike connectors such as [Notion](/use-groundcover/connectors/notion) and [GitHub](/use-groundcover/connectors/github), groundcover communicates with this connector through the Claude Managed Agents API rather than an MCP server. The configured Claude agent can still use its own MCP servers and tools.

## Prerequisites

* Access to [Claude Managed Agents](https://platform.claude.com/docs/en/managed-agents/overview)
* A Claude API key from the [Anthropic Console](https://platform.claude.com/)
* An [Agent ID](https://platform.claude.com/docs/en/managed-agents/agent-setup)
* An [Environment ID](https://platform.claude.com/docs/en/managed-agents/environments)
* Optional [vault IDs](https://platform.claude.com/docs/en/managed-agents/vaults) and a [memory store](https://platform.claude.com/docs/en/managed-agents/memory)

## Admin: Enable the Connector

1. Go to **Integrations → Connectors → Org Connectors**.
2. Select **Claude Managed Agents**.
3. Toggle the connector to **Enabled**.

No organization-level Claude credentials are required. Enabling the connector makes it available under **My Connections**.

## Connect a Claude Managed Agent

1. Go to **Integrations → Connectors → My Connections**.
2. Select **Claude Managed Agents**.
3. Enter the required values:
   * **API Key**: Your Claude API key.
   * **Agent ID**: The Claude agent that should run tasks.
   * **Environment ID**: The environment in which the agent should run.
4. Optionally add one or more **Vault IDs**.
5. Optionally expand **Memory Store** and configure a memory store.
6. Click **Connect**.

You can add multiple Claude Managed Agents connections. After the first connection, each additional connection requires a unique name so Agent Mode can distinguish them.

## Optional Configuration

{% hint style="warning" %}
Memory-store access initially defaults to **Read & write**. Explicitly select **Read only** for shared reference data and whenever the agent does not need to modify memory. Only keep **Read & write** when the agent deliberately needs to update the store: prompt injection from untrusted input can otherwise persist malicious instructions and affect later sessions.
{% endhint %}

| Setting             | Description                                                                                              |
| ------------------- | -------------------------------------------------------------------------------------------------------- |
| **Vault IDs**       | Secrets that the configured Claude agent can use while running tasks                                     |
| **Memory Store ID** | A memory store that persists context across agent runs                                                   |
| **Access**          | Whether the agent has **Read & write** (the initial default) or **Read only** access to the memory store |
| **Instructions**    | Guidance describing how the agent should use the memory store                                            |

## Manage a Connection

Open a Claude connection under **My Connections** to:

* Edit its Agent ID, Environment ID, vaults, or memory store.
* Rename it.
* Update the API key without recreating the connection.
* Click **Test connection** to send a short probe and confirm that Claude can reach the configured agent and environment.
* Configure tool restrictions to run actions automatically, require approval, or deny them.
* Add another Claude connection or disconnect the current one.

## Troubleshooting

| Issue                                                     | Resolution                                                                                                  |
| --------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------- |
| Claude Managed Agents is unavailable under My Connections | Ask a workspace admin to enable it under **Org Connectors**                                                 |
| Test connection fails                                     | Confirm the API key is active and the Agent ID and Environment ID belong to resources the key can access    |
| A task cannot access a secret                             | Confirm the correct vault ID is attached and that the configured agent has permission to use it             |
| Memory is not retained                                    | Confirm the Memory Store ID and access level, and check that the store is available to the configured agent |


# Notion

Connect Notion to search workspace knowledge and publish Agent Mode findings

Connect Notion to let Agent Mode find operational knowledge and turn new findings into durable documentation.

## What You Can Do

**Read and investigate**

> Find the checkout latency runbook in Notion and summarize the mitigation and rollback steps.
>
> Search our incident database for similar payment failures and compare their root causes.

**Take action**

> Create an RCA page under Engineering Incidents from this investigation. Include the timeline, root cause, evidence, and follow-up actions.
>
> Update the checkout runbook with the verified recovery steps we used today, and add a comment explaining the change.

Agent Mode can search and read pages and databases, and—when the relevant tools are allowed—create or update pages, databases, and comments.

## Setup

1. A workspace admin goes to **Integrations → Connectors → Org Connectors**, selects **Notion**, and enables it.
2. The admin clicks **Connect**, completes OAuth, and reviews the [tool restrictions](/use-groundcover/connectors/mcp#managing-tool-restrictions).
3. Each user goes to **Integrations → Connectors → My Connections** and selects **Notion**.
4. Click **Connect** and authorize the Notion workspace and pages that Agent Mode may access.

{% hint style="info" %}
Agent Mode uses the connected user's Notion permissions and the pages shared during OAuth, subject to groundcover's tool restrictions. Write actions may require approval.
{% endhint %}


# Pylon

Connect Pylon to investigate customer issues and take support actions from Agent Mode

Connect Pylon to give Agent Mode customer context during an investigation and let it turn technical findings into support work.

## What You Can Do

**Read and investigate**

> Find open Pylon issues mentioning checkout timeouts, group them by customer and priority, and summarize the common symptoms.
>
> Read the latest conversation for this account and show me the relevant knowledge-base guidance before I respond.

**Take action**

> Create a high-priority Pylon issue for the customer impact shown in this monitor investigation. Include the affected service, evidence, and investigation link, then assign it to the support engineering team.
>
> Add an internal note with the root cause and workaround to the related issue, then move it to waiting on customer.

Available actions depend on the tools and permissions enabled for the connection. They can include creating, assigning, and updating issues as well as adding customer-facing or internal context.

## Setup

1. A Pylon admin grants **MCP Access** to the Member or Admin account that will authorize the groundcover organization connection, and to every user who will create a personal connection. Viewer and Integration accounts cannot connect.
2. The Pylon admin enables **Settings → AI Controls → MCP Server** in Pylon.
3. A groundcover workspace admin goes to **Integrations → Connectors → Org Connectors**, selects **Pylon**, and enables it.
4. The admin clicks **Connect**, completes OAuth with an account that has **MCP Access**, and reviews the [tool restrictions](/use-groundcover/connectors/mcp#managing-tool-restrictions).
5. Each user goes to **Integrations → Connectors → My Connections**, selects **Pylon**, and clicks **Connect**.
6. Complete the Pylon OAuth flow with a Pylon Member or Admin account that has **MCP Access**.

Without **MCP Access**, authorization fails and Agent Mode cannot discover Pylon tools.

{% hint style="info" %}
Agent Mode acts with the connected user's Pylon permissions, subject to groundcover's tool restrictions. Write actions may require approval.
{% endhint %}


# GitHub

Connect GitHub to investigate code and manage development work from Agent Mode

Connect GitHub to let Agent Mode correlate production behavior with code, issues, pull requests, and Actions runs—and take follow-up actions without losing investigation context.

## What You Can Do

**Read and investigate**

> Find recent pull requests that changed the checkout service, inspect their diffs, and identify changes that could explain this latency regression.
>
> Summarize the failing GitHub Actions run for this repository and check whether an open issue or pull request already tracks it.

**Take action**

> Create a GitHub issue for this monitor problem with the impact, relevant telemetry, suspected code path, and a link to the groundcover investigation.
>
> Add the confirmed root cause and reproduction steps to the existing issue, then draft a pull request with the fix for me to review.

The available repositories and actions depend on the connected account, token scopes, and tool restrictions.

## Setup with a Personal Access Token

GitHub uses personal access tokens by default.

1. A workspace admin goes to **Integrations → Connectors → Org Connectors**, selects **GitHub**, and enables it.
2. The admin enters a GitHub personal access token to discover the available tools and reviews the [tool restrictions](/use-groundcover/connectors/mcp#managing-tool-restrictions).
3. Each user goes to **Integrations → Connectors → My Connections**, selects **GitHub**, and clicks **Connect**.
4. Enter a personal access token with access to the repositories and operations Agent Mode should use.

## Optional: Switch to OAuth

A workspace admin can switch GitHub to OAuth after creating a GitHub OAuth App:

1. Open [GitHub Developer settings](https://github.com/settings/developers), select **OAuth Apps**, and click **New OAuth App**.
2. Complete the form:
   * **Application name**: Any recognizable name, such as `groundcover`.
   * **Homepage URL**: Any valid URL, such as your top-level groundcover URL (<https://app.groundcover.com> unless you use a custom domain).
   * **Application description**: Optional.
   * **Authorization callback URL**: `https://app.groundcover.com/connectors/mcp/oauth/callback` (replace the base URL if you use a custom domain).
3. Leave **Enable Device Flow** disabled and click **Register application**.
4. Copy the **Client ID** and generate a **Client Secret**.
5. In groundcover, open **Integrations → Connectors → Org Connectors → GitHub** and click **Edit** under **General Information**.
6. Set **Authentication Method** to **OAuth**, enter the Client ID and Client Secret, and click **Save**.
7. Click **Connect** under **Tool Restrictions** and complete authorization.

<div data-with-frame="true"><figure><img src="/files/65HtO0zJZp3i06CRc44q" alt="GitHub form for registering a new OAuth app"><figcaption><p>Set the Homepage URL to your UI domain, the Authorization callback URL to https://app.groundcover.com/connectors/mcp/oauth/callback and leave Device Flow disabled.</p></figcaption></figure></div>

groundcover includes the tenant UUID and backend ID in the OAuth `state` parameter. The callback URL must be set to `https://app.groundcover.com/connectors/mcp/oauth/callback` exactly GitHub will reject the OAuth flow if the registered redirect URI does not match.

New personal connections use OAuth after the switch; existing token connections continue to work. Switching back to **Token** removes the stored OAuth client, but existing OAuth connections continue until they expire and must then reconnect with a token.

{% hint style="info" %}
Agent Mode acts with the connected user's GitHub permissions, subject to groundcover's tool restrictions. Write actions may require approval.
{% endhint %}


# PagerDuty

Connect PagerDuty to investigate and manage incidents from Agent Mode

Connect PagerDuty to bring incident, service, and on-call context into Agent Mode and take response actions from the same investigation.

This connector is for Agent Mode. To send monitor notifications to PagerDuty, configure a [PagerDuty Destination](/integrations/connected-apps/pagerduty).

## What You Can Do

**Read and investigate**

> Who is on call for the payments service, what incidents are currently open, and have we seen this failure mode before?
>
> Summarize this incident's alerts, responder notes, affected service, and recent incident history.

**Take action**

> Acknowledge the PagerDuty incident linked to this monitor, add a note with the investigation summary, and tell me who owns the next escalation.
>
> Create an incident for this confirmed checkout outage with the relevant service, urgency, and groundcover investigation link.

The actions Agent Mode can take depend on the connected user's PagerDuty permissions and the enabled tools.

## Setup with an API Token

PagerDuty uses API tokens by default.

1. A workspace admin goes to **Integrations → Connectors → Org Connectors**, selects **PagerDuty**, and enables it.
2. The admin enters a PagerDuty API token to discover the available tools and reviews the [tool restrictions](/use-groundcover/connectors/mcp#managing-tool-restrictions).
3. Each user goes to **Integrations → Connectors → My Connections**, selects **PagerDuty**, and clicks **Connect**.
4. Enter a personal PagerDuty API token.

## Optional: Switch to OAuth

After registering an OAuth client in PagerDuty, a workspace admin can switch the connector to OAuth:

1. Copy the OAuth client's **Client ID** and **Client Secret**.
2. In groundcover, open **Integrations → Connectors → Org Connectors → PagerDuty**.
3. Under **General Information**, click **Edit** and set **Authentication Method** to **OAuth**.
4. Enter the Client ID and Client Secret, then click **Save**.
5. Click **Connect** under **Tool Restrictions** and complete authorization.

New personal connections use OAuth after the switch; existing token connections continue to work. Switching back to **Token** removes the stored OAuth client, but existing OAuth connections continue until they expire and must then reconnect with a token.

{% hint style="info" %}
Agent Mode acts with the connected user's PagerDuty permissions, subject to groundcover's tool restrictions. Write actions may require approval.
{% endhint %}


# Rootly

Connect Rootly to investigate incidents and coordinate response from Agent Mode

Connect Rootly to give Agent Mode incident history and on-call context, then use that context to coordinate and document the response.

This connector is for Agent Mode. To send monitor notifications to Rootly, configure a [Rootly Destination](/integrations/connected-apps/incident.io-1).

## What You Can Do

**Read and investigate**

> Find Rootly incidents similar to this database saturation event and summarize their causes, mitigations, and time to resolution.
>
> Show the current and next on-call responders for the database team and summarize incidents from the current shift.

**Take action**

> Create a Rootly incident for this confirmed production impact with the correct service, severity, summary, and investigation link.
>
> Update the incident with the confirmed root cause and create action items for the remediation work, including suggested owners.

Rootly tools can cover incidents, alerts, on-call schedules, retrospectives, and action items. The exact actions available depend on permissions and tool restrictions.

## Setup

1. A workspace admin goes to **Integrations → Connectors → Org Connectors**, selects **Rootly**, and enables it.
2. The admin clicks **Connect**, completes OAuth, and reviews the [tool restrictions](/use-groundcover/connectors/mcp#managing-tool-restrictions).
3. Each user goes to **Integrations → Connectors → My Connections**, selects **Rootly**, and clicks **Connect**.
4. Complete the Rootly OAuth flow.

{% hint style="info" %}
Agent Mode acts with the connected user's Rootly permissions, subject to groundcover's tool restrictions. Write actions may require approval.
{% endhint %}


# incident.io

Connect incident.io to analyze operations and manage incidents from Agent Mode

Connect incident.io to investigate incident history, alerts, on-call coverage, and service ownership, then take response and follow-up actions.

This connector is for Agent Mode. To send monitor notifications to incident.io, configure an [incident.io Destination](/integrations/connected-apps/incident.io).

## What You Can Do

**Read and investigate**

> Show incidents affecting the payments service in the last 30 days and summarize recurring root causes and outstanding follow-ups.
>
> Who is on call for platform right now, and which alerts caused the most overnight responder work this month?

**Take action**

> Declare a high-severity incident for the checkout degradation shown in this investigation and include the affected service and evidence.
>
> Acknowledge the page for this incident, update its status and severity, and create follow-ups for the remediation items we identified.

Available tools include incident and escalation management as well as alert, schedule, catalog, and investigation analysis.

## Setup

1. An incident.io admin enables **Settings → MCP** in incident.io.
2. A groundcover workspace admin goes to **Integrations → Connectors → Org Connectors**, selects **incident.io**, and enables it.
3. The admin clicks **Connect**, completes OAuth, and reviews the [tool restrictions](/use-groundcover/connectors/mcp#managing-tool-restrictions).
4. Each user goes to **Integrations → Connectors → My Connections**, selects **incident.io**, and clicks **Connect**.
5. Complete the incident.io OAuth flow.

{% hint style="info" %}
Agent Mode acts with the connected user's incident.io permissions, subject to groundcover's tool restrictions. Write actions may require approval. incident.io OAuth access expires periodically, so users may need to authorize again.
{% endhint %}


# Atlassian Rovo

Connect Atlassian Rovo to work with Jira and Confluence from Agent Mode

Connect Atlassian Rovo to give Agent Mode access to Jira work and Confluence knowledge, and to turn production findings into assigned, durable follow-up work.

## What You Can Do

**Read and investigate**

> Search Jira for open or past incidents matching this checkout failure, then find the related Confluence runbook and summarize the response steps.
>
> Read the Jira issue for this service degradation and compare its evidence with the latest Confluence architecture notes.

**Take action**

> Create a Jira ticket for this monitor issue. Include the impact, relevant telemetry, investigation link, and acceptance criteria, then assign it to the service owner.
>
> Create or update a Confluence RCA page with the investigation summary, timeline, root cause, evidence, and follow-up actions so the team can learn from it later.

Depending on enabled tools and permissions, Agent Mode can search and read Jira issues and Confluence pages, create and update issues or pages, add comments, and move Jira work through its workflow.

## Setup

1. If your Atlassian organization restricts Rovo MCP access, an Atlassian admin opens **Atlassian Administration → Rovo → Rovo MCP server**, selects **Add domain**, and adds the HTTPS URL pattern for the deployed groundcover host. For example, use `https://app.groundcover.com/**` for the standard groundcover environment. Include the `https://` protocol; Atlassian rejects bare domains.
2. A groundcover workspace admin goes to **Integrations → Connectors → Org Connectors**, selects **Atlassian Rovo**, and enables it.
3. The admin clicks **Connect**, completes OAuth, and reviews the [tool restrictions](/use-groundcover/connectors/mcp#managing-tool-restrictions).
4. Each user goes to **Integrations → Connectors → My Connections**, selects **Atlassian Rovo**, and clicks **Connect**.
5. Complete the Atlassian OAuth flow and select the Atlassian site to authorize.

{% hint style="info" %}
Agent Mode acts with the connected user's Jira and Confluence permissions, subject to groundcover's tool restrictions. Write actions may require approval.
{% endhint %}


# Remote MCP Servers

Allowlist Remote MCP Servers to extend Agent Mode with custom tools and capabilities

The Remote MCP Servers connector lets [Agent Mode](/use-groundcover/agent-mode) use tools from external [Model Context Protocol](https://modelcontextprotocol.io) (MCP) servers. Admins allowlist Remote MCP Servers at the organization level and choose token or OAuth authentication. Individual users then connect their own account to access the tools those servers expose.

For services that already appear as their own connector card, follow that service's setup guide in [Connectors](/use-groundcover/connectors). Use **Remote MCP Servers** for any other compatible server.

{% hint style="info" %}
Remote MCP Servers must be allowlisted by a workspace admin before users can connect. See [Connectors](/use-groundcover/connectors) for the two-tier model.
{% endhint %}

## Admin Setup

### Allowlisting a Remote MCP Server

1. Go to **Integrations → Connectors** and open the **Org Connectors** tab
2. Select **Remote MCP Servers** and click **Allowlist a Remote MCP Server**
3. Fill in the server details:
   * **Connector Name**: A unique name to identify this server
   * **URL**: The MCP server endpoint (e.g., `https://mcp.example.com/mcp`)
   * **Authentication Method**: Select **Token** or **OAuth**
   * **Client ID** and **Client Secret**: Optional fields shown for OAuth; see [OAuth Client Registration](#oauth-client-registration)
   * **Headers**: Optional key-value pairs sent with every request to the server
4. Click **Save**

The server is now available for users to connect to.

### OAuth Client Registration

When **Authentication Method** is **OAuth**, choose the setup that matches the MCP server:

* **Dynamic client registration**: Leave **Client ID** and **Client Secret** empty. During authorization, groundcover asks the server to register an OAuth client automatically.
* **Pre-registered client**: If the server does not support dynamic client registration, create an OAuth application with the server's provider and enter both its **Client ID** and **Client Secret**. Register these two redirect URLs, replacing the example domain with your groundcover domain:
  * `https://app.groundcover.com/settings/ai-agents/connectors/mcp/oauth/callback`
  * `https://app.groundcover.com/connectors/mcp/oauth/callback`

The Client ID and Client Secret are a pair. If you enter one, you must enter the other.

{% hint style="info" %}
OAuth servers must publish compatible authorization metadata. Dynamic registration registers both groundcover redirect URLs automatically; a pre-registered client must allow both URLs before admins and individual users can authorize.
{% endhint %}

Custom headers are independent of the authentication method and are sent with both token and OAuth requests.

{% hint style="info" %}
Users can view the server's name, URL, and headers from their manage page, but cannot modify any details set by the admin.
{% endhint %}

{% hint style="info" %}
You can allowlist multiple Remote MCP Servers. Each one is added independently with its own Connector Name, URL, and tool restrictions. Repeat the steps above for each server you want to make available.
{% endhint %}

### Authorize the Admin Connection

groundcover uses an admin connection to discover the server's tools and configure organization-wide restrictions:

1. Click **Manage** on an allowlisted Remote MCP Server
2. In the **Tool Restrictions** section:
   * For a token connector, enter an admin API token.
   * For an OAuth connector, click **Connect** and complete the provider's authorization flow.
3. After authorization, the full list of available tools appears.

{% hint style="info" %}
Admin credentials are stored securely. A token can be updated at any time; an OAuth connection can be authorized again if access expires or is revoked.
{% endhint %}

### Managing Tool Restrictions

Each tool can be set to one of three permission levels:

| Permission         | Behavior                                   |
| ------------------ | ------------------------------------------ |
| **Automatically**  | Tool executes without user approval        |
| **Needs Approval** | User must approve each execution (default) |
| **Deny**           | Tool is completely blocked                 |

**Default Behavior**: Sets the permission for all tools that don't have a custom override, and applies automatically to new tools the server adds in the future.

**Per-tool override**: Click the permission selector on an individual tool to override the default. A reset button appears on customized tools to revert them back to the default.

Use the search field to filter tools by name.

{% hint style="warning" %}
Admin tool restrictions act as the **ceiling** for user permissions. Users can make their own permissions more restrictive, but never more permissive than what the admin allows. See [How Tool Restrictions Work](#how-tool-restrictions-work) for details.
{% endhint %}

### Disabling Remote MCP Servers

Admins can disable the Remote MCP Servers connector at any time from **Integrations → Connectors** under the **Org Connectors** tab. When disabled, Agent Mode cannot use any MCP tools, even if users have already connected their tokens. User credentials and tool restriction settings are preserved, not deleted. Once the admin re-enables Remote MCP Servers, all previously configured user connections resume working as before.

### Removing a Remote MCP Server

Click **Remove** on the server in the credentials list. This disconnects all users from the server.

## User Setup

### Connecting to a Remote MCP Server

1. Go to **Integrations → Connectors → My Connections**
2. Select **Remote MCP Servers** and browse the list of servers your admin has allowlisted
3. Click **Connect** on the server you want to use
4. Complete the configured authentication flow:
   * For a token connector, enter your personal API token and click **Connect**.
   * For an OAuth connector, authorize your account with the external provider.

You'll be taken to the server's manage page where you can view its configuration and manage your tool restrictions.

{% hint style="info" %}
The server's name, URL, and headers are configured by your admin and displayed as read-only. Contact your admin if any details need to change.
{% endhint %}

### Listing Available Tools

The **Tool Restrictions** section at the bottom of the manage page shows every tool the server exposes, including each tool's name and description. Use the search field to filter the list.

### Managing Personal Tool Restrictions

You can set your own tool permissions using the same three levels as the admin, with one key difference, **your choices are capped by the admin's restrictions**.

| Permission         | Behavior                                   |
| ------------------ | ------------------------------------------ |
| **Automatically**  | Tool executes without user approval        |
| **Needs Approval** | User must approve each execution (default) |
| **Deny**           | Tool is completely blocked                 |

* **Default Behavior**: Sets your personal default for all tools without a custom override.
* **Per-tool override**: Customize individual tools. A reset button reverts a tool to your personal default.

Tools where the admin has set a more restrictive ceiling show a **lock icon** with a "Restricted by admin" tooltip. You cannot select a permission more permissive than what the admin allows.

### Updating or Reauthorizing Credentials

For a token connector, click **Update Token** on the manage page to replace your personal API token. For an OAuth connector, click **Connect** again when groundcover prompts you to reauthorize.

### Disconnecting

Click the **Disconnect** button to remove your connection. You can reconnect at any time.

## How Tool Restrictions Work

Tool restrictions use a **least-permissive** model with two tiers:

1. **Admin tier**: The admin sets org-wide permissions (Automatically, Needs Approval, or Deny) per Remote MCP Server. These act as the maximum allowed permission.
2. **User tier**: Each user sets personal permissions that can only be **equal to or more restrictive** than the admin's setting.

The effective permission for any tool is always the most restrictive of the two tiers.

### Examples

| Admin Permission | User Permission | Effective Permission           |
| ---------------- | --------------- | ------------------------------ |
| Automatically    | Automatically   | Automatically                  |
| Automatically    | Needs Approval  | Needs Approval                 |
| Automatically    | Deny            | Deny                           |
| Needs Approval   | Automatically   | Needs Approval                 |
| Needs Approval   | Needs Approval  | Needs Approval                 |
| Needs Approval   | Deny            | Deny                           |
| Deny             | Any             | Deny, tool is blocked entirely |

This means:

* If the admin denies a tool, no user can use it.
* If the admin sets a tool to Needs Approval, users can keep it at Needs Approval or deny it, but cannot set it to Automatically.
* The admin's default permission also applies to **new tools** the server adds after initial setup.

## Troubleshooting

| Issue                                  | Resolution                                                                                                                                                           |
| -------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Remote MCP Server not listed           | Ask a workspace admin to allowlist the server under **Integrations → Connectors → Org Connectors → Remote MCP Servers**                                              |
| "Restricted by admin" on a tool        | The admin has set a more restrictive permission; contact your admin to change it                                                                                     |
| Tools not loading                      | Confirm that the admin connection is authorized and that the MCP server is reachable                                                                                 |
| Token connection failed                | Check that the personal token is valid and that the MCP server is accessible                                                                                         |
| OAuth connection failed before consent | Confirm the server publishes OAuth discovery metadata; if it does not support dynamic client registration, recreate the connector with a Client ID and Client Secret |
| OAuth callback fails                   | Confirm the OAuth application's callback URL is correct and that both Client ID and Client Secret belong to the same application                                     |


# Insights

Quickly understand your data with groundcover

groundcover insights give you a clear snapshot of notable events in your data. Currently, the platform supports **Error Anomalies**, with more insight types on the way.

<figure><img src="/files/zRCMwvKnNQqYB7vn04SM" alt=""><figcaption><p>Logs Insights</p></figcaption></figure>

#### Error Anomalies

Error Anomalies instantly highlight workloads, containers, or environments experiencing unusual spikes in **Error or Critical logs**, as well as **Traces marked with an error status**. These anomalies are detected using statistical algorithms, continuously refined through user feedback for accuracy.

Each insight surfaces trends based on the entity’s error signals (e.g., workload, container, etc.):

On **Logs**, anomalies are based on logs filtered by `level:error` or `level:critical`, and grouped by:

* `workload`
* `container`
* `namespace`
* `environment`
* `cluster`

On **Traces**, anomalies are based on traces filtered by `status:error`, and grouped by a more granular set of dimensions:

* `protocol_type`
* `return_code`
* `role` (client/server)
* `workload`
* `container`
* `namespace`
* `environment`
* `cluster`


# AI Observability

For an overview of capabilities, supported providers, and cost tracking, see [AI Observability](/capabilities/ai-observability).

***

## Getting Started

### eBPF (Automatic)

If your services call a [supported provider](/capabilities/ai-observability#supported-providers), groundcover is already capturing AI API calls automatically — no SDKs, no code changes. By default, you get model identification, token counts, cost, latency, and full prompt/response content.

Open [AI Observability](https://app.groundcover.com/ai-observability) in your console to see what's there. If the page is empty, see [Troubleshooting](/use-groundcover/ai-observability/troubleshooting).

### SDK (Full Visibility)

To see agent workflows, tool execution chains, conversation threading, and span trees with parent-child relationships, add OpenTelemetry GenAI instrumentation to your services.

groundcover ingests traces using [OpenTelemetry GenAI Semantic Conventions](https://opentelemetry.io/docs/specs/semconv/gen-ai/). Three attributes are all you need to unlock the core experience:

| Attribute               | What it enables                           |
| ----------------------- | ----------------------------------------- |
| `gen_ai.operation.name` | Span appears in AI Observability          |
| `gen_ai.provider.name`  | Provider identification and filtering     |
| `gen_ai.request.model`  | Model identification and cost attribution |

`gen_ai.operation.name` is required for spans to appear in AI Observability. `gen_ai.provider.name` and `gen_ai.request.model` are strongly recommended — without them, you lose provider filtering and cost attribution.

Everything else — messages, agent name, tool definitions, conversation ID — adds depth but isn't required to get started.

***

## SDK Instrumentation

### Approaches

| Approach                                                                         | Best for                                                                                                                                 |
| -------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------- |
| **Auto-instrumentation SDKs** (Traceloop/OpenLLMetry, LangSmith, etc.)           | Fastest path to visibility. Add a package, initialize, and your agent workflows appear. Works well for most teams.                       |
| **Native OTel provider libraries** (OpenAI, Anthropic, Google, Claude Agent SDK) | The standard the ecosystem is converging on. Provider-maintained and actively developing.                                                |
| **Manual instrumentation**                                                       | Full control. You emit `gen_ai.*` attributes directly — the data is exactly what you put in. Precise cost attribution, clean span trees. |

**Not sure where to start?** Most teams begin with **Traceloop** — one init call covers OpenAI, Anthropic, and agent frameworks (LangGraph, CrewAI, Pydantic AI) with full content capture out of the box.

```python
# pip install traceloop-sdk opentelemetry-exporter-otlp-proto-http
from opentelemetry.exporter.otlp.proto.http.trace_exporter import OTLPSpanExporter
from traceloop.sdk import Traceloop

Traceloop.init(
    app_name="my-service",
    exporter=OTLPSpanExporter(endpoint="<COLLECTOR_ENDPOINT>/v1/traces"),
)
# Your existing code runs unchanged — OpenAI, Anthropic, LangGraph, CrewAI all work
```

Replace `<COLLECTOR_ENDPOINT>` with your endpoint from the [**Data Sources**](https://app.groundcover.com/data-sources) page, which also has setup instructions for other SDKs and agent frameworks.

### Adding Depth

Once you have the core attributes above, more attributes unlock more value:

| Add this                                                      | Get this                                  |
| ------------------------------------------------------------- | ----------------------------------------- |
| `gen_ai.usage.input_tokens` + `gen_ai.usage.output_tokens`    | Token counts and cost attribution         |
| `gen_ai.conversation.id`                                      | Multi-turn session grouping               |
| `gen_ai.agent.name`                                           | Per-agent identification and analytics    |
| `gen_ai.input.messages` + `gen_ai.output.messages`            | Prompt/response content in trace drawer   |
| Framework spans (from your agent framework's SDK integration) | Full span trees and behavioral visibility |

For the complete list of supported attributes, see [Span Attributes](/use-groundcover/ai-observability/attributes). For query patterns using these attributes, see [Example Queries](/use-groundcover/ai-observability/example-queries).

***

## Verify Setup

After deployment, verify data is flowing:

1. **Generate AI traffic** — make a few calls to a [supported provider](/capabilities/ai-observability#supported-providers) from a service running on a monitored node
2. **Open AI Observability** — navigate to [AI Observability](https://app.groundcover.com/ai-observability) in your console
3. **Check for spans** — within 1-2 minutes, GenAI spans should appear
4. **Verify data** — confirm `gen_ai.operation.name` is populated on individual spans. You can also run this query in [Data Explorer](/use-groundcover/querying-your-groundcover-data/explore-and-monitors-query-builder) to check programmatically:

   ```gcql
   span_type:gen_ai | stats count()
   ```

   If `gen_ai.request.model` and `gen_ai.provider.name` are also present, cost attribution and provider filtering will be available.

If no data appears after several minutes, see [Troubleshooting](/use-groundcover/ai-observability/troubleshooting).

***

## OTel GenAI Conventions

The OpenTelemetry GenAI semantic conventions have evolved through three generations. groundcover normalizes all of them automatically — your data displays consistently regardless of which convention your SDK emits.

| Era       | Key Attributes                                                          | Status  | Used By            |
| --------- | ----------------------------------------------------------------------- | ------- | ------------------ |
| **Era 1** | `gen_ai.prompt` `gen_ai.completion` `gen_ai.system`                     | Legacy  | Most SDKs today    |
| **Era 2** | Per-message log events                                                  | Legacy  | Near-zero adoption |
| **Era 3** | `gen_ai.input.messages` `gen_ai.output.messages` `gen_ai.provider.name` | Current | Official OTel      |

{% hint style="info" %}
Most SDKs still emit Era 1 attributes — this is normal. groundcover maps them to the current convention automatically. You don't need to change anything.
{% endhint %}

***

## Data Privacy

By default, groundcover captures full prompt and response content for supported providers. For teams that need to disable AI data collection or strip content from storage, see [Privacy Controls](/use-groundcover/ai-observability/privacy-controls).


# Span Attributes

groundcover adheres to the [OpenTelemetry GenAI Semantic Conventions](https://opentelemetry.io/docs/specs/semconv/gen-ai/). The more attributes your spans carry, the richer the experience — cost analytics, prompt debugging, execution chain visibility.

For how groundcover handles older SDK attribute conventions (Era 1 `gen_ai.prompt`/`gen_ai.completion` and Era 2 log events), see [OTel GenAI Conventions](/use-groundcover/ai-observability#otel-genai-conventions).

***

## Core Attributes

| Attribute                        | Type      | Description                      |
| -------------------------------- | --------- | -------------------------------- |
| `gen_ai.provider.name`           | string    | GenAI provider                   |
| `gen_ai.operation.name`          | string    | Operation type                   |
| `gen_ai.request.model`           | string    | Model requested                  |
| `gen_ai.response.model`          | string    | Model that responded             |
| `gen_ai.response.id`             | string    | Completion ID                    |
| `gen_ai.usage.input_tokens`      | int       | Tokens consumed by the prompt    |
| `gen_ai.usage.output_tokens`     | int       | Tokens generated in the response |
| `gen_ai.response.finish_reasons` | string\[] | Why the model stopped            |
| `gen_ai.conversation.id`         | string    | Session/thread identifier        |

***

## Additional Attributes

### Token Usage

| Attribute                                  | Type | Description                             |
| ------------------------------------------ | ---- | --------------------------------------- |
| `gen_ai.usage.cache_read.input_tokens`     | int  | Input tokens served from provider cache |
| `gen_ai.usage.cache_creation.input_tokens` | int  | Input tokens written to provider cache  |

### Content

| Attribute                    | Type | Description                               |
| ---------------------------- | ---- | ----------------------------------------- |
| `gen_ai.input.messages`      | JSON | Structured input messages (role + parts)  |
| `gen_ai.output.messages`     | JSON | Structured output messages (role + parts) |
| `gen_ai.system_instructions` | JSON | System prompt, separate from chat history |
| `gen_ai.tool.definitions`    | JSON | Tool schemas passed to the model          |

eBPF captures content automatically for supported providers. For SDK instrumentation, content capture is controlled by your SDK configuration — see [Privacy Controls](/use-groundcover/ai-observability/privacy-controls) for details.

### Agent

| Attribute                  | Type   | Description               |
| -------------------------- | ------ | ------------------------- |
| `gen_ai.agent.name`        | string | Human-readable agent name |
| `gen_ai.agent.id`          | string | Unique agent identifier   |
| `gen_ai.agent.description` | string | Agent description         |
| `gen_ai.agent.version`     | string | Agent version             |

### Tool Execution

| Attribute                    | Type   | Description                          |
| ---------------------------- | ------ | ------------------------------------ |
| `gen_ai.tool.name`           | string | Name of the tool                     |
| `gen_ai.tool.type`           | string | `function`, `extension`, `datastore` |
| `gen_ai.tool.call.id`        | string | Tool call identifier                 |
| `gen_ai.tool.call.arguments` | JSON   | Parameters passed to the tool        |
| `gen_ai.tool.call.result`    | JSON   | Result returned by the tool          |

For request parameters, response metadata, and the complete specification, see the [OpenTelemetry GenAI Semantic Conventions](https://opentelemetry.io/docs/specs/semconv/gen-ai/).

***

## groundcover Enrichment

groundcover adds the following attributes at ingestion time. These are not part of the OTel specification — they're computed by groundcover for cost analytics and filtering.

### Cost

| Attribute            | Type   | Description                                                                                |
| -------------------- | ------ | ------------------------------------------------------------------------------------------ |
| `gc.llm_cost.input`  | float  | Cost of input tokens, computed from the model's pricing table                              |
| `gc.llm_cost.output` | float  | Cost of output tokens, computed from the model's pricing table                             |
| `gc.llm_cost.status` | string | `computed` when pricing is available, `unpriced` when the model isn't in the pricing table |

Cost is calculated per-span using `gen_ai.request.model` and `gen_ai.provider.name` to look up the correct pricing. When cache tokens are present (`gen_ai.usage.cache_read.input_tokens`), cached and non-cached portions are priced separately.

{% hint style="info" %}
If `gc.llm_cost.status` shows `unpriced`, the model isn't in groundcover's pricing table yet. This typically resolves automatically as the table is updated. [Contact us](https://www.groundcover.com/contact) if it persists.
{% endhint %}

### Filtering

| Attribute                   | Type       | Description                                                                                                                         |
| --------------------------- | ---------- | ----------------------------------------------------------------------------------------------------------------------------------- |
| `span_type:gen_ai`          | expression | [GCQL](/use-groundcover/querying-your-groundcover-data/groundcover-query-language) filter — use in search, dashboards, and monitors |
| `protocol_type == "gen_ai"` | expression | [Traces Pipeline](/use-groundcover/data-pipelines/traces-pipeline) condition — use in OTTL pipeline rules                           |

groundcover assigns `span_type:gen_ai` automatically — eBPF spans from supported providers always receive it; SDK spans receive it when `gen_ai.operation.name` is present (best-effort). You do not need to set it yourself.

See [Example Queries](/use-groundcover/ai-observability/example-queries) for patterns using these attributes.


# Example Queries

AI Observability data is stored as traces with `gen_ai.*` attributes and [groundcover enrichment](/use-groundcover/ai-observability/attributes#groundcover-enrichment) (`gc.llm_cost.*`). Filter GenAI spans using `span_type:gen_ai` in any [GCQL](/use-groundcover/querying-your-groundcover-data/groundcover-query-language) query.

Run these queries in the [Data Explorer](/use-groundcover/querying-your-groundcover-data/explore-and-monitors-query-builder), the search bar in [AI Observability](https://app.groundcover.com/ai-observability), or as the basis for [dashboard widgets](/use-groundcover/dashboards-and-alerts/create-a-dashboard).

***

## Call Volume & Distribution

**Total LLM calls by model:**

```gcql
span_type:gen_ai | stats by(gen_ai.request.model) count()
```

**Calls by provider:**

```gcql
span_type:gen_ai | stats by(gen_ai.provider.name) count()
```

**Operation type breakdown:**

```gcql
span_type:gen_ai | stats by(gen_ai.operation.name) count()
```

***

## Token Usage

**Total tokens (input vs output):**

```gcql
span_type:gen_ai | stats sum(gen_ai.usage.input_tokens) as input_tokens, sum(gen_ai.usage.output_tokens) as output_tokens
```

**Average tokens per request by model:**

```gcql
span_type:gen_ai | stats by(gen_ai.request.model) avg(gen_ai.usage.input_tokens) as avg_input, avg(gen_ai.usage.output_tokens) as avg_output
```

**Top services by token consumption:**

```gcql
span_type:gen_ai | stats by(service.name) sum(gen_ai.usage.input_tokens) as total_tokens | sort by(total_tokens desc) | limit 10
```

***

## Cost Analysis

**Total cost by model:**

```gcql
span_type:gen_ai | stats by(gen_ai.request.model) sum(gc.llm_cost.input) as input_cost, sum(gc.llm_cost.output) as output_cost | math input_cost + output_cost as total_cost | sort by(total_cost desc)
```

**Cost by service:**

```gcql
span_type:gen_ai | stats by(service.name) sum(gc.llm_cost.input) as input_cost, sum(gc.llm_cost.output) as output_cost | math input_cost + output_cost as cost | sort by(cost desc)
```

**Most expensive individual calls:**

```gcql
span_type:gen_ai | math gc.llm_cost.input + gc.llm_cost.output as total_cost | sort by(total_cost desc) | limit 20
```

***

## Latency & Performance

**Average latency by provider:**

```gcql
span_type:gen_ai | stats by(gen_ai.provider.name) avg(duration_seconds) as avg_latency
```

**P99 latency by model:**

```gcql
span_type:gen_ai | stats by(gen_ai.request.model) quantile(0.99, duration_seconds) as p99_latency
```

**Latency distribution:**

```gcql
span_type:gen_ai | stats quantile(0.50, duration_seconds) as p50, quantile(0.95, duration_seconds) as p95, quantile(0.99, duration_seconds) as p99
```

***

## Errors & Issues

**Error count by model:**

```gcql
span_type:gen_ai status:error | stats by(gen_ai.request.model) count() as errors
```

**Errors with details:**

```gcql
span_type:gen_ai status:error | stats by(gen_ai.request.model, error.type) count()
```

***

## Agent & Tool Tracking

**Calls by agent:**

```gcql
span_type:gen_ai gen_ai.agent.name:* | stats by(gen_ai.agent.name) count()
```

**Tool execution frequency:**

```gcql
span_type:gen_ai gen_ai.operation.name:execute_tool | stats by(gen_ai.tool.name) count()
```

**Cost breakdown by agent and model:**

```gcql
span_type:gen_ai gen_ai.agent.name:* | stats by(gen_ai.agent.name, gen_ai.request.model) sum(gc.llm_cost.input) as input_cost, sum(gc.llm_cost.output) as output_cost | math input_cost + output_cost as total_cost | sort by(total_cost desc)
```

***

## Visualization Best Practices

| Use Case            | Recommended Display       | Group By               |
| ------------------- | ------------------------- | ---------------------- |
| Call volume trends  | Time Series               | `gen_ai.request.model` |
| Cost breakdown      | Pie Chart                 | `gen_ai.provider.name` |
| Top expensive calls | Top List                  | —                      |
| Token usage stats   | Stat                      | -                      |
| Model comparison    | Table                     | `gen_ai.request.model` |
| Latency percentiles | Time Series               | -                      |
| Error trends        | Time Series (Stacked Bar) | `error.type`           |

To build dashboard widgets with these queries, see [Creating Dashboards](/use-groundcover/dashboards-and-alerts/create-a-dashboard) — select **Traces** mode when adding a chart widget.


# Privacy Controls

## Privacy Controls

With groundcover's [BYOC (Bring Your Own Cloud)](/architecture/byoc) deployment, all AI telemetry — including prompts, responses, and model metadata — stays within your own infrastructure. groundcover processes AI data exclusively in your environment; nothing is sent to external services.

For teams that need additional control, groundcover supports two approaches: **disable collection entirely** — a hard guarantee that no AI spans reach storage — or **metadata-only mode**, which keeps performance metadata (tokens, cost, latency, model) while stripping prompts and responses (best-effort; see scope note below). Both use [Traces Pipeline](/use-groundcover/data-pipelines/traces-pipeline) rules.

***

### Disable GenAI Data Collection

To prevent all GenAI spans from reaching storage, add the following rule to your traces pipeline. Non-GenAI traffic (HTTP, gRPC, DB, etc.) is completely unaffected.

1. Open [**Traces Pipeline Settings**](https://app.groundcover.com/settings/traces-pipeline)
2. Switch to **YAML mode**
3. Add the rule below
4. Save — changes take effect on new spans immediately, no sensor restart required

{% code title="Traces Pipeline YAML" %}

```yaml
ottlRules:
  - ruleName: gc-genai-off
    conditions:
      - 'protocol_type == "gen_ai"'
    statements:
      - 'set(drop, true)'
```

{% endcode %}

This drops all GenAI spans — both eBPF-captured and SDK-instrumented — before they reach storage.

To re-enable GenAI data collection, remove the rule and save.

#### Custom Providers

groundcover auto-detects **OpenAI**, **Anthropic**, and **AWS Bedrock** traffic. If you use a provider that isn't auto-detected (e.g., Cohere, Gemini, or a self-hosted model), those calls appear as regular HTTP spans and are not affected by the rule above.

To include custom providers in the drop rule, add a rule matching their hostname:

{% code title="Traces Pipeline YAML" %}

```yaml
ottlRules:
  - ruleName: gc-genai-custom-providers
    conditions:
      - 'attributes["http.host"] == "api.cohere.com"'
      - 'attributes["net.peer.name"] == "api.cohere.com"'
      - 'attributes["http.host"] == "llm.internal.yourcompany.com"'
      - 'attributes["net.peer.name"] == "llm.internal.yourcompany.com"'
    statements:
      - 'set(drop, true)'
    conditionLogicOperator: or
```

{% endcode %}

Replace the hostnames with your provider endpoints. Each hostname needs both `http.host` and `net.peer.name` conditions because different instrumentation sources use different attribute names for the same host.

***

### Metadata-Only Mode

To keep performance metadata (tokens, cost, latency, model) while stripping prompts and responses from storage, use two layers:

#### SDK Spans

Most OTel GenAI instrumentation libraries support disabling content capture at the source. The standard environment variable is:

```bash
OTEL_INSTRUMENTATION_GENAI_CAPTURE_MESSAGE_CONTENT=false
```

Check your SDK's documentation — some libraries use different configuration keys. When content capture is off, spans arrive with all performance metadata intact and no prompt or response text.

#### eBPF Spans

Add this pipeline rule to clear content fields from eBPF-captured GenAI spans. It covers Era 3 attributes (`gen_ai.input.messages`, etc.), Era 1 attributes (`gen_ai.prompt`, `gen_ai.completion`), and the indexed variants emitted by Traceloop/OpenLLMetry SDKs.

{% code title="Traces Pipeline YAML" %}

```yaml
ottlRules:
  - ruleName: gc-genai-reduce-content
    conditions:
      - 'protocol_type == "gen_ai"'
    statements:
      - 'delete_key(attributes, "gen_ai.input.messages")'
      - 'delete_key(attributes, "gen_ai.output.messages")'
      - 'delete_key(attributes, "gen_ai.system_instructions")'
      - 'delete_key(attributes, "gen_ai.tool.definitions")'
      - 'delete_key(attributes, "gen_ai.prompt")'
      - 'delete_key(attributes, "gen_ai.completion")'
      - 'delete_matching_keys(attributes, "^gen_ai\\.(prompt|completion)\\.[0-9]+\\.")'
      - 'set(request_body, "")'
      - 'set(response_body, "")'
    statementsErrorMode: ignore
```

{% endcode %}

If you already have an `ottlRules:` block in your pipeline YAML, add only the `- ruleName:` entry to the existing list — do not add a second `ottlRules:` key.

{% hint style="warning" %}
**Scope of pipeline-side stripping:** This rule covers content fields for auto-detected providers across all SDK convention versions. It does not guarantee zero content in storage — tool call arguments, error messages that echo input, and spans from providers not yet auto-detected are not covered. For a hard guarantee, use the [kill switch](#disable-genai-data-collection) above. For SDK spans, disabling content capture at the source is the definitive control.
{% endhint %}

For more surgical control — replacing specific JSON keys within request/response bodies with a placeholder rather than stripping entire fields — see [Obfuscate Traces](https://docs.groundcover.com/use-groundcover/data-pipelines/traces-pipeline/obfuscate-traces).

{% hint style="info" %}
Pipeline rules apply to **new spans only**. Existing data in storage is not affected.
{% endhint %}


# Troubleshooting

## I'm seeing duplicate spans

This is expected when both eBPF and SDK instrumentation are active for the same service. Each source captures the same LLM call independently — you'll see one span from eBPF and one from the SDK. Both are correct; cost is not double-counted. When groundcover detects an SDK span for the same call, eBPF cost and token data is excluded from aggregations.

***

## I don't see AI data

Work through the path you're using.

### AI coding-tool integrations (Claude Code, Claude Cowork, Codex)

AI Observability shows GenAI **traces** — spans with `gen_ai.*` attributes. The Claude Code, Claude Cowork, and Codex data-source integrations ship tool-usage telemetry (logs, plus metrics for Claude Code), not GenAI traces, so they don't appear here. See [AI Tools Observability](/integrations/data-sources/ai-tools-observability) for where to find the data.

### eBPF Path

* **Unsupported provider.** eBPF auto-detects OpenAI, Anthropic, and AWS Bedrock only. Calls to Vertex AI, Azure OpenAI, Cohere, or a self-hosted model appear as regular HTTP spans, not in AI Observability. Check [Supported Providers](/capabilities/ai-observability#supported-providers) for the current list. If your provider isn't listed, add [SDK instrumentation](/use-groundcover/ai-observability#sdk-instrumentation) instead, or [contact us](https://www.groundcover.com/contact) to request eBPF support for your provider.
* **API gateway or proxy in the path.** If your service calls a cloud-hosted gateway (OpenRouter, Azure OpenAI Service, LiteLLM Cloud) instead of OpenAI, Anthropic, or AWS Bedrock directly, eBPF doesn't recognize the hostname and the call appears as regular HTTP. Self-hosted proxies (LiteLLM, Portkey running in your cluster) are not affected — eBPF sees both hops. To get GenAI visibility through an unsupported gateway, add SDK instrumentation to your calling service. [Contact us](https://www.groundcover.com/contact) to request native support for your gateway.
* **Sensor version.** eBPF provider support ships in sensor releases. Check that you're running a current version — [Installation & Updating](/getting-started/installation-and-updating).

### SDK Path

* **Missing required attributes.** Your SDK must emit `gen_ai.operation.name` for spans to appear in AI Observability. Valid values: `chat`, `text_completion`, `embeddings`, `invoke_agent`, `execute_tool`, `create_agent`. See [required attributes](/use-groundcover/ai-observability#sdk-full-visibility) and use the [**Data Sources**](https://app.groundcover.com/data-sources) page for setup instructions specific to your SDK.
* **OTel exporter not configured.** Confirm your service's OTLP exporter points to groundcover's collector endpoint. The [**Data Sources**](https://app.groundcover.com/data-sources) page has your endpoint and setup steps.
* **SDK not listed.** Check the [**Data Sources**](https://app.groundcover.com/data-sources) page for supported SDKs. If yours isn't listed, [contact us](https://www.groundcover.com/contact) — we're actively expanding SDK support.

***

## Spans Disappeared

Check whether a kill-switch rule is active in [Traces Pipeline Settings](https://app.groundcover.com/settings/traces-pipeline). A `set(drop, true)` rule on `protocol_type == "gen_ai"` drops all GenAI spans before storage. See [Disable GenAI Data Collection](/use-groundcover/ai-observability/privacy-controls#disable-genai-data-collection).

***

## No Content in Spans

* **eBPF:** Content is captured automatically. If it's missing, check for a content-reduction or kill-switch rule in [Traces Pipeline Settings](https://app.groundcover.com/settings/traces-pipeline).
* **SDK:** Check whether `OTEL_INSTRUMENTATION_GENAI_CAPTURE_MESSAGE_CONTENT=false` is set in your environment. With this flag off, spans arrive with all performance metadata but no content. Remove it or set it to `true` to restore content capture. See [Metadata-Only Mode](/use-groundcover/ai-observability/privacy-controls#metadata-only-mode) for details.

***

## Cost Shows Unpriced

The `gc.llm_cost.status` attribute shows `unpriced` when groundcover doesn't have pricing for the model in its pricing table. This typically happens with newly released models, fine-tuned models, or custom model names. Cost will populate automatically once the pricing table is updated. See [groundcover Enrichment](/use-groundcover/ai-observability/attributes#groundcover-enrichment) for details.


# Workflows

{% hint style="warning" %}
**Workflows are deprecated.** We recommend migrating to [Notification Routes](/use-groundcover/monitors/notification-routes) and [Destinations](/integrations/connected-apps) (formerly Connected Apps) for a more reliable and easier-to-configure notification system. New features and bug fixes are only being applied to the new system.
{% endhint %}

## Migrating from Workflows to Notification Routes

The new notification system (Destinations + Notification Routes) offers several advantages:

* **Simpler configuration** - No YAML required, use the UI wizard
* **Better reliability** - Improved notification delivery and link handling
* **Automatic label passing** - Labels are automatically included without complex templating
* **Testing capabilities** - Test notifications before enabling on production monitors
* **Terraform support** - Fully supported via Terraform provider

### Migration Steps

1. **Create Destinations** (Integrations → Destinations)
   * Set up your Slack webhook, PagerDuty, OpsGenie, or other destinations
   * Admin permissions required for creating Destinations
2. **Create Notification Routes** (Monitors → Notification Routes)
   * Define scope using gcQL (e.g., `env:prod AND severity:S1`)
   * Add rules for Firing and/or Resolved states
   * Select your Destinations
3. **Disable old Workflows**
   * To disable a workflow without deleting it, add a filter that matches nothing (e.g., `key: _never_match_`)

### Template Variable Changes

The new system uses a simplified variable syntax:

| Old Workflows Syntax                   | New Destinations Syntax |
| -------------------------------------- | ----------------------- |
| `{{ alert.labels.workload }}`          | `{{ labels.workload }}` |
| `{{ alert.annotations._gc_severity }}` | `{{ severity }}`        |
| `{{ alert.alertname }}`                | `{{ alertname }}`       |
| `{{ alert.fingerprint }}`              | `{{ fingerprint }}`     |

{% hint style="info" %}
**Note:** The old `alert.labels.*` format is still supported in Destinations, so you don't need to update your existing monitor templates when migrating. However, the new simplified syntax is recommended for new configurations.
{% endhint %}

### Common Migration Issues

**Issue links not working**: Old workflows generate links that may not resolve correctly. Add `?resolve_timerange=true` to your workflow URL templates, or migrate to Notification Routes which handle this automatically.

**Labels not appearing**: Ensure you're using the correct variable syntax with the `labels.` prefix.

***

## Legacy Workflows Documentation

Workflows are YAML-based configurations that are executed whenever a monitor is triggered. They enable you to integrate with third-party systems, apply custom logic to format and transform data, and set different conditions to handle your monitoring alerts intelligently.

## Workflow components

### Triggers

Triggers apply filtering conditions that determine whether a specific workflow is executed. In groundcover, the trigger type is always set to "alert".

**Example**: This trigger ensures that only monitors fired with telemetry data from the Prod environment will actually execute the workflow. Note that the "env" attribute needs to be provided as a context label from the monitor:

```yaml
triggers:
  - type: alert
    filters:
    - key: env
      value: prod
```

> **Note**: Workflows are "pull based" which means they will try to match monitors even when these monitors did not explicitly add a specific workflow. Therefore, the filters need to accurately define the condition to be used for a monitor.

### Consts

Consts is a section where you can declare predefined attributes based on data provided with the monitor context. A set of functions is available to transform existing data and format it for propagation to third-party integrations. Consts simplify access to data that is needed in the actions section.

**Example**: The following example shows how to map a predefined set of severity values to the monitor severity as defined in groundcover - here, any potential severity in groundcover is translated into one of P1-P5 values.

The function `keep.dictget` gets a value from a map (dictionary) using a specific key. In case the key is not found - P3 will be the default severity:

```yaml
consts:
    severities: '{"S1": "P1","S2": "P2","S3": "P3","S4": "P4","critical": "P1","error": "P2","warning": "P3","info": "P4"}'
    severity: keep.dictget({{ consts.severities }}, {{ alert.annotations._gc_severity }}, "P3")
```

### Actions

Actions specify what happens when a workflow is triggered. Actions typically interface with external systems (like sending a Slack message). Actions can be an array of actions, they can be executed conditionally and include the integration in their config part as well as a payload block which is typically dependent on the exact integration used for the notification.

Actions include:

1. **Provider part (provider:)** - Configures the integration to be used
2. **Payload part (with:)** - Contains the data to submit to the integration based on its actual structure

**Example**: In this example you can see a typical Slack notification. Note that the actual integration is referenced through the 'providers' context attribute. The integration name is the exact string used to [create the integration](/integrations/workflow-integrations) (in this case "groundcover-alerts").

```yaml
actions:
- name: slack-action-firing
  provider:
    config: '{{ providers.groundcover-alerts }}'
    type: slack
    with:
      attachments:
      - color: '{{ consts.red_color }}'
        footer: '{{ consts.footer_url }}'
        text: '{{ consts.slack_message }}'
        title: 'Firing: {{ alert.alertname }}'
        type: plain_text
      message: ' '
```


# Create a new Workflow

{% hint style="warning" %}
Workflows are getting an upgrade: meet [Notification Routes](/use-groundcover/monitors/notification-routes)
{% endhint %}

### Creation

Creating new workflows is currently supported through the groundcover app in two ways from the Monitors menu:

<div align="left"><figure><img src="/files/EEYr6KUkExZNhyBDr0XX" alt="Create Workflow" width="800"><figcaption></figcaption></figure></div>

#### 1. "Create Notification Workflow" button - The quick way

This provides a guided approach to create a workflow. When creating a notification workflow, you will be asked to give your workflow a name, add filters, and select the specific integration to use.

**Filter Rules By Labels** - Add key-value attributes to ensure your workflow executes under specific conditions only - for example, `env = prod` only.

**Delivery Destinations** - Select one or more integrations to be used for notifications with this workflow.

**Scope - When The Workflow Will Run** - This setting allows you to limit this workflow execution only to monitors that explicitly select to route their triggers to this workflow, as opposed to "Handle all issues" that catches triggers from any monitor.

Once you create a workflow using this option, you can later edit the workflow to apply any configuration or logic by using the editor option (see next).

#### 2. "Create Workflow" button

Clicking the button will open up a text editor where you can add your workflow definition in YAML format by applying any valid configuration, logic, and functionality.

> **Note**: Make sure to create your integration prior to creating the workflow as it requires using an existing integration.

### View

Upon successful workflow creation it will be active immediately, and a new workflow record will appear in the underlying table.

For each existing workflow, we can see the following fields:

* **Name**: Your defined workflow name
* **Description**: If you've added a description of the workflow
* **Creator**: Workflow creator email
* **Creation Date**: Workflow creation date
* **Last Execution Time**: Timestamp of last workflow execution (depends on workflow trigger type)
* **Last Execution Status**: Last execution status (failure or success)

### Editing

From the right side of each workflow record in the display, you can access the menu (three dots) and click "Edit Workflow". This will open the editor so you can modify the YAML to conform to the available functionality. See examples below.


# Workflow Examples

{% hint style="warning" %}
Workflows are getting an upgrade: meet [Notification Routes](/use-groundcover/monitors/notification-routes)
{% endhint %}

This page provides practical examples of workflows for different use cases and integrations.

## Triggers Examples

### Filter by Monitor Name

This example shows how to create a workflow that only triggers for a specific monitor (by its name):

```yaml
workflow: 
  id: specific-monitor-workflow
  description: Workflow triggered only by Workload Pods Crashed Monitor
  triggers:
    - type: alert
      filters:
        - key: alertname
          value: Workload Pods Crashed Monitor
```

### Filter by Environment

Execute only on the Prod environment. The "env" attribute needs to be part of the monitor context attributes (either by using it in the group by section or by explicitly adding it as a context label):

```yaml
workflow: 
  id: prod-only-workflow
  description: Workflow triggered only by production environment alerts
  triggers:
    - type: alert
      filters:
        - key: env
          value: prod
```

### Filter by Multiple Conditions

This example shows how to combine multiple filters. In this case it will match events from the prod environment and also monitors that explicitly routed the workflow with the name "actual-name-of-workflow":

```yaml
workflow: 
  id: multi-filter-workflow
  description: Workflow triggered by critical alerts in production
  triggers:
    - type: alert
      filters:
        - key: env
          value: prod
        - key: annotations.multi-filter-workflow
          value: enabled
```

### Filter by Regex

In this case we will use a regular expression to filter on events coming from the groundcover OR monitoring namespaces. Note that any regular expression can be used:

```yaml
workflow: 
  id: regex-filter-workflow
  description: Workflow triggered by alerts from groundcover or monitoring namespaces
  triggers:
    - type: alert
      filters:
        - key: namespace
          value: r"(groundcover|monitoring)"
        - key: annotations.regex-filter-workflow
          value: enabled          
```

## Consts Examples

The consts section is the best location to create pre-defined attributes and apply different transformations on the monitor's metadata for formatting the notification messaging.

### Map Severities

Severities in your notified destination may not match the groundcover predefined severities. By using a dictionary, you can map any groundcover severity value to another, and extract it by using the actual monitor severity. Use the "keep.dictget" function to extract from a dictionary and apply a default in case the value is missing.

<pre class="language-yaml"><code class="lang-yaml">workflow:
  id: severity-mapping-example
  description: Example of mapping severities using consts
  triggers:
    - type: alert
<strong>      filters:
</strong>        - key: annotations.severity-mapping-example
          value: enabled            
  consts:
    severities: '{"S1": "P1","S2": "P2","S3": "P3","S4": "P4","critical": "P1","error": "P2","warning": "P3","info": "P4"}'
    severity: keep.dictget({{ consts.severities }}, {{ alert.annotations._gc_severity }}, "P3")
</code></pre>

### Best Practice for Accessing Monitor Labels

When accessing a context label via `alert.labels`, if this label is not transferred within the monitor - the workflow might crash. Best practice to pre-define labels is to declare them in the consts section with a default value, using "keep.dictget" so the value is gracefully pulled from the labels object.

```yaml
workflow:
  id: labels-best-practice-example
  description: Example of safely accessing monitor labels
  triggers:
    - type: alert
      filters:
        - key: annotations.labels-best-practice-example
          value: enabled            
  consts:
    region: keep.dictget({{ alert.labels }}, "cloud.region", "")
```

> **Note**: Label names that are dotted, like "cloud.region" in this example, cannot be referenced in the monitor itself and can only be retrieved using this technique of pulling the value with "keep.dictget" from the alert.labels object.

### Additional Useful Functions

* **`keep.dict_pop({{alert.labels}}, "_gc_monitor_id", "_gc_monitor_name", "_gc_severity", "backend_id", "grafana_folder", "_gc_issue_header")`** - "Clean" a key-value dictionary from some irrelevant values (keys). In this case, the groundcover labels dictionary has some internal keys that you might not want to include in your notification content.
* **`keep.join(["a", "b", "c"], ",")`** - Joins a list of elements into a string using a given delimiter. In this case the output is "a,b,c".

## Action Examples

### Conditional Statements

Use "if" condition to apply logic on different actions.

Create a separate block for a firing monitor (a resolved monitor can use different logic to change formatting of the notification):

```yaml
workflow:
  id: conditional-actions-example
  description: Example of conditional actions based on alert status
  triggers:
    - type: alert
      filters:
        - key: annotations.conditional-actions-example
          value: enabled                
  actions:
    - if: '{{ alert.status }} == "firing"'
      name: slack-action-firing
      provider:
        config: '{{ providers.groundcover-alerts-dev }}'
        type: slack
        with:
          attachments:
          - color: '{{ consts.red_color }}'
            footer: '{{ consts.footer_url }}'
            footer_icon: '{{ consts.footer_icon }}'
            text: '{{ consts.slack_message }}'
            title: 'Firing: {{ alert.alertname }}'
            title_link: '{{ consts.title_link }}'
            ts: keep.utcnowtimestamp()
            type: plain_text
          message: ' '
```

"If" statements can include and/or logic for multiple conditions:

```yaml
workflow:
  id: multi-condition-actions-example
  description: Example of multiple conditions in actions
  triggers:
    - type: alert
      filters:
        - key: annotations.multi-condition-actions-example
          value: enabled                    
  actions:
    - if: '{{ alert.status }} == "firing" and {{ alert.labels.namespace }} == "namespace1"'
      name: slack-action-firing
      provider:
        config: '{{ providers.groundcover-alerts-dev }}'
        type: slack
        with:
          attachments:
          - color: '{{ consts.red_color }}'
            footer: '{{ consts.footer_url }}'
            footer_icon: '{{ consts.footer_icon }}'
            text: '{{ consts.slack_message }}'
            title: 'Firing: {{ alert.alertname }}'
            title_link: '{{ consts.title_link }}'
            ts: keep.utcnowtimestamp()
            type: plain_text
          message: ' '
```

### Notification by Specific Hours

Use the function `keep.is_business_hours` combined with an "if" statement to trigger an action within specific hours only.

In this example the action block will execute on Sundays (6) between 20-23 (8pm to 11pm) or on Mondays (0) between 0-1am:

```yaml
workflow:
  id: time-based-notification-example
  description: Example of time-based conditional actions
  triggers:
    - type: alert
      filters:
        - key: annotations.time-based-notification-example
          value: enabled                    
  actions:
    - if: '({{ alert.status }} == "firing" and (keep.is_business_hours(timezone="America/New_York", business_days=[6], start_hour=20, end_hour=23) or keep.is_business_hours(timezone="America/New_York", business_days=[0], start_hour=0, end_hour=1)))'
      name: time-based-notification
      provider:
        type: slack
        config: '{{ providers.slack_webhook }}'
        with:
          message: "Time-sensitive alert: {{ alert.alertname }}"
```


# Integration Examples with Workflows

{% hint style="warning" %}
Workflows are getting an upgrade: meet [Notification Routes](/use-groundcover/monitors/notification-routes)
{% endhint %}

This page provides examples of how to integrate workflows with different notification systems and external services.

## Slack Notification

This workflow sends a simple Slack message when triggered:

```yaml
workflow: 
  id: slack-notification
  description: Send Slack notification for alerts
  triggers:
    - type: alert
      filters:
        - key: annotations.slack-notification
          value: enabled                
  actions:
    - name: slack-notification
      provider:
        type: slack
        config: '{{ providers.slack_webhook }}'
        with:
          message: "Alert: {{ alert.alertname }} - Status: {{ alert.status }}"
```

## Slack with Rich Formatting

This workflow sends a formatted Slack message using Block Kit:

```yaml
workflow: 
  id: slack-rich-notification
  description: Send formatted Slack notification
  triggers:
    - type: alert
      filters:
        - key: annotations.slack-rich-notification
          value: enabled                
  actions:
    - name: slack-rich-notification
      provider:
        type: slack
        config: '{{ providers.slack_webhook }}'
        with:
          blocks:
          - type: header
            text:
              type: plain_text
              text: ':rotating_light: {{ alert.alertname }} :rotating_light:'
              emoji: true
          - type: divider
          - type: section
            fields:
            - type: mrkdwn
              text: |-
                *Cluster:*
                {{ alert.labels.cluster}}
            - type: mrkdwn
              text: |-
                *Namespace:*
                {{ alert.labels.namespace}}
            - type: mrkdwn
              text: |-
                *Status:*
                {{ alert.status}}
```

## PagerDuty Integration

This workflow creates a PagerDuty incident:

```yaml
 workflow:
  id: pagerduty-incident-workflow
  description: Create PagerDuty incident for alerts
  name: pagerduty-incident-workflow
  triggers:
    - type: alert
      filters:
        - key: annotations.pagerduty-incident-workflow
          value: enabled
  consts:
    severities: '{"S1": "critical","S2": "error","S3": "warning","S4": "info","critical": "critical","error": "error","warning": "warning","info": "info"}'
    severity: keep.dictget( '{{ consts.severities }}', '{{ alert.annotations._gc_severity }}', 'info')
    description: keep.dictget( {{ alert.annotations }}, "_gc_description", "")
    title: keep.dictget( {{ alert.annotations }}, "_gc_issue_header", '{{ alert.alertname }}')
    redacted_labels: keep.dict_pop({{ alert.labels }}, "_gc_monitor_id", "_gc_monitor_name", "_gc_severity", "backend_id", "grafana_folder")
    env: keep.dictget( {{ alert.labels }}, "env", "- no env -")
    namespace: keep.dictget( {{ alert.labels }}, "namespace", "- no namespace -")
    workload: keep.dictget( {{ alert.labels }}, "workload", "- no workload -")
    pod: keep.dictget( {{ alert.labels }}, "podName", "- no pod -")
    issue: https://app.groundcover.com/monitors/issues?backendId={{ alert.labels.backend_id }}&selectedObjectId={{ alert.fingerprint }}
    monitor: https://app.groundcover.com/monitors?backendId={{ alert.labels.backend_id }}&selectedObjectId={{ alert.labels._gc_monitor_id }}
    silence: https://app.groundcover.com/monitors/create-silence?keep.replace(keep.join({{ consts.redacted_labels }}, "&", "matcher_"), " ", "+")
  actions:
  - name: pagerduty-alert
    provider:
      config: '{{ providers.pagerduty-integration-name }}'
      type: pagerduty
      with:
        title: '{{ consts.title }}'
        severity: '{{ consts.severity }}'
        dedup_key: '{{alert.fingerprint}}'
        custom_details:
          01_environment: '{{ consts.env }}'
          02_namespace: '{{ consts.namespace }}'
          03_service_name: '{{ consts.workload }}'
          04_pod: '{{ consts.pod }}'
          05_labels: '{{ consts.redacted_labels }}'
          06_monitor: '{{ consts.monitor }}'
          07_issue: '{{ consts.issue }}'
          08_silence: '{{ consts.silence }}'

```

## Opsgenie Integration

This workflow creates an Opsgenie alert:

* Alias is used to group identical events together in Opsgenie (`alias` key in the payload)
* Severities must be mapped to Opsgenie valid severities (`priority` key in the payload)
* Tags are a list of string values (`tags` key in the payload)

```yaml
workflow:
  id: Opsgenie Example
  description: "Opsgenie workflow"
  triggers:
    - type: alert
      filters:
        - key: annotations.Opsgenie Example
          value: enabled
  consts:
    description: keep.dictget( {{ alert.annotations }}, "_gc_description", "")
    redacted_labels: keep.dict_pop({{ alert.labels }}, "_gc_monitor_id", "_gc_monitor_name", "_gc_severity", "backend_id", "grafana_folder", "CampaignName")
    severities: '{"S1": "P1","S2": "P2","S3": "P3","S4": "P4","critical": "P1","error": "P2","warning": "P3","info": "P4"}'
    severity: keep.dictget({{ consts.severities }}, {{ alert.annotations._gc_severity }}, "P3")
    title: keep.dictget( {{ alert.annotations }}, "_gc_issue_header", '{{ alert.alertname }}')
    region: keep.dictget( {{ alert.labels }}, "cloud.region", "")
    TenantID: keep.dictget( {{ alert.labels }}, "tenantID", "")
  name: Opsgenie Example
  actions:
  - if: '{{ alert.status }} == "firing"'
    name: opesgenie-alert
    provider:
      config: "{{ providers.Opsgenie }}"
      type: opsgenie
      with:
        alias: '{{ alert.fingerprint }}'
        description: '{{ consts.description }}'
        details: '{{ consts.redacted_labels }}'
        message: '{{ consts.title }}'
        priority: '{{ consts.severity }}'
        source: groundcover
        tags:
        - '{{ alert.alertname }}'
        - '{{ consts.TenantID }}'
        - '{{ consts.region }}'
  
```

## Jira Ticket Creation

This workflow creates a Jira ticket using webhook integration:

```yaml
workflow:
  id: jira-ticket-creation
  description: Create Jira ticket for alerts
  triggers:
    - type: alert
      filters:
        - key: annotations.jira-ticket-creation
          value: enabled    
  consts:
    description: keep.dictget({{ alert.annotations }}, "_gc_description", '')
    title: keep.dictget({{ alert.annotations }}, "_gc_issue_header", "{{ alert.alertname }}")
  actions:
    - name: jira-ticket
      provider:
        type: webhook
        config: '{{ providers.jira_webhook }}'
        with:
          body:
            fields:
              description: '{{ consts.description }}'
              issuetype:
                id: 10001
              project:
                id: 10000
              summary: '{{ consts.title }}'
```

## Multiple Actions

This workflow performs multiple actions for the same alert:

```yaml
workflow:
  id: multi-action-workflow
  description: Perform multiple actions for critical alerts
  triggers:
    - type: alert
      filters:
        - key: annotations.multi-action-workflow
          value: enabled      
        - key: severity
          value: critical
  actions:
    - name: slack-notification
      provider:
        type: slack
        config: '{{ providers.slack_webhook }}'
        with:
          message: "Critical alert: {{ alert.alertname }}"
    - name: pagerduty-incident
      provider:
        type: pagerduty
        config: '{{ providers.pagerduty_prod }}'
        with:
          title: "Critical: {{ alert.alertname }}"
    - name: jira-ticket
      provider:
        type: webhook
        config: '{{ providers.jira_webhook }}'
        with:
          body:
            fields:
              summary: "Critical Alert: {{ alert.alertname }}"
              description: "Critical alert triggered in {{ alert.labels.namespace }}"
              issuetype:
                id: 10001
              project:
                id: 10000
```


# Full Webhook Examples

{% hint style="warning" %}
Webhook Notification Channel is getting an upgrade: meet [Webhook Destination](/integrations/connected-apps/generic-webhook)
{% endhint %}

This section contains comprehensive examples of webhook integrations with various third-party services. These examples provide step-by-step instructions for setting up complete workflows with external systems.

## Available Examples

* [**incident.io**](/use-groundcover/workflows/full-webhook-examples/incident.io) - Integrate with incident.io for incident management
* [**MS Teams**](/use-groundcover/workflows/full-webhook-examples/ms-teams) - Send notifications to Microsoft Teams channels
* [**Email via Zapier**](broken://pages/AnUw8JjLjI9n2Uw0k24k) - Route alerts to email using Zapier
* [**Slack App with Bot Tokens**](/use-groundcover/workflows/full-webhook-examples/slack-app-for-channel-routing) - Route alerts to different slack channels with a single Webhook

Each example includes:

* Prerequisites and setup requirements
* Step-by-step configuration instructions
* Complete workflow YAML configurations
* Integration-specific considerations and best practices

These examples demonstrate advanced webhook usage patterns and can serve as templates for other webhook integrations.


# incident.io

{% hint style="warning" %}
incident.io Notification Channel is getting an upgrade: meet [incident.io Destination](/integrations/connected-apps/incident.io)
{% endhint %}

To integrate groundcover with [incident.io](https://incident.io), follow the steps below. Note that you’ll need a **Pro** incident.io account to view your incoming alerts.

1. **Generate an Alerts configuration for groundcover**\
   Log in to your [incident.io](https://incident.io) account. Go to "On-call" -> "Alerts" -> "Configure" and add a new source.
2. **On the "Create Alert Source" screen** the answer to the question "Where do your alerts come from?" should be "HTTP". Select this source and give it a unique name. Hit "continue".
3. **incident.io will create your configuration now** from which you will need to copy the following items for the Webhook integration\\

   <figure><img src="/files/GUGzVC6KcLZLyPsQq4dH" alt=""><figcaption></figcaption></figure>
4. **Set Up the Webhook in groundcover**
   * Head out to the integrations section: Settings -> Integrations, to create a new [Webhook](https://docs.groundcover.com/integrations/workflow-integrations/webhook-integration)
   * Start by giving your Webhook integration a name. This name will be used below in the provider block sample .
   * Set the `Webhook URL` to the url you copied from field (1)
   * Keep the HTTP method as `POST`
   * Under headers add `Authorization`, and paste the "Bearer \<token>" copied from field (2).
5. **Create a Workflow**\
   Go to `Monitors --> Workflows --> Create Workflow`, and paste the YAML configuration provided below.\
   **Note:** The `body` section is a dictionary of keys that will be sent as a JSON payload to the incident.io platform
6. **Configure the `provider` Block**\
   In the `provider` block, replace `{{ providers.your-incident-io-integration-name }}` with your actual Webhook integration name (the one you created in step 4)\
   For example, if you named your integration `test-incidentio`, the config reference would be: `{{ providers.test-incidentio }}`
7. **Required Parameters for Creating an alert**\
   When triggering an alert, the following keys are required:
   1. `title` - Alert title that can be pulled from groundcover as seen in the example
   2. `status` - One of "firing" or "resolved" that can also be pulled from groundcover as the example shows.
8. You can include additional parameters for richer context (optional):
   1. `description`
   2. `deduplication_key` - unique attribute used to group identical alerts, groundcover provides this through the fingerprint attribute
   3. `metadata` - Any additional metadata that you've configured within your monitor in groundcover. Note that these set should actively reflect your monitor definition in groundcover

{% hint style="info" %}
The attributes shown in the yaml block of the metadata section below are an example only! Alert labels can only be attributes used in the group by section of the actual monitor
{% endhint %}

**Example code for your groundcover workflow:**

```yaml
workflow:
  id: incident-io-alerts-workflow
  name: incident-io-alerts-workflow
  description: Sends an API to incident.io alerts endpoint
  triggers:
  - type: alert
    filters:
      - key: annotations.incident-io-alerts-workflow
        value: enabled      
  consts:
    description: keep.dictget( {{ alert.annotations }}, "_gc_description", '')
    issue: https://app.groundcover.com/monitors/issues?backendId={{ alert.labels.backend_id }}&selectedObjectId={{ alert.fingerprint }}
    monitor: https://app.groundcover.com/monitors?backendId={{ alert.labels.backend_id }}&selectedObjectId={{ alert.labels._gc_monitor_id }}
    redacted_labels: keep.dict_pop({{alert.labels}}, "_gc_monitor_id", "_gc_monitor_name", "_gc_severity", "backend_id", "grafana_folder", "_gc_issue_header")
    silence: https://app.groundcover.com/monitors/create-silence?keep.replace(keep.join({{ consts.redacted_labels }}, "&", "matcher_"), " ", "+")
    title: keep.dictget( {{ alert.annotations }}, "_gc_issue_header", "{{ alert.alertname }}")
    cluster: keep.dictget( {{ alert.labels }}, "cluster", "[no-cluster]")
    namespace: keep.dictget( {{ alert.labels }}, "namespace", "[no-namespace]")
    workload: keep.dictget( {{ alert.labels }}, "workload", "[no-workload]")
  actions:
  - name: webhook
    provider:
      config: ' {{ providers.your-incident-io-integration-name }} '
      type: webhook
      with:
        body:
          title: '{{ alert.alertname }}'
          description: '{{ alert.description }}'
          deduplication_key: '{{ alert.fingerprint }}'
          status: '{{ alert.status }}'
          # To use metadata attributes that refer to alert.labels, the attributes 
          # must be used in the group by section of the monitor - the example below
          # assumes that cluster, namespace and workload were used for group by
          metadata:
            cluster: '{{ consts.cluster }}'
            namespace: '{{ consts.namespace }}'
            service: '{{ consts.workload }}'
            severity: '{{ alert.annotations._gc_severity }}'
```


# MS Teams

{% hint style="warning" %}
Webhook Notification Channel is getting an upgrade: meet [Webhook Destination](/integrations/connected-apps/generic-webhook)
{% endhint %}

To integrate groundcover with MS Teams, follow the steps below. Note that you’ll need at least a **Business** subscription of MS Teams to be able to create workflows.

1. **Create a webhook workflow for your dedicated Teams channel**\
   Go to Relevant Team -> Specific Channel -> "Workflows", and create a webhook workflow
2. **The webhook workflow is associated a URL** which is used to trigger the MS Teams integration on groundcover - make sure to copy this URL
3. **Set Up the Webhook in groundcover**
   * Head out to the integrations section: Settings -> Integrations, to create a new [Webhook](https://docs.groundcover.com/integrations/workflow-integrations/webhook-integration)
   * Start by giving your Webhook integration a name. This name will be used below in the provider block sample .
   * Set the `Webhook URL` to the url you copied from field (2)
   * Keep the HTTP method as `POST`
4. **Create a Workflow**\
   Go to `Monitors --> Workflows --> Create Workflow`, and paste the YAML configuration provided below.
5. **Configure the `provider` Blocks (There are two of them)**\
   In the `provider` block, replace `{{ providers.your-teams-integration-name }}` with your actual Webhook integration name (the one you created in step 3)\
   For example, if you named your integration `test-ms-teams`, the config reference would be: `{{ providers.test-ms-teams }}`

{% hint style="info" %}
The following example shows a pre-configured MS Teams workflow template. You can easily modify workflows to support different formats based on the MS Teams workflow schema.
{% endhint %}

**Sample code for your groundcover workflow:**

```yaml
workflow:
  id: ms-teams-alerts-workflow
  description: Sends an API to MS Teams alerts endpoint
  name: ms-teams-alerts-workflow
  triggers:
  - type: alert
    filters:
    - key: annotations.ms-teams-alerts-workflow
      value: enabled
  consts:
    silence_link: 'https://app.groundcover.com/monitors/create-silence?keep.replace(keep.join(keep.dict_pop({{ alert.labels }}, "_gc_monitor_id", "_gc_monitor_name", "_gc_severity", "backend_id", "grafana_folder"), "&", "matcher_"), " ", "+")'
    monitor_link: 'https://app.groundcover.com/monitors?backendId={{ alert.labels.backend_id }}&selectedObjectId={{ alert.labels._gc_monitor_id }}'
    title_link: 'https://app.groundcover.com/monitors/issues?backendId={{ alert.labels.backend_id }}&selectedObjectId={{ alert.fingerprint }}'
    description: keep.dictget( {{ alert.annotations }}, "_gc_description", '')
    redacted_labels: keep.join(keep.dict_pop({{alert.labels}}, "_gc_monitor_id", "_gc_monitor_name", "_gc_severity", "backend_id", "grafana_folder", "_gc_issue_header"), "-\n")
    title: keep.dictget( {{ alert.annotations }}, "_gc_issue_header", "{{ alert.alertname }}")

  actions:
  - if: '{{ alert.status }} == "firing"'
    name: teams-webhook-firing
    provider:
      config: ' {{ providers.your-teams-integration-name }} '
      type: webhook
      with:
        body:
          type: message
          attachments:
          - contentType: application/vnd.microsoft.card.adaptive
            content:
              $schema: http://adaptivecards.io/schemas/adaptive-card.json
              type: AdaptiveCard
              version: "1.2"
              body:
              - type: TextBlock
                text: "\U0001F6A8 Firing: {{ consts.title }}"
                weight: bolder
                size: large
              - type: TextBlock
                text: "[Investigate Issue]({{consts.title_link}})"
                wrap: true
              - type: TextBlock
                text: "{{ consts.description }}"
                wrap: true
              - type: TextBlock
                text: "[Silence]({{consts.silence_link}})"
                wrap: true
              - type: TextBlock
                text: "[See monitor]({{consts.monitor_link}})"
                wrap: true
              - type: TextBlock
                text: "{{ consts.redacted_labels }}"
                wrap: true
  - if: '{{ alert.status }} != "firing"'
    name: teams-webhook-resolved
    provider:
      config: ' {{ providers.your-teams-integration-name }} '
      type: webhook
      with:
        body:
          type: message
          attachments:
          - contentType: application/vnd.microsoft.card.adaptive
            content:
              $schema: http://adaptivecards.io/schemas/adaptive-card.json
              type: AdaptiveCard
              version: "1.2"
              body:
              - type: TextBlock
                text: "\U0001F7E2 Resolved: {{ consts.title }}"
                weight: bolder
                size: large
              - type: TextBlock
                text: "[Investigate Issue]({{consts.title_link}})"
                wrap: true
              - type: TextBlock
                text: "{{ consts.description }}"
                wrap: true
              - type: TextBlock
                text: "[Silence]({{consts.silence_link}})"
                wrap: true
              - type: TextBlock
                text: "[See monitor]({{consts.monitor_link}})"
                wrap: true
              - type: TextBlock
                text: "{{ consts.redacted_labels }}"
                wrap: true
  
```


# Slack App for Channel Routing

{% hint style="info" %}
Webhook Notification Channel is getting an upgrade: meet [Webhook Destination](/integrations/connected-apps/generic-webhook)
{% endhint %}

groundcover supports sending notifications to Slack using a **Slack App with bot tokens** instead of static webhooks. This method allows dynamic routing of alerts to any channel by including the channel ID in the payload. In addition to routing, messages can be enriched with formatting, blocks, and mentions — for example including <@user\_id> in the payload to directly notify specific team members. This provides a flexible and powerful alternative to fixed incoming webhooks for alerting.

Make sure you created a [webhook for a Slack App with Bot Tokens](https://docs.groundcover.com/~/revisions/ETrLpNk6KtHjyaVUTLoE/integrations/workflow-integrations/slack-app-with-bot-tokens).

Use the following workflow as an example. You can later enrich your workflow with additional functionality.

Here are a few tips for using the example workflow:

1. In the consts section, the `channels` attribute defines the mapping between Slack channels and their IDs. Use a clear, readable label to identify each channel (for example, the channel’s actual name in Slack), and map it to the corresponding channel ID.
2. To locate a channel ID, open the channel in Slack, click the channel name at the top, and scroll to the About section. The channel ID is shown at the bottom of this section.

<figure><img src="/files/zvQyTxJ1MohXBV0tfSRd" alt="" width="375"><figcaption></figcaption></figure>

1. The channel name should be included in the monitor’s **Metadata Labels**, or you can fall back to a default. See the channel\_id attribute in the workflow example.\\

   <figure><img src="/files/kPnV4lZsSaMtRhJZT2UB" alt=""><figcaption></figcaption></figure>
2. Finally, replace the integration name in `{{ providers.slack-routing-webhook }}` with the actual name of the Webhook integration you created.

```yaml
workflow:
  id: slack-channel-routing-workflow
  description: workflow for all channels with dynamic routing
  triggers:
  - type: alert
    filters:
    - key: annotations.slack-channel-routing-workflow
      value: enabled
  name: slack-channel-routing-workflow
  consts:
    channels: '{"devops":"C0111111111", "alerts":"C0222222222", "incidents":"C0333333333"}'
    channel_id: keep.dictget( '{{ consts.channels }}', '{{ alert.labels.channel_id }}', 'C09G9AFHLTB')
    env: keep.dictget({{ alert.labels }}, 'env', 'no-env')
    upper_env: "keep.uppercase({{consts.env}})"
    severity: keep.dictget({{ alert.annotations }}, '_gc_severity', 'unknown-severity')
    summary: keep.dictget({{ alert.labels }}, 'summary', 'no-summary')
    slack_message: "<https://app.groundcover.com/monitors/create-silence?keep.replace(keep.join(keep.dict_pop({{ alert.labels }}, \"_gc_monitor_id\", \"_gc_monitor_name\", \"_gc_severity\", \"backend_id\", \"grafana_folder\", \"_gc_issue_header\"), \"&\", \"matcher_\"), \" \", \"+\")|Silence> :no_bell: | \n<https://app.groundcover.com/monitors/issues?backendId={{ alert.labels.backend_id }}&selectedObjectId={{ alert.fingerprint }}|Investigate> :mag: | \n<https://app.groundcover.com/monitors?backendId={{ alert.labels.backend_id }}&selectedObjectId={{ alert.labels._gc_monitor_id }}|See Monitor> :chart_with_upwards_trend:\n\n*Labels:*  \n- keep.join(keep.dict_pop({{alert.labels}}, \"_gc_monitor_id\", \"_gc_monitor_name\", \"_gc_severity\", \"backend_id\", \"grafana_folder\", \"_gc_issue_header\"), \"\\n- \")\n"
    title_link: "https://app.groundcover.com/monitors/issues?backendId={{ alert.labels.backend_id }}&selectedObjectId={{ alert.fingerprint }}"
    red_color: "#FF0000"
    green_color: "#008000"
    footer_url: "groundcover.com"
    footer_icon: "https://app.groundcover.com/favicon.ico"
  actions:
  - if: "{{ alert.status }} == 'firing'"
    name: webhook-alert
    provider:
      type: webhook
      config: "{{ providers.slack-routing-webhook }}"
      with:
        body:
          channel: "{{ consts.channel_id }}"
          attachments:
          - color: "{{ consts.red_color }}"
            footer: "{{ consts.footer_url }}"
            footer_icon: "{{ consts.footer_icon }}"
            text: "{{ consts.slack_message }}"
            title: "\U0001F6A8 Firing: {{ alert.alertname }} [{{ consts.upper_env}}]"
            title_link: "{{ consts.title_link }}"
            type: plain_text
  - if: "{{ alert.status }} != 'firing'"
    name: webhook-alert-resolved
    provider:
      type: webhook
      config: "{{ providers.slack-routing-webhook }}"
      with:
        body:
          channel: "{{ consts.channel_id }}"
          text: "\u2705 [RESOLVED][{{ consts.upper_env}}] {{ consts.severity }} {{ alert.alertname }}"
          attachments:
          - color: "{{ consts.green_color }}"
            text: "*Summary:* {{ consts.summary }}"
            fields:
            - title: "Environment"
              value: "{{ consts.upper_env}}"
              short: true
            footer: "{{ consts.footer_url }}"
            footer_icon: "{{ consts.footer_icon }}"
```


# Alert Structure

Fields description in the alert you can use in your workflows

{% hint style="warning" %}
Workflows are getting an upgrade: meet [Notification Routes](/use-groundcover/monitors/notification-routes)
{% endhint %}

## Structure

<table><thead><tr><th width="245.015625">Field Name</th><th>Description</th><th>Example</th></tr></thead><tbody><tr><td>labels</td><td>Map of key:values derived from monitor definition.</td><td><p>{<br>"workload": "frontend",</p><p>"namespace": "prod"<br>}</p></td></tr><tr><td>status</td><td>Current status of the alert</td><td><ul><li>firing - Active alert indicating an ongoing issue.</li><li>resolved - The issue has been resolved, and the alert is no longer active.</li><li>suppressed - Alert is suppressed.</li><li>pending - No Data or insufficient data to determine the alert state.</li></ul></td></tr><tr><td>lastReceived</td><td>Timestamp when the alert was last received</td><td>This alert timestamp</td></tr><tr><td>firingStartTime</td><td>Start time of the firing alert</td><td>First timestamp of the current firing state.</td></tr><tr><td>source</td><td>Sources generating the alert</td><td>grafana</td></tr><tr><td>fingerprint</td><td>Unique fingerprint of the alert, this is a hash of the labels</td><td>02f5568d4c4b5b7f</td></tr><tr><td>alertname</td><td>Name of the monitor</td><td>Workload Pods Crashed Monitor</td></tr><tr><td>_gc_severity</td><td>The defined severity of the alert</td><td>S3, error</td></tr><tr><td>trigger</td><td>Trigger condition of the workflow</td><td>alert / manual / interval</td></tr><tr><td>values</td><td><p>A map containing two values that can be used:</p><p>The numeric value that triggered the alert (threshold_input_query) and the actual threshold that was defined for the alert (threshold_1)</p></td><td>"values": { "threshold_1": 0, "threshold_input_query": 99.507}</td></tr></tbody></table>

## Usage

When crafting your workflows you can use any of the fields above using templating in any workflow field. Encapsulate you fields using double opening and closing curly brackets.

### Examples

#### Using Label Values

You can access label values by `alert.labels.*`

```
message: "Pod Crashed - Pod: {{ alert.labels.pod_name }} Namespace: {{ alert.labels.namespace }}"
```


# Search & Filter

### Search and filter

To help you slice and dice your data, you can use our dynamic filters (left panel) and/or our powerful filter bar, which supports key:value pairs, as well as free text search. The Query Builder works in tandem with our filters.

To further focus your results, you can also restrict the results to specific time windows using the time picker on the upper right of the screen.

## Query Builder

The Query Builder is the default search option wherever search is available. Supporting advanced autocomplete of keys, values, and our discovery mode that across values in your data to teach users the data model.

The following syntaxes are available for you to use in Query Builder:

<table><thead><tr><th width="154">Syntax</th><th width="243">Description</th><th width="224">Examples</th><th>Sections</th></tr></thead><tbody><tr><td><code>key:value</code></td><td><p><strong>Search attributes</strong>:</p><p>Both groundcover built-ins custom attributes.</p><p>Use <code>*</code> for wildcard search.<br><br><em>Note: Multiple filters for the same key act as 'OR' conditions, whereas multiple filters for different keys act as 'AND' conditions.</em></p></td><td><code>namespace:prod-us</code><br><code>namespace:prod-*</code></td><td>Logs<br>Traces<br>K8s Events<br>API Catalog<br>Issues</td></tr><tr><td><code>term</code></td><td><strong>Free text:</strong><br>Search for single-word terms.<br><br><em>Tip: Expand your search results by using wildcards.</em></td><td><code>Exception</code><br><code>DivisionBy*</code></td><td>Logs</td></tr><tr><td><code>"term"</code></td><td><strong>Phrase Search (case-insensitive):</strong><br>Enclose terms within double quotes to find results containing the exact phrase.<br><br><em>Note: Using double quotes does not work with <code>*</code> wildcards.</em></td><td><code>"search term"</code></td><td>Logs</td></tr><tr><td><code>-key:value</code></td><td><strong>Exclude:</strong><br>Specify terms or filters to omit from your search; applies to each distinct search.</td><td><code>-key:value</code><br><code>-term</code><br><code>-"search term"</code></td><td>Logs<br>Traces<br>K8s Events<br>API Catalog<br>Issues</td></tr><tr><td><code>*:value</code></td><td><p><strong>Search all attributes:</strong></p><p>Search any attribute for a value, you can use double quotes for exact match and wildcards.</p></td><td><code>*:error</code><br><code>*:"POST /api/search"</code><br><code>*:erro*</code></td><td>Logs<br>Traces<br>Issues</td></tr></tbody></table>

### How to use filters

Filters are very easy to add and remove, using the filters menu on the left bar. You can combine filters with the Query Builder, and filters applied using the left menu will also be added to the Query Builder in text format.

<div align="left"><figure><img src="/files/kGRFIScPZa4dBRWJSGzL" alt="" width="299"><figcaption></figcaption></figure></div>

* **Select / deselect a single filter** - click on the checkbox on the left of the filter. (You can also deselect a filter by clicking the 'x' next to the text format of the filter on the search bar).
* **Deselect all but one filter** (within a filter category, such as 'Level' or 'Format') - hover over the filter you want to leave on, then click on "ONLY".
  * You can switch between filters you want to leave on by hovering on another filter and clicking "ONLY" again.
  * To turn all other filters in that filter category back on, hover over the filter again and click "ALL".
* **Clear all filters within a filters category** - click on the funnel icon next to the category name.
* **Clear all filters currently applied** - click on the funnel icon next to the number of results.


# Saved Views

Save the view of any groundcover page exactly the way you like it, then jump back in a click.

A Saved View captures your current page layout: filters, columns, toggles, etc.. so you and your team can reopen the page with the same context every time. Each groundcover page maintains its own catalogue of views, and every user can pick their personal **Favorites**.

***

### Where to find them

On the pages: Traces, Logs, API Catalog, Events.

Look for the Views selector next to the time‑picker. Click it to open the list, create a new view, or switch between existing ones.

***

### Saving a view

1. Configure the page until it looks exactly right—filters, columns, panels, etc.
2. Click **➕Save View**.
3. Give the view a clear **Name**.
4. Hit **Save**. The view is now listed and available to everyone in the project.

> **Scope** – Saved Views are **per‑page**. A Logs view appears only in Logs; a Traces view only in Traces.

***

### What a view stores

#### Common to all pages

| Category                            | Details                                                    |
| ----------------------------------- | ---------------------------------------------------------- |
| **Filters & facets**                | All query filters plus facet open/closed state             |
| **Columns**                         | Chosen columns, and their order, sort, width               |
| **Filter** **Panel & Added Facets** | <p>Filter panel open/closed</p><p>Facets added/removed</p> |

#### Page‑specific additions

| Page            | Extra properties saved                                                                                                                                                               |
| --------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| **Logs**        | <ul><li><code>logs</code> / <code>patterns</code></li><li><code>textWrap</code></li><li><code>Insight</code> show / hide</li></ul>                                                   |
| **Traces**      | <ul><li><code>traces</code> / <code>span</code></li><li><code>table</code> / <code>drilldown</code></li><li><code>textWrap</code></li><li><code>Insight</code> show / hide</li></ul> |
| **API Catalog** | <ul><li><code>protocol</code></li><li>Kafka role: <code>Fetcher</code> / <code>Producer</code></li></ul>                                                                             |
| **Events**      | <ul><li><code>textWrap</code></li></ul>                                                                                                                                              |

***

### Updating a view

The **Update View** button appears only when **you are the creator** of the view. Click it to overwrite the view with your latest changes.

Underneath every View you can see which user created it.

***

### Managing views (row operations)

| Action            | Who can do it?          |
| ----------------- | ----------------------- |
| **Edit / Rename** | Creator                 |
| **Delete**        | Creator                 |
| **Star / Unstar** | Any user for themselves |

***

### Searching, Sorting, and Filtering the list

Searching the Views will look up based on View names and the creators.

The default sorting pins the favorites views at the top, and the rest of the views below. Each group of views is sorted from A→Z.

In addition, 3 filtering options are available:

1. **All Views** - The entire workspace's views for a specific page
2. **My Favorites** – The favorite views of the user for a specific page
3. **Created By Me** - The views created by the user


# Role-Based Access Control (RBAC)

{% hint style="info" %}
This capability is only available to organizations subscribed to our [Enterprise plan](https://www.groundcover.com/pricing).
{% endhint %}

Role-Based Access Control (RBAC) in groundcover gives you a flexible way to manage who can access certain features and data in the platform. By defining both **default roles** and **policies**, you ensure each team member only sees and does what their level of access permits. This approach strengthens security and simplifies onboarding, allowing administrators to confidently grant or limit access.

{% embed url="<https://www.youtube.com/watch?v=HMDAcYCxMfc>" %}

### Policies

Policies are the **foundational elements** of groundcover’s RBAC. Each policy defines:

1. A **permission level** – which actions the user can perform (Admin, Editor, or Viewer-like capabilities).
2. A **data scope** – which clusters, environments, or namespaces the user can see.

By assigning one or more policies to a user, you can precisely control **both** what they can do **and** where they can do it.

#### Default Policies

groundcover provides three **default policies** to simplify common use cases:

1. **Default Admin Policy**
   * **Permission**: Admin
   * **Data Scope**: Full (no restrictions)
   * **Behavior**: Unlimited access to groundcover features and configurations.
2. **Default Editor Policy**
   * **Permission**: Editor
   * **Data Scope**: Full (no restrictions)
   * **Behavior**: Full creation/editing capabilities on observability data, but no user or system management.
3. **Default Viewer Policy**
   * **Permission**: Viewer
   * **Data Scope**: Full (no restrictions)
   * **Behavior**: Read-only access to all data in groundcover.

These default policies allow you to quickly onboard new users with typical Admin/Editor/Viewer capabilities. However, you can also create **custom policies** with narrower data scopes, if needed.

#### Policy Structure

A policy’s **data scope** can be defined in two modes: **Simple** or **Advanced**.

1. **Simple Mode**
   * Uses **AND** logic across the specified conditions.
   * Applies the *same* scope to all entity types (e.g., logs, traces, events, workloads).
   * **Example**: “Cluster = `Dev` **AND** Environment = `QA`,” restricting **all** logs, traces, events, etc. to the Dev cluster and QA environment.
2. **Advanced Mode**
   * Lets you set a **separate scope per data type**, each enabled independently:
     * **Workload / Infrastructure**
     * **Logs**
     * **Traces**
     * **Events**
     * **Metrics** — also covers **ingestion** data (volume and measurements); there is no separate ingestion scope.
   * Each scope can use **OR** logic among conditions, allowing more fine-grained control.
   * **Example**:
     * **Logs**: “Cluster = `Dev` **OR** `Prod`,”
     * **Traces**: “Namespace = `abc123`,”
     * **Events**: “Environment = `Staging` **OR** `Prod`.”

When creating or editing a policy, you select **permission** (Admin, Editor, or Viewer) *and* a **data scope mode** (Simple or Advanced).

#### Multiple Policies

A user can be associated with **multiple** policies. When that occurs:

1. **Permission Merging**
   * The user’s final permission level is the **highest** among all assigned policies.
   * **Example**: If one policy grants Editor and another grants Viewer, the user is effectively an Editor overall.
2. **Data Scope Merging**
   * Data scopes merge via **OR** logic, broadening the user's overall data access.
   * **Example**: Policy A => "Cluster = `A`," Policy B => "Environment = `B`," so final scope is "Cluster A **OR** Environment B."
   * This applies to all data types including logs, traces, events, workloads, and metrics.

A user may be assigned a policy granting the Editor role with a data scope relevant to specific clusters, and simultaneously be assigned another policy granting the Viewer role with a different data scope. The user's effective access is determined by the highest role across all assigned policies and by the union (OR) of scopes.

***

**In summary:**

* **Policies** define both **permission** (Admin, Editor, or Viewer) and **data scope** (clusters, environments, namespaces).
* **Default Policies** (Admin, Editor, Viewer) provide no data restrictions, suitable for quick onboarding.
* **Custom Policies** allow more granular restrictions, specifying exactly which entities a user can see or modify.
* **Multiple Policies** can co-exist, merging permission levels and data scopes via OR logic across all data types.

This flexible system gives you robust control over observability data in groundcover, ensuring each user has precisely the access they need.


# Remote Access & APIs

groundcover has various authentication key types for remotely interacting with our platform, whether to ingest observability data or to automate actions via our APIs:

1. [**API Keys**](/use-groundcover/remote-access-and-apis/api-keys)- An API key in groundcover provides secure, programmatic access to the API on behalf of a [service account](/use-groundcover/remote-access-and-apis/service-accounts). It inherits that account’s permissions and should be stored safely. This is also the key you need when working groundcover’s terraform provider. See:
   1. groundcover's [APIs documentation](https://docs.groundcover.com/use-groundcover/remote-access-and-apis/api-examples).
   2. groundcover's Terraform provider: <https://github.com/groundcover-com/terraform-provider-groundcover>.
2. [**Ingestion Keys**](/use-groundcover/remote-access-and-apis/ingestion-keys)- Ingestion Keys let sensors, integrations and browsers send observability data to your groundcover backend. These keys are the counterpart of API Keys, which are optimized for reading data or automating dashboards and monitors.
3. [**Datasources (ds) API Key**](/use-groundcover/remote-access-and-apis/querying-you-data-using-an-api)- A key used to connect to groundcover as a datasource, querying Clickhouse and VictoriaMetrics directly.
4. [**Grafana Service Account Token**](/use-groundcover/remote-access-and-apis/build-alerts-and-dashboards-with-grafana-terraform-provider)- Used to remotely configure create Grafana Alerts & Dashboards via Terraform.

\\


# Service Accounts

A service account is a non-human identity for API access, governed by RBAC and supporting multiple API keys.

## Summary

Service accounts in groundcover are non-human identities used for programmatic access to the API. They’re ideal for CI pipelines, automation, and backend services, and are governed by groundcover’s RBAC system.

### Identity and Permissions

A service account has a name and email, but it cannot be used to log into the UI or via SSO. Instead, it functions purely for API access. Each account must have at least one RBAC policy assigned, which defines its permission level (Admin, Editor, Viewer) and data scope. Multiple policies can be attached to broaden access; effective permissions are the union of all policies.

### Creation and Management

Only Admins can create, update, or delete service accounts. This can be done via the UI (Settings → Access → Service Accounts) or API. During creation, Admins define the name, email, and initial policies. You can edit service account, changing email address and assigned policies, but can't rename.

### API Key Association

A service account can have multiple API keys. This makes it easy to rotate credentials or issue distinct keys for different use cases. All keys are tied to the same account and carry its permissions. Any action taken using a key is logged as performed by the associated service account.


# API Keys

An API key in groundcover provides secure, programmatic access to the API on behalf of a service account. It inherits that account’s permissions and should be stored safely.

An API key in groundcover provides secure, programmatic access to the API on behalf of a service account. It inherits that account’s permissions and should be stored safely

### Binding and Permissions

Each API key is tied to a specific service account. It inherits the permissions defined by that account’s RBAC policies. Optionally, the key can be limited to a subset of those policies for more granular access control. An API key can never exceed the permissions of its parent service account.

### **Creation and Storage**

Only Admins can create or revoke API keys. To create an API key:

1. Navigate to the Settings page using the settings button located in the bottom left corner
2. Select "Access" from the sidebar menu
3. Click on the "API Keys" tab
4. Create a new API Key, ensuring you assign it to a service account that is bound to the appropriate RBAC policy

When a key is created, its value is shown once—store it securely in a secret manager or encrypted environment variable. **If lost, a new key must be issued.**

### Authentication and Usage

To use an API key, send it in the Authorization header as bearer token:

```sh
curl 'https://api.groundcover.com/api/k8s/v3/clusters/list' \
  -H 'accept: application/json' \
  -H 'authorization: Bearer <YOUR_API_KEY>' \
  -H 'X-Backend-Id: <YOUR_BACKEND_ID>' \
  --data-raw '{"sources":[]}'
```

The key authenticates as the service account, and all API permissions are enforced accordingly.

API Key authentication will work using `https://api.groundcover.com/` only.

### Validity and Revocation

API keys do not expire automatically. Revoking a key immediately disables its access.

### Scope of Use

API keys are valid only for requests to `https://api.groundcover.com`. They do not support data ingestion or Grafana integration—those require dedicated tokens.

### API Keys vs Ingestion Keys

|                           | **Ingestion Key**                        | **API Key**                           |
| ------------------------- | ---------------------------------------- | ------------------------------------- |
| Primary purpose           | Write data (ingest)                      | Read data / manage resources via REST |
| Permissions capabilities  | Write‑only + optional remote‑config read | Mirrors service‑account RBAC          |
| Visibility after creation | Always revealable                        | Shown **once** only                   |
| Typical lifetime          | Tied to integration lifecycle            | Rotated for CI/CD automations         |
| Revocation effect         | Data stops flowing immediately           | API calls fail                        |

### Security Best Practices

**Store securely:** Use secrets managers like AWS Secrets Manager or HashiCorp Vault. Never commit keys to source control.

**Follow least privilege:** Assign the minimal required policies to service accounts and API keys. Avoid defaulting to admin-level access.

**Rotate regularly:** Periodically generate new keys, update your systems, and revoke old ones to limit exposure.

**Revoke stale keys:** Remove keys that are no longer in use or suspected to be compromised.


# Ingestion Keys

Secure, write‑focused credentials for streaming data into groundcover

Ingestion Keys let sensors, integrations and browsers **send** observability data to your groundcover backend.\
They are the counterpart of API Keys, which are optimized for **reading** data or automating dashboards and monitors.

***

### Key types

| **Sensor\***    | Install the eBPF sensor on Kubernetes or Hosts/VMs                                                   |
| --------------- | ---------------------------------------------------------------------------------------------------- |
| **RUM**         | Send Real‑User‑Monitoring events using JS snippet embedded in web pages                              |
| **Third Party** | Integrate 3rd-party data sources that *push* data (e.g. OpenTelemtry, AWS Firehose, FluentBit, etc.) |

\*Only the `Sensor` has limited read capability in order to support pulling [remote configuration](/use-groundcover/fleet-manager#remote-configuration) such as [OTTL parsing rules](https://github.com/groundcover-com/docs/tree/main/use-groundcover/broken-reference/README.md) applied from the UI. `RUM` and `Third Party` have write-only configurations.

***

### Creating an Ingestion Key

**It is recommended to create a dedicated Ingestion Key for every data source**, so that they can be managed and rotated appropriately, minimize exposure or risk, and allow groundcover to identify the datasource of all the ingested data.

1. Open **Settings → Access → Ingestion Keys** and click **Create key**.
2. Give the key a clear, descriptive **Name** (for example `k8s-prod‑eu‑central‑1`).
3. Select the **Type** that matches your integration.
4. Click **Click & Copy Key**.
   1. Unlike API Keys, Ingestion Keys stay visible on the page. Treat every reveal as sensitive and follow the same secret‑handling practices.
5. **Store** they Key securely, and continue to **integrate** your data source.

***

### Using an Ingestion Key

#### Kubernetes sensor example

```
helm upgrade --install groundcover groundcover/groundcover \
  --set global.groundcover_token=<INGESTION_KEY>,clusterId={cluster-name}
```

#### OpenTelemetry integration (OTel/HTTP) example

```
exporters:
  otlphttp/groundcover:
    endpoint: https://{GROUNDCOVER_MANAGED_OPENTELEMETRY_ENDPOINT}
    headers: 
      apikey: {INGESTION_KEY}

pipelines:
  traces:
    exporters:
    - otlphttp/groundcover
```

***

### Viewing keys

The Ingestion Keys table lets you:

* **Reveal** the key at any time.
* See **who created** the key and **when**.
* Sort by **Type** or **Creator** to locate specific credentials quickly.

***

### Revoking a key

Click **⋮ → Revoke** next to the key. Revocation **permanently deletes** the key, unlike API Keys which only disables it:

* The key will disappear from the list.
* Any service using it will receive 403 / PERMISSION\_DENIED and will not be able to continue to send data or pull latest configurations.

This operation cannot be undone — create a new key and update your deployments if you need access again.

***

### Ingestion Keys vs. API Keys

|                           | **Ingestion Key**                        | **API Key**                           |
| ------------------------- | ---------------------------------------- | ------------------------------------- |
| Primary purpose           | Write data (ingest)                      | Read data / manage resources via REST |
| Permissions capabilities  | Write‑only + optional remote‑config read | Mirrors service‑account RBAC          |
| Visibility after creation | Always revealable                        | Shown **once** only                   |
| Typical lifetime          | Tied to integration lifecycle            | Rotated for CI/CD automations         |
| Revocation effect         | Data stops flowing immediately           | API calls fail                        |

***

### Best Practices

* **One key per integration** – simplifies rotation and blast radius.
* **Store securely** – AWS Secrets Manager, GCP Secret Manager, HashiCorp Vault, Kubernetes Secrets.
* **Rotate regularly** – create a new key, roll it out, then revoke the old one.
* **Monitor for 403 errors** – a spike usually means a revoked or expired key.

***


# Datasource API Keys (Legacy)

groundcover provides a robust user interface that allows you to view and analyze all your observability data from inside the platform. However, there may be cases in which you need to query the data from outside our platform using API communication.

Our proprietary eBPF sensor automatically captures granular observability data, which is stored via our integrations with two best-of-breed technologies. VictoriaMetrics for metrics storage, and ClickHouse for storage of logs, traces, and Kubernetes events.

Read more about our architecture [here](https://github.com/groundcover-com/docs/tree/main/use-groundcover/remote-access-and-apis/broken-reference/README.md).

{% hint style="warning" %}
Using the datasource API is not recommended (Legacy) as there are newer ways to use the API. You can refer to [Metrics and Logs API](/use-groundcover/remote-access-and-apis/raw-prometheus-and-clickhouse)
{% endhint %}

### Generate the API key

Run the following command in your CLI, and select tenant:

`groundcover auth get-datasources-api-key`

### **Querying ClickHouse**

Example for querying ClickHouse database using POST HTTP Request:

```bash
curl 'https://ds.groundcover.com/' \
        --header "X-ClickHouse-Key: ${API_KEY}" \
        --data "SELECT count() from traces where start_timestamp > now() - interval '15 minutes' "
```

#### Command parameters

* `X-ClickHouse-Key` (header): API Key you retrieved from the groundcover CLI. Replace `${API_KEY}` with your actual API key, or set `API_KEY` as env parameter.
* `SELECT count() FROM traces WHERE start_timestamp > now() - interval '15 minutes'` (data): The SQL query to execute. This query counts the number of traces where the `start_timestamp` is within the last 15 minutes.

Learn more about the ClickHouse query language [here](https://clickhouse.com/docs/en/sql-reference/statements/select).

### **Querying VictoriaMetrics**

Example for querying the VictoriaMetrics database using the [`query_range`](https://prometheus.io/docs/prometheus/latest/querying/api/#range-queries) API:

```bash
curl 'https://ds.groundcover.com/datasources/prometheus/api/v1/query_range' \
    --get \
    --header "apikey: ${API_KEY}" \
    --data 'query=sum(rate(groundcover_resource_total_counter{type="http"}))' \
    --data 'start=1715760000' \
    --data 'end=1715763600'
```

#### Command parameters

* `apikey` (header): API Key you retrieved from the groundcover CLI. Replace `${API_KEY}` with your actual API key, or set `API_KEY` as env parameter.
* `query` (data): The promql query to execute. In this case, it calculates the sum of the rate of `groundcover_resource_total_counter` with the `type` set to `http`.
* `start` (data): The start timestamp for the query range in Unix time (seconds since epoch). Example: `1715760000`.
* `end` (data): The end timestamp for the query range in Unix time (seconds since epoch). Example: `1715763600`.

Learn more about the promql syntax [here](https://prometheus.io/docs/prometheus/latest/querying/basics/).

Learn more about VictoriaMetrics HTTP API [here](https://prometheus.io/docs/prometheus/latest/querying/api/).


# Grafana Service Account Token

## Step 1 - generate Grafana Service Account Token

* Make sure you have [groundcover cli installed](/getting-started/installation-and-updating/connect-kubernetes-cluster#installing-groundcover-cli)

```bash
groundcover version
```

* Generate service account token

```bash
groundcover auth generate-service-account-token
```

{% hint style="warning" %}
Service Account Token are only accessible once, so make sure you keep them somewhere safe, running the command again will generate a new service account token
{% endhint %}

{% hint style="warning" %}
Only groundcover tenant admins can generate Service Account Tokens
{% endhint %}

## Step 2 - Use Grafana Terraform provider

* make sure you have [Terraform installed](https://developer.hashicorp.com/terraform/tutorials/aws-get-started/install-cli)
* Use the [official Grafana Terraform provider](https://registry.terraform.io/providers/grafana/grafana/latest/docs) with the following attributes

```hcl
terraform {
  required_providers {
    grafana = {
      source = "grafana/grafana"
    }
  }
}

provider "grafana" {
  url  = "https://app.groundcover.com/grafana"
  auth = "{service account token}"
}
```

Continue to create Alerts and Dashboards in Grafana, see: [Build alerts & dashboards with Grafana Terraform provider](/use-groundcover/embedded-grafana/embedded-grafana-dashboards/build-alerts-and-dashboards-with-grafana-terraform-provider).

You can read more about what you can achieve with the Grafana Terraform provider in the [official docs](https://registry.terraform.io/providers/grafana/grafana/latest/docs)


# Metrics and Logs API

This page describes the available API endpoints for querying logs and metrics in groundcover, including how to authenticate and structure requests for general data retrieval.

## Authentication

Authentication is performed using an API key generated in the [API Keys section](/use-groundcover/remote-access-and-apis/api-keys) of the groundcover console.

All API requests must include the API key in the Authorization header using the following format:

```
Authorization: Bearer <YOUR_API_KEY>
```

## Rate Limits

groundcover APIs enforce rate limits on a per-client basis to ensure fair usage and system stability. Each client is allowed up to **150 requests per second and 1,000 requests per minute**. A “client” is identified by the Authorization header, which typically represents an API key. This means that rate limits are applied independently for each API key used

## Raw Prometheus API for Metrics

Use the following endpoint to query the groundcover Prometheus API:

```bash
https://app.groundcover.com/api/prometheus/api/v1/query
```

To see usage example on how to query metrics in groundcover see: [Query metrics examples](/use-groundcover/remote-access-and-apis/api-examples/query-metrics)

For a complete Prometheus REST API documentation, and available operations, see: [Prometheus HTTP API](https://prometheus.io/docs/prometheus/latest/querying/api/)

## Logs API

Use the following endpoint to query the groundcover logs API:

```bash
https://app.groundcover.com/api/logs/v2/query
```

To see usage examples on how to query logs in groundcover see: [Query logs examples](/use-groundcover/remote-access-and-apis/api-examples/query-logs)

## Legacy API (Clickhouse)

The following configurations are deprecated but may still be in use in older setups.

{% hint style="warning" %}
The legacy datasources is using a different API key than described above. The API key can be obtained by running: `groundcover auth get-datasources-api-key`
{% endhint %}

You can query the legacy API to execute SQL statements directly on the Clickhouse database.

Try the following structure:

```bash
curl https://ds.groundcover.com/ \
    -H "X-ClickHouse-Key: <DS-API-KEY-VALUE>" \
    --data "SELECT count() FROM logs LIMIT 1 FORMAT JSON" 
```


# API Examples

Welcome to the API examples section. Here, you’ll find practical demonstrations of how to interact with our API endpoints using cURL commands. Each example is designed to help you quickly understand h

### Structure of the Examples

* cURL-based examples: Every example shows the exact cURL command you can copy and run directly in your terminal.
* Endpoint-specific demonstrations: We walk through different API endpoints one by one, highlighting the required parameters and common use cases.
* Request & Response clarity: Each section contains both the request (what you send) and the response (what you get back) to illustrate expected behavior.

### Prerequisites

Before running any of the examples, make sure you have:

1. [API Key](https://docs.groundcover.com/use-groundcover/remote-access-and-apis/api-keys)
2. Backend ID

The Backend ID can be obtained from your groundcover Admin panel.

**To locate it:**

* Navigate to Settings → Access
* Open the API Keys tab
* The Backend ID appears in the section header

**When to Use**

The Backend ID is required when your account is associated with multiple backends.

In such cases, include it in your API requests using the appropriate request header to ensure the request is routed to the correct backend.

```yaml
X-Backend-Id: <YOUR_BACKEND_ID>
```


# List Workloads

Retrieve a list of Kubernetes workloads with their performance metrics, resource usage, and metadata.

### Endpoint

```
POST /api/k8s/v3/workloads/list
```

### Authentication

This endpoint requires API Key authentication via the Authorization header.

### Headers

| Header          | Required | Description                    |
| --------------- | -------- | ------------------------------ |
| `Authorization` | Yes      | Bearer token with your API key |
| `X-Backend-Id`  | Yes      | Your backend identifier        |
| `Content-Type`  | Yes      | Must be `application/json`     |
| `Accept`        | Yes      | Must be `application/json`     |

### Request Body

| Parameter    | Type    | Required | Default  | Description                                                     |
| ------------ | ------- | -------- | -------- | --------------------------------------------------------------- |
| `conditions` | Array   | No       | `[]`     | Filter conditions for workloads                                 |
| `limit`      | Integer | No       | `100`    | Maximum number of workloads to return (1-1000)                  |
| `skip`       | Integer | No       | `0`      | Number of workloads to skip for pagination                      |
| `order`      | String  | No       | `"desc"` | Sort order: `"asc"` or `"desc"`                                 |
| `sortBy`     | String  | No       | `"rps"`  | Field to sort by (e.g., `"rps"`, `"cpuUsage"`, `"memoryUsage"`) |
| `sources`    | Array   | No       | `[]`     | Filter by data sources                                          |

### Response

The response contains a paginated list of workloads with their metrics and metadata.

#### Response Fields

| Field       | Type    | Description                         |
| ----------- | ------- | ----------------------------------- |
| `total`     | Integer | Total number of workloads available |
| `workloads` | Array   | Array of workload objects           |

#### Workload Object Fields

| Field             | Type    | Description                                                                     |
| ----------------- | ------- | ------------------------------------------------------------------------------- |
| `uid`             | String  | Unique identifier for the workload                                              |
| `envType`         | String  | Environment type (e.g., `"k8s"`)                                                |
| `env`             | String  | Environment name (e.g., `"prod"`, `"ga"`, `"alpha"`)                            |
| `cluster`         | String  | Kubernetes cluster name                                                         |
| `namespace`       | String  | Kubernetes namespace                                                            |
| `workload`        | String  | Workload name                                                                   |
| `kind`            | String  | Kubernetes resource kind (e.g., `"ReplicaSet"`, `"StatefulSet"`, `"DaemonSet"`) |
| `resourceVersion` | Integer | Kubernetes resource version                                                     |
| `ready`           | Boolean | Whether the workload is ready                                                   |
| `podsCount`       | Integer | Number of pods in the workload                                                  |
| `p50`             | Float   | 50th percentile response time in seconds                                        |
| `p95`             | Float   | 95th percentile response time in seconds                                        |
| `p99`             | Float   | 99th percentile response time in seconds                                        |
| `rps`             | Float   | Requests per second                                                             |
| `errorRate`       | Float   | Error rate as a decimal (e.g., 0.004 = 0.4%)                                    |
| `cpuLimit`        | Integer | CPU limit in millicores (0 = no limit)                                          |
| `cpuUsage`        | Float   | Current CPU usage in millicores                                                 |
| `memoryLimit`     | Integer | Memory limit in bytes (0 = no limit)                                            |
| `memoryUsage`     | Integer | Current memory usage in bytes                                                   |
| `issueCount`      | Integer | Number of issues detected                                                       |

### Examples

#### Basic Request

```bash
curl 'https://api.groundcover.com/api/k8s/v3/workloads/list' \
  -H 'accept: application/json' \
  -H 'authorization: Bearer <YOUR_API_KEY>' \
  -H 'content-type: application/json' \
  -H 'X-Backend-Id: <YOUR_BACKEND_ID>' \
  --data-raw '{"conditions":[],"limit":100,"order":"desc","skip":0,"sortBy":"rps","sources":[]}'
```

#### Response Example

```json
{
  "total": 6314,
  "workloads": [
    {
      "uid": "824b00bf-db68-47b5-8a53-9abd98bf7c0a",
      "envType": "k8s",
      "env": "ga",
      "cluster": "akamai-lk41ok",
      "namespace": "groundcover-incloud",
      "workload": "groundcover-incloud-vector",
      "kind": "ReplicaSet",
      "resourceVersion": 651723275,
      "ready": true,
      "podsCount": 5,
      "p50": 0.0005824280087836087,
      "p95": 0.005730729550123215,
      "p99": 0.0327172689139843,
      "rps": 5526.0027359781125,
      "errorRate": 0,
      "cpuLimit": 0,
      "cpuUsage": 50510.15252730218,
      "memoryLimit": 214748364800,
      "memoryUsage": 46527352832,
      "issueCount": 0
    }
  ]
}
```

### Pagination

To retrieve all workloads, use pagination by incrementing the `skip` parameter:

#### Fetching All Results

```bash
# First batch (0-99)
curl 'https://api.groundcover.com/api/k8s/v3/workloads/list' \
  -H 'accept: application/json' \
  -H 'authorization: Bearer <YOUR_API_KEY>' \
  -H 'content-type: application/json' \
  -H 'X-Backend-Id: <YOUR_BACKEND_ID>' \
  --data-raw '{"conditions":[],"limit":100,"order":"desc","skip":0,"sortBy":"rps","sources":[]}'

# Second batch (100-199)
curl 'https://api.groundcover.com/api/k8s/v3/workloads/list' \
  -H 'accept: application/json' \
  -H 'authorization: Bearer <YOUR_API_KEY>' \
  -H 'content-type: application/json' \
  -H 'X-Backend-Id: <YOUR_BACKEND_ID>' \
  --data-raw '{"conditions":[],"limit":100,"order":"desc","skip":100,"sortBy":"rps","sources":[]}'

# Continue incrementing skip by 100 until you reach the total count
```

#### Pagination Logic

To fetch all results programmatically:

1. Start with `skip=0` and `limit=100` (or your preferred page size)
2. Check the `total` field in the response
3. Continue making requests, incrementing `skip` by your `limit` value
4. Stop when `skip` >= `total`

**Example calculation:**

* If `total` is 6314 and `limit` is 100
* You need ⌈6314/100⌉ = 64 requests
* Last request: `skip=6300`, `limit=100` (returns 14 items)


# List Namespaces

Retrieve a list of Kubernetes namespaces within a specified time range.

### Endpoint

```
POST /api/k8s/v2/namespaces/list
```

### Authentication

This endpoint requires API Key authentication via the Authorization header.

### Headers

| Header          | Required | Description                    |
| --------------- | -------- | ------------------------------ |
| `Authorization` | Yes      | Bearer token with your API key |
| `X-Backend-Id`  | Yes      | Your backend identifier        |
| `Content-Type`  | Yes      | Must be `application/json`     |
| `Accept`        | Yes      | Must be `application/json`     |

### Request Body

| Parameter | Type   | Required | Description                                          |
| --------- | ------ | -------- | ---------------------------------------------------- |
| `sources` | Array  | No       | Filter by data sources (empty array for all sources) |
| `start`   | String | Yes      | Start timestamp in ISO 8601 format (UTC)             |
| `end`     | String | Yes      | End timestamp in ISO 8601 format (UTC)               |

#### Time Range Parameters

* **Format**: ISO 8601 format with milliseconds: `YYYY-MM-DDTHH:mm:ss.sssZ`
* **Timezone**: All timestamps must be in UTC (denoted by 'Z' suffix)

### Response

The response contains an array of namespaces for the specified time period.

#### Response Fields

| Field        | Type  | Description                                   |
| ------------ | ----- | --------------------------------------------- |
| `namespaces` | Array | Array of namespace names or namespace objects |

### Examples

#### Basic Request

```bash
curl 'https://api.groundcover.com/api/k8s/v2/namespaces/list' \
  -H 'accept: application/json' \
  -H 'authorization: Bearer <YOUR_API_KEY>' \
  -H 'content-type: application/json' \
  -H 'X-Backend-Id: <YOUR_BACKEND_ID>' \
  --data-raw '{"sources":[],"start":"2025-01-24T06:00:00.000Z","end":"2025-01-24T08:00:00.000Z"}'
```

#### Response Example

```json
{
  "namespaces": [
    "groundcover",
    "monitoring",
    "kube-system",
    "default"
  ]
}
```

### Time Range Usage

#### Last 24 Hours

```bash
# Get current time and subtract 24 hours for start time
start_time=$(date -u -v-24H '+%Y-%m-%dT%H:%M:%S.000Z')
end_time=$(date -u '+%Y-%m-%dT%H:%M:%S.000Z')

curl 'https://api.groundcover.com/api/k8s/v2/namespaces/list' \
  -H 'accept: application/json' \
  -H 'authorization: Bearer <YOUR_API_KEY>' \
  -H 'content-type: application/json' \
  -H 'X-Backend-Id: <YOUR_BACKEND_ID>' \
  --data-raw "{\"sources\":[],\"start\":\"$start_time\",\"end\":\"$end_time\"}"
```


# List Clusters

Retrieve a list of Kubernetes clusters with their resource usage metrics, metadata, and health information.

### Endpoint

```
POST /api/k8s/v3/clusters/list
```

### Authentication

This endpoint requires API Key authentication via the Authorization header.

### Headers

| Header          | Required | Description                    |
| --------------- | -------- | ------------------------------ |
| `Authorization` | Yes      | Bearer token with your API key |
| `X-Backend-Id`  | Yes      | Your backend identifier        |
| `Content-Type`  | Yes      | Must be `application/json`     |
| `Accept`        | Yes      | Must be `application/json`     |

### Request Body

| Parameter | Type  | Required | Description                                          |
| --------- | ----- | -------- | ---------------------------------------------------- |
| `sources` | Array | No       | Filter by data sources (empty array for all sources) |

### Response

The response contains an array of clusters with detailed resource usage and metadata.

#### Response Fields

| Field        | Type    | Description              |
| ------------ | ------- | ------------------------ |
| `clusters`   | Array   | Array of cluster objects |
| `totalCount` | Integer | Total number of clusters |

#### Cluster Object Fields

| Field               | Type    | Description                                                 |
| ------------------- | ------- | ----------------------------------------------------------- |
| `name`              | String  | Cluster name                                                |
| `env`               | String  | Environment (e.g., "prod", "ga", "beta", "alpha", "latest") |
| `creationTimestamp` | String  | When the cluster was created (ISO 8601)                     |
| `cloudProvider`     | String  | Cloud provider (e.g., "AWS", "GCP", "Azure")                |
| `kubernetesVersion` | String  | Kubernetes version                                          |
| `nodesCount`        | Integer | Number of nodes in the cluster                              |
| `issueCount`        | Integer | Number of issues detected                                   |

**CPU Metrics**

| Field                          | Type    | Description                               |
| ------------------------------ | ------- | ----------------------------------------- |
| `cpuUsage`                     | Integer | Current CPU usage in millicores           |
| `cpuLimit`                     | Integer | CPU limits set on resources in millicores |
| `cpuAllocatable`               | Integer | Total allocatable CPU in millicores       |
| `cpuRequest`                   | Integer | Total CPU requests in millicores          |
| `cpuUsageAllocatablePercent`   | Float   | CPU usage as percentage of allocatable    |
| `cpuRequestAllocatablePercent` | Float   | CPU requests as percentage of allocatable |
| `cpuUsageRequestPercent`       | Float   | CPU usage as percentage of requests       |
| `cpuUsageLimitPercent`         | Float   | CPU usage as percentage of limits         |
| `cpuLimitAllocatablePercent`   | Float   | CPU limits as percentage of allocatable   |

**Memory Metrics**

| Field                             | Type    | Description                                  |
| --------------------------------- | ------- | -------------------------------------------- |
| `memoryUsage`                     | Integer | Current memory usage in bytes                |
| `memoryLimit`                     | Integer | Memory limits set on resources in bytes      |
| `memoryAllocatable`               | Integer | Total allocatable memory in bytes            |
| `memoryRequest`                   | Integer | Total memory requests in bytes               |
| `memoryUsageAllocatablePercent`   | Float   | Memory usage as percentage of allocatable    |
| `memoryRequestAllocatablePercent` | Float   | Memory requests as percentage of allocatable |
| `memoryUsageRequestPercent`       | Float   | Memory usage as percentage of requests       |
| `memoryUsageLimitPercent`         | Float   | Memory usage as percentage of limits         |
| `memoryLimitAllocatablePercent`   | Float   | Memory limits as percentage of allocatable   |

**Pod Information**

| Field  | Type   | Description                                                   |
| ------ | ------ | ------------------------------------------------------------- |
| `pods` | Object | Pod counts by status (e.g., {"Running": 157, "Succeeded": 4}) |

### Examples

#### Basic Request

```bash
curl 'https://api.groundcover.com/api/k8s/v3/clusters/list' \
  -H 'accept: application/json' \
  -H 'authorization: Bearer <YOUR_API_KEY>' \
  -H 'content-type: application/json' \
  -H 'X-Backend-Id: <YOUR_BACKEND_ID>' \
  --data-raw '{"sources":[]}'
```

#### Response Example

```json
{
  "clusters": [
    {
      "name": "production-cluster",
      "env": "prod",
      "cpuUsage": 126640,
      "cpuLimit": 289800,
      "cpuAllocatable": 302820,
      "cpuRequest": 187975,
      "cpuUsageAllocatablePercent": 41.82,
      "cpuRequestAllocatablePercent": 62.07,
      "cpuUsageRequestPercent": 67.37,
      "cpuUsageLimitPercent": 43.70,
      "cpuLimitAllocatablePercent": 95.70,
      "memoryUsage": 242994409472,
      "memoryLimit": 604262891520,
      "memoryAllocatable": 1227361431552,
      "memoryRequest": 495549677568,
      "memoryUsageAllocatablePercent": 19.80,
      "memoryRequestAllocatablePercent": 40.38,
      "memoryUsageRequestPercent": 49.04,
      "memoryUsageLimitPercent": 40.21,
      "memoryLimitAllocatablePercent": 49.23,
      "nodesCount": 6,
      "pods": {
        "Running": 109,
        "Succeeded": 3
      },
      "issueCount": 1,
      "creationTimestamp": "2021-11-01T14:37:31Z",
      "cloudProvider": "AWS",
      "kubernetesVersion": "v1.30.14-eks-931bdca"
    }
  ],
  "totalCount": 116
}
```


# List Deployments

Get a list of Kubernetes deployments with status information, replica counts, and operational conditions for a specified time range.

### Endpoint

**POST** `/api/k8s/v2/deployments/list`

### Authentication

This endpoint requires API Key authentication via the Authorization header.

### Headers

| Header          | Required | Description                    |
| --------------- | -------- | ------------------------------ |
| `Authorization` | Yes      | Bearer token with your API key |
| `X-Backend-Id`  | Yes      | Your backend identifier        |
| `Content-Type`  | Yes      | Must be `application/json`     |
| `Accept`        | Yes      | Must be `application/json`     |

#### Request Body

The request body requires a time range and supports filtering by fields:

```json
{
  "start": "2025-08-24T07:21:36.944Z",
  "end": "2025-08-24T08:51:36.944Z", 
  "namespaces": ["groundcover"],
  "sources": []
}
```

**Parameters**

| Parameter    | Type   | Required | Description                                                                |
| ------------ | ------ | -------- | -------------------------------------------------------------------------- |
| `start`      | string | Yes      | Start time in ISO 8601 UTC format (e.g., `"2025-08-24T07:21:36.944Z"`)     |
| `end`        | string | Yes      | End time in ISO 8601 UTC format (e.g., `"2025-08-24T08:51:36.944Z"`)       |
| `namespaces` | array  | No       | Array of namespace names to filter by (e.g., `["groundcover", "default"]`) |
| `sources`    | array  | No       | Source filters                                                             |

### Response

#### Response Schema

```json
{
  "deployments": [
    {
      "name": "string",
      "namespace": "string", 
      "workloadName": "string",
      "creationTime": "2023-08-30T18:27:01Z",
      "cluster": "string",
      "env": "string",
      "available": 1,
      "desired": 1,
      "ready": 1,
      "conditions": [
        {
          "type": "string",
          "status": "string",
          "lastProbeTime": null,
          "lastHeartbeatTime": null,
          "lastTransitionTime": "string",
          "reason": "string",
          "message": "string"
        }
      ],
      "warnings": [],
      "id": "string",
      "resourceVersion": 0
    }
  ]
}
```

**Field Descriptions**

| Field                             | Type        | Description                                           |
| --------------------------------- | ----------- | ----------------------------------------------------- |
| `deployments`                     | array       | Array of deployment objects                           |
| `name`                            | string      | Deployment name                                       |
| `namespace`                       | string      | Kubernetes namespace                                  |
| `workloadName`                    | string      | Associated workload name                              |
| `creationTime`                    | string      | Deployment creation timestamp in ISO 8601 format      |
| `cluster`                         | string      | Kubernetes cluster name                               |
| `env`                             | string      | Environment name (e.g., `"prod"`, `"staging"`)        |
| `available`                       | integer     | Number of available replicas                          |
| `desired`                         | integer     | Number of desired replicas                            |
| `ready`                           | integer     | Number of ready replicas                              |
| `conditions`                      | array       | Array of deployment condition objects                 |
| `conditions[].type`               | string      | Condition type (e.g., `"Available"`, `"Progressing"`) |
| `conditions[].status`             | string      | Condition status (`"True"`, `"False"`, `"Unknown"`)   |
| `conditions[].lastProbeTime`      | string/null | Last time the condition was probed                    |
| `conditions[].lastHeartbeatTime`  | string/null | Last time the condition was updated                   |
| `conditions[].lastTransitionTime` | string      | Last time the condition transitioned                  |
| `conditions[].reason`             | string      | Machine-readable reason for the condition             |
| `conditions[].message`            | string      | Human-readable message explaining the condition       |
| `warnings`                        | array       | Array of warning messages (usually empty)             |
| `id`                              | string      | Unique identifier for the deployment                  |
| `resourceVersion`                 | integer     | Kubernetes resource version                           |

**Common Condition Types**

| Type          | Description                                         |
| ------------- | --------------------------------------------------- |
| `Available`   | Deployment has minimum availability                 |
| `Progressing` | Deployment is making progress towards desired state |

**Common Condition Reasons**

| Reason                     | Description                                         |
| -------------------------- | --------------------------------------------------- |
| `MinimumReplicasAvailable` | Deployment has minimum number of replicas available |
| `NewReplicaSetAvailable`   | New ReplicaSet has successfully progressed          |

### Examples

#### Basic Request

Get deployments for a specific time range:

```bash
curl 'https://api.groundcover.com/api/k8s/v2/deployments/list' \
  -H 'accept: application/json' \
  -H 'authorization: Bearer <YOUR_API_KEY>' \
  -H 'content-type: application/json' \
  -H 'X-Backend-Id: <YOUR_BACKEND_ID>' \
  --data-raw '{"start":"2025-08-24T07:21:36.944Z","end":"2025-08-24T08:51:36.944Z","namespaces":[],"sources":[]}'
```

#### Filter by Namespace

Get deployments from specific namespaces:

```bash
curl 'https://api.groundcover.com/api/k8s/v2/deployments/list' \
  -H 'accept: application/json' \
  -H 'authorization: Bearer <YOUR_API_KEY>' \
  -H 'content-type: application/json' \
  -H 'X-Backend-Id: <YOUR_BACKEND_ID>' \
  --data-raw '{"start":"2025-08-24T07:21:36.944Z","end":"2025-08-24T08:51:36.944Z","namespaces":["groundcover","monitoring"],"sources":[]}'
```

#### Response Example

```json
{
  "deployments": [
    {
      "name": "db-manager",
      "namespace": "groundcover",
      "workloadName": "db-manager",
      "creationTime": "2023-08-30T18:27:01Z",
      "cluster": "karma-cluster",
      "env": "prod",
      "available": 1,
      "desired": 1,
      "ready": 1,
      "conditions": [
        {
          "type": "Available",
          "status": "True",
          "lastProbeTime": null,
          "lastHeartbeatTime": null,
          "lastTransitionTime": "2025-08-22T06:18:27Z",
          "reason": "MinimumReplicasAvailable",
          "message": "Deployment has minimum availability."
        },
        {
          "type": "Progressing",
          "status": "True",
          "lastProbeTime": null,
          "lastHeartbeatTime": null,
          "lastTransitionTime": "2023-08-30T18:27:01Z",
          "reason": "NewReplicaSetAvailable",
          "message": "ReplicaSet \"db-manager-867bc8f5b8\" has successfully progressed."
        }
      ],
      "warnings": [],
      "id": "f3b1f4a5-f38a-4c63-a7c0-9333fcbf1906",
      "resourceVersion": 747039184
    }
  ]
}
```

### Time Range Guidelines

* Use ISO 8601 UTC format for timestamps
* Typical time ranges: 1-24 hours for operational monitoring
* Maximum recommended range: 7 days
* Format: `YYYY-MM-DDTHH:MM:SS.sssZ`


# List Monitors

Get a list of all configured monitors in the system with their identifiers, titles, and types.

### Endpoint

**POST** `/api/monitors/list`

### Authentication

This endpoint requires API Key authentication via the Authorization header.

### Headers

| Header          | Required | Description                    |
| --------------- | -------- | ------------------------------ |
| `Authorization` | Yes      | Bearer token with your API key |
| `Content-Type`  | Yes      | Must be `application/json`     |
| `Accept`        | Yes      | Must be `application/json`     |

### Request Body

The request body supports filtering by sources:

```json
{
  "sources": []
}
```

**Parameters**

| Parameter | Type  | Required | Description                                       |
| --------- | ----- | -------- | ------------------------------------------------- |
| `sources` | array | No       | Source filters (empty array returns all monitors) |

### Response

#### Response Schema

```json
{
  "monitors": [
    {
      "uuid": "string",
      "title": "string",
      "type": "string"
    }
  ]
}
```

**Field Descriptions**

| Field      | Type   | Description                            |
| ---------- | ------ | -------------------------------------- |
| `monitors` | array  | Array of monitor objects               |
| `uuid`     | string | Unique identifier for the monitor      |
| `title`    | string | Monitor name/description               |
| `type`     | string | Monitor type (see monitor types below) |

**Monitor Types**

| Type        | Description                                   |
| ----------- | --------------------------------------------- |
| `"metrics"` | Metrics-based monitoring                      |
| `"traces"`  | Distributed tracing monitoring                |
| `"logs"`    | Log-based monitoring                          |
| `"events"`  | Event-based monitoring                        |
| `"infra"`   | Infrastructure monitoring                     |
| `""`        | (empty string) General/unspecified monitoring |

### Examples

#### Basic Request

Get all monitors:

```bash
curl -L \
  --request POST \
  --url 'https://api.groundcover.com/api/monitors/list' \
  --header 'Authorization: Bearer <YOUR_API_KEY>' \
  --header 'Content-Type: application/json' \
  --data-raw '{"sources":[]}'
```

#### Response Example

```json
{
  "monitors": [
    {
      "uuid": "xxxx-xxxx-xxxx-xxxx-xxxx",
      "title": "PVC usage above threshold (90%)",
      "type": "metrics"
    },
    {
      "uuid": "xxxx-xxxx-xxxx-xxxx-xxxx",
      "title": "HTTP API Errors Monitor",
      "type": "traces"
    },
    {
      "uuid": "xxxxx-xxxx-xxxx-xxxx-xxxx",
      "title": "Error Logs Monitor",
      "type": "logs"
    },
    {
      "uuid": "xxxxx-xxxx-xxxx-xxxx-xxxx",
      "title": "Node CPU Usage Average is Above 85%",
      "type": "infra"
    },
    {
      "uuid": "xxxx-xxxx-xxxx-xxxx-xxxx",
      "title": "Rolling Update Triggered",
      "type": "events"
    },
    {
      "uuid": "xxxx-xxxx-xxxx-xxxx-xxxx",
      "title": "Deployment Partially Not Ready - 5m",
      "type": "events"
    }
  ]
}
```


# Get Monitor

Retrieve detailed configuration for a specific monitor by its UUID, including queries, thresholds, display settings, and evaluation parameters.

### Endpoint

**GET** `/api/monitors/{uuid}`

### Authentication

This endpoint requires API Key authentication via the Authorization header.

### Headers

| Header          | Required | Description                    |
| --------------- | -------- | ------------------------------ |
| `Authorization` | Yes      | Bearer token with your API key |
| `Content-Type`  | Yes      | Must be `application/json`     |
| `Accept`        | Yes      | Must be `application/json`     |

### Path Parameters

| Parameter | Type   | Required | Description                                      |
| --------- | ------ | -------- | ------------------------------------------------ |
| `uuid`    | string | Yes      | The unique identifier of the monitor to retrieve |

**Field Descriptions**

| Field                           | Type    | Description                                                        |
| ------------------------------- | ------- | ------------------------------------------------------------------ |
| `title`                         | string  | Monitor name/title                                                 |
| `display.header`                | string  | Alert header template with variable substitution                   |
| `display.resourceHeaderLabels`  | array   | Labels shown in resource headers                                   |
| `display.contextHeaderLabels`   | array   | Labels shown in context headers                                    |
| `display.description`           | string  | Monitor description                                                |
| `severity`                      | string  | Alert severity level (e.g., `"S1"`, `"S2"`, `"S3"`)                |
| `measurementType`               | string  | Type of measurement (`"state"`, `"event"`)                         |
| `model.queries`                 | array   | Query configurations for data retrieval                            |
| `model.thresholds`              | array   | Threshold configurations for alerting                              |
| `executionErrorState`           | string  | State when execution fails (`"OK"`, `"Alerting"`, `"Error"`)       |
| `noDataState`                   | string  | State when no data is available (`"OK"`, `"NoData"`, `"Alerting"`) |
| `evaluationInterval.interval`   | string  | How often to evaluate the monitor                                  |
| `evaluationInterval.pendingFor` | string  | How long to wait before alerting                                   |
| `isPaused`                      | boolean | Whether the monitor is currently paused                            |

### Examples

#### Basic Request

Get monitor configuration by UUID:

```bash
curl -L \
  --url 'https://api.groundcover.com/api/monitors/xxxx-xxxx-xxx-xxxx-xxxx' \
  --header 'Authorization: Bearer <YOUR_API_KEY>' \
  --header 'Accept: application/json'
```

#### Response Example - Metrics Monitor (MetricsQL, uses `rollup`)

```yaml
title: 'PVC usage above threshold (90%)'
display:
  header: PV usage above 90% threshold - {{ labels.cluster }}, {{ labels.name }}
  contextHeaderLabels:
  - cluster
  - namespace
  - env
severity: S2
measurementType: state
model:
  queries:
  - name: threshold_input_query
    dataType: metrics
    expression: >
      avg by (cluster, env, name, namespace) (
        groundcover_pvc_usage_percent{name!~"object-storage-cache-groundcover-incloud-clickhouse-shard.*"}
      )
    datasourceType: prometheus
    queryType: instant
    rollup:
      function: avg
      time: 5m
  thresholds:
  - name: threshold_1
    inputName: threshold_input_query
    operator: gt
    values:
    - 90
noDataState: OK
evaluationInterval:
  interval: 1m0s
  pendingFor: 1m0s
```

#### Response Example - Traces Monitor (gcQL, uses `instantRollup`)

```yaml
title: gRPC API Errors Monitor
display:
  header: gRPC API Error {{ labels.status_code }}
  resourceHeaderLabels:
  - span_name
  - role
  contextHeaderLabels:
  - env
  - cluster
  - namespace
  - workload
severity: S3
measurementType: event
model:
  queries:
  - name: threshold_input_query
    dataType: traces
    expression: >
      span_type:grpc status_code:!=0 status:error source:eBPF
      | stats by (env, cluster, namespace, workload, status_code, span_name, role) count() errors_total
    instantRollup: 1 minutes
  thresholds:
  - name: threshold_1
    inputName: threshold_input_query
    operator: gt
    values:
    - 0
executionErrorState: OK
noDataState: OK
evaluationInterval:
  interval: 1m
  pendingFor: 0s
```

The query language is [MetricsQL](https://docs.victoriametrics.com/metricsql/) for `dataType: metrics`/`infra` and [gcQL](/use-groundcover/querying-your-groundcover-data/groundcover-query-language/groundcover-query-language-gcql-reference) for `logs`/`traces`/`events`/`rum`/`entities`/`issues`. The two examples also show the two different rollup fields: `rollup` (object with `function` + `time`) is for MetricsQL, and `instantRollup` (a duration string) is for gcQL. See [Monitor YAML structure](/use-groundcover/monitors/monitor-yaml-structure) for the full schema reference.




---

[Next Page](/llms-full.txt/1)

