> For the complete documentation index, see [llms.txt](https://docs.groundcover.com/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://docs.groundcover.com/collect-data/data-sources/aws/ingest-cloudwatch-metrics.md).

# Ingest CloudWatch Metrics

groundcover supports ingesting CloudWatch metrics directly into our platform, allowing you to visualize them using dashboards and create monitors.

Before setting thresholds on `aws_*` series, see [Alerting on CloudWatch metrics](#alerting-on-cloudwatch-metrics) - the CloudWatch period and the polling lag both affect which monitor settings work.

## How does it work

CloudWatch integration is done by deploying a service called `integrations-agent` which is responsible for pulling metrics from CloudWatch using periodic polling of these APIs:

* [ListMetrics](https://docs.aws.amazon.com/AmazonCloudWatch/latest/APIReference/API_ListMetrics.html)
* [GetMetricData](https://docs.aws.amazon.com/AmazonCloudWatch/latest/APIReference/API_GetMetricData.html)
* [GetMetricStatistics](https://docs.aws.amazon.com/AmazonCloudWatch/latest/APIReference/API_GetMetricStatistics.html)

The integration setup is done directly through the App by following these steps:

1. Navigate to the Data Sources page or follow [this link](https://app.groundcover.com/data-sources). Note that only users with Admin permissions can navigate to this page.
2. Select Amazon Web Services and follow the wizard steps. Note that in step 1 you'll need to provide an ARN, granting groundcover with permissions to poll metrics. To do that, please follow the guidelines in [this section](#setting-up-the-integration).

## Things to know

### Ingestion interval

The integration pulls data from CloudWatch according to this interval. The lower the interval, the higher the polling rate and as a result, the overall costs will be higher. The interval can be altered in the Advanced Settings in the configuration wizard.

### Data storage

Data fetched is stored in the Victoria Metrics database, meaning metrics are queried via the CloudWatch API only one time per data point.

### Metric Statistics

Each metric has a label called `stat` which denotes the [AWS statistic ](https://docs.aws.amazon.com/AmazonCloudWatch/latest/monitoring/Statistics-definitions.html)used during querying. Some metrics have multiple stats which are useful for different cases.

### Supported AWS services

<details>

<summary>Click to Open</summary>

```json
/aws/sagemaker/Endpoints
/aws/sagemaker/ProcessingJobs
/aws/sagemaker/TrainingJobs
/aws/sagemaker/TransformJobs
AWS/ACMPrivateCA
AWS/AOSS
AWS/AmazonMQ
AWS/ApiGateway
AWS/AppRunner
AWS/AppStream
AWS/AppSync
AWS/ApplicationELB
AWS/Athena
AWS/AutoScaling
AWS/Backup
AWS/Bedrock
AWS/Bedrock/Agents
AWS/Bedrock/Guardrails
AWS/Billing
AWS/Cassandra
AWS/CertificateManager
AWS/ClientVPN
AWS/CloudFront
AWS/Cognito
AWS/DDoSProtection
AWS/DMS
AWS/DX
AWS/DataSync
AWS/DirectoryService
AWS/DocDB
AWS/DynamoDB
AWS/EBS
AWS/EC2
AWS/EC2Spot
AWS/ECR
AWS/ECS
AWS/EFS
AWS/EKS
AWS/ELB
AWS/EMRServerless
AWS/ES
AWS/ElastiCache
AWS/ElasticBeanstalk
AWS/ElasticMapReduce
AWS/Events
AWS/FSx
AWS/Firehose
AWS/GameLift
AWS/GatewayELB
AWS/GlobalAccelerator
AWS/IPAM
AWS/IoT
AWS/KMS
AWS/Kafka
AWS/KafkaConnect
AWS/Kinesis
AWS/KinesisAnalytics
AWS/Lambda
AWS/Logs
AWS/MWAA
AWS/MediaConnect
AWS/MediaConvert
AWS/MediaLive
AWS/MediaPackage
AWS/MediaTailor
AWS/MemoryDB
AWS/NATGateway
AWS/Neptune
AWS/Network Manager
AWS/NetworkELB
AWS/NetworkFirewall
AWS/PrivateLinkEndpoints
AWS/PrivateLinkServices
AWS/Prometheus
AWS/QuickSight
AWS/RDS
AWS/RUM
AWS/Redshift
AWS/Route53
AWS/S3
AWS/SES
AWS/SNS
AWS/SQS
AWS/SageMaker
AWS/Sagemaker/ModelBuildingPipeline
AWS/SecretsManager
AWS/States
AWS/StorageGateway
AWS/Timestream
AWS/Transfer
AWS/TransitGateway
AWS/TrustedAdvisor
AWS/Usage
AWS/VPN
AWS/WAFV2
AWS/WorkSpaces
AmazonMWAA
CWAgent
CloudWatchSynthetics
ContainerInsights
ECS/ContainerInsights
Glue
LambdaInsights
```

</details>

### Resource Discovery Methods

{% hint style="success" %}
groundcover seamlessly integrates both methods below to avoid duplicate metric fetching.
{% endhint %}

The integration uses two methods to discover the AWS resources to fetch metrics for:

1. Tagging-based discovery - this method uses the AWS tagging mechanism to discover resources across all metric namespaces.\
   This method supports all AWS namespaces but only works for resources which are tagged with at least one AWS tag.
2. List-based discovery - this method uses standard AWS APIs to list the resources in each namespace. It works for all resources regardless of tags, but the coverage is limited to specific namespaces as listed below:
   1. AWS/RDS
   2. AWS/S3
   3. AWS/SQS
   4. AWS/Lambda
   5. AWS/ElastiCache
   6. AWS/DynamoDB
   7. AWS/ELB
   8. AWS/NetworkELB
   9. AWS/ApplicationELB

{% hint style="info" %}
If you're not seeing metrics for a specific resource, it likely has no tags and is not in the list of services above. Contact us on Slack to help with resolving the issue.
{% endhint %}

### Scoping tagging-based discovery to specific resources

By default, tagging-based discovery picks up every tagged resource in a namespace. To narrow this down, add one or more **search tags** in the wizard's tag step: each is a tag key plus a value, where the value is matched as a **regular expression** rather than an exact string. Only resources with a tag matching both the key and the value pattern are scraped. For example, a tag key `team` with value pattern `platform.*` matches resources tagged `team=platform`, `team=platform-core`, and so on, but not `team=data`.

## Create an IAM role and policy

{% hint style="info" %}
This guide assumes your BYOC backend is set up on AWS. If it's not, please follow the guidelines in this [guide](/collect-data/data-sources/aws/adding-aws-integration-with-a-backend-on-another-cloud-provider.md).
{% endhint %}

Follow the below guidelines to set up a role manually, or configure it using CloudFormation using this [script](https://console.aws.amazon.com/cloudformation/home#/stacks/create/review?stackName=groundcover-integratios-agent\&templateURL=https://groundcover-public-cloudformation-templates.s3.us-east-1.amazonaws.com/integrations-agent/cloudformation.yaml).

{% hint style="info" %}
The following part requires two parameters:

* `YOUR_GROUNDCOVER_ACCOUNT_ID` - the AWS account id hosting the groundcover backend, [created during onboarding](/collect-data/byoc/setup-byoc-with-aws.md#step-1-allocate-an-aws-account-to-groundcover)
* `GROUNDCOVER_SITE_ID` - the groundcover site ID as extracted from your BYOC endpoint:
  * Fetch your `BYOC endpoint` from [these docs](/collect-data/byoc/ingestion-endpoints.md#fetching-the-byoc-endpoint)\
    It will look like `<SITE_ID>.platform.grcv.io`
  * The `GROUNDCOVER_SITE_ID` is the first part marked above as `<SITE_ID>`
  * For example, if your BYOC endpoint address is m234r1.platform.grcv.io, then the GROUNDCOVER\_SITE\_ID will be m234r1.
    {% endhint %}

1. Go to[ Amazon IAM](https://console.aws.amazon.com/iam/)
2. Click on **Roles** in the side bar
3. Click on **Create Role**
   1. Select **Custom trust policy**
   2. Paste the following policy:

      ```json
      {
        "Version": "2012-10-17",
        "Statement": [
          {
            "Effect": "Allow",
            "Principal": {
              "AWS": "arn:aws:iam::<YOUR_GROUNDCOVER_ACCOUNT_ID>:role/groundcover-integrations-agent-<GROUNDCOVER_SITE_ID>-sa"
            },
            "Action": "sts:AssumeRole"
          }
        ]
      }

      ```
   3. Click on **Next** twice (we'll attach permissions later)
   4. Provide a name for the role
   5. Click on **Create Role**
4. Go to your newly created role
   1. In the **Permissions** section, click on **Add permissions** and then **Create inline policy**
   2. Click on **JSON** and paste the following:

      ```json
      {
          "Version": "2012-10-17",
          "Id": "groundcover-integrations-agent",
          "Statement": [
              {
                  "Action": [
                      "tag:GetResources",
                      "storagegateway:ListTagsForResource",
                      "storagegateway:ListGateways",
                      "shield:ListProtections",
                      "iam:ListAccountAliases",
                      "ec2:DescribeTransitGatewayAttachments",
                      "ec2:DescribeSpotFleetRequests",
                      "dms:DescribeReplicationTasks",
                      "dms:DescribeReplicationInstances",
                      "cloudwatch:ListMetrics",
                      "cloudwatch:GetMetricStatistics",
                      "cloudwatch:GetMetricData",
                      "autoscaling:DescribeAutoScalingGroups",
                      "aps:ListWorkspaces",
                      "apigateway:GET",
                      "s3:ListAllMyBuckets",
                      "s3:GetBucketLocation",
                      "s3:GetBucketTagging",
                      "sqs:ListQueues",
                      "sqs:GetQueueAttributes",
                      "rds:DescribeDBInstances",
                      "rds:DescribeDBClusters",
                      "lambda:ListFunctions",
                      "elasticache:DescribeCacheClusters",
                      "elasticache:DescribeServerlessCaches",
                      "elasticloadbalancing:DescribeLoadBalancers",
                      "dynamodb:ListTables",
                      "dynamodb:ListTagsOfResource",
                      "dynamodb:DescribeTable",
                      "airflow:GetEnvironment",
                      "airflow:ListEnvironments",
                      "ecs:ListClusters",
                      "ecs:DescribeClusters",
                      "ecs:ListServices",
                      "ecs:DescribeServices",
                      "ecs:ListTasks",
                      "ecs:DescribeTasks",
                      "es:ListDomainNames",
                      "cloudfront:ListDistributions",
                      "kinesisanalytics:ListApplications",
                      "kinesisanalytics:ListTagsForResource"
                  ],
                  "Effect": "Allow",
                  "Resource": "*"
              }
          ]
      }
      ```
   3. Click on **Next**
   4. Give the policy a name
   5. Click on **Create Policy**

## Alerting on CloudWatch metrics

CloudWatch stamps each datapoint at the **start of the period** it covers - often 5 minutes - and the datapoint only reaches groundcover on the next poll, one [ingestion interval](#ingestion-interval) later. A monitor evaluates the last few minutes of data **at the moment it runs**, not the chart you look at afterwards, so monitors built on `aws_*` metrics need settings that account for both.

### Recommended settings

| Setting                 | Recommendation                                                                                                                                                                                                                                                                                                                                                                                                                                     |
| ----------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Evaluation delay**    | At least the metric's CloudWatch period plus the [ingestion interval](#ingestion-interval) set for the integration, so the datapoint has arrived by the time the monitor looks for it. `900` (15 minutes) is what the wizard prefills for metrics named `aws_*` and covers a 5-minute period at the default interval - raise it if you lengthened the ingestion interval. Monitors created from YAML, the API or Terraform must set it explicitly. |
| **Evaluation interval** | Match the CloudWatch period of the metric - 5 minutes for most, 1 minute where detailed monitoring is enabled.                                                                                                                                                                                                                                                                                                                                     |
| **Rollup window**       | Match the same period.                                                                                                                                                                                                                                                                                                                                                                                                                             |
| **Treat No Data As**    | Avoid **Normal** if you need to know when the evaluation window is empty.                                                                                                                                                                                                                                                                                                                                                                          |

### Match the evaluation interval to the CloudWatch period

CloudWatch datapoints are sparse compared to a 1-minute evaluation interval, which makes an interval shorter than the rollup window the default trap here: the same datapoint is counted on consecutive evaluations, and a `sum` rollup returns a multiple of the value CloudWatch reported - a metric that reported 30 evaluated as 150, with a threshold tuned against a number the metric never had. Matching the interval to the rollup window removes the duplication. See [Dashboard shows different values than monitor](/automate/monitors/create-a-new-monitor.md#dashboard-shows-different-values-than-monitor) for the general case.

A 1-minute chart of a 5-minute metric also stays flat between datapoints. That plateau is one value carried forward, not extra samples.

### Expect issues to lag the CloudWatch timestamp

With an evaluation delay of 15 minutes and no pending period, an issue opens roughly 15 minutes after the timestamp CloudWatch assigned to the datapoint that crossed the threshold, and proportionally later if a longer delay covers a longer ingestion interval. That gap is the cost of waiting for data that arrives late, and it is expected. A [pending period](/automate/monitors/create-a-new-monitor.md#section-1-query) adds its own duration on top.

### Decide what an empty window means

A poll that lands later than usual leaves the evaluation window empty. With **Treat No Data As** set to **Normal** that is indistinguishable from a healthy evaluation and the monitor stays green, so a missed datapoint passes silently. See [Treat No Data As](/automate/monitors/create-a-new-monitor.md#section-1-query) for the alternatives and [Alerting on No Data](/automate/monitors/notification-routes.md#alerting-on-no-data) for notifying on them.

No Data applies to the query as a whole. A query grouped by resource keeps returning rows while any one resource still reports, so a single load balancer that stops publishing is not a No Data evaluation.

### Example monitor

CloudWatch metric names are ingested as `aws_<namespace>_<metric>` in lower snake case, and each series carries a `stat` label for the [AWS statistic](https://docs.aws.amazon.com/AmazonCloudWatch/latest/monitoring/Statistics-definitions.html) it was queried with. Filter on `stat` so a metric with multiple statistics does not return several series per resource.

```yaml
title: ALB 502 Errors
display:
  header: ALB 502 errors on {{ labels.load_balancer }}
  description: The load balancer returned 502 responses in the last 5 minutes.
  resourceHeaderLabels:
    - load_balancer
  contextHeaderLabels:
    - region
    - account_id
severity: S2
measurementType: event
model:
  queries:
    - name: elb_502_query
      expression: sum(aws_applicationelb_httpcode_elb_502_count{stat="sum"}) by (load_balancer, region, account_id)
      datasourceType: prometheus
      queryType: instant
      rollup:
        function: sum
        time: 5m
      evaluationDelay: 900
  thresholds:
    - name: threshold_1
      inputName: elb_502_query
      operator: gt
      values:
        - 0
executionErrorState: OK
noDataState: NoData
evaluationInterval:
  interval: 5m
  pendingFor: 0s
```

See [Monitor YAML structure](/automate/monitors/monitor-yaml-structure.md) for the full schema.


---

# Agent Instructions
This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com.

## Querying This Documentation
If you need additional information that is not directly available in this page, you can query the documentation by asking a question.

Perform an HTTP GET request on the following URL with the `ask` and `goal` query parameters:

```
GET https://docs.groundcover.com/collect-data/data-sources/aws/ingest-cloudwatch-metrics.md?ask=<question>&goal=<user_goal>
```

`ask` is the immediate question: it should be specific, self-contained, and written in natural language.
`goal` is what the user is ultimately trying to achieve, the reason they need the answer. Sharing it helps GitBook give you a better, more relevant answer. A goal is most helpful when it describes the outcome the user wants rather than restating the question. For example, with `ask=how do I create an API token`, a goal like `automate deployments from our CI pipeline` lets GitBook tailor the answer to that use case.

The response will contain a direct answer to the question and relevant excerpts and sources from the documentation.

Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections.
