Set a retention period on every CloudWatch log group first: AWS keeps log data indefinitely by default, and stored logs are billed. For a small app, AWS logging and monitoring is five more moves after that: one CloudTrail trail, structured logs with a level, a few custom metrics, a short list of alarms, and a test that each alarm fires.

What AWS logging and monitoring is: four services, and the rest you can leave for later

AWS logging and monitoring is four native services doing four jobs: CloudWatch Logs stores what the app wrote, CloudWatch metrics count and time things, CloudWatch alarms tell a person, and CloudTrail records who did what in the account. The other services add to those four, and most small apps can leave them for later.

The three nouns are easy to blur. A log is a line of text the app wrote, with a time on it. A metric is “a time-ordered set of data points that are published to CloudWatch”, in AWS’s own definition. An alarm watches a single metric, or a math expression built on metrics, against a threshold and acts when its state changes.

The table lists the AWS logging and monitoring services by the job each does, with AWS’s own one-line description in the second column. The last column is my reading of what one small app should do with each this year.

ServiceWhat it records, in AWS’s wordsOn by default?What a small app does with it this year
CloudWatch Logs”helps you centralize the logs from all your systems, applications, and AWS services so you can monitor them and archive them securely”Lambda function logs arrive on their own when the execution role has the log permissionsKeep every line the app writes, with a retention value on each group
CloudWatch metricsCloudWatch “helps you analyze logs and, in real time, monitor the metrics of your AWS resources and hosted applications”Many AWS services provide metrics at no chargeAdd three or four business metrics of your own
CloudWatch alarmsA metric alarm “watches a single CloudWatch metric or the result of a math expression based on CloudWatch metrics”Only the alarms you createCreate the six alarms below
CloudTrail”recording the actions taken by a user, role, or an AWS service”Event history, the past 90 days of management eventsCreate one multi-Region trail
X-Ray”collects data about requests that your application serves”Lambda sends no traces until you turn on Active tracingLater, when one request crosses several services
VPC Flow Logs”captures information about the IP traffic going to and from network interfaces in your VPC”Not stated in AWS’s docsWhen a customer or an incident asks
GuardDuty”a threat detection service that continuously monitors, analyzes, and processes AWS data sources and logs”Starts ingesting its data sources when you enable itWhen a customer or an incident asks
Security Hub CSPMConsumes, aggregates, organizes and prioritizes “your security findings from multiple AWS services”When enabled without Security Hub, needs AWS Config enabled and resource recording turned onWhen a customer’s security review asks
AWS Config”provides a detailed view of the configuration of AWS resources in your AWS account”Not stated in AWS’s docsWhen an audit asks for configuration history

The AWS logging and monitoring architecture for one app, as I read how the pieces join, fits in one line: the app writes JSON lines to stdout, the platform ships them to a log group, metric filters or the embedded metric format turn some lines into metrics, alarms watch those metrics, one SNS topic carries alarm messages to a person, and CloudTrail writes the account’s own history to an S3 bucket. Monitoring and logging in AWS rarely needs more than that until a second team or a compliance review arrives.

The AWS logging and monitoring whitepaper that still turns up is the monitoring and logging page of “Introduction to AWS Security”, and AWS marks it: “This whitepaper is for historical reference only.” The current material is AWS’s guide to its logging and monitoring services, written for application owners. The AWS monitoring and logging tools above are the cloud version of the general list for any app’s logging and monitoring.

Why it matters for a small app: what AWS keeps, leaves off or does not record by default

A new AWS account comes with defaults that decide what exists on the day something breaks. Each row below comes from AWS’s own documentation.

DefaultWhat it means in practiceThe setting that changes it
”By default, log data is stored in CloudWatch Logs indefinitely”Every log group keeps growing, and stored data is billed, until you actA retention value per log group
Lambda sends function logs to /aws/lambda/<function-name> when the execution role has the log permissionsLogs exist with no setup, and like any log group they never expire until a retention is setA retention value on that group, set in the template that creates the function
The instance metrics EC2 lists in the AWS/EC2 namespace include CPU, disk operations and network, but no memory or disk-space metricA server running out of memory or disk shows nothing in the default metricsThe CloudWatch agent, which collects system-level metrics from inside the instance
An alarm invokes actions only when it changes state (Auto Scaling actions excepted), and CloudWatch “doesn’t test or validate the actions that you specify”A topic nobody subscribed to fails quietly, and an alarm that stays in ALARM sends its message onceA forced test of every alarm (test 4 below)
CloudTrail event history covers the past 90 days of management events and shows no data eventsAfter 90 days the account has no record, and object reads from S3 never appear in itA trail or an event data store
EKS control plane logs are not sent to CloudWatch Logs by defaultA cluster’s audit log does not exist until each log type is turned onControl plane logging, per log type, per cluster

The console shows the first row plainly: a log group with no retention reads “Never Expire” in its Retention column, and AWS’s page on CloudWatch Logs retention settings gives the click path to change it.

The reason I lead with defaults is what my audits found. In the apps I audited, 17 of the 21 third-party apps had no error tracking or alerting: when a user hits an error, nothing records it. The Deployment and Operations pillar averages 37.0 out of 100 across the same 21 apps, scored on all 21. Those are my June and July 2026 audits of 21 third-party apps, a selected set, not a random sample and not a rate for AI-built apps in general.

As I read it, an app exported from a builder arrives on AWS without the builder’s log viewer, so these defaults are the whole of its monitoring until someone changes them. The stack-agnostic list of what to watch, whatever the host, belongs to application monitoring best practices; this page is its AWS translation.

How it works: the settings, the alarms and the bill

Three parts make the setup work: eight settings, six alarms, and a bill you can predict.

AWS logging and monitoring best practices, sized for one small app

AWS logging and monitoring best practices for one small app come down, by my working rule, to 8 settings: retention on every log group, JSON log lines, a level chosen per environment, one multi-Region trail, a handful of custom metrics, alarms only on what needs a person, one notification topic, and a test of each alarm.

The list is my working rule, not an AWS standard.

#SettingWhere it is setThe mistake it prevents
1A retention value on every log groupIn the same template that creates the function or service, so a new group never misses itLogs kept and billed indefinitely by default
2JSON log lines with a request id, a level and a routeIn the app’s loggerLines that Logs Insights cannot query field by field
3A level chosen per environmentAn environment variable read by the logger, or Lambda’s logging configurationDebug output shipped and stored in production
4One multi-Region trail writing to its own S3 bucket, with public access blockedTerraform or the CloudTrail consoleNo account history beyond 90 days
5Three or four custom metrics that describe the business: sign-ups completed, payments failed, jobs waiting, model calls failedThe app, through one of the three routes belowAlarms that only see CPU while checkout fails
6Alarms only on what needs a personCloudWatch alarmsNoise that trains everyone to ignore the topic
7One SNS topic with a confirmed subscriptionAmazon SNSAlarms that fire into a topic nobody receives
8A test for every alarmset-alarm-state and a controlled failureAlarms that have never fired once

Row 1’s how-long question is a policy decision about how long to keep application logs; this page only gives the AWS setting. For row 2, Logs Insights reads JSON fields with dot notation, though for Lambda logs it discovers fields only in the first embedded JSON fragment of each event. What the event itself should carry, and what never goes into a log, is set out in error logging best practices for a solo AI-built app.

For row 3, AWS’s own guidance says “we recommend that you disable info and debug levels for production environments because these can generate excessive logging data”. That is AWS’s recommendation, quoted from AWS’s logging best practices for application owners; for what each level is for in your own app, the level table in the error logging article above is the one I follow.

For row 4, the trail and its bucket policy are in the Terraform block further down; new S3 buckets have all four Block Public Access settings on by default, so leave them on for that bucket. For row 7, an alarm’s actions show “Pending confirmation” until every email recipient on the topic has confirmed the subscription. Routing beyond email, to a channel the team reads, is a job for Slack alerting.

The six alarms worth the bill

The alarms worth paying for on a small AWS app number six, by my working rule: error rate, p95 latency, zero traffic in hours that should have some, database connections or CPU, dead-letter queue depth, and the monthly bill. Each one names a first step. Everything else is a dashboard line, not an alarm.

The starting rules are my working rule, written as starting points to tune, never as a standard.

AlarmThe metric behind itA starting ruleThe first step when it firesHow to test it
Error rate5xx responses from the load balancer, API Gateway’s 5XXError, or a custom metric, divided by requestsAbove about 2 percent of requests for 5 minutesOpen Logs Insights on the last 15 minutes of ERROR linesForce it with set-alarm-state, then return a 500 from a test route
p95 latencyThe p95 statistic of the load balancer, API Gateway, Lambda or RDS latency metricAbove about twice normal, or about 2 seconds, for 10 minutesFind the slowest route in the logsset-alarm-state, then add a delay to a test route
Zero trafficRequest count, with missing data treated as breachingNo requests for about 15 minutes in hours that normally have trafficCheck that the app answers from outsideset-alarm-state, then stop a test copy of the service
Database pressureRDS DatabaseConnections or CPUUtilizationAbove about 80 percent of the limit for 5 minutesLook for a connection leak or a slow queryset-alarm-state
Dead-letter queueApproximateNumberOfMessagesVisible on the dead-letter queueAbove zeroRead the failed message and its errorSend one message to a test queue’s dead-letter queue
Monthly billThe billing alarm or budgetYour monthly budget, set up as in the paragraph belowFind the meter that grewConfirm the alarm exists and its subscription is confirmed

p95 works because CloudWatch supports percentile statistics on Application Load Balancer, API Gateway, Lambda and Amazon RDS metrics, among others. The zero-traffic alarm needs a choice about missing data: CloudWatch offers four treatments (notBreaching, breaching, ignore and missing) and the default is missing, so a metric that stops reporting leaves the alarm in INSUFFICIENT_DATA instead of ALARM. Silence is the signal here, so treat missing data as breaching; AWS’s page on how CloudWatch alarms treat missing data shows where the option sits. To limit it to busy hours, wrap the metric in a metric math expression using the time functions HOUR and DAY, which return values based on each data point’s timestamp, in UTC. Without that condition the alarm also fires in the small hours of a quiet night.

For the dead-letter queue, AWS’s SQS documentation names ApproximateNumberOfMessagesVisible as the metric to watch, because messages moved there automatically are not counted in NumberOfMessagesSent. The bill alarm gets one clause here: it uses the same setup you would use to cap monthly usage on a metered API. Composite alarms, which watch the states of other alarms and can reduce noise, can wait until the six above are steady. Tuning each threshold after the first week is its own subject: how to alert on error rate spikes.

Testing without breaking production: set-alarm-state sets an alarm to any state, and that temporary state lasts only until the next alarm comparison, so you exercise the notification path once and the alarm returns on its own. CloudWatch keeps alarm history for 30 days, which is where the evidence of each test lives.

What drives the CloudWatch bill

CloudWatch cost for a small app comes mostly from 4 meters: log data ingested per GB, log data stored per GB-month, custom metrics per metric per month, and alarms per alarm metric. Debug logging left on in production grows the first two, and a high-cardinality dimension such as a request id grows the third.

These are the prices on the Amazon CloudWatch pricing page for the US East (Ohio) Region, Standard log class.

MeterUnitFree tier allowancePriceWhat makes it grow
Log ingestionper GB5 GB of data a month, shared by ingestion, archive storage and Logs Insights scans$0.50 per GBDebug and info lines in production
Log storage (archive)per GB compressed, per monthShared 5 GB above$0.03 per GB compressedNo retention value on the group
Logs Insightsper GB of data scannedShared 5 GB above$0.005 per GB scannedWide time ranges over many groups
Custom metricsper metric, per month10 metrics (custom and detailed monitoring)$0.30 per metric for the first 10,000Each unique dimension combination, and every agent metric
Standard alarmsper alarm metric, per month10 alarm metrics, for alarms that list metrics directly and do not use a Metrics Insights query$0.10 per alarm metricEvery metric listed in a math-expression alarm (up to 10 metrics using metric math)
High-resolution alarmsper alarm metric, per monthNone (the free alarms are standard resolution only)$0.30 per alarm metricPeriods of 10 or 30 seconds
API requestsper request1 million requests (GetMetricData, GetInsightRuleReport and GetMetricWidgetImage are always charged)Not one figure: the page’s text and its own worked example state the PutMetricData unit differentlyFrequent PutMetricData calls
Dashboardsper dashboard, per month3 custom dashboards of up to 50 metrics each$3.00 per dashboardExtra custom dashboards

Prices checked on 2026-10-03, US East (Ohio); AWS notes that pricing varies by Region.

What makes each meter grow comes from AWS’s own caveats. CloudWatch treats every unique combination of dimensions as a separate metric. The embedded metric format, used with a high-cardinality dimension such as a request id, will “by design create a custom metric corresponding to each unique dimension combination”. Metrics the CloudWatch agent collects are billed as custom metrics. And every PutMetricData call for a custom metric is charged, so publishing a high-resolution metric more often can cost more. My working rule follows from those: no debug level left on in production, dimensions with a handful of values (never a user id or request id), and retention set on every group.

A worked line for one small app, with my assumptions of 2 GB of logs a month, 4 custom metrics and 6 alarms that list 7 metrics between them (the error-rate alarm divides two): before the free tier, ingestion is 2 x $0.50 = $1.00, a month of stored logs is 2 GB x 0.15 compression = 0.3 GB x $0.03, about $0.01, metrics are 4 x $0.30 = $1.20, and alarms are 7 x $0.10 = $0.70, about $2.91 a month in all. Inside the free tier (5 GB of data, 10 custom metrics, 10 standard alarm metrics) every one of those lines is $0, as long as nothing else in the account uses up the allowance first. In that arithmetic, custom metrics and log ingestion are the two largest lines.

CloudWatch log levels, the agent, and writing a custom metric

These are the three things people look up after the overview: where a log level lives, whether a server needs the CloudWatch agent, and how a number from the app becomes a metric.

CloudWatch log level: where the level really comes from, and the Lambda logging level setting

A CloudWatch log level is not a CloudWatch setting: CloudWatch Logs stores whatever line it receives, and the level is a field your logger writes. Lambda is the exception, with application and system log level settings that can filter logs before they reach CloudWatch, and only when the function’s log format is JSON.

Filtering by level is then a Logs Insights query on that field, or a metric filter that counts matching lines as they arrive. The pair below is written from the syntax in AWS’s Logs Insights and Lambda documentation; the values are an example, not output from a real app.

{"timestamp":"2026-10-03T09:15:02.114Z","level":"ERROR","requestId":"req-7f3a","route":"POST /api/checkout","msg":"payment provider timeout"}
fields @timestamp, route, msg
| filter level = "ERROR"
| sort @timestamp desc

The AWS Lambda logging level settings can filter, in AWS’s words, “your function’s system logs (the logs that Lambda generates) and application logs (the logs that your function code generates)” separately. The function must use the JSON log format, and “The default log format for all Lambda managed runtimes is currently plain text.” Application levels run TRACE, DEBUG, INFO, WARN, ERROR and FATAL, with INFO the default; system levels are DEBUG, INFO and WARN, also defaulting to INFO. Output whose “level” field is invalid or missing is assigned INFO, and a level set in the function’s own code takes precedence over the console setting.

In the console: open the function, choose Monitoring and operations tools, then Edit in the Logging configuration pane, make sure Log format is JSON, and pick the two levels. From the CLI, aws lambda update-function-configuration takes --logging-config LogFormat=JSON,ApplicationLogLevel=ERROR,SystemLogLevel=WARN. The “environment variable” people ask about is the result, not the switch: behind the scenes Lambda sets the application log level in the runtime using AWS_LAMBDA_LOG_LEVEL, and the format using AWS_LAMBDA_LOG_FORMAT, which a custom runtime can read. AWS’s page on Lambda’s log-level filtering has both forms.

What each level should mean is a matter of logs severity levels; a Python app can get there with the Flask logger and one JSON format.

The amazon-cloudwatch-agent: what it is for, and when a small app needs it

The amazon-cloudwatch-agent is software you install on a server to collect system-level metrics beyond EC2’s basic monitoring and to ship log files to CloudWatch Logs. A small app on Lambda needs no agent: Lambda sends function logs to CloudWatch Logs itself when the execution role has the log permissions.

In the words of the CloudWatch agent documentation, it can collect internal system-level metrics from EC2 instances (which can include in-guest metrics), logs from EC2 instances and on-premises servers, and custom metrics over the StatsD and collectd protocols (collectd on Linux only), and from version 1.300025.0 it can collect traces from OpenTelemetry or X-Ray client SDKs. Its default namespace is CWAgent, its metrics are billed as custom metrics, and it does not collect logs from FIFO pipes.

The install runs in the order AWS gives:

  1. 01 Create an IAM role for the instance and attach the CloudWatchAgentServerPolicy managed policy.
  2. 02 Download the agent package.
  3. 03 Create or edit the agent configuration file with the metrics and log files you want.
  4. 04 Install and start the agent on the server.
  5. 05 Confirm the CWAgent namespace appears under Metrics and the log group you named appears under Log groups.

The first four steps are AWS’s; the confirm step is my check. If you would rather click than type, AWS documents a console route through Systems Manager: open Run Command, choose the AWS-ConfigureAWSPackage document, set Action to Install and Name to AmazonCloudWatchAgent, and keep Version at latest; the instance needs SSM Agent version 2.2.93.0 or later. You do not need the agent on Lambda, as above, nor on ECS with Fargate, where the awslogs log driver sends container output instead.

Custom CloudWatch metrics: PutMetricData, the embedded metric format and metric filters

A custom CloudWatch metric is a number your app publishes under its own namespace, in one of three ways: the PutMetricData API call, a JSON log line in the embedded metric format, or a metric filter that counts matching log lines. The two log-based routes keep the metric call out of the request path.

AWS’s own line comes first: “For new implementations, we recommend using OpenTelemetry to publish custom metrics.” The three routes below are the ones a small app already has without adding a collector.

RouteHow it worksWhat it costsWhen to use it
PutMetricData (CLI: aws cloudwatch put-metric-data)Sends a namespace, metric name, dimensions, value, unit and timestamp; standard resolution at one-minute granularity or high resolution at one secondEvery call for a custom metric is charged, plus the metric itselfOne-off tests and scripts
Embedded metric formatA JSON log event that CloudWatch turns into custom metrics asynchronously; needs logs:PutLogEvents, not cloudwatch:PutMetricData; delivery is at least onceLog ingestion and archival, plus the custom metrics it generatesLambda functions and containers already writing JSON logs
Metric filterTurns matching log data into a metric as it arrives; does not filter data retroactivelyThe custom metric it creates, plus the logs you already pay forCounting ERROR lines or a known message

The last column is my reading. A few rules apply to every custom CloudWatch metric whichever route you pick. There is no default namespace, and AWS’s own services use AWS/<service>, so the PutMetricData reference says you “should not specify a namespace that begins with AWS/”. Each unique combination of dimensions is a separate metric, with up to 30 dimensions on one metric. An alarm on a high-resolution metric can use a period of 10 or 30 seconds, at a higher charge. After the first data point, statistics can take up to 2 minutes to be retrievable, and the metric can take up to 15 minutes to appear in list-metrics. Which four business metrics to publish first is row 5 of the settings table above. On price, each custom metric costs $0.30 a month for the first 10,000 in US East (Ohio), after the 10 the free tier covers, as the bill table above lists it.

A custom metric example as a copy-pasteable pair, written from AWS’s publishing custom metrics page and the specification for the embedded metric format: the first line is the CloudWatch put metric data call for a one-off test, the second is the same number written as one log line. The timestamp is the one in AWS’s own example.

aws cloudwatch put-metric-data --namespace MyApp --metric-name PaymentsFailed --unit Count --value 1 --dimensions Service=checkout
{"_aws":{"Timestamp":1574109732004,"CloudWatchMetrics":[{"Namespace":"MyApp","Dimensions":[["Service"]],"Metrics":[{"Name":"PaymentsFailed","Unit":"Count"}]}]},"Service":"checkout","PaymentsFailed":1}

CloudTrail and container logging on AWS

Two more records matter on AWS: the account’s own history, and the logs of whatever runs in containers.

CloudTrail: the account’s audit trail, and the AWS CloudTrail Terraform resource

CloudTrail is the account’s audit log: who did what through the console, the CLI and the SDKs, and when. Event history keeps the past 90 days of management events in each Region, with no CloudTrail charge to view it, and a trail keeps an ongoing record in S3. One multi-Region trail suits a small account, as my working rule.

CloudTrail event history is on from the start: AWS calls it “a viewable, searchable, downloadable, and immutable record of the past 90 days of management events in an AWS Region”. It does not show data events, Insights events or network activity events, and for an ongoing record AWS says to create a trail or an event data store. In my reading of those limits, a small account still wants a trail for three reasons: history beyond 90 days, every Region in one place, and a bucket the app’s own credentials cannot delete from.

In Terraform, the aws_cloudtrail resource needs only name and s3_bucket_name. Two of its defaults matter: is_multi_region_trail (“Whether the trail is created in the current region or in all regions”) and enable_log_file_validation (“Whether log file integrity validation is enabled”) both default to false, so a small account sets both to true. include_global_service_events defaults to true and should stay that way: the provider’s own basic example sets it to false, while the same page says “For capturing events from services like IAM, include_global_service_events must be enabled.” That example also sets force_destroy = true on the bucket; my block leaves it out, as my working rule for a bucket that holds the account’s history. The block below follows the provider’s example and AWS’s CloudTrail bucket policy, with the aws:SourceArn condition AWS recommends as a security best practice.

data "aws_caller_identity" "current" {}
data "aws_partition" "current" {}
data "aws_region" "current" {}

locals {
  account   = data.aws_caller_identity.current.account_id
  trail_arn = "arn:${data.aws_partition.current.partition}:cloudtrail:${data.aws_region.current.region}:${local.account}:trail/account-trail"
}

resource "aws_s3_bucket" "trail" {
  bucket = "my-app-cloudtrail-logs"
}

resource "aws_s3_bucket_policy" "trail" {
  bucket = aws_s3_bucket.trail.id
  policy = jsonencode({
    Version = "2012-10-17"
    Statement = [
      { Sid = "AWSCloudTrailAclCheck", Effect = "Allow", Principal = { Service = "cloudtrail.amazonaws.com" },
        Action = "s3:GetBucketAcl", Resource = aws_s3_bucket.trail.arn,
        Condition = { StringEquals = { "aws:SourceArn" = local.trail_arn } } },
      { Sid = "AWSCloudTrailWrite", Effect = "Allow", Principal = { Service = "cloudtrail.amazonaws.com" },
        Action = "s3:PutObject", Resource = "${aws_s3_bucket.trail.arn}/AWSLogs/${local.account}/*",
        Condition = { StringEquals = { "s3:x-amz-acl" = "bucket-owner-full-control", "aws:SourceArn" = local.trail_arn } } }
    ]
  })
}

resource "aws_cloudtrail" "account" {
  depends_on                    = [aws_s3_bucket_policy.trail]
  name                          = "account-trail"
  s3_bucket_name                = aws_s3_bucket.trail.id
  is_multi_region_trail         = true
  enable_log_file_validation    = true
  include_global_service_events = true
}

To alarm on the trail, send its events to a CloudWatch Logs group as well: cloud_watch_logs_group_arn takes the log group’s ARN (CloudTrail requires the log stream wildcard, :*), and cloud_watch_logs_role_arn is the role CloudTrail assumes to write to it. What a trail still does not record by default is data events, such as S3 object reads (GetObject): “By default, trails and event data stores do not log data events. Additional charges apply for data events.”

One alarm earns its place here, and it comes from AWS’s own CloudTrail examples. With the trail delivering to a log group (the prerequisite AWS names), you create a metric filter that matches IAM policy changes and an alarm on it, at a threshold of 1 change event in 5 minutes, as the example sets it. The example’s console sign-in failure alarm, at 3 failures in 5 minutes, is the second one worth adding.

The FTC’s complaint against Chegg, announced on October 31, 2022, alleges that Chegg stored users’ personal data in Amazon S3 and “failed to adequately monitor its networks and systems for unauthorized attempts to transfer or exfiltrate users’ and employees’ personal information outside of Chegg’s network boundaries”. It says that in or around April 2018 a former contractor used an AWS Root Credential (the complaint’s term for a single access key with full administrative privileges over the S3 data, which Chegg shared among employees and outside contractors) to exfiltrate a database with personal information of approximately 40 million users, and that “Had Chegg employed reasonable access controls and monitoring, it would have likely detected and/or stopped the attack more quickly.” The complaint’s next dated event is September 2018, when “a threat intelligence vendor informed Chegg that a file containing some of the exfiltrated information was available in an online forum.”

The lesson I take from it is about order: an owner who keeps no account history and no alarm on it may first hear of a misused credential from someone else. The complaint does not say which AWS logging Chegg had or whether CloudTrail was on, and reads of S3 objects are data events that a trail does not log unless you configure it, at extra cost, so neither a default trail nor the IAM alarm above would by itself have caught this exfiltration.

Two AWS tools handle security monitoring and the evaluation of what the logging records. GuardDuty is a threat detection service whose foundational data sources are CloudTrail management events, VPC flow logs and DNS logs, ingested once you enable it. Security Hub CSPM aggregates and prioritizes findings, and it “uses AWS Config rules to run security checks and generate findings for most controls”, so with Security Hub CSPM enabled without Security Hub “you must manually enable AWS Config and turn on resource recording”. Deciding what to detect is a separate job, part of how to prevent insufficient logging and monitoring.

AWS EKS logging: control plane logs, pod logs, and the ECS and Fargate routes

AWS EKS logging has 2 halves, set up separately. The control plane has 5 log types, each off by default, sent to CloudWatch Logs only when you turn them on for the cluster. Pod logs are a separate setup, for example through CloudWatch Container Insights or, on Fargate, the log router, which uses AWS for Fluent Bit.

EKS control plane logging covers five types: api, audit, authenticator, controllerManager and scheduler. Each is enabled per cluster, logs land in the group /aws/eks/<cluster>/cluster, delivery is “best effort”, and CloudWatch Logs ingestion, archive storage and data scanning rates apply. In the console: open the cluster, choose the Observability tab, then Manage logging in the Control plane logging section. From the CLI, aws eks update-cluster-config takes a --logging argument with the list of types and "enabled":true.

For pods, CloudWatch Container Insights “collects, aggregates, and summarizes metrics and logs from your containerized applications and microservices”, and on Fargate the log router streams logs to AWS services or partner tools using AWS for Fluent Bit. The table puts EKS logging and monitoring on AWS next to the other places a small app’s containers or functions might run.

PlatformHow stdout reaches CloudWatch LogsWhat you must turn on
LambdaLambda sends function logs itself, to /aws/lambda/<function-name>An execution role with logs:CreateLogGroup, logs:CreateLogStream and logs:PutLogEvents
ECS on FargateThe awslogs log driver passes the container’s STDOUT and STDERR to CloudWatch LogsThe logConfiguration parameters in the task definition
ECS, logs going elsewhereFireLens routes logs to an AWS service or partner destination, with Fluentd or Fluent BitFireLens task definition parameters
EKS control planeSent to /aws/eks/<cluster>/clusterEach of the five log types, per cluster
EKS podsContainer Insights, or the Fargate log router with AWS for Fluent BitContainer Insights or the log router
App RunnerStreams application output to /aws/apprunner/<service-name>/<service-id>/applicationNot stated in AWS’s docs; App Runner is closed to new customers

On ECS with Fargate, the awslogs log driver works only once its logConfiguration parameters are in the task definition. The App Runner row matters for one reason: AWS says “AWS App Runner is no longer open to new customers”, so a new customer cannot start on it. As I read it, a small app rarely needs EKS, and if a template put it there, the two switches above (control plane log types and a pod log route) are the minimum. Monitoring EKS with Prometheus and Grafana is a different setup, starting from the Prometheus metric types.

How to check your own AWS account

An AWS monitoring setup is checked with 6 tests you can fail: every log group has a retention value, a test request is traceable by its id, a published custom metric appears, a forced alarm reaches a phone, a console change shows up in CloudTrail, and the bill alarm exists with a confirmed subscription.

Run them on an account you own. Each starts in the console, and the commands for four of them, written from the AWS CLI reference, follow the list.

  1. 01 Retention on every log group. In the CloudWatch console, open Logs, then Log groups, and read the Retention column: pass is no group showing Never Expire. Keep the dated output of the describe-log-groups command below.
  2. 02 One request, traced end to end. Send one request, take the request id your app wrote for it (from its first log line, or a response header if the app returns one), and search for it in Logs Insights in every log group it touched. Use the app's own request-id field: a platform's own request id is a different value. Lambda logs can take 5 to 10 minutes to show up, so wait before you search. Pass is one query that returns the whole path; keep the query and its result.
  3. 03 A test custom metric. Publish one data point with the put-metric-data command below, then retrieve it with get-metric-statistics after 2 minutes. Do not judge by the console graph: a new metric can take up to 15 minutes to appear in the metric list. Keep the get-metric-statistics output.
  4. 04 Every alarm, forced. Set each alarm to ALARM with set-alarm-state, confirm the message reaches a phone or an inbox, then confirm the alarm returns at its next comparison. Keep the alarm history entry and the message you received.
  5. 05 A console change in CloudTrail. Change a harmless setting in the console, such as a tag, then find the event in CloudTrail Event history, or with lookup-events filtered on your user name, in the same Region where you made the change. Keep the event record.
  6. 06 The bill alarm. Confirm it exists and, if it notifies through an SNS topic, that the subscription shows as confirmed and not Pending confirmation. Keep a screenshot of the subscription status.

The commands for tests 1, 3, 4 and 5, with names and times you replace:

# 1. A group with no "retentionInDays" in the output keeps its data indefinitely
aws logs describe-log-groups
aws logs put-retention-policy --log-group-name /aws/lambda/my-function --retention-in-days 30

# 3. Publish, wait 2 minutes, then read it back
aws cloudwatch put-metric-data --namespace MyApp --metric-name TestMetric --unit Count --value 1
aws cloudwatch get-metric-statistics --namespace MyApp --metric-name TestMetric --statistics Sum --period 60 --start-time 2026-10-03T09:00:00Z --end-time 2026-10-03T10:00:00Z

# 4. Force one alarm; it returns at the next comparison
aws cloudwatch set-alarm-state --alarm-name error-rate --state-value ALARM --state-reason "alarm test"

# 5. Events by your IAM user name, in the Region you changed
aws cloudtrail lookup-events --lookup-attributes AttributeKey=Username,AttributeValue=my-user

Event history records an event in the Region where it happened, which is why test 5 names the Region. Tests 1 and 4 sit first and fourth for a reason that is my working rule, not a statistic: they are the two people most often skip, and they are the two that turn a configured setup into a proven one.

Where the sprint fits

In area 08 of the Production Hardening Sprint, deliverable 8.1 adds structured request logs with correlation IDs and appropriate user references, excluding passwords, tokens and unnecessary personal data, and is verified by tracing a test request across services and checking log content for sensitive fields ; 8.3 alerts on error spikes, latency, connection pressure and queue backlog, verified by triggering each configured condition and recording its alert threshold and behavior ; 8.4 routes alerts to the designated Slack or email destination and tunes thresholds to reduce noise, verified by sending test alerts and checking their destination, context and response instructions ; and 8.5 sets and documents log retention across the application’s services, verified by inspecting the configured windows and verifying retention behavior in supported systems. Deliverable 13.1, the production readiness report, delivers the result for every scope item, the work completed and its verification evidence, accounting for all 123 IDs, keeping failures visible until resolved and explaining genuine non-applicable items. We start from your app’s current framework and hosting setup and refactor or replace components where the production work requires it ; hosting, paid tools and API usage remain in your accounts, and we explain any required third-party costs before enabling them. Each deliverable and its verify line is listed among the logging and monitoring items in the sprint’s scope.

Common questions about CloudWatch, CloudTrail and AWS logs

Is AWS CloudWatch metrics free?

Partly. Many AWS services publish their own metrics at no charge, while detailed monitoring and your own application metrics cost extra, and the CloudWatch agent’s metrics are billed as custom metrics. The free tier covers 10 metrics of custom and detailed monitoring metrics a month, as AWS’s pricing page stated it on 2026-10-03. Logs have a separate free allowance of 5 GB a month, which covers log ingestion, archived log storage and the data Logs Insights queries scan.

How long are metrics kept in CloudWatch?

Up to 455 days, at falling resolution. Data points under 60 seconds are kept for 3 hours, 1-minute data for 15 days, 5-minute data for 63 days, and 1-hour data for 455 days (15 months); shorter-period data is aggregated into the coarser resolution as it ages.

Is AWS CloudWatch a SIEM?

Not by AWS’s own description: its logging and monitoring guide describes CloudWatch as a service for analyzing logs and watching the metrics of AWS resources and hosted applications in real time, and does not call it a SIEM. For threat detection AWS offers GuardDuty, and Security Hub CSPM aggregates findings from several services.

Does Lambda automatically create a log group?

Yes, when the function’s execution role has logs:CreateLogGroup, logs:CreateLogStream and logs:PutLogEvents, Lambda captures its logs into /aws/lambda/<function-name> by default. That group keeps data indefinitely until you set a retention value on it.

What are the key differences between CloudTrail and GuardDuty?

CloudTrail records activity in the account, the actions taken by a user, role or AWS service; GuardDuty is a threat detection service that analyzes sources including CloudTrail management events, VPC flow logs and DNS logs. One keeps the record, and the other reads records like it for signs of an attack.