Set a retention period on every CloudWatch log group first: AWS keeps log data indefinitely by default, and stored logs are billed. For a small app, AWS logging and monitoring is five more moves after that: one CloudTrail trail, structured logs with a level, a few custom metrics, a short list of alarms, and a test that each alarm fires.
What AWS logging and monitoring is: four services, and the rest you can leave for later
AWS logging and monitoring is four native services doing four jobs: CloudWatch Logs stores what the app wrote, CloudWatch metrics count and time things, CloudWatch alarms tell a person, and CloudTrail records who did what in the account. The other services add to those four, and most small apps can leave them for later.
The three nouns are easy to blur. A log is a line of text the app wrote, with a time on it. A metric is “a time-ordered set of data points that are published to CloudWatch”, in AWS’s own definition. An alarm watches a single metric, or a math expression built on metrics, against a threshold and acts when its state changes.
The table lists the AWS logging and monitoring services by the job each does, with AWS’s own one-line description in the second column. The last column is my reading of what one small app should do with each this year.
| Service | What it records, in AWS’s words | On by default? | What a small app does with it this year |
|---|---|---|---|
| CloudWatch Logs | ”helps you centralize the logs from all your systems, applications, and AWS services so you can monitor them and archive them securely” | Lambda function logs arrive on their own when the execution role has the log permissions | Keep every line the app writes, with a retention value on each group |
| CloudWatch metrics | CloudWatch “helps you analyze logs and, in real time, monitor the metrics of your AWS resources and hosted applications” | Many AWS services provide metrics at no charge | Add three or four business metrics of your own |
| CloudWatch alarms | A metric alarm “watches a single CloudWatch metric or the result of a math expression based on CloudWatch metrics” | Only the alarms you create | Create the six alarms below |
| CloudTrail | ”recording the actions taken by a user, role, or an AWS service” | Event history, the past 90 days of management events | Create one multi-Region trail |
| X-Ray | ”collects data about requests that your application serves” | Lambda sends no traces until you turn on Active tracing | Later, when one request crosses several services |
| VPC Flow Logs | ”captures information about the IP traffic going to and from network interfaces in your VPC” | Not stated in AWS’s docs | When a customer or an incident asks |
| GuardDuty | ”a threat detection service that continuously monitors, analyzes, and processes AWS data sources and logs” | Starts ingesting its data sources when you enable it | When a customer or an incident asks |
| Security Hub CSPM | Consumes, aggregates, organizes and prioritizes “your security findings from multiple AWS services” | When enabled without Security Hub, needs AWS Config enabled and resource recording turned on | When a customer’s security review asks |
| AWS Config | ”provides a detailed view of the configuration of AWS resources in your AWS account” | Not stated in AWS’s docs | When an audit asks for configuration history |
The AWS logging and monitoring architecture for one app, as I read how the pieces join, fits in one line: the app writes JSON lines to stdout, the platform ships them to a log group, metric filters or the embedded metric format turn some lines into metrics, alarms watch those metrics, one SNS topic carries alarm messages to a person, and CloudTrail writes the account’s own history to an S3 bucket. Monitoring and logging in AWS rarely needs more than that until a second team or a compliance review arrives.
The AWS logging and monitoring whitepaper that still turns up is the monitoring and logging page of “Introduction to AWS Security”, and AWS marks it: “This whitepaper is for historical reference only.” The current material is AWS’s guide to its logging and monitoring services, written for application owners. The AWS monitoring and logging tools above are the cloud version of the general list for any app’s logging and monitoring.
Why it matters for a small app: what AWS keeps, leaves off or does not record by default
A new AWS account comes with defaults that decide what exists on the day something breaks. Each row below comes from AWS’s own documentation.
| Default | What it means in practice | The setting that changes it |
|---|---|---|
| ”By default, log data is stored in CloudWatch Logs indefinitely” | Every log group keeps growing, and stored data is billed, until you act | A retention value per log group |
Lambda sends function logs to /aws/lambda/<function-name> when the execution role has the log permissions | Logs exist with no setup, and like any log group they never expire until a retention is set | A retention value on that group, set in the template that creates the function |
The instance metrics EC2 lists in the AWS/EC2 namespace include CPU, disk operations and network, but no memory or disk-space metric | A server running out of memory or disk shows nothing in the default metrics | The CloudWatch agent, which collects system-level metrics from inside the instance |
| An alarm invokes actions only when it changes state (Auto Scaling actions excepted), and CloudWatch “doesn’t test or validate the actions that you specify” | A topic nobody subscribed to fails quietly, and an alarm that stays in ALARM sends its message once | A forced test of every alarm (test 4 below) |
| CloudTrail event history covers the past 90 days of management events and shows no data events | After 90 days the account has no record, and object reads from S3 never appear in it | A trail or an event data store |
| EKS control plane logs are not sent to CloudWatch Logs by default | A cluster’s audit log does not exist until each log type is turned on | Control plane logging, per log type, per cluster |
The console shows the first row plainly: a log group with no retention reads “Never Expire” in its Retention column, and AWS’s page on CloudWatch Logs retention settings gives the click path to change it.
The reason I lead with defaults is what my audits found. In the apps I audited, 17 of the 21 third-party apps had no error tracking or alerting: when a user hits an error, nothing records it. The Deployment and Operations pillar averages 37.0 out of 100 across the same 21 apps, scored on all 21. Those are my June and July 2026 audits of 21 third-party apps, a selected set, not a random sample and not a rate for AI-built apps in general.
As I read it, an app exported from a builder arrives on AWS without the builder’s log viewer, so these defaults are the whole of its monitoring until someone changes them. The stack-agnostic list of what to watch, whatever the host, belongs to application monitoring best practices; this page is its AWS translation.
How it works: the settings, the alarms and the bill
Three parts make the setup work: eight settings, six alarms, and a bill you can predict.
AWS logging and monitoring best practices, sized for one small app
AWS logging and monitoring best practices for one small app come down, by my working rule, to 8 settings: retention on every log group, JSON log lines, a level chosen per environment, one multi-Region trail, a handful of custom metrics, alarms only on what needs a person, one notification topic, and a test of each alarm.
The list is my working rule, not an AWS standard.
| # | Setting | Where it is set | The mistake it prevents |
|---|---|---|---|
| 1 | A retention value on every log group | In the same template that creates the function or service, so a new group never misses it | Logs kept and billed indefinitely by default |
| 2 | JSON log lines with a request id, a level and a route | In the app’s logger | Lines that Logs Insights cannot query field by field |
| 3 | A level chosen per environment | An environment variable read by the logger, or Lambda’s logging configuration | Debug output shipped and stored in production |
| 4 | One multi-Region trail writing to its own S3 bucket, with public access blocked | Terraform or the CloudTrail console | No account history beyond 90 days |
| 5 | Three or four custom metrics that describe the business: sign-ups completed, payments failed, jobs waiting, model calls failed | The app, through one of the three routes below | Alarms that only see CPU while checkout fails |
| 6 | Alarms only on what needs a person | CloudWatch alarms | Noise that trains everyone to ignore the topic |
| 7 | One SNS topic with a confirmed subscription | Amazon SNS | Alarms that fire into a topic nobody receives |
| 8 | A test for every alarm | set-alarm-state and a controlled failure | Alarms that have never fired once |
Row 1’s how-long question is a policy decision about how long to keep application logs; this page only gives the AWS setting. For row 2, Logs Insights reads JSON fields with dot notation, though for Lambda logs it discovers fields only in the first embedded JSON fragment of each event. What the event itself should carry, and what never goes into a log, is set out in error logging best practices for a solo AI-built app.
For row 3, AWS’s own guidance says “we recommend that you disable info and debug levels for production environments because these can generate excessive logging data”. That is AWS’s recommendation, quoted from AWS’s logging best practices for application owners; for what each level is for in your own app, the level table in the error logging article above is the one I follow.
For row 4, the trail and its bucket policy are in the Terraform block further down; new S3 buckets have all four Block Public Access settings on by default, so leave them on for that bucket. For row 7, an alarm’s actions show “Pending confirmation” until every email recipient on the topic has confirmed the subscription. Routing beyond email, to a channel the team reads, is a job for Slack alerting.
The six alarms worth the bill
The alarms worth paying for on a small AWS app number six, by my working rule: error rate, p95 latency, zero traffic in hours that should have some, database connections or CPU, dead-letter queue depth, and the monthly bill. Each one names a first step. Everything else is a dashboard line, not an alarm.
The starting rules are my working rule, written as starting points to tune, never as a standard.
| Alarm | The metric behind it | A starting rule | The first step when it fires | How to test it |
|---|---|---|---|---|
| Error rate | 5xx responses from the load balancer, API Gateway’s 5XXError, or a custom metric, divided by requests | Above about 2 percent of requests for 5 minutes | Open Logs Insights on the last 15 minutes of ERROR lines | Force it with set-alarm-state, then return a 500 from a test route |
| p95 latency | The p95 statistic of the load balancer, API Gateway, Lambda or RDS latency metric | Above about twice normal, or about 2 seconds, for 10 minutes | Find the slowest route in the logs | set-alarm-state, then add a delay to a test route |
| Zero traffic | Request count, with missing data treated as breaching | No requests for about 15 minutes in hours that normally have traffic | Check that the app answers from outside | set-alarm-state, then stop a test copy of the service |
| Database pressure | RDS DatabaseConnections or CPUUtilization | Above about 80 percent of the limit for 5 minutes | Look for a connection leak or a slow query | set-alarm-state |
| Dead-letter queue | ApproximateNumberOfMessagesVisible on the dead-letter queue | Above zero | Read the failed message and its error | Send one message to a test queue’s dead-letter queue |
| Monthly bill | The billing alarm or budget | Your monthly budget, set up as in the paragraph below | Find the meter that grew | Confirm the alarm exists and its subscription is confirmed |
p95 works because CloudWatch supports percentile statistics on Application Load Balancer, API Gateway, Lambda and Amazon RDS metrics, among others. The zero-traffic alarm needs a choice about missing data: CloudWatch offers four treatments (notBreaching, breaching, ignore and missing) and the default is missing, so a metric that stops reporting leaves the alarm in INSUFFICIENT_DATA instead of ALARM. Silence is the signal here, so treat missing data as breaching; AWS’s page on how CloudWatch alarms treat missing data shows where the option sits. To limit it to busy hours, wrap the metric in a metric math expression using the time functions HOUR and DAY, which return values based on each data point’s timestamp, in UTC. Without that condition the alarm also fires in the small hours of a quiet night.
For the dead-letter queue, AWS’s SQS documentation names ApproximateNumberOfMessagesVisible as the metric to watch, because messages moved there automatically are not counted in NumberOfMessagesSent. The bill alarm gets one clause here: it uses the same setup you would use to cap monthly usage on a metered API. Composite alarms, which watch the states of other alarms and can reduce noise, can wait until the six above are steady. Tuning each threshold after the first week is its own subject: how to alert on error rate spikes.
Testing without breaking production: set-alarm-state sets an alarm to any state, and that temporary state lasts only until the next alarm comparison, so you exercise the notification path once and the alarm returns on its own. CloudWatch keeps alarm history for 30 days, which is where the evidence of each test lives.
What drives the CloudWatch bill
CloudWatch cost for a small app comes mostly from 4 meters: log data ingested per GB, log data stored per GB-month, custom metrics per metric per month, and alarms per alarm metric. Debug logging left on in production grows the first two, and a high-cardinality dimension such as a request id grows the third.
These are the prices on the Amazon CloudWatch pricing page for the US East (Ohio) Region, Standard log class.
| Meter | Unit | Free tier allowance | Price | What makes it grow |
|---|---|---|---|---|
| Log ingestion | per GB | 5 GB of data a month, shared by ingestion, archive storage and Logs Insights scans | $0.50 per GB | Debug and info lines in production |
| Log storage (archive) | per GB compressed, per month | Shared 5 GB above | $0.03 per GB compressed | No retention value on the group |
| Logs Insights | per GB of data scanned | Shared 5 GB above | $0.005 per GB scanned | Wide time ranges over many groups |
| Custom metrics | per metric, per month | 10 metrics (custom and detailed monitoring) | $0.30 per metric for the first 10,000 | Each unique dimension combination, and every agent metric |
| Standard alarms | per alarm metric, per month | 10 alarm metrics, for alarms that list metrics directly and do not use a Metrics Insights query | $0.10 per alarm metric | Every metric listed in a math-expression alarm (up to 10 metrics using metric math) |
| High-resolution alarms | per alarm metric, per month | None (the free alarms are standard resolution only) | $0.30 per alarm metric | Periods of 10 or 30 seconds |
| API requests | per request | 1 million requests (GetMetricData, GetInsightRuleReport and GetMetricWidgetImage are always charged) | Not one figure: the page’s text and its own worked example state the PutMetricData unit differently | Frequent PutMetricData calls |
| Dashboards | per dashboard, per month | 3 custom dashboards of up to 50 metrics each | $3.00 per dashboard | Extra custom dashboards |
Prices checked on 2026-10-03, US East (Ohio); AWS notes that pricing varies by Region.
What makes each meter grow comes from AWS’s own caveats. CloudWatch treats every unique combination of dimensions as a separate metric. The embedded metric format, used with a high-cardinality dimension such as a request id, will “by design create a custom metric corresponding to each unique dimension combination”. Metrics the CloudWatch agent collects are billed as custom metrics. And every PutMetricData call for a custom metric is charged, so publishing a high-resolution metric more often can cost more. My working rule follows from those: no debug level left on in production, dimensions with a handful of values (never a user id or request id), and retention set on every group.
A worked line for one small app, with my assumptions of 2 GB of logs a month, 4 custom metrics and 6 alarms that list 7 metrics between them (the error-rate alarm divides two): before the free tier, ingestion is 2 x $0.50 = $1.00, a month of stored logs is 2 GB x 0.15 compression = 0.3 GB x $0.03, about $0.01, metrics are 4 x $0.30 = $1.20, and alarms are 7 x $0.10 = $0.70, about $2.91 a month in all. Inside the free tier (5 GB of data, 10 custom metrics, 10 standard alarm metrics) every one of those lines is $0, as long as nothing else in the account uses up the allowance first. In that arithmetic, custom metrics and log ingestion are the two largest lines.
CloudWatch log levels, the agent, and writing a custom metric
These are the three things people look up after the overview: where a log level lives, whether a server needs the CloudWatch agent, and how a number from the app becomes a metric.
CloudWatch log level: where the level really comes from, and the Lambda logging level setting
A CloudWatch log level is not a CloudWatch setting: CloudWatch Logs stores whatever line it receives, and the level is a field your logger writes. Lambda is the exception, with application and system log level settings that can filter logs before they reach CloudWatch, and only when the function’s log format is JSON.
Filtering by level is then a Logs Insights query on that field, or a metric filter that counts matching lines as they arrive. The pair below is written from the syntax in AWS’s Logs Insights and Lambda documentation; the values are an example, not output from a real app.
{"timestamp":"2026-10-03T09:15:02.114Z","level":"ERROR","requestId":"req-7f3a","route":"POST /api/checkout","msg":"payment provider timeout"}
fields @timestamp, route, msg
| filter level = "ERROR"
| sort @timestamp desc
The AWS Lambda logging level settings can filter, in AWS’s words, “your function’s system logs (the logs that Lambda generates) and application logs (the logs that your function code generates)” separately. The function must use the JSON log format, and “The default log format for all Lambda managed runtimes is currently plain text.” Application levels run TRACE, DEBUG, INFO, WARN, ERROR and FATAL, with INFO the default; system levels are DEBUG, INFO and WARN, also defaulting to INFO. Output whose “level” field is invalid or missing is assigned INFO, and a level set in the function’s own code takes precedence over the console setting.
In the console: open the function, choose Monitoring and operations tools, then Edit in the Logging configuration pane, make sure Log format is JSON, and pick the two levels. From the CLI, aws lambda update-function-configuration takes --logging-config LogFormat=JSON,ApplicationLogLevel=ERROR,SystemLogLevel=WARN. The “environment variable” people ask about is the result, not the switch: behind the scenes Lambda sets the application log level in the runtime using AWS_LAMBDA_LOG_LEVEL, and the format using AWS_LAMBDA_LOG_FORMAT, which a custom runtime can read. AWS’s page on Lambda’s log-level filtering has both forms.
What each level should mean is a matter of logs severity levels; a Python app can get there with the Flask logger and one JSON format.
The amazon-cloudwatch-agent: what it is for, and when a small app needs it
The amazon-cloudwatch-agent is software you install on a server to collect system-level metrics beyond EC2’s basic monitoring and to ship log files to CloudWatch Logs. A small app on Lambda needs no agent: Lambda sends function logs to CloudWatch Logs itself when the execution role has the log permissions.
In the words of the CloudWatch agent documentation, it can collect internal system-level metrics from EC2 instances (which can include in-guest metrics), logs from EC2 instances and on-premises servers, and custom metrics over the StatsD and collectd protocols (collectd on Linux only), and from version 1.300025.0 it can collect traces from OpenTelemetry or X-Ray client SDKs. Its default namespace is CWAgent, its metrics are billed as custom metrics, and it does not collect logs from FIFO pipes.
The install runs in the order AWS gives:
- 01 Create an IAM role for the instance and attach the CloudWatchAgentServerPolicy managed policy.
- 02 Download the agent package.
- 03 Create or edit the agent configuration file with the metrics and log files you want.
- 04 Install and start the agent on the server.
- 05 Confirm the CWAgent namespace appears under Metrics and the log group you named appears under Log groups.
The first four steps are AWS’s; the confirm step is my check. If you would rather click than type, AWS documents a console route through Systems Manager: open Run Command, choose the AWS-ConfigureAWSPackage document, set Action to Install and Name to AmazonCloudWatchAgent, and keep Version at latest; the instance needs SSM Agent version 2.2.93.0 or later. You do not need the agent on Lambda, as above, nor on ECS with Fargate, where the awslogs log driver sends container output instead.
Custom CloudWatch metrics: PutMetricData, the embedded metric format and metric filters
A custom CloudWatch metric is a number your app publishes under its own namespace, in one of three ways: the PutMetricData API call, a JSON log line in the embedded metric format, or a metric filter that counts matching log lines. The two log-based routes keep the metric call out of the request path.
AWS’s own line comes first: “For new implementations, we recommend using OpenTelemetry to publish custom metrics.” The three routes below are the ones a small app already has without adding a collector.
| Route | How it works | What it costs | When to use it |
|---|---|---|---|
PutMetricData (CLI: aws cloudwatch put-metric-data) | Sends a namespace, metric name, dimensions, value, unit and timestamp; standard resolution at one-minute granularity or high resolution at one second | Every call for a custom metric is charged, plus the metric itself | One-off tests and scripts |
| Embedded metric format | A JSON log event that CloudWatch turns into custom metrics asynchronously; needs logs:PutLogEvents, not cloudwatch:PutMetricData; delivery is at least once | Log ingestion and archival, plus the custom metrics it generates | Lambda functions and containers already writing JSON logs |
| Metric filter | Turns matching log data into a metric as it arrives; does not filter data retroactively | The custom metric it creates, plus the logs you already pay for | Counting ERROR lines or a known message |
The last column is my reading. A few rules apply to every custom CloudWatch metric whichever route you pick. There is no default namespace, and AWS’s own services use AWS/<service>, so the PutMetricData reference says you “should not specify a namespace that begins with AWS/”. Each unique combination of dimensions is a separate metric, with up to 30 dimensions on one metric. An alarm on a high-resolution metric can use a period of 10 or 30 seconds, at a higher charge. After the first data point, statistics can take up to 2 minutes to be retrievable, and the metric can take up to 15 minutes to appear in list-metrics. Which four business metrics to publish first is row 5 of the settings table above. On price, each custom metric costs $0.30 a month for the first 10,000 in US East (Ohio), after the 10 the free tier covers, as the bill table above lists it.
A custom metric example as a copy-pasteable pair, written from AWS’s publishing custom metrics page and the specification for the embedded metric format: the first line is the CloudWatch put metric data call for a one-off test, the second is the same number written as one log line. The timestamp is the one in AWS’s own example.
aws cloudwatch put-metric-data --namespace MyApp --metric-name PaymentsFailed --unit Count --value 1 --dimensions Service=checkout
{"_aws":{"Timestamp":1574109732004,"CloudWatchMetrics":[{"Namespace":"MyApp","Dimensions":[["Service"]],"Metrics":[{"Name":"PaymentsFailed","Unit":"Count"}]}]},"Service":"checkout","PaymentsFailed":1}
CloudTrail and container logging on AWS
Two more records matter on AWS: the account’s own history, and the logs of whatever runs in containers.
CloudTrail: the account’s audit trail, and the AWS CloudTrail Terraform resource
CloudTrail is the account’s audit log: who did what through the console, the CLI and the SDKs, and when. Event history keeps the past 90 days of management events in each Region, with no CloudTrail charge to view it, and a trail keeps an ongoing record in S3. One multi-Region trail suits a small account, as my working rule.
CloudTrail event history is on from the start: AWS calls it “a viewable, searchable, downloadable, and immutable record of the past 90 days of management events in an AWS Region”. It does not show data events, Insights events or network activity events, and for an ongoing record AWS says to create a trail or an event data store. In my reading of those limits, a small account still wants a trail for three reasons: history beyond 90 days, every Region in one place, and a bucket the app’s own credentials cannot delete from.
In Terraform, the aws_cloudtrail resource needs only name and s3_bucket_name. Two of its defaults matter: is_multi_region_trail (“Whether the trail is created in the current region or in all regions”) and enable_log_file_validation (“Whether log file integrity validation is enabled”) both default to false, so a small account sets both to true. include_global_service_events defaults to true and should stay that way: the provider’s own basic example sets it to false, while the same page says “For capturing events from services like IAM, include_global_service_events must be enabled.” That example also sets force_destroy = true on the bucket; my block leaves it out, as my working rule for a bucket that holds the account’s history. The block below follows the provider’s example and AWS’s CloudTrail bucket policy, with the aws:SourceArn condition AWS recommends as a security best practice.
data "aws_caller_identity" "current" {}
data "aws_partition" "current" {}
data "aws_region" "current" {}
locals {
account = data.aws_caller_identity.current.account_id
trail_arn = "arn:${data.aws_partition.current.partition}:cloudtrail:${data.aws_region.current.region}:${local.account}:trail/account-trail"
}
resource "aws_s3_bucket" "trail" {
bucket = "my-app-cloudtrail-logs"
}
resource "aws_s3_bucket_policy" "trail" {
bucket = aws_s3_bucket.trail.id
policy = jsonencode({
Version = "2012-10-17"
Statement = [
{ Sid = "AWSCloudTrailAclCheck", Effect = "Allow", Principal = { Service = "cloudtrail.amazonaws.com" },
Action = "s3:GetBucketAcl", Resource = aws_s3_bucket.trail.arn,
Condition = { StringEquals = { "aws:SourceArn" = local.trail_arn } } },
{ Sid = "AWSCloudTrailWrite", Effect = "Allow", Principal = { Service = "cloudtrail.amazonaws.com" },
Action = "s3:PutObject", Resource = "${aws_s3_bucket.trail.arn}/AWSLogs/${local.account}/*",
Condition = { StringEquals = { "s3:x-amz-acl" = "bucket-owner-full-control", "aws:SourceArn" = local.trail_arn } } }
]
})
}
resource "aws_cloudtrail" "account" {
depends_on = [aws_s3_bucket_policy.trail]
name = "account-trail"
s3_bucket_name = aws_s3_bucket.trail.id
is_multi_region_trail = true
enable_log_file_validation = true
include_global_service_events = true
}
To alarm on the trail, send its events to a CloudWatch Logs group as well: cloud_watch_logs_group_arn takes the log group’s ARN (CloudTrail requires the log stream wildcard, :*), and cloud_watch_logs_role_arn is the role CloudTrail assumes to write to it. What a trail still does not record by default is data events, such as S3 object reads (GetObject): “By default, trails and event data stores do not log data events. Additional charges apply for data events.”
One alarm earns its place here, and it comes from AWS’s own CloudTrail examples. With the trail delivering to a log group (the prerequisite AWS names), you create a metric filter that matches IAM policy changes and an alarm on it, at a threshold of 1 change event in 5 minutes, as the example sets it. The example’s console sign-in failure alarm, at 3 failures in 5 minutes, is the second one worth adding.
The FTC’s complaint against Chegg, announced on October 31, 2022, alleges that Chegg stored users’ personal data in Amazon S3 and “failed to adequately monitor its networks and systems for unauthorized attempts to transfer or exfiltrate users’ and employees’ personal information outside of Chegg’s network boundaries”. It says that in or around April 2018 a former contractor used an AWS Root Credential (the complaint’s term for a single access key with full administrative privileges over the S3 data, which Chegg shared among employees and outside contractors) to exfiltrate a database with personal information of approximately 40 million users, and that “Had Chegg employed reasonable access controls and monitoring, it would have likely detected and/or stopped the attack more quickly.” The complaint’s next dated event is September 2018, when “a threat intelligence vendor informed Chegg that a file containing some of the exfiltrated information was available in an online forum.”
The lesson I take from it is about order: an owner who keeps no account history and no alarm on it may first hear of a misused credential from someone else. The complaint does not say which AWS logging Chegg had or whether CloudTrail was on, and reads of S3 objects are data events that a trail does not log unless you configure it, at extra cost, so neither a default trail nor the IAM alarm above would by itself have caught this exfiltration.
Two AWS tools handle security monitoring and the evaluation of what the logging records. GuardDuty is a threat detection service whose foundational data sources are CloudTrail management events, VPC flow logs and DNS logs, ingested once you enable it. Security Hub CSPM aggregates and prioritizes findings, and it “uses AWS Config rules to run security checks and generate findings for most controls”, so with Security Hub CSPM enabled without Security Hub “you must manually enable AWS Config and turn on resource recording”. Deciding what to detect is a separate job, part of how to prevent insufficient logging and monitoring.
AWS EKS logging: control plane logs, pod logs, and the ECS and Fargate routes
AWS EKS logging has 2 halves, set up separately. The control plane has 5 log types, each off by default, sent to CloudWatch Logs only when you turn them on for the cluster. Pod logs are a separate setup, for example through CloudWatch Container Insights or, on Fargate, the log router, which uses AWS for Fluent Bit.
EKS control plane logging covers five types: api, audit, authenticator, controllerManager and scheduler. Each is enabled per cluster, logs land in the group /aws/eks/<cluster>/cluster, delivery is “best effort”, and CloudWatch Logs ingestion, archive storage and data scanning rates apply. In the console: open the cluster, choose the Observability tab, then Manage logging in the Control plane logging section. From the CLI, aws eks update-cluster-config takes a --logging argument with the list of types and "enabled":true.
For pods, CloudWatch Container Insights “collects, aggregates, and summarizes metrics and logs from your containerized applications and microservices”, and on Fargate the log router streams logs to AWS services or partner tools using AWS for Fluent Bit. The table puts EKS logging and monitoring on AWS next to the other places a small app’s containers or functions might run.
| Platform | How stdout reaches CloudWatch Logs | What you must turn on |
|---|---|---|
| Lambda | Lambda sends function logs itself, to /aws/lambda/<function-name> | An execution role with logs:CreateLogGroup, logs:CreateLogStream and logs:PutLogEvents |
| ECS on Fargate | The awslogs log driver passes the container’s STDOUT and STDERR to CloudWatch Logs | The logConfiguration parameters in the task definition |
| ECS, logs going elsewhere | FireLens routes logs to an AWS service or partner destination, with Fluentd or Fluent Bit | FireLens task definition parameters |
| EKS control plane | Sent to /aws/eks/<cluster>/cluster | Each of the five log types, per cluster |
| EKS pods | Container Insights, or the Fargate log router with AWS for Fluent Bit | Container Insights or the log router |
| App Runner | Streams application output to /aws/apprunner/<service-name>/<service-id>/application | Not stated in AWS’s docs; App Runner is closed to new customers |
On ECS with Fargate, the awslogs log driver works only once its logConfiguration parameters are in the task definition. The App Runner row matters for one reason: AWS says “AWS App Runner is no longer open to new customers”, so a new customer cannot start on it. As I read it, a small app rarely needs EKS, and if a template put it there, the two switches above (control plane log types and a pod log route) are the minimum. Monitoring EKS with Prometheus and Grafana is a different setup, starting from the Prometheus metric types.
How to check your own AWS account
An AWS monitoring setup is checked with 6 tests you can fail: every log group has a retention value, a test request is traceable by its id, a published custom metric appears, a forced alarm reaches a phone, a console change shows up in CloudTrail, and the bill alarm exists with a confirmed subscription.
Run them on an account you own. Each starts in the console, and the commands for four of them, written from the AWS CLI reference, follow the list.
- 01 Retention on every log group. In the CloudWatch console, open Logs, then Log groups, and read the Retention column: pass is no group showing Never Expire. Keep the dated output of the describe-log-groups command below.
- 02 One request, traced end to end. Send one request, take the request id your app wrote for it (from its first log line, or a response header if the app returns one), and search for it in Logs Insights in every log group it touched. Use the app's own request-id field: a platform's own request id is a different value. Lambda logs can take 5 to 10 minutes to show up, so wait before you search. Pass is one query that returns the whole path; keep the query and its result.
- 03 A test custom metric. Publish one data point with the put-metric-data command below, then retrieve it with get-metric-statistics after 2 minutes. Do not judge by the console graph: a new metric can take up to 15 minutes to appear in the metric list. Keep the get-metric-statistics output.
- 04 Every alarm, forced. Set each alarm to ALARM with set-alarm-state, confirm the message reaches a phone or an inbox, then confirm the alarm returns at its next comparison. Keep the alarm history entry and the message you received.
- 05 A console change in CloudTrail. Change a harmless setting in the console, such as a tag, then find the event in CloudTrail Event history, or with lookup-events filtered on your user name, in the same Region where you made the change. Keep the event record.
- 06 The bill alarm. Confirm it exists and, if it notifies through an SNS topic, that the subscription shows as confirmed and not Pending confirmation. Keep a screenshot of the subscription status.
The commands for tests 1, 3, 4 and 5, with names and times you replace:
# 1. A group with no "retentionInDays" in the output keeps its data indefinitely
aws logs describe-log-groups
aws logs put-retention-policy --log-group-name /aws/lambda/my-function --retention-in-days 30
# 3. Publish, wait 2 minutes, then read it back
aws cloudwatch put-metric-data --namespace MyApp --metric-name TestMetric --unit Count --value 1
aws cloudwatch get-metric-statistics --namespace MyApp --metric-name TestMetric --statistics Sum --period 60 --start-time 2026-10-03T09:00:00Z --end-time 2026-10-03T10:00:00Z
# 4. Force one alarm; it returns at the next comparison
aws cloudwatch set-alarm-state --alarm-name error-rate --state-value ALARM --state-reason "alarm test"
# 5. Events by your IAM user name, in the Region you changed
aws cloudtrail lookup-events --lookup-attributes AttributeKey=Username,AttributeValue=my-user
Event history records an event in the Region where it happened, which is why test 5 names the Region. Tests 1 and 4 sit first and fourth for a reason that is my working rule, not a statistic: they are the two people most often skip, and they are the two that turn a configured setup into a proven one.
Where the sprint fits
In area 08 of the Production Hardening Sprint, deliverable 8.1 adds structured request logs with correlation IDs and appropriate user references, excluding passwords, tokens and unnecessary personal data, and is verified by tracing a test request across services and checking log content for sensitive fields ; 8.3 alerts on error spikes, latency, connection pressure and queue backlog, verified by triggering each configured condition and recording its alert threshold and behavior ; 8.4 routes alerts to the designated Slack or email destination and tunes thresholds to reduce noise, verified by sending test alerts and checking their destination, context and response instructions ; and 8.5 sets and documents log retention across the application’s services, verified by inspecting the configured windows and verifying retention behavior in supported systems. Deliverable 13.1, the production readiness report, delivers the result for every scope item, the work completed and its verification evidence, accounting for all 123 IDs, keeping failures visible until resolved and explaining genuine non-applicable items. We start from your app’s current framework and hosting setup and refactor or replace components where the production work requires it ; hosting, paid tools and API usage remain in your accounts, and we explain any required third-party costs before enabling them. Each deliverable and its verify line is listed among the logging and monitoring items in the sprint’s scope.
Common questions about CloudWatch, CloudTrail and AWS logs
Is AWS CloudWatch metrics free?
Partly. Many AWS services publish their own metrics at no charge, while detailed monitoring and your own application metrics cost extra, and the CloudWatch agent’s metrics are billed as custom metrics. The free tier covers 10 metrics of custom and detailed monitoring metrics a month, as AWS’s pricing page stated it on 2026-10-03. Logs have a separate free allowance of 5 GB a month, which covers log ingestion, archived log storage and the data Logs Insights queries scan.
How long are metrics kept in CloudWatch?
Up to 455 days, at falling resolution. Data points under 60 seconds are kept for 3 hours, 1-minute data for 15 days, 5-minute data for 63 days, and 1-hour data for 455 days (15 months); shorter-period data is aggregated into the coarser resolution as it ages.
Is AWS CloudWatch a SIEM?
Not by AWS’s own description: its logging and monitoring guide describes CloudWatch as a service for analyzing logs and watching the metrics of AWS resources and hosted applications in real time, and does not call it a SIEM. For threat detection AWS offers GuardDuty, and Security Hub CSPM aggregates findings from several services.
Does Lambda automatically create a log group?
Yes, when the function’s execution role has logs:CreateLogGroup, logs:CreateLogStream and logs:PutLogEvents, Lambda captures its logs into /aws/lambda/<function-name> by default. That group keeps data indefinitely until you set a retention value on it.
What are the key differences between CloudTrail and GuardDuty?
CloudTrail records activity in the account, the actions taken by a user, role or AWS service; GuardDuty is a threat detection service that analyzes sources including CloudTrail management events, VPC flow logs and DNS logs. One keeps the record, and the other reads records like it for signs of an attack.
If you have a working app built with these tools and need it ready for real customers, this is what we do.
Built it with AI. Now it has to hold up for real customers.
The Production Hardening Sprint takes the app you already have and builds the production foundation underneath it. Authentication and access rules, payments that stay consistent, error handling, monitoring, backups, automated tests and a documented handover. Our engineers work inside your existing codebase for ten working days. All 123 deliverables are included, and you get the evidence for each one.
See the Production Hardening Sprint →
$2,500 fixed price · 10 working days · One codebase