A cache is a saved copy of an answer, kept so the work is not repeated. The caching strategies web applications rely on sit at 4 layers: the browser, the CDN, the framework’s data cache and a store such as Redis. The dangerous failure is a page built for one signed-in customer and served to the next.

What a web cache is, and the four layers your app already has one at

A web cache is a stored copy of a response, and a typical web app has 4 of them: the browser, the CDN, the framework’s data cache and an application store such as Redis. The browser’s is private to one user. The other three are shared: each holds only what every user may see, unless its key names the customer.

The HTTP half of that split comes from MDN’s guide to HTTP caching. A private cache is tied to one client, usually the browser, so it can hold a personalized response; a shared cache sits between clients and the server and “can store responses that can be shared among users”. Put personalized content anywhere but a private cache and, in MDN’s words, “other users may be able to retrieve those contents”.

Caching is one part of web performance optimization for an AI-built app, next to bundle size, images and load testing. This page stays on the cache.

The table maps each layer to who sets its expiry, from each layer’s documentation. The last column, how it goes wrong, is my reading of where each layer usually bites a small app.

LayerWhat it storesWho sets the expiryHow it goes wrong
Browser cacheResponses for one user, private to that browserYour server, through Cache-Control, with validators such as ETag to revalidate a stale copyAn old JavaScript bundle keeps running after a deploy
CDN or edge cacheResponses shared by every visitor who asks for the same URLCache-Control directives such as public and s-maxage, plus the CDN’s own cache rulesOne user’s response is served to everyone who asks for that URL
Framework data or route cacheFetched data and rendered output on the serverThe framework: in Next.js, cacheLife on a use cache function, or the fetch cache option in the previous modelData that never updates after a write
Application store (Redis, Valkey, Memcached)Whatever your code writes, under keys your code buildsYour code, through a TTL on each keyA key that leaves out the customer

What the framework caches by default depends on the version and its settings. In the Next.js docs for version 16.3.8, an app with Cache Components turned on (cacheComponents: true) caches a function or component where you add the use cache directive; without it, the previous model applies, where fetch requests are not cached by default and you can cache one by setting its cache option to 'force-cache'. The database keeps its own hot data in memory too; that cache rewards better queries and has no setting on this page.

The caching strategies web applications use: cache-aside, write-through and expiry

Caching strategies for a web application come down to 2 patterns and a safety net: cache-aside reads the cache and fills it on a miss, write-through updates the cache on every write, and a TTL under both limits how long a missed invalidation can hurt.

AWS’s caching strategies page weighs lazy loading against write-through, the two main strategies of caching, and adds a TTL to each; it also lists reading from cache replicas, which scales the cache rather than filling it. Lazy loading usually goes by the name cache-aside, though AWS’s page never uses it. AWS’s trade-offs are blunt: lazy loading caches only requested data and a node failure is not fatal to the app, but every miss costs three trips and the data “can become stale”; write-through keeps the cached data current, but a new node starts with data missing, and in AWS’s words “Most data is never read”.

The types of caching in a web application differ less by product than by when the copy is written: on a read, on a write, or never past its expiry. The “how it works” column follows AWS and MDN; the stale-risk and fits columns are my reading.

StrategyHow it worksStale riskFits what data
Lazy loading (cache-aside)Read the cache; on a miss, read the database and write the result to the cacheStale until the TTL ends or the key is deleted, because a database write does not touch the cacheReads that repeat far more often than the data changes: a plan list, a settings lookup, a slow report
Write-throughAdd or update the cache whenever the database is writtenAWS calls the cached data never stale; it goes stale when anything writes the database outside that code path, such as an admin edit or a scriptData read right after it is written: a profile, a project name
TTLA lifetime on each entryBounded by the TTLEvery key, as the floor
stale-while-revalidate (HTTP)A cache “could reuse a stale response while it revalidates it”Visitors see the old copy for up to the window you setPublic pages where a slightly old copy is fine

Whether to cache a given read at all, and which data suits a cache, is worked through in when caching raises the ceiling and when it hides the bug, including fixing the query first. I won’t repeat that test here. What I’d add is a short list of things that never go in a shared cache:

  • anything behind a login, unless the customer is part of the key
  • authorization decisions: who may see or do what
  • payment and subscription state
  • one-time tokens: password resets, magic links, invitation codes

Where a cache fits among other changes to a growing app follows the order for scaling web applications.

What goes wrong without it

Performance & Scale averages 53.3 out of 100 across the 21 third-party apps scored on it in my June and July 2026 audits. Those 21 were 11 public vibe-coded apps audited across all 12 pillars plus 10 held-out apps audited blind, a selected set of apps I audited rather than a random sample, so the figure is not a rate for AI-built apps in general.

Caching problems show up as four complaints. When an app shows stale data after an update, the write usually worked and the read came from a copy that nothing told to change. A cached page showing another user’s data is the serious version of the same mistake. The cause and first-check columns below are my reading.

What the customer seesThe causeThe layerThe first check
”I saved it, and the old value is still there”The write succeeded, but nothing deletes the cached read on write and the TTL is longFramework cache or application store, sometimes the CDNThe response’s Age header, or the CDN’s cache-status header
The old version of the app after a deployHTML cached for longer than the hashed assets it points toBrowser or CDNCompare the bundle file names in the page source with the latest build
One customer’s dashboard shown to anotherA response to a signed-in request landed in a shared cache: the route allowed shared caching, a CDN rule cached it regardless, or the application key was dashboard:summary with no tenant in itCDN or application storeThe two-account test in the verify section
A bigger bill and a slower app as traffic growsEvery repeated expensive read is paid for againNo cache where one is safeQuery counts per page view in the database dashboard

The bill row has a concrete form on Supabase: uncached and cached egress are charged at different rates past the included amount on paid plans, and cached versus uncached egress on Supabase has the figures. The third row has security names, web cache deception and cache poisoning: OWASP has a community page on cache poisoning, and Cloudflare’s cache docs carry pages under both names. I won’t describe either technique; the key and header rules in the next section are the defense this page covers.

Repeated expensive work has a plain shape in code. In one app I audited, a health-data API imported a patient’s history one database round trip per item, while the same code did it in one batch a few hundred lines away. The lesson I take from it: that is the work a cache is meant to spare, but the batch is the fix. The pattern behind it is the N+1 query problem and its fix, and a cache in front of such a loop moves the cost rather than removing it, as the caching section linked above argues.

A plan that stays on the old tier after a payment is often this same gap, covered as stale UI after a Stripe webhook. The cold cache after a deploy is covered in that same caching section.

How to do it: keys, headers, invalidation and the store

Safe caching is 4 decisions made in order: what goes in the key, which responses a shared cache may store, what deletes an entry when the data changes, and where the store runs so nothing outside the app can reach it.

The header, framework and store behavior below comes from each project’s documentation, read on October 4, 2026; the rules built on it are mine.

Tenant-safe cache keys, and the header for signed-in pages

A tenant-safe cache key starts with the tenant id taken from the verified session, then the user or role, then the resource and a version. My working rule for headers: a response to a signed-in request carries Cache-Control: private, no-store unless that route has been decided otherwise, so no shared cache keeps it.

The key order, in full, is my working rule: tenant, then the user id or role when the content differs by user, then the resource and its parameters, then a version, as in t:42:u:7:dashboard:v3. The ids come from the session your server verified, never from a query parameter or a request body, because a key built from client input lets a client ask for another tenant’s entry.

// Tenant, then user, then resource, then version
const CACHE_VERSION = 'v3'
export function cacheKey(s: { tenantId: string; userId: string }, resource: string) {
  return `t:${s.tenantId}:u:${s.userId}:${resource}:${CACHE_VERSION}`
}
// On every response to a signed-in request
res.setHeader('Cache-Control', 'private, no-store')

The header rule sits on MDN’s definitions: private means the response “can be stored only in a private cache”, and no-store means “any caches of any kind (private or shared) should not store this response”. MDN also names the cost. It advises against using no-store liberally, because the page loses the browser’s back/forward cache, and it says to prefer no-cache in combination with private. I keep no-store on signed-in pages that show account data, since OWASP’s testing guide checks that the app “correctly instructs the browser to not retain sensitive data”, and I’d use private, no-cache on a route where that trade has been weighed.

The header matters because, without one, the CDN’s own rules decide. Cloudflare’s default cache behavior does not cache HTML or JSON by default, and skips any response marked private or no-store or carrying Set-Cookie. Its cache rules can override the origin’s Cache-Control, though, and a “cache everything” rule added for speed is one way signed-in HTML ends up shared (my reading).

A Vary header on the cookie or the Authorization header is no substitute. MDN says apps that use cookies to keep personalized content from being reused “should specify Cache-Control: private instead of specifying a cookie for Vary”, and Cloudflare’s docs say that by default it “does not consider vary values in caching decisions”.

My rule for framework caches is the same: a cached function that uses a user’s token or cookies stays out of any cache shared between users. In Next.js, the arguments of a use cache function and the values it captures become part of its cache key, so pass the tenant id in as an argument; a function that reads cookies directly belongs in use cache: private, whose results live only in the browser. In the database, the same rule goes by the name multi tenant data isolation.

What a cache invalidation strategy is, and how to invalidate a cache key

A cache invalidation strategy is the written rule for when a cached copy stops being served. The methods are expiry, delete on write, versioned keys and tags. My working rule: a small app puts a TTL on everything and deletes the key after the database commit for anything a user edits and then looks at.

How you invalidate a cache depends on the change: each method below fits a different kind. The “when it fits” column is mine; the Next.js and Cloudflare cells come from their docs.

MethodHowWhen it fits
Expiry (TTL)A lifetime on every entry, sized to how stale that data may beEverything, as the floor under the other three
Delete on writeThe function that writes the row deletes the key, after the commitRecords a user changes and expects to see changed on the next load
Versioned keysBump the version in the key (v3 to v4); old entries are never read again and age outA deploy that changes what is cached or its shape
TagsGroup entries under a tag and invalidate the group: revalidateTag in Next.js, purge by tag on Cloudflare; revalidatePath and purge by URL clear one path insteadMany pages built from the same data

Versioned keys are the third part of my rule, bumped at each deploy that changes a cached shape. Tags behave differently per tool. In the Next.js caching docs, revalidateTag with the recommended max profile marks tagged data stale, and the next request “is served stale content while it runs”; when the data must be gone at once, the docs name updateTag in Server Actions or a { expire: 0 } profile. Cloudflare lists purge by URL, hostname, tag and prefix on all four plans, Free included, and notes that a purge’s HTTP 200 only means the request arrived: request the asset again and confirm CF-Cache-Status is no longer HIT.

Two traps sit around invalidation. Switching a route to no-store does not clear what a shared cache already holds; MDN says it “does not delete any already-stored response for the same URL”, so purge as well. The mistake I’d look for first in application code is deleting the key before the database commit: a read that lands between the delete and the commit misses, reads the old row and puts it back, where it stays until the TTL.

Where the cache lives: Redis, Valkey, Memcached, the managed options, and the ports nobody should reach

Redis, Valkey and Memcached are 3 common cache stores, and Amazon ElastiCache offers Valkey, Memcached and Redis OSS as a managed service. Pick the one your host runs for you. A store you run yourself stays off the public internet: Memcached’s own docs say it must not be exposed directly to the internet.

On AWS, Memcached is one of the three engines that what Amazon ElastiCache is describes, next to Valkey and Redis OSS, run either as a serverless cache or as a node-based cluster you size. The key difference between ElastiCache and Memcached is that one is AWS’s managed service and the other is an engine it runs. Node-based clusters and serverless caches on a VPC endpoint are reached from inside AWS, so they fit an app that already runs there; a serverless cache can instead take a public endpoint, which needs IAM authentication and TLS 1.3 on every connection and, as of October 4, 2026, a cache on Valkey 9.0 or later.

Valkey is an open source (BSD) key/value datastore backed by the Linux Foundation. Valkey’s cluster tutorial describes a Valkey cluster as a way to run it “where data is automatically sharded across multiple Valkey nodes”, with every node needing a client port and a cluster bus port.

StoreWhat it isManaged on AWSDefault portAccess control
RedisA data store used as a database, cache or messaging systemElastiCache runs Redis OSS6379; 16379 in cluster mode, 26379 for SentinelPort denied to all but trusted clients; ACL users since Redis 6 or the legacy requirepass; protected mode
ValkeyOpen source (BSD) key/value datastore, Linux FoundationElastiCache runs Valkey, serverless included6379 on AWS’s page; the cluster bus port is the data port plus 10000Port denied to all but trusted clients; ACLs or requirepass; protected mode
MemcachedAn in-memory key-value cacheElastiCache runs Memcached, serverless on 1.6.22 and later11211 on AWS’s page; TCP only by default since 1.5.6, UDP offSpends little, if any, effort defending itself from random internet connections; SASL “helps, but should not be totally trusted”

My choice for a small app: the store your host already manages; Redis or Valkey when you also need queues, locks or rate-limit counters, and Memcached only for a pure cache.

The ports are the part that goes wrong. Memcached’s configuring guide says Memcached “does not spend much, if any, effort in ensuring its defensibility from random internet connections”, that you “must not expose memcached directly to the internet, or otherwise any untrusted users”, and that since 1.5.6 it listens only on TCP by default, with UDP off. The UDP default has history: CISA’s alert on UDP-based amplification attacks added Memcached as an attack vector on February 27, 2018. Redis’s security docs say access to its port “should be denied to everybody but trusted clients in the network”.

So a store you run yourself sits on a private network with its port closed to the outside, and a managed store reached over the internet sits behind the provider’s TLS and authentication. The check, my reading: probe the port from outside your network, and try a managed store without its password. The edge in front of the app is a separate job, how to set up a web application firewall, and so are the network rules in hardening a managed database.

cache-manager and the framework caches in Node

The cache-manager package on npm puts one interface over an in-memory store and other stores: by default everything sits in memory, and you can add any Keyv-compatible storage adapter, such as @keyv/redis. Its wrap helper runs a function the first time and serves later calls from the cache, which is cache-aside in one call (my reading); TTLs are in milliseconds, and version 7 is current, 7.2.9 on the npm registry. The cache-manager README lists the adapters.

NestJS’s caching docs build on it: you install @nestjs/cache-manager with cache-manager, and with no ttl set, entries never expire. Its CacheInterceptor keys cached responses by the request URL unless you override trackBy(), and the docs’ own example of when to do that is the Authorization header on profile endpoints. A URL-only key on a signed-in route is the missing-tenant key from the failure table earlier on this page.

This is not npm’s own download cache. npm cache clean and npm cache verify manage the folder of downloaded packages and do nothing for an app’s speed.

One more reading of mine: an in-memory store lives inside one running instance, so on a serverless host that starts and stops instances it is not shared between requests that land on different ones. Next.js says the same of its default store, which “doesn’t persist across serverless requests”. The shared store, or a framework cache that is durable across instances, is the one that holds.

How to verify it: a cache test for hit rate, invalidation and tenant separation

A cache is verified by 3 tests: compare hits and misses over a normal day, change a record and time how long the old value survives, and load the same URL as two different customers and with no session at all. Keep the output of each.

Cache testing here means checks you can fail, each with what you should see. The first needs only a browser.

  1. 01 Browser first: open DevTools, go to the Network tab, and load a static asset or a page you meant the CDN to cache, twice. Read the CDN cache-status header (CF-Cache-Status on Cloudflare) and the Age header. Seen: HIT on the second load of the asset, and never HIT on a signed-in page.
  2. 02 Test the cache hit rate in production: read the hit and miss counters from the store itself (keyspace_hits and keyspace_misses in the stats section of INFO on Redis and Valkey, get_hits and get_misses from stats on Memcached) or from the CDN analytics, over a normal day. Seen: two counts with a date next to them.
  3. 03 The stale-data test: change a record, start a stopwatch, and reload as the same user. Seen: the seconds until the new value appears. With delete on write it shows on the first reload; with a Next.js tag revalidated on the max profile, the first request after the change can still get the old copy while it refreshes, so time the second one too.
  4. 04 The tenant-separation test: sign in as customer A in one browser and customer B in another, load the same URL in both, and compare. Seen: each sees only their own data. Then, right after A has loaded the page signed in, request the same URL through the CDN hostname, not the origin, with no cookie, using the command below. Seen: the status line and a body with none of A's data in it.
  5. 05 After a deploy: confirm the HTML is fresh and the app reads the new keys. Seen: the new asset file names in the page source, and a miss on the first read under the new key version, while the old entries wait out their TTL.
  6. 06 The store is closed to strangers: probe a store you run yourself from outside its network, and try a managed store reached over the internet without its password. Seen, for your own store: connection refused. A filtered result counts only if the same probe reaches the port from inside the private network, since a wrong address also shows as filtered. Seen, for the managed store: every query refused until the client authenticates.

The counters in check 2 are defined on Redis’s INFO command page: keyspace_hits is the number of successful key lookups and keyspace_misses the number of failed ones.

The no-cookie request for check 4:

# Through the CDN hostname, no cookie, right after customer A's signed-in load
curl -sS -i https://app.example.com/dashboard | head -n 20

A redirect to the login page passes, so does a 401 or 403, and so does an app that renders in the browser and answers 200 with an empty shell: on Supabase, a read whose row a policy filters out “raises nothing, matches zero rows”, while a missing table grant raises error 42501, and which status reaches the browser depends on how your build handles each (my reading). The fail is any of A’s name, email or records in that body.

Keep the evidence together: the headers, the hit and miss counts with their date, the stopwatch result, both accounts’ responses, the refused request with no session, and the store probe.

In the sprint, we verify deliverable 9.1, Appropriate caching, this way: compare cache hits and misses, test invalidation, and verify tenant separation. Deliverable 1.4, Data isolation across every table, is verified one layer down: run read and write tests as anonymous users, different roles, and separate tenants.

What the cache bought shows up in a load test run before and after the change: load testing a web application for pages, API load testing for endpoints.

The caching review checklist for a SaaS app

Nine lines a reviewer can tick, my list:

  • Every cached thing is listed with its layer, its TTL and how it is invalidated.
  • No signed-in response can be stored by a shared cache unless that route is on the list.
  • Every key that holds customer data starts with the tenant.
  • The ids in keys come from the verified session, never from the request.
  • Writes delete or update the cached copy after the commit, not before.
  • Keys carry a version that a deploy bumps.
  • The store is on a private network, or is a managed store that needs its password and TLS, and requires authentication either way.
  • Hit and miss counts are visible somewhere the team looks.
  • The two-account test and the stale-data test each have a date next to them.

Where the sprint does this

Deliverable 9.1 of the Production Hardening Sprint, Appropriate caching, means we cache repeated reads where it is safe, with explicit invalidation and customer-data boundaries; its verification is the line quoted in the verify section above. The load test that shows the before and after is deliverable 9.2. Both are recorded in the production readiness report, deliverable 13.1, which accounts for all 123 IDs, keeps failures visible until resolved and explains genuine non-applicable items. Hosting, paid tools, and API usage remain in your accounts. Every deliverable is listed in the published scope.

Common questions about caching a web app

Is Memcached still used?

Yes. The project still publishes current server documentation, including its TCP-only default since 1.5.6, and AWS runs it as one of ElastiCache’s three engines, with a serverless option.

Is ElastiCache the same as Redis?

No. ElastiCache is AWS’s managed service, and the engine inside it is Valkey, Memcached or Redis OSS, chosen when you create the cache. So to run a Redis cache on AWS without managing servers, you pick the Redis OSS engine, or Valkey, in ElastiCache.

Can I use ElastiCache with Memcached?

Yes. Memcached is one of its engines, either as a node-based cluster or as a serverless cache on Memcached 1.6.22 and later.

When not to use Redis cache?

Skip it when nobody will own invalidation for that data, or when the data must never be even seconds stale, such as payment state and permissions. Whether a cache helps at all is worked through in the article on why AI apps stall under concurrent users.

How to calculate cache hit rate?

Divide hits by hits plus misses over a stated period. On Redis and Valkey the counts are keyspace_hits and keyspace_misses in the stats section of INFO; on Memcached they are get_hits and get_misses from the stats command.