How do you cap monthly usage on a metered API? Keep the key on a signed-in server route, limit each user per month, set an app-wide ceiling that refuses calls, and test the spend alert. Of the 14 AI apps I audited in June and July 2026, 12 had a confirmed denial-of-wallet path: a stranger or free account can burn the owner’s paid AI or compute bill without limit.

What these controls are, together: how to cap monthly usage on a metered API

Capping monthly usage on a metered API takes 3 layers in your own app: the provider key stays on the server behind a signed-in route, each user has a monthly limit, and the whole app has a ceiling that refuses calls instead of reaching the provider. A spend alert below that ceiling tells a person first.

Those 14 apps were third-party AI apps from my June and July 2026 audits, a selected set of apps I audited rather than a random sample, so the figure is not a rate for AI-built apps in general. This page takes the controls for paid keys from secrets management, the area guide and goes one level deeper.

LayerWhat it stopsWhere it livesHow you know it works
The key never leaves the serverA key copied out of the browser bundle and used from anywhereA server route or edge function reading a server-only environment variableThe bundle search below finds no trace of the key
The route requires a signed-in userScripts calling your AI, email or SMS route without an accountA session check at the top of the route, before any other workThe anonymous-call check: the route refuses and the provider is never called
Each user has a monthly limitOne account, free or scripted, spending everyone’s budgetA counter per user per month, taken before the provider callThe over-limit and race checks: the limit refuses, and two racing requests cannot both pass
The whole app has a monthly ceilingMany accounts together crossing what you can affordOne counter for the app, checked in the same step as the user’sThe ceiling check: once it is crossed, a second user is refused too
The provider’s own limit, where one existsYour own code failing openThe provider’s console (the alert-or-stop table below says which ones stop anything)The provider’s documentation, then the lowered-threshold check for its alert
An alert fires below the ceilingNobody noticing until the invoiceProvider alerts, or a daily job over your own usage tableThe threshold and delivery checks: the alert fires and a named person receives it

The server-side key usage checklist is these six rows with a date and an owner beside each one: who checked it, and when it was last proven.

These two controls are deliverables 2.5 and 2.7 of the Production Hardening Sprint. For 2.5 we move private AI, email, SMS and other service keys server-side and implement usage caps; for 2.7 we configure spend or usage alerts for every metered API, using provider alerts or application-side measurement.

API usage and API key pricing: what a key costs when somebody else is holding it

API usage is what a metered provider bills, by token, request, message or map load, and Google’s pricing page says an API key itself is free of charge. Anyone holding your key, or calling your open endpoint, spends at your prices, so the usage page is where abuse shows before the invoice does.

Google’s API keys pricing page is one short paragraph: “API Keys is free of charge”, and if you use Cloud Endpoints to manage your API, “you might incur charges at high traffic volumes”. What a key is, and whether keys cost anything at other providers, is covered in what an API key is. Google API key pricing questions therefore come down to the APIs the key unlocks: the Google Maps Platform pricing list bills per SKU, with “Costs listed per 1000 events”. Google’s Maps documentation also says who pays when a key is misused: “You are financially responsible for charges caused by abuse of unrestricted API keys.”

For OpenAI API key usage, OpenAI’s Usage API guide is the reference: usage can be pulled per key and per project, which is how you spot a key spending where it should not. Anthropic tracks usage and cost per workspace through its Usage and Cost API, and each workspace other than the Default Workspace can carry its own monthly spend limit, covered in the alert-or-stop table further down. If the Anthropic key in question sits on a developer’s laptop rather than in your app, what an API key does to a Claude Code bill covers that case.

The price of abuse is your provider’s price list multiplied by the other person’s appetite, and the other person does not stop when the month’s budget runs out. This page keeps no price table, because prices change; the pricing pages linked above were checked on September 28, 2026. What usage costs per customer, and what to charge for it, is a separate margin question, answered in AI API costs and gross margin.

What goes wrong without them

What happenedHow the owner found outWhich layer would have stopped it
A script found an AI route that answers anyone and looped itThe invoice, or the provider’s own limit cutting the app offThe signed-in route, then the per-user monthly limit
Spend jumped overnight with no attacker: a batch job or one heavy customerThe usage page, if someone happened to open itThe per-user monthly limit and the app-wide ceiling
The month’s bill was far larger than anyone expectedThe invoice itselfAn alert below the ceiling, sent to a person

An unexpected API bill from abuse

The shape to look for is an endpoint that calls the model for anyone who asks: no sign-in at all, or a free account with no limit behind it. A script can find that route, call it in a loop, and spend on your account until something outside the app stops it. That is the path behind the figure in the opening, from the same audits. The wider class, where a stranger makes your AI feature spend, is covered under prompt injection and denial of wallet. A private key shipped in the browser bundle is the same failure by a shorter road, since anyone can lift it and skip your route entirely; API keys exposed on the frontend explains which keys may be public and which may not.

AI API costs spiked overnight

Whether an overnight jump means somebody got in is its own triage, covered under a bill that spiked overnight. What I’d add, as my reading: a spike can have no attacker in it at all. A background job that reprocesses every row after a deploy, or one customer pasting a very long document into an AI feature, spends exactly like an attacker would. Both are caught the same way by the monthly cap on each user and the ceiling over the whole app, since a counter sees units spent, not intent. If the spike came from a key that was exposed, stop reading here and work through what to do when an API key leaked first.

A surprise cloud bill at the end of the month

The metered lines a small app pays for are listed in the metered lines on a monthly app bill. The point here is narrower. Without an alert, the first signal is the invoice, and the provider alerts in the alert-or-stop table below are settings somebody has to create; each row’s documentation says how. Hosting platforms bill this way too, and what a Replit bill can do if nobody sets a limit walks through one of them. The alert also has to reach a person who still works on the product and can still sign in to the account, which is not a given when a developer owns your hosting account.

How to set them up

Build them in this order: the proxy first, since every later control lives inside it, then the counter, then the provider settings, then the services that are easy to forget.

How to proxy AI API calls server side

Proxying AI API calls server side takes 5 steps: move the call into a server route, read the key from the server environment, require a session, bound the input and the output tokens so one call has a known maximum cost, and return only the result the client needs.

  1. 01 Move the provider call out of the browser and into a server route or edge function
  2. 02 Read the key from the server environment, never from a variable your framework exposes to the browser
  3. 03 Require a valid session and reject anonymous calls before any other work happens
  4. 04 Validate and bound the input with a maximum character count, a maximum output token count and an allow list of models, so one call has a known maximum cost
  5. 05 Return only what the client needs, not the raw provider response

The handler below is pseudocode for any stack; the model id and the output cap are placeholders you fill in.

handler POST /api/summarize(request):
  user = getSession(request)
  if user is null: return 401 "Sign in first"
  text = request.body.text
  if length(text) > MAX_INPUT_CHARS: return 400 "Input too long"
  units = worstCaseUnits(text, MAX_OUTPUT_TOKENS)
  if not takeFromMonthlyLimits(user.id, currentMonth(), units):
    return 429 "Monthly limit reached"
  result = provider.call(
    key: env.PROVIDER_API_KEY,
    "model": "<model id>",
    max_tokens: <your cap>,
    input: text)
  recordUsage(user.id, currentMonth(), result.usage)
  return { summary: result.text }

In the HTTP specifications, a 401 means the request “lacks valid authentication credentials”, and a 429 means the caller “has sent too many requests in a given amount of time”. Streaming works through the same proxy: the checks run before the first token goes out, and the route passes the stream on. An API gateway can do the same job for a team that already runs one.

The per-user limit and the monthly ceiling

A monthly cap is a counter per user and one for the whole app, incremented atomically before the provider is called, so two simultaneous requests cannot both take the last unit. Count tokens or messages where the price varies per call. At the limit, return 429 and skip the provider call entirely.

OpenAI’s own rate limits guide makes the same recommendation for protection against misuse: “set a usage limit for individual users within a specified time frame (daily, weekly, or monthly)”. In Postgres, the counter is one statement per key, one row per user per month plus one row for the whole app:

UPDATE usage
SET used = used + $3
WHERE account_id = $1 AND period = $2 AND used + $3 <= monthly_limit;

Read how many rows it updated: one means the units are yours, none means refuse. Create the month’s row before the first call, with an insert that does nothing when the row already exists (in Postgres, ON CONFLICT DO NOTHING, which “simply avoids inserting a row”); otherwise the first call of every month is refused. Run the user’s row and the app’s row in one transaction and roll back if either one updates no row. Two requests racing for the last unit cannot both pass, because under the default Read Committed level described in PostgreSQL’s transaction isolation docs, the second update waits for the first to commit or roll back, and if the first committed, “The search condition of the command (the WHERE clause) is re-evaluated to see if the updated version of the row still matches the search condition.” An atomic increment in a key-value store does the same job, provided the check and the increment happen in one step. The 429 should carry a plain message the user can act on, such as when the limit resets.

The “fits when” and “the catch” columns below are my reading, except the gateway’s catch, which is AWS’s own documentation.

Counter storeFits whenThe catch
A row in your Postgres databaseThe app already runs on Postgres and traffic is modestEvery paid call adds a write to your main database
Redis or a hosted key-value storeRequest volume is high, or the relational database is not on the hot pathA second system to run, and the monthly totals still need a durable home
An API gateway usage planThe app already sits behind a gatewayOn AWS API Gateway, usage plans are for REST APIs, and throttling and quotas “are not hard limits, and are applied on a best-effort basis”

AWS goes further: “Don’t rely on usage plan quotas or throttling to control costs or block access to an API.” So a usage plan is never the ceiling; the app’s own counter is. Reserving the worst-case cost before the call and giving back the rest afterwards is covered step by step in check the limit before the expensive work, and the per-route limiter for a model endpoint in how to rate limit an AI endpoint. A short-window limit on every route is a different control (rate limiting in an API), per-account model budgets inside an AI feature sit with what prompt injection is, and plan tiers you sell to your own customers are a separate question: how to enforce plan limits on the backend.

In an AI coding workspace I audited, the cheap AI chat route was rate-limited but the two routes that start cloud sandboxes and run commands, which cost far more, were not. Each sandbox call started a fresh cloud sandbox. My reading is that a limit belongs on the route that costs the most, not the one that is easiest to find, and the counter should count what the provider bills.

What a billing alert threshold is, and which provider limits actually stop usage

A billing alert threshold is an amount at which a provider notifies you, and it stops nothing unless the provider says so. Vercel states that setting a spend amount does not stop usage on its own, and Google Cloud says an alerts-only budget does not automatically cap usage. Check each provider you pay before trusting its limit.

The threshold can be a fixed amount or a percentage of a budget. A cloud billing alert is the same notice on a cloud account: Google Cloud, for example, sends alert emails when actual or forecasted costs exceed a percentage of the budget you set, to the recipients you specify. The table puts the providers a small app most often pays side by side, every cell from the provider’s own documentation.

ProviderThe setting’s nameAlert or hard stopWhat happens at the limitWhere it is setDate checked
OpenAISpend alert; hard spend limitBoth: alerts notify, a hard spend limit stopsA spend alert “Sends a notification; API traffic continues”; under a hard limit, “Affected API requests return a 429 error”, and enforcement “is not instantaneous”Organization limits, or Project settings then Limits2026-09-28
AnthropicSpend limit (organization or workspace); tier spend capBoth: workspace limits can also send alerts at thresholdsAt your own limit, requests return HTTP 400 with invalid_request_error; at the tier cap, usage “pauses until 00:00 UTC on the first day of the next month, unless you request a higher limit sooner”The Billing page for the organization; the workspace’s Spend limits tab in the Claude Console2026-09-28
Google CloudAlerts-only budget; spend cap budget (preview)Alerts-only budget: alert. Spend cap budget: stop, for eligible servicesAn alerts-only budget “doesn’t automatically cap” usage; an enforced spend cap pauses the service “until you manually lift the spend cap”A Cloud Billing budget; a spend cap covers one project and one eligible service, such as the Gemini API or Cloud Run2026-09-28
Google Maps PlatformQuota limits; budgets and budget alertsQuota: stop. Budget: alertQuotas cap “the number of requests your project can make”, and a quota set too low “could cause a user-facing outage”The quotas for each API in the Google Cloud console; Cloud Billing budgets2026-09-28
AWSCloudWatch billing alarm; AWS Budgets; budget actionsAlarm and budget: alert. Actions: apply an IAM policy or SCP, or target EC2 or RDS instancesThe alarm notifies an Amazon SNS topic, which can include your email; Budgets notify an SNS topic, an email address or bothCloudWatch in US East (N. Virginia) after billing alerts are enabled; AWS Budgets in Billing and Cost Management2026-09-28
VercelSpend Management (spend amount)Alert, plus a stop only with Pause Production Deployments onNotifications at 50%, 75% and 100% of the amount; with pausing on, production deployments stop serving traffic, though AI Gateway key usage and v0 usage continueTeam settings, Billing, Spend Management, on Pro or Enterprise (Flexible Commitment) with an Owner or Billing role2026-09-28
ReplitUsage limitStop”usage-based services are blocked until the next billing cycle or until you increase the limit”For Core accounts: Settings, Account, Usage, then Manage limits in the Account usage section2026-09-28
TwilioUsage triggerAlert: a webhook to your applicationThe trigger calls your callback URL, once per period for a recurring trigger; stopping usage is not stated in Twilio’s usage trigger docsThe Usage Triggers API, with a callback URL2026-09-28

Where each row comes from: OpenAI’s rate limits guide covers usage tiers and the organization-wide usage limit that sits apart from your own spend limits (the last question below explains it). Anthropic’s workspace docs cover workspace spend limits, which can be set lower than the organization’s but not higher. Google Cloud’s budget docs name the spend cap budget as the alternative “if supported for your service”, and add that Pub/Sub notifications can automate cost tasks “such as programmatically disabling Cloud Billing on a project”; before wiring that up, read what a Google Cloud billing kill switch costs you. Google Maps Platform’s cost controls warn that the quota and billing systems are separate, so the two numbers can disagree. AWS’s billing alarm guide and AWS Budgets cover the two AWS routes. Vercel’s spend management docs say that with Pause Production Deployments enabled, “When your team reaches the spend amount, Vercel automatically pauses the production deployment for all projects on your team”, that “Pausing is not instantaneous”, that the feature “is available on Enterprise and Pro plans (Enterprise teams require the Flexible Commitment plan)”, and that enabling it takes “an Owner or Billing role on your team”. Replit’s spend docs say usage limits “cap your spending beyond your monthly credits”. Twilio usage triggers are “a webhook that notifies your application of usage thresholds”, so your app has to pass the message on to a person.

My working rule for the thresholds: about half, most and all of a normal month’s spend, every one of them below the app’s own ceiling, sent to a shared inbox and a chat channel rather than one person’s address.

Email, SMS and the other metered services

A public “send me a code” or “invite a friend” form is a metered API call that anyone on the internet can trigger. The named fraud here is SMS pumping, which Twilio’s explanation of SMS pumping describes as fraudsters taking “advantage of a phone number input field to receive a one-time passcode, an app download link, or anything else via SMS”, sent to numbers where “the fraudsters get a share of the generated revenue”. The same layers apply: the provider key on the server, a limit per user and per destination number, and a daily ceiling for the whole app. Bot protection on the form is a separate control and does not replace the ceiling. Where a provider has no alert of its own, measure in the app: a daily job that sums the usage table and messages the owner when the total crosses a threshold.

How to verify each one

Usage caps and spend alerts are verified by triggering them: 7 checks, from an anonymous call that is refused to a lowered threshold whose email reaches the named owner. A cap nobody has hit and an alert nobody has received are only settings so far.

  1. 01 Call the proxied route with no session and expect a refusal (a 401, or whatever your route returns for no session) with no provider call; evidence: the response, and no provider call in your own usage table or logs
  2. 02 Search the built client bundle and the browser network tab for the first dozen or so characters of each real key and find nothing (a bare prefix such as sk- also matches ordinary words like task-, so it proves nothing); evidence: the empty search
  3. 03 As a test user, go past the per-user limit and expect a 429 with the provider never called; evidence: the 429, and no new provider call in your usage table
  4. 04 In staging, set the app-wide ceiling to a tiny number, cross it with one test user, then confirm a second test user is refused too; evidence: both refusals
  5. 05 Set a test user counter so exactly one request worth of units remains, fire two requests at once, and confirm only one passes; evidence: the counter row and the two responses
  6. 06 Lower each provider alert threshold below current usage, or use the provider test action where its docs name one, wait out the reporting delay before calling it failed, and restore the threshold afterwards; evidence: the alert and its timestamp
  7. 07 To test that the spend alert email is delivered, check the named owner inbox, the spam folder and the chat channel; evidence: who received it and when, repeated whenever the owner changes

The reporting delay in check 6 is real: AWS Budgets information “is updated up to three times a day”, and Google says the first budget email or notification “may take several hours” after you create a budget. Run checks 3 to 5 in staging or with a test account, since they spend real units.

On the Production Hardening Sprint, we verify deliverable 2.5 by testing unauthorized calls, per-user limits, and configured usage ceilings, and deliverable 2.7 by triggering a test threshold and confirming the alert reaches the designated owner.

Where the sprint does this

Both deliverables are recorded in the production readiness report, deliverable 13.1, which accounts for all 123 IDs, keeps failures visible until resolved and explains genuine non-applicable items. The fee covers our engineering work: hosting, paid tools and API usage are paid through your own accounts, and we explain any required costs before enabling them. Each deliverable and its verify line is listed in the published scope.

Common questions about capping API usage and spend

How can I see my API key usage on OpenAI?

Use OpenAI’s Usage API with an Admin key: the completions usage endpoint accepts optional lists of API key IDs and project IDs, and can group results, so you can break usage down per key. OpenAI’s cookbook adds that the default usage and cost dashboards are enough for most users, and that the API is for more detailed data or a custom dashboard.

How to set an AWS billing alert?

Enable billing alerts first (Billing and Cost Management console, Billing Preferences, Receive CloudWatch Billing Alerts), switch the console to US East (N. Virginia), where AWS stores billing metric data, and create a CloudWatch alarm on the Total Estimated Charge metric that notifies an Amazon SNS topic with your email on it. After billing alerts are first enabled, AWS says it takes about 15 minutes before you can see billing data and set alarms. AWS Budgets is the other route, with notifications to an email address, an SNS topic or both; a budget action can apply an IAM policy or a service control policy, or target EC2 or RDS instances.

What does API abuse mean?

API abuse is use of an endpoint beyond what it was built for: scraping data in bulk, looping a paid AI call from a script, or pushing texts to number ranges that pay a fraudster, which Twilio calls SMS pumping. The endpoint does what it was coded to do; what is missing is a limit on who can call it and how often.

What is the API usage limit?

It is two different limits. The provider’s limit sits on your account: OpenAI, for example, sets “an approved monthly usage limit for each organization” based on its usage tier, separate from any spend limit you configure yourself. Your app’s limit sits on each of your users, and you set it: the provider’s limit protects the provider and your total, while only yours stops one account from using everyone’s share.