13 of the 21 third-party apps I audited in June and July 2026 had no rate limiting on their most expensive endpoint. Rate limiting in API routes and public forms is a counter per caller, per time window, that refuses the excess with a 429 status. The work is choosing the key, the starting number and the test.

What is rate limiting in API and web app endpoints, and where does the limit go?

Rate limiting in an API or a web app is a counter per key per time window with a defined refusal: a 429 Too Many Requests status, which may carry a Retry-After header. The key is the user, the address, the endpoint or the API key, and choosing which key protects which endpoint is most of the work.

That opening number comes from apps I chose to audit: the 21 are 11 public third-party apps I audited exhaustively across all 12 pillars and 10 held-out third-party apps audited blind, in June and July 2026. They are a selected set, not a random sample, and 13 of 21 is not a rate for AI-built apps in general. Rate limiting is one of the controls in web app security, the one that stops a single caller from using the app as much as it likes.

MDN’s page on 429 Too Many Requests defines the status as a client that “has sent too many requests in a given amount of time”, and notes that restrictions “may be server-wide or per resource”. A limit is there to prevent API abuse of three kinds: spam through public forms, outages from one runaway client, and bills from someone looping an expensive call. An API rate limit is only as useful as the key it counts, and the table below is my reading of what each key protects and what it cannot see, not a vendor statement.

KeyWhat it protectsWhat it cannot see
Per authenticated userExpensive and personal endpoints: AI calls, exports, account dataAnonymous abuse before login
Per address (IP)Public forms, signup, the login URLWho is behind a shared address: one office or mobile network shares one count
Per endpointThe backend behind one route, as a global ceilingWhich caller caused the load, so it punishes everyone at once
Per API keyA public API with paying customersAny caller without a key; it only exists where keys exist

Missing API rate limiting sits under OWASP API4:2023 Unrestricted Resource Consumption. MITRE’s name for the same weakness is CWE-770, Allocation of Resources Without Limits or Throttling, so a report that finds no rate limiting can cite that CWE. Testing the other categories on that list is its own job, covered in how to test OWASP Top 10 vulnerabilities.

Rate limiting best practices for a small app: starting numbers by endpoint

The table is my working rule for a small SaaS before its first load test, not a standard and not any provider’s default; I move the numbers once load testing a web application shows what real traffic looks like. The best practice I hold to for API rate limiting is simple: a good rate limit for an API is one a normal user never reaches and a script reaches within seconds.

For login, the rate limit best practice I follow counts failed attempts per account and per address together and slows the caller down instead of locking the account.

EndpointKeyStarting limit (my working rule)What the refusal says
LoginPer account and per addressAbout 5 to 10 failed attempts in 15 minutes, then a delay rather than a lockout429, with a message that does not say whether the account exists
Signup and password resetPer addressAbout 3 to 5 an hour429, with Retry-After where the limiter sets one
Contact and other public formsPer addressAbout 5 to 10 an hour429, and a short message the form can show
AI or export endpointPer userAbout 10 to 20 a minute, plus a daily cap sized to what one user’s calls cost you429, plus a plain message when the daily cap is hit
Public read APIPer API keyAbout 60 to 120 a minute429, with Retry-After
Everything else behind loginPer userA ceiling of a few hundred a minute, so a runaway script stops before the database does429

MDN says a Retry-After header “may be included” with a 429, so it is optional, and edge rules on Cloudflare and Vercel return 429 by default. The 15-minute and hourly rows of my working rule need a longer counting window than the free and entry edge plans offer. Cloudflare’s Free plan counts over 10 seconds and Pro over periods up to 1 minute. Vercel’s WAF rate limiting counts over 10 seconds to 10 minutes on Hobby and Pro, and Hobby is for non-commercial, personal use only. On those plans the login, signup and form rows live in the app, with a shared store (the Redis section below).

The two form rows are my rate limiting checklist for public forms: count per address, refuse with a 429 and log each refusal. That is also how to rate limit a public endpoint with no login in front of it, because the address is the only key it has. What each field may hold is a separate control, the input validation web app forms need.

For the AI row, the route-level middleware and the per-user token cap are in how to rate limit an AI or LLM endpoint; the rest of AI feature hardening, from prompt handling to timeouts, starts at what is prompt injection.

What goes wrong without it

Rate limiting helps prevent three costs an open app pays in spam, outages and bills, and one kind of caller can cause all three. The table below answers how to protect a backend API from being abused, one failure at a time.

FailureWhat the owner seesThe first limit to add
The billA free account loops the AI or export endpoint, and the provider’s invoice is the first signPer user, with a daily cap
The spamA contact form or signup fills the database with junk and burns the email quotaPer address
Credential stuffingA script tries password after password against the login formPer account and per address, the login row above
The outageOne client’s retry loop slows the app for everyoneA per-user ceiling

To add rate limiting to login, count failed attempts per account as well as per address. A script that rotates addresses stays under a per-address count, as the case below shows. The rest of the login flow, from sessions to resets, is in authentication best practices. The outage row is the quietest: a lack of rate limiting on routes behind login means one buggy client retrying in a tight loop can starve everyone else.

The UK Information Commissioner’s penalty notice to 23andMe, dated 5 June 2025, shows the credential-stuffing row at scale. 23andMe told the regulator its rate-limiting rules were set up to limit traffic from a given IP address, and that they were not triggered because the attacker rotated thousands of unique IP addresses. The Commissioner found the company had failed to implement effective rate-limiting rules and alerts. The attack ran between April and September 2023 and led to unauthorized access to personal information of 155,592 UK residents. The lesson I take from it: a missing or bypassable limit shows up as someone else’s success, a bill or an outage before it shows up in any review.

The bill row has a number from my own work. In the same June and July 2026 audits, 12 of the 14 AI apps had a confirmed denial-of-wallet path, where a stranger or free account can burn the owner’s paid AI or compute bill without limit. The 14 are the third-party apps I audited that have an AI or LLM surface. The 13 of 21 figure at the top of this page comes from those audits too, and both counts describe the apps I picked, not AI-built apps at large.

Two neighbors of a rate limit are different controls. The Supabase built-in email sender has its own limit, which is a separate thing from the limits you set, covered in the Supabase email rate limit. Usage caps and spend alerts on a metered provider sit beside a rate limit: the limit slows one caller, the cap stops the month’s total, and the steps are in cap monthly usage on a metered API.

IP-based rate limiting that trusts a header the caller sets

An app behind a proxy or a hosting platform sees the proxy’s address on every request, so IP-based rate limiting reads the client address from a forwarding header instead. MDN’s X-Forwarded-For reference calls the header a de-facto standard for identifying the originating client through a proxy, and its security section says any security-related use, rate limiting included, “must only use IP addresses added by a trusted proxy”. It names the result of ignoring that: “rate-limiter avoidance”.

I audited an AI coding workspace that keyed its AI rate limit on the x-forwarded-for header, which the caller sets, so unless the host’s edge rewrote that header, a new fake value on each request got a new bucket. The whole limiter also returned “not limited” whenever its key-value store settings were missing: it failed open.

My take: a limit is only as strong as the key it counts and what it does when its store is gone. My fix is to key per-address limits on the address your host’s docs say to trust, never the raw header, and to decide the store-down behavior on purpose, which the Redis section covers. The spoofed-header check in the burst test below catches this version.

How do you handle rate limiting?

Rate limiting in an API is handled in three places: per-address rules for public forms and the login URL, at the edge where the plan’s window allows it, else in the app; per-user limits in the app for expensive endpoints; and per-key limits at a gateway for a public API. Which key each layer counts matters more than the algorithm.

Handling rate limits and preventing abuse on a small app takes a library or an edge rule, not a service you build. An API rate limiter of your own, the kind system-design interviews ask for, solves a scale problem one app does not have. The rate limiting strategies below come in the order a small app meets them, and together they cover how to implement rate limiting without new infrastructure.

Rate limiting vs throttling, and the rate limiting algorithms that matter

Rate limiting refuses the excess and throttling slows or queues it. Fixed window, sliding window and token bucket differ mainly in how they treat bursts and how much they store; for one app, the library’s default is fine, and the choice starts to matter at gateway scale.

API throttling vs rate limiting comes down to what the caller experiences: a throttled request waits and is served later, while a rate-limited one is told no, sometimes with a time to try again. Rate limiting vs spike control is a MuleSoft distinction. MuleSoft’s Spike Control policy limits the number of messages an API processes within a set time and queues a request over the limit for retry, as the policy is configured. It rejects the request after the set number of retry attempts, and it does not enforce quotas.

The types of rate limiting algorithms, as Redis’s tutorial describes and stores them:

AlgorithmHow it countsHow it treats a burstWhat it stores
Fixed window counterCounts requests in fixed, non-overlapping time intervalsAllows 2x burst at window boundariesOne key per window (a string)
Sliding window logRecords the timestamp of every request and prunes the old onesNo burstsOne entry per request in the window (a sorted set)
Sliding window counterBlends the current and previous window counts by a weighted averageSmoothed boundariesTwo counters
Token bucketTokens refill at a steady rate up to a maximum; each request spends oneAllows controlled bursts up to the bucket sizeToken count and last refill time (a hash)
Leaky bucketThe bucket drains at a fixed rate; overflow is rejected (policing) or delayed (shaping)No bursts (steady drain)Fill level and last drain time, or the next free slot (a hash)

The best rate limiting algorithm for one app is the one your library documents as its default. Redis’s own tutorial says “There’s no single best algorithm” and names the sliding window counter as the best balance of accuracy, simplicity and low memory for most APIs. The token bucket is the rate limiting algorithm behind Go’s rate package and AWS API Gateway’s throttling. In my reading, the choice starts to matter when one gateway counts for many services, not for one small app.

Rate limiting with Redis when the app runs on more than one server

Redis is where the counter goes when the app runs on more than one server or on a serverless host: every instance reads and increments the same key, so a burst split across instances is still counted against one limit.

A counter in memory is not enough on a serverless host or with two servers, because each instance counts on its own. express-rate-limit’s README says it comes with “a built-in memory store” and takes a custom store “to share hit counts across multiple nodes”; that a memory count splits across instances is my reading of that line.

Redis’s rate limiting tutorial builds its fixed window from INCR and EXPIRE wrapped in one Lua script run as a single atomic EVAL call. It says why: “Without this, a crash between the two commands could leave a key with no expiry, permanently blocking that client.” For exact counts, the tutorial’s sliding window log gives better rate limiting with Redis sorted sets: it uses ZREMRANGEBYSCORE, ZCARD, ZADD, ZRANGE and EXPIRE, at the cost of memory that grows with each client’s request volume.

Here is a Redis rate limiting example for an Express route, using the same script shape and client call as the tutorial; the limit and the window in it are placeholders for the numbers in the starting table. It assumes a connected client named redis and an auth middleware that has already set req.user.

const SCRIPT = `
local count = redis.call('INCR', KEYS[1])
if count == 1 then redis.call('EXPIRE', KEYS[1], ARGV[1]) end
return { count, redis.call('PTTL', KEYS[1]) }`;

async function limitPerUser(req, res, next) {
  const [count, pttl] = await redis.eval(SCRIPT, { keys: [`rl:${req.user.id}`], arguments: ['60'] });
  if (count <= 10) return next();
  res.set('Retry-After', String(Math.ceil(pttl / 1000)));
  res.status(429).send('Too Many Requests');
}

On Next.js and Vercel, Upstash’s rate limiting SDK reaches Redis over HTTP, and its docs list serverless functions, Vercel Edge and Next.js among the places it is designed for.

Then decide what happens when the store is unreachable. My working rule: the AI, export and login endpoints fail closed, ordinary pages fail open, and the choice is written down where the next developer will find it. express-rate-limit exposes this as passOnStoreError, which blocks traffic by default when the store becomes unavailable, so a fail-open route sets it on purpose. Upstash’s SDK documents the opposite default: when the Redis call does not resolve within its timeout, it allows the request. So Redis does become a single point of failure, but only for the routes you chose to fail closed. The fail-open limiter in the IP section above is what the other choice looks like when nobody makes it on purpose.

Rate limiting in Next.js and other frameworks (Express, NestJS, Laravel, Rails, Flask, FastAPI, Go, Spring, C#)

Each row names the package or built-in and where it keeps the count by default. The last column is the one that matters: whether that count survives a second instance is my reading of each project’s docs, not their wording.

FrameworkThe package or built-inWhere the count lives by default
Next.js on VercelUpstash’s rate limiting SDK, over HTTPRedis (Upstash)
Express (Node.js)express-rate-limitA built-in memory store; a custom store shares hit counts across nodes
NestJS@nestjs/throttlerA built-in in-memory cache; alternate storage providers supported
LaravelRateLimiter::for and the throttle middlewareTypically the application’s default cache
RailsThe controller’s rate_limit, or Rack::AttackAn ActiveSupport::Cache store (config.action_controller.cache_store); Rack::Attack uses Rails.cache if present
FlaskFlask-LimiterMemory in the quick start; Redis, Memcached or MongoDB through the limits library
FastAPISlowAPIRedis, Memcached or memory, with memory as a fallback
Gogolang.org/x/time/rateA Limiter value in your process; shared storage not stated in Go’s docs
Spring BootBucket4jLocal in-memory buckets; clustered back ends include JCache, Hazelcast and Redis
ASP.NET Core (C#)Microsoft.AspNetCore.RateLimitingIn the app’s memory: each key creates and caches its own limiter

For Next.js API rate limiting on Vercel, the count has to leave the function, which is why that row points at an HTTP Redis client rather than a middleware package. For rate limiting in Node.js with Express, the express-rate-limit package on npm takes windowMs and limit, answers with 429 by default and keys on the IP address unless you pass a keyGenerator. The Express rate limit best practice I would put first: give it a Redis store before the second instance exists, and key expensive routes on the user, not the address.

Rate limiting in NestJS comes from the throttler module, which identifies users by their IP address by default (NestJS rate limiting docs). Laravel API rate limiting defines named limiters with RateLimiter::for and attaches them with the throttle middleware, backed by the cache (Laravel’s rate limiting docs). Rails API rate limiting can use Rails’ rate_limit in a controller, which refuses with 429 Too Many Requests by default, or Rack::Attack, which describes itself as Rack middleware for blocking and throttling abusive requests.

In Python, Flask API rate limiting runs through Flask-Limiter, and to rate limit FastAPI or Starlette API calls you use SlowAPI, which its docs say was adapted from flask-limiter and is still alpha quality code. Rate limiting in Golang starts with Go’s rate package, whose Limiter implements a token bucket. To implement rate limiting in Spring Boot, Bucket4j is a Java library based on the token-bucket algorithm, with clustered back ends for more than one node. To implement rate limiting in an API in C#, use ASP.NET Core’s rate limiting middleware: you configure policies, attach them to endpoints and can set RejectionStatusCode to 429 for the refusal. Microsoft says apps using it should be carefully load tested and reviewed before deploying.

To rate limit Supabase apps, treat your own routes and Edge Functions like any other backend: the same middleware or edge rule applies, and the Supabase email limit mentioned above is a separate thing.

API gateway rate limiting and edge rules (AWS, Azure APIM, Cloudflare, nginx, Vercel)

API gateway rate limiting and edge rules limit by API key or by address before the request reaches the app. They suit a public API and the public forms. Neither sees the logged-in user, so per-user limits on expensive endpoints stay in the app.

Cloudflare’s rate limiting rules match requests with an expression and act when the rate is reached; a Cloudflare rate limiting rule example in its own syntax matches http.request.uri.path eq "/login", with a period, a request count and a duration set beside the expression. On Cloudflare rate limiting pricing, the plan decides what the rule can do. Per the docs updated August 25, 2026, the Free plan gets 1 rule with an expression on the path, counting by IP over 10 seconds. Pro gets 2 rules and periods up to 1 minute, and Business gets 5 rules and periods up to 10 minutes. The best practice I follow for Cloudflare API rate limiting on the Free plan is to spend the one rule on the login or signup path and keep the per-user limits in the app.

AWS WAF rate-based rules count incoming requests, group them by your criteria and rate limit each group over an evaluation window, and a rule can take a scope-down statement that narrows which requests it tracks. The AWS WAF rate limiting best practice I would add is that scope-down, to the login and form paths, so the rule counts only what it protects. Google’s Cloud Armor rate limiting has two rule types. Throttle caps requests per client or across all clients at a threshold, and rate-based ban bans a client for a set time after it exceeds one. Rate limiting with Cloud Armor can answer with a 429, the code Google recommends.

Vercel’s WAF rate limiting is available on all plans. Hobby gets 1 rule per project and Pro gets 40, both with a 10-second to 10-minute window and fixed-window counting keyed on IP or JA4 digest. Counters are tracked per region, so traffic from several regions can exceed the limit you set for any single region. Vercel rate limiting on Hobby also carries the plan condition noted under the starting numbers. Rate limiting with nginx, the self-hosted edge, uses nginx’s limit_req module, which limits the request rate per key, in particular per client address, with the leaky bucket method, and refuses with 503 by default. The nginx rate limit best practice I apply is setting limit_req_status to 429 so clients read the refusal as a rate limit.

On AWS, API gateway rate limiting goes by the name throttling, so API gateway throttling vs rate limiting is a naming difference there, not a different control. AWS API Gateway throttling uses a token bucket. The default is 10,000 requests per second per account per Region, with a burst bucket of up to 5,000 requests (lower in some Regions), and clients over it may receive 429 Too Many Requests. Per-client limits on AWS come from usage plans and API keys, which are for REST APIs. AWS attaches its own conditions: usage plan throttling and quotas “are not hard limits, and are applied on a best-effort basis”, and AWS says not to rely on them to control costs or block access, or to use API keys for authentication or authorization.

APIM rate limiting in Azure uses two policies. Azure API Management’s rate-limit policy limits calls per subscription, rate-limit-by-key limits them per any key a policy expression computes, and both return 429 Too Many Requests. Microsoft notes that rate limiting “is never completely accurate”, and it treats quotas, usually used over longer periods, as a separate policy. For Azure App Service rate limiting, the Microsoft pages I read cover the APIM policies and the ASP.NET Core middleware; a limit built into App Service itself is not stated in Microsoft’s docs. For GCP API Gateway rate limiting, GCP API Gateway quotas lists the gateway’s own default of 10,000,000 quota units per 100 seconds per service producer project. Per-caller quotas work as in Cloud Endpoints, which counts requests per minute per consumer Google Cloud project, identified by the API key each request sends.

Kong API gateway rate limiting runs through Kong’s Rate Limiting plugin, which counts from per second up to per year and keeps counters in memory on each node, in Kong’s data store or in Redis. The in-memory option diverges as nodes are added. A rate limit service for Envoy is the step after that: Envoy’s global rate limiting calls a gRPC rate limiting service, with a reference implementation in Go on a Redis backend. Kubernetes gateway API rate limiting and rate limiting in a service mesh are for teams already running that stack.

LayerWhat it limits byWhat it cannot see
Edge (Cloudflare, AWS WAF, Cloud Armor, Vercel, nginx)The client address by default, plus the path; AWS WAF and Cloud Armor can also count by a header or cookieThe logged-in user
Gateway (AWS API Gateway, Azure APIM, Kong)The API key or subscription, or a key a policy computesAnything the key does not carry
App (framework middleware)The user, the account or the keyA flood large enough to take the app down before the limit runs

The “cannot see” column is my reading. It is why API gateway rate limiting per user means the gateway reading a user claim into its key, as rate-limit-by-key allows, or the app counting per user itself. An application limit cannot absorb a flood either: Microsoft’s ASP.NET Core docs say rate limiting is “not a comprehensive solution for Distributed Denial of Service (DDoS) attacks”, and stopping one is the edge’s job, starting with how to set up a web application firewall.

How to fix rate limit exceeded on someone else’s API

Rate limit exceeded on someone else’s API means your app sent more than the provider allows in its window. Read the status and headers, wait the time they state, retry with backoff and a cap, cache what you can, and move the work off the request. Ask for a higher limit last.

Rate limit exceeded is the provider’s 429, or its own error, telling you your app passed its window; “rate limit exceeded, try again later” is the same answer in friendlier words. How long a rate limit lasts is whatever the provider’s window is, and the response states it when the provider sends a Retry-After header, in seconds or as a date, or a reset header with a time.

GitHub’s REST API rate limits are the worked example: 60 requests per hour unauthenticated and 5,000 per hour for an authenticated user, with x-ratelimit-limit, x-ratelimit-remaining, x-ratelimit-used and x-ratelimit-reset (in UTC epoch seconds) on each response. When the GitHub API rate limit is reached you get a 403 or 429 with x-ratelimit-remaining at 0, and GitHub says not to retry until the x-ratelimit-reset time. It also warns that continuing to send requests while limited may get the integration banned. Jira Cloud rate limiting on its REST API combines an hourly points quota, per-second burst limits and per-issue write limits, and returns 429 Too Many Requests with a Retry-After header to respect.

Backoff and its cap are their own topic (what is exponential backoff), and so is moving the call into a queue (how to run long tasks in the background). Caching is the cheapest way to avoid rate limiting on data that changes rarely, since a cached answer needs no new call to the provider.

To keep your own calls under the provider’s limit, Bottleneck, on npm, is a lightweight, zero-dependency task scheduler and rate limiter for Node.js and the browser. p-throttle throttles promise-returning and async functions and rate-limits calls without discarding them. The Supabase email version of “exceeded” has its own page, linked above.

How to test API rate limiting: the burst test and the evidence to keep

API rate limiting is tested six ways: a burst from one key gets a 429 at the limit, a second account is unaffected, a spoofed address header changes nothing, a normal user is never blocked, the window recovers on time, and the count survives a deploy. Keep the script and its output.

Run every check against staging or your own production endpoint, never against someone else’s.

  1. 01 A burst above the limit from one account gets 200s up to the limit, then 429, with a Retry-After header where the limiter sets one. Keep the script output showing the status sequence and any header. On an edge rule the cut is not exact. Cloudflare says excess requests could still reach the origin before its counters update, and Vercel counts per region. There, a pass is a 429 within a few requests of the limit, with every later request in the window refused.
  2. 02 The same burst from a second account, sent with that account's own token or cookie (swap the Authorization header or cookie, never reuse the first account's), in the same window gets 200s. Keep both outputs side by side. This is how you test per-user API limits.
  3. 03 On a per-address limit, the burst with a different X-Forwarded-For value on each request still gets 429 at the limit. Keep the output. A 200 past the limit is the header-trusting failure from the IP section.
  4. 04 A normal user's pattern, a page load with its requests and a login with one typo then the right password, never sees a 429. Keep the browser's network log. This is the test that the login rate limit works without locking real people out.
  5. 05 After the window, the next request succeeds. Keep the timestamps.
  6. 06 Redeploy, or wait for a cold start, in the middle of a window, then continue the burst: it is still refused. Keep the output around the deploy.

The burst script, with a placeholder URL and token. Set the loop count above your limit, and for the spoofed-address check add the commented header to the curl line.

URL="https://staging.example.com/api/<your-endpoint>"
TOKEN="<a test account's token>"
# check 3: add  -H "X-Forwarded-For: 203.0.113.$i"
for i in $(seq 1 30); do
  code=$(curl -s -o /dev/null -D - -H "Authorization: Bearer $TOKEN" "$URL" \
    | grep -i -E '^HTTP/|^retry-after' | tr -d '\r' | tr '\n' ' ')
  echo "$(date +%T) request $i: $code"
done

Three rate limiting testing tools cover these checks: the short script above, Postman and k6. To test rate limiting in Postman, Postman’s collection runner sends a collection’s requests in the order you choose, for a set number of iterations with an optional delay, and logs the test results for each request. k6 is what Grafana calls an open-source, developer-friendly and extensible performance testing tool, which suits a longer burst. The same burst is one of the checks in a wider website security check.

In the Production Hardening Sprint, we verify deliverable 3.6 this way: simulate bursts and verify enforcement without blocking ordinary use; for deliverable 1.6 we simulate repeated attempts and verify limits, responses, and normal-user recovery.

Where the sprint does this

In the sprint, we protect public forms, signup flows, and resource-consuming endpoints with appropriate limits (deliverable 3.6), and implement rate limits and brute-force protections on authentication and recovery endpoints (1.6). Deliverable 3.12 places a CDN or web application firewall in front of the origin, with rate and abuse rules that absorb floods before they reach the application. Each result goes into the production readiness report with its evidence, and the report accounts for all 123 IDs, keeps failures visible until resolved and explains genuine non-applicable items. Third-party hosting, service subscriptions and API consumption remain in the client’s own accounts, and required service costs are explained before anything is enabled. Every deliverable and its verify line is in the published scope.

Common questions about rate limiting

How to rate limit an API to 10 requests per minute?

Use a 60-second window with a limit of 10 per key. In express-rate-limit that is rateLimit({ windowMs: 60 * 1000, limit: 10 }), which refuses the eleventh request in the window with a 429. If more than one instance serves the API, give it a shared store such as Redis so every instance counts against the same ten.

Is it good practice to rate limit API calls?

Yes, on every public endpoint and every endpoint that costs you money, with limits set where a normal user never reaches them. Without one, a single script can fill your forms, try passwords or run up your provider bill.

What are the downsides of rate limiting?

A limit set too low blocks real users, and everyone behind one shared address, such as a whole office or a phone carrier’s network, shares one per-address count. That is why expensive routes count per user and why the burst test includes a normal-user check.

Is HTTP 429 a rate limit?

Yes. 429 Too Many Requests is the status MDN defines for a client that has sent too many requests in a given amount of time, and a Retry-After header may be included to say how long to wait.

What is a good API rate limit?

For a public read API on a small app, I start at about 60 to 120 requests a minute per API key, as my working rule rather than a standard, and move it after a load test shows real traffic.

How do I know if I am rate limited?

You get a 429 status or the provider’s own error, and the rate limit headers say so: GitHub, for example, sends x-ratelimit-remaining at 0 and an x-ratelimit-reset time in UTC epoch seconds.