OpenFeature offers one vendor-neutral API for feature flags, and comparisons of open source feature flags are often written by a vendor whose tool is in the comparison. What both leave out is the smallest version that works: one table with a key and a boolean, read on the server, which turns a broken feature off with no service to run and no deploy.
What is a feature flag: the smallest useful version
A feature flag is a named condition in code that decides at run time whether a feature is on, so it can change without a deploy. The smallest useful version is one database row with a key and a boolean, read on the server and cached briefly.
A flag is one control in the release path of a live app, and that path as a whole, from repository to production, is DevOps for startups.
Pete Hodgson’s “Feature Toggles” on martinfowler.com sorts flags into four categories by how long they live and how dynamic the decision is:
- Release toggles let “incomplete and un-tested codepaths” ship to production as latent code that may never be turned on.
- Experiment toggles run multivariate or A/B tests, sending each user down one path according to their cohort.
- Ops toggles “control operational aspects of our system’s behavior”, and Hodgson notes it is “not uncommon” for systems to keep a small number of long-lived ones, which he calls kill switches.
- Permissioning toggles change “the features or product experience that certain users receive”, such as premium or beta features.
The rest of this article is about the ops toggle. Hodgson describes kill switches as flags that “allow operators of production environments to gracefully degrade non-vital system functionality” when the system is “enduring unusually high load”, and adds that “needing to roll out a new release in order to flip an Ops Toggle is unlikely to make an Operations person happy.” That second line is the whole case for keeping the value outside the build, and the same switch serves a broken feature or a failing dependency, not only high load.
The smallest version is a table with four columns and one check at the server boundary. The table:
create table public.feature_flags (
key text primary key,
enabled boolean not null default true,
updated_at timestamptz not null default now(),
updated_by text not null
);
And the check, in the API route or server action, before the risky work starts:
if (!(await isEnabled('ai_generation', false))) {
return featurePaused(); // your designed off state
}
An environment variable looks like a flag but cannot do this job on Vercel or Netlify. Vercel’s environment variables docs say changes “only apply to new deployments”, and Netlify’s environment variables docs say “Environment variable changes require a build and deploy to take effect.” A value that needs a deploy to change is no help when the deploy is the slow part.
The flag is read on the server or at the edge, and the browser is only told the result: whether to show the entry point. A flag that only the browser reads hides a button and stops nothing, which matters again in step 5 of the build below.
Open source feature flags: a table you own, a flag server, or OpenFeature
Open source feature flags come in 3 shapes: a table you own, a self-hosted flag server such as Unleash, Flagsmith, GrowthBook or GO Feature Flag, and OpenFeature, which is not a flag server at all but a vendor-neutral SDK standard that sits in front of any of them.
Open source still means something runs. Each server below is one more service to run, with its own upgrades and its own outage mode, and most need a database of their own (GO Feature Flag needs none). The “what you run” column comes from each project’s docs and repository, read in October 2026. The “when it is worth it” column is my working rule, not a vendor claim.
| Option | What it is | License | What you run | Beyond on and off | When it is worth it |
|---|---|---|---|---|---|
| A config table you own | One database table and a server helper | Your own code | Nothing new: the database you already have | A percentage rollout, if you hash the user id | First, and for every kill switch |
| Unleash Open Source | A feature management server with an admin UI | AGPL-3.0 | The Unleash server and PostgreSQL | Activation strategies, over 30 SDKs | Targeting rules, roles, people outside engineering flipping flags |
| Flagsmith open source | A flag and remote config platform | Mostly BSD-3-Clause, some MIT; enterprise governance features need an Enterprise license | The Flagsmith API and PostgreSQL, on Docker or Kubernetes | Remote config, experimentation | The same needs as Unleash, plus remote config values |
| GrowthBook | Flags plus experiment analysis | Open core: mostly MIT, some directories under a commercial license | The app and MongoDB | An experiment stats engine | When you will run real experiments |
| GO Feature Flag | Flags in a configuration file, served by a relay proxy | MIT | The relay proxy; the file sits in S3, GitHub or another store | Targeting rules, rollouts, OpenFeature SDKs | When you want flags in a file and no database |
| PostHog feature flags | Flags inside PostHog, next to its events and session replays | MIT expat, except the ee directory | PostHog’s cloud, or a hobby self-host on Docker | Targeting by person property, cohort or group | When PostHog already runs in your app |
What an SDK does when its flag server is unreachable decides whether a flag outage becomes an app outage. Unleash Open Source documents that once initialized, “all Unleash clients continue to function” through a server outage, and that a client with no flag data falls back to disabled or to your default. Flagsmith’s server SDKs call a default flag handler “when a flag cannot be found or if the network request to the API fails”. GrowthBook’s JavaScript SDK docs say its init call does not throw on network issues; the SDK stays in a default state where every feature evaluates to null. GO Feature Flag’s docs promise a value back every time, the default one if anything goes wrong. With local evaluation, PostHog’s server SDKs fetch flag definitions every 30 seconds by default (5 minutes in the Go SDK), and PostHog’s docs say an unreachable PostHog “doesn’t affect flags that are already cached”; per-flag defaults cover a cold start. The table version fails the way your helper says it does, which is the next section’s job.
Feature flags without a paid service therefore have two honest answers: the table, or a self-hosted server whose cost is your time plus an instance you pay your host for. Free tiers of hosted flag services, where they exist, change too often to quote here. PostHog’s own repository says its open-source deployments “should scale to approximately 100k events per month” and come with no customer support.
My working rule: start with the table, and move to a server when you need targeting rules, an audit trail with roles, or people outside engineering flipping flags.
OpenFeature: an open standard in front of the flags, not a flag service
OpenFeature is an open standard and a set of SDKs for evaluating feature flags, hosted by the Cloud Native Computing Foundation. It stores no flags. Your code calls one evaluation API, and a provider plugged in behind it fetches the value from whichever flag system you run.
OpenFeature’s documentation calls it “an open specification that provides a vendor-agnostic, community-driven API for feature flagging”. OpenFeature’s CNCF project page records its acceptance on June 17, 2022 and its move to the Incubating maturity level on November 21, 2023. The name causes some confusion: OpenFeature is an open standard, and your feature flags still live in a table or a flag server behind it.
Three of its concepts matter first. The evaluation API is what your code calls. A provider is the “translation layer” between that API and your flag system, and it may wrap a vendor SDK, call a REST API or read a local file. Hooks add behavior at points in the evaluation life cycle, such as logging or validation.
The OpenFeature SDK comes in server versions for C++, Dart, .NET, Go, Java, Node.js, NestJS, PHP, Python, Ruby and Rust, and client versions for Dart, Kotlin, iOS, Web, Angular and React. For a small team, the gain is that the flag calls in your code stay the same when the backend changes, so the table can become a server later without touching every call site. The specification also says evaluation calls “must always return the default value in the event of abnormal execution”. What it does not do is store, target or serve flags.
The simple providers are the bridge from the table version: the Node.js SDK ships a TypedInMemoryProvider, and OpenFeature’s contributed providers include an environment variable provider, @openfeature/env-var-provider. My working rule is to adopt OpenFeature on day one only if a second flag backend is likely; otherwise the isEnabled helper is the same seam with less to learn.
import { OpenFeature } from '@openfeature/server-sdk';
// Register the provider for whichever flag system you run, once at startup.
await OpenFeature.setProviderAndWait(new YourProviderOfChoice());
const client = OpenFeature.getClient();
const aiOn = await client.getBooleanValue('ai_generation', false); // false = the default
The package and method names follow OpenFeature’s Node.js SDK docs, where YourProviderOfChoice is the placeholder for a real provider.
What goes wrong without it
When you need to disable a feature without deploying and there is no switch, the only way back runs through the build. These are the four cases where the fastest fix is “off”. The cells are my reading of each case, not measured times.
| What broke | The fix without a switch | The fix with one | What reaches users first |
|---|---|---|---|
| A new feature throws errors for some users | Revert, rebuild, redeploy | Flip the feature’s flag off | Without: errors until the new build is live. With: a paused notice |
| A third-party service is down or rate-limiting, and its widget or call blocks the page | Wait for the provider, or ship code that removes the call | Switch the integration off and show the fallback | Without: a page that hangs or breaks. With: the page, minus that piece |
| An AI feature starts costing far more than expected | Ship a change that disables it, or rotate the API key and break it | Switch it off while the cause is found | Without: more spend for every minute of build. With: a paused notice |
| A release cannot simply be rolled back because a migration already ran | Write and ship a fix forward under pressure | Switch the new path off and leave the migrated data in place | Without: the broken path until the fix ships. With: the rest of the app |
As I read it, a host’s instant rollback covers the first case and does nothing for the second and third, because the old build calls the same failing provider and the same expensive model. Rollback steps by host belong to how to roll back a deployment on Vercel or Netlify. For the cost case, the switch buys time, and the rest of the response is how to stop a runaway API bill. The migration case is where code and data part ways, and planning for it is part of what a rollback plan is.
The number to compare a flip against is your own pipeline’s time from commit to live. Measure it once, then work on how to speed up CI build times if it is long.
How to do it: a kill switch per risky feature and per dependency
A kill switch needs 4 parts: a flag store the server reads, a check wrapped around the risky feature, a designed off state the user sees, and a way to flip it that does not depend on the thing that is broken. A deploy should never be one of the steps.
The config-table flag on Supabase, Next.js or any server
The steps below are written from Supabase’s and Next.js’s documentation and plain SQL, not from a run in a production app; rename the pieces to fit your stack.
- 01 Create the feature_flags table so only the server can read it and only admins can change it
- 02 Seed one row per switch with enabled set to true and updated_by naming who added it
- 03 Write the isEnabled helper: read through a cache of a few seconds, and return the default you chose for that flag if the read fails
- 04 Wrap the feature at the server boundary (the API route, the server action, the background job), then hide the UI entry point from the same value
- 05 Never rely on a flag the browser reads to stop a server action: the server checks the flag again on every call
- 06 Log every flip with who and when; the updated_by and updated_at columns already hold both
The helper, server only:
const cache = new Map<string, { value: boolean; at: number }>();
const TTL_MS = 5_000; // a few seconds; tune it to how fast a flip must land
export async function isEnabled(key: string, fallback: boolean): Promise<boolean> {
const hit = cache.get(key);
if (hit && Date.now() - hit.at < TTL_MS) return hit.value;
try {
// select enabled from feature_flags where key = $1 with the server's key; throw on a query error
const row = await readFlagRow(key);
const value = row ? row.enabled : fallback;
cache.set(key, { value, at: Date.now() });
return value;
} catch {
return fallback; // the flag store is down: use this flag's chosen default
}
}
The cache window and the defaults are my working rules: a few seconds of cache lets a flip reach users quickly without a query on every request, and the default should fail open for a core feature and fail closed for a risky or costly one. Write down which default each flag uses, and why.
On Supabase, turn row level security on for the table and create no policy for the browser’s roles. Supabase’s row level security guide says that once RLS is on, “no data is accessible through the API when using a publishable key, until you create policies”, and that on existing projects a new table starts with privileges granted to anon and authenticated that policies do not take back. Revoke those grants as well: a browser request then fails with error 42501 instead of quietly getting no data. Read the table on the server with a secret key, which uses the service_role role with bypassrls, from a client created without the user’s session, because a secret key “bypasses RLS only when the request carries no user access token.”
On Next.js without Cache Components, a page that shows or hides the feature must not keep a stale value. Next.js caching docs describe export const dynamic = 'force-dynamic', which renders a route “for each user at request time”, and unstable_cache for database reads, whose revalidate option is “the number of seconds before the cache is revalidated.” Use the first, or keep the revalidate window as short as the helper’s cache. On other frameworks, check your framework’s caching docs before you trust a flip to reach a cached page.
Step 5 is the one that fails quietly. An AI coding workspace I audited sold usage-based tiers, but the server never checked them: the plan tier was read only in browser components. A kill switch read the same way hides a button and stops nothing, and the server check that closes the gap for plan tiers is how to enforce plan limits on the backend.
Which features get a switch: the risky ones and every third-party dependency
Give one switch to each risky feature and one to each third-party dependency, so any single one can go off alone. The starter list below is my working list, and the last column is what the helper returns when the flag store itself cannot be read.
| Switch | What off looks like | Default if the flag store fails |
|---|---|---|
| AI generation | The button says generation is paused; saved results still open | Off: it costs money on every call |
| Outbound email sends | Sends are skipped with a log line; the action that triggered them still completes | On: password resets depend on it |
| The payment upgrade path | ”Upgrades paused” on the pricing page; existing plans and access untouched | On: a core path for revenue |
| File uploads | The upload control becomes a short notice; existing files still load | On: a core task in most apps |
| The third-party chat or analytics widget | The widget is not loaded; the page works without it | Off: the page loses nothing it needs |
| New-signup intake | A waitlist form replaces signup; existing users sign in as usual | On: signups keep coming in |
Authentication checks, authorization checks and billing webhooks never get a switch. A switch on sign-in or on a permission check is a way to open the app to everyone, and a paused webhook handler lets your billing records drift away from the payment provider’s. A kill switch is not a permission system either: who may use a feature belongs in your authorization code, not in a global on and off row.
The off state is designed, not blank: a sentence that says the feature is paused, and the rest of the page still working around it. Designing that state well is most of what graceful degradation is.
Flipping it during an incident, and cleaning up after
The flip path must not depend on the broken thing. When the app’s own admin screen is down with the app, a direct update from your database provider’s SQL editor still works, so put that one SQL line in the runbook before you need it, and make sure two people know how to run it.
- 01 Open the database provider's SQL editor, not the app's admin screen
- 02 Run the runbook line: update feature_flags set enabled = false, updated_at = now(), updated_by = 'your name' where key = 'ai_generation';
- 03 Wait out the cache window, then confirm as a signed-in user that the feature shows its paused state
- 04 Note the time, the flag and the reason in the incident log
- 05 Once the cause is fixed, turn the flag back on deliberately with the same line and enabled = true, never as a side effect of a deploy
Kill switches are permanent: list them in one file in the repository so whoever is on call can find them. Release flags are temporary and get a removal date on the day they are created. Hodgson’s article says savvy teams view their toggles as “inventory which comes with a carrying cost” and notes that some teams put “expiration dates” on their toggles. Remove a release flag once its feature is fully on for everyone; keep the kill switches.
A dependency upgrade that changes a flagged path still goes through CI like any other change, including the update pull requests a bot opens, which is what Renovate bot is for.
Flags versus experiments: where A/B testing platforms fit
A/B testing platforms and feature flags share one mechanism, a condition that picks a code path, and differ in purpose. A kill switch is on or off for everyone. An experiment assigns users to variants, keeps the assignment stable, and measures an outcome, which needs analytics a flag table does not have.
Hodgson’s experiment toggle sends each user down one path “based upon which cohort they are in”, and it has to stay in place “long enough to generate statistically significant results”. The table below sets the three jobs side by side; its last row is my reading.
| Need | Kill switch | Percentage rollout | Experiment |
|---|---|---|---|
| Who gets it | Everyone, or no one | A share of users | Users assigned to variants |
| Assignment | None | Stable bucketing by user id | Stable assignment, recorded per user |
| What you log | Who flipped it, and when | Who changed the share, and when | Each user’s exposure, plus the outcome metric |
| What else it needs | A designed off state | A hash of the user id | Analytics and enough traffic to read a result |
| The table version | Does it | Can do it with a hash of the user id | Should not attempt it |
A/B testing platforms, and the open-source tools that include experiment analysis, exist for the last column. GrowthBook, from the table above, is one: its repository lists an experiment stats engine beside the flags. A caution of my own: many apps at this stage may not have the traffic for a test to conclude, so a decision taken from a short test can be noise.
How to verify it
A kill switch is proven by a test: turning one flagged feature off in production at a quiet hour, confirming the feature is hidden and every other journey still works, then turning it back on. Record who flipped it, when, and how long the change took to reach users.
Pick a feature that is safe to pause, and before you start, write down what its server endpoint should return when paused, since that response comes from your code, not from the framework. Each check names the evidence to keep:
- 01 Note the time and flip one flag off in production with the incident path, not a deploy. Evidence: the SQL line and the time
- 02 Time how long until a signed-in user stops seeing the feature. Evidence: two timestamps, with a gap inside the cache window you chose plus any page revalidation
- 03 Call the feature's server endpoint directly with a signed-in user's session; it must refuse or do nothing, not only hide. Evidence: the response, matched against the one you wrote down
- 04 Walk signup, the main task and billing with the feature off; nothing else breaks. Evidence: a short log of each journey
- 05 In staging or a local copy, make the flag read fail (revoke the server role's select on the table, or point the helper at a missing table) and confirm each flag falls to its chosen default. Evidence: each flag's state
- 06 Flip it back on and confirm updated_by and updated_at show who and when. Evidence: the row
On Next.js, the revalidate window in check 2 is the one set on the page or on unstable_cache. On Supabase, revoking the server role’s grant in check 5 makes the read fail with error 42501; have readFlagRow throw when the query returns an error, so the helper’s catch returns the default.
Repeat the drill once per switch, then again whenever a switch is added. Running the same idea against a whole provider outage is close to what chaos testing is.
In the Production Hardening Sprint, deliverable 7.12, Feature flags and kill switches, is verified this way: disable a flagged feature in production and confirm the product stays usable with the feature hidden.
Where the sprint does this
In the Production Hardening Sprint, deliverable 7.12 adds a flag mechanism for risky features and third-party dependencies so any one of them can be turned off without a deployment. Deliverable 6.8 keeps unaffected functions usable when an external service is unavailable. Deliverable 7.6 documents and rehearses release rollback, including how database changes are handled safely. New features that change the product’s core capabilities are separate work. Hosting, paid tools, and API usage remain in your accounts. All three deliverables are listed in the published sprint scope.
Common questions about feature flags
What are the downsides of using feature flags?
The downsides are more code paths to test, stale flags that pile up as debt, and a flag store that becomes one more thing that can fail. Hodgson’s article says toggles “introduce a significant testing burden” and come with “a carrying cost”. The table version keeps the third risk small, because the store is a database you already run.
Can you give me some examples of feature flags?
Yes. Five common ones, each named by what the user sees when it is off: AI generation paused with old results still readable, a support chat widget gone while the page carries on, upgrades on hold with current plans unchanged, uploads swapped for a short notice, and new signups sent to a waitlist.
What are some alternatives to feature flags?
The alternatives are a host’s instant rollback, a revert and redeploy, and a preview deployment checked before release. Each covers some failures, and none of them can switch off a third-party dependency that breaks after the release is live, because the old build calls the same service.
What happens if I turn off all feature flags?
In an app built this way, every optional feature pauses and the core keeps working: people can still sign in, do the main task and pay. If turning everything off breaks sign-in or billing, a switch sits on something that should never have had one, which makes this a useful test of where the switches were placed.
What are feature flags on my iPhone?
Those are a different thing: Safari’s Feature Flags settings, which Apple describes as a way to “test new web platform features before they ship in Safari”. They have nothing to do with the flags in your own app.
If you have a working app built with these tools and need it ready for real customers, this is what we do.
Built it with AI. Now it has to hold up for real customers.
The Production Hardening Sprint takes the app you already have and builds the production foundation underneath it. Authentication and access rules, payments that stay consistent, error handling, monitoring, backups, automated tests and a documented handover. Our engineers work inside your existing codebase for ten working days. All 123 deliverables are included, and you get the evidence for each one.
See the Production Hardening Sprint →
$2,500 fixed price · 10 working days · One codebase