Why Consolidating Feature Flags and Experiments Cuts Hidden Costs
Feature flags and experiments do two different jobs.
A flag controls delivery. It decides whether a feature is on for a given user, and lets you switch it off fast when something breaks.
An experiment is about measurement. Split users into groups, show them different variations, and find out which performs better on a metric that matters.
Take a new checkout button. The flag delivers it: on for these users, off for those. The experiment measures it: does the new button convert better than the old one?
Both halves describe a single idea - try a different checkout button and see if it helps - but because the jobs are different, teams run them in different places.
Delivery lives in a flagging tool or a config service. Measurement lives somewhere else: a dedicated A/B platform, or a home-grown mix of an SDK and an analytics pipeline.
Describe that one idea in two systems, and keeping them in agreement about which users are in which group becomes your job, not the tooling's.
That's the seam - the join between delivery and measurement - and it carries a cost in money and engineering time that's easy to overlook.
Now that AWS AppConfig runs A/B tests natively on the same feature flags you already deploy, the seam is worth a second look.
This post is about the hidden tax of keeping delivery and measurement apart, and what you get back when they share one control plane.
The multi-service tax
You pay in four ways.
-
Integration. Two systems to wire into your application, each with its own setup, authentication, and configuration. Two things to provision and keep correct across every environment.
-
SDK sprawl. A flag-evaluation SDK here, an experiment or analytics SDK there, multiplied across every language and every service you run. More dependencies to ship, more versions to keep current, more surface area when something breaks.
-
Operational overhead. Two control planes. Two sets of dashboards. Two permission models to manage and audit. Two things every engineer has to learn before they can safely change what customers see. And the gap between "what's flagged" and "what's being measured" becomes its own category of confusion.
-
Assignment drift. This is the subtle one. When a flagging system and an experiment system bucket users independently, keeping a single user in a consistent variation across both gets fiddly. A small mistake here quietly pollutes your experiment data. It's the worst kind of bug, because nothing errors - you just end up deciding on bad numbers.
None of these is dramatic on its own. Together they're a standing tax on every change worth measuring.
What changed
In mid-2026, AWS announced the general availability of experimentation tools in AWS AppConfig. A/B and multivariate testing are now built into the feature flags you already deploy, with no separate experimentation infrastructure to build or run.
You define variations, target an audience with a rule builder, set the traffic split, ramp exposure, and promote the winner - all anchored to the flag itself - through the console, CLI, API, or CDK.
It's available in all commercial AWS Regions.
The timing matters for a second reason. Amazon CloudWatch Evidently, AWS's dedicated A/B testing service, reached end of support on October 17, 2025. So "where does A/B testing live now?" has been a live question for anyone who relied on Evidently, not a hypothetical one.
The practical effect is simple: the flag that delivers a feature is now the same flag that runs the experiment measuring it.
What you get back
Walk back through the same four costs.
One control plane
The flag is the experiment surface. An experiment selects an existing feature flag, and each treatment is a variation of that flag's configuration. The thing that controls the rollout is the thing that runs the test, and defining a variation and its traffic split is just part of the experiment definition.
Here it is in the CDK, where each treatment carries the flag values it sets and the weight that splits traffic:
from aws_cdk import aws_appconfig as appconfig
AttrValue = appconfig.CfnExperimentDefinition.AttributeValueProperty
Treatment = appconfig.CfnExperimentDefinition.TreatmentProperty
# app, env, and profile are the L1 CfnApplication, CfnEnvironment, and
# CfnConfigurationProfile for your feature flags, defined elsewhere in the stack.
# The checkout_button flag must already define a "color" attribute. Treatments
# set the value of an attribute the flag defines; they don't create new ones.
appconfig.CfnExperimentDefinition(
self,
"CheckoutButtonExperiment",
name="checkout-button-color",
application_identifier=app.ref,
environment_identifier=env.ref,
configuration_profile_identifier=profile.ref,
flag_key="checkout_button",
# Editor-syntax expression. $country is a caller-supplied context attribute,
# not an AWS Region; here, everyone in the US is eligible.
audience_rule='(eq $country "
Comments
No comments yet. Start the discussion.