Experimentation · Case study
Spotify: close every experiment
Carry each experiment from a checked launch to a recorded ship-or-stop decision and a cleaned-up flag.
By OneShot · Updated
Read the solution ↓What the workflow produces
Experiment decision record- Setup
- Hypothesis, metrics, audience and planned duration as launched.
- Guardrails
- Alerts raised, readings and the owner’s response.
- Decision
- Ship, iterate or stop, with the owner’s reasoning.
- Cleanup
- Flag and losing variant removed, or scheduled with an owner.
The documented problem
Spotify’s engineering blog describes growing from fewer than 20 priority experiments a year to thousands per year by 2020. As volume grew, teams were slowed by restarted experiments, manual statistical work in notebooks and test groups coordinated in spreadsheets, and experiments increasingly collided with one another.
Spotify built its Confidence platform in response, and says more than 300 teams now use it to run more than 10,000 experiments a year. These figures are Spotify’s own. The OneShot workflow below is a proposed design for the work around each experiment; it does not replace an experimentation platform or its statistics.
- Spotify Engineering: Confidence, an experimentation platform from SpotifyPrimary source · Checked 2026-10-08
- Spotify Confidence product pagePrimary source · Checked 2026-10-08
The OneShot solution
Experiments that end in a recorded decision
Keep every experiment checked before launch, watched while live and closed with its owner’s recorded decision.
- 01
Check the launch
Compare the experiment’s configuration with the team’s launch checklist: hypothesis, primary metric, guardrail metrics, audience and planned duration. Return incomplete requests to the owner and flag overlap with experiments already running on the same audience.
- Email send
- Email Inbox
- Experiment platform and metrics
- 02
Watch the guardrails
The OneShot agent reads the platform’s results on the agreed schedule. When a guardrail metric crosses its threshold, it sends the owner the reading; pausing or rolling back stays the owner’s call unless the team has approved it in advance.
- Email send
- Email Inbox
- Experiment platform and metrics
- 03
Prepare the decision
At the planned end, assemble the platform’s results, the guardrail history and any open data-quality questions into one packet. Link to the platform’s statistics; do not restate or recompute them.
- 04
Record and clean up
Store the owner’s decision to ship, iterate or stop, with the reasoning. Chase removal of the feature flag and the losing code path, and keep the experiment open until that is confirmed.
- Email send
- Email Inbox
- Experiment platform and metrics
- Source repositories and CI
What counts as complete
The owner’s decision is recorded with its evidence, and the flag and losing variant are confirmed removed or scheduled with a named owner.
Decisions and exceptions
- The agent does not judge statistical significance; the platform and the owner do.
- Rollbacks happen without the owner only where the team has approved that in advance.
- Conflicting experiments on one audience go to both owners before either launches.
What to measure
- Experiments past their end date
- Time from end to decision
- Launch requests returned incomplete
- Flags left in code
Establish the baseline and review period before launch. Compare completed cases, unresolved work and reviewer corrections against the same scope.
The agent’s tool basket
OneShot gives its agents the tools to analyze records, communicate, research and act on authorized business data. The platform carries the workflow from the first action through to the recorded result.
Experiment platform and metrics
Read experiment configuration, results and guardrail metrics; record the owner's decision.
Source repositories and CI
Read in-scope repositories, open pull requests and read build and test results for each change.
Workflow pricing
Scope the full workflow around your volume, business systems and required outcome.
Discuss workflow pricing →