Run AI experiments
AI Experimentation Framework
The AI Experimentation Framework tests an AI intervention in eight steps: problem, hypothesis, AI intervention, baseline, experiment, metric, result and decision, so that teams measure the effect of AI instead of assuming it.
The model
- 1Problem
- 2Hypothesis
- 3AI intervention
- 4Baseline
- 5Experiment
- 6Metric
- 7Result
- 8Decision
Developed by Markaigen · Updated
What each part means
- Problem
- The specific pain, in numbers: how long, how much, how often.
- Hypothesis
- A statement that can be wrong, for example: AI-assisted briefs cut production time without reducing organic performance.
- AI intervention
- Exactly what changes: which step, which tool, which prompt, who checks.
- Baseline
- The current numbers over a representative period, measured before anything changes.
- Experiment
- How you run it: duration, sample, control group or before-and-after, who is involved.
- Metric
- The primary metric that decides, and the guardrail metrics that must not get worse.
- Result
- The numbers against the baseline, with the uncertainty stated honestly.
- Decision
- Scale, adjust or stop, written down with the reason and the date.
How to use it
- Write the decision rule before the experiment starts: what result leads to scaling, and what stops it.
- Always include a guardrail metric for quality, so time savings cannot hide a drop in results.
- Run experiments long enough to cover normal variation, usually at least two to four weeks.
Write an experiment card
Fill in the steps. The card is the one-page agreement the team signs off before the experiment starts.
Answers are saved in this browser so you can return to them.
Download or e-mail this framework
Get a branded PDF or Word document, with your answers or as an empty template to fill in.
