Fictura
Experiments

Metrics & decisions

A Bayesian engine that recomputes from the raw event ledger on every read — and one metric class it will never let auto-decide.

Metrics

Every experiment has a primary metric — trial_started by default, or any event name you type. That's not a separate thing you configure elsewhere: it's the exact string you already pass as the first argument to logEvent() in your app. Named a milestone generation_success? Type generation_success as the primary metric and the engine starts counting it immediately — no extra wiring, no separate event you have to declare for experiments specifically.

Metrics split into two kinds, and the split matters for what happens next:

KindExamplesCan decide a winner
Ratetrial_started, purchase_completed, paywall_cta_tappedYes
Descriptive-onlyrevenue_per_user, trial_conversion_rate, active_subscribers, churned_subscribersNo — see below

The decision engine

Bayesian, Beta-Binomial, recomputed from the event ledger on every read rather than stored — so changing which metric you're watching never needs a backfill. 10,000 Monte Carlo draws per evaluation give each arm a prob_best, a 95% credible interval, and a lift-vs-control. Evaluation runs every 15 minutes; there's no penalty for checking early, because the machinery is built to be looked at continuously rather than waited on until a fixed sample size.

A winner requires all of: every arm has at least 100 exposures, and the leading arm's prob_best clears your confidence threshold.

Revenue and descriptive-only metrics never auto-decide — by design

If your primary metric is revenue_per_user, or any metric flagged descriptive-only, the experiment can never become decisive, no matter how confident the numbers get. This isn't a bug to work around — heavy-tailed money through a win/loss posterior manufactures false confidence, and a metric like churn would happily declare the worst arm the winner if left to decide on its own.

You still get everything else: prob_best, credible intervals, lift, the full billing-outcomes table. What you don't get is auto-ship, and in notify mode you don't get an alert either — the leader-found notification is gated on the same decisiveness check. Concluding a revenue-metric experiment is a manual action: read the numbers yourself, then promote a variant or stop the test.

Alerts

Two notifications, both fanned out to whichever of Slack, email, and outgoing webhook you've connected:

  • experiment_winner_shipped — auto mode, decisive, and it actually promoted.
  • experiment_leader_found — notify mode, decisive, and this is the first time it crossed the bar. It fires exactly once per experiment; if the leader later flips while the test keeps running, there's no second alert. Re-check the dashboard rather than waiting on a notification for anything past the first signal.

Promotion

Promoting a variant sets its config live for that placement and demotes whatever was live before — the losing arms need no rollback, since not promoting them is enough. A winning holdout promotes nothing, because the point of a holdout is that the baseline already won.

On this page