an agent skill · A/B tests, funnels, retention, pricing — sized to your real traffic

/growth

growth improves the numbers a product lives on — signups, activation, retention, referrals, revenue. Ask it to design an A/B test and it first tells you how much traffic that test honestly needs; ask why conversions dropped and it walks the funnel and the cohorts; ask what a winning test means and it tells you how much of the lift is real before you ship. Pricing tests, referral loops, onboarding, product-led growth — 19 references and four runnable calculators behind one router. What makes it different: most A/B-testing advice silently assumes hundreds of thousands of visitors. This pack is built for products with real but limited traffic — it tells you what you can actually learn at your size, and what to do instead when the honest answer is "not this test."

# natural language — no flags, no fixed pipeline /growth is a 3% lift on our signup form worth testing, given we get 400 signups a month

runs onClaude CodeCodexCursorAntigravityopencodeGrok BuildHermes

The router is the skill. There is no fixed test-to-ship pipeline to run start-to-finish — each job stands alone and enters where your request is. The animation traces one path (the flagship gate, deciding what a "significant" result can honestly support); the sections below map the whole surface it routes across.

Define and evaluate the experiment

Improve acquisition-to-revenue conversion by deciding what's worth testing, whether it can be answered at your scale, and what a result actually means once it exists. growth owns the feasibility gate, the design, and the interpretation — not the measurement machinery underneath it, and not the demand that fills the top of the funnel. The scope is funnel-wide: diagnosis, prioritization, feasibility, readout, activation, conversion optimization, retention, referral loops, pricing experiments, product-led growth, and quasi-experimental methods for when randomization isn't available. growth writes experiment designs, feasibility verdicts, readouts and requirements. It writes no production code, and it does not compose the variant it tests.

growth owns

everything between "should we test this" and "what did the result actually prove"

  • what a test can honestly prove ⭐ — the feasibility gate matched to your real traffic, and the Ambition Tax a significant result on it pays
  • the diagnosis and the design — funnel and cohort diagnosis, opportunity prioritization, experiment design, and the quasi-experimental substitute when randomization isn't available
  • the interpretation — peek-safe readout, the winner's-curse haircut, and the decision rule committed to before the result exists
  • activation, retention and referral experimentation — does a resurrection nudge work, does a loop actually compound, on top of a pipeline `data` has certified

hands off to

the "Not this skill" table — ten asks this skill declines by design

  • data — measurement validity: SRM, CUPED's mechanics, assignment integrity, peek-safe method selection. Growth cites it by name and never re-teaches it
  • marketing — demand, channels, positioning; marketing proposes the hypothesis, growth tests it and never reports the readout as marketing's finding
  • product — pricing tiers, the roadmap, which metric a strategy optimizes for; growth moves the metric, never sets it
  • operate — rollout ramps, feature-flag lifecycle and debt, a standing production threshold — proceeds on the null, where an experiment proceeds on a rejected one
  • design · frontend · backend · ai — page composition, in-product voice, and implementing the winning variant. Growth ships a spec and a readout, never code
  • sales · success — pipeline and cold outbound; onboarding education and retention communications a growth test showed were worth sending

The default reader runs a self-serve product with real but limited traffic. Enough to test something, rarely enough to power a standard conversion-rate experiment the way a 200,000-visit playbook assumes. What remains at that scale is real: bolder tests, upstream metrics, and honest acceptance of a higher false-positive rate — never "you can't test."

The faceted router

SKILL.md is a router, not a script. Every request selects the smallest sufficient route: one primary job — the twelve below — combined with at most one base surface that reshapes how the job applies to how the product is actually bought, plus additive overlays when traffic is too small or a model runs the experiment. Read the selected references completely; load one or two, never the whole pack.

facetoptionsrule
① Primary job experiment-design-and-feasibility ⭐ · growth-model-and-loops · funnel-and-cohort-diagnosis · opportunity-and-prioritization · experiment-readout-and-learning · activation-and-onboarding · conversion-optimization · retention-and-resurrection · referral-and-product-loops · monetization-and-pricing-experiments · product-led-growth · quasi-experiments Exactly one. Pick the single job the request needs. The flagship answers the question the rest of the field avoids — can this be answered at your scale, and what would a significant result actually be worth. Often the honest answer changes the test, not just the result.
② Base surface surface-selfserve (default) · surface-b2b-sales-assisted · surface-mobile-subscription · surface-marketplace-network At most one. The surface reshapes what each job means for how the product is bought and used — a marketplace's interference risk, a mobile app's IAP economics. Many questions (the feasibility gate's own arithmetic) are surface-independent.
③ Additive overlays overlay-small-sample — traffic or users too small for standard power
overlay-agentic — a model designs, runs, or reads an experiment
Additive. Each stacks on top of a base surface, never instead of one. Read growth-model-and-loops as a fourth, orthogonal move whenever a claim invokes a named growth-model figure (Balfour, Winters, Kwok, Verna, Eyal, McClure) — it grades and attributes them so the other references need not re-derive it.
Every claim on this page is true of the shipped pack. Where a number is cited — the power-table constant, the winner's-curse haircut, the CUPED variance-reduction figures — it carries its named source; calculators are executable and unit-validated against published anchors at import time, never a prose formula this page would then repeat after it drifted.

The twelve primary jobs

Each job is one reference, read fully only when its route is selected. Feasibility is the flagship because every other growth reference that proposes running a test assumes this file's gate has already been cleared. This is the whole funnel-wide surface, not a headline slice.

I need to…ReadContribution
Decide whether a test can be answered at your scale, and what a significant result would actually mean experiment-design-and-feasibility.md Flagship. The feasibility gate; the derived power table against vendor floors; the metric-skew rule; the Bayes posterior on a "significant" winner; the Ambition Tax; where CUPED fails
Understand growth as a system before touching a single funnel stage — loops, funnels, and the family's AARRR split growth-model-and-loops.md Loops vs funnels (Balfour/Winters/Kwok/Chen, 2018) and the Racecar frame correctly attributed to Hockenmaier + Rachitsky, not Balfour; Verna's Five Laws and 3×3 matrix, dated; the K-factor/cycle-time/compounding math the ecosystem never built; reciprocates marketing's AARRR split
Diagnose where a funnel is actually leaking, or which cohort definition to trust funnel-and-cohort-diagnosis.md The 3-way retention definition — N-day, rolling, survival — and which question each one actually answers
Decide what to work on next across a backlog of growth ideas opportunity-and-prioritization.md ICE read skeptically as throughput tooling, not evidence; Verna's combinatorial counter to single-lever prioritization
Read out a finished experiment, or decide whether a "significant" result is safe to act on experiment-readout-and-learning.md The one-curve peeking reconciliation; the winner's-curse haircut applied at readout; guardrails-before-shipping-a-win; the learning ledger; Twyman's law as a Bayesian prior
Improve activation or onboarding, or evaluate a claimed "aha moment" activation-and-onboarding.md Why CUPED fails for new users — load-bearing for every onboarding test's power math; "aha moment" labeled a folklore term with no traceable origin
Run or brief a conversion-rate test on a page, form, or button conversion-optimization.md The red-button-on-blue-page external-validity lesson; the copy→placement→color test-order rule; form-field-count as a variable, not a rule; the Trustworthy A/B Patterns project
Reduce churn, or design a resurrection or win-back approach retention-and-resurrection.md The Duolingo streak specimen as a fully worked, resolvable example; Sarah Tavel's retention framing; a habit-ethics pointer into monetization's ethics section
Design a referral program, or diagnose a growth loop's own math referral-and-product-loops.md The loop math the ecosystem lacks entirely, applied to a live referral design; "loops" fully disambiguated from loops-over-moats
Test a price, a plan structure, or a packaging change monetization-and-pricing-experiments.md Booking's own pricing-test refusal; Van Westendorp's provenance and its stated-preference critique; the subscription dark-patterns table on Mathur CSCW 2019 plus FTC 2022 enforcement
Evaluate or benchmark a product-led-growth motion product-led-growth.md Five-field benchmark-provenance discipline (source, sample, method, date, caveat); the Sean Ellis 40% test with the creator's own generalizability caveat
Answer a causal question when randomization isn't available quasi-experiments.md Precondition checklists in place of numeric floors; Abadie's own warning that a large pre-period can't fix a bad fit; the staggered-DiD structural-invalidity warning

Full router table & invariants: SKILL.md.

Four surfaces + two additive overlays

One base surface, at most, reshapes every job for how the product is actually bought and used — the same conversion-test job resolves differently for a marketplace with spillover between arms than for a mobile app's IAP economics. The two overlays are additive — they stack on top of whichever base you picked, never replace it — and carry a distinct violet identity throughout this page, the same convention automation, operate, quality, data and marketing use for their own additive overlays.

Self-serve SaaS surface-selfserve.md

The default: anyone can sign up or trial without talking to a human. If a request names no business model, this is what the skill assumes — and it says so rather than stalling for a clarification it doesn't need.

reshapesthe funnel as the whole product experience · what "activation" means with no human in the loop

B2B sales-assisted surface-b2b-sales-assisted.md

A rep, a demo, or procurement gates the deal — a human conversation shapes the funnel. This surface names the seam growth draws with sales: pipeline and qualification signal in, an experiment readout out.

reshapeswhich stage of a sales-assisted funnel is even eligible for a randomized test

Mobile subscription surface-mobile-subscription.md

In-app-purchase economics reshape the test — store fees, platform review cycles, and refund/chargeback dynamics that don't exist on the web.

reshapespower math against the RevenueCat-layer dataset — high-N, low-external-validity, vendor+edition required

Marketplace / network surface-marketplace-network.md

Supply and demand sit on the same product, so a treatment can leak into control through the marketplace itself — interference, not just noise.

reshapeswhether randomization is even valid — diagnose the interference mechanism before picking a design

Small-sample additive ⭐

Traffic or users too small for standard power. Stacks on, does not replace. The overlay this pack's flagship gate exists to serve — the three levers (move upstream, swing bigger, skip the test) and the skip-the-test decision rule, taught as equipping an existing practice, never as a refusal.

reshapeswhich effect sizes are even detectable · what an honest "this needs a bigger bet" looks like as a deliverable

Agentic additive ⭐

A model designs, runs, or reads an experiment — not a human reviewing one draft a model produced. Stacks on, does not replace. The deterministic envelope around an agent-decided test: guardrails a model cannot silently loosen mid-run.

reshapeswho can peek and when · what's ai's (the model's own cognition) vs ours (the envelope around a test it runs)

Experiment design and feasibility — the flagship

Ask an agent whether a test is worth running and it either skips straight to statistics with nowhere to land, or straight to "run an A/B test" with no check on whether the traffic can support one. The gate comes first: can this specific comparison reach a power you'd trust, before a hypothesis or a flag exists. Skip it and you inherit the ecosystem's most common failure mode — even careful incumbent suites go straight from hypothesis to flag to metrics with no sample-size or MDE step in between. The teaching that matters most: halving the effect you want to detect quadruples the traffic you need — at a 5% baseline conversion rate, a +20% relative MDE needs 7,457 exposures per arm; +10% needs 29,826 (4.0×); +5% needs 119,303 (4.0× again). Metric choice alone can move that requirement two more orders of magnitude — a skewed metric like revenue-per-user can need ~114,000 where a low-skew metric needs ~1,550, on the same underlying change (Kohavi's skewness rule, KDD 2014).

The Ambition Tax — paid twice, and the two effects multiply, not cancel

Two facts that compound in exactly the wrong direction

  • fact onewinning estimates are biased upward: 13% at 80% power with one treatment arm, 21% with two, 25% with three — and 30% if you Bonferroni-correct, because a stricter threshold selects even more extreme draws from the noise (Kohavi, 2024-10-26)
  • fact twoa small sample forces bigger bets: from the power table, detecting a 2% relative lift needs roughly 4–20× the traffic of detecting a 20% relative lift, at every baseline
  • the taxthe two effects don't cancel, they multiply — the smaller your sample, the more ambitious your test must be, and the more ambitious your test, the less a significant result on it actually means

Kohavi's own floor for e-commerce conversion work is roughly 200,000 users for adequate power, stated directly: "the statistics do not support A/B tests with 5,000 users under common goals and assumptions" (2024-10-30). He then names three levers rather than stopping at "you can't test": swing for the fences, move upstream to an earlier metric, or consciously accept a higher false-positive rate — see overlay-small-sample.md for the fuller operating model.

What a "significant" result is actually worth

π=50%, 1/3 — 94.1%, 88.9% posterior. A coin-flip prior, or Microsoft's reported average (2009, no published denominator) π=20%, 10% — 80.0%, 64.0% posterior. Bing's high and low ends (KDD 2014) π=5% — 45.7% posterior. A dramatic, surprising result deserves a low prior — which makes it less trustworthy at an identical p-value, not more impressive
P(true positive | significant) = (1−β)π / [(1−β)π + α(1−π)] — at α=.05, β=.20 (80% power). A "statistically significant winner" is not a fact; it is a posterior, and its trustworthiness is set entirely by a hit rate you knew before you ran the test (Kohavi, Seven Rules of Thumb, KDD 2014, Rule #2).
the same discipline `data`'s SRM check protects at the assignment layertwo experiments, identical p-value, different priors, different truths

CUPED will not rescue you here. Three real Microsoft experiments showed 45%, 52%, and 49% variance reduction with one week of pre-period data — but the same paper's own §5.2.3 reports revenue-per-user reduced by less than 5%, "due to the low correlation." It fails hardest exactly where growth spends most of its effort: new users have a covariate-metric correlation of 0.2–0.4 versus ~40% for existing ones (CUPED's own paper, Eppo, Statsig, and Netflix's KDD 2016 case study all agree), and retention sees small variance reduction for both new and existing users. The small-sample problem cannot be variance-reduced away precisely where activation and retention testing lives — sequential monitoring changes when you find out, not how much data the question needs.

What makes this different

Almost every individual fact in this pack exists somewhere else. What does not exist anywhere else is a pack that puts the growth surface and the validity layer in the same repo, and grades every benchmark it leans on rather than repeating it because it sounds measured — so the wedge is named as convergence and provenance discipline, not discovery.

Two literatures that never shared a repo — until the gate demanded it

The genre teaches statistical rigor with no growth surface attached, or the growth surface — activation, retention, loops, PLG — with no validity layer underneath. Nowhere in that genre does anyone tell an underpowered reader anything but "gather more data" or "redesign" — a refusal, not a redirect.

  • the feasibility gate is asked first — before a hypothesis or a flag is written, not discovered later as a diagnostic step you should already have used
  • a redirect, never a refusal — move upstream, swing for a bigger effect, or document a deliberate skip; the three levers plus the skip-the-test decision rule
  • the growth-vs-operate boundary, decided not discovered — a standing threshold on the live system is operate's; the same metric bound to one tested change's decision is growth's

Every benchmark carries five fields, or it doesn't travel

Growth's canon is a small, tightly self-citing cluster of practitioners — the corpus found the same four or five names co-citing each other repeatedly, convenient for consistency and weak as independent evidence. This pack names who actually said what, and when.

  • misattribution corrected, on the record — the Racecar Growth Framework is Hockenmaier + Rachitsky, not Balfour; a corrected canon attribution, stated so it isn't repeated
  • win rates never blend — only Microsoft's ~1/3 (all four qualifiers, no published denominator) and Bing's 10–20% are first-party citable; every other famous win-rate figure is never-ship
  • a disclaimed figure is still a figure — never name a never-ship number in order to forbid it; state the mechanism and point at the primary source instead
  • calculators are executable, not prose — power_calc.py, srm_check.py, peeking_table.py, skew_check.py each self-test against a published anchor; two of three incumbent treatments checked this run were wrong, in opposite directions
Open at scale, not empty. Competent prior art exists — rampstack, GrowthBook, PostHog, coreyhaines31's incumbent pack — and is cited respectfully throughout. The gap this pack fills is that the growth surface and the validity layer have never shared a repo, and that data's own experiment-measurement-foundations.md explicitly cedes design and interpretation to whoever picks it up. An honest "nobody has measured this, and here is why" is a deliverable, not a failure to produce one.

The universal invariants

These govern every route, whichever references it loads.

What a pass produces

Designs, verdicts and readouts are delivered as documents. Neither is delivered as production code. A growth artifact is incomplete unless it carries the metric and who picked it, the feasibility verdict and what it assumes about scale, every claim with a named proof source and its evidence tier, the OEC and guardrails if an experiment is being designed, the decision rule committed to before the result exists, what the result cannot answer, and the sibling packs it hands to or depends on.

step 1routeone primary job + one base surface; overlays only if they apply; state any surface you assumed
step 2name the decisionwhat's actually being decided, and who owns that decision once the result exists
step 3run the gatebefore designing anything further — can this question be answered at your scale, and what would a significant result be worth
step 4gradeevery benchmark, win rate or "best practice" relied on, and where its evidence stops
step 5designthe experiment (or quasi-experimental substitute), OEC and guardrails named up front, on `data`'s certified pipeline
step 6producethe artifact — design, feasibility verdict, readout, or requirements — with a named proof source behind every claim
step 7commit the rulestate the decision rule before the result exists, and the haircut to apply once it does
step 8hand offa compact handoff when downstream work is expected — paths, owners, residual risk

power_calc.py

assets/power_calc.py

Sample-size and power, unpooled by default (the exact two-sample formula), with a `--pooled` mode. Validated against PostHog's and Booking.com's own published anchors.

srm_check.py

assets/srm_check.py

Sample-ratio-mismatch chi-square test — is the observed assignment split trustworthy before any result is read. Validated against Fabijan et al.'s worked example.

peeking_table.py

assets/peeking_table.py

The sequential-testing inflated-α curve — how many looks (K), at what nominal α, actually costs. Validated against the Armitage sequential-testing table.

skew_check.py

assets/skew_check.py

Kohavi's 355·s² normality floor for skewed metrics, computed from a real sample rather than eyeballed. Validated against Bing's post-erratum skewness table.

Eval suites

evals/routing · evals/stats-cases · evals/never-ship

Does the router reach the smallest sufficient reference set (including declining to sibling packs); does the content actually change model behavior on real feasibility/peeking prompts; does a disclaimed figure ever resurface, even inside a correction of it.

The handoff seams

growth operates independently when invoked alone, and uses compatible upstream artifacts without silently overriding them. It is rarely terminal: handoff.md maps every seam between this pack and the rest of the family from growth's side — what it produces for each sibling, what it consumes, and the line it does not cross.

growth emits

metric: <who picked it>
feasibility_verdict: <can this be
  answered, at what cost>
claims: [claim, named proof source,
  evidence tier]
oec_and_guardrails: <if an
  experiment is designed>
decision_rule: <committed before
  the result exists>
cannot_answer: <what the result
  does not prove>

data

An experiment design naming its assignment population, exposure logic, and OEC/guardrails for data to certify. Growth does not re-teach or re-implement SRM, CUPED mechanics, or peek-safe method selection — data cedes design and interpretation to growth explicitly, and growth cites it by name rather than re-deriving the mechanics.

marketing (lateral)

Experiment readouts and results on a hypothesis marketing proposed. Growth never creates demand, chooses channels, or writes positioning — it evaluates the experiments marketing's hypotheses generate, and never reports a readout as marketing's own finding.

operate

The evidence-class scope of a rollout that is also an experiment, and the readout once it concludes. The crisp rule: a standing threshold on the live system is operate's; the same metric bound to one tested change's decision is growth's — a ramp proceeds on the null, an experiment proceeds on a rejected null.

The AARRR note. A practitioner who knows Acquisition → Activation → Retention → Referral → Revenue will notice this pack starts at Activation, not Acquisition — a deliberate seam, not a gap. AARRR spans three packs: marketing owns Acquisition — channels, positioning, demand; growth owns the experiments that improve conversion across the whole funnel — Activation, Referral loop design and math, and Retention experimentation; success owns Retention execution — the onboarding education and retention communications a growth test showed were worth sending. See growth-model-and-loops.md for the full reciprocation of marketing's own shipped note.

The rest of the family — each an independently installable pack with its own guide:

/product /design /architecture /frontend /backend /ai /data /automation /quality /operate /marketing sales · success — seams declared, packs not yet shipped

Start here

Install once. It's a plain SKILL.md router — no flags, no config, no fixed pipeline — so it activates on natural-language phrasing ("is a 3% lift worth testing at our traffic," "we peeked at this four times, is it safe to ship," "design a referral program and check the loop math") rather than a fixed command.

# skills.sh ecosystem npx skills add gabros20/growth-skill -g -y # clone + manual copy git clone https://github.com/gabros20/growth-skill cp -R growth-skill/skills/growth ~/.claude/skills/growth # use — natural language, any host /growth is a 3% lift on our signup form worth testing, given we get 400 signups a month /growth diagnose why activation drops between signup and first value /growth we can't hit standard sample-size floors — what's actually worth testing at our scale

The same install runs on any Agent Skills host. Codex triggers with $growth; manual copy into any client's skills directory also works.

what's in the repo
skills/growth/ the skill: SKILL.md (router) + 19 references/ + 4 self-testing assets/ + evals/ research/ multi-channel research corpora + build-gate synthesis site/ this guide — deploys to growthskill.vercel.app README.md · SOURCES.md · LICENSE

More detail: SKILL.md · SOURCES.md — source attribution, licensing rule, and the numbers this pack refuses to ship.