growth improves the numbers a product lives on — signups, activation, retention, referrals, revenue. Ask it to design an A/B test and it first tells you how much traffic that test honestly needs; ask why conversions dropped and it walks the funnel and the cohorts; ask what a winning test means and it tells you how much of the lift is real before you ship. Pricing tests, referral loops, onboarding, product-led growth — 19 references and four runnable calculators behind one router. What makes it different: most A/B-testing advice silently assumes hundreds of thousands of visitors. This pack is built for products with real but limited traffic — it tells you what you can actually learn at your size, and what to do instead when the honest answer is "not this test."
runs onClaude CodeCodexCursorAntigravityopencodeGrok BuildHermes
experiment-design-and-feasibility.mdsurface-selfserve.md
2 of 19 loaded · read fully
Route before acting. One job, at most one base surface, overlays added only when they apply. If the request names no business model, the skill does not stall — it assumes self-serve SaaS and says which it assumed.
The router is the skill. There is no fixed test-to-ship pipeline to run start-to-finish — each job stands alone and enters where your request is. The animation traces one path (the flagship gate, deciding what a "significant" result can honestly support); the sections below map the whole surface it routes across.
Improve acquisition-to-revenue conversion by deciding what's worth testing, whether it can be
answered at your scale, and what a result actually means once it exists. growth owns
the feasibility gate, the design, and the interpretation — not the measurement machinery
underneath it, and not the demand that fills the top of the funnel. The scope is funnel-wide:
diagnosis, prioritization, feasibility, readout, activation, conversion optimization, retention,
referral loops, pricing experiments, product-led growth, and quasi-experimental methods for when
randomization isn't available. growth writes experiment designs, feasibility
verdicts, readouts and requirements. It writes no production code, and it does not compose
the variant it tests.
everything between "should we test this" and "what did the result actually prove"
the "Not this skill" table — ten asks this skill declines by design
The default reader runs a self-serve product with real but limited traffic. Enough to test something, rarely enough to power a standard conversion-rate experiment the way a 200,000-visit playbook assumes. What remains at that scale is real: bolder tests, upstream metrics, and honest acceptance of a higher false-positive rate — never "you can't test."
SKILL.md is a router, not a script. Every request selects the smallest sufficient route: one primary job — the twelve below — combined with at most one base surface that reshapes how the job applies to how the product is actually bought, plus additive overlays when traffic is too small or a model runs the experiment. Read the selected references completely; load one or two, never the whole pack.
| facet | options | rule |
|---|---|---|
| ① Primary job | experiment-design-and-feasibility ⭐ · growth-model-and-loops · funnel-and-cohort-diagnosis · opportunity-and-prioritization · experiment-readout-and-learning · activation-and-onboarding · conversion-optimization · retention-and-resurrection · referral-and-product-loops · monetization-and-pricing-experiments · product-led-growth · quasi-experiments | Exactly one. Pick the single job the request needs. The flagship answers the question the rest of the field avoids — can this be answered at your scale, and what would a significant result actually be worth. Often the honest answer changes the test, not just the result. |
| ② Base surface | surface-selfserve (default) · surface-b2b-sales-assisted · surface-mobile-subscription · surface-marketplace-network | At most one. The surface reshapes what each job means for how the product is bought and used — a marketplace's interference risk, a mobile app's IAP economics. Many questions (the feasibility gate's own arithmetic) are surface-independent. |
| ③ Additive overlays ⭐ | overlay-small-sample — traffic or users too small for standard power overlay-agentic — a model designs, runs, or reads an experiment |
Additive. Each stacks on top of a base surface, never instead of one. Read growth-model-and-loops as a fourth, orthogonal move whenever a claim invokes a named growth-model figure (Balfour, Winters, Kwok, Verna, Eyal, McClure) — it grades and attributes them so the other references need not re-derive it. |
Each job is one reference, read fully only when its route is selected. Feasibility is the flagship because every other growth reference that proposes running a test assumes this file's gate has already been cleared. This is the whole funnel-wide surface, not a headline slice.
| I need to… | Read | Contribution |
|---|---|---|
| Decide whether a test can be answered at your scale, and what a significant result would actually mean ⭐ | experiment-design-and-feasibility.md |
Flagship. The feasibility gate; the derived power table against vendor floors; the metric-skew rule; the Bayes posterior on a "significant" winner; the Ambition Tax; where CUPED fails |
| Understand growth as a system before touching a single funnel stage — loops, funnels, and the family's AARRR split | growth-model-and-loops.md |
Loops vs funnels (Balfour/Winters/Kwok/Chen, 2018) and the Racecar frame correctly attributed to Hockenmaier + Rachitsky, not Balfour; Verna's Five Laws and 3×3 matrix, dated; the K-factor/cycle-time/compounding math the ecosystem never built; reciprocates marketing's AARRR split |
| Diagnose where a funnel is actually leaking, or which cohort definition to trust | funnel-and-cohort-diagnosis.md |
The 3-way retention definition — N-day, rolling, survival — and which question each one actually answers |
| Decide what to work on next across a backlog of growth ideas | opportunity-and-prioritization.md |
ICE read skeptically as throughput tooling, not evidence; Verna's combinatorial counter to single-lever prioritization |
| Read out a finished experiment, or decide whether a "significant" result is safe to act on | experiment-readout-and-learning.md |
The one-curve peeking reconciliation; the winner's-curse haircut applied at readout; guardrails-before-shipping-a-win; the learning ledger; Twyman's law as a Bayesian prior |
| Improve activation or onboarding, or evaluate a claimed "aha moment" | activation-and-onboarding.md |
Why CUPED fails for new users — load-bearing for every onboarding test's power math; "aha moment" labeled a folklore term with no traceable origin |
| Run or brief a conversion-rate test on a page, form, or button | conversion-optimization.md |
The red-button-on-blue-page external-validity lesson; the copy→placement→color test-order rule; form-field-count as a variable, not a rule; the Trustworthy A/B Patterns project |
| Reduce churn, or design a resurrection or win-back approach | retention-and-resurrection.md |
The Duolingo streak specimen as a fully worked, resolvable example; Sarah Tavel's retention framing; a habit-ethics pointer into monetization's ethics section |
| Design a referral program, or diagnose a growth loop's own math | referral-and-product-loops.md |
The loop math the ecosystem lacks entirely, applied to a live referral design; "loops" fully disambiguated from loops-over-moats |
| Test a price, a plan structure, or a packaging change | monetization-and-pricing-experiments.md |
Booking's own pricing-test refusal; Van Westendorp's provenance and its stated-preference critique; the subscription dark-patterns table on Mathur CSCW 2019 plus FTC 2022 enforcement |
| Evaluate or benchmark a product-led-growth motion | product-led-growth.md |
Five-field benchmark-provenance discipline (source, sample, method, date, caveat); the Sean Ellis 40% test with the creator's own generalizability caveat |
| Answer a causal question when randomization isn't available | quasi-experiments.md |
Precondition checklists in place of numeric floors; Abadie's own warning that a large pre-period can't fix a bad fit; the staggered-DiD structural-invalidity warning |
Full router table & invariants: SKILL.md.
One base surface, at most, reshapes every job for how the product is actually bought and
used — the same conversion-test job resolves differently for a marketplace with spillover between
arms than for a mobile app's IAP economics. The two overlays are additive — they stack on
top of whichever base you picked, never replace it — and carry a distinct violet identity
throughout this page, the same convention automation, operate,
quality, data and marketing use for their own additive
overlays.
Ask an agent whether a test is worth running and it either skips straight to statistics with nowhere to land, or straight to "run an A/B test" with no check on whether the traffic can support one. The gate comes first: can this specific comparison reach a power you'd trust, before a hypothesis or a flag exists. Skip it and you inherit the ecosystem's most common failure mode — even careful incumbent suites go straight from hypothesis to flag to metrics with no sample-size or MDE step in between. The teaching that matters most: halving the effect you want to detect quadruples the traffic you need — at a 5% baseline conversion rate, a +20% relative MDE needs 7,457 exposures per arm; +10% needs 29,826 (4.0×); +5% needs 119,303 (4.0× again). Metric choice alone can move that requirement two more orders of magnitude — a skewed metric like revenue-per-user can need ~114,000 where a low-skew metric needs ~1,550, on the same underlying change (Kohavi's skewness rule, KDD 2014).
Kohavi's own floor for e-commerce
conversion work is roughly 200,000 users for adequate power, stated directly: "the
statistics do not support A/B tests with 5,000 users under common goals and assumptions"
(2024-10-30). He then names three levers rather than stopping at "you can't test": swing for
the fences, move upstream to an earlier metric, or consciously accept a higher false-positive
rate — see overlay-small-sample.md for the fuller operating model.
CUPED will not rescue you here. Three real Microsoft experiments showed 45%, 52%, and 49% variance reduction with one week of pre-period data — but the same paper's own §5.2.3 reports revenue-per-user reduced by less than 5%, "due to the low correlation." It fails hardest exactly where growth spends most of its effort: new users have a covariate-metric correlation of 0.2–0.4 versus ~40% for existing ones (CUPED's own paper, Eppo, Statsig, and Netflix's KDD 2016 case study all agree), and retention sees small variance reduction for both new and existing users. The small-sample problem cannot be variance-reduced away precisely where activation and retention testing lives — sequential monitoring changes when you find out, not how much data the question needs.
Almost every individual fact in this pack exists somewhere else. What does not exist anywhere else is a pack that puts the growth surface and the validity layer in the same repo, and grades every benchmark it leans on rather than repeating it because it sounds measured — so the wedge is named as convergence and provenance discipline, not discovery.
The genre teaches statistical rigor with no growth surface attached, or the growth surface — activation, retention, loops, PLG — with no validity layer underneath. Nowhere in that genre does anyone tell an underpowered reader anything but "gather more data" or "redesign" — a refusal, not a redirect.
Growth's canon is a small, tightly self-citing cluster of practitioners — the corpus found the same four or five names co-citing each other repeatedly, convenient for consistency and weak as independent evidence. This pack names who actually said what, and when.
data's own experiment-measurement-foundations.md explicitly cedes
design and interpretation to whoever picks it up. An honest "nobody has measured this, and here
is why" is a deliverable, not a failure to produce one.These govern every route, whichever references it loads.
Designs, verdicts and readouts are delivered as documents. Neither is delivered as production code. A growth artifact is incomplete unless it carries the metric and who picked it, the feasibility verdict and what it assumes about scale, every claim with a named proof source and its evidence tier, the OEC and guardrails if an experiment is being designed, the decision rule committed to before the result exists, what the result cannot answer, and the sibling packs it hands to or depends on.
Sample-size and power, unpooled by default (the exact two-sample formula), with a `--pooled` mode. Validated against PostHog's and Booking.com's own published anchors.
Sample-ratio-mismatch chi-square test — is the observed assignment split trustworthy before any result is read. Validated against Fabijan et al.'s worked example.
The sequential-testing inflated-α curve — how many looks (K), at what nominal α, actually costs. Validated against the Armitage sequential-testing table.
Kohavi's 355·s² normality floor for skewed metrics, computed from a real sample rather than eyeballed. Validated against Bing's post-erratum skewness table.
Does the router reach the smallest sufficient reference set (including declining to sibling packs); does the content actually change model behavior on real feasibility/peeking prompts; does a disclaimed figure ever resurface, even inside a correction of it.
growth operates independently when invoked alone, and uses compatible upstream artifacts
without silently overriding them. It is rarely terminal: handoff.md maps every seam
between this pack and the rest of the family from growth's side — what it produces for each
sibling, what it consumes, and the line it does not cross.
metric: <who picked it> feasibility_verdict: <can this be answered, at what cost> claims: [claim, named proof source, evidence tier] oec_and_guardrails: <if an experiment is designed> decision_rule: <committed before the result exists> cannot_answer: <what the result does not prove>
An experiment design naming its assignment population, exposure logic, and OEC/guardrails
for data to certify. Growth does not re-teach or re-implement SRM, CUPED
mechanics, or peek-safe method selection — data cedes design and interpretation
to growth explicitly, and growth cites it by name rather than re-deriving the mechanics.
Experiment readouts and results on a hypothesis marketing proposed. Growth never creates demand, chooses channels, or writes positioning — it evaluates the experiments marketing's hypotheses generate, and never reports a readout as marketing's own finding.
The evidence-class scope of a rollout that is also an experiment, and the readout once it
concludes. The crisp rule: a standing threshold on the live system is operate's;
the same metric bound to one tested change's decision is growth's — a ramp
proceeds on the null, an experiment proceeds on a rejected null.
The AARRR note. A practitioner who knows Acquisition → Activation → Retention → Referral → Revenue will notice this pack starts at Activation, not Acquisition — a deliberate seam, not a gap. AARRR spans three packs: marketing owns Acquisition — channels, positioning, demand; growth owns the experiments that improve conversion across the whole funnel — Activation, Referral loop design and math, and Retention experimentation; success owns Retention execution — the onboarding education and retention communications a growth test showed were worth sending. See growth-model-and-loops.md for the full reciprocation of marketing's own shipped note.
The rest of the family — each an independently installable pack with its own guide:
Install once. It's a plain SKILL.md router — no flags, no config, no fixed
pipeline — so it activates on natural-language phrasing ("is a 3% lift worth testing at our
traffic," "we peeked at this four times, is it safe to ship," "design a referral program and check
the loop math") rather than a fixed command.
The same install runs on any Agent Skills
host. Codex triggers with $growth; manual copy into any client's skills directory
also works.
More detail: SKILL.md · SOURCES.md — source attribution, licensing rule, and the numbers this pack refuses to ship.