A/B testing usually means a platform, a script, and a dashboard. But the most common marketing experiments — which landing page, which link copy, which send time — only need two links and honest measurement. This guide covers the experimental discipline: what to test, how to split, how big the sample needs to be, and how to read the numbers.
The three tests that need only links
1. Destination tests. Two landing pages, two links, same channel slot:
https://yas.sh/signup-a → landing page A (short form)
https://yas.sh/signup-b → landing page B (long form)
Measure clicks and destination-side conversions (conversion tracking); the click-through tells you which link wins attention, the conversion tells you which wins business.
2. Copy/alias tests. The same destination, two aliases — this tests the label, which matters in SMS and print where the URL is read first:
https://yas.sh/get-started vs https://yas.sh/free-trial
3. Timing tests. Two sends (Wednesday vs Thursday, 9am vs 4pm) with identical content — the daily series in analytics makes the comparison.
The setup pattern
The discipline that keeps tests honest:
1. One question per test (no double-testing)
2. Two variants only (A/B, not A/B/C/D)
3. Parallel timing (same day/week — seasonality is real)
4. One difference between variants
5. Pre-declared sample size
Splitting traffic happens upstream (email tool, social scheduler, SMS platform) — each variant gets its own link; the links do the measuring.
Sample sizes: the math that prevents self-deception
The shortcut table for binary outcomes (clicked vs not):
| Expected difference | Clicks per variant needed |
|---|---|
| 10% → 11% | ~16,000 |
| 10% → 15% | ~1,300 |
| 20% → 25% | ~1,600 |
| 50% → 60% | ~400 |
Rule of thumb: 600–1,000 clicks per variant detects a meaningful (≥5-point) difference at 80% power. Under ~200 clicks per variant, "wins" are indistinguishable from noise — label the result a pilot and re-test at volume.
Reading the results
- Compare click-through rates, not totals. Total clicks scale with list size; the rate is the decision.
- Check the series, not just the total. A variant that wins on day one and dies by day three is a spike, not a signal.
- One metric per test. If clicks were the question, conversions are a follow-up test, not a bonus finding.
- Document the cutoff date. Results after the decision date don't retroactively change it — and cherry-picking the "right" window is how teams fool themselves.
Variant A: 512 clicks / 2,000 sent = 25.6%
Variant B: 428 clicks / 2,000 sent = 21.4%
Δ = +4.2 points — above the noise floor at this sample size → declare A
When not to A/B test
- Sample size impossible (B2B lists under ~500) — run a pilot and iterate qualitatively instead.
- One-off moments (launch day) — split attention and you underpower both variants; run a control/rollout instead.
- Infrastructure questions — whether redirects are fast or links are secure is not a popularity contest; those are measured differently.
Choosing the right metric for the test
Before you run anything, decide what "winning" means. There are two very different kinds of winners, and confusing them produces misleading conclusions:
- Click-through winner — which link gets more people to click. This measures attention and is the natural metric for copy and destination experiments where a click is the action you want.
- Conversion winner — which link leads to the actual business outcome (a signup, a purchase, a booking). This is the metric that pays the bills.
A variant can win clicks and lose conversions, or vice versa. The discipline is to pre-declare the primary metric and the decision threshold before you split traffic. If clicks are the question, measure clicks; if revenue is the question, wire up conversion tracking so the clicks connect to the outcome. Running a test without a pre-declared metric is how teams pick the variant that happens to match the conclusion they wanted.
Controlling the variables you can't split
Traffic-splitting tools handle the A/B split, but several variables silently corrupt results if you do not control them:
- Time of day. A variant shown at 9am vs one shown at 4pm is not a fair test. Run both variants in parallel over the same clock window, or use the daily series to compare like-for-like days.
- Audience composition. If variant A goes to your engaged subscribers and variant B to your cold list, differences are about the list, not the link. Randomize or stratify the split.
- Placement. A link at the top of a page outperforms one at the bottom. Give both variants the same placement.
- Seasonality. A Monday-send test compared with a Sunday-baseline test mixes in day-of-week effects. Compare within the same day type.
Naming the uncontrolled variables up front is most of the work. When you can name them, you can usually control or at least account for them.
Reading the series, not just the total
A single aggregate number can hide the truth. Plot the daily click series for each variant and look at the shape:
- A variant that wins every day by a consistent margin is a reliable winner.
- A variant that surges on day one and flatlines by day three is a spike, not a signal — often an artifact of when the send landed.
- A variant that trades the lead back and forth across days is probably measuring noise; you need more clicks before you conclude.
This is why the analytics daily series matters for testing: the shape over time is more honest than the total, and it lets you spot a transient effect that an aggregate would happily disguise.
A/B testing checklist
Before you launch any link test, run this quick checklist: one question pre-declared, two variants only, parallel timing, a single difference between them, a pre-declared sample size and primary metric, and a documented decision cutoff date. When all six are in place, the results are trustworthy enough to act on. When even one is missing, label the output a pilot and plan a re-test at volume — an honest pilot is worth more than a misleading "winner."
Conclusion
Two links, one question, a pre-declared sample size, and a series — that's an honest A/B test. The dashboard gives each variant its own analytics; the UTM guide keeps the variants comparable; conversion tracking closes the loop on what the winner actually earned.
