← All articles

    Strategy

    Baseball Needed 910 At-Bats

    BetterAug 3, 20267 min read
    Baseball Needed 910 At-Bats

    Baseball has a useful habit, it measures before it believes.

    In 2007, Russell Carleton showed that some hitter stats become readable after a little data, while others need far more. Strikeout rate became useful after about 60 plate appearances. Batting average needed 910 at-bats. The number people quote most often, the one that looks most decisive, was the least trustworthy one.

    That inversion is the point. Founders do the same thing in marketing. They kill a newsletter after four sends, a LinkedIn series after six posts, or a cold outbound motion after a few dozen replies. They read the outcome metric before the sample can support it.

    This is not a case for waiting forever. It is a case for reading the right metric at the right time.

    What stabilizes

    Carleton’s original work used split-half reliability on Retrosheet data from 2001 to 2006. In plain English, it asked whether the first half of a player’s season predicts the second half well enough to count as stable signal. Later revisions used Kuder-Richardson 21 and Cronbach’s alpha, but the practical lesson stayed the same, not all stats are equally readable.

    FanGraphs summarizes the canonical thresholds this way, a statistic does not stabilize, it becomes more stable, and that distinction matters. The label gets treated as a cliff. The data say otherwise.

    Approximate hitter thresholds, using the published baseball work, look like this:

    • Strikeout rate, about 60 plate appearances
    • Walk rate, about 120 plate appearances
    • Home run rate, about 170 plate appearances
    • Isolated power, about 160 at-bats
    • Slugging, about 320 at-bats
    • On-base percentage, about 460 plate appearances
    • Batting average, about 910 at-bats
    • BABIP, about 820 balls in play

    The full table is startling because the loudest statistic is often the slowest to read. The thing fans obsess over, batting average, is much less informative than the component rates underneath it.

    Why order matters

    The ordering is the lesson. The more a metric is a direct count of repeated, mostly binary events, the sooner it becomes useful. The more it mixes luck, sequencing, context, and compound effects, the longer you need before the signal separates from noise.

    Marketing has the same structure.

    Open rate is closer to strikeout rate. It is a simple event, did the person open or not. Reply rate, click-through rate, and positive meeting rate behave similarly. They are crude, but they move fast.

    Customer conversion, revenue per post, and channel CAC are closer to batting average. They are the numbers you care about. They are also the ones most likely to mislead you early.

    That is why founders often do the reverse of what they should. They ignore the readable early metric, then overreact to the unreadable late one.

    Marketing translation

    Here is the practical transfer.

    • Newsletter, open rate and click rate are readable quickly, customer conversion is not.
    • LinkedIn or X, impression to follow rate and engagement rate are readable quickly, pipeline per post is not.
    • Cold outreach, reply rate becomes readable early, meetings booked to closed revenue does not.
    • Landing pages, conversion rate can be estimated, but small deltas need far more traffic than most early-stage teams have.

    The consequence is uncomfortable. You can usually tell whether a channel has life in it before you can tell whether it produces revenue efficiently. That means the first decision is not, is this channel winning, it is, which signal am I entitled to read yet?

    A worked example

    Suppose a SaaS landing page converts at 3.8 percent, which is in the range reported in industry benchmark data. You want to know whether an observed 5.0 percent rate is real improvement or just noise.

    For a rough two-proportion check, the standard error of a proportion is:

    SE = sqrt(p(1-p)/n)

    If p is 0.038 and n is 1,000 visits, the standard error is about 0.006, or 0.6 percentage points. A 95 percent interval is roughly ±1.2 points. That means a measured 5.0 percent rate sits inside the range you could plausibly see from a true 3.8 percent page with ordinary variation.

    At 10,000 visits, the standard error shrinks to about 0.002, or 0.2 points. Now a move from 3.8 percent to 5.0 percent is more likely to reflect a real difference.

    The exact threshold depends on how certain you want to be, and how costly the decision is. That is the real rule. Not a magic number, a decision standard.

    Do the arithmetic

    When you evaluate a channel, do not borrow baseball’s numbers. Use the same logic.

    For a binary rate, the sample size needed to detect a change is driven by three things, baseline rate, desired lift, and confidence level. If you want a simple working estimate, use this approximation:

    n ≈ 16 × p(1-p) / d²

    Where p is the baseline rate, and d is the smallest lift you care about.

    If your newsletter open rate is 40 percent and you want to detect a 5 point change, the estimate is:

    n ≈ 16 × 0.4 × 0.6 / 0.05² = 1,536 sends

    That does not mean you need to wait for perfection. It means you should stop pretending that 48 sends can tell you much.

    For lower-frequency outcomes, the requirement rises quickly. If your close rate from reply to customer is 10 percent and you want to detect a 4 point shift, the sample needed becomes much larger. This is why most early-stage teams should separate leading indicators from final outcomes, then read each on its own schedule.

    Where it breaks

    The analogy is useful, but it is not perfect.

    First, baseball assumes a relatively stationary process. A founder’s skill changes mid-season. Product improves, messaging changes, targeting changes. A 2025 critique in arXiv argues that many stabilization heuristics miss exactly this, they treat all variation as sampling noise, even when the underlying process has changed. That critique is serious, and in marketing it may be more serious than in baseball.

    Second, teams can change the product while they measure. Baseball players cannot change their bat, their swing, and their opponent pool in the same way, at least not without leaving the analogy behind.

    Third, Carleton has spent years warning people not to turn his work into a simplistic cutoff chart. FanGraphs notes, correctly, that most of us picked up the word stabilize and ran with it. He has also written pieces with titles like Please Stop Talking About Statistical 'Stabilization'. The warning is worth respecting. These thresholds are not laws of nature, they are decision aids.

    A statistic does not stabilize, it becomes more stable.

    That sentence should be carried into marketing. Most founders are not asking whether a channel is objectively good. They are asking whether they have enough signal to keep going.

    When to wait

    If the outcome metric is unreadable, wait.

    Do not kill a newsletter on four sends because revenue is too low to justify the motion. Read open rate, click rate, reply quality, unsubscribes, and cost to produce. Those are the metrics that can speak early.

    Do not kill outbound on 30 emails because you did not close a deal. Read reply quality and meeting rate first. If those are weak, the motion may be wrong. If they are decent, the sample is probably too small to judge revenue.

    Do not kill content after nine posts because pipeline is flat. Read distribution, saves, engagement, and assisted traffic before concluding the channel is dead.

    The job is not patience as virtue. The job is sequencing.

    What to do

    1. Pick the leading metric that should move first.
    2. Set the review date before you start.
    3. Write the decision rule in advance.
    4. Cap the spend while the outcome is unreadable.
    5. Kill on cost and leading signals, not on thin outcome data.

    This is how you avoid the common founder error, confusing early silence with failure.

    Takeaway

    Baseball solved a hard version of your problem years ago. It showed that the numbers people care about most are often the slowest to become meaningful. Marketing is the same. Rate metrics speak early, outcome metrics speak late.

    If you are about to abandon a channel, ask a simpler question first, is this the right metric for the sample I have? If the answer is no, you are not making a decision. You are making a guess with a spreadsheet attached.

    That is the real lesson, and the reason this piece exists. Not to import baseball into marketing for style, but to close the gap between what founders feel and what their data can actually say.

    Share: Twitter LinkedIn
    sample size
    newsletter
    founder marketing
    metrics
    baseball