I keep running into the same argument whenever I suggest testing a product change at a small company:

“We do not have enough traffic for statistical significance.”

And, technically, that may be true.

If your website gets 20, 50, or even a few hundred visitors a week, a traditional split A/B test can take forever to produce a trustworthy result. Divide that already-small audience between two versions, wait for enough conversions, account for normal variation, and you may be looking at months before the data tells you anything useful.

Sometimes it never will.

But I think people take the correct statistical observation and turn it into the wrong product conclusion.

Not having enough traffic for a formal A/B test does not mean you cannot experiment. It does not mean you cannot learn. It definitely does not mean the safest decision is to leave everything exactly as it is.

For a startup, a new product, or a low-traffic website, waiting for perfect statistical certainty can be its own form of bad decision-making.

“A/B Test” Has Become Too Loaded

Part of the problem is that people use the phrase “A/B test” to describe several different things.

At a large company, an A/B test usually means a controlled experiment. Users are randomly assigned to different variants. The team defines the success metric in advance. The test runs until it reaches an acceptable confidence level. Then someone decides whether the observed difference is likely to be real or just noise.

That is a useful standard when you have enough traffic.

It is also important when a small percentage change could affect millions of dollars, when the change is difficult to reverse, or when hundreds of teams need to agree on what happened.

But when I say “let’s test this” at a startup, I often mean something much simpler.

Let’s change the thing.

Let’s measure what happens.

Let’s pay attention.

Let’s see whether people respond differently.

Let’s use what we learn to decide what to try next.

That may not qualify as a rigorous A/B test. That is fine. The goal is not to publish a research paper. The goal is to make the product better.

The Donation Amount Example

A recent example was a discussion about suggested donation amounts.

The existing donation options appeared high relative to the audience and the context. I suggested trying a lower range, something like $5 to $20, seeing how conversion changed, establishing a baseline, and potentially testing higher amounts from there.

The immediate pushback was that the site had very little traffic and had historically received few meaningful online donations. Did we really need to run an A/B test with such small sample sizes?

Honestly, no.

A formal split test would probably be a poor use of the traffic. If only 20 or 50 people visit in a week, cutting that group into two variants makes an already noisy situation even noisier.

But the actual idea was still worth trying.

We could simply replace the suggested amounts for everyone. Then we could track how many people viewed the donation page, how many started the process, how many completed it, how much they donated, and whether the overall pattern changed.

Maybe nothing happens.

Maybe the lower amounts generate more completed donations but reduce the average donation size.

Maybe people who start with a smaller suggested amount ultimately choose a larger custom amount.

Maybe the page begins converting at all, which would already be more information than we had before.

None of those outcomes would be statistically conclusive after a handful of visitors. They would still give us more information than doing nothing.

Small Samples Are Noisy, but They Are Not Always Meaningless

There is a reasonable warning buried inside the statistical-significance argument.

Small samples can absolutely mislead you.

If one version gets three conversions and another gets one, that does not prove the first version is three times better. A single unusual visitor can completely distort the results. Traffic quality may change from one week to the next. A newsletter, campaign, news event, or social post may send visitors with very different intent.

You should not pretend weak evidence is strong evidence.

But there is a large gap between “this result is conclusive” and “this result tells us nothing.”

Suppose a signup flow normally gets one or two completions a month. You simplify the form, and suddenly ten people complete it over the next two weeks. That does not automatically prove causation, but it is worth investigating.

Suppose you lower a product’s starting price and support immediately starts receiving more questions from serious buyers. That is evidence.

Suppose you change the homepage message and sales calls begin with prospects repeating the exact language from the new headline. That is evidence too.

Early product decisions are almost never based on one clean dataset. They come from a combination of behavioral data, customer conversations, support requests, recordings, sales feedback, conversion patterns, and judgment.

The smaller the sample, the more carefully you should interpret each signal. You do not have to ignore the signal altogether.

Directional Learning Still Matters

At a startup, many experiments are not trying to detect a 2 percent improvement.

They are trying to answer much larger questions.

Do people understand what this product does?

Will anyone click this button?

Are visitors willing to start the checkout process?

Does the lower price remove an obvious barrier?

Do users complete the shorter form?

Does anyone use the new feature twice?

These are often closer to product discovery than conversion optimization.

If a product has almost no usage, debating whether one button color outperforms another would be ridiculous. But changing the onboarding flow, the offer, the pricing, or the primary call to action can still produce valuable directional information.

The expected effect size matters.

A low-traffic company probably cannot reliably measure a tiny improvement. It may still be able to observe a dramatic one.

When the baseline is close to zero, you are usually looking for a meaningful change in behavior, not a fractional optimization.

The Cost of Waiting Is Usually Ignored

People tend to focus heavily on the risk of drawing the wrong conclusion from limited data.

They spend less time thinking about the risk of not changing anything.

The current version of a product is not automatically correct just because it already exists.

It may be based on an assumption someone made six months ago. It may have been copied from a competitor. It may reflect a design preference rather than user behavior. It may have never been evaluated at all.

Leaving the existing version in place is still a decision.

If the current donation amounts are too high, keeping them for another six months has a cost. If the signup form is intimidating, waiting for more traffic before simplifying it has a cost. If users do not understand the homepage, continuing to send traffic there has a cost.

Statistical caution can sound neutral, but the status quo is not neutral.

Sometimes the cost of being wrong is low, the change is easy to reverse, and the potential upside is substantial. In those situations, insisting on a highly powered experiment before acting can be less rational than making the change and watching closely.

Reversibility Should Shape the Standard of Evidence

I think the amount of evidence required should depend partly on how risky and reversible the decision is.

Changing suggested donation amounts is easy to reverse.

Rewriting a headline is easy to reverse.

Reducing the number of fields in a contact form is usually easy to reverse.

Changing the company’s entire pricing model, removing a heavily used feature, or sending a major campaign to millions of customers requires a higher standard.

This seems obvious, but teams often apply the same experimentation expectations to every decision.

A low-risk interface change does not need the same proof as a permanent strategic commitment.

When a change is reversible, I would rather ship a reasonable hypothesis, monitor it, and correct quickly than spend weeks debating whether we will eventually have enough data to validate it.

The ability to reverse a decision is itself a form of risk management.

Use All the Traffic When Splitting It Makes No Sense

One practical mistake low-traffic teams make is assuming experimentation requires splitting the audience.

It does not.

When traffic is scarce, a before-and-after test may be more useful than a simultaneous A/B test.

Change the experience for everyone. Record the date. Track the same metrics before and after. Note any campaigns, seasonal effects, traffic changes, or outside events that could influence the comparison. Let the new version run long enough to encounter a reasonable range of users.

This still has limitations. A before-and-after comparison cannot control for time-based differences as cleanly as a randomized test.

But it lets the new version receive all available traffic, which may be far more practical when every conversion counts.

You can also combine the quantitative results with qualitative observation.

Watch session recordings.

Read support messages.

Ask users what confused them.

Look at where people stop.

Listen for changes in the questions prospects ask.

Check whether the people converting are actually good customers.

This produces a broader view of the change than a single conversion percentage ever could.

Establishing a Baseline Is a Valid Goal

Another thing that gets lost in these conversations is that sometimes there is no meaningful baseline yet.

A new product may have only a handful of users. A donation page may have received almost no gifts. A new feature may have launched without clear instrumentation.

In that situation, the immediate goal may not be to prove that version B beats version A.

The goal may be to establish what normal behavior even looks like.

How many people see the page?

How many interact with it?

Where do they leave?

What amount do they choose?

How long does the process take?

What questions do they ask afterward?

A baseline does not need to be permanent or statistically perfect to be useful. It gives the team a starting point. It helps improve instrumentation. It exposes obvious gaps. It makes the next experiment better informed.

You cannot optimize a process you barely understand.

Be Honest About the Strength of the Evidence

The solution is not to lower the standard and declare every small movement a win.

The solution is to describe the evidence accurately.

Do not say:

“The new version increased conversions by 200 percent.”

Say:

“We saw three conversions after the change compared with one in the previous period. The sample is too small to draw a firm conclusion, but the result is encouraging enough to continue observing.”

Do not say:

“Users prefer the new pricing.”

Say:

“We received more checkout starts after lowering the entry price, although traffic volume and audience quality varied between the two periods.”

Do not say:

“The experiment proved the hypothesis.”

Say:

“The initial behavior supports trying the next iteration.”

This kind of language matters. It allows a team to learn from weak evidence without overstating what it knows.

The problem is not acting with uncertainty. Every startup acts with uncertainty.

The problem is forgetting that the uncertainty exists.

Repeated Weak Signals Can Become Stronger

A single small experiment may not tell you much.

A sequence of experiments often does.

Suppose you lower the suggested donation amounts and see more donation starts. Then you simplify the form and see more completions. Then you adjust the page copy and hear donors mention that the process felt easier.

No individual result may be statistically significant. Together, they begin to tell a coherent story.

This is especially useful when the same pattern appears across different sources.

Behavioral data points in one direction.

Customer interviews point in the same direction.

Support requests decline.

Sales conversations become easier.

Retention improves.

The evidence chain becomes more persuasive even when no single test reaches the standard used by a large experimentation team.

Startups learn through accumulation. Each change reduces uncertainty a little. Each observation shapes the next hypothesis.

The Real Question Is Whether the Experiment Improves the Decision

The best question is not always, “Can this test reach statistical significance?”

A better question is:

“Will trying this help us make a better decision than we could make today?”

Sometimes the answer will be no.

The traffic may be too low. The metric may be too rare. The change may be too subtle. The surrounding conditions may be too chaotic. In those cases, the team may need customer research, usability testing, direct outreach, or a more substantial product change instead.

But sometimes the experiment is cheap, reversible, and capable of producing a useful directional signal.

Then it is probably worth doing.

The purpose of experimentation is to reduce uncertainty. Statistical significance is one tool for doing that. It is an important tool, but it is not the only one.

A small company should not imitate the experimentation process of a company with millions of users while ignoring the completely different environment it operates in.

Use rigor where rigor is possible.

Use judgment where it is necessary.

Measure what you can.

Be honest about what the data proves.

And keep learning.

Because having too little traffic for a perfect experiment is not a good reason to stop experimenting altogether.