Article
How much traffic you need before an A/B test means anything
The real statistical threshold for a valid A/B test and why many Romanian sites never reach it. What to do below it so you don't mistake noise for a result.
“We changed a button, it seemed to do better for a week, then everything went back to normal and we don’t know what happened.” If you have been through that, the problem was almost certainly not the test but the lack of statistical discipline behind it. A test run below the required threshold does not give you greater uncertainty, it gives you a false positive with the appearance of certainty, which is more dangerous than not testing at all.
Why most companies in Romania do not have enough traffic
The sample size needed for an A/B test at 95% confidence and 80% statistical power depends on the baseline conversion rate of the page being tested and on the minimum effect you want to be able to detect (MDE, minimum detectable effect). The relationship is quadratic, not linear: halving the MDE quadruples the traffic required, and a 1% baseline rate needs roughly 11 times more traffic than a 10% one to detect the same relative effect.
Concretely, at a 3% baseline rate, close to the Romanian market average, detecting a 10% relative effect takes approximately 106,000 visitors per variant. A 20% effect takes around 28,000 visitors per variant. A simpler threshold, common in the industry, is a minimum of 5,000 visitors a week and at least 500 conversions a week on the page being tested, that is 250 per variant.
Those figures explain honestly why few companies in Romania reach the threshold. According to the European Commission’s 2025 Digital Decade report, CRM systems are present in under 18% of Romanian companies, against an EU average of around 28.5%, and only about 12% of Romanian SMEs sell online. The average conversion rate on Romanian online shops sits around 1%, ranging between 1% and 3%, significantly below the above 2% recorded in markets such as the UK, the US or Australia. With a low baseline rate and low total site traffic, very few individual pages ever reach the volume needed for a trustworthy result within a reasonable timeframe.
The errors that inflate false positives
Even with apparently sufficient traffic, testing badly produces wrong conclusions. The most frequent: stopping the test too early, known as “peeking”. Repeatedly checking the results and stopping at the first sign of significance inflates the real false positive rate well past the nominal 5% threshold; on a sample of 1,000 per variant it can exceed 20%, four times the stated level.
The second error is uncorrected multiple testing: tracking several metrics or segments at once, with no statistical correction, raises the chance of finding “something significant” by accident. With 10 metrics tracked, the chance reaches around 40%; with 20, over 64%. The third is misunderstanding statistical significance: p<0.05 does not mean “a 95% chance that the variant is better”, it means that with no real effect, a result this extreme would appear only 5% of the time. It says nothing about the size or the commercial relevance of the effect; a “statistically significant” increase of 0.1 percentage points can be entirely irrelevant to the business. On top of those comes a rule independent of sample size: any test has to run for at least one full business cycle, 2-4 weeks, so that day of week variation or seasonality is not mistaken for a real effect, even if the sample threshold looks reached earlier.
A factor that distorts the data before you even start testing
The cookie consent banner genuinely changes the numbers in Google Analytics 4, not just your legal obligations. With advanced Consent Mode, which models lost conversions, at a refusal rate of 40% of visitors, modelling recovers about 70% of the lost traffic but leaves a gap of about 12% of total conversions. At 60% refusal, the gap rises to about 27%. The higher the refusal rate, which is common in European markets, the more systematically GA4 underestimates real performance, and any testing decision based on it alone starts from an incomplete picture. Google Consent Mode v2 has, in any case, been mandatory for visitors from the European Economic Area since March 2024, with strict enforcement from 2025, and in Romania the ANSPDCP, the national data protection authority, has already fined a site in 2025 for placing cookies without prior consent.
What to do if you are below the threshold
This is where the approach differs from one that sells A/B testing to everyone regardless of traffic. Below the statistical threshold there are real alternatives, not “weaker” compromises.
Proxy metrics or micro-conversions, goals with higher volume than the final conversion, CTA clicks, scroll depth, add to cart, validated later through their correlation with real conversion. Qualitative research instead of quantitative testing: heatmaps, session recordings, surveys, which identify obvious friction without a statistical sample. A larger MDE, with bolder changes instead of micro-tweaks, because tests with a large expected effect need a smaller sample. Controlled before and after comparison, with explicit control for seasonality and other simultaneous changes, instead of a classic A/B. Aggregation at page type level, where a test applied simultaneously across all pages in a category builds aggregate volume far faster than testing them in isolation. And, crucially, an explicit reassessment threshold: when aggregate traffic justifies classic testing, you move to it; until then the method above is the correct one for the volume available.
What happens to attribution on long sales cycles
Even when traffic is sufficient for testing, another limitation affects decision quality: for B2B and high ticket services, the typical sales cycle runs between 3 and 18 months, and 40-60% of the buyer’s journey happens in offline or untrackable channels, calls, meetings, demos. It is no surprise that around 90% of B2B marketing teams report real attribution problems: a last-click model over-rewards the last channel touched, often branded organic, effectively free, and ignores the channel that started the interest, often paid, with a real cost. An A/B test run correctly in statistical terms, but interpreted through broken attribution, can point to the wrong winner.
The practical fix is not complicated, but it takes discipline: a UTM convention the whole team respects, storing the first touch source with an extended lifetime, 6-12 months, aligned to the real sales cycle rather than the 90 day default, and hidden form fields that carry the source straight into the CRM, visible at individual lead level. Without that level of discipline, every testing conclusion stays partly suspect, however correctly the sample was calculated.
Why the difference between reporting activity and reporting results matters
An activity report, impressions, clicks, raw pageviews, is easy to produce and hard to argue with. A results report asks for more: qualified leads, not just submitted ones, cost per qualified lead, a qualification rate confirmed in the CRM. The difference is not stylistic. Activity reporting holds someone accountable for volume, results reporting holds them accountable for commercial impact, and that requires CRM access and a shared definition, agreed in advance, of a “qualified lead”.
If your traffic is below the threshold needed for a statistically valid test, the best thing you can do is not to test anyway and draw unsafe conclusions, but to find out exactly where you stand relative to the threshold and which method makes sense for your real volume. The CRO & analytics page describes exactly that calculation, including the case where the correct answer is “not yet”.
Sources
- Digitalizarea economiei României în 2025, raport CE DESI/Deceniul Digital, Webactiv.ro
- Rate de conversie: ce sunt, cât sunt, cum se cresc, Creadiv.ro
- Optimizare Rata Conversie CRO eCommerce, Roweb Media
- A/B Test Sample Size Guide, Atticus Li
- A/B Testing Alternatives for Low-Traffic Websites, Rich Page
- A/B Testing Alternatives for Low-Traffic Websites, CXL
- A/B Testing Mistakes That Invalidate Results, MetricGate
- The more the merrier? The problem of multiple comparisons in A/B Testing, Statsig
- Correct me if I’m wrong: Navigating multiple comparison corrections in A/B Testing, Statsig
- GA4 Cookie Consent: Data You Lose When Users Opt Out, Kukie.io
- Google Consent Mode V2 Setup Guide, CookieHub
- Amendă GDPR de 10.000 EUR pentru Călin Georgescu, după ce site-ul său a folosit cookie-uri fără consimțământ, StartupCafe
In a similar situation?
Tell us where you stand. If the right answer is not the one you just read here, we will say so.
Request a quote