A/B Testing for Data Analysts — Complete Guide India
Experiment design, sample size calculation, randomisation, running the test, statistical analysis, interpreting results, and communicating findings to stakeholders — with Indian e-commerce, fintech, and product examples throughout.
The A/B Test Lifecycle — 7 Stages
Sample Size and MDE — Before You Start
Sample Ratio Mismatch (SRM) — The Silent Killer
SRM occurs when the actual user split differs significantly from the intended split (50/50). It indicates a problem in the experiment setup — biased assignment, bot traffic, or a logging bug. Never analyse results with SRM present; the comparison is invalid.
A/B Testing Pitfalls — India Context
Frequently Asked Questions
How long should an A/B test run in India?
An A/B test should run for at least 1–2 full business cycles — typically a minimum of 7 days and ideally 14 days. This captures weekday/weekend variation, which is significant in Indian consumer markets (weekend orders are typically 30–40% higher than weekdays). For monthly salary-driven spikes (1st–5th of the month), a 3-week run ensures you do not over-sample or under-sample salary-date traffic. Never stop a test the moment it becomes significant — run for the full planned duration. Two additional Indian-specific rules: (1) Do not run tests over Diwali, Navratri, or Holi unless you are specifically testing festive-season behaviour — these periods have abnormal traffic composition and the results will not generalise. (2) Be cautious running tests during IPL months (April–May) for sports or entertainment apps — traffic profile changes significantly. Plan experiment calendars around these events.
What is novelty effect and how does it affect A/B test results in India?
Novelty effect occurs when users in the variant group engage more with a new feature simply because it is new — not because it is better. This inflates variant performance in the first few days, gradually returning to the true steady-state effect as users habituate. This is particularly common in Indian mobile apps where users are highly engaged and tend to explore new UI elements. To mitigate: (1) Run tests for at least 2 weeks so the novelty wears off; (2) Analyse day-by-day metric trends — a genuine improvement maintains its lift over time, while novelty shows a declining trend; (3) If you have a returning user segment, compare the novelty-prone "first-week" cohort separately from "returning users who saw the variant multiple times" — the latter gives a more reliable signal. The opposite problem — change aversion — occurs when users initially resist a change but adopt it over time. Both effects mean short tests mislead.
What is the difference between A/B testing and multivariate testing?
A/B testing compares two versions of a single element: the control (A) vs one variant (B). It is simple, requires less traffic, and produces a clear go/no-go decision. Multivariate testing (MVT) simultaneously tests multiple elements and their combinations — e.g. headline colour (red vs blue) × CTA text ("Buy Now" vs "Add to Cart" vs "Shop Now") × banner image (3 options) = 18 combinations. MVT identifies interaction effects between elements (CTA text works better with a specific banner) but requires much more traffic because each combination needs statistical power. For most Indian startups and mid-size companies, A/B testing is the right tool — traffic is not high enough to power multivariate experiments. MVT is practical only at very high traffic volumes (millions of daily sessions). A sequential A/B approach — test headline first, then CTA on the winner — achieves similar learning with feasible sample sizes.
How do data analysts communicate A/B test results to non-technical stakeholders?
Lead with the business outcome, not the statistic. Instead of "p = 0.03, significant at α = 0.05," say: "The new checkout page increased order completion by 0.8 percentage points — from 3.2% to 4.0%. Over a month, this translates to approximately 2,400 additional orders and ₹34 lakh extra revenue." Structure every test readout as: (1) What we tested and why; (2) What we measured (primary metric); (3) What we found — absolute lift, relative lift, confidence interval; (4) Whether this is statistically reliable; (5) Business impact in INR or unit terms; (6) Recommendation with caveats. Always show confidence intervals rather than just point estimates — saying "the true lift is likely between 0.3% and 1.3%" is more honest than reporting "0.8% lift" as if it were exact. For inconclusive results, frame it as a decision about sample size or MDE, not as "the test failed." Quantify the cost of being wrong: if we ship this based on 80% confidence and it is a false positive, what does that cost?
EVIKA ACADEMY · NOIDA SECTOR 51
Learn A/B Testing on Real Product Data
Our curriculum covers experiment design, statistical significance testing, and result communication — applied to real Indian e-commerce and product datasets.
Book Free Demo Class →