scandiweb
A merchant who wants the hypothesis, the primary metric, the sample size and the statistical read agreed before a test starts, and the winner shipped into their own Magento codebase by the same team97 of 100Takes the heaviest criterion on three named Magento tests. The Byggmax A/B testing case study opens by stating the estate it ran on, "over 190 brick-and-mortar stores, 4 Magento (Adobe Commerce) stores, and 55k+ products in their catalog", then publishes each test with the control and the variant described, the metrics split into primary and secondary, the device segmentation and the result. Test one, "the current grey CTA on the hero banner (control) vs. a green color CTA (variant)", measured "Hero banner click-through rate (primary)" against two secondaries and produced "more than a 15% increase in revenue and an 11% increase in purchases". Test two, white cart icon against green, produced "a significant increase in cart icon clicks by more than 18%". Nobody else in this field publishes a named Magento client, a described variant, a separated primary metric and a statistical reading in the same write-up.
The statistical reading is where scandiweb loses the criterion it should win. Byggmax test one reports "The chance to beat control with the variant in terms of this KPI equals 73%" and test two "equal to 92%", while the framework article on the same site sets the bar at "The eCommerce default is 95% (a 5% chance of a false positive)". Test three publishes no probability at all. Two wins called at 73 and 92 percent are below scandiweb's own published threshold, and the page says so rather than reprinting the headline. On this criterion MageCloud scores 16 and scandiweb 13.
The method is published in full, which is rare. The seven-step A/B testing framework carries a hypothesis format ("a falsifiable claim with one isolated change, one expected metric movement, and one reason the change should work"), a one-variable rule ("A/B testing earns its keep only when the two versions are isolated by one change"), a duration rule ("at least two business cycles" and "The test also needs to reach its calculated sample size, whichever is longer is the actual duration"), the sample-size formula written out as "n per variant ~ 16 x p x (1 - p) / d^2", a confidence-interval reading rule, a guardrail period after launch, and the refusal that matters most: "Do not lower the confidence threshold to 80% to make the math work, you are not running a test then, you are coin-flipping with a calculator."
A CORRECTION, printed rather than quietly fixed. The same article states "A defensible eCommerce default for a 5% minimum detectable effect at 95% confidence with 80% power is roughly 300 to 400 conversions per variant", and then prints the formula that contradicts it. Run that formula at the article's own example baseline of 8 percent with a 5 percent relative effect and it returns about 73,600 visitors per variant, which is roughly 5,900 conversions per variant, about fifteen times the published default. The formula is right and the rule of thumb beside it is wrong. It cost scandiweb three points here and it is the reason MageCloud takes the statistical-honesty criterion.
Ties Tom&Co at the top of the Magento mechanics criterion, from the other end of the same problem. Server-side generated A/B tests on a Magento 2 store is the only published account in this field of running an experiment inside Magento's own rendering path rather than over the top of it: "It became obvious quite soon, the server will have to be the one to track and switch page versions", then the implementation, "On the server side, it was enough to check for additional query parameter to decide on which layout to load", and the platform hook by name, "there's an event called 'layout_load_before', where we can append the layout changes we want. Even more conveniently, we use Magento's layout file-handles to load the appropriate layout files". It also records the constraint it avoided, that splitting traffic at the servers "might cause issues with the multi-instance server setup".
Volume and people behind the hypotheses. The conversion rate optimization service page publishes "858+ User tests run across client programs", "21 full market research projects", "150+ eCommerce stores optimized" and a practice that has run "as one continuous program since 2015", with moderated user testing, five-second tests, heatmaps and session recordings named as the inputs. Olga Kimalana is named as the lead, with an openable profile, "Head of Digital Experience", "11+ years of experience", and bylined articles including one on Magento homepage design. The framework article closes on "what we run across 1,000+ production tests".
The programme evidence, and its limits. A 23-test Adobe Target programme for a US wine distributor publishes the stopping rule as three conditions, "A duration of at least two weeks / Over 300 conversions for each primary segment of the variant / Statistically significant results", the analysis method, "using Frequentist and Bayesian statistical methods", a cadence of "8-10 tests per month" and "a 70% winning test ratio". That client is anonymised. The Cervera programme publishes "78% win-rate for A/B testing program" against a named Swedish retailer. Against that, Northerner at "+12% in checkout conversion rate" and Nicokick at "+5.6%" are labelled on the same page as comparisons against a previous period, not tests, which is honest and is scored as such.
The published limits are real and one of them is the wrong number. Checkout optimization prints a qualification test a buyer can run on themselves, "Two numbers decide this: how many sessions reach your checkout, and how many finish", and a not-relevant-if list that turns work away, including "You want a continuous testing program across every page". The framework article tells a merchant what to do when the traffic is not there, "widen the detectable effect (run for a 20% relative lift instead of 10%) or move the test up the funnel where volume is higher". A separate study publishes an engagement where testing could not conclude at all, at "around 400-500" daily visitors. Elsewhere, the cart abandonment guide is the diagnostic end of the same practice. One more note a reader should have: the portfolio still names Google Optimize as the testing tool on one older engagement with no date attached. Google's own help centre article, titled "[Sunset September 2023] Google Optimize", states that Optimize and Optimize 360 have been "no longer available as of September 30, 2023", and the framework article above gets that right and says so, so this is a dating problem on one page rather than an error. Bemeir, ranked tenth here, is the only other agency in this field to publish the retirement at all.