The Experiment RegisterMagento and Adobe Commerce agencies, scored on the test evidence they publish Updated 29 September 2026

Magento CRO agencies, ranked on the test evidence they publish

Ten Magento and Adobe Commerce agencies were read on their own websites on 29 September 2026 and scored out of 100 on one question: do they publish a concluded A/B test with the variant described, the client named, the lift measured and a statistical reading attached to it. Two of the ten do. scandiweb scores 97 and ranks first, publishing three named Magento A/B tests with primary and secondary metrics and a chance-to-beat-control figure, a seven-step framework carrying the sample-size formula, and a server-side test implemented inside Magento's own layout system. MageCloud ranks second on 75 and beats scandiweb outright on statistical honesty, because its one published test is read at 99.75 percent confidence with both sample sizes and the non-significant metrics disclosed, while a sample-size default published on scandiweb.com is wrong by more than an order of magnitude against the formula printed next to it. The field is thin: page one for this search is blog posts, a tool vendor and an extension vendor dated 2013, and not one page-one result is an agency service page with a concluded test on it.

1 The shortlist

Every agency on this page, in order

1
scandiweb A merchant who wants the hypothesis, the primary metric, the sample size and the statistical read agreed before a test starts, and the winner shipped into their own Magento codebase by the same team 97 of 100.
2
MageCloud A merchant who wants to read one complete test report before hiring anybody, including the sample sizes, the dates, the confidence figure per metric and the metrics that did not reach significance 75 of 100.
3
Blue Acorn iCi An Adobe Commerce and Adobe Experience Manager estate that wants a multivariate test designed, arbitrated and graduated to full traffic by a team that publishes its own significance threshold 73 of 100.
4
Absolute Design A smaller Magento or Hyva store that wants the CRO workshop, the hypothesis list and the prioritisation matrix written down first, and is comfortable that the published test evidence is from the agency's Shopify side 60 of 100.
5
Tom&Co An Adobe Commerce estate already licensing Adobe Target, or about to, that needs the implementation audited for flicker and delivery before anyone argues about what to test 51 of 100.
6
Vervaunt A Magento merchant who wants the hypothesis discipline and the research sources taught to an in-house team rather than outsourced, and who will ask for one concluded test with a number on it 42 of 100.
7
Hatimeria A Magento 2 store choosing between Luma, Hyva and a React checkout, who wants that decision made on live traffic rather than on a vendor comparison table 37 of 100.
8
Krish Technolabs An Adobe Commerce or Magento merchant who wants a free audit first and a retainer with experiment governance attached, and who will ask what statistical bar each published figure cleared 36 of 100.
9
Gene Commerce A Magento or Adobe Commerce merchant whose conversion problem is specifically inside the payment step, and who values a published operating limit on the fix more than a test report 26 of 100.
10
Bemeir A Magento store that has outgrown visual-overlay testing and needs the experiment routing built into the platform, and whose buyer is willing to take the delivery record on trust 25 of 100.

Ten agencies that publish both a Magento or Adobe Commerce capability and something checkable about experimentation, scored out of 100 against six weighted criteria. Read on their own websites on 29 September 2026. The heaviest criterion, worth 30, asks for one thing: a concluded test published with the variant, the client, the lift and a statistical reading. Two agencies take it. Three of the ten publish no conversion result from a test at all.

2 How these were judged

The six criteria, and what each one is worth

CriterionWhat a pass looks likeWhat a fail looks likeWeight
A concluded test published with the number, the variant and the clientThe heaviest criterion, because it is the only one that separates a test from a redesign. 30 where the variant, the client, the measured lift and a statistical reading for that lift are all published. 26 where the variant, the client and the lift are published with the test design visible, the arms, the allocation or the losing arm reported, but no statistical reading. 22 for a named client and a measured lift from a test where the variant is identified but the design is not published. 16 where the client or the variant is anonymised, or where the headline figure is a period comparison after rollout rather than the in-test delta.10 for a programme-level figure only, a win rate or a typical uplift band with no individual test behind it. 5 where a test is described with a direction but no figure, or where conversion figures appear with no client attached. Nothing where no conversion result of any kind was found on the pages read, including where the agency sells A/B testing and publishes only adjectives. A figure inside a testimonial is scored as a testimonial, and a figure in an unattributed stat band is scored on the bottom rung however large it is.30
Test design published as a method: hypothesis, primary metric, one variable, a duration or sample-size rule, prioritisationThis is the part a merchant cannot check after signing, so it is scored on what is written down before the money moves. Five things counted, four points each: a hypothesis format rather than the word hypothesis; a named primary metric, ideally with secondaries separated from it; a one-variable rule; a duration or sample-size rule fixed before the test starts; and a prioritisation model with its inputs named. 20 for all five, 16 for four, 12 for three, 8 for two, 4 for one.Nothing where the method is described only as testing and iterating, or where the process is a list of stage names with nothing inside them. A framework acronym with no published content behind it scores nothing. Scores nothing for a page that explains what A/B testing is in general terms without saying how this agency runs one. A test with no pre-registered primary metric cannot be wrong, which means it cannot be right either.20
Statistical honesty: significance, power, sample size, inconclusive and losing testsA lift with no confidence attached is a number, not a result. 16 where a numeric confidence or significance figure is published against a real test, a non-significant or inconclusive outcome is disclosed rather than dropped, and no published statistical claim is wrong on its own terms. 13 for the same with one published statistical claim that is wrong on its own terms, or for a numeric significance threshold published as the agency's own standard plus a stated policy on losing or inconclusive tests. 10 for a published win rate or a stated commitment to report what every test won and lost, with no numeric threshold.6 where significance is discussed as a requirement with no number and no losing-test policy. 3 for the words statistically significant or statistically validated with nothing behind them. Nothing where no statistical statement was found on the pages read. Publishing losses is the single strongest signal that the tests are real, so a page that shows only winners is scored as showing only winners.16
Experimentation on Magento specifically, with the platform mechanics namedMagento breaks client-side testing in ways a hosted platform does not. 14 where the mechanics that constrain experimentation on this platform are named and explained with the remedy for each: flicker or flash of original content, full page cache or Varnish, client-side against server-side delivery, or the theme and checkout constraints. 11 for two or more of those named with the remedy specified, or for a test actually run against Magento-specific components, a checkout, a theme, a layout or a storefront, and published as such. 8 where the platform's testing options are stated correctly, that Magento Open Source ships no native A/B testing and Adobe Target is the Adobe route, with the tools that fill the gap named.5 where Magento or Adobe Commerce is named as a platform the agency tests on with no platform-specific mechanic. 2 where Magento is a service on the site but never appears in connection with testing. Nothing where no Magento or Adobe Commerce connection to testing was found on the pages read. A Magento page whose case studies all run on another platform scores on what the cases say, not what the heading says.14
Named research inputs that generate the hypothesesTest ideas have to come from somewhere, and research volume is the input to a win rate. 12 for a published volume of research, a count of user tests, sessions or studies, plus the named methods behind it and a named person who leads it with a profile a reader can open. 9 for named research methods, moderated user testing, surveys, session recordings or heatmaps, with either a named lead or a published volume but not both. 6 for named research tools only, with no volume and no named lead.3 where research is described generically as data analysis or understanding user behaviour with no method and no tool named. Nothing where no research input was found on the pages read. A tool logo strip scores nothing. A named person scores only where the profile opens: a byline with a joke biography, or a role label with no surname, is scored as no named person.12
Published limits: traffic floors, what will not be tested, when not to testThe sentence a sales page cuts. A merchant on 8,000 sessions a month needs to be told a test will not conclude. 8 where a numeric floor is published, traffic, sessions, conversions or revenue, together with what the agency does or recommends instead below it. 6 for a published qualification test the buyer can apply to themselves, or an explicit statement that the agency will say when a test cannot reach significance, with no number. 4 for a published not-relevant-if or we-do-not-do-this list that touches testing.2 for a caveat that results take time or depend on traffic. Nothing where no limit was found on the pages read. This criterion carries the smallest weight because it is the easiest to write, not because it matters least: it is the only line on any of these sites that can save a merchant from buying a programme their traffic cannot support.8

3 The ranking

The ten Magento CRO agencies, scored out of 100 in 2026

1

scandiweb

A merchant who wants the hypothesis, the primary metric, the sample size and the statistical read agreed before a test starts, and the winner shipped into their own Magento codebase by the same team97 of 100

Takes the heaviest criterion on three named Magento tests. The Byggmax A/B testing case study opens by stating the estate it ran on, "over 190 brick-and-mortar stores, 4 Magento (Adobe Commerce) stores, and 55k+ products in their catalog", then publishes each test with the control and the variant described, the metrics split into primary and secondary, the device segmentation and the result. Test one, "the current grey CTA on the hero banner (control) vs. a green color CTA (variant)", measured "Hero banner click-through rate (primary)" against two secondaries and produced "more than a 15% increase in revenue and an 11% increase in purchases". Test two, white cart icon against green, produced "a significant increase in cart icon clicks by more than 18%". Nobody else in this field publishes a named Magento client, a described variant, a separated primary metric and a statistical reading in the same write-up.

The statistical reading is where scandiweb loses the criterion it should win. Byggmax test one reports "The chance to beat control with the variant in terms of this KPI equals 73%" and test two "equal to 92%", while the framework article on the same site sets the bar at "The eCommerce default is 95% (a 5% chance of a false positive)". Test three publishes no probability at all. Two wins called at 73 and 92 percent are below scandiweb's own published threshold, and the page says so rather than reprinting the headline. On this criterion MageCloud scores 16 and scandiweb 13.

The method is published in full, which is rare. The seven-step A/B testing framework carries a hypothesis format ("a falsifiable claim with one isolated change, one expected metric movement, and one reason the change should work"), a one-variable rule ("A/B testing earns its keep only when the two versions are isolated by one change"), a duration rule ("at least two business cycles" and "The test also needs to reach its calculated sample size, whichever is longer is the actual duration"), the sample-size formula written out as "n per variant ~ 16 x p x (1 - p) / d^2", a confidence-interval reading rule, a guardrail period after launch, and the refusal that matters most: "Do not lower the confidence threshold to 80% to make the math work, you are not running a test then, you are coin-flipping with a calculator."

A CORRECTION, printed rather than quietly fixed. The same article states "A defensible eCommerce default for a 5% minimum detectable effect at 95% confidence with 80% power is roughly 300 to 400 conversions per variant", and then prints the formula that contradicts it. Run that formula at the article's own example baseline of 8 percent with a 5 percent relative effect and it returns about 73,600 visitors per variant, which is roughly 5,900 conversions per variant, about fifteen times the published default. The formula is right and the rule of thumb beside it is wrong. It cost scandiweb three points here and it is the reason MageCloud takes the statistical-honesty criterion.

Ties Tom&Co at the top of the Magento mechanics criterion, from the other end of the same problem. Server-side generated A/B tests on a Magento 2 store is the only published account in this field of running an experiment inside Magento's own rendering path rather than over the top of it: "It became obvious quite soon, the server will have to be the one to track and switch page versions", then the implementation, "On the server side, it was enough to check for additional query parameter to decide on which layout to load", and the platform hook by name, "there's an event called 'layout_load_before', where we can append the layout changes we want. Even more conveniently, we use Magento's layout file-handles to load the appropriate layout files". It also records the constraint it avoided, that splitting traffic at the servers "might cause issues with the multi-instance server setup".

Volume and people behind the hypotheses. The conversion rate optimization service page publishes "858+ User tests run across client programs", "21 full market research projects", "150+ eCommerce stores optimized" and a practice that has run "as one continuous program since 2015", with moderated user testing, five-second tests, heatmaps and session recordings named as the inputs. Olga Kimalana is named as the lead, with an openable profile, "Head of Digital Experience", "11+ years of experience", and bylined articles including one on Magento homepage design. The framework article closes on "what we run across 1,000+ production tests".

The programme evidence, and its limits. A 23-test Adobe Target programme for a US wine distributor publishes the stopping rule as three conditions, "A duration of at least two weeks / Over 300 conversions for each primary segment of the variant / Statistically significant results", the analysis method, "using Frequentist and Bayesian statistical methods", a cadence of "8-10 tests per month" and "a 70% winning test ratio". That client is anonymised. The Cervera programme publishes "78% win-rate for A/B testing program" against a named Swedish retailer. Against that, Northerner at "+12% in checkout conversion rate" and Nicokick at "+5.6%" are labelled on the same page as comparisons against a previous period, not tests, which is honest and is scored as such.

The published limits are real and one of them is the wrong number. Checkout optimization prints a qualification test a buyer can run on themselves, "Two numbers decide this: how many sessions reach your checkout, and how many finish", and a not-relevant-if list that turns work away, including "You want a continuous testing program across every page". The framework article tells a merchant what to do when the traffic is not there, "widen the detectable effect (run for a 20% relative lift instead of 10%) or move the test up the funnel where volume is higher". A separate study publishes an engagement where testing could not conclude at all, at "around 400-500" daily visitors. Elsewhere, the cart abandonment guide is the diagnostic end of the same practice. One more note a reader should have: the portfolio still names Google Optimize as the testing tool on one older engagement with no date attached. Google's own help centre article, titled "[Sunset September 2023] Google Optimize", states that Optimize and Optimize 360 have been "no longer available as of September 30, 2023", and the framework article above gets that right and says so, so this is a dating problem on one page rather than an error. Bemeir, ranked tenth here, is the only other agency in this field to publish the retirement at all.

2

MageCloud

A merchant who wants to read one complete test report before hiring anybody, including the sample sizes, the dates, the confidence figure per metric and the metrics that did not reach significance75 of 100

Publishes the single best test report found anywhere in this research, and it is the reason MageCloud beats scandiweb on statistical honesty. Its Thermos delivery progress bar test states the hypothesis under its own heading, the window, "we ran a controlled A/B test from 23 June to 4 July 2026", the allocation, "the original cart (12,792 visitors) and the variation with the progress bar (13,058 visitors) 25,850 visitors in total", the pre-declared metric, "The primary metric was completed checkouts", and the result with the baseline, the variant and the confidence on one line: "1.74% to 2.37% conversion rate, 99.75% confidence".

It then does the thing almost nobody does. Secondary metrics carry their own confidence figures, "+66.40% revenue per visitor (£0.50 to £0.83, 99.36% confidence)" and "+20.03% checkout starts (95.36% confidence)", and the metrics that failed are disclosed rather than dropped: "Later checkout steps (contact and shipping details) showed directional improvements that did not reach significance on their own". It also separates proof from interpretation, "the mechanisms themselves are interpretation rather than something the data directly proves", and states the operating principle, "Ideas earn their place on the site by winning a test, not by being someone's preference".

Where it loses is Magento specificity and depth of practice. The Magento CRO page publishes "12 YEARS OF EXPERIENCE", "MAGENTO CERTIFIED TEAM" and A/B testing as a service line, and the testing offer covers "Magento, Shopify, and WooCommerce", but no platform-specific testing mechanic was found on the pages read, and the Thermos page does not state which platform thermos.co.uk runs on. The case cards on the Magento page itself carry no test numbers. No testing tool is named for the Thermos test, no research volume and no named optimisation lead were found on the pages read, and the only published limit is that "Initial gains can appear within weeks".

3

Blue Acorn iCi

An Adobe Commerce and Adobe Experience Manager estate that wants a multivariate test designed, arbitrated and graduated to full traffic by a team that publishes its own significance threshold73 of 100

The only agency here that publishes both a real Adobe Commerce practice and a fully designed experiment. Its SeaWorld conversion test publishes the allocation, "Each customer had an equal chance (33 percent chance each) of being in the Control, Variation 1, or Variation 2 group", both variants in words, the results by device, "a 4.7% lift in Conversion Rate, a 5.92% lift in RPV" on mobile against "2.74%" and "3.83%" on desktop, and the arm that lost: "Users in Variation 2 who were directed immediately to the Tickets landing page did not perform as well as those in Variation 1 or the Control." The stopping rule is stated, the experiment was complete "once the metrics reached statistical significance and traffic volume was high enough", but no confidence figure is attached to the lift, which is the four points between this entry and the top rung.

The method article behind it is the most complete published by any agency in the pool that also works on Adobe Commerce. Conversion testing: what brands really need to know publishes a hypothesis format with a worked example, the choice between A/B, split-URL and multivariate testing with the rule for each, the sample-size arithmetic as "Sample Size x Number of Variations in Your Experience = Total Number of Visitors You Need", and a numeric threshold as the house standard: "Optimization experiments typically run at a 90 percent statistical significance... Anything below 90 percent will increase the risk of deploying a 'loser' page or funnel." It also publishes what happens when nothing wins, "If there is no clear winner, then put a hold on deployment", and a seasonal refusal, "We do not recommend any significant UX changes during the holiday season".

The Adobe Commerce credentials are current and the testing content sits on separate pages from them. The commerce practice page publishes a solution "Built on Adobe Commerce and Adobe Experience Manager" that "has been accredited by Adobe", and the about page carries "Blue Acorn iCi Recognized as the 2020 Adobe Emerging Solution Partner of the Year". What is missing is the join: no Magento or Adobe Commerce testing mechanic was found on the pages read, and the named optimisation leads, a Head of Strategy and Analytics and a Director of Insights and UX, are quoted in body copy without individual profile pages. Most other case studies are download-gated, so SeaWorld is the only numeric test a reader can open.

4

Absolute Design

A smaller Magento or Hyva store that wants the CRO workshop, the hypothesis list and the prioritisation matrix written down first, and is comfortable that the published test evidence is from the agency's Shopify side60 of 100

Takes the test-design criterion outright, which nobody else in the top three manages. Its Splash About A/B test case study publishes the sequence, "a full CRO workshop, research and data collection, hypothesis generation, prioritisation matrix, design, A/B tests and implementation", a hypothesis in the exact words it uses, "Changing the Buy Now CTA button colour to a more vibrant colour will increase click-through rate because it stands out to the user more", and a prioritisation model with its three inputs named: "business value... implementation complexity... and certainty of improvement". The one-variable rule is published too, on the how-to article: "Ensure that only one element is changed at a time to isolate the impact of the tested variable."

The results are real and the platform is the catch. The same case study reports "a revenue increase of 4%, conversion rate increase of 15.44% and a revenue per customer increase of 22.55%", with the tool named, and states what happened next, "these tests... have been promoted to the main variant". But the client was moved to Shopify Plus before the testing began, so the published tests are not Magento tests. Magento work is published separately, including "Magento Hyva Theme & Hyva Checkout for Charles Bentley", and the CRO service page names "Adobe Silver Solution Partner and Shopify Plus Partner", but no Magento testing mechanic was found on the pages read.

The statistical half is the weak half. Significance appears as a requirement with no number, "set a sufficient test duration to achieve statistically significant results", and the low-traffic problem is named as a challenge rather than a floor: "achieving statistically significant results, especially for stores with lower traffic volumes... businesses can focus on high-impact areas". There is a published observation minimum, "Typically we need at least one month so that we can capture week-over-week trends", and a policy on tests that fail, "even inconclusive tests offer valuable learning opportunities", but no confidence figure, no sample size and no win rate were found on the pages read. Heatmap and behaviour tools are named; no research volume and no named CRO lead were.

5

Tom&Co

An Adobe Commerce estate already licensing Adobe Target, or about to, that needs the implementation audited for flicker and delivery before anyone argues about what to test51 of 100

Ties scandiweb at the top of the Magento mechanics criterion and is the only agency in the pool to name the failure mode that kills client-side testing on a commerce storefront. Its Adobe Target page puts it first: "Flicker, first. Nothing damages confidence in a testing programme faster than customers watching the page change under them. Getting the implementation right, pre-hiding, async loading tuned properly, server-side delivery where it matters, is the difference between a programme people trust and one the brand team quietly kills." It carries the headless case too, "On a decoupled storefront, client-side delivery gets awkward. Server-side and hybrid delivery are usually the right answer", and the inherited-instance diagnosis, "We usually start by auditing the implementation for flicker and delivery problems".

It also takes the published-limits criterion outright, with a number. "A site with 50,000 monthly sessions cannot run six concurrent tests and detect a 2% lift in anything. We size the roadmap to the traffic you have, and say so when the honest answer is that a test will never reach significance." And the guidance below that floor, which no other agency here supplies: "as a rough guide, tens of thousands of sessions a month per test if you are looking for single-digit conversion lifts. Below that, personalisation and targeting rules deliver more than classical A/B testing. We will tell you which camp you are in before you buy anything." On null results it is equally plain: "A test that shows no difference is a real result and it saves you building the thing."

What is missing is the evidence. No concluded A/B test, no per-test lift, no confidence figure, no sample size and no test duration were found on any page read. The only figure attached to testing is a band, "Typical clients see conversion uplifts in the 15-30% range over the first six months. That figure comes from the programme, not the software", and every named-client number, including "45% increase in mobile conversion rate" for one Adobe Commerce retailer, is a replatform or build outcome. The Adobe credentials are strong, "an official Adobe Solutions Partner and certified Adobe Commerce partner, with 30+ certified specialists", and a dedicated CRO service page returns 404. A buyer should ask for one concluded test read with its confidence figure before the first invoice.

6

Vervaunt

A Magento merchant who wants the hypothesis discipline and the research sources taught to an in-house team rather than outsourced, and who will ask for one concluded test with a number on it42 of 100

Publishes the clearest hypothesis discipline of any consultancy in the pool. How to generate ideas for CRO tests sets the format, "usually in an 'If, Then, Because' format", and then works a real one end to end, from the insight, "From GA4, we noticed a high frequency of sessions who add an item to the cart, but do not complete the checkout", through the question and the experiment idea to the hypothesis itself. It names its research partner and its ideation sources, "Baymard Institute, Guess the Test, GoodUI and Evidoo", and publishes a corroboration rule, "We often try to corroborate any insight found by one method with another". Two named authors carry openable biographies.

Magento-specific testing is published, and it is old. Improving the conversion rate of Magento stores describes a real Magento checkout experiment, "we found that removing the header and footer and various other unnecessary elements of the checkout pages had a positive impact on abandonment rates, and the one step checkout did perform slightly better than our alternative (three steps, email, delivery, payment, executed via AJAX)", which is a test run against a Magento component and published as such. It also states the precondition nobody else states so plainly, "realistically you need to have a certain amount of traffic and trade to effectively use testing". The article carries a 2015 byline and recommends a Magento Enterprise full-page-cache module, so a reader should treat the tooling as historical and the mechanics as current.

The gap is numbers. No concluded test of theirs is published with a figure: the results that do carry figures are anonymised and attached to other work, "a luxury retailer... managed to increase their conversion rate by 50% through significant performance improvements". The service page publishes two tiers, CRO Establish and CRO Embed, with what each one tests, and an impartiality pledge, "we work with all technology partners and do not favour solutions based on commission". The only win rates on the site are quoted from a third party and clearly attributed as such. No significance figure, no sample size, no stopping rule and no losing-test policy were found on the pages read.

7

Hatimeria

A Magento 2 store choosing between Luma, Hyva and a React checkout, who wants that decision made on live traffic rather than on a vendor comparison table37 of 100

Publishes the most Magento-native experiment in the field: three checkouts tested head to head on one store. How we increased checkout conversion by 67.5% names the client, "Bedre Naetter is a Danish company, one of Europe's largest manufacturers and suppliers of beds, mattresses, sofas, quilts, pillows and bedding", and the arms, "We ran 3 simultaneous tests... React Checkout / Hyva Checkout (based on Magewire) / Luma Checkout". The design is published in four steps including "Random assignment: We then randomly assigned customers to one of three payment options", the measurement source is named, "data from New Relic", and the winner is reasoned rather than asserted.

The headline number is not the test result, and the page says which it is. "Between June and August 2024, Bedre Naetter achieved a 67.5% increase in checkout conversion rate compared to the same period in 2023" is a year-on-year comparison after the winner was rolled out. No in-test delta between the three checkouts is published, and the only statistical statement is "we used statistical tools to accurately compare the performance of the three payment options". That is why a test with genuinely good design scores 16 rather than 26 on the heaviest criterion.

The Adobe Commerce integration route is published as a tutorial, which is useful and undisciplined in equal measure. Starting A/B testing with Optimizely and Adobe Commerce walks a merchant through the snippet, the variation and the click tracking, then reads the result with no waiting period and no significance gate: "The results are visible immediately so you can quickly and easily check what works best for your business." Read against every vendor's own documentation, including Optimizely's own, that is the single most expensive habit in experimentation. Hatimeria is a Hyva Gold Partner with a headless Magento history going back to 2017 and three named authors; no traffic minimum, no win rate and no losing-test policy were found on the pages read.

8

Krish Technolabs

An Adobe Commerce or Magento merchant who wants a free audit first and a retainer with experiment governance attached, and who will ask what statistical bar each published figure cleared36 of 100

Publishes three CRO case studies with numbers, one platform sentence nobody else writes, and no statistics at all. The CRO and A/B testing service page is the only page in this entire field to say that the platform bounds the recommendation: "we have deep experience working within the constraints and architectures of Shopify Plus, Adobe Commerce, and Magento, ensuring our CRO recommendations are scoped to what your specific platform allows". It names its experimentation stack, "Optimizely, GrowthBook, VWO", and its retainer scope includes "Experiment Design & Implementation", "Experiment Monitoring & Reporting" and "Experimentation Governance".

The Perfect Water case study names the client and the method, "Key pages were audited to identify friction points, followed by hypothesis-led A/B testing across product and landing pages using VWO", and reports "38% improvement in conversions achieved following optimized layout and messaging" and "18% reduction in form abandonment". A second case names XLfeet and reports "27% increase in Add to Cart following validated PDP fit content restructuring" and "19% lift in checkout completion", with "Controlled A/B tests were conducted" and "Only validated improvements with measurable conversion gains were deployed".

Two things a buyer should press on. The headline figure on the Perfect Water case is "11,179% lift in high-impact conversion over the baseline experience (VWO A/B Testing)", published with no baseline value, no sample size, no duration and no confidence figure, and a lift of that magnitude on a purchase-adjacent metric is unusual enough to need its arithmetic shown. And the XLfeet page opens "XLfeet is a niche U.S. retailer specializing in large-size footwear" while a pull-quote on the same page describes "A footwear brand from Europe", with a returns figure that appears as 32 percent in one place and 41 percent in another, so a reader cannot reconcile the two from what is published. No significance level, no sample size, no test duration, no win rate and no traffic minimum were found on the pages read, and none of the three CRO cases names its platform.

9

Gene Commerce

A Magento or Adobe Commerce merchant whose conversion problem is specifically inside the payment step, and who values a published operating limit on the fix more than a test report26 of 100

The strongest payment-step evidence in the pool, and it is not a test. The 3DS case study names the client and the problem, "Choice Shops operate a number of Magento sites" whose healthcare customers "found the 3DS authentication process confusing or challenging", then the result with its window: "In the last two months, we noticed a 77% reduction in 3DS rejections after implementing the exemption. This translated to an increase of around 300 orders per week, or a 4% increase in sales." That is a before-and-after over two months, which the page states plainly, so it is scored as a period comparison rather than as a test.

It publishes something rarer than a lift: the risk it took on. "It's important to note that requesting a 3DS exemption comes with the shift of liability for any fraudulent transactions back to the merchant. To mitigate this risk, we set a cap of £100 for exemption requests." Two further checkout results appear with the client withheld, "a 3% lift in conversion by deleting just two fields" and "20% of failed submissions resolved at the first attempt simply by improving error handling", alongside a fully worked uplift model on 100,000 sessions a month and a published engineering threshold, "Keep the checkout route under 150KB of JavaScript".

Nothing behind it is a published testing practice. No hypothesis format, no primary metric, no one-variable rule, no duration or sample-size rule, no prioritisation model and no statistical statement of any kind were found on the pages read, and the sitemap contains no CRO or experimentation service page. Adobe Target is covered at length in an Adobe Target article and the advice on Magento checkouts is to "A/B test different optimisation strategies and listen to what your data tells you", which is the right instruction with no method attached. Worth knowing: three of the insight posts carry a joke byline rather than a real author, so those are not scored as named people.

10

Bemeir

A Magento store that has outgrown visual-overlay testing and needs the experiment routing built into the platform, and whose buyer is willing to take the delivery record on trust25 of 100

Publishes the best architectural thinking about testing on Magento and not one test result. Customization flexibility and conversion optimization opens on the mistake it sees most: "Most conversion optimizers think of A/B testing tools as a layer you add on top of your platform. Install Optimizely or VWO, paste a snippet, start testing. That works for simple visual tests, but it breaks down when you need to test structural changes to the buying experience." It then publishes a four-level table, from "Visual overlay... Client-side JS snippet" through "Component-level... Theme-level branching with feature flags" and "Flow-level... Platform API integration with server-side routing" to "Architectural... Backend customization with data layer segmentation", and states its own build practice, "feature flagging at the theme level, server-side experiment routing for checkout flow tests".

The tool review adds the platform-specific caveat that matters and the retirement nobody else names. On tooling for a customised storefront: "For heavily customized Magento or Shopware storefronts with complex JavaScript-driven interactions, the visual editor sometimes produces unreliable variations. In those cases, developer-implemented variations using VWO's code editor are more reliable." And on the tool half this field still recommends, it is the only agency here to state the fact: "Google Optimize was retired in September 2023, leaving a gap in the free A/B testing market." It also publishes a numeric revenue floor with the alternative below it, behaviour analytics under roughly five million dollars and "adding VWO for A/B testing" above it, plus test-volume bands by tool.

The scoring consequence is blunt. No concluded test, no client with a conversion figure, no hypothesis, no primary metric, no duration or sample-size rule, no significance statement and no win rate were found on the pages read, and both a CRO service page and a services page return 404, so the testing content lives entirely in articles. One published passage reads like a record and is not: a contrast between a "flexible team" that "has run fifty tests, found eight winners" and a rigid one is an illustrative scenario in context, and it is not scored as Bemeir's own result. The Optimizely partner page still carries a 2022 copyright line and describes the product under its pre-rebrand name.

4 Which one fits

Pick by situation, not by ranking

If this is youShortlistWhy
You have a Magento or Adobe Commerce store doing six figures of sessions a month and you want a testing programme that produces a decision you can defend to a boardscandiweb, then Blue Acorn iCiThis is the one situation where the ranking and the recommendation agree. scandiweb is the only agency here publishing named Magento tests with the primary metric separated from the secondaries, a sample-size formula, a duration rule and a guardrail period after launch, and it ships the winner into the store's own codebase. Blue Acorn iCi is the alternative if the estate is Adobe Commerce plus Adobe Experience Manager, because it publishes a 90 percent significance threshold as its own standard and a rule for what happens when nothing wins. Ask scandiweb which of the two conversion-per-variant figures on its framework article applies to your baseline, and ask Blue Acorn iCi for a confidence figure on the SeaWorld lift.
You want to read one complete, honest test report before you talk to anybody, so you know what a deliverable looks likeMageCloudMageCloud's Thermos write-up is the only published report in this field that carries the hypothesis, the dates, both sample sizes, the baseline and variant conversion rates, a confidence figure per metric, and the metrics that failed to reach significance. Read it as a template for what to demand, whoever you hire. Then ask MageCloud two things it does not publish: which testing tool ran that test, and what platform the store was on.
You already license Adobe Target, or your Adobe rep is pushing it, and your last testing programme quietly diedTom&Co, then Gene CommerceTom&Co is the only agency here that names why programmes die and what to check first, auditing an inherited instance for flicker and delivery problems before arguing about the backlog, and it publishes the traffic guidance that tells you whether classical A/B testing is the right purchase at all. Gene Commerce is the second read if the problem is concentrated in payments. Neither publishes a concluded test with a confidence figure, so make that the first thing you ask for.
You are choosing between a Luma checkout, Hyva Checkout and a React checkout on Magento 2, and you do not want to decide from a feature tableHatimeriaHatimeria is the only agency in this field that has published running all three against each other on one live store with random assignment. Treat the headline 67.5 percent as what the page says it is, a year-on-year comparison after the winner was rolled out, and ask for the in-test difference between the three arms, which is not published. Ask also how long it waits before reading a result, because its own tutorial says the results are visible immediately.
Your store does under about 20,000 sessions a month and someone is selling you a testing retainerTom&Co, then scandiwebRead the sessions section on this page first, then ask both to run the arithmetic on your numbers in writing. Tom&Co publishes the clearest statement that it will tell you when a test cannot reach significance and what to do instead. scandiweb publishes the same instruction, widen the detectable effect or move the test up the funnel, plus a study of an engagement at 400 to 500 daily visitors where testing could not conclude. Any agency that takes the retainer without doing that arithmetic is selling you months of inconclusive dashboards.
Your problem is structural, a different checkout flow or a different information hierarchy, not a button or a headlineBemeir, then scandiwebBemeir publishes the clearest account of why a JavaScript overlay cannot answer a structural question and what has to be built instead, level by level, up to server-side experiment routing. scandiweb publishes the only worked example of actually doing it on Magento 2, switching layouts server side through the platform's own layout event rather than rewriting the page in the browser. Bemeir publishes no test results, so if the deciding factor is evidence rather than architecture, the order reverses.
You want the discipline built into your own team rather than bought as a service, and you have a designer and a developer alreadyVervaunt, then Absolute DesignVervaunt publishes a hypothesis format with a fully worked example and names its ideation libraries, and one of its two service tiers is explicitly about embedding a test-and-learn culture rather than running the tests for you. Absolute Design publishes the prioritisation matrix and the one-variable rule in plain terms. Neither publishes a confidence figure against a test of its own, so use them for the method and hold the evidence question separately.

5 Evidence

The tests behind the entries, printed as published

ClientWhat was doneResultSource
ByggmaxHero banner CTA colour, grey control against green variant, on a Magento (Adobe Commerce) estate. scandiwebPrimary metric hero banner click-through rate; secondaries PLP view rate and eCommerce conversion rate; analysed for all users and split by mobile and desktop. "The variant showed an increase in PLP view rate by 0.7% and a significant uptick with more than a 15% increase in revenue and an 11% increase in purchases. The chance to beat control with the variant in terms of this KPI equals 73%." Note the 73 percent is below the 95 percent default the same publisher sets elsewhere.Source
ByggmaxCart icon, white control against green variant. scandiwebPrimary metric rate of sessions with checkout; secondary eCommerce conversion rate per user; device segment analysis on both. "The variant showed a significant increase in cart icon clicks by more than 18%. The chance to beat control for the green counter in terms of this KPI was equal to 92%." The published reading of the result is cautious: the decision further down the funnel "could have been influenced by factors such as the mini-cart popup design".Source
ByggmaxAdd-to-cart CTA colour on product listing pages, grey control against green variant. scandiwebPrimary metric PDP view rate; secondaries add-to-cart rate from PLP and eCommerce conversion rate. "The variant showed an increase in add-to-cart from PLPs by 5% and a higher than 3% improvement in average order value (AOV), with the effect for mobile devices being more pronounced." No probability figure is published for this third test, which is the gap a reader should notice across the set.Source
ThermosA free-delivery progress bar added to the mini cart, against the original cart. MageCloud"we ran a controlled A/B test from 23 June to 4 July 2026", 12,792 visitors on the control and 13,058 on the variation, 25,850 in total. Primary metric completed checkouts: "1.74% to 2.37% conversion rate, 99.75% confidence", a 35.74 percent increase. Secondaries: revenue per visitor £0.50 to £0.83 at 99.36 percent confidence, products ordered per visitor at 97.43 percent, checkout starts at 95.36 percent. "Later checkout steps (contact and shipping details) showed directional improvements that did not reach significance on their own."Source
SeaWorld San DiegoA three-arm multivariate test on the ticket medallion: control, a redesigned ticket drawer with filtering, and a direct route to the ticket page. Blue Acorn iCi"Each customer had an equal chance (33 percent chance each) of being in the Control, Variation 1, or Variation 2 group." Mobile: "a 4.7% lift in Conversion Rate, a 5.92% lift in RPV, and a 1.11% lift in Start Checkout rate". Desktop: "a 2.74% lift in Conversion Rate, a 3.83% lift in RPV". The losing arm is published: "Users in Variation 2 who were directed immediately to the Tickets landing page did not perform as well as those in Variation 1 or the Control." Graduated to full traffic "once the metrics reached statistical significance". No confidence figure is attached to the lifts.Source
Bedre NaetterThree Magento 2 checkouts run simultaneously against each other: React Checkout, Hyva Checkout based on Magewire, and Luma Checkout. Hatimeria"We then randomly assigned customers to one of three payment options, as if we were flipping a coin to decide which door they would go through." Monitored "for several weeks" on speed and usability, measured with "data from New Relic". Hyva Checkout won on speed and ease of use. The published number is a post-rollout comparison, not the in-test delta: "Between June and August 2024, Bedre Naetter achieved a 67.5% increase in checkout conversion rate compared to the same period in 2023, representing a 17.76 percentage point difference in finalized purchases."Source
Splash AboutA CRO programme of A/B tests following a workshop, hypothesis generation and a prioritisation matrix, on a store the agency had moved to Shopify Plus from Magento. Absolute Design"The results of the tests have achieved a revenue increase of 4%, conversion rate increase of 15.44% and a revenue per customer increase of 22.55%." The winners were promoted: "As these tests have both shown positive results, they have been promoted to the main variant." No confidence figure, sample size or duration is published. Included here because the design is published in full, and flagged because the platform is not Magento.Source
A major wine distributor in the US market, client not namedTwenty-three A/B tests on Adobe Target across five websites, at eight to ten tests a month. scandiwebThe stopping rule is published as three conditions that must all be met: "A duration of at least two weeks / Over 300 conversions for each primary segment of the variant / Statistically significant results." Analysis method: "manual and segmented analysis, looking at experiment performance split by device types and using Frequentist and Bayesian statistical methods." Outcome across the programme: "23 (and counting) A/B tests using Adobe Target, with a 70% winning test ratio." Individual tests are published with their variants and per-site numbers, including a free-gift threshold in the cart at "+25% average order value (website A)" and "+31% average order value (website B)".Source
CerveraAn ongoing A/B testing programme on a Swedish retailer with around 70 physical stores, including a platform decision settled by test. scandiweb"78% win-rate for A/B testing program". The platform decision is published with three metric movements: "A/B testing supported our decision to use Nosto's eCommerce personalization platform / 14.72% in product list clicks / 29.18% in add-to-cart rates / 36.53% in unique purchases". The client's own digital growth manager describes the arrangement as assisting "with our own A/B-testing", so read the win rate as a joint programme rather than a sole agency record.Source
Choice ShopsA 3DS exemption implemented on Magento sites whose healthcare customers were failing authentication. Gene Commerce"In the last two months, we noticed a 77% reduction in 3DS rejections after implementing the exemption. This translated to an increase of around 300 orders per week, or a 4% increase in sales." Published with the risk it created and the limit set on it: "requesting a 3DS exemption comes with the shift of liability for any fraudulent transactions back to the merchant. To mitigate this risk, we set a cap of £100 for exemption requests." A two-month before-and-after, not a split test, which the page states.Source
The Perfect WaterHypothesis-led A/B testing on product and landing pages using VWO. Krish Technolabs"38% improvement in conversions achieved following optimized layout and messaging", "25% improvement in page load performance", "18% reduction in form abandonment after streamlining inputs and checkout flow". The headline figure, "11,179% lift in high-impact conversion over the baseline experience (VWO A/B Testing)", is published with no baseline value, no sample size, no duration and no confidence figure, and a buyer should ask for all four before quoting it.Source
An eCommerce store selling bicycle covers, client not namedA homepage value-proposition and category-block test. scandiweb"Across all segments, we reached the statistical significance in the Value proposition CTA clicks for the Variation. Overall variation reached a 35% increase in CTA clicks compared to the original." And on the second element: "Overall variation reached a 32% increase in clicks compared to the original", with "the conversion rate improved by 89% for users who saw this block compared to those who didn't". Significance is asserted across segments without a figure, which is the weaker half of the same publisher's record.Source
A B2B store at 400 to 500 daily visitors, client not namedA testing programme that could not reach a conclusion, published as such. scandiweb"It meant that in order to achieve conclusive test results we needed a lot more daily visitors or had to leave tests on for months which was not good for the specific case. We discovered that we don't have the needed user amount to have significant test results / accurate statistical conclusions. And in order to test very notable changes e.g. whole menu or page layout, we had to have significantly more users." The same article works the arithmetic for a store at 10,000 monthly visitors and a 5 percent conversion rate. This is the only published account in the field of an engagement where the honest answer was that the traffic was not there.Source
Nuclear BlastA complete checkout redesign requiring server-side variant assignment. scandiweb"scandiweb's CRO and UX specialists wanted to do a complete checkout redesign for our client Nuclear Blast and involved the Analytics team in the process due to the rather complex A/B testing it requires." The mechanism is published: the experiment tool assigned the audience, and the cookie it set was read server side so that the server, not the browser, decided which version to build. Included as mechanics evidence rather than as a result: no lift figure is published for this engagement.Source

6 In detail

How many sessions a Magento store needs before a test can conclude anything

This is the question that decides whether a testing retainer is a good purchase or an expensive way to generate dashboards, and it has an arithmetic answer that nobody selling the retainer leads with. For a test comparing two conversion rates, the sample size per variation is (z + z)squared times the two variances, divided by the square of the difference you want to detect. Written out: n per variation equals (z at 1 minus alpha over 2, plus z at 1 minus beta) squared, times p1 times (1 minus p1) plus p2 times (1 minus p2), all over (p2 minus p1) squared. At the conventional settings, 95 percent two-sided significance and 80 percent power, that bracket is 7.85, and doubling it for two arms gives the rule of thumb published by Evan Miller as n = 16 sigma squared over delta squared, where for a conversion rate the variance is p times (1 minus p). The full two-proportion form, with the textbook citation attached, is published by Sealed Envelope, and the general relation appears in the NIST and SEMATECH e-Handbook of Statistical Methods. At 90 percent power the constant becomes 21 rather than 16.

Now put a Magento store's numbers into it. Take a 2 percent order conversion rate, 95 percent two-sided, 80 percent power, an even split and every session eligible. To detect a 10 percent relative lift, taking the rate from 2.0 to 2.2 percent, the test needs 80,679 sessions per variation, 161,357 in total, which is about 1,614 orders per variation. To detect a 5 percent relative lift it needs 315,203 per variation, 630,406 in total. To detect a 20 percent relative lift, 21,106 per variation, 42,211 in total. Halving the effect you want to detect roughly quadruples the sample, which Convert states as a rule: "MDE is usually the largest statistical lever because halving it roughly quadruples required sample." The baseline matters nearly as much: for the same 10 percent relative lift, a 1 percent baseline needs 326,184 sessions in total and a 3 percent baseline needs 106,415.

Turn that into months, at a 2 percent baseline and a 10 percent relative lift. A store at 500,000 sessions a month concludes in under a fortnight. At 200,000 it takes 0.8 of a month. At 100,000, 1.6 months. At 50,000, 3.2 months. At 30,000, 5.4 months. At 20,000, 8.1 months. At 10,000 sessions a month, 16.1 months, which is not a test, it is a year and a bit of drift during which the store, the catalogue, the traffic mix and the season all change underneath the experiment. Turned round the other way, a store doing 30,000 sessions a month that runs a test for 28 days collects 27,632 sessions, 13,816 per variation, and at 80 percent power that detects nothing smaller than a 25 percent relative lift. Very few page changes produce a 25 percent relative lift in order conversion.

One thing the arithmetic cannot tell you, and it decides everything above: the real duration is the longer of the calculated sample and whole business cycles, never whichever arrives first. That, and where the baseline comes from, are in the next section.

Check 1

The three inputs the arithmetic needs, and the floors the tools enforce

Every number in the section above depends on three inputs, and only one of them is a choice. The baseline conversion rate is measured, not chosen. The minimum detectable effect is the choice, and it is the expensive one. Statistical significance and power are conventions, 95 percent and 80 percent, and lowering them to make a test fit is the one move that turns a test back into an opinion.

Two vendor floors are worth holding beside the arithmetic, because they are what the tools themselves enforce, and because they are frequently mistaken for the answer. Optimizely will not call a binary metric at all below a stated minimum: "Binary metrics - Require at least 100 visitors or sessions and 25 conversions in both the variation and the baseline before Optimizely declares a winner." Adobe Target's automated allocation holds every experience at an even split until "each experience in the activity has a minimum of 1,000 visitors and 50 conversions". Neither is a sample size for a decision. Both are the point below which the software refuses to guess, and they sit one to two orders of magnitude below what the formula asks for.

Then there is calendar time, which is a separate constraint from sample size and the one most often skipped. Optimizely states it plainly: "You should run tests for a minimum of one business cycle (seven days) to ensure all kinds of user behavior are accounted for." VWO says the same: "It is recommended to run your campaigns for at least 7 days to capture natural fluctuations in metric performance across weekdays and weekends." Adobe Target rounds up rather than down: "it is recommended that the required time always be rounded up to the nearest whole week, so day-of-week effects are avoided". Adobe also prices the alternative metric, which matters if you would rather read the result on money than on rate: "using RPV as a metric requires 20-30% longer to achieve the same level of statistical confidence for the same level of measured lift."

One honest caveat on the baseline, because it is the input everything else rests on and there is no authoritative Magento-specific figure to use. The closest published series is IRP Commerce's monthly Ecommerce Market Data, which reported "The average conversion rate in the Ecommerce Market increased by 20.14% from 1.85% to 2.23% in August 2026 compared to August 2025", with sector rates from 0.57 percent for baby and child up to 6.87 percent for toys and games. It is platform-agnostic, UK-weighted, and it does not publish the size of its panel. Littledata's Shopify benchmark states its sample, "Littledata benchmarked 2,800 Shopify sites in 2023", and found a 1.4 percent average with "more than 4.7% would put you in the best 10%", but that is Shopify. Use your own store's rate, measured on the template you intend to test, and notice that both Optimizely and Adobe Target build their illustrative examples on a 5 percent baseline, roughly double the market rate either series reports.

Check 2

Where in the funnel you test changes the arithmetic more than the change itself does

The same store, the same power, the same significance level: a test on a step with a high baseline rate needs a fraction of the traffic a sitewide order-rate test needs, because the variance term p times (1 minus p) peaks at a 50 percent rate while the squared difference in the denominator grows much faster. Work it through. A 5 percent relative lift on a 45 percent checkout completion rate, taking it to 47.25 percent, needs 7,701 sessions per variation, 15,403 checkout entrants in total. A 10 percent relative lift on the same step needs 1,928 per variation, 3,856 in total. A 10 percent relative lift on an 8 percent add-to-cart rate needs 18,869 per variation. Against 80,679 per variation for a 10 percent lift on a 2 percent order rate, the checkout-step test is roughly forty times cheaper in traffic.

That is why the strongest published limit in this field is a funnel question rather than a traffic question. scandiweb's checkout page frames it as two numbers a merchant can read off their own analytics, "how many sessions reach your checkout, and how many finish", and says what follows if both are weak: "a full-funnel testing program comes first". Its framework article gives the same instruction as an escape route when the sample will not arrive, "move the test up the funnel where volume is higher". Convert publishes the trap on the other side: "Use traffic for the exact tested flow; site-wide traffic makes the duration look artificially short." A test on the checkout must be sized on checkout entrants, not on sessions.

The cost of moving up the funnel is that you stop measuring money. A test read on add-to-cart rate can win while orders fall, which is why the discipline worth buying reads the result on revenue per visitor and treats the upper-funnel metric as a diagnostic. scandiweb's framework article makes the trade explicit, "a 2% lift at checkout converts to more revenue than a 10% lift at the top of the funnel", and its guardrail rule covers the rest: monitor the primary metric for two full business cycles after launch, because "Real wins hold their lift within the confidence interval, false positives revert". Ask any agency which metric it will declare on before the test is built, and get the answer in writing.

Check 3

Why Magento breaks client-side testing, and what the ranked agencies do about it

A Magento storefront serves most of its pages from cache. Adobe's own documentation is explicit that "Adobe Commerce and Magento Open Source use full-page caching on the server to display category, product, and CMS pages quickly" and that "with full-page caching enabled, a fully generated page can be read directly from the cache", and it recommends Varnish in front of that: "Varnish sits in front of the web server and proxies these requests to the web server... Any subsequent requests for those assets are fulfilled by Varnish". See Adobe on cache management and Adobe on configuring Varnish. The consequence for experimentation is structural: the HTML a client-side variant script has to rewrite has already been rendered and cached before the script runs.

That produces the failure mode Adobe names itself. In Adobe's flicker documentation: "Flicker, also called FOOC (Flash of Original Content), is when an original content is briefly displayed before the alternative appears during testing/personalization." The standard remedy is to hide the page until the variant resolves, and Adobe publishes its cost without flinching: setting the pre-hiding style to hide the whole body "has the downside of leading to worse page rendering performance reported by tools like Lighthouse, Web Page Tests, etc.", with a timer that "by default removes the snippet after 3000 milliseconds". So the anti-flicker fix a merchant buys to protect a test is also a deliberate delay on the page the test is measuring.

Only two of the ten ranked agencies publish anything about this, from opposite ends. Tom&Co names the failure and the remedy stack, "pre-hiding, async loading tuned properly, server-side delivery where it matters", and the headless case, "On a decoupled storefront, client-side delivery gets awkward. Server-side and hybrid delivery are usually the right answer". scandiweb publishes the implementation, switching layouts on the server through Magento's own layout event and layout file-handles so the variant is built rather than rewritten. Bemeir publishes the architecture in between, four levels from a client-side snippet up to "Platform API integration with server-side routing", and the platform-specific tooling caveat, that on "heavily customized Magento or Shopware storefronts with complex JavaScript-driven interactions, the visual editor sometimes produces unreliable variations".

Worth knowing before a vendor conversation: Adobe publishes no native A/B testing capability inside the Adobe Commerce admin. Experimentation is delivered either by Adobe Target, a separately licensed product, or by the experimentation plugin on the Edge Delivery Services storefront, which Adobe describes as "The optional AEM experimentation integration", with the implementation reference at Adobe's own repository. If your store runs the classic PHP storefront, a third-party tool on a cached page is the default, and the flicker question above is your problem to solve rather than a detail. None of the ten ranked agencies publishes anything about full page cache or Varnish as a testing constraint, which leaves the biggest platform-specific gap in this lane unfilled by everybody.

Check 4

What a concluded test report has to contain, and what to ask for before you sign

Across ten agencies and several hundred pages, exactly one published report contains everything a reader needs to check a result: MageCloud's Thermos write-up. Use it as the specification. A concluded test report should carry the hypothesis as it was written before the test, the single change that was made, the primary metric declared in advance and the secondaries kept separate from it, the start and end dates, the number of visitors in each arm, the baseline and variant rates rather than only the percentage difference between them, a confidence or significance figure for the primary metric, and the metrics that moved in the right direction without reaching significance. If a report gives you a percentage uplift and nothing else, it is a claim, not a result.

Two of those items do more work than the rest. The first is the declared primary metric, because a test with no pre-registered metric cannot be wrong: whichever number moved becomes the headline afterwards. scandiweb's Byggmax write-ups are the clearest example in this field of the right shape, each test naming one primary and two secondaries before the numbers. The second is the baseline and variant rates in absolute terms. A 35.74 percent uplift means one thing at 1.74 to 2.37 percent and something entirely different at 0.1 to 0.14 percent, and only one of those is worth building.

Four questions get you most of the way, whoever you are talking to. What sample size did you calculate before the test, and from what baseline? What is the primary metric and who signs it off before the build? At what confidence do you declare, and what do you do with a result below it? And how many of your last ten tests won? The last question is the one the field answers least: of the ten agencies here, two publish a win rate for their own programme. scandiweb publishes 78 percent on one named engagement and 70 percent across a 23-test programme, and Bemeir's fifty-tests-eight-winners passage is an illustration rather than a record, which this page says rather than counting it.

One reading discipline to take into the meeting, from the same framework that publishes the formula: a headline win with a wide confidence interval is not the same purchase as a headline win with a tight one. "a 95%-confidence win with an interval of +2% to +22% is safe to launch, the same headline win with an interval of -1% to +25% is borderline and warrants a re-test." Nobody in this field publishes an interval against a client result, scandiweb included. Ask for one anyway: an agency that can produce it on request is running the tests properly, and an agency that cannot is reading a dashboard.

Check 5

What the AI assistants say about Magento test sample size, and where the arithmetic goes wrong

Three assistants were asked the buyer's version of this question on 29 September 2026, with web search on, through a single API so the prompts and models are on record. Perplexity, model sonar, was asked which agencies run A/B tests on Magento or Adobe Commerce and which publish real test results with numbers. It named ten agencies, cited twenty sources, and reached the same conclusion this page reaches from the other direction: "real test results with numbers are not clearly present in the snippets you provided for these agencies. The results mention capabilities, platforms, and process, but not concrete outcome metrics like uplift percentages, revenue gains, or conversion-rate deltas." ChatGPT, model gpt-5.4-mini, named five and correctly identified two that publish confidence figures. Gemini, model gemini-3.5-flash, named six. scandiweb was named by none of the three and cited by none of them, which is a finding about this page's own publisher and is recorded rather than hidden.

The Gemini answer went further and published the arithmetic, and the arithmetic is wrong in the direction that makes a testing retainer look affordable. It printed the correct formula, "n approximately equals 2 times (z at alpha over 2 plus z at beta) squared times p times (1 minus p), over delta squared", with the right critical values, 1.96 for 95 percent confidence and 0.842 for 80 percent power. It then published three worked scenarios whose figures do not follow from it. For a 2 percent baseline and a 5 percent relative lift it gave about 155,000 sessions per variant; the printed formula gives 315,203. For a 2 percent baseline and a 10 percent relative lift it gave about 38,000; the formula gives 80,679. Both figures match that same formula with its factor of two dropped, which returns 157,604 and 40,341. Its third scenario, a 15 percent relative lift on a 15 percent checkout baseline, came out about 29 percent high at 5,400 against 4,190.

The consequence is in the answer itself, not in the abstract. Gemini told the reader that the first scenario takes "31 days" at 10,000 sessions a day and that the second reaches significance in "15 to 16 days" at 5,000 a day. Run the same session rates against the correct totals and the answers are 63 days and 32 days. A merchant who plans a quarter on the first pair of numbers has budgeted for half the test they need, and will either stop early on a lead that has not settled or conclude that testing does not work. Every figure in this section was recomputed twice, exactly and with the pooled approximation, and the check script is named in the methodology.

The wider point for anyone researching this purchase through an assistant: the sample-size question is exactly the kind an assistant answers fluently and gets wrong, because the formula is easy to recall and easy to misapply, and because almost every source it can reach has an interest in the number being small. Ask for the formula, the power, the significance level and the baseline together, then check the multiplication yourself. And note what all three probes agree on without being asked: in this lane, the agencies an assistant names are not the agencies that publish the test evidence.

7 Methodology

How this was put together, and how to rebuild it

Ten agencies were read on their own websites on 29 September 2026 and scored out of 100 across six weighted criteria: a concluded test published with the number, the variant and the client at 30; test design published as a method at 20; statistical honesty at 16; experimentation on Magento specifically with the platform mechanics named at 14; named research inputs that generate the hypotheses at 12; and published limits at 8. The point ladder inside each criterion is published in the criteria table on this page, with the score each rung corresponds to, and the ranking is the score order with no adjustment. The weights were written down before a single agency was scored and were not changed afterwards.

Which agencies were eligible was decided by a rule fixed at the same time: an agency is ranked only if it publishes both a Magento or Adobe Commerce capability and something checkable about experimentation. That rule excluded the best-documented experimentation practices in eCommerce, and the exclusion is the most useful finding here, so it is named rather than hidden. Several specialist conversion agencies publish tighter gates than anything in this ranking, including one that requires at least 100 conversions per variation, 95 percent significance and three to four full weeks covering two business cycles before it will read a test, and another that publishes confidence figures, allocation, duration and sample size together on its case studies. A probe of the same path on each of their sites returned nothing for Magento, and one states outright that it tests exclusively on another platform. A buyer on Magento who wants that level of statistical discipline should read those sites for method and then ask the agencies on this page to match it.

Where a figure sits decides what it is worth. A number inside a testimonial is scored as a testimonial. A number in an unattributed stat band, a review aggregate, a return-on-investment calculator or an illustrative scenario is scored on the bottom rung however large it is, and one passage that reads like a win rate was identified as a hypothetical in context and not counted. A headline figure that turns out to be a before-and-after against a previous period is scored as a period comparison, not as a test, and three of the results on this page are labelled that way. Blogs, case studies and insight articles counted for every agency equally: sitemaps of between 129 and 1,227 URLs were swept for each, and the strongest evidence for several agencies, this page's publisher included, was found in articles rather than on service pages.

Pages were fetched as plain HTTP requests through Python urllib with a browser User-Agent and read as served text, so no figure here depends on JavaScript executing or sits inside an image. Where a page returned an empty shell, the fact is recorded as not checked rather than not published. Four sites could not be read from this machine at all and are excluded rather than scored: two returned a Cloudflare or regional block to every route tried, including an independent crawler, one serves an incomplete TLS certificate chain, and one was intercepted by a local network filter. One agency that ranks first in organic search for this term returned 403 to every plain client and was therefore not ranked. A gap is always written as not found on the pages read, never as absent from a site, because nobody has read every page of any of these sites.

The order was tested against every integer weighting that sums to 100, gives each criterion a floor of 5 points, and keeps the heaviest criterion heaviest: 2,779,797 weightings, and the order holds in all of them, with the narrowest winning margin at 3.15 points. It was then tested against a harder space, the same floor with the heaviest criterion free to be outweighed: 17,259,390 weightings, of which the leader holds 17,258,694, and the second-placed agency takes 696, every one of them putting statistical honesty at two thirds or more of the total. One weighting in that space produces an exact tie. Removing any single criterion entirely and rescaling the rest does not change the leader, and the narrowest such margin is 15.1 points. No weight, ladder or cell value was changed after a standing was checked.

Two disclosures that belong in the methodology rather than in an entry. First, two criteria are tied at the top rung and the page says so in both places: the heaviest criterion is taken by two agencies, whose statistical readings are not equally strong, 99.75 percent confidence against a 73 percent chance to beat control, and the Magento mechanics criterion is taken by two agencies working opposite halves of the same problem. Second, this page publishes a correction against its own publisher: one sample-size default on scandiweb.com is wrong by roughly a factor of fifteen against the formula printed beside it, it is printed in that entry rather than quietly omitted, and it is the reason the statistical-honesty criterion goes to another agency. The arithmetic in the sessions section and in the assistant check was computed twice, exactly and with the pooled approximation, with the scripts retained alongside the fact bank.

8 Questions

Questions a Magento merchant asks before shortlisting

Which Magento agency publishes real A/B test results with numbers?

Two of the ten agencies on this page publish a concluded test with the variant described, the client named, the lift measured and a statistical reading attached. scandiweb publishes three named Magento tests on Byggmax, each with a primary metric, secondaries and a chance-to-beat-control figure. MageCloud publishes one test on Thermos with both sample sizes, the dates, the baseline and variant rates and 99.75 percent confidence on the primary metric. Blue Acorn iCi publishes a three-arm test on SeaWorld with the allocation and the losing arm but no confidence figure. Three of the ten publish no conversion result from a test at all.

How many sessions does my Magento store need before an A/B test can conclude?

At a 2 percent order conversion rate, 95 percent significance and 80 percent power, detecting a 10 percent relative lift takes 80,679 sessions per variation, 161,357 in total. That is 1.6 months at 100,000 sessions a month, 5.4 months at 30,000, and 16.1 months at 10,000. Detecting a 5 percent relative lift takes 630,406 sessions in total. If your store is below roughly 50,000 sessions a month, sitewide order-rate testing is not a realistic purchase, and the honest alternatives are to test further down the funnel where the baseline rate is higher, to accept a larger detectable effect, or to spend the money on research and known fixes instead.

Does Magento have built-in A/B testing?

No. Adobe publishes no native A/B testing capability inside the Adobe Commerce admin. Experimentation comes from Adobe Target, which is a separately licensed product, from the experimentation plugin on the Edge Delivery Services storefront, which Adobe itself describes as optional, or from a third-party tool such as Optimizely, VWO, AB Tasty, Convert or GrowthBook loaded onto the storefront. If you are on the classic PHP storefront, a third-party tool on a cached page is the default route, and the flicker question becomes yours to solve.

Why does A/B testing behave differently on Magento than on a hosted platform?

Because most Magento pages are served from cache. Adobe's own documentation states that Adobe Commerce and Magento Open Source use full-page caching for category, product and CMS pages, and recommends Varnish in front of that. The page a client-side variant script has to rewrite has therefore already been rendered and delivered, which produces flicker, what Adobe calls a flash of original content, and the standard remedy is to hide the page until the variant resolves. Adobe publishes the cost of that remedy: hiding the whole body makes the page render measurably slower in Lighthouse and similar tools. The alternative is to build the variant on the server rather than rewrite it in the browser.

How long should I run a Magento A/B test?

The longer of two things, never whichever comes first. The calculated sample size for the effect you want to detect, and whole business cycles. Optimizely publishes a minimum of one business cycle, seven days, to account for weekday and weekend behaviour. VWO publishes the same seven-day minimum. Adobe Target goes further and rounds the required time up to the nearest whole week so that day-of-week effects are avoided. Stopping a test on the day it looks good is the single most common reason a published win does not repeat in production.

What is a statistically significant result in an eCommerce A/B test?

The eCommerce convention is 95 percent significance with 80 percent power, meaning a 5 percent chance of calling a winner that is not one and a 20 percent chance of missing a real effect. Two figures in this field sit below that and the page names both: one agency publishes a 90 percent threshold as its own standard, and one published win on a named Magento client is reported at a 73 percent chance to beat control, which is below the 95 percent default the same publisher sets elsewhere. Ask an agency what confidence it declares at, and what it does with a result below it.

How much traffic do the testing tools themselves require?

Far less than a decision needs, which is worth knowing so you do not mistake one for the other. Optimizely will not declare a winner on a binary metric until there are at least 100 visitors and 25 conversions in both the variation and the baseline. Adobe Target's automated allocation holds an even split until each experience has at least 1,000 visitors and 50 conversions. Those are the points below which the software refuses to guess. The sample size for a decision you would act on is one to two orders of magnitude larger.

Is it cheaper to test the checkout than the whole funnel?

Much cheaper in traffic, and that is the most useful arithmetic in this section. A 5 percent relative lift on a 45 percent checkout completion rate needs 15,403 checkout entrants in total. A 10 percent relative lift on the same step needs 3,856. Against 161,357 sessions for a 10 percent lift on a 2 percent order rate, a checkout-step test is roughly forty times cheaper. The catch is that a checkout test must be sized on checkout entrants rather than on sessions, and that upper-funnel metrics can win while orders fall, so read the result on revenue per visitor.

What should a test report contain before I accept it?

The hypothesis as written before the test, the single change made, the primary metric declared in advance with the secondaries kept separate, the start and end dates, the number of visitors in each arm, the baseline and variant rates in absolute terms rather than only the percentage difference, a confidence figure for the primary metric, and the metrics that moved without reaching significance. Exactly one report across the ten agencies on this page contains all of that. If you receive a percentage uplift and nothing else, you have a claim rather than a result.

Should I ask an agency for its win rate?

Yes, and expect most of the field not to have one published. Two of the ten agencies here publish a win rate for their own programme: 78 percent on one named engagement and 70 percent across a 23-test programme on the same site. One passage elsewhere that reads like a win rate is an illustrative scenario in context, not a record, and this page says so rather than counting it. Published industry averages sit far lower than any agency claim, which is why a win rate on its own is weak evidence and a win rate beside a published stopping rule is strong evidence.

What happens to a test that loses or comes out flat?

Ask, because it separates a programme from a sales exercise. The better published answers in this field are explicit. One agency publishes that a dedicated consultant reports what every test won or lost. One publishes that a result showing no difference is a real result that saves you building the thing. One publishes that where there is no clear winner, deployment is put on hold and the hypothesis refined. An agency whose published record contains only winners is either very lucky or not publishing its losses, and industry win rates make the first unlikely.

Can I A/B test Hyva Checkout against Luma Checkout on Magento 2?

One agency in this field has published doing exactly that, running React Checkout, Hyva Checkout based on Magewire, and Luma Checkout simultaneously on one live store with random assignment, monitored for several weeks with New Relic as the measurement source. Hyva Checkout won on speed and ease of use. The published headline of 67.5 percent is a year-on-year comparison after the winner was rolled out, not the difference measured inside the test, and the in-test deltas between the three arms are not published. Ask for those before treating the number as a test result.

Is a Magento CRO agency different from a Magento analytics agency?

Yes, and the difference is what each one is accountable for. An analytics practice is judged on whether the numbers are correct, which means reconciling what the platform reports against the order table. A CRO practice is judged on whether a change caused an effect, which means designing an experiment, powering it and reading it against a pre-declared metric. The two are sequential rather than interchangeable: a test read on instrumentation nobody has reconciled will produce a confident answer to the wrong question. Buy the reconciliation first if you are not sure yours is sound.

What does a Magento CRO programme cost?

None of the ten agencies on this page publishes a price for a testing programme, so the honest answer is that the cost is scoped after an audit. What several do publish is a floor of a different kind. One publishes that below roughly tens of thousands of sessions a month per test, personalisation and targeting rules deliver more than classical A/B testing. One publishes revenue bands for when to add a testing tool at all. One publishes a free audit as the entry point, and one publishes that its checkout project has an end date and does not become a retainer. Use those to work out whether you are buying the right thing before you ask what it costs.

Do AI assistants recommend the right Magento CRO agencies?

Not on the evidence of this lane. Three assistants were asked the buyer's question on 29 September 2026 with web search on, and between them named seventeen agencies. Two of the three named no agency that publishes a concluded test with a confidence figure, and one of the three said outright that concrete outcome metrics were not present for any of the agencies it had named. One assistant also published sample-size figures roughly half the size the formula it printed in the same answer produces, in the direction that makes testing look affordable. Check the arithmetic and check whether the named agency publishes a test.

Which agency should I pick if my store is too small to test?

Pick the one that tells you so in writing before taking the work. Two agencies on this page publish the instruction: one says it will size the roadmap to the traffic you have and say plainly when a test will never reach significance, and recommends personalisation and targeting rules instead below that line. The other publishes both the escape route, widen the detectable effect or move the test up the funnel, and a study of an engagement at 400 to 500 daily visitors where the honest conclusion was that the traffic was not there. Neither of those answers is a sale, which is why they are worth more than a testing proposal.

How can this ranking be checked?

Every score is the sum of six published cells, each cell is a rung named in the criteria table, and every claim carries the URL it was read from inside the sentence that makes it. Open the source link on any entry and the quoted sentence is on that page. The weighting is published, the ladders are published, and the sensitivity result names the narrowest margin found across 2,779,797 weightings and the harder 17,259,390. If a quoted sentence has changed since 29 September 2026, the page that carries it is the authority and this page is out of date.

Why are only ten agencies on this page?

Because ten is the number that publish both a Magento or Adobe Commerce capability and something checkable about experimentation. More than thirty candidates were read. Several were dropped for publishing no test evidence of any kind, several for having no Magento connection despite excellent experimentation practice, one because its domain now belongs to an unrelated company, one because it redirects to another firm, and four because they could not be read from this machine at all and are recorded as not checked rather than not published. The exclusions are described in the methodology, including the agency that ranks first in organic search for this term.