Skip to content

AB Tasty vs Optimizely vs VWO in 2026: Three of Them Stopped Publishing Prices, and Shopify Made It Free

We read all four pricing pages on the same day. Only one publishes a number. Meanwhile Shopify shipped native A/B testing in June 2026. Here is what actually decides whether testing works for you, and it is not the tool.

The Sellarix team · 22 Jul 2026 · 13 min read

Here's the most expensive mistake in conversion optimisation, and it has nothing to do with which tool you pick. You run a test on traffic that could never have detected the effect you were looking for, get a result, believe it, and ship it.

You didn't learn anything. You bought a coin flip and gave it a name.

That's the old problem. There's a newer one, and I only noticed it while researching this piece. We opened the pricing pages for AB Tasty, Optimizely, VWO and Intelligems on the same day. Three of the four have stopped publishing a price. In the same year Shopify made basic A/B testing free and in-plan.

What we found: the pricing pages went dark

All four checked on 22 July 2026, reading only what each vendor publishes.

VendorPublishes a price?What the page actually says
OptimizelyNo"Every Optimizely plan is individually packaged. Tell us a bit about your digital needs, and we'll create a plan together."[1]
AB TastyNo"No fixed plans. Just a custom proposal built around your goals and scope." Priced on traffic volume, domains, modules and implementation scope[2]
VWONoGrowth, Pro and Enterprise tiers named. No figures, no stated pricing metric[3]
IntelligemsYesUnlimited at $2,314/mo list, $1,279/mo discounted, based on 10,000 orders a month, up to 200,000[4]

The VWO one is the change that matters. VWO's whole historical position in this market was being the one you could buy without a procurement process. Comparison articles still describe it that way. The pricing page no longer supports that description[3].

I'd read this as the category moving upmarket rather than as anyone behaving badly. Custom pricing genuinely fits enterprise experimentation, where scope varies wildly. But it does mean that every comparison table you'll find with VWO or Optimizely prices in it is quoting numbers the vendors no longer stand behind, and most of those tables are on sites selling an alternative.

The other half of the squeeze

On 5 June 2026 Shopify expanded Rollouts to schedule, publish and A/B test themes and checkout and customer account configurations, natively[5]. Theme testing is broadly available; the deeper checkout component testing is a Plus-tier capability[6].

So the market now looks like this. The cheap end got absorbed into the platform. The expensive end went quote-only. The middle, where most mid-market retailers actually live, got thinner.

Does this actually impact you? The traffic test comes first

Before you compare anything, work out whether you can run a test at all.

The arithmetic nobody wants to do

Sample size depends on your baseline conversion rate, the smallest effect you'd care about detecting, and your tolerance for false positives and false negatives. Evan Miller's calculator is free, has been the reference for over a decade, and takes two minutes[7].

Put your real numbers in. A 2% baseline conversion rate and a desire to detect a 10% relative improvement needs tens of thousands of visitors per variant. Most stores discover at this point that they cannot detect anything smaller than a large change, ever.

What that means practically

Monthly conversionsWhat you can honestly do
Under 200No A/B testing. Qualitative research, session recordings, talking to customers
200 to 1,000Big swings only. Whole page redesigns, price changes, offer changes. Long run times
1,000 to 5,000Real testing on meaningful changes. One test at a time
Over 5,000A genuine programme. Multiple concurrent tests, if you separate them properly

This is the table I'd want a founder to see before any pricing page. Buying an experimentation platform for a store doing 300 conversions a month is buying a microscope to look at the moon.

The four ways tests lie to you

The statistical literature on this is excellent and almost entirely ignored by the marketing side of the industry. Kohavi, Deng and Vermeer's KDD 2022 paper on common misunderstandings in online controlled experiments is the single most useful thing you can read on the subject, and it's explicit that some of the misleading intuitions are promoted by A/B testing vendors and agencies[8][9].

1. Peeking

You check the dashboard daily and stop when it goes green. This inflates your false positive rate substantially, because you've turned a fixed-horizon test into an unplanned sequential one. Johari, Pekelis and Walsh set out the problem and the fix, which is always-valid inference designed for continuous monitoring[10]. Evan Miller's older piece on the same problem is shorter and just as persuasive[11].

If your tool shows a live significance figure that updates all day, it is inviting you to do the wrong thing. Decide the sample size and the end date before you start, and write them down where somebody else can see them.

2. Testing the wrong metric

Conversion rate goes up. Revenue goes down. This happens constantly when the winning variant sells more of a cheaper thing. Pick an overall evaluation criterion that includes value, not only count, before you launch.

3. Segment mining after the fact

The test lost. You slice by mobile, then by new visitors, then by mobile new visitors from paid, and find a winner. You have found noise. Kohavi's paper covers this and it's the most common way a genuinely negative result gets shipped anyway[8].

4. Novelty and primacy effects

Returning customers react to change itself, not to the change's merits. A week-one result on a familiar audience often reverses. Run for full business cycles, and if you can, look at new visitors separately as a sanity check rather than as the headline.

Multiple tests at once

Running concurrent experiments is fine and Google published the canonical architecture for doing it safely with overlapping layers[12]. Doing it without that structure means your tests interact and you can't attribute anything.

What to do instead when you haven't got the traffic

Most stores reading this fail the sample-size test, and the honest answer isn't "test anyway with lower confidence". It's to use methods that don't need statistical power.

Session recordings and funnel drop-off

You don't need significance to see somebody fail to find the size selector four times. Ten recordings of people abandoning at the same step is a finding. Pair it with the specific fixes in the checkout piece.

Fix the things that are just broken

Page speed doesn't need an A/B test to justify. Neither does a search box that returns nothing for a plural. The page speed piece and the site search piece both cover work that's worth doing without proving it first.

Ask five customers

Unfashionable, free, and consistently more informative than a underpowered test. Five people who bought last week, on a call, describing what nearly stopped them.

The trap to avoid is running tests anyway and treating the results as knowledge. An underpowered programme is worse than no programme, because it manufactures confident wrong beliefs that then get defended in meetings.

Price testing is a different product, and it's the underrated one

Most experimentation tools test layout. Intelligems built its business on testing price, shipping thresholds and bundles, and reporting on profit per visitor rather than conversion rate[4][13].

The difference is bigger than it sounds. A layout test asks which button converts better. A price test asks whether selling at £48 instead of £45 made more money after the conversion drop. For a margin-constrained DTC brand the second question is worth more, and almost nobody asks it because the tooling grew up around layout.

The honest caveats

Price testing has consequences layout testing doesn't. Two customers see different prices for the same product, which is defensible when it's a genuine experiment with a defined end and much less defensible when it drifts into permanent personalised pricing. Free-shipping threshold tests interact directly with your shipping rate structure, so run them together or not at all.

And if the pricing decision is driven by a model rather than a random assignment, you've moved from experimentation into automated decision-making about individuals, which carries transparency obligations in the EU that most CRO teams have never looked at. Our EU AI Act compliance guide for ecommerce covers where the line sits.

If you are on Shopify

Start with Rollouts

Native, in-plan, and it covers the common case of testing a theme change[5]. A rollout runs either as a launch, which applies to everyone, or as an experiment, which splits traffic. Only the second is an A/B test, and it's worth being clear internally about which one you're running.

What it doesn't do

Server-side experimentation, feature flags, complex targeting rules and cross-device identity resolution are not in scope. If you need those, you're in the paid category and you should expect a sales process[1].

The performance trap

Client-side testing tools inject a script that must run before the page paints, or the visitor sees the original version flicker to the variant. The usual fix is a synchronous blocking script, which makes the page slower for every visitor including the control group. You're then measuring a change plus a slowdown, and attributing all of it to the change.

Native rollouts avoid this because the split happens before the response leaves Shopify. That's a genuine technical advantage and it barely gets mentioned.

Not on Shopify? The other platforms

WooCommerce

No native experimentation. You're running a client-side tool or building server-side splits yourself against the REST API[14]. WordPress caching is the specific hazard: a full-page cache will happily serve the same variant to everyone and your test will report a dead heat forever. Check your cache configuration before you trust a single result.

Magento and Adobe Commerce

Adobe will sell you Target, which is genuinely capable and generally needs a licence separate from Commerce itself, which is the recurring pattern with Adobe's stack. The platform's full-page cache raises the same variant-caching problem as WordPress, and multi-store and multi-website scope makes traffic allocation more complicated than it looks[15]. Price the third-party route in parallel.

BigCommerce

Third-party tools via script manager, clean APIs to build against[16]. Same client-side flicker considerations as everywhere else.

Headless

The best position for experimentation and the most work. You can split server-side at the edge with no flicker and no blocking script, and you own the assignment logic, which means you own getting it wrong. Write the assignment down as a spec, and be aware that headless stores are undercounted in every platform statistic you'll read[17].

The pre-flight checklist

Steal this. Every test, every time, before it goes live.

#QuestionAnswer must be
1What's the hypothesis?One sentence with a mechanism, not "we think this looks better"
2What's the primary metric?One metric. Includes value, not only count
3What's the minimum detectable effect?A number you'd act on[7]
4What's the required sample size?Calculated, written down, per variant[7]
5What's the end date?Fixed in advance. Full business cycles[10]
6What are the guardrail metrics?Revenue per visitor, refund rate, page speed
7Which segments will you look at?Named before launch, or none[8]
8Does the tool add latency?Measured, on both variants
9Is caching serving one variant?Verified in a private window
10What will you do if it loses?Decided now, while you're honest

Question ten catches more bad decisions than the other nine combined.

What agents change about experimentation

Something structural, and I don't think the CRO industry has noticed yet.

A/B testing measures human responses to visual and copy changes on a rendered page. When an AI assistant reads structured product data and returns a recommendation, there is no page, no layout and no button colour. Your variant never gets seen.

The emerging commerce protocols describe product, price, availability and fulfilment as structured fields[18][19]. What an agent evaluates is your data, not your design. Which means the experimentation surface for agent-mediated traffic is product data quality, price and delivery promise. Not layout.

How much does this matter right now? Not much. Kaiser and Schulze's peer-reviewed work across 973 sites and $20 billion of revenue puts ChatGPT under 0.2% of ecommerce traffic[20], while Shopify reports AI-referred orders up roughly 13 times year on year[21]. Small and growing fast.

The practical note is narrower and it applies today: if a meaningful share of your traffic isn't rendering your page, exclude it from your experiment population, or it'll dilute every result you get. The AI in ecommerce piece has the current traffic numbers and what to do about the data layer.

What to do this week

  • Run the sample size calculation on your real numbers. Before any tool conversation[7].
  • If you're on Shopify, open Rollouts. It's already there and it's already paid for[5].
  • Measure your current testing tool's latency on a real device. If it blocks render, it's biasing every test you've run.
  • Check whether your cache is serving one variant. Private window, both variants, ten minutes[14].
  • Write down the sample size and end date for your current live test. If you can't, stop it.
  • If you're renewing, ask for the price in writing including next year's tier, given three of the four vendors no longer publish anything[1][2][3].
  • Consider whether your real question is about price rather than layout[4].

The takeaway

Three of the four major experimentation vendors stopped publishing prices[1][2][3], one publishes $1,279 a month discounted[4], and Shopify made the entry-level version free in June 2026[5].

None of that changes the thing that decides whether testing works for you, which is whether you have the traffic and the discipline to stop a test when you said you would.

I called a test at day four once because it was up 12% and I wanted to tell someone. It finished flat. The change shipped anyway, because by then it was in the sprint, and I never worked out whether it hurt us or not. That's the actual cost of peeking. Not the false positive. The fact that you stop being able to tell.

What's the sample size on your current live test? If you don't know it, what are you going to do with the result?

Sources

  1. Optimizely, "Pricing". Accessed 22 July 2026. Vendor page; no published figures.
  2. AB Tasty, "Pricing". Accessed 22 July 2026. Vendor page; custom proposal only.
  3. VWO, "Pricing". Accessed 22 July 2026. Vendor page; tiers named, no figures.
  4. Intelligems, "Pricing". Accessed 22 July 2026. Unlimited at $2,314/mo list, $1,279/mo discounted, 10,000 orders/month base.
  5. Shopify, "Rollouts now supports online store and checkout and customer accounts", 5 June 2026. Shopify developer changelog.
  6. Shopify, "Pricing". Accessed 22 July 2026.
  7. Evan Miller, "Sample size calculator". Independent, free, no vendor interest.
  8. Ron Kohavi, Alex Deng and Lukas Vermeer, "A/B Testing Intuition Busters: Common Misunderstandings in Online Controlled Experiments". KDD 2022. Peer-reviewed.
  9. ACM Digital Library, "A/B Testing Intuition Busters", Proceedings of KDD 2022. DOI 10.1145/3534678.3539160.
  10. Ramesh Johari, Leo Pekelis and David J. Walsh, "Always Valid Inference: Bringing Sequential Analysis to A/B Testing". arXiv:1512.04922.
  11. Evan Miller, "How Not to Run an A/B Test". Independent.
  12. Google Research, "Overlapping Experiment Infrastructure: More, Better, Faster Experimentation". Peer-reviewed.
  13. Intelligems, Shopify App Store listing. Accessed 22 July 2026.
  14. WooCommerce, "REST API". Accessed 22 July 2026.
  15. Adobe, "General configuration". Adobe Commerce documentation. Accessed 22 July 2026.
  16. BigCommerce, "Orders API". Accessed 22 July 2026.
  17. HTTP Archive, "Web Almanac 2025: Ecommerce". Accessed 22 July 2026.
  18. Google, "Under the Hood: Universal Commerce Protocol (UCP)". Google Developers Blog.
  19. Agentic Commerce Protocol, specification repository. Maintained by OpenAI and Stripe.
  20. Maximilian Kaiser and Christian Schulze, "ChatGPT Referrals to E-Commerce Websites", Marketing Science. 973 sites, $20B revenue, 12 months. Peer-reviewed.
  21. Shopify, financial reports, Q1 2026. AI-referred orders up roughly 13x YoY.