Rezolve Ai
Using Rezolve Ai

Experiments

Test your content against a control group, so you know what a change was worth, and keep a record of every result.

An experiment compares new content with content you left alone, over the same period, so you can tell what the content did from what the season or a promotion did. Open Experiments in the sidebar, under Grow.

What an experiment is

Ordinary reporting shows what happened after your content went live. It cannot say why, because seasons, promotions and site changes move the same numbers. An experiment adds a control group: products, or shoppers, deliberately left on your original copy. The difference between the two groups is what your content did.

There are two ways to split:

SplitHow it worksWhere it works
Split by shopperThe same product shows different copy to different shoppers, at the same moment.A Shopify or WordPress storefront.
Split by productSome products get the new content and some are left unchanged.Every project, including Amazon, Merchant Center and AI crawler data.

A "Feeds and files" project has no storefront to run code on, so the screen says Splitting by shopper is not available here and offers Split by product only.

What you need first

An experiment can only measure what a traffic source reports. Storefront visits come from one of two switches, and both start switched off:

PlatformWhere to turn it on
ShopifyIn the Rezolve Ai Shopify app, open Settings and find Storefront analytics, then choose Turn on storefront analytics. If Shopify needs your approval first, the button reads Enable storefront analytics. The app's Dashboard shows the same button in a banner while it is off.
WordPress and WooCommerceIn the plugin's Settings, under Visitor pageviews, tick Count visitor pageviews on this site and Send daily pageview and sales totals to my Rezolve Ai project.

Google Merchant Center, Amazon and the AI crawler watchlist each add history of their own. They are enough for a product split, but not for a shopper split.

The screen leads with the strongest figure it has and says which one that is:

  • Visits, when a storefront or Amazon is reporting.
  • Shopping clicks, when Google Merchant Center is reporting and no storefront or Amazon source is. A banner says so.
  • AI crawler visits, when none of those is reporting. A banner explains that this is what answer engines pulled from your pages, not shopper traffic.

Revenue columns appear only when a source reports sales. They are hidden otherwise, never shown as zero.

The Setup tab

Setup shows Where your numbers come from: one card for each source, marked Reporting or Partial, with Days reported, Pieces seen and the date Range. A source with no days has not reported yet, which is not the same as reporting zero.

When nothing is reporting, Setup lists each connected store with its state (Not measuring, Stopped reporting, Waiting for the first visit or Cannot tell) and the one step that fixes it. A project with no traffic source and no experiments opens on this tab.

The Experiments screen on the Setup tab of a project with no storefront connected, saying no traffic source is reporting yet

The Experiments screen

TabWhat it shows
ExperimentsYour experiments, grouped as In progress and Finished, and the New experiment button. Below the list, Which version won and Which content types pay back summarise results across experiments.
HistoryWhat testing has been worth so far, and every test with how it ended.
ResultsYour traffic over time, with a numbered marker on each day enriched content went live.
Pages and productsEvery page or product a source has reported on, with its trend and the number of content updates.
ForecastA ranked list of what to work on next.
SetupWhich traffic sources are reporting.

The 30 days, 90 days and 180 days control sets the period for Results, Pages and products and Setup.

Build a deck opens the New deck dialog with Experiments already chosen. See Decks.

Create an experiment

On the Experiments tab choose New experiment. The wizard has five steps: Type, Products, Versions, Goal and Launch.

  1. Type. Under How should the traffic be split?, choose Split by shopper or Split by product. A choice your project cannot use is marked Not available here, with the reason.
  2. Products. Choose the products the test runs over. See Choosing products.
  3. Versions. Choose which content is tested and how it is written. See What can be tested.
  4. Goal. Under What counts as better?, pick the measure the result is read on, then choose a length under Run for: 14, 28, 42, 56 or 90 days. See Goals.
  5. Launch. Give the experiment a name, check the Split, Products, Combinations and Window it will use, and choose Start experiment.

From the Goal step onwards, a panel checks whether the design can reach an answer. See The design check.

Step 1 of a new experiment, with Split by shopper unavailable without a storefront and Split by product selected

Choosing products

The Products step builds one selection, and every control on it adds to or narrows that selection:

  • Saved group. Load a product group you saved earlier. The group is copied in, so you can change the selection without changing the group.
  • Every product. Select the whole catalogue. Filtering is off while this is selected.
  • Filters. Narrow by Categories, Collections, Product type, Brand or Tags. Collections are for Shopify stores only.
  • Describe them instead. Describe the products in your own words and choose Find products. Untick any you do not want, then choose the button that starts Use these to keep the rest. This uses credits.
  • Review and adjust. See the products the selection matches, untick any to leave them out, and search your catalogue to add one by hand. You can hand-pick up to 500 products.

A summary under the filters says how many products match. For a product split it also says how many will be held back as the control group, or how many more products you need before a control group is big enough to compare against.

The products are fixed when you start the test. Products added to a category later do not join it.

Step 2 of a new experiment, with the Saved group menu, Every product and the product filters

Saved product groups

To reuse a selection, choose Save as group and name it. Manage groups opens Settings, Product Groups, where you can rename or delete a group.

A group is a saved way of choosing products, not a saved list. It re-reads your catalogue each time it is used, so products added to a category since you saved it are included. The count beside a group's name shows the day it was taken. Experiments already running keep the products they started with, even if you change or delete the group.

What can be tested

In a shopper split, the step asks What should we write more than one way? Only Product description can be swapped for one shopper and not another, so it is the only content on offer. Everything else is listed as Not in a shopper split, with the reason. Choose at least two versions: House style, which is the copy already on your page, and one or more other writing angles.

In a product split, the step asks What should this experiment measure? Each piece of content you include goes live on the products in the test and stays off the ones held back. Include one at a time to measure each on its own: content that goes live together cannot be told apart.

ContentShopper splitProduct split
Product descriptionYes, in any writing angleYes, in any writing angle
Meta title and descriptionNoYes, in any writing angle
IRL scenario cards, Highlights, Product Q&ANoYes, in house style
Product title, Schema markup, Image alt text, Product attributesNoYes, in house style
Google Shopping feed, Agentic commerce feed, Microsoft Shopping feedNoYes, in house style
Internal linksNoNo

Internal links cannot be tested either way, because links added to one product point shoppers at others, including products in the control group.

The writing angles are:

AngleWhat it does
House styleYour usual voice and structure, exactly as a normal run would write it.
Specification-ledOpens with materials, measurements and concrete facts before any persuasion.
Use-case-ledOpens with who it is for and the situation they are buying for.
Story-ledOpens with origin, craft or provenance, where the product data supports it.

A named angle is used on purpose. "Specification-led descriptions won" is a finding you can apply to the next thousand products. "Version B won" is not.

One experiment covers at most 4 pieces of content and at most 8 combinations. Every extra version splits your shoppers further, so each comparison takes longer to settle.

Goals

SplitGoalWhat it measures
ShopperReached checkoutThe share of shoppers who reached checkout.
ShopperRevenue per shopperWhat each shopper is worth. Needs more traffic than a rate does.
ProductOrdersOrders for the products in the test, compared with the products held back.
ProductRevenueSales value for the products in the test. Cannot be used when sales are in more than one currency.
ProductVisitsShopper sessions on the product page.
ProductShopping clicksClicks from Google Merchant Center listings.
ProductAI crawler pullsHow often answer engines fetched the page. Needs 21 days to settle.

A shopper split also offers Watch added to cart as a second check beside the main goal. A version can win on cart adds and lose at checkout, so it is never the main goal.

If your project has nothing reporting for the goal you pick, the experiment is refused before anything is written, and the message says what to connect.

The design check

Before you spend credits, the wizard says whether the design can reach an answer:

MessageWhat it means
This can be measuredYour traffic is enough to tell a real change from noise over the length you chose.
Checked when it launchesA product split is sized against the products you selected when you start it.
This will take longer than plannedThe design works, but not inside the length you set. You can still start it.
Too many versions for your trafficFewer versions would let the same test reach an answer.
Not enough shopper traffic to measure thisThere are too few shoppers behind these pages at any length. Split by product instead.
No shopper traffic reported yetYour storefront is not reporting visits. Turn on the switch described under What you need first.
No checkouts measured yetShoppers are counted, but there were no checkouts in the last four weeks to compare against.
This run is too small to measureA product split needs more products, or products with more measured activity.

For a shopper split the panel also draws how shoppers will be split and the smallest change you could spot at each length. When the design cannot work, Start experiment is unavailable.

What happens after you start

Starting an experiment writes the new versions and sends them to Review. Nothing is split until you approve a version and it goes live on your store, so the experiment begins as a Draft. If another content run is already in progress for the project, the experiment is not created and you are asked to wait for that run to finish.

The experiment's page shows five steps. A step that is waiting on you is shown in amber, with the reason underneath.

StepWhat it means
Set upThe products are enrolled and the new versions are being written. Nothing is shown to shoppers until you approve them.
Live on your storeWaiting for the new versions to be approved in Review and reach your store. In a shopper split, both versions must be live before anything is compared. Your store checks for a new test once a day, so a split usually starts within 24 hours of approval.
CollectingVisits are being counted. The caption shows the day, for example "Day 10 of 28".
Enough evidenceResults are computed once a night, and only once enough days have passed for a change to mean something.
DecidedThere is an answer to act on, or the test is closed.

An experiment moves through these statuses: Draft, Running, Paused, Control group released, Finished and Abandoned.

The buttons on the page depend on the status:

  • Start now (Draft). Available once at least one product has a second version live. You do not have to use it: the experiment starts by itself when the first version goes live.
  • Discard (Draft). No shopper has seen anything, and the content already written stays in your review queue.
  • Pause and Resume (Running, Paused). While paused, every shopper sees your original copy. Resuming continues the same experiment, and shoppers keep the version they were seeing.
  • Stop (Running, Paused). Closes the measurement window for good. Results already computed stay on the page. A stopped experiment cannot be resumed.

Reading the results

The banner at the top of an experiment's page says what to do today:

BannerWhat it meansButton
Waiting for a second version to go liveNothing is being split and no days are counting.None
Ready to startA second version is live on some products.Start now
Still collectingNo clear answer yet.None
This may not settle on its ownAt the current traffic, the answer may not become clear.None
A version's name followed by "won"The version beat your original copy by more than chance explains.A button to keep that version for everyone
Your original copy wonThe new version performed worse. That is a real finding.Finish and keep the original
The versions performed the sameEnough shoppers saw each version to rule out a difference worth acting on.Finish this test
This test cannot be trusted yetOne of the checks below is broken, so any difference could be the fault and not the content.Stop this test

Keeping a version finishes the test and makes that version the copy every shopper sees. Splitting stops at once, and the new wording reaches your store the next time approved content is applied.

Below the banner, the page shows:

  • How each version is doing, or Visits, new content against the control group for a product split.
  • When will I know? The range around the difference, narrowing as shoppers arrive. When it stops touching zero, the result can be read.
  • Where shoppers drop off and Is this being measured properly?
  • What each version is worth, when your store reports order values.
  • Can this be trusted? Five things that can be wrong while the numbers still look fine.
  • What each version was shown to: counted shoppers and checkouts. These are facts, so they show whether or not a result can be read yet.
  • Which way of writing won, when more than one piece of content was tested.
  • What it found: every combination compared, including the ones that did not win.

The five checks under Can this be trusted? are:

CheckWhat it looks at
SplitWhether each version received its share of shoppers.
Control groupWhether every control product is still unchanged. On a shopper split this check is called Products.
Versions liveWhether at least two versions are live, and whether your store is serving the split.
AttributionWhether every shopper counted could be matched to this test.
MeasurementWhich traffic sources are reporting inside the test's window.

Each result carries a label saying how much it can claim:

LabelWhat you can conclude
Measured against a control groupThe difference between the groups is the change your content caused.
No clear differenceAgainst an unchanged control group, this did not move the figure either way. That is a result, not a gap in the data.
Too early to readNot enough days have passed.
Not enough trafficToo little traffic to tell a real change from day-to-day noise.
Control group brokenSome control products were updated anyway, so the groups are no longer comparable.
Groups were already driftingThe groups were moving apart before the change. Treat the figure as directional.
Went live togetherThis content went live alongside other content on the same products, so its own effect cannot be separated. The result for the whole experiment still stands.
Too few went liveToo few products received this content at around the same time to compare.
Mixed currenciesSales were recorded in more than one currency. Measure orders instead.
No control group, so not a test resultThis shows what happened, not why.

Words used on these pages

Each experiment page ends with What these words mean:

  • Range. The band the true difference is very likely to sit inside. A range that still crosses zero means the version could be better or worse, and more shoppers are needed before you act on it.
  • Control group. The products, or the shoppers, deliberately left on your original copy. Without them a rise in sales could be the season, a promotion, or anything else that happened at the same time.
  • Even split. Each version should reach roughly the same number of shoppers. When one gets far fewer, something is filtering them, and the comparison is between two different audiences.
  • Conclusive. Enough shoppers have seen each version that the difference between them is bigger than the day-to-day noise.
  • Not measurable yet. Not the same as no difference. Too little has been counted to say anything either way, which is why no number is shown instead of a zero.

History

History keeps every test and what it found, so you can see what you have already learned before deciding what to test next.

What testing has been worth counts money only from tests that reached a conclusive answer with real order values behind them:

  • Banked so far: the yearly value of the decisions you have taken, with the cautious end of the range underneath.
  • Running tests, if the current estimate holds: shown separately and not counted in the banked figure.
  • Counts of Winners kept, Rollouts avoided, Measured the same and No clear answer.

A test where your original copy won counts as value too: not rolling out the weaker version is what the test was worth.

Every test you have run lists each experiment with its dates, how it ended, the measured change and its range. Search by name, or filter by outcome: New version won, Original won, No difference, Stopped, not measurable, Stopped early, Never started or Discarded. A test whose winner was rolled out is marked Applied.

Choose Add what you learned on any finished test to record a note. With more than one measured result, Everything you have measured, side by side draws them on one axis.

Forecast

Forecast ranks the content most worth enriching next. A chip says how much the ranking rests on:

ChipWhat it means
Based on your own resultsFigures come from measured results in your own completed runs. Expected revenue is shown.
Based on results elsewhereFigures come from measured results across other stores, because none of your own runs have finished yet. Treat the order as a steer and the sizes as rough.
Ranked by opportunityAn order only: the pages with the most traffic and the most room to improve. There are no figures, because nothing has been measured yet.

What to work on next lists each piece of content, what to add, and where figures are available, Extra visits, Typical uplift and Extra revenue. Select a row to see its projection over the forecast period, with the range you should not be surprised by. Show working opens the inputs behind a row.

A forecast is an expectation on average, not a promise. If the tab reads Nothing to forecast yet, it needs a few weeks of traffic first and fills in by itself.

Results and Pages and products

These two tabs describe what happened. They have no control group, so the Results chart and each piece's own page are marked No control group, so not a test result.

  • Results shows your headline figure, Content updates live and Days measured, then Traffic and content changes over time. Numbered markers are the days enriched content went live.
  • Pages and products lists each content piece with its figure, trend, number of updates and when it was last seen. Open a piece to see its own history and its figures by source.

Control groups and Review

In a product split, about a fifth of the products you chose are held back as the control group. Their new content is written with everyone else's, and you are charged once, but it must stay off your store while the experiment is open.

Review does not hold control products back for you, and it does not mark which products they are. If content for a control product goes live while the experiment is open, that product is dropped from the comparison, the Control group check reports it, and the result is labelled Control group broken with no figure.

A few details:

  • The hold begins when the experiment is created, not when it starts running. It also applies while the experiment is paused.
  • Only the content the experiment measures is held. Other content for the same products can go live as usual.
  • A control group is held for four weeks. After that the experiment is marked Control group released, and those products can be approved and pushed like any others.
  • A shopper split holds no products back. Every product is in the test, and the split happens between shoppers. Its new version is served to a share of shoppers and is not written over the product's own copy until you keep it for everyone.

You can also measure an ordinary optimization run. On the last step of Product Optimization, Run this as an experiment is ticked when the run is large enough to compare. It creates a product split with the same four week control group.

In the Shopify app and the WordPress plugin

Both have their own Experiments screen, where you can start a New experiment and follow the experiments in your project with the same progress captions.

Shopify. A shopper split needs the A/B test app embed turned on in your theme. In the theme editor it is listed as A/B test product description. Without it no shopper is ever assigned a version, and the experiment reports nothing. When it is off, the app says so on the experiment's page and points you to its Schema page, where you can turn it on. If your theme does not use the standard Dawn layout, set the embed's Description block selector to the block that holds your product description. A product split does not need the embed.

WordPress. In the plugin's Settings, under A/B test your product copy, tick Split visitors between versions on this site and Send results to my Rezolve Ai project. The split stores a cookie, so it only runs for visitors who have given consent through your cookie banner. With no consent plugin, nothing is stored and every visitor sees your normal copy. If your theme does not use the standard WooCommerce layout, set Description block selector to the block that holds your product description, or the test will run and change nothing.

See Shopify app and WordPress plugin.

Credits

Two things on this screen use credits: writing the new versions when you start an experiment, and Find products on the Products step. Your original copy is not written again. Checking a design before launch is free. See Credits.

On this page