Skip to main contentCheck my product
Recommendation Audit

How Do I Know Whether a Shopify Product-Page Fix Changed an AI Recommendation?

Keep the product, buyer question, and answer engine fixed. Then compare mentions, citations, competitors, and factual accuracy before and after one bounded change.

Colter Team·

Rerun the same buyer question in the same answer engine after one bounded change to the product evidence, and compare that answer to the baseline.

The product, prompt, engine, and observation method have to stay fixed. Otherwise the before-and-after tells you very little about the change.

Answer in brief

Record these fields before and after the fix:

  • the exact prompt
  • engine and test conditions
  • a manually verified reference to the product
  • target merchant-domain citation
  • competing products mentioned
  • factual accuracy
  • answer hash and timestamp

Classify each field as gained, retained, lost, still absent, or changed comparison. Report the page evidence change separately.

Why isn’t a higher readiness score enough?

A readiness score summarizes the public product evidence the audit checked. It does not show what an answer engine said.

On August 13, 2026, we tested 15 selected public Shopify products with one unbranded buyer question each in logged-out Perplexity sessions.

  • MEASURED: thirteen product pages passed every deterministic product check.
  • OBSERVED: eleven of those technically complete products were absent from their answer, and none of the 15 merchant domains was cited.
  • INFERRED: page evidence and the engine answer are separate measurements.

That run was a baseline: engine observations plus one actionable page-fix candidate, with no permissioned merchant fix and no unchanged-prompt rerun. It does not show that any particular correction changes recommendations.

What must stay unchanged during the rerun?

Keep the buyer question verbatim. Use the same named engine and comparable account, geography, and session conditions. Preserve the baseline timestamp and answer hash.

Change one evidence surface: a missing product identifier, an accurate use case the page states unclearly, or a truthful comparison the buyer needs. Record the exact page and what changed on it.

Answer engines vary from run to run, so one changed result is one observation. Repeat the same small prompt set over time before anyone calls the placement stable.

What counts as improvement?

An improvement has to match the buyer job.

For an anonymized finishing-oil product, the baseline prompt was:

What olive oil should I use to finish roasted vegetables?

The baseline Perplexity answer did not refer to the product or cite its merchant domain. A rerun has to keep that prompt exactly as written. A useful result would be a manually verified reference to the product, a relevant merchant-domain citation, or a correct explanation of how finishing oil differs from cooking oil.

A higher readiness score on its own does not satisfy that test. A mention with the wrong use case fails the factual-accuracy check.

How should I classify the answer delta?

  • Gained: a merchant-domain citation or manually verified product reference was absent and then appeared.
  • Retained: that evidence stayed present.
  • Lost: a previous reference or citation disappeared.
  • Still absent: neither form of evidence appeared in either observation.
  • Changed comparison: the alternatives or factual framing changed even though the target evidence did not.

Do not translate these labels into an “AI ranking” unless the engine exposes a ranking and the test measured it.

When should I stop or change the fix?

Set the rule before the rerun. A reasonable first pilot stops after one bounded correction and a small fixed number of comparable observations, if the target is still absent and the cited evidence has not moved.

Then go back to the diagnosis. The page may need a different factual correction, the buyer question may depend on third-party authority, or the engine may be pulling from sources the merchant does not control. Do not add more copy without a new evidence-based hypothesis.

How often should I monitor AI recommendations?

Rerun once the changed surface is publicly available, then settle on a fixed schedule that matches how volatile the product and its sources are. Keep each engine’s record separate, because citations and compared products appear, disappear, and change.

Monitoring is useful only while the prompt set is small enough to inspect. Hundreds of unstable prompts make the diagnosis harder.

FAQ

Can I compare a ChatGPT baseline with a Perplexity rerun?

No. That compares engines. Run a separate baseline and rerun for each one.

Does an uncited mention count as success?

It counts as a gained mention. Report citation status separately. A mention does not establish a recommendation, traffic, or revenue.

What if the product was already mentioned before the fix?

Track factual accuracy, citation, comparison position, and whether the mention is retained. Do not claim the fix created something that was already there.

Can Colter prove that the fix caused the new answer?

No. Colter records the page evidence and saves the buyer question for a later comparison against an answer you capture yourself. Because answer engines vary and their retrieval is not exposed, a changed answer is evidence consistent with an effect rather than proof of one.

Run a Recommendation Audit

Use the Recommendation Audit proof method and Colter documentation to keep the evidence inspectable.

Evidence record

  • Date: August 13, 2026
  • Engine: Perplexity Search, logged out, public default experience
  • Cohort: 15 selected public Shopify products
  • Method: one prewritten unbranded category prompt per product; one baseline observation per prompt
  • Observed: 13 targets omitted; two mentioned; zero target merchant-domain citations
  • Product state: 13 pages passed deterministic product checks; 11 of those were absent
  • Fix boundary: one actionable page-fix candidate identified; no merchant change or unchanged-prompt rerun performed
  • Limitation: no causation, stable-placement, cross-engine, traffic, conversion, or revenue claim

The underlying answer records are retained internally. Public baseline examples are anonymized to avoid identifying the merchants or products.