Rerun the same buyer question in the same answer engine after one bounded product-evidence change. Compare the answer to the baseline rather than relying on a new readiness score.
The product, prompt, engine, and observation method must remain fixed. Otherwise the before-and-after result cannot tell you much about the change.
Answer in brief
Record these fields before and after the fix:
- exact prompt;
- engine and test conditions;
- target product mention;
- target merchant-domain citation;
- competing products mentioned;
- factual accuracy; and
- answer hash and timestamp.
Classify each field as gained, retained, lost, or still absent. Report the product-page evidence change separately.
Why isn’t a higher readiness score enough?
A readiness score measures the public product evidence selected by the audit. It does not show what an answer engine chose to say.
On August 13, 2026, we tested 15 selected public Shopify products with one unbranded buyer question each in logged-out Perplexity sessions.
- MEASURED: thirteen product pages passed every deterministic product check.
- OBSERVED: eleven of those technically complete products were absent from their answer, and none of the 15 merchant domains was cited.
- INFERRED: a page can become easier to retrieve without earning a mention or citation. The unchanged-prompt rerun is needed to observe that separate outcome.
The selected run was a baseline only. It did not include a permissioned merchant fix followed by a rerun, so it does not prove that any specific correction changes recommendations.
The field cohort yielded one actionable Fix candidate and baseline engine observations. No merchant change or unchanged-prompt rerun was performed. The method below describes the next test.
What must stay unchanged during the rerun?
Keep the buyer question verbatim. Use the same named engine and comparable account, geography, and session conditions. Preserve the baseline timestamp and answer hash.
Change one evidence surface. Examples include correcting a missing product identifier, clarifying an accurate use case, or adding a truthful comparison the buyer needs. Record the exact page and bytes changed.
Answer engines are variable, so one changed result is still one observation. Repeat the same small prompt set over time before calling the placement stable.
What counts as improvement?
An improvement must match the buyer job.
For Graza Drizzle, the baseline prompt was:
What olive oil should I use to finish roasted vegetables?
The baseline Perplexity answer did not mention Graza or cite its domain. A future rerun must keep this prompt unchanged. A useful observed improvement could be an accurate Drizzle mention, a relevant Graza citation, or a correct explanation that separates finishing oil from cooking oil.
A readiness-score increase alone would not satisfy that test. A mention with the wrong use case would also fail the factual-accuracy check.
How should I classify the answer delta?
Use direct labels:
- Gained: the target mention or merchant citation was absent and then appeared.
- Retained: the target remained present.
- Lost: a previous mention or citation disappeared.
- Still absent: the target did not appear in either observation.
- Changed comparison: the alternatives or factual framing changed even though target status did not.
Do not translate these labels into an “AI ranking” unless the engine exposes a ranking and the test actually measured it.
When should I stop or change the fix?
Set the rule before the rerun. A reasonable first pilot stops after one bounded correction and a small fixed number of comparable observations if the target remains absent and the cited evidence is unchanged.
At that point, inspect the diagnosis. The page may need a different factual correction, the buyer question may depend on third-party authority, or the engine may be retrieving sources outside the merchant’s control. Do not keep adding copy to the page without a new evidence-based hypothesis.
How often should I monitor AI recommendations?
Rerun after the changed surface is publicly available, then use a fixed schedule appropriate to the product and source volatility. Preserve each engine separately because citations and compared products can appear, disappear, or change.
Monitoring is useful only when the prompt set remains small enough to inspect. Hundreds of unstable prompts can produce a dashboard without a diagnosis.
FAQ
Can I compare a ChatGPT baseline with a Perplexity rerun?
No. That compares engines, not the effect of the fix. Establish a separate baseline and rerun for each engine.
Does an uncited mention count as success?
It is a gained mention. Report citation status separately. A mention does not prove recommendation, traffic, or revenue.
What if the product was already mentioned before the fix?
Track factual accuracy, citation, comparison position, and whether the mention is retained. Do not claim the fix created a result that already existed.
Can Colter prove that the fix caused the new answer?
Colter can preserve a bounded before-and-after observation. Because answer engines are variable and their internal retrieval is not exposed, the result is evidence consistent with an effect, not definitive causal proof.
Start a Recommendation Audit baseline
Use the Recommendation Audit proof method and Colter documentation to keep the evidence inspectable.
Evidence record
- Date: August 13, 2026
- Engine: Perplexity Search, logged out, public default experience
- Cohort: 15 selected public Shopify products
- Method: one prewritten unbranded category prompt per product; one baseline observation per prompt
- Observed: 13 targets omitted; two mentioned; zero target merchant-domain citations
- Fix boundary: one actionable Fix candidate identified; no merchant change or unchanged-prompt rerun performed
- Limitation: no causation, stable-placement, cross-engine, traffic, conversion, or revenue claim
Public baseline examples: Graza Drizzle, Azuna, and PerTronix wire set.