An ecommerce agency can sell a bounded AI product recommendation audit built around evidence the client can inspect.
Start with one client product, one buyer question, and one answer engine. Record the public product evidence, the products and sources in the answer, one correction, and the unchanged rerun.
Answer in brief
A client-ready recommendation audit has four stages:
- Baseline: verify the public product facts available for retrieval.
- Fix: identify one material evidence gap and correct it with the client’s permission.
- Prove: record what a named engine says for the exact buyer question.
- Monitor: rerun the unchanged question and report the answer delta.
The agency should deliver the evidence behind each stage. A readiness score, screenshot, or claim that a product is “AI optimized” is not enough.
Which client product should an agency test first?
Choose a product with a clear buyer job and enough commercial value to justify the work. Avoid branded prompts because they only test whether an engine can repeat a name the user supplied.
Good prompts express a purchase constraint:
What braiser works on induction and can go in a 500 degree oven?
What loose-leaf black tea should I buy for a strong everyday breakfast brew?
What stock-look spark plug wire set fits a 1967 Volkswagen Beetle?
Each question can be tied to a specific product, factual requirements, and a comparison set.
What should the agency establish before changing the page?
Run two separate baselines.
The product-page baseline
Record identity, price, availability, identifiers, structured data, policies, and the product claims needed to answer the question.
The answer baseline
Record the engine, exact prompt, date, mentioned products, merchant citations, and an answer hash. The client should be able to see what the engine returned without treating the output as stable placement.
Our August 13, 2026 test covered 15 selected Shopify products in logged-out Perplexity sessions.
- MEASURED: thirteen product pages passed every deterministic product check. The field review selected one actionable Fix candidate from the remaining evidence.
- OBSERVED: 13 target products were absent from their answer. The rendered answers cited none of the target merchant domains.
- INFERRED: some clients need a technical Fix; others need a Prove test focused on comparison, authority, or buyer-job evidence. The baseline distinguishes those jobs, but it does not establish causation.
The field cohort yielded one actionable Fix candidate and baseline engine observations. No merchant change or unchanged-prompt rerun was performed.
How does the agency choose the first fix?
Use the smallest correction that addresses the buyer question.
For a product with missing or inconsistent facts, correct the factual surface first. For a technically complete page, inspect whether the product explains the buyer job and differentiates itself truthfully from the alternatives the engine selected.
In the selected cohort, Misen’s braiser had a concrete deterministic product gap and was absent from an answer that named four alternatives. Graza Drizzle passed the deterministic checks but was still absent from a finishing-oil question. Those are different client jobs. The first is a Fix candidate; the second is a Prove candidate.
Do not change schema, product copy, site architecture, and off-site distribution at the same time. The rerun will not tell the client which change mattered.
What should the client report contain?
The deliverable should include:
- store, product URL, and target product;
- exact buyer question and named engine;
- verified, inferred, and runtime-required product evidence;
- baseline answer hash, mentions, citations, and compared products;
- one approved correction and a record of the changed surface;
- unchanged-prompt rerun evidence; and
- gained, unchanged, or lost mention and citation status.
Keep commercial outcomes separate. A changed answer does not prove qualified traffic, conversion, or revenue. Those require attribution beyond the recommendation audit.
Can an agency sell this before it has a before-and-after case?
It can sell a bounded pilot with accurate expectations. The first goal is to complete one permissioned Fix and unchanged-prompt rerun. Until then, the agency has a measurement method and baseline evidence, not proof that its work changes placement.
Colter’s current 15-product cohort establishes that recommendation misses occur in one named-engine test. It does not yet establish willingness to pay or a repeatable fix-to-placement effect.
FAQ
How many products should be in the first client audit?
One to three products is enough to establish the workflow. Choose distinct buyer jobs and keep each prompt fixed.
Should an agency combine ChatGPT, Gemini, and Perplexity into one score?
No. Record each engine separately. They can use different sources and return different products for the same question.
Does a client need to install an app before the audit?
No. The initial Recommendation Audit uses public storefront evidence and read-only observations. Any later store change requires the client’s authority.
What is the first successful agency outcome?
A permissioned correction followed by an unchanged-prompt rerun with a clearly recorded answer delta. A reply, meeting, or readiness-score increase is not product proof.
Run a client Recommendation Audit
Review the Recommendation Audit proof method and evidence documentation before presenting the result to a client.
Evidence record
- Date: August 13, 2026
- Engine: Perplexity Search, logged out, public default experience
- Cohort: 15 selected public Shopify products
- Method: one prewritten unbranded category prompt per product; one observation per prompt
- Observed: 13 targets omitted; two mentioned; zero target merchant-domain citations
- Product state: 13 pages passed deterministic product checks; field review selected one actionable Fix candidate
- Fix boundary: one actionable Fix candidate identified; no merchant change or unchanged-prompt rerun performed
- Limitation: no permissioned merchant fix-to-rerun result, cross-engine rate, traffic, conversion, or revenue evidence
Public examples: Misen braiser, Graza Drizzle, and PerTronix wire set.