The short answer
Look for six capabilities: multi-provider measurement, daily benchmark cadence, citation-level attribution, a served substrate, protocol coverage, and honest statistics. In practice that means: measurement across at least ChatGPT, Gemini, and Perplexity (provider concentration is the biggest source of bias); daily automated runs; attribution down to which URLs the AI actually read, not just whether you were mentioned; structured data and live endpoints the system serves for you, not just recommendations to your content team; MCP support today with ACP and UCP feed-readiness as checkout moves in-chat; and fixed question cohorts with held-out controls and confound checks. If a vendor demo can't show you the exact AI response text and the URLs it cited, keep looking.
The capability checklist
| Capability | What good looks like | The question that exposes weak vendors |
|---|---|---|
| Provider coverage | ChatGPT, Gemini, Perplexity minimum; Claude a plus | "Which providers, and can I see per-provider numbers side by side?" |
| Cadence | Daily automated runs, weekly at absolute minimum | "Show me last Tuesday vs today for the same question." |
| Citation attribution | Every response stores its cited source URLs | "Which of MY pages did the AI read when it recommended me?" |
| Served substrate | JSON-LD, llms.txt, MCP endpoint, AI feeds — generated and hosted | "What do you publish, or do you only give recommendations?" |
| Protocol readiness | MCP server for the catalog; ACP/UCP-shaped feeds | "Can Claude or ChatGPT query my catalog through you today?" |
| Statistical honesty | Fixed cohorts, held-out controls, flagged-run hygiene | "How do you tell my lift from a provider-wide shift?" |
Why the measurement details matter more than the dashboard
Two failure modes dominate this category, and both look fine in a demo:
- Composition drift. Composition drift is when a changing question set moves your trend line instead of real visibility: the tool adds questions over time, the denominator changes, and the "trend" is an artifact. Ask whether trends are anchored to a fixed question cohort. A vendor who can't answer this is reporting noise.
- Attribution theater. A rising mention rate means nothing if the rise also happened for brands that did nothing. The honest test is a control: comparable buyer queries deliberately left untreated, or untreated competitor brands tracked on the same window. In our own published pilot work, roughly a third of an apparent lift was provider-wide drift that a naive before/after would have claimed as a win — the claimable part was only the answers that cite the published pages.
A falsifiable rule of thumb: any vendor claiming a specific visibility lift should be able to show the response-level evidence — the answers, the cited URLs, and a control that didn't move. "Visibility up 40%" with none of those three is marketing, not measurement.
A worked example with real numbers
An Australian B2B office-equipment brand went from 7% to 32% ChatGPT recommendation rate on generic buyer questions in two weeks after publishing five buyer-intent guides on its own domain. The reason we can claim that number: the winning responses cite the new pages (zero such citations existed before publication, 96 after), while answers citing nothing moved in line with untreated brands on the same window. That is the standard of evidence to demand from any system you evaluate — including ours.
Red flags
- Screenshot-based "monitoring" or manual spot checks instead of automated, stored response data.
- Monthly snapshots sold as trend data — AI providers change retrieval behavior week to week.
- Guarantees of the form "we'll make you #1 in ChatGPT". Nobody controls the model; the honest promise is substrate + measurement + iteration.
- No served output. If the deliverable is a PDF of recommendations, you're buying consulting, not a system.
- B2B catalogs treated like consumer retail: quote-based pricing, compatibility constraints, and multi-SKU configurations need first-class handling, not a price field left at zero — agents read a zero as a price.
FAQ
What should an agent commerce system cost? Pricing in the category ranges from roughly $100/month for single-brand visibility tracking to four figures monthly for full substrate-plus-measurement platforms with protocol endpoints. The cost question that matters is coverage: how many providers, how many tracked questions, how much of the served layer is included. A cheap tool that measures one provider weekly is expensive per useful datapoint.
Do I need this if I already rank well in Google? Yes — organic rank and AI recommendation share are correlated but regularly diverge. AI answers weight machine-readable specificity and citable structure, and they synthesize across sources rather than ranking pages; brands that dominate classic search are routinely absent from generic AI buying answers in their own category. The only way to know your position is to measure it directly.
Should the system publish content for me automatically? Automated publishing works when it's gated: scored against a quality rubric, published to surfaces you control, and measured against held-out controls so dead content gets caught and rerouted. Ungated auto-publishing at volume risks thin pages on your domain. Look for systems that treat every published piece as an experiment with a measurable outcome, and that tell you when content is NOT the right lever.
How fast should I expect results? When the mechanism works, it's visible quickly: in our published pilot, the first ChatGPT citation of new pages landed within about 24 hours of indexing and the recommendation-rate step was measurable within two weeks. If a system shows no citation movement after four to six weeks, the content is targeting the wrong queries, the pages aren't being crawled, or the category's answers are dominated by third-party sources — a good system tells you which, rather than asking for more time.
What about Gemini and Google AI Mode specifically? Treat per-provider claims separately. Providers differ in how they ground answers: movement on one does not imply movement on another, and a system should show you per-provider evidence rather than a blended number. Blended numbers hide the fact that a lift may exist only where the system's published surfaces are actually being cited.