This page is the complete rulebook for how AIBUY measures AI shopping assistants. Every published number can be traced to prompts, dates, model versions, and a logged answer. When our method changes, the version changes — and old reports keep their original version label.
100 purchase-intent prompts per category, built from three layers: category words (what), scenario words (when/why), and persona words (who's buying). Prompts are phrased the way real buyers ask an assistant — not the way marketers write ads.
Every prompt goes to every assistant in the panel for that market, fresh conversation per prompt, browser or official API, answers captured verbatim with timestamps and model version.
Answers are parsed into brand mentions, positions, stated specs and prices. Scoring is rule-based — the counting code is deterministic and versioned.
A person re-reads a random 15% of raw answers against the automated scoring, plus 100% of any number destined for a headline. Discrepancies fix the rule, not the data.
Drafts and charts are generated automatically; conclusions and opinions are written by a person. Every report opens with the verification card stating exactly which parts were human-checked.
Defined once, used everywhere. All four are computed per brand, per assistant, per test window.
| Metric | Question it answers | Definition |
|---|---|---|
| Mention Rate | How often does AI name you at all? | Share of prompts where the brand appears in the answer. Also reported as top-3 mention rate. |
| Rank Position | How early does AI name you? | Average position of first mention in ordered lists (1 = first recommended). Lower is better. |
| Info Accuracy | Does AI describe you correctly? | Share of stated facts (price, key specs, availability) that match the merchant's current official data. Errors are logged verbatim. |
| Suppression Ratio | Who wins your questions? | For prompts where your brand is mentioned, how often a named competitor appears earlier. Reported as a multiple (e.g. 2.4× = the leader appears first 2.4× more often). |
The US-market panel covers the first four; the China-market panel covers the rest. The cross-market project runs identical prompts through both panels. Panels are reviewed each quarter as assistants launch or change materially.
For every run: the full prompt set with version hash; per-prompt raw answers (verbatim), assistant identity, model version, UTC timestamp; the scoring output; the spot-check record. Reports link to a public excerpt of this log. Raw logs stay on our servers for 24 months.
Assistants personalize, A/B test and change without notice — our numbers are a structured sample of a moving target, not a census. We state sample sizes and dates everywhere, we don't smooth over week-to-week noise, and we never project GMV impact from mention data.
Method v1.0 (this document) governs all 2026 Q3–Q4 reports. Material changes bump the version and are listed on this page with dates. Historical reports are never re-scored under a new method silently — if we re-run them, both versions are shown.