Pilot report · Method v1.0 · US market

What AI recommends when you ask for a mechanical keyboard.

Forty purchase-intent prompts — budgets, switch types, use cases, personas — asked to eight assistants. This pilot validates our pipeline; the headline tables below come from the dry run and are labeled as such. The final version ships with the complete prompt log.

HUMAN-VERIFIED · AIBUY-VB-1.0
40 prompts × 8 assistants · window Sep 05–12, 2026 · scoring rules v1.0 Automated: data, charts, draft · Human: spot-check 15% of answers, all headline numbers, conclusions

Headline findings dry-run sample, final tables pending

62%prompts with one brand leading all assistants
31%answers containing at least one spec error
2.4×mention gap, #1 vs #5 brand
17%recommended products already outdated
AssistantMention rate (top-3)Avg. first positionSpec accuracyTypical behavior
ChatGPT78%1.691%Concise list, heavy consensus on 2 brands
Claude64%2.188%Adds reasoning per pick; widest niche coverage
Gemini58%2.482%Occasional price staleness
Perplexity52%2.886%Cites reviews; most current availability
Doubao44%3.374%US-market specs weakest; CN market names differ
Kimi40%3.579%Leans on translated CN reviews

“Spec accuracy” = share of stated facts matching current official merchant data at test time. Full per-prompt log ships with the final version.

Cross-market excerpt

The same 20 core prompts were also run through the China panel (Doubao, Yuanbao, Kimi, Wenxin). Early pattern: US and China panels agree on the “safe default” brand less than half the time — and CN assistants cite domestic retail ecosystems the US panel never mentions.

US PANEL AGREEMENT

Top pick identical across all four US assistants in 26% of prompts. Consensus cluster of 2–3 brands in 62%.

CN PANEL AGREEMENT

Top pick identical across all four CN assistants in 31% of prompts. Brand-name overlap with US panel: 38%.

What this means for merchants: being absent from the consensus cluster is not “ranking #6” — it is not existing for a buyer who never scrolls past the first answer. That is the gap the diagnostic measures.

What happens next

This pilot was built with the production pipeline; final publication adds the complete prompt log, per-answer excerpts, and model version strings. Category #2 ships next month — current candidates: coffee gear, baby strollers, upright vacuums.