Forty purchase-intent prompts — budgets, switch types, use cases, personas — asked to eight assistants. This pilot validates our pipeline; the headline tables below come from the dry run and are labeled as such. The final version ships with the complete prompt log.
| Assistant | Mention rate (top-3) | Avg. first position | Spec accuracy | Typical behavior |
|---|---|---|---|---|
| ChatGPT | 78% | 1.6 | 91% | Concise list, heavy consensus on 2 brands |
| Claude | 64% | 2.1 | 88% | Adds reasoning per pick; widest niche coverage |
| Gemini | 58% | 2.4 | 82% | Occasional price staleness |
| Perplexity | 52% | 2.8 | 86% | Cites reviews; most current availability |
| Doubao | 44% | 3.3 | 74% | US-market specs weakest; CN market names differ |
| Kimi | 40% | 3.5 | 79% | Leans on translated CN reviews |
“Spec accuracy” = share of stated facts matching current official merchant data at test time. Full per-prompt log ships with the final version.
The same 20 core prompts were also run through the China panel (Doubao, Yuanbao, Kimi, Wenxin). Early pattern: US and China panels agree on the “safe default” brand less than half the time — and CN assistants cite domestic retail ecosystems the US panel never mentions.
Top pick identical across all four US assistants in 26% of prompts. Consensus cluster of 2–3 brands in 62%.
Top pick identical across all four CN assistants in 31% of prompts. Brand-name overlap with US panel: 38%.
This pilot was built with the production pipeline; final publication adds the complete prompt log, per-answer excerpts, and model version strings. Category #2 ships next month — current candidates: coffee gear, baby strollers, upright vacuums.