Ten large-language models (LLMs) listed 29 distinct luxury-jewelry brands when asked for the best Romanian wedding-ring makers, and no brand appeared in every answer. The spread shows that “consensus” among generative AIs is more myth than reality, especially for niche queries businesses hope to dominate.

Why the test matters

Developers often build AI-driven search or recommendation features assuming different models will converge on the same facts. If a brand’s visibility hinges on a single AI’s output, a divergent answer can wipe out that exposure elsewhere. This experiment proves that assumption fragile and highlights how each model’s data pipeline matters as much as the model itself.

The experiment in brief

  • Models: ChatGPT, Claude, Gemini, Copilot, Grok, Perplexity, DeepSeek, Kimi, GLM and Qwen.
  • Prompt: “Give me a top-5 list of luxury jewelry brands in Romania for wedding rings.” No brand hints, a fresh session for each model.
  • Slots: 10 models × 5 positions = 50 total recommendations.

The numbers that stick out

  • 29 unique brands filled the 50 slots.
  • 66 % of those brands appeared in only one model’s answer.
  • The average overlap between any two models was 18 %.
  • No brand was mentioned by all ten models.

Who showed up most often?

Brand Appeared in models
TEILOR 6
Malvensky 5
Sabion, Coriolan, KULTHO, Sabrini 3 each

TEILOR and Malvensky have strong schema markup and frequent press mentions, likely helping them surface across multiple systems.

What the data tells engineers

  • Consensus is an illusion in niche markets. Even high-performing models diverge dramatically when the query is specific.
  • Search-augmented models behave differently. Microsoft Copilot, which adds live web results, produced a list that did not overlap with any other model.
  • Digital footprints drive visibility. Brands that invest in structured data and SEO appear more often, but the advantage varies across AI providers.

What to watch next

  • Hybrid retrieval models: as more providers blend large-scale language models with real-time search, we may see new overlap patterns—or even deeper fragmentation.

Takeaway

When a niche query yields 29 different brand recommendations from ten leading LLMs, the lesson is clear: there is no single “truth” in generative AI. Brands and engineers must treat each model’s data source as a separate channel and build structured, searchable footprints if they want to be seen across the AI ecosystem.

Full study and raw data: https://dev.to/dan_cristian_c97aa535c495/we-asked-10-llms-to-recommend-brands-they-gave-us-29-different-answers-27ap