EXPERIMENT_016 // AI.IMPRESSION.SIMULATOR

AI Impression Simulator

Type a brand name and watch ChatGPT, Gemini, and Claude answer the same three questions a real buyer would ask — discovery, consideration, comparison — then see each model's visibility score, sentiment, and consistency compared side by side.

TEXT INPUTCREATED 2026.07.10BETA

LOADING EXPERIMENT...

AI models are now commercially influential buying surfaces: a Research and Markets estimate puts AI influence on product comparisons at 62% of decisions, but AI-attributed checkouts at only 23%. This experiment demonstrates that gap directly — the same three questions a real buyer would ask, sent in parallel to three frontier models, scored against a fixed rubric for visibility, sentiment, and internal consistency.

HOW IT WORKS

Three questions, three models, one brand

Enter a brand name. The API sends identical discovery, consideration, and comparison queries to ChatGPT (gpt-5.4-mini), Gemini (gemini-3.5-flash), and Claude (claude-haiku-4-5) in parallel, streaming each model's card as it resolves rather than waiting for all three.

A scoring pass, not a fact-check

Once a model's three responses return, a separate Claude Haiku call at temperature zero grades them against a fixed rubric: a 0–100 visibility score, a sentiment read from the consideration answer, a consistency badge across the three responses, and an internal-consistency flag — explicitly not a verified accuracy check, since the experiment has no ground-truth brand database.

Side-by-side, shareable

Results render as a card per model plus a Recharts bar chart comparing visibility scores. The brand name is encoded into the URL, so a shared link re-runs the same live query rather than replaying a cached snapshot.

WHAT THIS PROVES

The gap between AI mentions and AI recommendation is visible and measurable in real time, not just in aggregate market research — the same brand can score very differently across three models on the same day.

A structured LLM-as-judge scoring pass is a workable substitute for a ground-truth database when the goal is internal consistency, not fact-checking — the rubric is explicit about that boundary rather than presenting a confidence score it cannot back up.

← BACK TO PLAYGROUND

SYSTEM.INT // 2026 LABS_CORE v2.78.2

LATENCY: STATUS: NOMINAL