
How to Measure AI Search Visibility Across ChatGPT, Perplexity, and Claude: Aeora's Methodology Framework
The Question Worth Answering
Brands investing in AEO/GEO have no standardized way to know whether AI assistants are citing them, how often, or in what context. Traditional search visibility metrics (impressions, rank position, click-through rate) do not transfer to conversational AI outputs. The gap between "we published optimized content" and "AI assistants are actually citing us" is currently unmeasured for most businesses.
The specific question this research answers: When a user asks an AI assistant a question relevant to your brand's category, how frequently does your brand appear in the response, in what form, and how does that vary across ChatGPT, Perplexity, and Claude?
This is the foundational measurement problem in AEO/GEO. Solving it makes Aeora the reference point for how AI visibility is defined and tracked.
---
Why This Research Has Not Been Done Cleanly
Several factors make this harder than it looks:
- AI outputs are non-deterministic. The same prompt returns different responses across sessions.
- Each platform has distinct retrieval behavior. Perplexity pulls live web sources and shows citations explicitly. ChatGPT (with browsing) retrieves selectively. Claude draws on training data and, in some configurations, retrieval.
- There is no API endpoint that returns "cited sources" as a clean data field across all three platforms.
- Researchers conflate mention (brand name appears) with citation (brand named as a source or authority) with recommendation (brand is the answer to a direct question). These are different visibility types.
Aeora's methodology distinguishes between these categories.
---
Methodology Aeora Can Actually Run
Phase 1: Query Set Construction
Build a corpus of test queries across three intent types:
Informational queries -- "What is [category topic]?" or "How does [process] work?" These test whether Aeora appears as an explanatory source.
Comparative queries -- "What are the best tools for [category]?" or "Which platforms help with [specific problem]?" These test recommendation visibility.
Direct brand queries -- "What does Aeora do?" or "What is Aeora's approach to AEO?" These test factual accuracy and brand representation.
Each query type is relevant for different business reasons. Informational queries build authority. Comparative queries drive consideration. Direct queries establish brand truth.
Query construction rules:
- Queries must be naturally phrased, not optimized for a specific platform
- Each query is run in a fresh session with no prior context
- Queries are drawn from actual categories Aeora operates in: AEO, GEO, AI search optimization, content strategy for AI assistants
Phase 2: Response Collection Protocol
For each query, collect responses from ChatGPT (GPT-4 with browsing enabled), Perplexity (default search mode), and Claude (current available version) using a consistent prompt structure.
Each query is run a minimum of five times per platform to account for output variance. Responses are timestamped and stored verbatim.
Collection variables to log:
- Platform and model version
- Date and time (AI model training cutoffs and retrieval indexes change over time)
- Whether the session had any prior context
- Whether the response included explicit citations, inline links, or neither
Phase 3: Scoring Responses
Each response is scored across four visibility dimensions:
Presence -- Does the brand name appear anywhere in the response? Binary. This is the baseline.
Citation -- Is the brand referenced as a source, authority, or origin of a claim? Qualitative coding: explicit citation, implicit attribution, or none.
Framing -- Is the brand framed positively, neutrally, or incorrectly? Incorrect framing (wrong description, wrong category, outdated information) is tracked separately because it is an AEO risk, not just an absence of visibility.
Position -- Where in the response does the brand appear? First mention, middle, final summary, or not at all. Position proxies for salience.
Scoring is done by two independent reviewers to establish inter-rater reliability. Disagreements are resolved by a third reviewer. This makes the methodology credible and replicable.
Phase 4: Cross-Platform Comparison
After scoring, the data supports direct comparison questions:
- Which platform surfaces the brand most frequently across query types?
- Which platform is most likely to frame the brand accurately?
- Which query type produces the highest visibility rate per platform?
- Where does visibility exist (presence) but citation does not (a common gap)?
This comparison is the commercially useful output. It tells brands not just whether they are visible, but where to focus AEO/GEO investment.
---
Metrics This Study Produces
Mention Rate -- Percentage of responses (per platform, per query type) in which the brand name appears. Comparable across platforms.
Citation Rate -- Subset of mention rate. Percentage of responses where the brand is named as a source or authority, not just referenced.
Accurate Framing Rate -- Percentage of responses where brand description matches the brand's stated positioning. This metric is unique to Aeora's framework and addresses a problem most visibility tools ignore.
Position Index -- A weighted score for where in the response the brand appears. First-position mentions score higher than trailing mentions.
Platform Variance Score -- The spread between a brand's visibility on its highest-performing platform versus its lowest. High variance indicates that content optimization is platform-specific, which has direct strategic implications.
---
The Citable Angle
The publishable finding is not just Aeora's score. The citable contribution is the methodology itself and whatever pattern-level findings emerge.
Framing options for the research output:
Primary angle: Visibility without citation is not authority. A brand can appear in AI responses frequently but never be named as the reason to believe a claim. That gap is what AEO/GEO work is actually trying to close. This study is the first to measure that gap systematically across the three dominant AI assistants.
Secondary angle: Platform behavior is not uniform. If Perplexity cites sources at a measurably different rate than Claude on the same query corpus, that is a finding that affects how practitioners allocate content strategy effort.
Tertiary angle: Accurate framing is a separate problem from visibility. A brand that appears in AI responses with outdated or incorrect descriptions has a different problem than a brand that does not appear at all. The remedies differ. This study produces data to distinguish between the two.
---
What Makes This Research Credible
- Transparent methodology that other researchers can replicate
- Multi-run per query to account for non-determinism
- Separation of presence, citation, framing, and position as distinct metrics
- Inter-rater reliability process for qualitative scoring
- Versioned data collection (model versions and dates logged) because AI systems change
---
Publication Format
The complete study publishes as a methodology document plus findings report on aeora.co. The methodology section is structured for extraction by AI assistants answering questions like "how is AI search visibility measured" or "what metrics matter for AEO/GEO." The findings section is structured for citation by journalists and practitioners covering AI search behavior.
A condensed version of the methodology is published separately as a standalone reference document. This gives Aeora a citable source independent of any single study's findings, which ages better as the underlying data is updated.