How to Measure AI Search Visibility

Cape Wired · GEO & AI Search Guides

How to Measure AI Search Visibility

Measure AI search visibility with a repeatable prompt baseline, citations, Google AI impressions, ChatGPT referrals, competitor visibility and commercial outcomes.

In this guide

AI search visibility is not one number.

A business can be mentioned without being cited, cited without being recommended, recommended without receiving a click, and receive referral traffic without being able to see the exact prompt that caused it.

The practical measurement job is therefore to combine several signals without pretending they are interchangeable.

Measure exposure, representation, source use, traffic and outcomes separately. Then look for patterns across them.

That is the Cape Wired measurement principle.

What exactly are you trying to measure?

Before choosing tools, define the event you care about.

Useful distinctions include:

  • Presence: the brand, product or service appears in the answer.
  • Mention: the system names the brand in explanatory text.
  • Citation: a page from the business is cited or linked as a source.
  • Comparison: the brand appears alongside alternatives or competitors.
  • Recommendation: the answer presents the brand or product as an option that fits the user's request.
  • Platform impression: a platform reports that a URL from the site appeared in one of its AI features.
  • Referral visit: a user clicks from an AI/search experience to the website.
  • Commercial outcome: the visit or journey contributes to a lead, purchase, booking, signup or other valuable action.

Those events are related, but they are not synonyms.

Why one AI visibility score is rarely enough

A single percentage can be useful as a headline, but it hides what actually changed.

For example, a brand may move from being absent to being cited as a source without being recommended. That is a different change from moving from 'one of several options' to the first product suggested in a shopping-style answer.

Cape Wired therefore uses a dashboard of separate measures and only creates summary percentages where the calculation is explicit.

These are Cape Wired reporting conventions. They are not official metrics supplied by Google, OpenAI, Perplexity or another AI platform.

The Cape Wired AI Visibility Measurement Stack

1. Platform-native exposure

2. Prompt-set visibility

3. Citation and source use

4. Referral and engagement data

5. Commercial outcomes

Each layer answers a different question.

1. Platform-native exposure: use first-party platform data where it exists

Google Search Console now has dedicated Generative AI performance reports

In June 2026, Google announced dedicated Generative AI performance reports in Search Console for Search and Discover.

The Search report includes impressions from AI Overviews and AI Mode.

Google is rolling the reports out to a subset of websites, so not every property will see them yet.

For Search, the dedicated report currently lets site owners analyse generative-AI impressions by:

  • page
  • country
  • date
  • device

The report tells you how often links from your property were shown in supported generative AI Search features.

It does not currently provide a prompt or query dimension showing the exact question that triggered the impression. The dedicated report currently exposes impressions rather than clicks, CTR, average position or a recommendation score.

Google says this data is also included in the overall Search Performance report, while the dedicated view isolates generative-AI visibility.

Google also provides a separate Search generative AI control

For properties where the control is available, site owners can include or exclude their links and content from supported Search generative AI features such as AI Overviews and AI Mode.

Google says excluding the site prevents its content from being visible or used to ground those supported generative AI features, so the site will not receive traffic or impressions from them.

Google also says this control is not used as a ranking or inclusion signal for other parts of Search.

This is a participation control, not a visibility optimisation tactic. If a business chooses to exclude itself, the absence of generative-AI impressions is expected and should not be interpreted as an SEO failure.

What can Google's AI impression report tell you?

  • whether URLs from the property are appearing in AI Overviews or AI Mode
  • which pages receive those impressions
  • how visibility changes over time
  • which countries generate those impressions
  • which devices are involved

What can it not tell you by itself?

  • whether Google verbally recommended your brand
  • the exact user prompt behind each impression
  • whether your brand was compared positively or negatively
  • whether a page was the decisive source for the generated answer
  • whether the impression produced a sale or lead
  • whether a user clicked a generative-AI result in the dedicated report

Platform-native impressions are valuable exposure data. They are not a complete AI-visibility measurement system.

Be careful with Search Console aggregation

Google's reporting can aggregate data by property or by page depending on the view, which can produce differences between chart totals and page-level tables.

Search Console also applies row and privacy limitations to some performance data.

When comparing periods, use the same view, filters and aggregation method.

What about ChatGPT?

OpenAI documents a useful referral measurement mechanism, but it is different from Google's platform-impression reporting.

OpenAI says publishers that allow OAI-SearchBot can track referral traffic from ChatGPT in analytics platforms such as Google Analytics.

OpenAI also says ChatGPT automatically adds the parameter utm_source=chatgpt.com to referral URLs from ChatGPT search results.

That makes clicked traffic from ChatGPT easier to identify.

Referral tracking measures visits after a click. It does not measure every time ChatGPT mentioned, cited, compared or recommended the brand without a click.

ChatGPT Search can display inline citations and a Sources panel, so manual prompt testing can separately record whether your pages are being used as sources.

What about Perplexity and other AI answer engines?

Perplexity publicly documents its use of web search and citations, which makes manual citation tracking possible.

For platforms that do not provide a Search Console-style site-owner impression report, measurement normally relies more heavily on repeatable prompt testing, source inspection and website referral data.

Do not assume that different platforms expose equivalent metrics.

2. Build a repeatable prompt-set baseline

Platform reports are useful, but they do not answer every commercial question.

A repeatable prompt set lets you observe how the brand is represented for customer questions that matter under defined test conditions.

Cape Wired typically starts a focused project with around 10 to 20 representative prompts.

That range is a practical starting sample, not an OpenAI, Google or Perplexity requirement.

A broader brand, product catalogue or multi-market business may need a larger set.

Choose prompts by customer decision, not by keyword volume

Include a mix of:

  • informational prompts
  • problem or need-based prompts
  • comparison prompts
  • recommendation prompts
  • product or service suitability prompts
  • brand-specific verification prompts where useful
  • local or geographic prompts when location materially affects the decision

For example:

  • What causes beard dandruff?
  • What should I look for in a beard shampoo for flakes?
  • Best beard shampoo for a dry, flaky beard in the UK
  • Beard shampoo vs normal shampoo for dandruff
  • Which beard-care brands make anti-dandruff shampoo?
  • Is [brand/product] suitable for beard dandruff?

Do not fill the baseline with twenty near-identical prompts merely to make the sample larger.

Record the testing conditions

AI answers can vary.

For each test, record where practical:

  • platform or product
  • date
  • prompt wording
  • country or market
  • device or environment where relevant
  • logged-in or signed-out state
  • whether memory or personalisation may affect the result
  • model or search mode where the platform exposes it
  • brand outcome
  • cited sources
  • competitors mentioned
  • notes on material differences from the previous run

The objective is not perfect laboratory control. It is enough consistency to make later comparisons meaningful.

How do you measure when answers are personalised?

You cannot remove every source of personalisation or context from modern AI products, but you can reduce avoidable variation and document what remains.

For repeatable benchmark testing:

  • Use the same market or geographic setting for comparison periods.
  • Use the same logged-in or signed-out state where practical.
  • Keep memory, saved preferences and prior conversation context consistent, or use a fresh session where the platform allows it.
  • Avoid mixing mobile, desktop, app and web results unless device experience is part of the test.
  • Record the exposed model, search mode or product when the platform shows it.
  • Test local prompts from the relevant market rather than assuming one country's result represents another.
  • If a commercially important result looks highly personalised, repeat it in a cleaner session or controlled comparison before reporting it as a wider brand trend.
The goal is not to claim a universal answer. It is to create a documented, repeatable observation under defined conditions.

AI answers are variable: do not treat one run as permanent truth

The same prompt can produce different wording, sources or recommendations across runs.

Changes can also come from model updates, search-index changes, new web sources, personalisation, location or the evolving conversation context.

For that reason, avoid making a major business conclusion from one isolated answer.

If a result matters commercially and appears volatile, repeat the prompt or a representative subset before calling the movement a trend.

Do not turn natural answer variation into false decimal-point precision.

A practical result classification

Cape Wired records each prompt using a simple outcome ladder.

1. Absent — the brand is not present.
2. Mentioned — the brand appears but is not being used as a cited source or recommendation.
3. Cited — the business website is cited or linked as a source.
4. Compared — the brand is discussed alongside alternatives.
5. Recommended — the brand or product is presented as a suitable option for the user's request.

More than one label can apply. A brand can be both cited and recommended in the same answer.

Do not assume that 'recommended' means the platform formally endorses the business. It is a reporting label for how the answer was phrased in that test.

Track representation quality as well as presence

Two answers can both mention the brand while representing it very differently.

One may describe the business accurately and positively. Another may contain outdated information, emphasise a limitation, repeat criticism or describe the offer incorrectly.

Cape Wired therefore recommends adding a simple qualitative Representation Quality field to the raw prompt results:

  • Positive: accurate representation with clearly favourable context.
  • Neutral: accurate mention without meaningful positive or negative framing.
  • Mixed: both strengths and limitations are presented.
  • Negative: the answer contains materially unfavourable framing.
  • Factually incorrect: important information about the brand, product or service is wrong or outdated.

Representation Quality is a Cape Wired observation field, not a platform sentiment score or ranking metric.

Record the evidence behind the classification so the label can be reviewed rather than treated as an unexplained number.

Cape Wired metric 1: Brand Visibility Rate

Brand Visibility Rate answers a simple question: in what proportion of the tracked prompts did the brand appear at all?

Brand Visibility Rate = prompts where the brand is present ÷ total prompts tested × 100

If the brand appears in 12 of 20 representative prompts, the Brand Visibility Rate is 60%.

This does not tell you whether the brand was cited, compared favourably or recommended. Use it as a presence metric only.

Cape Wired metric 2: Citation Rate

Citation Rate = prompts where the business website is cited ÷ total prompts tested × 100

This helps distinguish brand awareness in an answer from direct use of the business's own pages as sources.

A citation can still be neutral or even used to support a limitation, so inspect the context rather than celebrating the percentage alone.

Cape Wired metric 3: Recommendation Rate

Recommendation Rate = prompts where the brand or product is presented as a suitable option ÷ recommendation-eligible prompts tested × 100

The denominator should include only prompts where a recommendation is a realistic outcome.

Do not divide recommendation appearances by a set dominated by purely informational questions such as 'What is beard dandruff?'.

Cape Wired metric 4: Prompt Share of Voice

Share of Voice is useful when you want to compare the tracked brand with a defined competitor set.

Cape Wired uses a prompt-level approach so verbose answers do not gain extra weight simply by repeating one brand several times.

Prompt Share of Voice = brand appearances across the tracked prompt set ÷ appearances of all tracked brands across the same prompt set × 100

Count each tracked brand no more than once per prompt for this calculation.

The competitor list, prompt set and test window must stay consistent for the comparison to be meaningful.

This is a Cape Wired comparative metric, not a platform-defined share-of-voice statistic.

Should you create one weighted AI visibility score?

You can, but use caution.

Assigning more points to a recommendation than a citation can create a convenient executive score, but the weights are a business choice.

Unless the weighting has a clear decision purpose, separate metrics are usually more transparent.

A business should be able to see whether improvement came from more appearances, more citations, more recommendations or simply a change in the prompt sample.

3. Track citations and source use

For every prompt that uses search or displays sources, record which domains and pages are cited.

Useful fields include:

  • your own cited URL
  • third-party source URL
  • competitor source URL
  • source type: product page, guide, review, directory, retailer, media, forum or other
  • whether the cited information is current
  • whether the source supports, contradicts or merely mentions the brand
  • whether the same source recurs across several prompts

This is where Article 14's third-party evidence audit becomes measurable.

Do not count a citation as a recommendation

A source may be cited because it contains a useful fact, not because the platform prefers the business.

Likewise, a competitor may be recommended while your page is cited for background information.

Record source use and recommendation language separately.

4. Track referral traffic and on-site engagement

Website analytics show what happens when AI visibility produces a click.

OpenAI's documented utm_source=chatgpt.com parameter makes ChatGPT Search referrals comparatively easy to isolate in analytics.

In Google Analytics 4, traffic-source dimensions such as source and medium can be used to understand where sessions came from.

Useful referral measures include:

  • sessions
  • users
  • landing pages
  • engaged sessions
  • engagement rate
  • key events
  • lead submissions
  • purchases
  • revenue where ecommerce tracking is configured

You may also create custom reports or explorations grouping relevant AI referrals together.

Referral traffic is a lower-funnel signal, not a visibility census

A person can see your brand in an AI answer and never click.

A platform may mention your product as part of a comparison without sending traffic.

Referral sessions therefore answer 'who clicked through?', not 'how often were we visible?'.

Watch for post-search demand indicators without claiming attribution

AI visibility can influence a later action without producing an immediately identifiable referral click.

A customer might see the brand in an AI answer, remember it, then search for the brand directly, return through another channel or mention the AI tool during a later sales conversation.

Useful supporting observations can include:

  • changes in branded search demand
  • direct or returning visits to priority pages
  • customers mentioning ChatGPT, AI Mode, Perplexity or another AI tool during enquiries
  • CRM notes about how a prospect first encountered the brand
  • assisted conversions where the analytics setup can support them
  • changes in lead quality alongside broader AI visibility movement

Treat these as post-search demand indicators, not attributable AI conversions.

Correlation can help identify a pattern, but it does not prove that an AI answer caused the later visit, search, enquiry or purchase.

5. Track commercial outcomes

Visibility is valuable only in the context of a business goal.

For ecommerce, useful outcomes may include:

  • product views
  • add-to-cart events
  • checkout starts
  • purchases
  • revenue
  • repeat purchase where measurable

For service businesses, outcomes may include:

  • contact-form submissions
  • qualified leads
  • consultation bookings
  • phone calls
  • quote requests
  • closed work where CRM attribution exists

Google Analytics can report key events and attribute credit across marketing touchpoints, but attribution models determine how credit is distributed.

Do not describe every sale after an AI referral as being caused solely by AI search.

Build a baseline before making major changes

Measurement becomes much more useful when you record the starting position before changing the site.

A practical baseline includes:

  • the fixed prompt set
  • current Brand Visibility Rate
  • Citation Rate
  • Recommendation Rate for recommendation-eligible prompts
  • Prompt Share of Voice against a defined competitor set
  • the pages most often cited
  • third-party domains most often cited
  • Google generative-AI impressions if the Search Console report is available
  • current ChatGPT and other identifiable AI referral traffic
  • current commercial outcomes from those referrals

Save the raw test results, not only the summary percentages.

That lets you inspect what changed later rather than relying on a dashboard number with no evidence behind it.

How should you compare competitors?

Use the same prompt set, market and test window for every tracked brand.

Do not compare your UK brand results from one month with a competitor's US results from another model or date.

Track only competitors that are genuinely relevant to the customer decision.

For each prompt, record which competitors are:

  • mentioned
  • cited
  • compared
  • recommended

The objective is to understand the competitive information landscape, not to create a universal AI league table.

What does a useful measurement table look like?

For each prompt test, a working spreadsheet can include:

  • prompt ID
  • prompt
  • intent type
  • market
  • platform
  • date
  • brand present: yes/no
  • brand cited: yes/no
  • brand compared: yes/no
  • brand recommended: yes/no
  • own URL cited
  • third-party sources cited
  • competitors present
  • first-listed brand where meaningful
  • answer notes
  • representation quality: positive / neutral / mixed / negative / factually incorrect
  • material factual error: yes/no
  • follow-up action

For referral reporting, maintain a separate table or analytics view for sessions, landing pages, key events and revenue.

Should you track first position or first mention?

Only when it has a clear meaning in the experience you are testing.

Some answers use ordered lists; others present several options without implying a strict ranking.

If you record 'first listed' or 'first mentioned', describe it exactly that way.

Do not rename it 'AI ranking position' unless the platform itself presents an explicit ranked result.

What counts as a meaningful change in AI visibility?

There is no universal percentage change that automatically counts as meaningful.

A useful movement is one that is large enough, persistent enough and commercially relevant enough to survive the normal variability of the platform and the limits of your sample.

Before calling a change meaningful, check:

  • Sample stability: was the same prompt set, market and competitor group used?
  • Repeatability: does the movement appear again on a rerun or in a representative subset?
  • Breadth: did several prompts move, or did one unusual answer create most of the change?
  • Direction across layers: did mentions, citations, recommendations, platform impressions or referrals move in a consistent direction?
  • Commercial relevance: did the change occur on prompts that matter to the business rather than low-value informational questions?
  • Source change: are new citations or third-party sources appearing that help explain the movement?
  • Platform change: did a model, search product, geography, rollout or reporting method change at the same time?

A five-point increase can matter if it represents several important prompts moving consistently. A twenty-point jump can be noise if it comes from one small or changed sample.

Use trend bands rather than false precision

For executive reporting, it can be useful to classify movement as stable, directional or material rather than treating every decimal change as equally significant.

Cape Wired does not apply a universal threshold to those labels. The threshold should reflect the prompt-set size, volatility, business importance and evidence available.

How often should you measure?

There is no universal cadence.

Choose a rhythm that matches how quickly you are making changes and how much natural variation the platform exhibits.

During an active project, a weekly or fortnightly pulse on a representative subset can help spot material movement without rerunning the entire benchmark constantly.

For a stable programme, a monthly full baseline may be enough.

Those are workflow examples, not platform requirements.

Avoid reacting to daily fluctuations unless the business has enough data and a specific operational reason to monitor at that frequency.

Article 16 covers the separate question of how long meaningful improvement can take.

What if Google Search Console does not show the Generative AI report?

That does not prove the site has zero AI visibility.

Google says the report is still being rolled out and may also be absent when a property has not received enough generative-AI impressions.

Continue using the standard Search Performance report, prompt testing, source checks, referral analytics and commercial outcomes.

If the dedicated report becomes available later, add it as another measurement layer rather than rewriting the historical methodology.

What if you have very little AI referral traffic?

Low referral volume does not automatically mean low visibility.

AI answers can satisfy the user's question without a click, and many appearances will never create a referral session.

For a small site, the prompt baseline and citation/source audit may therefore show useful movement before referral traffic becomes statistically meaningful.

Keep the commercial objective in view, but do not demand large traffic samples before acknowledging earlier-stage visibility changes.

A worked measurement example

Imagine a business tracks 20 representative prompts around one product problem.

At baseline:

  • brand present in 6 prompts: Brand Visibility Rate = 30%
  • own website cited in 2 prompts: Citation Rate = 10%
  • 10 prompts are recommendation-eligible; brand recommended in 1: Recommendation Rate = 10%
  • tracked competitors appear 22 times across the prompt set; the brand appears in 6 prompts: Prompt Share of Voice = 21.4% when each brand is counted once per prompt
  • ChatGPT sends 8 identifiable referral sessions during the reporting period
  • one referral session triggers a lead-form key event

After a period of technical, content and evidence work, the same prompt set is tested again under comparable conditions.

Suppose the results become:

  • brand present in 11 prompts: 55%
  • own site cited in 6 prompts: 30%
  • brand recommended in 3 of the same 10 recommendation-eligible prompts: 30%
  • prompt share of voice increases against the same tracked competitors
  • ChatGPT referral sessions increase
  • Google's generative-AI impressions increase on two priority pages

That is evidence that several measured outcomes moved in the same direction under the comparison conditions.

It still does not prove that one specific page edit caused the change. The measurement tells you what moved; causal attribution requires more caution.

What not to do

  • Do not treat one prompt as representative of the whole market.
  • Do not call every mention a recommendation.
  • Do not call every citation positive endorsement.
  • Do not compare different prompt sets and present the percentages as a clean trend.
  • Do not change the competitor set halfway through a share-of-voice comparison without documenting it.
  • Do not mix markets or languages without separating the results.
  • Do not claim Google Generative AI impressions are ChatGPT or Perplexity visibility.
  • Do not use ChatGPT referral traffic as proof of every no-click brand appearance.
  • Do not classify all direct traffic as 'dark AI' without evidence.
  • Do not use arbitrary weights to create a proprietary score and present it as a platform metric.
  • Do not overreact to one day's answer variation.
  • Do not report conversion attribution more precisely than your analytics setup supports.

The Cape Wired AI visibility measurement checklist

  • We have defined what counts as presence, citation, comparison and recommendation.
  • We have a fixed representative prompt set.
  • Prompts are grouped by intent, not just keyword variants.
  • We record the platform, date, market and important test conditions.
  • We use the same competitor set for comparative reporting.
  • Brand Visibility Rate uses the same denominator across comparison periods.
  • Recommendation Rate includes only recommendation-eligible prompts.
  • Prompt Share of Voice is labelled as a Cape Wired metric, not a platform statistic.
  • We record the actual URLs and third-party domains cited.
  • We review Google Generative AI impressions where Search Console provides the report.
  • We track ChatGPT referrals using analytics and OpenAI's utm_source=chatgpt.com signal.
  • We track relevant engagement, key events, leads, purchases or revenue.
  • We retain raw test results behind summary percentages.
  • We do not interpret a single run as a permanent ranking.
  • We separate visibility movement from claims about what caused it.

The simplest AI visibility measurement rule

Do not ask whether one score went up. Ask where the brand appeared, how it was represented, which sources were used, whether people clicked, and whether those visits contributed to the business.

Google's dedicated generative-AI Search Console reports now provide a valuable native exposure layer for eligible properties.

OpenAI provides identifiable ChatGPT referral tracking.

Prompt testing fills part of the gap between those platform signals by showing mentions, citations, comparisons and recommendations for customer questions you care about.

No single layer tells the whole story.

Use them together, keep the definitions stable and resist the temptation to turn an evolving information environment into a false universal ranking score.

Want a baseline of how your brand appears in AI search today?

Cape Wired's AI Visibility Review builds a representative prompt set, records brand and competitor visibility, checks citations and source use, and turns the results into a practical baseline for future comparison.

Need to improve the underlying foundation as well?

The GEO Foundation Project combines visibility measurement with crawlability, page clarity, internal linking, structured data, content planning and evidence work.

How Long Does It Take to Improve AI Search Visibility?Understand why recrawling, source changes, model behaviour and evidence growth do not move on one fixed schedule. Brand Mentions, Reviews and Third-Party Sources: Do They Affect AI Visibility?Audit the external sources behind citations and brand representation. What Is AI Search Visibility?Return to the core definitions behind presence, source use and recommendation visibility. Can ChatGPT and AI Search Engines Crawl Your Website?Check whether technical access problems are limiting the pages available to search systems. Why Internal Linking Matters for AI Search and GEOConnect the pages you want customers and crawlers to discover without relying on arbitrary linking formulas. GEO vs AEO vs SEO: What Does Your Business Actually Need?Place AI-visibility measurement within a wider organic search and answer-engine strategy.

Sources and further reading

This website uses cookies

We use cookies to personalise content, provide social media features, and analyse our traffic. We also share information about your use of our site with our analytics partners. You can change your preferences at any time. For more information, please see our Privacy Policy and Cookie Policy.