Book a call

How to Measure AI Search Visibility

AI search visibility is a pattern to measure across many signals. The useful work is checking whether answer engines can find your business, describe it accurately, cite credible sources, and send better-informed buyers toward the right next step.

Start with the buyer problem

The commercial question is rarely "are we visible in AI?" in the abstract. It's usually sharper than that. A buyer is asking ChatGPT, Perplexity, Google AI Overviews, AI Mode, or another answer engine who they should shortlist. Your business may be omitted, mentioned with thin proof, described in old language, or grouped with competitors that look clearer because their sites explain the offer better.

AI visibility measurement starts with diagnosis. A dashboard helps only when the team knows what it's measuring. The aim is to see whether AI systems can find the right evidence, connect it to the buyer problem, and represent the business accurately.

For a service business or professional firm, the audit should answer whether the business is mentioned for the prompts that matter, whether it's cited by name, whether the answer is accurate about services and geography, whether useful pages can be accessed, and whether AI-influenced searches or enquiries look commercially better over time.

This avoids a vague score that moves every time a model, prompt, location, or source set changes.

Measure across several signals

Generated answers vary by platform, prompt wording, location, timing, account context, source access, and retrieval method. A single screenshot proves only that something happened once. A stable measurement system needs repeated checks.

Google's guidance for AI features says AI Overviews and AI Mode use the same broad foundations as the rest of Google Search. Standard indexing and snippet eligibility apply. The guidance also explains query fan-out, where the system may issue related searches across subtopics and data sources before forming an answer.

One exact keyword is too narrow. A useful audit measures groups of buyer prompts, related questions, comparison scenarios, local modifiers, and proof-seeking searches. The goal is to see whether the business is consistently findable, accurately represented, and supported by sources a buyer would trust.

Use four measurement levels.

Level What it shows Confidence
Answer output What the model said for a prompt on a date Low to medium
Citation and source set Which pages, brands, and third parties supported the answer Medium
Access and log data Whether crawlers, referrals, and user-triggered fetches reached the site Medium to high
Commercial signals Whether branded demand, enquiries, and lead quality changed High when repeated

The higher levels are slower to collect, but more useful. Screenshots can start a conversation. Logs, citations, Search Console trends, referral traffic, and enquiry quality help decide what to change.

Dedicated Google reporting now gives teams a stronger Google-specific layer. It works best beside prompts, citations, crawler logs, analytics, and commercial signals.

When that reporting reveals falling clicks or CTR, the AI Overviews traffic-drop diagnostic shows how to investigate the anomaly without treating feature exposure as proof of cause.

Build a prompt set around buyer decisions

Start with prompts a real buyer would use before they know what to buy or who to trust. Keep the set small enough to repeat monthly. A bloated prompt library becomes hard to maintain and easy to over-interpret.

Use six groups.

Prompt type Example prompt What to watch
Navigational "What does [business name] do?" Accuracy, old positioning, wrong services
Commercial "Who helps [audience] with [service]?" Mentions, competitors, service fit
Local "Best [service] studio for [location] businesses" Geography, local proof, local competitors
Comparison "[Business] vs [competitor] for [need]" Differentiation, proof, positioning
Diagnostic "Why is my [problem] not converting leads?" Whether your expertise is connected to the problem
Proof-seeking "Show examples of [business] work or results" Case studies, reviews, cited sources

Run each prompt across the platforms that matter for your buyers. For some businesses that means ChatGPT, Perplexity, Google AI Overviews, AI Mode, and Microsoft Copilot. For others, Google plus one answer engine is enough. When Perplexity is the problem surface, use the Perplexity-specific visibility diagnostic before expanding the measurement programme.

Test more than the brand name. Branded prompts show whether the system understands you. Unbranded and comparison prompts show whether it considers you before the buyer has chosen you.

Record what the answer actually does

Measurement improves when the team records the same fields each time. The audit log should capture the prompt, answer, citations, and likely buyer takeaway.

Field Record
Platform ChatGPT, Perplexity, Google, Copilot, or another tool
Date The day the prompt was run
Prompt Exact wording used
Brands named Your business and competitors mentioned
Sources cited Pages, third-party sites, directories, reviews, or no citation
Accuracy Correct, partly correct, wrong, or missing
Service fit Whether the answer links you to the right work
Geography Locations named or implied
Proof Case studies, reviews, credentials, examples, or claims
Sentiment Positive, neutral, cautious, or negative
Evidence gap What the answer needed but could not find
Google AI impressions Impression total and relevant page, country, device, and date dimensions
Next action Fix content, source quality, access, positioning, or tracking

Visibility tools can help when they track prompts over time. They still need a prompt set that reflects how buyers decide. A neat chart built on weak prompts creates false confidence.

Separate mentions, citations, and recommendations

A brand mention, a citation, and a recommendation are different signals.

A mention means the system knows the business exists or has seen it in a source. A cited source means the answer is leaning on a page, profile, review, article, or directory. A recommendation means the answer is actively positioning the business as a fit for the buyer's need.

Google reporting can show whether pages appeared in Google's generative AI features. One impression confirms that a link to the site appeared in the combined AI Overviews and AI Mode reporting surface. It doesn't identify the query, distinguish the feature, or reveal what the answer said. Answer quality still needs prompt checks, citation review, analytics, and sales feedback because the report can't show whether an answer recommended the business with the right proof, described the offer accurately, or sent a qualified buyer.

Result What it means Likely fix
Mentioned with weak support The brand is known, but the answer lacks a strong supporting source Improve source-worthy pages, third-party proof, and internal evidence
Cited but misrepresented The source exists, but the page or external profile is unclear or stale Update service language, page structure, and business profiles
Recommended with thin proof The answer sounds positive but gives the buyer little reason to trust it Add case studies, reviews, outcomes, process, and author context
Competitors cited instead Other sources explain the category or decision better Build stronger comparison, service, and proof content

This is where AI-ready website foundations become practical. If an audit shows weak source quality, unclear service pages, blocked crawlers, or thin proof, the fix usually sits in the content system and site structure.

If the pattern is inaccurate representation, use the AI entity trust audit before changing the whole content plan. It separates business identity, service clarity, profile consistency, schema, proof, third-party corroboration, and access issues.

Check whether AI systems can access the site

If answer engines have limited access to important pages, the audit will keep finding gaps. Check crawlability, indexability, robots.txt, CDN settings, WAF rules, server responses, and whether key pages are available as text.

OpenAI's crawler documentation separates its crawlers by job. OAI-SearchBot is used for search products, GPTBot is used for model training, and ChatGPT-User supports user-triggered requests. Perplexity's crawler documentation also separates PerplexityBot, which supports website surfacing in search results, from Perplexity-User, which handles user-triggered fetches. It recommends using current published IP ranges when configuring WAF rules.

Those distinctions matter because a business may choose different access rules for search visibility, training, and live user retrieval. Treating every AI crawler as one category creates noisy conclusions.

Infrastructure data is more reliable than a one-off answer check. Cloudflare's AI Crawl Control documentation describes tools for monitoring AI crawler activity and robots.txt compliance, setting crawler-level rules, and inspecting crawler behaviour. Server logs, CDN logs, Search Console, and analytics data can fill in the same picture.

When the audit finds access issues, link the finding to the likely cause. A key service page may be missing from the index. Useful proof may be hidden in images, scripts, PDFs, or inaccessible components. A WAF rule may block a crawler needed for retrieval. robots.txt may allow one class of access while blocking another. The page may load for humans while returning poor text to crawlers. Internal links may also make priority pages harder to discover. For the setup decisions behind those findings, use the AI crawler access and robots.txt guide to separate each crawler's role before changing policy.

If the issue is structure, the repair may sit with website design alongside content work. If the issue is search interpretation, internal linking, citations, or topic coverage, it belongs in SEO.

Use Google's generative AI report with the standard Performance report

Google measurement needs its own handling because the dedicated report answers a narrow visibility question. As of 31 August 2026, Google says it has rolled out generative AI performance insights to websites worldwide. An individual property may still lack a visible report or useful data while rollout finishes or when the site hasn't received enough generative AI impressions.

The Search report combines impressions from AI Overviews and AI Mode. It groups those impressions by page, country, device, and date. It doesn't expose the triggering query, clicks, or feature-level attribution, so an increase can't tell you which feature appeared, what the answer said, or whether a buyer visited. These details reflect Google's documentation checked on 7 September 2026.

Google's AI feature eligibility guidance still rests on normal Search foundations. A supporting page must be indexed and eligible to show with a snippet. Google doesn't require separate AI-only technical work for eligibility.

Evidence source What it answers What to check next
Generative AI performance report Are links to our pages appearing in Google's combined generative AI features? Compare page, country, device, and date patterns
Standard Performance report How are clicks, CTR, queries, and wider Web search trends changing? Compare the same period and priority pages
Manual result samples What did the answer say, cite, and recommend? Record accuracy, proof, competitors, and buyer takeaway
Analytics and branded demand Did people arrive directly or search for the brand later? Review landing pages and assisted journeys
Enquiry and CRM evidence Did visibility influence suitable opportunities? Record source context, service fit, and lead quality

Start with the dedicated report and note which priority pages gained or lost impressions. Open the standard Search Performance report for the same period, then compare clicks, CTR, queries, and broader Web search movement. Sample the priority queries manually to inspect wording and citations. Finish with analytics and CRM notes to see whether the change influenced real demand.

This workflow keeps an impression in proportion. A page can gain AI impressions while clicks fall. It can also earn fewer visits while generating better branded follow-up or stronger enquiries. None of those movements establishes cause alone.

For broader platform context, read our guide to AI Overviews and AI Mode. If page-level results reveal gaps across the buyer journey, the query fan-out content planning guide helps map those missing questions to the right pages.

Check reporting quality before diagnosing change

Check the reporting record before treating a sharp movement as a content or technical problem. Google's Search Console data anomalies record lists known logging and reporting issues. Use it to see whether a relevant entry covers the affected dates. A generic anomaly notice doesn't prove that your movement is a logging error.

Compare the dates, pages, countries, and devices affected. Note recent site releases, migrations, content changes, and exclusion settings. Then compare the standard Performance report over the same window. When AI impressions, clicks, and CTR move in different directions, use the Google traffic-drop diagnostic to test plausible causes before choosing a fix.

Look for referral and lead-quality signals

AI visibility reaches beyond a search report. Some buyers click a cited source. Others remember the brand and search later. Some arrive through answer-engine referrals or mention in an enquiry that they already compared providers.

Track signals that show better-informed demand: identifiable referral traffic from answer engines, server log events from AI crawlers and user-triggered fetches, branded search increases after content or PR activity, enquiries that mention specific services or case studies, higher conversion rates on service and proof pages, and sales calls where prospects arrive with better context.

SparkToro's 2026 zero-click research is useful context because it reinforces a practical reporting problem. Many search journeys never become a clean organic session, so sessions alone understate visibility and influence. Use that as context, then bring the report back to citations, branded demand, referral quality, and leads.

A site can lose some low-value informational clicks while gaining better-qualified enquiries. Every traffic decline still deserves investigation, but visibility should be judged against the work the website is meant to do: build trust, clarify fit, and help the right buyer take action.

Keep llms.txt as supporting evidence

An llms.txt file can clarify priority pages, canonical URLs, service language, and machine-readable context. Keep it in a supporting role while citations, crawler logs, referrals, branded demand, and lead quality do the heavier measurement work.

The evidence is still cautious. Ahrefs' analysis of 137,189 websites with valid llms.txt files found that 97% received no requests during May 2026. Contentful's review of llms.txt and search visibility found no validated evidence that the file reliably improves AI citation frequency, referral traffic, or answer inclusion.

For measurement, llms.txt belongs in the low-confidence layer. Record whether the file exists, whether it's current, whether it's requested, and whether it points to the strongest pages. Then keep watching actual citations, crawler logs, referrals, branded demand, and lead quality.

Off Piste's own llms.txt and llms-full.txt are examples of support files. They're useful housekeeping alongside crawlable pages, clear service content, structured proof, accessible HTML, and earned authority.

Turn the audit into fixes

The audit's value is the repair map. Every finding should point to an action.

Finding Fix path
AI tools omit the business for obvious buyer prompts Strengthen service pages, internal links, topical coverage, and third-party proof
Answers describe old positioning Update site copy, business profiles, structured data, and cited pages
Competitors are cited for category explanations Publish stronger decision content, comparisons, examples, and practical frameworks
The business is cited but weakly recommended Add proof, outcomes, process detail, reviews, and clearer fit signals
Important content is hard to parse Improve headings, semantic HTML, accessible content, and page structure
Crawler logs show blocked access Review robots.txt, WAF rules, CDN settings, and bot-specific controls
Priority pages have few AI impressions or repeated citation absence Use the citation-absence diagnostic to test eligibility, relevance, and source gaps
AI impressions, clicks, and CTR diverge Use the traffic-drop diagnostic before assigning cause
Page visibility exposes gaps across the buyer journey Map the missing query families with the AI Mode fan-out guide
Enquiries are low quality despite mentions Refine positioning, service boundaries, pricing context, and next steps

When the audit points to weak source quality, use the citation-worthy content framework to decide what each page needs before rewriting. For a systematic check of claims and corroboration, run a website content evidence audit. The repair is usually more specific proof, clearer scope, better examples, named sources, or entity signals rather than more keyword coverage.

The overlap between accessibility, search, and AI visibility is strongest when structure is the problem. Clear headings, descriptive links, text alternatives, readable content, and logical sections help people and machines understand the same page. Our guide to website accessibility and SEO covers that foundation in more detail.

Some fixes are content-led. Some are technical. Some are positioning decisions that need sharper service language. Measurement stops the team guessing which problem they have.

Review one monthly scorecard

Repeat the audit on a cadence the business can maintain. Monthly is enough for most service businesses. Fortnightly can make sense during a launch, repositioning, migration, or active content campaign.

Use the same prompt set, logging fields, and commercial signals. Add new prompts only when buyer behaviour changes or a new service matters. Keep a short notes field for model changes, website updates, PR mentions, reviews, and major search shifts.

Review Google AI impressions on the same cadence. Note significant reporting changes beside content refreshes, technical releases, PR mentions, reviews, and service-page updates so the trend has context. Keep the measures separate so a large impression number can't hide poor answers or weak enquiries.

Measure Monthly review
Google AI impressions Direction by priority page, country, and device
Cited-page coverage Share of the sampled prompt set that cites a useful page
Answer accuracy Repeated errors or missing service and geography details
Referred sessions Identifiable answer-engine visits and useful landing pages
Branded follow-up Changes in branded queries and direct return visits
Qualified enquiries Leads with the right need, fit, and buying context

A single odd answer is a watch item. Act when a movement repeats across at least two evidence layers. Use citations and manual samples to understand an answer gap, logs and Search Console to test access and visibility, then analytics and lead quality to judge commercial effect. If those systems need to be joined or the fix crosses content and technical work, SEO support can turn the evidence into an implementation plan.