01

Evidence record for an AI visibility audit

Evidence record for an AI visibility audit
FieldWhat to recordWhy it matters
ScopeMarket, language, audience, competitors and collection datesPrevents an aggregate score from hiding the comparison
PromptExact wording, intent class and approved variantsMakes the observation repeatable
SurfacePlatform, model or mode, account state and location where relevantAnswers differ by product and context
AnswerFull response, screenshot and timestampPreserves evidence after the live answer changes
Brand outcomeMention, recommendation, rank/order, accuracy and qualificationSeparates visibility from correctness
SourcesCited URL, domain, claim supported and source typeShows what evidence entered the answer
DecisionGap, confidence, proposed action, owner and recheck dateConnects research to a testable next step
02

Define the decision and prompt universe first

An audit is meaningful only inside a fixed market, audience and set of buyer decisions.

Build prompts from real discovery, comparison, suitability, location, pricing and risk questions. Include branded checks for accuracy, but do not let them dominate the score. Record why each prompt matters and which competitors a buyer could reasonably compare.

Executive Intelligence’s partner page proposes agreeing the market, competitors, buyer questions, AI tools and proof limits before collection, then retaining answers, screenshots and sources. That is a strong audit record; the provider’s conclusions still need to be evaluated against the underlying evidence.[1]

  • Market and language
  • Buyer stage and decision
  • Prompt family and exact wording
  • Eligible competitors
  • AI surfaces and collection dates
  • Claim the audit is allowed to make
03

Repeat prompts and preserve answer-level evidence

One answer is an example, not a stable visibility estimate.

Run the same prompt multiple times where the product allows it, and use predefined paraphrases to test intent rather than one lucky wording. Record model or mode, date, account state and location when relevant. Do not silently replace failed, empty or contradictory runs.

Recent primary research argues that AI-search visibility should be treated as a distribution because answers vary across runs, prompts and time. A 2026 critical review similarly warns that discoverability, citation, prominence, factual use and business outcomes are different stages with heterogeneous evidence.[4][5]

04

Measure distinct outcomes instead of one opaque score

A brand can be cited without recommendation, mentioned inaccurately or recommended without a visible citation.

For every answer, record whether the brand appears, how it is described, whether it is shortlisted, its order, qualifications or warnings, and which URLs are cited. Add factual-accuracy review by a person who knows the product or service.

Publish the numerator and denominator behind rates. For example: mentioned in 12 of 30 eligible answers, recommended in 5, correctly described in 10, and supported by a visible citation in 4. Keep prompt-level rows available so another reviewer can challenge the summary.

Read the SEO content architecture article
05

Check access and source evidence before rewriting pages

A missing mention can originate before writing quality enters the decision.

OpenAI says OAI-SearchBot is used to surface websites in ChatGPT search results and recommends allowing it in robots.txt and through relevant network controls. Google says eligibility for its AI search features uses the same search fundamentals: crawl access, index eligibility, useful textual content and structured data that matches the visible page.[2][3]

Check robots rules, CDN or firewall blocks, canonical URLs, indexability, internal links and important content rendered in HTML. Then review whether the page states current product facts, location, availability, limitations and evidence clearly. A special AI schema or an llms.txt file is not a substitute for these fundamentals; Google says no special AI markup is required for its AI features.[3]

06

Turn repeated gaps into measured experiments

Prioritise gaps that recur across commercially important prompts and have a plausible evidence-based cause.

Classify the proposed action: technical access, entity or business-fact correction, answer-page improvement, primary proof, or independent source coverage. Assign an owner and recheck date. Preserve the baseline prompt set while adding a separate exploratory set for new questions.

The original GEO paper reported visibility gains within its experimental benchmark, but those results do not guarantee organic discovery or durable commercial impact on live platforms. Treat content changes as hypotheses, repeat the same audit and connect visibility movement to qualified visits, leads or other business outcomes.[6][5]

Compare AI-search specialists
07

FAQ

Frequently asked questions

What does an AI visibility audit measure?

It measures how a brand appears for a fixed set of buyer questions across defined AI surfaces, including mentions, recommendations, citations, accuracy, competitors and source patterns.

How many prompts should an audit use?

There is no universal number. Use enough predefined prompts and repeats to cover important decisions without creating a sample too large to review at answer level. Publish the exact scope and denominator.

Why should each prompt be repeated?

Generative answers can vary across runs, wording, model versions and time. Repetition helps distinguish a recurring pattern from a single observation.

Does allowing OAI-SearchBot guarantee a ChatGPT citation?

No. It removes one possible access barrier. OpenAI does not guarantee that an accessible page will be selected, summarised, cited or recommended.

Does a high AI visibility score prove revenue impact?

No. Visibility is an intermediate outcome. Connect changes to qualified referral visits, assisted demand, leads, sales or another defined business measure while retaining attribution limits.

Can an agency guarantee AI recommendations?

No. AI systems and their retrieval processes are outside an agency’s control. A credible provider can document the baseline, improve relevant evidence and measure change without guaranteeing placement.

08

Sources and further reading

Editorial method: English and Portuguese SERP patterns were reviewed on 26 July 2026. Partner pages supplied practitioner questions; factual claims were checked against the primary and official sources below. Commercial and partner material was not treated as independent proof.

  1. How to run an AI visibility audit — Executive Intelligence

    Partner methodology for scoping questions, retaining evidence and prioritising gaps. It is a practitioner source, not independent validation.

  2. OpenAI — Overview of OpenAI Crawlers

    Official documentation for OAI-SearchBot, search discovery, robots controls and published crawler IP ranges.

  3. Google Search Central — AI features and your website

    Official crawl, index, content and structured-data guidance for AI Overviews and AI Mode.

  4. Schulte, Bleeker and Kaufmann — Don’t Measure Once

    Primary 2026 research on repeated measurement and run-to-run variability in AI-search visibility.

  5. Martinez — Critical Survey of Generative Engine Optimization

    Primary 2026 review separating discoverability, citation and outcome claims and documenting evidence limits.

  6. Aggarwal et al. — GEO: Generative Engine Optimization

    Foundational benchmark study; cited with its experimental scope and limits retained.