Evidence record for an AI visibility audit
| Field | What to record | Why it matters |
|---|---|---|
| Scope | Market, language, audience, competitors and collection dates | Prevents an aggregate score from hiding the comparison |
| Prompt | Exact wording, intent class and approved variants | Makes the observation repeatable |
| Surface | Platform, model or mode, account state and location where relevant | Answers differ by product and context |
| Answer | Full response, screenshot and timestamp | Preserves evidence after the live answer changes |
| Brand outcome | Mention, recommendation, rank/order, accuracy and qualification | Separates visibility from correctness |
| Sources | Cited URL, domain, claim supported and source type | Shows what evidence entered the answer |
| Decision | Gap, confidence, proposed action, owner and recheck date | Connects research to a testable next step |
Define the decision and prompt universe first
An audit is meaningful only inside a fixed market, audience and set of buyer decisions.
Build prompts from real discovery, comparison, suitability, location, pricing and risk questions. Include branded checks for accuracy, but do not let them dominate the score. Record why each prompt matters and which competitors a buyer could reasonably compare.
Executive Intelligence’s partner page proposes agreeing the market, competitors, buyer questions, AI tools and proof limits before collection, then retaining answers, screenshots and sources. That is a strong audit record; the provider’s conclusions still need to be evaluated against the underlying evidence.[1]
- Market and language
- Buyer stage and decision
- Prompt family and exact wording
- Eligible competitors
- AI surfaces and collection dates
- Claim the audit is allowed to make
Repeat prompts and preserve answer-level evidence
One answer is an example, not a stable visibility estimate.
Run the same prompt multiple times where the product allows it, and use predefined paraphrases to test intent rather than one lucky wording. Record model or mode, date, account state and location when relevant. Do not silently replace failed, empty or contradictory runs.
Recent primary research argues that AI-search visibility should be treated as a distribution because answers vary across runs, prompts and time. A 2026 critical review similarly warns that discoverability, citation, prominence, factual use and business outcomes are different stages with heterogeneous evidence.[4][5]
Measure distinct outcomes instead of one opaque score
A brand can be cited without recommendation, mentioned inaccurately or recommended without a visible citation.
For every answer, record whether the brand appears, how it is described, whether it is shortlisted, its order, qualifications or warnings, and which URLs are cited. Add factual-accuracy review by a person who knows the product or service.
Publish the numerator and denominator behind rates. For example: mentioned in 12 of 30 eligible answers, recommended in 5, correctly described in 10, and supported by a visible citation in 4. Keep prompt-level rows available so another reviewer can challenge the summary.
Read the SEO content architecture articleCheck access and source evidence before rewriting pages
A missing mention can originate before writing quality enters the decision.
OpenAI says OAI-SearchBot is used to surface websites in ChatGPT search results and recommends allowing it in robots.txt and through relevant network controls. Google says eligibility for its AI search features uses the same search fundamentals: crawl access, index eligibility, useful textual content and structured data that matches the visible page.[2][3]
Check robots rules, CDN or firewall blocks, canonical URLs, indexability, internal links and important content rendered in HTML. Then review whether the page states current product facts, location, availability, limitations and evidence clearly. A special AI schema or an llms.txt file is not a substitute for these fundamentals; Google says no special AI markup is required for its AI features.[3]
Turn repeated gaps into measured experiments
Prioritise gaps that recur across commercially important prompts and have a plausible evidence-based cause.
Classify the proposed action: technical access, entity or business-fact correction, answer-page improvement, primary proof, or independent source coverage. Assign an owner and recheck date. Preserve the baseline prompt set while adding a separate exploratory set for new questions.
The original GEO paper reported visibility gains within its experimental benchmark, but those results do not guarantee organic discovery or durable commercial impact on live platforms. Treat content changes as hypotheses, repeat the same audit and connect visibility movement to qualified visits, leads or other business outcomes.[6][5]
Compare AI-search specialistsFAQ
Frequently asked questions
What does an AI visibility audit measure?
It measures how a brand appears for a fixed set of buyer questions across defined AI surfaces, including mentions, recommendations, citations, accuracy, competitors and source patterns.
How many prompts should an audit use?
There is no universal number. Use enough predefined prompts and repeats to cover important decisions without creating a sample too large to review at answer level. Publish the exact scope and denominator.
Why should each prompt be repeated?
Generative answers can vary across runs, wording, model versions and time. Repetition helps distinguish a recurring pattern from a single observation.
Does allowing OAI-SearchBot guarantee a ChatGPT citation?
No. It removes one possible access barrier. OpenAI does not guarantee that an accessible page will be selected, summarised, cited or recommended.
Does a high AI visibility score prove revenue impact?
No. Visibility is an intermediate outcome. Connect changes to qualified referral visits, assisted demand, leads, sales or another defined business measure while retaining attribution limits.
Can an agency guarantee AI recommendations?
No. AI systems and their retrieval processes are outside an agency’s control. A credible provider can document the baseline, improve relevant evidence and measure change without guaranteeing placement.
Sources and further reading
Editorial method: English and Portuguese SERP patterns were reviewed on 26 July 2026. Partner pages supplied practitioner questions; factual claims were checked against the primary and official sources below. Commercial and partner material was not treated as independent proof.
- How to run an AI visibility audit — Executive Intelligence
Partner methodology for scoping questions, retaining evidence and prioritising gaps. It is a practitioner source, not independent validation.
- OpenAI — Overview of OpenAI Crawlers
Official documentation for OAI-SearchBot, search discovery, robots controls and published crawler IP ranges.
- Google Search Central — AI features and your website
Official crawl, index, content and structured-data guidance for AI Overviews and AI Mode.
- Schulte, Bleeker and Kaufmann — Don’t Measure Once
Primary 2026 research on repeated measurement and run-to-run variability in AI-search visibility.
- Martinez — Critical Survey of Generative Engine Optimization
Primary 2026 review separating discoverability, citation and outcome claims and documenting evidence limits.
- Aggarwal et al. — GEO: Generative Engine Optimization
Foundational benchmark study; cited with its experimental scope and limits retained.
