Marketers: Measure AI Visibility with a 30–50 Prompt Test
Published by The Branded Agency · · 12 min read

Measure AI visibility as the combined rate of selection (being retrieved or cited) and absorption (how much an AI answer actually uses your page), tracked with a fixed prompt set plus human reviewed quality checks. The recommended approach pairs automated monitoring with periodic manual audits so scores stay grounded in what AI answers actually say, not just whether your brand appears.
TL;DR:
- Regularly track both selection rate and absorption score because some platforms cite fewer sources but rely on them more heavily, affecting overall visibility quality.
- Build a representative prompt dataset from search logs and customer interactions, dividing prompts into categories to accurately reflect how questions are posed across decision stages.
- Conduct manual reviews alongside automated scoring to verify answer accuracy, support quality, and citation relevance, ensuring metrics reflect true content influence.
- Use a multidimensional scorecard focusing on citation breadth, influence, support quality, and coverage to identify visibility weaknesses and fragility risks.
- Repeated, scheduled measurements with human calibration help detect shifts caused by platform updates or model changes, maintaining accurate visibility assessment over time.
Table of Contents
- Core AI visibility metrics to track
- How to build a representative prompt dataset and baseline
- Measurement workflow: manual testing to automation to continuous evaluation
- Data sources and tools for tracking AI exposure
- Designing a multidimensional AI visibility scorecard
- Mapping visibility to business outcomes and experiments to prove impact
- Human review, quality rubrics, and measuring citation absorption
- Monitoring cadence, drift detection, and alerting
- How the Brand Health Audit maps to these measurement steps
- Realistic expectations and strategic priorities for AI visibility
- Run a free Brand Health Audit scan or request a sample report
- Sources
- FAQ
Core AI visibility metrics to track
A single visibility score hides more than it reveals. Marketers need a set of metrics that separate whether a brand gets mentioned from whether that mention carries any weight.
Start with mention volume, the raw count of times a brand name surfaces across tracked prompts. Selection rate follows: the share of prompts in your test set where an AI answer retrieves or cites your domain at all. Citation rate narrows this further to the proportion of answers that include a direct link or named reference. None of these tell you whether the citation mattered, which is where absorption or influence score comes in: a measure of how much the AI's actual response content draws on your page rather than simply listing it among sources.
- Selection rate = (prompts where your domain is cited) ÷ (total prompts tested), tracked against a fixed set run on a schedule.
- Citation rate = (answers with a direct, named reference to your page) ÷ (total answers where your brand was selected).
- Absorption or influence score rates how much of the answer's substance traces back to your content, not just its source list.
- Prominence captures where your citation sits: early paragraph, supporting footnote, or buried mention.
- Sentiment and citation quality flag whether the surrounding text is accurate, neutral, or damaging.
- AI referral traffic tracks visits arriving from AI platforms as a downstream confirmation signal.
A two-stage measurement framework built on 602 prompts and 21,143 citations found that selection and absorption diverge sharply by platform: some engines cite fewer sources but lean on them more heavily, while others cite broadly with shallow use of each one. Selection and absorption measured separately reveal this gap, which a single visibility number cannot. Tracking both is the only way to know whether your brand is merely present or actually shaping the answer.
How to build a representative prompt dataset and baseline
A prompt set that only tests your brand name will flatter you and tell you nothing useful. The dataset needs to mirror how real buyers phrase questions at each stage of their decision.
- Pull prompts from search logs, sales call transcripts, and customer support tickets to capture how people actually phrase questions before they know your brand name.
- Sort prompts into four categories: brand queries (direct name mentions), category queries (generic product or service questions), decision queries (comparison and shortlist language), and edge cases (niche or unusual framing that stress-tests coverage).
- Build a golden test set of 30 to 50 prompts with documented expected outputs, agreed by a human reviewer, so automated scoring has something fixed to calibrate against.
- Run the full set once on each target platform to establish a baseline before making any content changes.
OpenAI's evaluation guidance recommends this same structure: define the objective, build a representative dataset, specify scoring rules, and compare future runs against that baseline rather than judging any single snapshot in isolation.
Pro Tip: Freeze your golden test set in version control before your first baseline run, so later score changes reflect real shifts in AI answers and not a quietly edited prompt list.
Measurement workflow: manual testing to automation to continuous evaluation
Treat AI visibility measurement as a pipeline with three stages, each handing off to the next rather than replacing it.
- Manual capture. Run a sample of prompts by hand across target platforms monthly, recording full response text, not just whether a citation appeared, so reviewers can assess tone and accuracy.
- Automation. Schedule the full prompt set on a recurring basis using API access or permitted capture methods, then normalize outputs into a consistent format for scoring across platforms.
- Continuous evaluation. Maintain versioned prompt sets, run automated scoring on each cycle, and use the human-reviewed labels from manual capture to calibrate and validate the automated scores over time.
Anthropic's evaluation documentation warns that frequency metrics alone can mask inaccurate or negative descriptions, which is why the manual stage never fully disappears even once automation is running. The handoff point matters most: automated scores should be spot-checked against human review monthly, not assumed accurate indefinitely.
Data sources and tools for tracking AI exposure
First-party and third-party sources answer different questions, and neither replaces the other.
- Google Search Console's generative AI reports show impressions, pages, countries, devices, and dates for appearances in AI Overviews and AI Mode, giving a ground-truth view of Google-specific exposure.
- These reports are narrower than a full dashboard: they report page-level impressions but not prompt-level context or how your citation compares to competitors in the same answer, per Search Console's own documentation.
- Third-party visibility platforms typically add cross-engine comparison, scheduled prompt runs, and citation tracking across ChatGPT, Perplexity, and other assistants that Search Console does not cover.
- Triangulation means reconciling the two: when Search Console shows strong Google AI Overview impressions but a third-party tracker shows weak ChatGPT citations, the gap tells you where to focus content work, not which tool is wrong.
Neither source alone gives a full account of visibility across platforms, which is why the scorecard in the next section treats them as inputs rather than final answers.
Designing a multidimensional AI visibility scorecard
A single dashboard number invites false confidence. A brand can post a high mention count while every mention misrepresents its pricing or positioning, which is exactly the blind spot a scorecard is built to catch.
Build a panel of four to six elements instead: selection rate, citation breadth (how many distinct platforms cite you), absorption or influence score, support quality (accuracy and sentiment from human review), and coverage equity (whether citations cluster on one page or spread across your site). Weight them based on what matters to your business. A company selling through long sales cycles might weight support quality and absorption above raw selection rate, since a single accurate, well-placed citation in a decision-stage answer often matters more than ten shallow brand mentions.
- Selection rate and citation breadth flag whether you are showing up at all, across which platforms.
- Absorption score flags whether your content is shaping the answer or just listed alongside it.
- Support quality from human review catches factual errors and negative framing that counts don't show.
- Coverage equity shows whether one page is doing all the work, a fragility risk if that page changes or drops out.
A GEO framework analysis of citation patterns across ChatGPT, Google, and Perplexity found that absorption and citation breadth move independently: a page can be widely cited and rarely absorbed, or narrowly cited but heavily relied upon. Monthly reports should separate new citations, lost citations, and top-performing queries, with short excerpts showing exactly how the AI used the content, rather than a single trend line.
Mapping visibility to business outcomes and experiments to prove impact
Visibility metrics only matter if they connect to pipeline. The connection has to be built deliberately through experiments, since no platform hands you a clean attribution path from AI answer to signed deal.
- Run before-and-after content experiments: update a page's evidence density or structure, then re-run the prompt set two to four weeks later to measure selection and absorption changes.
- Test landing page variants aimed at decision-stage prompts and track whether citation prominence shifts.
- Track branded search volume as a downstream proxy: rising searches for your brand name after a visibility improvement suggest the AI answer is driving discovery.
- Use assisted conversions and estimated referral visits from AI platforms as directional, not exact, attribution signals.
Executive reporting should lead with one headline metric (typically the scorecard's overall trend), follow with a business impact estimate grounded in branded search or referral data, and close with the next content or technical fix planned. A B2B buyer journey framework offers a useful template for connecting an AI answer to a later signed deal across a longer sales cycle.
Human review, quality rubrics, and measuring citation absorption
Automated scoring cannot tell you whether an AI answer got your pricing wrong or described your product in a way that would make a prospect hesitate. That judgment requires a human reading the actual text.
Build a rubric around four dimensions: accuracy (does the answer state facts correctly), supporting evidence (does the AI's claim match what your page actually says), excerpt reuse (does the answer quote or closely paraphrase your content), and sentiment (neutral, favorable, or unfavorable framing). Anthropic's evaluation guidance explicitly recommends pairing quantitative metrics with qualitative rubrics, since frequent mentions alone can still be inaccurate or negative.
- Repeated textual overlap between your page and the AI's answer suggests meaningful absorption, not just a passing citation.
- Early paragraph coverage signals that the AI treated your content as a primary source rather than a minor supplement.
- Claim-level matching checks whether specific facts or figures in the answer trace directly to your page.
Pro Tip: Run a monthly deep sample of 20 to 30 full-text responses by hand, backed by automated weekly flags for sudden citation loss, so you catch quality problems the counts alone would miss.
Monitoring cadence, drift detection, and alerting
AI answers change as models update, competitors publish new content, and platforms adjust their retrieval logic. A measurement program needs a cadence built around that instability.
- Daily automated captures for a small, high-priority prompt subset catch sudden shifts fast.
- Weekly snapshots across the full prompt set track short-term trend direction.
- Monthly full audits combine automated scoring with human review to confirm quality, not just presence.
- Set alert thresholds for drift signals: a sudden citation loss, a new dominant competitor source, or signs of a model version change, each triggering a defined investigation and content review playbook.
How the Brand Health Audit maps to these measurement steps
A brand health audit evaluates a business's public presence across multiple categories, scored against a fixed checklist rather than subjective opinion. A free scan can provide a starting read on where a brand stands before committing to deeper work.
The GEO audit specifically examines how AI assistants describe a brand, which maps directly to the selection and absorption tracking described above, while the audit's scoring structure separates score, data coverage, and confidence, echoing the scorecard approach this guide recommends. Readers building their own prompt sets and rubrics can use the audit's category breakdown as a reference checklist, and the agency's GEO strategy post and brand description analysis go deeper on the content side of absorption.
Realistic expectations and strategic priorities for AI visibility
AI answers are not deterministic, so a single prompt run tells you little. Methodology, a fixed prompt set, a documented baseline, repeated runs, is what makes a score mean anything at all.
Prioritize pages built as modular evidence containers: clear claims, supporting data, and well-structured sections that an AI can lift cleanly. Resist the pull toward vanity mention counts. A brand cited often but inaccurately is worse off than one cited rarely but well, and the only way to tell the difference is tying every score back to an actual business outcome.
— Quincy
Run a free Brand Health Audit scan or request a sample report
Building the prompt set, scorecard, and review rubric described in this guide takes real setup time, and most marketing teams are already stretched across a dozen other priorities. A free scan can provide a starting baseline across website, search, and AI presence without requiring a sales call or credit card, allowing brands to see where they stand before deciding how deep to go.
- A free scan can cover public data across website, social, and review signals as a starting diagnostic. Other audits may focus on how AI assistants describe and cite a brand, or extend evidence-based scoring to additional categories, some available for a fee.
For teams that want a preview before committing, the sample brand audit report shows the format, and the GEO audit landing page is the direct next step for AI visibility specifically. A further tactical reference worth bookmarking is this 90 day plan for AI search visibility, useful alongside the audit's own recommendations. Start with the free Brand Health Audit scan to see your current standing.
Sources
- From Citation Selection to Citation Absorption: A Measurement Framework for Generative Engine Optimization Across AI Search Platforms
- Evaluation best practices | OpenAI API
- Define success criteria and build evaluations | Anthropic
- Introducing Search Generative AI performance reports in Search Console | Google Search Central Blog
FAQ
What is a good AI visibility score?
There is no universal benchmark because scores depend on your prompt set, industry, and which platforms you track, which is exactly why methodology matters more than the raw number. A useful approach is comparing your own selection and absorption rates over time against your own baseline, rather than chasing an external standard.
How do you track visibility across multiple AI platforms?
Track it by running a fixed, versioned prompt set across each target platform on a recurring schedule, combining automated capture with periodic human review of full response text. Pair Search Console's generative AI reports for Google-specific data with a third-party tracker for cross-engine coverage, then reconcile the two.
What is the best AI visibility tool?
No single tool covers every platform and metric, so most measurement programs combine Search Console's generative AI reports for Google exposure with a third-party tracker for cross-engine citation monitoring. The right combination depends on which platforms matter most to your audience and how much human review capacity you have to validate automated scores.
How do you measure AI visibility in practice?
Measure it as a combination of selection (whether an AI answer cites your brand) and absorption (how much the answer actually relies on your content), tested against a fixed prompt set with a documented baseline. Add human review for accuracy and sentiment, since frequency of mentions alone can mask inaccurate or unfavorable descriptions.
How often should you re-run AI visibility measurements?
A workable cadence is daily automated captures on a small priority prompt subset, weekly snapshots across the full set, and a monthly audit that pairs automated scoring with human review. This cadence catches sudden drift, like a citation loss or a model update, before it affects a full reporting cycle.
Recommended
Where this shows up in your audit
These scored categories cover what this article talks about.
Industry brand audits
Accounting & bookkeeping firms brand audit · Architecture & design studios brand audit · Automotive brand audit
Want this handled for you?
Getting your brand recommended inside AI assistants and AI search.
AI Content Studio at The Branded AgencyGoing deeper on the strategy behind it: The strategy thinking behind the audit, written for brand and marketing leads. Insights from The Branded Agency.
Measured against real data
Every figure we publish comes from completed audits, reported as anonymised averages.
Related articles
- The 30-Day Audit Remediation Plan: Fast Wins vs Strategic Fixes
A pragmatic 30-day remediation sprint guide to prioritising brand health audit recommendations, separating quick technical wins from deep strategic positioning fixes.
- High Overall Score, Low Trust Signals: Reading the Discrepancy
Discover how to read the discrepancy between a high composite Brand Health Audit score and low trust signals, and why trust deficits create critical conversion risks.
- Mapping the B2B Brand Journey From AI Answer to Signed Deal
Learn how modern B2B buyers move from AI assistant synthesis to signed contracts, and how to eliminate narrative drift across every digital touchpoint.
Stay sharp
Get the next brand breakdown in your inbox
Practical brand strategy, messaging and AI-search insights. No fluff, no daily sends — just the work that moves brands.
Written by
Quincy Samycia
Founder & Brand Strategist, The Branded Agency
Quincy leads brand strategy at The Branded Agency, where he has spent over a decade helping founders and B2B teams sharpen their positioning, messaging and creative systems so growth stops depending on guesswork.
More from Quincy Samycia →See where your brand actually stands
Run the Brand Health Audit and get a scored diagnostic of your messaging, positioning and visibility.
Brand Audit