Skip to content
Back to Resources
Guide

AI Visibility: How to Measure Whether AI Recommends You

Skopx Team
August 21, 2026
16 min read

AI visibility is measured by running a fixed set of buyer-intent prompts through AI answer engines on a repeating schedule, then counting how often your brand is named, cited, or linked versus how often competitors are named instead. That count, expressed as share of voice per prompt set per engine per run, is the practical signal available today, because AI answers are generated fresh for each user and no engine publishes a ranking report you can download the way you can pull a position report from Search Console.

Everything else in this article is about doing that counting well: how to build prompt sets that match real purchase questions, which engines to sample, how to define a "mention" so your numbers mean something over time, how to compute share of voice without fooling yourself, and how to report the sampling error honestly instead of pretending a sample of thirty answers is a census. If you are earlier in the journey and still deciding what to change on your site, start with the generative engine optimization guide and come back here for the measurement layer.

What does AI visibility actually measure?

Classic SEO measured position: your URL sat at rank four for a query, and rank four had a knowable click-through curve. AI answers do not have positions. They have a synthesized paragraph, a handful of named entities inside it, and a citation list that may or may not include the sources that shaped the wording. So the unit of measurement changes from "where do I rank" to "how often am I in the answer, and in what role."

There are three distinct things people mean when they say a brand shows up in AI answers, and conflating them produces dashboards that move for no reason:

Mention. Your brand name appears in the generated prose. This is the strongest signal because it is what the reader actually sees. A user who asks "what tools help a small team automate reporting" and reads five vendor names in a paragraph has just formed a consideration set.

Citation. Your domain appears in the source list, footnotes, or link chips attached to the answer. Citation without mention is common and much weaker: the engine used your page to learn a fact and then recommended someone else. Tracking those cases separately is the single highest-value split in the whole discipline, and it is covered in depth in AI citation tracking.

Recommendation. Your brand is named in a position of endorsement, not just listed. "X is a good fit for teams that need Y" is different from "alternatives include X, Y, and Z." If you only count string matches you will treat these as identical, which quietly overstates your position.

A measurement program that separates these three, and reports each one, gives you a story you can act on. A program that reports one blended number gives you a chart nobody trusts by month three.

How do you build a prompt set that reflects real buyer intent?

The prompt set is the entire experiment. A bad prompt set produces confident, useless numbers, and the most common failure is writing prompts that already contain your brand name. "Is Skopx good for social scheduling" will return an answer about Skopx. That measures nothing except the engine's willingness to discuss a named entity.

Buyer-intent prompts are unbranded and phrased the way someone types into a chat box when they have a problem and no shortlist. Build them from four sources:

  1. Your own site. Every problem statement on your product and pricing pages implies a question a buyer asked before landing there. A page about consolidating reporting across tools implies "how do I get one report from tools that do not talk to each other."
  2. Search Console queries. Pull the non-branded queries that already bring impressions, then rewrite them as full sentences. Chat prompts are longer and more conversational than search queries, so "social scheduling tool" becomes "what is the best way to schedule the same post to LinkedIn, Reddit, and Bluesky without doing it three times."
  3. Sales and support transcripts. The objection phrasing in a real conversation is closer to prompt phrasing than anything a marketer writes.
  4. Competitor comparison language. "Alternatives to," "instead of," "versus," and "cheaper than" prompts surface the exact answers where your absence costs the most.

Aim for a set large enough to be stable and small enough to run often. Thirty to sixty prompts per product line is a workable range. Below roughly twenty, a single answer flipping moves your share of voice enough to look like a trend. Above a hundred, cost and run time push you toward sampling less frequently, which is worse: cadence beats breadth when the underlying system is nondeterministic.

Then freeze the set. The strongest discipline in this whole practice is refusing to edit prompts mid-quarter. If you change the questions, you cannot compare this month to last month, and every improvement you report is partly an artifact of rewording. Version the set, date it, and start a new series when you change it.

Skopx generates buyer-intent prompts automatically from a customer's own site as part of its AI visibility feature set, then runs them through search-grounded AI and reports share of voice plus the citation gaps where competitors were named instead. The generation step matters less than the freezing step: however you produce prompts, treat the frozen set as the instrument.

Which AI engines should you sample, and how often?

Different engines behave differently enough that a blended cross-engine number hides more than it reveals. Report per engine, always, and only aggregate for an executive summary that carries the per-engine breakdown underneath it.

SurfaceWhat a "mention" looks likeCitation behaviorPractical sampling notes
ChatGPT with browsingBrand named in prose, sometimes with a link chipLinks appear when the answer is grounded in a live fetch; ungrounded answers cite nothingAnswers vary run to run; sample repeatedly per prompt. See ChatGPT SEO optimization
PerplexityBrand named in prose with numbered inline referencesConsistently attaches a numbered source listThe most legible surface for citation counting. See the Perplexity SEO guide
Google AI OverviewsBrand named in the summary block above resultsLinks out to supporting pages in the panelAppearance is query-dependent and can differ by location and device
Claude with searchBrand named in prose, sources listedGrounded answers cite; reasoning-only answers may notGood comparison surface because phrasing style differs from the others
Copilot / GeminiBrand named in the generated answerVaries by mode and groundingTreat as secondary unless your audience concentrates there

On cadence: weekly is the sweet spot for most teams. Daily runs generate noise that looks like signal and cost more than they teach. Monthly runs are too sparse to catch a drop caused by a competitor publishing a strong comparison page. Weekly gives you roughly thirteen observations a quarter, enough to see a trend line rather than a pair of dots.

One nuance that trips up new programs: run each prompt more than once per cycle. A single generation is a sample of size one from a stochastic system. Three runs per prompt per engine per cycle costs three times as much and gives you a per-prompt hit rate instead of a coin flip. If budget forces a choice, prefer three runs on thirty prompts over one run on ninety.

How do you score share of voice across AI answers?

Share of voice is the proportion of measured answers in which your brand appears, relative to the total brand appearances counted across your competitive set. Two formulas are common and they answer different questions.

Presence rate answers "how often do I show up at all." It is your mentions divided by the number of answers sampled. If you appear in 18 of 90 sampled answers, your presence rate is 20 percent. This number is easy to explain and it moves when you publish.

Share of voice answers "when a vendor is named, how often is it me." It is your mentions divided by all vendor mentions across the same answers. If those 90 answers contained 240 vendor mentions total and 18 were yours, your share of voice is 7.5 percent. This is the harder number and the more honest one, because it accounts for answers that name six competitors and answers that name one.

Track both. Presence rate tells you whether the category conversation includes you. Share of voice tells you whether you are winning inside it.

Then add the diagnostic splits that make the numbers actionable:

  • Citation gap rate. The share of answers where your domain is cited but your brand is not mentioned in the prose. A high gap rate means your content is good enough to inform the answer and not framed clearly enough to earn the recommendation. Usually the fix is explicit, self-contained statements about who the product is for, rather than more content.
  • Competitor concentration. Which two or three competitors take the majority of mentions in your prompt set. If one competitor dominates a specific prompt cluster, read the pages they own for that cluster.
  • Answer position. Whether you appear in the first sentence of the recommendation or in a trailing "other options" list. Simple string position within the answer body is a rough proxy and it is better than nothing.
  • Prompt cluster performance. Group prompts by job to be done and score each group separately. A brand can be strong on "how do I do X" prompts and invisible on "best tool for X" prompts, and those two failures have completely different fixes.

Define your counting rules once and write them down. Does "Skopx.com" count as a mention of Skopx? Does a mention inside a quoted user review count? Does a plural or possessive form count? These decisions are individually trivial and collectively decide whether your quarter-over-quarter comparison is real. A regex plus a documented exception list beats a vague rule applied by a different person each month.

What are the honest sampling caveats?

This is the section most measurement dashboards leave out, and leaving it out is how a measurement program loses credibility with the people funding it.

Generation is nondeterministic. The same prompt, same engine, same day can produce different brand lists. Your numbers have variance. Report a range or a rolling average, not a single precise-looking decimal. A presence rate that moved from 19.4 percent to 21.1 percent has probably not moved.

Grounding changes the answer more than your content does. Whether the engine performed a live search, and which pages that search returned, dominates the outcome. A competitor publishing one strong page can shift a whole prompt cluster in a week, independent of anything you did.

Personalization and context are invisible to you. Real users have chat history, memory, location, and account context. Your API-based or clean-session sampling has none of that. Your measurement is a controlled proxy for user reality, not user reality itself. Say so in the report.

API results and consumer-app results can differ. The model behind a consumer product may be configured with different tools, system prompts, and retrieval behavior than the same model exposed through an API. If you sample by API, label the numbers as API-sampled.

Absence of mention is not absence of influence. A user may read an ungrounded answer, then search your brand separately. Your measured number does not capture that path, which is one reason brand search volume and direct traffic belong on the same report. Brand mentions monitoring in the AI era covers the adjacent tracking that fills this gap.

Prompt sets encode your assumptions. You wrote the questions. If your prompt set omits the phrasing your actual buyers use, you will measure yourself as visible in a conversation nobody is having.

None of these caveats invalidate the measurement. They set the correct interpretation: this is a directional, repeatable index, useful for detecting change and comparing yourself with named competitors, not a precise audience metric.

How do you turn citation gaps into a content plan?

Measurement earns its budget when it changes what gets published. The most direct path runs through the citation gap list, because each gap is a specific answer where the engine already trusted your page and still recommended someone else.

Work the gaps in this order:

First, fix framing on pages already being cited. If an engine reads your page and cannot state plainly who the product is for, what it costs, and what it connects to, it will summarize your fact and recommend a vendor whose page says those things outright. Add a short, unambiguous statement of fit near the top. Put pricing on a page a crawler can read. State integrations by name.

Second, build the pages for prompt clusters where you have zero presence. If ten prompts about a job to be done never name you, you are missing the artifact that would make naming you easy. That is usually a comparison page, a specific how-to, or an integration page for a named tool.

Third, pursue third-party surfaces for the prompts where the cited sources are all review sites and forums. If every answer in a cluster cites the same three aggregators, no amount of your own content will move it. Your listings and reviews on those aggregators are the lever.

Fourth, re-run and wait. Indexing and re-grounding are not instant. Give a change two to four weekly cycles before you judge it. The AI search optimization checklist is a reasonable execution list once you know which cluster to attack, and what changes with LLM SEO explains why the framing fixes matter more than volume.

Keep technical health on the same report. A page that is slow, blocked, or broken cannot be cited regardless of how well written it is. Core Web Vitals and crawlability are table stakes for this work, not a separate project, and core web vitals monitoring plus a periodic SEO health score review keep that side honest.

What does a working measurement cadence look like?

A program that survives contact with a real calendar looks roughly like this:

Weekly, automated. Run the frozen prompt set across your chosen engines, three generations per prompt. Store the full answer text, the citation list, the engine, the timestamp, and the run identifier. Never store only the computed metric. Storing raw answers means you can recompute history when you improve your counting rules, and you will improve them.

Weekly, human, fifteen minutes. Read five answers where a competitor was named and you were not. Not the aggregate, the actual paragraphs. This is where you notice that the engine consistently frames the category in language your site never uses.

Monthly. Recompute presence rate, share of voice, citation gap rate, and cluster performance. Compare against the rolling three-run average, not against a single prior point. Publish one page: the trend, the two clusters that moved, and the one thing you are changing.

Quarterly. Review the prompt set itself. If buyer language has genuinely shifted, version the set, note the break in the series, and start a new baseline. Do not silently patch prompts.

Continuously. Watch competitor movement. Sitemap and pricing-page diffs tell you when a competitor ships the page that is about to take a cluster from you, often before the answers change. Skopx runs that competitor pulse alongside its share of voice reporting, and surfaces live Reddit and Hacker News threads where the category is being discussed, which is frequently where the aggregator citations originate.

Should you build this yourself or run it in a platform?

Both are legitimate. The build-versus-buy line falls in a predictable place.

Building it yourself is reasonable if you have engineering time, a small prompt set, and one or two engines to cover. The core is not complex: a scheduler, API calls, answer storage, a matcher, and a chart. The hidden costs are the parts that are not the core. Keeping API integrations working, handling rate limits and partial failures, deduplicating brand aliases, normalizing citation URLs, and keeping a year of raw answers queryable turn a weekend project into a maintained system. Teams that already run technical SEO automation and pull from the Search Console API usually have the muscle for it.

Running it in a platform makes sense when you want the measurement adjacent to the work it should trigger. Skopx bundles AI visibility with Site Health, which pulls Lighthouse scores through Google PageSpeed Insights, real-user Core Web Vitals through CrUX, Search Console performance, and an in-house on-page audit that produces a 0 to 100 score with a fix list. The same workspace connects nearly 1,000 business tools, so the output of a measurement run can feed a workflow, a document, or a publishing queue rather than sitting in a spreadsheet. Social Autopilot handles the distribution side, publishing to LinkedIn, Facebook Pages, Reddit, Instagram, X, Threads, Bluesky, Mastodon, Telegram, Discord, an email newsletter through your own Resend account, and the Skopx community feed, with content adapted per network character limit. Plans are $5 per month for Solo and $16 per seat per month for Team, and the AI runs on your own key with zero markup or on the included allowance. Details are on the pricing page.

Whichever route you take, the deciding factor is not the dashboard. It is whether the numbers reach the person who writes the next page, with the specific answer text attached. A share of voice chart with no linked evidence changes nothing. A citation gap list with five real paragraphs attached changes the editorial calendar on Monday.

Frequently Asked Questions

How many prompts do I need for a reliable measurement?

Thirty to sixty unbranded, buyer-intent prompts per product line, run three times per engine per cycle, is a defensible baseline for most companies. Fewer than twenty makes single-answer variance look like a trend. The number that actually matters is total generations per cycle, since that is what determines how stable your percentages are: sixty prompts at three runs across two engines is 360 observations, which is enough for weekly trend reading and not enough to justify reporting decimals.

Is AI visibility different from tracking AI citations?

They are related and worth separating. Citation tracking counts whether your domain appears in an answer's source list. Visibility measurement, as most teams use the term, counts whether your brand is named in the answer text itself, with citations as one input. The gap between the two is the useful part: pages cited without the brand being recommended point at a framing problem rather than a content-volume problem.

Can I track competitors as well as my own brand?

Yes, and you should, because share of voice is meaningless without a denominator. Define a competitive set of five to ten brands at the start, apply the same matching rules to all of them, and keep the list frozen alongside the prompt set. Competitor numbers also serve as a sanity check: if every brand in your set drops the same week, the engine changed, not your marketing.

How long does it take to see a change after publishing?

Plan on two to four weekly cycles before judging a change, and longer for pages that need to be discovered and indexed first. Grounded answers can only cite what retrieval surfaces, so the lag includes indexing time plus however long it takes the engine's retrieval to prefer the new page. Judging a publish after one run is the fastest way to conclude that nothing works.

Do numbers sampled through an API match what real users see?

Not exactly, and the report should say so. Consumer chat products carry personalization, memory, location, and product-specific system configuration that clean API sampling does not reproduce. Treat API-sampled numbers as a controlled index for detecting change and comparing brands, and label them as API-sampled rather than presenting them as a measurement of user experience.

What should I do first if my brand appears in almost no answers?

Check whether your pages can be read and cited at all before changing any copy: crawlability, speed, and indexation come first, since an uncitable page cannot be recommended. Then read ten answers from your weakest prompt cluster and note the exact language the engine uses to frame the category. Most zero-presence cases resolve into one of two things, a technical block or a vocabulary mismatch between how you describe the product and how buyers describe the problem.

Share this article

Skopx Team

The Skopx engineering and product team

Related Articles

Stay Updated

Get the latest insights on AI-powered code intelligence delivered to your inbox.