Skip to content

AI Visibility: What It Is, How to Measure It and How to Improve It

What AI visibility is, how to measure it with a repeatable baseline across ChatGPT, Gemini and Perplexity, how to monitor brand mentions week to week, and how to improve it.

Lourdes Paul Agilan, founder of Aquinel

Founder, Aquinel · Published · 11 min read · Updated

AI visibility is how often, where and how accurately AI tools like ChatGPT, Gemini and Perplexity name your business when buyers ask questions. To measure it, run 20 to 40 real buyer questions several times on each engine and log whether you are named. Repeat weekly. To improve it, get mentioned on other sites, clarify your pages and answer those exact questions.

Measure first, fix second. That order matters. Here is why.

What is AI visibility?

In August 2026 the IAB published a framework called Measuring Visibility in the AI Era. One line in the announcement is worth sitting with: more than 20 companies now sell AI visibility measurement tools, each using a different method, so the same brand can get different answers from different vendors on the same day.

The IAB splits visibility into four parts. Presence is whether you get mentioned at all. Prominence is where in the answer you land. Portrayal is whether the AI describes you correctly. Persuasion is whether any of it makes someone act. Most tools sell you Presence and call it visibility.

The framework also draws a useful line. Some measurement is directional, good enough to show a trend but not to move budget on. Some is decision grade, meaning it meets a bar for sample size and repeatability. Caroline Giegerich, the IAB's VP of AI, put it plainly: people are discovering brands inside AI tools, and measurement has not kept up.

Why a single check tells you nothing

If you ask ChatGPT a question today and it names you, that is not a result. It is one draw from a dice roll.

The cleanest test of this came from the University of St. Gallen. Julius Schulte, Malte Bleeker and Philipp Kaufmann tracked four industries across ChatGPT, Gemini, Google AI Mode and Perplexity for about six weeks in early 2026. Day to day, the sources the AI pulled from overlapped only 34 to 42 percent. Brand mentions were steadier at 45 to 59 percent. Then they ran the same prompts again minutes apart and saw the same instability, so the wobble comes from the model itself, not from anything changing in the world.

A second paper, from Ronald Sielinski in July 2026, asked how much data you need before a ranking of brands stops moving. Across 30 platform and topic combinations on Gemini, SearchGPT and Perplexity, he found it takes between 33 and 94 answers before the order settles. Three of the 30 tests never settled at all, even after 125 questions, and all three were on SearchGPT. Worth noting: his questions came from ChatGPT rather than real searches, and he sells software in this space.

So when a tool hands you a tidy score, ask how many runs sit behind it. If the answer is one run per prompt, the number is noise wearing a suit.

How many runs? The experts disagree

The St. Gallen team says run each prompt at least seven times a day for brand tracking, eight for source tracking, and look at a rolling two to four week window. Nick Lafferty, who keeps a public metrics reference updated through August 2026, says the same thing in plainer words: a 50 prompt set run 10 times tells you more than a 500 prompt set run once.

Dmitrij Zatuchin disagrees. In July 2026 he broke 12,933 AI responses covering 20 brands, eight languages and three models into buckets of variation. Repeating the same prompt explained 34.8 percent of the wobble and the language of the question explained 31.6 percent, while the model used explained only 1.7 percent. His conclusion: a repeat past the fifth buys almost nothing, and those queries are better spent on more languages and contexts. He is affiliated with a company in this market, and his brands were all Central and Eastern European, so it is not a like for like test.

A second Zatuchin study, in September 2026, pulls the other way. He put 50 questions through six engines 15 times each, 4,500 responses in all. A single run showed only 62% to 77% of the brands that five runs found, and on engines answering without web search, new brands were still appearing at run 15.

Nobody has reconciled these. Our read: five runs per prompt is enough to stop fooling yourself, and anything past ten is money better spent widening your question list.

How to measure AI visibility yourself

You do not need a subscription to start. Open a spreadsheet.

List 20 to 40 questions a real buyer types, not the ones you wish they typed. Your sales team knows them. Run each one five times in ChatGPT, Gemini and Perplexity, in a fresh chat each time, with memory and personalisation off so you are not scoring yourself against your own history. Log two things: were you named, and were you linked. Repeat next week.

A paper published on 14 September 2026 by Edward Malthouse and colleagues tested brand recommendations across six models and five product categories. Broad category questions left out well established brands surprisingly often, but those same brands often reappeared once the question carried a specific situation. So test both shapes. Ask "best IT support company" and also "IT support for a 40 person law firm that has to stay HIPAA compliant."

After a month you have a baseline built on a method you understand, and you will read any tool you later buy far better for it. If you would rather someone else run that first pass, that is what our audit is for.

How to monitor brand mentions week to week

Measuring once gives you a baseline. Monitoring is the weekly habit that tells you whether anything changed.

Freeze your prompt set. The most common mistake is rewriting your questions between runs. Do that and you measure your own editing, not the market. Peec AI, which sells a tracking tool, ran 37,804 AI responses across 1,754 prompts on five engines in June 2026. Once wording drifted far enough apart, visibility fell by 2.40 percentage points, about half. Kazem Faghih and five co-authors found in May 2026 that 13 models flipped between right and wrong answers on reworded versions of the same question, with mismatch rates above 23%. Write your list once. If you must add a prompt, add a new row rather than editing an old one.

Run it on the same day, at the same time. Paul Tschisgale and Peter Wulff sent the same question to GPT-4o every three hours for nearly 88 days, 6,930 queries in total. About 20% of the variation was periodic, with cycles at 7.3 days and 5.5 days. It was one physics question on one model, not a brand question. But the lesson is cheap: run your set at a fixed hour on a fixed weekday so time is not one of the variables.

Log one row per answer. Columns: date and time, engine, the prompt copied exactly, named yes or no, where in the answer you appeared, which competitors were named, which domains were cited, and one line on how you were described. Keep the raw answer text in a second sheet. The citation column is the one that pays. It tells you which pages the engines lean on, and those pages you can go and influence.

Decide when a change is real. Treat a change as real only if it does at least two of three things: holds for two weeks running, shows up on more than one engine, and is bigger than the swing you see in quiet weeks. That last one needs a baseline. Run your set for four weeks while changing nothing on your site. Whatever range you see is your noise floor. Anything inside it is weather.

Tool scores need the same caution. David Nelson ran the same branded prompt through six visibility tools in August 2026. The four that gave a percentage said 100%, 90%, 75% and 16.7%. A number labelled "visibility" is not always measuring the same thing from one product to the next. We compare buying a tool with hiring help in tool or agency.

The numbers on your own site that nobody can fudge

Prompt logs tell you about the AI. Your own analytics tell you about your business. Wil Reynolds of Seer Interactive argued in February 2026 that visibility scores are a vanity metric until you tie them to something real, and pointed at direct traffic, branded search and where AI visitors land.

Search Console has a generative AI performance report, rolled out worldwide on 31 August 2026. Impressions only, no clicks, and no way to pull AI Overviews apart from AI Mode. Limited, but free and yours.

Your analytics referral report shows visits from chatgpt.com, perplexity.ai and gemini.google.com. Expect the volume to look tiny. Seer's widely quoted case study found AI traffic was 0.07 percent of organic sessions on the one site they studied, though it converted at 15.9 percent against 1.76 percent for Google organic. One client, 2025 data, and the multiples the industry quotes run from 4x to 23x depending on who is selling. Something is there. The size is unsettled.

Branded search volume is third. If more people hear your name inside AI answers, more of them later type it into Google. Slow signal, but hard to fake.

How to improve AI visibility

Once you have a baseline, the work is less exotic than the tools suggest.

Get mentioned elsewhere. Ahrefs, in vendor research covering 75,000 brands, found branded web mentions had the strongest link to AI Overview visibility of anything they measured, ahead of backlinks. They stress it is correlation, not proof. It points the same way as the Malthouse paper, which found prominence tracked how much a brand was discussed and searched for rather than how big it was.

Answer the exact questions in your log. Not variations. If three prompts keep naming your competitor, write the page that answers those three prompts better than the page the AI is currently quoting. Our guide on how to optimise your website for AI search covers the page level work, and our services page covers what we do when you want help with it.

Then wait properly. Data from about 900 marketing pages tracked between March and May 2026 put the median time to a first AI citation at about seven days, with a long tail past a month. That comes from a vendor platform, so hold it loosely. Checking on day two proves nothing.

Where to start this week

Build the question list and freeze it. Run each prompt five times across three platforms, at the same hour on the same weekday. Log named, linked and cited domains. Open your Search Console AI report and write down today's impressions. Then leave it alone for a week. You now have what almost nobody in your market has: a before.

Questions

How often should I check my AI visibility?

Weekly is enough for most businesses, monthly if your market moves slowly. Daily checking shows you swings that are just the model being random. The St. Gallen research found day to day source overlap as low as 34 percent when nothing had changed.

Do I need to buy an AI visibility tool?

Not to start. A spreadsheet and a few hours give you a real baseline. Buy a tool when the manual work gets too big, and when you do, ask how many times they run each prompt and whether they report a margin of error. The IAB calls anything failing that bar directional rather than decision grade.

Why does my AI brand visibility score differ between two tools?

Because they ask different questions, different numbers of times, on different engines, and add them up differently. The IAB noted in August 2026 that more than 20 vendors each use their own method. One test that month got scores from 16.7% to 100% for the same brand and prompt. Pick one method and track your own trend.

My brand mentions dropped 10% this week. What do I do?

Nothing yet. Check whether it held the next week and whether it shows on more than one engine. If it did both, look at the cited domains in your log rather than your own pages, because usually the source the engine leans on changed, not you.

Is AI visibility worth measuring if the traffic is so small?

The traffic is small today, the quality signals look good, and the numbers are contested. Measure it because measuring is cheap, and because branded search and direct traffic will tell you whether it is landing. Do not rebuild your marketing around a percentage nobody can reproduce yet.

Sources

  • Measuring Visibility in the AI Era, announcement and framework, IAB, 3 August 2026. iab.com
  • Don't Measure Once: Measuring Visibility in AI Search (GEO) (Julius Schulte, Malte Bleeker and Philipp Kaufmann, University of St. Gallen), arXiv 2604.07585, 10 April 2026.
  • From Stochastic to Stable: Rank Stability and Structural Sufficiency in AI Visibility Measurement (Ronald Sielinski, co-founder of IQRush), arXiv 2607.10341, 11 July 2026. Headline figures read via Search Engine Journal, AI Visibility Rankings Aren't Stable, 11 July 2026.
  • Where Does the Noise Come From? A Variance-Components Decomposition of Non-Determinism in LLM Brand Answers (Dmitrij Zatuchin, affiliated with Rankfor.AI), arXiv 2607.13304, 14 July 2026.
  • Sampling Completeness in Generative Search: Brand and Cited-Domain Accumulation under Repeated Queries (Dmitrij Zatuchin, affiliated with a company selling in this market), arXiv 2609.05059, September 2026.
  • Evaluating Brand Retrieval and Ranking in Large Language Model Recommendations (Edward Malthouse, Kun-Yu Lee, Jing Yang, Sanchary Pal and Xueyan Feng), arXiv 2609.16304, 14 September 2026.
  • AI Visibility Metrics: Formulas, Benchmarks and Sample Sizes (Nick Lafferty), published 9 June 2026, updated 20 August 2026. Time to first citation figures drawn from Profound datasets, a vendor source.
  • Peec AI prompt variance study, reported by Search Engine Journal, 15 June 2026. Vendor research.
  • Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy (Faghih, Cheng, Saha, Pournemat, Gerami and Feizi), arXiv 2607.22554, May 2026.
  • Daily and Weekly Periodicity in Large Language Model Performance and Its Implications for Research (Tschisgale and Wulff), arXiv 2602.15889.
  • Can You Trust AI Search Visibility Scores? I Tested the Same Prompt Across 6 Tools (David Nelson), marketingwithdave.com, 26 August 2026.
  • AI Visibility Is a Vanity Metric (Wil Reynolds), Seer Interactive, 3 February 2026.
  • Case Study: How Traffic from ChatGPT Converts (Nick Haigler and Garman Chan), Seer Interactive, 3 June 2025. Single client, GA4 data from 1 October 2024 to 30 April 2025.
  • An Analysis of AI Overview Brand Visibility Factors (Louise Linehan and Xibeijia Guan), Ahrefs, 26 May 2025. Vendor research, 75,000 brands, Spearman correlation.
  • Generative AI performance report, Google Search Console Help, accessed 18 September 2026. Global rollout stated as 31 August 2026.
  • September 2026 Google Webmaster Report, Search Engine Roundtable, accessed 18 September 2026.

If you would like us to walk you through how this would work for your firm, the call is free.

Lourdes Paul Agilan, founder of Aquinel

About the author

Lourdes Paul Agilan

Founder of Aquinel. Works with MSPs and IT services firms from Chennai, India, on getting recommended by ChatGPT and found on Google.

LinkedIn

See how this would work for your firm.

Free Consultation Call, No Obligation