AI Search Jul 16, 2026 16 min read

AI Visibility Rankings Are Mostly Noise (Until You Prove Otherwise): A Practical Playbook For SMEs, Brands, And Agencies

AI search “visibility” dashboards can swing between runs, making winners and losers look real when it’s often just randomness. Here’s how to measure AI visibility like an analyst, report it like a business leader, and operationalize improvements with approved execution.

Featured image for AI Visibility Rankings Are Mostly Noise (Until You Prove Otherwise): A Practical Playbook For SMEs, Brands, And Agencies

AI search visibility has a measurement problem—and it’s bigger than most dashboards will admit. If you’ve ever refreshed a report and watched your “rank” move, you weren’t necessarily watching your brand rise or fall. You were often watching randomness.

Generative AI systems (ChatGPT-style answers, AI Overviews-style summaries, and AI Search engines that cite sources) frequently produce different citations and different brand recommendations across repeated runs. That means “AI visibility rankings” can drift between measurements even when nothing changed on your website.

New research covered by Search Engine Journal highlights what many practitioners suspected: a single reading is mostly statistical noise. The more important contribution is practical: a method (a “stopping rule”) for deciding when you have enough repeated measurements to treat a Ranking as trustworthy.

This matters for every operator—SME owners, ecommerce teams, multi-location brands, and agencies—because when measurement is unstable, decision-making becomes unstable. You end up “optimizing” based on a mirage, overreacting to normal variance, and attributing wins (or losses) to work that didn’t cause them.

In this editorial, I’ll translate the research into a plain-English playbook: what changed, why it matters, how to report AI visibility honestly, and how to build an execution system that improves what AI systems are likely to cite—without chasing noise. I’ll also show where AYSA fits: as a monitored, approval-based execution layer that prepares changes, asks for your sign-off, and then implements accepted improvements.

Concise summary

Three AI search runs showing different cited sources for the same question, illustrating variability.
Same question. Different citations. That’s why a single AI visibility reading is a snapshot—not a fact.
  • AI visibility rankings are unstable by default because generative systems include randomness and can cite different sources run-to-run.
  • A single number is rarely safe to interpret. Treat it like a snapshot, not a fact.
  • Trust requires two things at once: the ranking order stabilizes and top sites are separated by more than the margin of error.
  • Different platforms require different sampling. The same budget can be “enough” for one engine and meaningless for another.
  • Execution still matters—but the workflow must include repeatable measurement, careful reporting, and approved implementation.

Key takeaways (for busy leaders)

Whiteboard sketch showing stable rankings and a clear gap beyond uncertainty bands.
Trust comes from repeatability and a gap that’s larger than the uncertainty—not from a clean-looking rank.
  1. Stop presenting AI visibility like traditional SEO rank. Present ranges, uncertainty, and “insufficient data” when appropriate.
  2. Measure before/after more than once. One baseline and one post-change reading cannot prove improvement.
  3. Over-Index on what you can defend: leadership positions and clear gaps; treat “middle ranks” as approximate.
  4. Build a repeatable workflow: prompt sets, controls, run cadence, and a stopping rule—then connect that to Approved Execution.

Table of contents

Checklist workflow for deciding when AI visibility sampling is sufficient to trust rankings.
A stopping rule is a business control: it tells you when to keep sampling and when to trust the story.

The Problem: AI Answers Are Stochastic, So Rankings Bounce

In classic SEO, rankings change because Google changes the results, competitors change their pages, or you change yours. There’s volatility—but it’s anchored to something observable in the SERP.

In AI-driven experiences, there’s an extra layer: even if the underlying information sources and the model were identical, the system may still vary the output between runs. Many generative systems intentionally include randomness so responses aren’t identical every time. That creates a specific measurement trap: if you ask the same question repeatedly, you can get different answers and different cited sources—even when nothing “happened.”

That’s the core reason AI visibility dashboards can be misleading when they present a single clean number. A clean number implies certainty. But the system you’re measuring is probabilistic.

Search Engine Journal’s coverage of the latest IQRush preprint describes repeated-query testing across platforms and topics, showing that it can take dozens of cited answers—and sometimes more—to reach a point where rankings are stable and meaningfully separated. In some tests, even a large number of queries didn’t produce a clean separation at the top (especially where top sites were very close). That’s not a failure of your brand. It’s a feature of the measurement environment.

What Changed: Measurement Finally Became the Story

Most “AI visibility” content over the past year has focused on tactics: be cited, be summarized, be recommended. But there’s a more fundamental question: Can we even trust what the tracker is telling us?

This is where the new wave of work is useful. Instead of treating AI citations as fixed, it treats them like samples from a distribution. That framing matters because it changes your KPI philosophy:

  • You stop asking, “What is our rank today?”
  • You start asking, “Given repeated samples, what range of outcomes is plausible—and is the difference between us and competitors bigger than the uncertainty?”

That’s the difference between a dashboard and a measurement system.

It also aligns with something Rand Fishkin has emphasized publicly in related discussions: vendors should “show their math.” If your AI visibility provider can’t explain sampling, variability, and confidence, they may be selling certainty they don’t actually have. (SEJ references this point in the context of the new research.)

The New Insight: “Stability + Separation,” Not a Single Number

The most practical idea in the SEJ-covered paper is that a ranking becomes trustworthy only when two conditions are true at the same time:

  1. Stability: the order stops changing as you collect more answers.
  2. Separation: the top sites are far enough apart that the gaps are larger than the margin of error.

These conditions sound simple, but they correct a widespread reporting mistake: teams often accept “stable order” as proof. But stability alone can still be misleading if the leaders are bunched together and the uncertainty overlaps.

Separation matters because it distinguishes “real advantage” from “coin flip.” If your visibility share is 9% and a competitor’s is 6%, that looks like a 50% lead. But if the uncertainty range overlaps, that apparent lead isn’t a business fact—it’s an artifact of sampling.

In other words: a rank can look decisive while being statistically non-decisive.

The implication is uncomfortable but healthy: sometimes the honest answer is, “We can’t call it yet.” If your measurement system can’t say that, it’s not measuring—it’s performing.

A Plain-English “Stopping Rule” You Can Use Without Being a Statistician

A stopping rule is a decision rule for when you have collected enough samples to stop sampling—because additional runs are unlikely to change the conclusion in a meaningful way.

Here’s a plain-English stopping rule that matches the spirit of the research discussed by SEJ:

  • Keep running the same prompt set repeatedly (with consistent controls) and record which sites are cited/recommended.
  • After each batch (e.g., every 10–20 cited answers), recompute your ranking and the uncertainty for top entities.
  • Stop only when (a) the order of the leaders no longer changes across additional batches and (b) the gaps between leaders are larger than your uncertainty.
  • If you can’t reach that point within budget, report “insufficient data” rather than publishing a forced ranking.

The research summary cited by SEJ indicates that across multiple tests, the number of usable, citation-containing answers required to reach that trustworthy zone varied substantially. That’s the key operational lesson: there is no universal “just do 20 prompts” rule that works everywhere.

For business owners, the takeaway isn’t to become a statistician. It’s to demand that your measurement process:

  • uses repeated sampling,
  • shows a range (not just a point estimate), and
  • has a concept of “not enough data.”

Why Platforms Differ: Citations Aren’t Equal to Information

One of the most counterintuitive points highlighted in the SEJ write-up is that the platform you measure changes how much data you need, and not necessarily in the way you’d guess.

It’s tempting to assume: “If an engine gives more citations per answer, then I need fewer answers.” But the research suggests what matters is the amount of independent information each answer provides.

If a platform tends to cite the same sources repeatedly within a single response, those citations can be redundant. You get “more citations,” but fewer new insights per run. Another platform might cite fewer sources per answer but vary them more across runs, providing more independent signal.

So when someone tells you “we ran 50 prompts,” a smart follow-up is: 50 prompts on which platform, in which topic, and how many of those runs produced usable citations?

This also means budgets should be framed around confidence, not activity. Measuring is not producing a number; measuring is estimating a distribution with enough precision to make decisions.

What Goes Wrong in Real Businesses (And Why It’s Expensive)

When measurement is noisy, the biggest risk isn’t that you’ll be wrong on a spreadsheet. The risk is that you’ll build a business process around reacting to noise.

Here’s what that looks like in practice:

1) False wins create the wrong incentives

You publish a new “AI-optimized” page, check a dashboard once, and see your citation share increase. Everyone celebrates. You double down on the same tactic. But if that change was within natural variability, you’ve just rewarded a random outcome. Over time, this trains teams to chase patterns that aren’t causal.

2) False losses trigger unnecessary pivots

You change nothing, but your brand drops from #2 to #6 in the report. Leadership panics. You pull resources from conversion work, email, or paid search to “fix AI.” Weeks later, the number rebounds on its own. You didn’t fix anything—you paid a volatility tax.

3) Bad attribution breaks your experimentation culture

Teams start “testing” without controls. If the metric can’t separate effect from noise, you can’t learn. You can only narrate.

4) Agency-client conflict escalates

Agencies are judged on movement. Clients see movement and assume it’s proof. When movement reverses, trust breaks. Both sides would be better served by a measurement contract that recognizes uncertainty and defines what “enough evidence” looks like.

5) Reputation risk: publishing rankings you can’t defend

If you’re a publisher, SaaS company, or agency producing “AI visibility reports,” unstable rankings can create reputational damage. Your audience will notice when your list changes dramatically without explanation. The right move is to publish uncertainty and methodology—like mature analytics disciplines do.

SME Scenario: A Clinic That “Won” AI Visibility (But Didn’t)

Let’s make this tangible with a scenario that mirrors what I see across SMEs.

Business: a local dermatology clinic with two locations.

Goal: show up when people ask AI tools questions like “best dermatologist for acne scars near me” or “what clinic treats melasma.”

What happens:

  1. The clinic updates a service page, adds FAQs, and improves their location pages.
  2. They run an AI visibility tool once and see they’re cited 3 times in a set of prompts. A competitor is cited once.
  3. They conclude the work “paid off” and allocate budget to produce 30 more pages in the same style.

The problem: Without repeated sampling, this is indistinguishable from randomness—especially if citation counts are low. If the platform is volatile, a single run can produce a flattering sample. The next run might flip the outcome.

The correct business approach:

  • Run a baseline repeatedly (multiple batches) and record a range.
  • Make the change.
  • Run post-change repeatedly and record a range.
  • Only claim improvement when the post-change range is clearly better than the baseline range.

That sounds “slower,” but it’s actually faster at the business level because you stop building strategy on false signal. You invest where you can prove movement.

What Agencies Need to Rethink (Contracts, Reporting, Accountability)

For agencies, AI visibility volatility isn’t just a measurement issue—it’s a commercial issue. If you sell “rank improvements,” you’re selling something that may not be stable enough to productize the way classic SEO ranking reports were.

Here’s what should change in agency operations if you want to stay credible:

Report ranges and uncertainty, not just rank

Leadership-friendly reporting doesn’t mean simplified to the point of dishonesty. It means you translate uncertainty into risk-aware decisions.

Instead of:

  • “You are #3 in Perplexity for X.”

Use:

  • “Across repeated runs this month, you typically appear between #2 and #6 for X; the top 3 brands are not clearly separated yet.”

That kind of statement is defensible, and it reduces churn when the next run looks different.

Define “confidence” in the SOW

Put the idea of sampling into the agreement:

  • What constitutes a “measurement period”?
  • How many runs are required before declaring a result?
  • What do you do when the result is inconclusive?

Without this, agencies are set up to be punished for randomness.

Separate monitoring from execution (but connect them)

A lot of “AI visibility” offerings are monitoring-only. But clients don’t pay for monitoring; they pay for outcomes. You need an execution layer that turns insights into site changes safely and efficiently.

This is where tooling and process matter, and it’s exactly why we built AYSA around monitored, approval-based execution.

A Practical Measurement Framework: Prompts, Controls, Runs, Ranges

If you’re an SME or an agency, you don’t need a research lab. You need a repeatable process. Here’s the framework I recommend, grounded in the research logic described by SEJ.

Step 1: Define prompt sets that mirror real intent

Create a prompt library that maps to your revenue:

  • “Best [service] near [city]” (local intent)
  • “Which [product type] is best for [use case]” (commerce intent)
  • “Alternatives to [competitor]” (consideration intent)
  • “How to [problem]” (education intent)

Keep prompts stable across runs so you’re measuring change, not rewriting the test.

Step 2: Add controls so you don’t fool yourself

Controls are the difference between measurement and storytelling.

  • Control prompts: prompts you don’t expect to change (or that target a stable topic) to detect platform-wide drift.
  • Holdout set: a subset you don’t use for day-to-day reporting, reserved for validation.
  • Standardized settings: same language, region, and any available personalization toggles.

The SEJ source context also references the idea of enterprise testing methodologies for multiple platforms (ChatGPT, Perplexity, Claude, Google) and prompt control groups in the ecosystem around this discussion. Even if you’re not enterprise, the principle scales down: control what you can.

Step 3: Run repeatedly and track distributions

Don’t “check” AI visibility. Sample it.

  • Run the same prompt set multiple times per measurement window.
  • Record whether you were cited, where you were cited (if the platform provides ordering), and who else was cited.
  • Compute a range (or at least show variability) rather than a single point estimate.

Step 4: Apply a stopping rule before you publish conclusions

Only publish rankings when order stabilizes and leaders are separated beyond uncertainty. Otherwise, report the result as “inconclusive” and keep sampling.

Yes, this can feel uncomfortable. But it’s also what mature measurement cultures do. In paid media and analytics, we accept that some tests are underpowered. AI visibility needs the same honesty.

Step 5: Tie AI visibility to business outcomes carefully

A major limitation raised in the SEJ write-up is that the ecosystem still lacks perfect attribution plumbing—particularly because traditional tools don’t always label “AI” as a referrer or channel in a clean way.

So do two things:

  • Use AI visibility as a leading indicator, not the only KPI.
  • Track business outcomes (leads, sales, bookings) and annotate changes in GA4 so you can correlate responsibly.

If you’re unsure how to structure that monitoring, start here: AYSA Monitoring.

From Measurement to Execution: What to Actually Change (That AI Systems Can Use)

Once you accept that AI visibility is probabilistic, the goal changes from “win the rank” to “increase the probability that you’re selected as a source.” That probability is shaped by the inputs AI systems can access: your pages, your structured data, your authority signals, and your consistency across the web.

Here are the execution areas that tend to matter for AEO/GEO work—without pretending any single tactic guarantees a citation.

1) Build content that is citeable

  • Clear definitions, direct answers, and scannable sections.
  • Unique, first-party details (pricing approach, process steps, policies, constraints).
  • Evidence and specificity (but don’t invent claims; cite sources when you make factual assertions).

AI systems gravitate toward content that is easy to extract and summarize. That’s not “AI tricks.” That’s good publishing.

2) Strengthen entity consistency (brand, products, people, locations)

For local and multi-location businesses, inconsistent location data is a visibility killer across both classic search and AI summaries. Align name/address/phone, services, and key attributes across your site.

Learn more about how we think about AI search visibility and entity signals: AYSA AI Search Visibility.

3) Ensure technical accessibility and clarity

  • Indexable pages, clean canonicals, and logical internal linking.
  • Structured data where appropriate (organization, product, FAQ, local business) implemented correctly.
  • Fast, reliable pages that users can actually consume if they click through.

Even if AI answers reduce clicks, being the cited source still sends qualified traffic and brand lift. And when clicks happen, experience matters.

4) Build authority the boring way (and don’t fake it)

AI citation systems often reflect broader web signals. If you want to be cited, you typically need to be a credible source. That comes from:

  • Real PR and mentions,
  • Partner pages and associations,
  • Reviews and reputation where relevant,
  • Original research or genuinely useful resources.

Be careful with shortcuts. If your visibility system is noisy, it’s easy for bad actors to sell you “wins” that are actually sampling artifacts.

5) Adopt update discipline: fewer changes, better measured

Noisy measurement punishes teams that change everything at once. If you can’t isolate cause, you can’t learn. Make fewer, higher-quality changes and measure them with repeat sampling.

Where AYSA Fits: Monitor → Prepare → Approve → Execute

Most businesses don’t fail at AI visibility because they lack ideas. They fail because they lack an execution system that turns insights into changes safely, consistently, and at speed.

AYSA is designed to be that execution system:

  • Monitors your search presence and site signals over time (so you can observe change with context): AYSA Monitoring
  • Prepares recommended website improvements that are relevant to SEO/AEO/GEO work.
  • Asks for approval before implementing changes—so you maintain control and governance.
  • Executes accepted changes to reduce the lag between “we learned something” and “we fixed something.”

That approval-based model is critical in AI search because volatility increases the temptation to overreact. A good system slows down conclusions (by requiring enough evidence) while speeding up implementation (once you decide to act).

If you’re exploring the toolset, start with: AYSA AI SEO Tools. If you want to understand how we approach visibility in AI answers specifically: AI Search Visibility. For commercial details: Pricing. And for more editorials like this: Blog.

What to do next (action list)

  1. Audit your current AI visibility reporting. If you’re using single readings, mark them as directional only.
  2. Build a prompt library tied to revenue. Keep it stable for at least a month to establish baseline variability.
  3. Run repeated samples. Establish ranges, not point estimates.
  4. Adopt a stopping rule. Don’t publish a ranking until it’s stable and</em meaningfully separated.
  5. Change your leadership narrative. Replace “rank” with “probability of citation” and “confidence level.”
  6. Implement improvements with governance. Use an approval-based execution workflow so you don’t thrash.
  7. Track outcomes alongside visibility. Use GA4 annotations and business KPIs so you don’t optimize for citations alone.

Sources and further reading

Note: The SEJ coverage references additional papers and vendor research, but those primary documents are not included in the supplied context here. Where you need hard numbers or formal methodology for your industry, insist on primary-source documentation from your measurement provider and treat single-run dashboards as directional.

Related AI SEO resources

Continue the AI search topic inside AYSA.

Use these pages to connect the article with AI SEO tools, AI visibility monitoring, AI Overviews and approved website execution.

Execution hubs

Turn this topic into a website action plan.

Use these AYSA hubs to move from reading to technical fixes, AI visibility monitoring, research, glossary context and approval-first SEO execution.

Marius Dosinescu, author at AYSA.ai

Written by

Marius Dosinescu

Marius Dosinescu is the founder of AYSA.ai, an entrepreneur focused on SEO automation, ecommerce growth, authority building and approved website execution for businesses that want organic growth without specialist overhead.

SEO execution, not more busywork

Turn SEO reading into approved website action.

AYSA monitors your website, prepares the work, asks for approval, and executes approved changes inside your website.

Start now View pricing

Only €29 to €99 per month, depending on the size of your business.

AYSA SEO Magazine

Latest search intelligence.

View all articles