AI Search

11 min read

An AI Visibility Score and an AI Visibility Audit Answer Different Questions

An AI visibility score ranks you against peers on comparable terms. An AI visibility audit tells you whose framing of your category the engines are citing.

The Industry Just Standardized How to Measure AI Visibility. The Score Still Won't Tell You Why You're Losing.

IAB just released "Measuring Visibility in the AI Era," a framework built with a working group that includes Walmart, Acxiom, Microsoft, WPP Media, eMarketer, and Tinuiti (IAB, 2026). The pitch: one set of definitions, comparable across brands, in a category where more than 20 vendors have been selling the same idea under 20 different methodologies. That part is real progress. Marketing leaders finally get a shared vocabulary instead of a stack of incompatible dashboards, each claiming to measure the same thing.

But a standardized score answers one question: where you rank. It was never built to answer the question that actually costs pipeline: why is an AI engine citing a competitor's framing of your category instead of yours. Standardizing the number doesn't explain the gap underneath it. It just makes the gap easier to compare across companies who are all equally in the dark about its cause.

That distinction, score versus diagnosis, is the whole story here.

What does IAB's new 4 P's AI-visibility framework actually measure?

IAB's framework scores four dimensions: Presence (are you mentioned at all), Prominence (how often, how high), Portrayal (what the AI says about you), and Persuasion (does the mention change behavior). It exists because 60 to 75% of senior buy-side leaders say advanced measurement, attribution, incrementality testing, MMM, falls short on rigor, timeliness, and trust (IAB, State of Data 2026).

The push for a standard didn't come from nowhere. Organic click-through for queries that trigger an AI Overview fell 61%, from 1.76% to 0.61%, across 3,119 tracked terms and 42 organizations (Seer Interactive, 2025). That number is why visibility stopped being a nice-to-have metric and became something marketing leaders are asked to defend in board meetings. When the AI answer itself absorbs the click, ranking on a traditional results page stops mapping to anything that shows up in pipeline.

IAB's four dimensions give every brand shared vocabulary for that reality. Presence tells you if you exist in the answer at all. Prominence tells you how much room you occupy relative to competitors. Neither tells you what the answer actually says about you, which is where the real exposure sits. That's the same gap between production speed and measurement rigor covered in AI Made Marketing Production Faster. It Didn't Make Measurement Any Smarter. AI search just added a second layer on top of it.

None of that is a knock on the framework itself. Turning 20 competing methodologies into four shared definitions is real progress, and it gives procurement teams a way to compare vendors without taking each vendor's own scorecard at face value. The problem isn't what IAB built. It's what a standardized score was never designed to hold. A score is a snapshot: where you rank today, on a defined set of queries, against a defined set of competitors. It says nothing about the sentence an AI engine actually returns when a buyer asks it a direct question, and that sentence is where the category gets defined, correctly or not.

A score that responds to a real traffic problem is still just a score. The next question is what it can't see when a competitor starts pulling ahead.

Why doesn't a visibility score explain why you're losing ground to a competitor?

Brands that show up in ChatGPT recommendations get 2.5 times more site visits within a week than unrecommended peers (Similarweb, 2026). That's Presence and Persuasion appearing together in the data. But 55.9% of those visits arrive through branded search, so they read as organic in your dashboards while the AI conversation that actually drove them stays invisible.

That masking matters because the stakes underneath a visibility score are steeper than a ranking. Gartner's framing, cited in a recent Demand Gen Report feature, is blunt: if a brand isn't visible to answer engines, "you don't make the shortlist" at all, you're "being removed from consideration altogether" (Demand Gen Report, 2026). A 4 P's score can confirm you're losing Share of Voice to a named competitor. It can't tell you whether that competitor is winning because their content is genuinely better, because a review site trusts them more, or because the AI learned an outdated version of your positioning from a directory nobody at your company has looked at in a year.

Those are three different fixes. The score returns one symptom for all three. Read the AI buying agents piece for what that shortlist mechanism looks like from the buyer's side.

There's a reason this matters more than a missed click. A B2B buying committee doesn't run one AI query and stop. They run several, across different tools, over weeks, forming an impression before anyone on your sales team knows a deal exists. A visibility score tracks whether you showed up in that process. It doesn't track whether the impression the committee formed matches what you'd want them to walk away believing. Those are different failure modes, and they require different fixes: one is a distribution problem, the other is a message problem.

That branded-search masking is one symptom of a bigger mechanism: the AI is letting someone else define the category you compete in.

What is categorization drift, and why can't a scoring framework catch it?

About 85% of AI citations for broad B2B category queries come from third-party sources, review sites, analyst reports, directories, not the vendor's own site (Rampiq, 2026). That means someone else wrote the category definition the AI repeats to a buyer. Categorization drift is what happens when that outside definition stops matching how you describe your company.

This is the mechanism a 4 P's score structurally can't see. Portrayal measures sentiment: is the mention positive, neutral, negative. It doesn't measure whether the sentence describing what you do still matches your own homepage. A company selling workflow automation might find an AI engine describing it as "a project management tool" because that's the framing a widely cited analyst report used a while back, and that report outranks the company's own site in the citation pool by the 85% margin above, across more than 30 B2B brands analyzed. The sentiment score on that mention could read positive.

The company would still be losing every deal where the buyer's actual need was workflow automation, not project management. That's not a hypothetical edge case. It's the default state of any category where the loudest third-party source got there first. The same content-versus-credibility gap shows up in how AI redistributes competitor content without redistributing trust.

This is also why fixing categorization drift isn't a content-volume problem. Publishing more pages about your own category doesn't outweigh an entrenched third-party source if that source is what the AI's retrieval layer keeps pulling from. The fix has to start with identifying which specific source is doing the defining, then either correcting what that source says or building a citation profile strong enough to compete with it. A 4 P's score, run again next quarter, will just report the same Portrayal number with no indication of which lever to pull.

Categorization drift explains why the words are wrong. The next question is why paying to be seen doesn't fix that.

Why doesn't showing up in an AI answer (Presence) guarantee an accurate Portrayal?

In one analysis of Google AI Mode results, a text ad appeared on 29% of commercial keywords tested, but the advertiser's own domain showed up among the cited sources only 11% of the time, and the exact advertiser URL just 1.95% of the time (Search Engine Journal, 2026). Paying for Presence didn't buy Portrayal.

That gap between paying to appear and controlling what appears is the sharpest illustration of why the 4 P's need to be read together, not banked as one composite score. A brand could register strong Presence, the ad shows, and weak everything else, the AI's actual answer text draws almost entirely from other sources describing the category. The buyer sees the brand's name. The buyer does not see the brand's own explanation of what makes it different.

A composite score would average these into something that looks fine on a dashboard and explains nothing to the person who has to fix it. That distinction matters more than it sounds, because it's exactly the gap that shows up when buyers go back and fact-check what the AI told them: they're checking the AI's characterization, not the ad. If the characterization came from a competitor's content, the ad spend bought exposure to the wrong pitch.

None of this shows up as a line item on a 4 P's dashboard. It shows up as a shortlist quietly forming without you.

What should you audit instead of just re-running the score next quarter?

About 90% of B2B buyers purchase from the shortlist they'd already formed before formal evaluation began (Bain, 2026). If an AI engine is quietly building that shortlist with a competitor's framing of your category, a quarterly score confirms the erosion without explaining it. An audit reads the actual sentence the AI returns and traces where it came from.

Running the score again next quarter won't surface any of that. It'll tell you the number moved, not why. What actually closes the gap is the same audit work we run underneath any AI-visibility engagement at Moving Parade: pull the exact sentence an AI engine returns for a direct question about the company, trace which third-party sources fed that sentence, and diff it against the company's own category language line by line. That work identifies which competitor's definition is winning, which source is amplifying it, and which specific page needs to change to close the gap.

It's slower than reading a dashboard. It's also the only version of this that tells a CMO what to do Monday morning instead of what happened last quarter. If your shortlist is already forming inside an AI conversation before your sales team ever hears about it, that's the diagnostic that matters, not the score covering it up.

Score vs. audit, side by side

IAB 4 P's dimension

What the score tells you

What only an audit tells you

Presence (Mention Rate)

Whether you're named in the AI's answer at all

Which third-party source got cited instead of your own site

Prominence (Citation Rate, Share of Voice)

How often you appear, relative to named competitors

Where the AI's phrasing diverges from your own category language

Portrayal (Sentiment)

Whether the tone of the mention is positive, neutral, or negative

Which competitor's definition of the category the AI is anchoring to

Persuasion (Recommendation Strength)

Whether the mention is associated with downstream action

Whether the click that follows a mention ever reaches your actual page

Frequently asked questions

What are the 4 P's in IAB's AI visibility framework?

Presence, Prominence, Portrayal, and Persuasion. Presence measures whether a brand is mentioned in an AI answer at all. Prominence measures how often and how high. Portrayal measures the sentiment and accuracy of what's said. Persuasion measures whether the mention correlates with downstream action, like a site visit or purchase consideration.

Does a high AI-visibility score mean an AI engine is describing your company accurately?

No. A high score on Presence or Prominence only confirms you're being mentioned often, relative to competitors. It says nothing about whether the description matches your own positioning. A brand can score well on visibility while an AI engine repeats an outdated or third-party-defined version of what the company actually does.

What is categorization drift in AI search visibility?

Categorization drift is when the definition of your category that an AI engine repeats to buyers no longer matches how your company describes itself, because the citation sources feeding that answer, mostly third-party sites, are setting the framing instead of your own site. A score can't detect it; only a direct comparison can.

Why do third-party sites control more of an AI engine's answer about a company than the company's own site does?

Roughly 85% of AI citations for broad B2B category queries come from third-party sources, review platforms, analyst reports, and publications, rather than the vendor's own site (Rampiq, 2026). AI retrieval layers weight independent, frequently cited sources more heavily than self-published claims, which is a credibility signal, not a bug.

How is a visibility score different from a visibility audit?

A score tells you where you rank on defined metrics, Presence, Prominence, Portrayal, Persuasion, against a defined set of competitors. An audit reads the actual sentence an AI engine returns, traces which sources produced it, and identifies exactly which competitor's framing or which outdated source is winning instead of yours.

One move: Ask an AI engine directly, "What does [your company] do," then diff its answer, phrase by phrase, against the category-defining sentence on your own homepage. Every mismatch is categorization drift, and it's the fastest single diagnostic a composite score will never produce.

Ready to build pipeline?

Tell us where you are.
We'll tell you what we can do.

Ready to build pipeline?

Tell us where you are.
We'll tell you what we can do.

Ready to build pipeline?

Tell us where you are.
We'll tell you what we can do.

Ready to build pipeline?

Tell us where you are.
We'll tell you what we can do.