AI visibility metrics: diagnostic vs decision-grade
20+ vendors sell AI visibility scores and none agree. The IAB two-tier standard separates numbers fit for a forecast from ones that stay diagnostic.
Ask three AI visibility vendors to score the same brand and you get three numbers that disagree. Each comes from a proprietary framework, a different prompt set, and a different sampling cadence, and no external body audits any of them. More than 20 companies sell these tools. Your CFO will ask which number is real, and “the one from the vendor we bought” is not an answer that survives a finance review.
The IAB’s “Measuring Visibility in the AI Era” guidelines, published August 3, 2026, give you the vocabulary to sort this out. Two pieces matter in practice.
The four Ps: what a visibility metric can measure. Presence asks whether your brand appears in an answer at all. Prominence asks where and how large. Portrayal asks what the model says about you, accurate or not, positive or not. Persuasion asks whether the mention moves a buyer. Most tools on the market measure presence and call it visibility. Portrayal is where B2B deals are won or lost, because a model that names you with the wrong positioning does more damage than one that skips you. Ask any vendor which of the four their score covers. If the answer is one, you now know what the other three cost extra or don’t exist.
Two tiers: what a metric is allowed to touch. The IAB splits measurement into directional and decision-grade. Directional metrics spot trends and flag competitive movement. Decision-grade metrics meet defined standards on query volume, sample size, prompt-type coverage, testing cadence, reproducibility, and platform coverage, and only decision-grade numbers belong in budget or strategy decisions. Almost everything sold today is directional. That includes composite visibility scores, share-of-AI-voice percentages, and crawler traffic counts.
Crawler counts deserve their own warning. Server-log tools now report which AI agents fetched which pages, and the growth numbers look spectacular, with vendors citing measured cases in the hundreds of percent. A crawler requests pages whether or not a single buyer asked about your category. Rising crawler traffic means your site became legible to machines. It says nothing about demand, and user agent strings are self-declared, so the counts are spoofable on top of being ambiguous. Coverage diagnostic: yes. Demand signal: no.
The practical build for a GTM team:
Sort your current metrics into two lists. Anything from a proprietary unaudited framework goes in the diagnostic list, and diagnostic metrics never appear in a forecast, a board deck, or a budget justification. They steer experiments.
Run your own presence baseline before buying anything. Ten to fifteen buyer-shaped prompts across ChatGPT, Perplexity, Gemini and Claude, re-run monthly, logged in a sheet. This is directional by the IAB’s definition too, but it is your prompt set against your competitors, it costs an afternoon, and it tells you whether a vendor’s score even correlates with what buyers see.
Make portrayal the metric you act on. When a model describes you wrongly, trace which sources it cites for the claim and fix those sources. That work routes into the same system as citation building; the mechanics are in citation infrastructure.
Hold budget decisions to the decision-grade bar. If a vendor wants their score in your planning cycle, send them the IAB’s criteria and ask which they meet, in writing. The ones with real methodology will answer. The ones selling a black-box composite will send a deck.
The pattern behind all of this is older than AI. A new channel produces metrics before it produces standards, vendors price the metrics, and teams that treat unaudited numbers as reported numbers fund the wrong things for two years until the standard catches up. Display went through it, social went through it. The IAB tiers exist so you can skip that phase: measure now, spend on what’s decision-grade, and keep everything else in the lab.