Measuring Brand Visibility in AI Answers: A 10-Level GEO Ladder Test

Forest Liu · Data Marketing Lead for Multiple Companies#geo#ai-search#brand-visibility

Summary

A 10-level difficulty-ladder test across four AI models — about 120 calls per round at roughly $0.23 — shows AI visibility is layered: the example brand kept exposure at moderate list lengths but hit 0% across all models once prompts demanded 30+ brands.

Quick answer

Build a 10-level ladder of category questions — from “recommend a product” to “list 30 brands” — and run each level repeatedly against ChatGPT, Claude, Gemini and Perplexity, recording mentions and positions. One full round of about 120 calls costs about $0.23, so it can run as a standing monitor. In our test the example brand kept exposure at shorter list lengths and hit 0% at 30+ brands — visibility in long AI lists has to be deliberately built.

There is a question we have always wanted a clean answer to: when a user asks ChatGPT, Claude, Gemini or Perplexity something about our product category, how likely is our brand to be mentioned — and where does it show up? Traditional brand monitoring never answers that; it counts mentions on the web, not citations inside a model's answer. So we built a test of our own, and it now runs as a standing monitor.

Designing a 10-level difficulty ladder

We sliced the problem into 10 difficulty levels. Level 1 is the friendliest — “recommend a product in this category”. The deepest level is deliberately brutal — “list 30 brands in this category”. The logic behind a ladder is that AI visibility is not a single number: the same brand can be cited in easy answers and invisible in hard ones, and those two situations demand completely different work. We run one fixed prompt template per level, repeat each prompt multiple times on every model, and record whether our brand is mentioned and where — first answer, front of the list, or absent. Mentions and position are logged separately, because being named and being listed first are very different outcomes.

What a round of testing costs — and finds

One full round is roughly 120 calls and costs about $0.23. At that price, AI-answer visibility is something you run as a standing monitor, not a one-off research project — and repeating the same prompts matters, because model answers carry variance, and a single absence proves nothing. The first run was already informative: the example brand kept some exposure while prompts only demanded shorter lists (up to around 20 names), then dropped to 0% exposure across every model once the ladder asked for 30 brands or more. Visibility in long AI lists does not happen by itself.

Layered visibility, layered work

AI visibility is layered: the easier levels are cheap to enter, the deeper ones require deliberate cultivation. What actually helped us was turning the test into instruments — a brand × model exposure matrix that shows, at a glance, which model under-cites you; competitor tracking on the same queries so you know whose brand is cited instead of yours; and trend alerts that fire when your share of AI answers rises or collapses. The numbers in the matrix are only the start — the value is in the actions they point to, like adding structured data, rewriting a citable summary block, or anchoring the page models actually quote.

The 5-step GEO ladder checklist

  • Define the category question levels — 10 levels from easy to brutal; the deepest one should hurt.
  • Freeze the prompt template — only identical prompts produce comparable results.
  • Repeat across models — every level, every model, multiple runs; AI answers have variance.
  • Record exposure, position and competitors — who is cited, where, and who is cited instead of you.
  • Rerun on a schedule — AI visibility shifts, so the monitor has to be continuous.

This ladder, the exposure matrix and the trend alerts are exactly what our GEO competitive monitoring module runs for us every week.

Note on system details: The iport platform is under active development. Any product features, interfaces, or workflows described in this article reflect the version in use at the time of writing and may differ from the latest release. For the most current capabilities, refer to the official platform documentation.

Frequently asked questions

How do you measure whether a brand appears in AI answers?

We run a 10-level difficulty ladder of category questions against ChatGPT, Claude, Gemini and Perplexity — from “recommend a product” to “list 30 brands” — repeating each prompt several times and recording whether the brand is mentioned and where: first, front of list, or absent. A full round of about 120 calls costs about $0.23.

What does 0% visibility in AI answers mean?

It means your brand is not in the set of names models reliably produce for your category. In our test the example brand kept some exposure while prompts asked for up to ~20 names and dropped to 0% across all models when the demand reached 30+ names. Easy levels are cheap to enter; deep-list visibility has to be built deliberately.

How much does GEO monitoring cost to run?

Measurement is nearly free — a round of about 120 model calls costs about $0.23, so it can run as a standing monitor. The real investment is the cultivation work the results point to, not the measuring itself.

Keep reading