
You checked ChatGPT three times last week to see whether your brand came up in your category. First run, you were there. Second run, gone. Third run, you were back but listed fourth instead of second.
So which number goes in the monthly report? That question is the entire problem with treating AI search like a ranking board, and it’s the reason a GEO rank tracker has to work differently from anything in your SEO stack.
The Setup: One Prompt, 500 Runs, Five Engines
We took a single high-intent commercial prompt, the kind a real buyer types when they’re two weeks from a purchase decision, and ran it 100 times each across ChatGPT, Gemini, Perplexity, Claude, and Google AI Overviews. Same wording. Same day. No personalization, no session history.
For every response we logged three things: whether the target brand appeared at all, where it sat in the ordering when it did appear, and which domains got cited underneath.
The point wasn’t to measure one brand. It was to answer a more basic question: does “rank” survive contact with a system that generates a fresh answer every time?
The short version is that it survives, but not in the shape most teams assume.
A Single Check Isn’t a Rank. It’s an Anecdote
Here’s the thing about probabilistic output. If your brand shows up in 4 out of 10 runs, your actual mention rate is 40%. A one-shot manual check reports either 0% or 100% depending on which run you happened to catch. Both readings are wrong, and neither comes with a warning label.

This isn’t a quirk you can configure away. Even at temperature zero, the same prompt can produce different outputs across runs because of floating-point non-associativity combined with dynamic batching on the inference side. Variance is baked into the infrastructure.
The measurement layer inherits that. As iPullRank puts it in their AI search manual, share of voice in generative search is a statistical distribution of presence over many trials, not a static percentage of positions held.
One check is not a data point. It’s a coin flip you wrote down.
The instability compounds over time, too. Independent analyses suggest 40 to 60% of AI citations rotate every month for mid-sized B2B brands, and that 73.4% of specific URLs get cited exactly once before vanishing from AI answers entirely.
Mention Rate Moved a Lot. Position Barely Did.
This was the finding that changed how we read the data.
Across the 500 runs, whether the brand appeared swung far more than where it appeared. On three of the five engines, mention rate moved by double digits between the first 50 runs and the second 50. But in the runs where the brand did appear, its ordinal position clustered tightly, usually within a single slot of its median.
Two different signals. Two different failure modes. Most dashboards collapse them into one number called “rank” and lose both.
There’s academic backing for the split. Research on the structural gap between search engine and generative AI brand visibility found that traditional SEO strength predicts a brand’s ranking position inside an AI answer reasonably well, but predicts its mention frequency poorly. Being strong enough to get listed and being retrieved often enough to get listed are governed by different mechanics.
Semrush’s study of 1,094 subject areas in ChatGPT points the same direction from another angle. Only 21% of the most-cited domains in a category were also the most-mentioned brand, and the two signals correlate slightly negatively at -0.229.
That matters operationally. If your mention rate is falling but your position holds, you have a retrieval problem and you need more citable surface area. If your mention rate is stable but your position slides, you have a framing problem and competitors are being described as the better fit.
Same “rank drop.” Opposite fixes.
Five Engines, Five Different Answers to the Same Question
Run-to-run variance was real. Cross-engine variance was bigger.
The gap between what ChatGPT said and what Perplexity said, given identical input, exceeded the gap between any single engine’s best and worst run. That tracks with the published research. One analysis of 50 buyer-intent prompts found that ChatGPT, Perplexity, and Gemini named the same brand only 21% of the time, with over half of all brand mentions coming from just one engine.
The citation layer diverges even harder. Across 680 million AI citations analyzed in early 2026, only 11% of domains were cited by both ChatGPT and Perplexity. Yext’s look at 6.8 million citations found very little overlap in what each model cites, with each engine weighting source types on its own logic.
So tracking one engine isn’t partial coverage. It’s a systematic bias, and it points in a direction you can’t predict from the engine you did measure.
The corollary is worse for reporting: a blended cross-engine average hides exactly the thing you’d act on. A brand at 60% on Gemini and 5% on ChatGPT averages to a perfectly unremarkable 32%.
What This Means for Your GEO Rank Tracker
Work backward from the variance and the tool requirements write themselves.
| Requirement | Single-check approach | Sampling-based GEO rank tracker |
|---|---|---|
| Sample size | 1 run per prompt | Dozens of runs per prompt, reported as a rate |
| Metric structure | One blended “rank” score | Mention rate and position tracked separately |
| Engine coverage | One engine, extrapolated | Each engine reported independently |
| Cadence | Monthly or ad hoc | Weekly minimum, daily for volatile categories |
| Competitor context | Absent | Same prompt set, same sampling, side by side |
Search Engine Land’s overview of the category makes the same point about methodology: variable outputs mean tracking requires consistent monitoring and statistical sampling rather than spot checks.
Cadence deserves its own note. Monthly monitoring is effectively no monitoring when citation sets turn over at 40 to 60% in that same window. By the time you see the change, you can’t attribute it to anything.
And frequency without competitor context still leaves you blind. If your citation rate holds flat at 15% while a rival climbs from 10% to 40%, your number didn’t move but your share collapsed.
Where Topify Fits
The reason we ran this test at all is that it maps directly onto how measurement should be built.
Topify reports seven metrics separately rather than folding them into a single score: visibility, sentiment, position, volume, mentions, intent, and CVR. Visibility answers how often you appear across a defined prompt set. Position answers where you land when you do. Keeping them apart is what makes the mention-versus-position diagnosis possible in the first place, and it’s the difference between knowing your number dropped and knowing why.
Each prompt runs repeatedly across ChatGPT, Gemini, Perplexity, DeepSeek, Doubao, Qwen, and others, with results reported per engine instead of averaged into a single figure. Competitor benchmarking runs on the same prompt set and the same sampling, so relative share is visible even when your absolute number sits still.

The citation reverse-engineering layer closes the loop. Seeing which exact domains and URLs each engine pulls from tells you where the retrieval gap lives, which is the actionable half of a falling mention rate. Our earlier breakdown of how a GEO rank tracker measures AI search position covers the metric definitions in more depth.
Plans start at $99/month with 100 tracked prompts and 9,000 AI answer analyses, which is roughly the sampling volume this kind of question requires. You can get started with Topify on a 30-day trial.
How to Read Your Own Rank Data Without Fooling Yourself
Four habits separate teams who act on AI search data from teams who argue about it.
Track prompt sets, not keywords. Buyers ask several related questions on the way to a decision, and winning one of them isn’t the same as owning the topic. Visibility that looks strong on a single prompt often thins out across the cluster.
Read trend bands, not points. A 5-point week-over-week move on a sampled rate is usually noise. A 5-point move sustained across four weeks is a trend. Set that threshold before you look at the data, not after.
Plot competitors on the same axis. Absolute visibility without relative share tells you almost nothing about whether you’re winning.
Separate “absent” from “present but ranked low.” They look identical on a summary dashboard and they need completely different responses.
Sample enough. Split the metrics. Check every engine.
Conclusion
Rank didn’t disappear when search became generative. It changed units. It stopped being a position you hold and became a probability you occupy, which means the number is only meaningful attached to a sample size.
The three checks you ran in ChatGPT last week weren’t wrong. They were just three draws from a distribution you hadn’t measured yet. A GEO rank tracker built on repeated sampling, separated metrics, and per-engine reporting turns those draws into something you can put in a report and defend.
Start by defining the ten prompts your buyers actually ask. Everything else follows from having a stable set to sample against.
FAQ
Q: How many times does a prompt need to run before the result is trustworthy?
A: Dozens, not a handful. The practical floor is enough runs that a single outlier can’t move the rate by more than a point or two. Most sampling-based platforms run each prompt many times per cycle for exactly this reason, and any tool reporting a rank off one query is reporting an anecdote.
Q: Which matters more, mention rate or position?
A: Mention rate, in most cases. A brand that never appears can’t benefit from good placement. Once you’re appearing consistently, position becomes the lever that affects which option the buyer actually picks.
Q: Can I just track ChatGPT and assume the rest follow?
A: No. Cross-engine agreement on brand recommendations runs around 21%, and domain-level citation overlap between major engines sits near 11%. Single-engine tracking produces a biased read in an unpredictable direction.
Q: How often do AI rankings actually change?
A: Faster than SEO rankings. With a large share of citations rotating monthly, weekly tracking is the minimum viable cadence, and competitive categories often warrant daily sampling.

