← Blog ChatGPT AI visibility Measurement GEO SMEs
What ChatGPT can't tell you about your AI visibility
2026-08-04 · Martin Nymann · 7 min reading
ChatGPT is an excellent adviser on AI visibility — by all means use it. But three things it cannot do: measure whether you are mentioned over time, know how you are developing compared to your competitors, or verify whether a fix worked. ChatGPT gives you advice. Geoa measures whether it works — plus the honest answer to when do-it-yourself is actually enough.
Key takeaways
- ChatGPT is a good adviser on AI visibility — the advice is often sensible, and the five free guides in this series show that do-it-yourself is a real path. By all means use it.
- The three gaps: one answer is a sample (tomorrow it may be different), it doesn't know your position over time — nor your competitors' — and it cannot verify whether a fix worked.
- All three gaps are about the same thing: repetition under constant conditions. It is infrastructure ChatGPT lacks, not intelligence — a conversation cannot stand still for six weeks and measure itself.
- Honestly: a small local business in one town with the discipline for the weekly routine can cover its needs manually. Automation only starts to make sense with greater scale, slipping discipline or more expensive decisions.
- The division of labour is simple: ChatGPT gives you advice, Geoa measures whether it works — a free account monitors your first 5 questions, paid plans from 299 kr/month.
By all means ask ChatGPT for advice about your AI visibility — it is good at it, and the advice is usually sensible. But there are three things it cannot do: it cannot measure whether you are mentioned over time, it doesn't know how you are developing compared to your competitors, and it cannot verify whether a fix worked. In short: ChatGPT gives you advice. Geoa measures whether it works. Here is exactly where the line runs — and when do-it-yourself is actually enough.
Why is ChatGPT actually a good adviser here?
Let's start by granting the premise in full: ChatGPT is an excellent sparring partner on AI visibility. It can explain what an llms.txt is, suggest structured data, rewrite your about page so it is quotable, and go through your copy with fresh eyes. The advice is free, patient and often better than what many would get at an hourly rate. We mean that seriously enough to have written a whole series of free guides teaching you to do the work yourself: the test of whether you are mentioned, llms.txt in 30 minutes, the bank of customer questions, the difference between cited and recommended and the weekly manual routine.
And the stakes are real enough: Ahrefs' analysis of 300,000 keywords measured a 34.5 % lower click-through rate for the top result in Google when an AI answer sits above it — and 58 % in the December 2025 re-run. Customer research is moving into the answers, so it is wise to work on your visibility there. The question is not whether you should use ChatGPT in that work. You should. The question is what it cannot tell you afterwards.
Because advising and measuring are two different disciplines. The Princeton/Georgia Tech research (KDD '24) measured up to a 40 % visibility lift from targeted optimisation — but “up to” hides the fact that some moves work, others don't, and that it depends on your starting point. Which moves worked for you is something only a measurement can settle. And that is exactly where the chain breaks.
What are the three gaps in ChatGPT as an adviser?
None of the three gaps is a flaw in ChatGPT. They follow from what a conversation is: a snapshot with no memory of your history and no way to repeat itself systematically. Here are the three, and why they matter for your decisions:
| Gap | What happens in practice? | What does it take instead? |
|---|---|---|
| Sample ≠ measurement | You ask whether your business is visible and get one answer. Tomorrow the answer may be different — same phrasing, new draw. | The same fixed phrasings, asked systematically over weeks, so patterns can be told apart from noise. |
| No position over time | ChatGPT doesn't know whether you were mentioned more often in May than in August — or whether a competitor has overtaken you. It has no time series for you. | A log of who is mentioned on your questions, week by week — competitors included. |
| Cannot verify a fix | You follow a piece of advice — new llms.txt, rewritten page — and ask: “did it work?” It cannot know; it can only give you yet another snapshot. | A before/after in the same measuring setup: same phrasings, same method, measured across the fix. |
Notice that all three gaps are about the same thing: repetition under constant conditions. It is not intelligence ChatGPT lacks — it is infrastructure. A conversation cannot stand still for six weeks and measure itself.
Why does ChatGPT answer differently from one time to the next?
Because the answers are not deterministic: the same phrasing can produce different answers from one time to the next, the model is updated continuously, and memory or earlier conversations can colour the answer in your particular account. That doesn't make the answers worthless — it makes them samples. One sample is a fine place to start; the fifteen-minute test is built on exactly that. But decisions — should I spend money on content or reviews? did last month's work pay off? — require you to tell a fluctuation from a trend.
That is also why you cannot “ask your way” to a competitor analysis: today's answer to “who are the best [industry] in [town]?” doesn't say whether the picture was the same last month, or whether it is slipping. The pattern over time is the whole information — and it doesn't exist in any single conversation.
When is do-it-yourself actually enough?
Honest answer: more often than you'd think, and we are happy to say so openly. If you are a small local business in one town, with a handful of fixed customer questions and the discipline to keep the weekly routine running, you can cover your needs manually: five fixed phrasings, a log sheet, a monthly follow-up. It costs about 45 minutes a week and no money — and it is real measurement, not a discount version.
Automation starts to make sense when one of three points is reached:
- The scale grows: you want to measure more than five phrasings, more towns, or more AI engines than the one your fifteen minutes can cover.
- Discipline slips: gaps in the log sheet ruin the time series — and then the manual measurement has effectively started over.
- The decisions get more expensive: if you are about to spend serious money on content or review work, you want to know whether it moves anything — not guess.
If none of the three is true for you yet, stick with the manual approach with a clear conscience. The guides are written so that you can.
What can the measurement do that the conversation cannot?
Systematics — the boring part that settles the matter. Geoa asks your fixed customer questions again and again, week after week, across AI engines, logs who is mentioned and cited, and notifies you when something moves. On top of the measurement sits auto-fix in all paid plans: concrete fixes generated from what the measurement finds. It is not smarter advice than ChatGPT's — it is the same kind of advice, tied to a before/after, so you can see what worked.
So the division of labour is simple, and you can start it for free: use ChatGPT as your adviser, and put your five most important customer questions under automatic monitoring — a free account measures the first 5, paid plans from 299 kr/month cover 25, 100 or 300 with a 14-day free trial, no card required. ChatGPT gives you advice. Geoa measures whether it works. Use both.
Frequently asked questions
Can't I just ask ChatGPT whether my business is visible?
Yes — and it is a perfectly good first step. But the answer is one sample: tomorrow it may be different, and you cannot see whether the picture is changing. The difference is measurement over time: the same phrasings, asked systematically, so you can tell a random fluctuation from a real trend.
Why does ChatGPT answer differently from one time to the next?
Because the answers are not deterministic: the same phrasing can produce different answers from one time to the next, the model is updated continuously, and the conversation's context and memory can colour the answer in your account. That is why a single answer should be read as a sample — and why repetition under constant conditions is the entire difference between an impression and a measurement.
What can Geoa do that ChatGPT cannot?
The systematics: the same customer questions, measured week by week across AI engines, with a log of who is mentioned and cited, a notification when something moves — and auto-fix, which turns the findings into concrete fixes. ChatGPT is the adviser; the measurement shows whether the advice worked for you specifically.
What does it cost to get started?
Nothing: a free account automatically monitors your first 5 customer questions. Paid plans start at 299 kr/month and cover 25, 100 or 300 monitored questions depending on tier — all with a 14-day free trial, no card required. And the manual guides in this series stay free no matter what.
Get your free GEO Score — 60 seconds, no credit card required.
Get started for free