Right now: we’ll set up your AI employee for you — free of charge (worth €249)Find out how →

← Blog AI visibility Technical SEO GEO Data SMEs

The search layer behind the AI answer: what decides whether your page gets retrieved

2026-08-25 · Martin Nymann · 9 min reading

It is rarely the language model itself that finds your pages — there is a search layer in front of it. The providers of these layers (Tavily, Exa, Brave, Perplexity Search) have now themselves described what they rank by: authority and evidence quality, contradictions between sources, and recency. Here’s what each of these three signals means for your website — including the one that almost nobody takes into account: when your own pages contradict one another.

Key takeaways

  • An AI response is generated in two stages: a search layer retrieves and sorts candidates, and only then does the model analyse the selection. The model cannot cite a page that the search layer did not return.
  • Search engine providers form a market of their own — Tavily, Exa, Brave Search API, Perplexity Search, You.com, Parallel — and they compete on public benchmarks such as SealQA and SimpleQA Verified.
  • Three criteria recur: authority and quality of evidence (does the page actually answer the question?), handling of conflicting excerpts, and up-to-date information on time-sensitive topics.
  • Contradiction is what goes unnoticed: if two of your pages say different things about the same fact, hedge their answer or downplay your position — and each page seems correct when read on its own.
  • Our own measurements highlighted this as a caveat: out of 576 AI responses, only Perplexity used its own engine; the other three retrieved sources via the OpenRouter plugin (Exa).
  • The pages that were ultimately used as sources had higher technical scores (77.9 versus 72.7) and more frequently featured structured data (68 per cent versus 57 per cent), p=0.000. Of the 4,034 sources to which the responses referred, 61 per cent were the companies’ own websites.

When ChatGPT or Perplexity answers a question about your sector, it is rarely the model itself that has found your pages. It has a search layer — a separate search API that retrieves candidates from the web, sorts them and passes a small selection on to the model. That layer is the real bottleneck for your visibility, and the providers behind it have now themselves set out what they sort by: the authority of the source, how well the page actually answers the question, whether your information contradicts itself, and how recent the page is. Here’s what this means for your website — including the signal that almost no one is working on.

What is the search layer — and why haven’t you heard of it?

An AI answer with sources is generated in two steps, not one. First, a search layer retrieves a set of web pages that might be relevant. Then, the language model reads only that selection and formulates the answer. The model cannot quote a page that the search layer didn’t provide — no matter how good the page is.

The layer has its own providers, and they constitute a market in their own right: Tavily, Exa, Brave Search API, Perplexity Search, You.com and Parallel all sell search services built for AI agents rather than for humans. They compete on benchmarks such as SealQA and SimpleQA Verified – and in August 2026, Tavily publicly described the changes they’d made to move into first place. This is the sort of insight you don’t usually get: a provider explaining its own ranking.

The reason you haven’t heard of the layer is simple: it’s invisible from the user’s perspective. You see an answer with source references and assume the model has carried out the search. In practice, an intermediary has made the choice for it.

What does the search layer prioritise when selecting your pages?

Tavily describes three areas they have been working on. These are not specific to this particular provider — they are the three fundamental problems that any AI search layer must solve, and each one points to something specific on your website.

What the search engine doesWhat this means for your pageWhat you can do about it
Ranks by authority and quality of evidence — not just whether the words match, but whether the page actually contains the answerA page that answers the question directly and early on outperforms a page that beats about the bushAnswer first: the conclusion in the first paragraph, the question as the headline
Removes duplicates and resolves conflicts between the extracts foundIf your pages say different things about the same fact, you become the source that the answer takes with a grain of saltOne figure, one price, one telephone number — the same information everywhere
Prioritises recency for topics that changeAn undated page looks old to a search engine, regardless of when you last updated itVisible dates and genuine updates on pages about prices, rules and markets

Note what isn’t on the list: how many words the page contains, how many keywords it repeats, or how well-designed it is. This is consistent with existing research in the field. The Princeton and Georgia Tech study (KDD ’24) measured a 30–40 per cent increase in visibility in generative search engines by adding cited sources, statistics and direct quotes — in other words, by making the content more evidence-based, not longer.

Why does contradiction penalise you more severely than you think?

The middle point is the one almost nobody works on, and it’s worth sticking with. A search query typically retrieves a handful of snippets from several pages. If two of them say different things about the same fact — your price is listed as 995 DKK on the homepage and 1,195 DKK on the pricing page, your opening hours are one thing on the contact page and something else in your Googleprofile, your staff is listed as “12 specialists” in a case study and “over 20” in a blog post — then the search layer has a problem it needs to resolve.

It won’t be resolved in your favour. The model will either present conflicting evidence and hedge its bets (“some sources state …”), or it will downplay the source that doesn’t match the rest. Both outcomes cost you exactly what you were after: a clear, named reference.

The troublesome thing about this error is that it doesn’t look like an error. Every single page looks correct when you look at it on its own. The contradiction only becomes apparent when someone reads two of your pages at the same time — and that is exactly what a machine does. The classic examples are the same time and time again: old promotional prices still listed on a landing page; a telephone number that was changed in one place; year references in text that say ‘this year’ about something that happened last year; and figures about your own business that grow slightly every time they’re written.

How important is freshness?

Freshness isn’t a general requirement to write every week. It’s a signal of relevance that carries significant weight for time-sensitive issues and almost none for timeless ones. A question about what an accountant typically costs in 2026 is time-sensitive. A question about how double-entry bookkeeping works is not.

The practical approach, therefore, is not more content, but datable content: a visible ‘last updated’ date on pages containing figures and regulations, a dateModified in your structured data that is actually kept up to date, and a fixed routine for reviewing the five to ten pages containing prices and legislation. A page that has been updated without stating so receives no credit for it.

What did our own measurements show?

We ourselves came across the ‘search layer’ as a caveat long before it became a topic of discussion. In our analysis of 576 AI responses with web search enabled — four search engines, twelve Danish cities, four unbranded customer queries, measured against 465 real companies across three sectors (our own measurement, July 2026) — only Perplexity searched using its own engine. The other three ran the web search through OpenRouter’s plugin, which uses Exa and provided five sources per response. In other words: for three out of four search engines, we effectively measured what the search layer chose to deliver.

This makes the results more relevant here, not less so. The pages that ended up as sources in the results had a significantly higher technical score than the rest (77.9 versus 72.7). They also more frequently contained structured data (68 per cent versus 57 per cent). Both differences are highly significant, p=0.000. It is technical readability and structure that determine whether you even make it through the eye of the needle. Conversely, the same indicators predicted nothing regarding whether the company was recommended by name (74.9 versus 74.9, p=0.988) — the two outcomes are different, and it is worth keeping them separate, as we have written about in ‘cited or recommended’.

A third figure from the same dataset is also relevant. The responses referred to a total of 4,034 sources, and 61 per cent of these were the companies’ own websites. The role of a source is not reserved for industry portals and directories. Your own website is the largest single platform — provided the search engine can read it.

How do you put this into practice?

Four things, in the order of their effectiveness:

  • Make sure you can be crawled at all. If your robots.txt AI crawlers, or if your content requires JavaScript to be displayed, the rest is irrelevant. This is less of a problem than feared: in a scan of 1,252 Nordic SME websites in July 2026, only 3.6 per cent (±1.0pp) were actually unreadable — but 88.2% (±1.8pp) were missing an llms.txt file.
  • Sort out the inconsistencies. Make a list of the facts about your business that appear in multiple places — prices, numbers, years, contact details, services — and correct them so that they say the same thing everywhere, including in your structured data and your Google profile.
  • Write the answer first. Use the question as the heading, the answer in the first paragraph, and figures with sources. This is the format recommended by both research and the suppliers’ own descriptions.
  • Date anything that may become out of date. A visible date, correctly dateModified, and a routine for pages containing figures.

If you want to see how your own site measures up in terms of its technical foundation, the free test takes less than a minute and looks at exactly the signals that search engines use to rank sites.

Frequently asked questions

What is the search layer behind an AI response?

A separate retrieval and sorting stage ahead of the language model. It searches the web, ranks the candidates and passes a small selection on to the model, which formulates the answer based on that specific selection. The layer has its own providers — including Tavily, Exa, Brave Search API and Perplexity Search — which sell search solutions built for AI agents rather than for humans.

Does ChatGPT use Tavily for searching?

We don’t know, and we won’t speculate. Providers rarely disclose who their customers are, and AI providers rarely reveal which search layer they use. The point is the mechanism, not the name: there’s a retrieval stage before the model, and it sorts according to criteria you can work with. In our own measurements, it was OpenRouter’s plugin (Exa) that retrieved the sources for three out of four search engines.

What does the search layer prioritise most when selecting pages?

Three factors recur in the providers’ own descriptions: whether the source is authoritative and actually contains the answer, not just the right words; whether the extracts found contradict one another; and how up-to-date the page is on topics that change. The Princeton and Georgia Tech study (KDD ’24) points in the same direction: sources, statistics and citations boosted visibility in generative search engines by up to 30–40 per cent.

Why is it a problem if my pages contradict each other?

Because the search engine retrieves several snippets at a time. If your price differs in two places, or your number of employees varies from page to page, the model receives conflicting information — and the result is either a tentative response or a downgrading of you as a source. The error is hard to spot because each individual page looks correct on its own.

Is classic technical SEO still relevant?

Yes — much of what the search layer weighs is technical SEO under a different name: readable HTML without JavaScript requirements, a clear structure, structured data and correct dates. What’s new is that inconsistencies between your own pages are taken into account, and this has rarely been a factor in a classic SEO review.

Do I need to update all my pages regularly to keep them looking fresh?

No. Freshness carries significant weight for time-sensitive issues and almost none for timeless ones. Identify the pages that include prices, rules, years or market conditions, and keep them up to date with a visible date and a correct ‘dateModified’ value. The rest can stay as they are — and an update where only the date changes is neither honest nor effective.

Get your free GEO Score — 60 seconds, no credit card required.

Get started for free

Related articles

Stay up to date with more GEO insights for smaller businesses.

View all articles