Ask an AI engine a buying question like "best AI SEO service" or "affordable monthly SEO" and it answers in a paragraph, then names a few sources underneath. Those named sources are the new front page. We wanted to know, from real data rather than opinion, which pages actually earn that spot for SEO and GEO queries in 2026. So on July 22, 2026, we logged which sources three engines named across six buyer and definitional queries on the sites we run. This is what that snapshot showed.

One thing up front, because it decides how much weight to give any of this: this is a small first-party sample from a single collection run, gathered by our own automated tracking, and it is not running on a repeating schedule right now. We say more about that in the method section below and again in the limitations note at the end. Read the magnitudes here as directional, not as a poll of the whole web.

Method: what we measured and how

ConceptSEO ran this citation tracking across the domains it manages (hyperliquidguide.com, carsnipe.com, drcritter.com, immlawcenter.com) and its own site. For this benchmark we pulled the SEO and GEO slice of that data: four commercial buyer queries ("done-for-you SEO", "affordable monthly SEO", "best AI SEO service", "best GEO agency") and two definitional queries ("what is GEO", "what is generative engine optimization").

For each query we recorded which sources three engines named in their answers: ChatGPT search, Perplexity, and Google AI Mode (the same surface that feeds AI Overviews). The tracking was automated and AI-assisted, which is how we covered the full query set in one pass instead of by hand, and it is the same approach described in our write-up on GEO citation tracking. What follows is counts of cited-source types, not rankings and not traffic.

Which sources win the commercial queries?

For the buyer queries, the answers rarely cited the vendors themselves. The sources the engines named were overwhelmingly third-party aggregators: review directories that carry a dedicated GEO category, and "best tools" or "best agencies 2026" roundup listicles. In the results we logged, the two review directories that kept surfacing were G2's Generative Engine Optimization category and Clutch's GEO agencies directory. The roundups came from a rotating cast of publishers, among them theStacc, Nightwatch, Thrive, SE Ranking, Yotpo, Semrush, and OGTool.

Put plainly: across the domains and queries in this run, directories and roundup listicles were the dominant cited source type for commercial SEO queries, and a vendor's own site was the exception rather than the rule. That is the direction of a small single-run sample, not a measured rate. The source analyses behind this collection are no longer stored, so we describe the lean in words rather than quote a ratio we could not now recompute.

The definitional queries broke the other way. "What is GEO" and its variants pulled far more from high-authority editorial publications: Search Engine Land, Backlinko, and Inc. showed up where the vendor directories had dominated the buyer queries. So the source mix is not fixed. It tracks intent.

Was the cited source the ranking page, or a third party?

This is the part most site owners get wrong, so it is worth stating directly. For the commercial queries, the primary-site citation share was low. When an engine wanted to recommend a service, it mostly did not quote that service's own page. It quoted a directory or a roundup that listed the service. The vendor's homepage and pricing page were usually not the cited source, even when the vendor was arguably the best fit for the question.

We watched one case that makes the mechanism obvious. A vendor whose own product model matched the AI's framing of the ideal answer almost exactly was still left out of the answer. Not because it was irrelevant, but because there was no self-contained, retrievable page stating who the product is for, what it costs, and what you get. The engine had nothing clean to lift, so it reached for a source that did. Recommendation followed retrievable fit-facts, not relevance.

Informational queries were friendlier to primary sources. Original pages, especially ones publishing their own data, earned citations at a noticeably higher rate than they did on the commercial side. If you publish something first-hand and specific, an engine is more willing to name you as the origin of it.

Cited-source types by query intent

Here is the pattern in one table. The labels are directional (dominant, common, rare) and come from the source types we logged, not from invented percentages.

Cited source type Commercial buyer queries Definitional queries
Review directories (G2, Clutch GEO categories) Dominant Rare
"Best tools / agencies 2026" roundups Dominant Common
High-authority editorial publications Common Dominant
Vendor's own site (primary source) Rare Common
Original-data / first-party research pages Rare Common

How do ChatGPT, Perplexity, and Google AI Mode differ?

All three drew mainly on the aggregator layer, but they weighted it differently. Perplexity reached for roundup listicles most often and named them openly. ChatGPT search mixed directories and roundups inside a single answer, less predictably. Google AI Mode picked whichever passage answered the sub-question best, including passages on pages ranking below position one, and it was the only one of the three that regularly reached past the top-ranked result.

Perplexity leaned hardest on roundup listicles and, usefully, showed its sources openly. When a "best X 2026" article existed, Perplexity was the most likely of the three to pull from it and cite it by name. If you are trying to see the citation pattern for yourself, it is the easiest engine to read.

ChatGPT search blended the two aggregator types, mixing directories and roundups in a single answer. It was less predictable than Perplexity about which one it would favor for a given query, but the sources still came from the aggregator layer far more often than from vendor sites.

Google AI Mode leaned on the passage that best answered the sub-question, even when that passage lived on a page ranking below position one. This is the one to internalize: AI Mode was the most willing to reach past the top-ranked result for a cleaner answer chunk lower down. A page that is not winning the blue-link race can still win the citation if it states the answer better.

What did the cited pages have in common?

Four traits recurred across the cited sources whatever their type: self-contained passages that answer a question without the rest of the page, high information density, structured data tying the page to a real entity, and dates that agree with one another. The sources that kept getting named, whether directory, roundup, or editorial, shared a short list of properties, and they line up with what we already argue in our guide to earning citations in AI search, and with the dimension-level results in the GEO ranking ledger. One trait we did not test for is sentence-level word order, which Google Research has since put a number on, and which we counted across our own pages in subject-object entity order and what AI answers say.

  • Self-contained passages. A two to five sentence chunk that answers the sub-question on its own, without needing the rest of the page for context. This was the single most consistent trait.
  • High information density. Named entities and concrete numbers packed into each sentence, rather than hedged or general copy. Specific text gives an engine something to trust and lift.
  • Schema with a real graph. Organization or Article structured data with a populated sameAs that ties the page to a known entity, not an empty markup shell.
  • Freshness that agrees with itself. Recent dates where datePublished, the sitemap entry, and the visible date on the page all match. Contradictory or stale dates correlated with pages that got skipped.

None of these is exotic. The reason directories and roundups win is not that they are more authoritative in some abstract sense. It is that a "best 10 tools" listicle is built entirely out of self-contained, dense, dated passages. It is the ideal shape for retrieval, by accident of its format.

What does this mean for a small site trying to get cited?

Three things, none of which require being a big brand. Get listed accurately in the directories and roundups your category already has, because for commercial queries those are what the answer cites. Publish the self-contained page that states who your product is for and what it costs, so there is something worth lifting. And publish original data, because a small site can own the citation for a finding nobody else has.

First, for commercial intent, being in the answer often means being in the aggregators the answer cites, not just fixing your own pages. If your category has a G2 or Clutch listing and a set of "best of 2026" roundups, your presence in those is part of the citation surface. This is not a license to chase spammy link placements. It is a reason to make sure the legitimate directories and roundups in your space have you listed accurately.

Second, build the retrievable fit-facts page you are missing. The vendor that got left out despite being a perfect match lost the citation to a missing page, not to a stronger competitor. A clean, self-contained page stating who your product is for, what it costs, and what a buyer gets is the single highest-leverage fix for commercial-query visibility. Write the passage an engine would want to lift, and lift it yourself first.

Third, on the informational side, original data is your opening. Primary and first-party pages did better for definitional queries, and pages publishing their own numbers did best of all. A small site cannot outrank an encyclopedia on "what is GEO", but it can own the citation for a specific finding nobody else has published. That is a large part of why we built this benchmark in the first place, and it is the strategy behind the rest of our GEO work.

Methodology and limitations

Treat this as a first-party, ongoing benchmark, not a definitive study. A few things you should know before you cite it:

  • This is a single collection run from July 22, 2026, not a continuing series. The tracking is not running on a schedule right now, so nothing here has been refreshed since that date.
  • The sample is small. It covers a handful of domains we manage plus our own site, and a focused set of SEO and GEO queries, not a broad cross-section of industries.
  • The tracking was automated and AI-assisted. That is what let us cover the whole query set at once, and it also means the source classification is machine-produced and imperfect, not hand-audited line by line.
  • The magnitudes here (dominant, common, rare) are directional summaries of what we logged, not precise statistics. We have deliberately avoided quoting percentages or citation rates we cannot stand behind, and the raw analyses from this run are no longer retained, so none can be recomputed after the fact.
  • AI answers change. Engines rerank sources, roundups get republished, and directories update. A pattern that held in this snapshot can shift, which is why the findings here are dated rather than presented as a standing result.

Anything we add to this page will be dated, so you can see what changed and when. If you want to see where your own site currently stands in AI answers for your buyer queries, a free audit is the simplest place to start. It will show you which of the traits above your important pages already have, and which ones are keeping you out of the answer.