How ChatGPT Picks Sources (And What SMEs Must Do to Get Cited in 2026)
Network-level evidence suggests ChatGPT doesn’t “rank your site” the way Google does. It classifies the query, triggers (or skips) web retrieval, then fetches pages through specific pipelines—often commercial scrapers—and only sometimes turns those pages into citations. Here’s what changed, why it matters, and what to do next with an execution-first system like AYSA.
AI Search visibility is quickly becoming a board-level concern for small and mid-sized businesses. Not because “SEO is dead,” but because the way answers get assembled is changing. In classic search, you fought for rankings. In ChatGPT-style experiences, you fight for inclusion (being fetched), Attribution (being cited), and recognition (being mentioned). Those are related—but not the same.
This editorial is based on network-level observations published by Search Engine Journal about how ChatGPT selects and labels sources underneath the UI—an approach that focuses on what the system requests and fetches, not just what it prints in the final answer. The original reporting is here: Search Engine Journal: “How ChatGPT Actually Picks Sources (I Read The Network Traffic, Not The Outputs)”.
I’m writing this as Marius Dosinescu, from AYSA.ai, with a practical bias: businesses don’t need more theory about “GEO” or “AEO.” They need to know what changed, what that does to demand capture, and what to implement this month—on the website and off it—so they can be eligible for citations when AI systems decide they need outside sources.
Concise summary

- AI answers start with Query Classification. Some queries trigger web retrieval; others may not. If your target query doesn’t fetch the web, your on-page SEO won’t matter for that moment.
- There are multiple “source pipelines.” Evidence suggests ChatGPT fetches web pages via different backends (including commercial scraping providers) and labels each source internally.
- Being fetched is not being cited. Your page can be retrieved but never appear as a citation; your brand can be mentioned without your site being used as evidence.
- Scrapable facts beat polished design. If key facts (pricing, specs, eligibility, hours, shipping) aren’t in plain HTML, you’re harder to verify and easier to exclude.
- Third-party validation is the moat. For “best X” claims, the system tends to cite external reviews, communities, and reference sources more than vendor self-claims.
- Execution is the bottleneck. The winners will be the companies that can monitor AI visibility, ship website changes quickly, and build credible third-party coverage—without endless audit decks.
Table of contents

- What changed: why source mechanics matter more than GEO slogans
- The big shift: ChatGPT is a query router, not a traditional search engine
- What the network evidence suggests: multiple retrieval pipelines and internal labels
- The uncomfortable truth: some queries may never reach the web
- Fan-out queries: one question can trigger dozens of sub-searches
- Fetched vs cited vs mentioned: the three outcomes you must separate
- What tends to win citations: verifiable facts + independent validation
- Technical foundation: make your site easy to fetch, parse, and quote
- Content foundation: write for claims, not keywords
- Authority foundation: PR, reviews, communities, and why you can’t cite yourself
- A concrete SME scenario: an ecommerce brand trying to win “best” and “pricing” queries
- What agencies must rethink: deliver implementation, not just recommendations
- Where AYSA fits: monitoring + approved execution for AI Search
- What to do next: a 30–60–90 day action plan
- Sources and further reading
What changed: why source mechanics matter more than GEO slogans

For the past year, “GEO” (Generative Engine Optimization) has been marketed like a new set of ranking hacks: post more on Reddit, publish listicles, sprinkle schema, and you’ll “rank in ChatGPT.” It’s comforting because it sounds like SEO with different nouns.
But the network-level evidence described by Search Engine Journal points to something less glamorous and more operational: modern AI answer systems behave like decision engines. They decide whether to retrieve. They decide which retrieval pipeline to use. They decide which fetched pages are “good enough” to support a specific sentence. And then they decide whether your brand deserves a mention even if your site didn’t earn a citation.
That means your old mental model—“publish page → get indexed → rank → get clicks”—is incomplete. The new model looks more like:
- Classify the question (what kind of task is this?)
- Retrieve or don’t retrieve (do we need the web?)
- Fetch candidates (from one of several pipelines)
- Extract facts (pricing, specs, definitions, comparisons)
- Assemble an answer (with optional citations)
The business takeaway is blunt: if you only optimize “content quality” without understanding how retrieval and citations work, you can spend months improving pages that are never eligible to be used.
The big shift: ChatGPT is a query router, not a traditional search engine
One of the most important implications in the SEJ reporting is that ChatGPT appears to assign each user turn (question) into a use case category. That category then governs whether it browses the web, uses specialized tools (like shopping or local), or answers from internal knowledge.
This is a different “front door” than Google Search. Google nearly always starts from web retrieval and then applies ranking. In AI chat, the engine can decide that web retrieval is unnecessary, or too slow, or too risky—and proceed without it.
For a business owner, this changes how you evaluate opportunity:
- If the query type often does not retrieve, then your strategy is less about “publishing a page” and more about long-run brand authority (getting into training corpora, being referenced widely, being in datasets that models learn from). That’s slower and harder.
- If the query type usually does retrieve, then you can compete today by ensuring your pages are fetchable, extractable, and quotable—and by ensuring independent third parties talk about you.
In other words: your first job is to learn which of your money queries are retrieval-triggering queries. If you’re not sure, you’re guessing where to invest.
What the network evidence suggests: multiple retrieval pipelines and internal labels
According to the SEJ analysis, network traffic inspection revealed an internal field that labels where each fetched result came from. The article describes multiple values (for example, an open-web “SERP” baseline, a publisher/reference allowlist tier, and two commercial scraping-provider pipelines).
I’m intentionally not repeating any percentages here, because (as the original author clearly notes) those are directional observations based on a limited sample and a single account. The structural insight—the existence of internal source labels and multiple retrieval paths—is the part that matters strategically.
Why should an SME care about which pipeline fetches them?
- Different pipelines have different failure modes. A page that looks great to a human might fail in a scraping pipeline if key content requires heavy JavaScript, geo-blocks, cookie walls, or delayed rendering.
- Commercial pipelines tend to over-index on “extractable facts.” If the system is trying to verify a price, a spec, hours, or availability, it needs content it can parse quickly and confidently.
- Publisher/reference tiers are not easily “optimized into.” If some sources are effectively privileged (licensed, allowlisted, or historically trusted), then SMEs should focus on earning mentions in those sources rather than trying to become them.
This is why generic GEO advice often disappoints. You can “write better content” all day, but if your pricing is inside an image, your product comparison table is rendered client-side, or your pages block bots, you’re invisible to the exact systems that need your facts.
The uncomfortable truth: some queries may never reach the web
The SEJ piece highlights that some questions appear to be handled without any web retrieval. That’s not a moral judgement; it’s a product choice. Web browsing is slower, more expensive, and introduces additional risk (bad sources, paywalls, errors). For certain tasks, the model may decide it can answer directly.
That reality forces a new kind of keyword research for AI Search:
- “Will the system browse?” becomes as important as “How many people search?”
- Query wording matters. Two queries on the same topic can trigger different behavior depending on how they’re phrased (how-to vs “latest,” generic vs “near me,” product vs “with reviews”).
- Risky industries are not immune. Even sensitive topics may sometimes be answered without retrieval, which makes third-party validation and clear disclosures on your own site even more important when retrieval does happen.
From an execution standpoint, I recommend teams build a simple list:
- Top 25 revenue-driving queries (by leads, calls, bookings, carts).
- Top 25 “comparison” and “best” queries that influence buying decisions.
- Top 25 support/education queries that reduce churn.
Then test and monitor which ones appear to trigger retrieval behavior across AI platforms you care about. AYSA’s direction is exactly this: track AI search visibility as a system, not as anecdotes. See AI search visibility at AYSA and monitoring.
Fan-out queries: one question can trigger dozens of sub-searches
The most underappreciated behavior described in the SEJ reporting is what happens when the AI system decides to do deeper research (often called “thinking” modes, or simply more intensive reasoning). A single comparison prompt can fan out into many sub-queries: direct vendor checks, “site:” probes, and verification searches for specific claims.
This matters because it breaks the simplistic “optimize for the exact keyword” mindset. The system may rewrite the query, decompose it, and then execute multiple targeted lookups like:
- “site:domain.com/pricing” style probes
- verification searches for a specific price, plan name, or feature
- broader exploration to find alternatives the user didn’t mention
In practical terms: you are no longer optimizing for a single query string—you are optimizing to survive a fact-checking process.
So if your pricing page:
- loads prices only after a JS toggle,
- renders prices in an image,
- hides plan details behind accordions that don’t render server-side,
- requires a region selector or blocks certain user agents,
…you’ve introduced friction precisely where the AI system wants speed and certainty. That friction doesn’t just reduce “SEO.” It can remove you from the candidate set entirely when the system is trying to verify a number.
Fetched vs cited vs mentioned: the three outcomes you must separate
Most businesses lump everything into “we showed up in ChatGPT.” That’s a mistake. The SEJ analysis stresses a distinction that should reshape your measurement:
- Fetched: your page is retrieved into the model’s context. Invisible to users.
- Cited: your page is attached as the source for a specific sentence (a footnote-style link).
- Mentioned: your brand appears in the answer, possibly linked, without being the evidence for the claim.
These outcomes can diverge wildly:
- You can be fetched because the system checked your pricing page, but not cited because it used a third-party review to justify the recommendation.
- You can be mentioned because you’re a known brand, but not cited because the system didn’t need your site to support any sentence.
- You can be cited for a narrow, verifiable claim (a spec, a policy, a price) even if you’re not mentioned prominently as a recommendation.
This is why “AI mentions” and “AI citations” should not be treated as the same KPI. They map to different mechanics—and different strategies.
If you’re building a measurement stack, start by separating:
- Query coverage: which of your target queries produce retrieval-based answers?
- Citation coverage: which domains get cited for which claim types?
- Brand presence: do you appear in recommended shortlists?
AYSA is built around this separation: monitor visibility, identify what’s missing, prepare changes, ask for approval, and then execute the accepted changes. See AYSA AI SEO tools and monitoring.
What tends to win citations: verifiable facts + independent validation
When AI systems choose citations, they’re solving a problem: “What source should I attach to this specific sentence so the user can verify it?” That pushes them toward sources with two qualities:
- Extractability: the relevant claim is present as plain text and easy to quote.
- Credibility: the source is a suitable witness for that claim.
Here’s the nuance that most marketing teams miss:
- Your own site is a credible witness for your own policies and facts (pricing, hours, shipping terms, product specs—assuming they’re clear and consistent).
- Your own site is not a credible witness for why you’re the best. “We are the best CRM for contractors” is marketing; the AI system will often prefer third-party validation to support that sentence.
This aligns with what the SEJ reporting suggests: independent sources—publishers, review hubs, community threads—often play an outsized role in citations for comparative claims, while vendor pages may be used for narrow, factual verification.
So the strategy becomes two-lane:
- Lane 1: Own your facts. Make your site the easiest place to verify your price, features, eligibility, and policies.
- Lane 2: Earn your reputation elsewhere. Get reviewed, listed, discussed, and compared on sources likely to be retrievable and quotable.
Technical foundation: make your site easy to fetch, parse, and quote
Technical SEO used to be about crawlability for Googlebot, indexation, and performance. Those still matter. But AI retrieval introduces a stricter requirement: your key claims must survive a fast, automated fetch and extraction process.
1) Put critical facts in plain HTML (not images, not PDFs, not JS-only)
The SEJ reporting mentions price verification behavior and “find” style extraction (looking for currency symbols and plan names). Even without reproducing exact mechanics, the principle holds:
- Put prices, plan names, model numbers, dosage information, hours, addresses, and key constraints in plain text HTML.
- Do not rely on images for tables that matter (comparison charts, pricing grids).
- Use PDFs as downloadable references, not as the only location of critical facts.
2) Reduce bot friction (without harming user privacy)
Many SMEs unintentionally add barriers:
- cookie walls that block content before acceptance
- geo redirects and forced language interstitials
- anti-bot systems that challenge anything unfamiliar
- heavy client-side rendering that hides content on first fetch
You don’t need to remove security. You need to ensure that legitimate retrieval can access your public facts. For regulated industries, publish a clear public “facts” layer (policies, services, pricing ranges) even if personalized flows require logins.
3) Build one strong page per claim (not 50 thin pages)
A key operational point in the SEJ piece: systems may deduplicate by domain when selecting results. If that’s true (and it’s consistent with how many retrieval systems reduce redundancy), then mass-producing thin pages for every long-tail variation is counterproductive.
Instead:
- Create a small number of authoritative, comprehensive pages.
- Make each page clearly responsible for a set of claims (pricing, comparisons, policies, product specs, how-to).
- Ensure each page has clean internal structure (headings, tables, short definitional blocks).
4) Use structured content (and schema) as an assist, not a crutch
Schema can help disambiguate entities and properties, but it doesn’t replace text. If the retrieval system can’t easily fetch your content, schema won’t save you. Treat schema like packaging: useful, but not the product.
If you need help operationalizing these changes across many pages, AYSA is designed to prepare change sets and execute them after approval—so technical fixes don’t die in a backlog. See AI SEO Tools.
Content foundation: write for claims, not keywords
Classic content SEO often starts with a keyword list and a publishing calendar. For AI citations, your starting point should be:
What claims do customers need to believe in order to buy?
Examples of claim types that are citation-friendly:
- “This product supports X standard”
- “This clinic offers Y treatment for Z condition”
- “This tool integrates with A, B, C”
- “Shipping takes 2–4 business days in the U.S.”
- “Refund policy is 30 days”
Then structure pages so those claims are easy to extract:
- Put the claim in a short, explicit sentence.
- Follow with a brief explanation and constraints.
- Include a “last updated” date where appropriate (especially for pricing, policies, and guidelines).
- Use consistent terminology across the site so verification doesn’t contradict itself.
Comparison pages still matter—but only if they’re honest
Many brands publish “X vs Y” pages that are thin and biased. In an AI world, that can backfire: if the system cross-checks and sees your page as self-serving or inconsistent, it may use it only for factual extraction (or ignore it entirely) and cite a third party for the conclusion.
Better approach:
- Write comparisons that admit tradeoffs.
- Use clear criteria.
- Back factual statements with documentation (your own docs for your features; external sources for independent claims).
Authority foundation: PR, reviews, communities, and why you can’t cite yourself
If you sell something, you have an inherent conflict of interest. That doesn’t mean your site is untrustworthy—it means it’s not the best witness for claims like “best,” “top,” “most reliable,” or “worth it.”
The SEJ reporting reinforces a practical truth: third-party validation is a primary lever for AI citations in recommendation contexts.
What “third-party validation” looks like for SMEs
You do not need a Wall Street Journal profile to build validation. You need a portfolio of credible references that match your industry:
- Industry review sites and comparison hubs (where applicable).
- Community discussions that are text-rich and accessible.
- Partner pages (integrations, resellers, associations).
- Local and regional press for location-based businesses.
- Academic or standards references if you operate in regulated or technical domains.
The goal is not “backlinks” in the old sense. The goal is to appear in sources that retrieval systems can fetch and quote—and that a model would view as appropriate evidence for a specific sentence.
A note on communities (including Reddit): opportunity and risk
Community content can be influential because it’s conversational, text-first, and contains first-hand experiences. But it also introduces brand risk and volatility.
If you choose to invest there:
- Don’t astroturf. It’s detectable and reputationally expensive.
- Support customers publicly with real help.
- Encourage satisfied customers to share their experience naturally (without scripts).
- Use those discussions to inform your FAQ and documentation so your site becomes the best factual witness.
A concrete SME scenario: an ecommerce brand trying to win “best” and “pricing” queries
Let’s make this real with an SME scenario. Imagine a direct-to-consumer ecommerce brand selling specialty coffee equipment: grinders, kettles, and accessories. The team is small—founder plus one marketer—and they’re seeing fewer clicks from classic SEO because shoppers are getting “best grinder under $200” answers in AI tools.
They decide to “do GEO” and publish 40 blog posts. Traffic barely moves. Why? Because they optimized content volume, not retrieval/citation readiness.
What’s broken (typical SME issues)
- Pricing is displayed in a JS component and doesn’t render cleanly on first load.
- Key specs (burr size, materials, warranty) are in product images, not text.
- Shipping and returns are scattered across multiple pages with inconsistent wording.
- There are no third-party reviews that can be used as independent evidence.
What they do instead (execution-first)
- Build a “facts layer” on product pages. A simple HTML block: price, warranty, materials, dimensions, compatibility, and returns—consistent across the catalog.
- Create one authoritative buying guide. Not 40 posts. One strong guide with clear categories and criteria, linked to supporting product pages.
- Publish a transparent comparison page. Include tradeoffs and who each product is for.
- Earn third-party validation. Send products to a handful of credible reviewers; encourage verified purchasers to leave reviews on platforms that are accessible and text-based.
- Monitor mentions and citations. Track whether AI answers cite their product specs and whether they are mentioned in shortlists.
Notice the strategy split:
- For pricing/spec verification, they want their own pages to be the best source.
- For “best” recommendations, they want independent sources to do the praising.
This is the kind of program AYSA is designed to support: monitor where you stand, generate a prioritized change plan, ask for approval, and execute changes so the site becomes citation-ready. Start here: AI Search Visibility and AYSA Pricing.
What agencies must rethink: deliver implementation, not just recommendations
Agencies are feeling pressure from two sides:
- Clients want results faster because AI answers compress the funnel.
- Traditional SEO deliverables (audits, content calendars) don’t guarantee citations.
The agencies that win in 2026 will change their operating model:
- From audits to shipping. Every month must include implemented technical changes and content improvements, not just analysis.
- From keyword lists to claim maps. A living inventory of claims, where they live on-site, and which third-party sources validate them.
- From link building to evidence building. PR, reviews, partner pages, and community credibility—built in ways retrieval systems can use.
- From “rank tracking” to AI visibility monitoring. Mentions and citations are not vanity metrics if they’re tied to revenue-driving query sets.
AYSA’s “approved execution” model is built for this new reality: the system monitors, proposes changes, asks for approval, and executes accepted updates—reducing the gap between strategy and implementation. If you’re an agency, the operational question becomes: how many approved changes can you ship per month, with quality control?
To explore that approach, see AI SEO tools and the latest operational guidance in our blog.
Where AYSA fits: monitoring + approved execution for AI Search
Most “AI SEO” tools today stop at reporting: they show you mentions, citations, or prompt outputs. Useful, but incomplete. Visibility isn’t the end goal—implementation is.
AYSA is positioned as an execution system for AI Search and modern SEO:
- Monitors your AI search visibility and technical readiness signals. (Monitoring)
- Prepares recommended website changes aligned to citation mechanics (HTML-first facts, consolidated pages per claim, internal linking, content restructuring).
- Asks for approval so humans stay in control—especially important for regulated industries and brand messaging.
- Executes accepted changes so improvements go live, not into a backlog.
This isn’t a philosophical difference. It’s a throughput difference. The teams that win will be the ones that can iterate: test what queries retrieve, improve fetchability, strengthen claims, build third-party validation, and re-measure. That’s a loop—AYSA is built to run it.
If you want to see how we frame this end-to-end, start here:
What to do next: a 30–60–90 day action plan
This is the practical plan I’d give an SME (or an agency serving SMEs) that wants measurable progress without wasting quarters on vague GEO tactics.
Days 1–30: establish eligibility (can you be fetched and quoted?)
- Inventory your money pages: pricing, product/service pages, location pages, comparison pages.
- Move critical facts into HTML: pricing, specs, warranty, shipping, returns, eligibility, hours, address.
- Reduce extraction friction: remove or adjust blockers that hide main content (without compromising compliance).
- Consolidate thin content: replace 10 weak pages with 1 strong “source-of-truth” page for a claim cluster.
- Set a baseline: monitor where you are being mentioned/cited for your key query set. (Use a system like AYSA Monitoring.)
Days 31–60: build claim architecture (make answers easy to assemble)
- Create a claim map: list the 20–50 claims that close deals (features, policies, benefits, constraints).
- Assign each claim a home page on your site where it’s stated clearly and updated regularly.
- Write “verification-friendly” sections: short paragraphs, tables, FAQs that state facts plainly.
- Strengthen internal linking so those pages are easy to discover and interpret.
Days 61–90: earn third-party evidence (so you can be cited for conclusions)
- Pick 5–10 credible third-party targets in your industry (review sites, associations, partners, local press).
- Run an evidence campaign: product reviews, expert commentary, case studies with partners, community support.
- Encourage real customer narratives where appropriate—authentic, text-based, and specific.
- Measure again: which claims now earn citations, and where are you still missing independent validation?
What to do next (action list)
- Choose 10 high-intent queries you care about and document whether AI tools retrieve and cite sources for them.
- Audit your pricing/spec pages: can a fast fetch read the key numbers in HTML?
- Consolidate thin pages into one authoritative “source-of-truth” page for a core topic.
- List 10 third-party sources you want to be referenced in (or reviewed by) over the next 90 days.
- Set up monitoring so you can see mentions/citations shifting over time. Start with AYSA AI Search Visibility.
Sources and further reading
- Search Engine Journal — How ChatGPT Actually Picks Sources (network traffic analysis)
- Search Engine Journal — SEO section (ongoing coverage)
- Search Engine Journal — SEO News
- Search Engine Journal — Link Building (context for third-party validation)
Related AYSA resources
Continue the AI search topic inside AYSA.
Use these pages to connect the article with AI SEO tools, AI visibility monitoring, AI Overviews and approved website execution.
Turn this topic into a website action plan.
Use these AYSA hubs to move from reading to technical fixes, AI visibility monitoring, research, glossary context and approval-first SEO execution.