
How to Optimize Your Homepage for AI Traffic

Wikipedia and Reddit dominate LLM citations in the U.S., each accounting for roughly 12–13% of all ChatGPT citations. In Google AI Mode, Fandom.com leads by a significant margin, followed by Wikipedia and YouTube. Our analysis of nearly 600,000 citation events (instances where an AI engine references and links to a specific URL in a generated response) across ChatGPT and Google AI Mode reveals which domains AI engines trust most, and the patterns that explain why. Instead of returning a ranked list of links, LLM systems collect pages from across the web, blend them with internal knowledge and produce a single coherent answer.
Only a handful of sources are cited in each response, often the only opportunity for a website to receive a click. Zero‑click searches for news-related queries increased from 56% to 69% year over year following the rollout of Google’s AI Overviews. In this clickless environment, understanding which domains are cited and how to influence those citations is essential.
Below, I analyzed our data to see which domains ChatGPT and Google AI mode cite most often, discuss how AI engines choose sources, and outline a process for performing citation analysis.
The table below shows the top 20 domains cited by ChatGPT (web browsing mode) and their share of total citations, for the U.S., between January and February 2026.
Data note: Citation counts are drawn from Similarweb’s AI Citation Analysis tool, tracking citation events across monitored prompts in the United States. A “citation event” is recorded each time ChatGPT (web browsing mode) references a domain in a response. Data covers January–February 2026. Citation share is a percentage of total citation events in the measurement window.
| Domain | Citations # | Citations % |
| wikipedia.org | 118,285 | 13.15% |
| reddit.com | 107,680 | 11.97% |
| openai.com | 55,876 | 6.21% |
| walmart.com | 26,118 | 2.90% |
| youtube.com | 23,976 | 2.67% |
| linkedin.com | 21,736 | 2.42% |
| reuters.com | 20,451 | 2.27% |
| nih.gov | 19,962 | 2.22% |
| google.com | 19,478 | 2.17% |
| media-amazon.com | 17,477 | 1.94% |
| wikimedia.org | 17,329 | 1.93% |
| facebook.com | 15,813 | 1.76% |
| ebay.com | 15,730 | 1.75% |
| amazon.com | 15,343 | 1.71% |
| github.com | 14,569 | 1.62% |
| apple.com | 13,342 | 1.48% |
| yahoo.com | 12,942 | 1.44% |
| forbes.com | 12,454 | 1.38% |
| fandom.com | 11,630 | 1.29% |
| squarespace-cdn.com | 11,575 | 1.29% |
The gap between the top two domains and the rest of the list is striking. Wikipedia and Reddit are in a league of their own, each pulling significantly more citations than any other source. This tells you that when ChatGPT needs to ground an answer, it reaches first for broad encyclopedic coverage and then for real-world human opinion, often in the same response.
One of the most unexpected findings is that OpenAI appears third overall, ahead of Reuters, NIH, LinkedIn, and every major retailer in this dataset. ChatGPT is actively pulling from its own documentation, research, and blog content to answer user questions.
This is not a neutral technical quirk. When an AI system preferentially cites its own publisher, it shapes which narratives about AI get amplified and which get buried. For brands operating in or adjacent to the AI category, the strategic implication is direct: getting your content published or cited on high-authority AI-adjacent domains, academic repositories, technology publications, and OpenAI’s own community platforms is more likely to influence ChatGPT’s answers than publishing the same content exclusively on your own site. The engine trusts the sources it already knows.
Retail and marketplace domains are scattered across the top 20, suggesting ChatGPT is fielding a high volume of product, shopping, and purchasing queries, and that commercial pages with structured, detailed content are holding their own in an AI-first world.
LinkedIn, Reuters, Forbes, and Google sit comfortably in the mid-range. They’re not at the top of the list, but they’re a steady presence. For brands in B2B, finance, or news, this is encouraging, authoritative, well-structured content in these spaces is clearly getting picked up.
GitHub’s presence confirms that AI isn’t just for general consumers. Developer-focused, code-heavy content is actively being cited, which matters for any brand operating in the tech space.
Our data also tracked Google AI Mode. The table below lists the top 20 domains cited in AI Mode and their share of total citations for the U.S. during the same period.
Data note: Citation counts are drawn from Similarweb’s AI Citation Analysis tool, tracking citation events across monitored prompts in the United States for Google AI Mode. A “citation event” is recorded each time AI Mode references a domain in a response. Data covers January–February 2026. Citation share is a percentage of total citation events in the measurement window.
| Domain | Citations # | Citations % |
| fandom.com | 42,332 | 7.16% |
| wikipedia.org | 30,792 | 5.21% |
| youtube.com | 29,032 | 4.91% |
| reddit.com | 24,764 | 4.19% |
| google.com | 16,878 | 2.85% |
| facebook.com | 14,410 | 2.44% |
| amazon.com | 10,886 | 1.84% |
| nih.gov | 8,821 | 1.49% |
| github.com | 8,728 | 1.48% |
| apple.com | 8,127 | 1.37% |
| instagram.com | 7,007 | 1.19% |
| microsoft.com | 6,950 | 1.18% |
| quora.com | 6,130 | 1.04% |
| ebay.com | 5,641 | 0.95% |
| linkedin.com | 5,479 | 0.93% |
| imdb.com | 4,784 | 0.81% |
| clevelandclinic.org | 4,529 | 0.77% |
| irs.gov | 4,485 | 0.76% |
| walmart.com | 4,068 | 0.69% |
| medium.com | 3,970 | 0.67% |
The most striking finding is that a fan wiki platform sits at the very top of AI Mode’s citation list, above Wikipedia, YouTube, and Reddit. The instinctive explanation, that AI Mode sees lots of entertainment and gaming queries, is true, but incomplete.
Fandom pages are structurally optimized for exactly what AI engines prefer: individual pages run to thousands of words covering a single specific subject, answers are organized under precise headings, content is maintained by communities with encyclopedic precision, and every page exists to answer one specific question about one specific thing. That is the template. A brand does not need to be in entertainment to apply it, any organization that publishes deep, single-topic reference pages on its area of expertise is building the kind of content AI Mode is designed to surface.
If there’s one universal truth in AI citation behavior, it’s that well-structured reference content and video explanations consistently earn their place, regardless of which engine is generating the answer.
Google.com is among the top five most cited domains in its own AI product. This self-referential pattern is even more pronounced than what we see with OpenAI in ChatGPT. The practical implication goes beyond “try to rank in Google.” Getting your brand featured in Google’s own properties, your Google Business Profile, Google Knowledge Panel, YouTube channel, Google Scholar citations, now feeds directly into AI Mode’s source pool. For many brands, optimizing those owned Google properties is a faster path to AI Mode citations than competing on third-party publishers.
Apple, Microsoft, and Google together form a notable cluster of corporate tech citations, suggesting AI Mode leans heavily on official product documentation and support content. This has real implications for any brand competing in the tech space: publishing clear, authoritative product documentation is a citation strategy, not just a support resource.
NIH, Cleveland Clinic, and the IRS all earn citations in AI Mode, a sign that government and institutional sources carry strong trust signals with the engine, spanning health, science, and personal finance queries.
Community platforms like Reddit, Facebook, Instagram, and Quora together account for nearly 9% of AI Mode citations, showing the engine balances authoritative sources with real-world community knowledge depending on the query type.
AI tools like ChatGPT and Google AI Mode don’t choose sources the same way traditional search engines do. They don’t just look at the top-ranked pages. Instead, they scan a wide range of content across the web and pick a small number of sources that best answer the question.
In the section below, I’ll share a few insights from our Generative AI Landscape report to answer the big question. Note that citation behavior varies across platforms and changes over time, the patterns below reflect consistent signals we’ve observed in our data, though individual results may differ by topic and query type.
ChatGPT and Google AI Mode rely on many of the same content categories. In both tools, news and publisher sites make up the largest share of citations, followed by reviews, business services, and social platforms. This shows that authoritative and widely referenced sources are important across AI systems.
That said, there are small but meaningful differences. Google AI Mode places relatively more weight on reviews, user-generated content, and social media, while ChatGPT leans slightly more toward business services and ecommerce brands.
Our dataset also reveals a meaningful scale difference between the two engines. ChatGPT’s top cited domain alone, Wikipedia, at 118,285 events, generated nearly three times more citation events than Fandom.com, the entire leader in AI Mode at 42,332. This volume gap matters for strategy: ChatGPT is currently the larger citation surface by a significant margin. Brands that treat ChatGPT and AI Mode as a single optimization target are likely leaving citation share on the table in both.

The data shows that AI rarely cites homepages. Most ChatGPT citations come from pages that are several folders deep within a site. These include blog posts, FAQs, how-to guides, product detail pages, and research articles. AI favors detailed, specific pages over broad landing pages.

AI is especially good at answering very specific questions. Content that targets long-tail prompts, like detailed product questions or niche use cases, is more likely to be cited. Creating content that directly answers what users are asking gives AI something clear and useful to reference.
For a given prompt type (like running gear or hotel booking), AI repeatedly pulls from a small group of familiar sources, such as Reddit, Wikipedia, and well-known publishers. Once a domain is trusted for a topic, it tends to show up again and again.
For opinion- or comparison-driven prompts (like running gear), AI favors community platforms and niche publishers. For decision-focused prompts (like booking hotels), it leans more toward major publishers, aggregators, and brand sites. This shows that AI changes its sources based on what the user is trying to do.

To understand and improve my citation footprint, I used Similarweb’s AI Citation Analysis tool and AI Prompts analysis tool. Here’s the process I followed, using OpenAI as the example brand. One finding stood out immediately: the domains with the highest influence scores were not the same domains that drove the most traffic. High citation volume and high citation value are not the same thing, and conflating them leads to misallocated effort. The steps below are structured to keep those two metrics separate.
First, I clarified my goal (e.g., “increase my AI citation share on a specific topic by 5 % in the next quarter”) and chose one or two competitors with overlapping offerings.
The Prompt Analysis report revealed which user questions generated citations for me and where I was missing. In the OpenAI example, prompts like “Which industries benefit most from machine learning?” generated many citations but did not mention the brand, while “What are the most promising AI research collaborations?” produced positive mentions.

The Citation Analysis tab shows where generative engines get their sources from. I sorted sources by influence score to identify high‑authority domains. For example, in the OpenAI case study, top citing domains included arxiv.org, medium.com, en.wikipedia.org, and geeksforgeeks.org.

The Cited URLs table provided granular insights. It showed influence scores for individual pages, such as Wikipedia’s Word Embedding article or specific arXiv papers, and how many prompts referenced each page. I noted which URLs my competitors were cited on and identified gaps where my own site was absent. For a deeper guide on running this process competitively, see our citation gap analysis walkthrough.

Finally, I integrated AI Traffic Analytics to see how many visits each citing domain actually drove. Some domains contribute many citations but little traffic, while others send high‑value visitors despite few citations. This step ensured I focused on sources that matter for my business rather than vanity metrics.
The domains at the top of these lists share something in common: they publish content that’s specific, well-structured, and easy for AI to extract. Wikipedia answers questions directly. Reddit surfaces real opinions. YouTube explains things visually. Fandom goes deep on niche topics. These aren’t coincidences, they’re patterns you can learn from.
The Fandom finding, a niche wiki platform topping AI Mode’s citation list above Wikipedia and Reddit, is the clearest signal that depth on a specific topic beats breadth every time. If Fandom can outrank Wikipedia for entertainment queries, a well-structured product page can outrank Amazon for niche product questions. The formula is the same: structured content, specific answers, and consistent topic ownership.
Citation sources shift frequently, and there’s very little overlap between AI platforms. That means the landscape is still fluid, and there’s a real opportunity for brands that start tracking their citation footprint now. As AI engines become the primary discovery layer for millions of queries, Similarweb’s AI Brand Visibility tools help you see where you stand, which domains and prompts matter most in your category, and how citations connect to actual traffic and business results.
How do AI engines like ChatGPT and Google AI Mode choose which sites to cite?
AI engines scan content from across the web, combine it with their internal knowledge, and select a small number of sources that best answer the user’s question. They prioritize relevance, clarity, structure, and trust over traditional search rankings.
Why are citations more important than rankings in AI search?
In generative search, users often see only one answer and a few citations. With zero-click searches increasing rapidly, citations may be the only way to gain visibility, authority, and traffic.
Which domains are cited most often by ChatGPT?
Wikipedia and Reddit dominate, each accounting for roughly 12-13% of all citations, far ahead of any other source. After them, ChatGPT draws from a broad mix of sources, including its own documentation, major retailers, news outlets, and professional platforms.
Which domains are cited most often by Google AI Mode?
Fandom.com leads by a notable margin, followed by Wikipedia and YouTube. Together, these three domains account for over 17% of all citations in AI Mode, making them the most consistently referenced sources across entertainment, reference, and video content.
How do citation sources differ between ChatGPT and Google AI Mode?
Both engines cite many of the same domains, but the weighting differs. Wikipedia and Reddit appear in both top 20 lists, but ChatGPT leans more heavily toward commerce, news, and professional sources, while AI Mode shows a stronger pull toward entertainment, fan communities, and big tech documentation. Fandom.com, topping AI Mode’s list, while not even leading in ChatGPT, is the clearest example of how differently the two engines approach the same web.
Why does AI cite Reddit and other community platforms so often?
Community platforms offer real-world opinions, comparisons, and validation. AI uses them heavily for subjective or experience-based prompts, such as product recommendations.
Does AI prefer deep pages over homepages?
Yes. AI rarely cites homepages. Most citations come from pages several levels deep, such as blog posts, FAQs, research articles, and detailed product pages. These pages answer specific questions more clearly, which aligns with AEO principles.
What kind of content is most likely to earn AI citations?
Hyper-specific content performs best. Pages that address long-tail questions, niche use cases, or detailed comparisons are more likely to be cited. This is closely tied to AI prompt analysis, which helps identify what users are actually asking.
Why do the same domains appear repeatedly for certain topics?
AI tends to reuse sources it already trusts for a topic. Once a domain proves reliable for a category, running gear or hotel booking, it often reappears across similar prompts. This creates a strong incentive to analyze citation gaps and identify where competitors are winning.
How can brands measure and improve their AI citation visibility?
Brands can use citation and prompt-level data to track where they appear, where competitors dominate, and which sources matter most. Tools like Similarweb’s AI Brand Visibility help connect citations to traffic and business outcomes.
Is optimizing for AI citations a one-time effort?
No. Citation sources change frequently, and overlap between AI platforms is low. Brands need ongoing monitoring, content updates, and distribution across trusted domains. Investing in AI visibility now helps future-proof your strategy as search continues to evolve.
Senior SEO Specialist at Similarweb
Maayan is a senior SEO specialist with 7+ years of experience in SEO. She loves complex research projects, creating SEO strategies and performing technical audits.
Give it a try or talk to our insights team — don’t worry, it’s free!