
How to Optimize Your Homepage for AI Traffic

TL;DR: Information gain is no longer a content quality trophy. It is the entry requirement for AI citation eligibility.
A new study by On-Page.ai scored 150 pages holding top-3 positions on Google for content originality. One in four graded ‘mostly shared’: the content was semantically nearly identical to the other pages ranking alongside it. And yet they rank. That happens because Google weighs authority, backlinks, and domain age alongside content originality, which means that unoriginal pages can still hold top positions.
That is not how it works in AI search. AI engines don’t weigh authority or backlinks when selecting sources. They select based on what a page contributes that other pages don’t. Which means a quarter of the current top-3 pages are ranking on signals that AI systems don’t recognize, and their visibility in AI-generated answers reflects that.
In this piece, I’ll walk you through what information gain actually is and how to calculate it, what the data shows about where most pages currently stand, why AI search changes the stakes, what Similarweb’s own brand-level AI visibility data reveals about who’s winning citations and why, and what you can do about it.
Information gain, as a concept applied to SEO, measures how much new semantic content a page adds beyond the other pages ranking for the same query. It is not about word count, keyword density, or comprehensiveness in the traditional sense. It is a set-difference calculation: what does your page say that the cumulative SERP does not?
The concept entered the SEO discourse through a Google patent granted in June 2022 titled “Contextual Estimation of Link Information Gain”. The patent describes a system for scoring documents based on the additional information they contribute beyond previously viewed results, not for an abstract user but for a specific user who has already been exposed to certain documents. A page that restates what the user just read scores near zero. A page that introduces genuinely new content scores high.
The mathematical intuition is straightforward. Think of your page P and the cohort of pages already ranking for your query B:
Information Gain (G) = P \ B
In plain terms, G is everything in your page that is not semantically present in the union of competitor pages. The score (0-100) grades the size and quality of that residual.
What falls into the baseline (P ∩ B, the overlap):
What counts as gain (P \ B, your contribution):
Note what is explicitly absent from the gain column: length. More words do not move the score. More original words do.

Google has not published a tool for this, and its patent describes a capability rather than a confirmed live ranking system. What SEOs can do (and should) is approximate the calculation using available methods. There are two approaches, one manual and one tool-assisted.
I’ve run this on probably 40 pages over the last six months. The part that surprises people most is Step 4: most pages have fewer than three unique data points when you actually count. It takes about 45 minutes per target keyword and produces an honest picture of where you stand. Copy this free audit template for repetitive use.
Copy the audit template for your own use
For teams comfortable with Python, a more precise approximation uses text embeddings and cosine similarity. The cosine similarity formalization of IGS reduces to:
IGS ≈ 1 − max(cosine_similarity(embed(P), embed(Ci)))
Where P is your page’s main content, C₁ through Cₙ are the competitor pages, and the maximum cosine similarity represents how close your page is to its nearest competitor. A score of 1.0 means total uniqueness. A score of 0.0 means complete overlap.
In practice:
A page scoring above 0.50 on this metric (meaning its maximum similarity to any single competitor is below 0.50) is meaningfully differentiated. Below 0.30: rework before publishing.
The study’s median top-3 page scored 52/100 on its semantic scoring. Calibrate accordingly.
The 0.5 threshold sounds precise until you run it the first time. Most competitive topics cluster between 0.35 and 0.55, which makes that cutoff a genuine differentiator rather than an arbitrary line.
On-Page.ai’s free Information Gain Checker runs the semantic comparison against the live ranking cohort for any URL. It is the fastest option for a semantic spot check.
For AEO gap analysis, rather than semantic scoring, Similarweb’s citation analysis tool does something different: it identifies which topic questions the SERP leaves unanswered, which is the complementary layer that follows once you know your semantic score.
One in four pages holding a top-3 position scores below 40/100 on semantic originality. That is the current state of the SERP. The study that produced this finding scored 150 pages across 50 keywords in 10 verticals, and it is currently the only published large-scale measurement of SERP originality available.
The score is a semantic metric, not a confirmed Google signal, but it is the most useful benchmark we have.
That means roughly half of a typical top-ranking page’s content is semantically present in the other pages ranking alongside it. The most-covered content (definitions, standard advice, widely cited statistics) repeats verbatim-equivalent across most pages in a SERP cohort. Only the residual contributes to the score.
One in four pages currently holding a top-3 position adds little beyond what the SERP already contains. It ranks on authority, not originality. This is the number most practitioners find counterintuitive: you can rank near the top of page one while contributing almost nothing new.
Among the top 3, information gain does not distinguish between positions. Both positions share that median, which is the counterintuitive finding. Whatever combination of signals sustains those positions, content originality is not what separates them.
This is the most important number in the study for practitioners who already rank: if you are at position 2 and optimizing for information gain to reach position 1, the data does not support that investment. You are competing on other signals.
Below the top 3, mostly-shared content becomes far more common. An exploratory extension of the study scored pages at positions 4, 7, and 10. The share of mostly-shared pages jumped to 37-40%, nearly twice the 24% rate in the top 3. The pattern is consistent: originality correlates with getting into the top 3, not with where you sit within it.
Legal content scored highest (median 62/100), reflecting the specificity requirements of jurisdiction-based advice. Technology and dev scored 60.
At the bottom: health (42), ecommerce (43), and B2B SaaS (44).

For those in B2B SaaS specifically (and Similarweb’s audience sits squarely here): The average B2B SaaS page ranking today adds almost nothing new compared to the other pages ranking for the same query. The competitive gap for a page with genuine original data is larger than it appears.
The practical read from these numbers: the floor for the top-3 Google ranking is lower than most people think. You can clear it without being particularly original. The question is whether clearing the Google floor is still sufficient. That question brings us to why this matters more in 2026 than it did in 2023.

AI answer engines are, architecturally, information-gain engines. This is not a metaphor. It describes how retrieval-augmented generation (RAG) works.
When a user submits a query to ChatGPT, Perplexity, or Google AI Mode, the system does not select one authoritative source and summarize it. It retrieves passages from multiple sources, synthesizes them into a coherent answer, and cites the sources it drew from. The selection logic at the retrieval stage asks a version of the same question: does this passage add something to the synthesis that other retrieved passages don’t?
A page that restates what five other pages already say is not worth citing. The synthesis already has that content. A page with a statistic nobody else has, a framework nobody else named, or an answer to a question nobody else answered: that is what citation engines pull. AirOps’ April 2026 citation analysis (covered by Kevin Indig) found that content with 5 to 7 statistics has a 20% higher likelihood of citation than content without them. Adding quantified claims is not decoration. It is the mechanism.
This creates a gap between what it takes to rank in Google and what it takes to be cited in AI search. Google can sustain a mostly-shared page in the top 3 by weighing domain authority, link profiles, and user signals alongside content novelty. AI answer engines weigh only what they can extract and synthesize. There is no link equity for a passage. There is no domain authority that makes a restatement worth citing.
The implication: a page can hold a top-3 Google position and still be effectively invisible in AI-generated answers for the same query, because the information gain it provides to the synthesis is zero. In a world where AI Overviews appear on 48% of all Google queries, up from 31% a year ago, and ChatGPT processes billions of monthly queries, that gap has direct commercial consequences.

It looks like a brand that punches well above its weight in AI citations. Not the biggest name in the category, not the highest domain authority, but the one with the most specific, structured answers to the questions users actually ask AI systems.
Similarweb’s 2026 AI Brand Visibility Report tracked citation patterns across six sectors and identified exactly this pattern: brands whose AI citation rank is substantially higher than their branded search rank.
| Brand | Sector | AI citation rank | Branded search rank | Delta |
|---|---|---|---|---|
| NerdWallet | Finance | 7 | 73 | +66 |
| WhoWhatWear | Fashion | 27 | 96 | +69 |
| Travelmath | Travel | 31 | 91 | +60 |
| Bankrate | Finance | 13 | 81 | +68 |
| eCosmetics | Beauty | 26 | 84 | +58 |
| Dermstore | Beauty | 20 | 73 | +53 |
The common trait across all overachievers: specialist content, structured informational depth, comparison utility, and specific answers.
Looking at NerdWallet’s live data in the AI citation tool today, it appears across car insurance queries: cheap coverage comparisons, bundling auto with renters or homeowners insurance, and policy review guides.

The pattern is consistent: every citation traces to a page that answers a specific, comparative question with structured guidance that generic insurance content does not provide. The cheap full coverage comparison, the bundling guide, and the policy add-ons breakdown.
This is information gain as a citation strategy, and it is working at scale.
NerdWallet and Bankrate are not ranking in AI answers because they have the most brand awareness in finance. They are ranking because their content contains specific, structured, calculable answers to questions users ask AI systems. WhoWhatWear provides specific editorial guidance that functions as a direct answer to styling queries. Travelmath’s entire value proposition is original route-calculation data that no generalist travel site carries.

These are, in content terms, high-information-gain pages operating in categories where the established brands are ‘mostly-shared’. The information gain study found B2B SaaS and ecommerce at the bottom of the originality distribution. The visibility index shows that the brands overperforming in those same categories are the ones with the deepest, most specific content.
B&H Photo’s AI visibility momentum score of 296.9 tells the same story from a different angle. Similarweb’s AI visibility momentum score measures citation acceleration: a score of 296.9 means citation frequency grew nearly 3x over the tracked period, in a category where Apple sits at 100/100 on brand visibility.
A specialist retailer accelerating that fast against a brand with perfect visibility is not a fluke. It is what happens when product comparison pages, technical specification guides, and specific buying advice are exactly what AI systems extract when users ask “which camera should I buy” or “what’s the best mirrorless under $2,000.”

The argument for investing in information gain is not purely about content quality. It is about a revenue channel that most teams’ analytics cannot currently see.
Similarweb’s Downstream Impact of AI Visibility study tracked real user journeys across Finance, Travel, and Beauty, following users who asked ChatGPT a category-relevant question and received a specific brand recommendation.
Methodology: The study tracked whether those users visited the recommended brand’s website within 7 days of the AI conversation. Users who had not visited the site in the previous four weeks were the sample, to isolate genuine new acquisitions.
The headline finding: users who were recommended a brand by ChatGPT were 2.5 times more likely to visit that brand’s website in the following seven days than users who were recommended a competitor brand. The effect was symmetrical: when a competitor received the recommendation, the traffic went to them.
Specific brand pairs from the study:

The Ulta data is particularly instructive in light of what Similarweb’s live AI visibility tracker shows today (June 2026): across tracked beauty prompts covering cruelty-free makeup, skincare for specific concerns, vegan products, and salon services, Ulta is entirely absent from AI-cited sources. Specialty sources, brand-specific pages, and editorial outlets are capturing the citations.

The absence is not a brand awareness problem. Ulta has enormous brand awareness. It is an information-gain problem: the category questions AI systems answer in beauty are not the ones Ulta’s content currently answers with specificity and depth.
The visit does not arrive via an AI referral link. It arrives via search or direct. 55.9% of AI-influenced visits reach the brand site through search, compared to 40.4% for standard visits. The user remembers the recommendation, opens a new browser tab, searches the brand name, and navigates to the site.

Standard analytics attributes this as branded search or direct traffic. The AI recommendation is invisible in the funnel. This attribution blind spot has a predictable consequence: SEO teams systematically underestimate the ROI of AI visibility investment because they cannot see the channel.
AI-influenced visitors also arrive with higher intent: they view approximately twice as many pages and spend approximately twice as long on site compared to standard visitors. They have already completed the research phase in the AI conversation. The website visit is the last step before purchase, not the first step in discovery.
This is what makes information gain a commercial floor rather than an editorial preference. If your page’s information gain is too low to earn an AI citation, a competitor’s page earns it instead. The visit happens, and it goes to them. The zero-sum framing is not rhetorical. Similarweb’s data makes it literal.
Information gain theory without a workflow is just another SEO concept to feel vaguely guilty about. Here is the five-step audit framework I use. I call it DELTA. It operationalizes the concepts above into a repeatable process.

Pull the top 3-5 pages ranking for your target keyword. These are your comparison set, not your inspiration. Strip them to their main content (remove navigation, headers, and footers).
Run a quick scan of each H2: what specific claims, statistics, and examples does each one make? This is your baseline: everything already in the cohort.
For your own page (or the page you are planning), identify every claim that is not present in the baseline. Be rigorous. “You should create original research” is baseline advice: every page says it.
Your own proprietary Similarweb data shows that the B2B SaaS median information gain score is 44/100 and that 37-40% of positions 4-10 are mostly-shared: that is gain. Work through your page section by section, flagging each non-baseline element.
Use PAA, Perplexity’s related queries, and Similarweb’s answer engine optimization tools to identify the questions readers have that none of the top-3 pages answer.
The information gain study found at least one such question in 90% of SERPs.
These gaps are your highest-value content investments: unclaimed territory where a single well-sourced answer can produce a citable content chunk no competitor has.
Count the numeric figures in your page that appear nowhere in the competitor cohort. The median is 4. The threshold that separates the mostly-shared grade band from the moderately-original band is roughly 5-6 unique data points. To reach the highly-original grade (70+), target 15 or more.
This is not about padding. Junk statistics that could be fabricated add numbers, not information gain. It is about having a genuine quantitative moat: survey data, platform data, proprietary analysis, or original measurements.
For each H2 section, ask: could an AI system extract this section as a standalone, citable passage? Each section must be independently answerable, with no pronouns referencing earlier sections, no “as we discussed above,” no context dependencies. If a section fails this test, it cannot be cited even if it contains original content.
This is the AEO structural layer of information gain: the content can be original and still be uncitable if it is not retrievable as an atomic chunk. A competitor with a lower semantic score but cleaner extractable sections will get cited more often. The AI visibility data consistently bears this out.
| Dimension | 0 | 1 | 2 |
|---|---|---|---|
| D: Cohort baseline mapped | Not done | Partial (1-2 competitors) | Full (3-5 competitors, section by section) |
| E: Unique claims identified | 0 unique claims | 1-4 unique claims | 5+ unique claims |
| L: Unanswered questions covered | None addressed | 1-2 addressed | 3+ addressed |
| T: Unique data points | 0 to 1 | 2 to 6 | 7+ (15+ for highly original) |
| A: AEO chunk compliance | No sections standalone | Some sections standalone | All H2s independently extractable |
Maximum score: 10. Target 7+ before publishing. A page scoring below 5 on DELTA is almost certain to land in the mostly-shared grade band and below the top-3 threshold for AI citation eligibility.
E-E-A-T (Experience, Expertise, Authoritativeness, Trustworthiness) evaluates the credibility of the source producing the content. A board-certified cardiologist writing about heart disease has high E-E-A-T. Information gain quantifies the marginal contribution of the page’s content relative to its cohort.
The same cardiologist can write a high-E-E-A-T article that scores 25/100 on information gain if it restates what every other cardiology article already says. I see teams optimize E-E-A-T meticulously, author bios, trust signals, credential pages, and then publish content that says exactly what the other five credible pages already said. The March 2026 core update was Google’s answer to that pattern.
E-E-A-T is a prerequisite for being trusted enough to rank. Information gain determines whether your trusted content is worth citing. Google’s March 2026 core update did not reduce the importance of E-E-A-T. Based on Similarweb’s tracking of pages that lost visibility in that update, the pattern was consistent: technically credible pages with high E-E-A-T signals that were substantively redundant lost ground. The update penalized competent restatement rather than incompetent content.
For AI systems, the relationship is similar but weighted differently. E-E-A-T signals (named authors, institutional affiliation, consistent brand entity presence) increase citation trustworthiness. Information gain determines whether the content is worth citing at all.
A highly credible page that says nothing new will not be cited, because there is nothing to extract that the synthesis doesn’t already have. A moderately credible page with a unique data point will be cited for that specific claim, regardless of how the rest of the page scores.
The practical implication: optimize E-E-A-T as your credibility infrastructure (author bios, source attribution, brand entity consistency, structured data). Optimize information gain as your content moat. Neither alone is sufficient. In practice, most teams optimize E-E-A-T diligently and then publish content that says exactly what the other five credible pages already said. The March 2026 update was Google telling them that the combination is no longer a strategy.
Measuring your position in the AI search landscape requires tracking citation frequency, share of model, and citation gaps as separate metrics. The gap between your information gain score and your actual citation rate is usually where the story lives.
I track this using Similarweb’s AI Search Intelligence suite, which monitors citation frequency across AI engines for my tracked prompts. Not because I work here (well, I do), but because it is the only setup I have found that connects citation frequency to actual query clusters and website traffic rather than giving you a vanity score with no actionable signal.
The metrics that matter here are not generic. Citation frequency tells you how often your content is pulled as a source across the engines you track. In practice, this is where you see the information gain problem made visible: a page that ranks top 3 on Google can show near-zero citation frequency in AI if it is sitting in the mostly-shared band. The AI Traffic tracker can show you which cited pages are actually earning clicks to your website, which lets you work backwards from what is performing to understand why.
Share of model (your citation share relative to competitors across your tracked prompt cluster) is the zero-sum metric. Every citation your competitor earns is one your page did not. When I run this for a client in a competitive category, the share of model number is usually the one that gets people’s attention in a way that an information gain score by itself does not.
Perform a full citation gap analysis, powered by Similarweb’s AI citation analysis tool, identifies which topic questions in your category are currently being answered by other sources. This is where you find the unclaimed questions from Step L of the DELTA audit confirmed by real citation data. It shows not just that a gap exists, but which competitors are filling it and what type of content they are using.
A page that scores 45/100 on information gain with an 8% citation share across tracked prompts has a quantified problem and a quantified target. Track your AI brand visibility alongside your information gain audits. The two signals together tell you what fixing alone cannot.
The information gain study puts a number on what most SEOs already suspect: a substantial portion of the SERP is occupied by pages that say nothing more than the other pages do. The median top-3 page carries 4 unique data points. The majority of top-ranked B2B SaaS and ecommerce content scores in the mostly-shared range.
This was survivable when Google was the only channel. It is not survivable as AI systems become the first point of contact for the queries that matter. Those systems select sources based on what they can extract. Redundant content has nothing to extract.
The floor here is real and measurable. A page scoring low on information gain will not be cited by AI systems, regardless of how well it ranks on Google. A page that scores highly is playing in territory most competitors have left empty.
The gap is not technical. It is not a schema problem or a crawlability problem. It is a content investment problem: whether your team is producing the proprietary data, original frameworks, and specific answers that AI systems are actually built to prefer.
To measure where your brand currently stands in the AI citation landscape, including which prompts are sending traffic to competitors, which sources AI engines are pulling from in your category, and where the unanswered questions live: use Similarweb AI Search Intelligence.
What is information gain in SEO?
Information gain in SEO measures how much new, semantically distinct content a page adds beyond what the other pages ranking for the same query already say. Derived from a Google patent granted in 2022 (US11354342B2), it scores the additional content a document contributes relative to a comparison set of previously ranked pages. A score near zero means the page is a semantic duplicate of its cohort. A score near 100 indicates the page contains content not found elsewhere in the SERP. It is not about length or keyword density. It is specifically about what your page says that the competition does not.
Does information gain affect Google rankings?
Yes, but with an important nuance. The On-Page.ai study of 150 top-3 pages found that ‘mostly-shared’ pages appear in positions 4 to 10 at nearly twice the rate they appear in positions 1 to 3: 37 to 40% versus 24%. Getting into the top 3 correlates with higher group information gain. Within the top 3, however, both positions share a median of 52/100, so information gain does not separate positions once you are already in the top tier.
How does information gain relate to AI citations?
AI engines are structurally information-gain systems. When a generative AI builds an answer, it retrieves passages that contribute novel content to the synthesis. A page restating what five other pages already say provides no marginal contribution and will not be cited.
How do I calculate my information gain score?
Manual: extract claims section by section from the top-3 competitors, compare against your page, count unique elements, and unanswered questions covered. Cosine similarity approximation: embed your page and each competitor using a sentence-transformer model. Your IGS is approximately 1 minus the maximum cosine similarity between your page and any single competitor (target above 0.5).
What is a good information gain score?
Below 40 is mostly shared content (24% of top-3 pages fall here), 40 to 69 is moderately original content (55%), and 70 to 100 is highly original content (21%). For AI citation eligibility, target the moderately-original range as a minimum and highly-original as the goal for anchor content. Pages with 15 or more unique data points averaged 60+. Run a DELTA audit before publishing, aiming for a score of at least 7 out of 10.
Is information gain the same as E-E-A-T?
No. E-E-A-T measures whether your content is trustworthy enough to rank: experience, expertise, authority, and trust signals. Information gain measures whether your content is original enough to be cited: the marginal contribution of your page relative to what the SERP cohort already contains. A well-written, credible page that repeats the consensus passes E-E-A-T and fails information gain. Both are necessary; they respond to different interventions.
What is the downstream impact of earning an AI citation?
According to Similarweb’s Downstream Impact of AI Visibility study, users who received a brand recommendation from ChatGPT were 2.5 times more likely to visit that brand’s website within the next 7 days. Those visitors viewed twice as many pages and stayed twice as long on site. The traffic arrives via branded search or direct, not as an AI referral, making it invisible to standard attribution and systematically underreported by last-click models.
Is the skyscraper technique still relevant now that information gain matters?
For link building, yes. For AI citation strategy, no. The skyscraper technique still attracts backlinks, but length and comprehensiveness without original content produce zero information gain. For AI citation eligibility, a 600-word page with one proprietary data point outperforms a 4,000-word guide that rehashes published statistics. The two strategies serve different goals; only information gain serves AI visibility.
Can a page rank on Google but be invisible in AI Overviews?
Yes. Google’s organic ranking algorithm weights domain authority, links, and engagement signals. AI citation engines weigh content novelty and the relevance of extractable passages. A page can meet the organic ranking threshold without meeting the information gain threshold required for AI citation. Tracking both metrics separately is now a practical necessity for any SEO team with AI visibility as a business goal.
How do you improve your information-gain score?
Use the DELTA audit framework: map the existing cohort (top 3–5 ranking competitors), extract what’s unique to your page, list unanswered questions in the SERP, tally unique data points (target 15+ for highly original status), and check that every H2 section can be extracted as a standalone citable chunk. The most reliable lever is original data: proprietary surveys, platform-specific measurements, or case-level analysis that no competitor has access to. Length and rewrites don’t move the score. Original data points do.
Director of SEO & AI Search at Similarweb
Limor brings 20 years of expertise in SEO and AI Search. She thrives on solving complex problems, creating scalable strategies, and building amazing dashboards.
Give it a try or talk to our insights team — don’t worry, it’s free!