← Back to blog

Strategy · September 22, 2026 · 8 min read

Keyword Clustering by SERP Overlap: One Page or Five?

Shared URLs in the top 10 decide whether five keywords need five pages or one - here are the thresholds worth acting on and the four cases where the count misleads.

By FluxWriter Team

Keyword Clustering by SERP Overlap: One Page or Five?

Keyword clustering by SERP overlap settles the one-page-or-five question with evidence, and it is the only grouping method that does. Most lists get split on how the terms read — same root, same page — which is backwards, because the words two queries share predict almost nothing about whether one document can win both. This covers what overlap measures, the thresholds worth acting on, and the four cases where the number lies.


Why Grouping by Words Fails Both Ways

The default method is a read-through. Same root, same page — and that rule is wrong in both directions at once, which is why it survives. It merges terms that wanted separate pages and splits terms that would have won as one. Both errors look reasonable on the spreadsheet.

"Running shoes for flat feet" and "flat feet running shoes" are one page, and the word test gets that right. "SEO audit" and "SEO audit checklist" are not, and it gets that wrong. Two of the three words match. The results barely do — the first query returns agency service pages and long explainers, the second returns downloadable lists and templates.

Difficulty scores do not rescue it. A score says how hard a term is to rank for, not whether two terms belong on one URL — a pair with identical scores can need one page or four.

The only vote that counts is which pages already rank for each term. Everything else is a guess dressed up as a metric, and at 30 or 40 briefs a quarter the guessing gets expensive — every wrong split buys you a page that never gathers enough signal to win.

What SERP Overlap Actually Measures

Overlap is a count. Run both queries, take the top 10 organic results for each, and count how many exact URLs appear in both lists.

That number is a proxy for something you cannot observe directly: whether Google treats the two queries as asking for the same thing. High overlap means the ranking system has already decided one document can satisfy both. Low overlap means it has decided otherwise, and on-page work rarely reverses that inside a quarter.

Compare URLs, not domains. Domain-level matching inflates every score, because a handful of large sites rank somewhere in the top 10 for almost everything in their category — Wikipedia, Reddit and the two biggest publishers in your niche appear on both lists regardless of intent.

Position weighting matters too. Three shared URLs sitting at 1, 3 and 5 is a much stronger signal than three shared URLs at 8, 9 and 10, where the results are loose matches anyway.

Hold the conditions steady while you sample. Run both queries within the same day and from the same country, and keep the device type constant — mobile and desktop results diverge enough to move a borderline pair across your threshold.

The Thresholds Worth Acting On

Thresholds here are conventions rather than laws, but the set below holds up well enough as a default. Score each pair on shared URLs in the top 10:

Shared URLs (top 10) What it means Action
0–2 Different intents Separate pages
3 Ambiguous Recheck positions 1–5 only
4–5 Same intent, soft signal One page, secondary term in an H2
6–8 Same page in Google's view One page, no exceptions
9–10 Effectively the same query One page, drop the duplicate from the plan

The row that decides most real cases is the middle one. At exactly 3 shared URLs, re-run the comparison across the top 5 positions only — if two of those five are shared, treat the pair as a cluster, and if none are, split it.

Pick one threshold and apply it to the whole list. Sliding it per keyword turns the exercise back into a series of opinions. Write the number at the top of the sheet before you open the first result page.

Where Overlap Lies to You

Four situations produce a number you should not act on without a second look.

Giant-publisher results. In categories dominated by 3 or 4 huge sites, overlap runs high because those sites rank for everything nearby. Discount any domain appearing in more than half your comparisons, then recount.

Volume asymmetry. A pair can overlap at 7 and still deserve different handling when one term draws 8,000 searches a month and the other draws 90. Keep them on one page, but past roughly a 10x gap the smaller term is a heading inside the bigger one's article. It does not get the title or the opening, and internal links still point at the bigger term.

Format splits. Two queries can share 5 URLs where every shared result is a generic overview, and the unshared results split cleanly into video on one side and long text on the other. That is a format signal hiding inside an intent signal. Look at what the unshared results are before you cluster.

Unstable results. News-adjacent, seasonal and product-release queries re-rank weekly, so a score taken today may not survive the month. Sample twice, 2 to 3 weeks apart, and use the lower number.

How Many Keywords One Page Can Carry

In most Search Console accounts a single well-built page earns clicks for 40 to 200 distinct queries over a year, which makes the ceiling feel unlimited. It is not. The terms you deliberately target behave nothing like the ones you collect by accident.

Target one primary term and 3 to 8 secondaries that cleared your threshold. Past that the page starts hedging — it covers each variant shallowly to cover them all, and it loses to a competitor page that picked one and committed.

The tail arrives on its own. Open the Search Console Performance report, filter to a single strong page, and switch to the Queries tab — a page you built around 4 deliberate keywords is usually collecting impressions for many times that number. Those extra queries came from covering the primary topic properly, not from bolting on sections.

More keywords in a cluster does not mean a longer page. A page serving 6 tight variants runs about the length of a page serving 1. What changes is the headings. Each variant has to be visibly answered where a skimmer will look, and that is a phrasing job rather than 400 extra words.

Running the Check Without Buying a Tool

Pairwise comparison explodes fast. Forty keywords make 780 possible pairs, and nobody is opening 780 result pages by hand.

The count is misleading, and the misreading is what sends most people straight to a paid tool. You need one capture per keyword, not one per pair. Record the top 10 organic URLs for each of the 40 terms once, and every comparison after that is spreadsheet arithmetic — at 2 to 3 minutes a capture, the full set costs about 2 hours.

Ordering makes the sorting quick. Sort by search volume, seed on the highest-volume term, and assign everything that clears your threshold to cluster one. Re-seed on the highest-volume keyword still unassigned and repeat until the pool empties. A 200-keyword list usually resolves into 15 to 40 clusters, so expect 15 to 40 passes — seconds each, on captures you already took.

Dedicated tools do the same job in one run. Keyword Insights and Keyword Cupid both group on live results rather than on wording, and both typically meter by the keyword processed rather than by the seat, which suits a one-off list. Semrush and Ahrefs fold grouping into their existing subscriptions, and Ahrefs comes at it from the other end by showing which terms a single ranking page already covers. Check what any of them compares before you trust it. Some group on live results. Some group on wording.

FAQ

How often should I re-run clustering on a list I already built?

Once a year for stable topics, and every 6 months for anything tied to product releases or changing regulation. Results drift as new formats win, so a pair you sent to two pages two years ago can read as one intent now. Re-check before a large refresh rather than on a fixed calendar.

Two keywords overlap at 6 but feel like different intents to me. Which wins?

The overlap wins. Your read of intent is a hypothesis, and the shared URLs are the outcome of enormous amounts of live testing. The one exception is a results page stacked with giant generalist domains — strip those out, recount, and if the number drops below 3, your instinct was right.

Can I cluster keywords with an AI model instead of checking results?

Not reliably. A language model groups by meaning, which is the word-similarity mistake in a more convincing wrapper. It earns its place on a first pass over 500 raw keywords, cutting obvious junk before the real check — but the grouping decision itself needs live results behind it.

The Practical Takeaway

Pick one threshold and hold it. Record the top 10 organic results for your seed keyword, compare every other term on the list against that seed, and count shared URLs — 4 or more joins the cluster, 0 to 2 earns its own page, and exactly 3 gets a second look at the top 5 positions only. Discount any domain that shows up in more than half your comparisons, and where a clustered pair's volumes differ by 10x or more, let the bigger term own the title while the smaller one lives in a heading. Start with the 20 keywords you were about to brief this month.

If you are producing a page for every cluster the check produces, tools like FluxWriter can help keep primary and secondary terms placed consistently across a run of posts — but deciding where the cluster boundary sits, and re-checking it when the results move, stays your call.



← All posts