Semalt series

From Keyword List to Topic Map: Structuring Research in Semalt

A list of 800 keywords is not a strategy. How to cluster by intent, turn clusters into pages, and know when to stop researching and start writing.

Updated: 2026-08-07 9 min read 2.033 words
A topic map and organised notes on a planning wall

Key takeaways

  • A long keyword list is raw material, not a plan. The deliverable is a map: which page targets which cluster, and why.
  • Cluster by intent and by what already ranks, not by word similarity. Two phrases with the same words can need two different pages.
  • One page per intent. Splitting an intent across three pages guarantees all three underperform.
  • Stop researching when the next hour of work stops changing the plan. That point arrives far earlier than most people think.

Ask for keyword research and you will usually receive a spreadsheet: eight hundred phrases, volume, difficulty, sorted by volume descending. It looks like work and it is nearly useless, because it answers a question nobody asked. Nobody needs to know which phrases exist. They need to know which pages to build, in what order, and what each one has to say.

This article covers turning research into that structure using the Semalt platform: how to gather candidates without drowning, how to cluster them by intent rather than by wording, and how to convert clusters into a page map that a writer can actually work from.

Gathering: three sources, not one

Tool exports are the least interesting source. They are built from what people already search in volumes large enough to record, which means they systematically miss the newest and most specific phrasing — often exactly where the winnable opportunities are.

  1. The tool export

    Start here for breadth and volume estimates. Include the competitor gap report, which surfaces terms your rivals rank for and you do not — evidence-based rather than imagined.

  2. Your own search data

    Search Console shows queries you already receive impressions for, including many nobody would have guessed. Terms where you rank on positions eight to twenty are the fastest opportunities on the entire list, because you are already close.

  3. The words customers actually use

    Sales emails, support tickets, call notes, the site's internal search log. This is where the phrasing that converts lives, and it rarely matches the industry vocabulary marketing departments prefer.

3 sourcesTools, your own search data, and what customers actually say
8–20The position band where the cheapest wins are hiding
1 pagePer intent — never one intent split across several pages

Working parameters from the method described below.

The single highest-return filter. Before doing anything else, pull the queries where you already rank between positions eight and twenty with meaningful impressions. Those pages exist, are already relevant, and usually need one round of improvement rather than a new page. Most research projects discover this list at the end, after committing the budget elsewhere.

Clustering by intent, not by words

The common mistake is grouping phrases that look alike. "Price of X" and "X cost" belong together; "how X works" and "buy X" do not, even though they share a word. What determines whether two phrases belong on the same page is whether the same content would satisfy both searchers.

The reliable test is empirical rather than linguistic: look at what currently ranks for each phrase. If the same pages appear on both results pages, one page can serve both. If the results are entirely different, they are different intents regardless of how similar the wording is — and forcing them onto one page produces something that serves neither.

IntentWhat the searcher wantsPage type
InformationalTo understand somethingGuide or explainer
ComparativeTo choose between optionsComparison, alternatives, "best of"
CommercialTo find a provider or productService or product page
NavigationalTo reach a specific placeBrand, login, contact
LocalTo find something nearbyLocation page

Assign each cluster exactly one of these. Clusters that resist classification are usually two clusters that have not been separated yet, and the ambiguity will resurface later as a page nobody knows how to write.

From clusters to a page map

The output of research should be a table that a writer can pick up without asking questions. Six columns are enough.

ColumnWhat goes in it
ClusterThe group of phrases this page serves
Primary termThe one that defines the title and heading
IntentOne of the five above
Target URLExisting page to improve, or new page to create
What must be answeredThe three or four questions the page has to cover
PriorityBased on commercial value and realistic difficulty

The fifth column is what separates a usable brief from a keyword handed to a writer. It comes from reading the results page: the questions the ranking pages all answer are the questions your page must answer to be considered comparable.

We cover this in detail in Is It Worth Competing? Reading the Results Page Before Investing in a Topic.

Watch for cannibalisation before you create anything. If two clusters map to two new pages that would target near-identical intent, merge them now. Two pages competing for the same intent split their signals and both underperform, and the problem is much cheaper to prevent at the mapping stage than to unpick a year later.

Turning a cluster into a brief a writer can use

The gap between a page map and published content is the brief, and it is where most content programmes quietly stall. A writer handed a keyword and a word count will produce something generic, because a keyword contains no information about what the page has to accomplish.

A brief that works fits on half a page and contains five things. The intent, stated plainly — someone comparing options, someone ready to buy, someone trying to understand a term. The questions the page must answer, taken from what the currently ranking pages all cover, because a page that omits what everyone else includes reads as incomplete to both readers and search engines. The angle that makes this version worth publishing: original data, a practitioner's experience, a specific market, a level of honesty competitors avoid. The internal links in and out, decided in advance rather than added later if someone remembers. And the one action the reader should be able to take by the end.

See also: Your Own Results Page.

What deliberately does not belong in the brief: a target keyword density, a required number of headings, a list of phrases to include a certain number of times. These instructions actively degrade writing and address a version of search that has not existed for years. If the brief names the questions and the angle, the vocabulary follows naturally, because someone answering a question properly uses the words that question implies.

One test before handing it over: could a competent writer who knows the subject but not SEO produce a good page from this brief alone? If not, the missing piece is usually the angle — you have told them what to cover but not why anyone should read your version rather than the four that already exist.

Prioritising without guessing

Volume is the worst available sorting criterion and the most commonly used one. A better ordering weighs three things together: how close the intent is to a transaction, how realistic the position is given what currently ranks, and whether you already have a page that is nearly there.

Do first

  • Existing pages ranking 8–20 with commercial intent
  • Clusters where current results are thin or outdated
  • Terms your competitors rank for and you have no page at all
  • Anything with clear buying intent and modest competition

Do later, or never

  • High-volume informational terms owned by major publishers
  • Clusters where every result is a different content type than you can make
  • Phrases with volume but no plausible path to revenue
  • Anything you cannot write about better than what already ranks
The most valuable output of keyword research is the list of things you decided not to pursue. Everything else is a to-do list that will never be finished.Why prioritisation is the deliverable, not the phrase list

Working in Portuguese: what changes

Two practical adjustments apply to Portuguese-language research and are worth stating because most published methodology assumes English.

The first is data density. Volumes are lower, so estimates for long-tail terms are noisier and should be treated as orders of magnitude rather than numbers. Build the plan on the terms with confirmed demand and treat the tail as upside rather than as forecast.

The second is variant handling. Accented and unaccented spellings, European and Brazilian phrasing, and formal versus everyday vocabulary all fragment what is really one intent. Group them deliberately at the clustering stage: they belong on one page, and treating them as separate opportunities is how sites end up with three thin pages competing for the same searcher.

There is also an upside worth naming. Because fewer people do this properly in Portuguese, a genuinely well-structured cluster map produces results here that the same effort would not buy in a saturated English-language niche. The competition is often not better research — it is no research at all.

Knowing when to stop

Research expands to fill whatever time it is given, and the last hours produce almost nothing. The stopping rule is simple: when another hour of investigation stops changing the priority order, you are finished. Additional phrases at that point are decoration.

For most small and mid-sized sites, that point arrives after roughly a day of work, producing perhaps fifteen to thirty clusters. That is enough for a year of content at a realistic publishing rate, and revisiting it quarterly is more useful than extending it now, because the business will change and the results pages will change with it.

There is a full walkthrough in Generating Pages at Scale Without Producing Junk.

A final note on tooling. The reason this method is practical rather than theoretical is that the clustering test — comparing which pages rank for each phrase — used to require opening dozens of tabs by hand. With tracked keywords and competitor overlap in the same project, the comparison is a single view, which is what makes it realistic to do properly for thirty clusters rather than skipping it and grouping by wording like everyone else.

What good research looks like when it is done

It fits on one screen. Fifteen to thirty rows, each with a page, an intent, the questions to answer and a priority. Someone who was not in the process can read it and start writing. Nothing in it is there because it had volume; everything is there because a decision was made about it.

What it is not is eight hundred rows sorted by search volume, which is the artefact of a process that ran out of time before the thinking started.

If you have a keyword list but no map, the conversion takes an afternoon. Open the dashboard, pull the queries where you already rank eight to twenty, and start the map from those — they are the rows most likely to move first.

Frequently asked questions

How many keywords should we target per page?

One intent per page, which usually means one cluster of anywhere from three to forty related phrasings. The number of phrases does not matter; what matters is that the same content genuinely satisfies all of them. If two phrases in your cluster return completely different results pages, they are different intents and need different pages.

Should we target high-volume or low-volume keywords?

Neither by itself. Sort by commercial value and realistic difficulty rather than volume. A term with 40 monthly searches and clear buying intent, where current results are thin, is a better first project than a term with 4,000 searches owned by national publishers. Volume matters only after you have established that the position is achievable and the traffic is worth having.

How do we know if two keywords need separate pages?

Compare the results pages. If largely the same pages rank for both phrases, one page can serve both. If the results are substantially different, they are different intents and need separate pages, regardless of how similar the words look. This test is faster and more reliable than any judgement about wording.

How often should keyword research be redone?

Review quarterly rather than rebuilding. Businesses add services, drop products and shift markets, and results pages change format around your terms. A quarterly pass that adds new clusters and retires irrelevant ones keeps the map accurate for a fraction of the effort of starting again, which almost nobody actually does anyway.

Try it

Open your Semalt dashboard

Audits, rank tracking, competitor data and reporting in one place. Sign in and you will be looking at real numbers for your own domain within minutes.

Sign in to Semalt

Or browse the service overview at semalt.com.