Find Topics to Write: Research, PAA Mining, Clustering, and Gap Analysis
The correct sequence for topic discovery is: understand what people search, mine what they ask, cluster related questions, map against what already exists. This produces a prioritized list with zero guesswork.
Step 1: Seed Keyword Research
Start with 3-5 seed keywords that represent the core of the business or niche:
Sources for seed keywords:
- Primary services or products offered
- Geographic + service combinations ("roof repair [city]")
- Problem-based searches ("how to fix [problem]")
- Comparative searches ("[service] vs [alternative]")
For each seed keyword, expand to:
- The primary keyword (high volume, competitive)
- 3-5 long-tail variants (more specific, lower competition)
- 2-3 question variants (PAA candidates)
Step 2: PAA Mining
Run PAA Researcher in Deep Mode for the top 2-3 seed keywords. (See paa-researcher-modes.mdx for full protocol)
From PAA mining, extract:
- All questions at awareness stage
- All questions at consideration stage (these have highest content ROI)
- All questions at decision stage
- Questions that competitors are NOT currently answering well
Quick identification of unanswered PAA:
- Search the PAA question in Google
- If position 1-3 results only partially answer the question, or the content is 3+ years old → gap confirmed
- If position 1-3 results are from low-authority sites → opportunity confirmed
Step 3: Topic Clustering
Group related topics and questions into content clusters. Each cluster gets one primary article (the pillar) and 2-4 supporting pieces.
Cluster structure:
CLUSTER: [Topic name]
Primary keyword: [main keyword — pillar content]
Pillar article: [title + target word count]
Supporting pieces:
1. [Supporting keyword] → [content type: FAQ / comparison / how-to]
2. [Supporting keyword] → [content type]
3. [Supporting keyword] → [content type]
Internal link plan:
- Pillar links to all supporting pieces
- Supporting pieces all link back to pillar
- Supporting pieces link to each other where contextually relevant
Step 4: Gap Analysis Against Existing Content
Before adding any topic to the production queue, check:
- Does any existing page already target this keyword?
- Is the existing page ranking (positions 1-20)?
- Does the existing page fully cover the topic?
Decision matrix:
| Existing Content? | Ranking? | Action | |------------------|---------|--------| | No | N/A | Create new | | Yes | No | Update existing (refiner workflow) or rewrite | | Yes | Positions 4-20 | Update and expand existing page | | Yes | Positions 1-3 | Leave it — don't disturb a winner |
Step 5: Priority Output
Score each topic and output a production list:
TOPIC PRODUCTION LIST:
Tier 1 — Produce immediately:
1. [Title] | [Keyword] | [Word count] | [Content type] | Priority: [score]
2. ...
Tier 2 — Produce this quarter:
...
Tier 3 — Monitor / Low priority:
...
Tier criteria:
- Tier 1: High commercial intent + easy to medium difficulty + no existing content
- Tier 2: Informational with traffic potential + medium difficulty + cluster value
- Tier 3: Long-tail informational + low volume + nice-to-have coverage
Topic Discovery Sources
Beyond search data, mine these for high-signal topic ideas:
| Source | What to Look For | |--------|----------------| | Reddit (subreddits in niche) | Recurring questions, complaints, "what I wish I knew" | | YouTube comments on competitor videos | Unanswered questions, complaints about existing content | | Google autocomplete | Long-tail variants of seed keywords | | Answer the Public | Visual map of question + preposition variants | | Client sales calls / CRM | Actual questions customers ask before buying | | Review mining (GMB, Yelp, G2) | Problems solved, language customers use |
Customer language beats algorithmic data — when both sources agree on a topic, it's highest priority.
#content-sop #topic-research #keyword-research #content-planning