Key Takeaways
- Ranking keywords and revenue keywords are different assets; selection should account for intent match, claim substantiation, and downstream conversion metrics, not just volume and difficulty.
- Label every candidate keyword by informational, navigational, or transactional intent before scoring, because intent distribution varies by topic and misclassified clusters miss traffic forecasts 4, 5.
- Prioritize specificity over volume: longer queries and noun-modifier constructions signal narrower, serviceable intent, while head terms rarely convert for challenger clients 7.
- Read the top ten results as graded evidence of page type, entity fit, and SERP features to decide whether a client can compete or should route effort elsewhere 8.
- Apply substantiation, HIPAA, accessibility, and disclosure checks before production, since unsupported claims, PHI misuse, or inaccessible pages cap what a cluster can legally or practically earn 1, 2, 3, 9, 10.
- Every cluster entering production needs a defined page type, conversion action, and client-tracked KPI so reporting reflects business outcomes rather than position changes.
- At portfolio scale, automate discovery, clustering, SERP pulls, and scoring drafts while versioning taxonomies per vertical to shift throughput without adding analyst headcount 6.
- Keep human judgment on four gates: claim documentation, cluster-to-service fit, SERP competitiveness, and marketing data authorization, regardless of how much upstream work is automated 2, 6, 8.
What separates a ranking keyword from a revenue keyword
A ranking keyword moves a URL up a results page. A revenue keyword moves a qualified buyer into a client's pipeline. Agency SEO leads managing a portfolio already know these are not the same asset, yet keyword lists are still filtered primarily by volume and difficulty scores that describe the SERP, not the client's economics.
The distinction sits in three places. First, intent: the query has to match a page the client can build and a decision the client can serve. Query-log research classifies web search into informational, navigational, and transactional categories, and intent distribution shifts by topic, which means a legal client and a home services client cannot share the same funnel assumptions 4, 5. Second, substantiation: in health, dental, behavioral health, and legal work, the keyword is only viable if the client can back the implied claim with evidence the FTC would accept 1, 9. Third, measurement: relevance at the top of the SERP is a graded signal, not a binary rank, and it still has to connect to calls, bookings, and cost per lead before it counts as a result 8.
What follows is a selection workflow built for delivery leads who are scored on client outcomes, not keyword counts.
Classify intent before you touch a volume column
The three-way split and why topic changes the math
Every keyword shortlist should be intent-labeled before a volume or difficulty column enters the conversation. The reason is empirical, not stylistic. Jansen's query-log analysis of live search behavior found that more than 80% of web queries were informational, with roughly 10% each navigational and transactional 4. That distribution is the baseline any agency SEO lead should hold in mind when a client hands over a wish list dominated by bottom-funnel service terms.
The practical consequence: most of the demand a client sees in a keyword tool is people learning, comparing, and diagnosing, not people ready to book. A page built for a transactional query will underperform against an informational SERP no matter how well it converts, because the searchers arriving on it were never in a buying posture to begin with.
The 80/10/10 figure is a starting point, not a rule. It comes from a general web query log, and Jansen's own follow-up work classifying more than 20,000 queries by topic and intent found that intent distribution shifts by topic 5. A dental client's branded emergency queries skew far more transactional than a general baseline would predict. A behavioral-health client's symptom queries skew even more informational. A home-services client sees seasonal transactional spikes that swamp the informational base for weeks at a time.
Delivery leads should treat the three-way split as a diagnostic frame: label every candidate keyword by dominant intent, then compare the client's cluster mix to the vertical's expected mix. A shortlist that is 95% transactional in a topic where the real search population is 60% informational is a shortlist that will miss its traffic forecast.
Building vertical taxonomies for legal, dental, home services, and behavioral health
Because intent distributions shift by topic 5, a single agency-wide intent taxonomy will misclassify half the accounts it touches. Each vertical needs its own labeling scheme, ideally built once and reused across every client in that category.
- For a legal client, informational queries dominate the top of the funnel and carry unusually high research depth: statutes, procedure explanations, jurisdictional differences, cost expectations. Transactional queries cluster tightly around "lawyer near me," practice-area plus city, and consultation modifiers. Navigational queries often carry firm names or attorney names and belong on branded pages, not competitive landing pages.
- Dental taxonomies split cleanly between routine care, cosmetic, and emergency. Routine queries lean informational and price-aware. Cosmetic queries mix informational research with high-intent comparison terms. Emergency queries collapse the funnel: the searcher is transactional within minutes, and the taxonomy should route every emergency modifier to a same-day-contact page.
- Home services runs on job-type plus geography, with strong seasonality. The taxonomy should tag every cluster with a service category, a geographic qualifier, and a season, so the delivery team can time production against demand curves rather than treating the calendar as flat.
- Behavioral health carries the heaviest informational skew and the tightest regulatory constraints. Symptom, condition, and "is this normal" queries dominate. Transactional queries are rarer and higher-stakes. The taxonomy should separate condition-education pages from treatment-service pages so the two can be reviewed under different substantiation standards.
Reusable taxonomies are what let a delivery team onboard a new client in a known vertical in hours instead of weeks. Build them once per category and version them like code.
Using clustering to surface intent groups a human analyst would miss
Hand-built taxonomies cover the intent patterns an analyst already expects. They miss the ones that emerge only when a large query set is examined without preconceptions. K-means and related clustering methods have been evaluated as ways to group queries by intent characteristics rather than by pre-selected categories 6, and that property is what makes them useful for portfolio-scale keyword work.
Applied to a client's exported search-console data, a competitor's ranking set, or a scraped SERP corpus, clustering surfaces intent groupings that don't map cleanly to the vertical taxonomy: hybrid research-and-price queries in dental cosmetic, jurisdictional-comparison queries in legal, warranty-plus-brand queries in home services. These are the clusters that a manual review skips because they don't fit an existing bucket.
Two operational cautions apply. First, unsupervised clusters are not business categories. An analyst still has to interpret each cluster, decide whether it maps to a page type the client can produce, and validate that the implied intent matches what actually ranks. Second, clustering scales the discovery step, not the judgment step. A delivery lead who runs clustering across a 40-client portfolio will surface hundreds of candidate groups; the value comes from the review workflow that filters them, not the algorithm itself.
Use clustering to widen the intent map before scoring. Keep human interpretation as the gate before any cluster reaches a production queue.
Visualize the three-way intent taxonomy and how vertical topic shifts the mix, directly supporting the section's argument that intent labeling must precede volume scoring
Prioritize specificity over volume
Once intent is labeled, the next filter is specificity. A high-volume head term almost always describes a broad information-gathering audience; a longer, qualified query almost always describes a narrower need the client can actually serve. Delivery leads who default to volume as the primary scoring input consistently overweight the wrong end of the tail.
The empirical basis is straightforward. A query-specificity study of 5,115 unique queries found that query length and parts of speech reliably distinguish narrow queries from general ones, with longer queries and noun-modifier constructions signaling tighter intent 7. In practical terms, "invisalign" is a research bucket, "invisalign cost with insurance ann arbor" is a shortlist candidate, and the second query tells the delivery team what the page has to answer, who it has to answer it for, and where the searcher expects service.
Two operational corrections follow. Demote head terms in the scoring rubric unless the client owns a brand or category position that already commands them; head-term rankings for challenger clients rarely convert at a rate that justifies the production cost. And treat local qualifiers, procedure modifiers, insurance and pricing modifiers, and comparison phrasings as first-class targets, not scraps left over after the head-term pass.
Specificity is a filter, not a guarantee. A precise query can still carry weak demand, poor service fit, or a SERP the client cannot compete on 7. The scoring rubric in a later section handles those failure modes; the point here is that specificity earns a cluster its way into scoring, and raw volume does not.
Test AI-driven keyword strategies on live campaigns
See measurable keyword impact on client projects before making a commitment.
Validate against the live SERP with graded relevance thinking
Intent labels and specificity filters produce a candidate list. The live SERP decides whether that list is real. Before any cluster reaches production, a delivery analyst should pull the top ten results for each seed query and read them the way TREC evaluators read retrieval outputs: not as a binary win/loss against position one, but as graded evidence of what the engine has already decided the query means.
TREC's Deep Learning Track formalizes this with four-point human relevance judgments—Irrelevant, Relevant, Highly Relevant, Perfectly Relevant—and uses NDCG@10 as a primary quality metric because a single ranking slot is a poor summary of how well a result set serves the query 8. The transferable idea for client SEO is simple: the current top ten is a labeled dataset the engine built for that query, and its composition tells the analyst which page types, formats, and depth levels are eligible to compete.
Three checks belong in every SERP validation pass.
- Page-type consistency: if nine of ten results are educational guides and the client can only offer a service page, the cluster is misclassified and belongs in a content brief, not a conversion queue.
- Entity fit: if the ranking domains are national publishers, insurers, or associations, a local service business will not displace them on the head term and should route effort to qualified variants instead.
- Feature footprint: video carousels, local packs, and forum threads each cap the addressable click share and change what "ranking" is worth.
Graded relevance thinking also protects reporting. A cluster that moves from position 14 to position 6 without producing calls has improved on the SERP's terms and failed on the client's terms. The scoring rubric in the next section ties every cluster to a conversion action so those two signals stop getting confused.
The substantiation and privacy filter that caps opportunity
FTC substantiation for health, legal, dental, and behavioral-health claims
A cluster that survives intent, specificity, and SERP validation can still fail one more test: whether the client can legally make the claim the target page will imply. The FTC requires that advertising be truthful, non-deceptive, and backed by evidence the advertiser holds before the ad runs, and health-related claims require adequate scientific substantiation for both express and implied messages a reasonable consumer would take away 1, 9.
For delivery leads, that translates into a claim-check step that happens before a keyword enters the production queue. "Fastest recovery," "guaranteed outcome," "most effective treatment," "top-rated attorney," and "proven results" are not keyword problems in isolation; they become keyword problems when the page built to rank for them implies a claim the client cannot document. Substantiation review should confirm that outcome language, comparative superiority, and any success-rate framing are tied to evidence on file, with the type and depth of evidence matching what the claim would suggest to a reasonable reader 1.
The operational rule is direct. If the strongest ranking angle for a cluster requires a claim the client cannot substantiate, the cluster ships with a reframed angle or does not ship. Opportunity does not override the substantiation requirement.
HIPAA constraints on keyword and call-intelligence workflows
Healthcare, dental, and behavioral-health accounts add a second filter on top of substantiation. HHS defines marketing under HIPAA as a communication that encourages the recipient to purchase or use a product or service, and with limited exceptions, protected health information cannot be used or disclosed for marketing without written authorization 2.
For agency SEO leads, the consequence lands in two workflows. Keyword research and audience modeling cannot ingest identifiable patient records, appointment histories, or diagnostic data pulled from a client's clinical systems, and call-intelligence programs that transcribe and tag inbound calls cannot repurpose PHI captured on those calls to build lookalike targeting or content briefs. Aggregated, de-identified search demand and SERP data remain fair game; PHI-derived signals do not.
Build the intake so clinical and marketing data streams stay separated by default, and require written authorization documentation before any workflow touches PHI for marketing purposes 2.
Landing-page accessibility and PPC disclosure as part of keyword ROI
Substantiation caps what a page can claim. Accessibility and disclosure caps what a page can convert. Both belong in the keyword decision, not downstream cleanup.
DOJ guidance recommends that websites use accessible headings, keyboard navigation, readable text, captions, accessible forms, and clear paths to report accessibility issues 3. A keyword cluster routed to a landing page that fails those basics wastes the spend that generated the traffic: users who cannot navigate the form do not become leads, regardless of ranking position. Accessibility review belongs on the same checklist as the substantiation review.
When keyword research feeds paid, disclosure geometry becomes part of the ROI calculation. FTC guidance requires disclosures to be clear and conspicuous, with placement and proximity evaluated against the claim they qualify 10. A high-intent PPC keyword tied to a landing page that buries required disclosures below the fold is a compliance exposure, not a conversion asset. Score disclosure placement into the paid keyword rubric before spend is authorized.
Score each cluster against a page type, a conversion action, and a KPI
Every cluster that clears intent, specificity, SERP, and substantiation review should enter production with three fields already filled: the page type it will live on, the conversion action the page must produce, and the KPI the delivery team will be measured against. A cluster that cannot answer all three does not belong in the sprint.
Page type comes from the SERP validation pass. Educational guide, comparison page, service landing page, location page, or FAQ hub are the working options for most client verticals; the current top ten dictates which one is eligible to compete. Conversion action follows from intent: informational clusters convert on newsletter capture, resource downloads, or assisted-chat handoffs; transactional clusters convert on calls, booked consultations, quote requests, or scheduled appointments. Navigational clusters convert on branded engagement and should not be scored against acquisition targets.
KPI assignment is where most agency scorecards break. Ranking position is a graded relevance signal, not a business outcome, and TREC-style evaluation treats it as one input among several rather than the endpoint 8. Tie each cluster to a downstream metric the client already tracks: qualified calls for legal and home services, booked appointments for dental and behavioral health, form-completed leads for senior living. When the cluster ships, the report shows movement on that metric, not just the SERP.
See How Leading Agencies Operationalize Keyword Selection at Scale
Get a walkthrough of advanced workflows for finding, prioritizing, and deploying high-impact keywords across multi-client portfolios—without increasing headcount or sacrificing governance.
If you manage multiple client accounts: running selection at portfolio scale
This section shifts the audience: from a delivery lead selecting keywords for a single client to an SEO director running the same workflow across a book of 20, 50, or 100 accounts. The mechanics do not change. The math does.
At portfolio scale, keyword selection stops being a research task and becomes a throughput problem. Every account needs its own intent taxonomy, its own SERP validation pass, its own substantiation check, and its own scoring rubric. Clustering can widen the intent map faster than manual review, but the algorithm surfaces candidates rather than approving them 6. Judgment still gates production, and judgment does not scale linearly with headcount.
The variables that decide whether a portfolio workflow is sustainable are analyst hours per cluster and clusters shipped per account per quarter. A traditional agency workflow, where an analyst runs discovery, clustering interpretation, SERP validation, substantiation review, and scoring by hand, typically absorbs 4 to 6 analyst hours per cluster. An approval-first automated workflow, where discovery, clustering, SERP pulls, and scoring drafts are produced by software and the analyst reviews and approves, compresses that to roughly 1 to 1.5 hours per cluster. The delta is not a productivity slogan; it is what determines whether a 40-account portfolio ships 8 clusters per account per quarter or 2.
Two operating rules keep the model honest. Version the taxonomy per vertical, not per client, so onboarding a new dental account inherits the labeling scheme instead of rebuilding it. And keep the approval gate on every cluster before production, regardless of how much of the upstream work is automated. Speed compounds only when the review step stays intact.
Support the section's comparison between traditional and approval-first automated workflows using the per-cluster hour ranges and throughput logic described in adjacent prose
Where automation ends and human judgment holds the line
Automation handles the parts of keyword selection that are pattern recognition at volume: pulling SERPs, generating cluster candidates, drafting scoring rows, flagging queries that trip substantiation triggers. It does not handle the parts that require a person accountable to the client. Four decisions stay with a human reviewer, every time.
- Whether an implied claim on the target page is one the client can document.
- Whether a cluster surfaced by unsupervised grouping maps to a service the client actually sells 6.
- Whether the current SERP composition means the client is competing or spectating 8.
- Whether a health, dental, or behavioral-health workflow is drawing on data the client is authorized to use for marketing 2.
A delivery lead who protects those four gates can let software absorb the rest and still ship work that survives a compliance review and a quarterly client meeting. Vectoron is built around that split: automated production, approval-first release, human judgment on every cluster before it goes live.
Frequently Asked Questions
References
- 1.Health Products Compliance Guidance.
- 2.Marketing.
- 3.Guidance on Web Accessibility and the ADA.
- 4.Determining the informational, navigational, and transactional intent of Web queries.
- 5.Classifying Web Queries by Topic and User Intent.
- 6.Classifying the user intent of web queries using k-means clustering.
- 7.Understanding the specificity of web search queries.
- 8.TREC Deep Learning Track: Reusable Test Collections in the Large Data Regime.
- 9.Advertising FAQ's: A Guide for Small Business.
- 10.How to Make Effective Disclosures in Digital Advertising.
