Key Takeaways
- Volume-first keyword lists underperform in regulated verticals because phrase viability depends on what the client can legally substantiate on the landing page, not impression potential 3, 11.
- A six-variable scoring model—intent, specificity, commercial relevance, competitive difficulty, geographic fit, and downstream lead quality—should be weighted differently across awareness, research, decision, and purchase stages 1, 2.
- Substantiation thresholds, disclosure feasibility, and HIPAA constraints belong as pre-shortlist filters, eliminating non-compliant phrase patterns before strategists invest research hours 3, 6, 11.
- Scaling across 20+ accounts depends on vertical-level artifacts—cleared pools, stage weight vectors, and compliant template libraries—so client work reduces to pool selection, authority ceiling, and geographic constraints.
Why Volume-First Keyword Lists Underperform in Regulated Verticals
Agency SEO leads often default to sorting keywords by monthly search volume, filtering by difficulty, and then handing the list to a strategist. While this workflow produces output, in regulated sectors like healthcare, legal, dental, home services, and senior living, it frequently generates phrases that cannot be legally matched to a client's landing page.
The issue isn't that volume is irrelevant, but rather that it's a weak predictor of qualified-buyer behavior in regulated verticals. Here, the viability of a phrase is limited by what the client can substantiate on the page. For instance, the FTC mandates that objective product claims must have adequate substantiation before dissemination, and health or safety claims typically require competent and reliable scientific evidence 3, 11. A phrase like "guaranteed addiction recovery program" might show strong commercial intent and low competition, but no behavioral health client could legally defend it on a landing page without significant risk.
Academic research on keyword selection supports this from a different angle. The Penn State framework, which models organic clicks and rankings, indicates that content relevance becomes crucial when consumers are further along the buying journey and actively seeking to purchase. Conversely, online authority is the primary driver of clicks at the awareness stage 1. Volume alone fails to inform strategists which factor is most important for a given query.
This article introduces a six-variable scoring model that integrates volume as just one of several inputs. This model routes phrases to page templates based on their substantiation capacity and is designed to withstand portfolio-level scrutiny across numerous client accounts.
The Six-Variable Scoring Model for Phrase Selection
Intent, Specificity, and Commercial Relevance
Intent is the foundational variable, as it dictates the potential for all subsequent actions. A query cluster that is purely informational cannot be forced into a consultation-booking template without negatively impacting both conversion rates and ranking stability. A k-means classification of 130,000 web queries revealed that over 75% were informational, with approximately 12% navigational and 12% transactional 9. This distribution is critical for agency planning: the transactional inventory a single client can realistically pursue represents a small fraction of the total addressable query space, and this share diminishes further after applying specificity and geographic filters.
Specificity provides crucial insights that volume cannot. A phrase like "dental implants" has high volume but ambiguous intent, attracting price comparison shoppers, insurance researchers, and even clinicians. In contrast, "All-on-4 dental implants near [city] cost" eliminates ambiguity: the searcher has a specific procedure in mind, a location constraint, and a clear purchase intent. Specificity also highlights the importance of content relevance, particularly for consumers in later stages of their journey who are actively looking to make a purchase 1.
Commercial relevance is the third filter, often overlooked by agency strategists. A phrase can be transactional and specific yet still commercially irrelevant to the client's actual service offerings. For example, a personal injury firm that does not handle maritime claims should not consider "Jones Act attorney" as a candidate, regardless of CPC data. Commercial relevance ensures the client can convert the inquiry into revenue through services they are equipped to deliver, rather than just sounding valuable.
Competitive Difficulty, Geographic Fit, and Downstream Lead Quality
Competitive difficulty is scored next, and it should be assessed against the client's realistic authority ceiling, not an absolute scale. A regional dental group with 40 locations and a modest backlink profile is unlikely to outrank WebMD or the Mayo Clinic for "what causes tooth sensitivity" within a relevant planning horizon. The Penn State framework identifies online authority as the key driver of organic clicks at the awareness stage 1. This means difficulty scoring must weigh domain-level authority against the specific SERP, not solely rely on a third-party tool's keyword difficulty metric.
Geographic fit transforms national keyword lists into actionable targets for multi-location clients. A home services franchise with 60 branches requires phrases scored based on whether the local SERP favors service-area pages, Google Business Profile listings, or city-specific landing templates. Similarly, a behavioral health network licensed in only four states cannot pursue "inpatient rehab [state]" phrases for the other 46 states, regardless of search volume. Geographic scoring is where agency SEO leads must enforce discipline to prevent strategists from wasting time on phrases the client cannot legally or operationally serve.
Downstream lead quality is the sixth variable, distinguishing qualified-buyer phrase selection from mere form-fill optimization. The FTC defines lead generation as identifying and cultivating consumers potentially interested in purchasing a product or service, emphasizing informed inquiries 5. For scoring, this means weighting phrases by what sales or intake teams report as qualified calls, booked consultations, or pipeline contributions, rather than just the volume of forms a page generates.
Weighting the Six Variables by Journey Stage
The six variables do not hold equal weight across all stages of the buyer journey. The common agency tendency to over-emphasize purchase-stage phrases often fails scrutiny when revenue data is considered. A 33-month analysis of a $56 million sponsored search campaign, encompassing nearly 7 million records, categorized phrases into awareness, research, decision, and purchase stages. The study found statistically different behavior across all four stages, and notably, awareness phrases had a lower cost per click and generated more sales revenue in the examined campaign than purchase-stage queries 2. While this campaign was paid search, not organic, and the authors caution that the funnel model classifies query types more cleanly than it describes actual consumer sequencing, the directional finding challenges the default agency assumption that bottom-funnel phrases always yield the best results.
Practical weighting derives from these academic findings. At the awareness stage, authority and specificity are more heavily weighted because content relevance is not yet the primary driver, and the client's ability to rank depends significantly on domain strength 1. Commercial relevance is scored lower here, as awareness phrases should be judged on their capacity to build topical coverage that compounds over time, rather than on direct conversion.
During the research and decision stages, content relevance and specificity become paramount, and competitive difficulty is re-evaluated against the specific SERP rather than the overall domain. At the purchase stage, geographic fit and downstream lead quality dominate, with clear intent being a prerequisite. A single scoring model with stage-specific weight vectors provides agency strategists with a defensible explanation when an account director questions the deprioritization of a high-volume phrase: it scored poorly on the variables most critical for its actual journey stage.
Show how the six scoring variables are weighted differently across the four buyer journey stages, reinforcing the section's central framework
Substantiation-Gated Selection in YMYL Verticals
In regulated verticals, the primary limit on phrase viability is not search volume or SERP difficulty, but rather what the client can legally defend on the landing page without regulatory exposure. Agency SEO leads who treat substantiation as a post-shortlist legal review step often waste strategist hours on phrases that will ultimately be removed before publication. A more efficient workflow applies substantiation capacity as a filter before any research begins.
Three key thresholds perform most of this filtering. The FTC requires objective product claims to be adequately substantiated before dissemination, with health or safety claims generally needing competent and reliable scientific evidence 3. The FTC's health-products guidance further clarifies that endorsements cannot make claims the advertiser couldn't directly substantiate, closing a loophole often exploited by commercial pages with testimonial-heavy layouts 11. For covered entities, the HIPAA Privacy Rule generally mandates written authorization before protected health information is used or disclosed for marketing, with limited exceptions 6. Applying these three thresholds as a pre-shortlist filter eliminates entire categories of phrases before a strategist even opens a research tool.
Figure 3.1 — The three substantiation thresholds that gate phrase viability in YMYL verticals. Objective health or safety claims require competent and reliable scientific evidence 3. Endorsements and testimonials cannot carry claims the advertiser could not substantiate directly 11. HIPAA generally requires written authorization before protected health information is used for marketing 6. A phrase that would force a page past any of these thresholds is removed before shortlisting, not after legal review.
The practical outcome is a concise list of phrase patterns that should be considered non-starters in behavioral health, dental, legal, and senior living accounts, regardless of their commercial intent. Guarantee language regarding treatment outcomes fails the first threshold. Phrases implying typical results from a procedure or legal case fail the first and third thresholds simultaneously, as the FTC's endorsement guidance requires disclosure of material connections and representations of typical consumer experience, and typicality claims themselves require substantiation 8. Phrases that would direct a visitor into an intake flow collecting condition-specific information for remarketing fail the HIPAA threshold if the client is a covered entity and has not secured written authorization 6.
Disclosure placement forms the next filter. The FTC advises placing disclosures as close as possible to the triggering claim and ensuring they are sufficiently noticeable for consumers to see and understand 4. For phrase selection, this rules out candidates whose associated landing page template cannot physically accommodate a prominent disclosure near the headline without disrupting the conversion layout. A phrase like "same-day crown no pain" might score well on intent and specificity, but if the only available template buries pain-management qualifications in a footer, the phrase fails disclosure feasibility and should not be shortlisted.
Operationally, substantiation-gated selection alters the brief a strategist works from. Instead of a raw keyword list filtered by volume and difficulty, the strategist receives a phrase pool pre-cleared against the client's documented evidence base, authorized testimonial inventory, and template library. Phrases outside this cleared pool are not scored. Phrases within it proceed to the six-variable model. The hours saved on legal rework and page revisions typically outweigh the hours spent building and maintaining the cleared-phrase pool. This is why agencies managing multiple YMYL accounts establish this filter once at the vertical level, rather than re-litigating it for each client.
Visualize the three substantiation thresholds that gate phrase viability before any scoring begins, directly supporting the section's pre-shortlist filter concept
Test AI-driven SEO key phrase strategies live
Validate qualified keyword selection by publishing real content and measuring short-term impact before any commitment.
Measurement and Attribution Under Privacy Constraints
Phrase selection can silently fail when the measurement system cannot differentiate qualified calls from irrelevant form submissions. The FTC's perspective on lead generation emphasizes informed inquiries from consumers genuinely interested in purchasing a product or service, with clear disclosures about data usage 5. For agency SEO leads, this shifts the attribution question from "which phrase generated the form" to "which phrase produced an inquiry the client's intake team actually valued." These two questions often yield different scoring outcomes, which is significant.
Privacy constraints then dictate what answers an agency can even collect. HHS defines online tracking technology as code or scripts used to gather information about user interactions with websites or mobile applications, outlining HIPAA implications for covered entities and business associates when this data involves health information 7, 12. For behavioral health or dental clients, a conventional analytics stack that pipes query strings, page URLs containing condition names, or form field contents into an advertising audience can create exposure that no phrase-level ROI justifies. The practical consequence for scoring: phrases whose value is only proven through condition-level attribution should be deprioritized if the attribution path itself is unavailable, rather than being reported as unmeasurable and quietly dropped.
Three measurement patterns effectively navigate these constraints:
- Call intelligence that records qualification outcomes without transmitting Protected Health Information (PHI) to ad platforms provides phrase-level insights on booked consultations.
- Server-side event forwarding with field-level filtering supports conversion counts without exposing the content of intake forms.
- Documented offline-conversion imports from the client's practice management or CRM, keyed to call or appointment IDs rather than user identifiers, allow an agency to link phrase clusters to pipeline without expanding the covered entity's tracking footprint.
Each pattern influences which phrases are scoreable, providing the necessary measurement input for the six-variable model.
Phrase Selection When AI Overviews Absorb the Click
The economics of phrase selection change when the SERP answers a query before the user reaches a website. McKinsey's analysis of AI-mediated search indicates that approximately 50% of Google searches already return AI summaries, with a projection that this share could exceed 75% by 2028 10. While this is a scenario, not a definitive forecast, and depends on product rollout and adoption, it fundamentally reshapes scoring. Phrases whose answers can be fully synthesized from public sources lose click value, even if they retain impression value. Conversely, phrases requiring vendor-specific evidence, local operational details, or credentialed judgment tend to retain their click value more effectively.
Agency SEO leads should re-score two categories of phrases in light of this shift. The first includes definitional and comparative informational queries, which AI overviews absorb most completely because the answer is a synthesis task. Phrases like "what is subacute rehabilitation" or "PPO vs HMO dental coverage" may still generate impressions, but their click-through rate diminishes as summaries become more comprehensive. These phrases retain their value as topical-authority building blocks, supporting the Penn State finding that authority drives awareness-stage visibility 1, but they lose their effectiveness as direct conversion drivers and should not be scored for lead quality.
The second category, queries where the answer requires facts only the client can provide, gains relative value. Information such as pricing within a specific market, wait times for a particular location, intake requirements for a specific insurance panel, availability of a specific procedure code, or attorney admissions in a specific jurisdiction cannot be answered by a general-purpose summary without citing a source. Scoring should assign higher weight to these phrases for both specificity and downstream lead quality. This is because the AI layer tends to cite rather than replace such information, and clicks following a citation arrive with more purchase context than a typical organic click did previously.
Two scoring adjustments follow from this. First, incorporate a citation-likelihood signal into the specificity variable: phrases that compel the summary to attribute a source increase the probability of a branded citation, which builds authority over time. Second, discount commercial relevance for phrases whose landing pages directly compete with the AI summary's own answer surface. A page designed to merely restate what the overview already provides will not earn the click, regardless of how well it scored on volume and difficulty in a pre-AI model.
See How Leading Agencies Operationalize Qualified Keyword Targeting at Scale
Connect with our team to review data-driven frameworks for selecting, ranking, and deploying SEO key phrases that align with high-intent buyer journeys across multiple client accounts.
Scaling the Framework Across an Agency Portfolio
Templating the Scoring Model for 20+ Client Accounts
This section addresses the challenge of applying the phrase selection framework across an agency portfolio, specifically for SEO leads managing 20 or more accounts across multiple verticals. The core question here is not whether the six-variable model works, but whether it can be consistently enforced by strategists who weren't involved in its initial development.
Templating begins with vertical-level artifacts, not client-specific ones. For each vertical an agency serves, three assets are built once and versioned:
- A cleared-phrase pool pre-filtered against the substantiation thresholds applicable to that vertical.
- A weight vector specifying how the six variables are scored at each journey stage.
- A page-template library that maps phrase patterns to layouts capable of prominently displaying required disclosures near the triggering claim 4.
For example, a behavioral health template library will differ from a home services one because the former must accommodate HIPAA-constrained intake flows 6, while the latter does not. Agencies that develop these artifacts at the vertical level avoid repeatedly addressing the same compliance issues for each individual account.
Client-level work then simplifies to three decisions:
- Identifying the appropriate cleared-pool subset for the client's service mix.
- Determining the authority ceiling for difficulty scoring based on the client's domain.
- Establishing the geographic constraints for the shortlist.
Strategists score phrases within the templated model rather than rebuilding it from scratch. The output is a shortlist that an account director can confidently defend without requiring a legal review cycle, because the filters that would have triggered such a review were applied before research even began.
Portfolio Economics: Strategist Time Versus Phrase Throughput
The viability of the templated model across a portfolio of 20 to 80 accounts hinges on portfolio economics. Key variables include hours per client for phrase research, phrases evaluated per strategist hour, the percentage of a shortlist that passes substantiation review by vertical, and the resulting throughput per strategist per week. While dollar figures vary by market, the comparison below uses only variables an agency SEO lead can populate from their own time-tracking data.
| Variable | Manual per-client workflow | Templated scoring workflow |
|---|---|---|
| Vertical artifacts built once (cleared pool, weight vectors, template library) | None; rebuilt per account | Built once per vertical, versioned |
| Phrases evaluated per strategist hour | Lower; each phrase re-scored from scratch | Higher; phrases scored inside pre-filtered pool |
| Shortlist survival rate after substantiation review | Variable; depends on strategist's compliance familiarity | Higher; non-compliant patterns removed pre-shortlist 3, 11 |
| Legal rework hours per client per quarter | Higher; triggered by shortlist, not by filter | Lower; filter catches patterns before research |
| Hours per new client onboarding | Full phrase research cycle | Vertical artifact selection plus client-specific inputs |
| Attribution scope for scoring | Form fills, often decoupled from pipeline | Qualified calls and booked consultations where privacy constraints permit 5, 7 |
The throughput formula for agency leads is straightforward: phrases scored per strategist per week multiplied by the shortlist survival rate yields usable phrases per strategist per week. The templated workflow increases both factors in this equation. It also reallocates hours away from compliance rework, a cost often invisibly absorbed because it's logged against the triggering account rather than the underlying methodology. Agencies aiming to expand their client base without proportional hiring are, in essence, shifting work from repetitive client-level tasks to scalable vertical-level artifacts. Platforms like Vectoron, which streamline specialist recommendations through a single approval workflow, operate on this same principle.
A Defensible Phrase Strategy, Not a Longer Keyword List
Agency SEO leads who excel in regulated verticals are not those who conduct the most keyword research cycles. Instead, they are those who have clearly defined which variables matter at each journey stage, identified phrase patterns that are non-starters before research begins, and established which measurement signals truly indicate a qualified buyer. A six-variable scoring model, weighted by journey stage, applied within a pre-cleared phrase pool, measured against pipeline rather than just form fills, and re-evaluated for how AI summaries redistribute click value, provides strategists with a shortlist that an account director can confidently defend without a legal review cycle.
At the portfolio level, this approach generates capacity. Hours are reallocated from client-by-client compliance rework to vertical-level artifacts that compound across accounts. Strategists score phrases within an established model instead of rebuilding one. Shortlists survive review because the filters that would have caused them to fail were applied upfront. For agency SEO leads seeking to expand their client base without increasing headcount, the shift is from simply generating longer keyword lists to implementing governed workflows. Platforms like Vectoron are designed around this approval-first operating model.
Frequently Asked Questions
References
- 1.Keyword Selection Strategies in Search Engine Optimization.
- 2.Bidding on the buying funnel for sponsored search and keyword advertising.
- 3.Advertising FAQ's: A Guide for Small Business.
- 4.How to Make Effective Disclosures in Digital Advertising.
- 5.Staff Perspective: "Follow the Lead" workshop.
- 6.Marketing.
- 7.Resource for Health Care Providers on Educating Patients about Using Online Tracking Technologies.
- 8.16 CFR Part 255: Guides Concerning the Use of Endorsements and Testimonials in Advertising.
- 9.Classifying the user intent of web queries using k-means clustering.
- 10.New front door to the internet: Winning in the age of AI search.
- 11.Health Products Compliance Guidance - Federal Trade Commission.
- 12.Use of Online Tracking Technologies by HIPAA Covered Entities and Business Associates.
