Key Takeaways

  • A single difficulty score collapses too many variables to guide agency decisions, since Google ranks page by page and identical scores can hide very different competitive realities 11.
  • Pull location-specific demand from the Google Ads API and inspect monthly series for seasonality, so clusters aren't mislabeled evergreen based on an annual average 10, 9.
  • Treat the paid competition index as advertiser pressure on ad slots, not organic difficulty, keeping it in a separate column from the six organic signals 9.
  • Score intent clarity against the live SERP by checking whether the top ten results align to one intent, downweighting mixed SERPs that won't fit the client's page type 11, 4.
  • Record SERP composition feature by feature, because local packs, AI overviews, and PAA units change the clicks available to an organic winner regardless of nominal rank 8.
  • Rate incumbent strength page by page across the top ten URLs, since mixed SERPs with beatable pages are the real opportunity band, not strong incumbent lineups.
  • Measure the client's authority gap using Search Console impressions and positions on neighboring queries rather than third-party domain scores, as existing footprint shortens the path to ranking 7, 12.
  • Use Core Web Vitals thresholds and URL Inspection as a binary technical gate so strategist time isn't spent scoring clusters a failing template cannot support 6, 12.

Why a single difficulty score breaks at agency scale

A keyword difficulty score compresses a dozen variables into one integer. This works for a solo consultant but fails when an agency manages 40 clients across diverse verticals and geographies. The single number doesn't reveal which signal drives the score, if the client can compete on that signal, or if ranking would generate revenue.

The commercial pressure to optimize search strategies is increasing. U.S. search advertising revenue is projected to reach $114.2 billion of a $294.6 billion digital ad market in 2025, with search growing approximately 11% year over year 3. More budget chasing the same results means organic incumbents defend harder, SERP features consume more real estate, and the cost of mis-prioritized keyword clusters compounds across every client.

Google's ranking documentation explicitly states that scoring occurs page by page across hundreds of billions of pages, with site-wide classifiers playing a secondary role 11. A domain-level difficulty proxy cannot capture what Google actually evaluates. Two keywords with identical third-party scores can hide a thin-content incumbent in one SERP and a hardened category leader in another.

For an agency head of SEO, the solution isn't a better single score. It's replacing the score with a composite rubric that junior strategists and automation can apply consistently. This frees senior review time for prioritization and approval, rather than per-keyword forensics. The following section defines the eight signals this rubric must measure.

Chart showing US Digital Ad Revenue Breakdown (2025)US Digital Ad Revenue Breakdown (2025)

Based on the IAB/PwC 2025 annual report, this shows the breakdown of the $294.6 billion total US digital advertising revenue, with search advertising comprising $114.2 billion.

The eight signals that actually predict ranking feasibility

Location-specific demand pulled from the Ads API

Demand is the initial filter because a keyword without searchers in the client's service area cannot produce business outcomes, regardless of SERP winnability. The Google Ads API provides historical metrics with geographic, language, and network parameters, including average monthly searches and monthly search volumes over the past year 10. Pulling demand via the API, rather than manually, ensures consistency across 40 clients with varying geo-targets.

Strategists should note two operational aspects. First, volumes are estimates, not exact counts, and serve as a sizing input rather than a precise forecast 10. Second, the KeywordPlanHistoricalMetrics resource returns monthly volumes as a series, allowing the rubric to flag seasonality before a cluster is labeled evergreen 9. A query averaging 1,200 searches monthly but concentrating 70% of that volume in two months behaves differently from a consistently flat 1,200-per-month term, and the composite score should reflect this.

The paid competition index is frequently misinterpreted in keyword research. Google defines it on a 0 to 100 scale, calculated as the number of filled ad slots divided by the total available ad slots for a given keyword and location 9. It measures advertiser competition for paid inventory, not the difficulty of organic ranking.

Confusing these two signals leads to errors. A home services query in a mid-size metro might show a competition index above 80 because many local providers bid on it, yet the top organic positions could be held by directories with thin, templated pages that a competent client site can easily outrank. Conversely, a specialized B2B term might have a low paid index due to fewer advertisers, while its organic SERP is dominated by category incumbents with extensive topical coverage.

The rubric separates these into distinct columns. The paid competition index, along with 20th and 80th percentile bid estimates, comes directly from the Ads API, indicating commercial interest and ad-slot pressure 10. Organic difficulty is assessed using the subsequent six signals. This distinction prevents strategists from dismissing winnable organic opportunities based on misleading paid metrics.

Intent clarity scored against the live SERP

Intent is evaluated twice: first from the query itself, then against the actual results Google serves. The second evaluation is crucial because Google's ranking systems operate at the page level, and the live SERP clearly indicates what the algorithm currently rewards as relevant 11.

Historical research provides a taxonomy. A Penn State study from 2007, analyzing over five million queries, found approximately 80% informational, 10% navigational, and 10% transactional intent, with automated classification reaching 74% accuracy 4. While this study predates modern search advancements, it highlights the inherent uncertainty in intent labels and that most queries don't explicitly state commercial purpose.

Operationally, the rubric assigns a clarity score (high, mixed, or low) based on how uniformly the top ten results align to one intent. A SERP with nine how-to guides and one product page indicates high-clarity informational intent. A SERP mixing calculators, product pages, local packs, and definitions suggests low-clarity, warranting downweighting even if the query appears commercial. Mixed-intent SERPs are often where junior strategists overcommit, targeting keywords whose top positions won't accommodate the client's page type.

SERP composition: features, formats, and real estate

The composition signal records the actual appearance of the SERP, feature by feature. Local packs, AI overviews, shopping modules, video carousels, image blocks, and People Also Ask units all push the first organic link further down the page, altering the realistic click share available to an organic winner. The European Commission's DMA documentation on Alphabet describes ranking data as a URL's position, its ordinal relationship to other results, and its relative or absolute prominence, which is a useful framework for scoring real estate rather than assuming uniform value for position one 8.

The rubric captures three composition inputs:

  • presence of non-organic features,
  • number of organic positions above the fold on mobile, and
  • whether the dominant feature accommodates the client's content type.

A SERP where a local pack and an AI overview consume the top of the viewport presents a different competitive environment than a ten-blue-link SERP with the same difficulty score. The composite rating should reflect this disparity between nominal rank and attainable clicks.

Incumbent organic strength assessed page by page

Domain-level authority scores often obscure the critical signal: whether the specific ranking pages are strong, average, or beatable. Google's ranking guidance states that its systems primarily operate at the page level, with site-wide signals contributing but not dominating 11. This makes page-level inspection the appropriate unit of analysis.

The strategist examines each of the top ten URLs, noting:

  • content depth relative to query scope,
  • freshness of primary date signals,
  • presence of original data or expert sourcing,
  • internal link context from the hosting site, and
  • whether the page addresses the full intent or only a facet.

A thin aggregator in position three is a different competitor than a well-maintained category page by a known publisher, even if both sites have similar third-party authority scores.

Scoring across the top ten yields an incumbent strength rating of weak, mixed, or strong. The rubric identifies "mixed" as the opportunity band: at least two or three incumbent pages that a competent client page could surpass in depth, recency, or specificity. Strong incumbent SERPs are deprioritized regardless of what a difficulty tool reports.

Client authority gap measured in Search Console, not DA proxies

Third-party authority scores describe a domain in isolation. The signal that predicts a client's ability to compete for a specific cluster is the gap between their existing topical footprint and the incumbent set on that SERP. Search Console is the primary source for this measurement, as it reports what Google already associates with the client's pages.

The Search performance report details impressions, clicks, CTR, and average position by query, page, country, device, and search appearance, accessible via the interface, Search Analytics API, Looker Studio connector, or spreadsheet exports 7. For agency workflows, the API or Looker Studio connector ensures repeatable data pulls across the client roster. URL Inspection confirms current index status and allows live testing when a specific page is evaluated as a responder for a target cluster 12.

The gap score addresses three questions for each candidate cluster:

  1. Does the client already receive impressions for semantically related queries?
  2. If so, what is the average position range?
  3. Do the pages receiving these impressions address the target intent, or is a new page required?

A client already ranking at positions 11 to 25 for neighboring queries has a significantly shorter path than one with no footprint, irrespective of any domain-level metrics.

Technical eligibility as a precondition for competing

Technical eligibility acts as a gate, not a tiebreaker. While content and relevance drive rankings, a page failing basic page experience thresholds wastes the opportunity created by strong content. Google defines Core Web Vitals as real-world metrics for loading, interactivity, and visual stability, recommending LCP within 2.5 seconds, INP below 200 milliseconds, and CLS below 0.1 for a good user experience 6.

The rubric records these three measurements for the client page intended for the target cluster, or for the template if the page doesn't exist. A template failing INP on mobile is flagged before any content work is scoped, as fixing interactivity is a prerequisite for keyword investment to yield returns. Indexing status from URL Inspection is checked concurrently 12.

Good vitals don't guarantee rankings; Google explicitly states that many ranking signals and content relevance also matter 6. The eligibility gate's purpose is narrower: a page that cannot clear the measurable technical floor should not consume strategist time on competition scoring until that floor is met.

Conversion value weighted by the client's revenue model

The final signal connects the rubric to the client's P&L. A high-volume, low-competition cluster that doesn't align with the client's conversion path is a poorer investment than a smaller cluster reaching qualified buyers. Conversion value is a weight applied to the composite score, not a replacement for other signals.

Weighting inputs are client-specific, derived from live account data:

  • average deal value or lifetime value by service line,
  • qualified lead rate from organic by cluster (where historical data exists), and
  • the cluster's position in the buying path (awareness, consideration, comparison, or decision).

The intent signal from the SERP pass informs this calculation but doesn't supersede it, as a transactional SERP for a low-margin service is less valuable than a comparison-stage SERP for a high-margin one.

Clusters are tagged with a value tier that modifies the composite priority. A mid-score cluster linked to the highest-value service often outranks a top-score cluster tied to a loss leader during prioritization.

Chart showing Breakdown of web query intent (historical study)Breakdown of web query intent (historical study)

A foundational academic study of over five million web queries found that approximately 80% were informational, 10% were navigational, and 10% were transactional.

Codifying the signals into a scoring rubric any strategist can run

The rubric's purpose is to remove subjectivity from data collection and concentrate judgment at the prioritization stage. Each of the eight signals is scored on a fixed scale with a defined data source, ensuring consistent composite calculations. A strategist following the rubric for a dental client on Tuesday and another for a home services client today should produce comparable scores for similar SERPs.

A practical scoring sheet uses three bands per signal:

  • Demand is scored low, mid, or high against vertical-specific thresholds from the Ads API historical metrics, with monthly series flagged for seasonality 9.
  • The paid competition index converts the 0-100 Ads field into low (0–33), mid (34–66), or high (67–100) 9.
  • Intent clarity, SERP composition, and incumbent strength are scored low, mixed, or strong based on live SERP inspection, which Google's ranking guidance supports as the primary evaluation unit 11.
  • Authority gap is derived from Search Console segmentation of existing impressions and positions on neighboring queries 7.
  • Technical eligibility is a binary gate against Core Web Vitals thresholds.
  • Conversion value is a weight tier, not a band.

The composite score combines the first seven signals into a single priority, which the value tier then modifies. Clusters below a cutoff are archived with recorded reasons. Those above the cutoff enter the approval queue with their scoring sheet, providing senior reviewers with inputs rather than just a conclusion to accept or reject.

Test Real Keyword Competition Insights in Production

Trial live workflows for competitive keyword analysis and publish actionable content on actual client projects.

Start Free Trial

Where the rubric compresses senior strategist time

The benefit of codifying these eight signals is not the elimination of work, but its reallocation. The process shifts from senior strategists' calendars to a pipeline executable by interns, junior analysts, and scripted pulls against a defined specification.

The table below maps each signal to its data source and highlights two throughput variables: scriptability via API and requirement for human SERP inspection. Time estimates are omitted as they depend on vertical complexity and tooling maturity, not on figures this article can provide.

SignalData sourceScriptable pullHuman SERP review required
Location demandGoogle Ads API historical metrics 10YesNo
Paid competition indexGoogle Ads API KeywordPlanHistoricalMetrics 9YesNo
Intent clarityLive SERP inspection 11PartialYes
SERP compositionLive SERP inspection 11PartialYes
Incumbent strengthPage-by-page review of top ten URLs 11NoYes
Authority gapSearch Console API or Looker Studio 7YesNo
Technical eligibilityCore Web Vitals plus URL Inspection 12YesNo
Conversion valueClient CRM and account dataYesNo

Five of the eight signals can be processed without senior involvement once the pipeline is established. Senior time is then concentrated on the three SERP-inspection steps and the approval of ranked outputs, which is the appropriate application of scarce expert judgment.

Running the workflow inside Google's scaled-content boundaries

Codifying the rubric into a repeatable pipeline naturally leads to the question of how much scoring, drafting, and page production can be automated without Google classifying the output as scaled content abuse. Google's policy defines scaled content abuse as generating many pages primarily to manipulate search rankings rather than help users, citing generating many low-value pages with generative AI as an example 1. The violation is not the use of automation itself, but producing pages whose primary purpose is ranking over user value.

Google's generative AI guidance reinforces this distinction. AI can assist content creation, but output must meet Search Essentials and spam-policy standards. Producing many pages without added value may trigger the scaled-content rule 2. Both documents were updated in September 2026, with current wording addressing manipulation of generative AI responses alongside traditional rankings 1.

For the eight-signal workflow, this translates into three design constraints:

  • Signal collection, scoring, and prioritization can utilize scripted pulls and model-assisted analysis without policy exposure, as this work doesn't produce published pages.
  • Content production triggered by a high-priority cluster must add something not already present in the SERP: original client data, first-hand expertise, proprietary analysis, or synthesis that incumbent pages lack.
  • Output volume is governed by the number of clusters clearing the rubric's value tier, not by the capacity of the generation layer.

An agency that allows production capacity to dictate publishing cadence inverts the policy and risks the exact classification it aims to avoid.

See How Leading Agencies Benchmark Keyword Competition at Scale

Get a walkthrough of data-driven workflows for assessing keyword difficulty across multiple clients—designed for agencies seeking efficient, repeatable competitive analysis without expanding team resources.

Contact Sales

Approval-first operations: machine speed below, human judgment at the gate

The economic efficiency of the rubric relies on keeping machine-speed and human judgment layers distinct. Signal collection, scoring, Search Console data segmentation, and ranked outputs can run continuously across the client roster 7. Prioritization, cluster selection, and the decision to commission a page remain with a senior reviewer. This split is not merely philosophical; it's the only configuration that maintains high throughput without exposing the agency to the scaled-content classification the rubric is designed to prevent 1.

Three gates concentrate senior strategist time where it has the most impact:

  1. The SERP inspection pass, where intent clarity, composition, and incumbent strength are evaluated against pages Google actively ranks 11.
  2. The ranked output review, where composite scores and value tiers are weighed against account strategy, client appetite, and portfolio conflicts.
  3. The production approval, where the brief for each winning cluster is checked for the original data, expertise, or synthesis that distinguishes a useful page from a scaled one 2.

Everything upstream of these gates runs on automation and junior execution. Everything downstream of approval ships. Agencies that scale client wins without increasing senior headcount are those that shift from treating keyword competition as a research question to managing it as a governed pipeline—precisely the operating model the Vectoron platform is built to support.

If the agency manages multi-location or franchise portfolios

Agencies managing multi-location or franchise portfolios face a distinct scoring challenge compared to single-brand accounts. While the eight signals remain constant, demand, SERP composition, and incumbent strength vary by location. This necessitates running the rubric per market rather than per client.

Three adjustments simplify the pipeline:

  • Demand pulls use the Ads API with each location's geographic parameter instead of a national average, as franchise markets rarely share identical monthly search series 10.
  • Search Console segmentation runs by country and page group, ensuring the authority gap is measured against specific location pages competing in each market, not the aggregated brand domain 7.
  • SERP inspection samples a representative set of markets—typically the top-revenue, median, and lowest-performing locations—rather than every market, with the sample refreshed quarterly.

The output is a priority list per location, rolling up to a portfolio view. This allows senior reviewers to approve clusters where both the composite score and the location's revenue weight meet the defined cutoff.

Infographic showing Year-over-year growth of US search advertising revenue (2025)Year-over-year growth of US search advertising revenue (2025)

Year-over-year growth of US search advertising revenue (2025)

Frequently Asked Questions