What evidence should I require before increasing our GEO budget?
Short answer
Require four things before you scale: citation rate on tracked commercial prompts up 10 percentage points over a quarter, share of voice flat or gaining against three named competitors, branded search or direct traffic up 5% with no competing explanation, and survey attribution at 15% or higher of new customers naming AI surfaces. Miss any one and you hold. Two red flags override all four and mean you hold regardless of how good the trend charts look.
Why "AI is the future" is not evidence
Every GEO vendor leads with a narrative that feels compelling in the room and evaporates three months later when someone asks what changed.
The confusion is between two very different claims. "AI-search matters" is probably true. "Increasing our GEO budget will produce return" requires actual measurement. Scale on the first and you've made GEO the expensive experiment it doesn't need to be.
Measure for a full quarter before you commit incremental spend. See Is AI-search visibility commercially meaningful enough to deserve its own budget? for whether to fund it at all.
The four thresholds
| Signal | Threshold to scale | Where it comes from |
|---|---|---|
| Citation rate trend | Up 10 percentage points over the quarter on worked prompts | Prompt tracking across two or more surfaces |
| Share of voice | Flat or gaining vs three named competitors | Same prompt set, same period |
| Branded search or direct lift | Up 5% or more, seasonally adjusted, no competing cause | GSC and GA4 over 24 months |
| Survey attribution | 15% or higher of new customers naming an AI surface | Signup survey, 100+ responses |
Below 10% on survey attribution, the channel hasn't reached your buyers yet. Between 10 and 15% you're at the edge, so hold current spend one more quarter rather than scaling into uncertainty.
The citation-rate threshold is set at 10 points because our own state of AI search research found that improvements tend to persist rather than reverse. Once a brand is mentioned for a prompt, next-observation mention probability is 83.2% on ChatGPT, 83.3% on Perplexity, and 84.2% on Google AI Mode. A real 10-point gain is likely to hold, which is what makes it worth scaling behind. Measure it as a series though, because researchers argue visibility should be measured as a distribution across runs, prompts, and time rather than as one snapshot, and a 10-point gain read from one observation is not a trend.
The two red flags that override everything
Red flag one: gains concentrated on a single surface. Our research found mean domain overlap between ChatGPT and Perplexity is 0.05, so winning on one surface doesn't propagate to the others. Scaling behind a single-surface win means paying for work that only pays off in a narrow slice of buyer behaviour.
Red flag two: visibility improved but branded search and direct traffic didn't move. This is the strongest evidence that your AI-search work isn't reaching buyers who convert. Either you're winning prompts your buyers don't use, or the surface isn't sending the attention you assumed. Adding budget amplifies the miss.
Either flag means hold, even with all four thresholds green.
Run the review
1. Confirm the prompt list still matches buyer intent. Refresh quarterly. Stale prompt lists produce confident numbers about questions nobody asks anymore.
2. Pull citation rate by surface. Weekly data over 12 weeks, split by prompt category.
3. Pull share of voice against three named competitors. Same prompts, same window.
4. Pull branded search and direct traffic monthly. 24 months of history for seasonal context.
5. Pull survey answers. 100 minimum for the quarter.
6. Apply the thresholds, then check both red flags. All four green and neither flag raised means scale. Anything else means hold.
Automate the quarterly review
The four thresholds and two red flags are only a discipline if something applies them consistently, and a quarterly spreadsheet rebuilt by a different person each time is not that.
model-blind-spots is what catches the first red flag, because it surfaces per-surface weakness rather than averaging your providers into one number that hides a single-surface win. rising-threats catches the case where your citation rate improved and competitors improved faster, which is a negative return that looks positive in isolation.

The quarterly agent:
Start (schedule, quarterly) → Visibility Score (per provider, weekly over 12 weeks) → Citation Share → brand-vs-competitor recipe → rising-threats recipe → model-blind-spots recipe → GSC Top Keywords for Site (branded regex, 24 months) → GA4 AI Traffic Overview → HubSpot Search Contacts (survey answers) → workflow-memory (prior quarterly runs) → Code (evaluate all four thresholds numerically, then test both red flags, and return a scale, hold, or reduce verdict) → Conditional (if survey responses under 100, force a hold verdict regardless of the other signals) → Prompt LLM (state the verdict and which specific threshold failed) → DOCX export → Send Email to CMO and CFO.
Forcing a hold on thin survey data is the safeguard that matters most. Three of the four thresholds can look excellent while the channel is reaching nobody who converts, and the survey is the only signal that catches it.
FAQ
Related answers
- Is AI-search visibility commercially meaningful enough to deserve its own budget?
- How should I budget for SEO when AI answers are reducing organic clicks?
Want the evidence pipeline that tells you when to scale AI-search budget? Start a free Analyze AI trial and run the quarterly review from live data.
