Is Your GEO Work Actually Working? How to Tell Signal From Noise

AI answers change from one run to the next even when nothing about your brand has changed. Here's how much they wobble, how many answers you need before trusting a number, and how to judge whether a specific fix worked.

2 min readAnovox team

The same question, different answers

Ask an AI assistant the same question twice and you will often get two different lists of recommended brands. Models sample their answers, search results shift, and small changes in wording move things around. That means one answer is an anecdote. If you check once, see your brand missing, publish a page, check once more and see it appear, you have learned very little — it might have appeared anyway.

This is the most common way GEO results get misread, in both directions: celebrating noise as a win, or abandoning work that was helping.

How big the noise is

Share of voice is a proportion: of the answers collected, how many named you. Proportions from small samples are uncertain in a precise, calculable way. With 20 answers, if 10 name you, the honest statement isn't "50%" — it's "somewhere around 30% to 70%". If none of 20 name you, your true rate could still plausibly be as high as 16%.

Those ranges come from a Wilson score interval, a standard statistical method that behaves well with small counts and at 0% or 100%. They shrink as answers accumulate: at 200 answers, the same 50% narrows to roughly 43% to 57%.

How many answers you need

A practical rule: don't read week-to-week movement unless each week has at least a few dozen answers behind it. Get there by asking each question several times per check, and by pooling a rolling window — a month of weekly checks, or a week of daily ones — instead of reacting to a single run.

When you compare two periods, ask whether the difference is bigger than the noise. A two-proportion test at 95% confidence is a sensible bar. On 10 answers each, going from 3 to 5 mentions — a 20-point jump — is not significant. The same proportions on 200 answers each are.

Judging one specific fix

Most GEO work is targeted: you publish a comparison page because you were missing from "[competitor] alternatives". So measure that prompt, not your overall number. Compare how often that exact question named you in the weeks before the page went live against the answers since. Wait for at least a handful of new answers before judging. And look for the most direct evidence of all: an engine citing the page you published.

A fix that moved its own prompt from 1 in 12 answers to 9 in 12 has clearly worked. A fix that moved it from 2 in 10 to 3 in 10 hasn't shown anything yet — keep measuring.

How Anovox applies this

Every share-of-voice figure in Anovox shows its likely range and the number of answers behind it, pools a rolling window, and only reports a change when it passes a 95% significance test; otherwise it says "no clear change". When you mark a fix published, the dashboard compares that prompt before and after, labels it "too early" until enough answers arrive, and tells you when an engine cites your new page. The full scoring model is published at /resources/how-anovox-scores-visibility.

See this on your own brand

Run the free check — same scoring the dashboard uses, no account needed.

Run free check