Most SEO work is never verified. A change is made, something happens to traffic over the following weeks, and the two are connected by narrative. That is not a criticism of practitioners — verification is genuinely hard, the data is noisy, and the platforms that hold it are not built to answer "did this edit do anything". But it does mean the field accumulates advice that has never survived a test, and it means an agency and a client can disagree about outcomes for a year with no way to settle it.
So we built a verdict step. Every change we apply gets judged, on a schedule, by the same procedure, and you are shown the judgment whether or not it flatters us. It has one precondition, and it is worth stating plainly because everything below depends on it: the verdict is computed from Google Search Console, so a site that has not connected Search Console gets no verdicts at all — and, since autonomy is earned from verdicts, nothing on such a site can ever go automatic.
The procedure
For each changed page we compare Google Search Console metrics across two windows: a before window ending the day before the change went live, and an after window starting a few days later. The gap is deliberate. A change does not take effect the moment it is applied — it takes effect when the page is re-crawled and re-processed — and including those in-between days would mix pre-change and post-change behaviour into the "after" number.
We run the comparison twice, at two horizons:
- d14 — fourteen-day windows
- An early read. Enough to catch something going badly wrong, rarely enough to be confident about a modest gain.
- d28 — twenty-eight-day windows
- The one we actually lean on. Slower, less noisy, and long enough that a single unusual week does not decide the answer.
Both are stored. We do not overwrite the early verdict with the later one, because the disagreements are informative: a change that helped at d14 and didn’t move at d28 usually caught a novelty bump rather than a durable improvement, and you only see that if both numbers survive.
The four answers
A verdict is one of four values, and the fourth is the important one. These are the words themselves — the four labels below are the strings the app puts on the change, in the weekly email and in the Work queue, not a paraphrase of them.
- It helped
- Click-through rate rose by at least 15% relative, or average position improved by at least 1 rank without impressions collapsing alongside it.
- It did worse
- The mirror image of the same two tests.
- It didn’t move
- Neither threshold was crossed, or the two signals disagreed with each other. A change that improved position while CTR fell has not clearly done anything, and saying so is more useful than picking the flattering half.
- It couldn’t be measured
- The two windows together carry fewer than 50 impressions, or one of them has no data at all.
Note what that fourth answer is not: it is not “still measuring”. A change whose window has not closed yet says so and prints the day its verdict can land. “couldn’t be measured” is the opposite — the comparison ran, the window is closed, and the answer is final. Nothing further arrives for that change, and the app does not show you a date pretending otherwise.
That last one is a real answer, not a failure state, and we show it as one. A page with almost no search impressions cannot produce a statistically meaningful before-and-after; any verdict computed from it would be a coin flip wearing a label. The honest output is "there was not enough signal here to judge", and a system that never emits that verdict is a system that is guessing on its quietest pages and telling you it isn't.
It is also, in practice, the most common verdict on a small site, and we would rather say so up front than have it arrive as a disappointment. If most of your pages sit below the impression floor, the useful work is probably not meta description tuning — it is having pages that people search for.
The guards, and why they exist
Two of the thresholds above carry conditions that look fussy and are not.
The position signal is guarded by impression volume. Average position is an average over impressions, so it moves mechanically when the mix of queries changes: a page that suddenly picks up a long tail of low-ranking impressions will show its average position get worse while nothing about the page got worse. So a position gain only counts as helped if impressions did not drop more than 20% alongside it, and a position loss only counts as “did worse” if impressions did not surge by the same margin. Without that guard, the most common “did worse” verdict would be a page that started ranking for more things.
Clicks and impressions are compared as per-day rates rather than window totals, because the two windows are rarely equally populated — Search Console data arrives late and unevenly, and comparing a fourteen-day total against an eleven-day total would manufacture a decline out of nothing.
What a verdict is not
It is not causation, and we label it as correlation everywhere it appears. Two windows of search data around an edit is evidence. Between those windows the engine also updated its systems, your competitors also changed their pages, the season also moved, and demand for your subject also drifted. We control for none of that, and neither does anyone else with access to the same data.
It is also narrow. A verdict speaks for Google search only, because that is the only source that reports impressions and clicks — it says nothing about whether an assistant started citing the page, which is a separate lane with separate evidence and separate limits (crawlers, referrals and agents). A change that helped one and not the other is common, and neither number can stand in for the other.
What the verdict does give you is a discipline. It makes the claim falsifiable, it makes it dated, and it makes the record cumulative — after thirty changes you have thirty verdicts, and the pattern across them says considerably more than any single one. It also means a change that made things worse gets found by a scheduled process rather than by someone eventually noticing.
That is the whole bet: a measured verdict that is often inconclusive is worth more than a confident story that was never checked.