Most SEO work is never verified. A change is made, something happens to traffic over the following weeks, and the two are connected by narrative. That is not a criticism of practitioners — verification is genuinely hard, the data is noisy, and the platforms that hold it are not built to answer "did this edit do anything". But it does mean the field accumulates advice that has never survived a test, and it means an agency and a client can disagree about outcomes for a year with no way to settle it.
So we built a verdict step and made it mandatory. Every change we apply gets judged, on a schedule, by the same procedure, and the judgment is published whether or not it flatters us.
The procedure
For each changed page we compare Google Search Console metrics across two windows: a before window ending the day before the change went live, and an after window starting a few days later. The gap is deliberate. A change does not take effect the moment it is applied — it takes effect when the page is re-crawled and re-processed — and including those in-between days would mix pre-change and post-change behaviour into the "after" number.
We run the comparison twice, at two horizons:
- d14 — fourteen-day windows
- An early read. Enough to catch something going badly wrong, rarely enough to be confident about a modest gain.
- d28 — twenty-eight-day windows
- The one we actually lean on. Slower, less noisy, and long enough that a single unusual week does not decide the answer.
Both are stored. We do not overwrite the early verdict with the later one, because the disagreements are informative: a change that reads positive at d14 and flat at d28 usually caught a novelty bump rather than a durable improvement, and you only see that if both numbers survive.
The four answers
A verdict is one of four values, and the fourth is the important one.
- Helped
- Click-through rate rose by at least 15% relative, or average position improved by at least 1 rank without impressions collapsing alongside it.
- Hurt
- The mirror image of the same two tests.
- Flat
- Neither threshold was crossed, or the two signals disagreed with each other. A change that improved position while CTR fell has not clearly done anything, and saying so is more useful than picking the flattering half.
- We couldn't measure it
- The two windows together carry fewer than 50 impressions, or one of them has no data at all.
That last one is a real answer, not a failure state, and we show it as one. A page with almost no search impressions cannot produce a statistically meaningful before-and-after; any verdict computed from it would be a coin flip wearing a label. The honest output is "there was not enough signal here to judge", and a system that never emits that verdict is a system that is guessing on its quietest pages and telling you it isn't.
It is also, in practice, the most common verdict on a small site, and we would rather say so up front than have it arrive as a disappointment. If most of your pages sit below the impression floor, the useful work is probably not meta description tuning — it is having pages that people search for.
The guards, and why they exist
Two of the thresholds above carry conditions that look fussy and are not.
The position signal is guarded by impression volume. Average position is an average over impressions, so it moves mechanically when the mix of queries changes: a page that suddenly picks up a long tail of low-ranking impressions will show its average position get worse while nothing about the page got worse. So a position gain only counts as helped if impressions did not drop more than 20% alongside it, and a position loss only counts as hurt if impressions did not surge by the same margin. Without that guard, the most common "hurt" verdict would be a page that started ranking for more things.
Clicks and impressions are compared as per-day rates rather than window totals, because the two windows are rarely equally populated — Search Console data arrives late and unevenly, and comparing a fourteen-day total against an eleven-day total would manufacture a decline out of nothing.
What a verdict is not
It is not causation, and we label it as correlation everywhere it appears. Two windows of search data around an edit is evidence. Between those windows the engine also updated its systems, your competitors also changed their pages, the season also moved, and demand for your subject also drifted. We control for none of that, and neither does anyone else with access to the same data.
What the verdict does give you is a discipline. It makes the claim falsifiable, it makes it dated, and it makes the record cumulative — after thirty changes you have thirty verdicts, and the pattern across them says considerably more than any single one. It also means a change that made things worse gets found by a scheduled process rather than by someone eventually noticing.
That is the whole bet: a measured verdict that is often inconclusive is worth more than a confident story that was never checked.