Almost everything written about SEO and GEO tactics is argued from experience rather than from records. Someone fixed a thing, the numbers went up a few weeks later, and the story gets told as cause and effect. The advice may well be right. The problem is that nobody publishes the ledger, so nobody can check it.
We happen to keep one. ConceptSEO audits a portfolio of live sites every week and stores, for each site and each week, every recommendation it generated, whether that recommendation was actually completed, and the dimension score recorded on the next scan. That is a longitudinal intervention log rather than a survey or a scrape of public search results. This page reports what is in it.
Every figure below comes from that log. Where a number could not be computed we left the claim out rather than softening it, and the sample size sits beside each result so you can judge the weight yourself. The parts that did not work are here too, in the same detail as the parts that did. There is a separate piece on content produced at volume: what did not work, and why.
Every dimension in the ledger, and the one that reverses. Figures as tabulated below.
What is in the ledger?
The platform runs a set of independent weekly audits against each site it manages. Each audit covers one dimension and returns a score out of 100 along with a list of recommendations. Those recommendations then get worked, and their status is recorded when they are finished. Because the audits repeat on a weekly cadence, each site accumulates a run of scores per dimension, and each score sits at a known point in time relative to the work completed before it.
That structure is what makes the question answerable. For any two consecutive scans of the same dimension on the same site, we know the score at the start, the score at the end, and exactly how many recommendations belonging to that dimension were completed in between. Group every interval by whether it contained completed work, and the comparison falls out.
The sample as it stands:
- 11 sites receiving weekly audits
- 2026-03-09 to 2026-08-03, a window of 21 weeks
- 1,048 completed audits carrying a score
- 2,331 recommendations generated, of which 2,035 were completed
- 966 consecutive-scan intervals usable for the comparison, 563 containing completed work and 403 containing none
The portfolio is a mix: several client sites, a handful of reference sites we own, and our own marketing pages. Client domains are described rather than named unless a case study for that client is already public here.
What the audits asked for
Recommendations are tagged by category when they are generated, and the split is lopsided enough to be worth stating plainly before we get anywhere near outcomes.
| Category | Generated | Completed | Completion rate |
|---|---|---|---|
| Technical | 851 | 719 | 84% |
| Content | 773 | 729 | 94% |
| SEO | 464 | 409 | 88% |
| Internal links | 180 | 138 | 77% |
| UX | 14 | 14 | 100% |
| Strategy | 14 | 7 | 50% |
| Local | 14 | 7 | 50% |
| E-E-A-T | 12 | 11 | 92% |
| Paid search | 8 | 0 | 0% |
| Authority | 1 | 1 | 100% |
| Total | 2,331 | 2,035 | 87% |
Three categories account for 89% of everything generated. The paid-search row is the honest embarrassment in the table: eight recommendations produced, none completed, because the portfolio has no meaningful ad spend to act on and the audit kept generating them anyway.
Did completing the work move the score?
Each row below covers one dimension. The left pair is intervals where at least one recommendation from that dimension was completed; the right pair is intervals where none were. Mean is the average score change from one scan to the next, in points out of 100. Improved is the share of intervals where the score went up at all.
| Dimension | Intervals with work | Mean change | Improved | Intervals with none | Mean change | Improved |
|---|---|---|---|---|---|---|
| Page speed | 57 | +3.14 | 56% | 82 | −0.52 | 39% |
| GEO / AI visibility | 80 | +2.15 | 55% | 53 | +1.13 | 34% |
| Images | 75 | +1.96 | 72% | 58 | +0.72 | 36% |
| Technical SEO | 95 | +1.94 | 44% | 57 | +0.33 | 46% |
| Schema | 77 | +1.84 | 57% | 64 | +0.34 | 20% |
| Content | 108 | +1.63 | 47% | 28 | −0.04 | 39% |
| Internal links | 58 | +1.52 | 47% | 27 | +1.74 | 26% |
| All dimensions | 563 | +2.01 | 53% | 403 | +0.51 | 34% |
Three dimensions are missing from that table because the sample is too thin to say anything: local search had only 13 intervals containing completed work, and paid search and the combined weekly review had none at all. Reporting a mean on 13 intervals next to one built on 108 would give both the same visual weight, so they are excluded and named here instead.
Page speed separates hardest
Page speed produces the widest gap in the ledger: +3.14 points when work landed, −0.52 when it did not. It is also the only dimension where the no-work average is meaningfully negative, which fits how performance behaves. Left alone, a site gets slower. Images get added, a script gets embedded, a third-party tag arrives. The score decays rather than holding steady, so the comparison is not really "work versus nothing" but "work versus drift".
Schema shows the same story from a different angle. The mean gap is smaller (+1.84 against +0.34), but the frequency gap is the largest in the table: scores improved in 57% of intervals with completed work against 20% without. Structured data does not rot on its own the way performance does, so when the schema score moves, something moved it.
Technical SEO is not what it looks like
The technical row is the strangest in the table and the easiest to misread. The mean change is clearly better with work (+1.94 against +0.33), yet scores improved slightly less often with work than without: 44% against 46%. The averages and the frequencies point in opposite directions.
The reading that fits the data is that technical work is lumpy. Most technical recommendations are small and change the score by nothing at all, while a few structural fixes move it a long way at once. Averaging rewards the rare large win; counting improvements does not. Judge technical work by how often the number ticks up and you'd conclude it does nothing, and you'd be wrong for a purely arithmetic reason.
How long before the score moved?
The tables above answer whether the work moved anything. They do not answer the question a prospect actually asks first, which is how many weeks before you see it. So we ran the same records again with a clock on them.
The construction: every day on which at least one recommendation was completed for a site in a given dimension counts as one intervention, whatever the batch size, because a release of ten items shipped together is one event and not ten. The baseline is that dimension's score at the last scan on or before the completion date. The clock stops at the first later scan whose score sits at least 2 points above that baseline and is still there at the following scan. Interventions with fewer than 8 weeks of scans after them are dropped, so a null reads as "did not move" rather than "not observed yet".
Two points on that definition, because it changes the answer more than anything else here. A 2-point threshold rather than any upward tick: with weekly scans and scores that wobble, "first scan that is higher at all" returns a median of 0.9 weeks for almost everything, which is measuring noise. And the sustained requirement discards a spike that falls back the following week.
| Dimension | Interventions | Reached +2 and held | 25th pct | Median weeks | 75th pct |
|---|---|---|---|---|---|
| Content | 49 | 59% | 0.8 | 0.9 | 1.9 |
| GEO / AI visibility | 42 | 74% | 0.8 | 0.9 | 2.9 |
| Images | 35 | 91% | 0.8 | 1.9 | 3.2 |
| Page speed | 33 | 79% | 0.9 | 2.8 | 7.2 |
| Schema | 39 | 87% | 0.9 | 2.9 | 10.6 |
| Technical SEO | 39 | 74% | 0.9 | 3.5 | 8.1 |
| All dimensions | 237 | 76% | 0.8 | 1.9 | 5.1 |
| No work shipped (control) | 90 | 61% | 1.0 | 1.0 | 3.5 |
Across 237 interventions, the median time from shipping a fix to a sustained 2-point gain in that dimension was 1.9 weeks, with a quarter landing inside 0.8 weeks and a quarter taking longer than 5.1 weeks. Three dimensions are absent because they fall under the 20-intervention floor this page uses throughout: internal links (14), local search (7) and the combined weekly review (1).
The spread between dimensions is larger than the middle. Content and GEO reach the threshold at the next weekly scan, median 0.9 weeks. Schema and technical SEO take three to four weeks at the median, and their 75th percentiles run to 10.6 and 8.1 weeks. If you ship a schema fix and check a fortnight later, you are inside the window where a quarter of our schema interventions still show nothing.
Read the control row before you use any of this
The bottom row is the same clock started on scans where nothing shipped. Those baselines reached a sustained 2-point gain 61% of the time against 76% with work, so the honest gap is in how often a dimension improves, not how fast. The control's median is 1.0 weeks, nominally quicker than the 1.9 with work, because scores drift upward on their own often enough that the fastest-moving cases are not the ones we caused.
This is the same limitation as the rest of the page, and it bites harder here. Nobody randomized these weeks. Read the medians as a planning expectation of when a change tends to show up, not as proof the change caused it. The share column is the one carrying the signal.
Why this is faster than the lag we publish elsewhere
We describe a 4 to 6 week Health Score lag in our own material, and a 1.9 week median looks like it contradicts that. It does not, because the two measure different things, and the difference is worth being precise about.
The figures above are dimension audit scores. Those measure whether the fix is present and correct on the page, so they update as soon as the page is scanned again. The Health Score is 60% those foundations and 40% a performance component built from search and analytics signals, and that half moves at the speed of crawling, indexing and ranking rather than the speed of deployment. A fix can be visible in its dimension score within a fortnight and still be weeks away from showing up in the blended number.
So both figures stand, and neither is the one to quote on its own. If you want to know when you can confirm the work landed, use the table above. If you want to know when the composite starts reflecting it, the longer figure is the right one.
What did not work?
The internal-links row is the one we'd leave out if this were a sales page. It's also the most useful line in the ledger.
Two individual cases show the same thing at the level of a single site, which is harder to wave away than an average.
On our own marketing site at concept211.com, 11 internal-link recommendations were completed over four weeks. The dimension scored 78 on July 6, 78 on July 13, 78 on July 20, 78 on July 27, and 79 on August 3. Eleven completed items, one point.
On a reference site we run, asterpedia.com, 14 internal-link recommendations were completed across the same period. The dimension scored 80 on every one of the six weekly scans between June 29 and August 3. Fourteen completed items, no movement whatsoever.
We're not concluding that internal linking is worthless. There are at least three explanations we can't separate with this data, and honesty requires listing them rather than picking the flattering one.
- The scoring may be saturated. Both sites were already at 78 to 80 before the work started, and a dimension near its ceiling has little room to register anything. The audit may simply stop rewarding additional links past some threshold.
- The work may be real but invisible to this particular metric. Internal links plausibly act on crawl paths and on how authority distributes across a site, and those effects would surface in rankings and impressions rather than in an on-page audit score.
- Or the work may genuinely not matter at this margin. Adding the twelfth contextual link to a site that already has decent structure may be worth approximately nothing.
What we can say without hedging is narrower and still worth publishing: on this portfolio, in this window, completing internal-link recommendations did not predict a rise in the internal-links score, while completing work in six other dimensions did. If your reporting counts completed internal-link tasks as progress, that assumption is not supported here.
How much weight should you put on this?
This is observational data, not an experiment. Nobody randomized which weeks received work. That matters more than any single number in the tables, because the weeks with completions differ from the weeks without in ways beyond the completions themselves.
The most likely confound runs in a specific direction. A dimension scoring badly generates more recommendations, gets more attention, and has more headroom to improve. A dimension already scoring 95 generates little, receives little, and cannot move much regardless. So some of the gap between the two columns is regression toward the mean rather than the effect of the work, and the honest position is that the ledger shows an association of a certain size, not a proven causal effect of that size.
Two things push back the other way. The association survives inside each dimension separately rather than only in aggregate, so it is not an artifact of one dimension dominating the pool. And the internal-links result shows the method is capable of returning a null, which is the main thing you want to know about any measurement that keeps agreeing with the person running it.
What we changed because of this
Three of our own habits changed because of the numbers above rather than because of anyone's opinion.
We stopped counting completed internal-link items as reportable progress. If the score doesn't move, the client shouldn't be shown a number that says it did.
We stopped judging technical work by how often the score ticks up. That 44%-against-46% frequency split would have retired a category the averages show is working, and it would have been the measure at fault rather than the work.
And we stopped generating paid-search recommendations for sites with no ad spend. Eight generated against zero completed is a generator producing work nobody could act on.
Republish the results table
Every heading and every headline figure on this page carries a stable id, so a single statistic can be linked directly rather than by pointing at the article and hoping. As an example, the internal-links null result is at #finding-internal-links and the two site-level cases are at #stat-null-agency-site and #stat-null-reference-site.
The main results table is below as plain HTML with the source line already attached. Paste it as it stands.
Methodology. Every figure is computed from the ConceptSEO platform's own audit records. No figure is modelled, estimated, or drawn from any external dataset.
Sample. 11 sites under weekly audit; 1,048 scored audits; 2,331 recommendations generated and 2,035 completed; 966 consecutive-scan intervals compared, of which 563 contained completed work and 403 did not.
Window. March 9, 2026 to August 3, 2026, a span of 21 weeks.
Construction. A dimension is defined by the audit type that generated the recommendation, so a recommendation is only ever credited against the score it was raised from. An interval is a pair of consecutive scored audits of the same dimension on the same site. Intervals longer than 21 days are excluded, so a gap in coverage is never read as a week of inactivity. A recommendation counts toward an interval when its completion timestamp falls inside it. Dimensions with fewer than 20 intervals in either group are excluded from the results table and named in the text.
Time-to-effect construction. Computed from the same audit records, extended to August 10, 2026. One intervention is one site-dimension-day on which at least one recommendation was completed, so a batch shipped together counts once. Baseline is the dimension's score at the last scan on or before that day. The clock stops at the first later scan at least 2 points above baseline that is still above it at the following scan; a spike that falls back does not count. Interventions with under 8 weeks of subsequent scans are excluded so an unresolved case reads as no movement rather than no observation. The control row applies the identical clock to scans with no completed work within 7 days. 237 interventions qualified, against a 90-scan control.
Last updated. August 10, 2026 for the time-to-effect section; August 3, 2026 for the sections above it. The window is open and these figures will move, so the date above is the version you are reading.
Reuse. You may republish these figures with attribution and a link to https://seo.concept211.com/resources/geo-ranking-ledger.