SEO results from the playbook
Results are credible only when the baseline, timeframe, intervention, evidence, and limits travel with the number. Explore published outcomes, how we measure them, and the use cases they can—and cannot—support.
How we measure these results
A result is a defined change from a recorded baseline during a stated observation window, supported by evidence we can reproduce. We separate visibility signals from clicks, qualified actions, and commercial outcomes.
- ✓Baseline — use the period before the intervention, a matched group, or another named comparator; never an unlabeled high-water mark.
- ✓Attribution window — state when work began, when measurement ended, and the lag expected before crawling, ranking, or citation systems respond.
- ✓Exclusions — remove or disclose paid traffic, brand campaigns, migrations, seasonality, tracking changes, and other events that can distort the comparison.
- ✓Evidence — retain exports or screenshots, query and page scopes, filters, dates, and calculation steps. Read how we measure results before interpreting any card below.
How we measure results
These figures come from the two published case studies currently available. Different units describe different parts of the search journey, so they belong in separate tiles rather than a single blended score.
What these numbers do and do not prove
The HZ-Containers evidence includes a 16-month Google Search Console view, an Ahrefs trend, the traffic-growth baseline, and the documented one-person operating model. The TarmacView evidence includes its Search Console and Ahrefs views for a new domain over roughly eight months.
What every result includes
A reader should be able to understand the comparison without asking what the chart leaves out. That is why every publishable result carries its scope, baseline, intervention, window, and evidence as part of the claim itself.
Results by business type
Business type changes the unit of analysis. A catalog needs page-cohort and product-discovery evidence; a consultancy needs qualified-demand and buying-committee evidence. Use these facets to choose the right measurement design, not to borrow another company's target.
Results by playbook component
“We did SEO” is not an intervention. Record whether the team launched a page cohort, retrofitted an element, or completed a process phase, then choose a comparison that matches that unit of change.
| Component changed | Why it can matter | Measurement design | Decision the result supports |
|---|---|---|---|
| Post-type rollout | A repeated document shape lets a team satisfy the same class of intent consistently across a page cohort. | Freeze the launch list; compare cohort-level indexation, query coverage, citations, clicks, engagement, and qualified actions with the baseline and an unchanged group where possible. | Continue the rollout, revise the specification, narrow the eligible cohort, or stop. |
| Element retrofit | An element can remove a specific comprehension or extraction problem without requiring a full rewrite. | Record changed URLs and dates. Compare the same pages before and after adding features such as comparison tables or direct answer blocks, while noting other edits. | Standardize the element, change its acceptance rules, or reserve it for contexts where it resolves a real reader need. |
| Process phase | Publishing more cannot compensate for pages that crawlers cannot access, duplicate architecture, or work aimed at the wrong demand. | Capture defect counts and affected URLs before and after the phase, then observe downstream indexation and visibility only after the expected recrawl lag. | Clear the production gate, continue remediation, or investigate a different constraint before creating more pages. |
For process work, start with the technical baseline audit and AI accessibility audit. For production, use keyword and prompt research to define demand and a topical map to prevent overlap before pages are commissioned.
Case studies
Each card names the business model, starting point, work applied, timeframe, and measured outcome. The component links describe the closest playbook specifications; they do not assign isolated causal credit.
When publishing the next result, use the case study specification. It requires a sourced baseline, timed intervention, evidence ledger, attribution limits, client approval, and a next step that fits the reader's decision stage.
Anonymized use cases
The scenarios below describe implementation and measurement designs for situations where a client name or outcome cannot be published. They are not case studies, benchmarks, or performance claims; no numerical gain should be inferred from them.
Turn a result into your next decision
The purpose of measurement is not to decorate a report. Before work begins, define what each possible pattern will trigger: scale, revise, wait, investigate, or stop.
Hold agent-written content to the same evidence bar
Content produced by an agent is not automatically good, and saying otherwise would undermine every number on this page. It is measured exactly like everything else.
An agent-produced page gets a baseline before publication, a stated comparison window, and the same honest account of what else moved in that period.
Nothing about the authoring method changes the measurement contract. If anything it raises the bar, because volume makes it easier to mistake activity for progress.
Results here are tied to the part of the system that produced them — a post type rolled out, an element retrofitted across an existing corpus, a phase run in the right order.
That is a more useful claim than crediting automation in general, and it is one you can actually test on your own site.
Because elements are typed, you can measure the corpus directly: share of pages carrying a sources block, spec conformance by post type, element coverage, freshness distribution.
These move before traffic does, which makes them the earliest honest signal that a content system is working.
A win from a single page presented as a sitewide result. A window chosen after the fact because it looked good. A metric selected because it moved.
Those exclusions apply identically to human and agent output, and stating them is the reason the rest of the figures are worth reading.
Read the measurement rules before the outcomes. A result is only useful if you can reconstruct how it was produced and decide whether the same approach would work on your site.
Start with your baseline
Define the scope, source, filters, observation window, exclusions, and decision rule. Then run the visibility check and preserve the first export. A result becomes defensible when another person can reconstruct it.
Establish the baseline before you change the site
Free check · 7-day trial · no credit card