How We Measure SEO and AI Visibility Results
Use this SEO results measurement methodology to define baselines, comparison windows, attribution, AI visibility metrics, rounding, and publication rules.
A result is a measured change against a dated starting point, calculated with a definition chosen before the work began and interpreted in the context of everything else that changed. It is not a favorable number discovered after delivery. This methodology is the contract behind every result we publish: a reader should be able to take the same source data, apply the stated formula and reach the same number.
That standard matters because search performance is an observational system. Pages change, competitors react, demand moves, search engines update, and marketing campaigns overlap. We can prove what we shipped and when. We can measure what happened next. We can claim that our work caused the entire movement only when the research design supports that claim.
The measurement record every result must include
Every published result must provide enough information to be audited. At minimum, its record contains:
- the business objective and the metric selected before implementation;
- the metric formula, source, timezone, currency where relevant, and attribution model;
- the site, directory, template, page cohort, market, device, engine, and prompt scope;
- baseline and comparison dates, including the number of complete days in each window;
- the absolute starting value, absolute final value, and calculated change;
- implementation dates and the exact work included in the claim;
- material annotations and known competing explanations;
- missing data, tracking changes, exclusions, and sample or denominator size;
- the strength of the attribution language we are entitled to use; and
- the reviewer who approved the client-facing figures.
If one of these fields is unavailable, we either disclose the limitation beside the result or do not publish it. A private spreadsheet containing the caveat does not repair an unqualified public claim.
1. What counts as a result
The purpose of measurement is to show whether the work changed something valuable. We organize metrics as a chain, because each stage answers a different question:
- Delivery: Was the intended change published correctly? Examples include pages released, templates updated, redirects resolved, and crawler access restored.
- Discovery and visibility: Did search or AI systems find and surface the work? Examples include valid indexed pages, impressions, ranking distribution, AI visibility score, mentions, and citations.
- Selection and engagement: Did people or answer engines select it? Examples include organic clicks, qualified sessions, cited URLs, engaged visits, and visits to a next-step page.
- Business outcome: Did that behavior create value? Examples include qualified leads, trials, purchases, subscription revenue, booked appointments, or another outcome agreed with the client before work began.
Delivery is evidence that the work happened. Visibility is a leading indicator: a measure expected to move before a business outcome. Engagement shows response. The business outcome is the result most organizations ultimately fund.
Traffic alone is therefore not a result for most business models. A SaaS company does not benefit merely because 20,000 extra visitors read an irrelevant article; it benefits when suitable prospects start or influence trials, demos, and subscriptions. An ecommerce business needs profitable orders, not sessions detached from revenue and margin. A publisher may legitimately treat qualified page views as a business outcome when those views produce advertising or subscription value, but even then bot traffic, accidental visits, and audience geography can change the value of the same visit count.
We report traffic when it is the metric’s proper place in the chain, and we pair it with quality or downstream evidence whenever the business model permits. We refuse to present impressions, a single ranking, page count, raw traffic, engagement rate, or a proprietary score as final business impact when no validated connection has been shown.
The outcome is selected in advance. If the agreed primary outcome is qualified demo requests, we do not replace it after delivery with impressions simply because impressions rose and demo requests did not.
2. Baselines: freeze the starting point
A baseline is a dated, immutable measurement of the relevant scope before the first in-scope change ships. The P5 baseline measurement phase explains how to capture it. The baseline must use the same formula, sources, filters, segments, timezone, currency, attribution model, prompt set, competitor set, and exclusions that the later checkpoint will use.
For rate metrics, we save the numerator and denominator, not only the rate. For sitewide totals, we also save the affected cohort. If 40 product pages are being optimized, their baseline remains separate from the other 8,000 URLs; otherwise unrelated site growth can disguise what happened to the work.
A retroactively reconstructed baseline is not a baseline. Once the team knows the outcome, choices about dates, filters, segments, and exclusions are vulnerable to hindsight. Live tools may also have changed their processing, retention, attribution, or competitor data. Calling a reconstruction a baseline would imply a level of control that did not exist.
When a client has no clean history, we use one of three honest alternatives:
- Prospective baseline: connect and validate measurement, make no in-scope change during a defined observation period, then freeze that period before implementation.
- Reconstructed reference: export the best available history, document source and tracking changes, label it “reconstructed,” and restrict the claim to what the surviving data supports.
- No before-and-after claim: begin measurement now and report absolute checkpoint performance until a valid prospective comparison exists.
Missing data is unknown, not zero. We do not interpolate it merely to create a clean chart. If the missing interval overlaps a material change or makes the comparison non-equivalent, the result is not publishable as a before-and-after claim.
3. Comparison windows and the reporting clock
The two windows must contain the same number of complete days, use the same day-of-week mix, and exclude the incomplete current day. We state both ranges in the result rather than using labels such as “last month” that change meaning over time.
Year-over-year comparison compares a window with the equivalent calendar window one year earlier. It is our default when demand has seasonal patterns and the earlier data is clean. Fixed holidays and moving holidays are annotated; when a holiday shifts between windows and materially affects demand, we align trading weeks or disclose the mismatch.
Period-over-period comparison compares adjacent equal-length windows, such as 28 complete days after implementation with the preceding 28 complete days. We use it when the site lacks a clean prior year, when the business is not materially seasonal, or when we need a recent operational signal. It is more exposed to campaigns, launches, holidays, and market changes, so we never present it as seasonally controlled.
Where sufficient history exists, we show both. Year over year answers “are we ahead of the equivalent season?” Period over period answers “did the recent direction change?” If they disagree, the disagreement is information to explain, not a reason to hide one.
No performance movement is reportable as a result before 28 complete days have elapsed after the relevant change. Earlier observations may verify that a page is live, crawlable, indexed, or cited, but they are implementation checks rather than results. Twenty-eight days is a floor, not an automatic approval. We extend the window when:
- the normal purchase or lead cycle is longer than 28 days;
- the change took time to reach all pages or be processed by search and AI systems;
- weekly volume is too small for a few events to stop dominating the rate;
- the business has strong monthly, quarterly, or annual seasonality; or
- an outage, campaign, migration, tracking change, or major external event contaminates much of the window.
We pre-register the first checkpoint and the intended reporting window. We do not keep moving the start or end dates until the curve looks favorable.
4. Attribution honesty
Attribution is the assignment of an observed effect to a cause. Timing alone does not establish it: “after” does not necessarily mean “because of.” We use four levels of claim.
Verified delivery means we can attribute the implemented change itself to the work: for example, 120 specified pages received valid canonical tags on the logged release date. It does not establish the performance effect.
Experiment-supported effect requires a design capable of estimating an incremental effect, such as a randomized or credibly matched holdout with a pre-specified hypothesis, metric, sample, and end date. We describe the design and uncertainty; we do not convert an imperfect content experiment into laboratory certainty.
Contributory evidence applies when the affected cohort moves after the logged change, the direction matches the pre-specified hypothesis, and an unchanged reference cohort or competitive benchmark does not show the same movement. We may say the result is “consistent with,” “associated with,” or “supports the contribution of” the work.
Correlation only applies when the metric moved after the change but no valid counterfactual exists. We report the sequence and the competing explanations. We do not say “generated,” “caused,” or “delivered” for the outcome.
In ordinary client work, we frequently cannot separate the effect of:
- search engine or AI provider algorithm updates;
- seasonality, holidays, weather, and changes in underlying demand;
- paid search, paid social, retargeting, and affiliate activity;
- PR coverage, creator activity, brand campaigns, and offline promotion;
- product launches, pricing, availability, merchandising, and conversion-rate changes;
- migrations, redesigns, tracking changes, consent changes, and outages;
- competitor launches, losses, promotions, and changes in media spend; or
- economic, regulatory, cultural, and category-level market shifts.
We name the factors that occurred, say whether each probably affects the measured cohort, and lower the strength of the claim accordingly. If paid and organic land on the same page during a launch, we do not assign all conversions to organic content. If an algorithm update overlaps the checkpoint, we can report relative movement against unaffected segments or competitors, but that comparison is evidence, not perfect isolation.
5. Annotations make interpretation possible
An annotation is a dated record of a change or external event placed on the same timeline as the metric. Without annotations, attribution becomes a memory exercise months later.
We log the event when it happens, not when reporting begins. Each annotation records the date and timezone, owner, affected URLs or segment, category, description, deployment or campaign reference, expected metric and direction, expected delay, checkpoint date, and known overlapping activity. Multi-day releases have start and completion dates. Reversals and rollbacks receive their own entries rather than overwriting the original.
We annotate content releases, technical deployments, migrations, redirects, template changes, analytics and consent changes, outages, paid campaigns, PR, product and pricing changes, major competitor events, and confirmed platform updates. Annotation Outcomes keeps those changes, expectations, checkpoints, and observed outcomes together while explicitly treating association as evidence rather than proof of causation.
6. How we measure AI visibility
AI visibility metrics are calculated separately because they answer different questions. We never blend visibility, voice, citations, sentiment, and rank into an estimated “AI impact” number.
Before comparing periods, we freeze the prompt text, topic, locale, engine, run cadence, monitored brand set, and monitored domain set. A prompt observation is one stored answer from one defined prompt, engine, locale, and scheduled run. The answer and its cited URLs are retained so the aggregate can be audited.
AI visibility score
The visibility score answers: “For how much of the tracked prompt set did this domain earn at least one citation?”
AI visibility score = prompt observations citing the domain at least once ÷ all valid prompt observations in the frozen cohort × 100
Multiple citations in one answer do not make that observation more visible; it counts once in the numerator. A collection failure is not a zero-citation answer. It is missing data, reported through coverage and rerun or excluded consistently from both windows. The AI Visibility view exposes the underlying answers needed to inspect the score.
Share of voice
Share of voice answers: “Of the monitored brands mentioned in these answers, what share belonged to this brand?”
Share of voice = mentions of the brand ÷ mentions of every monitored brand in the fixed competitive set × 100
We state whether the counting unit is one mention per answer or every mention occurrence and keep that rule unchanged. We do not renormalize the history when a competitor is added; a competitor-set change starts a new comparable series or causes the earlier window to be recalculated from stored answers.
Citation share
Citation share answers: “Of the citations assigned to the monitored domain set, what share pointed to this domain?”
Citation share = citations to the domain ÷ citations to every monitored domain in the fixed set × 100
The numerator counts citations, not brand mentions. One answer can cite several URLs from the same domain, and each citation counts under the declared rule. Domain-level reporting must not be presented as page-level performance. Source & Citation Intelligence retains the domain, URL, prompt, and engine detail behind the aggregate.
These metrics have known limitations. AI answers can vary between runs; providers change models, retrieval systems, interfaces, and citation behavior; location, language, personalization, and freshness can alter an answer; and the chosen prompt and competitor sets define the measured universe. A tracked set is a repeatable sample of the questions we chose, not a census of everything potential customers ask. Citation presence also does not prove that a person saw, trusted, or visited the source. We therefore report engine, locale, cohort, run frequency, coverage, and absolute counts with every material AI visibility claim.
7. Rounding, ranges, and absolute values
Percentages are never published alone. Every percentage appears with the absolute values used to calculate it.
For a count change, we report the baseline, final value, absolute difference, and relative change: “qualified organic leads increased from 40 to 52, up 12 leads or 30%.” For a rate, we report both rates in percentage points and include their numerators and denominators: “conversion rate moved from 2.0% (40/2,000) to 2.4% (52/2,167), up 0.4 percentage points.” Calling that a 0.4% increase would be wrong; the relative increase is 20%.
Counts are whole numbers. Percentages are normally rounded to one decimal place, except when that would turn a non-zero value into 0.0%, in which case we add a second decimal. We calculate changes from unrounded inputs and round only the displayed result. Values under 1,000 are not abbreviated. Larger abbreviations may appear in headings, but the exact number remains in the body or accessible table.
We use ranges only when the source or method genuinely produces uncertainty. The range must name its method, such as a confidence interval, attribution-model range, or data-coverage bound. We do not invent a range to make a point estimate look scientific, and we do not replace small exact counts with percentages that exaggerate movement.
8. Client anonymization without removing meaning
An anonymous result must remain testable. We can replace the client name, domain, product names, and commercially sensitive figures, but the published account must preserve:
- business model, broad category, market, and material constraints;
- site or page-cohort scope and the kind of pages changed;
- exact baseline and comparison dates and window lengths;
- metric definition, source, attribution model, and denominator;
- absolute baseline and final values, or an explicitly indexed series when contractual confidentiality prohibits absolutes;
- the intervention, release dates, and elapsed time before measurement;
- known concurrent activity and the resulting attribution limitation; and
- enough method detail for another team to apply the calculation.
If exact commercial values cannot be shown, an index may set the baseline to 100 and apply the same scale to every later point. We label it as indexed and still show the underlying event or sample counts where disclosure permits. We never combine several clients into a fictional “typical client,” alter dates to strengthen the curve, or conceal that the result came from one directory rather than the whole site.
If preserving this context would identify the client, we do not publish the result. Confidentiality is a reason to withhold a case, not permission to publish an impressive but unverifiable percentage.
9. What we will not publish
Publication is a selection decision, so its rules have to be as strict as its calculations. We will not publish:
- a cherry-picked start date, end date, search engine, country, device, prompt, keyword, or segment;
- a single-page win described as sitewide growth;
- a metric selected after the fact because it moved favorably;
- a percentage without its baseline, final absolute value, and denominator where applicable;
- a comparison that mixes different formulas, scopes, sources, attribution models, currencies, or prompt cohorts;
- a partial month compared with a full month, or unequal windows presented as equivalent;
- missing observations treated as zero, or tracking breaks hidden inside a continuous chart;
- a result reported before the minimum window or before the normal conversion cycle completes;
- an AI score presented without its prompt, engine, locale, competitor, and coverage definitions;
- a correlational result written as proof that our work alone caused the outcome;
- only the best page, query, engine, or checkpoint when the full declared cohort is available; or
- a client result that the authorized reviewer has not approved for accuracy, context, and disclosure.
We also do not delete unfavorable checkpoints from a series. A hypothesis that missed can still teach the team what to change. Hiding it would make the later success less credible, not more.
Publication sign-off
Before publication, the person accountable for client-facing numbers must reproduce the calculation from the frozen exports, confirm the comparison scope and dates, inspect the annotation ledger, verify every absolute and percentage, approve the attribution wording, and confirm that anonymization complies with the client agreement. A second editorial check must ensure the headline says no more than the evidence.
The final question is simple: could a skeptical reader identify what changed, what did not, what else might explain the movement, and how every number was calculated? If not, the result is not ready to publish.
Ready to put it into practice?
Free check · 7-day trial · no credit card