Technical SEO Audit: Crawl and Indexation
Run a technical baseline audit that finds crawl, indexation, canonical, rendering, and internal-link problems before you invest in new SEO content at scale.
Technical baseline audit
Phase P2 · Stage A — Understand
Timebox: 2–4 hours for a light pass, 1–2 working days for a standard pass, or 3–8 working days for a deep pass.
Owner: the technical SEO lead. Engineering, analytics, content, and localization owners contribute evidence and accept fixes in their areas.
A technical baseline audit establishes whether search engines can reach, interpret, and select the URLs the business expects them to show. Its scope covers crawl controls, HTTP responses, indexation, canonicals, links, rendering, international targeting, and secure delivery. The result is a prioritized findings register with named owners and acceptance tests, not a score.
Why this phase comes here
Publishing onto a site with crawl or indexation problems compounds the damage. Search engines often discover a repeated defect across new URLs faster than they evaluate and reward the content. A broken canonical template can point every article elsewhere; a robots rule can hide a directory; client-rendered navigation can create orphans for a non-JavaScript client. Every new page enlarges the affected set and makes repair riskier.
Fix the foundation first. The order is crawlability → indexability → content quality → performance because each layer is a gate. Crawlability means a crawler can discover and request a URL; indexability means the reachable URL is eligible for inclusion. Only then should content quality and performance be judged. A fast page blocked by robots.txt cannot compete, and title tags do not matter on unreachable pages.
This phase consumes the scope, priority journeys, markets, and risks from Discovery and goals . Running it earlier produces a crawl with no business context. Skipping it lets research and production target templates that cannot reliably enter the index.
Inputs and outputs
Inputs define the intended site, not merely what a crawler finds. Outputs tell the next owner which URLs are safe to test and which remain blocked.
| Direction | Item | Acceptance condition |
|---|---|---|
| Input | Production origins and canonical host | Includes protocol, www decision, subdomains, international hosts, and known legacy domains. |
| Input | Intended indexable URL inventory | Lists templates, directories, locales, sitemap sources, and exclusions such as filters, account pages, and internal search. |
| Input | Access and evidence | Production crawl permission, Google Search Console, Bing Webmaster Tools, analytics, log files when available, deployment history, and CMS rules. |
| Input | Discovery brief | Names priority journeys, revenue or lead value, markets, launch constraints, and accountable owners. |
| Input | Recent change register | Records migrations, redesigns, JavaScript framework changes, canonical or pagination changes, incidents, and release dates. |
| Output | Prioritized findings register | Every finding has affected scope, evidence, root cause, impact, effort estimate, confidence, owner, deadline, and done-when test. |
| Output | Crawl and index baseline | Records eligible URLs, crawled URLs, status distribution, sitemap coverage, indexation ratio, orphan count, and depth distribution. |
| Output | Blocking-dependency decision | States whether publishing may proceed, proceed only for unaffected templates, or pause until named blockers pass retest. |
| Output | Handoff package | Gives the next phase a clean URL sample, unresolved exclusions, rendering evidence, and accepted limitations. |
Choose the audit depth
Select depth before crawling. Estimates assume access is ready and exclude implementation.
| Mode | Choose it when | Honest timebox | Coverage and limitations |
|---|---|---|---|
| Light | Under roughly 500 indexable URLs, one main template and language, no recent migration, and no JavaScript-dependent primary content | 2–4 hours | Controls, sitemaps, responses, representative crawl, priority inspection, basic canonicals, and mobile/HTTPS samples. It can miss long-tail orphans, rare loops, near-duplicates, template-specific rendering failures, and hreflang defects. It is triage, not migration assurance. |
| Standard | Up to roughly 50,000 intended URLs, several templates, routine JavaScript, or a substantial content program | 1–2 working days | Full crawl, sitemap reconciliation, sampled inspection, duplicates, depth, rendering, and template rules. This is the default for an established site. |
| Deep | Over roughly 50,000 URLs, faceted navigation, multiple locales, separate mobile behavior, heavy rendering, a migration, unexplained index loss, or material revenue risk | 3–8 working days | Adds segmented crawls, logs, parameters, pagination, broader rendering comparisons, release correlation, and systematic hreflang samples. Large migrations may take longer. |
The checklist
Work in order. A failed gate can invalidate later samples, so record the failure and its scope before continuing.
1. Confirm the target is production
What to do: verify scheme, host, robots file, analytics property, Search Console property, and sitemap host. Why it matters: staging can look clean while production remains broken. How to do it: resolve the agreed canonical host, compare priority pages and response headers, and record the crawl origin. Tool: browser, crawler configuration, Search Console selector. Done when: the register names the confirmed production origin and property, with no staging hostname in seeds or exports.
2. Test robots.txt before crawling
What to do: inspect each production host’s /robots.txt and referenced sitemaps. Why it matters: a disallow rule prevents crawling before content can be evaluated. How to do it: compare Disallow patterns with the intended inventory, test matching and non-matching URLs, and distinguish a crawl block from noindex. Tool: raw response and robots tester. Done when: robots returns 200, intended blocks have reasons, indexable samples are allowed, and one unintended block triggers a critical finding.
3. Reconcile sitemaps with real URLs
What to do: compare submitted sitemaps with the canonical, indexable inventory. Why it matters: a sitemap should name URLs the site wants selected, not redirects, errors, or duplicates. How to do it: normalize entries, compare counts by template, then sample additions and omissions in Sitemaps and Indexing
. Tool: https://app.amicited.com/reports/google-search/sitemaps-indexing and crawl exports. Done when: coverage is at least 95%, 0 entries redirect or error, and every gap has a reason or owner.
4. Measure status-code distribution
What to do: classify responses as 2xx, 3xx, 4xx, or 5xx. Why it matters: errors stop retrieval and redirects add hops. How to do it: follow and report redirects, segment by template, and compare with Bing Crawl
. Tool: https://app.amicited.com/reports/bing-webmasters/crawl, crawler, and monitoring. Done when: indexable URLs return 200; internal errors, loops, and chains are zero; and intentional redirects are documented.
5. Remove redirect chains and loops
What to do: trace redirects to their final response. Why it matters: hops slow discovery; a loop never reaches content. How to do it: export paths, update internal links to final canonicals, and consolidate rules. Tool: redirect report and header checks. Done when: internal links go direct, legacy redirects take one hop, and no loop or chain remains.
6. Establish the eligible indexation ratio
What to do: compare Google’s index state with deliberately eligible URLs. Why it matters: including redirects, filters, duplicates, or noindex pages makes the ratio meaningless. How to do it: build the eligible denominator, inspect priority samples in URL Inspection
, and group exclusions by template. Tool: https://app.amicited.com/reports/google-search/url-inspection, Search Console, and inventory. Done when: at least 90% are indexed or every gap has a root-cause owner; below 80% is a major finding.
7. Verify canonical correctness
What to do: compare declared, final, and Google-selected canonicals. A canonical is the preferred version among similar URLs. Why it matters: a wrong one consolidates signals away from the intended page. How to do it: test unique-page self-references, deliberate cross-canonicals, and consistency across HTML, sitemaps, redirects, and links. Tool: canonical report and https://app.amicited.com/reports/google-search/url-inspection. Done when: 100% of unique indexable pages name one absolute, 200, indexable canonical, with every selected mismatch explained.
8. Find duplicate and near-duplicate clusters
What to do: group URLs with identical or substantially overlapping main content and the same search purpose. Why it matters: duplicates split internal signals and force search engines to choose a version the business may not prefer. How to do it: compare exact hashes, normalized text similarity, titles, canonicals, parameters, and template purpose; then choose consolidation, differentiation, noindex, or removal. Tool: crawler duplicate reports, page inventory, and Google Search Pages
. Done when: no cluster contains more than one unexplained canonical indexable URL serving the same intent, and each accepted variant has a distinct purpose recorded.
9. Find orphans and measure link depth
What to do: join crawler URLs with sitemaps, analytics, Search Console, CMS exports, and backlinks to find pages with no crawlable internal link. Measure the shortest click path from the homepage. Why it matters: an orphan may appear in a sitemap yet receive little internal context or authority; excessive depth makes discovery fragile. How to do it: compare sources, inspect directory patterns in Directory View
, and trace navigation, breadcrumbs, hubs, and contextual links. Tool: https://app.amicited.com/reports/directory and a multi-source crawl. Done when: intended orphan count is zero, priority pages are within three clicks of the homepage, other intended indexable pages are within five, and every exception has a deliberate discovery route.
10. Compare rendered and non-JavaScript HTML
What to do: compare the initial server response with the page after JavaScript executes. Why it matters: a browser may display content and links that a non-JavaScript client never receives. How to do it: fetch representative pages with scripts disabled, inspect raw HTML, then compare headings, main copy, links, canonical, robots directives, structured data, and status after rendering. Tool: crawler in HTML and rendered modes plus browser developer tools. Done when: the initial response contains the primary content, canonical, index directives, and crawlable navigation needed to discover priority pages; any JavaScript-only dependency is explicitly accepted and tested across templates.
11. Validate hreflang where applicable
What to do: verify annotations that connect language or regional equivalents. Why it matters: incomplete or conflicting clusters can make search engines ignore targeting and show the wrong market version. How to do it: test valid language-region codes, absolute canonical URLs, self-references, reciprocal return links, x-default where it has a real fallback role, and indexability of every target. Tool: crawler hreflang report and URL samples. Done when: invalid codes, missing self-references, missing returns, non-canonical targets, redirects, and errors are all zero. If the site has no localized equivalents, record “not applicable” rather than inventing annotations.
12. Test pagination and crawl paths
What to do: verify that multi-page category or archive sequences expose crawlable links and useful unique URLs. Why it matters: infinite scroll or button-only loading can hide deeper items, while canonicalizing every page to page one can remove distinct inventory from discovery. How to do it: disable JavaScript, follow next and numbered links, inspect status, canonical and robots directives, and test the last page and out-of-range parameters. Tool: non-rendered crawl and browser. Done when: every intended item is reachable through anchor links, each useful page self-canonicalizes, invalid page numbers return an appropriate error rather than a soft 200, and no sequence creates an unbounded URL space.
13. Check mobile parity
What to do: compare mobile and desktop delivery for content, links, metadata, directives, structured data, and response status. Why it matters: Google primarily evaluates the mobile representation; hiding meaningful content or links only on mobile changes what it can understand. How to do it: crawl with desktop and smartphone user agents and compare representative templates, not only visual screenshots. Tool: paired crawls, mobile URL inspection, and browser responsive mode. Done when: all indexable content and crawlable links required for meaning and discovery are equivalent, with zero mobile-only blocks, canonical differences, or error responses.
14. Enforce HTTPS and remove mixed content
What to do: verify secure delivery, host redirects, certificates, canonical scheme, internal URLs, and resources loaded over HTTP. Mixed content means an HTTPS page requests an insecure resource. Why it matters: insecure requests can be blocked, expose users, and create inconsistent URL signals. How to do it: crawl all HTTP variants, inspect certificate coverage and browser security errors, and search rendered resource requests. Tool: crawler, browser security panel, and server configuration. Done when: every HTTP page redirects once to its matching HTTPS URL, all canonicals and internal links use HTTPS, certificates are valid for every live host, and active or passive mixed-content requests are zero.
Tools in AmICited
Use product reports as checklist evidence, not as a crawl replacement.
- Sitemaps and Indexing
at
https://app.amicited.com/reports/google-search/sitemaps-indexingshows submitted sitemap state, warnings, errors, and indexing actions. - URL Inspection
at
https://app.amicited.com/reports/google-search/url-inspectiongives Google’s live verdict for sampled URLs and the selected canonical. - Bing Crawl
at
https://app.amicited.com/reports/bing-webmasters/crawlexposes Bing’s crawler activity and URL-level issues. - Google Search Pages
at
https://app.amicited.com/reports/pageshelps select high-value landing pages and separates pages with visibility from pages absent from search data. - Directory View
at
https://app.amicited.com/reports/directoryreveals section-level patterns and supports depth and orphan investigations. - Data Health
at
https://app.amicited.com/features/data-health/records whether connected evidence is complete enough to support confident decisions.
Decision rules
Thresholds create findings; they do not replace judgment. Segment by template and business importance: ten checkout-category failures may matter more than a thousand broken archive tags.
| Check | Finding threshold | Default severity |
|---|---|---|
| Robots | One intended indexable URL blocked, or robots unavailable/non-200 | Critical when scope is a priority template |
| Sitemap coverage | Less than 95% of intended canonical indexable URLs included; any redirect, 4xx, 5xx, blocked, or non-canonical entry | Major; critical for systemic omission |
| Indexation | Less than 90% of eligible URLs without explained exclusions; less than 80% always a finding | Major; critical when a release caused the drop |
| Canonicals | Any unique page with no canonical, multiple canonicals, a non-200 target, or an unintended target; any systemic self-reference error | Major or critical by scope |
| Responses | Any internal 4xx or 5xx; more than 5% of crawlable internal URLs redirect | Major; any widespread 5xx is critical |
| Redirects | Any loop or chain of two or more hops; any internal link to a redirect | Major for loops/chains, minor for isolated stale links |
| Duplication | More than one unexplained canonical indexable URL serving substantially the same intent | Major when template-wide |
| Orphans and depth | Any intended orphan; priority URL deeper than 3 clicks; other intended URL deeper than 5 | Major for priority or template patterns |
| JavaScript | Primary content, canonical, index directive, or discovery links absent from initial HTML without an accepted tested dependency | Critical for affected templates |
| Hreflang | Any invalid code, missing reciprocal link, non-indexable target, redirect, or error | Major when localization applies |
| Pagination | Items unreachable without JavaScript, all pages canonicalized to page one, or unbounded parameter combinations | Major |
| Mobile parity | Any missing primary content/link, conflicting directive/canonical, or mobile-only error | Critical when systemic |
| HTTPS | Any invalid certificate, HTTPS downgrade, or active mixed content; any internal HTTP link | Critical for certificate/active content; major otherwise |
Prioritize with impact × effort × confidence. Score impact from 1–5 based on affected eligible URLs and business journeys. Score effort from 1–5 as an ease factor, where 5 means a small, reversible change and 1 means a large risky program; also record the honest estimate in hours or days. Score confidence as 0.5 for a plausible hypothesis, 0.75 for repeated evidence, or 1.0 for a reproduced root cause. The product gives an ordering aid, not false precision.
Apply the dependency override: a fix that unblocks other work outranks a larger score that does not. Removing a robots block before launch comes ahead of polishing indexed title tags. At the same dependency level, address template-wide causes before symptoms.
Deliverable: the prioritized findings register
Hand over one shared register, not a crawler export. Use one row per root cause and attach URL samples separately.
| Field | Required content |
|---|---|
| Finding ID and title | Stable identifier plus a plain description of the defect |
| Gate | Crawlability, indexability, content quality, or performance |
| Root cause | The rule, template, component, deployment, or configuration creating the symptom |
| Scope and evidence | Affected template/count, representative URLs, report links, crawl timestamp, and reproduction steps |
| Impact | Expected change to discovery, eligibility, consolidation, or user journey; impact score 1–5 |
| Effort | Named team, estimate in hours/days, ease score 1–5, dependencies, and rollback risk |
| Confidence | 0.5, 0.75, or 1.0 with the evidence supporting that choice |
| Priority | Calculated score plus any dependency override and its reason |
| Owner and due date | One accountable person and an agreed delivery date |
| Done when | Exact retest, threshold, sample, and evidence required for closure |
The register is complete when critical and major findings have owners and estimates, blockers have a sequence, hypotheses are labeled, and the publishing decision is explicit.
What goes wrong
A 200-item report nobody can act on
Crawler exports confuse observations with decisions. Group repeated URLs under the template or rule causing them, provide a representative sample, and assign one owner. Two hundred broken URLs produced by one navigation component are one root-cause finding with a measurable scope, not two hundred tasks.
Reporting symptoms instead of causes
“Page not indexed” is a symptom. The cause may be an unintended canonical, an orphaned template, thin parameter variants, a mobile error, or a JavaScript-only link. A finding is not ready for prioritization until it identifies the controllable cause or clearly labels the next diagnostic test.
Auditing staging by accident
Staging may have different robots rules, authentication, data, templates, feature flags, and host behavior. Record the production origin and Search Console property at the top of every export. If a crawl must run against staging for release assurance, label it as a separate comparison and never merge its metrics into the production baseline.
Also avoid counting deliberate exclusions as losses, treating sitemap inclusion as proof of indexing, testing only the homepage, or prioritizing by URL count alone. Define the eligible set, segment by template, and retain acceptance evidence.
Next phase
The next phase, AI accessibility and agent readiness , needs a technically stable sample. Hand over the intended indexable inventory, clean representative URLs for every priority template, raw and rendered HTML comparisons, robots and response evidence, canonical decisions, known exclusions, and the open-findings register.
Do not claim the site is “technically healthy.” State which templates passed the crawl and index gates, which remain blocked, and whether publishing can proceed. The next owner accepts when they can test AI-specific user agents and extraction without rediscovering unresolved search-crawl defects.
FAQ
How often should we repeat a technical baseline audit?
Run it before a migration, redesign, domain change, or large publishing program, then repeat affected checks after release. Monitor continuously and repeat a standard pass when templates, navigation, rendering, or canonical rules change.
What indexation ratio should a healthy site have?
For deliberately eligible URLs, 90% or more is the starting expectation, 80–90% needs explanation, and below 80% is a finding. Exclude redirects, duplicates, filters, and intentional noindex pages from the denominator.
Can we publish content while technical fixes are in progress?
Only when new URLs are crawlable, indexable, canonicalized, internally linked, and unaffected by the defect. If discovery or selection is blocked, pause; new URLs only expand the cleanup.
Do we need a crawler if Search Console is connected?
Yes. Search Console reports what Google observed; a crawler tests the current site and reveals links, responses, depth, canonicals, and duplicates. Neither substitutes for the other.
Who owns fixes found in the audit?
The SEO lead owns the register and acceptance criteria. Engineering usually owns server, rendering, redirect, canonical, and HTTPS fixes; content teams may own duplication and linking. Every item needs one named person.
More tutorials in this section
Ready to put it into practice?
Free check · 7-day trial · no credit card