Academy

Technical SEO Audit: Crawl and Indexation

Run a technical baseline audit that finds crawl, indexation, canonical, rendering, and internal-link problems before you invest in new SEO content at scale.

16 min read

Technical baseline audit

Phase P2 · Stage A — Understand
Timebox: 2–4 hours for a light pass, 1–2 working days for a standard pass, or 3–8 working days for a deep pass.
Owner: the technical SEO lead. Engineering, analytics, content, and localization owners contribute evidence and accept fixes in their areas.

A technical baseline audit establishes whether search engines can reach, interpret, and select the URLs the business expects them to show. Its scope covers crawl controls, HTTP responses, indexation, canonicals, links, rendering, international targeting, and secure delivery. The result is a prioritized findings register with named owners and acceptance tests, not a score.

Why this phase comes here

Publishing onto a site with crawl or indexation problems compounds the damage. Search engines often discover a repeated defect across new URLs faster than they evaluate and reward the content. A broken canonical template can point every article elsewhere; a robots rule can hide a directory; client-rendered navigation can create orphans for a non-JavaScript client. Every new page enlarges the affected set and makes repair riskier.

Fix the foundation first. The order is crawlability → indexability → content quality → performance because each layer is a gate. Crawlability means a crawler can discover and request a URL; indexability means the reachable URL is eligible for inclusion. Only then should content quality and performance be judged. A fast page blocked by robots.txt cannot compete, and title tags do not matter on unreachable pages.

This phase consumes the scope, priority journeys, markets, and risks from Discovery and goals . Running it earlier produces a crawl with no business context. Skipping it lets research and production target templates that cannot reliably enter the index.

The sequence is a control, not a preference
Do not use a publishing deadline as permission to skip a blocking crawl or indexation defect. A fix that restores access to an entire template outranks a larger-looking optimization that affects pages already eligible to rank.

Inputs and outputs

Inputs define the intended site, not merely what a crawler finds. Outputs tell the next owner which URLs are safe to test and which remain blocked.

DirectionItemAcceptance condition
InputProduction origins and canonical hostIncludes protocol, www decision, subdomains, international hosts, and known legacy domains.
InputIntended indexable URL inventoryLists templates, directories, locales, sitemap sources, and exclusions such as filters, account pages, and internal search.
InputAccess and evidenceProduction crawl permission, Google Search Console, Bing Webmaster Tools, analytics, log files when available, deployment history, and CMS rules.
InputDiscovery briefNames priority journeys, revenue or lead value, markets, launch constraints, and accountable owners.
InputRecent change registerRecords migrations, redesigns, JavaScript framework changes, canonical or pagination changes, incidents, and release dates.
OutputPrioritized findings registerEvery finding has affected scope, evidence, root cause, impact, effort estimate, confidence, owner, deadline, and done-when test.
OutputCrawl and index baselineRecords eligible URLs, crawled URLs, status distribution, sitemap coverage, indexation ratio, orphan count, and depth distribution.
OutputBlocking-dependency decisionStates whether publishing may proceed, proceed only for unaffected templates, or pause until named blockers pass retest.
OutputHandoff packageGives the next phase a clean URL sample, unresolved exclusions, rendering evidence, and accepted limitations.
Logo

Ready to Monitor Your AI Visibility?

Track how AI chatbots mention your brand across ChatGPT, Perplexity, and other platforms.

Choose the audit depth

Select depth before crawling. Estimates assume access is ready and exclude implementation.

ModeChoose it whenHonest timeboxCoverage and limitations
LightUnder roughly 500 indexable URLs, one main template and language, no recent migration, and no JavaScript-dependent primary content2–4 hoursControls, sitemaps, responses, representative crawl, priority inspection, basic canonicals, and mobile/HTTPS samples. It can miss long-tail orphans, rare loops, near-duplicates, template-specific rendering failures, and hreflang defects. It is triage, not migration assurance.
StandardUp to roughly 50,000 intended URLs, several templates, routine JavaScript, or a substantial content program1–2 working daysFull crawl, sitemap reconciliation, sampled inspection, duplicates, depth, rendering, and template rules. This is the default for an established site.
DeepOver roughly 50,000 URLs, faceted navigation, multiple locales, separate mobile behavior, heavy rendering, a migration, unexplained index loss, or material revenue risk3–8 working daysAdds segmented crawls, logs, parameters, pagination, broader rendering comparisons, release correlation, and systematic hreflang samples. Large migrations may take longer.

The checklist

Work in order. A failed gate can invalidate later samples, so record the failure and its scope before continuing.

1. Confirm the target is production

What to do: verify scheme, host, robots file, analytics property, Search Console property, and sitemap host. Why it matters: staging can look clean while production remains broken. How to do it: resolve the agreed canonical host, compare priority pages and response headers, and record the crawl origin. Tool: browser, crawler configuration, Search Console selector. Done when: the register names the confirmed production origin and property, with no staging hostname in seeds or exports.

2. Test robots.txt before crawling

What to do: inspect each production host’s /robots.txt and referenced sitemaps. Why it matters: a disallow rule prevents crawling before content can be evaluated. How to do it: compare Disallow patterns with the intended inventory, test matching and non-matching URLs, and distinguish a crawl block from noindex. Tool: raw response and robots tester. Done when: robots returns 200, intended blocks have reasons, indexable samples are allowed, and one unintended block triggers a critical finding.

3. Reconcile sitemaps with real URLs

What to do: compare submitted sitemaps with the canonical, indexable inventory. Why it matters: a sitemap should name URLs the site wants selected, not redirects, errors, or duplicates. How to do it: normalize entries, compare counts by template, then sample additions and omissions in Sitemaps and Indexing . Tool: https://app.amicited.com/reports/google-search/sitemaps-indexing and crawl exports. Done when: coverage is at least 95%, 0 entries redirect or error, and every gap has a reason or owner.

4. Measure status-code distribution

What to do: classify responses as 2xx, 3xx, 4xx, or 5xx. Why it matters: errors stop retrieval and redirects add hops. How to do it: follow and report redirects, segment by template, and compare with Bing Crawl . Tool: https://app.amicited.com/reports/bing-webmasters/crawl, crawler, and monitoring. Done when: indexable URLs return 200; internal errors, loops, and chains are zero; and intentional redirects are documented.

5. Remove redirect chains and loops

What to do: trace redirects to their final response. Why it matters: hops slow discovery; a loop never reaches content. How to do it: export paths, update internal links to final canonicals, and consolidate rules. Tool: redirect report and header checks. Done when: internal links go direct, legacy redirects take one hop, and no loop or chain remains.

6. Establish the eligible indexation ratio

What to do: compare Google’s index state with deliberately eligible URLs. Why it matters: including redirects, filters, duplicates, or noindex pages makes the ratio meaningless. How to do it: build the eligible denominator, inspect priority samples in URL Inspection , and group exclusions by template. Tool: https://app.amicited.com/reports/google-search/url-inspection, Search Console, and inventory. Done when: at least 90% are indexed or every gap has a root-cause owner; below 80% is a major finding.

7. Verify canonical correctness

What to do: compare declared, final, and Google-selected canonicals. A canonical is the preferred version among similar URLs. Why it matters: a wrong one consolidates signals away from the intended page. How to do it: test unique-page self-references, deliberate cross-canonicals, and consistency across HTML, sitemaps, redirects, and links. Tool: canonical report and https://app.amicited.com/reports/google-search/url-inspection. Done when: 100% of unique indexable pages name one absolute, 200, indexable canonical, with every selected mismatch explained.

8. Find duplicate and near-duplicate clusters

What to do: group URLs with identical or substantially overlapping main content and the same search purpose. Why it matters: duplicates split internal signals and force search engines to choose a version the business may not prefer. How to do it: compare exact hashes, normalized text similarity, titles, canonicals, parameters, and template purpose; then choose consolidation, differentiation, noindex, or removal. Tool: crawler duplicate reports, page inventory, and Google Search Pages . Done when: no cluster contains more than one unexplained canonical indexable URL serving the same intent, and each accepted variant has a distinct purpose recorded.

What to do: join crawler URLs with sitemaps, analytics, Search Console, CMS exports, and backlinks to find pages with no crawlable internal link. Measure the shortest click path from the homepage. Why it matters: an orphan may appear in a sitemap yet receive little internal context or authority; excessive depth makes discovery fragile. How to do it: compare sources, inspect directory patterns in Directory View , and trace navigation, breadcrumbs, hubs, and contextual links. Tool: https://app.amicited.com/reports/directory and a multi-source crawl. Done when: intended orphan count is zero, priority pages are within three clicks of the homepage, other intended indexable pages are within five, and every exception has a deliberate discovery route.

10. Compare rendered and non-JavaScript HTML

What to do: compare the initial server response with the page after JavaScript executes. Why it matters: a browser may display content and links that a non-JavaScript client never receives. How to do it: fetch representative pages with scripts disabled, inspect raw HTML, then compare headings, main copy, links, canonical, robots directives, structured data, and status after rendering. Tool: crawler in HTML and rendered modes plus browser developer tools. Done when: the initial response contains the primary content, canonical, index directives, and crawlable navigation needed to discover priority pages; any JavaScript-only dependency is explicitly accepted and tested across templates.

11. Validate hreflang where applicable

What to do: verify annotations that connect language or regional equivalents. Why it matters: incomplete or conflicting clusters can make search engines ignore targeting and show the wrong market version. How to do it: test valid language-region codes, absolute canonical URLs, self-references, reciprocal return links, x-default where it has a real fallback role, and indexability of every target. Tool: crawler hreflang report and URL samples. Done when: invalid codes, missing self-references, missing returns, non-canonical targets, redirects, and errors are all zero. If the site has no localized equivalents, record “not applicable” rather than inventing annotations.

12. Test pagination and crawl paths

What to do: verify that multi-page category or archive sequences expose crawlable links and useful unique URLs. Why it matters: infinite scroll or button-only loading can hide deeper items, while canonicalizing every page to page one can remove distinct inventory from discovery. How to do it: disable JavaScript, follow next and numbered links, inspect status, canonical and robots directives, and test the last page and out-of-range parameters. Tool: non-rendered crawl and browser. Done when: every intended item is reachable through anchor links, each useful page self-canonicalizes, invalid page numbers return an appropriate error rather than a soft 200, and no sequence creates an unbounded URL space.

13. Check mobile parity

What to do: compare mobile and desktop delivery for content, links, metadata, directives, structured data, and response status. Why it matters: Google primarily evaluates the mobile representation; hiding meaningful content or links only on mobile changes what it can understand. How to do it: crawl with desktop and smartphone user agents and compare representative templates, not only visual screenshots. Tool: paired crawls, mobile URL inspection, and browser responsive mode. Done when: all indexable content and crawlable links required for meaning and discovery are equivalent, with zero mobile-only blocks, canonical differences, or error responses.

14. Enforce HTTPS and remove mixed content

What to do: verify secure delivery, host redirects, certificates, canonical scheme, internal URLs, and resources loaded over HTTP. Mixed content means an HTTPS page requests an insecure resource. Why it matters: insecure requests can be blocked, expose users, and create inconsistent URL signals. How to do it: crawl all HTTP variants, inspect certificate coverage and browser security errors, and search rendered resource requests. Tool: crawler, browser security panel, and server configuration. Done when: every HTTP page redirects once to its matching HTTPS URL, all canonicals and internal links use HTTPS, certificates are valid for every live host, and active or passive mixed-content requests are zero.

Tools in AmICited

Use product reports as checklist evidence, not as a crawl replacement.

  • Sitemaps and Indexing at https://app.amicited.com/reports/google-search/sitemaps-indexing shows submitted sitemap state, warnings, errors, and indexing actions.
  • URL Inspection at https://app.amicited.com/reports/google-search/url-inspection gives Google’s live verdict for sampled URLs and the selected canonical.
  • Bing Crawl at https://app.amicited.com/reports/bing-webmasters/crawl exposes Bing’s crawler activity and URL-level issues.
  • Google Search Pages at https://app.amicited.com/reports/pages helps select high-value landing pages and separates pages with visibility from pages absent from search data.
  • Directory View at https://app.amicited.com/reports/directory reveals section-level patterns and supports depth and orphan investigations.
  • Data Health at https://app.amicited.com/features/data-health/ records whether connected evidence is complete enough to support confident decisions.

Decision rules

Thresholds create findings; they do not replace judgment. Segment by template and business importance: ten checkout-category failures may matter more than a thousand broken archive tags.

CheckFinding thresholdDefault severity
RobotsOne intended indexable URL blocked, or robots unavailable/non-200Critical when scope is a priority template
Sitemap coverageLess than 95% of intended canonical indexable URLs included; any redirect, 4xx, 5xx, blocked, or non-canonical entryMajor; critical for systemic omission
IndexationLess than 90% of eligible URLs without explained exclusions; less than 80% always a findingMajor; critical when a release caused the drop
CanonicalsAny unique page with no canonical, multiple canonicals, a non-200 target, or an unintended target; any systemic self-reference errorMajor or critical by scope
ResponsesAny internal 4xx or 5xx; more than 5% of crawlable internal URLs redirectMajor; any widespread 5xx is critical
RedirectsAny loop or chain of two or more hops; any internal link to a redirectMajor for loops/chains, minor for isolated stale links
DuplicationMore than one unexplained canonical indexable URL serving substantially the same intentMajor when template-wide
Orphans and depthAny intended orphan; priority URL deeper than 3 clicks; other intended URL deeper than 5Major for priority or template patterns
JavaScriptPrimary content, canonical, index directive, or discovery links absent from initial HTML without an accepted tested dependencyCritical for affected templates
HreflangAny invalid code, missing reciprocal link, non-indexable target, redirect, or errorMajor when localization applies
PaginationItems unreachable without JavaScript, all pages canonicalized to page one, or unbounded parameter combinationsMajor
Mobile parityAny missing primary content/link, conflicting directive/canonical, or mobile-only errorCritical when systemic
HTTPSAny invalid certificate, HTTPS downgrade, or active mixed content; any internal HTTP linkCritical for certificate/active content; major otherwise

Prioritize with impact × effort × confidence. Score impact from 1–5 based on affected eligible URLs and business journeys. Score effort from 1–5 as an ease factor, where 5 means a small, reversible change and 1 means a large risky program; also record the honest estimate in hours or days. Score confidence as 0.5 for a plausible hypothesis, 0.75 for repeated evidence, or 1.0 for a reproduced root cause. The product gives an ordering aid, not false precision.

Apply the dependency override: a fix that unblocks other work outranks a larger score that does not. Removing a robots block before launch comes ahead of polishing indexed title tags. At the same dependency level, address template-wide causes before symptoms.

Deliverable: the prioritized findings register

Hand over one shared register, not a crawler export. Use one row per root cause and attach URL samples separately.

FieldRequired content
Finding ID and titleStable identifier plus a plain description of the defect
GateCrawlability, indexability, content quality, or performance
Root causeThe rule, template, component, deployment, or configuration creating the symptom
Scope and evidenceAffected template/count, representative URLs, report links, crawl timestamp, and reproduction steps
ImpactExpected change to discovery, eligibility, consolidation, or user journey; impact score 1–5
EffortNamed team, estimate in hours/days, ease score 1–5, dependencies, and rollback risk
Confidence0.5, 0.75, or 1.0 with the evidence supporting that choice
PriorityCalculated score plus any dependency override and its reason
Owner and due dateOne accountable person and an agreed delivery date
Done whenExact retest, threshold, sample, and evidence required for closure

The register is complete when critical and major findings have owners and estimates, blockers have a sequence, hypotheses are labeled, and the publishing decision is explicit.

What goes wrong

A 200-item report nobody can act on

Crawler exports confuse observations with decisions. Group repeated URLs under the template or rule causing them, provide a representative sample, and assign one owner. Two hundred broken URLs produced by one navigation component are one root-cause finding with a measurable scope, not two hundred tasks.

Reporting symptoms instead of causes

“Page not indexed” is a symptom. The cause may be an unintended canonical, an orphaned template, thin parameter variants, a mobile error, or a JavaScript-only link. A finding is not ready for prioritization until it identifies the controllable cause or clearly labels the next diagnostic test.

Auditing staging by accident

Staging may have different robots rules, authentication, data, templates, feature flags, and host behavior. Record the production origin and Search Console property at the top of every export. If a crawl must run against staging for release assurance, label it as a separate comparison and never merge its metrics into the production baseline.

Also avoid counting deliberate exclusions as losses, treating sitemap inclusion as proof of indexing, testing only the homepage, or prioritizing by URL count alone. Define the eligible set, segment by template, and retain acceptance evidence.

Next phase

The next phase, AI accessibility and agent readiness , needs a technically stable sample. Hand over the intended indexable inventory, clean representative URLs for every priority template, raw and rendered HTML comparisons, robots and response evidence, canonical decisions, known exclusions, and the open-findings register.

Do not claim the site is “technically healthy.” State which templates passed the crawl and index gates, which remain blocked, and whether publishing can proceed. The next owner accepts when they can test AI-specific user agents and extraction without rediscovering unresolved search-crawl defects.

FAQ

How often should we repeat a technical baseline audit?

Run it before a migration, redesign, domain change, or large publishing program, then repeat affected checks after release. Monitor continuously and repeat a standard pass when templates, navigation, rendering, or canonical rules change.

What indexation ratio should a healthy site have?

For deliberately eligible URLs, 90% or more is the starting expectation, 80–90% needs explanation, and below 80% is a finding. Exclude redirects, duplicates, filters, and intentional noindex pages from the denominator.

Can we publish content while technical fixes are in progress?

Only when new URLs are crawlable, indexable, canonicalized, internally linked, and unaffected by the defect. If discovery or selection is blocked, pause; new URLs only expand the cleanup.

Do we need a crawler if Search Console is connected?

Yes. Search Console reports what Google observed; a crawler tests the current site and reveals links, responses, depth, canonicals, and duplicates. Neither substitutes for the other.

Who owns fixes found in the audit?

The SEO lead owns the register and acceptance criteria. Engineering usually owns server, rendering, redirect, canonical, and HTTPS fixes; content teams may own duplication and linking. Every item needs one named person.

Fix the crawl and index foundation before you scale content
Open AmICited's sitemap and indexing report, capture the baseline, and turn every blocker into an owned, testable finding.

← All Academy tutorials

Ready to put it into practice?

Free check · 7-day trial · no credit card