Content Inventory and Audit: Keep, Improve, Merge, or Prune
Build a content inventory that classifies every URL as keep, improve, merge, or prune using traffic, rankings, conversions, links, and strategic value.
A content inventory is a row-by-row record of the URLs a site currently exposes. A content audit adds the decision: every URL is classified as keep, improve, merge, or prune, with evidence, an owner, and a completion condition. No page stays in “review later.”
Phase: P9, Content Inventory and Audit. Stage: B — Decide. Timebox: 3–5 working days for up to 1,000 indexable URLs; 1–2 weeks for a larger or multi-market site, using sampled manual review within template groups. Owner: SEO strategist or content lead, with analytics, engineering, legal, and product owners approving affected URLs.
Keep leaves the purpose and URL intact. Improve retains the URL but changes its content, targeting, experience, or links. Merge moves useful material and signals into a named survivor. Prune removes a page with no unique, necessary job from the indexable content set; its technical treatment depends on whether it must remain accessible.
Why this phase, and why here
P9 consumes the approved topical map and information architecture . That earlier phase defines which reader need each page should own, where it belongs, and which URL should be canonical. The inventory now tests the live site against that intended system. Without the map, reviewers can see weak numbers but cannot tell whether a page is unnecessary, poorly executed, or strategically essential.
It also consumes earlier tracking, technical, measurement, research, and gap-analysis outputs. These distinguish “no performance” from “measurement failed,” and one shared search intent from two different reader jobs.
Run the audit before the content production system. If skipped, briefs are opened for topics the site already covers, weak URLs are abandoned rather than repaired, redirects are improvised after publication, and writers link to pages already scheduled for retirement. If run before analytics and the topical map are reliable, the team tends to prune low-volume pages that serve a real product or conversion need and keep busy pages that compete with a better destination.
Inputs and outputs
The inventory joins performance and business evidence to one URL set. Its outputs let the next phase create work without reopening each disposition.
| Direction | Item | Acceptance condition |
|---|---|---|
| Input | Canonical URL list and crawl export | Includes status code, indexability, declared and selected canonical where available, title, headings, word count, template, directory, and last modified date. |
| Input | Topical map | Names the intended page owner for each audience need, target intent, post type, and planned internal links. |
| Input | Search performance | Supplies at least 12 months of impressions, clicks, click-through rate, average position, and ranking queries by URL, plus a recent-period comparison. |
| Input | Business outcomes | Supplies leads, purchases, sign-ups, assisted outcomes, and the event definitions used to count them. |
| Input | Link evidence | Includes internal links to each URL and external referring pages or domains, with broken and redirected links identified. |
| Input | Content and ownership facts | Records product necessity, legal requirements, accuracy, audience, author or subject owner, publication date, and meaningful update date. |
| Output | Complete disposition register | Every in-scope URL has exactly one of keep, improve, merge, or prune and a written evidence-based rationale. |
| Output | Merge map | Every source names one survivor, material to transfer, redirect behavior, link updates, and validation owner. |
| Output | Improvement backlog | Each retained URL has a specific problem, action, priority, acceptance test, and accountable owner. |
| Output | Prune implementation list | Each URL has an approved technical treatment, dependency check, removal date, and rollback evidence. |
| Output | Cannibalization register | Each suspected intent collision names the competing URLs, supporting query evidence, chosen owner, and remedy. |
The checklist
1. Build and reconcile the URL universe
What: Create one deduplicated row for every URL found through sitemaps, a site crawl, analytics, search performance data, paid landing-page data, and known campaign or application routes.
Why: A sitemap-only inventory omits orphaned and legacy pages; an analytics-only inventory omits URLs with no visits. Missing URLs cannot receive a disposition, which makes “complete” meaningless.
How: Normalize protocol, host, case, fragments, trailing slashes, and tracking parameters without collapsing distinct pages. Preserve discovered variants. Reconcile redirects, canonicals, non-HTML files, staging routes, and language variants. Keep excluded account or legal pages as rows with reasons.
Tool: Crawl export, XML sitemaps, analytics, search-console exports, CMS export, and AmICited Directory View at app.amicited.com/reports/directory .
Done when: Counts reconcile by source and directory; every URL is in scope or has an exclusion reason, owner, and review date; variants attach to one normalized record.
2. Join performance, conversion, and link evidence
What: Add organic traffic , impressions, rankings, paid activity, conversions, internal links, and backlinks to each row. A backlink is an external page linking to the audited URL; keep it separate from internal links.
Why: Traffic alone rewards broad informational pages and hides pages that close sales, support customers, or carry external authority. A page can be quiet and valuable, or busy and strategically useless.
How: Use a trailing 12 months, plus recent versus previous 90-day comparisons. Join by normalized canonical URL. Record raw conversions and conversion rate ; flag missing data separately from zero.
Tool: Organic vs Paid Pages at app.amicited.com/reports/pages , analytics, conversion tracking, crawl link data, and backlink exports.
Done when: Every indexable URL has a value or “unavailable” state for clicks, impressions, conversions, referring domains, inbound internal links, and last meaningful update; ranges and attribution rules are recorded.
3. Assign the page job and test strategic necessity
What: State the single job of each URL: the audience, need, useful outcome, and next action it supports.
Why: Metrics can show whether a page performs, but not whether the site needs it. Pricing, legal, documentation, store-location, onboarding, and campaign pages may be necessary with little organic demand.
How: Match the URL to the topical map, product catalog, customer journey, support flow, and legal obligations. Ask whether another page completes the same job. Record protected pages and their approver.
Tool: Topical map, CMS ownership fields, product documentation, conversion paths, and stakeholder review.
Done when: Every URL has one stated job or is explicitly marked as having no defensible job, and every protected exception has a named approver and review date.
4. Detect cannibalization while matching queries to URLs
What: Identify cases where two or more URLs repeatedly compete for the same query set and satisfy substantially the same intent. This is content cannibalization , not merely two related pages ranking for different needs.
Why: Cannibalization is easiest to see during inventory work because query, URL, intent, links, and content are already side by side. Ignoring it leaves authority and maintenance split across competing destinations.
How: Group rows by intent and query overlap. Flag pairs when the same material query shows both URLs, the ranking URL repeatedly switches, or both make the same promise and next action. Inspect results and content manually. A category and specific item can both rank usefully; keyword overlap alone does not justify a merge.
Tool: Page-level search data, query exports, Organic vs Paid Pages , crawl content comparison, and the topical map.
Done when: Every suspected collision is confirmed or dismissed with a reason; every confirmed collision has one chosen owner URL and a keep-distinct, improve, or merge remedy.
5. Assign exactly one disposition
What: Label each in-scope URL keep, improve, merge, or prune.
Why: A mixed label such as “improve/merge” transfers uncertainty to delivery, where deadlines encourage the easiest action rather than the correct one.
How: Apply the threshold matrix, then consider necessity, data confidence, seasonality, and risk. Keep distinct successful pages; improve valid but weak pages; merge into the better owner of the same need; prune only when no unique job or transferable value remains.
Tool: The completed inventory, topical map, page report, content-freshness report, and stakeholder approvals.
Done when: A filter for blank or multiple dispositions returns zero rows, every decision cites at least one quantitative and one qualitative reason, and every merge names its survivor.
6. Specify merge and prune mechanics before approval
What: Turn each merge or prune label into an implementation instruction.
Why: Deleting a source before consolidation can lose useful content, interrupt users, and strand links. A decision is not deliverable until engineering knows the exact response and editors know what must move.
How: Choose a merge survivor by intent fit, then performance, conversions, links, URL stability, and maintainability. Move unique accurate material into it, make its canonical URL self-referential, update internal linking , and permanently redirect retired sources. Remove sources from sitemaps and request direct updates to important external links where practical. Use a canonical without a redirect only when a duplicate must remain accessible.
For pruning, redirect only to a genuinely equivalent destination. Use 410 Gone or 404 Not Found when no replacement exists, and noindex when people still need the page but it should not be indexed. Do not redirect removals to the homepage.
Tool: CMS, redirect specification, crawl link report, backlink export, sitemap owner, and analytics annotation log.
Done when: Every source has a destination or explicit no-destination response, consolidation notes, internal-link changes, sitemap action, launch owner, QA test, and approval from affected product, legal, or content stakeholders.
7. Validate, annotate, and hand over
What: Quality-check the register and package work into executable batches.
Why: Bulk changes without a baseline make losses hard to diagnose. Mixing content improvements, migrations, and deletion in one unannotated release hides which decision caused the result.
How: Save pre-change performance, links, index state, and crawl status. Batch by directory or intent cluster, annotate launches, and schedule checks. Review revenue, legal, and heavily linked pages separately.
Tool: Inventory database, analytics annotations, crawl comparison, change tracker, and AmICited reports.
Done when: Each action has an owner and due date, baselines are attached, risky changes are approved, implementation order respects dependencies, and the next phase accepts the backlog without unresolved classifications.
Tools in AmICited
AmICited supplies three complementary views. They narrow and validate decisions; they do not replace the crawl, business context, or manual page review.
| Product view | Use it for | Deep link | Evidence to save |
|---|---|---|---|
| Organic vs Paid Pages | Compare impressions, clicks, CTR, position, paid activity, and landing-page reach by URL; identify pages with demand, weak realization, or channel dependence. | Open Pages | Exported row or captured filter, date range, comparison period, and the metric used in the disposition. |
| Content Freshness | Find directories with old or rarely updated URLs and inspect additions, updates, removals, and freshness coverage. Freshness is a review signal, not a reason to rewrite dates. | Open Content Freshness | Age or churn evidence, coverage confidence, and the exact URLs selected for manual review. |
| Directory View | Compare section size with clicks and impressions, locate large low-return sections, and spot template-level patterns before judging pages individually. | Open Directory View | Directory, depth, comparison window, affected URL count, and whether the issue is sectional or page-specific. |
Decision rules: what bad looks like in numbers
Thresholds create consistent triage, not universal ranking laws. Adjust them before review for extreme seasonality, low-volume enterprise demand, or large product catalogs. These defaults assume 12 months of valid data and a recent 90-day comparison.
| Disposition | Default quantitative trigger | Required qualitative test | Decision |
|---|---|---|---|
| Keep | At least 100 organic clicks, 1 attributed conversion, 3 linking root domains, or a sustained top-10 ranking for a materially relevant query in 12 months. | The page owns a distinct job, is accurate, and has no critical technical or experience defect. | Keep the URL and purpose. Minor housekeeping does not change the disposition. |
| Improve | The job is valid, but the page has 10–99 clicks, ranks mostly in positions 11–30 with at least 500 impressions, has lost 20% or more clicks in the recent 90 days versus the preceding 90 days, or has not been substantively reviewed in 12 months where facts can change. | Demand or business need exists, and no stronger URL already owns the same intent. | Retain the URL; specify the weak section, evidence, remedy, and target outcome. |
| Merge | Two or more URLs share the same primary intent and either receive impressions for the same material query set, alternate as the ranking URL, or duplicate most of the useful answer. | One survivor can satisfy the combined need without confusing the reader or changing the required page type. | Consolidate unique material and signals into the named survivor, update links, then permanently redirect sources. |
| Prune | Over 12 months: 0 organic clicks, 0 conversions, 0 linking root domains, and no meaningful paid or referral use; or the URL is obsolete, duplicated, inaccurate, or empty. | It has no necessary legal, product, support, navigation, brand, or campaign job, and no unique material worth moving. | Remove it from the indexable set using the approved redirect, 410, 404, or noindex treatment. |
These are OR conditions for review and AND conditions for action. A 25% decline triggers review, not an automatic rewrite. Pruning normally requires all zero-value signals plus no strategic job. A page ranking 18th with 5,000 impressions may deserve improvement; five annual visits may still matter if one becomes a qualified lead.
Apply three additional gates:
- Data quality gate: if tracking, canonicalization, or indexability failed, fix measurement or technical access before judging content.
- Seasonality gate: compare like-for-like periods. Do not prune a tax, holiday, admissions, or event page during its off-season.
- Risk gate: require manual approval for any URL with conversions, backlinks, legal duties, product dependencies, or material direct traffic, even if another rule suggests merge or prune.
Why pruning can help — and why it is oversold
Content pruning can improve a site when it removes inaccurate pages, duplicate destinations, expired inventory with no successor, or low-value archives that consume editorial and linking attention. Consolidation can make one page more complete, direct internal links toward the intended owner, and give users fewer conflicting answers. On very large sites, removing endless low-value URL combinations may also make crawling and monitoring easier.
The oversold version says that deleting low-traffic pages raises the quality of the whole domain and therefore rankings. That is too broad. Low traffic is not the same as low quality, search engines do not promise a site-wide reward for deleting pages, and many before-and-after pruning stories combine redirects, rewrites, internal-link changes, technical fixes, and seasonal recovery. They cannot isolate deletion as the cause.
Use pruning to correct a known content-system defect, not to manufacture a ratio. State the mechanism you expect: fewer duplicate intents, removal of false information, a cleaner product catalog, stronger consolidated pages, or less crawl waste in a specific large section. Then measure that mechanism. A successful prune may preserve total conversions and qualified traffic while reducing indexed URLs; raw URL count is not the outcome.
Deliverable: the disposition and migration register
The deliverable is a versioned table or database with one row per normalized URL. Preserve the original data snapshot and add decisions in reviewable fields:
URL ID | Current URL | Declared canonical | Status/indexability | Directory/template
Page job | Primary intent | Topical-map node | Owner | Last meaningful review
12m impressions/clicks | Recent 90d change | Top queries/positions
Conversions | Paid/referral use | Internal links | Linking root domains
Disposition | Quantitative reason | Qualitative reason | Confidence
Survivor/destination | Material to consolidate | Redirect or index treatment
Internal-link changes | Sitemap action | Approver | Implementer | Due date
Baseline snapshot | Launch annotation | QA result | Review date
Provide four filtered views: keep register, improvement backlog, merge map, and prune list. Several sources may feed one survivor, but redirect chains are not acceptable. Attach the cannibalization register and threshold exceptions.
The deliverable is accepted when URL-source counts reconcile, blank dispositions equal zero, every merge has one destination, every prune has a technical treatment, and every action has an owner and done-when test.
What goes wrong
- Zero traffic becomes an automatic delete rule. Necessary support, product, legal, and conversion pages disappear. Treat zero as a review trigger and test the page job.
- The inventory starts and ends with the sitemap. Orphans, old campaign URLs, redirects, and untracked landing pages remain unclassified. Reconcile every discovery source.
- Missing data is stored as zero. Tracking failures masquerade as poor performance. Use explicit unavailable states and resolve critical gaps first.
- Freshness means changing the date. A timestamp moves while facts, screenshots, and advice remain stale. Record substantive changes, not cosmetic republishing.
- The highest-traffic page always survives. Choose by intent fit first, then use performance and links as evidence.
- A canonical tag substitutes for migration. The duplicate remains linked and accessible while signals stay ambiguous. Redirect a retired source and update links; use canonicals when duplicates must coexist.
- Every removal redirects to the homepage. Users reach an irrelevant destination and the audit hides the absence of a true replacement. Redirect only to a close equivalent.
- Inbound links are forgotten. Record important external links and request direct updates where worthwhile.
- Cannibalization is diagnosed from one shared keyword. Related pages are merged despite serving different tasks. Confirm overlapping intent, result behavior, and content before acting.
- Changes ship in one opaque batch. Release by cluster or directory and annotate each batch.
- No one owns post-launch QA. Redirects chain, sources remain in sitemaps, and internal links still point at retired URLs. Assign a validator before approval.
Next phase
The content production system needs this phase’s approved improvement backlog, merge instructions, protected keep register, prune dependencies, topical-map IDs, and acceptance tests. It should not ask writers to decide whether a URL survives while they are drafting it.
Handoff is complete when the production owner can create work packages that name the existing URL, intended page job, evidence, required content changes, consolidation sources, links to update, and done-when condition. Engineering receives redirects and index treatments as a separate but sequenced batch. Production begins only after merge destinations are stable, so new content never links to a URL scheduled for retirement.
FAQ
Does every URL really need one disposition?
Yes. Every in-scope, indexable or intentionally accessible URL must be marked keep, improve, merge, or prune. Excluded technical URLs still need an exclusion reason and owner so they do not disappear from the audit by accident.
Should a page with zero organic traffic always be pruned?
No. Zero traffic is a review trigger, not a deletion instruction. Keep a page when it serves a necessary product, support, legal, navigation, campaign, or conversion job; improve it when demand exists but execution is weak; merge it when another URL serves the same intent; prune only when it has no defensible job or unique value.
How long should the measurement window be?
Use a trailing 12 months by default so seasonality and slow B2B buying cycles are represented, then compare the most recent 90 days with the prior 90 days to detect direction. Use a longer window for highly seasonal or low-volume sites and document the exception.
Is a canonical tag enough when two pages are merged?
Usually not when the source URL is being retired. Consolidate the useful material into the survivor, update internal links, remove the source from sitemaps, and permanently redirect it. A canonical tag is a hint for duplicates that must remain accessible, while a redirect moves users and crawlers to the survivor.
Does pruning automatically improve rankings?
No. Removing obsolete, duplicate, or misleading pages can simplify maintenance, linking, and crawling, but deletion is not a general ranking tactic. Measure the affected cluster before and after, and judge success by retained demand, clean redirects, fewer conflicting URLs, and better business outcomes.
More tutorials in this section
Ready to put it into practice?
Free check · 7-day trial · no credit card