Academy

GEO and AEO Readiness Checklist

Use this GEO and AEO readiness checklist to test AI crawler access, extractable answers, schema, llms.txt, prompt visibility, and citation evidence today.

16 min read

This checklist decides whether a site is technically reachable, easy to extract, unambiguous about its entities, and measurably present in AI answers. Generative engine optimization (GEO) improves the likelihood that generative systems retrieve, use, and cite a source. Answer engine optimization (AEO) makes a page capable of supplying a direct answer. Neither is a promise of inclusion: readiness removes avoidable obstacles, while prompt and citation data show what actually happened.

Checklist: GEO and AEO readiness. Timebox: 3–5 working days for a representative audit and remediation plan, followed by a minimum 28-day measurement window. Owner: SEO lead, with engineering accountable for access and rendering, editorial for passage quality, and an analyst for prompt and citation measurement.

Test the homepage, one page from each revenue-critical template, the ten pages mapped to priority prompts, and pages already earning or unexpectedly missing citations. Record every URL for retesting.

Why this checklist, and why here

This gate consumes the AI accessibility and agent readiness audit, which identifies crawler and extraction constraints; baseline measurement , which freezes the pre-change state; and keyword and prompt research , which defines the real questions and engines to test. It also needs the approved page inventory, entity facts, schema ownership, server or CDN access, and a release log.

The order matters because reachability, answer quality, and visibility are different layers. Prompt tracking before crawl checks can report that a brand is missing without explaining whether the cause is access, relevance, authority, or simple retrieval delay. Rewriting passages before confirming business policy can expose content the organization intended to withhold. Adding schema before the visible facts and entity model are stable can make a contradiction machine-readable rather than correct it.

If this checklist is skipped, teams tend to assert that a page is “AI optimized” because it has short paragraphs, FAQ schema, or an llms.txt file. Those are implementation facts, not outcomes. The contract here is stricter: state what crawlers may access, prove representative pages can be fetched and understood, then compare completed prompt runs and source citations against a dated baseline.

Inputs and outputs

The outputs are the contract with publishing, engineering, and measurement. “Ready” without a URL set, evidence, thresholds, and observation dates cannot be reproduced.

DirectionItemAcceptance condition
InputRepresentative URL manifestIncludes each critical template, ten priority-prompt destinations, currently cited pages, and strategically important missing pages; every URL has an owner.
InputCrawler policyLists relevant crawler families, allowed or blocked state, business rationale, approver, review date, and any path-level exception.
InputPrompt baselineStores exact prompt wording, country, language, provider, cadence, brand set, and at least one completed pre-change run.
InputEntity and evidence recordNames the organization, products, people, locations, identifiers, canonical URLs, approved claims, and the source of each fact.
InputTechnical accessProvides read access to robots rules, CDN or firewall behavior, rendered HTML, sitemaps, headers, and deployed structured data.
OutputReadiness test matrixOne row per URL and check, with observed result, evidence, severity, owner, due date, retest date, and pass or fail.
OutputCrawler decision registerRecords policy separately from technical reachability so an intentional block is not misreported as an implementation defect.
OutputPassage and schema remediation queueIdentifies the exact page, section, target prompt, entity, required change, acceptance test, and accountable owner.
OutputMeasurement planFreezes prompts, providers, baseline dates, deployment annotation, 28-day observation window, and comparison rules.
OutputSigned handoffNames remaining risks, accepted exceptions, failed checks, release decision, and the person authorized to reopen the gate.
Logo

Ready to Monitor Your AI Visibility?

Track how AI chatbots mention your brand across ChatGPT, Perplexity, and other platforms.

The checklist

Every item ends with a Done when condition. Attach the response, rendered extract, validator result, screenshot, or report row rather than recording an unsupported green check.

1. Make an explicit crawler-access decision

  • What: Decide which AI user agents may crawl which public paths. A user agent is the identifier a crawler presents in its request; it is a signal, not strong authentication.
  • Why: Allowing a crawler can improve discovery but may also enable reuse, add server load, conflict with licensing, or expose material that was public only by accident. Blocking may be a valid commercial decision, but it must not be mistaken for an SEO bug.
  • How: List relevant crawler families and group them by purpose: search or answer retrieval, model training, and general web archives. For each, document allow, block, or path-limited access; business rationale; approver; and review date. Compare that decision with robots.txt , CDN rules, web application firewall rules, authentication, and origin behavior.
  • Tool: Policy register, robots parser, CDN and firewall configuration, server logs, and legal or content-owner review.
  • Done when: 100% of in-scope crawler families have an owner-approved decision; every intentional block is labeled as policy; and zero live rules contradict the recorded decision on the representative URL set.

2. Test reachability as the crawler, not as a browser session

  • What: Verify that allowed crawlers receive the canonical content with a successful response and without a challenge, login, consent wall, or empty client-side shell.
  • Why: A permissive robots rule does not prove delivery. A CDN can return 403, 429, a CAPTCHA, or different HTML to a non-browser request while a signed-in employee sees a normal page.
  • How: Fetch each representative URL with the relevant user-agent string from a clean request. Record status, redirects, response time, content type, canonical, index directives, final body size, and whether the main answer appears in the returned or rendered HTML. Compare bot and normal-browser responses for material differences.
  • Tool: AI Accessibility and Agent Readiness in AmICited, response and header inspection, server logs, and a rendered HTML comparison.
  • Done when: Every intentionally allowed test returns the expected canonical page with 200 status; redirect chains contain no more than one hop; no allowed request receives 401, 403, 429, challenge HTML, or a blank main region; and differences have a documented, non-deceptive reason.

3. Audit llms.txt as a map, not a magic switch

  • What: Review /llms.txt, an emerging voluntary text file intended to point language-model tools toward useful site resources. Treat it as guidance, not an access control or guaranteed ranking signal.
  • Why: A concise map can help an agent find canonical documentation, but a stale file can send it to redirects, duplicate pages, or retired claims. Its presence cannot compensate for blocked crawling or weak content.
  • How: If the business adopts the file, keep the title and description clear, link only to canonical public URLs, group resources by real user purpose, and prefer durable pages over a dump of the entire sitemap. Test every listed URL. If the business chooses not to publish it, record that decision without failing the whole readiness gate.
  • Tool: AmICited llms.txt check, link checker, URL inventory, and content-owner review.
  • Done when: The decision to publish or omit is recorded; if present, the file returns 200 as plain text, contains zero broken, redirected, blocked, duplicate, or noncanonical links, and has a named owner and review date.

4. Make priority passages self-contained and front-loaded

  • What: Give each priority question a self-contained passage: a section that states the answer early and includes enough nouns, scope, conditions, and evidence to remain accurate when extracted from surrounding copy.
  • Why: Retrieval systems often select a passage rather than the whole page. “It depends” or “this method” loses meaning when detached from the heading; a delayed answer forces the system to assemble facts across multiple sections and raises the chance of omission or distortion.
  • How: Put the direct answer in the first one or two sentences under the matching heading. Name the entity and topic instead of relying on pronouns. Follow with qualifications, evidence, examples, and exceptions. Keep necessary context with the claim; do not reduce complex legal, medical, financial, or safety advice to an unconditional snippet.
  • Tool: Prompt-to-section map, editorial extraction test, plain-text reader, and subject-matter review.
  • Done when: Each of the ten priority prompts maps to one canonical page and one answer section; the answer appears within the first 80 words of that section; and a reviewer can copy the passage alone without losing the subject, scope, condition, or evidence source.

5. Use extractable formats for the job

  • What: Represent sequences as numbered steps, alternatives as comparison tables, specifications as labeled values, and short sets as lists. Keep the same facts available in meaningful HTML, not only in images, video, canvas, or interaction-only tabs.
  • Why: Format encodes relationships. A prose paragraph can hide which value belongs to which product, while a table exposes the comparison. Content that exists only after a click or inside an image may be missed or detached from its labels.
  • How: Inspect the document outline and raw HTML. Give tables headers, lists one idea per item, figures captions, images useful alternative text, and interactive content a server-rendered summary. Ensure hidden tabs do not contain the only copy of a critical answer.
  • Tool: Accessibility tree, HTML source inspection, keyboard-only review, and a JavaScript-disabled or plain-text rendering.
  • Done when: 100% of critical facts remain available and correctly labeled without interaction; every comparison has explicit row and column labels; every sequence has ordered steps; and no priority answer exists only in media or a client-rendered widget.

6. Align visible facts, entities, and structured data

  • What: Clarify the people, organizations, products, places, and relationships on the page, then express supported facts through valid structured data . An entity is a distinct real-world thing that can be named and disambiguated from similar things.
  • Why: Ambiguous names and contradictory identifiers make attribution unreliable. Schema markup can reduce ambiguity, but markup that is broader, newer, or more promotional than the visible page creates conflict rather than trust.
  • How: Use one canonical name, URL, logo, and stable identifier set for the organization. Connect authors and reviewers to real profile pages. Choose the most specific applicable schema type, include only visible and verified facts, and connect related nodes with consistent identifiers. Validate syntax and compare every material property with the rendered page.
  • Tool: Entity record, JSON-LD inspection, Schema.org validator, rich-result testing where applicable, and template-level QA.
  • Done when: Every representative page has one unambiguous primary entity; zero material schema properties contradict or exceed visible claims; zero syntax errors remain; and each revenue-critical template has an approved schema owner and test fixture.

7. Protect citation quality with sources and freshness

  • What: Support claims that require evidence with identifiable primary or authoritative sources, and expose when the page was substantively reviewed.
  • Why: Extractability without evidence can make an unsupported claim easier to repeat. Stale prices, policies, benchmarks, and product capabilities are especially risky because a fluent passage may still look current.
  • How: Trace decision-relevant claims to sources, place citations near the claim, use descriptive anchor text, and state the relevant measurement or effective date. Remove dead evidence or rewrite the claim. Change an “updated” date only after a real review changes or revalidates the content.
  • Tool: Claim-source ledger, link checker, content inventory, and subject-matter approval.
  • Done when: 100% of high-risk and decision-relevant claims have a current source or named accountable owner; zero citations lead to dead or unrelated pages; and the displayed review date matches the recorded review.

8. Freeze a representative prompt set before release

  • What: Establish a repeatable set of buyer questions used to measure mentions, citation rank, cited URLs, and provider differences before and after changes.
  • Why: Changing prompts after deployment can manufacture apparent improvement. One hand-picked answer is an anecdote because generative responses and source selection can vary between runs and providers.
  • How: Select at least 20 prompts across discovery, comparison, evaluation, and brand-specific intent. Include prompts where the brand is currently cited, mentioned without a citation, and missing. Fix wording, country, language, provider, tags, and schedule; record the mapped destination and business priority.
  • Tool: Prompt Tracking and Management in AmICited and the approved prompt-to-page map.
  • Done when: At least 20 prompts have one or more completed baseline runs; 100% retain fixed wording and settings during the observation window; each has an intended page and intent; and failed or pending runs are excluded from outcome rates rather than counted as missing.

9. Measure citations at domain, URL, prompt, and provider level

  • What: Track whether the brand is named, whether its domain is cited, which exact URL is cited, its position, and which other sources win for the same prompt.
  • Why: A domain total can hide that the wrong page is earning citations. A mention can rise while owned-source citations fall, meaning engines know the brand but trust another source for the answer.
  • How: Preserve the pre-change export, annotate the release, and compare equivalent completed runs over the agreed 28-day window. Segment by provider and prompt intent. Review full answers for important movements and separate owned citations from third-party citations that mention the brand.
  • Tool: Source and Citation Intelligence in AmICited, prompt history, and the release annotation.
  • Done when: Every priority prompt has a completed pre-change baseline and post-change observation record; cited domains and exact URLs are stored; provider differences are visible; and every claimed improvement can be reproduced from the same prompt set and date window.

10. Retest failures and sign the readiness decision

  • What: Consolidate technical, editorial, schema, and measurement results into pass, conditional pass, intentional block, or fail.
  • Why: An average score can hide a critical access failure. Readiness belongs to exact URLs and templates under a recorded policy, not to the site as an unsupported label.
  • How: Retest every corrected item from a clean request. Keep intentional policy blocks separate from defects. A conditional pass must name the exception, affected URLs, risk, approver, correction owner, and expiry date. Preserve raw evidence and the template or deployment version.
  • Tool: Readiness test matrix, issue tracker, release record, and owner sign-off.
  • Done when: Zero critical failures remain on intentionally allowed priority URLs; every other failure has an owner and due date; every exception has an expiry; and the accountable SEO, engineering, and content owners sign the same dated record.

Tools in AmICited

Use product checks as evidence inside the matrix, not as a substitute for business policy or editorial judgment.

  1. Open the Agent Accessibility audit to inspect llms.txt, accessibility structure, sitemap coverage, and the separate facts behind crawler access. Record a robots allow, a live CCBot-style request, and confirmed Common Crawl presence independently; an unknown result is not a pass or a failure.
  1. Open Prompt Tracking to import or create the frozen prompt set, select providers, country, tags, and cadence, and preserve completed baseline runs. Queued, processing, and failed runs are operational states, not “missing citation” outcomes.
  1. Open Sources to move from domain totals to exact cited pages and the prompts each page wins. Compare owned pages with third-party sources instead of assuming a brand mention came from the brand’s website.

Decision rules: what bad looks like

These are operational gates, not claims about how an engine ranks pages. Tighten them for regulated, safety-critical, or high-value content.

SignalPassWarningFail or stop
Crawler-policy coverage100% of in-scope crawler families have a recorded decisionA decision is older than its review dateAny live allow or block contradicts policy
Allowed URL fetches100% return expected canonical contentMore than one redirect hop or materially slower bot responseAny 401, 403, 429, challenge, blank main content, or unexpected noindex
llms.txt, if adopted200 plain text; all listed URLs canonical and reachableOwnership or review date missingAny broken, redirected, blocked, duplicate, or noncanonical listed URL
Priority answer coverage10 of 10 mapped prompts have a self-contained answer sectionAnswer begins after 80 words or depends on vague pronounsNo canonical destination, contradictory answer, or essential context missing
ExtractabilityAll critical facts survive plain-text and no-interaction reviewLabels are understandable only with nearby visual contextCritical fact exists only in image, video, canvas, or interaction state
Schema qualityZero syntax errors and zero visible-data conflictsApplicable template lacks an owner or fixtureMarkup invents, exaggerates, or contradicts a material fact
Prompt baselineAt least 20 fixed prompts with completed runsProvider, country, or intent coverage is unbalancedWording or settings change during comparison without restarting baseline
Outcome evidenceEquivalent completed runs compared for 28 daysToo few completed runs for the scheduled cadenceImprovement is claimed from one answer, a different prompt set, or domain totals without URL evidence

Do not blend the rows into one score. One blocked revenue-critical template is not canceled out by nine well-structured articles. Conversely, an intentional, approved training-crawler block is not a technical defect if retrieval crawlers needed for the chosen strategy can still reach the approved content.

Deliverable

Hand over a versioned table plus an evidence folder. Required fields are: audit ID, URL, template, target prompt, provider, crawler family, policy decision, fetch result, extractability result, schema result, llms.txt inclusion, baseline mention and citation, deployment date, post-change result, severity, owner, due date, retest date, evidence link, exception, expiry, and final decision.

Include the crawler register, frozen prompt export, entity and claim-source record, and release annotation. Store raw responses or rendered extracts alongside screenshots so reviewers can verify what the machine received.

The accountable owner signs one of four outcomes:

  • Pass: all intentionally allowed priority URLs clear critical access, extraction, entity, and evidence gates.
  • Conditional pass: no critical failure remains, but time-bounded noncritical exceptions are accepted by named owners.
  • Intentional block: a crawler or path is unavailable by approved policy and the expected visibility tradeoff is recorded.
  • Fail: a critical template is unreachable, misleading, contradictory, or unmeasurable; release or promotion stops until retest.

What goes wrong

Robots policy is treated as proof of access. The file says “allow,” but the CDN challenges the request. Fix this by storing both policy and a real fetch result.

Every crawler is allowed without a business owner. The SEO team optimizes discovery while legal or content owners intended to restrict training reuse. Separate crawler purposes and get an explicit decision rather than making one blanket rule.

llms.txt becomes a second sitemap. Hundreds of unprioritized URLs create noise, stale links, and competing canonical choices. Keep it curated and useful, or omit it deliberately.

The copy is shortened until it becomes wrong. Front-loaded answers lose qualifications, dates, or audience constraints in pursuit of a snippet. Keep the direct answer early, then include the conditions required for it to stand alone accurately.

FAQ schema is added to invisible or unsupported answers. Valid syntax does not make fabricated or hidden claims trustworthy. Align markup with visible content and remove properties the page cannot prove.

A homepage test is generalized to the whole site. Documentation, product, category, and JavaScript-heavy templates can behave differently under the same domain. Audit representative templates and priority destinations.

One favorable AI answer becomes the success story. The team reruns or rephrases until the brand appears, then reports the screenshot. Freeze prompts first and compare equivalent completed runs across the observation window.

Mention share is confused with citation share. The engine names the brand but cites a review site or competitor. Report brand presence and owned-source citation separately, then inspect the exact cited URL.

Next phase

Next is continuous refresh and iteration . It needs the readiness matrix, crawler register, frozen prompts, cited URLs, deployment annotation, exceptions, and review dates.

Do not repeatedly rewrite a page merely because a citation has not appeared. First recheck access, prompt validity, cited-source changes, and completed-run volume. Then choose the smallest evidence-backed intervention: technical repair, clearer passage, stronger entity support, fresher evidence, or no change.

Ready to verify GEO and AEO readiness?

Start with the SEO process , then run the Agent Accessibility audit and attach its evidence to the checklist. Readiness is complete only when the access decision is explicit, representative pages pass extraction and entity checks, and prompt and citation reporting can verify the outcome.

← All Academy tutorials

Ready to put it into practice?

Free check · 7-day trial · no credit card