Improve visibility
Domain audit and AI accessibility
Check whether AI crawlers and agents can reach, read and act on your site: robots.txt, sitemaps, llms.txt, the accessibility tree, WebMCP, agentic commerce and Common Crawl, plus a full site audit crawl.
An engine cannot cite a page it cannot fetch, and an agent cannot use a page it cannot read. The Domain audit page (Audit → Domain audit) collects the checks that decide whether your domain can be found, trusted and read by AI crawlers and agents. The Site audit page (Audit → Site audit) goes deeper: it crawls every URL in your sitemap and reports technical issues and how internal links distribute authority.
Run the domain audit first when a domain has zero visibility or when a page you expect to be cited never is. A single Disallow line for GPTBot or a CDN rule that blocks crawlers explains more missing citations than any content problem.
Agent readiness summary#
The top of the page summarizes every section in one row: llms.txt score, accessibility score, crawler access, sitemap URLs, sitemap vs. robots, WebMCP, agentic commerce and domain expiry. Your tracked competitors appear in comparison tables beside you, so you can see whether a rival is simply easier for agents to read.
robots.txt and sitemaps#
Shows whether the important search and AI crawlers are allowed to read the site, whether sitemaps are declared in robots.txt, and how many URLs those sitemaps list. The crawlers checked are:
GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot, Claude-User, Claude-SearchBot, PerplexityBot, Google-Extended, Googlebot, Bingbot, Applebot, Amazonbot, Meta-ExternalAgent and CCBot.
It also flags conflicts between your sitemap and robots.txt: sitemap URLs your own robots.txt blocks, crawlers blocked from URLs the sitemap lists, sitemap files that are themselves disallowed, Sitemap: directives pointing at another host, and sitemap files over the 50,000-URL or 50 MB limit. No robots.txt at all means every crawler is allowed by default.
llms.txt review#
/llms.txt is a Markdown file that gives language models a curated overview of your site. AmICited fetches yours and validates it against the llmstxt.org proposal: transport (served at the root over HTTPS with HTTP 200, a text content type, not your HTML app shell, UTF-8, under 128 KB), structure (one H1 first, a summary blockquote right after it, H2 sections that are link lists, no deeper headings) and link quality (absolute URLs in - [Title](URL): description form, no raw HTML, no placeholder text).
- Required checks weigh 1, recommended checks weigh 0.5, and advisory tips weigh 0 (shown as guidance, never penalized).
- Each check shows whether it passed, what it expects and an example.
- Re-check now fetches the file again (0.005 credits). Re-check competitors rescores theirs.
Worked example#
A full review runs 5 required and 19 recommended checks (plus 4 advisory tips), worth 5 + 9.5 = 14.5 points. A file that passes every required check and 15 of the recommended ones scores (5 + 7.5) ÷ 14.5 × 100 = 86.
Accessibility tree#
AI agents navigate a page through its accessibility tree (roles, names, structure), not its pixels. AmICited checks your homepage’s HTML, and you can Check URL for any other page (0.005 credits per review). The checks: page reachable, interactive elements named, no hidden interactive elements, images have alt text, form controls labeled, single H1, heading order, a main landmark, page language declared, and a document title.
WebMCP, agentic commerce and Common Crawl#
- WebMCP. Detects whether your homepage exposes tools that agents can call directly (search, add to cart, book) instead of guessing from the page. Declarative tools are listed with read or write flags; tools registered from JavaScript are noted but cannot be listed from static HTML.
- Agentic commerce. Detects whether the homepage advertises a protocol that lets agents browse, check out and pay (ACP, UCP), with a link to the manifest.
- Common Crawl. Whether CCBot can fetch the site (allowed, disallowed in robots.txt, or blocked at the CDN edge), whether the domain is in the latest index, how many URLs were captured, and your footprint over time against competitors. Common Crawl is the corpus most open training datasets are built from; it is not what GPTBot, OAI-SearchBot or PerplexityBot fetch, so read it separately from live AI search.
- Domain registration and certificate. Days until the domain registration and the TLS certificate expire, and the registrar.
Site audit#
The site audit crawls every URL in the domain’s sitemap, records each page’s technical facts and links, and computes the internal link graph.
Configure the scan
Go to Audit → Site audit and click Schedule scan. Set the pace (requests per second, concurrent requests, maximum wait per request), when to stop (maximum URLs, maximum duration, error rate), which URLs to include or exclude (up to 20 regular expressions each), whether to store links to external sites, and whether to Repeat scan every N days.
Get a price and confirm
Click Get price. The quote counts your sitemap URLs, applies your filters and plan cap, and prices the scan at 0.002 credits per URL (1,000 URLs = 2 credits). The quote is valid for one hour, and a scan cannot start unless your balance covers all of it.
Read the report
Tabs: Issues, Pages & links, Outbound domains, Runs, Server response and Crawl log. You are notified when the run finishes.
The crawler is polite by design: it halves its rate and pauses when the site answers 429 or 503, and slows down when response times double. A stopped scan keeps everything already fetched and charges only for it. Each page gets a PageRank (plain, and weighted by where the link sits: main content counts more than navigation or footer) next to its organic clicks and paid clicks; the traffic columns need Search Console or an ad platform connected.