Free AI Crawler Inspector
Inspect readable page content and robots.txt permissions for ChatGPT, Claude, and Perplexity agents. Copy or download text and Markdown.
Why inspect your page’s readable content?
A page can look complete in your browser while its server response contains very little readable text. Use this preview to check whether your product description, article, or key answers are present in the HTML before reviewing crawler access.
Check your main message
Compare the extracted text with the page you intended to publish. Look for your main heading, a clear description of the topic, and the details a reader needs to understand it.
Review page metadata
Confirm the title, meta description, and canonical declaration describe the right page. A readable preview and accurate metadata answer different questions; review both.
Keep a copy for comparison
Download the Markdown or plain text before updating your page. Compare it with a later inspection to see which information appears in the server response. Reports may be cached for up to one hour.
What to review when page content is missing
Start with the HTML your server returns. This simulation does not run JavaScript or evaluate external stylesheets, and it focuses on a visible main element when one exists.
Content loaded with JavaScript
Check whether important copy appears in the initial page source or arrives only after a script runs. If you want it available without JavaScript, use your framework’s server rendering or static generation for that content.
Content outside the main element
Check whether the main element contains the actual page content. Navigation, footers, forms, scripts, and hidden elements are excluded from this preview, so compare the result with the page’s HTML structure.
Login walls or server restrictions
Check whether the page loads publicly and whether automated requests receive a challenge, an error, or different HTML. The inspector requests the page as HyperGSCBot; other visitors and crawlers may receive different responses.
How to interpret AI crawler permissions
The inspector evaluates robots.txt for the final page URL after redirects. Permission describes the rule found for that URL; it does not confirm a crawler visit or an AI citation.
Review each agent separately
OAI-SearchBot and GPTBot are separate agents, and Google-Extended is an AI-use control. Search access and training policies can differ. Decide which uses you intend to permit before changing any rules.
Resolve unknown responses first
An unknown status means this check could not establish permission. Inspect the robots.txt response, redirects, or server errors before treating it as allowed or blocked.
Example: from page HTML to readable Markdown
This fictional page puts its core message in the main element. The extraction keeps that message and removes the navigation. The page title is reported separately as metadata. This is an illustration, not a live scan.
<title>Acme Projects</title>
<nav>Home · Pricing · Contact</nav>
<main>
<h1>Project management for small teams</h1>
<p>Plan projects and track deadlines in one place.</p>
</main># Project management for small teams
Plan projects and track deadlines in one place.If the paragraph is added only after JavaScript runs, it will be absent from this preview. Check the server-returned HTML to understand what is missing before changing your content.
How to inspect your page’s readable content
Preview the content a simple crawler can extract from your server-returned HTML.
Enter a public HTML page URL and click Inspect page.
Review the readable content, metadata, headings, links, and robots.txt permissions.
Copy or download the text or Markdown, then review important content missing from the preview.
Continue checking your website
If important content or metadata needs review, use the AI SEO checker to inspect page signals and optional agent discovery files.
To compare Google and Bing access with AI-use permissions, use the Robots.txt tester for the same page.
If your site publishes guidance for agents, use the llms.txt checker to review the file’s headings and links.
Frequently asked questions
No. This is HyperGSC’s crawler simulation, not a reproduction of an AI provider’s private extraction process. It reads server-returned HTML without executing JavaScript and removes common navigation, forms, hidden elements, scripts, and styles. External stylesheets are not evaluated.
Content added by JavaScript, protected by a login, or blocked from automated requests may be absent. When the HTML contains a visible main element, the preview focuses on that element; otherwise it uses the readable page content. Check the page source and your server response before making changes.
Yes. Switch between Markdown and plain text, then copy the content or download an .md or .txt file. Previews are limited to 60,000 characters and up to 100 unique links; downloads contain the same bounded preview.
No. The permissions reflect robots.txt rules for the inspected URL. They do not confirm a crawler visit, indexing, model training, or an AI citation. Different agents handle search, user-requested retrieval, and training controls.
Reports may be cached for up to one hour. The tool makes a public request as HyperGSCBot, so a server may return different content to other crawlers or visitors.
Skip
Search ConsoleJust ask your AI
Google Search Console
Google Analytics 4
SERP Analysis
Keyword Research
- Backlink Analysis
- Technical SEO
Free to start · Takes 2 minutes