AI crawlers read the HTML your server sends and nothing else, so the structure of that HTML decides what an answer engine can use. The bottom line: put your real content in the raw HTML, give each page one H1 and a clear H2 per question, put the answer first, use lists and tables where the content is a list or a table, and test by fetching the page without a browser.
For a full definition, read our answer engine optimization guide.
What AI crawlers actually receive
Vercel and MERJ analyzed the crawler traffic across Vercel's network, including OpenAI's GPTBot at 569 million requests in the month studied and Anthropic's crawler at 370 million, and found that none of the major AI crawlers render JavaScript. Googlebot, and Gemini through it, does, and so does Apple's crawler. The practical result is that a page built by a JavaScript framework in the browser can look complete to a person and arrive as an empty shell to ChatGPT, Claude, and Perplexity. The example of a restaurant site with empty containers that fill in after the page loads is exactly this failure.
Which crawler does what
| Company | Crawler | Purpose | robots.txt note |
|---|---|---|---|
| OpenAI | GPTBot | Content that may train models | Independent of the search crawler |
| OpenAI | OAI-SearchBot | Surfacing sites in ChatGPT search | Block it and you are not cited as a source in ChatGPT search |
| OpenAI | ChatGPT-User | Pages a user asks ChatGPT to open | OpenAI says robots.txt may not apply |
| Anthropic | ClaudeBot | Training data | Blocking it does not block the search bot |
| Anthropic | Claude-SearchBot | Search indexing | Anthropic says blocking may reduce visibility in results |
| Anthropic | Claude-User | Pages a user asks Claude to fetch | Honors robots.txt, per Anthropic |
| Perplexity | PerplexityBot | Search index | Honors robots.txt |
| Perplexity | Perplexity-User | Pages a user asks about | Perplexity says it generally does not apply |
How to structure the HTML
- Put the content in the server response. View the page source, not the browser inspector. If your text is not there, a crawler will not see it.
- Use one H1 and an H2 for each question or section. Our scan requires exactly one H1 and at least three H2 sections.
- Write the answer under the heading, first. The AirOps finding that 44 percent of citations come from the first 30 percent of a page, and the "Lost in the Middle" result that models use the start and end of their input best, both favor answers near the top.
- Use real lists and tables. An ordered list for steps, a table for comparisons. They give a system clean units to extract.
- Do not hide content behind tabs or accordions that load on click. If the text is only fetched after an interaction, it is not in the HTML.
- Use real links. Anchor tags with href values, not click handlers, so a crawler can follow them.
- Serve the page to everyone. Content behind a login, a cookie wall, or a bot challenge cannot be read. Check that your firewall is not challenging the search bots.
- Write descriptive titles and alt text. They are part of what is read.
Test it in one minute
# What an AI crawler receives: the raw HTML, with no JavaScript run
curl -sL -A "Mozilla/5.0 (compatible; OAI-SearchBot/1.0; +https://openai.com/searchbot)" \
https://www.example.com/services/emergency-plumbing | head -c 4000
# Count the words in the raw HTML body
curl -sL https://www.example.com/services/emergency-plumbing | sed 's/<[^>]*>/ /g' | wc -w
If the word count is under about 450, the page is thin as far as a crawler is concerned, and under 30 words it is effectively an empty shell. Those are the thresholds our scan uses.
What Google says
For Google's own AI features, a page needs to be indexed and eligible to appear with a snippet, and there are no additional technical requirements. Google's guidance lists the fundamentals: crawling allowed in robots.txt and by any CDN or hosting layer, important content available as text, and internal links so pages can be found.
Questions people ask
Do AI crawlers run JavaScript?
Vercel's analysis found that the major AI crawlers, including OpenAI's, Anthropic's, and Perplexity's, do not render JavaScript. Googlebot and therefore Gemini do.
How do I check what an AI crawler sees?
Fetch the page with curl and read the raw HTML, or use view source in your browser. If your content is not in that HTML, it is not visible to crawlers that do not run JavaScript.
Should the answer be at the top of the page?
Yes. AirOps found 44 percent of ChatGPT citations came from the first 30 percent of a page, and research shows language models use the start and end of their input better than the middle.
Does blocking GPTBot stop ChatGPT from citing me?
No. GPTBot is the training crawler. OAI-SearchBot controls ChatGPT search, and the two are independent settings.