A Next.js site on Vercel can be very easy for AI crawlers to read or completely invisible to them, and the difference is whether your content is in the HTML the server sends. The bottom line: use static generation or server rendering for every page that carries content you want cited, keep client side rendering for interactive parts only, make robots.txt and the sitemap explicit, and verify with a plain HTTP request instead of trusting the browser.
Related: our guide to answer engine optimization defines the term and lists eight steps.
Why the rendering mode decides visibility
Vercel and MERJ studied AI crawler traffic across Vercel's network, including Next.js applications and traditional web apps, and concluded that none of the major AI crawlers currently render JavaScript. Googlebot does, which is why a client rendered page can rank in Google and still be empty to ChatGPT, Claude, and Perplexity. Next.js gives you three ways to produce a page, and they behave very differently for a crawler.
| Rendering mode | What the crawler receives | Fit for content pages |
|---|---|---|
| Static generation, HTML built ahead of time | Complete HTML | Best. Fast and fully readable |
| Server side rendering, HTML built per request | Complete HTML | Good, when content changes often |
| Client side rendering, JavaScript builds the page in the browser | An empty shell and a script | Poor. The content is not in the response |
How our own site is built
The Omni Online Strategies site runs on Next.js on Vercel, and most of its content pages are plain static HTML files: more than 1,800 blog pages and about 80 top level pages sit in the project's public folder. Rewrites in the Next.js configuration map clean URLs such as /blog/example to the matching file, with a catch all for other pages, and a permanent redirect list covers renamed pages so old URLs still resolve. That choice has a cost, because a flat file has no templating and every site wide change has to be scripted across thousands of files. It also has a large benefit for AI visibility: every page is complete HTML by definition, and there is no rendering step that can fail.
The supporting pieces are ordinary. A sitemap route walks the public folder and lists every HTML page. A dynamic llms.txt route rebuilds its markdown from a Supabase table every hour. The robots.txt file explicitly allows the AI crawlers by name and ends with a Sitemap line. A build script generates the sitemap and submits to IndexNow before the Next.js build runs, and a scan API route powers the free visibility check.
Steps to set up a site this way
- Decide the rendering mode per page type. Static generation or server rendering for anything you want cited. Client rendering only for interactive widgets.
- Put the title, headings, and body text in server rendered output. Data fetched in a client component after load will not be in the response.
- Write a robots.txt that names the crawlers you want. Allow the search and user fetch bots for OpenAI, Anthropic, and Perplexity, and make a separate decision about training bots.
- Generate a sitemap from the real page list. Use accurate lastmod dates.
- Set canonical URLs and permanent redirects. Keep one URL per page and redirect renamed pages instead of deleting them.
- Add structured data in the server output. Organization, Article, and BreadcrumbList generated with the page.
- Check firewall and bot protection settings. Any challenge or block that stops a search bot removes you from answers.
- Verify with plain requests. Fetch pages with curl and read the response.
# Confirm the content is in the HTML the server sends
curl -sL https://www.example.com/blog/my-post | grep -c "<h2"
# Confirm robots.txt and the sitemap are served
curl -s https://www.example.com/robots.txt | head -20
curl -s https://www.example.com/sitemap.xml | head -20
Pitfalls we watch for
- Content fetched after load. A related links block or a pricing table loaded by a client script is not in the response, so we bake content that matters into the HTML at build time.
- Redirect chains. An extra hop between an old and a new URL slows a crawler, so we point old URLs straight at the final destination.
- Generated files that drift. A sitemap or an llms.txt that is not regenerated goes stale, so both are built by code, not by hand.
What Google says about all of this
For Google's AI features, a page must be indexed and eligible to appear with a snippet, and there are no additional technical requirements. The fundamentals it lists are the ones above: crawling allowed in robots.txt and by any CDN or hosting layer, important content available as text, and internal links.
Questions people ask
Is Next.js good for AI visibility?
It can be, if pages are statically generated or server rendered so the content is in the HTML. Vercel's analysis found the major AI crawlers do not render JavaScript, so client rendered pages are invisible to them.
Do I need a special file for AI crawlers?
Google says no additional files or markup are required for AI Overviews and AI Mode. A robots.txt that allows the search bots and an accurate sitemap are the basics.
How do I test what a crawler sees on my Next.js site?
Request the page with curl and read the raw response, or use view source. If your text is not in that output, crawlers that do not run JavaScript cannot read it.
Should I use static files or a framework for content pages?
Either works if the output is complete HTML. Flat files are simple for crawlers but need scripts for site wide changes. Framework rendering gives templating at the cost of making sure the render mode is right.