Next.js and React SEO for AI search: how to make your content readable without JavaScript
Published 7 October 20269 min readBy the SmoothSeen editorial team
For AI to read a site built with React, the main content has to arrive in the HTML the server sends. Google runs JavaScript in a second phase, but OpenAI, Anthropic and Perplexity do not document their crawlers doing so. In Next.js, Server Components already solve this; a Vite SPA needs prerendering or a framework with SSR.
Key points
- Google runs JavaScript in a second phase, but OpenAI, Anthropic and Perplexity do not document their crawlers doing so; a Vercel study from December 2024 saw none of them execute it.
- In Next.js with the App Router, pages are Server Components and arrive with their text in the HTML; the risk lies in loading content in a useEffect or with ssr false.
- In a test with Next.js 16.4, slow metadata in generateMetadata ended up in the body, not the head, for OAI-SearchBot, ClaudeBot and PerplexityBot.
- A React SPA built with Vite needs prerendering or a move to a framework with server rendering; Google calls dynamic rendering a workaround, not a solution.
To check it on your own site: AI visibility audit
On this page
A SPA (single-page application) sends an almost empty HTML document and builds the content in the browser with JavaScript. For a person there is no difference. For a crawler that reads the HTML and does not run the code, the page says "Loading…". This guide explains which crawler sees what, how Next.js handles it (version 16.4, released on 6 October 20261) and what to do if your site is plain React with Vite.
Who runs JavaScript and who does not?
Google processes pages in three phases: crawling, rendering and indexing. Pages wait in a render queue that can last seconds or longer, and then a headless Chromium runs the JavaScript2. Even so, Google recommends server-side rendering or prerendering because it makes the site faster and because not all bots can run JavaScript2.
AI crawlers are those bots. Neither OpenAI, nor Anthropic, nor Perplexity says in its crawler documentation that it runs JavaScript345. The most cited data point is a Vercel study from 17 December 2024, which analysed a month of crawler requests on its network (569 million from GPTBot and 370 million from Claude, among others). Its conclusions, for that period6:
- Crawler
- Googlebot (and Gemini, which uses its infrastructure)
- Did it run JavaScript in the Vercel study?
- Yes
- Crawler
- Applebot
- Did it run JavaScript in the Vercel study?
- Yes
- Crawler
- OAI-SearchBot, ChatGPT-User, GPTBot
- Did it run JavaScript in the Vercel study?
- No: they downloaded JS files without executing them
- Crawler
- ClaudeBot
- Did it run JavaScript in the Vercel study?
- No
- Crawler
- PerplexityBot
- Did it run JavaScript in the Vercel study?
- No
It is a study by a vendor, from 2024, and crawlers change. But no official documentation contradicts it, so the working rule is: if a piece of text is not in the initial HTML, assume ChatGPT, Claude and Perplexity do not see it. The wider picture, beyond React, is in the guide to JavaScript SEO and AI crawlers.
Next.js App Router: what arrives in the HTML and what does not
In the App Router, pages and layouts are Server Components by default: they read data on the server and Next.js generates the HTML with the content7. Client Components, the ones marked 'use client', are also prerendered to HTML on the first load7. The problem is not 'use client': it is fetching the text from the browser.
We tested each pattern in an application running Next.js 16.4.0 in production mode, requesting the pages with each crawler's user agent:
- Pattern
page.tsxas a Server Component that reads the data- Does the text arrive in the HTML?
- Yes
- Pattern
- Client Component that receives the text through props
- Does the text arrive in the HTML?
- Yes, prerendered on the server
- Pattern
- Client Component that fetches the text in a
useEffect - Does the text arrive in the HTML?
- No: the HTML says "Loading…"
- Pattern
- Component loaded with
dynamic(..., { ssr: false }) - Does the text arrive in the HTML?
- No: it only renders in the browser8
This is the anti-pattern in the third row. GPTBot received <article>Loading…</article> and nothing else:
'use client'
import { useEffect, useState } from 'react'
// Anti-pattern: the main text arrives later, from the browser
export default function Page() {
const [text, setText] = useState('')
useEffect(() => {
fetch('/api/article')
.then((r) => r.json())
.then((d) => setText(d.body))
}, [])
return <article>{text || 'Loading…'}</article>
}The fix is to read the data in the page's Server Component and pass only the interactive parts, such as a "like" button, to a Client Component. Keep ssr: false for things nobody needs to read: a map, a chat widget, an editor.
Metadata and canonical with generateMetadata
Metadata is declared with the metadata object or the generateMetadata function, which only work in Server Components9. Set the base URL once in the root layout with metadataBase: new URL('https://www.example.com'), and each page can then use relative paths, including for the canonical9. Here is an article page with title, description, canonical and structured data:
import type { Metadata } from 'next'
import { getPost } from '@/lib/posts'
type Props = { params: Promise<{ slug: string }> }
export async function generateMetadata({ params }: Props): Promise<Metadata> {
const { slug } = await params
const post = await getPost(slug)
return {
title: post.title,
description: post.summary,
alternates: { canonical: `/blog/${slug}` },
}
}
export default async function Page({ params }: Props) {
const { slug } = await params
const post = await getPost(slug)
const jsonLd = {
'@context': 'https://schema.org',
'@type': 'BlogPosting',
headline: post.title,
datePublished: post.publishedAt,
dateModified: post.updatedAt,
}
return (
<article>
<script
type="application/ld+json"
dangerouslySetInnerHTML={{
__html: JSON.stringify(jsonLd).replace(/</g, '\\u003c'),
}}
/>
<h1>{post.title}</h1>
<p>{post.summary}</p>
<div>{post.body}</div>
</article>
)
}The JSON-LD block follows the official Next.js guide: a native <script> tag on the page, with the < character replaced by its Unicode escape to prevent code injection10. Google also accepts JSON-LD in the <body>11.
The metadata streaming catch
Since version 15.2, if generateMetadata is slow, Next.js sends the page first and appends the metadata afterwards, at the end of the <body>. It only waits to put them in the <head> for a list of bots that do not run JavaScript, the htmlLimitedBots option9. The default list, checked in the Next.js source code, includes Bingbot, DuckDuckBot, Applebot and Twitterbot, but no AI crawler12.
We tested it with a generateMetadata that took a second and a half: OAI-SearchBot, GPTBot, ClaudeBot, Claude-SearchBot and PerplexityBot received the <title> and description inside the <body>; Bingbot got them in the <head>. No AI crawler documents whether it reads metadata outside the <head>. There are two ways out:
- Make sure metadata is not slow: if the data is cached or the page is prerendered, metadata goes into the initial HTML9. In the same test, with fast data, every bot received it in the
<head>. - Turn off metadata streaming with the setting Next.js documents, at the cost of a slightly slower initial response for everyone12:
import type { NextConfig } from 'next'
const config: NextConfig = {
htmlLimitedBots: /.*/,
}
export default configIf you would rather keep your own list, bear in mind that htmlLimitedBots replaces the default list rather than extending it12.
robots.ts and sitemap.ts
Next.js generates robots.txt and sitemap.xml from two files at the root of app1314. This app/robots.ts lets everyone through, including Googlebot and AI search engines, and only blocks model training:
import type { MetadataRoute } from 'next'
export default function robots(): MetadataRoute.Robots {
return {
rules: [
// Everyone else, including Googlebot and AI search engines
{ userAgent: '*', allow: '/', disallow: '/admin/' },
// Model training: blocking it does not remove you from answers
{
userAgent: ['GPTBot', 'ClaudeBot', 'Google-Extended', 'Applebot-Extended', 'CCBot'],
disallow: '/',
},
],
sitemap: 'https://www.example.com/sitemap.xml',
}
}OpenAI separates GPTBot, used for training, from OAI-SearchBot, which decides whether you appear in ChatGPT search3, and Anthropic does the same with ClaudeBot versus Claude-SearchBot and Claude-User4. Google-Extended does not affect Google Search15. The full breakdown is in the guide to AI crawlers and robots.txt.
And app/sitemap.ts, built from the articles:
import type { MetadataRoute } from 'next'
import { getPosts } from '@/lib/posts'
export default async function sitemap(): Promise<MetadataRoute.Sitemap> {
const posts = await getPosts()
return [
{ url: 'https://www.example.com', lastModified: '2026-10-01' },
...posts.map((post) => ({
url: `https://www.example.com/blog/${post.slug}`,
lastModified: post.updatedAt,
})),
]
}The Next.js documentation includes changeFrequency and priority in its examples, but Google ignores both; it does use lastmod if it is accurate16. Both files compiled without type errors on Next.js 16.4.0 and produced this robots.txt:
User-Agent: *
Allow: /
Disallow: /admin/
User-Agent: GPTBot
User-Agent: ClaudeBot
User-Agent: Google-Extended
User-Agent: Applebot-Extended
User-Agent: CCBot
Disallow: /
Sitemap: https://www.example.com/sitemap.xmlReact with Vite (SPA): three ways out
If your site is a React SPA built with Vite, the HTML a crawler receives is an empty <div id="root">. React recommends starting new projects with a framework, and names Next.js and React Router17. Your options:
- Prerender at build time. You generate the HTML for each route when you build. If you use React Router in framework mode, listing the routes in
react-router.config.tsis enough18:
import type { Config } from "@react-router/dev/config";
export default {
async prerender() {
return ["/", "/about", "/contact"];
},
} satisfies Config;- Move to server rendering: Next.js, React Router with
ssr: true18, or Astro, which renders on the server and ships zero JavaScript by default, with React components where you need them19. - Do not use dynamic rendering (serving prerendered HTML only to bots). Google says it was a workaround and not a long-term solution, and recommends server-side rendering, static rendering or hydration instead20.
How to check it
- Request the page as an AI crawler and look for a sentence from your content and for the title:
curl -s -A "Mozilla/5.0 (compatible; OAI-SearchBot/1.3; +https://openai.com/searchbot)" https://www.example.com/blog/hello-world | grep -o "<title>.*</title>"
curl -s -A "Mozilla/5.0 (compatible; OAI-SearchBot/1.3; +https://openai.com/searchbot)" https://www.example.com/blog/hello-world | grep -c "a sentence from your article"- "View page source" versus "Inspect". The page source (Ctrl+U) is the HTML that arrives from the server; the inspector shows the page after JavaScript has run. If the text is only in the inspector, AI crawlers do not see it.
- URL Inspection in Search Console. The live test shows a screenshot of how Google's tool sees the page, plus the HTML, headers and JavaScript console messages21. It tells you what Google sees after rendering, not what a crawler that does not run JavaScript sees.
What SmoothSeen does with this
In its AI visibility analysis, SmoothSeen measures how much text arrives in the HTML without JavaScript compared with the page once rendered in a browser, and whether the answer comes first. It also reads your robots.txt, separating search crawlers from training crawlers, and requests the page with each AI crawler's user agent to detect a 403. In the SEO analysis it checks titles, canonical, sitemap and structured data.
What to do this week
Pick your most important page and run the curl above with a sentence from its second paragraph. If it does not show up, find the component that loads it in the browser and move it to the server. If you want to measure the gap between the HTML and the rendered page, analyse it with SmoothSeen.
Frequently asked questions
Does Google index websites built with React?
Yes. Googlebot runs JavaScript with a headless Chromium in a rendering phase after crawling, and the page may wait in the queue for anything from a few seconds to longer. Even so, Google recommends server-side rendering or prerendering, because the site loads faster and because other bots do not run JavaScript.
Does ChatGPT run JavaScript when it reads a page?
OpenAI does not say so in its crawler documentation. The Vercel study from December 2024 found that OAI-SearchBot, ChatGPT-User and GPTBot downloaded JavaScript files without executing them. Until documentation says otherwise, work on the basis that they only read the HTML your server sends.
Does 'use client' in Next.js hurt SEO?
Not on its own. Client Components are also prerendered to HTML on the server on first load, so text they receive through props reaches crawlers. What hurts is fetching the content from the browser, with a fetch inside useEffect, or loading the component with ssr false.
Does dynamic rendering help bots see my SPA?
Google describes it as a workaround that is not a long-term solution and recommends server-side rendering, static rendering or hydration instead. It also forces you to maintain a list of bots and two versions of every page. Prerendering at build time or moving to a framework with SSR solves the same problem for every visitor.
Sources
- 1Next.js 16.4, Next.js Blog (Vercel), 6 October 2026.
- 2Understand JavaScript SEO basics, Google Search Central, updated 4 March 2026.
- 3Overview of OpenAI Crawlers, OpenAI, accessed 7 October 2026.
- 4Does Anthropic crawl data from the web, and how can site owners block the crawler?, Anthropic, updated 7 April 2026.
- 5Perplexity Crawlers, Perplexity, accessed 7 October 2026.
- 6The rise of the AI crawler, Giacomo Zecchini, Alice Alexandra Moore, Malte Ubl and Ryan Siddle, Vercel, 17 December 2024.
- 7Server and Client Components, Next.js documentation, updated 5 October 2026.
- 8How to lazy load Client Components and libraries, Next.js documentation, updated 10 March 2026.
- 9generateMetadata, Next.js documentation, updated 19 August 2026.
- 10How to implement JSON-LD in your Next.js application, Next.js documentation, updated 2 March 2026.
- 11Introduction to structured data markup in Google Search, Google Search Central, updated 10 December 2025.
- 12htmlLimitedBots, Next.js documentation, updated 3 October 2025; default list in html-bots.ts, accessed 7 October 2026.
- 13robots.txt, Next.js documentation, updated 1 May 2026.
- 14sitemap.xml, Next.js documentation, updated 18 August 2026.
- 15Google's common crawlers, Google Crawling Infrastructure, updated 14 July 2026.
- 16Build and submit a sitemap, Google Search Central, updated 8 July 2026.
- 17Creating a React App, React documentation, accessed 7 October 2026.
- 18Rendering Strategies, React Router documentation, accessed 7 October 2026.
- 19Why Astro?, Astro documentation, accessed 7 October 2026.
- 20Dynamic rendering as a workaround, Google Search Central, updated 10 December 2025.
- 21URL Inspection tool, Search Console Help, accessed 7 October 2026.
How to cite this article
SmoothSeen. (2026, October 7). Next.js and React SEO for AI search: how to make your content readable without JavaScript. https://smoothseen.com/en/blog/nextjs-react-seo-ai-search/
Keep reading
AI search optimization: a guide to AEO and GEO for getting cited by ChatGPT, Gemini and Google
What AI search optimization (AEO and GEO) is, how ChatGPT, Gemini and Google pick sources, which bots to allow, plus a block-by-block checklist.
AEO vs GEO vs SEO: what each one is and how they differ
AEO vs GEO vs SEO: what each one aims for, where the result appears, which signals matter and how each is measured. With a table and an example.
AI crawlers and robots.txt: how to block GPTBot without dropping out of AI answers
Which AI crawlers OpenAI, Anthropic, Google, Perplexity, Apple, Meta and Amazon use, which to block in robots.txt and how to see if your server stops them.