---
title: "How to Set Up robots.txt in Nuxt"
description: "Generate robots.txt automatically with the Nuxt Robots module, and avoid the mistakes that block Googlebot or leak your admin paths."
canonical_url: "https://nuxtseo.com/learn-seo/nuxt/controlling-crawlers/robots-txt"
last_updated: "2026-07-16"
---

<key-takeaways>

- robots.txt is advisory: crawlers can ignore it, so never rely on it for security
- The Nuxt Robots module generates an environment-aware file automatically, blocking staging without extra config
- Content-Usage and Content-Signal directives let you allow search indexing while opting out of AI training
- Include a sitemap reference and keep your search-engine rules separate from your AI-crawler rules

</key-takeaways>

The `robots.txt` file controls which parts of your site crawlers can access. [Officially adopted as RFC 9309](https://datatracker.ietf.org/doc/html/rfc9309) in September 2022, it's used to [manage crawl budget](https://developers.google.com/search/docs/crawling-indexing/large-site-managing-crawl-budget) on large sites and to tell AI crawlers whether they can use your content for training or real-time answers.

## Quick Setup

### Static robots.txt

For simple static rules, add the file in your public directory:

```dir
public/
  robots.txt
```

Add your rules:

```robots-txt [robots.txt]
# Allow all crawlers
User-agent: *
Disallow:

# Optionally point to your sitemap
Sitemap: https://mysite.com/sitemap.xml
```

### Server Route

For custom dynamic generation, create a server route:

```ts [server/routes/robots.txt.ts]
export default defineEventHandler((event) => {
  const isDev = process.env.NODE_ENV !== 'production'
  const robots = isDev
    ? 'User-agent: *\nDisallow: /'
    : 'User-agent: *\nDisallow:\nSitemap: https://mysite.com/sitemap.xml'

  setHeader(event, 'Content-Type', 'text/plain')
  return robots
})
```

### Automatic Generation with Module

For environment-aware generation, including automatic staging blocks, use the Nuxt Robots module:

<module-card className="w-1/2" slug="robots">



</module-card>

Install the module:

```bash
npx nuxi@latest module add robots
```

The module generates `robots.txt` with zero config. For environment-specific rules:

```ts [nuxt.config.ts]
export default defineNuxtConfig({
  modules: ['@nuxtjs/robots'],
  robots: {
    disallow: process.env.NODE_ENV !== 'production' ? '/' : undefined
  }
})
```

## Robots.txt Syntax

The `robots.txt` file consists of directives grouped by user agent. Google [uses the most specific matching rule](https://developers.google.com/search/docs/crawling-indexing/robots/robots_txt) based on path length:

```robots-txt [robots.txt]
# Define which crawler these rules apply to
User-agent: *

# Block access to specific paths
Disallow: /admin

# Allow access to specific paths (optional, more specific than Disallow)
Allow: /admin/public

# Point to your sitemap
Sitemap: https://mysite.com/sitemap.xml
```

### User-agent

The `User-agent` directive specifies which crawler the rules apply to:

```robots-txt [robots.txt]
# All crawlers
User-agent: *

# Just Googlebot
User-agent: Googlebot

# Multiple specific crawlers
User-agent: Googlebot
User-agent: Bingbot
Disallow: /private
```

Common crawler user agents:

- [Googlebot](https://developers.google.com/search/docs/advanced/crawling/overview-google-crawlers): Google's search crawler, roughly half of the AI-and-search-crawler traffic [Cloudflare measured in May 2025](https://blog.cloudflare.com/from-googlebot-to-gptbot-whos-crawling-your-site-in-2025/)
- [Google-Extended](https://developers.google.com/crawling/docs/crawlers-fetchers/google-common-crawlers): not a separate crawler, a control token that governs whether Googlebot-crawled content can train [Gemini](https://gemini.google.com); blocking it doesn't affect Search
- [GPTBot](https://developers.openai.com/api/docs/bots): OpenAI's training crawler, about 7.7% of that same May 2025 sample
- [ClaudeBot](https://support.claude.com/en/articles/9906653-claude-bot-and-crawling): Anthropic's training crawler
- [CCBot](https://commoncrawl.org/ccbot): Common Crawl's dataset builder, widely blocked because its data feeds many third-party AI models
- [Bingbot](https://ahrefs.com/seo/glossary/bingbot): Microsoft's search crawler
- [FacebookExternalHit](https://developers.facebook.com/docs/sharing/webmasters/web-crawlers/): Facebook's link preview crawler

Not every [OpenAI](https://openai.com) bot needs blocking. [ChatGPT-User](https://developers.openai.com/api/docs/bots) only fires when a person asks [ChatGPT](https://chatgpt.com) about your page live; it doesn't crawl proactively for training the way GPTBot does.

### Allow / Disallow

The `Allow` and `Disallow` directives control path access:

```robots-txt [robots.txt]
User-agent: *
# Block all paths starting with /admin
Disallow: /admin

# Block a specific file
Disallow: /private.html

# Block files with specific extensions
Disallow: /*.pdf$

# Block URL parameters
Disallow: /*?*
```

Wildcards supported ([RFC 9309](https://datatracker.ietf.org/doc/html/rfc9309)):

- `*`: matches zero or more characters
- `$`: matches the end of the URL
- Paths are case-sensitive and relative to domain root

### Sitemap

The `Sitemap` directive tells crawlers where to find your [sitemap.xml](/learn-seo/nuxt/controlling-crawlers/sitemaps):

```robots-txt [robots.txt]
Sitemap: https://mysite.com/sitemap.xml

# Multiple sitemaps
Sitemap: https://mysite.com/products-sitemap.xml
Sitemap: https://mysite.com/blog-sitemap.xml
```

The [Nuxt Sitemap module](/docs/sitemap/getting-started/introduction) automatically adds the sitemap URL to your `robots.txt`.

### Crawl-Delay (Non-Standard)

`Crawl-Delay` is not part of [RFC 9309](https://datatracker.ietf.org/doc/html/rfc9309). Google ignores it. Bing and Yandex support it:

```robots-txt [robots.txt]
User-agent: Bingbot
Crawl-delay: 10  # seconds between requests
```

For Google, you [manage crawl rate in Search Console](https://developers.google.com/search/docs/crawling-indexing/large-site-managing-crawl-budget).

## Security: Why robots.txt Fails

[Robots.txt is not a security mechanism](https://developer.mozilla.org/en-US/docs/Web/Security/Practical_implementation_guides/Robots_txt). Malicious crawlers ignore it, and listing paths in `Disallow` [reveals their location to attackers](https://www.searchenginejournal.com/robots-txt-security-risks/289719/).

A common mistake:

```robots-txt
# ❌ Advertises your admin panel location
User-agent: *
Disallow: /admin
Disallow: /wp-admin
Disallow: /api/internal
```

Use [proper authentication](https://developers.google.com/search/docs/crawling-indexing/control-what-you-share) instead. See our [security guide](/learn-seo/nuxt/routes-and-rendering/security) for details.

## Crawling vs Indexing

Blocking a URL in `robots.txt` prevents crawling but [doesn't prevent indexing](https://developers.google.com/search/docs/crawling-indexing/robots/intro). If other sites link to the URL, Google can still index it without crawling, showing the URL with no snippet. Use [meta robots tags](/learn-seo/nuxt/controlling-crawlers/meta-tags) for page-level indexing control.

To prevent indexing:

- Use [`noindex` meta tag](/learn-seo/nuxt/controlling-crawlers/meta-tags) (requires allowing crawl)
- Use password protection or authentication
- Return 404/410 status codes

Don't block pages with `noindex` in `robots.txt`. Google can't see the tag if it can't crawl.

## Common Mistakes

### 1. Blocking JavaScript and CSS

[Google needs JavaScript and CSS to render pages](https://developers.google.com/search/docs/crawling-indexing/javascript/javascript-seo-basics). Blocking them breaks indexing:

```robots-txt [robots.txt]
# ❌ Prevents Google from rendering your Nuxt app
User-agent: *
Disallow: /assets/
Disallow: /*.js$
Disallow: /*.css$
```

Nuxt apps are JavaScript-heavy. Never block `.js`, `.css`, or `/assets/` from Googlebot.

### 2. Blocking Dev Sites in Production

Copy-pasting a dev `robots.txt` to production blocks all crawlers:

```robots-txt [robots.txt]
# ❌ Accidentally left from staging
User-agent: *
Disallow: /
```

The [Nuxt Robots module](/docs/robots/getting-started/introduction) handles this automatically based on environment.

Also easy to get wrong: confusing `robots.txt` with `noindex`. Blocking a page in `robots.txt` doesn't remove it from search results, robots.txt controls crawling, not indexing. Use [`noindex` meta tags](/learn-seo/nuxt/controlling-crawlers/meta-tags) to deindex a page instead.

To confirm your robots.txt is live and correct:

1. Visit `https://yoursite.com/robots.txt` to confirm it loads
2. Run it through the [Google Search Console robots.txt tester](https://search.google.com/search-console/robots-txt), which validates syntax and tests individual URLs
3. Check server logs for a 200 status on `/robots.txt` to confirm crawlers can access it

## Common Patterns

### Allow Everything (Default)

```robots-txt
User-agent: *
Disallow:
```

### Block Everything

Useful for staging or development environments.

```robots-txt
User-agent: *
Disallow: /
```

See our [security guide](/learn-seo/nuxt/routes-and-rendering/security) for more on environment protection.

### Block AI Training Crawlers

GPTBot is one of the most blocked AI crawlers: [Cloudflare's 2025 analysis](https://blog.cloudflare.com/from-googlebot-to-gptbot-whos-crawling-your-site-in-2025/) found it fully disallowed by 250 domains and partially disallowed by another 62. Blocking AI training bots doesn't affect search rankings:

```robots-txt
# Block AI model training (doesn't affect Google search)
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: CCBot
User-agent: Google-Extended
Disallow: /
```

`Google-Extended` doesn't have its own crawler: it's a control token that governs whether Googlebot-crawled content trains Gemini, and blocking it doesn't affect Search visibility.

### AI Directives: Content-Usage & Content-Signal

Blocking user agents outright isn't always what you want. Content-Usage and Content-Signal let you express granular preferences about how AI systems use your content without blocking crawlers entirely: stay indexed for search while opting out of model training.

- **Content-Usage** (IETF AI Preferences): a robots.txt directive and HTTP header using `y`/`n` values for `train-ai`
- **Content-Signal** (Cloudflare): uses `yes`/`no` values for `search`, `ai-input`, and `ai-train`

```robots-txt [robots.txt]
User-agent: *
Allow: /

# Allow AI for real-time answers, block training
Content-Usage: train-ai=n
Content-Signal: search=yes, ai-input=yes, ai-train=no
```

The [Nuxt Robots module](/docs/robots/guides/ai-directives) configures these signals programmatically in `nuxt.config.ts` instead of you hand-writing the directives:

```ts [nuxt.config.ts]
export default defineNuxtConfig({
  robots: {
    groups: [{
      userAgent: '*',
      allow: '/',
      contentUsage: {
        'train-ai': 'n'
      },
      contentSignal: {
        'ai-train': 'no',
        'ai-input': 'yes',
        'search': 'yes'
      }
    }]
  }
})
```

<tip>

Allow `ai-input` and real-time AI tools like [Perplexity](https://perplexity.ai) or ChatGPT Search can cite you. Block it and you drop out of AI-generated answers entirely. See [AI-optimized content](/learn-seo/nuxt/launch-and-listen/ai-optimized-content) for more on visibility in AI search results.

</tip>

<warning>

AI directives rely on voluntary compliance: crawlers can ignore them, so combine them with User-agent blocks for stronger protection.

</warning>

### Block Search, Allow Social Sharing

For private sites where you still want [link previews](/learn-seo/nuxt/mastering-meta/open-graph):

```robots-txt
# Block search engines
User-agent: Googlebot
User-agent: Bingbot
Disallow: /

# Allow social link preview crawlers
User-agent: facebookexternalhit
User-agent: Twitterbot
User-agent: Slackbot
Allow: /
```

### Optimize Crawl Budget for Large Sites

If you have 10,000+ pages, [block low-value URLs](https://developers.google.com/search/docs/crawling-indexing/large-site-managing-crawl-budget) to focus crawl budget on important content:

```robots-txt
User-agent: *
# Block internal search results
Disallow: /search?
# Block infinite scroll pagination
Disallow: /*?page=
# Block filtered/sorted product pages
Disallow: /products?*sort=
Disallow: /products?*filter=
# Block print versions
Disallow: /*/print
```

Sites under 1,000 pages don't need crawl budget optimization.

## Checklist

<checklist id="nuxt-robots-txt">

- robots.txt lives at your domain root and returns 200
- Sitemap directive points to your sitemap.xml
- JavaScript, CSS, and `/assets/` stay unblocked for Googlebot
- Staging and dev environments block all crawlers automatically
- Sensitive paths use authentication, not `Disallow`, to stay hidden
- AI training bots are blocked with User-agent rules or Content-Usage/Content-Signal directives

</checklist>

Building with Vue instead? [Robots.txt in Vue →](/learn-seo/vue/controlling-crawlers/robots-txt)
