---
title: "Control Web Crawlers and Crawl Budget in Vue"
description: "Configure robots.txt, sitemaps, canonicals, and redirects in Vue to control crawl budget and block unwanted AI bots."
canonical_url: "https://nuxtseo.com/learn-seo/vue/controlling-crawlers"
last_updated: "2026-10-05"
---

Web crawlers discover and fetch pages. Search engines decide which pages to index. Controlling them affects your [crawl budget](https://developers.google.com/search/docs/crawling-indexing/large-site-managing-crawl-budget): the set of URLs Google can and wants to crawl on your site. It also means guiding AI bots to your best content while blocking the ones that scrape it for training.

Most sites don't need to worry about crawl budget. Large, rapidly changing sites may need crawl-budget controls. AI crawler preferences are a separate concern.

## Types of Crawlers

**Search engines**: Index your pages for search results

- [Googlebot](https://developers.google.com/search/docs/advanced/crawling/overview-google-crawlers)
- [Bingbot](https://www.bing.com/webmasters/help/help/which-crawlers-does-bing-use-8c184ec0)
- [Applebot](https://support.apple.com/en-us/119829) (Spotlight and Siri suggestions)

**Social platforms**: Generate link previews when shared

- [FacebookExternalHit](https://developers.facebook.com/docs/sharing/webmasters/web-crawlers/)
- Twitterbot, Slackbot, Discordbot

**AI training crawlers and usage controls**: Providers use crawlers and robots.txt control tokens for different purposes

- [GPTBot](https://developers.openai.com/api/docs/bots) (OpenAI)
- [ClaudeBot](https://support.claude.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler) (Anthropic)
- [Google-Extended](https://developers.google.com/search/docs/crawling-indexing/overview-google-crawlers) (a control token for training and grounding in specified [Gemini](https://gemini.google.com) products, not a separate crawler)
- [Applebot-Extended](https://support.apple.com/en-us/119829) (a control token for Apple foundation-model training, not a separate crawler)

**AI search and user fetches**: Search crawlers build indexes. User agents can fetch pages after a user request. Provider rules differ

- [OAI-SearchBot and ChatGPT-User](https://developers.openai.com/api/docs/bots) (ChatGPT)
- [Claude-SearchBot and Claude-User](https://support.claude.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler) (Claude)
- [PerplexityBot and Perplexity-User](https://docs.perplexity.ai/guides/bots) (Perplexity)

**Malicious**: Ignore robots.txt, spoof user agents, scan for vulnerabilities. Block these at the [firewall level](/learn-seo/vue/routes-and-rendering/security), not with robots.txt.

::tip{title="Training bots vs. search bots"}
Training crawlers collect content for model development. Search crawlers support search discovery. User-triggered fetchers serve a separate role. You can block one without the other: disallow `GPTBot` in robots.txt while leaving `OAI-SearchBot` untouched if you want your site cited in ChatGPT Search.
::

## Control Mechanisms

| Mechanism                                                                                              | Use When                                                              |
| ------------------------------------------------------------------------------------------------------ | --------------------------------------------------------------------- |
| [robots.txt](/learn-seo/vue/controlling-crawlers/robots-txt)                                           | Block site sections, manage crawl budget, block AI bots               |
| [Sitemaps](/learn-seo/vue/controlling-crawlers/sitemaps)                                               | Help crawlers discover pages, especially on large sites               |
| [Meta robots](/learn-seo/vue/controlling-crawlers/meta-tags)                                           | Control indexing per page (noindex, nofollow)                         |
| [Canonical URLs](/learn-seo/vue/controlling-crawlers/canonical-urls)                                   | Consolidate duplicate content, handle URL parameters                  |
| [Redirects](/learn-seo/vue/controlling-crawlers/redirects)                                             | Send moved pages to relevant replacement URLs                         |
| [llms.txt](/learn-seo/vue/controlling-crawlers/llms-txt)                                               | Guide AI tools to your documentation (MCP servers, coding assistants) |
| [X-Robots-Tag](https://developers.google.com/search/docs/crawling-indexing/robots-meta-tag#xrobotstag) | Control non-HTML files (PDFs, images)                                 |
| [Firewall](/learn-seo/vue/routes-and-rendering/security)                                               | Block malicious bots at network level                                 |

## Quick Recipes

**Block page from indexing**: [Full guide](/learn-seo/vue/controlling-crawlers/meta-tags)

```vue [pages/admin.vue]
<script setup lang="ts">
import { useSeoMeta } from '@unhead/vue'

useSeoMeta({ robots: 'noindex, follow' })
</script>
```

**Block AI training bots**: [Full guide](/learn-seo/vue/controlling-crawlers/robots-txt)

```robots-txt [public/robots.txt]
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: Applebot-Extended
User-agent: Google-Extended
Disallow: /
```

**Fix duplicate content**: [Full guide](/learn-seo/vue/controlling-crawlers/canonical-urls)

```vue [pages/products/[id\\].vue]
<script setup lang="ts">
import { useHead } from '@unhead/vue'
import { useRoute } from 'vue-router'

const route = useRoute()

useHead({
  link: [{ rel: 'canonical', href: () => new URL(route.path, 'https://mysite.com').href }]
})
</script>
```

**Redirect moved page**: Express middleware excerpt; install Express and register this before your page handler. [Full guide](/learn-seo/vue/controlling-crawlers/redirects)

```ts [server.ts]
app.get('/old-url', (req, res) => res.redirect(301, '/new-url'))
```

## When Crawler Control Matters

Most small sites don't need to optimize crawler behavior, but it matters when:

**Crawl budget concerns**: Large sites with rapidly changing pages may need crawl inventory controls. Google's size estimates are examples, not exact thresholds. Block low-value pages (search results, filtered products, admin areas) so crawlers focus on what matters.

**Duplicate content**: URLs like `/about` and `/about/` can expose duplicate content, as can `?sort=price` variations. [Canonical tags](/learn-seo/vue/controlling-crawlers/canonical-urls) consolidate these.

**Staging environments**: Protect private staging with authentication. Use crawlable noindex when a public staging page must stay out of search. robots.txt alone does not prevent indexing.

**AI training opt-out**: GPTBot is [the most frequently blocked AI crawler](https://blog.cloudflare.com/from-googlebot-to-gptbot-whos-crawling-your-site-in-2025/), disallowed by more than 300 of the top 10,000 domains per Cloudflare's 2025 analysis. Configure each provider's training controls separately from its search controls.

**Server costs**: Bots consume CPU, and heavy pages (maps, infinite scroll, SSR) cost money per request. Blocking unnecessary crawlers reduces load.

Using Nuxt? [Nuxt SEO](/docs/nuxt-seo/getting-started/introduction) handles much of this automatically. See the [Nuxt controlling crawlers guide](/learn-seo/nuxt/controlling-crawlers) for details.

## Sitemap

See the full [sitemap](/sitemap.md) for all pages.
