How to Configure robots.txt for AI Crawlers in 2026 (GPTBot, PerplexityBot, ClaudeBot)

SEO
GEO
robots.txt
AI crawlers

A practical guide to configuring robots.txt for AI crawlers in 2026 - the difference between training bots and citation bots, and how to allow one without the other.

If you have a robots.txt file, there's a good chance it still treats every crawler the same way: one User-agent: * block, allow or disallow everything. That made sense when the only thing reading your site was Googlebot. It doesn't hold up anymore, because "an AI crawler" isn't one thing.

Training Bots vs. Search Bots: Not the Same Job

AI companies run separate crawlers for separate jobs, and conflating them is the most common mistake I see in robots.txt files right now:

  • Training bots (GPTBot, ClaudeBot, Google-Extended, CCBot) collect content to train future models. Allowing these gives your content away for free, with no guarantee it ever gets attributed back to you.
  • Search/citation bots (OAI-SearchBot, PerplexityBot, Claude-SearchBot) crawl live, specifically to answer a real query someone just asked, and cite the source. This is the traffic-driving, attribution-carrying half of the equation.

A single wildcard rule can't tell these apart. If you disallow *, you block both. If you allow *, you allow both. Neither is obviously right, which is exactly why treating it as one decision is the mistake.

The configuration I'd point most people toward, and the one that's become the most common single setup among sites that care about this, blocks the training crawlers and explicitly allows the citation crawlers:

User-agent: GPTBot Disallow: / User-agent: ClaudeBot Disallow: / User-agent: Google-Extended Disallow: / User-agent: CCBot Disallow: / User-agent: OAI-SearchBot Allow: / User-agent: PerplexityBot Allow: / User-agent: Claude-SearchBot Allow: / User-agent: * Allow: / Sitemap: https://yourdomain.com/sitemap.xml

This isn't the only reasonable answer. If you're fine with your content training future models, in exchange for whatever visibility that might eventually bring, a blanket allow is simpler and not wrong. It's a genuine trade-off, not a technical bug to be fixed, and it's worth actually deciding rather than defaulting into it by not thinking about it.

Why This Matters More Than llms.txt

A lot of 2026 GEO advice leads with adding an llms.txt file. I'd deprioritise it. Google's own guidance from May 2026 states plainly that it isn't used for AI Overviews or AI Mode, and in practice the major AI search crawlers skip it and read your HTML directly instead. Where it has found a real use is coding-agent tools like Cursor and Claude Code, which do fetch it routinely when pointed at documentation sites. Useful for a very different audience, not the one most sites writing about GEO actually have.

Getting the robots.txt split right, on the other hand, directly controls whether the bots that generate real citations can even reach your content in the first place. It's a five-minute change with a much clearer mechanism than most of what gets recommended.

Checking It Worked

Once it's deployed, load yourdomain.com/robots.txt directly and confirm the bot-specific blocks are there. Most parsers apply the most specific matching User-agent block rather than the first one in the file, so a bot-specific rule further down still takes precedence over an earlier wildcard rule, but keeping them clearly separated (as above) avoids relying on that behaviour being implemented consistently everywhere.

Frequently asked questions

What's the difference between GPTBot and OAI-SearchBot?

GPTBot scrapes content to train OpenAI's foundation models. OAI-SearchBot is a completely separate crawler that indexes pages specifically to answer live ChatGPT Search queries and generate citations. Blocking GPTBot stops your content being used for training; blocking OAI-SearchBot stops it appearing in ChatGPT's search results and citations. They need separate directives, not one blanket rule.

Will blocking training bots hurt my AI search visibility?

No, they're different bots doing different jobs. Blocking GPTBot, ClaudeBot or Google-Extended (all training crawlers) doesn't affect whether OAI-SearchBot, PerplexityBot or Claude-SearchBot (the citation/answer crawlers) can still index and cite your pages, as long as you allow those separately.

Is there a standard, recommended robots.txt setup for 2026?

The most common configuration among sites that want AI visibility without giving away free training data is to block the training bots (GPTBot, ClaudeBot, Google-Extended, CCBot) and explicitly allow the search/citation bots (OAI-SearchBot, PerplexityBot, Claude-SearchBot). Roughly 30% of top sites now run this split configuration rather than a single blanket rule.

Do I need an llms.txt file as well?

Probably not for AI search visibility specifically. Google's own May 2026 guidance states llms.txt isn't used for AI Overviews or AI Mode, and in practice GPTBot, ClaudeBot, PerplexityBot and OAI-SearchBot overwhelmingly skip it and crawl your HTML directly. It's turned out to be more useful for coding-agent tools like Cursor and Claude Code reading documentation sites than for AI search citation, so it's low-priority unless that's your actual audience.

Let's chat about your next project

Whether you need a single contractor or a full team, I run every project through Storm Digital, the studio I co-founded. Having delivered large-scale projects for businesses of every size, I bring that same commitment to clients locally in Surrey, London, and worldwide.

See also...

CnCNet - Website & App Development

Comprehensive design and development of responsive web and desktop apps for CnCNet, enhancing user experience and accessibility.

AR Configurator for E-Bikes

Bespoke augmented reality web configurator for E-Bikes, offering full customization options, including color choices and feature configurations.

Stock Investment App

Real-time stock investment app for web and mobile, built with TypeScript, React, and WebSockets, delivering fast, responsive, and data-driven experiences.

KickTown Football - Website

Custom-built website and API's for KickTown Football, integrating a merchandise store, booking system, and seamless user experience.

Tempest Rising - Official Website

Website development for Slipgate Ironworks' Tempest Rising, crafted to deliver a sleek, immersive experience for fans and players.

Cosmetic Visualiser - Web App

React and TypeScript-powered web app for visualizing facial cosmetic treatments, offering a cutting-edge, interactive user experience.

C&C Community Website

Command & Conquer Community platform with improved SEO, a sleek interface, and content integrations like Twitch and Steam, creating a hub for fans and creators.

React Native Health App

Custom Android and iOS health app developed with React Native and TypeScript, tailored for a health-tech startup, ensuring cross-platform compatibility.

Oriental Garden Restaurant - Website Redesign

Website redesign featuring an online ordering system that boosts sales and enhances customer engagement with a modern, intuitive interface.

Brands I've had the privilege to contribute to...

Logo for Slipgate Studios
Logo for Evolve
Logo for Disney
Logo for Heathrow
Logo for BAE Systems
Logo for University of Surrey
Logo for Allergan
Logo for OKA
Logo for Ribena
Logo for NARS