robots.txt in 2026: Why Blanket Blocking or Allowing Is Both Wrong
Simon Heistermann
Owner
This article was written with AI assistance and editorially reviewed.
Most sites have not touched their robots.txt since it was first set up. For the AI crawlers that have appeared since 2023, that means they are not being handled by a deliberate decision, but by a default that predates ChatGPT. In practice, that leaves many sites either silently blocking every new bot or just as thoughtlessly leaving everything open - both are the wrong call, just for different reasons.
In short
Robots.txt decides which AI crawlers are even allowed to read your content. Block every bot and you disappear from AI answers. Allow every bot and you also feed scrapers that give you nothing back. The right configuration distinguishes bots by purpose, not by gut feeling.
Why robots.txt now decides AI visibility
Until a few years ago, robots.txt was a footnote: set up once to keep search engines out of admin areas, then never touched again. Generative AI changed that. Every major provider now runs its own bots that fetch content live, either to answer questions or to train models - and every one of them reads robots.txt before it accesses anything. One wrong or outdated line there directly decides whether your business can even be surfaced as a source in a ChatGPT or Perplexity answer.
The check takes a minute: open yourdomain.com/robots.txt in a browser. If it shows a blanket block for every bot, an empty shell left over from a hosting template, or nothing at all, that is not a neutral starting point - it is a silent decision against AI visibility that nobody actually made on purpose.
What robots.txt technically does, and does not do
Robots.txt is a plain text file in a domain's root directory that tells individual or all bots which paths they may crawl. The core syntax is two lines per block:
User-agent: GPTBot
Allow: /
User-agent names the bot, Allow or Disallow set which paths it may visit. A block only applies to a bot that matches its user-agent name exactly - the catch-all rule under User-agent: * only applies to bots that find no more specific block of their own. The limit worth understanding: robots.txt is a request, not access control. Reputable bots from OpenAI, Anthropic, Perplexity or Google honour it because doing so serves their own interest in a functioning bot ecosystem. An actor that does not want to comply simply will not - robots.txt does not protect against malicious access, it only steers the behaviour of cooperative systems.
The AI crawlers worth knowing
Not every bot serves the same purpose, and that is exactly what most robots.txt files ignore:
| Bot | Operator | Purpose |
|---|---|---|
| GPTBot, OAI-SearchBot, ChatGPT-User | OpenAI | Live fetching for ChatGPT answers and search |
| ClaudeBot, Claude-SearchBot, Claude-User | Anthropic | Live fetching for Claude answers and model training |
| PerplexityBot, Perplexity-User | Perplexity | Live fetching for answers with source attribution |
| Google-Extended | Separate switch for AI training, independent of classic Googlebot | |
| Applebot, Applebot-Extended | Apple | Spotlight search and Apple Intelligence training |
| CCBot, Bytespider, Diffbot | Common Crawl, ByteDance, Diffbot | Bulk scraping for third-party training sets, usually with no citation or traffic in return |
The first four groups give something back that matters for a business: a chance at being cited with a link, or at least brand presence in the memory of a system potential customers use daily. The last group typically gives back neither - the data disappears into a training set with no user ever finding their way back through a link.
Blocking everything or allowing everything are both wrong
The most common reaction to AI crawlers is one of two overcorrections. Some block every bot with "AI" or "GPT" in the name out of caution - and cut themselves off from exactly the channels where potential customers now look for answers before they ever open a classic search engine. Others leave robots.txt wide open because that is the path of least resistance, and end up feeding unfiltered scrapers that provide no return whatsoever.
The considered position sits between the two and follows a single question: does your business get something back for allowing access - citation, traffic, brand presence in a system customers actually use? Bots that answer yes belong in robots.txt with an explicit Allow, not just under the silent *. Bots that answer no belong just as explicitly blocked. Both require actually naming the bots rather than relying on one catch-all line.
- Give search and answer bots with citation potential (GPTBot, ClaudeBot, PerplexityBot, Google-Extended) an explicit Allow
- Explicitly Disallow pure bulk scrapers with no backlink value (CCBot, Bytespider, Diffbot)
- Do not forget the sitemap line at the end of the file, so crawlers can even find the site's structure
- Re-check robots.txt after every major structural change to the website
Concrete steps for the next 90 days
- Days 1-30: open the current robots.txt in a browser, check every existing block against your own list of wanted and unwanted bots
- Days 31-60: add missing Allow rules for answer bots and missing Disallow rules for pure scrapers, verify the sitemap line
- Days 61-90: check server logs for actual traffic from the bots you allowed, and update robots.txt whenever new bot names appear
Conclusion
Robots.txt is the smallest technical decision with the largest effect on AI visibility, which is exactly why it is so often left to chance. A configuration that distinguishes answer bots from pure scrapers takes an hour to implement and keeps working for as long as the site exists. How Google and AI crawlers technically process the content behind that access is covered in our article on JavaScript SEO for SPAs. The sitemap line at the end of the file deserves just as much attention as the bot rules themselves, as we show in our article on XML Sitemaps 2026. What actually turns that access into a citable result is covered in our article on Schema.org for local businesses, and how solid HTTPS fundamentals fit in is covered in our article on HTTPS and HSTS. For a no-obligation review of your current robots.txt, feel free to get in touch.
Want to know which AI crawlers your website is currently blocking?
Get in touchYou might also like
WordPress or Custom Build? Run the Five-Year Numbers
WordPress powers a large share of the web, and for good reasons. What it actually costs to run, where it wins outright, and when a custom build makes sense.
Websites for IT Service Providers: Your Own Site Is the Work Sample
An IT provider with a slow, insecure website refutes its own pitch. What an IT manager checks in the first few minutes, and what follows from it.
Website Maintenance in 2026: What It Costs and What Must Be In It
What website maintenance actually covers, what the market charges for it, and how to spot an empty maintenance contract before you sign it.
Website Hosting for Businesses: What Actually Matters in 2026
Shared hosting, managed hosting or a platform: what the difference means for load time and resilience - and who actually owns the domain at the end.
SEO Costs 2026: What Visibility Really Costs
What SEO realistically costs small and mid-sized businesses: one-off optimisation versus ongoing management, and what should be included in the price.
A GDPR Check for Your Website: The Gaps That Are Almost Always There
Fonts from someone else's server, maps without consent, analytics before agreement: the typical gaps on SME websites, as a list you can actually check.
Frequently asked questions

Simon Heistermann
Owner
Heistermann Solutions is the web studio run by Simon Heistermann. We build custom websites for small and medium-sized businesses that want to achieve more online.
Every article grows out of day-to-day project work and is reviewed editorially before publication.
- Borken, Münsterland region
- simon@heistermann-solutions.de
Get it for free
Enter your email address. You'll immediately receive a confirmation link - after clicking it the checklist is available right away.
Let's talk about your project
Free introductory call