robots.txt is one of the oldest and simplest web standards still in daily use, but it's also one of the most misunderstood — mostly because its name sounds like it controls something stronger than it actually does.

It's a Request, Not a Lock

robots.txt is a voluntary convention, not an enforced security boundary. Well-behaved crawlers — Google, Bing, and most legitimate search bots — read the file and respect its rules before crawling your site. But nothing about the file technically prevents a browser, a scraper, or a malicious bot from ignoring it entirely and requesting a disallowed page directly. If a URL genuinely needs to stay private, robots.txt is the wrong tool — that's a job for authentication or server-side access control instead.

The "Hiding From Google" Misconception

Disallowing a path in robots.txt stops crawlers from fetching that page's content, but it doesn't stop the URL itself from potentially appearing in search results. If enough other sites link to that disallowed URL, Google can index the URL as a bare link — sometimes shown with no description, just the address — because it never needed to crawl the page to know the URL exists. Truly preventing a page from appearing in search results requires a noindex meta tag or HTTP header on the page itself, which ironically means the page needs to stay crawlable for Google to see and honor that instruction.

Tip: If you actually want a page out of search results entirely, use noindex on the page (and leave it crawlable) rather than disallowing it in robots.txt — combining both often backfires, since a disallowed page can't be crawled to discover the noindex tag telling Google to drop it.

Reading Allow, Disallow, and Wildcards

Each rule set starts with a User-agent line specifying which crawler it applies to (an asterisk means "all crawlers"), followed by Disallow lines listing paths that crawler shouldn't fetch, and optionally Allow lines carving out exceptions within a disallowed area. An empty Disallow line (just "Disallow:" with nothing after it) means nothing is blocked — full access — which is different from "Disallow: /", which blocks the entire site starting from the root.

Where the Sitemap Line Fits In

The Sitemap directive is unrelated to the Allow/Disallow rules — it doesn't control access at all. It's simply a pointer telling any crawler where to find your sitemap.xml file, which lists your important URLs directly rather than making the crawler discover them purely by following links around your site. Including it is a small efficiency courtesy that can help large or newer sites get crawled more completely.

Generate One Instantly

Choose Allow All, Custom Rules, or Block All, add your sitemap URL, and get a correctly formatted file with our free Robots.txt Generator — copy the result and upload it to your site's root as robots.txt.

FAQ

Where do I put the generated robots.txt file? Upload it to the root of your domain so it's reachable at yourdomain.com/robots.txt — search engines specifically look for it at that exact location and won't find it anywhere else.

Does blocking a page in robots.txt guarantee it stays out of Google's index? Not entirely — robots.txt only tells crawlers not to fetch a page's content. If other sites link to that URL, Google can sometimes still index the URL itself (without its content) unless you also use a noindex meta tag on the page, which does require it to be crawlable.

What does the Sitemap line actually do? It's a courtesy pointer telling any crawler exactly where your sitemap.xml lives, so it can find and process all your important URLs efficiently, rather than relying purely on discovering links by crawling around your site.

Is the information I enter sent anywhere? No — the file is built entirely in your browser using JavaScript, so nothing you enter is sent to a server.

Need a robots.txt file right now? Try the free Robots.txt Generator — no sign-up, correctly formatted every time.