robots.txt
A small text file in the root directory of your website that tells search engines which areas they may crawl and which they may not.
robots.txt is a small text file in the root directory of your website that tells search engine crawlers which areas they may index and which they may not.
Where the file lives
Always at the root URL, i.e. bytebrise.com/robots.txt. Not in subfolders, not at www.bytebrise.com/de/robots.txt. Every host has its own robots.txt. If the file is missing, crawlers assume that everything is allowed.
What the syntax can do
User-agent specifies which crawler the following rules apply to; an asterisk acts as a wildcard for all. Disallow blocks specific paths. Allow explicitly permits sub-areas of a blocked path. Sitemap points to the site's XML sitemap. Crawl-delay limits how quickly the crawler makes successive requests; Google ignores it, but Bing respects it.
Typical contents
Blocking admin areas such as /admin/, /backend/, /wp-admin/. Blocking lead magnet PDFs that should only be delivered after an email address is entered. Blocking search result pages or filter combinations that would create duplicate content. A mandatory Sitemap directive for every active sitemap file.
What robots.txt cannot do
It cannot remove a page from Google's index. If a page that is already indexed is blocked via Disallow, it stays in the index, just without a snippet. Deindexing requires a noindex meta tag in the HTML or an X-Robots-Tag in the HTTP header. Nor can it protect private data. Anyone who reads your robots.txt can see straight away which URLs look interesting and can visit them manually.
Related terms
Want to actually improve your website?
We analyse your site and tell you honestly what is worth fixing. The initial consultation is free.
Free initial consultation