What Is a robots.txt File?
A robots.txt file is a plain text file placed at the root directory of a website that tells web crawlers which areas of the site they are allowed to access and which areas they should avoid. It contains a list of user-agent strings (the name of the bot), crawling directives, and the URL paths to which access is restricted.
When you create a website, you may want to prevent search engines from indexing certain pages — such as thank-you pages, confirmation pages, staging content, or duplicate pages. The robots.txt file is where web crawlers learn what they can and cannot access. It is also the mechanism for blocking non-search-engine crawlers (such as those from third-party tools) from accessing your pages.
Where Is the robots.txt File Located?
The robots.txt file must be placed in the root directory of your website. For example, if your site is https://www.example.com, the file must be accessible at https://www.example.com/robots.txt.
If a web crawler cannot find the robots.txt file at the root directory, it will assume no file exists and proceed to crawl all pages accessible via links on your site. How you upload the file depends on your website platform and server architecture — contact your hosting provider if you are unsure.
Why Does robots.txt Matter for SEO?
The robots.txt file is considered a fundamental part of technical SEO because search engines discover and understand websites entirely through their crawlers. The robots.txt file is the most direct way to communicate with those crawlers. Key benefits include:
- Improved crawl efficiency by directing bots toward valuable content
- Preventing low-value pages (e.g., confirmation pages, internal search results) from being indexed
- Reducing the risk of duplicate content issues
- Keeping unfinished or private content away from search results
How Does robots.txt Work?
When a search engine bot encounters a robots.txt file, it reads the file and follows the instructions. A basic entry looks like this:
User-agent: googlebot
Disallow: /confirmation-page/
This tells Googlebot not to crawl the /confirmation-page/ path or any URL within that subdirectory (e.g., /confirmation-page/order/, /confirmation-page/demo/). If a URL is not mentioned in the file, the crawler is free to access it as normal.
It is important to note that robots.txt is a directive, not a security mechanism. Most reputable crawlers follow it, but it does not technically prevent access to pages — it is a convention, not a lock. Do not rely on it to protect sensitive data.
Best Practices for robots.txt
- The file must be a plain text file encoded in UTF-8 format
- The filename is case sensitive — it must be named exactly
robots.txt - Place the file at the root directory of your website or subdomain
- Maintain only one robots.txt file per (sub)domain
- You can only have one group of directives per user agent within the file
- Be as specific as possible with disallow paths to avoid accidentally blocking entire sections of your site
- Do not use the noindex directive inside robots.txt — this directive is not reliably supported there; use meta robots tags on individual pages instead
- Remember that robots.txt is publicly visible — do not reveal the existence of confidential or sensitive sections of your site through it
- robots.txt is not a substitute for configuring robots meta tags on individual pages; use both appropriately
- Reference your XML sitemap within your robots.txt file, and ensure the sitemap is publicly accessible
Common robots.txt Issues and How to Fix Them
When a site's robots.txt file is misconfigured, it can negatively affect how search engines crawl and index your content. Search Atlas's Site Auditor can flag these issues automatically. Below are the most common problems and how to resolve them.
1. robots.txt File Not Present
If you do not have a robots.txt file, or if it is not located at the correct root-level URL, crawlers will assume no restrictions exist and crawl your entire site. Resolution: Create a valid robots.txt file and upload it to the root directory of your website.
2. robots.txt Present on a Non-Canonical Domain Variant
Best practice requires only one robots.txt file per canonical (sub)domain. If your robots.txt is located on a non-canonical variant of your domain (e.g., the non-www or HTTP version when the canonical is www/HTTPS), it may cause confusion for crawlers.
For example, if your canonical domain is https://www.example.com, your robots.txt should be at https://www.example.com/robots.txt — not at http://example.com/robots.txt or https://example.com/robots.txt.
Resolution: Move the file to the canonical domain location, or set up 301 redirects from non-canonical robots.txt URLs to the canonical version.
3. Invalid Directives or Syntax
Including unsupported or incorrectly formatted directives means crawlers may ignore your instructions entirely and access pages you intended to block. Common syntax errors include misspelled directives, incorrect spacing, or using directives that are not part of the robots.txt standard.
Resolution: Audit your robots.txt file carefully. Use only supported directives (User-agent, Disallow, Allow, Sitemap, Crawl-delay where supported). Validate your file using Google Search Console's robots.txt tester or a similar tool.
4. Sitemap Not Referenced or Not Accessible
Best practice is to include a reference to your XML sitemap within your robots.txt file so crawlers can easily find it. If the sitemap is referenced but returns a 4xx or 5xx error, it can hinder crawling efficiency.
Resolution: Add a Sitemap: directive pointing to your sitemap's full URL, and confirm the sitemap is publicly accessible and returns a 200 status code.
5. Crawl-Delay Directive Misuse
Some implementations incorrectly include crawl delay settings that are either unsupported by major search engines (Googlebot does not support Crawl-delay) or set too aggressively, which can slow down beneficial crawling. Resolution: Use Google Search Console to manage Googlebot's crawl rate instead of relying on the Crawl-delay directive. Remove the directive if it is causing issues with major search engine bots.
Using Search Atlas to Identify robots.txt Issues
Search Atlas's Site Auditor can automatically detect robots.txt problems as part of a full technical SEO audit — including missing files, non-canonical placement, invalid syntax, inaccessible sitemaps, and problematic directives. OTTO SEO can also assist with ongoing technical SEO monitoring to help keep your site's crawlability in good standing.
If you need further assistance, open the chat widget in the bottom-right corner of the platform and type human teammate to be connected with a member of our team.