The Search Atlas Site Auditor scans your robots.txt file for errors that can prevent crawlers from reaching your pages. Use this guide to understand each audit warning and how to fix it.
🤖 What Is a robots.txt File?
A robots.txt file tells web crawlers which parts of your website they can or cannot access. Crawlers fetch it before visiting any other page on your site. Each file contains:
- A user-agent string — the crawler's name (e.g.,
Googlebot,Applebot) - Directives — rules such as
AlloworDisallow - Paths or URLs — the pages or directories the rule applies to
Use robots.txt to prevent unready or private pages from being indexed, restrict specific crawlers, and focus your crawl budget on valuable content.
📂 Where to Place Your robots.txt File
Place the file in the root directory of your website — for example, https://www.website.com/robots.txt. If a crawler cannot find a robots.txt file, it assumes all pages are crawlable. On large sites this wastes crawl budget on low-value pages, so create an explicit file even if you intend to allow everything. If you don't have a robots.txt file yet, upload one to your root directory — your hosting provider can help.
🔍 Why robots.txt Matters for SEO
Robots.txt directly controls which pages crawlers visit and how efficiently they use your crawl budget. Key benefits include:
- Improves crawl efficiency
- Prevents indexing of low-value or duplicate pages (such as confirmation pages)
- Protects confidential areas of your site
- Keeps search results clean and relevant
⚙️ How robots.txt Works
When a crawler visits your site, it fetches /robots.txt first and reads the directives for its user-agent. It then applies those rules to every URL it tries to visit. If no restriction is specified for a path, crawlers access and index it freely.
Example:
User-agent: Googlebot Disallow: /confirmation-page/
🧠 robots.txt Best Practices
- Encode the file in UTF-8 format.
- Name the file exactly robots.txt (case-sensitive).
- Place it in the root directory.
- Maintain only one robots.txt file per (sub)domain.
- Give each user-agent its own directive group.
- Be specific — avoid accidentally blocking entire subdirectories.
- Do not use
noindexin robots.txt; configure robots meta tags on individual pages instead. - Remember: the file is publicly viewable.
🚨 robots.txt Issues Detected by the Search Atlas Site Auditor
The Search Atlas Site Auditor automatically detects the following robots.txt problems. Select each issue to see the details and fix.
1️⃣ robots.txt Not Present
Issue: No robots.txt file exists in the root directory.
Fix: Create a robots.txt file and upload it to https://www.website.com/robots.txt. The warning clears on your next Site Auditor crawl.
2️⃣ robots.txt on a Non-Canonical Domain Variant
Issue: Multiple robots.txt files exist across www/non-www or HTTP/HTTPS versions of your domain.
Fix:
- Keep only one canonical robots.txt at your preferred domain (e.g.,
https://www.website.com/robots.txt). - Set up 301 redirects from all other domain variants to point to the canonical version.
3️⃣ Invalid Directives or Syntax
Issue: Incorrect syntax prevents crawlers from following your rules.
Fix: Open your robots.txt file and correct the errors. The Site Auditor lists the specific invalid directives it detected so you know exactly what to address.
4️⃣ robots.txt Should Reference an Accessible Sitemap
Issue: Your robots.txt file does not include a reference to your XML sitemap.
Fix: Add this line at the end of your robots.txt file:
Sitemap: https://www.website.com/sitemap.xml
This helps crawlers discover and prioritize your pages faster.
5️⃣ Disallow: / Blocks All Crawlers from Your Entire Site
Issue: Your robots.txt contains Disallow: / under a user-agent directive, which blocks that crawler from accessing every page on your site. When applied to User-agent: *, all search engines are blocked — which can completely prevent your site from appearing in search results.
This is one of the most critical robots.txt misconfigurations.
⚠️ WordPress caveat: A regex detection flaw (tracked as WP-322) currently prevents this warning from triggering reliably on WordPress sites in the Site Auditor. The fix is awaiting release. Until it ships, manually verify your robots.txt file for a Disallow: / line under User-agent: * rather than relying solely on the Site Auditor.
Common WordPress cause: WordPress's Settings → Reading page includes a "Discourage search engines from indexing this site" checkbox. Enabling it automatically adds Disallow: / to your robots.txt.
Fix:
- Open your robots.txt file at
https://www.website.com/robots.txt. - Find any line that reads
Disallow: /under a broad user-agent (especiallyUser-agent: *). - Remove that line, or replace it with only the specific path you intend to restrict (e.g.,
Disallow: /private/). - If you use WordPress, go to Settings → Reading and confirm that "Discourage search engines from indexing this site" is unchecked.
- Save the file and verify the update is live at your robots.txt URL.
The warning clears on your next Site Auditor crawl once the file is corrected.
🔌 Managing robots.txt with the Search Atlas WordPress Plugin
If you run WordPress, the Search Atlas plugin includes a built-in robots.txt manager so you can edit, back up, and restore your file without leaving your WordPress admin.
🛠️ How to Use the WordPress robots.txt Manager
Access the manager:
- Log in to your WordPress admin dashboard.
- Open the Search Atlas plugin menu.
- Select robots.txt to open the manager.
Available actions:
- Edit — update the contents of your robots.txt directly in the editor and save changes.
- Backup — the plugin automatically saves a snapshot each time you save a change.
- Restore — open the backup history and roll back to a previous version with one click.
⏰ Timestamp note: Timestamps in the robots.txt manager currently display in UTC regardless of your WordPress timezone setting (tracked as WP-326); a fix is in progress. Convert manually when comparing backup times to your local timezone until the fix is released.
⚠️ Reminder: Because of the WP-322 detection issue described above, always double-check your saved file for an unintended Disallow: / directive after editing.
🛠️ Troubleshooting a Stalled Site Audit Crawl
If your Site Audit shows as processing for more than 30 minutes without completing, the crawl may be stalled. To recover:
- Stop the current crawl from the Site Auditor dashboard.
- Wait a few minutes for the queue to clear.
- Restart the crawl.
If the problem persists, contact Search Atlas support so the team can investigate.
🎯 You can now identify and resolve every robots.txt issue flagged by the Search Atlas Site Auditor, manage your file directly inside the WordPress plugin, and recover from a stalled Site Audit crawl. Run a new Site Auditor crawl to confirm your fixes — and remember to manually verify your Disallow: / directive on WordPress sites until WP-322 ships.