Troubleshooting: Crawl Failures & Blocked Crawlers

49 articles Camilo Aponte By Camilo Aponte

🔧 Fix WordPress Plugin Approvals, Robots.txt & Content Score

🔍 Fix Site Audit Zero Pages Crawled Issue

🔄 Trigger a Fresh OTTO SEO Recrawl

🛠️ Fix Empty OTTO AI Recommendations After Crawl

🛠️ Refunds for Failed Project Crawls

🔍 Why project crawls get stuck A project crawl can remain loading when Search Atlas cannot access or complete a scan of the website. Common causes include robots.txt rules, firewall or security settings, server timeouts, DNS issues, or a site that is temporarily unavailable. A crawl that has not finished after several hours may need attention, especially if the same project has failed more than once. 🔄 What happens after a crawl fails Search Atlas uses retry attempts to recover from temporary crawl problems. The system can make up to three attempts. If all attempts fail, the crawl should be treated as unsuccessful rather than left loading indefinitely. When an eligible crawl fails after the available retries, the related quota or credits may be returned automatically. Refresh the project and review its crawl status before starting another attempt. ✅ Check your website before retrying 1. Confirm the website loads normally in a browser. 2. Check that the site is not blocking Search Atlas or automated crawlers. 3. Review robots.txt and firewall settings for restrictions. 4. Confirm that DNS and hosting services are working. 5. Make sure the project uses the correct website URL and access details. Fixing the underlying access issue before retrying can help prevent another failed crawl and additional processing time. 💳 When to request a refund Request a refund review when a project remains stuck, repeated attempts fail, or the expected quota or credit return does not appear. Include the project name, website URL, approximate start time, number of attempts, and screenshots of the loading or error message. Do not repeatedly start new attempts while the original crawl is still processing. Multiple attempts can make the project status harder to review. 💬 Contact support about a refund If the crawl has been stuck for an extended period or you need a refund review, open the chat widget in the bottom-right corner of the platform and type human teammate to be connected with a member of our team. Our team can review the crawl history, confirm whether a retry or refund was processed, and investigate cases where a failed crawl did not receive the expected quota or credit return.

🛠️ Troubleshoot OTTO Site Audit Crawls

🔍 What a Failed Crawl Means An OTTO site audit crawl can fail because the website blocks the crawler, access details are incomplete, the site is temporarily unavailable, or a platform processing step encounters an error. A failed crawl does not automatically mean the issue is on your side or on OTTO’s side. Some crawls also take longer when JavaScript rendering is enabled or when a site is large. During this time, the crawl may appear stuck even though processing is still in progress. 📊 Check the Crawl Status 1. In the left sidebar, open OTTO SEO. 2. Select Crawl Monitoring to review recent crawl activity and status messages. 3. Open All Audits to check whether the audit completed, failed, or is still processing. 4. Review the affected site under All Sites and note the latest crawl time and error message. Messages such as “Crawl failed. We’ll try again soon.” indicate that the platform could not complete the crawl. The message alone may not identify whether the cause is website access, site performance, or an internal processing issue. ⚙️ Check Your Website Setup - Confirm the website is online and loads normally in a browser. - Check that robots.txt or security tools are not blocking Search Atlas or OTTO crawlers. - Review firewall, CDN, bot protection, and rate-limit settings for blocked requests or challenge pages. - Confirm that required OTTO installation or access settings are complete. Open OTTO SEO → Installation Guide to review the setup. - Check whether recent hosting, DNS, redirect, authentication, or security changes could affect crawling. - For JavaScript-heavy sites, allow additional processing time before treating a long-running crawl as failed. 🔄 Retry the Crawl Safely 1. Resolve any access or website availability issues first. 2. Open OTTO SEO → All Sites. 3. Select the affected site and choose Scan. 4. Monitor the new attempt in Crawl Monitoring. Avoid starting many scans in quick succession. Repeated retries can make it harder to identify the original cause and may increase processing time. 🧭 Tell Platform Issues From Site Issues The issue is more likely related to the website when the site is unavailable, crawler requests are blocked, security challenges appear, or the failure begins after a site configuration change. The issue may be platform-related when the website is reachable, crawler access is allowed, the same site repeatedly fails without configuration changes, or multiple sites show similar failures at the same time. Internal handoff and post-processing failures can also cause a crawl to fail after the crawling stage has completed. 💬 When to Contact Support Before contacting support, record the site name, approximate failure time, exact status message, recent website changes, and whether a retry produced the same result. This helps the team determine whether the failure occurred during crawling or later processing. If you need further assistance, open the chat widget in the bottom-right corner of the platform and type human teammate to be connected with a member of our team.

🛠️ Fix Repeated OTTO Site Audit Crawl Failures

🔍 What a failed crawl means OTTO must crawl your site before it can complete a site audit. A repeated failure can come from your website, access restrictions, slow JavaScript rendering, or an issue during OTTO processing. The message “Crawl failed. We’ll try again soon.” does not by itself identify the exact cause. A failed crawl does not automatically mean the problem is on OTTO’s side. Use the checks below to rule out common website and configuration issues first. 🧭 Check the crawl status 1. Open OTTO SEO from the left sidebar. 2. Open Crawl Monitoring to review the latest crawl status and timing. 3. Open All Audits to confirm whether the audit is still processing, failed, or completed. 4. Check whether multiple attempts failed for the same site and whether the failures occurred at similar stages. Long processing times do not always indicate a failure. Sites that rely heavily on JavaScript may take substantially longer to crawl. Wait for the status to update before starting repeated manual scans. 🌐 Check your website access - Open the site in a private browser window and confirm it loads without a login. - Check that the domain and preferred protocol, such as HTTPS, are correct in the project settings. - Review firewall, CDN, bot protection, and security tools for blocked or challenged crawler requests. - Confirm that robots.txt, noindex rules, or server permissions are not preventing access to important pages. - Check server logs for timeouts, repeated 403 or 429 responses, or 5xx errors during the crawl window. - Make sure the site is not undergoing maintenance or experiencing intermittent downtime. ⚙️ Review OTTO settings 1. Go to OTTO SEO → All Sites. 2. Select the affected site and confirm its URL and configuration. 3. Use Scan to start a new crawl only after checking access and server health. 4. Review OTTO SEO → Installation Guide if the site connection or installation has recently changed. Do not repeatedly start scans while an earlier attempt is still processing. Multiple overlapping attempts can make it harder to identify the original failure and may increase processing time. 📊 How to tell where the issue is The issue is more likely related to your site when the site is unavailable, protected by a challenge page, returning server errors, blocking crawlers, or taking a long time to respond. It is more likely related to platform processing when the site is publicly reachable, server logs show successful access, several attempts fail in the same way, and other site checks work normally. Customers may also see a failure when the crawl completes but the handoff to audit processing does not finish. In that case, the crawl itself may have succeeded even though the audit displays a failure. Share the exact status and timestamp when reporting the issue. 📝 What to collect before contacting support - Site name and domain. - The exact error message. - Approximate start and failure times, including your time zone. - Screenshots from Crawl Monitoring and All Audits. - Whether the site uses a CDN, firewall, bot protection, or JavaScript-heavy pages. - Relevant server log results, including status codes, timeouts, or blocked requests. - Whether a normal browser visit and other crawling tools can access the site. This information helps the team determine whether the failure is caused by site access, crawl duration, or an OTTO processing stage. 💬 Get help with a confirmed failure If the site is accessible, access controls are not blocking crawlers, and multiple attempts continue to fail, the issue may require platform investigation. Include the details above so the team can check the crawl and post-processing stages. If you need further assistance, open the chat widget in the bottom-right corner of the platform and type human teammate to be connected with a member of our team.

🔍 Fix Crawl Processing and URL Tracking

🔍 Overview Search Atlas audits crawl your site, then process the collected pages before adding them to audit results and page catalogs. A post-processing failure can prevent newly discovered pages from appearing. Tracking scripts may also report incorrect URLs when trailing slashes are handled inconsistently. 🧭 Check the Audit 1. Open AI SEO → OTTO SEO → All Audits. 2. Find the affected audit, such as audit ID 120474, and review its status and results. 3. Check whether the audit shows a processing, partial crawl, or page collection warning. 4. Compare the number of crawled pages with the number of cataloged pages. If the crawl completed but new pages are not cataloged, the issue may be in the handoff or post-processing stage rather than in the site crawl itself. ⚙️ Retry a Blocked Process Open AI SEO → OTTO SEO → Crawl Monitoring to review the crawl and processing status. Allow an active job to finish before starting another scan. If the process remains blocked, record the audit ID, site, approximate start time, and the affected page URLs. Do not repeatedly launch scans for the same site while a previous job is still processing. This can make it harder to identify the original failure. Search Atlas is improving post-processing reliability with safer retries, clearer partial-failure messages, and more reliable crawl-to-processing handoffs. 🔗 Check Trailing-Slash URLs A URL with a trailing slash and the same URL without one can be treated as different strings by tracking systems. For example, https://example.com/page/ and https://example.com/page may produce separate records even when the server treats them as the same page. 1. Confirm the canonical URL format used by your site. 2. Check whether the MetaSync or OTTO tracker script generates URLs in that same format. 3. Use one consistent format for page URLs, canonical references, internal links, and tracking events. 4. After updating the implementation, run a fresh scan and compare the reported URLs with the canonical format. Do not manually remove the slash from every URL unless your site officially uses the non-slash format. The correct format is the one enforced by your redirects and canonical URLs. 🛠️ Validate OTTO and MetaSync Setup Open AI SEO → OTTO SEO → Installation Guide and confirm that the current installation instructions are followed. Then review the affected site under AI SEO → OTTO SEO → All Sites. Confirm that the site is selected correctly and that its crawl and tracking configuration matches the live site. For deeper site-level comparison, use Site Metrics → Analyze and review the project history for duplicate URL variants or missing pages. 📊 What to Include in a Report - Audit ID and site name. - Examples of pages that were crawled but not cataloged. - Examples showing both trailing-slash URL formats. - The audit, crawl, or processing status displayed in the platform. - The approximate time the issue occurred. - Whether the issue affects one site or multiple sites. This information helps the team distinguish a crawl problem from a post-processing or URL-normalization problem. 💬 Get Help If you need further assistance, open the chat widget in the bottom-right corner of the platform and type human teammate to be connected with a member of our team.

🛠️ Investigate Crawl Post-Processing Issues

🔍 What this issue means Crawl post-processing happens after pages are discovered. During this stage, Search Atlas prepares crawl data for cataloging, reporting, and downstream tools such as MetaSync and OTTO. If processing is blocked or incomplete, newly discovered pages may not appear in the catalog even though the crawl started successfully. 📋 Review the affected audit 1. Open OTTO SEO from the left sidebar. 2. Select All Audits. 3. Find the affected audit, such as audit ID 120474, and open its results. 4. Check whether the audit shows processing, partial crawl, or page-level errors. 5. Compare the number of discovered pages with the number of cataloged pages. A difference between discovered and cataloged pages usually indicates a crawl-to-processing handoff or post-processing issue rather than a problem with the site’s page content. 🔄 Retry processing safely 1. Allow an audit that is still processing time to complete before starting another scan. 2. If processing is stalled, return to OTTO SEO → All Sites and select the site. 3. Use Scan to start a fresh crawl after confirming the site is accessible. 4. Review the new audit and confirm that discovered pages move into the catalog. Do not repeatedly launch scans for the same site while an audit is active. Repeated scans can make it harder to identify which crawl produced the incomplete results. 🔗 Check trailing-slash tracking URLs Tracking scripts must use the same URL format as the pages they reference. A trailing-slash mismatch can cause tracking, matching, or page updates to fail. For example, these are different URL strings: - https://example.com/page/ - https://example.com/page 1. Confirm the site’s preferred URL format by checking its canonical tags and redirect behavior. 2. Use that format consistently in MetaSync and OTTO tracker configuration. 3. Make sure the script does not add or remove a trailing slash inconsistently. 4. Test the script on the homepage and on a representative interior page. 5. Re-run the audit and verify that the tracked URL matches the cataloged URL exactly. ✅ Validate the result - Open OTTO SEO → Crawl Monitoring to review crawl activity. - Confirm that processing finishes without a blocker. - Check that newly discovered pages are cataloged. - Verify that page URLs in tracking requests use the site’s preferred trailing-slash format. - Use OTTO SEO → Overview to confirm that site status and page counts are updated. If pages remain missing after a fresh scan, record the audit ID, site URL, approximate start time, discovered page count, and cataloged page count. This information helps the team distinguish a processing failure from a site access or URL-format issue. 💬 Get help with a persistent blocker Some post-processing failures require platform-side investigation or a retry of the crawl handoff. If the issue continues after these checks, include the audit ID and the validation results when reporting it. If you need further assistance, open the chat widget in the bottom-right corner of the platform and type human teammate to be connected with a member of our team.

🔄 Fix Stuck OTTO Recrawl and Site Audit Jobs

🔍 Overview OTTO recrawl and Site Audit jobs normally complete within a few hours. If your job has been running for more than 24 hours without finishing, it is likely stuck due to a processing error. This article explains why this happens, what you can do right now, and when to contact support. ⚠️ Why Jobs Get Stuck Several known issues can cause a recrawl or audit job to stall indefinitely: - Status misread: OTTO incorrectly reads a completed task as still processing, causing it to wait indefinitely for a result that already exists. - Orphaned reprocess task: A background reprocessing task gets detached from its parent job and loops without ever finishing. - No staleness guard: Older jobs lacked a timeout mechanism, so a stuck job had no automatic way to recover. - Recrawl frequency not set: Some audits were missing a recrawl frequency value, which prevented them from progressing through the queue correctly. - Image recheck routing: OTTO image rechecks were not always routed correctly to the Site Audit queue, causing them to stall. The Search Atlas engineering team has resolved the most common root causes. If your job is still stuck, the steps below will help you recover it manually. 🛠️ How to Unstick a Recrawl or Audit Job 1. Check the current status. Go to Left sidebar → OTTO SEO → All Audits. Find the affected site and note how long the job has been running. If it has been more than 24 hours, it is safe to assume it is stuck. 2. Refresh the job. On the All Sites page (Left sidebar → OTTO SEO → All Sites), locate your site and click Scan. This triggers a fresh crawl request and can override a stalled job. 3. Wait up to 30 minutes. After triggering a new scan, allow up to 30 minutes for the status to update. Refresh the page periodically to check for progress. 4. Check Crawl Monitoring. Navigate to Left sidebar → OTTO SEO → Crawl Monitoring to see if the crawl is actively processing URLs. If the counter is moving, the job has recovered. 5. Review the Overview page. Go to Left sidebar → OTTO SEO → Overview to confirm that suggestions and audit results are being generated after the crawl completes. ✅ What a Healthy Job Looks Like - The crawl status updates within a few minutes of being triggered. - Crawl Monitoring shows a URL counter actively increasing. - The audit completes and new suggestions appear in the Overview within a few hours for most sites. - Sites with very large page counts (thousands of pages) may take longer but should still show visible progress. 💡 Tips to Prevent Future Stuck Jobs - Avoid triggering multiple scans in quick succession. Starting a new scan before the previous one finishes can create conflicting tasks in the queue. - Check your site is reachable. If your website is down or returning errors when OTTO tries to crawl it, the job may stall waiting for a valid response. Confirm your site loads correctly before starting a scan. - Keep your OTTO connection active. A disconnected or improperly installed OTTO may cause audit tasks to fail silently. Refer to Left sidebar → OTTO SEO → Installation Guide if you need to verify your setup. 💬 Still Stuck After Following These Steps? If your recrawl or audit job is still not progressing after you have tried the steps above, our team can manually investigate and reset the job on your behalf. If you need further assistance, open the chat widget in the bottom-right corner of the platform and type human teammate to be connected with a member of our team. Please have the following information ready to speed up the investigation: - Your website URL (e.g. sowebdesigns.co.uk) - The date and time you first noticed the job was stuck - Any error messages or unusual statuses shown in All Audits or Crawl Monitoring

🔍 Fix OTTO Site Audit Crawl Failures

🗺️ Overview When OTTO runs a site audit, it crawls your website to collect SEO data. Occasionally, this crawl may fail or return incomplete results. This article explains the most common causes of crawl failures, how to tell whether the issue is on OTTO's side or yours, and the steps to get your crawl running again. 🔎 How to Check Crawl Status in OTTO 1. Go to Left sidebar → OTTO SEO → All Sites. 2. Locate the site you want to investigate and check its status indicator. 3. For a more detailed view, navigate to Left sidebar → OTTO SEO → Crawl Monitoring. 4. Review the crawl log entries for error messages, timestamps, and any noted interruptions. The Crawl Monitoring page is your first stop for diagnosing what went wrong. It shows whether the crawl started, how far it got, and where it stopped. ⚠️ Common Reasons a Crawl Fails Crawl failures can originate either on OTTO's side or on your website's side. Use the list below to identify which applies to your situation. - Site was unreachable during the crawl: If your server was down, in maintenance mode, or returning 5xx errors at the time OTTO attempted to crawl, the crawl will fail. Check your server logs for the crawl timestamp. - Robots.txt is blocking the crawler: If your robots.txt file disallows all bots or specifically blocks OTTO's user agent, pages will not be crawled. Review your robots.txt at yourdomain.com/robots.txt. - Password protection or authentication walls: If your site requires a login, HTTP authentication, or is behind a firewall, OTTO cannot access pages without credentials. - Crawl budget exceeded or timeout: Very large sites may time out before the crawl completes. This appears as a partial or failed crawl in Crawl Monitoring. - OTTO installation issue: If the OTTO script is not correctly installed, audit features may not function as expected. Verify your setup via Left sidebar → OTTO SEO → Installation Guide. - Temporary platform-side issue: On rare occasions, a crawl may fail due to a transient issue within Search Atlas. These are typically resolved quickly and a recrawl will succeed. ✅ How to Rule Out Issues on Your End Follow these steps to confirm your website is accessible and correctly configured for crawling: 1. Check server uptime: Confirm your site was online at the exact time the crawl was attempted. Cross-reference with your hosting provider's uptime logs. 2. Review robots.txt: Visit yourdomain.com/robots.txt and confirm no rules are blocking all crawlers or OTTO specifically. 3. Disable password protection temporarily: If your site is behind a login or IP restriction, disable it briefly and trigger a recrawl to test access. 4. Verify OTTO installation: Go to Left sidebar → OTTO SEO → Installation Guide and confirm the OTTO script is correctly placed on your site. 5. Test with an external crawler: Use a tool like Google Search Console's URL Inspection to confirm Google can also access your pages. If Google is blocked, OTTO likely is too. 🔄 How to Trigger a Recrawl Once you have identified and resolved the issue, you can manually start a new crawl: 1. Go to Left sidebar → OTTO SEO → All Sites. 2. Select the site you want to recrawl by checking its checkbox. 3. Click the Scan button in the top action bar. 4. Navigate to Left sidebar → OTTO SEO → Crawl Monitoring to watch the crawl progress in real time. If the crawl completes successfully, your updated audit data will appear under Left sidebar → OTTO SEO → All Audits. 📋 What to Do If the Crawl Keeps Failing If you have checked all of the items above and the crawl still fails, gather the following information before reaching out to support: - The domain name of the affected site. - The date and time of the failed crawl (visible in Crawl Monitoring). - Any error messages shown in the crawl log. - Confirmation of whether your robots.txt and server access have been verified. If you need further assistance, open the chat widget in the bottom-right corner of the platform and type human teammate to be connected with a member of our team.

🔍 Fix OTTO Site Audit Crawl Failures and Recrawls

🔍 Overview If your OTTO site audit shows a "Crawl failed. We'll try again soon." message, or if a crawl appears stuck and never completes, this article explains what causes these failures, what OTTO does automatically, and what steps you can take to resolve the issue. ⚠️ What Causes a Crawl Failure? OTTO crawl failures can be caused by issues on either the platform side or your website. Common causes include: - Temporary server unavailability — your site was unreachable at the time of the crawl - Crawl budget or rate-limiting — your server blocked or throttled the crawl bot - robots.txt restrictions — crawling is blocked for the user agent OTTO uses - Post-processing interruptions — the crawl completed but data failed to process correctly afterward - Platform-side processing errors — an internal issue prevented the crawl from finishing or handing off results In most cases, the failure is not caused by something you did wrong. OTTO is designed to detect these failures and retry automatically. 🔄 What OTTO Does Automatically When a crawl fails, OTTO will attempt to retry the crawl on your behalf. You do not need to take any action immediately. However, if the retry does not resolve the issue within a reasonable time, or if the audit remains stuck, you should follow the manual steps below. 🛠️ How to Manually Trigger a Recrawl 1. Go to Left sidebar → OTTO SEO → All Sites. 2. Locate the site showing the crawl failure or stuck audit. 3. Click the Scan button associated with that site to trigger a fresh crawl. 4. Monitor the crawl progress under Left sidebar → OTTO SEO → Crawl Monitoring. 5. Once the crawl completes, check Left sidebar → OTTO SEO → All Audits to confirm audit results have been generated. ✅ How to Rule Out Issues on Your End Before escalating, check the following on your website: - robots.txt — confirm that crawlers are not blocked. Visit yourdomain.com/robots.txt and look for any Disallow rules that may apply to all bots or the OTTO crawl agent. - Server availability — confirm your site was online and responding at the time of the crawl failure. - Firewall or security plugins — tools like Cloudflare, Wordfence, or similar services may block automated crawlers. Temporarily whitelist the crawl or add an exception if needed. - Site speed and stability — if your server is slow or overloaded, the crawl may time out. Check your hosting performance during the crawl window. If none of the above apply and your site appears healthy, the failure is likely on OTTO's side and our team can investigate further. 📊 Check Crawl Status in Real Time Use Left sidebar → OTTO SEO → Crawl Monitoring to view the live status of your crawl. This page shows whether the crawl is in progress, completed, or failed. If the status appears stuck for more than 30 minutes without updating, proceed with a manual recrawl using the steps above. 💡 Tips to Prevent Future Crawl Failures - Ensure your site is accessible to crawl bots and not behind authentication walls during scheduled crawls. - Avoid running large server maintenance tasks or deployments during active crawls. - Keep your robots.txt updated and review it after any site changes. - Check your hosting plan to ensure it can handle crawl traffic without rate-limiting or throttling. 🆘 Still Experiencing Issues? If you have followed the steps above and your crawl is still failing or stuck, our team can investigate whether the issue is on the platform side and manually trigger reprocessing if needed. If you need further assistance, open the chat widget in the bottom-right corner of the platform and type human teammate to be connected with a member of our team.

🔄 Fix a Stalled OTTO Recrawl Investigation

🔍 What Is a Stalled OTTO Recrawl? When OTTO runs a recrawl, it scans your website to refresh audit data and generate updated SEO suggestions. In most cases, a recrawl completes within a few hours. If your recrawl has been running for more than 24 hours without finishing, it is considered stalled. This can happen due to a processing loop, an orphaned reprocess task, or a missing staleness guard that normally resets stuck jobs automatically. ⚠️ Common Causes - Async task loop: The recrawl job enters a repeated "still processing" state and never exits. - Orphaned reprocess task: A background task loses its connection to the crawl job and keeps running silently. - Missing recrawl frequency setting: Some audits were previously set to "never reprocess," which prevented automatic restarts. - Spinner never clears: The UI shows a loading spinner indefinitely, even though the underlying job has stalled. 🧭 How to Check Your Recrawl Status 1. In the left sidebar, click OTTO SEO, then select All Sites. 2. Find your website in the list and check the recrawl status indicator next to it. 3. Click on your site to open its detail view, then navigate to Crawl Monitoring in the left sidebar under OTTO SEO. 4. Review the crawl log for the most recent entry. If the status reads "processing" or "still processing" and the timestamp is older than 24 hours, the recrawl is stalled. 5. You can also go to All Audits to confirm whether audit data or suggestions appear outdated or missing. 🛠️ Steps to Attempt Recovery 1. Go to OTTO SEO → All Sites in the left sidebar. 2. Select the checkbox next to your stalled site. 3. Click Scan at the top of the page to trigger a fresh crawl manually. 4. Wait 10–15 minutes, then return to Crawl Monitoring to verify the new crawl has started and the previous stalled status has cleared. 5. If the spinner is still visible and no progress is shown after 30 minutes, do not trigger additional scans — multiple overlapping jobs can worsen the issue. 📊 What to Expect After a Successful Restart Once the crawl resumes correctly, you will see the status update in Crawl Monitoring and new audit data will appear under All Audits. OTTO suggestions including title tag recommendations and on-page fixes will regenerate automatically. Depending on your site size, this can take anywhere from 30 minutes to a few hours. 💡 How to Prevent Stalled Recrawls - Avoid triggering multiple manual scans in quick succession — let each scan complete before starting another. - Check Crawl Monitoring regularly if you have a large site, as bigger crawls are more prone to timeout issues. - If you notice suggestions are not updating even though no crawl appears to be running, use the Scan button to kick off a fresh recrawl rather than waiting for the automatic schedule. 🆘 When to Contact Support If your recrawl remains stalled after following the steps above, or if triggering a new scan does not resolve the issue, our team can investigate and manually clear the orphaned task on the backend. If you need further assistance, open the chat widget in the bottom-right corner of the platform and type human teammate to be connected with a member of our team. When reaching out, please have the following ready to speed up the investigation: your website URL, the approximate time the recrawl started, and a screenshot of the Crawl Monitoring page showing the stalled status.

🔄 Fix a Stalled OTTO Recrawl

🔍 What Is a Stalled OTTO Recrawl? When OTTO recrawls your website, it audits your pages and generates updated SEO suggestions. A stalled recrawl occurs when the crawl remains in a processing state without completing or producing new results. ⚙️ How to Check Your Recrawl Status 1. Open the affected website in OTTO. 2. Review the displayed recrawl status and note whether it remains in a processing state. 3. Record the project or website name and how long the recrawl has been running. 🛠️ Steps to Resolve a Stalled Recrawl Refresh the page and check the recrawl status again. If it still has not completed, avoid repeatedly starting new recrawls and contact support with the affected project or website name, the current status, and the approximate time the recrawl began. 📊 What to Expect After the Fix If the recrawl remains stalled after refreshing the status, support can investigate the affected project and determine the next step. If you need further assistance, open the chat widget in the bottom-right corner of the platform and type human teammate to be connected with a member of our team.

🛠️ Fix Shopify Title Tags and Crawler Whitelisting

🔍 Overview Two of the most common issues Search Atlas customers encounter with Shopify are: title tag changes not appearing on the live site and the Search Atlas crawler being blocked from accessing the site. This article walks you through the causes and the steps to resolve both problems. ⚙️ Why Title Tag Updates Are Not Appearing Shopify's architecture differs from traditional CMS platforms. Meta title changes made through Search Atlas are applied via your site's deployment channel, but several Shopify-specific factors can prevent those changes from going live. Common reasons title tags are not updating: - Theme not connected: Search Atlas must be connected to your active Shopify theme. If your theme was recently changed or republished, the connection may have been lost. - Deployment not visible: Shopify website deployments are sometimes not detected inside the platform. If your site does not appear under your Website Deployments list, the changes cannot be pushed correctly. - Metafield or app embed not enabled: Search Atlas delivers on-page optimisations through a theme app extension. If the app embed block is not active in your Shopify theme editor, updates will not render on the front end. - Caching: Shopify and third-party CDNs (such as Cloudflare) cache page content aggressively. Even after a successful update, you may see outdated title tags until the cache clears. ✅ How to Fix Title Tag Updates 1. Confirm your Shopify site is connected: Go to Settings → Websites in Search Atlas and verify your Shopify store appears as an active deployment. If it is missing, go to Settings → CMS Connectors and use the Edit, Reconnect, or Disconnect options for your Shopify connection. 2. Enable the app embed block: In your Shopify Admin, navigate to Online Store → Themes → Customize. Open App Embeds and make sure the Search Atlas embed is toggled on. Save the theme. 3. Re-publish your active theme: If you recently switched themes or made theme edits, re-saving and publishing the active theme re-triggers the app extension render. 4. Clear your cache: After confirming the above steps, purge any CDN or browser cache and reload the page. Use an incognito window or a tool like Google's URL Inspection in Search Console to check the live title tag as Google sees it. 5. Check OTTO deployment status: If you are using OTTO SEO, open the OTTO dashboard and confirm the site status shows as Active. A status of Undetected or Inactive means OTTO has not successfully fingerprinted your Shopify store. 🚫 Why the Search Atlas Crawler Is Being Blocked The Search Atlas Site Audit and OTTO use a dedicated crawler to scan your pages. Shopify stores can inadvertently block this crawler in several ways, which prevents audits from running and stops OTTO from reading your live page data. Common causes of crawler blocking: - Password-protected storefront: If your Shopify store is in development mode or behind a storefront password, all external bots — including the Search Atlas crawler — are blocked. - Cloudflare or firewall rules: WAF (Web Application Firewall) rules, bot-fight mode, or custom firewall policies can identify and block the crawler before it reaches your pages. - robots.txt restrictions: A restrictive robots.txt file may be disallowing the crawler user agent from accessing key paths. - Shopify bot protection: Shopify has built-in bot protection that can flag and block unfamiliar crawlers, especially on high-traffic or enterprise plans. ✅ How to Whitelist the Search Atlas Crawler 1. Remove the storefront password: In Shopify Admin, go to Online Store → Preferences and disable the storefront password if it is enabled. 2. Whitelist the crawler IP in Cloudflare: If your store uses Cloudflare, create a WAF Allow rule or an IP Access rule that permits the Search Atlas crawler. Contact our support team via the chat widget to obtain the current crawler IP range. 3. Disable Bot Fight Mode temporarily: In Cloudflare, go to Security → Bots and temporarily disable Bot Fight Mode, then re-run your Site Audit to confirm access. Re-enable it after creating an explicit allow rule. 4. Check your robots.txt: Visit yourstore.com/robots.txt and confirm there are no Disallow directives that would block the Search Atlas user agent. Shopify's default robots.txt is generally crawler-friendly, but custom apps or scripts may have modified it. 5. Re-run the Site Audit: After making these changes, go to OTTO SEO → Site Audit → Overview (Website Overview) in Search Atlas and click Recrawl Site to trigger a fresh scan. Monitor the crawl progress to confirm pages are now being indexed. 📊 How to Verify Everything Is Working - Your Shopify store appears as Active under Website Deployments in Search Atlas. - The Site Audit completes with a full page count and no crawler-blocked errors. - Title tag updates are visible when you inspect the page source or use Google's URL Inspection tool. - OTTO shows a detected CMS: Shopify status on the site dashboard. 💬 Still Need Help? If you have followed all the steps above and are still experiencing issues with title tags or crawler access, our team is ready to investigate your specific setup. If you need further assistance, open the chat widget in the bottom-right corner of the platform and type human teammate to be connected with a member of our team.

🕷️ Fix WordPress Crawl Blocked on SiteGround

🔍 Why Your Site Audit Crawl Gets Stuck If your Site Audit crawl has been running for more than 24 hours without completing — or shows no progress at all — SiteGround's built-in security tools are almost certainly blocking the Search Atlas crawler. This is one of the most common reasons WordPress sites hosted on SiteGround fail to crawl successfully. SiteGround uses two layers of protection that can silently block automated crawlers: - SiteGround Security plugin — flags and bans IP addresses that make rapid, repeated requests (exactly what a crawler does). - Cloudflare integration — SiteGround routes traffic through Cloudflare by default, which can trigger bot-protection rules that block crawl requests before they even reach your server. Because these blocks happen silently, your crawl appears to be running but never actually retrieves any pages. ✅ Before You Start: Confirm the Issue Check these two things first to confirm a block is the cause: 1. Your Site Audit project was added recently and the crawl has not completed after 24+ hours. 2. Your WordPress site is hosted on SiteGround (check your hosting dashboard or domain registrar if unsure). If both are true, follow the steps below in order. 🛠️ Step 1 — Whitelist the Crawler in the SiteGround Security Plugin If you have the SiteGround Security plugin installed on your WordPress site, it may have automatically banned the Search Atlas crawler IP range. Here is how to allow it: 1. Log in to your WordPress admin dashboard. 2. In the left menu, go to Security → Blocked IPs. 3. Look for any recently blocked IP addresses. If you see entries added around the time your crawl started, these are likely the Search Atlas crawler IPs. 4. Select those entries and click Unblock. 5. Next, go to Security → Activity Log Settings and add the Search Atlas crawler user agent (SearchAtlasBot) to your allowlist so it is not blocked in future crawls. If you are unsure which IPs belong to Search Atlas, open the chat widget and a teammate can provide the current IP ranges. ☁️ Step 2 — Adjust Cloudflare Bot Protection SiteGround enables Cloudflare by default for many hosting plans. Cloudflare's bot-fight mode or firewall rules can block the crawler independently of the SiteGround Security plugin. 1. Log in to your SiteGround hosting dashboard. 2. Go to Speed → Cloudflare and click Manage next to your domain. 3. Inside Cloudflare, navigate to Security → Bots. 4. If Bot Fight Mode is enabled, temporarily disable it and restart your crawl to test whether this resolves the issue. 5. Alternatively, go to Security → WAF → Firewall Rules and create an allow rule for the SearchAtlasBot user agent so Cloudflare does not challenge or block it. Note: You can re-enable Bot Fight Mode after whitelisting the Search Atlas user agent. Disabling it permanently is not necessary or recommended. 🔄 Step 3 — Restart Your Site Audit Crawl After completing the steps above, restart the crawl from Search Atlas: 1. In the left sidebar, go to OTTO SEO → Site Audit → Overview (Website Overview). 2. Find your affected project. 3. Click Recrawl Site (or delete and re-add the project if the crawl appears frozen). 4. Monitor progress over the next 15–30 minutes. Most sites complete an initial crawl within this window once the block is lifted. 💡 Tips to Prevent This in the Future - Whitelist SearchAtlasBot before adding new projects — if you host multiple sites on SiteGround, whitelist the crawler user agent proactively in both the SiteGround Security plugin and Cloudflare. - Check your SiteGround Activity Log regularly — the plugin logs blocked IPs in real time, making it easy to spot false positives. - Avoid aggressive rate-limiting rules — if you have custom Cloudflare firewall rules that throttle requests per IP, consider excluding known SEO tool bots from those rules. 💬 Still Seeing Issues? If your crawl is still stuck after following all the steps above, there may be a server-level firewall rule or a custom .htaccess restriction that requires closer inspection. If you need further assistance, open the chat widget in the bottom-right corner of the platform and type human teammate to be connected with a member of our team.

Cloudflare Crawler Whitelisting: Troubleshooting Guide and Alternatives

🔍 Overview When the Search Atlas crawler cannot access a Cloudflare-protected site, the recommended resolution is to have your client whitelist the Search Atlas crawler IPs in their Cloudflare configuration. This guide explains what to request from your client and what information to have ready if you need to escalate. ✅ Step 1: Confirm the Crawler Is Actually Being Blocked Before changing any Cloudflare settings, verify that Cloudflare is the true source of the block. A misconfigured project or incorrect URL can produce the same symptoms. - In your Search Atlas project settings, confirm the target domain is entered correctly, including any subdomain. - Ask your client to check their Cloudflare dashboard for blocked requests around the time the crawl was attempted. If Search Atlas IPs appear as blocked, Cloudflare is actively preventing access. - If no blocks appear in Cloudflare, the issue may be a robots.txt restriction, a server-side firewall, or a login/paywall rather than Cloudflare. Resolve those first before proceeding. ⚙️ Step 2: Review Cloudflare Protection Settings with Your Client Many crawl failures stem from Cloudflare protection settings that your client can adjust without a full IP whitelist. Work with your client to check whether any bot protection, firewall, or rate limiting configurations are blocking the crawler, and whether temporarily adjusting those settings during a scheduled crawl window is feasible. 📋 Step 3: Request IP Whitelisting from Your Client If reviewing and adjusting Cloudflare settings is not sufficient or practical, the resolution is to have your client whitelist the Search Atlas crawler IPs in their Cloudflare configuration. Contact our support team using the method below to obtain the current list of Search Atlas crawler IP addresses to share with your client. 📞 Need Help Escalating? If you are escalating this issue, please have the following ready: the affected project name, the target domain (including any subdomain), and a description of the crawl behavior observed (e.g., fully blocked, partial data returned, specific error messages). If you need further assistance, open the chat widget in the bottom-right corner of the platform and type human teammate to be connected with a member of our team.

🔍 Fix Site Audit Crawler Blocked by Robots.txt

🧭 Overview When you run a site audit in Search Atlas, our crawler needs to access your website's pages to collect SEO data. If your robots.txt file is misconfigured — especially after a WordPress update — it can block our crawler entirely, resulting in incomplete or missing audit results. This article explains why this happens and how to fix it. ⚠️ Why Is the Search Atlas Crawler Being Blocked? There are several common reasons our crawler cannot access your site: - WordPress is set to discourage search engines: WordPress has a built-in setting that adds a Disallow: / directive to your robots.txt, blocking all crawlers. - A plugin overwrote your robots.txt: SEO plugins (e.g., Yoast, Rank Math) or security plugins may have regenerated or modified your robots.txt after an update. - A custom robots.txt rule blocks all bots: A wildcard rule like User-agent: * followed by Disallow: / blocks every crawler, including ours. - Your hosting provider or firewall is blocking crawlers: Some security configurations or WAF (Web Application Firewall) rules may block unknown user agents. 🛠️ How to Fix Robots.txt on WordPress 1. Check the WordPress search engine visibility setting. Log in to your WordPress dashboard. Go to Settings → Reading. Make sure the checkbox labelled "Discourage search engines from indexing this site" is unchecked. Save your changes. WordPress will update your robots.txt automatically. 2. Review your robots.txt file directly. Open a browser and visit yourdomain.com/robots.txt. Look for any line that says Disallow: / under User-agent: *. If you see this, it means all crawlers — including Search Atlas — are blocked. 3. Edit robots.txt via your SEO plugin. If you use an SEO plugin such as Yoast SEO or Rank Math, check your plugin's settings for a robots.txt editing option. Consult your plugin's official documentation for the exact location of this feature, as menu paths may vary by version. Make sure the file contains at minimum: User-agent: * followed by Allow: /. 4. Edit robots.txt via FTP or cPanel (advanced). If you do not use an SEO plugin, connect to your server via FTP or cPanel File Manager. Locate robots.txt in your root directory. Remove any Disallow: / lines that apply to all user agents. Save and re-upload the file. 5. Allow the Search Atlas crawler specifically. If you want to keep other crawlers restricted but allow Search Atlas, you can add a dedicated rule to your robots.txt for the Search Atlas crawler user agent. Place this rule above any wildcard rules to ensure it takes priority. Contact our support team to confirm the exact user-agent string to use. If you need further assistance, open the chat widget in the bottom-right corner of the platform and type human teammate to be connected with a member of our team.

🤖 Fix Unexpected Disallow Entries in robots.txt

🔍 Overview If your live robots.txt file contains Disallow entries that do not appear in the Search Atlas robots.txt editor, those rules are almost certainly being added by your hosting provider, WordPress plugin, or another third-party tool — not by Search Atlas. This article explains why this happens and what you can do about it. ⚠️ Why Extra Disallow Entries Appear Search Atlas generates and previews the robots.txt rules that it controls. However, your website's final, publicly served robots.txt file is assembled by your hosting environment. Common sources of extra Disallow entries include: - Hosting provider defaults — Many managed hosts (e.g. WP Engine, Kinsta, SiteGround) automatically inject their own robots.txt rules, sometimes blocking AI crawlers or staging paths. - WordPress plugins — SEO plugins such as Yoast SEO, Rank Math, or All in One SEO can override or append to the robots.txt file independently of Search Atlas. - CMS-level settings — Some CMS platforms have a built-in robots.txt editor in their admin dashboard that takes priority over other tools. - Server-level configuration — A physical robots.txt file stored in your site's root directory will override virtually generated rules. Because Search Atlas only previews the rules it manages, any entries injected upstream will not appear in the Search Atlas editor — but they will appear when you fetch the live file directly. 🛠️ How to Identify the Source 1. Open your browser and go to https://yourdomain.com/robots.txt to see the live file. 2. Compare that output with what is shown in the Search Atlas robots.txt editor. 3. Any Disallow lines present in the live file but absent from the Search Atlas preview are coming from an external source. 4. Log in to your hosting control panel and check for a robots.txt management section or firewall/crawl settings. 5. In your WordPress dashboard (if applicable), check Settings → Reading and any active SEO plugin for robots.txt controls. 6. Use FTP, SFTP, or your host's file manager to check whether a physical robots.txt file exists in your site's root (/public_html/ or equivalent). ✅ How to Remove the Unwanted Entries Once you have identified the source, follow the appropriate steps below. - Hosting provider rules: Contact your host's support team and ask them to remove or disable any automatic robots.txt injection. Some hosts provide a toggle in their dashboard under crawl or bot-management settings. - Physical robots.txt file: If a static file exists in your root directory, delete it or edit it to remove the unwanted Disallow lines. Search Atlas can then manage the file virtually. - WordPress SEO plugin: Open the plugin's settings, locate its robots.txt editor, and remove any conflicting entries. If you want Search Atlas to be the sole manager, disable the plugin's robots.txt feature entirely. - CMS built-in editor: Navigate to your CMS robots.txt settings and clear any rules that conflict with your Search Atlas configuration. 🔄 Verifying the Fix 1. After making changes, wait a few minutes for your server cache to clear. 2. Reload https://yourdomain.com/robots.txt in your browser. 3. Confirm that the unwanted Disallow entries are gone. 4. Return to Search Atlas and verify that the preview matches the live file. 💡 Best Practices - Designate one tool only as the robots.txt manager for your site to avoid conflicts. - After any hosting migration or plugin update, re-check your live robots.txt file to catch newly injected rules early. - Keep a record of what each Disallow entry is for so you can quickly spot anything unexpected in the future. 🙋 Need More Help? If you have followed the steps above and are still seeing unexpected Disallow entries in your live robots.txt file, If you need further assistance, open the chat widget in the bottom-right corner of the platform and type human teammate to be connected with a member of our team.

🔍 Fix Crawler Blocking and Understand Crawl Scheduling

🗺️ Overview Search Atlas uses its own crawler to audit your site independently of Google. This means two separate processes are at work: the Search Atlas crawler (which powers your site audit data inside the platform) and Google's crawler (which determines your rankings). Understanding the difference — and making sure Search Atlas can actually access your pages — is the first step to accurate audit results. ⏱️ When Do Search Atlas Results Update After Site Changes? Search Atlas does not update automatically the moment you publish changes. Here is how the process works: 1. You trigger or schedule a crawl. Navigate to your site audit dashboard inside the platform and start a fresh crawl manually using the crawl option available there. 2. The Search Atlas crawler visits your pages. Depending on your site size, this can take a few minutes to several hours. 3. Your audit data refreshes. Once the crawl completes, all metrics — broken links, missing tags, page speed issues — reflect the current state of your site. This process is entirely separate from Google. Even after Google re-indexes your updated pages, your Search Atlas audit data will only reflect those changes after a new Search Atlas crawl is run. 🚫 Why Is the Search Atlas Crawler Being Blocked? If you see an error message indicating the crawler is blocked, it means your site's robots.txt file is preventing Search Atlas from accessing your pages. This is one of the most common reasons audit data appears incomplete or stale. A blocked crawler cannot collect data — regardless of what the platform status shows. Do not assume your site is configured correctly if you have received a blocking error. The error message is definitive. Follow the steps below to diagnose and fix the issue. 🔎 How to Check Your robots.txt File Your robots.txt file lives at the root of your domain. To view it, type the following into your browser address bar, replacing the example with your own domain: https://yourdomain.com/robots.txt Look for any lines that contain Disallow rules that could block the Search Atlas crawler. A blocking rule would look like one of these: - User-agent: * followed by Disallow: / — this blocks all crawlers, including Search Atlas. - A user-agent entry for the Search Atlas crawler followed by Disallow: / — this blocks Search Atlas specifically. - Disallow: / appearing under a wildcard rule with no corresponding Allow exception for the Search Atlas crawler. 🛠️ Fixing robots.txt Blocking on WordPress On WordPress, your robots.txt rules are often controlled by an SEO or security plugin rather than a static file. Follow these steps to locate and resolve the blocking rule: 1. Check your SEO plugin settings. If you use a plugin such as Yoast SEO or Rank Math, navigate to its settings in your WordPress admin dashboard and look for a robots.txt editor. Review any Disallow rules and remove or adjust any rule that blocks all bots or specifically blocks the Search Atlas crawler. 2. Check your security plugin settings. Security plugins such as Wordfence include bot-blocking features that can prevent crawlers from accessing your site. In the plugin's settings, review any bot or crawler blocking rules and add an exception to allow the Search Atlas crawler through. 3. Check for maintenance mode. If your site is in maintenance mode (via a plugin or theme setting), all external crawlers will typically be blocked. Disable maintenance mode and re-check your robots.txt file. 4. Verify the fix. After making changes, revisit https://yourdomain.com/robots.txt in your browser to confirm the blocking rule has been removed. Then return to the Search Atlas platform and trigger a new crawl from your site audit dashboard so the crawler can re-access your pages. 5. Confirm audit data updates. Once the new crawl completes, verify that your site audit data is now populating correctly. If the crawler-blocked error persists after following these steps, note your domain, the exact error message shown in the platform, and which plugin(s) you checked so our team can assist you efficiently. If you need further assistance, open the chat widget in the bottom-right corner of the platform and type human teammate to be connected with a member of our team.

🕷️ Fix a Crawler That Misses Pages Beyond the Homepage

🔍 Overview When the Search Atlas crawler only picks up your homepage and ignores the rest of your site, several different root causes could be responsible. This guide walks you through every common reason — in order of likelihood — so you can isolate and fix the problem quickly without guesswork. ⚙️ Step 1: Check Your Crawl Settings Start inside the platform before looking anywhere else. 1. Open Website Studio from the left sidebar (URL: /website-studio). 2. Select the project you want to audit. 3. Go to Crawl Settings and confirm the starting URL is set to your root domain (e.g., https://example.com), not a single page. 4. Make sure Follow Internal Links is enabled. If this toggle is off, the crawler will never leave the starting URL. 5. Save any changes and re-run the crawl. 📏 Step 2: Increase the Crawl Depth Limit Crawl depth controls how many links away from the homepage the crawler is allowed to travel. A depth of 1 means only the homepage; a depth of 2 means the homepage plus pages directly linked from it, and so on. 1. In Crawl Settings, locate the Crawl Depth field. 2. If it is set to 1, raise it to at least 3 for small sites or 5–10 for larger sites. 3. Save and re-run the crawl. If your important pages are buried several clicks from the homepage, a low depth limit is the most common reason they are missed. 🚫 Step 3: Review URL Exclusion Patterns Exclusion rules (sometimes called URL filters or blocklists) tell the crawler to skip certain paths. A pattern that is too broad can accidentally block large sections of your site. 1. In Crawl Settings, open the Excluded URLs or URL Patterns section. 2. Look for wildcard rules such as /blog/*, /products/*, or /*?* that could be sweeping up pages you actually want crawled. 3. Remove or narrow any rule that unintentionally targets your missing pages. 4. Save and re-run the crawl. ✅ Step 4: Confirm Your Domain Is Whitelisted The crawler only follows links that belong to the whitelisted domain. If your site uses subdomains (e.g., shop.example.com) or a www vs. non-www variation, those must be explicitly allowed. 1. In Crawl Settings, find the Allowed Domains or Whitelist field. 2. Add every variation of your domain that hosts content — including subdomains, www, and non-www versions. 3. Save and re-run the crawl. 🤖 Step 5: Audit Your robots.txt File A misconfigured robots.txt file on your server can instruct crawlers to skip entire sections of your site. This is a server-side issue, so you will need to check it outside the platform. 1. Visit https://yourdomain.com/robots.txt in your browser. 2. Look for Disallow rules that cover the paths where your missing pages live (e.g., Disallow: /category/). 3. If the Search Atlas crawler is listed under a specific User-agent with restrictive rules, those rules apply to it. 4. Update your robots.txt to allow the paths you want crawled, then re-run. 🗺️ Step 6: Check Your XML Sitemap The crawler uses your sitemap as a discovery aid. If pages are missing from it, they may never be found — especially if internal linking is sparse. 1. Visit https://yourdomain.com/sitemap.xml to confirm it exists and loads correctly. 2. Verify that the pages you expect to be crawled are listed. 3. If your sitemap is empty, outdated, or missing, regenerate it through your CMS or SEO plugin. 4. In Crawl Settings, enter your sitemap URL under Sitemap so the crawler can reference it directly. 🔗 Step 7: Inspect Internal Linking on the Homepage If the crawler finds links correctly but your internal pages are still missing, the homepage itself may not be linking to them — or those links may not be crawlable. - Links inside JavaScript-rendered menus or single-page application (SPA) frameworks may not be visible to the crawler. - Links using onclick events instead of standard <a href> tags are typically not followed. - Ensure key pages are reachable through standard HTML anchor links from at least one crawlable page. 🌐 Step 8: Rule Out Server-Side Blocking Some hosting environments or security plugins (e.g., Cloudflare firewall rules, rate limiters, or bot-protection services) can block crawlers by IP or User-agent before they even reach your pages. - Check your hosting control panel or CDN dashboard for any rules that might block automated traffic. - Temporarily whitelist the Search Atlas crawler User-agent if your host allows it. - If you see 403 Forbidden or 429 Too Many Requests errors in the crawl report, server-side blocking is likely the cause. 🔁 Quick Troubleshooting Checklist - Follow Internal Links toggle is on - Crawl depth is set to 3 or higher - No overly broad URL exclusion patterns - All domain variants are whitelisted - robots.txt does not block the target paths - XML sitemap is valid and entered in crawl settings - Key pages are linked via standard HTML anchors - Server or CDN is not blocking crawler traffic If you need further assistance, open the chat widget in the bottom-right corner of the platform and type human teammate to be connected with a member of our team.

🛡️ Site Audit Blocked Competitor Domains Explained

🔍 Overview When running a Site Audit on a competitor's domain, you may occasionally find that the audit cannot complete or returns an error. This is not a bug in Search Atlas — it is caused by the competitor's website actively blocking the audit crawler. This article explains why this happens and what steps you can take. 🚧 Why Some Competitor Domains Are Blocked Many large websites use a content delivery network (CDN) called Amazon CloudFront to serve their content and protect their infrastructure. CloudFront includes security rules that can detect and block automated crawlers, including the Search Atlas Site Audit bot. When a domain has these protections active, our crawler is denied access before it can gather any data. This is a deliberate security decision made by the website owner — it is entirely outside Search Atlas's control. ⚙️ What Happens During a Blocked Audit If you attempt to audit a competitor domain that is protected by CloudFront or a similar security layer, the audit may fail to return meaningful data. This outcome does not indicate a problem with your Search Atlas account or project setup — it reflects the competitor's server actively rejecting crawler access. ✅ What You Can Do While you cannot force access to a domain that actively blocks crawlers, there are several productive steps you can take: 1. Confirm the block is on their end. Try accessing the competitor's sitemap directly in your browser (e.g., example.com/sitemap.xml). If it loads normally for you but the audit fails, the block is targeting automated tools specifically. 2. Use alternative competitor research tools within Search Atlas. Other tools in the platform rely on Search Atlas's existing index rather than live crawling, meaning they are not affected by the competitor's crawler-blocking configuration. Explore the available research and analysis features to gather competitor intelligence without needing to crawl their domain directly. 3. Monitor the domain over time. Security configurations change. You can retry the Site Audit on the competitor domain at a later date — some websites relax or update their crawler rules periodically. 4. Focus on auditable competitors. Identify alternative competitors in your niche whose domains are not blocked, and use those audits to inform your strategy instead. ❓ Frequently Asked Questions Is this a Search Atlas bug or limitation? No. CloudFront blocking is a security measure applied by the website owner. Search Atlas has no ability to override or bypass another website's security infrastructure. Can Search Atlas fix this for me? Because the block is enforced by the competitor's server, there is no action our team can take to grant crawler access to a third-party domain. If you are unsure whether the issue is caused by blocking or something else, contact our support team with your project name and the domain you are trying to audit so we can investigate. If you need further assistance, open the chat widget in the bottom-right corner of the platform and type human teammate to be connected with a member of our team.

🔄 Fix Site Report Refresh and Crawl Quota Errors

🔍 Overview If your site report appears to have finished but the data hasn't updated — or you're seeing a message that you've run out of crawl quota — this guide explains what may be happening and how to get the right help. ⚠️ Why Your Site Report May Not Refresh A site report can display a completed status while still showing outdated data. This is a known issue that our team can investigate on a per-account basis. Common reported scenarios include: - The report status shows completed but data looks unchanged. This may indicate a processing or sync issue on the account. - A crawl quota error message appears. This means the account has reached its crawl limit for the current period and no new crawls can run. - The report appears stalled or stuck in progress. This may require a support team member to review the crawl status on the backend. 📊 Understanding Crawl Quota Search Atlas plans include a crawl quota — a limit on the number of pages that can be crawled within a given period. When this quota is reached, new crawls cannot be initiated until it resets or is increased. If you are seeing a quota-exceeded message, our team can confirm your current limit and advise on next steps. 🛠️ What to Do If Your Site Report Is Stalled or Showing a Quota Error 1. Note the exact error message or status displayed in your site report. 2. Note the domain or project name affected. 3. Note the approximate time the issue started or the crawl was triggered. 4. Contact our support team with this information so a teammate can review your account and crawl status directly. If you need further assistance, open the chat widget in the bottom-right corner of the platform and type human teammate to be connected with a member of our team.

🌐 Fix Your Website Not Globally Crawlable Flag

🔍 What Does 'Not Globally Crawlable' Mean? When Google flags your website as not globally crawlable, it means Googlebot is being blocked from accessing your site from certain regions or entirely. This prevents Google from fully indexing your pages, which can seriously hurt your search visibility and rankings. This flag typically appears in Google Search Console and signals that something in your site's configuration is restricting crawl access at a global level. ⚠️ Common Causes of This Issue - IP-based geo-blocking: Your server or CDN is blocking requests from Google's IP ranges, which are routed through US-based data centers but crawl on behalf of all regions. - Robots.txt misconfiguration: A disallow rule is unintentionally blocking Googlebot from crawling key pages or your entire site. - Firewall or security rules: Web application firewalls (WAFs) or DDoS protection tools (e.g., Cloudflare) may be flagging Googlebot as suspicious traffic and blocking it. - CDN or hosting restrictions: Your content delivery network may have regional access rules that prevent crawlers from non-local IP ranges. - HTTP authentication: Password-protected staging environments or pages requiring login block all crawlers by default. - VPN or proxy detection: Some security plugins block requests that appear to come from proxy servers or data centers, which includes Googlebot. 🛠️ How to Diagnose the Problem 1. Check Google Search Console: Navigate to Settings → Crawling in Google Search Console. If the "Not globally crawlable" warning is active, note any additional details provided. 2. Test your robots.txt: Go to https://yourdomain.com/robots.txt and review the rules. Use Google Search Console's robots.txt Tester to check if Googlebot is being blocked. 3. Use the URL Inspection Tool: In Google Search Console, use the URL Inspection Tool on your homepage and key pages. Select Test Live URL to see if Googlebot can currently access those pages. 4. Review your firewall and CDN settings: Log in to your hosting panel, CDN dashboard (e.g., Cloudflare), or WAF settings and check for any rules that block bot traffic or restrict access by IP range or geography. 5. Check for HTTP authentication: Confirm that no password protection is active on your live site, especially if you recently migrated from a staging environment. ✅ How to Fix the Issue - Update robots.txt: Remove any Disallow: / rules that block all crawlers. Make sure the rule targeting Googlebot does not restrict access to important pages. A correctly configured robots.txt should include User-agent: * followed only by specific disallow rules for pages you intentionally want excluded. - Whitelist Googlebot in your firewall or CDN: Add Google's verified crawler IP ranges to your allowlist. You can find the current list at developers.google.com/search/apis/ipranges/googlebot.json. In Cloudflare, create a firewall rule that allows requests where the verified bot equals Googlebot. - Disable geo-blocking for crawlers: If you use server-side geo-restrictions, configure an exception for verified search engine bots. Work with your hosting provider or developer to implement this safely. - Remove HTTP authentication from live pages: Ensure your production environment does not have basic auth enabled. This is a common oversight after launching from a staging setup. - Disable aggressive bot-blocking plugins: If you use a WordPress security plugin (e.g., Wordfence, iThemes Security), check that it is not blocking requests from data center IPs, as Googlebot crawls from data centers. 📊 Verify the Fix in Search Atlas After making changes, use Search Atlas to monitor your site's crawlability and indexing health. Navigate to Left sidebar → Site Metrics (Site Explorer) to review your site's overall performance and spot any remaining technical issues affecting your search visibility. You should also return to Google Search Console and re-run the URL Inspection Test on your homepage to confirm Googlebot can now access your site. It may take a few days for the flag to clear after Google re-crawls your site. 💡 Prevention Tips - Always review your robots.txt before and after a site migration or relaunch. - Test crawl access in a staging environment before going live. - Set up Google Search Console alerts so you are notified immediately if crawl issues are detected. - Regularly audit your firewall and CDN rules to ensure legitimate bots are not accidentally blocked. 🙋 Still Need Help? If you need further assistance, open the chat widget in the bottom-right corner of the platform and type human teammate to be connected with a member of our team.

🕷️ Search Atlas Crawler Blocked by SiteGround

🔍 Overview SiteGround's built-in security firewall can intercept Search Atlas crawler requests and return a challenge page instead of your actual page content. This means your site may appear to crawl successfully inside the Search Atlas dashboard while the crawler is actually receiving a blocked or challenged response — not real content. This article explains why this happens, how to confirm it, and what steps to take to resolve it. ⚠️ Why the Dashboard May Show a Successful Crawl Search Atlas records a crawl as successful when the server returns a successful HTTP status code. In some cases, SiteGround's firewall may return a response code that the platform interprets as a success — even when the response body is a CAPTCHA or challenge page rather than your actual content. This is a known limitation that our engineering team is aware of and is reviewing. 🛠️ How to Confirm SiteGround Is Blocking the Crawler 1. Log in to your SiteGround Site Tools dashboard. 2. Review your site's security and traffic logs to look for repeated requests from Search Atlas crawler IP addresses that are being blocked or challenged by the firewall. 3. Check your server access logs to see the exact responses being served to the Search Atlas crawler. 4. If you are unsure where to find these logs in your SiteGround account, consult SiteGround's own support documentation or contact their support team for guidance on locating blocked traffic records. ⚙️ About the Search Atlas Crawler User-Agent and IP Addresses Customers have asked us to clarify two specific technical points: - User-agent value: The Search Atlas crawler identifies itself with a user-agent string that includes SearchAtlas. To confirm the precise value being sent for your site, contact our support team and provide your project name so we can look up the exact crawler configuration for your account. - IP addresses: Search Atlas uses crawler IP addresses that may need to be whitelisted in SiteGround's firewall. We do not publish the full IP range publicly, as it is subject to change. Our support team can provide the current IP list relevant to your account so you can whitelist them in SiteGround. ✅ How to Whitelist the Search Atlas Crawler in SiteGround 1. Contact our support team (see below) and provide your project name and the domain being crawled. We will supply you with the current crawler IP addresses to whitelist. 2. Log in to SiteGround Site Tools and navigate to the security or IP management section of your account. If you are unsure of the exact location, SiteGround's support documentation can guide you to the correct panel for your plan. 3. Add each Search Atlas crawler IP address provided by our support team to your whitelist and save your changes. 4. Trigger a new crawl in Search Atlas and verify that real page content is now being returned. 📋 What to Have Ready When You Contact Support Because resolving a SiteGround firewall block requires backend investigation, please have the following information ready when you reach out: - Your project name in Search Atlas - The domain that is being crawled - Any error messages or status codes you have observed in your SiteGround logs or in the Search Atlas dashboard - The date and approximate time when the crawl issue was first noticed If you need further assistance, open the chat widget in the bottom-right corner of the platform and type human teammate to be connected with a member of our team.

🛠️ Why Does Crawl Monitoring Show 'No Data'? Troubleshooting Guide

Crawl Monitoring in Search Atlas may display No data for several expected reasons — a new or low-traffic domain, a mismatched date range, or a plan that does not include the feature. Use this guide to identify the cause and determine the right next step. 📊 What Crawl Monitoring Tracks Crawl Monitoring shows how search engines and AI bots interact with your website. It captures: - Total crawl requests - Average response time - Total download size - Crawler distribution - Historical crawl activity - Site indexation percentage - Crawl purpose (discovery vs. refresh) - Device / screen size - Load time distribution - Crawl frequency - Crawled pages 🤖 Crawl Analysis by Bot The Crawl analysis by bot subsection lets you filter crawl data by a specific bot and view dedicated donut charts for: - Site indexation % - Crawl purpose - By device - Load time distribution A healthy state shows properly sized donuts, a scrollable bot selector, and aligned legends — all of which render correctly after a recent UI fix. 🔍 Common Reasons Crawl Monitoring Shows No Data The most common causes fall into two scenarios depending on whether the dashboard has ever shown data before. Case 1: The dashboard has never shown data If Crawl Monitoring has never displayed data, one of the following is likely the cause: - The domain is new or recently launched - The domain has not had enough crawl activity yet — newer or low-traffic sites may need several weeks to months before meaningful data appears - The site has limited visibility or low crawler traffic - The account is not on a plan that includes Crawl Monitoring — confirm your tier (Pro, Agency, or Enterprise) on the Search Atlas pricing page - Search engines or AI bots have not crawled the site during the selected Daily, Weekly, or Monthly tab range in OTTO SEO → Crawl Monitoring Case 2: The dashboard previously had data but now shows none If Crawl Monitoring was working before and has suddenly gone empty, check the following: - Confirm the selected Daily, Weekly, or Monthly tab in OTTO SEO → Crawl Monitoring covers a period with known crawler activity - Check whether the domain was recently re-configured, re-verified, or changed in your account - In Crawl analysis by bot, set the bot selector dropdown to All bots or a bot with known activity - Confirm the site is still indexable and has not been blocked by an updated robots.txt - If none of the above explains the drop, this may indicate a data pipeline issue worth escalating to support ✅ What to Check First Run through this checklist before reporting an issue: 1. Confirm the selected tab in OTTO SEO → Crawl Monitoring 2. Try switching between Daily, Weekly, and Monthly tabs 3. Check whether the domain is new or recently launched 4. Confirm the site is indexable and not blocked by robots.txt 5. Confirm the account plan includes Crawl Monitoring 6. In Crawl analysis by bot, confirm the bot selector is set to All bots or a bot with known activity — not a specific low-traffic bot 7. Allow more time for crawler activity to accumulate if the site is new or low-traffic After completing these steps, you will know whether the empty dashboard is expected or signals a real issue that needs investigation. 💡 Crawl Monitoring and Domain Age Crawl Monitoring becomes more useful as your domain builds crawl history. For newer domains, some charts may remain empty until search engines and AI bots begin crawling the site more frequently. ❓ Frequently Asked Questions Does 'No data' mean Crawl Monitoring is broken? Not necessarily. No data is expected for new domains or sites that have not yet accumulated enough crawler activity. How old does a domain need to be before Crawl Monitoring shows data? There is no fixed threshold. Newer or low-traffic domains may need several weeks to months of active crawl activity before meaningful data appears. Is Crawl Monitoring available on every plan? No. Crawl Monitoring is available on Pro, Agency, and Enterprise plans. Check the Search Atlas pricing page to confirm whether your plan includes this feature. 🎯 You now know how to diagnose a "No data" state in Crawl Monitoring and what to do next. If the dashboard remains empty after waiting several weeks and ruling out every cause above, contact the Search Atlas support team so we can investigate further.

🗺️ Fix Empty Video and News Sitemaps

🔍 Overview If your video-sitemap.xml or news-sitemap.xml is returning an empty URL set despite having published video pages or recent news articles on your site, the most common cause is that the relevant custom post types have not been enabled in your sitemap settings. This article walks you through how to identify and fix the issue. ❓ Why Are My Sitemaps Empty? Search Atlas generates video and news sitemaps based on the post types you explicitly select in the sitemap configuration. By default, only standard post types (such as Posts and Pages) may be active. If your video content or news articles live inside a custom post type, those URLs will be excluded from the sitemap until you enable them manually. Common scenarios where this happens include: - Video content is stored in a custom post type such as Videos, Media Posts, or a plugin-generated type. - News articles are published under a custom post type rather than the default Posts type. - A recent site migration or plugin change created new post types that were never added to the sitemap settings. 🛠️ How to Enable Custom Post Types in Your Sitemap Settings 1. Log in to your WordPress admin panel for your connected site. 2. Navigate to Search Atlas → Settings to open the sitemap settings area. 3. Locate the section for the sitemap type that is empty — either the video sitemap or the news sitemap. 4. Find the option to select or enable post types for that sitemap. You should see a list of post types registered on your site. 5. Enable each custom post type that contains your video or news content. 6. Save your settings using the available save option in the interface. 7. If a sitemap regeneration option is available, trigger it; otherwise, wait for the next scheduled refresh. 8. Once regenerated, visit your video-sitemap.xml or news-sitemap.xml URL directly in your browser to confirm URLs are now appearing. ✅ How to Verify the Fix Worked After saving and regenerating, confirm the following: - Open yourdomain.com/video-sitemap.xml or yourdomain.com/news-sitemap.xml in your browser. You should now see a list of URLs matching your published content. - The number of URLs listed should reflect your published posts in those custom post types. - If you use Google Search Console, resubmit your sitemap after the fix to prompt Google to recrawl. 💡 Tips to Prevent This Issue - Whenever you install a new plugin that registers additional post types, review your sitemap settings to ensure those types are included if needed. - After any site migration, audit your sitemap configuration to confirm all relevant post types are still enabled. If you need further assistance, open the chat widget in the bottom-right corner of the platform and type human teammate to be connected with a member of our team.

🔍 Fix Crawler Blocks and Recrawl After Updates

🧭 Overview Search Atlas does not automatically recrawl your site when you make website changes. To see updated audit results, you must trigger a manual recrawl. Additionally, if your WordPress robots.txt file is blocking the Search Atlas crawler, your audit will fail or return incomplete data. This article walks you through both fixes. 🔄 How Site Audits Work in Search Atlas Search Atlas runs site audits based on snapshots taken at the time of the crawl. Unlike Google, the platform does not passively monitor your site for changes. This means: - Updates you make to your website are not reflected automatically in your audit results. - You must manually trigger a new crawl after making changes to see updated data. - Results shown in your dashboard represent the state of your site at the time of the last completed crawl. Waiting for a Google crawl will not update your Search Atlas audit — these are independent processes. 🚀 How to Manually Recrawl Your Site 1. Log in to your Search Atlas account. 2. Navigate to the site audit section of the platform for the project you want to update. 3. Locate the project for the website you have updated. 4. Open the site audit for that project. 5. Look for the option to run a new crawl or trigger a recrawl for that project and select it. 6. Wait for the crawl to complete — this may take a few minutes depending on site size. 7. Once finished, your audit results will reflect the current state of your website. Tip: Run a manual recrawl any time you publish significant content changes, fix technical SEO issues, or update your site structure to keep your audit data accurate. 🛠️ Why the Search Atlas Crawler Gets Blocked The most common reason a site audit fails to run is that the Search Atlas crawler is being blocked by your website's robots.txt file. This happens frequently on WordPress sites, especially after plugin updates or security setting changes. Signs that your crawler is blocked include: - The site audit fails to complete or returns no data. - You see an error indicating the crawler could not access your site. - Only a very small number of pages are returned in the audit results. ⚙️ How to Fix a WordPress robots.txt Crawler Block 1. Access your WordPress site and navigate to your robots.txt file. You can view it directly by going to yourdomain.com/robots.txt in your browser. 2. Look for any Disallow rules that may be blocking all crawlers or specifically blocking the Search Atlas user agent. A common culprit looks like this: Disallow: / under a User-agent: * rule. 3. To edit your robots.txt file in WordPress, access your WordPress dashboard and use an SEO plugin (such as Yoast SEO or Rank Math) that includes a robots.txt editor, or edit the file directly via your hosting file manager or FTP. 4. Remove or update the Disallow rule that is blocking crawlers so that the Search Atlas crawler can access your site. If you only want to block specific bots, replace the broad User-agent: * rule with more targeted entries. 5. Save your changes and verify the updated robots.txt by revisiting yourdomain.com/robots.txt in your browser. 6. Once the block is removed, return to Search Atlas and trigger a manual recrawl as described above to get updated audit results. Note: If you are unsure which plugin or setting is generating your robots.txt rules, temporarily deactivating recently updated security or SEO plugins can help identify the source of the block. If you need further assistance, open the chat widget in the bottom-right corner of the platform and type human teammate to be connected with a member of our team.

🔍 Find and Fix Crawl Issues in Site Auditor

🗺️ Overview Site Auditor scans your website and flags technical issues — including broken links (404 errors) — that can hurt your search rankings. This article explains how to locate crawl issues and 404 errors inside Site Auditor and what you can do to resolve them. 🧭 How to Navigate to Site Auditor Follow these steps to open Site Auditor from any page in Search Atlas: 1. Log in to your Search Atlas account. 2. Look at the left sidebar and click OTTO SEO, then select Site Audit (Site Auditor) from the nested menu. 3. If you have more than one project, select the correct website for the project you want to audit. 4. Wait for the dashboard to load. If a crawl has already been run, your results will appear immediately. If this is your first visit, initiate a new crawl using the option available on the dashboard. Note: A crawl may take a few minutes to complete depending on the size of your website. Wait for the crawl to finish before reviewing results. 📋 How to Find Crawl Issues Once the crawl is complete, your results are displayed on the Site Auditor dashboard. Here is how to locate your issues: 1. On the main Site Auditor dashboard, look for the section that lists detected problems. This is where every crawl issue is reported after scanning completes. 2. Browse through the issues list to surface 404 or broken link errors specific to your site. 3. Click on any issue to expand it and see further details about the affected URLs. Tip: If you see no issues listed but expect problems, check whether the crawl finished successfully. If the crawl is still running, wait for it to complete and then refresh the page. 🔗 How to View 404 Broken Link Details Once you have located 404 errors in the dashboard, Site Auditor provides additional information to help you understand the scope of the problem: - Broken destination URL: The page that returned a 404 error. - Referring page: The page on your site that contains the broken link. Reviewing these details tells you exactly which pages you need to edit in order to resolve each error. 🛠️ How to Fix 404 Errors There are two main ways to fix a broken link, depending on whether the destination page exists or not: 1. If the destination page has moved: Set up a 301 redirect from the old URL to the new one. This preserves link equity and automatically sends visitors to the correct page. 2. If the destination page no longer exists: Edit the linking page and remove or update the broken link so it points to a valid, relevant URL instead. After making your fixes, return to Site Auditor and run a new crawl to confirm the 404 errors have been resolved. If you need further assistance, open the chat widget in the bottom-right corner of the platform and type human teammate to be connected with a member of our team.

🛠️ Fix Shopify Crawl-Delay in robots.txt

🔍 What Is the Crawl-Delay Issue? If your Search Atlas dashboard is flagging a crawl-delay directive as a critical error in your robots.txt file, you are not alone. This is a common issue for Shopify store owners. Shopify automatically inserts a crawl-delay: 1 directive into every store's robots.txt file — and it cannot be removed or edited by the store owner. The directive looks like this inside your robots.txt file: - User-agent: * - Crawl-delay: 1 This means Shopify is instructing all bots to wait 1 second between each page request when crawling your site. While this protects server performance, it can slow down how quickly search engines index your content. ⚠️ Why Can't I Delete It? Shopify does not allow merchants to directly edit the robots.txt file in the same way a WordPress or custom-built site would. The platform generates this file automatically, and the crawl-delay line is hard-coded into Shopify's default template. Even if you use Shopify's robots.txt customisation feature (available on certain plans), fully removing the crawl-delay directive is not straightforward and depends on your Shopify plan and theme setup. This is a platform-level limitation, not a misconfiguration on your part. 📊 Does Crawl-Delay Actually Hurt My SEO? The impact depends on your site size and your crawl budget: - Small stores (under 500 pages): The 1-second crawl-delay is unlikely to cause meaningful indexing delays. Googlebot and other major crawlers often ignore or deprioritise this directive. - Large stores (500+ pages, frequent updates): A crawl-delay can slow down how quickly new or updated pages are discovered and indexed, which may affect time-sensitive content like product launches or price changes. - Google specifically: Google has publicly stated that Googlebot does not support the crawl-delay directive. However, other bots (Bingbot, and various third-party crawlers) do respect it. Search Atlas flags this as a critical error because it represents a technical SEO risk that should be reviewed — even if the real-world impact varies by site. 🛠️ What Can You Do About It? While you cannot fully remove the directive on standard Shopify plans, there are steps you can take to manage and mitigate the issue: 1. Verify the directive is present: Go to yourdomain.com/robots.txt in your browser and confirm the crawl-delay line exists. This helps you understand exactly what search engines are seeing. 2. Check your Shopify plan: Shopify Plus merchants have access to a liquid-based robots.txt template that allows more granular customisation. If you are on Shopify Plus, work with a developer to modify the robots.txt.liquid template and remove or override the crawl-delay directive. 3. Submit your sitemap manually: Compensate for slower crawling by ensuring your XML sitemap is submitted in Google Search Console and Bing Webmaster Tools. This gives search engines a direct map of your pages regardless of crawl speed. 4. Use Google Search Console's URL Inspection tool: For high-priority pages (new products, landing pages), use the Request Indexing feature in Google Search Console to prompt faster indexing without relying on a crawl. 5. Monitor crawl coverage regularly: Inside Search Atlas, use the OTTO SEO section (left sidebar → OTTO SEO → All Sites (SEO Automation)) to track crawl errors and indexing status over time. This helps you catch any pages that are falling behind on indexation. ✅ How to Mark This Error as Acknowledged in Search Atlas If you have reviewed the issue and determined the impact is low for your store size, you can acknowledge the error inside Search Atlas to keep your dashboard clean while you monitor the situation: 1. Navigate to the site audit section of your Search Atlas dashboard. 2. Locate the crawl-delay critical error in your robots.txt findings. 3. Review the recommendation details and confirm you have followed the mitigation steps above. 4. Use the available acknowledgement or dismiss option if your workflow allows for it, and add a note referencing this article for your records. Acknowledging the error does not fix the underlying directive, but it signals that your team is aware and has taken appropriate action given the platform constraints. 💡 Key Takeaways - The crawl-delay directive in Shopify robots.txt is added automatically by Shopify and cannot be removed on standard plans. - Google ignores the crawl-delay directive, but other bots respect it. - The real-world SEO impact is low for small stores but worth monitoring for large or frequently updated catalogues. - Shopify Plus merchants can customise the robots.txt.liquid file to remove the directive. - Submitting your sitemap and using URL inspection tools in Google Search Console are the best workarounds for all Shopify plans. If you need further assistance, open the chat widget in the bottom-right corner of the platform and type human teammate to be connected with a member of our team. Additional Notes Method 1: Check for Apps Injecting the Directive — third-party SEO/security apps (e.g. SEO Manager, Plug in SEO) may inject crawl-delay. Disable apps one at a time, check live robots.txt, or find crawl settings within the app and set value to 0. Wait 24–48 hours then re-crawl in Search Atlas.

🔍 Fix Site Audit Sitemap Detection Discrepancies

🗺️ Why This Discrepancy Happens When Site Audit reports that a page is missing from your XML sitemap, but you can see the page listed in your live sitemap, the most common cause is a data freshness issue. Site Audit captures a snapshot of your site at the time the crawl runs. If your sitemap was updated after the crawl began — or after the last completed crawl — the audit results will not reflect those changes until the next crawl is processed. Other frequent causes include: - Sitemap caching: Your server or CDN may serve a cached version of the sitemap to crawlers. The cached copy Site Audit fetched may have been outdated at crawl time, even if your live sitemap is now correct. - Multiple sitemap files: If your site uses a sitemap index file pointing to several child sitemaps, Site Audit may have successfully fetched some child sitemaps but not others due to a temporary fetch error or timeout. - Conditional serving: Some sitemap plugins or CMS configurations serve different sitemap content based on the requesting user-agent or IP address. The crawler may receive a different response than a browser does. - Redirect chains on the sitemap URL: If the sitemap URL itself redirects before returning XML content, the crawler may not follow all redirect hops, resulting in an incomplete or empty sitemap read. - Recently added pages: Pages added to your sitemap shortly before or during an active crawl may fall outside the crawl window and only appear in the next scheduled audit. 🛠️ How to Troubleshoot the Discrepancy 1. Re-run the Site Audit crawl. Navigate to your Site Audit dashboard and trigger a fresh crawl. Once the crawl completes, check whether the page is still flagged as missing. This is the quickest way to rule out a stale-data issue. 2. Verify the sitemap URL used by Site Audit. Open your Site Audit project settings and confirm the sitemap URL entered matches your actual live sitemap URL exactly, including the protocol (https://) and any subdirectory path. A mismatch here means the tool is checking a different file than the one you are viewing in your browser. 3. Fetch the sitemap URL directly. Paste your sitemap URL into a browser or a tool like Google Search Console's Sitemap report. Confirm the page in question is listed and that the sitemap returns a valid HTTP 200 status with correct XML markup. 4. Check for sitemap caching. Add a cache-busting query string (e.g., ?nocache=1) to your sitemap URL in a browser to see if you get a fresher version with more entries. If the cached and uncached versions differ, work with your hosting provider or CDN to reduce the sitemap cache TTL. 5. Inspect child sitemaps in a sitemap index. If you use a sitemap index, open each child sitemap URL individually and confirm the page appears in the correct child file. If a child sitemap returns an error or is empty, that file needs to be fixed at the CMS or plugin level. 6. Review the page's canonical tag and index status. Occasionally, Site Audit flags a page as missing from the sitemap because the page itself contains a noindex directive or a canonical tag pointing to a different URL. Site Audit may exclude noindexed URLs from sitemap coverage checks by design. Confirm the page is indexable before expecting it to appear as covered. 7. Check for redirect chains on the page URL. If the URL listed in your sitemap redirects to another URL, Site Audit may record the final destination URL as the canonical version. Compare the URL in your sitemap against the URL Site Audit tracked after following redirects. 📋 Best Practices to Prevent Future Discrepancies - Keep your sitemap URL stable and ensure it returns a 200 OK status at all times. - Avoid making large sitemap changes while a crawl is actively running. - Set your CDN or server cache TTL for sitemap files to 1 hour or less so crawlers always receive a near-current version. - Use a single, authoritative sitemap index file rather than submitting multiple disconnected sitemap URLs. - After publishing new pages, allow at least one full audit cycle before expecting them to appear as covered in Site Audit reports. 💡 Understanding the Audit Timestamp Every Site Audit report displays a last crawled timestamp at the top of the dashboard. Always compare this timestamp against the date you last updated your sitemap. If your sitemap was modified after the crawl timestamp, the current report cannot reflect those changes. Re-running the crawl will bring the report up to date. 🙋 Still Seeing the Discrepancy? If you have re-run the crawl and verified all the steps above but the page is still showing as missing, there may be a crawl configuration or account-level issue that requires investigation. If you need further assistance, open the chat widget in the bottom-right corner of the platform and type human teammate to be connected with a member of our team. Additional Notes Step 2: Clear Your Browser Cache and Reload — A cached version of the Site Audit dashboard can display outdated results even after a new crawl has completed. Clear cached images/files and cookies (Ctrl+Shift+Delete / Cmd+Shift+Delete), set time range to All time, then reopen browser and log back in.

🛠️ Fix WordPress Crawls and Onboard Agency Clients

🔍 Overview This article covers two common questions from new agency users: why a WordPress site may not crawl after connecting, and how to get the most out of Search Atlas when onboarding new clients. Follow the steps below to resolve crawl issues and explore the platform's core capabilities. ⚙️ Why Your WordPress Site Isn't Crawling After connecting WordPress to Search Atlas, a crawl may fail to start or return no data. This is almost always caused by one of the following: - Plugin not activated: The Search Atlas WordPress plugin must be installed and activated on your site. Without it, the platform cannot communicate with your WordPress installation. - Incorrect site URL: The URL entered during connection must exactly match your WordPress site address, including whether it uses www or not, and https vs http. - Firewall or security plugin blocking crawlers: Plugins such as Wordfence, Cloudflare firewall rules, or password-protected staging environments can block the Search Atlas crawler. Temporarily whitelist the crawler IP or disable bot-blocking rules during the crawl. - robots.txt disallowing crawl: Check your robots.txt file to ensure it does not block all bots or the Search Atlas user agent. - Site is not publicly accessible: If your WordPress site is in maintenance mode or behind a login wall, the crawler cannot access it. 🔧 Steps to Fix a WordPress Crawl 1. Log in to your WordPress admin dashboard and confirm the Search Atlas plugin is installed and active under Plugins → Installed Plugins. 2. In Search Atlas, navigate to your site settings and verify the connected URL matches your live WordPress address exactly. 3. Check your robots.txt file by visiting yourdomain.com/robots.txt. Make sure no Disallow: / rule is blocking all bots. 4. If you use a security plugin or Cloudflare, temporarily pause bot-protection rules and attempt the crawl again. 5. Ensure the site is publicly accessible — disable maintenance mode and remove any login requirements on the front end. 6. Return to Search Atlas, open your project, and trigger a new crawl manually. If the crawl still does not complete after following these steps, if you need further assistance, open the chat widget in the bottom-right corner of the platform and type human teammate to be connected with a member of our team. 🚀 How Agencies Onboard New Clients in Search Atlas Search Atlas is built with agencies in mind. Here is how to structure your workflow when bringing on a new client: 1. Create a new project: Each client gets their own project inside Search Atlas. Navigate to your dashboard (Home) and select Create Project, then enter the client's domain. 2. Connect their WordPress site (if applicable): Install the Search Atlas plugin on the client's WordPress site and connect it to the project following the steps above. 3. Run a site audit: Trigger a crawl to identify technical SEO issues such as broken links, missing meta tags, slow pages, and crawl errors. Share the audit report with your client as a baseline. 4. Set up keyword tracking: Add the client's target keywords to the Rank Tracker to monitor positions over time and measure the impact of your work. 5. Use OTTO SEO for automation: Navigate to Left sidebar → OTTO SEO → All Sites (SEO Automation) (URL: /seo-automation-v3) to activate OTTO, Search Atlas's AI-powered SEO automation. OTTO can automatically fix on-page issues, deploy schema, and optimise content at scale — saving agencies significant time across multiple clients. 🔗 Search Atlas Backlink Capabilities for Agencies Search Atlas includes a comprehensive backlink intelligence suite that agencies can use to analyse, monitor, and build links for clients: - Backlink Analyzer: Enter any domain to see its full backlink profile, including referring domains, anchor text distribution, domain authority, and link types (dofollow vs nofollow). - Competitor backlink gap: Compare your client's backlink profile against competitors to identify link-building opportunities they are missing. - Link monitoring: Track newly acquired and lost backlinks over time so you can demonstrate link-building ROI to clients. - Prospecting data: Use the platform's database to find high-authority sites in your client's niche for outreach campaigns. These tools are available directly within each client project, making it easy to generate white-label-ready reports and insights for your agency deliverables. 💡 Agency Best Practices - Use separate projects for each client to keep data, crawls, and reports isolated and organised. - Schedule regular crawls so you catch technical regressions before they affect rankings. - Combine the backlink analyzer with keyword tracking to build a full-picture SEO strategy for each client. - Use OTTO SEO to handle repeatable optimisation tasks so your team can focus on strategy and client communication. If you need further assistance, open the chat widget in the bottom-right corner of the platform and type human teammate to be connected with a member of our team.

🔍 Fix Site Audit Not Crawling All WordPress Pages

🧩 Why This Happens Site Audit relies on being able to discover and access every page on your site. When pages are missing from the crawl, it usually means one or more common barriers are preventing the crawler from reaching them. This is especially common on WordPress sites due to plugin configurations, crawl limits, or technical settings that block or restrict access. ✅ Step 1: Check Your Crawl Settings Before investigating deeper, confirm your project is configured correctly inside Search Atlas. 1. Navigate to OTTO SEO → Site Audit from the left sidebar. 2. Select your site and open Settings or Edit Project. 3. Verify the Start URL is set to your root domain (e.g., https://yourdomain.com) and not a subfolder or inner page. 4. Check the crawl limit. If your site has hundreds or thousands of pages, make sure the limit is set high enough to cover all of them. 5. Ensure crawl subdomains is enabled if you have content hosted on subdomains. 🚫 Step 2: Review Your robots.txt File A misconfigured robots.txt file is one of the most common reasons pages are skipped during a crawl. If your robots.txt is blocking the Search Atlas crawler — or all crawlers — entire sections of your site will be invisible to Site Audit. 1. Open a browser and go to https://yourdomain.com/robots.txt. 2. Look for any Disallow rules that block User-agent: * or specific paths. 3. If important pages or directories are disallowed, update your robots.txt to allow access to them. 4. After updating, return to Site Audit and re-run the crawl. 🔌 Step 3: Disable or Configure WordPress Plugins Several WordPress plugins can interfere with crawling without you realising it. The most frequent offenders include: - Security plugins (e.g., Wordfence, iThemes Security) — these can block crawlers that resemble bots. Whitelist the Search Atlas crawler IP or temporarily disable bot blocking rules to test. - Caching plugins (e.g., WP Rocket, W3 Total Cache) — some caching rules can cause redirect loops or serve incomplete pages to crawlers. - Coming Soon / Maintenance mode plugins — if your site or any section is in maintenance mode, those pages will not be crawled. - Membership or login-required plugins — pages behind a login wall cannot be crawled unless you configure authenticated crawling. Test by temporarily deactivating suspected plugins, then re-running the Site Audit crawl to see if more pages are discovered. 🗺️ Step 4: Verify Your XML Sitemap Site Audit can use your XML sitemap as a reference to find pages. If your sitemap is missing pages, outdated, or broken, the crawl may not discover all content. 1. Go to https://yourdomain.com/sitemap.xml (or your custom sitemap URL). 2. Confirm all important pages are listed. 3. If you use a WordPress SEO plugin like Yoast SEO or Rank Math, check the sitemap settings to ensure all post types and taxonomies are included. 4. Regenerate the sitemap if needed, then re-submit it in Site Audit settings. 🔗 Step 5: Fix Internal Linking Gaps The Site Audit crawler discovers pages by following links. If some pages have no internal links pointing to them (orphan pages), the crawler cannot find them through link traversal alone. - Add internal links from other pages to any orphaned content. - Ensure your navigation menus, category pages, and footer links cover all key areas of your site. - Use your XML sitemap alongside crawling to maximise page discovery. 🔄 Step 6: Re-Run the Crawl After making any of the changes above, always trigger a fresh crawl so Site Audit can process the updated site. 1. Go to OTTO SEO → Site Audit → Overview (Website Overview) in the left sidebar. 2. Select your project and click Recrawl Site. 3. Wait for the crawl to complete, then review the Page Explorer report to confirm all expected pages are now listed. 4. Check the Issues report to view and action any errors that were previously missed. 💬 Still Need Help? If you have followed all the steps above and Site Audit is still not crawling all pages on your WordPress site, If you need further assistance, open the chat widget in the bottom-right corner of the platform and type human teammate to be connected with a member of our team.

🗺️ Site Audit XML Sitemap Detection Discrepancy Fix

🔍 Why Does This Discrepancy Happen? The Site Audit tool crawls your XML sitemap at a specific point in time. If your sitemap was updated after the last crawl — or if a caching, formatting, or submission issue exists — the tool may show a page as missing even though it appears in your live sitemap. This is a common and fully resolvable issue. Work through the steps below in order. Most customers resolve this in under five minutes. ✅ Step 1: Force a Re-Crawl of Your Site Audit The most common cause is simply a stale crawl. The tool captures a snapshot of your sitemap; any pages added after that snapshot will not appear until you re-crawl. 1. Navigate to OTTO SEO → Site Audit from the left sidebar. 2. Select the project that contains the affected site. 3. Click the Re-Crawl or Recrawl Site button, usually located in the top-right area of the audit dashboard. 4. Wait for the crawl to complete — crawl time varies depending on site size. 5. Once finished, return to the sitemap report and check whether the page now appears. If the page is still missing after a fresh crawl, continue to the next step. 🧹 Step 2: Clear Your Browser Cache and Reload A cached version of the Site Audit dashboard can display outdated results even after a new crawl has completed. 1. Press Ctrl + Shift + Delete (Windows) or Cmd + Shift + Delete (Mac) to open your browser's cache-clearing settings. 2. Select Cached images and files and Cookies and site data. 3. Set the time range to All time and confirm. 4. Close and reopen your browser, then log back into Search Atlas. 5. Navigate back to the Site Audit sitemap report and verify the result. 🔗 Step 3: Validate the URL Format in Your Sitemap The Site Audit tool performs exact URL matching. A mismatch in URL format between what the tool expects and what is declared in your sitemap will cause a page to be flagged as missing even if the content is live. Check for these common formatting issues in your XML sitemap file: - HTTP vs. HTTPS: Ensure all URLs use https://. Mixed protocols cause missed matches. - Trailing slash inconsistency:https://example.com/page/ and https://example.com/page are treated as different URLs. Pick one format and apply it consistently. - www vs. non-www: Confirm your sitemap uses the same domain variant (with or without www) as your canonical URLs. - URL encoding: Special characters in URLs must be properly encoded. Spaces, ampersands, or accented characters that are not encoded will cause detection failures. - Sitemap URL accessibility: Open the sitemap URL directly in your browser (e.g., https://yourdomain.com/sitemap.xml) and confirm the page in question is listed exactly as expected. Correct any formatting issues in your sitemap file, save the changes, and then force a re-crawl following Step 1 again. 📡 Step 4: Verify Your Sitemap Is Submitted in Google Search Console If your sitemap is not submitted to Google Search Console (GSC), or if a recently updated sitemap has not been re-submitted, the Site Audit tool may reference an outdated or incomplete version. 1. Log in to Google Search Console at search.google.com/search-console. 2. Select the correct property for your site. 3. In the left menu, go to Indexing → Sitemaps. 4. Confirm your sitemap URL is listed and shows a Success status. 5. If the sitemap is not listed, click Add a new sitemap, enter your sitemap URL, and submit. 6. If the sitemap is listed but shows errors or a stale last-read date, remove it and resubmit to prompt Google to re-fetch it. After verifying or resubmitting in GSC, return to Search Atlas and run a fresh Site Audit re-crawl (Step 1). ⚙️ Step 5: Check for Multiple or Nested Sitemaps Some sites use a sitemap index file that references multiple child sitemaps. If the page in question lives in a child sitemap that is not linked from the index, the Site Audit tool may not crawl it. - Open your sitemap index file (commonly at /sitemap_index.xml or /sitemap.xml) and confirm it includes a reference to the child sitemap containing the missing page. - If the child sitemap is missing from the index, add the reference and re-submit the sitemap index to GSC. - Then trigger a re-crawl in Site Audit. 📋 Quick Diagnostic Checklist - Re-crawl completed after the page was added to the sitemap - Browser cache cleared before checking results - URL uses HTTPS, consistent trailing slash, and correct www/non-www format - URL is properly encoded with no special character issues - Sitemap is submitted and shows Success status in Google Search Console - Page appears in the correct child sitemap if a sitemap index is in use Working through this checklist resolves the vast majority of sitemap detection discrepancies in Site Audit. If you complete all steps and the page is still showing as missing, the issue may require a deeper technical review. 💬 Still Seeing the Discrepancy? If you need further assistance, open the chat widget in the bottom-right corner of the platform and type human teammate to be connected with a member of our team.

🌐 Fix Google's 'Not Globally Crawlable' Flag

🔍 What Does 'Not Globally Crawlable' Mean? When Google flags your site as not globally crawlable, it means Googlebot is being blocked from accessing your pages from one or more geographic regions. This flag appears in Google Search Console and signals that your site may not be indexed consistently worldwide. This flag has two possible root causes that require very different solutions. Understanding which applies to your situation will save you significant troubleshooting time. ⚠️ Two Root Causes — One Flag Before taking any action, it is important to distinguish between the two categories below: - Category 1 — Search Atlas / OTTO actionable: Issues related to how your site is configured within the Search Atlas ecosystem, such as incorrect crawl settings or OTTO directives that unintentionally restrict access. - Category 2 — External configuration (requires your developer or hosting provider): Issues that exist at the server, CDN, firewall, or hosting level and are entirely outside the scope of Search Atlas or OTTO. These require intervention from your technical team or hosting provider. The vast majority of not globally crawlable flags reported through Google Search Console fall into Category 2. Search Atlas does not control your server infrastructure, firewall rules, or CDN access policies. ✅ What Search Atlas Can Help You Check If you believe the issue may be related to your Search Atlas or OTTO configuration, verify the following before escalating externally: 1. Confirm that OTTO has not applied any unintended robots.txt directives that block Googlebot. Review your active OTTO tasks and look for any robots.txt changes that may have been deployed. 2. Check that no noindex or nofollow tags have been pushed site-wide by mistake through OTTO optimisations. 3. Ensure your sitemap submitted in Google Search Console is current and reflects the correct URLs for your site. If none of the above apply, or if OTTO has not made changes to your robots.txt or meta directives, the root cause is almost certainly external to Search Atlas. 🛠️ External Causes That Require Developer or Hosting Intervention The following issues are outside the scope of Search Atlas and OTTO. Your developer or hosting provider must resolve them: - Geo-blocking or IP filtering: Your server, firewall, or security software is blocking requests from certain countries or IP ranges, including Google's crawl infrastructure. - CDN misconfiguration: Your content delivery network (for example, Cloudflare, Fastly, or AWS CloudFront) may be restricting access based on region or flagging Googlebot as suspicious traffic. - Hosting provider restrictions: Some hosting plans restrict traffic from outside specific regions by default. Your hosting provider's control panel or support team can confirm and adjust this. - DDoS protection or bot mitigation tools: Security services that block automated bots may inadvertently block Googlebot. Your developer needs to whitelist Google's verified crawlers in your security settings. - Server-side redirects or authentication walls: If your server returns different responses based on the visitor's geographic location, Google may be served a blocked or redirect response in certain regions. 📋 How to Diagnose the Issue Quickly Follow these steps to identify whether the problem is internal or external: 1. Open Google Search Console and navigate to Settings → Crawl Stats. Look for errors tied to specific response codes (403, 429, or connection timeouts) that indicate server-level blocking. 2. Use Google's URL Inspection Tool inside Search Console to test a sample URL. If Google reports it cannot reach the page, the block is at the server or network level. 3. Use a third-party tool such as GeoPerf or Uptrends to test your site's availability from multiple global locations. If the site fails to load from certain regions, confirm a server or CDN-level block is in place. 4. Ask your hosting provider or developer to review your firewall logs for blocked requests from Googlebot IP ranges, which are published in Google's official documentation. 🚀 Recommended Next Steps Based on Your Situation - If OTTO or Search Atlas changes are the suspected cause: Review your OTTO task history and revert any robots.txt or meta tag changes. Contact our support team via the chat widget for guided assistance. - If the issue is external (most common scenario): Share the diagnostic information from Google Search Console and your geo-availability test results with your developer or hosting provider. Ask them specifically to review firewall rules, CDN geo-restrictions, and bot mitigation settings. - If you are unsure which category applies: Start with the OTTO review steps above, then run the URL Inspection Tool in Google Search Console. If Google still cannot reach the page and OTTO has made no relevant changes, escalate externally. If you need further assistance, open the chat widget in the bottom-right corner of the platform and type human teammate to be connected with a member of our team.

🔍 Fix Auto Scan Failures and Run Manual Crawls

🧭 Overview If your auto scan did not run on its scheduled date, there are several possible causes. This article walks you through the most common reasons a scan is skipped and shows you exactly how to run a manual crawl so you never lose an audit cycle. ⚠️ Why Your Auto Scan May Have Been Skipped A missed auto scan is rarely caused by a single issue. Work through each of the following checks before concluding the root cause. ✅ Check 1: Scheduling Settings The most common reason a scan does not run is that the schedule was never saved correctly or was changed accidentally. 1. Open the project that missed its scan. 2. Navigate to the scheduling or crawl settings for that project. 3. Confirm the scan frequency (daily, weekly, monthly) is set to the option you expect. 4. Check that the scheduled date and time match your intended configuration. 5. If anything looks incorrect, update the settings and save your changes. ✅ Check 2: Account Permissions and Crawl Credits Auto scans may require the correct role permissions and sufficient crawl credits on your plan. - Permissions: If your account role was recently changed, your scheduled scan may have been affected. Check with your account owner if you are unsure about your current permissions. - Crawl credits: If your account reached its crawl credit limit, the scan may have been skipped. Check your account's usage or billing information to verify your remaining credits. ✅ Check 3: Rate Limiting on Your Website Search Atlas respects server-side rate limits. If your website's server returned too many errors or throttled the crawler, the scan may have stopped or been skipped entirely. - Check your website's server logs around the expected scan time for any error spikes or throttling activity. - If rate limiting is active, consider adjusting your server's crawl rate settings or reducing the crawl depth in your project's settings. ✅ Check 4: Bot-Blocking or Firewall Rules (Including Cloudflare) Security tools such as Cloudflare, Sucuri, or a custom Web Application Firewall (WAF) can block the Search Atlas crawler, causing a scan to fail silently. - Log in to your firewall or CDN provider and review the firewall activity log for blocked requests around the scheduled scan time. - Look for blocked requests from the Search Atlas crawler user agent. - If you find blocks, add the Search Atlas crawler to your allowlist, or contact the person who manages your website's security settings. - Note: If you do not have access to these settings yourself, you will need to ask your website administrator or developer to make this change. 🔄 How to Run a Manual Crawl If you need to run a crawl immediately rather than waiting for the next scheduled scan, you can trigger a manual crawl from within your project's crawl or audit settings. Look for an option to start or run a new crawl and confirm the action to begin. 📞 Still Need Help? If you have worked through all of the checks above and your auto scan is still not running, or if you are unsure which step applies to your situation, please reach out to our support team. When you do, have the following ready: your project name, the date and time the scan was expected to run, and any error messages you have seen. If you need further assistance, open the chat widget in the bottom-right corner of the platform and type human teammate to be connected with a member of our team.

🔍 Why Site Audit Isn't Crawling All Pages

🧭 Overview Search Atlas Site Audit crawls your website by following the URLs listed in your XML sitemap. If your sitemap does not include all of your pages, Search Atlas will only crawl the pages it can find there — even if those pages exist on your live site and are accessible to other tools. This is the most common reason customers see a lower crawled page count than expected. 📋 How Search Atlas Discovers Pages Unlike some crawlers that spider every link on a page, Search Atlas primarily relies on your submitted XML sitemap to discover and audit pages. This means: - Pages listed in your sitemap will be crawled and audited. - Pages not listed in your sitemap will be skipped, even if they are live and indexable. - A third-party tool that crawls by following internal links may report a higher page count simply because it uses a different discovery method. If you have 500+ pages on your site but only 114 appear in your sitemap, Search Atlas will audit approximately 114 pages — not 500+. ✅ Step 1 — Check How Many URLs Are in Your Sitemap Before investigating anything else, verify your sitemap is complete: 1. Open your sitemap directly in a browser. The most common location is yourdomain.com/sitemap.xml. 2. Count the total number of <url> entries, or use a free sitemap checker tool to get an accurate count. 3. Compare that number to the total number of pages you expect to be crawled. 4. If the sitemap count is significantly lower than your total page count, your sitemap is incomplete — this is the root cause. 🛠️ Step 2 — Fix an Incomplete Sitemap If your sitemap is missing pages, you need to regenerate or update it so all intended pages are included. How you do this depends on your platform: - WordPress: Use an SEO plugin such as Yoast SEO, Rank Math, or All in One SEO. Navigate to the plugin's sitemap settings and ensure all post types, pages, and taxonomies you want crawled are enabled. Save and regenerate the sitemap. - Shopify: Shopify generates a sitemap automatically. If products or pages are missing, confirm they are published and not set to hidden. - Custom or headless sites: Work with your developer to ensure your sitemap generation script includes all relevant URLs and is not filtering out pages unintentionally. - Large sites with sitemap indexes: If your site has more than 50,000 URLs, you may need a sitemap index file that references multiple individual sitemaps. Confirm all child sitemaps are linked from the index. After updating your sitemap, re-submit it in Google Search Console under Sitemaps to keep your data consistent across tools. 🔄 Step 3 — Re-Submit Your Sitemap in Search Atlas Once your sitemap is updated and includes all the pages you want audited, re-run your Site Audit in Search Atlas: 1. Navigate to your project in Search Atlas. 2. Open Site Audit and locate the sitemap or crawl settings for your project. 3. Confirm the correct sitemap URL is entered. If you have updated the sitemap URL or added a sitemap index, update this field accordingly. 4. Start a new crawl and wait for it to complete. 5. Verify the crawled page count now matches the number of URLs in your updated sitemap. ⚠️ Other Reasons Pages May Be Missing If your sitemap is complete but pages are still not appearing in your crawl results, check the following secondary causes: - Noindex tags: Pages with a <meta name="robots" content="noindex"> tag or an X-Robots-Tag: noindex HTTP header may be excluded from audit results intentionally. - robots.txt blocks: Open yourdomain.com/robots.txt and confirm that the Search Atlas crawler or the Disallow rules are not blocking sections of your site. - Canonical tags pointing elsewhere: Pages with a canonical tag pointing to a different URL may be treated as duplicates and consolidated in your results. - Crawl limits on your plan: Some Search Atlas plans include a maximum number of pages per crawl. If your site exceeds this limit, upgrade your plan or prioritise the most important sections of your sitemap. - Sitemap URL not reachable: Confirm your sitemap returns a 200 HTTP status and is not behind a login, password protection, or firewall that blocks crawlers. 🔑 Quick Checklist - Sitemap URL is accessible at a public address (no login required). - Sitemap contains all pages you want audited — count the <url> entries. - No unintended noindex tags or robots.txt blocks are in place. - Correct sitemap URL is entered in your Search Atlas project settings. - A fresh crawl has been started after any sitemap changes. 💬 Need More Help? If you need further assistance, open the chat widget in the bottom-right corner of the platform and type human teammate to be connected with a member of our team.

🔄 Cancel and Restart Stuck Site Crawls

🧭 Overview Site crawls in Search Atlas usually complete within a few hours, depending on the size of your website. If a crawl has been running for more than 24 hours without finishing, it is likely stuck. This article explains how to cancel a stuck crawl and restart it so your site audit can complete successfully. ⚠️ Signs Your Crawl May Be Stuck - The crawl progress has not moved for several hours. - The crawl status still shows as in progress after more than 24 hours. - No new data or pages have appeared in the audit report. - The crawl has been running for multiple days with no completion. 🛑 How to Cancel a Stuck Crawl 1. In the left sidebar, navigate to your project and open Site Auditor. 2. Locate the crawl session that has been running without completing. 3. Find the option to cancel or stop the active crawl session — this is typically available alongside the crawl session entry in the auditor interface. 4. Confirm the cancellation and wait a few seconds for the status to update before proceeding. Note: Cancelling a crawl does not delete any previously collected data. Any pages already crawled before cancellation remain available in your audit history. 🔁 How to Restart the Crawl 1. Once the previous crawl is cancelled, look for the option to initiate a new crawl session within the Site Auditor for your project. 2. Review your crawl settings before confirming — check the crawl scope, crawl limit, and any URL exclusion rules. 3. Start the new crawl session and monitor its progress from within the Site Auditor interface. If you made no changes to your crawl settings, the recrawl will use the same configuration as your previous session. 🛠️ Things to Check Before Restarting Before you restart the crawl, consider the following adjustments to help prevent it from getting stuck again: - Reduce the crawl limit if your site has a very large number of pages — crawling a smaller subset can help complete the audit faster. - Review URL exclusion rules to make sure the crawler is not being caught in redirect loops or crawling unintended sections of your site. - Check your crawl scope to ensure it is set appropriately for the size of the site you are auditing. 📋 What to Have Ready If You Need Support If your crawl continues to get stuck after cancelling and restarting, our team can investigate on the backend. Please have the following ready when you reach out: - Your project name and the URL of the site being crawled - The approximate date and time the crawl started - A description of what you see in the Site Auditor (e.g. whether the status appears frozen or unchanged) - Any error messages displayed in the interface If you need further assistance, open the chat widget in the bottom-right corner of the platform and type human teammate to be connected with a member of our team.

🔍 Fix a Stuck or Slow Site Audit Crawl

🧭 Overview A Site Audit crawl typically completes within a few minutes to a couple of hours, depending on your site's size and crawl settings. In some cases, crawls can take longer or appear stuck entirely. This article explains the most common causes and how to fix them. ⏱️ How Long Should a Crawl Take? Crawl duration varies based on several factors: - Small sites (under 500 pages): Usually 5–30 minutes - Medium sites (500–5,000 pages): Typically 1–4 hours - Large sites (5,000+ pages): Can take several hours If your crawl has been running for more than 24 hours without completing, something is likely preventing it from finishing normally. ⚠️ Common Causes of a Stuck or Slow Crawl - JavaScript rendering is enabled: Crawling JavaScript-heavy pages is significantly more resource-intensive. When JS rendering is on and the crawl budget is set to a very low number (1–2), crawls can take 17–31 hours or more. - Low crawl budget setting: A crawl budget of 1–2 pages per second drastically slows progress on larger sites. - Site errors or redirect loops: Pages that loop, return errors, or are unreachable can cause the crawler to stall on a single URL. - Server-side blocking: Some servers detect and block automated crawlers, causing requests to hang indefinitely. - Very large or complex sitemaps: Sitemaps with thousands of URLs or nested structures can slow the initial crawl setup. 🛠️ Step-by-Step Troubleshooting 1. Check your JavaScript rendering setting. If your site does not rely heavily on client-side rendering, disable JS rendering in your audit settings. This is the single most common cause of extremely long crawls. 2. Increase your crawl budget. Navigate to your Site Audit settings and raise the crawl budget above 1–2. A setting of 5–10 is a good starting point for most sites. Higher budgets complete crawls faster. 3. Cancel and restart the crawl. If the crawl has been running for more than 24 hours, cancel it and launch a fresh crawl with the adjusted settings above. Stalled crawls do not self-recover. 4. Verify your site is reachable. Open your website in a browser and confirm it loads without errors, redirects, or login walls. The crawler must be able to access your pages freely. 5. Check for crawl errors in the audit report. After restarting, monitor the audit dashboard for error messages. A message indicating a crawl is failing or looping on a single URL points to a specific page or redirect issue—review that URL directly. 6. Review your robots.txt file. Make sure your robots.txt is not inadvertently blocking the Search Atlas crawler from accessing key sections of your site. 7. Reduce scope for large sites. If your site has tens of thousands of pages, consider limiting the crawl to a specific subdirectory or URL prefix to isolate and test performance before running a full crawl. ✅ Best Practice Settings for Faster Crawls - Disable JavaScript rendering unless your site requires it for content to load. - Set your crawl budget to 5 or higher for most sites. - Keep your sitemap clean and free of redirect chains or broken URLs before initiating a crawl. - Schedule crawls during off-peak hours to reduce the chance of server-side throttling. 💬 Still Need Help? If you have tried the steps above and your crawl is still stuck or failing, our team can investigate directly. If you need further assistance, open the chat widget in the bottom-right corner of the platform and type human teammate to be connected with a member of our team.

🤖 Fix Shopify robots.txt Crawl-Delay Directive

🔍 Why This Error Appears in Search Atlas When Search Atlas crawls your site, it flags a crawl-delay directive in your robots.txt file as a critical error. This directive instructs search engine bots to pause between requests, which can significantly slow down how quickly Google and other search engines index your pages. If your store runs on Shopify, you have likely already discovered that you cannot simply open and edit your robots.txt file the way you would on a self-hosted site. This is expected behavior — it is a platform-level restriction, not a problem with your account or Search Atlas. 🛍️ Why Shopify Locks Your robots.txt File Shopify automatically generates a robots.txt file for every store. On most plans, this file is managed by Shopify's infrastructure and is not directly editable through the Shopify admin panel or theme editor by default. The crawl-delay directive you see is typically injected by Shopify itself or by a third-party app installed on your store. Because of this, standard advice such as "open your robots.txt file and delete the line" does not apply to Shopify stores. You need to use one of the platform-specific methods described below. ✅ Method 1: Check for Apps Injecting the Directive The most common cause of a crawl-delay directive on Shopify is a third-party SEO or security app that modifies your robots.txt output. Follow these steps to identify and remove the source: 1. Log in to your Shopify Admin panel. 2. Go to Apps in the left sidebar. 3. Review any installed SEO, site speed, or security apps such as SEO Manager, Plug in SEO, or bot-protection tools. 4. Open each relevant app's settings and look for a robots.txt or crawl settings section. 5. If you find a crawl-delay setting, disable it or set the value to 0 and save. 6. Wait 24–48 hours, then re-run a crawl in Search Atlas to confirm the error is resolved. If you are unsure which app is responsible, temporarily disable SEO-related apps one at a time and check your live robots.txt file at yourdomain.com/robots.txt after each change to identify the source. ✅ Method 2: Use a Custom robots.txt Template (Shopify 2.0 Themes) If your store uses a Shopify Online Store 2.0 theme, Shopify allows you to override the default robots.txt file using a Liquid template. This method gives you full control over the file's content. 1. In your Shopify Admin, go to Online Store → Themes. 2. Next to your active theme, click Actions → Edit code. 3. In the left file panel, look for a folder called Templates. 4. If a file named robots.txt.liquid already exists, click it to open it. If it does not exist, click Add a new template, select robots.txt from the dropdown, and click Create template. 5. Review the template content and remove any line that contains crawl-delay. 6. Click Save. 7. Visit yourdomain.com/robots.txt to confirm the directive is no longer present. Important: When creating a custom robots.txt.liquid template, Shopify replaces its automatically generated file entirely with your template. Make sure the template retains all necessary Disallow and Sitemap directives from the original. You can copy the current content of your live robots.txt as a starting point before making edits. ✅ Method 3: Contact Shopify Support If your theme is not a 2.0 theme, the robots.txt.liquid template option is unavailable, and no installed app appears to be the source of the directive, the crawl-delay may be hard-coded by Shopify at the infrastructure level. In this case, direct editing is not possible without Shopify's involvement. Contact Shopify Support directly and provide the following information: - Your store URL and the current content of your robots.txt file. - The specific crawl-delay line you need removed. - A note that the directive is causing indexation issues flagged by your SEO tooling. Shopify support can advise whether an upgrade to a 2.0-compatible theme or a platform-level change is required. 📊 Verifying the Fix in Search Atlas After applying any of the methods above, confirm the error has been resolved in the Search Atlas dashboard: 1. In the left sidebar, navigate to OTTO SEO V3. 2. Open your active project and locate the site audit or technical SEO section. 3. Trigger a new crawl of your site. 4. Once the crawl completes, confirm the crawl-delay critical error no longer appears in your issue list. If the error persists after the crawl, revisit your robots.txt file directly at yourdomain.com/robots.txt to verify the directive has been fully removed before re-crawling. 💬 Need More Help? If you need further assistance, open the chat widget in the bottom-right corner of the platform and type human teammate to be connected with a member of our team.

🔍 How OTTO Crawls Your Site for Errors

🤖 What Is an OTTO Site Scan? When you activate OTTO on a website, OTTO automatically crawls your site to discover technical SEO issues — such as broken links, missing meta tags, slow pages, and crawlability errors. This scan gives OTTO the data it needs to begin suggesting and applying fixes. 🚀 How to Start a Site Scan 1. Navigate to the OTTO section of the platform. 2. If you have not connected a site yet, follow the on-screen prompts to add your website and verify ownership. 3. Once your site is connected, OTTO will automatically initiate the first crawl. No manual trigger is required for the initial scan. 4. To run a fresh crawl at any time, locate your site on the OTTO dashboard and use the option available to re-scan your property. OTTO begins processing your site after the crawl is triggered. You can return to the OTTO dashboard to check on the status of your scan. ⏱️ How Long Does a Crawl Take? Crawl duration depends on the size and complexity of your website. Larger sites with many pages will naturally take longer to fully process than smaller ones. You do not need to keep the browser window open while the crawl runs — OTTO processes the scan in the background, and your results will be ready when you return to the dashboard. If your crawl appears to be taking an unexpectedly long time, it may be related to your site's size, server response times, or crawl configuration. See the troubleshooting section below for guidance. 📊 What Happens After the Crawl? Once the scan is complete, OTTO will display the detected on-page and technical SEO issues it has found. You can review all findings directly inside the OTTO section of the platform. New scans can be triggered at any time to reflect recent changes you have made to the site. 💡 Tips for Faster, More Accurate Scans - Submit a sitemap: Providing an XML sitemap helps OTTO discover all your pages more efficiently. - Check robots.txt: Make sure your robots.txt file does not accidentally block OTTO from crawling important sections of your site. - Avoid scanning during peak traffic: On very large sites, scheduling re-crawls during low-traffic periods can improve speed and reduce server load. 🛠️ Troubleshooting Common Crawl Issues - Scan appears stuck: Wait some time for larger sites before assuming an issue. If there has been no progress after an extended wait, try refreshing the page or initiating a new scan. - Pages not appearing in results: Check that your robots.txt is not blocking the crawler and that you have submitted an up-to-date sitemap. - Crawl did not start: Confirm that your site is fully connected and ownership has been verified in the platform. If you need further assistance, open the chat widget in the bottom-right corner of the platform and type human teammate to be connected with a member of our team.

🔍 Fix Stale Cached Content Still Showing in Crawls

🧩 Why This Happens When Search Atlas crawls your website, it stores a snapshot of your page content in its cache. If you remove text, deactivate a plugin, or make other changes to your site, Search Atlas may still display the old cached version of that page until a fresh crawl is completed and the cache is cleared. This is a common reason customers see removed content — such as phrases like "24/7" — still appearing in Search Atlas reports even after the content has been deleted from the live page. ⚠️ Common Triggers - Plugin deactivation: Deactivating OTTO SEO or another plugin does not immediately force a recrawl. The previous crawl data remains until refreshed. - Page edits without recrawl: Saving changes in your CMS updates your live site but does not automatically notify Search Atlas to recrawl that URL. - Conflicting SEO plugins: Tools like Rank Math can override meta deployments. Search Atlas may show "Deployed" while the live page still renders the original meta from a competing plugin. - Server-side or CDN caching: Your hosting provider or CDN may be serving a cached version of your page to Search Atlas crawlers, making it appear as if old content is still live. ✅ How to Fix Stale Crawl Data 1. Confirm the live page is updated: Open the URL in an incognito/private browser window to verify the content you removed is no longer visible on the actual page. 2. Clear your site cache: If you use a caching plugin (e.g., WP Rocket, W3 Total Cache) or a CDN (e.g., Cloudflare), purge the cache for the affected page so Search Atlas crawlers receive the latest version. 3. Trigger a manual recrawl in Search Atlas: Navigate to the relevant project in Search Atlas, locate the affected URL, and use the Recrawl or Refresh option to queue a fresh crawl of that page. 4. Wait for the crawl to complete: Depending on your site size and crawl queue, this may take a few minutes. Once finished, the updated content will replace the stale snapshot. 5. Verify the result: After the recrawl, reopen the page report in Search Atlas to confirm the outdated content is no longer displayed. 🛠️ If You Use OTTO SEO OTTO SEO includes a Cache Self-Verification system that monitors the health of deployed changes and checks whether your server is accessible to Search Atlas crawlers. If connectivity issues prevent OTTO from verifying a deployment, it will flag this in the platform so you can take action. To check the status of your OTTO deployments: - Go to OTTO SEO → All Sites (SEO Automation) in your Search Atlas dashboard. - Review the deployment status for affected pages. A status of "Deployed" confirms Search Atlas successfully applied the change — but if a conflicting plugin is active, it may override the output on the live page. - If you suspect a conflict, deactivate competing SEO plugins (such as Rank Math or Yoast) and trigger a recrawl to let OTTO's settings take full effect. 🚫 Avoid These Mistakes - Do not assume deactivating a plugin instantly updates Search Atlas data — a manual recrawl is always needed. - Do not run multiple SEO plugins simultaneously. Conflicting plugins can silently override each other's meta settings, causing discrepancies between what Search Atlas reports and what actually renders on your page. - Do not skip clearing your server or CDN cache before triggering a recrawl, as crawlers may still receive the old cached page. 💬 Still Seeing Outdated Content? If you have completed all the steps above and Search Atlas is still displaying stale content, there may be a deeper crawl or connectivity issue that requires investigation. If you need further assistance, open the chat widget in the bottom-right corner of the platform and type human teammate to be connected with a member of our team.

🕷️ Fix OTTO Crawler Failures on Homepage Redirects

🔍 Overview When you create an OTTO SEO project, OTTO crawls your homepage to verify the pixel installation and index your site. If your domain automatically redirects to a subdirectory — for example, example.com redirecting to example.com/news — OTTO may follow that redirect and land on a page that breaks the crawl. This can cause project creation to stall, return zero pages crawled, or prevent pixel detection from completing successfully. ⚠️ Why This Happens OTTO's crawler starts at your root domain and follows any redirects it encounters. Redirects to subdirectories can cause two specific problems: - Pixel not detected: If the subdirectory page does not carry the OTTO pixel in its HTML, the crawler cannot confirm a successful installation, and project creation fails. - Crawl returns zero pages: In some configurations, the redirect strips internal parameters OTTO needs to complete the crawl, resulting in no pages being indexed even though the site is reachable. ✅ Step-by-Step: Resolve the Redirect Issue 1. Confirm where your domain redirects. Open a browser and type your bare domain (e.g., example.com). Note the final URL in the address bar. If it ends in a subdirectory such as /news or /en, a redirect is in place. 2. Place the OTTO pixel on the redirect destination. Make sure the OTTO tracking pixel is installed on the page your domain actually lands on after the redirect — not just on the root-level template. If you are using a CMS, add the pixel to the header template of that specific section. 3. Use the exact landing URL when creating your project. In Search Atlas, go to Left sidebar → OTTO SEO → All Sites (SEO Automation) (URL: /seo-automation-v3). When prompted to enter your domain during project creation, enter the full subdirectory URL that is the true homepage — for example, example.com/news — rather than the bare domain. This gives OTTO a stable starting point that does not redirect. 4. Verify the pixel using your browser. Before retrying project creation, open the landing page, right-click, and select View Page Source. Search for your OTTO pixel code to confirm it is present. If it is missing, reinstall it and save your changes before proceeding. 5. Create or re-create the OTTO project. Return to Left sidebar → OTTO SEO → All Sites (SEO Automation) and select Create. Enter the subdirectory URL from step 3 and complete the setup wizard. OTTO will now crawl from that stable endpoint instead of chasing a redirect. 6. Monitor the crawl. After creation, check that the project status moves past the initial setup phase and that pages are being indexed. If the project shows 0 pages crawled after several minutes, proceed to the troubleshooting section below. 🛠️ Additional Troubleshooting If the steps above do not resolve the issue, consider the following: - Check for redirect chains. A single domain can have multiple hops — for example, example.com → www.example.com → example.com/news. Each hop increases the chance of parameter loss. Use a redirect-checker tool to map every hop and ensure the final destination carries the pixel. - Confirm there is no redirect loop. Misconfigured redirects can cause a page to redirect back to itself. If your site went offline or slowed down after adding OTTO, this may be the cause. Contact your hosting provider or developer to audit your redirect rules. - Check that the subdirectory is publicly reachable. Some subdirectories are gated behind a login or return a non-200 HTTP status for bots. OTTO requires a publicly accessible 200 OK response to crawl successfully. - Allow up to 15 minutes after pixel installation before triggering project creation. Caching layers on your server or CDN may serve an older version of the page without the pixel for a short period after you save changes. ❓ Frequently Asked Questions Can OTTO crawl a subdirectory like /news instead of the root domain? Yes. When creating your OTTO project, enter the subdirectory URL as your project domain. OTTO will treat that URL as the starting point for the crawl and will not require the bare root domain. Why did my project domain change automatically after creation? If you entered a root domain that immediately redirects, OTTO may have auto-resolved it to the redirect destination. This is expected behavior. Verify the project now points to the correct subdirectory and that the pixel is installed there. Will using a subdirectory as my project URL affect SEO recommendations? OTTO will scope its recommendations to the pages it can reach from your starting URL. If your entire site lives under /news, entering that subdirectory will give OTTO full access. If your site spans multiple subdirectories, contact support to discuss the best configuration for your setup. 💬 Need More Help? If you need further assistance, open the chat widget in the bottom-right corner of the platform and type human teammate to be connected with a member of our team.

🔍 Fix Site Audit Crawler Failures and Config Errors

🔍 Overview If your Site Audit is returning 0 crawled pages or you are seeing unexpected crawl configuration behavior, this article explains what to prepare and what to expect when our team investigates. Site Audit crawler failures and account configuration errors require backend review by a support agent. ⚠️ Common Symptoms - Site Audit completes but shows 0 pages crawled - Crawl settings cannot be saved or updated - Unexpected behavior in crawl configuration options - Audit results appear inconsistent with your site content 📋 What to Know Before Escalating Site Audit crawler failures and account configuration errors are typically caused by backend issues that cannot be resolved through the interface alone. Our support team will need to review your project and account configuration directly. To help us resolve your issue as quickly as possible, please have the following ready when you contact support: - Your project name as it appears in Search Atlas - The exact error message displayed in the Site Audit dashboard, if any - A timestamp of when the crawl failure occurred or when you first noticed the configuration issue - A description of any settings you changed before the issue appeared - Your current plan so our team can verify which features are available to your account 🛠️ Preliminary Checks You Can Do 1. Confirm your website is publicly reachable. The crawler cannot access password-protected, staging, or locally hosted sites. Verify your domain loads correctly in a browser without a login prompt. 2. Check that your domain is entered correctly in the project. A typo or missing protocol can prevent the crawler from reaching your site. 3. Note any error messages in the Site Audit dashboard. Record the exact wording of any errors shown so you can share them with the support team. 4. Do not repeatedly re-run the crawl while the issue is unresolved, as this may compound configuration conflicts that require backend correction. ⚙️ What Happens Next Once you contact support with the details above, a member of our team will review your account and project configuration on the backend and work with you to restore normal crawl behavior. We will confirm the outcome and any changes made to your account directly in the support chat. If you need further assistance, open the chat widget in the bottom-right corner of the platform and type human teammate to be connected with a member of our team.

🔄 Chrome User Experience Report Not Updating After Recrawl

🧭 Overview The Chrome User Experience (CrUX) report inside the Search Atlas Site Audit displays real-world performance data collected by Google from actual Chrome users visiting your site. Because this data originates from an external Google dataset rather than the crawl engine itself, it follows a different refresh cycle than other audit metrics. Understanding that distinction is the key to resolving most data-freshness issues. 🔍 Why the CrUX Report Does Not Update Immediately After a Recrawl Running a new site audit recrawl refreshes data that Search Atlas collects directly — such as broken links, indexability signals, and on-page elements. The Chrome User Experience report is different: - External data source: CrUX data is sourced from Google's public CrUX dataset, which Google updates on a rolling 28-day window, published once per month (typically in the first week of each month). Search Atlas cannot force Google to release new data ahead of that schedule. - Recrawl scope: A site audit recrawl instructs the Search Atlas crawler to re-examine your site's pages. It does not trigger a new pull from the Google CrUX dataset unless the dataset itself has been refreshed since your last audit. - Cached dataset: If the Google CrUX dataset has not been updated since your previous audit, the report will display the same figures regardless of how many recrawls you run. - Insufficient real-world traffic: Google only includes URLs in the CrUX dataset that receive a minimum threshold of Chrome user visits. Pages with low traffic may show no CrUX data at all, or data may appear intermittently. ⚠️ Common Scenarios and What They Mean - Recrawl completed but CrUX scores are identical: The Google dataset has not released a new monthly update yet. Wait until the dataset refreshes and then run a new audit. - Some pages show CrUX data and others do not: Pages without data likely fall below Google's minimum traffic threshold. This is a Google-side limitation and cannot be changed from within Search Atlas. - Scores changed in Google Search Console but not in Search Atlas: Google Search Console can display more granular or provisional data. The CrUX dataset used by Search Atlas reflects the full monthly aggregation, which may lag by a few days after Google publishes it. - CrUX data shows for the domain but not individual URLs: Google aggregates CrUX at both the origin (domain) level and the URL level. If a specific URL lacks sufficient traffic data, only origin-level data will appear. ✅ Step-by-Step Troubleshooting 1. Confirm the Google CrUX dataset has been updated. Check the Google CrUX release schedule — updates are typically published in the first week of each calendar month. If you ran your recrawl before the latest release, no new CrUX data will be available yet. 2. Run a fresh recrawl after the dataset releases. Navigate to Left sidebar → OTTO SEO → Site Audit → All Audits, open your project audit, and trigger a new recrawl. The report will then pull the most recently available CrUX data. 3. Check the audit timestamp. Inside your audit, confirm the date the CrUX data was last fetched. If it predates the most recent Google CrUX release, a new recrawl will resolve the discrepancy. 4. Verify your site has sufficient Chrome traffic. If pages consistently show no CrUX data, cross-reference your Google Search Console → Core Web Vitals report. If data is absent there too, the pages likely do not meet Google's minimum traffic threshold. 5. Review the Page Explorer for individual URL data. Navigate to Left sidebar → OTTO SEO → Site Audit → Page Explorer to inspect CrUX availability at the individual page level. This helps identify which URLs are included in the dataset and which are not. 6. Check Crawl Monitoring for any crawl errors. Navigate to Left sidebar → OTTO SEO → Crawl Monitoring to confirm the recrawl completed without errors. An interrupted crawl may result in incomplete data across all report sections, including CrUX. 📋 What to Expect Going Forward - CrUX data inside Search Atlas refreshes once per month, aligned with Google's public dataset release cycle. - After Google releases a new dataset, run a fresh recrawl to pull the updated figures into your audit report. - Improvements you make to Core Web Vitals will take at least one full 28-day collection window before they appear in CrUX scores. - Origin-level CrUX data is available for most sites with any meaningful traffic, but URL-level data requires higher individual page traffic volumes. 💬 Still Need Help? If you need further assistance, open the chat widget in the bottom-right corner of the platform and type human teammate to be connected with a member of our team.

🔄 Fix a Site Audit Stuck in Recrawling

🔍 What Is the Recrawling Status? When you trigger a manual crawl in Site Audit, the tool sets the crawl status to Recrawling while it processes your pages. In most cases this completes within a reasonable amount of time, depending on your site size and page limit. Occasionally, a crawl can get stuck in Recrawling status for an extended period without progressing to completion. During this time, some controls may become unavailable. This is a known platform issue that our engineering team is actively working to resolve. ⚠️ Why Does This Happen? A crawl can become stuck due to internal processing conditions that prevent it from completing or resetting automatically. Importantly, a stuck crawl does not overwrite or delete your existing audit data. Your previously collected data remains intact while the crawl is stalled. ✅ What You Can Do Right Now 1. Wait and allow time to pass — Short stalls occasionally resolve on their own once the background process catches up. If the crawl has not been stuck for very long, allow it more time before escalating. 2. Refresh the page — After waiting, do a hard refresh (Ctrl + Shift + R on Windows, Cmd + Shift + R on Mac) to reload the latest crawl status from the server. 3. Check if results have populated — Even if the status still shows Recrawling, scroll through your Site Audit results. Data often finishes processing before the status label updates correctly. 4. Do not delete the audit — Deleting the audit is irreversible and will permanently remove all previously collected data. Avoid this step unless instructed by the support team. 🛠️ When to Contact Support If your crawl has been stuck in Recrawling status for an extended period without any progress, our support team can intervene on the backend to cancel or force-complete the crawl without deleting your audit data. Before reaching out, gather the following information so the team can act quickly: - The name or URL of the affected site audit project. - How long the crawl has been stuck (approximate start time and current page count). - A screenshot showing the Recrawling status and any unavailable controls, if possible. If you need further assistance, open the chat widget in the bottom-right corner of the platform and type human teammate to be connected with a member of our team.

🛡️ Fix SEO Health Score Volatility and Exclude Pages from OTTO Crawls

🔍 Why Your SEO Health Score Changes After Fixing Issues Your SEO Health Score is calculated by a continuous audit, not a one-time snapshot. Every time OTTO crawls your site, it re-evaluates all pages against the latest SEO rules. This means the score can shift up or down even after you have fixed issues — and that is expected behavior, not a bug. Common reasons your score may show more issues after a fix: - New pages were added — freshly published or indexed pages are crawled and audited automatically. - Template-driven issues surfaced — if your site uses templates (e.g., WooCommerce product pages, archive pages), a single template problem can multiply across hundreds of URLs in one crawl cycle. - Rule updates — Search Atlas periodically refines its SEO audit rules, which can surface issues that were previously uncategorized. - Previously skipped URLs — OTTO discovers new URLs over time as your sitemap or internal links grow. The score reflects the real-time health of everything OTTO can see. If the number of detected issues rises, focus on which issue types are flagged rather than reacting to the score alone. Fixing one high-impact template issue can resolve hundreds of individual instances at once. 🚫 How to Exclude Pages from OTTO Crawls Excluding low-value or template pages from OTTO crawls prevents unnecessary credit consumption and keeps your audit results focused on pages that matter. There are three supported methods. ⚙️ Method 1 — Exclude Pages in the OTTO UI Settings This is the fastest method and does not require any changes to your website. 1. In the left sidebar, click OTTO SEO → SEO Automation (All Sites) (URL: /seo-automation-v3). 2. Select your project (e.g., designoneprinting.com) from the project list. 3. Open the Settings tab inside the project dashboard. 4. Locate the Crawl Exclusions or Excluded URLs section. 5. Enter the URL patterns you want to exclude. You can use exact URLs or wildcard patterns (e.g., /shop/page/* to exclude all paginated shop pages). 6. Click Save. Exclusions take effect on the next crawl cycle. Best for: Quickly excluding specific URL patterns without touching your website files. 🤖 Method 2 — Use the Robots Meta Tag on Individual Pages Add a noindex or nofollow robots meta tag directly to the HTML of pages you want OTTO to skip. OTTO respects standard robots directives. 1. Open the page editor for the page you want to exclude (in WordPress, Shopify, or your CMS of choice). 2. Add the following tag inside the <head> section of the page: <meta name="robots" content="noindex, nofollow"> 3. Publish or save the page. 4. OTTO will honor this directive during its next crawl and will not audit that page. Best for: Excluding individual pages or a small set of specific pages where you want the directive to live at the page level. 📄 Method 3 — Block URL Patterns via robots.txt For large groups of template-driven URLs (such as all tag pages, all paginated URLs, or all print-preview pages), updating your robots.txt file is the most efficient approach. 1. Access your robots.txt file via your hosting control panel, FTP, or CMS (in WordPress, plugins like Yoast SEO or Rank Math include a robots.txt editor). 2. Add a Disallow rule under the relevant user-agent. To block all crawlers from a pattern, use: User-agent: * followed by Disallow: /tag/ (replace /tag/ with your target path pattern). 3. Save and publish the updated file. Verify it is live by visiting yourdomain.com/robots.txt. 4. OTTO will respect the Disallow directives on its next crawl. Best for: Blocking large groups of template URLs (archive pages, paginated results, print versions) site-wide in a single edit. Tip for designoneprinting.com users: If your print-preview or product-configurator pages are generating repeated audit issues, use Method 3 with a wildcard path pattern to exclude the entire URL group in one robots.txt update, then verify the exclusion in OTTO UI Settings using Method 1 as a secondary layer. 💳 Understanding OTTO Credit Usage OTTO credit consumption follows a pattern that surprises some users: the first crawl of a project is the heaviest. Here is why and what to expect: - Initial crawl: OTTO discovers and audits every accessible URL on your site. If your site has 500 pages, all 500 are crawled and credits are consumed upfront. - Subsequent crawls: OTTO uses incremental crawling where possible, prioritizing changed or newly discovered pages. Credit usage typically drops after the first cycle. - Template pages multiply costs: A site with 10 product templates but 300 product pages will crawl all 300 URLs. Excluding template-driven pages before the first crawl saves the most credits. - Crawl frequency: More frequent crawl schedules consume more credits. If credit usage is a concern, reduce crawl frequency in your project settings until exclusions are configured. To maximize credit efficiency, set up your URL exclusions using Method 1 (UI Settings) before triggering a manual crawl or enabling an automatic crawl schedule on a new project. ✅ Quick Reference — Which Method Should I Use? - A few specific pages: Robots meta tag (Method 2). - URL patterns or folders: robots.txt (Method 3). - Fastest setup with no site edits: OTTO UI Settings (Method 1). - Maximum coverage: Combine Method 1 and Method 3. If you need further assistance, open the chat widget in the bottom-right corner of the platform and type human teammate to be connected with a member of our team.

🕷️ What Are Crawl Errors in SEO and How to Fix Them?

Crawl errors in SEO are issues that prevent search engine bots from accessing and indexing web pages. Crawl errors occur when search engines, such as Google, encounter issues while attempting to crawl your website. Crawl errors hurt your search engine optimization (SEO) performance. They block important pages from appearing in search results. Fixing crawl errors is one of the most important tasks in technical SEO because it directly impacts ranking potential and site visibility. Best practices for crawl health involve running crawl audits, fixing broken links, consolidating redirects, updating XML sitemaps, and using tools like Search Atlas OTTO SEO for automation. A consistent crawl error management strategy improves site health, crawl depth, and search performance. ❓ What Are Crawl Errors in SEO? Crawl errors in SEO are access or retrieval failures that prevent search engine bots (crawlers) from reaching, crawling, or indexing your web pages. Crawlers follow links from page to page across the internet. They analyze content and add it to search indexes. Crawl errors interrupt the crawling process. They prevent crawlers from reaching your pages or understanding your content structure. Crawl errors disrupt the crawling process by returning invalid responses like DNS failures, server errors, 404 not found pages, or redirect loops. Google Search Console flags crawl errors when bots cannot access or interpret a URL correctly. Crawl errors affect page discoverability, reduce index coverage, and waste crawl budget, especially when they recur across critical parts of a site. Search engines group crawl errors into site-level crawl errors and URL-level crawl errors. Site-level crawl errors occur when the crawler cannot reach the domain (e.g., DNS failures or server timeouts). URL-level crawl errors happen when specific resources like pages, images, or scripts are blocked, broken, or misconfigured. 🔍 What Does a Crawl Error Look Like? The most common crawl error examples are below. - HTTP/1.1 404 Not Found - HTTP/1.1 503 Service Unavailable - - Disallow: /checkout/ - Each crawl error either blocks discovery, delays indexing or causes duplicate or conflicting signals. Crawl diagnostics require matching status codes, directives, and canonical tags to the actual page intent. 🗂️ What Types of Crawl Errors Exist? There are six main types of crawl errors. The types of crawl errors include DNS issues, server errors, robots.txt blocks, not found (404) errors, soft 404s, and redirect loops. - DNS failures. Googlebot cannot resolve the IP address of your domain. - Server errors (5xx). The web server returns 500, 502, 503, or 504 errors. - Robots.txt restrictions. Disallowed folders prevent access to important assets or templates. - 404 not found. Page no longer exists, URL is broken, or deleted without a redirect. - Soft 404. The page loads with a 200 status but has no content or displays an error message. - Redirect chains or loops. Multiple chained 301s or infinite loops delay crawling. Crawl errors affect every type of page, such as home, product, category, article, media, and sitemap entries. Google Search Console flags these under Index > Pages > Why pages aren’t indexed. 📉 How Do Crawl Errors Impact SEO? Crawl errors negatively impact SEO by blocking indexation, wasting crawl budget, and breaking internal link equity. Search engines fail to discover or refresh content when crawl errors exist. These failures reduce visibility, traffic, and ranking stability. Crawl errors reduce index coverage. When Googlebot encounters broken pages, it skips them and may remove them from the index. This leads to fewer indexed URLs, weaker topical authority, and lower keyword presence across the site. Crawl errors disrupt internal linking. Broken links stop PageRank flow, which weakens authority signals passed between pages. Redirect loops and invalid anchor paths confuse crawlers and lower trust in the site structure. Crawl errors waste the crawl budget. Google allocates a fixed number of crawl requests per site based on performance, authority, and health. Redirect chains, blocked assets, or unreachable pages consume that budget and prevent important URLs from being crawled or refreshed. Crawl error frequency influences crawl stats and crawl scheduling. High error rates signal instability. Search engines deprioritize sites with persistent crawl failures, which delay content updates and degrade SERP performance. Pages with crawl errors lose ranking power. The entire website structure becomes less effective for SEO purposes. ⚠️ What Causes Crawl Errors? You find crawl errors using Google Search Console (GSC) or technical site audits. These methods reveal crawl failure patterns and support ongoing monitoring. The 2 best methods to find crawl errors are below. 1. Use Google Search Console Google Search Console provides the most authoritative crawl error data. Google reports exactly which pages it cannot crawl. The Coverage report shows the indexing status for all discovered URLs. GSC identifies specific error types and affected page counts. To check for crawl errors using Google Search Console, open the Index section and navigate to the Pages report to view crawl-related issues. Focus on sections labeled “Not Indexed” with reasons like blocked by robots.txt, soft 404, redirect error, or server error (5xx). Use the URL Inspection Tool to test specific URLs and view crawlability, indexing status, and active robots/meta directives. Use the Crawl Stats report in Google Search Console to see how often Googlebot visits your site, which URLs it requests, and whether those requests succeed. The Crawl Stats report helps you detect crawl bottlenecks, server issues, and wasted crawl budget. The Crawl Stats report includes the elements below. - Total crawl requests. The number of URLs Googlebot attempted to fetch. - Total download size. The volume of data downloaded. - Average response time. The average speed of the server during crawling. - Crawl response codes. A breakdown of status codes (200, 404, 5xx, etc.). - File types crawled. HTML, CSS, images, JavaScript, etc. - Crawl purpose. Whether the request was for discovery or refresh. - Googlebot types. Desktop, smartphone, AdsBot, etc. - Host status. Your server’s health, including DNS and robots.txt availability. Use the Crawl Stats report if your website has over 1,000 pages, you’ve noticed indexation delays or crawl errors, you recently changed robots.txt, DNS, or hosting infrastructure, or you want to diagnose why crawl rates spiked or dropped. The Crawl Stats report feature only works on domain-level or root-level URL prefix properties (e.g., https://example.com). 2. Run a Site Audit with Search Atlas Search Atlas includes a dedicated Site Audit Tool that simulates how search engines crawl and process your website. The Search Atlas Site Audit Tool identifies crawlability issues like broken links, 3XX chains, blocked resources, and content rendered via JavaScript. To start a crawl, follow the steps below. - Enter your domain and configure crawl depth (recommended crawl depth is the number of total pages plus 10%). - Set crawl frequency (weekly crawl for dynamic content, monthly crawl for static sites). - Choose the crawler user agent (Googlebot Mobile is recommended). - Adjust speed and rendering options. 1. The Overview section within the Search Atlas Site Audit Tool will show you all page-type crawl errors within your website. 2. Navigate to the Page Explorer section of the Search Atlas Site Auditor to find the exact links that are causing crawl errors within your site and fix them immediately. 3. Use the “Crawl Monitoring” dashboard to track how bots interact with your pages over time. The Crawl Monitoring dashboard is connected to the OTTO SEO agent and allows you to view crawl activity across Google, Bing, GPTBot, ClaudeBot, and other search engines or AI bots. 4. Monitor which pages get visited, how often, and which bots prioritize which sections, and spot crawl rate anomalies, status code spikes, or signs of crawl 5. Search Atlas is the first SEO platform to track multi-bot crawl activity in one unified view. You gain a clear competitive edge by understanding exactly how human and AI bots process your site. 🛠️ How to Fix Crawl Errors? Fix crawl errors by replacing broken links, collapsing redirect chains, correcting crawl directives, and submitting clean sitemaps for re-indexing. Use both Google Search Console and Search Atlas tools to apply and validate crawl error fixes. The nine steps to fix crawl errors are below. 1. Open Google Search Console. Navigate to Index > Pages and filter by “Not Indexed.” 2. Run a crawl audit. Use the Search Atlas Site Auditor to generate a full site crawl, flagging 404s, 5xxs, and disallowed resources. 3. Fix broken internal links. Replace dead links with valid URLs or remove them completely. 4. Update redirects. Collapse redirect chains into a single 301 and eliminate loops. 5. Clean your sitemap. Use Search Atlas Page Explorer to remove 404s and verify each URL returns a 200 status. 6. Review meta robots tags. Ensure public pages use “index, follow” and no conflicting “noindex” exists. 7. Verify robots.txt. Open “https://yourdomain.com/robots.txt” and confirm crawlable sections are not disallowed. 8. Submit the sitemap. Resubmit a clean sitemap to Google Search Console. 9. Reinspect fixed URLs. Use GSC’s URL Inspection Tool to request reindexing after fixes. Fix propagation takes 2 to 14 days, depending on crawl frequency. ⚡ How to Fix Crawl Errors Automatically with OTTO SEO? Search Atlas OTTO SEO automates crawl error resolution by detecting, prioritizing, and correcting issues that prevent indexing or efficient crawling. OTTO SEO fixes broken links, redirect chains, and crawl-blocking directives in real-time. OTTO SEO applies five automated corrections to common crawl issues. Each fix updates live without CMS access or developer support. - Replace broken internal links. OTTO SEO detects deleted or outdated internal links and replaces them with valid final destination URLs that return a 200 status. issues with links search atlas OTTO SEO - Remove redirect chains. The system identifies multi-step redirection paths and collapses them into a single direct link to the canonical target. - Correct meta robot conflicts. OTTO SEO finds and flags crawlable URLs that include conflicting noindex or nofollow directives and corrects them to use index, follow. - Fix canonical mismatches. OTTO SEO updates pages with inconsistent or incorrect canonical declarations to reflect accurate source-target relationships. - Trigger indexing. OTTO SEO submits fixed or newly generated URLs to Google using its Dynamic Indexing system after resolving crawl errors. dynamic indexing search atlas OTTO SEO OTTO SEO “Issues With Links” and “Indexing” modules allow one-click deployment. Search Atlas OTTO SEO fixes apply instantly without developer assistance. All corrections sync with audit logs. Search Atlas specifically enhances crawl error management by efficiently resolving redirect chains. It identifies all instances of outdated URLs across your website and ensures they are replaced with final destination URLs, which preserves link equity and improves the user journey. By dynamically analyzing redirects, Search Atlas eliminates unnecessary hops, ensuring a more streamlined crawl process. 🧭 What Are the Best Practices for Crawl Health? The best practices for crawl health help search engines navigate your site, avoid crawl traps, and prioritize indexable content. The list below includes the crawl health best practices that improve crawlability and indexability, reduce error rates, and support ranking performance. The nine best practices for crawl health are below. Use 301 Redirects Instead of 302 Using 301 redirects instead of 302 helps preserve ranking signals and improves crawl efficiency, especially during site migrations or URL restructuring. Permanent redirects (301) pass full link equity and signal to search engines that the destination URL has replaced the original. Avoid using 302 redirects unless the destination is temporary. Tools like the Search Atlas Site Auditor can detect redirect types and flag misconfigured links during audits. Keep Your Sitemap Clean and Accurate Your XML sitemap should only include live, indexable pages that return a 200 status code. Remove all 404s, redirects, noindexed pages, and disallowed URLs to prevent search engines from wasting the crawl budget. Validate your sitemap using Google Search Console’s sitemap submission tool or the Search Atlas Site Auditor (OTTO SEO → Site Audit → Issues) to confirm proper formatting and coverage. Avoid Redirect Chains in Internal Links Every internal link should go directly to the target page without passing through intermediate redirects. Redirect chains dilute link equity, delay crawler access, and may lead to incomplete indexing. The OTTO SEO “Issues with Links” module inside Search Atlas scans your site for chains and loops, then suggests one-click fixes to point each internal link to its final URL. Validate Your Robots.txt File Syntax issues or aggressive disallow rules in your robots.txt file can unintentionally block pages that need to be crawled and indexed. Test your robots.txt file using Google’s Robots.txt Tester and verify that only irrelevant or sensitive sections (like admin or login pages) are disallowed. You can view the crawl status by page in the Search Atlas Crawl Monitoring dashboard for deeper analysis. Avoid Excessive Disallow Rules Use Disallow directives only for non-public pages that should never appear in search results. Overblocking common paths or entire directories may cut off important navigation flows from crawler access. Review crawl paths and disallow directives using the Search Atlas crawl configuration, and refine exclusions to reduce friction for bot navigation without overexposure. Enable Self-Referencing Canonical Tags Every indexable page should include a canonical tag that points to its own URL unless a specific canonical version exists. Self-referencing canonical tags helps consolidate link signals and clarify indexing decisions. Missing or incorrect canonical tags can lead to duplicate content issues. Use Absolute URLs in Canonicals and Sitemaps Always use absolute URLs with the full protocol and domain in canonical tags and XML sitemaps. Relative URLs can be misinterpreted by search engines, leading to indexing inconsistencies. For example, use https://yourdomain.com/page instead of /page. Confirm correct usage through sitemap validation tools, which show canonical paths and sitemap inclusion side by side. Run Weekly Site Audits Running regular site SEO audits helps to detect crawl issues early and maintain optimal technical SEO health. Use the Search Atlas Site Audit Tool to scan for 404s, broken links, redirect chains, and meta tag conflicts. Schedule weekly SEO audits for dynamic or content-heavy websites to catch new errors quickly and reduce the time between error emergence and resolution. Monitor Crawl Stats in Google Search Console Google Search Console’s Crawl Stats report provides insights into how often Googlebot visits your site, how many URLs it crawls per day, and how long it takes to fetch each page. Watch for spikes in crawl errors, sudden drops in activity, or elevated fetch times that may signal deeper issues. Pair these reports with Search Atlas Crawl Monitoring to cross-reference bot behavior across Google, Bing, and AI crawlers like GPTBot and ClaudeBot. 🧠 What to Know About Crawl Errors Besides Technical SEO? Crawl errors reduce technical SEO performance by blocking discovery, delaying indexing, and misusing crawl resources. Technical SEO depends on crawlability, indexation, and site accessibility. Crawl errors break that flow. Google prioritizes sites that load fast, stay secure, and serve content without structural friction. Frequent crawl errors (e.g., 404s, 5xx, blocked JS) signal low-quality site health. Crawl error prevention is one of the most important segments of technical SEO. This optimization pillar improves site performance, strengthens architecture, and enhances search engine trust. 🔁 What Is the Difference Between Crawl Errors and Indexing Errors? Crawl errors prevent page discovery. Indexing errors occur after discovery, blocking search result inclusion. A crawl error example would be that robots.txt blocks /blog/, therefore, pages are not crawled. An indexing error example would be that the page is crawlable but marked noindex in the meta robots. You need to fix the crawl barriers first. Then address indexing directives and content quality signals. 📊 How Does the Crawl Budget Relate to Crawl Errors? The crawl budget defines how many pages search engines will crawl. Crawl errors waste that crawl budget. Every broken link, redirect chain, or disallowed path consumes crawl budget allocation. If 70% of crawl attempts fail, then critical pages never get discovered or refreshed. Improving crawl efficiency directly increases index coverage, freshness, and SERP visibility. Crawl errors may seem like a technical detail, but they are one of the biggest factors shaping your site’s visibility. When search engines cannot crawl your content, they cannot index or rank it. By diagnosing issues with Google Search Console and Search Atlas—and applying fixes manually or automatically with OTTO SEO—you ensure that your site remains discoverable, efficient, and competitive. Proactive crawl health management isn’t just maintenance—it’s the foundation of sustainable SEO growth.

🕷️ What Are Crawl Errors in SEO and How to Fix Them? (Part 2 of 2)

← Back to Part 1 — Part 2 of 2 🧠 What to Know About Crawl Errors Besides Technical SEO? Crawl errors reduce technical SEO performance by blocking discovery, delaying indexing, and misusing crawl resources. Technical SEO depends on crawlability, indexation, and site accessibility. Crawl errors break that flow. Google prioritizes sites that load fast, stay secure, and serve content without structural friction. Frequent crawl errors (e.g., 404s, 5xx, blocked JS) signal low-quality site health. Crawl error prevention is one of the most important segments of technical SEO. This optimization pillar improves site performance, strengthens architecture, and enhances search engine trust. 🔁 What Is the Difference Between Crawl Errors and Indexing Errors? Crawl errors prevent page discovery. Indexing errors occur after discovery, blocking search result inclusion. A crawl error example would be that robots.txt blocks /blog/, therefore, pages are not crawled. An indexing error example would be that the page is crawlable but marked noindex in the meta robots. You need to fix the crawl barriers first. Then address indexing directives and content quality signals. 📊 How Does the Crawl Budget Relate to Crawl Errors? The crawl budget defines how many pages search engines will crawl. Crawl errors waste that crawl budget. Every broken link, redirect chain, or disallowed path consumes crawl budget allocation. If 70% of crawl attempts fail, then critical pages never get discovered or refreshed. Improving crawl efficiency directly increases index coverage, freshness, and SERP visibility. Crawl errors may seem like a technical detail, but they are one of the biggest factors shaping your site’s visibility. When search engines cannot crawl your content, they cannot index or rank it. By diagnosing issues with Google Search Console and Search Atlas—and applying fixes manually or automatically with OTTO SEO—you ensure that your site remains discoverable, efficient, and competitive. Proactive crawl health management isn’t just maintenance—it’s the foundation of sustainable SEO growth.