Crawling & Robots
By Camilo Aponte
By Camilo Aponte
🔍 Check Crawl Status and Understand Crawl Timelines
🔄 How to Recrawl Your Website
This article explains how to start a new website crawl, manually reprocess the results, and confirm that updated data is available. 🛠️ Step-by-Step 1. Open the project containing the website you want to refresh. 2. Open the website's crawl or processing controls. 3. Select the option to start a new recrawl, then confirm the action. 4. Wait for the crawl to finish so the latest pages and website data are collected. 5. After the crawl completes, select the option labeled Manual reprocess or equivalent to process the newly collected data. 6. Allow processing to complete before reviewing reports or other results. ✅ How to Confirm It Worked Check that the crawl and processing statuses show as completed, the latest crawl or processing timestamp has updated, and refreshed website data appears in the relevant results. If an error appears or the timestamp does not change, retry after the current job finishes. If you need further assistance, open the chat widget in the bottom-right corner of the platform and type human teammate to be connected with a member of our team.
🔍 Scheduled URL Crawl Frequency Explained
📅 What the crawl frequency means A scheduled crawl frequency tells Search Atlas how often a URL should be considered for crawling. For example, a setting of every 7 days means the URL is eligible to be crawled approximately every 7 days. It is not a guaranteed completion time. ⏳ Why a crawl may be delayed A URL can go beyond its selected frequency when crawl work is delayed or cannot be completed. Common reasons include: - Other scheduled crawl jobs are being processed first. - Temporary platform capacity limits or traffic bursts require crawl requests to be spread out. - The previous crawl was unsuccessful or did not complete. - The URL is being checked as part of another crawl or processing workflow. - Results and related recommendations are still waiting for downstream processing. These delays can cause a URL set to every 7–14 days to show a last crawl date from 30–90 days ago. ⚙️ Frequency settings are not normally removed If you manually set a recrawl frequency, the setting should remain active. A missed crawl does not mean the frequency was intentionally deactivated. The next eligible crawl may occur after the delayed work is processed. 🔎 How to check a delayed URL 1. Open the URL or project area where the crawl schedule is managed. 2. Confirm that the URL still has the expected frequency, such as every 7 or 14 days. 3. Review the last crawl date and any available crawl status or error details. 4. Check whether the URL is included in another active crawl workflow. 5. Allow time for queued work to complete, then refresh the crawl status. ✅ What you should expect Scheduled crawling is designed to prioritize recurring checks while managing crawl capacity safely. The selected frequency is the intended cadence, but actual crawl completion can vary. A delayed crawl does not necessarily indicate a problem with your URL or settings. 🛠️ When to request help Contact support if a URL remains unprocessed for an extended period, its frequency changes unexpectedly, or the crawl shows a persistent error. Include the project, URL, selected frequency, last crawl date, and any visible error message so the issue can be reviewed efficiently. If you need further assistance, open the chat widget in the bottom-right corner of the platform and type human teammate to be connected with a member of our team.
🔍 Understanding URL Crawl Frequency
🔍 What recrawl frequency means A recrawl frequency is the preferred interval for checking a URL again, such as every 7, 14, or 30 days. It is a scheduling preference, not a guarantee that every URL will be crawled exactly on that date. 🗓️ Why a URL may be overdue Several factors can delay a recrawl, even when a URL has a frequent schedule: - Crawl capacity: Available crawl resources are shared across projects and may affect timing. - Project eligibility: The project or URL must be active and eligible for recrawling. - Previous crawl results: Errors, blocked access, timeouts, or unavailable pages can prevent a successful recrawl. - Scheduling delays: Recrawls may run later than the selected interval when queued work or system processing affects execution. - Configuration changes: A URL may have been added, removed, or updated after the current schedule was created. ⚙️ How schedules are applied The selected frequency tells Search Atlas how often to prioritize a URL for crawling. It does not reset the URL's crawl date immediately or force a crawl at the exact interval. For example, a 7-day setting means the URL becomes eligible for recurring processing around that interval, subject to availability and access. Changing a frequency also does not necessarily trigger an immediate crawl. The next run depends on the URL's current crawl state and the platform's scheduling process. 📊 What to check when crawls are delayed 1. Open the relevant project and confirm that the project is active. 2. Check that the URL is still included in the project and has the expected recrawl frequency. 3. Review the URL's latest crawl date and any crawl status or error details. 4. Confirm that the page is publicly accessible and does not block Search Atlas or search-engine crawlers. 5. Allow additional time if the URL is queued or the latest crawl attempt was unsuccessful. 🚀 Best practices for recurring crawls - Use shorter intervals for pages that change often or are business-critical. - Use longer intervals for stable pages to focus crawl capacity where it matters most. - Keep project settings and URL lists up to date. - Fix access, redirect, server, and robots.txt issues that prevent successful crawling. - Review crawl history regularly instead of relying only on the selected frequency. 💬 Need help investigating a URL? If a URL remains uncrawled well beyond its expected interval, review its project status, crawl history, and access settings first. If you need further assistance, open the chat widget in the bottom-right corner of the platform and type human teammate to be connected with a member of our team.
🔍 Understanding Per-Page Recrawl and Site Metrics
🗺️ Overview When you request a recrawl for a single page in Search Atlas, you may notice that the platform crawls your entire site rather than just the one URL. This article explains why that happens and how to get the most out of your Site Metrics data once the crawl is complete. 🔄 Why Does a Single-Page Recrawl Trigger a Full Site Crawl? Search Atlas crawls your site at the domain level, not on a page-by-page basis. This is by design. Here is why: - Accurate internal linking data: To correctly evaluate any single page, the crawler needs to understand how it connects to the rest of your site. Internal links, anchor text, and crawl depth all influence a page's SEO metrics. - Consistent site-wide metrics: Page authority, crawl health scores, and indexability signals are calculated relative to your full site structure. Crawling only one page would produce incomplete or misleading data. - Dependency on fresh site context: Changes you make to one page — such as updating a title tag or adding a canonical — can affect how other pages are evaluated. A full crawl ensures every metric reflects your site's current state. In short, a "per-page" recrawl request is processed as a trigger to refresh the entire site's crawl data so that all metrics remain accurate and in sync. ⏱️ How Long Will the Crawl Take? Crawl duration can vary depending on your site. Larger or more complex sites may take more time to process. You do not need to stay on the page — your results will be ready when you return. 📊 Viewing Results in Site Metrics Once the crawl is complete, your updated data is available in Site Metrics. Navigate to OTTO SEO → Site Audit → Crawl Monitoring within the platform and select the site you recrawled to review updated crawl health, page-level data, and any flagged issues. If you are unsure where to find your results or your crawl data does not appear to have updated, please reach out to our support team with your project name, the URL you attempted to recrawl, and the approximate time you submitted the recrawl request. If you need further assistance, open the chat widget in the bottom-right corner of the platform and type human teammate to be connected with a member of our team.
⚡ OTTO Crawl Limits for Bulk Page Additions
Overview When you add a large number of pages to OTTO at once, you may notice that only a portion of them appear in your crawl results immediately. This is expected behavior, not a bug. OTTO's crawl function has built-in processing limits for bulk additions, and full visibility across all newly added pages may not happen instantly. How OTTO Processes Newly Added Pages OTTO separates two distinct processes that customers often expect to happen simultaneously: - Page addition: The act of submitting new URLs to OTTO for tracking and optimization. - Crawl visibility: The appearance of those pages within OTTO's crawl results. Adding pages successfully does not mean they will all be visible in the crawl results right away. OTTO queues pages for analysis and processes them over time. When a large number of pages are submitted at once, only a portion may appear immediately — the rest are queued and will surface as OTTO continues to process them. Why Only Some Pages Are Appearing If you added a large batch of pages and only a fraction are showing in your crawl results, this reflects OTTO's processing behavior for bulk submissions. The pages that do not appear immediately are still queued and tracked — none are lost. They will become visible as OTTO works through the queue. This behavior is by design to maintain platform performance across all users. What You Can Do If you are waiting for pages to appear in OTTO after a bulk addition, the following steps can help manage your workflow: 1. Wait for processing to complete: Pages submitted in bulk will appear progressively as OTTO processes the queue. Check back periodically to see newly visible pages. 2. Submit smaller batches when possible: If you need certain pages to be processed sooner, consider submitting your highest-priority pages in a smaller, separate batch before submitting the rest. Smaller submissions are generally processed more quickly. 3. Confirm pages are queued: Verify within OTTO that your submitted URLs are listed and tracked, confirming they have been accepted into the system even if not yet fully visible in crawl results. 4. Contact support if pages remain missing: If pages have not appeared after an extended period and you believe there may be an issue, reach out to the support team with the specific URLs and the date you submitted them so the team can investigate. If you need further assistance, open the chat widget in the bottom-right corner of the platform and type human teammate to be connected with a member of our team.
🔄 How to Recrawl and Reprocess Your Website Audit
This article explains how to initiate a recrawl of your website and trigger a manual reprocess of your audit data in Search Atlas. Use these steps when your audit results feel outdated, changes you've made aren't reflected, or you want to force a fresh analysis of your site. 🛠️ Step-by-Step 1. Log in to your Search Atlas account and navigate to your project's site audit area. Depending on your plan and setup, this is typically accessible from your main project dashboard. 2. Open the audit report for the website you want to recrawl. Look for the site or domain you wish to update in your project list. 3. Navigate to OTTO SEO → Site Audit → Overview (Website Overview) and locate the Recrawl Site button. This is usually represented by a refresh button near the top of the audit report page, often next to the last crawl date or crawl status indicator. 4. Click the recrawl button to initiate a new crawl. The platform will begin re-scanning your website pages. Crawl time varies depending on your site's size — larger sites may take several minutes to complete. 5. Once the crawl has finished, locate the Reprocess option. This step re-analyzes the newly collected crawl data and updates all audit scores, issue counts, and recommendations. The reprocess button is typically found in the same audit controls area, often labeled Reprocess or Recrawl Site. 6. Click Reprocess (or its equivalent) to trigger the manual reprocessing. Wait for the platform to confirm the reprocess has completed before reviewing your updated results. ✅ How to Confirm It Worked After completing both steps, confirm the recrawl and reprocess were successful by checking the following: - The Last Crawled or Last Updated timestamp on your audit report should reflect today's date and the approximate time you initiated the recrawl. - Your audit scores, issue counts, and page-level data should reflect any changes you made to your site since the previous crawl (for example, fixed issues should no longer appear as errors). - If a progress bar or crawl status indicator was visible during the process, it should now show a Completed or Done status rather than a pending or in-progress state. - If issue counts or scores appear unchanged and the timestamp has not updated, wait a few minutes and refresh the page, as large sites may take additional time to fully reprocess. If you need further assistance, open the chat widget in the bottom-right corner of the platform and type human teammate to be connected with a member of our team.
🕷️ Optimise Crawler Settings for Your Server
🔍 Understanding How the Crawler Uses Server Resources The Search Atlas crawler is designed to be thorough and accurate, which means it performs multiple operations per page visit — fetching HTML, resolving redirects, analysing on-page elements, and storing results. Even when the crawl speed is set to 1 page per second, these parallel background processes can place a noticeable load on your server, especially on resource-constrained hosting environments. The key thing to understand is that crawl speed (pages per second) is not the only factor that determines server load. Concurrent connections, page complexity, server response times, and the size of your site all contribute to total resource consumption during a crawl. ⚙️ What Affects Server Resource Usage - Concurrent connections: Even at 1 page per second, the crawler may open multiple simultaneous connections to fetch resources like images, scripts, and stylesheets referenced on each page. - Page complexity: Pages with many internal links, large HTML payloads, or heavy JavaScript increase the processing work required per page. - Server response time: A slow server causes connections to remain open longer, stacking up alongside new requests and increasing peak load. - Redirect chains: Each redirect adds an extra request, multiplying the total number of HTTP calls the crawler makes. - Site size: Larger sites mean the cumulative effect of even a conservative crawl rate lasts much longer, sustaining elevated resource usage over an extended period. 🖥️ Recommended Configuration for an 8-Core Dedicated Server An 8-core dedicated server has more headroom than shared hosting, but it is still important to configure the crawler thoughtfully to avoid impacting live site performance for real visitors. Follow these guidelines: 1. Set crawl speed to 2–4 pages per second as a starting point. An 8-core server can generally handle this comfortably, but monitor CPU and memory usage during the first crawl to confirm. 2. Schedule crawls during low-traffic periods — typically late night or early morning in your visitors' primary time zone. This ensures the crawler competes with as little live traffic as possible. 3. Limit concurrent connections if your crawler settings expose this option. A value of 4–6 concurrent connections is a sensible ceiling for most dedicated servers. 4. Exclude unnecessary URLs from the crawl scope. Use the URL exclusion or path filtering settings to skip admin paths, media upload directories, and any dynamically generated pages that do not need SEO analysis. 5. Pause and resume crawls if you notice performance degradation. You do not need to restart from the beginning — the crawler can resume where it left off. 6. Review your server's error logs after the first crawl. A spike in 503 (Service Unavailable) or 429 (Too Many Requests) responses is a reliable signal that you need to reduce speed or concurrency further. 💡 Tips to Reduce Crawl Impact Without Sacrificing Coverage - Use a crawl budget wisely: Prioritise crawling your most important URLs — product pages, blog posts, and landing pages — and exclude low-value pages like tag archives, search result pages, and session-ID URLs. - Enable caching on your server: A properly configured server-side cache (e.g. Redis, Varnish, or a CDN) dramatically reduces the processing overhead of each crawler request because the server returns a cached response instead of regenerating the page. - Check your robots.txt: Ensure the Search Atlas crawler is not blocked from pages you want crawled, and that it is blocked from areas you do not — such as /wp-admin/ or /checkout/. - Keep the crawl scope focused: If you are running a targeted audit, crawl a specific subdirectory rather than the entire domain to reduce total requests. 🚀 Monitoring Resource Usage During a Crawl While a crawl is running, use your server's monitoring tools (such as htop, Netdata, or your hosting control panel's resource graphs) to watch CPU load, RAM usage, and active connections in real time. A healthy crawl on an 8-core dedicated server should keep average CPU usage below 60–70%. If you consistently see CPU peaking above 90% or RAM approaching its limit, reduce the crawl speed setting and restart the crawl. If your hosting provider enforces rate limits or automatically throttles connections from a single IP, contact them to whitelist the Search Atlas crawler IP ranges so legitimate crawl traffic is not misidentified as an attack. 🛠️ Need Further Help If you need further assistance, open the chat widget in the bottom-right corner of the platform and type human teammate to be connected with a member of our team.
🔄 How to Recrawl Your Website After a Site Migration
📋 Overview After completing a site migration—whether you changed domains, switched to HTTPS, restructured URLs, or moved platforms—your Search Atlas data may still reflect your old site. Running a fresh crawl ensures your audits, Site Explorer data, and on-page recommendations match your current website. This guide walks you through how to recrawl your entire site and what to do after the crawl finishes. 🕒 When You Should Recrawl Recrawl your website whenever you have made significant changes, including: - Migrating to a new domain or subdomain - Switching from HTTP to HTTPS - Changing your URL structure or permalinks - Moving to a new CMS or hosting platform - Publishing or removing large amounts of content ✅ Before You Start To get accurate results, confirm the following first: - Migration is complete. Wait until all redirects, DNS changes, and content updates are live. - The site is publicly accessible. Remove any password protection, staging restrictions, or "noindex" directives that should not be there. - Your robots.txt allows crawling. Make sure it does not block the pages you want analyzed. - The correct domain is set in your project. If you migrated to a new domain, you may need to add it as a new project (see below). 🔁 Recrawling Within the Same Domain If your site stayed on the same domain (for example, you only changed URL structures or content), you can recrawl your existing project: 1. Log in to your Search Atlas Home. 2. Open the Site Audit tool from the left-hand menu. 3. Select the project for your migrated website. 4. Click the Recrawl (or Crawl Now) button at the top of the audit page. 5. Confirm the crawl settings, including crawl depth and page limits, then start the crawl. The crawl runs in the background. Larger sites take longer, so allow time for it to finish before reviewing your results. 🌐 Recrawling After a Domain Change If you migrated to a brand-new domain, create a new project so your data is tied to the correct URL: 1. From your Search Atlas Home (Hi, CSM. What will we get done today?), click the Create Project button — or go to OTTO SEO → All Sites (OTTO SEO Projects) and click Create. 2. Enter your new domain exactly as it appears live (including HTTPS and www, if applicable). 3. Complete the project setup steps and connect any integrations, such as Google Search Console. 4. Open Site Audit for the new project and start your first crawl. Domain reprocessing for a migrated site now runs automatically once your project is set up—you no longer need to submit a manual request to our support team to have the domain recrawled. ⚙️ Adjusting Crawl Settings For larger or recently migrated sites, review these settings before crawling: - Page limit: Increase this if your site has more pages than the default allows, so the full site is captured. - Crawl depth: Set a depth that reaches deeper pages in your site structure. - Subdomains: Enable subdomain crawling if your content lives across multiple subdomains. 🏁 After the Crawl Finishes Once the crawl completes, take these steps to confirm everything migrated correctly: - Check the crawled page count. Compare it to your expected number of live pages. - Review redirects. Look for broken redirect chains or 404 errors from old URLs. - Scan for indexing issues. Confirm important pages are not blocked or marked "noindex." - Look at site health scores. Use the audit recommendations to fix any new issues introduced during migration. 🛠️ Troubleshooting 🔍 The crawl returns very few pages This usually means crawling is being blocked. Check your robots.txt file, remove any unintended "noindex" tags, and confirm the site is not behind a login or staging firewall. 🔗 Old URLs still appear in reports Historical data is retained for reference. After a fresh crawl, new reports reflect your current site. If you changed domains, make sure you are viewing the new project. 🤖 The AI agent says the recrawl finished, but results haven't updated If you triggered the recrawl through the in-app AI agent, confirm the crawl status directly in the Site Metrics tool rather than relying on the agent's confirmation. We are aware of an issue where the agent can report a recrawl as "completed" before it has actually verified the crawl's status. Treat the crawl status shown in Site Audit as the source of truth. ⏳ The crawl is taking a long time Large sites naturally take longer. Crawls run in the background, so you can leave the page and return later. If a crawl appears stuck for an unusually long time, contact our support team. 💬 Need More Help? If your results still look incorrect after recrawling, reach out to our support team through the in-app, typing "Human Teammate". Include your project name, your old and new domains, and the date of your migration so we can assist you quickly.
🔄 Trigger a Site Auditor Recrawl After Restructuring
🗺️ Why a Recrawl Is Needed After Architecture Changes When you make significant changes to your website — such as restructuring your hub-and-spoke architecture, reorganising URL hierarchies, adding new content hubs, or updating internal linking — the Site Auditor retains data from its most recent crawl. Until a new crawl runs, the audit results will reflect your old site structure, not the updated one. Triggering a fresh recrawl tells the system to re-examine every accessible page on your domain and rebuild its understanding of your site's current architecture, links, and technical health. 🛠️ How to Trigger a Recrawl in Site Auditor Follow these steps to start a new crawl for your project: 1. Log in to your Search Atlas account and navigate to Site Auditor from the left-hand menu. 2. Select the project (domain) you want to recrawl from your project list. 3. Once inside the project dashboard, locate the option to initiate a new crawl and confirm the action when prompted. The crawl will begin and may take several minutes to complete depending on the size of your site. 4. Once finished, your audit results will refresh to reflect your updated site structure, internal links, and technical health data. 💡 Best Practices When Recrawling After Major Changes - Wait until changes are fully live. Run the recrawl only after all your architecture updates have been published and are accessible to crawlers. Recrawling a half-deployed site will produce mixed results. - Submit an updated sitemap first. Before triggering the recrawl, submit your refreshed XML sitemap to Google Search Console. This ensures both Google and Site Auditor are working from the same URL set. - Check your robots.txt. Confirm that no sections of your new architecture are accidentally blocked from crawling before you initiate the audit. If you need further assistance, open the chat widget in the bottom-right corner of the platform and type human teammate to be connected with a member of our team.
🖥️ Search Atlas Crawler Resource Optimization for Dedicated Servers
🔍 Why the Crawler Can Overload a Server at 1 Page/Second Setting the crawler to 1 page per second does not mean your server processes only 1 lightweight request per second. Each crawl request triggers a chain of backend operations: DNS resolution, TCP handshake, HTML download, parsing, link extraction, redirect following, and database writes. On a typical server, this backend pipeline takes 200–800 ms per page depending on page complexity, server configuration, and whether caching is active. The actual CPU load is calculated as: CPU threads consumed ≈ Requests per second × Backend processing time (in seconds) For example, at 1 request/second with 600 ms average backend processing time, the crawler holds approximately 0.6 threads busy at all times. If you increase to 5 requests/second, that becomes 3 threads continuously occupied — consuming 75% of a 4-core server before accounting for your website, database, or other services running simultaneously. This is the root cause of excessive resource usage even at seemingly low crawl rates. ⚙️ How Caching Status Changes Resource Consumption Caching is the single biggest variable in crawler resource usage. When server-side caching (e.g., Redis, Varnish, or full-page caching plugins) is active and warm, backend processing time per page drops dramatically — often to 50–150 ms. When caching is cold or disabled, processing time rises to 400–900 ms per page, multiplying CPU load significantly. - Warm cache: Higher crawl speeds are safe; backend overhead is minimal. - Cold or no cache: Reduce crawl speed significantly; each request hits PHP, the database, and templating engines at full cost. - After a cache flush: Always lower crawl speed temporarily and allow the cache to rebuild before increasing again. 📋 Recommended Configuration Templates by Server Specs Use the table below as a starting point. Adjust based on your actual CPU load readings during a crawl (see Monitoring section below). 4-Core Server (16–32 GB RAM) - Max crawl speed (cache warm): 2–3 pages/second - Max crawl speed (cache cold): 1 page/second - Max concurrent connections: 2 - Crawl delay between requests: 400 ms minimum - Recommended crawl window: Off-peak hours only 8-Core Server (32–64 GB RAM) - Max crawl speed (cache warm): 5–7 pages/second - Max crawl speed (cache cold): 2–3 pages/second - Max concurrent connections: 4 - Crawl delay between requests: 150–200 ms - Recommended crawl window: Can run during business hours; monitor CPU for first 10 minutes 16-Core+ Server (64 GB+ RAM) - Max crawl speed (cache warm): 10–15 pages/second - Max crawl speed (cache cold): 5–6 pages/second - Max concurrent connections: 6–8 - Crawl delay between requests: 80–100 ms - Recommended crawl window: Flexible; maintain CPU headroom above 30% for other services Important: These are starting baselines. Always monitor live CPU and RAM usage during the first crawl session and scale back if utilization exceeds 70% sustained CPU on any core. 🛠️ How to Adjust Crawler Settings in Search Atlas 1. Log in to Search Atlas and navigate to Left sidebar → Site Explorer. 2. Select the project or domain you want to crawl. 3. Open the crawler configuration panel for that project. 4. Set Pages per second using the recommended values above for your server tier. 5. Set Concurrent connections to match the recommended value for your core count. 6. Enable Crawl delay and input the minimum millisecond value for your tier. 7. Save your configuration and start a test crawl on a small URL batch (50–100 pages) before running a full site crawl. 📊 Monitoring Tools and Metrics to Watch During and after a crawl, track the following server metrics to confirm your configuration is safe: - CPU utilization per core: Use top or htop on Linux. No individual core should sustain above 80% during a crawl. If it does, reduce pages/second immediately. - Load average: A 1-minute load average exceeding your core count (e.g., above 8.0 on an 8-core server) signals the server is queueing work faster than it can process it. - RAM usage: Watch for memory pressure above 85% used. Swap usage during a crawl indicates the server is under-resourced for your current settings. - Database query time: Slow query logs spiking during a crawl confirm the database is a bottleneck. Lower crawl speed and verify caching is routing repeated lookups away from the database. - Server response time: If your site's Time to First Byte (TTFB) rises above 1 second during crawling, the crawler is negatively impacting live visitors. Reduce speed immediately. ✅ Quick Troubleshooting Checklist - Verify caching is active and warm before starting a crawl at higher speeds. - Confirm no other resource-intensive processes (backups, database optimizations) are running simultaneously. - Start every new crawl configuration with a 50-page test batch before scaling to the full site. - After any server restart or cache flush, treat the server as "cold cache" and use the lower speed values until the cache rebuilds. - Review crawl logs for a high proportion of redirect chains or error responses — these increase backend processing time per request and require lower crawl speeds. 💬 Need More Help? If you need further assistance, open the chat widget in the bottom-right corner of the platform and type human teammate to be connected with a member of our team.
🔄 How to Initiate a Website Recrawl and Reprocess
This article explains how to trigger a fresh crawl of your website and manually reprocess its data inside Search Atlas — so that any recent changes to your site's content, structure, or metadata are picked up and reflected in your reports and tools. 📍 Where to Find It Navigate to OTTO SEO → Site Audit → Overview (Website Overview) using the left sidebar. Select the website you want to recrawl from your list of projects. 🛠️ Step-by-Step 1. Open OTTO SEO → Site Audit → Overview (Website Overview) from the left sidebar and select the target website. 2. Locate the crawl or site settings area within your selected project. Look for the button labeled Recrawl Site and click it to initiate a fresh crawl of your site. 3. Once the crawl has completed (you will typically see a progress indicator or a completion status message), look for a Reprocess or Manual Reprocess option in the same settings area. Click it to trigger the reprocessing of the newly crawled data. 4. If both actions are available as separate buttons, complete the recrawl first and wait for it to finish before initiating the manual reprocess — running them in sequence ensures the reprocess works on the most up-to-date crawl data. ✅ How to Confirm It Worked After initiating the recrawl and reprocess, confirm success by checking the following: - The crawl status indicator should update from In Progress to Completed (or a similar finished state) — this confirms the recrawl finished successfully. - The Last Crawled or Last Updated timestamp on your website should reflect today's date and time. - Navigate to any report or tool that pulls data from this site (such as site audit results or page-level data) and verify that recent changes you made to your site — such as new pages, updated titles, or fixed issues — are now appearing correctly. - If a reprocess was completed, previously flagged issues that you have since fixed should no longer appear, and any newly added content should now be visible in your data. If you need further assistance, open the chat widget in the bottom-right corner of the platform and type human teammate to be connected with a member of our team. Additional Notes If you cannot locate an option to trigger a recrawl or reprocess, a member of our support team can initiate a backend reprocess on your behalf. Contact via chat widget and type 'human teammate' to be connected.
⚙️ Site Audit Crawl Settings and Analysis Timeline
🗺️ Overview When running a Site Audit in Search Atlas, two settings have a direct impact on the quality of your results and how long the audit takes: user agent and crawl speed. Choosing the wrong user agent can reduce page capture accuracy, while an uncalibrated crawl speed may strain your server. This article also clarifies realistic timelines so you know when to expect your audit to finish. 🤖 Choosing the Right User Agent The user agent tells the Site Auditor how to identify itself when crawling your pages. Search Atlas offers several options, and selecting the correct one affects how accurately your pages are rendered and captured. Recommended user agent: Chrome Desktop - Chrome Desktop — The best choice for most sites. It renders pages the way a real browser does, capturing JavaScript-dependent content, Open Graph tags, and dynamic elements accurately. Use this for the most complete and reliable audit results. - Search Atlas Bot — A lightweight crawler that does not render JavaScript. It is faster but will miss dynamically loaded content, metadata, and tags generated client-side. This is not recommended if your site relies on JavaScript rendering. - Googlebot — Mimics Google's crawler. Useful if you specifically want to see how Googlebot views your site, but Chrome Desktop remains the better general-purpose choice. If your previous audits used Search Atlas Bot and returned incomplete Open Graph or metadata results, re-run the audit with Chrome Desktop selected to get a fuller picture of your site's issues. 🚀 Optimising Crawl Speed Crawl speed controls how many pages per second the auditor requests from your server. For most sites — particularly those on shared hosting or with limited server resources — a lower speed is recommended to reduce the risk of overloading your server and to help ensure the audit completes without triggering rate limits or returning incomplete data. - If your site is large and hosted on a dedicated or enterprise-grade server, you may use a higher crawl speed setting. - If you notice server errors or timeouts during a crawl, reduce the speed further to a lower value. To adjust crawl speed, open your Site Audit settings before starting or re-running a crawl and update the crawl speed to your preferred value before launching the audit. ⏱️ How Long Does a Site Audit Take? Site Audit analysis time depends on the size of your site and the crawl speed you have selected. Larger sites and lower crawl speeds will naturally require more time to complete. If your audit appears stuck in analysis, check that your crawl speed is appropriately set for your server capacity and consider re-running the audit with Chrome Desktop as the user agent for the most reliable results. If you need further assistance, open the chat widget in the bottom-right corner of the platform and type human teammate to be connected with a member of our team.
⏱️ Understanding Site Crawl Duration and Speed
🔍 Why Crawl Speed Varies Site crawls in Search Atlas do not run at a fixed speed. Several factors affect how quickly your pages are processed, and it is completely normal for a crawl to take hours — or even longer — depending on your site's conditions. - Site size: Larger sites with hundreds or thousands of pages naturally take more time to crawl than smaller ones. - Server response time: If your web server responds slowly, the crawler waits for each response before moving to the next page. A slow host directly slows down your crawl. - Platform-wide queue: Crawls from all Search Atlas users are processed through a shared queue. During peak usage periods, your crawl may progress more slowly as system resources are distributed across active jobs. - Page complexity: Pages with a large number of internal links, redirects, or heavy JavaScript can take longer to fully process. - Crawl rate limits: To protect your server from overload, the crawler respects rate limits. This prevents large sites from being flooded with rapid requests. 📊 What Is a Normal Crawl Speed? There is no single "normal" speed because every site and server environment is different. As a general guide: - Small sites (under 50 pages) may complete within 15–60 minutes. - Medium sites (50–500 pages) typically take a few hours. - Larger sites (500+ pages) can take several hours to over a day. Seeing only 24 out of 110 pages crawled after several hours is within the range of expected behaviour, particularly if your server has slower response times or the platform queue is busy. The crawl is still running — it has not stalled. ⚙️ How to Check Your Crawl Progress You can monitor an active crawl directly inside Search Atlas: 1. In the left sidebar, click Site Metrics (Site Explorer). 2. Select your site from the list at /site-explorer/list. 3. Review the crawl status indicator to see how many pages have been processed and whether the crawl is still active. As long as the crawl status shows as in progress, the job is running normally. Refreshing the page periodically will update the page count. 💡 Tips to Help Your Crawl Complete Faster While you cannot directly control platform queue times, there are steps you can take to reduce delays caused by your own site: - Improve server response time: Use caching, a CDN, or speak to your hosting provider about performance optimisation. Faster server responses mean faster crawls. - Fix redirect chains: Multiple redirects on a single URL slow down the crawler. Resolve any unnecessary redirect chains in advance. - Avoid triggering a re-crawl unnecessarily: Starting a new crawl before a previous one finishes can reset your progress. Wait for the current crawl to complete first. - Crawl during off-peak hours: Scheduling crawls overnight or during low-traffic periods on your site can sometimes result in faster completion. 🚨 When Should You Be Concerned? In most cases, a slow crawl is not a problem — it just requires patience. However, you should reach out for support if: - The crawl status has shown no progress at all for more than 24 hours. - The crawl appears to be stuck on the same page count over an extended period without any movement. - You receive an explicit error message or the crawl shows a failed status. If you are unsure whether your crawl is progressing normally, it is always fine to check in with our team. 🛠️ Need Further Help? If you need further assistance, open the chat widget in the bottom-right corner of the platform and type human teammate to be connected with a member of our team.
⏱️ Manual Site Crawl Duration and Speed Settings Explained
🔍 Overview When you trigger a manual site crawl in Search Atlas, the process can take anywhere from a few minutes to up to 24 hours, depending on your site's size and server conditions. This article explains what to expect, what the crawl speed setting actually does, and why your crawl duration is influenced by factors beyond that setting. ⏳ How Long Does a Manual Crawl Take? There is no single fixed duration for a manual crawl. The time it takes depends on several variables, but the ranges below give you a reliable reference point: - Small sites (under 500 pages): Typically completes within 15–60 minutes. - Medium sites (500–10,000 pages): Usually takes 1–6 hours. - Large sites (10,000+ pages): Can take anywhere from 6 to 24 hours. If your crawl has been showing as loading for several hours but falls within the range above for your site size, this is expected behaviour. A crawl is only considered stuck if it has exceeded 24 hours without completing. ⚙️ What Does the Crawl Speed Setting Actually Control? The crawl speed setting controls the rate at which the crawler sends requests to your server — measured in requests per second. In simple terms, it determines how aggressively the crawler asks your site for pages. - Lower speed: Fewer requests per second, which is gentler on your server and reduces the risk of triggering rate limits or temporary blocks. - Higher speed: More requests per second, which can help the crawler move through pages faster — but only when your server can respond quickly enough to keep up. Think of crawl speed as setting the maximum pace the crawler tries to move at — not the pace it will always achieve. 🚧 Why Doesn't Crawl Speed Guarantee Total Duration? Even with the highest crawl speed selected, your total crawl time can still be long. This is because crawl speed only controls one part of the process. Several server-side and backend factors determine the actual time to completion: - Server response times: If your web server is slow to respond to requests — due to hosting limitations, high traffic, or page complexity — the crawler must wait before it can move on. No speed setting overrides this. - Rate limiting and throttling: Some servers automatically slow down or temporarily block rapid requests. When this happens, the crawler backs off to avoid being blocked, regardless of your speed setting. - Redirect chains and crawl depth: Sites with deep page hierarchies or long redirect chains require more individual requests to fully map, extending total time. - Backend processing time: After the crawler collects raw data, Search Atlas processes and analyses that data on the backend. This processing phase runs after the crawl itself and adds to the total time before results appear. - Total number of unique URLs: The more unique pages discovered during a crawl, the longer the full process takes, independent of speed. 📋 What to Do If Your Crawl Seems Stuck 1. Check how long the crawl has been running. If it is under 24 hours, allow it to continue — this is within the expected window. 2. Verify that the site being crawled is publicly accessible and not returning widespread errors (such as 5xx server errors), which can cause the crawler to stall. 3. If the crawl has exceeded 24 hours with no change in status, try cancelling and restarting the crawl from OTTO SEO → Site Audit → Crawl Monitoring. 4. Consider lowering the crawl speed setting if your server has limited capacity — this can prevent throttling and actually result in a more stable, complete crawl. 💡 Tips for Faster, More Reliable Crawls - Schedule manual crawls during low-traffic periods on your site so the server can respond more quickly. - Use a moderate crawl speed if you are on shared hosting — very high speeds can trigger server-side blocking that slows the overall crawl. - Ensure your robots.txt file is not inadvertently blocking the Search Atlas crawler from key sections of your site. - If you only need data for a specific section, consider whether a targeted crawl of a subdirectory is more efficient than a full-site crawl. 🙋 Need Further Help? If you need further assistance, open the chat widget in the bottom-right corner of the platform and type human teammate to be connected with a member of our team.
🖼️ Force Search Atlas to Recrawl and Reprocess Images
🔍 Why Search Atlas Still Shows Old Image Data When you replace images on your website — for example, swapping PNG files for JPEGs — Search Atlas may continue displaying the previous image attributes, including outdated alt text, file types, or URLs. This happens because Search Atlas stores a cached snapshot of your site from its last crawl. Clearing your CMS cache or CDN cache does not automatically trigger a new crawl inside Search Atlas. You need to manually instruct the platform to re-fetch and reprocess your images. 📋 What You Will Need - Access to your Search Atlas account - Your CMS and CDN caches already cleared before triggering a new crawl - The target site added and verified inside Search Atlas 🚀 Step 1 — Trigger a Manual Recrawl To force Search Atlas to pick up your updated images, you need to trigger a new crawl of your site from within the platform. Navigate to your site inside Search Atlas and look for the option to manually start or re-run a crawl. Initiate that crawl and allow it to complete before moving to the next step. Depending on your site size, the recrawl can take anywhere from a few minutes to several hours. Wait until the crawl status shows as complete before reviewing your image data. ⚙️ Step 2 — Reprocess Image Data After the recrawl completes, Search Atlas needs to reprocess the updated image data so that new alt texts, file types, and attributes are correctly reflected across all reports. Within your site's reporting area, look for an option to re-run or refresh the analysis. Trigger that reprocessing action and wait for it to finish before checking your results. 🧹 Step 3 — Confirm Your External Cache Is Cleared Before verifying results inside Search Atlas, confirm that your external caches are genuinely cleared and serving the new image files. Stale responses from a CDN or browser cache can cause the Search Atlas crawler to re-fetch the old image even after a recrawl is triggered. - CDN (e.g., Cloudflare): Use your CDN provider's cache purge feature to clear the affected image URLs or purge all cached assets. - CMS (e.g., WordPress): Clear both your plugin-level cache and any server-side page cache. - Browser: Open the image URL directly in an incognito window to verify the new file is being served before re-triggering the crawl. 📞 Still Seeing Stale Image Data? If your images are still showing outdated attributes after completing the steps above, our support team can investigate further. When escalating, please have the following ready: your site's project name in Search Atlas, the specific image URLs affected, the approximate time you replaced the files, and confirmation that your external caches have been cleared. If you need further assistance, open the chat widget in the bottom-right corner of the platform and type human teammate to be connected with a member of our team.
🔄 OTTO Recrawl vs Weekly Analysis Cycle Explained
🎯 Overview OTTO offers two distinct ways to evaluate your site after changes: an immediate recrawl you can trigger manually, and a weekly analysis cycle that runs automatically in the background. Knowing which one to use — and what each one detects — ensures you get accurate feedback after making code or structural updates. ⚡ Immediate Recrawl: What It Does and When to Use It The immediate recrawl button lets you re-check your site on demand without waiting for the scheduled cycle. Use this option after making surface-level fixes such as updating meta tags, adjusting title tags, fixing on-page copy, or resolving schema markup issues that do not involve changes to your site's architecture or codebase. What the immediate recrawl detects: - Updated meta titles and descriptions - Fixed heading tags (H1, H2, etc.) - Corrected image alt text - Resolved internal linking issues on existing pages - Schema and structured data updates on live pages - Changes to on-page content and keyword usage What the immediate recrawl does NOT detect: - Deleted or redirected pages - Codebase-level modifications - New pages added to the site - Changes to site architecture or navigation structure - Server-side rendering updates or framework-level changes Timeline: Results from an immediate recrawl are typically reflected in OTTO within a few minutes to a couple of hours, depending on your site's size. 🗓️ Weekly Analysis Cycle: What It Does The weekly analysis is an automated, deep-scan cycle that OTTO runs on a recurring schedule. This cycle is designed to catch structural and architectural changes that a surface-level crawl cannot detect. If you have made significant updates to your codebase, deleted pages, added new sections, or restructured your site navigation, you will need to wait for — or be aware of — the weekly cycle to see those changes reflected in OTTO's recommendations. What the weekly analysis detects: - Deleted or newly added pages across the site - Codebase and framework-level modifications - Changes to site architecture and internal linking patterns - New or removed URL structures and redirects - Shifts in crawl depth and page indexability - Broad technical SEO changes across the full site Timeline: The weekly analysis runs automatically every seven days. There is no manual trigger for the full weekly cycle — it operates on a fixed schedule set by OTTO. After the cycle completes, OTTO updates its recommendations to reflect all structural changes detected. 🖱️ How to Trigger an Immediate Recrawl 1. In the left sidebar, click OTTO SEO → All Sites (SEO Automation) to open the OTTO dashboard. 2. Select the site project you want to recrawl. 3. Navigate to Site Audit → Overview (Website Overview) and locate the Recrawl Site button. 4. Click Recrawl Site to start the immediate crawl process. 5. Wait for the crawl to complete — a progress indicator will confirm when it is finished. 6. Once complete, review the updated recommendations to confirm your fixes have been recognised. ✅ Choosing the Right Option After Making Changes Use the table below as a quick reference to decide which process applies to your situation: - Fixed a meta tag, title, or on-page content? → Use the immediate recrawl button. Results will appear within hours. - Updated schema or structured data on an existing page? → Use the immediate recrawl button. - Deleted pages, added new pages, or restructured your site architecture? → Wait for the next weekly analysis cycle for accurate results. - Made changes to your codebase or server-side rendering? → Wait for the weekly analysis cycle, as these changes require a deep structural scan. - Not sure which category your change falls into? → Run the immediate recrawl first. If OTTO's recommendations do not update as expected, the change likely requires the weekly analysis cycle to be fully detected. ⏳ Timeline Summary - Immediate recrawl: Triggered manually; results typically within a few minutes to a few hours. - Weekly analysis: Runs automatically every 7 days; no manual trigger available; covers all structural and codebase-level changes. 💬 Need Help? If you need further assistance, open the chat widget in the bottom-right corner of the platform and type human teammate to be connected with a member of our team.
🔄 How to Initiate a Recrawl and Manual Reprocess
This article explains how to initiate a recrawl of your website and manually reprocess your site data. Use these steps when your site's data feels stale, when recent changes aren't reflected in your reports, or when you've been advised by support to force a fresh crawl and reprocess cycle. 🛠️ What to Do If your site data appears outdated or recent changes are not showing up in your reports, you can request a recrawl and manual reprocess through the platform. The exact controls for triggering a recrawl and reprocess vary depending on your project setup and the tools you have access to. To get started, navigate to the relevant project or website in Search Atlas. From there, look for crawl or audit-related options within your project area. If you are unsure where to find these controls or if you cannot locate an option to trigger a recrawl or reprocess, a member of our support team can initiate a backend reprocess on your behalf. ✅ What to Have Ready When You Reach Out To help our team resolve this as quickly as possible, please have the following information ready before contacting support: - The name of the project or domain you need recrawled or reprocessed - A description of what data appears stale or incorrect (e.g., page counts, scores, issue lists) - Approximately when you last saw the data update correctly - Any recent changes you made to your site that should now be reflected If you need further assistance, open the chat widget in the bottom-right corner of the platform and type human teammate to be connected with a member of our team.
⏳ Why Recrawled Sites Take Time to Appear in Active Projects
🔍 Overview After triggering a recrawl for a site, you may notice that it does not immediately appear — or reappear — under the Active tab in OTTO SEO → All Sites (SEO Automation). This is expected behavior. Search Atlas requires some time to fully process a recrawl and reflect the updated site data in your Active Projects list. This article explains why the delay happens and what to expect while the process completes. ⚙️ Why the Delay Happens When you trigger a recrawl, Search Atlas begins collecting and indexing fresh data from your site. Until this background processing is complete, the site will not appear in the Active tab under OTTO SEO → All Sites (SEO Automation). The time required can vary depending on your site's size and complexity. ✅ What to Do While You Wait There are a few things to keep in mind while the recrawl is being processed: - The delay is tied to backend data processing, not a display issue. Refreshing the page or logging out and back in will not speed up the process. - Check back after giving the process adequate time to complete. The site should appear in the Active tab under OTTO SEO → All Sites (SEO Automation) once processing finishes. - If the site still has not appeared after an extended wait, note the exact site URL and the approximate time you triggered the recrawl, as this information will help our team investigate. 🛠️ When to Contact Support If your recrawled site has not appeared in the Active tab under OTTO SEO → All Sites (SEO Automation) after an extended wait, please reach out to our support team. To help us investigate as quickly as possible, have the following ready: - The exact site URL you recrawled - The approximate date and time you triggered the recrawl - A description of what you currently see in OTTO SEO → All Sites (SEO Automation) If you need further assistance, open the chat widget in the bottom-right corner of the platform and type human teammate to be connected with a member of our team.
⚡ JavaScript Rendering for LLM Crawler Visibility
Why Wix Sites Can Be Invisible to LLM Crawlers Wix and other modern website builders render page content client-side using JavaScript. This means your text, headings, and body copy are not present in the raw HTML — they are injected into the page only after a browser executes JavaScript. LLM crawlers such as ChatGPT, Perplexity, and Google's AI Overviews often fetch pages without fully executing JavaScript. If your Wix site relies on client-side rendering, these crawlers may see a nearly empty page and miss all of your content entirely. The solution is to enable JavaScript rendering inside Search Atlas's Site Auditor so the crawler processes your pages the same way a browser would — executing JavaScript before reading the content. Enabling JavaScript Rendering in Site Auditor To enable JavaScript rendering for your Wix site, follow these steps: 1. Open Site Auditor in your Search Atlas dashboard and select the project for your Wix site. 2. Before starting or restarting a crawl, locate the crawl settings for that project. 3. Find the JavaScript rendering option within the crawl settings and enable it. 4. Save your settings and launch a new crawl. Once JavaScript rendering is active, Site Auditor will fully render each page — including all JavaScript-injected content — before analysing it. This mirrors how a real browser loads your Wix site and ensures body content, headings, and metadata are visible in the audit results. If you are unable to locate the JavaScript rendering setting in your account, please contact our support team via the chat icon in your Search Atlas dashboard and an agent will guide you through enabling it for your specific project. Confirming Your Content Is Now Visible After the crawl completes with JavaScript rendering enabled, review the audit results to confirm that your page body text, meta titles, descriptions, and heading tags are now correctly detected. If pages that previously showed no content now display full text, the JavaScript rendering configuration is working correctly. Troubleshooting When Results Are Still Incomplete If the crawl still returns incomplete results after enabling JavaScript rendering, consider the following: - Crawl was not restarted after the change: The setting only applies to new crawls. Make sure you launched a fresh crawl after saving the JavaScript rendering configuration. - Wix site may be blocking the crawler: Wix platforms can block unrecognised crawlers through firewall or rate-limiting rules. Check your Wix site's security or access settings for any bot-protection rules that could be rejecting the crawler, and consult your Wix documentation for guidance on allowing specific crawlers. - Still seeing empty content: Reach out to our support team with your project details so we can investigate further. Use the chat widget in the bottom-right corner of the platform and type human teammate to be connected with a member of our team.
⚙️ Configure Site Auditor Max Page Limit and Crawl Depth
Overview The Site Auditor in Search Atlas gives you control over how much of your website is crawled during each audit. Two key settings shape this behavior: Max Page Limit and Max Crawl Depth. Configuring these correctly helps you focus the audit on the pages that matter most and keep crawl times manageable. What These Settings Mean - Max Page Limit: The maximum number of pages the crawler will visit during a single audit. Once this limit is reached, the crawl stops — even if more pages exist on the site. - Max Crawl Depth: How many links deep the crawler will follow from your start URL. A lower depth value limits the crawl to pages closer to the root; a higher value allows the crawler to follow links further into the site structure. Setting these values thoughtfully ensures your audit is thorough without consuming unnecessary crawl credits or time. Setting Max Page Limit and Crawl Depth Both Max Page Limit and Max Crawl Depth can be configured within the Site Auditor when setting up a new project or updating an existing one. Look for these fields in the crawl configuration area of the Site Auditor project setup or project settings. Enter your desired numeric values for each setting and save your changes before running or re-running the audit so the updated settings take effect. If you are unsure where to locate these fields in the current version of the platform, or if the options do not appear as expected, please reach out to our support team for guided assistance. When escalating, have the following ready: your project name or URL, the values you are trying to set, and any error message or unexpected behavior you observed. If you need further assistance, open the chat widget in the bottom-right corner of the platform and type human teammate to be connected with a member of our team.
🔄 Understanding Automatic and Manual Site Recrawls
🗓️ How the Weekly Automatic Recrawl Works Search Atlas automatically recrawls your domain once per week. This recrawl refreshes your site audit data, updates page-level insights, and ensures your SEO recommendations reflect the current state of your website — no action required on your part. The weekly recrawl runs as part of a scheduled batch process. You do not need to manually trigger this recrawl — it happens automatically in the background, and updated results will appear in your dashboard after the process completes. 🛠️ How to Manually Reprocess Your Domain If you need fresher data before the next scheduled recrawl — for example, after publishing a batch of pages or implementing technical fixes — you can trigger a manual reprocess at any time. 1. Log in to your Search Atlas account. 2. Navigate to the Site Audit section from the left-hand menu. 3. Select the domain you want to reprocess. 4. Locate and trigger the manual reprocess option within the audit dashboard. 5. Wait for the crawl to complete. Depending on the size of your site, this may take a few minutes to several hours. Once the manual reprocess is complete, your audit scores, issue lists, and page-level data will all reflect the latest crawl results. ⚠️ Known Limitations to Be Aware Of In some cases, you may notice that certain target pages were not crawled or that suggestions were not generated after a reprocess. This can happen when: - Pages are newly added and have not yet been indexed by the crawler. - A reprocess was triggered very shortly after a previous crawl completed. If your manual reprocess appears to be stuck or if suggestions are still missing after the crawl finishes, please refer to the troubleshooting steps below. 🔍 Troubleshooting a Stuck or Incomplete Reprocess If your reprocess does not complete as expected, try the following steps before reaching out for support: - Refresh your browser and return to the Site Audit dashboard to check whether the crawl status has updated. - Wait 30 minutes — large sites can take time to fully process after a crawl completes. - Trigger a second reprocess — occasionally, initiating a fresh crawl resolves an incomplete previous run. - Confirm that the domain you selected is the correct and verified domain in your account settings. If none of the above steps resolve the issue, our support team can investigate directly and escalate if needed. If you need further assistance, open the chat widget in the bottom-right corner of the platform and type human teammate to be connected with a member of our team.
🚀 How to recrawl an existing site in Search Atlas (Site Auditor)
If you've made changes to your website and want Search Atlas to scan it again, you can easily launch a new crawl using the Recrawl option inside Site Auditor. Running a recrawl helps refresh your audit data so you can review updated insights, monitor site health, and confirm whether recent site changes are being reflected in your audit results. 🧭 Where to Find the Recrawl Option All Audits list showing site cards To recrawl an existing site: go to the left sidebar and click OTTO SEO, then under Site Audit click All Audits, and find your site in the list. Each site appears as a card with key metrics and a View Audit button. 🛠️ How to Start a Recrawl Recrawl option in the site card dropdown Once you're on the OTTO SEO → Site Audit → Overview (Website Overview) page, click the Recrawl Site button directly. This immediately starts a new crawl for the selected site. 🧩 You Can Also Recrawl from the Audit Overview Recrawl Site button on the Audit Overview page If you click View Audit, the audit Overview page opens. From there, you can also find the Recrawl Site option in the top area of the audit view — handy if you're already reviewing your site audit and want to relaunch the crawl without returning to the All Audits page. 📈 What Happens After You Click Recrawl Once the recrawl begins, Search Atlas starts a new crawl for that site. You can remain on the page and monitor the crawl as it progresses — the site card updates its crawl status, and during an active crawl you may see Live Crawl Logs appear on the card. ⚙️ Optional but Recommended: Update Crawl Settings First Crawl Settings menu option Crawl Settings panel with Max Pages and Crawl Frequency Before launching a recrawl, you may want to adjust your crawl settings to better match your current audit needs. Click the site's three-dot menu and select Crawl Settings. From there, you can configure Max. Pages to Crawl (for example, 100 pages) and Crawl Frequency (for example, Weekly). Click Show more to reveal advanced settings such as User Agent, an Ignore robots.txt toggle, a JS Rendering toggle, a Crawl Speed slider, and URL Exclusion conditions. After making your changes, click Update, return to the three-dot menu, and click Recrawl to launch the new crawl with your updated settings. Use Recrawl whenever you want Search Atlas to refresh the audit for an existing site and reflect the latest version of your website. If you want even better results, review your Crawl Settings first so the next crawl runs with the exact scope and behavior you need. If you need further assistance, open the chat widget in the bottom-right corner of the platform and type human teammate to be connected with a member of our team.
🔄 Why SEO Tasks Stay Empty After Recrawls
🧭 Overview After deploying WordPress updates and triggering a recrawl in Search Atlas, you may notice that your SEO task lists appear empty or unchanged. This is expected behavior — and understanding why will save you from unnecessary troubleshooting. The key distinction is this: recrawling detects your site's current state, but task generation depends on OTTO's separate algorithmic analysis cycle. These are two different processes that run on different timelines. 🔍 Recrawls vs. OTTO Analysis Cycles Many users assume that triggering a recrawl will immediately produce new or updated SEO tasks. Here is what actually happens: - Recrawl: Scans your website to detect its current technical state — page structure, metadata, content, and any changes you have made. It updates OTTO's awareness of your site. - OTTO Analysis Cycle: A separate algorithmic process that evaluates your site's SEO performance, compares it against best practices and ranking signals, and generates actionable tasks. This cycle does not run automatically every time you recrawl. A recrawl feeds new data into OTTO, but task generation only happens when OTTO's next analysis cycle runs. Triggering a recrawl does not trigger a new task generation cycle. 📅 How Often Does OTTO Generate New Tasks? OTTO's task refresh schedule is tied to SEO meta and best-practice changes and Google indexing periods, not to a fixed weekly schedule or your recrawl frequency. In practice, this means task updates occur approximately monthly or once per indexing period. This cycle exists because meaningful SEO improvements require time to be indexed and evaluated by Google. Generating new tasks before your previous fixes have been indexed would produce inaccurate or redundant recommendations. - Task updates reflect actual changes in your site's indexed state — not just changes to your site files. - Running recrawls more frequently will not accelerate the task generation cycle. - OTTO will not surface new tasks simply because you deployed updates, even significant ones. ✅ OTTO Fixes Persist Until Marked Outdated When OTTO applies or marks a fix as resolved, that fix remains in its completed state until the algorithm determines it has become outdated. This means: - Completed tasks will not reappear after a recrawl unless OTTO's next analysis cycle identifies a regression or new issue. - You do not need to re-apply fixes after every WordPress update. - Fixes are durable across recrawls unless your site changes cause OTTO to flag them again in a future cycle. 🟦 Why WordPress Sites Often Show Blank Task Sections If you are managing a WordPress site and certain task sections appear completely empty, this is normal and expected. WordPress sites commonly have configurations — such as page builders, caching plugins, or crawl restrictions — that can act as site blockers during OTTO's scan. Blank task sections on WordPress sites typically indicate one of the following: - OTTO has not yet completed its next analysis cycle since your updates were deployed. - The relevant pages have not yet been re-indexed by Google since your fixes were applied. - A site blocker or plugin is limiting what OTTO can evaluate in that section. An empty task section does not mean something is broken. It means OTTO has no new recommendations to surface at this time for that category. 📋 Expected Workflow After Deploying WordPress Updates Follow this sequence to work effectively with OTTO's task cycle rather than against it: 1. Deploy your WordPress fixes or updates as normal through your CMS or hosting environment. 2. Trigger a recrawl inside Search Atlas by navigating to Left sidebar → OTTO SEO → All Sites (SEO Automation), selecting your active site, then going to Site Audit → Overview (Website Overview) and clicking the Recrawl Site button. This ensures OTTO has the latest snapshot of your site. 3. Allow time for Google to re-index your changes. This typically takes several days to a few weeks depending on your site's crawl budget and the scope of changes. 4. Wait for OTTO's next algorithmic analysis cycle — approximately monthly or aligned with the current indexing period. New tasks will appear in your dashboard once this cycle runs and detects actionable opportunities. 5. Review updated tasks in the OTTO SEO section after the cycle completes. You will see refreshed recommendations based on your site's newly indexed state. Do not interpret an empty task list immediately after a recrawl as an error. If tasks were present before your updates and are now resolved, their absence confirms OTTO registered your fixes successfully. ⚠️ When to Investigate Further While empty tasks after a recrawl are normal, there are situations where you should take a closer look: - Tasks that were active before your updates have disappeared but you did not apply any fixes — this may indicate a crawl issue. - Your site has been re-indexed by Google but no new OTTO cycle has run after several weeks. - Task sections remain blank across multiple OTTO cycles with no explanation from your site's configuration. In these cases, verify that your WordPress site does not have crawl restrictions (such as disallow rules in robots.txt or password-protected staging environments) that could prevent OTTO from analyzing your pages. 💬 Need Help? If you need further assistance, open the chat widget in the bottom-right corner of the platform and type human teammate to be connected with a member of our team.
🔍 How to Exclude Pages from OTTO Crawls and Manage SEO Health Score Volatility
🧭 Overview When OTTO crawls your website, it scans every accessible page by default — including template pages, thank-you pages, tag archives, and other low-value URLs that you may not want audited. This can lead to two common frustrations: unexpected credit consumption and SEO Health Score volatility caused by issues surfacing on pages you never intended to optimize. This article explains how page exclusions work, how to configure them correctly, and what to expect from your SEO Health Score after making fixes. 🤔 Why Does My SEO Health Score Show More Issues After Fixing Them? Your SEO Health Score is recalculated every time OTTO runs a new crawl. If your score drops or new issues appear after a fix, there are a few common reasons: - New pages were discovered: Each crawl may find additional URLs — such as paginated pages, filtered product pages, or auto-generated tag archives — that were not part of the previous scan. These newly discovered pages contribute new issues to your score. - Template-level issues propagate widely: A single issue on a template (for example, a missing meta description on a category page template) can generate hundreds of individual issue instances, one per page that uses that template. Fixing one instance does not resolve the others. - Recrawl after deployment: If you published fixes and triggered a new scan before the changes fully propagated (such as a caching delay), OTTO may still detect the old state on some pages. - Dynamic content changes: Pages with dynamic titles, descriptions, or structured data can shift between crawls if the underlying content or database values changed. The most effective way to stabilize your SEO Health Score is to exclude pages that are not part of your core SEO strategy before running audits, so the score only reflects URLs you actively want to optimize. 📋 What Types of Pages Should Be Excluded? Not every page on your website deserves to be audited. Consider excluding the following page types to keep your audits focused and your credit usage efficient: - Template and archive pages: Tag archives, category index pages, author pages, and date-based archive URLs that are not individually optimized. - Utility and system pages: Thank-you pages, order confirmation pages, login and logout pages, cart and checkout pages. - Paginated URLs: Pages like /blog/page/2/ or /?page=3 that duplicate content from the primary page. - Parameter-based URLs: Filtered or sorted views such as /products?color=blue&sort=asc that are dynamically generated and not independently meaningful. - Staging or development paths: Any URL patterns that belong to test environments or internal tools not intended for public search visibility. ⚙️ How to Configure Page Exclusions in OTTO Follow these steps to exclude specific pages or URL patterns from OTTO's crawl and audits: 1. Log in to your Search Atlas account and navigate to OTTO SEO in the left-hand sidebar, then open SEO Automation. 2. Select the project (website) you want to configure. 3. Open the Crawl Settings or Site Settings for that project. 4. Locate the URL Exclusions field. This is where you enter the URL patterns you want OTTO to skip during future crawls. 5. Enter each URL pattern you want to exclude. You can use: - An exact URL path (e.g., /thank-you/) to exclude a single page. - A partial path or prefix (e.g., /tag/) to exclude all URLs that begin with or contain that string, such as all tag archive pages. - A query parameter pattern (e.g., ?page=) to exclude paginated or filtered URLs that share a common parameter. 6. Save your exclusion settings. 7. Trigger a new crawl so OTTO re-audits the site with the exclusions applied. Your SEO Health Score will be recalculated based only on the pages that remain in scope. Once exclusions are saved, OTTO will skip the matching URLs in all subsequent crawls, which reduces credit consumption and prevents low-value pages from affecting your SEO Health Score. 💡 Tips for Maintaining a Stable SEO Health Score - Set exclusions before your first full crawl on a new project whenever possible, so your baseline score is clean from the start. - Review newly discovered URLs after each crawl by checking the full URL list in your audit results. If unexpected page types appear, add their patterns to your exclusion list. - Fix template-level issues at the source — correcting a meta description or title tag in your CMS template will resolve the issue across all pages using that template in a single fix, rather than page by page. - Allow caching to clear before triggering a re-crawl after deploying fixes, so OTTO picks up the updated page state accurately. If you need further assistance, open the chat widget in the bottom-right corner of the platform and type human teammate to be connected with a member of our team. Additional Notes Rule updates — Search Atlas periodically refines its SEO audit rules, which can surface issues that were previously uncategorized. Also notes: fixing one high-impact template issue can resolve hundreds of individual instances at once.
🔍 How to Find Your Site Crawl Results
🚀 Where to Find Crawl Results Your site crawl results live inside the OTTO SEO section of Search Atlas. To access them: 1. Open the left sidebar menu in your dashboard. 2. Click on OTTO SEO → All Sites (SEO Automation). 3. Locate the site you crawled in the list. 4. Click on the site name to open its detail view. All crawl findings, recommendations, and reports appear on this page. If you have multiple sites, use the search or filter options at the top of the list to quickly locate the one you need. 📊 Checking Crawl Completion Status Each site in the OTTO SEO list shows a status that tells you where the crawl stands in its lifecycle. Check the status next to your site to determine whether the crawl has finished. If the status indicates the crawl is still in progress, results will continue updating until it completes. Once the crawl is finished, full results will be available in the site detail view. 💡 Tips for Accessing Results Quickly If you feel like you have lost your crawl results, they are almost always still in the platform. Here are a few things to check: - Verify you are in the right workspace. If your account has access to multiple teams or projects, make sure you have the correct one selected from your account menu. - Check your browser tab. If you navigated away during a crawl, the results page may have closed. Return to OTTO SEO → All Sites (SEO Automation) in the left sidebar to reopen it. - Allow time for processing. Larger sites take longer to crawl. If a site is still processing, give it time to finish before expecting full results. ⚙️ What to Do If Results Are Missing If you still cannot find your crawl results, try the following: 1. Refresh your browser page and navigate back to OTTO SEO → All Sites (SEO Automation) from the left sidebar. 2. Confirm the site was successfully submitted for a crawl. If the site does not appear in the list, the crawl may not have started. If you need further assistance, open the chat widget in the bottom-right corner of the platform and type human teammate to be connected with a member of our team.
🕷️ Crawl Monitoring
Use Crawl Monitoring to track which search engines and AI bots are crawling your site, how often they visit, and whether your pages are being discovered and indexed efficiently. 🖥️ Accessing Crawl Monitoring Go to Left Sidebar → OTTO SEO → Site Audit → Crawl Monitoring. 📊 Dashboard Components Crawler Distribution Table - Lists active crawlers (e.g., Googlebot, Bingbot, Google-Mobile) and their total request counts. - Includes interactive graphs for Historical Crawl Activity. What to do next: If a major crawler like Googlebot is missing or underrepresented, confirm your site is indexable and submit your XML sitemap in Google Search Console. Historical Crawl Activity Graph - Shows crawler activity trends over time, color-coded by crawler type. - Toggle between Daily, Weekly, or Monthly views. - Hover over data points to see detailed metrics for specific dates. What to do next: If you see a sudden drop in crawl activity, check for recent robots.txt changes, server errors, or sitemap issues in the Site Audit tool. Key Metrics - Site Indexation Percentage — Real-time representation of crawled vs. uncrawled pages. Use this to spot indexing gaps and opportunities. What to do next: If Site Indexation Percentage is below expectations, check your robots.txt and XML sitemap coverage in the Site Audit tool. - Crawl Purpose Analysis — Shows the split between: - Discovery Crawls: Search engines indexing new pages for the first time. - Refresh Crawls: Search engines revisiting existing pages for updates. What to do next: If Discovery Crawls are low after publishing new content, resubmit your sitemap and verify internal linking points to the new pages. - Device Distribution — Breaks down activity by Desktop vs. Mobile crawlers. Use this to verify mobile-first optimization coverage. What to do next: If Mobile crawler activity is low, run a mobile usability check in Site Audit to confirm responsive design and Core Web Vitals. - Crawl Frequency Metrics — View patterns across the last 7 days, 30 days, 6 months, or 1 year to identify long-term crawl trends. What to do next: If crawl frequency declines over time, audit page freshness, update key content, and check for crawl-budget waste from low-value URLs. 🤖 Crawl Analysis by Bot The Crawl analysis by bot section gives you a visual breakdown of crawler activity per bot type, so you can see exactly which search engines and AI bots are spending time on your site. - Donut chart — Shows each bot's proportional share of total crawl requests. Hover over a segment to see the exact request count and percentage for that crawler. - Bot selector dropdown — Scroll through the dropdown to view every available crawler (e.g., Googlebot, Bingbot, GPTBot, ClaudeBot, PerplexityBot). Select a bot to filter the donut chart and associated metrics to that crawler only. - Legend — Lists each crawler by name with a matching color swatch. Use the color coding to map chart segments to specific bots. To compare crawlers, switch between bots in the dropdown — the chart and legend update instantly. What to do next: If an important crawler (such as Googlebot or a target AI bot) shows an unexpectedly low share, verify the bot is not blocked in robots.txt and review your Site Audit report for crawl errors. If AI bots are missing entirely, confirm your site allows their user agents. 📡 AI Bot Activity and Data Tracked Crawl Monitoring tracks how search engines and AI bots interact with your website. Available data includes: - Total crawl requests - Average response time - Total download size - Crawler distribution - Historical crawl activity - Site indexation percentage - Crawl purpose - Device / screen size - Load time distribution - Crawl frequency - Crawled pages ❓ Why Crawl Monitoring Shows "No Data" A No data message does not always indicate a problem. Data may not appear if: - The domain is new or has been live for less than approximately 6 months. - The domain has not yet had enough crawl activity. - The site has limited visibility or low crawler traffic. - Your account plan does not include Crawl Monitoring access (Pro or higher required). - Search engines or AI bots have not crawled the site during the selected date range. ✅ What to Check Before Reporting an Issue 1. Confirm the selected Date Range covers a period with expected crawl activity. 2. Switch between Daily, Weekly, and Monthly views to see if data appears at a different granularity. 3. Check whether the domain is new or recently launched. 4. Confirm the site is indexable and not blocked by robots.txt. 5. Confirm your account plan includes Crawl Monitoring access (Pro or higher). 6. Allow more time for crawler activity to accumulate — data builds as bots visit the site. 🎯 You now have a complete view of how search engines and AI bots interact with your site — from high-level crawler distribution to per-bot visual analysis in the Crawl analysis by bot section. To act on what you discover, explore the Site Audit tool in OTTO SEO to resolve crawl gaps and indexing issues.
⚙️ Whitelist Search Atlas Crawler IPs in Your CDN or Firewall
If your site is protected by a CDN, firewall, or security service, Search Atlas crawlers may be blocked from accessing your website. When this happens, Site Audit and other crawling features may fail to collect data or report incomplete results. This article explains how to allow Search Atlas crawlers through your security layer. 🤔 Why This Happens Many Content Delivery Networks (CDNs), Web Application Firewalls (WAFs), and hosting security providers automatically block unknown bots or limit requests from crawler IP addresses. If Search Atlas crawler IPs are blocked, you may experience: - ❌ Site Audits that fail or never complete. - ❌ Missing crawl data. - ❌ Incomplete site health reports. - ❌ Pages not being discovered during scheduled crawls. ✅ Solution Allow all official Search Atlas crawler IP addresses in your CDN, firewall, or hosting provider. You can always find the most up-to-date IP list here: 👉 https://bots.searchatlas.com/ip_list.txt ⚠️ Important Do not manually maintain your own IP list. Always use the official Search Atlas list, as crawler IPs may be added or updated over time. ☁️ If You're Using Cloudflare Cloudflare makes this even easier. 1. Log in to your Cloudflare dashboard. 2. Open your website. 3. Navigate to your Bot Management or Security settings. 4. Select Search Atlas from the Verified Bots list. 5. Save your changes. 📘 If needed, refer to Cloudflare's documentation for additional information about Verified Bots. 🛡️ Other CDNs or Firewalls If you're using another provider (such as Akamai, Fastly, Imperva, Sucuri, AWS WAF, or another security platform): 1. Open your firewall or bot management settings. 2. Whitelist all IP addresses from the official Search Atlas IP list. 3. Save your configuration. 4. Run your Site Audit again. 🔍 Verify the Fix After updating your security settings: 1. Open Site Audit in Search Atlas. 2. Start a new crawl. 3. Wait for the crawl to begin. ✅ If the crawl starts successfully, your whitelist has been configured correctly. 🚨 Still Having Issues? If you've already whitelisted the official Search Atlas IP addresses but crawls are still failing: - ✔️ Verify that your CDN changes have fully propagated. - ✔️ Check whether another firewall, hosting provider, or security plugin is blocking requests. - ✔️ Confirm you're using the latest IP list from the official Search Atlas source. - ✔️ Contact Search Atlas Support if the issue persists. 📋 Quick Summary | ⚠️ Problem | Your CDN or firewall is blocking Search Atlas crawlers. | | --- | --- | | ✅ Solution | Whitelist the official Search Atlas crawler IPs or enable Search Atlas as a Verified Bot in Cloudflare. | | --- | --- | | 🔗 Official IP List | https://bots.searchatlas.com/ip_list.txt | | --- | --- |
What Factors Affect Your Website's Crawl Budget?
⚙️ What Factors Affect Your Website's Crawl Budget? Google does not crawl all websites equally. The amount of crawling attention Googlebot dedicates to your site — your crawl budget — is shaped by a combination of technical signals, authority metrics, and site health indicators. Understanding these factors helps you make informed decisions about how to improve your site's crawlability and search performance. 🔎 📐 The Two Core Components of Crawl Budget According to Google's official documentation, crawl budget is determined primarily by two interacting components: crawl rate limit and crawl demand. These two factors work together to define how frequently and how aggressively Googlebot will explore your site. ⚡ Crawl Rate Limit The crawl rate limit defines how fast Googlebot can request pages from your server without causing performance issues or overloading your infrastructure. Think of it as the speed governor on Googlebot's crawling engine — it prevents the crawler from sending so many simultaneous requests that your server slows down or crashes. ✅ Server response time is one of the most direct influences on crawl rate limit. If your server responds to requests quickly and consistently, Googlebot can safely crawl more pages per unit of time. ✅ Host load also matters — if your server is under heavy traffic from real users, Googlebot will back off to avoid making the situation worse. ✅ Server errors (5xx responses) are a strong negative signal. When Googlebot encounters repeated server errors, it reduces its crawl rate to protect your site, which means fewer pages are crawled overall. ⚡ You can influence the crawl rate limit directly through Google Search Console (Settings → Crawl Rate) by requesting that Google crawl your site more slowly or more quickly within a 90-day window. However, the most sustainable improvement comes from improving actual server performance. 📈 Crawl Demand Crawl demand reflects how much Google wants to crawl your URLs — independent of your server's capacity. Even if your server could handle millions of requests per day, Google will only invest crawl capacity in pages it considers worth revisiting. Crawl demand is driven by several factors: ✅ Link authority and PageRank — Pages with more high-quality inbound links signal greater importance to Google. Higher PageRank translates directly into higher crawl demand, as Google wants to keep authoritative pages fresh in its index. ✅ Content freshness and staleness — Pages that are updated frequently (news articles, product listings, blog posts) generate higher crawl demand because Google wants to capture the latest version. Pages that never change may be crawled less often over time. ✅ Overall site popularity — Websites that attract significant organic traffic, earn natural backlinks, and demonstrate strong engagement signals tend to receive more generous crawl budgets because Google recognizes them as valuable resources worth indexing thoroughly. ⚡ Key insight: You cannot directly instruct Google to increase crawl demand — it must be earned through building genuine authority, earning quality backlinks, and publishing content that is regularly updated and valuable to users. 🏆 🏗️ Site Size and Architecture The sheer size and structural complexity of your website has a direct impact on how efficiently Googlebot can spend its allocated crawl budget. ✅ Large websites with hundreds of thousands of URLs naturally take longer to crawl in full. If your crawl budget is limited relative to your URL count, some pages will inevitably be crawled less frequently or not at all. ✅ Deep site architecture — where important pages are buried many clicks away from the homepage — makes it harder for Googlebot to discover and prioritize content efficiently. Flatter architectures that keep key pages within three clicks of the homepage perform better. ✅ Orphan pages with no internal links pointing to them may never be crawled, regardless of their content quality, simply because Googlebot has no path to reach them. 🚨 Site Health Factors That Drain Crawl Budget Beyond the two core components, several common site health issues can silently consume large portions of your crawl budget without contributing any SEO value. ✅ 4xx error pages (page not found) waste a crawl request every time Googlebot tries to access a broken URL. At scale, a large number of broken URLs can meaningfully reduce the crawl capacity available for your live, indexable content. ✅ 5xx server errors not only waste individual crawl requests but actively signal to Googlebot that your server is unreliable, causing it to throttle its crawl rate further. ✅ Redirect chains — sequences of multiple redirects between an original URL and its final destination — force Googlebot to spend multiple crawl requests to reach a single piece of content. A single 301 redirect is acceptable, but chains of two or more hops should be flattened. ✅ Slow page load times reduce how many pages Googlebot can fetch in a given crawl window. If every page takes three seconds to load, Googlebot will crawl far fewer pages than it would on a server delivering sub-second responses. ✅ Duplicate content across multiple URLs forces Googlebot to spend crawl budget processing the same information repeatedly. Using canonical tags and avoiding unnecessary URL parameter variations helps consolidate crawl signals efficiently. 🎯 🤔 Does Crawl Budget Matter for Every Website? Google has noted that crawl budget is most critical for large websites — typically those with hundreds of thousands of URLs or more — as well as sites that update content very frequently. For smaller websites with a few hundred pages and a healthy site structure, crawl budget is rarely a limiting factor. ✅ If your site has fewer than a few thousand pages and all important content is being indexed regularly, crawl budget optimization is unlikely to be a priority concern. ✅ If your site has large-scale content (e-commerce catalogs, news archives, user-generated content), crawl budget management becomes a critical technical SEO discipline that directly impacts how much of your content appears in search results. 🔍 Bottom line: Crawl budget is shaped by factors you can control (server performance, site architecture, redirect efficiency, content quality) and factors you earn over time (authority and popularity). Addressing the controllable factors first creates the foundation for a more efficient, scalable crawl. ⚡
📉 Why Your Crawl Distribution Dropped Month-Over-Month
📊 What Crawl Distribution Means Crawl distribution shows how many of your site's URLs were crawled and processed within a given time period. The number you see for a month reflects activity inside that specific date range. Because it is tied to a window of time, the figure naturally rises and falls as crawl activity changes from one month to the next. A drop, such as going from 7,000 in May to 1,000 in June, does not automatically mean something is broken. In most cases there is a clear, expected reason. This article walks you through the common causes and how to check each one. 🗓️ Reason 1: The Month Is Still In Progress This is the most common cause. If you compare a completed month against one that is only partway through, the current month will always look smaller. May had a full set of days to accumulate crawls, while June may still be gathering data. - A figure of 1,000 mid-June compared to 7,000 for all of May is usually on track. - Wait until the month closes before comparing totals side by side. 🔁 Reason 2: Crawl Frequency and Scheduling Crawls run on a schedule rather than continuously. The number of crawls in a period depends on how often your project is set to crawl and when the last crawl completed. If a large crawl finished late in May, that volume is counted in May, leaving June lower until the next crawl runs. - Check when your most recent crawl completed. - Confirm whether a scheduled crawl is still pending for the current month. ⚙️ Reason 3: Project or Setting Changes Changes to your project configuration can affect how many URLs are crawled. Review whether any of the following changed recently: - Crawl limits were lowered, reducing the number of URLs processed. - URL scope or filters were narrowed, excluding pages that were previously included. - Sitemaps or seed URLs were updated or removed. - The project was paused or its schedule was adjusted. 📦 Reason 4: Your Monthly Crawled Pages Budget Was Reached Your plan includes a Number of Monthly Crawled Pages quota. Once the pages crawled in a month reach this budget, further crawling for that period stops until the quota resets at the start of the next month. If May consumed most of the budget late in the month, June can start lower until the allowance refreshes. - Look for the Crawl Budget Depletion Warning in the platform, which indicates the monthly page budget has been used up. - Compare the pages crawled against your plan's monthly crawled pages allowance to see if the quota was the limiting factor. - Note that the depletion warning appears once the budget is already used up. A proactive warning that flags when your crawl settings would exceed the monthly page budget before a crawl runs is currently in development and not yet available. 🌐 Reason 5: Changes On Your Website Sometimes the drop reflects real changes on your site rather than the platform. Consider whether any of these apply: - Pages were removed, redirected, or returned errors, leaving fewer URLs to crawl. - Your robots.txt file or meta directives now block sections of the site. - Server slowdowns or downtime prevented pages from being reached during the crawl. 🔍 How to Check Your Crawl Data Follow these steps to confirm what is happening: 1. Open the affected project in your dashboard. 2. Set the date range to a full, completed month so you compare equal time periods. 3. Review the crawl history to see when the last crawl ran and how many URLs it covered. 4. Check your crawl settings for recent changes to limits, scope, or scheduling. 5. Check whether your monthly crawled pages budget was reached, including any crawl budget depletion warning. 6. Verify your robots.txt and sitemap are accessible and unchanged. ✅ When the Drop Is Expected A lower number is normal when the current month is incomplete, when a large crawl landed in the prior month, when your monthly crawled pages budget was reached, or when you intentionally reduced crawl scope or limits. In these cases no action is needed, and numbers typically even out once the month completes, the quota resets, and the next crawl runs.
Crawl Budget
A crawl budget may seem like a foreign concept when you're first learning about how search engine bots work. While it's not the simplest SEO topic, it's less complicated than it appears. Once you understand how search engine crawling functions, you can begin to optimize your website for crawlability — helping your site reach its highest potential in Google's search results. 🔎 What Is a Crawl Budget? A crawl budget is the number of URLs Googlebot will crawl and process on your site within a given timeframe. Crawling is a prerequisite to (but not the same as) indexing — Google must first crawl a page before it can decide whether to index it, so the two processes are related but distinct. Rather than happening in a single discrete session, Google allocates crawl capacity continuously over time based on two main factors: the crawl rate limit (how fast Googlebot can request pages without overloading your server) and crawl demand (how often Google wants to revisit your URLs based on popularity, freshness, and authority). Other search engines such as Bing have analogous crawl budget concepts. ⚙️ What Factors Affect a Website's Crawl Budget? Google doesn't crawl all websites equally. According to Google's official documentation, crawl budget is shaped by two core components — crawl rate limit and crawl demand — along with the overall health and size of the site: - Crawl Rate Limit: How fast Googlebot can crawl your site without overloading the server. Server response time, host load, and server errors all influence this limit. - Crawl Demand: How much Google wants to crawl your URLs, based on link authority/PageRank, content freshness/staleness, and overall popularity. Pages with more inbound links and fresher content earn higher demand. - Site Size & Health: Large or complex websites naturally take longer to crawl. Crawlers also spend time budget processing 4xx/5xx error pages and following redirects, which consumes capacity that could otherwise go to indexable content. A single 301 redirect is generally acceptable, but redirect chains (multiple hops between the original and final URL) waste crawl budget and should be flattened. Slow load times further reduce how many pages Googlebot can fetch in a given crawl window. 📈 How Does Your Crawl Budget Affect SEO? If Googlebot can't find or index your content, your site won't appear in search results — resulting in lost search traffic. 🤖 Why Does Google Crawl Websites? Googlebot systematically explores a site's pages to understand their content and relevance. It categorizes and stores this information to decide which results appear (and in what order) in search results. 🧠 What Happens During a Crawl? Googlebot crawls sites within a limited time window. It prioritizes URLs based on robots.txt instructions and page importance. During a crawl, Google analyzes: - Meta tags and page meaning - Internal links and anchor text - Media files (for image/video search) - Schema and HTML markup Duplicate or canonicalized content gets lower crawl priority. ⏱️ Crawl Rate vs. Crawl Demand - Crawl Rate: How quickly Google crawls individual pages during a session. - Crawl Demand: How often Google returns to crawl your site based on its popularity and content updates. You can analyze crawl frequency via log file analysis. 🔍 How Can I Determine My Site's Crawl Budget? Since Google doesn't share exact crawl budget numbers, you can estimate it: 1. Get your site's total URL count (via sitemap or Yoast). 2. In Google Search Console, go to Settings > Crawl stats to see how many pages are crawled daily. 3. Divide total URLs by the average crawls per day. If the ratio is below 10, your crawl budget is healthy. Otherwise, consider optimization. 🚀 How Can You Optimize for Your Crawl Budget? When your site outgrows its crawl budget, focus on what you can control. Follow these best practices in order: 1️⃣ Increase Your Crawl Rate Limit 1. In Google Search Console, go to Settings to review crawl rate. 2. Increase the crawl limit for 90 days if needed. 2️⃣ Perform a Log File Analysis Request a server log file to analyze: - Crawl frequency - Top crawled pages - Unresponsive or missing URLs 3️⃣ Keep XML Sitemap and Robots.txt Updated 1. Ensure your sitemap lists only important URLs. 2. Use noindex tags in robots.txt for pages you don't want crawled. 4️⃣ Reduce Redirects & Redirect Chains Redirects (3xx codes) slow crawling. Minimize redirect chains to improve efficiency. 5️⃣ Fix Broken Links Update internal links that lead to 404 pages. Use Search Console → Index > Coverage report or the Site Audit tool to find broken links. 6️⃣ Improve Page Load Speeds Slow pages waste crawl time. Use PageSpeed Insights and follow Core Web Vitals guidelines. If needed, upgrade server resources (RAM, hardware, or hosting). 7️⃣ Use Canonical Tags Canonical tags prevent duplicate content from consuming crawl time. 8️⃣ Strengthen Internal Linking A solid internal link structure helps crawlers find important pages quickly. 9️⃣ Prune Unnecessary Content Remove outdated or low-traffic pages. Always redirect deleted URLs to relevant pages. 🔟 Accrue More Backlinks External links help Google discover your pages faster and crawl more often. 1️⃣1️⃣ Eliminate Orphan Pages Pages not linked from anywhere on your site can go undiscovered. Link them internally or intentionally keep them unlinked if they serve a limited purpose (e.g., campaign landing pages). 📜 Using Crawl Directives Strategically Crawl directives tell Googlebot which URLs to fetch, index, or consolidate — and they're some of the most powerful levers you have for preserving crawl budget: - robots.txt disallow rules: Block Googlebot from crawling low-value sections (faceted search URLs, internal search results, admin paths, infinite calendar pages). Disallowed URLs aren't crawled, so they don't consume budget. - noindex meta tags: Allow crawling but keep a page out of the index. Useful for thank-you pages, thin tag archives, or staging content. Note: noindex still costs crawl budget because Googlebot must fetch the page to see the directive. - Canonical tags: Consolidate duplicate or near-duplicate URLs (e.g., URL parameters, print views, sorted listings) into a single canonical version, so crawl demand concentrates on the version you want ranked. - XML sitemaps: List only canonical, indexable URLs so Googlebot can prioritize them. Keep sitemaps current — stale entries dilute crawl signals. 🧩 The Best Tools for Crawl Budget Optimization 1. Google Search Console — Track crawl stats and request indexing. 2. Google Analytics — Monitor internal link performance. 3. Site Audit Tools (Dashboard) — Identify crawl issues, index depth, duplicate content, and page speed. 🛠️ How Search Atlas Helps with Crawl Budget Search Atlas gives you two purpose-built tools for monitoring and improving crawl efficiency: - Site Audit surfaces the exact issues that drain crawl budget — broken links and 404 errors, redirect chains, slow-loading pages, duplicate content, orphan pages, and indexability problems. You can audit sites of up to 50,000 pages, making it suitable for large catalogs and complex content libraries. - OTTO can automate many of the fixes Site Audit identifies — applying canonical tags, resolving redirect chains, updating meta directives, and pushing changes live without manual developer work. OTTO supports projects up to the same 50,000-page ceiling. Used together, Site Audit pinpoints what's wasting Googlebot's time on your site, and OTTO helps you fix it at scale — freeing up crawl budget for the pages that actually drive search traffic. While you can't control how often search engines crawl your site, you can optimize your crawl efficiency. Start by reviewing your server logs and Search Console crawl stats, then fix crawl errors, redirects, and site speed issues. Keep refining your link structure, content quality, and technical SEO to boost your rankings over time.
📘 Why Am I Seeing Pages I Don't Recognize Being Crawled on My Site?
If you've noticed a large number of unfamiliar pages being crawled on your site—such as URLs with /author/, /category/, /tag/, or /date/—this is actually quite common. These URLs are typically auto-generated by your content management system (CMS), especially if you're using platforms like WordPress. This article will help you understand what these pages are, why they're being crawled, and how to control their visibility. 👣 Step-by-Step Instructions 1. What Are These Pages? Many CMSs, including WordPress, automatically generate archive pages you may not have manually created: - Author Archives – Lists all posts by a specific author (e.g., /author/admin/) - Date Archives – Groups content by month or year (e.g., /2023/07/) - Category Pages – Lists posts under a specific category (e.g., /category/news/) - Tag Pages – Shows content tagged with a specific keyword (e.g., /tag/seo/) These pages are public by default and can be discovered by any crawler or search engine. 2. Why Are These Pages Being Crawled? 1. Crawlers scan all publicly accessible pages. 2. If these archive pages are linked internally (e.g., sidebar, footer, sitemap), they're considered part of your site's structure. 3. Unless deliberately disabled or blocked, they will continue to exist and be crawled. 3. How Can I Control This? 1. Disable in CMS Settings - Many platforms let you turn off author archives, date archives, or other taxonomy pages directly from site settings. For example, in WordPress, navigate to Yoast SEO > Search Appearance > Archives to disable author and date archives. Other SEO plugins such as Rank Math offer similar settings under Titles & Meta > Archives. 2. Block Crawling with robots.txt - Add rules to disallow specific paths. For example, to block author and tag archives, add the following to your robots.txt file: User-agent: * Disallow: /author/ Disallow: /tag/ Disallow: /category/ - Important: robots.txt only prevents crawling—it does not guarantee de-indexing. If the URLs are linked externally, they can still appear in search results. To prevent indexing, you'll need a noindex meta tag or HTTP header. In some cases, you may need to use both methods together depending on your goal. - Add noindex Meta Tags - If you want the pages to exist but not appear in search results, configure your site to apply a noindex tag. - This lets them be crawled but prevents indexing. 3. Audit Your XML Sitemap - Customers often inadvertently include tag, category, or author URLs in their XML sitemap, which amplifies crawl budget waste. Review your sitemap and remove auto-generated archive URLs if they aren't intended to be indexed. Most SEO plugins (Yoast SEO, Rank Math, All in One SEO) let you exclude specific taxonomies from the sitemap under their sitemap settings. 4. Using Search Atlas OTTO & Site Lens - OTTO Site Audit and Site Lens may surface these archive/auto-generated URLs as orphan pages or unindexed pages. If they're intentional, you can mark them as excluded so they don't continue to appear as issues. - If available, use OTTO's robots.txt manager to apply the disallow rules described above without editing your server files directly. - Monitor crawl coverage in Site Lens over time to confirm that excluded URLs are no longer being prioritized by crawlers. ❓ FAQs ❓ Why do these pages exist if I didn't create them? 💡 Your CMS automatically generates them as part of its default functionality. ❓ Should I block or keep these pages? ⚡ If they don't serve SEO or user navigation purposes, it's often best to block or noindex them. ❓ What happens if I disable them? ✅ They'll no longer be accessible, meaning crawlers and users won't see them. ❓ I use Landing Page Studio with a custom domain — why are new pages appearing in crawl data so quickly? 🚀 If you're using Landing Page Studio with a custom domain and have Instant Indexing enabled, newly published pages are automatically submitted to Google. Seeing them appear quickly in crawl logs is expected and intentional behavior. ❓ Site Lens still shows orphan pages as unresolved even after I've fixed them — what should I do? 🔄 If Site Lens shows orphan pages as unresolved even after Google has indexed them, this may be a known display sync issue. Try performing a manual recrawl and verify the page's indexed status directly in Google Search Console. If the issue persists, contact support. Closing Note ✅ These unfamiliar URLs are usually nothing to worry about—they're simply auto-generated by your CMS. By adjusting settings, updating your robots.txt, or applying meta tags, you can take control over whether they're crawled or indexed.
🤖 Robots.txt: Best Practices and Site Audit Warnings
The Search Atlas Site Auditor scans your robots.txt file for errors that can prevent crawlers from reaching your pages. Use this guide to understand each audit warning and how to fix it. 🤖 What Is a robots.txt File? A robots.txt file tells web crawlers which parts of your website they can or cannot access. Crawlers fetch it before visiting any other page on your site. Each file contains: - A user-agent string — the crawler's name (e.g., Googlebot, Applebot) - Directives — rules such as Allow or Disallow - Paths or URLs — the pages or directories the rule applies to Use robots.txt to prevent unready or private pages from being indexed, restrict specific crawlers, and focus your crawl budget on valuable content. 📂 Where to Place Your robots.txt File Place the file in the root directory of your website — for example, https://www.website.com/robots.txt. If a crawler cannot find a robots.txt file, it assumes all pages are crawlable. On large sites this wastes crawl budget on low-value pages, so create an explicit file even if you intend to allow everything. If you don't have a robots.txt file yet, upload one to your root directory — your hosting provider can help. 🔍 Why robots.txt Matters for SEO Robots.txt directly controls which pages crawlers visit and how efficiently they use your crawl budget. Key benefits include: - Improves crawl efficiency - Prevents indexing of low-value or duplicate pages (such as confirmation pages) - Protects confidential areas of your site - Keeps search results clean and relevant ⚙️ How robots.txt Works When a crawler visits your site, it fetches /robots.txt first and reads the directives for its user-agent. It then applies those rules to every URL it tries to visit. If no restriction is specified for a path, crawlers access and index it freely. Example: User-agent: Googlebot Disallow: /confirmation-page/ 🧠 robots.txt Best Practices - Encode the file in UTF-8 format. - Name the file exactly robots.txt (case-sensitive). - Place it in the root directory. - Maintain only one robots.txt file per (sub)domain. - Give each user-agent its own directive group. - Be specific — avoid accidentally blocking entire subdirectories. - Do not use noindex in robots.txt; configure robots meta tags on individual pages instead. - Remember: the file is publicly viewable. 🚨 robots.txt Issues Detected by the Search Atlas Site Auditor The Search Atlas Site Auditor automatically detects the following robots.txt problems. Select each issue to see the details and fix. 1️⃣ robots.txt Not Present Issue: No robots.txt file exists in the root directory. Fix: Create a robots.txt file and upload it to https://www.website.com/robots.txt. The warning clears on your next Site Auditor crawl. 2️⃣ robots.txt on a Non-Canonical Domain Variant Issue: Multiple robots.txt files exist across www/non-www or HTTP/HTTPS versions of your domain. Fix: 1. Keep only one canonical robots.txt at your preferred domain (e.g., https://www.website.com/robots.txt). 2. Set up 301 redirects from all other domain variants to point to the canonical version. 3️⃣ Invalid Directives or Syntax Issue: Incorrect syntax prevents crawlers from following your rules. Fix: Open your robots.txt file and correct the errors. The Site Auditor lists the specific invalid directives it detected so you know exactly what to address. 4️⃣ robots.txt Should Reference an Accessible Sitemap Issue: Your robots.txt file does not include a reference to your XML sitemap. Fix: Add this line at the end of your robots.txt file: Sitemap: https://www.website.com/sitemap.xml This helps crawlers discover and prioritize your pages faster. 5️⃣ Disallow: / Blocks All Crawlers from Your Entire Site Issue: Your robots.txt contains Disallow: / under a user-agent directive, which blocks that crawler from accessing every page on your site. When applied to User-agent: *, all search engines are blocked — which can completely prevent your site from appearing in search results. This is one of the most critical robots.txt misconfigurations. ⚠️ WordPress caveat: A regex detection flaw (tracked as WP-322) currently prevents this warning from triggering reliably on WordPress sites in the Site Auditor. The fix is awaiting release. Until it ships, manually verify your robots.txt file for a Disallow: / line under User-agent: * rather than relying solely on the Site Auditor. Common WordPress cause: WordPress's Settings → Reading page includes a "Discourage search engines from indexing this site" checkbox. Enabling it automatically adds Disallow: / to your robots.txt. Fix: 1. Open your robots.txt file at https://www.website.com/robots.txt. 2. Find any line that reads Disallow: / under a broad user-agent (especially User-agent: *). 3. Remove that line, or replace it with only the specific path you intend to restrict (e.g., Disallow: /private/). 4. If you use WordPress, go to Settings → Reading and confirm that "Discourage search engines from indexing this site" is unchecked. 5. Save the file and verify the update is live at your robots.txt URL. The warning clears on your next Site Auditor crawl once the file is corrected. 🔌 Managing robots.txt with the Search Atlas WordPress Plugin If you run WordPress, the Search Atlas plugin includes a built-in robots.txt manager so you can edit, back up, and restore your file without leaving your WordPress admin. 🛠️ How to Use the WordPress robots.txt Manager Access the manager: 1. Log in to your WordPress admin dashboard. 2. Open the Search Atlas plugin menu. 3. Select robots.txt to open the manager. Available actions: - Edit — update the contents of your robots.txt directly in the editor and save changes. - Backup — the plugin automatically saves a snapshot each time you save a change. - Restore — open the backup history and roll back to a previous version with one click. ⏰ Timestamp note: Timestamps in the robots.txt manager currently display in UTC regardless of your WordPress timezone setting (tracked as WP-326); a fix is in progress. Convert manually when comparing backup times to your local timezone until the fix is released. ⚠️ Reminder: Because of the WP-322 detection issue described above, always double-check your saved file for an unintended Disallow: / directive after editing. 🛠️ Troubleshooting a Stalled Site Audit Crawl If your Site Audit shows as processing for more than 30 minutes without completing, the crawl may be stalled. To recover: 1. Stop the current crawl from the Site Auditor dashboard. 2. Wait a few minutes for the queue to clear. 3. Restart the crawl. If the problem persists, contact Search Atlas support so the team can investigate. 🎯 You can now identify and resolve every robots.txt issue flagged by the Search Atlas Site Auditor, manage your file directly inside the WordPress plugin, and recover from a stalled Site Audit crawl. Run a new Site Auditor crawl to confirm your fixes — and remember to manually verify your Disallow: / directive on WordPress sites until WP-322 ships.