🔍 Understanding How the Crawler Uses Server Resources
The Search Atlas crawler is designed to be thorough and accurate, which means it performs multiple operations per page visit — fetching HTML, resolving redirects, analysing on-page elements, and storing results. Even when the crawl speed is set to 1 page per second, these parallel background processes can place a noticeable load on your server, especially on resource-constrained hosting environments.
The key thing to understand is that crawl speed (pages per second) is not the only factor that determines server load. Concurrent connections, page complexity, server response times, and the size of your site all contribute to total resource consumption during a crawl.
⚙️ What Affects Server Resource Usage
- Concurrent connections: Even at 1 page per second, the crawler may open multiple simultaneous connections to fetch resources like images, scripts, and stylesheets referenced on each page.
- Page complexity: Pages with many internal links, large HTML payloads, or heavy JavaScript increase the processing work required per page.
- Server response time: A slow server causes connections to remain open longer, stacking up alongside new requests and increasing peak load.
- Redirect chains: Each redirect adds an extra request, multiplying the total number of HTTP calls the crawler makes.
- Site size: Larger sites mean the cumulative effect of even a conservative crawl rate lasts much longer, sustaining elevated resource usage over an extended period.
🖥️ Recommended Configuration for an 8-Core Dedicated Server
An 8-core dedicated server has more headroom than shared hosting, but it is still important to configure the crawler thoughtfully to avoid impacting live site performance for real visitors. Follow these guidelines:
- Set crawl speed to 2–4 pages per second as a starting point. An 8-core server can generally handle this comfortably, but monitor CPU and memory usage during the first crawl to confirm.
- Schedule crawls during low-traffic periods — typically late night or early morning in your visitors' primary time zone. This ensures the crawler competes with as little live traffic as possible.
- Limit concurrent connections if your crawler settings expose this option. A value of 4–6 concurrent connections is a sensible ceiling for most dedicated servers.
- Exclude unnecessary URLs from the crawl scope. Use the URL exclusion or path filtering settings to skip admin paths, media upload directories, and any dynamically generated pages that do not need SEO analysis.
- Pause and resume crawls if you notice performance degradation. You do not need to restart from the beginning — the crawler can resume where it left off.
- Review your server's error logs after the first crawl. A spike in 503 (Service Unavailable) or 429 (Too Many Requests) responses is a reliable signal that you need to reduce speed or concurrency further.
💡 Tips to Reduce Crawl Impact Without Sacrificing Coverage
- Use a crawl budget wisely: Prioritise crawling your most important URLs — product pages, blog posts, and landing pages — and exclude low-value pages like tag archives, search result pages, and session-ID URLs.
- Enable caching on your server: A properly configured server-side cache (e.g. Redis, Varnish, or a CDN) dramatically reduces the processing overhead of each crawler request because the server returns a cached response instead of regenerating the page.
- Check your robots.txt: Ensure the Search Atlas crawler is not blocked from pages you want crawled, and that it is blocked from areas you do not — such as /wp-admin/ or /checkout/.
- Keep the crawl scope focused: If you are running a targeted audit, crawl a specific subdirectory rather than the entire domain to reduce total requests.
🚀 Monitoring Resource Usage During a Crawl
While a crawl is running, use your server's monitoring tools (such as htop, Netdata, or your hosting control panel's resource graphs) to watch CPU load, RAM usage, and active connections in real time. A healthy crawl on an 8-core dedicated server should keep average CPU usage below 60–70%. If you consistently see CPU peaking above 90% or RAM approaching its limit, reduce the crawl speed setting and restart the crawl.
If your hosting provider enforces rate limits or automatically throttles connections from a single IP, contact them to whitelist the Search Atlas crawler IP ranges so legitimate crawl traffic is not misidentified as an attack.
🛠️ Need Further Help
If you need further assistance, open the chat widget in the bottom-right corner of the platform and type human teammate to be connected with a member of our team.