🔍 Why the Crawler Can Overload a Server at 1 Page/Second
Setting the crawler to 1 page per second does not mean your server processes only 1 lightweight request per second. Each crawl request triggers a chain of backend operations: DNS resolution, TCP handshake, HTML download, parsing, link extraction, redirect following, and database writes. On a typical server, this backend pipeline takes 200–800 ms per page depending on page complexity, server configuration, and whether caching is active.
The actual CPU load is calculated as:
CPU threads consumed ≈ Requests per second × Backend processing time (in seconds)
For example, at 1 request/second with 600 ms average backend processing time, the crawler holds approximately 0.6 threads busy at all times. If you increase to 5 requests/second, that becomes 3 threads continuously occupied — consuming 75% of a 4-core server before accounting for your website, database, or other services running simultaneously.
This is the root cause of excessive resource usage even at seemingly low crawl rates.
⚙️ How Caching Status Changes Resource Consumption
Caching is the single biggest variable in crawler resource usage. When server-side caching (e.g., Redis, Varnish, or full-page caching plugins) is active and warm, backend processing time per page drops dramatically — often to 50–150 ms. When caching is cold or disabled, processing time rises to 400–900 ms per page, multiplying CPU load significantly.
- Warm cache: Higher crawl speeds are safe; backend overhead is minimal.
- Cold or no cache: Reduce crawl speed significantly; each request hits PHP, the database, and templating engines at full cost.
- After a cache flush: Always lower crawl speed temporarily and allow the cache to rebuild before increasing again.
📋 Recommended Configuration Templates by Server Specs
Use the table below as a starting point. Adjust based on your actual CPU load readings during a crawl (see Monitoring section below).
4-Core Server (16–32 GB RAM)
- Max crawl speed (cache warm): 2–3 pages/second
- Max crawl speed (cache cold): 1 page/second
- Max concurrent connections: 2
- Crawl delay between requests: 400 ms minimum
- Recommended crawl window: Off-peak hours only
8-Core Server (32–64 GB RAM)
- Max crawl speed (cache warm): 5–7 pages/second
- Max crawl speed (cache cold): 2–3 pages/second
- Max concurrent connections: 4
- Crawl delay between requests: 150–200 ms
- Recommended crawl window: Can run during business hours; monitor CPU for first 10 minutes
16-Core+ Server (64 GB+ RAM)
- Max crawl speed (cache warm): 10–15 pages/second
- Max crawl speed (cache cold): 5–6 pages/second
- Max concurrent connections: 6–8
- Crawl delay between requests: 80–100 ms
- Recommended crawl window: Flexible; maintain CPU headroom above 30% for other services
Important: These are starting baselines. Always monitor live CPU and RAM usage during the first crawl session and scale back if utilization exceeds 70% sustained CPU on any core.
🛠️ How to Adjust Crawler Settings in Search Atlas
- Log in to Search Atlas and navigate to Left sidebar → Site Explorer.
- Select the project or domain you want to crawl.
- Open the crawler configuration panel for that project.
- Set Pages per second using the recommended values above for your server tier.
- Set Concurrent connections to match the recommended value for your core count.
- Enable Crawl delay and input the minimum millisecond value for your tier.
- Save your configuration and start a test crawl on a small URL batch (50–100 pages) before running a full site crawl.
📊 Monitoring Tools and Metrics to Watch
During and after a crawl, track the following server metrics to confirm your configuration is safe:
- CPU utilization per core: Use
toporhtopon Linux. No individual core should sustain above 80% during a crawl. If it does, reduce pages/second immediately. - Load average: A 1-minute load average exceeding your core count (e.g., above 8.0 on an 8-core server) signals the server is queueing work faster than it can process it.
- RAM usage: Watch for memory pressure above 85% used. Swap usage during a crawl indicates the server is under-resourced for your current settings.
- Database query time: Slow query logs spiking during a crawl confirm the database is a bottleneck. Lower crawl speed and verify caching is routing repeated lookups away from the database.
- Server response time: If your site's Time to First Byte (TTFB) rises above 1 second during crawling, the crawler is negatively impacting live visitors. Reduce speed immediately.
✅ Quick Troubleshooting Checklist
- Verify caching is active and warm before starting a crawl at higher speeds.
- Confirm no other resource-intensive processes (backups, database optimizations) are running simultaneously.
- Start every new crawl configuration with a 50-page test batch before scaling to the full site.
- After any server restart or cache flush, treat the server as "cold cache" and use the lower speed values until the cache rebuilds.
- Review crawl logs for a high proportion of redirect chains or error responses — these increase backend processing time per request and require lower crawl speeds.
💬 Need More Help?
If you need further assistance, open the chat widget in the bottom-right corner of the platform and type human teammate to be connected with a member of our team.