## **🔍 Why the Crawler Can Overload a Server at 1 Page/Second**

Setting the crawler to 1 page per second does not mean your server processes only 1 lightweight request per second. Each crawl request triggers a chain of backend operations: DNS resolution, TCP handshake, HTML download, parsing, link extraction, redirect following, and database writes. On a typical server, this backend pipeline takes **200–800 ms per page** depending on page complexity, server configuration, and whether caching is active.

The actual CPU load is calculated as:

**CPU threads consumed ≈ Requests per second × Backend processing time (in seconds)**

For example, at 1 request/second with 600 ms average backend processing time, the crawler holds approximately **0.6 threads busy at all times**. If you increase to 5 requests/second, that becomes **3 threads continuously occupied** — consuming 75% of a 4-core server before accounting for your website, database, or other services running simultaneously.

This is the root cause of excessive resource usage even at seemingly low crawl rates.

## **⚙️ How Caching Status Changes Resource Consumption**

Caching is the single biggest variable in crawler resource usage. When server-side caching (e.g., Redis, Varnish, or full-page caching plugins) is active and warm, backend processing time per page drops dramatically — often to **50–150 ms**. When caching is cold or disabled, processing time rises to **400–900 ms** per page, multiplying CPU load significantly.

- **Warm cache:** Higher crawl speeds are safe; backend overhead is minimal.
- **Cold or no cache:** Reduce crawl speed significantly; each request hits PHP, the database, and templating engines at full cost.
- **After a cache flush:** Always lower crawl speed temporarily and allow the cache to rebuild before increasing again.

## **📋 Recommended Configuration Templates by Server Specs**

Use the table below as a starting point. Adjust based on your actual CPU load readings during a crawl (see Monitoring section below).

**4-Core Server (16–32 GB RAM)**

- Max crawl speed (cache warm): **2–3 pages/second**
- Max crawl speed (cache cold): **1 page/second**
- Max concurrent connections: **2**
- Crawl delay between requests: **400 ms minimum**
- Recommended crawl window: Off-peak hours only

**8-Core Server (32–64 GB RAM)**

- Max crawl speed (cache warm): **5–7 pages/second**
- Max crawl speed (cache cold): **2–3 pages/second**
- Max concurrent connections: **4**
- Crawl delay between requests: **150–200 ms**
- Recommended crawl window: Can run during business hours; monitor CPU for first 10 minutes

**16-Core+ Server (64 GB+ RAM)**

- Max crawl speed (cache warm): **10–15 pages/second**
- Max crawl speed (cache cold): **5–6 pages/second**
- Max concurrent connections: **6–8**
- Crawl delay between requests: **80–100 ms**
- Recommended crawl window: Flexible; maintain CPU headroom above 30% for other services

**Important:** These are starting baselines. Always monitor live CPU and RAM usage during the first crawl session and scale back if utilization exceeds **70% sustained CPU** on any core.

## **🛠️ How to Adjust Crawler Settings in Search Atlas**

1. Log in to Search Atlas and navigate to **Left sidebar → Site Explorer**.
2. Select the project or domain you want to crawl.
3. Open the crawler configuration panel for that project.
4. Set **Pages per second** using the recommended values above for your server tier.
5. Set **Concurrent connections** to match the recommended value for your core count.
6. Enable **Crawl delay** and input the minimum millisecond value for your tier.
7. Save your configuration and start a test crawl on a small URL batch (50–100 pages) before running a full site crawl.

## **📊 Monitoring Tools and Metrics to Watch**

During and after a crawl, track the following server metrics to confirm your configuration is safe:

- **CPU utilization per core:** Use `top` or `htop` on Linux. No individual core should sustain above 80% during a crawl. If it does, reduce pages/second immediately.
- **Load average:** A 1-minute load average exceeding your core count (e.g., above 8.0 on an 8-core server) signals the server is queueing work faster than it can process it.
- **RAM usage:** Watch for memory pressure above 85% used. Swap usage during a crawl indicates the server is under-resourced for your current settings.
- **Database query time:** Slow query logs spiking during a crawl confirm the database is a bottleneck. Lower crawl speed and verify caching is routing repeated lookups away from the database.
- **Server response time:** If your site's Time to First Byte (TTFB) rises above 1 second during crawling, the crawler is negatively impacting live visitors. Reduce speed immediately.

## **✅ Quick Troubleshooting Checklist**

- Verify caching is active and warm before starting a crawl at higher speeds.
- Confirm no other resource-intensive processes (backups, database optimizations) are running simultaneously.
- Start every new crawl configuration with a 50-page test batch before scaling to the full site.
- After any server restart or cache flush, treat the server as "cold cache" and use the lower speed values until the cache rebuilds.
- Review crawl logs for a high proportion of redirect chains or error responses — these increase backend processing time per request and require lower crawl speeds.

## **💬 Need More Help?**

If you need further assistance, open the chat widget in the bottom-right corner of the platform and type **human teammate** to be connected with a member of our team.