🕷️ Fix a Crawler That Misses Pages Beyond the Homepage

Camilo Aponte

Camilo Aponte

Last updated on Sep 30, 2026

🔍 Overview

When the Search Atlas crawler only picks up your homepage and ignores the rest of your site, several different root causes could be responsible. This guide walks you through every common reason — in order of likelihood — so you can isolate and fix the problem quickly without guesswork.

⚙️ Step 1: Check Your Crawl Settings

Start inside the platform before looking anywhere else.

  1. Open Website Studio from the left sidebar (URL: /website-studio).
  2. Select the project you want to audit.
  3. Go to Crawl Settings and confirm the starting URL is set to your root domain (e.g., https://example.com), not a single page.
  4. Make sure Follow Internal Links is enabled. If this toggle is off, the crawler will never leave the starting URL.
  5. Save any changes and re-run the crawl.

📏 Step 2: Increase the Crawl Depth Limit

Crawl depth controls how many links away from the homepage the crawler is allowed to travel. A depth of 1 means only the homepage; a depth of 2 means the homepage plus pages directly linked from it, and so on.

  1. In Crawl Settings, locate the Crawl Depth field.
  2. If it is set to 1, raise it to at least 3 for small sites or 5–10 for larger sites.
  3. Save and re-run the crawl.

If your important pages are buried several clicks from the homepage, a low depth limit is the most common reason they are missed.

🚫 Step 3: Review URL Exclusion Patterns

Exclusion rules (sometimes called URL filters or blocklists) tell the crawler to skip certain paths. A pattern that is too broad can accidentally block large sections of your site.

  1. In Crawl Settings, open the Excluded URLs or URL Patterns section.
  2. Look for wildcard rules such as /blog/*, /products/*, or /*?* that could be sweeping up pages you actually want crawled.
  3. Remove or narrow any rule that unintentionally targets your missing pages.
  4. Save and re-run the crawl.

✅ Step 4: Confirm Your Domain Is Whitelisted

The crawler only follows links that belong to the whitelisted domain. If your site uses subdomains (e.g., shop.example.com) or a www vs. non-www variation, those must be explicitly allowed.

  1. In Crawl Settings, find the Allowed Domains or Whitelist field.
  2. Add every variation of your domain that hosts content — including subdomains, www, and non-www versions.
  3. Save and re-run the crawl.

🤖 Step 5: Audit Your robots.txt File

A misconfigured robots.txt file on your server can instruct crawlers to skip entire sections of your site. This is a server-side issue, so you will need to check it outside the platform.

  1. Visit https://yourdomain.com/robots.txt in your browser.
  2. Look for Disallow rules that cover the paths where your missing pages live (e.g., Disallow: /category/).
  3. If the Search Atlas crawler is listed under a specific User-agent with restrictive rules, those rules apply to it.
  4. Update your robots.txt to allow the paths you want crawled, then re-run.

🗺️ Step 6: Check Your XML Sitemap

The crawler uses your sitemap as a discovery aid. If pages are missing from it, they may never be found — especially if internal linking is sparse.

  1. Visit https://yourdomain.com/sitemap.xml to confirm it exists and loads correctly.
  2. Verify that the pages you expect to be crawled are listed.
  3. If your sitemap is empty, outdated, or missing, regenerate it through your CMS or SEO plugin.
  4. In Crawl Settings, enter your sitemap URL under Sitemap so the crawler can reference it directly.

🔗 Step 7: Inspect Internal Linking on the Homepage

If the crawler finds links correctly but your internal pages are still missing, the homepage itself may not be linking to them — or those links may not be crawlable.

  • Links inside JavaScript-rendered menus or single-page application (SPA) frameworks may not be visible to the crawler.
  • Links using onclick events instead of standard <a href> tags are typically not followed.
  • Ensure key pages are reachable through standard HTML anchor links from at least one crawlable page.

🌐 Step 8: Rule Out Server-Side Blocking

Some hosting environments or security plugins (e.g., Cloudflare firewall rules, rate limiters, or bot-protection services) can block crawlers by IP or User-agent before they even reach your pages.

  • Check your hosting control panel or CDN dashboard for any rules that might block automated traffic.
  • Temporarily whitelist the Search Atlas crawler User-agent if your host allows it.
  • If you see 403 Forbidden or 429 Too Many Requests errors in the crawl report, server-side blocking is likely the cause.

🔁 Quick Troubleshooting Checklist

  • Follow Internal Links toggle is on
  • Crawl depth is set to 3 or higher
  • No overly broad URL exclusion patterns
  • All domain variants are whitelisted
  • robots.txt does not block the target paths
  • XML sitemap is valid and entered in crawl settings
  • Key pages are linked via standard HTML anchors
  • Server or CDN is not blocking crawler traffic

If you need further assistance, open the chat widget in the bottom-right corner of the platform and type human teammate to be connected with a member of our team.