🔍 Overview
When the Search Atlas crawler only picks up your homepage and ignores the rest of your site, several different root causes could be responsible. This guide walks you through every common reason — in order of likelihood — so you can isolate and fix the problem quickly without guesswork.
⚙️ Step 1: Check Your Crawl Settings
Start inside the platform before looking anywhere else.
- Open Website Studio from the left sidebar (URL:
/website-studio). - Select the project you want to audit.
- Go to Crawl Settings and confirm the starting URL is set to your root domain (e.g.,
https://example.com), not a single page. - Make sure Follow Internal Links is enabled. If this toggle is off, the crawler will never leave the starting URL.
- Save any changes and re-run the crawl.
📏 Step 2: Increase the Crawl Depth Limit
Crawl depth controls how many links away from the homepage the crawler is allowed to travel. A depth of 1 means only the homepage; a depth of 2 means the homepage plus pages directly linked from it, and so on.
- In Crawl Settings, locate the Crawl Depth field.
- If it is set to 1, raise it to at least 3 for small sites or 5–10 for larger sites.
- Save and re-run the crawl.
If your important pages are buried several clicks from the homepage, a low depth limit is the most common reason they are missed.
🚫 Step 3: Review URL Exclusion Patterns
Exclusion rules (sometimes called URL filters or blocklists) tell the crawler to skip certain paths. A pattern that is too broad can accidentally block large sections of your site.
- In Crawl Settings, open the Excluded URLs or URL Patterns section.
- Look for wildcard rules such as
/blog/*,/products/*, or/*?*that could be sweeping up pages you actually want crawled. - Remove or narrow any rule that unintentionally targets your missing pages.
- Save and re-run the crawl.
✅ Step 4: Confirm Your Domain Is Whitelisted
The crawler only follows links that belong to the whitelisted domain. If your site uses subdomains (e.g., shop.example.com) or a www vs. non-www variation, those must be explicitly allowed.
- In Crawl Settings, find the Allowed Domains or Whitelist field.
- Add every variation of your domain that hosts content — including subdomains,
www, and non-wwwversions. - Save and re-run the crawl.
🤖 Step 5: Audit Your robots.txt File
A misconfigured robots.txt file on your server can instruct crawlers to skip entire sections of your site. This is a server-side issue, so you will need to check it outside the platform.
- Visit
https://yourdomain.com/robots.txtin your browser. - Look for
Disallowrules that cover the paths where your missing pages live (e.g.,Disallow: /category/). - If the Search Atlas crawler is listed under a specific
User-agentwith restrictive rules, those rules apply to it. - Update your
robots.txtto allow the paths you want crawled, then re-run.
🗺️ Step 6: Check Your XML Sitemap
The crawler uses your sitemap as a discovery aid. If pages are missing from it, they may never be found — especially if internal linking is sparse.
- Visit
https://yourdomain.com/sitemap.xmlto confirm it exists and loads correctly. - Verify that the pages you expect to be crawled are listed.
- If your sitemap is empty, outdated, or missing, regenerate it through your CMS or SEO plugin.
- In Crawl Settings, enter your sitemap URL under Sitemap so the crawler can reference it directly.
🔗 Step 7: Inspect Internal Linking on the Homepage
If the crawler finds links correctly but your internal pages are still missing, the homepage itself may not be linking to them — or those links may not be crawlable.
- Links inside JavaScript-rendered menus or single-page application (SPA) frameworks may not be visible to the crawler.
- Links using
onclickevents instead of standard<a href>tags are typically not followed. - Ensure key pages are reachable through standard HTML anchor links from at least one crawlable page.
🌐 Step 8: Rule Out Server-Side Blocking
Some hosting environments or security plugins (e.g., Cloudflare firewall rules, rate limiters, or bot-protection services) can block crawlers by IP or User-agent before they even reach your pages.
- Check your hosting control panel or CDN dashboard for any rules that might block automated traffic.
- Temporarily whitelist the Search Atlas crawler User-agent if your host allows it.
- If you see 403 Forbidden or 429 Too Many Requests errors in the crawl report, server-side blocking is likely the cause.
🔁 Quick Troubleshooting Checklist
- Follow Internal Links toggle is on
- Crawl depth is set to 3 or higher
- No overly broad URL exclusion patterns
- All domain variants are whitelisted
robots.txtdoes not block the target paths- XML sitemap is valid and entered in crawl settings
- Key pages are linked via standard HTML anchors
- Server or CDN is not blocking crawler traffic
If you need further assistance, open the chat widget in the bottom-right corner of the platform and type human teammate to be connected with a member of our team.