Overview
If OTTO is indexing significantly fewer pages than your site contains—for example, 16,000 of 65,000 pages—this is a crawl volume issue, not a plugin installation problem. This article explains the most common root causes of low crawl coverage and how to diagnose them.
Why Is OTTO Only Reaching a Subset of Pages?
There are five primary reasons OTTO may index fewer pages than expected. Work through each one before assuming a technical fault.
- Crawl budget limits: Large sites have a finite crawl budget—the number of URLs a crawler will fetch in a given period. If your site has 65,000 pages, OTTO prioritises the most discoverable and link-connected pages first. Pages buried deep in your architecture or receiving few internal links are reached in later crawl cycles. To expand coverage, improve internal linking to surface orphaned or deep pages.
- Noindex tags on large page groups: Template-level
noindexmeta tags or HTTP headers can silently exclude entire categories, tags, paginated archives, or product filter URLs. Audit your page templates—especially e-commerce faceted navigation, paginated series, and auto-generated tag pages—to confirm intended pages are not markednoindex. - Robots.txt restrictions: A
Disallowrule in yourrobots.txtfile can block OTTO from crawling entire URL path segments. Review yourrobots.txt(accessible atyourdomain.com/robots.txt) and confirm that no rules accidentally block large sections of the site such as/products/,/blog/, or/category/. - Sitemap gaps: If your XML sitemap does not include all expected pages—or if it contains broken URLs, incorrect priorities, or is not submitted—OTTO has no reliable map to follow. Verify your sitemap is complete, up to date, and submitted within the platform.
- Site structure and internal linking depth: Pages more than three to four clicks from the homepage are harder for any crawler to reach within a standard crawl budget window. Flatten your site architecture where possible and ensure key pages are reachable through prominent internal links.
What to Have Ready When Escalating
If you have reviewed all five areas above and crawl coverage remains unexpectedly low, please gather the following before contacting support:
- Your project name and the domain being crawled
- The total number of pages on your site and the number OTTO has indexed
- A copy of your
robots.txtfile contents - Your XML sitemap URL
- Any specific URL patterns or sections of the site that appear to be missing from the crawl
If you need further assistance, open the chat widget in the bottom-right corner of the platform and type human teammate to be connected with a member of our team.