🧠 Search Atlas Troubleshooting Guide (Diagnosis Layer)

Camilo Aponte

Camilo Aponte

Last updated on Sep 30, 2026

This guide helps you identify and understand issues across Search Atlas services.

⚠️ Scope: This article covers diagnosis only — it does not contain resolution steps. Once you have identified the issue category, proceed to the relevant fix guide or escalation path.

It answers:

  • What is happening?
  • Where is it happening?
  • Why is it happening?

Before fixing anything, this is your diagnosis layer.

🧭 How to Use This Article

Use this when:

  • A user reports an issue
  • A feature is not behaving as expected
  • You need to classify the problem before escalating

🛠️ Support Tool Health: The Intercom FIN / Search Atlas Connector instability issue (QP-5196) has been resolved. If FIN responses seem incorrect or incomplete, re-verify connector health before relying on its answers to diagnose customer issues.

🔍 SYSTEM DIAGNOSIS MAP

🧠 Root Categories

⚙️ Infrastructure Failures

  • Database unavailable
  • Redis down
  • Service unreachable
  • Memory crashes

👉 Symptoms:

  • 503 errors
  • Timeouts
  • Data not loading

👉 Next step: For resolution steps, see the Infrastructure Fix Guide or escalate via the Infrastructure / Platform engineering channel.

🔄 Async / Queue Failures

  • Celery tasks missing
  • Jobs stuck or never completed
  • Crawls not finishing

👉 Symptoms:

  • Spinners forever
  • “In progress” stuck
  • No visible error

📌 Known instance: The Atlas Brain indefinite spinner during analysis (LPS-442) is a confirmed example of this pattern — if you see it, check this category first. Note: this is distinct from the Atlas Brain 'no connection' freeze pattern, where the UI may report a false 'Success' state while the underlying deployment fails (see Atlas Brain / Website Studio section below). Both fall under async/queue failures but have different diagnostic indicators and escalation paths.

📌 Site Audit / OTTO sync note: Site Audit and OTTO now use the QP notify-data-change API (OTTO-2082). If data sync issues appear in these modules, verify notify-data-change pipeline health in addition to standard Celery/queue checks.

💳 Credit consumption note: Repeated automatic retries caused by connection errors can consume user credits. If a customer reports unexpected credit depletion, check for correlated Atlas Brain connection errors and retry loops in the same timeframe; if confirmed, escalate to the credits refund team with evidence of retry loops for manual credit restoration. Do not advise customers to re-run the tool until the underlying connection error is resolved.

👉 Next step: For resolution steps, see the Async / Queue Fix Guide or escalate via the relevant service's engineering channel (Linkgraph, Atlas Brain, or Data Pipelines).

🔗 External Dependency Failures

  • Gemini / OpenAI limits
  • SEMRush unavailable (now fetched asynchronously via Celery tasks per SE-536 — failures may surface as stuck jobs rather than immediate errors; check the Async/Queue Failures category in parallel)
  • DataForSEO throttling

👉 Symptoms:

  • Missing data
  • AI not generating
  • Partial outputs

👉 Next step: For resolution steps, see the External Dependency Fix Guide or escalate via the Data Pipelines engineering channel.

🖥️ Frontend / UI Failures

  • React errors
  • Graphs not rendering
  • reCAPTCHA issues
  • Billing Modal Freeze — The Add Payment Method modal may enter an indefinite loading state after clicking Add Card, with no error shown and no card added. This is a known frontend issue (QPB-697). Do not advise the customer to retry repeatedly; escalate to the billing team.
  • Press Release / Signal Genesys Publish Blocked by Image Validation — If a customer cannot publish and sees inline image errors in Step 3, the cause is an unsupported image format (WEBP, AVIF, HEIC, BMP, TIFF, ICO) or a corrupt/mislabelled file. This is expected validation behavior as of AB-970, not a system failure.

👉 Symptoms:

  • White screens
  • Broken UI elements
  • Actions not working

👉 Next step: For resolution steps, see the Frontend Fix Guide or escalate via the Frontend engineering channel.

🔀 Inter-Module Pipeline / Bridge Failures

  • API bridges between modules (e.g., Content Genius → Link Lab) can enter a blocked state
  • Actions complete on one side but do not propagate to the target module
  • Reference escalation: 215474220017778

👉 Symptoms:

  • Source module reports success
  • Target module never receives the handoff
  • No visible error in either module

👉 Next step: Escalate to engineering with pipeline telemetry data for the affected modules.

🧩 SERVICE-LEVEL DIAGNOSIS

🔧 Linkgraph (Core API)

🔴 Most Critical Patterns

  • Metrics not updating → queue issue
  • 503 errors → DB connection exhaustion
  • Jobs stuck → Celery worker version mismatch or task queue not processing (workers running an outdated task signature, or queue is backed up / not consuming — observable as tasks staying in PENDING/RECEIVED without progressing)
  • Crawls stuck → timeout on large sites

👉 Insight:

Linkgraph issues are usually backend or queue-related

🤖 Content Assistant

🔴 Most Critical Patterns

  • AI not generating (hard failure) → Gemini quota (429)
  • AI reports success but output not applied (silent failure) → backend sync issue; UI may show 'Success' while generated content is not persisted or activated
  • Content incomplete → keyword service failing
  • JSON errors → malformed LLM output

👉 Insight:

Most failures come from external AI dependencies, but watch for silent successes where the UI confirms an action that did not actually take effect on the backend.

🧠 Atlas Brain / Website Studio

🔴 Most Critical Patterns

  • Infinite spinner during analysis (LPS-442) → queue/async failure; see Async/Queue Failures category. Diagnostic indicator: UI shows a persistent spinner with no status update. Distinct from the 'no connection' freeze pattern below.
  • 'No connection' freeze with false success states → backend connection drop causing NLP terms and schema deployment to fail silently while the UI displays 'Success' (escalation 215474307923656). Diagnostic indicator: UI confirms an action that did not propagate to the backend. Different diagnostic path and escalation than LPS-442.
  • CMS Connector State Mismatch — Atlas Brain may display a disconnected status even when the connector is active on the backend (cases 215474386800374, 215474258075754). Verify via backend admin before advising the customer to reconnect.
  • AI Agent Schema/NLP Deployment Sync Failure — Success status displayed but terms not activated (case 215474307923656). Check backend sync state. This may require manual deployment and engineering escalation (reference pattern: SPE-692 / AIAGENT-1904).
  • Publish Blocked by False Disconnection Error — Atlas Brain reports website not connected despite the site being active (cases 215474258075754, 215474386800374). Verify connection state on backend before advising reconnection or manual publish workaround.
  • Unexpected credit depletion → connection errors can trigger URL Indexer or other jobs to retry automatically and burn credits (case 215474438977629); correlate with Atlas Brain or connector error events, and escalate to the credits refund team if retry loops are confirmed.

👉 Insight:

Atlas Brain failures often present as successful UI states masking backend faults — always verify actual backend state before trusting the UI.

🖥️ Customer Frontend

🔴 Most Critical Patterns

  • Login issues → reCAPTCHA
  • Graphs missing → null data
  • App crashes → chunk load errors
  • Billing Modal Freeze (QPB-697) → Add Payment Method modal hangs after clicking Add Card with no error and no card added; escalate to the billing team rather than advising retries.

👉 Insight:

Frontend errors are usually data-state or loading timing issues

🔍 Keyword Explorer

🔴 Most Critical Patterns

  • No keyword data → ClickHouse failure
  • Missing enrichment → SEMRush unavailable
  • Slow results → DataForSEO limits

👉 Insight:

This service depends heavily on external data pipelines

🔗 Backlink App

🔴 Most Critical Patterns

  • Data not loading → DB issues
  • No major user-facing errors → mostly infra/logging

🧠 KEY INSIGHT

👉 Most issues are not bugs
They are:

  • System limits
  • Dependencies
  • Scale problems

🔑 Golden Rule

Before fixing anything:

👉 Ask:

Is this: UI, Data, Queue, Infrastructure, or Inter-Module Pipeline?

➡️ Next Steps After Diagnosis

Once you have classified the issue:

  • Match the category to the relevant fix guide or runbook
  • Apply P-level severity classification (P0–P3) based on customer impact and scope
  • Escalate through the appropriate engineering channel (Linkgraph, Atlas Brain, Frontend, Data Pipelines, or Billing) with the diagnosis category, observed symptoms, and any referenced ticket IDs (e.g., LPS-442, SE-536, SPE-692, QP-5196, QPB-697, OTTO-2082, AB-970)
  • Hand off to the on-call engineer using the standard escalation protocol if the issue is P0/P1 or customer-blocking