This guide helps you identify and understand issues across Search Atlas services.
⚠️ Scope: This article covers diagnosis only — it does not contain resolution steps. Once you have identified the issue category, proceed to the relevant fix guide or escalation path.
It answers:
- What is happening?
- Where is it happening?
- Why is it happening?
Before fixing anything, this is your diagnosis layer.
🧭 How to Use This Article
Use this when:
- A user reports an issue
- A feature is not behaving as expected
- You need to classify the problem before escalating
🛠️ Support Tool Health: The Intercom FIN / Search Atlas Connector instability issue (QP-5196) has been resolved. If FIN responses seem incorrect or incomplete, re-verify connector health before relying on its answers to diagnose customer issues.
🔍 SYSTEM DIAGNOSIS MAP
🧠 Root Categories
⚙️ Infrastructure Failures
- Database unavailable
- Redis down
- Service unreachable
- Memory crashes
👉 Symptoms:
- 503 errors
- Timeouts
- Data not loading
👉 Next step: For resolution steps, see the Infrastructure Fix Guide or escalate via the Infrastructure / Platform engineering channel.
🔄 Async / Queue Failures
- Celery tasks missing
- Jobs stuck or never completed
- Crawls not finishing
👉 Symptoms:
- Spinners forever
- “In progress” stuck
- No visible error
📌 Known instance: The Atlas Brain indefinite spinner during analysis (LPS-442) is a confirmed example of this pattern — if you see it, check this category first. Note: this is distinct from the Atlas Brain 'no connection' freeze pattern, where the UI may report a false 'Success' state while the underlying deployment fails (see Atlas Brain / Website Studio section below). Both fall under async/queue failures but have different diagnostic indicators and escalation paths.
📌 Site Audit / OTTO sync note: Site Audit and OTTO now use the QP notify-data-change API (OTTO-2082). If data sync issues appear in these modules, verify notify-data-change pipeline health in addition to standard Celery/queue checks.
💳 Credit consumption note: Repeated automatic retries caused by connection errors can consume user credits. If a customer reports unexpected credit depletion, check for correlated Atlas Brain connection errors and retry loops in the same timeframe; if confirmed, escalate to the credits refund team with evidence of retry loops for manual credit restoration. Do not advise customers to re-run the tool until the underlying connection error is resolved.
👉 Next step: For resolution steps, see the Async / Queue Fix Guide or escalate via the relevant service's engineering channel (Linkgraph, Atlas Brain, or Data Pipelines).
🔗 External Dependency Failures
- Gemini / OpenAI limits
- SEMRush unavailable (now fetched asynchronously via Celery tasks per SE-536 — failures may surface as stuck jobs rather than immediate errors; check the Async/Queue Failures category in parallel)
- DataForSEO throttling
👉 Symptoms:
- Missing data
- AI not generating
- Partial outputs
👉 Next step: For resolution steps, see the External Dependency Fix Guide or escalate via the Data Pipelines engineering channel.
🖥️ Frontend / UI Failures
- React errors
- Graphs not rendering
- reCAPTCHA issues
- Billing Modal Freeze — The Add Payment Method modal may enter an indefinite loading state after clicking Add Card, with no error shown and no card added. This is a known frontend issue (QPB-697). Do not advise the customer to retry repeatedly; escalate to the billing team.
- Press Release / Signal Genesys Publish Blocked by Image Validation — If a customer cannot publish and sees inline image errors in Step 3, the cause is an unsupported image format (WEBP, AVIF, HEIC, BMP, TIFF, ICO) or a corrupt/mislabelled file. This is expected validation behavior as of AB-970, not a system failure.
👉 Symptoms:
- White screens
- Broken UI elements
- Actions not working
👉 Next step: For resolution steps, see the Frontend Fix Guide or escalate via the Frontend engineering channel.
🔀 Inter-Module Pipeline / Bridge Failures
- API bridges between modules (e.g., Content Genius → Link Lab) can enter a blocked state
- Actions complete on one side but do not propagate to the target module
- Reference escalation: 215474220017778
👉 Symptoms:
- Source module reports success
- Target module never receives the handoff
- No visible error in either module
👉 Next step: Escalate to engineering with pipeline telemetry data for the affected modules.
🧩 SERVICE-LEVEL DIAGNOSIS
🔧 Linkgraph (Core API)
🔴 Most Critical Patterns
- Metrics not updating → queue issue
- 503 errors → DB connection exhaustion
- Jobs stuck → Celery worker version mismatch or task queue not processing (workers running an outdated task signature, or queue is backed up / not consuming — observable as tasks staying in PENDING/RECEIVED without progressing)
- Crawls stuck → timeout on large sites
👉 Insight:
Linkgraph issues are usually backend or queue-related
🤖 Content Assistant
🔴 Most Critical Patterns
- AI not generating (hard failure) → Gemini quota (429)
- AI reports success but output not applied (silent failure) → backend sync issue; UI may show 'Success' while generated content is not persisted or activated
- Content incomplete → keyword service failing
- JSON errors → malformed LLM output
👉 Insight:
Most failures come from external AI dependencies, but watch for silent successes where the UI confirms an action that did not actually take effect on the backend.
🧠 Atlas Brain / Website Studio
🔴 Most Critical Patterns
- Infinite spinner during analysis (LPS-442) → queue/async failure; see Async/Queue Failures category. Diagnostic indicator: UI shows a persistent spinner with no status update. Distinct from the 'no connection' freeze pattern below.
- 'No connection' freeze with false success states → backend connection drop causing NLP terms and schema deployment to fail silently while the UI displays 'Success' (escalation 215474307923656). Diagnostic indicator: UI confirms an action that did not propagate to the backend. Different diagnostic path and escalation than LPS-442.
- CMS Connector State Mismatch — Atlas Brain may display a disconnected status even when the connector is active on the backend (cases 215474386800374, 215474258075754). Verify via backend admin before advising the customer to reconnect.
- AI Agent Schema/NLP Deployment Sync Failure — Success status displayed but terms not activated (case 215474307923656). Check backend sync state. This may require manual deployment and engineering escalation (reference pattern: SPE-692 / AIAGENT-1904).
- Publish Blocked by False Disconnection Error — Atlas Brain reports website not connected despite the site being active (cases 215474258075754, 215474386800374). Verify connection state on backend before advising reconnection or manual publish workaround.
- Unexpected credit depletion → connection errors can trigger URL Indexer or other jobs to retry automatically and burn credits (case 215474438977629); correlate with Atlas Brain or connector error events, and escalate to the credits refund team if retry loops are confirmed.
👉 Insight:
Atlas Brain failures often present as successful UI states masking backend faults — always verify actual backend state before trusting the UI.
🖥️ Customer Frontend
🔴 Most Critical Patterns
- Login issues → reCAPTCHA
- Graphs missing → null data
- App crashes → chunk load errors
- Billing Modal Freeze (QPB-697) → Add Payment Method modal hangs after clicking Add Card with no error and no card added; escalate to the billing team rather than advising retries.
👉 Insight:
Frontend errors are usually data-state or loading timing issues
🔍 Keyword Explorer
🔴 Most Critical Patterns
- No keyword data → ClickHouse failure
- Missing enrichment → SEMRush unavailable
- Slow results → DataForSEO limits
👉 Insight:
This service depends heavily on external data pipelines
🔗 Backlink App
🔴 Most Critical Patterns
- Data not loading → DB issues
- No major user-facing errors → mostly infra/logging
🧠 KEY INSIGHT
👉 Most issues are not bugs
They are:
- System limits
- Dependencies
- Scale problems
🔑 Golden Rule
Before fixing anything:
👉 Ask:
Is this: UI, Data, Queue, Infrastructure, or Inter-Module Pipeline?
➡️ Next Steps After Diagnosis
Once you have classified the issue:
- Match the category to the relevant fix guide or runbook
- Apply P-level severity classification (P0–P3) based on customer impact and scope
- Escalate through the appropriate engineering channel (Linkgraph, Atlas Brain, Frontend, Data Pipelines, or Billing) with the diagnosis category, observed symptoms, and any referenced ticket IDs (e.g., LPS-442, SE-536, SPE-692, QP-5196, QPB-697, OTTO-2082, AB-970)
- Hand off to the on-call engineer using the standard escalation protocol if the issue is P0/P1 or customer-blocking