This guide helps you **identify and understand issues across Search Atlas services**.

**⚠️ Scope:** This article covers diagnosis only — it does not contain resolution steps. Once you have identified the issue category, proceed to the relevant fix guide or escalation path.

It answers:

- What is happening?
- Where is it happening?
- Why is it happening?

Before fixing anything, this is your **diagnosis layer**.

## 🧭 How to Use This Article

Use this when:

- A user reports an issue
- A feature is not behaving as expected
- You need to classify the problem before escalating

**🛠️ Support Tool Health:** The Intercom FIN / Search Atlas Connector instability issue (QP-5196) has been resolved. If FIN responses seem incorrect or incomplete, re-verify connector health before relying on its answers to diagnose customer issues.

## 🔍 SYSTEM DIAGNOSIS MAP

## 🧠 Root Categories

## ⚙️ Infrastructure Failures

- Database unavailable
- Redis down
- Service unreachable
- Memory crashes

👉 Symptoms:

- 503 errors
- Timeouts
- Data not loading

👉 Next step: For resolution steps, see the Infrastructure Fix Guide or escalate via the Infrastructure / Platform engineering channel.

## 🔄 Async / Queue Failures

- Celery tasks missing
- Jobs stuck or never completed
- Crawls not finishing

👉 Symptoms:

- Spinners forever
- “In progress” stuck
- No visible error

**📌 Known instance:** The Atlas Brain indefinite spinner during analysis (LPS-442) is a confirmed example of this pattern — if you see it, check this category first. Note: this is distinct from the Atlas Brain 'no connection' freeze pattern, where the UI may report a false 'Success' state while the underlying deployment fails (see Atlas Brain / Website Studio section below). Both fall under async/queue failures but have different diagnostic indicators and escalation paths.

**📌 Site Audit / OTTO sync note:** Site Audit and OTTO now use the QP notify-data-change API (OTTO-2082). If data sync issues appear in these modules, verify notify-data-change pipeline health in addition to standard Celery/queue checks.

**💳 Credit consumption note:** Repeated automatic retries caused by connection errors can consume user credits. If a customer reports unexpected credit depletion, check for correlated Atlas Brain connection errors and retry loops in the same timeframe; if confirmed, escalate to the credits refund team with evidence of retry loops for manual credit restoration. Do not advise customers to re-run the tool until the underlying connection error is resolved.

👉 Next step: For resolution steps, see the Async / Queue Fix Guide or escalate via the relevant service's engineering channel (Linkgraph, Atlas Brain, or Data Pipelines).

## 🔗 External Dependency Failures

- Gemini / OpenAI limits
- SEMRush unavailable (now fetched asynchronously via Celery tasks per SE-536 — failures may surface as stuck jobs rather than immediate errors; check the Async/Queue Failures category in parallel)
- DataForSEO throttling

👉 Symptoms:

- Missing data
- AI not generating
- Partial outputs

👉 Next step: For resolution steps, see the External Dependency Fix Guide or escalate via the Data Pipelines engineering channel.

## 🖥️ Frontend / UI Failures

- React errors
- Graphs not rendering
- reCAPTCHA issues
- Billing Modal Freeze — The Add Payment Method modal may enter an indefinite loading state after clicking Add Card, with no error shown and no card added. This is a known frontend issue (QPB-697). Do not advise the customer to retry repeatedly; escalate to the billing team.
- Press Release / Signal Genesys Publish Blocked by Image Validation — If a customer cannot publish and sees inline image errors in Step 3, the cause is an unsupported image format (WEBP, AVIF, HEIC, BMP, TIFF, ICO) or a corrupt/mislabelled file. This is expected validation behavior as of AB-970, not a system failure.

👉 Symptoms:

- White screens
- Broken UI elements
- Actions not working

👉 Next step: For resolution steps, see the Frontend Fix Guide or escalate via the Frontend engineering channel.

## 🔀 Inter-Module Pipeline / Bridge Failures

- API bridges between modules (e.g., Content Genius → Link Lab) can enter a blocked state
- Actions complete on one side but do not propagate to the target module
- Reference escalation: 215474220017778

👉 Symptoms:

- Source module reports success
- Target module never receives the handoff
- No visible error in either module

👉 Next step: Escalate to engineering with pipeline telemetry data for the affected modules.

## 🧩 SERVICE-LEVEL DIAGNOSIS

## 🔧 Linkgraph (Core API)

## 🔴 Most Critical Patterns

- Metrics not updating → queue issue
- 503 errors → DB connection exhaustion
- Jobs stuck → Celery worker version mismatch or task queue not processing (workers running an outdated task signature, or queue is backed up / not consuming — observable as tasks staying in PENDING/RECEIVED without progressing)
- Crawls stuck → timeout on large sites

👉 Insight:

**Linkgraph issues are usually backend or queue-related**

## 🤖 Content Assistant

## 🔴 Most Critical Patterns

- AI not generating (hard failure) → Gemini quota (429)
- AI reports success but output not applied (silent failure) → backend sync issue; UI may show 'Success' while generated content is not persisted or activated
- Content incomplete → keyword service failing
- JSON errors → malformed LLM output

👉 Insight:

**Most failures come from external AI dependencies, but watch for silent successes where the UI confirms an action that did not actually take effect on the backend.**

## 🧠 Atlas Brain / Website Studio

## 🔴 Most Critical Patterns

- Infinite spinner during analysis (LPS-442) → queue/async failure; see Async/Queue Failures category. Diagnostic indicator: UI shows a persistent spinner with no status update. Distinct from the 'no connection' freeze pattern below.
- 'No connection' freeze with false success states → backend connection drop causing NLP terms and schema deployment to fail silently while the UI displays 'Success' (escalation 215474307923656). Diagnostic indicator: UI confirms an action that did not propagate to the backend. Different diagnostic path and escalation than LPS-442.
- CMS Connector State Mismatch — Atlas Brain may display a disconnected status even when the connector is active on the backend (cases 215474386800374, 215474258075754). Verify via backend admin before advising the customer to reconnect.
- AI Agent Schema/NLP Deployment Sync Failure — Success status displayed but terms not activated (case 215474307923656). Check backend sync state. This may require manual deployment and engineering escalation (reference pattern: SPE-692 / AIAGENT-1904).
- Publish Blocked by False Disconnection Error — Atlas Brain reports website not connected despite the site being active (cases 215474258075754, 215474386800374). Verify connection state on backend before advising reconnection or manual publish workaround.
- Unexpected credit depletion → connection errors can trigger URL Indexer or other jobs to retry automatically and burn credits (case 215474438977629); correlate with Atlas Brain or connector error events, and escalate to the credits refund team if retry loops are confirmed.

👉 Insight:

**Atlas Brain failures often present as successful UI states masking backend faults — always verify actual backend state before trusting the UI.**

## 🖥️ Customer Frontend

## 🔴 Most Critical Patterns

- Login issues → reCAPTCHA
- Graphs missing → null data
- App crashes → chunk load errors
- Billing Modal Freeze (QPB-697) → Add Payment Method modal hangs after clicking Add Card with no error and no card added; escalate to the billing team rather than advising retries.

👉 Insight:

**Frontend errors are usually data-state or loading timing issues**

## 🔍 Keyword Explorer

## 🔴 Most Critical Patterns

- No keyword data → ClickHouse failure
- Missing enrichment → SEMRush unavailable
- Slow results → DataForSEO limits

👉 Insight:

**This service depends heavily on external data pipelines**

## 🔗 Backlink App

## 🔴 Most Critical Patterns

- Data not loading → DB issues
- No major user-facing errors → mostly infra/logging

## 🧠 KEY INSIGHT

👉 Most issues are **not bugs**  
They are:

- System limits
- Dependencies
- Scale problems

## 🔑 Golden Rule

Before fixing anything:

👉 Ask:

**Is this: UI, Data, Queue, Infrastructure, or Inter-Module Pipeline?**

## ➡️ Next Steps After Diagnosis

Once you have classified the issue:

- Match the category to the relevant fix guide or runbook
- Apply P-level severity classification (P0–P3) based on customer impact and scope
- Escalate through the appropriate engineering channel (Linkgraph, Atlas Brain, Frontend, Data Pipelines, or Billing) with the diagnosis category, observed symptoms, and any referenced ticket IDs (e.g., LPS-442, SE-536, SPE-692, QP-5196, QPB-697, OTTO-2082, AB-970)
- Hand off to the on-call engineer using the standard escalation protocol if the issue is P0/P1 or customer-blocking