A healthy HTTP process does not prove the whole application is ready. Begin with GET /api/health for liveness, then test a signed-in read and a synthetic edit. Inspect the corresponding saved request, workflow run, or email state before assuming a background process succeeded. Keep private content and credentials out of shared logs.
| Symptom | Check first |
|---|---|
| Sign-in fails or loops | Matching Clerk keys, exact authorized origins, callback registration, and production/development application alignment |
| AI request stays queued | ATLAS_WORKER_MODE, supervised worker status, role selection, embedded startup flags, and worker database target |
| AI is unavailable | Provider configuration and saved failure state; choose the built-in helper explicitly only for supported tasks |
| Search misses an email | Mailbox ownership, completed sync, folder coverage, filters, and GET /api/search/status |
| A record edit conflicts | Workspace revision changed; refresh and review again instead of forcing the stale proposal |
| Attachments disappear after restart | Persistent data mount, ownership, and whether the correct private directory was mounted |
| Email connection succeeds but sending fails | Drafts/Sent capabilities, delegated Send As permission where relevant, and the saved provider error |
| Phone cannot reach the app | Device-reachable HTTPS API origin; physical-device localhost points to the device |
Recover without creating a second problem#
For an uncertain send, reconcile the original message before attempting an authorized resend. For a failed AI job, use the saved retry flow so context and staged work can be preserved. For a database migration failure, inspect the failing migration and checksum; create a corrective migration rather than editing applied history.
A pending semantic index should not block normal editing or keyword search. The SQLite search sidecar is rebuildable: with the server stopped, move aside search-v1.sqlite and its WAL/SHM companions, then restart to requeue indexing. Never remove atlas.sqlite, which contains canonical application data. PostgreSQL indexing has its own durable job and source-version checks.
Before considering recovery complete, repeat the original user flow, restart the relevant service, and verify persistence. If the issue involved files or storage, include a restore-verification check rather than relying only on a successful page load.