System Updates

Incident Alert

Incident Summary:

  • Downtime Window: 7:02 AM – 9:15 AM EDT (2h 13m)
  • Impact: Application degradation and intermittent 502 Bad Gateway / 504 Gateway Timeout errors affecting device queries and dashboard access.
  • Status: Resolved

Timeline & Root Cause Analysis:

  • 7:02 AM: Stalled overnight index creation led to severe database lock contention and connection pool exhaustion, causing application requests to queue and time out.
  • Mitigation Phase 1: Initial service restarts and attempts to revert the indexing changes failed to clear the locked queries and backlog.
  • Mitigation Phase 2: A restore from a healthy database snapshot was executed. However, the initial instance size was under-provisioned to absorb the spike in queued traffic and connection retries.
  • Resolution (9:15 AM): The database instance was resized to a higher tier to provide necessary compute and I/O capacity. All server environment configs were re-verified, and a coordinated restart across all application hosts brought all services back online.

Current Status:

All services are operating normally, query response times have returned to baseline, and error rates have cleared.

THERE ARE NO KNOWN CURRENT ISSUES

For updates please continue to check our website at allbridge.com/updates.

For support, please contact the appropriate number below.