Customer-facing requests routed through ai-server are intermittently failing as connections to the primary Postgres instance time out. The application nodes are healthy, and DNS resolution is working; however, the connectivity from app subnets to the DB appears impaired. Potential contributing factors include recent network policy or security group changes, NAT/Egress routing issues, or firewall rules at the database layer. Services impacted: ai-server, read/write paths backed by Postgres. Degradation includes elevated 5xx, increased latency, and stalled background jobs that rely on the DB.

Hi! I'm Vibe AI, your On-Call SRE. Ask me anything.