On May 12th and 13th at various times, a subset of Upstash Redis instances on Fly.io experienced intermittent hangs and elevated error rates. The Redis process would stall inside a logging syscall — alive but not making progress — which made the issue hard to spot from our usual telemetry. After investigating with Fly's team, we identified the root cause as a bad interaction between a recent guest kernel update on Fly's newer machines and an upstream Cloud Hypervisor bug (cloud-hypervisor#7672) affecting log writes from inside the VM. We mitigated by disabling the affected logging paths, and Fly has since rolled out a hypervisor-side patch, fully resolving the issue. No data was lost. Sorry for the disruption.
Posted May 15, 2026 - 12:49 UTC
Resolved
The incident has been resolved. We are working with Fly team on RCA.
Posted May 11, 2026 - 18:22 UTC
Update
We are working with Fly team to investigate the root cause.
Posted May 11, 2026 - 17:19 UTC
Update
We are continuing to investigate the issue.
Posted May 11, 2026 - 15:07 UTC
Investigating
Some databases may experience increased latency or timeouts in Fly.io’s FRA region.