Build goes red. You look at it for about four seconds, decide it's probably nothing, and hit rerun. It goes green. You move on. I do this. Most people I work with do this. And the uncomfortable thing…
Deep technical articles, product updates, and infrastructure thinking from the Frigga engineering team. Written by engineers who manage real production systems.
Build goes red. You look at it for about four seconds, decide it's probably nothing, and hit rerun. It goes green. You move on. I do this. Most people I work with do this. And the uncomfortable thing…

On 19 October 2025, at 11:48pm Pacific, a race condition in DynamoDB's DNS automation wiped the DNS records for the service in us-east-1. Two processes ran concurrently. A stale plan check let an…

At 11:20 UTC on 18 November 2025, Cloudflare's core network started failing. Within minutes X, ChatGPT, Spotify and Canva were returning errors, and a decent slice of the internet stopped working.…

There is a gap in incident response that gets almost no attention, and it sits in an awkward place. It is not detection. Monitoring is good now. It is not diagnosis either, which is where most of the…

Watch the first two minutes of any investigation and you will notice that almost none of it is investigation. An alert fires for a service. Before anybody can reason about what went wrong, somebody…

The thing nobody warns you about with your first few incidents isn't the pressure. It's the tab count. You start with an alert. By minute four you've got Grafana open, a terminal running kubectl ,…
