The disk was filling up. It was the logs
System journals had grown to 4 GB and kept going. Free space was shrinking and the applications had nothing to do with it.
4.0 GB → 31 MB, growth stopped, no restarts
Read the write-up
Four write-ups, each on its own page: the context, what I measured, what I did, how I verified it and what was left open. The last one is a write-up of my own mistake.
System journals had grown to 4 GB and kept going. Free space was shrinking and the applications had nothing to do with it.
4.0 GB → 31 MB, growth stopped, no restarts
Read the write-up
There were no metrics at all; service health was judged by whether the site opened.
A separate observation server; production exposed no new ports
Read the write-up
Two independent watchers reported the same thing, and some messages carried no cause at all.
Seven duplicates removed after coverage was proven
Read the write-up
The observation host lost network connectivity and declared five healthy sites down.
False alarms of this class stopped
Read the write-up
And what remains afterwards
In every case the result is not only a fixed system but also a document you can re-read six months later.