I’m curious as to which tools and technologies you all are using to keep track of all those services you are deploying, whether it be resource tracking, network traffic, logs, traces, or uptime.
As a bonus question, how have you organized your network or your services to reduce the overhead of implementing observability?


Things usually break on updates unless your disk fills up or worse some hardware fails…
So it’s pretty easy to fix: never upgrade automatically.
I always upgrade manually and check immediately what I have upgraded. Works fine so far.
And I was only partially joking… Yeah if a service stays broken for a long time, just uninstall it, it was useless.