I’m curious as to which tools and technologies you all are using to keep track of all those services you are deploying, whether it be resource tracking, network traffic, logs, traces, or uptime.
As a bonus question, how have you organized your network or your services to reduce the overhead of implementing observability?


Observawhat? My services just exist, if they fail and nobody complains… Well, they get removed. If the fail and somebody complains… The get restarted…
Pretty efficient but simple approach.
I got tired of only noticing something was broken when I needed to use it instead of when I was working on things. Now if I make a change I notice issues right away and can fix them.
Things usually break on updates unless your disk fills up or worse some hardware fails…
So it’s pretty easy to fix: never upgrade automatically.
I always upgrade manually and check immediately what I have upgraded. Works fine so far.
And I was only partially joking… Yeah if a service stays broken for a long time, just uninstall it, it was useless.