Prometheus
A CNCF open-source monitoring system using pull collection, labels, PromQL, rules, and Alertmanager for service health.
Why it’s here
It creates explainable metrics for rooms, request latency, queues, databases, and resource use, alerting before incidents expand.
Best fit
Analytics, feedback, support, and updates keep a released game healthy over time.
Before you commit
It is not for detailed events or unlimited long-term storage, and high label cardinality becomes unmanageable; design high availability, remote storage, and alert routing separately.