Prometheus
A CNCF open-source monitoring system using pull collection, labels, PromQL, rules, and Alertmanager for service health.
Why use it
It creates explainable metrics for rooms, request latency, queues, databases, and resource use, alerting before incidents expand.
Where it fits
Use analytics, feedback, player support, and updates to fix issues and decide what to improve after launch.
What to check
It is not for detailed events or unlimited long-term storage, and high label cardinality becomes unmanageable; design high availability, remote storage, and alert routing separately.