Datadog collects infrastructure metrics, logs, traces, and other signals so teams can investigate services and configure alerts.
Kubernetes monitoring tools for production clusters
Kubernetes monitoring must follow workloads that reschedule, scale and disappear. Start with cluster state, node health, application telemetry and the metadata needed to connect an incident to a namespace, workload, pod and deployment change.
For SREs, platform engineers and DevOps teams operating Kubernetes in production.
The decision to make
Instrument one representative cluster before selecting a platform. Measure collection overhead, cardinality, retention, query latency and alert quality while workloads scale and roll out.
9 matching tools
| Tool | Deployment | Source | API | Pricing model |
|---|---|---|---|---|
| Datadog | Web | Other license | Available | paid |
| Dynatrace | Self-hosted | Other license | Available | paid |
| Grafana | Self-hosted | Open source | Available | freemium |
| Netdata | Self-hosted | Open source | Not specified | freemium |
| New Relic | See profile | Other license | Not specified | freemium |
| OpenObserve | Self-hosted | Open source | Available | open source |
| Prometheus | Self-hosted | Open source | Available | open source |
| Sentry | Self-hosted | Open source | Available | freemium |
| SigNoz | Self-hosted | Open source | Not specified | freemium |
Dynatrace brings application and infrastructure telemetry into an observability platform, with capabilities governed by the selected subscription.
Grafana is an open source visualization and dashboarding platform that queries metrics, logs, and traces from Prometheus, Loki, Tempo, and dozens of other backends, with Grafana Cloud offering a hosted option.
Netdata collects system and service metrics and evaluates alerts near the monitored nodes.
New Relic brings application, infrastructure and log signals into observability workflows.
OpenObserve natively unifies logs, metrics, distributed traces, and real-user monitoring in a single self-hostable platform. Written in Rust, it stores data at a fraction of the cost of Elasticsearch.
Prometheus scrapes time-series metrics from instrumented targets, stores them locally, and evaluates alerting rules with PromQL, forwarding alerts to Alertmanager. It has no built-in dashboarding or long-term storage layer of its own.
Sentry provides error tracking, performance monitoring and session replay for web and mobile apps, open source and self-hostable or available as a hosted freemium service.
SigNoz unifies logs, metrics and traces natively on OpenTelemetry, available self-hosted for free or as SigNoz Cloud starting at $49/month.