← Back to match
GPUd
Self-contained daemon for monitoring, diagnosing, and identifying issues on GPU clusters
PlatformWebfreeglobal
A self-contained daemon that monitors and diagnoses GPU health and performance on Linux systems using NVIDIA GPUs, with integration for Docker, containerd, and Kubernetes. It helps teams running large GPU clusters detect issues before workloads fail and provides a unified view of key GPU metrics while keeping CPU and memory overhead low, aimed at ML infrastructure and platform engineering teams.
Categories
Infrastructure MonitoringGPU Monitoring

