Pulsegrid is the observability layer for GPU fleets and training runs — GPU health, job progress, and burn rate, unified in one dashboard your on-call engineer will actually open.
Trusted by infrastructure teams at
What you get
Every H100, A100, and MI300 across every cluster, in one pane, refreshed every two seconds. Zoom from datacenter down to a single die temperature.
Follow a run from queue to checkpoint with loss curves, step throughput, and gradient norms streamed inline — no more tailing logs over SSH.
Every GPU-hour split by team, project, and model so finance stops guessing and your Q3 budget review stops being a fistfight.
We flag thermal throttling, NCCL stalls, and slow OOM creep before your run does — usually before Slack does too.
Compare on-prem racks, AWS, and neoclouds side by side. One dashboard, whatever hardware you actually run on this month.
Slack, PagerDuty, and webhooks with noise-aware thresholds that learn your baseline instead of paging you at 3am for nothing.
How it works
Drop a single lightweight collector on your nodes. Pulsegrid does the rest — no manual dashboards, no YAML spelunking.
Pricing
For a first cluster
For teams shipping models
For serious fleets
Fourteen minutes from `curl` to your first dashboard. No credit card, no sales call, no Terraform module you'll regret.