The Physical Layer of AI: Monitoring Power, Heat, and GPU Health at Scale
Session Date: Aug 11, 2026
Time: 8:00am (PT) | 3:00pm (GMT) | 4:00pm (BST)
AI datacenters generate telemetry across layers that most monitoring stacks weren’t designed to handle together: GPU-level metrics (utilization, memory, temperature, ECC errors) and facility-level signals (power switch data, power draw, cooling load, thermal thresholds). Teams end up stitching together separate tools for each layer, losing the correlation between a GPU throttling event and the power or thermal conditions that caused it.
In this session, Ian Clark, Senior Sales Engineer at InfluxData, shows how InfluxDB unifies GPU and facility telemetry—including power switch data—into a single time series platform, capturing high-frequency metrics from GPU fleets and datacenter infrastructure in real time and querying across both layers without the cardinality or retention tradeoffs that fragment most monitoring setups.
Ian will cover
- How GPU, power, and facility metrics differ in shape, frequency, and cardinality, and why that breaks siloed monitoring
- A live demo of GPU, power switch, and thermal data flowing into InfluxDB
- How teams correlate hardware health with power and facility conditions to catch failures before they cascade