Table of Contents
Data historians compress data to save storage, keeping only enough points to redraw each signal within a set tolerance. That works for trends, reports, and compliance. It falls short when the signal lives in brief events, such as vibration spikes or early fault signatures, because compression discards those points.
Why do data historians use lossy compression?
Legacy data historians were designed when storage was expensive. Keeping every raw reading from every signal for years wasn’t economical, so historians were optimized to store process data efficiently rather than preserve every point. The standard approach is lossy compression: the historian keeps only the points needed to represent a signal within a defined tolerance and discards the rest.
For the work historians were built to support, that trade is sound. Trending, shift reports, and compliance records need a faithful, auditable approximation of the process, not every reading a sensor produced. A temperature that holds steady for an hour can be represented by a handful of points instead of thousands, and an engineer reviewing the trend sees the same shape.
How do deadband and swinging-door compression work?
Two techniques are common:
- Deadband. The historian ignores a new value unless it differs from the last stored value by more than a set amount (Δ). Fluctuations inside that band are never stored.
- Swinging door. The historian maintains a corridor of acceptable error around the line projected from recently stored points. A new point that falls inside the corridor is discarded, because the line already represents it within tolerance. A point outside the corridor is stored.
Consider a bearing whose vibration develops a small, rapid oscillation in the days before a fault. If each swing stays inside the deadband, the historian stores a flat line, and the early warning never reaches the archive. Both techniques did their job: they kept the signal’s shape, not its detail.
When does lost fidelity become a problem?
Approximation stops being acceptable when the value lies in short-lived or subtle behavior. Digital twins need high-fidelity data to capture transients and dynamic responses. Predictive maintenance, vibration analysis, and anomaly detection look for subtle deviations that compression can smooth over. Highly dynamic processes, with frequent excursions or rapid state changes, force a historian either to retain many more points or to risk smoothing over behavior that matters.
Time series databases typically separate storage from reduction. InfluxDB 3 Enterprise compresses with type-specific encodings, such as Gorilla encoding for floating-point values and delta-delta encoding for timestamps, which encode values rather than discard points (storage engine). When teams want smaller data, they downsample later, for example with a Downsampler plugin from the Processing Engine plugin library, so the decision about what detail to give up comes after they know what matters.
Where approximation works and where it doesn’t
| Use case | Compressed approximation | Why |
|---|---|---|
| Trends and dashboards | Usually sufficient | The shape of the signal is preserved |
| Production and shift reports | Usually sufficient | Summaries tolerate small errors |
| Compliance records | Usually sufficient | An auditable approximation within tolerance |
| Vibration analysis | Often insufficient | The signal lives in rapid oscillation |
| Predictive maintenance and anomaly detection | Often insufficient | Early fault signatures are subtle |
| Digital twins | Often insufficient | Models depend on transient responses |