< Back to customers

How ITER Monitors the Infrastructure Behind the World’s Largest Fusion Project with InfluxDB

ITER is an international fusion research facility under construction at Saint-Paul-lez-Durance in southern France, funded by seven members of the international community. Fusion is the nuclear reaction that powers the sun and stars, and a promising long-term option for sustainable, non-carbon-emitting energy. The tokamak at its center is an experimental device designed to produce 500 MW of fusion power from 50 MW of heating power injected into the plasma, a tenfold gain known as Q=10.

REGION

Europe

INDUSTRY

Scientific Research / Nuclear Fusion

Start building with InfluxDB

Start exploring InfluxDB and bring high-performance time series analytics to your applications.

Try InfluxDB

BUSINESS IMPACT

180

plant systems feeding the control network at full scale

5 years

telemetry retained and queryable

1 ns

timestamp resolution

Overview

Monitoring infrastructure built for nanosecond precision

Every team at ITER, from magnet testing to the experimental data chain, depends on the control system’s IT infrastructure. Servers, hypervisors, shared storage, and the networks between them all have to stay healthy through commissioning tests that can take significant preparation. When a server runs out of memory or a network link starts dropping packets, the control system group has to catch it before the teams relying on that infrastructure feel it.

Monitoring that infrastructure is inherently a time series problem. Readings arrive continuously from hundreds of machines, and every query is time bound, from one host over the last hour to the whole site over the last 90 days. InfluxDB 3 Enterprise, a purpose-built time series platform, ingests that telemetry, keeps it for five years, and stores the data at nanosecond resolution. Plasma diagnostic and experimental data travel through a separate data acquisition and archiving system, which Lana Abadie, Senior Database Engineer at ITER, owns. InfluxDB 3 Enterprise watches the infrastructure underneath it and across the control system, collecting CPU, memory, load, packet counts, and network card diagnostics. The team also tracks NFS, the network file system that lets machines across the site read and write shared storage as if it were local. Grafana sits on top, and its dashboards and alerts are how the control system group sees whether that infrastructure is healthy.

InfluxDB open source v1 carried that monitoring for five years. Then the deployment outgrew its server’s memory and began crashing about once a month, taking the dashboards and alerts down with it.

Challenge

Monitoring infrastructure for a machine still being built

ITER is being assembled and commissioned at the same time. At full scale, about 180 plant systems will feed the CODAC (control, data access, and communication) network, covering cryogenics, cooling water, magnets, and dozens of other subsystems. They come online in stages over years of commissioning ahead of Start of Research Operation (SRO), the milestone in ITER’s 2024 baseline when the research program begins. Each one needs monitoring from the day it first receives power, because the first hours under power are when problems show up. Today the team supports the magnet test facility and gyrotron commissioning teams.

Growth is what broke the old setup. Cardinality was the root cause: each new host, interface, and process adds series to track, and as more arrived, the v1 instance needed more memory than the server could give it. For a team whose job is keeping the infrastructure under every other system healthy, each crash was a serious problem. Lana relies on the same monitoring to track the health of her data chain, and her dense, multi-panel dashboards were views the previous deployment couldn’t render.

Xavier Mocquard, Control & Data Access Servers Engineer at ITER, runs the control system infrastructure and describes the work as IT with a different set of tolerances. His team supports storage, virtualization, directory services, and configuration management, like any large IT organization. What sets it apart is the bar for availability and performance, and a timing standard set in nanoseconds. ITER runs a precision time protocol (PTP) grandmaster clock fed by GPS antennas on the roof, and distributes a nanosecond-accurate signal across the control network. Fast controllers sitting in different buildings stay within 50 ns root mean square (RMS) of each other.

That precision is what makes the experimental data scientifically useful. Diagnostic systems watching an event inside the plasma sit in separate buildings and record it independently. If their clocks disagree, the recorded order of events is wrong, and a team can’t separate cause from effect.

We are doing exactly the same thing that IT is doing, but with a lot of constraints and a very high expectation about availability and performance. We are not doing something at the millisecond. We are at the nanosecond.

Xavier Mocquard

Control & Data Access Servers Engineer, ITER Organization

In looking for a replacement, native nanosecond timestamps were a non-negotiable requirement, and few databases offered them. Some options stored timestamps only in a whole second, which ruled them out. For example, Graphite, which stores timestamps in whole seconds, fell outside that requirement. Colleagues suggested ClickHouse, and the team evaluated it before staying with InfluxDB for the size of its install base and the depth of the ecosystem around it. Lana came to ITER from CERN and Xavier from CEA, and both checked with engineers running comparable data chains at other large scientific facilities first.

We try not to reinvent the wheel. Usually, before we make a choice, we go to our colleagues outside the organization, and we compare.

Xavier Mocquard

Control & Data Access Servers Engineer, ITER Organization

Moving to a commercial license followed the same long-term thinking as the rest of ITER’s infrastructure. ITER runs Red Hat Enterprise Linux across roughly 90 percent of its estate and buys hardware from vendors who certify chipset and driver support against a kernel that can’t be updated every few years. Xavier wanted the same footing for the time series layer, because ITER will operate for decades once research operations begin.

When I say partner, it is not a supplier, because we need to build a very strong relationship. We will run for the next 20 years.

Xavier Mocquard

Control & Data Access Servers Engineer, ITER Organization

Enter InfluxDB

A lightweight monitoring stack built for critical infrastructure

The systems ITER monitors can least afford interference from monitoring itself, so the team kept the collection layer lightweight. Real-time fast controllers run Collectd, sampling every 10 to 20 seconds with minimal CPU and memory overhead. Telegraf agents sample every 5 to 10 seconds through plugins, with hundreds carrying the Collectd telemetry back to InfluxDB.

Telegraf also serves as a collection gateway for vendor equipment, filtering and shaping device telemetry before it reaches InfluxDB. When ITER deployed HPE Alletra storage, for example, Xavier required its telemetry to feed into the same monitoring environment.

Several hundred nodes report core host metrics such as CPU, memory, and load, alongside measurements specific to ITER’s data chain. Those include reader and writer counts, archived and lost packet counts, NFS performance, and network card diagnostics. When packets go missing, that level of detail helps the team determine whether the loss occurred in the operating system kernel or on the network card, two problems that require different fixes.

ITER runs InfluxDB 3 Enterprise self-managed on HPE servers with NVMe storage, configured through Ansible playbooks like everything else in the control system. The team retains five years of telemetry, with most queries focused on a 90-day window, while Grafana provides dashboards and alerts for infrastructure health.

diagram-1-iter

The dashboards ITER’s engineers actually watch

diagram-2-iter

diagram-3-iter

diagram-4-iter

Result

No crashes, and the heaviest dashboards now load reliably

ITER’s first success criterion was straightforward: stop the crashes. InfluxDB 3 Enterprise eliminated the recurring out-of-memory failures, while preserving the nanosecond timestamp support the team needed.

Dense, multi-panel dashboards that had become unreliable on the previous deployment now respond consistently, while the team retains five years of telemetry and can query it reliably.

Now we can monitor our system reliably, and we can have very advanced dashboards. It is much more robust, much more reliable, and the performance is great.

Lana Abadie

Senior Database Engineer, ITER Organization

What’s Next

Building for high availability as ITER scales

Most of ITER’s 180 plant systems have yet to come online, and each one adds nodes and series to the monitoring layer as the site moves toward Start of Research Operation. Today the stack runs on a single server, so Xavier’s priority for 2027 is high availability, keeping monitoring and alerting available even if the underlying hardware fails.

ITER plans to build that expansion on InfluxDB 3 Enterprise.