InfluxDB 3 is a columnar time series database written in Rust and built on Apache Arrow and Apache DataFusion. It separates compute from storage, holding recent data in memory and persisting the rest as compressed columnar files in object storage. The architecture supports millions of writes per second, billions of series with no cardinality limit, and sub-10ms queries on recent data using SQL or InfluxQL.

Why do time series workloads need a purpose-built database?

Time series data arrives constantly and is almost never updated after it lands; for example, a wind farm reporting turbine vibration every hundred milliseconds or Eutelsat OneWeb ingesting one million points per second across more than 600 LEO satellites. Workloads like these write far more than they read, query recent data far more often than old, and add unique series with every asset that comes online.

General-purpose database design optimizes for a different shape. Row storage, mutable records, and index structures sized for a stable set of keys all pay off when reads and writes are balanced, and records change in place. Point those same assumptions at a stream of append-only, timestamped measurements, and the index becomes the constraint long before the disk does.

What are the components of the InfluxDB 3 architecture?

Component Role
Write buffer Validates incoming line protocol, holds the newest writes in memory
Write-ahead log (WAL) Flushes the buffer to object storage so accepted writes survive a restart
Object store Holds sorted, encoded, compressed columnar files on S3-compatible storage, GCS, Azure Blob, MinIO, or local disk
Catalog Tracks schema and file metadata so a query resolves to the specific files holding its data
Query engine Plans and executes SQL and InfluxQL through Apache DataFusion across memory, cache, and object storage
Compactor Merges small files into larger generations so long-range queries open fewer of them
Caches Keep last values and distinct tag values in memory for the most frequent query patterns
Processing Engine Runs Python plugins inside the database on write, on a schedule, or on request

Several of these layers are open Apache components rather than proprietary ones. Apache Arrow defines the in-memory layout, DataFusion plans and executes queries against it, and Arrow Flight SQL moves results across the network without a serialization step. Apache Parquet handles data at rest, persisting it in Core and Cloud products while serving as the export format for Enterprise. InfluxData named that combination the FDAP stack, and its engineers help maintain the projects behind it.

Essential functions of InfluxDB 3

Writes: How does a write move through InfluxDB 3?

Every point follows the same path, and the intervals below are shipping defaults, rather than tuned settings.

  1. Validation. The database checks incoming line protocol and rejects malformed or schema-incompatible points before they enter the system. Accepted points land in an in-memory write buffer.
  2. WAL persistence. Once per second, the write buffer flushes to the write-ahead log in object storage, and the server acknowledges the write. Latency-sensitive pipelines can set no_sync=true to acknowledge before persistence completes.
  3. Query availability. Data moves into the queryable buffer, where the server keeps up to 900 WAL files—roughly 15 minutes of data—available in memory.
  4. Columnar persistence. Every ten minutes, the oldest data in the queryable buffer is written to columnar files in object storage. The most recent five minutes stay in memory.
  5. Caching. Recently persisted files are held in an in-memory cache, so queries against the newest data skip the round trip to object storage.

A point becomes durable within about a second of arrival and queryable immediately after. That gap separates a database you monitor with from one you automate against.

See the InfluxDB 3 data durability documentation for the full write path and the configuration options at each stage.

Storage: How does InfluxDB 3 store data on disk?

Storage format decides what a query gets to skip. InfluxDB 3 Enterprise writes into a columnar format built specifically for time series workloads, sorted by column family key, series key, and timestamp, so the query path drops blocks that cannot satisfy a predicate before reading a value.

Compression is chosen per data type rather than applied uniformly. Timestamps use delta-delta run-length encoding, floats use Gorilla encoding, and low-cardinality strings use dictionary encoding.

Column families group related fields so a query reads only what it asked for. A double-colon delimiter in line protocol assigns the field, so cpu::usage_user, mem::free, and disk::read_bytes land in three separate families. Querying only mem::free reads the mem block and skips the rest. On a table with hundreds of fields, that is the difference between reading a slice of a file and reading all of it.

Schemas can reach millions of columns and take on new ones as they appear, without rewriting the table. Sensor fleets and satellite constellations produce exactly this shape, where every device reports a slightly different set of measurements and any single row is mostly empty. Persistence and compaction run against a fixed memory budget, which keeps resource use predictable during heavy ingest. Compaction merges the small files ingest produces into larger generations, so a query spanning a year opens far fewer files than ingest created.

Product On-disk storage
InfluxDB 3 Enterprise Columnar format built for time series query patterns, default for new clusters
InfluxDB 3 Core Apache Parquet
InfluxDB Cloud Dedicated, Cloud Serverless, and Clustered Apache Parquet

Compacted Enterprise data exports as Parquet for anything that needs to read it outside InfluxDB, so Spark, DuckDB, Snowflake, or a pandas script can reach the same data with no ETL pipeline in between.

Queries: How does InfluxDB 3 query recent and historical data together?

A dashboard showing the current reading and a query comparing this quarter to last year hit the same table with the same syntax. DataFusion builds a single plan spanning the in-memory buffer, the file cache, and object storage, using the catalog to narrow candidate files before any are opened.

Columnar layout does the rest of the pruning. Each column is stored separately with its own statistics, so a query touching two columns of a 200-column table reads two and skips blocks whose value ranges fall outside the predicate.

Two in-memory caches handle the query shapes that repeat constantly:

  1. The Last Value Cache stores the most recent N values for specified fields, which is what a real-time dashboard asks for on every refresh.
  2. Metadata queries listing hostnames or sensor IDs are served by the Distinct Value Cache, which keeps unique tag and field values in memory and returns them without a scan.

InfluxDB 3 Core reports last-value queries returning in up to 10 milliseconds and distinct metadata queries in up to 30 milliseconds.

Results move over Arrow Flight in Arrow’s in-memory format, landing them in BI tools with no conversion step in between.

How does InfluxDB 3 handle high cardinality?

Tags are columns, so connecting a new fleet of assets adds values to an existing column rather than entries to a structure that must be held in memory and searched. Cardinality growth costs storage, not query performance, and storage is cheap. LeoLabs tracks more than 25,000 objects in low Earth orbit on this model.

Retention follows from the same idea: compute scales independently of what sits in object storage, which is how Joby Aviation ingests flight data the moment an aircraft lands and still meets its retention requirements while holding storage costs down.

How does InfluxDB 3 act on data as it arrives?

Downsampling, anomaly detection, and alerting usually live in a separate stream processor that has to be deployed, secured, and kept in sync with the database. The InfluxDB 3 processing engine runs that logic as Python inside the database, triggered by a WAL write, a schedule, or an HTTP request.

Common patterns ship in the official plugin library, including downsampling, MAD anomaly detection, threshold and deadman checks, Prophet forecasting, and export to Apache Iceberg. Plugins are ordinary Python, so a team already running scikit-learn or a custom scoring function can move it next to the data instead of shuttling data out to reach it.

InfluxDB 3 deployment options

Every option runs the same engine and the same APIs. What changes is the operational shape, and each is tuned for a different set of demands.

Deployment Model Best suited for
InfluxDB 3 Core Open source, single node Edge collection, prototypes, and smaller workloads that do not need historical compaction
InfluxDB 3 Enterprise Self-managed, multi-node Production workloads needing high availability, read replicas, compaction, with separate ingest, query, compact, and process nodes
InfluxDB Cloud Serverless Fully managed, shared infrastructure Workloads that fit comfortably on multi-tenant capacity, with no infrastructure to run
InfluxDB Cloud Dedicated Fully managed, single tenant Scaling workloads that need isolation and dedicated resources without an operations team
Amazon Timestream for InfluxDB Fully managed on AWS AWS-standardized teams wanting provisioning, backups, and version management handled natively

Multi-node Enterprise deployments separate roles across nodes, so a heavy ingest period does not compete with query traffic for the same CPU. Single-node Core keeps the whole engine in one binary, which is what makes it practical at the edge.

Learn more about your deployment options for InfluxDB, or reach out if you’d like to discuss your options with our team.

Where architectures get tested

The useful question to ask of any time series database is what happens when the series count multiplies by ten and the retention window doubles. An index-based design has a ceiling, but a columnar design backed by cost-effective object storage scales linearly and can be tuned. InfluxDB was built for the specific needs of time series data and scales to fit any workload.

Frequently asked questions

What is InfluxDB 3 built on?

InfluxDB 3 is written in Rust and built on Apache Arrow for in-memory processing, Apache DataFusion for query planning and execution, and Apache Arrow Flight for moving results across the network. Building on open components means data stays reachable through Flight SQL and readable by other tools without custom integration work.

How does InfluxDB 3 store data?

Incoming writes land in an in-memory buffer, flush to a write-ahead log in object storage each second, and persist to compressed columnar files every ten minutes by default. InfluxDB 3 Enterprise uses a columnar format built for time series query patterns, which is the default for new clusters. Core and the Cloud products persist in Apache Parquet. Files are immutable once written, and a catalog tracks which files hold which data.

Does InfluxDB 3 still use Parquet?

Yes. InfluxDB 3 Core and the Cloud products persist data in Apache Parquet. Enterprise stores data on disk in a columnar format built for time series query patterns and uses Parquet as its export and interchange format, so compacted data can be read by external tools without a proprietary reader.

Does InfluxDB 3 have a cardinality limit?

No. InfluxDB 3 supports unlimited tag cardinality because it stores tags as columns rather than maintaining a series index. Adding unique tag values increases the volume of data stored rather than the size of an in-memory index, so high-cardinality workloads like per-device or per-trace identifiers do not hit a wall.

Which time series workloads does InfluxDB 3 support?

Time series workloads across infrastructure and network monitoring, industrial IoT and predictive maintenance, satellite telemetry, battery energy storage, and industrial AI are excellent applications for InfluxDB 3. Each writes far more than it reads, cares most about the last few minutes, accumulates unique series as it grows, and needs full-resolution history to be worth keeping.