The FDAP stack is a set of open source Apache components for building analytic data systems. It combines Apache Arrow Flight, Apache DataFusion, Apache Arrow, and Apache Parquet. InfluxData coined the term in 2023 and built InfluxDB 3 using it. Arrow defines the in-memory data format, Flight moves that data across the network, DataFusion plans and executes queries against it, and Parquet handles durable columnar exchange with other systems.

What does FDAP stand for?

Component Layer What it does
F for Apache Arrow Flight Transport Moves Arrow data over the network without serializing it into another format first
D for Apache DataFusion Query engine Parses, plans, optimizes, and executes SQL against Arrow data
A for Apache Arrow In-memory format A columnar memory layout that every other component reads and writes natively
P for Apache Parquet Interchange & bulk import format Compressed columnar files that other tools and query engines can read directly

The letters are ordered for pronunciation rather than for how data flows. In practice, Arrow comes first, because the other three are built around it.

Where did the term come from?

When Paul Dix announced the InfluxDB IOx project, he committed the rebuild of InfluxDB to four Apache projects that were, at the time, either young or unproven for this purpose. DataFusion in particular was still largely the work of its creator, Andy Grove, in his spare time.

The original FDAP architecture post argued that these components would do for analytic systems what LAMP did for interactive websites, giving builders a shared foundation so that engineering effort goes into what makes a system different rather than into re-implementing a query optimizer for the fifth time.

That bet has held up with DataFusion becoming a top-level Apache Software Foundation project, and it now underpins a growing number of analytic databases and data tools beyond InfluxDB. InfluxData engineers like Andrew Lamb remain deeply involved in both projects as PMC members of Apache Arrow and Apache DataFusion.

What does each component do?

Apache Arrow

Arrow is a columnar in-memory format. Every system that speaks Arrow represents a batch of records the same way in ram, which means data can pass between a query engine, a client library, and an analysis tool without anyone paying a conversion cost.

Columnar layout also suits how analytic queries actually read data. A query that touches three columns out of two hundred reads three columns, and modern CPUs can process those values in vectorized batches rather than row by row.

Apache Arrow Flight

Flight is an RPC framework for moving Arrow data across a network. Because both ends already hold data in Arrow, a Flight transfer skips the serialize-and-deserialize step that dominates the cost of most database wire protocols.

Flight SQL adds a standard interface for submitting SQL and retrieving results, so a client written for one Flight SQL database works against another with minimal change. InfluxDB 3 exposes Flight SQL alongside HTTP APIs, which is how the Python, Go, Java, C#, JavaScript, and Rust client libraries all reach the same query path.

Apache DataFusion

DataFusion is a query engine written in Rust that uses Arrow as its memory model. It handles SQL parsing, logical planning, optimization, and vectorized execution, and it exposes extension points at every stage.

Those extension points matter more than raw execution speed. Time series work raises planning problems a general SQL engine never has to solve. Overlapping data has to be deduplicated as it arrives, and time filters have to push deep enough into the plan to eliminate entire files before they are read. DataFusion lets InfluxDB supply its own operators and optimizer rules for that work while inheriting everything else. Andrew Lamb of InfluxData describes DataFusion as LLVM for databases. LLVM is the reusable compiler backend that Clang, Rust, and Swift are all built on, so the people designing those languages could focus on language design instead of writing an optimizer and code generator from scratch. DataFusion serves the same purpose for analytic systems. You bring the domain knowledge, and the framework handles the machinery every query engine needs.

Apache Parquet

Parquet is a compressed columnar file format that almost every analytic tool can read. Within the FDAP stack it serves as the interchange layer, the format you reach for when data needs to leave one system and be understood by another without an ETL pipeline in between.

Data exported from InfluxDB 3 as Parquet can be queried by Spark, DuckDB, Snowflake, Presto, or a pandas script, with no proprietary reader and no license attached to the files.

How does InfluxDB 3 use the FDAP stack?

Writes arrive as line protocol over HTTP. The database validates them against the catalog, records them in a write ahead log for durability, and holds them in memory in Arrow format so they are queryable immediately rather than after the next flush. Background compaction then organizes persisted data into larger files that queries can read efficiently.

Queries arrive as SQL or InfluxQL where DataFusion builds a plan that reads across both persisted files and the in-memory buffer, so a query covering the last ten seconds and the last ten months resolves as one plan rather than two. Results come back as Arrow and travel over Flight or HTTP.

The practical effect is a narrow gap between a measurement landing and that measurement being available to act on. InfluxDB 3 Core reports last-value queries returning in up to 10 milliseconds.

How does InfluxDB 3 Enterprise store data on disk?

InfluxDB 3 Enterprise writes data into a columnar file format built specifically for time series workloads. Within each file, data is sorted by column family key, series key, and timestamp, which lets the query path skip entire blocks that cannot satisfy a predicate before reading any values.

Compression is chosen per data type rather than applied uniformly. Timestamps use delta-delta run-length encoding, floats use Gorilla encoding, and low-cardinality strings use dictionary encoding.

Column families

Column families group related fields so a query reads only the fields it asked for. You assign a field to a family with a double-colon delimiter in line protocol, where the portion before the delimiter names the family.

metrics,host=sA cpu::usage_user=55.2,cpu::usage_sys=12.1,cpu::usage_idle=32.7 1000000000
metrics,host=sA mem::free=2048i,mem::used=6144i,mem::cached=1024i 1000000000
metrics,host=sA disk::read_bytes=50000i,disk::write_bytes=32000i 1000000000

That produces three families.

Family Fields
cpu usage_user, usage_sys, usage_idle
mem free, used, cached
disk read_bytes, write_bytes

A query referencing only mem::free reads the mem block and skips cpu and disk entirely. On a table with hundreds of fields, that is the difference between reading a slice of a file and reading all of it. Only the first delimiter is significant, so a field named a::b::c creates family a holding field b::c. Fields written without a delimiter are assigned to auto-generated families holding up to 100 fields each.

Wide and sparse schemas

Schemas can reach millions of columns and evolve as new columns appear, without rewriting the table. Sensor fleets, satellite constellations, and industrial telemetry tend to produce exactly this shape, where each device reports a slightly different set of measurements and the union across the fleet is enormous while any single row is mostly empty.

Persistence and compaction run against a fixed memory budget rather than growing with the size of the work, which keeps resource use predictable during heavy ingest.

Which products this applies to

Product On-disk storage
InfluxDB 3 Enterprise Native columnar format described above, default for new clusters
InfluxDB 3 Core Apache Parquet
InfluxDB Cloud Dedicated, Cloud Serverless, Clustered Apache Parquet

Where does Parquet fit today?

Parquet is the export and interchange format for InfluxDB 3 Enterprise. Compacted data can be exported as Parquet files for use in external tools, one 24-hour window at a time or across a range.

influxdb3 export data \
  --database mydb \
  --table cpu \
  --output-dir ./export_output

Data has to be compacted before it can be exported. Uncompacted data is not currently available for export.

This split lets each format do what it is good at. A format tuned for time series query patterns handles the hot path, and an open, universally supported format handles the moment data needs to be read by something that is not InfluxDB.

Why build on the FDAP stack?

Outcome How the stack delivers it
Engineering effort goes to differentiated work Query planning, vectorized execution, and columnar encoding come from shared components, so teams build domain features instead of rebuilding an optimizer
Data passes between layers without being rewritten Arrow in memory, Flight on the wire, and Parquet on disk share one columnar model, so moving data between them takes no serialization step and no format conversion
Other tools can read the data Parquet files and Flight SQL endpoints are open standards, so BI tools, notebooks, and warehouses connect without custom adapters
Improvements compound from outside Every organization building on DataFusion contributes optimizations that flow back to everyone using it
Governance is predictable All four projects sit in the Apache Software Foundation, with published decision-making and no single-vendor control

What does this look like at scale?

Eutelsat OneWeb runs real-time telemetry across more than 600 low Earth orbit satellites, ingesting one million points per second. LeoLabs tracks more than 25,000 objects in low Earth orbit. Seadrill moved to condition-based maintenance on its rigs and saved $55 million in asset lifecycle costs.

Those workloads share a shape. Data arrives continuously, the schema is wide and irregular, and the value of a reading decays quickly, which is the case the stack was assembled to serve.

Frequently asked questions

What does FDAP stand for?

FDAP stands for Flight, DataFusion, Arrow, and Parquet, four Apache Software Foundation projects used together to build analytic data systems. Arrow provides the columnar in-memory format, Flight moves Arrow data over the network, DataFusion executes SQL queries against it, and Parquet stores and exchanges columnar data on disk.

Who created the FDAP stack?

InfluxData coined the term. Paul Dix chose the four components for the InfluxDB IOx rebuild that became InfluxDB 3, and the name was introduced in an InfluxData engineering post that compared the stack's role in analytic systems to LAMP's role in web applications. The four projects themselves are independently governed by the Apache Software Foundation.

Does InfluxDB still use Parquet?

Yes. InfluxDB 3 Enterprise uses Parquet as its export and interchange format, so compacted data can be read by external tools. InfluxDB 3 Core and the Cloud products persist data in Parquet. Enterprise stores data on disk in a columnar format built for time series query patterns.

What is the difference between Apache Arrow and Apache Parquet?

Arrow is an in-memory format and Parquet is an on-disk format. Arrow is optimized for fast access and processing while data sits in ram, with no compression getting in the way. Parquet is optimized for size and durability on disk, using compression and encoding that trade some read speed for much smaller files.

Is the FDAP stack only useful for time series data?

No. The components are general purpose and are used in analytic databases, ETL pipelines, data lake query engines, and machine learning tooling. Time series happens to be a demanding test of the stack, since it combines continuous ingest, high cardinality, and queries that span both very recent and very old data.

Is the FDAP stack open source?

Yes. Arrow, Arrow Flight, DataFusion, and Parquet are all Apache Software Foundation projects released under the Apache 2.0 license. Governance runs through project management committees with published decision-making processes, and no single company controls any of them.

What makes DataFusion different from other SQL query engines?

DataFusion is designed to be embedded and extended rather than run as a finished database. It exposes extension points across parsing, logical planning, optimization, and physical execution, which lets a system like InfluxDB add its own operators and rules for time series work while inheriting the general SQL machinery.