Table of Contents
The FDAP stack is a set of open source Apache components for building analytic data systems. It combines Apache Arrow Flight, Apache DataFusion, Apache Arrow, and Apache Parquet. InfluxData coined the term in 2023 and built InfluxDB 3 using it. Arrow defines the in-memory data format, Flight moves that data across the network, DataFusion plans and executes queries against it, and Parquet handles durable columnar exchange with other systems.
What does FDAP stand for?
| Component | Layer | What it does |
|---|---|---|
| F for Apache Arrow Flight | Transport | Moves Arrow data over the network without serializing it into another format first |
| D for Apache DataFusion | Query engine | Parses, plans, optimizes, and executes SQL against Arrow data |
| A for Apache Arrow | In-memory format | A columnar memory layout that every other component reads and writes natively |
| P for Apache Parquet | Interchange & bulk import format | Compressed columnar files that other tools and query engines can read directly |
The letters are ordered for pronunciation rather than for how data flows. In practice, Arrow comes first, because the other three are built around it.
Where did the term come from?
When Paul Dix announced the InfluxDB IOx project, he committed the rebuild of InfluxDB to four Apache projects that were, at the time, either young or unproven for this purpose. DataFusion in particular was still largely the work of its creator, Andy Grove, in his spare time.
The original FDAP architecture post argued that these components would do for analytic systems what LAMP did for interactive websites, giving builders a shared foundation so that engineering effort goes into what makes a system different rather than into re-implementing a query optimizer for the fifth time.
That bet has held up with DataFusion becoming a top-level Apache Software Foundation project, and it now underpins a growing number of analytic databases and data tools beyond InfluxDB. InfluxData engineers like Andrew Lamb remain deeply involved in both projects as PMC members of Apache Arrow and Apache DataFusion.
What does each component do?
Apache Arrow
Arrow is a columnar in-memory format. Every system that speaks Arrow represents a batch of records the same way in ram, which means data can pass between a query engine, a client library, and an analysis tool without anyone paying a conversion cost.
Columnar layout also suits how analytic queries actually read data. A query that touches three columns out of two hundred reads three columns, and modern CPUs can process those values in vectorized batches rather than row by row.
Apache Arrow Flight
Flight is an RPC framework for moving Arrow data across a network. Because both ends already hold data in Arrow, a Flight transfer skips the serialize-and-deserialize step that dominates the cost of most database wire protocols.
Flight SQL adds a standard interface for submitting SQL and retrieving results, so a client written for one Flight SQL database works against another with minimal change. InfluxDB 3 exposes Flight SQL alongside HTTP APIs, which is how the Python, Go, Java, C#, JavaScript, and Rust client libraries all reach the same query path.
Apache DataFusion
DataFusion is a query engine written in Rust that uses Arrow as its memory model. It handles SQL parsing, logical planning, optimization, and vectorized execution, and it exposes extension points at every stage.
Those extension points matter more than raw execution speed. Time series work raises planning problems a general SQL engine never has to solve. Overlapping data has to be deduplicated as it arrives, and time filters have to push deep enough into the plan to eliminate entire files before they are read. DataFusion lets InfluxDB supply its own operators and optimizer rules for that work while inheriting everything else. Andrew Lamb of InfluxData describes DataFusion as LLVM for databases. LLVM is the reusable compiler backend that Clang, Rust, and Swift are all built on, so the people designing those languages could focus on language design instead of writing an optimizer and code generator from scratch. DataFusion serves the same purpose for analytic systems. You bring the domain knowledge, and the framework handles the machinery every query engine needs.
Apache Parquet
Parquet is a compressed columnar file format that almost every analytic tool can read. Within the FDAP stack it serves as the interchange layer, the format you reach for when data needs to leave one system and be understood by another without an ETL pipeline in between.
Data exported from InfluxDB 3 as Parquet can be queried by Spark, DuckDB, Snowflake, Presto, or a pandas script, with no proprietary reader and no license attached to the files.
How does InfluxDB 3 use the FDAP stack?
Writes arrive as line protocol over HTTP. The database validates them against the catalog, records them in a write ahead log for durability, and holds them in memory in Arrow format so they are queryable immediately rather than after the next flush. Background compaction then organizes persisted data into larger files that queries can read efficiently.
Queries arrive as SQL or InfluxQL where DataFusion builds a plan that reads across both persisted files and the in-memory buffer, so a query covering the last ten seconds and the last ten months resolves as one plan rather than two. Results come back as Arrow and travel over Flight or HTTP.
The practical effect is a narrow gap between a measurement landing and that measurement being available to act on. InfluxDB 3 Core reports last-value queries returning in up to 10 milliseconds.
How does InfluxDB 3 Enterprise store data on disk?
InfluxDB 3 Enterprise writes data into a columnar file format built specifically for time series workloads. Within each file, data is sorted by column family key, series key, and timestamp, which lets the query path skip entire blocks that cannot satisfy a predicate before reading any values.
Compression is chosen per data type rather than applied uniformly. Timestamps use delta-delta run-length encoding, floats use Gorilla encoding, and low-cardinality strings use dictionary encoding.
Column families
Column families group related fields so a query reads only the fields it asked for. You assign a field to a family with a double-colon delimiter in line protocol, where the portion before the delimiter names the family.
metrics,host=sA cpu::usage_user=55.2,cpu::usage_sys=12.1,cpu::usage_idle=32.7 1000000000
metrics,host=sA mem::free=2048i,mem::used=6144i,mem::cached=1024i 1000000000
metrics,host=sA disk::read_bytes=50000i,disk::write_bytes=32000i 1000000000
That produces three families.
| Family | Fields |
|---|---|
| cpu | usage_user, usage_sys, usage_idle |
| mem | free, used, cached |
| disk | read_bytes, write_bytes |
A query referencing only mem::free reads the mem block and skips cpu and disk entirely. On a table with hundreds of fields, that is the difference between reading a slice of a file and reading all of it. Only the first delimiter is significant, so a field named a::b::c creates family a holding field b::c. Fields written without a delimiter are assigned to auto-generated families holding up to 100 fields each.
Wide and sparse schemas
Schemas can reach millions of columns and evolve as new columns appear, without rewriting the table. Sensor fleets, satellite constellations, and industrial telemetry tend to produce exactly this shape, where each device reports a slightly different set of measurements and the union across the fleet is enormous while any single row is mostly empty.
Persistence and compaction run against a fixed memory budget rather than growing with the size of the work, which keeps resource use predictable during heavy ingest.
Which products this applies to
| Product | On-disk storage |
|---|---|
| InfluxDB 3 Enterprise | Native columnar format described above, default for new clusters |
| InfluxDB 3 Core | Apache Parquet |
| InfluxDB Cloud Dedicated, Cloud Serverless, Clustered | Apache Parquet |
Where does Parquet fit today?
Parquet is the export and interchange format for InfluxDB 3 Enterprise. Compacted data can be exported as Parquet files for use in external tools, one 24-hour window at a time or across a range.
influxdb3 export data \
--database mydb \
--table cpu \
--output-dir ./export_output
Data has to be compacted before it can be exported. Uncompacted data is not currently available for export.
This split lets each format do what it is good at. A format tuned for time series query patterns handles the hot path, and an open, universally supported format handles the moment data needs to be read by something that is not InfluxDB.
Why build on the FDAP stack?
| Outcome | How the stack delivers it |
|---|---|
| Engineering effort goes to differentiated work | Query planning, vectorized execution, and columnar encoding come from shared components, so teams build domain features instead of rebuilding an optimizer |
| Data passes between layers without being rewritten | Arrow in memory, Flight on the wire, and Parquet on disk share one columnar model, so moving data between them takes no serialization step and no format conversion |
| Other tools can read the data | Parquet files and Flight SQL endpoints are open standards, so BI tools, notebooks, and warehouses connect without custom adapters |
| Improvements compound from outside | Every organization building on DataFusion contributes optimizations that flow back to everyone using it |
| Governance is predictable | All four projects sit in the Apache Software Foundation, with published decision-making and no single-vendor control |
What does this look like at scale?
Eutelsat OneWeb runs real-time telemetry across more than 600 low Earth orbit satellites, ingesting one million points per second. LeoLabs tracks more than 25,000 objects in low Earth orbit. Seadrill moved to condition-based maintenance on its rigs and saved $55 million in asset lifecycle costs.
Those workloads share a shape. Data arrives continuously, the schema is wide and irregular, and the value of a reading decays quickly, which is the case the stack was assembled to serve.