Personal
dpv1 beside a laptop
An M.2 accelerator in a Thunderbolt enclosure with 1 GB of local resident memory.
Data. Accelerated.
dau maps high-value Polars, dataframe, time-series, and analytical operations into reconfigurable FPGA dataflow while the rest stays in familiar software—from a personal accelerator beside your laptop to data-center-scale systems.
One platform, every scale
dau selectively accelerates the parts of Polars, dataframe, time-series, and analytical workflows that benefit from streaming FPGA hardware. Reusable operations are composed into right-sized configurations, while unsupported work remains in familiar software. The same model scales from a personal accelerator to multi-card systems.
Deployment model
Work stays expressed through the same software model as capacity, memory, and physical placement change.
Personal
An M.2 accelerator in a Thunderbolt enclosure with 1 GB of local resident memory.
Scale-out
The host shards work across independent devices and merges results. Daisy-chained devices share one Thunderbolt tunnel.
Workstation
A PCIe x8 development platform with 8 GB of local DDR3 for larger resident datasets.
Concept
A planned fabric-linked array with distributed resident memory and direct card-to-card transport.
Personal computing
An accessible accelerator for laptop and workstation development. Load configurations matched to different analytical workflows without rebuilding applications around low-level hardware APIs.
Portable · Flashable · Developer-friendly
High-performance systems
Multiple accelerator cards partition data and workflows across independent local memory. The platform is designed to grow resident capacity and analytical throughput from a workstation to fabric-connected server arrays.
Parallel · Composable · Scalable
Accelerate supported operations when hardware fits their semantics, data shape, and performance needs. Keep everything else in software.
Combine filtering, transformation, aggregation, partitioning, and domain operations into streaming configurations shaped by measured workloads.
Connect reconfigurable compute to dataframe and analytical interfaces instead of rewriting workflows around low-level hardware APIs.
The dau flywheel
Workload evidence guides what runs in hardware today and what configuration should be built next. Every deployed design moves through the same verification path.
From proof to platform
dau began with a 5 GB NYSE TAQ market-data dataset and a concrete time-series analytical task: calculate OHLCV bars. After staging data on dpv1, repeated queries ran in 0.92 seconds versus 1.5 seconds on the CPU—39% lower per-query latency, with every result matched against the software golden. That measured workflow now informs the next platform configurations.
dpv1 · Proven
Initial platform
134K LUTs · 1 GB DDR3
Proved resident analytical queries can outperform CPU execution on accessible hardware.
dpv2 · Scaling now
Next platform
204K LUTs · 8 GB DDR3 · PCIe x8
Designed to expand resident capacity and host bandwidth for larger datasets and more parallel configurations.
dpv3 · Design program
Concept platform
663K LUTs · 32–128 GB DDR4 · PCIe Gen3 x8 · Fabric links
A concept design for fabric-connected arrays with distributed resident memory.
Build the next scale with us
We are working with early users and design partners to shape dau hardware, software, and deployment systems. Tell us what you compute, where it runs, and what performance would change for you.