Personal
dpv1 beside a laptop
An M.2 accelerator in a Thunderbolt enclosure with 1 GB of local resident memory.
Data. Accelerated.
dau maps selected Polars, dataframe, time-series, and analytical operations into reconfigurable FPGA dataflow and leaves the rest in software. The same model runs on a personal accelerator beside your laptop and on data-center-scale systems.
The platform
dau accelerates the parts of Polars, dataframe, time-series, and analytical workflows that suit streaming FPGA hardware. Reusable operations compose into right-sized configurations; unsupported work stays in software. The same model scales from a personal accelerator to multi-card systems.
Deployment model
The software model stays the same as capacity, memory, and physical placement change.
Personal
An M.2 accelerator in a Thunderbolt enclosure with 1 GB of local resident memory.
Scale-out
The host shards work across independent devices and merges results. Daisy-chained devices share one Thunderbolt tunnel.
Workstation
A PCIe x8 development platform with 8 GB of local DDR3 for larger resident datasets.
Concept
A planned fabric-linked array with distributed resident memory and direct card-to-card transport.
Personal computing
An accelerator for laptop and workstation development. Load a configuration that matches the workflow you are running, without rebuilding the application around low-level hardware APIs.
Portable · Flashable
High-performance systems
Multiple accelerator cards partition data and workflows across independent local memory. The platform is designed to grow resident capacity and analytical throughput from a workstation to fabric-connected server arrays.
Parallel · Composable
Accelerate supported operations when hardware fits their semantics, data shape, and performance needs. Keep everything else in software.
Combine filtering, transformation, aggregation, partitioning, and domain operations into streaming configurations shaped by measured workloads.
Connect reconfigurable compute to dataframe and analytical interfaces instead of rewriting workflows around low-level hardware APIs.
The dau flywheel
Measurements decide what runs in hardware today and what configuration gets built next. Every deployed design goes through the same verification path.
From proof to platform
dau began on dpv1 with a 5 GB NYSE TAQ dataset and a concrete task: calculate OHLCV bars. Staged resident, repeated queries ran in 0.92 seconds against 1.5 seconds on the CPU—39% lower per-query latency, every result matched to the software golden. dpv2 now carries that work further: a fused temporal-finance workflow—as-of join, derived features, time bars and rolling moments in a single pass, with no intermediate ever written back to memory—runs in 11.5 ms against 31.4 ms for the same workload on an 18-core laptop, bit-exact. The accelerator costs a fraction of the machine it outruns.
dpv1 · Where it started
First platform
134K LUTs · 1 GB DDR3
Established the core result: resident analytical queries can outperform CPU execution on hardware you can put beside a laptop.
dpv2 · Current platform
Running today
204K LUTs · 10 GB DDR3 · PCIe x8
Two memory systems (2 GB onboard for bandwidth, 8 GB SODIMM for resident capacity) and eight parallel lanes. Where the fused temporal workflow now beats the CPU baseline.
dpv3 · Planned
In design
663K LUTs · 32–128 GB DDR4 · PCIe Gen3 x8 · Fabric links
A concept design for fabric-connected arrays with distributed resident memory.
Design partner program
We are working with early users and design partners on dau hardware, software, and deployment. Tell us what you compute, where it runs, and how fast it needs to be.