dau.

Data. Accelerated.

dau maps high-value Polars, dataframe, time-series, and analytical operations into reconfigurable FPGA dataflow while the rest stays in familiar software—from a personal accelerator beside your laptop to data-center-scale systems.

One platform, every scale

Hardware shaped to your workload. Software that stays familiar.

dau selectively accelerates the parts of Polars, dataframe, time-series, and analytical workflows that benefit from streaming FPGA hardware. Reusable operations are composed into right-sized configurations, while unsupported work remains in familiar software. The same model scales from a personal accelerator to multi-card systems.

Deployment model

One platform, four deployment shapes.

Work stays expressed through the same software model as capacity, memory, and physical placement change.

Personal

dpv1 beside a laptop

An M.2 accelerator in a Thunderbolt enclosure with 1 GB of local resident memory.

Scale-out

A host-orchestrated dpv1 fleet

The host shards work across independent devices and merges results. Daisy-chained devices share one Thunderbolt tunnel.

Workstation

dpv2 in a desktop

A PCIe x8 development platform with 8 GB of local DDR3 for larger resident datasets.

Concept

Four dpv3 cards in a server

A planned fabric-linked array with distributed resident memory and direct card-to-card transport.

dpv1 is the first platform and where the approach was proven, dpv2 is the current platform running today, and dpv3 is a planned design.

Personal computing

An accelerator beside you.

An accessible accelerator for laptop and workstation development. Load configurations matched to different analytical workflows without rebuilding applications around low-level hardware APIs.

Portable · Flashable · Developer-friendly

High-performance systems

A fabric built around your data.

Multiple accelerator cards partition data and workflows across independent local memory. The platform is designed to grow resident capacity and analytical throughput from a workstation to fabric-connected server arrays.

Parallel · Composable · Scalable

Select the right work

Accelerate supported operations when hardware fits their semantics, data shape, and performance needs. Keep everything else in software.

Compose and right-size

Combine filtering, transformation, aggregation, partitioning, and domain operations into streaming configurations shaped by measured workloads.

Keep software familiar

Connect reconfigurable compute to dataframe and analytical interfaces instead of rewriting workflows around low-level hardware APIs.

The dau flywheel

A configuration loop that improves around real work.

Workload evidence guides what runs in hardware today and what configuration should be built next. Every deployed design moves through the same verification path.

  1. 01 Observe Profile real analytical workflows.
  2. 02 Select Choose operations worth accelerating.
  3. 03 Compose Build a right-sized configuration.
  4. 04 Verify Check software, simulation, and silicon.
  5. 05 ↺ Measure Deploy, reuse, and improve the next design.

From proof to platform

Measured on physical hardware.

dau began on dpv1 with a 5 GB NYSE TAQ dataset and a concrete task: calculate OHLCV bars. Staged resident, repeated queries ran in 0.92 seconds against 1.5 seconds on the CPU—39% lower per-query latency, every result matched to the software golden. dpv2 now carries that work further: a fused temporal-finance workflow—as-of join, derived features, time bars and rolling moments in a single pass, with no intermediate ever written back to memory—runs in 11.5 ms against 31.4 ms for the same workload on an 18-core laptop, bit-exact. The accelerator costs a fraction of the machine it outruns.

dpv1 · Where it started

Personal accelerator

First platform

134K LUTs · 1 GB DDR3

Established the core result: resident analytical queries can outperform CPU execution on hardware you can put beside a laptop.

dpv2 · Current platform

Desktop and server accelerator

Running today

204K LUTs · 10 GB DDR3 · PCIe x8

Two memory systems—2 GB onboard for bandwidth, 8 GB SODIMM for resident capacity—and eight parallel lanes. Where the fused temporal workflow now beats the CPU baseline.

dpv3 · Planned

Enterprise fabric card

In design

663K LUTs · 32–128 GB DDR4 · PCIe Gen3 x8 · Fabric links

A concept design for fabric-connected arrays with distributed resident memory.

Build the next scale with us

Bring us the workload that should be faster.

We are working with early users and design partners to shape dau hardware, software, and deployment systems. Tell us what you compute, where it runs, and what performance would change for you.