Business

Databricks Lakebase Postgres Loads 1 TB in Five Minutes

Databricks has updated Lakebase Postgres with an LTAP architecture that offloads bulk data loading to Spark, protecting live application performance while accelerating ingestion up to 147x.

Databricks AI19 hrs agoBusiness
Illustration generated for this story

Databricks has unveiled a new Lake Transactional/Analytical Processing (LTAP) architecture for Lakebase Postgres, designed to eliminate the performance bottlenecks of massive data ingestion. By offloading heavy bulk operations from the primary database compute to distributed Spark engines, the system can write directly to storage. Internal benchmarks show that this method can load 1 TB of data in less than 5 minutes, representing an ingestion speedup of up to 147x compared to traditional methods.

During the beta phase, a customer used the system's Synced Tables feature to ingest approximately 1 billion rows of data daily. Previously, this bulk load took more than 8 hours, completely saturating the CPU and memory of their database. Even with autoscaling, the customer had to overprovision their online transaction processing (OLTP) resources to survive the daily sync. Under the new LTAP architecture, the same load ran on separate Spark compute, leaving the primary database instance completely unaffected and free to serve live application traffic.

To achieve this, Spark executors run sandboxed Postgres instances in binary-upgrade mode to generate valid heap pages and indexes in parallel. The system uses a binary COPY with FREEZE command to build heap slices, while a custom table access method constructs B-tree indexes without downloading the entire heap. Once the files are written to object storage, the primary database node only needs to write a compact Write-Ahead Log (WAL) record to publish the final manifest, rather than routing the entire data volume through the primary writer.

For database administrators and developers, this architecture decouples ingestion scaling from query performance. A small 1 CU primary instance can continue serving live traffic uninterrupted while Spark handles massive parallel loads on temporary, disposable compute. Practitioners no longer need to schedule risky bulk loads during off-hours, sacrifice data freshness, or overprovision expensive database hardware just to handle periodic data syncs.

This is our own summary of reporting by Databricks AI

More in Business