Delta Lake
Working with Billions of Rows in Delta Lake
What changes when Delta tables reach billions of rows: partitioning, file sizing, MERGE, incremental processing, small files, compaction and concurrent writers.
- Status
- Proposed
- Type
- Talk
- Formats
- Webinar · Internal session
- Level
- Advanced
- Duration
- 45–60 minutes
This is a proposed session topic. No public event is currently scheduled.
Who it is for
- Data engineers operating large Delta tables in Fabric, Databricks or Spark
- Engineers whose MERGE and maintenance jobs are getting slower every month
What attendees will learn
- How to choose a partitioning strategy, including when not to partition
- Why file size matters for reads, writes and the transaction log
- How MERGE works internally and how to keep it narrow
- How to process only what changed at this scale
- How small files appear, and how to plan compaction and maintenance
- What happens when several jobs write to the same table
On this page
Session overview
Design choices that are harmless at a million rows become expensive at a billion. A partition column that once looked sensible creates millions of files. A MERGE that took minutes takes hours, and maintenance jobs collide with loads. This session explains the mechanics behind those symptoms and the design and operational patterns that keep very large Delta tables fast and predictable.
Outline
- How Delta stores data: Parquet files, the transaction log and checkpoints.
- Partitioning: cardinality, data per partition, and clustering as an alternative.
- File sizing and small files: where they come from and what they cost.
- MERGE at scale: narrowing the target, pruning, and deduplicating the source.
- Incremental processing: change detection and idempotent reruns.
- Compaction and maintenance: OPTIMIZE, VACUUM and scheduling around loads.
- Concurrency: optimistic concurrency, conflicts, and designing writers that do not collide.