Managing Data Files in Apache Iceberg

Everything was going great, your data was in your data lake, queries were fast and the SREs were happy. But then things started to slow down. Queries took longer, even specific queries which used to be fast now take a long time. The culprit? Small and unorganized files.

The solution? Apache Iceberg’s RewriteDatafile action. In this talk, Russell Spitzer will dive into how RewriteDataFiles can 1) right-size your files, merging small files and splitting large ones, ensuring that no time is waisted in query planning or in opening files; and 2) reorganize the data within your files, supporting hierarchal sort and multidimensional ordering algorithms, enabling you to make sure your data is optimally set out for your queries. With these two capabilities, any table can be kept at peak performance regardless of ingestion patterns and table size.

Download PDF

Topics Covered

Apache Iceberg
Table Formats

Ready to Get Started? Here Are Some Resources to Help

Using Data Mesh to Advance Distributed Data Access, Agility and Governance

Join this live fireside chat to learn about using Data Mesh to Advance Distributed Data Access, Agility and Governance.

read more


Smart Data – Smart Factory with Octotronic and Dremio

read more


What Is a Data Lakehouse?

The data lakehouse is a new architecture that combines the best parts of data lakes and data warehouses. Learn more about the data lakehouse and its key advantages.

read more

Get Started Free

No time limit - totally free - just the way you like it.

Sign Up Now

See Dremio in Action

Not ready to get started today? See the platform in action.

Watch Demo

Talk to an Expert

Not sure where to start? Get your questions answered fast.

Contact Us