InfoTechTarget and Informa Tech's Digital Businesses Combine.

Together, we power an unparalleled network of 220+ online properties covering 10,000+ granular topics, serving an audience of 50+ million professionals with original, objective content from trusted sources. We help you gain critical insights and make more informed decisions across your business priorities.

Managing Data Files In Apache Iceberg

Presented by

Russell Spitzer

About this talk

Everything was going great: your data was in your data lake, queries were fast, and the SREs were happy. But then things started to slow down. Queries took longer, even specific queries which used to be fast now take a long time. The culprit? Small and unorganized files. The solution? Apache Iceberg’s RewriteDatafile action. This talk will dive into how RewriteDataFiles can 1) right-size your files, merging small files and splitting large ones, ensuring that no time is wasted in query planning or in opening files; and 2) reorganize the data within your files, supporting hierarchal sort and multidimensional z-ordering algorithms, enabling you to make sure your data is optimally set out for your queries. With these two capabilities, any table can be kept at peak performance regardless of ingestion patterns and table size.
Dremio

Dremio

4486 subscribers103 talks
Dremio is the easy and open data lakehouse platform.
Dremio is the easy and open data lakehouse, providing self-service analytics with data warehouse functionality and data lake flexibility across all of your data. Dremio increases agility with a revolutionary data-as-code approach that enables Git-like data experimentation, version control, and governance.
Related topics