InfoTechTarget and Informa Tech's Digital Businesses Combine.

Together, we power an unparalleled network of 220+ online properties covering 10,000+ granular topics, serving an audience of 50+ million professionals with original, objective content from trusted sources. We help you gain critical insights and make more informed decisions across your business priorities.

Optimizing Data Files in Apache Iceberg: Performance strategies

Presented by

Dipankar Mazumdar, Developer Advocate, Dremio

About this talk

Querying 100s of petabytes of data demands optimized query speed specifically when data accumulates over time. We have to ensure that the queries remain efficient because over time you may end up with a lot of small files and your data might not be optimally organized. In this talk, we will cover: - Apache Iceberg table format - Problems in the data lake: small files, unorganized files - Techniques such as: partitioning, compaction, metrics filtering - Overlapping metrics problem - Solving it using sorting, Z-order clustering
Dremio

Dremio

4486 subscribers103 talks
Dremio is the easy and open data lakehouse platform.
Dremio is the easy and open data lakehouse, providing self-service analytics with data warehouse functionality and data lake flexibility across all of your data. Dremio increases agility with a revolutionary data-as-code approach that enables Git-like data experimentation, version control, and governance.
Related topics