Apache Iceberg

File Compaction's Real Impact on Query Speed, Tested Across Three SQL Workloads

Compacting 1,000 Apache Iceberg files into just six sounds drastic, but this benchmark shows why fewer, larger files can dramatically accelerate query speed.

3 min readTowards Data Science
File Compaction's Real Impact on Query Speed, Tested Across Three SQL Workloads

File compaction is one of those unglamorous chores that rarely gets the attention it deserves. The headline "I Compacted 1,000 Apache Iceberg Files Into 6" sounds like a minor housekeeping note, but the real story is about how storage structure dictates query speed. We think this is exactly the kind of practical, hands-on testing that moves the conversation forward, because it turns a vague best practice into a measurable decision. The author ran three SQL workloads against the before and after states, and the results reinforce what many of us suspected: fewer, larger files often mean faster queries, but not always for the reasons you might assume. This is not about chasing a magic number; it is about understanding the tradeoffs in your own data environment.

The practical takeaway here is that compaction is not a universal win. For some workloads, particularly those heavy on filtering or point lookups, the reduced metadata overhead can be substantial. For others, like full scans or aggregations, the benefit may shrink or even reverse if file sizes grow too large for your cluster's memory. The piece wisely avoids declaring a one-size-fits-all rule, and that restraint is refreshing. It aligns with a broader theme we have been tracking: performance tuning is contextual. For example, we recently covered how A sharper alignment: Jev's confidence accuracy jumps 68% showed that calibration techniques only help when applied to the right benchmark. Similarly, compaction only helps when your workload actually benefits from it. And just as Transform an Open LLM Into a Fast Classifier by Swapping Its Head demonstrates that architectural changes can have outsized effects without altering the model's core, compaction can be a structural lever that pays off if you measure first.

What we appreciate most is that the testing does not oversell. There is no "revolutionary" language here, no claim that this one test settles the debate. Instead, we get a clear methodology and honest results. That is the kind of evidence-based reporting we want to see more of. For practitioners, the immediate lesson is to benchmark your own workloads before and after compaction. Run your three most representative queries, measure the latency, and let the data decide. The cost of compaction itself, in compute and time, is also part of the equation, and that cost is not ignored. So the real question is not whether compacting 1,000 files into 6 is good, but whether it is good *for you*. The answer is out there, but you have to test for it. Watch for how your query planner's statistics change after compaction; that is often where the hidden wins, or hidden losses, will show up first.

From Towards Data Science

Benchmarking the impact of fewer, larger files across three SQL workloads

The post I Compacted 1,000 Apache Iceberg Files Into 6. Here’s What Happened to Query Performance. appeared first on Towards Data Science.

Read the original at Towards Data Science