Back to all articles

Analytics articles

Data analytics on open lakehouses — query performance, BI integration, interactive exploration, and engine selection for analytical workloads.

15 articles

Building an AI Agent Analytics System on Apache Iceberg Data
AIApache IcebergLakeOps

Building an AI Agent Analytics System on Apache Iceberg Data

AI agents are eliminating the BI bottleneck — translating natural language to SQL in seconds, routing queries to optimal engines, and producing analyst-grade insights autonomously. Learn how to build a production analytics system on Iceberg with multi-engine routing, guardrails, and self-optimizing storage.

Chris P
Chris P
19 min read
DuckDB and Apache Iceberg — query, write, and optimize lakehouse tables without a Spark cluster
Apache IcebergLakeOpsCompaction

DuckDB for Apache Iceberg

Query and write Apache Iceberg tables with DuckDB — no cluster required. Catalog setup, MERGE INTO, time travel, table layout, and when to route to DuckDB vs Spark or Trino.

Rob M
Rob M
19 min read
Apache IcebergLakeOpsCompaction

Why Your Iceberg Queries Are Slow (And How to Fix Them)

Slow Iceberg queries almost always trace back to five structural problems: small files, wrong sort order, manifest bloat, stale snapshots, or partition misalignment. This diagnostic guide shows you how to find each one, confirm it with SQL, and fix it — manually or with autonomous optimization.

Rob M
Rob M
17 min read
Faster Trino with Iceberg — Trino rabbit mascot with a speed gauge and layered Iceberg data blocks accelerating query performance
Apache IcebergTrinoLakeOps

Slow Trino Queries on Iceberg? 7 Fixes for Faster Trino Iceberg Performance

Slow Trino queries on Apache Iceberg are rarely a compute problem — they're a table layout problem. Unmaintained Iceberg tables turn sub-second Trino scans into minute-long full reads. This guide covers seven proven fixes for faster Trino Iceberg performance: compaction, query-aware sorting, partition strategy, manifest optimization, lifecycle cleanup, multi-engine routing, and continuous observability. Why Trino-native maintenance falls short — and how to automate each fix at scale.

Chris P
Chris P
25 min read
Iceberg Table Partitioning Strategies — A Practical Guide showing partition transforms splitting data into time, region, and category folders
PartitioningApache IcebergLakeOps

Iceberg Table Partitioning Strategies: A Practical Guide

How to choose, evaluate, and evolve partitioning strategies for Apache Iceberg tables — decision frameworks for real workloads, partition lifecycle management, and when to change your strategy.

Chris P
Chris P
16 min read
Hot and cold data tiering on Apache Iceberg with StarRocks — real-time analytics architecture
Data PlatformsApache IcebergAnalytics

Iceberg Hot and Cold Data Tiering: StarRocks + Iceberg for Real-Time Analytics

Hot/cold data tiering on Apache Iceberg — StarRocks for sub-second dashboards, Iceberg for petabyte history, one SQL surface via federation. Ingestion, tier transitions, dedup, and keeping the cold tier fast.

Chris P
Chris P
34 min read
Apache Iceberg query planning internals — predicate pushdown, manifest filtering, and data skipping
Apache IcebergAnalyticsLakeOps

Apache Iceberg Query Planning Explained: Predicate Pushdown, Manifest Filtering, and Data Skipping

Apache Iceberg query planning is the coordinator-bound bottleneck before any parallel scan starts. This guide covers predicate pushdown, manifest list pruning, file-level data skipping, and what data platform teams do to keep planning fast as tables grow.

Rob M
Rob M
25 min read
Apache Iceberg Data Quality and Table Health — where reliability actually breaks, healthy vs unhealthy comparison
Apache IcebergObservabilityLakeOps

Apache Iceberg Data Quality and Table Health: Where Reliability Actually Breaks

Data quality and table health are different failure modes — one breaks business trust, the other breaks performance silently. A practical guide to the metrics, monitoring queries, classification frameworks, and automated remediation that keep Iceberg tables reliable in production.

David W
David W
26 min read
Apache Iceberg Table Partitioning Best Practices — a geometric iceberg branching into date, region, and category partition columns, each with table and folder icons showing the partition hierarchy
Apache IcebergPartitioningLakeOps

Apache Iceberg Table Partitioning Best Practices

Partitioning determines how much data every query must scan. Apache Iceberg's hidden partitioning and partition evolution change the game — but choosing the wrong strategy still creates performance cliffs. A practical guide to transforms, sizing, evolution, and avoiding the small-files trap.

Chris P
Chris P
18 min read
Apache Iceberg Puffin Statistics — a puffin bird beside a statistics dashboard showing file counts, records, partitions, and data size, connected to a geometric iceberg
Apache IcebergLakeOpsAnalytics

Apache Iceberg Puffin Statistics: A Practical Guide

Puffin files store table-level statistics — NDV sketches and custom blobs — that query engines use for join ordering, split planning, and cost-based optimization. A practical guide to how they work, how to collect them, how they go stale, and how to keep them accurate at scale.

David W
David W
18 min read
Apache Iceberg with Trino Optimization — Trino logo with an optimization gauge sending query streams into a geometric iceberg, with performance metric icons for throughput, latency, and efficiency
Apache IcebergTrinoCompaction

Apache Iceberg with Trino: Performance Optimization Guide

A practical guide to optimizing Apache Iceberg queries and table maintenance with Trino — covering scan planning, predicate pushdown, file pruning, Trino-side tuning, maintenance procedures, physical layout optimization, and how a dedicated control plane eliminates JVM overhead while adding cross-engine intelligence.

Chris P
Chris P
18 min read
Iceberg Lake for Data Analytics: Optimization Guide — iceberg on water with analytics dashboard showing 9.4× query speed, 68% cost efficiency gain, and 82% less data scanned
Apache IcebergData PlatformsData Lake

Iceberg Lake for Data Analytics: Optimization Guide

Eight optimization layers for data platform engineers running BI, ad-hoc SQL, and aggregation pipelines on Apache Iceberg — from partition design and file sizing through compaction, routing, and continuous maintenance.

Jonathan Saring
Jonathan Saring
15 min read
Optimizing Iceberg Lakehouse Performance — problems (small files, fragmented manifests, unsorted data, delete files) flow through autonomous maintenance into faster queries, lower costs, higher throughput, and healthier data
Apache IcebergLakeOpsAnalytics

Optimizing Iceberg Lakehouse Performance

Iceberg tables degrade silently — small files from streaming, unsorted data, fragmented manifests, accumulated delete files. Each one caps query speed regardless of engine. Six concrete optimization layers, how they interact, and how autonomous maintenance keeps every table at peak performance.

David W
David W
11 min read
Data Lake vs Lakehouse vs Warehouse: A Practical Guide — watercolor illustration comparing a natural data lake (raw flexible storage), a lakehouse (open storage with analytics on the water), and a data warehouse (structured BI building with charts in the windows)
Data PlatformsData LakeLakehouse

Data Lake vs Lakehouse vs Warehouse: A Practical Guide

Data lakes, warehouses, and lakehouses are not interchangeable — each has hard limits the others cannot cover. A practical guide for platform leaders: where each architecture wins, where it fails, cost and governance trade-offs, and how to choose (or combine) them in 2026.

Chris P
Chris P
22 min read
Introducing QueryFlux: Open-Source Universal Multi-Engine Query Router and SQL Proxy
QueryFluxApache IcebergData Platforms

Introducing QueryFlux: Open-Source Universal Multi-Engine Query Router and SQL Proxy

QueryFlux is a universal SQL proxy and multi-engine query router in Rust—one access layer in front of Trino, DuckDB, StarRocks, and Athena with routing, dialect translation, and observability.

Jonathan Saring
Jonathan Saring
12 min read