Back to all articles

Streaming articles

Real-time and event-driven data engineering — streaming ingestion, CDC pipelines, and near-real-time analytics.

15 articles

Apache IcebergLakeOpsStreaming

Apache Iceberg CDC Pipeline: Change Data Capture Best Practices

Getting CDC data into Iceberg is solved — Debezium, Flink, and DMS handle ingestion. The hard part is maintaining CDC tables that receive continuous updates and deletes. A practical guide to ingestion patterns, delete file management, and autonomous maintenance.

Chris P
Chris P
21 min read
Apache Iceberg V3 for streaming — row-level lineage, schema evolution, and governance
Apache IcebergStreamingData Governance

Apache Iceberg V3 for Streaming: Row-Level Lineage, Schema Evolution, and Governance

Iceberg V3 brings row-level lineage, default column values, and deletion vectors — the features streaming pipelines need for governance without downtime. How V3 changes the streaming governance story on Flink, Kafka, and CDC sources.

Rob M
Rob M
27 min read
The rise of the open Apache lakehouse — modular vendor-neutral architecture with Iceberg, Polaris, and Fluss
Data PlatformsApache IcebergData Lakehouse

The Rise of the Open Apache Lakehouse: Modular Architecture for Vendor-Neutral Data Platforms

How Apache projects have assembled a fully modular, vendor-neutral lakehouse stack — covering table formats (Iceberg, Hudi, Paimon), REST catalogs (Polaris, Gravitino), compute engines (Spark, Trino, Flink), real-time ingestion (Fluss), and why the operational gap demands an autonomous control plane.

Jonathan Saring
Jonathan Saring
28 min read
Hot and cold data tiering on Apache Iceberg with StarRocks — real-time analytics architecture
Data PlatformsApache IcebergAnalytics

Iceberg Hot and Cold Data Tiering: StarRocks + Iceberg for Real-Time Analytics

Hot/cold data tiering on Apache Iceberg — StarRocks for sub-second dashboards, Iceberg for petabyte history, one SQL surface via federation. Ingestion, tier transitions, dedup, and keeping the cold tier fast.

Chris P
Chris P
34 min read
How LinkedIn scales Apache Iceberg CDC ingestion to billions of upserts per day
Data PlatformsApache IcebergStreaming

Iceberg CDC at Scale: How LinkedIn Ingests Billions of Upserts Per Day

LinkedIn runs Iceberg CDC on 10,000+ tables — billions of upserts daily, 2+ PB throughput. Equality vs position deletes, delete-file compaction, budgeted maintenance, and WAP branching lessons for any team running MERGE INTO at scale.

Rob M
Rob M
26 min read
Streaming lakehouse on Apache Iceberg — multi-engine data ingestion into Iceberg tables
Data PlatformsApache IcebergStreaming

Streaming Lakehouse on Apache Iceberg: Kafka, Flink, and Real-Time Pipelines Without Duplication

Build a streaming lakehouse on Apache Iceberg — unify Kafka/Flink ingestion and batch analytics without duplicating data. Production patterns, maintenance reality, and how to keep streaming Iceberg tables performant.

Jonathan Saring
Jonathan Saring
27 min read
Multi-table transactions in Apache Iceberg — cross-table atomicity for the open lakehouse
Apache IcebergLakeOpsData Governance

Iceberg Multi-Table Transactions: Cross-Table Atomicity for Production Lakehouses

Star schema ETL and multi-table CDC need atomic commits across Iceberg tables — not just single-table ACID. How the REST Catalog transaction API, Polaris, Nessie, and Gravitino enable cross-table atomicity, and what teams run in production today.

Chris P
Chris P
16 min read
Apache Iceberg Commit Conflicts — causes, prevention, and recovery with concurrent write paths
Apache IcebergStreamingApache Flink

Apache Iceberg Commit Conflicts: Causes, Prevention, and Recovery

Every concurrent write to an Apache Iceberg table risks a commit conflict. This guide covers how Iceberg's optimistic concurrency works, what triggers CommitFailedException, the common conflict scenarios in streaming and maintenance workloads, and the strategies — from partition isolation to branch-based writes — that eliminate conflicts in production.

Chris P
Chris P
33 min read
Kafka to Iceberg Compaction — Kafka events streaming into an Iceberg table, compacted through a gear process into optimized blocks.
CompactionApache IcebergApache Kafka

Kafka to Iceberg Compaction — Done Right

Streaming from Kafka into Apache Iceberg creates small files faster than any other write pattern. This guide covers why standard compaction approaches fail for streaming tables, how to measure compaction need, implement partition-aware compaction that avoids writer conflicts, tune rewriteDataFiles parameters, and run maintenance autonomously at scale.

Rob M
Rob M
26 min read
Kafka to Iceberg Ingestion Guide — Kafka logo with streaming data records flowing into a geometric iceberg lakehouse.
Apache IcebergApache KafkaApache Flink

Kafka to Iceberg: Ingestion Guide

A practical guide to streaming data from Apache Kafka into Apache Iceberg tables — covering Kafka Connect, Apache Flink, Spark Structured Streaming, and CDC with Debezium. Includes configuration examples, schema management, partitioning strategies, production pitfalls, and how to keep streaming tables healthy at scale.

Rob M
Rob M
27 min read
Apache Iceberg 1.11.0 What's New — Nessie mascot beside an iceberg with icons for performance, security, routing, and extensibility.
Apache IcebergLakehouseCompaction

Apache Iceberg 1.11.0 — What's New?

Apache Iceberg 1.11.0 lands V3 maturity with production-ready deletion vectors, a native Variant type for semi-structured data, server-side scan planning, built-in table encryption, and a pluggable File Format API that opens the door to next-generation storage formats.

Jonathan Saring
Jonathan Saring
10 min read
Apache Iceberg with Flink Optimization — Flink squirrel mascot with streaming data flowing through an optimization ring into a geometric iceberg, with performance metric icons
Apache IcebergApache FlinkStreaming

Apache Iceberg with Flink: Streaming Optimization Guide

Flink streaming into Iceberg creates thousands of small files per hour. This guide covers checkpoint tuning, write distribution modes, Flink SQL patterns, and why external maintenance is essential for production streaming tables.

Chris P
Chris P
15 min read
Apache Iceberg Delete Files — stacked data blocks with pink delete file markers funneled through compaction into clean, optimized data with a performance gauge showing improved read speed
Apache IcebergCompactionLakeOps

Apache Iceberg Delete Files: Reducing Merge-on-Read Overhead

Delete files let Iceberg avoid rewriting data on every UPDATE or DELETE — but every unresolved delete file forces readers to reconcile at query time. A deep guide to position deletes, equality deletes, measuring overhead, and resolving accumulation before it tanks performance.

David W
David W
17 min read
Fixing Small Files in Apache Iceberg — scattered small data cubes compacted into larger organized file blocks flowing toward a geometric iceberg
CompactionApache IcebergLakeOps

Fixing Small Files in Apache Iceberg: A Practical Guide

Small files silently degrade every Apache Iceberg lakehouse — inflating S3 costs, slowing query planning, and bloating metadata. This guide covers root causes, measurement, manual and automated fixes, and how to eliminate the problem at scale.

Rob M
Rob M
20 min read
Optimizing Iceberg Lake Compaction — scattered small data-block cubes funnel through a compaction machine onto a conveyor belt of optimized blocks, leading to a crystal-clear iceberg lakehouse
CompactionApache IcebergLakehouse

Optimizing Iceberg Lake Compaction: A Guide

Compaction is the most impactful operation in an Apache Iceberg lakehouse — and the hardest to get right at scale. File merging is the easy part. Knowing when to trigger it, what sort strategy to apply per table, how to avoid conflicting with other maintenance, and how to do it without spinning up expensive JVM clusters — that is the real problem. A breakdown of what modern compaction actually requires.

Jonathan Saring
Jonathan Saring
17 min read