
Apache Iceberg 1.12.0 was released on September 29, 2026. The short version: there is a new clustering strategy worth testing, deletion vectors finally have a migration path off equality deletes, the V4 format machinery landed in core, and four things were removed that will break an upgrade if you do not check for them first.
That last part is the one to plan around: 1.12 is as much a cleanup release as a feature release, and one of its removals can break reads on tables that work fine today.
The genuinely new capabilities break into three groups. First, physical layout: Spark 4.1 gains a Hilbert-curve clustering strategy for rewrite_data_files, and a Z-order encoding bug affecting floating-point columns was fixed. Second, row-level deletes: a Flink maintenance task now converts equality deletes into deletion vectors, and a new scan-based action removes dangling delete files. Third, the V4 groundwork — a V4 manifest reader, relative paths, and content stats in the spec — which tells you where the format is heading even though V4 is not yet the default.
Read those groups together and a pattern shows up that the release notes never state outright. Flink equality-delete conversion, fast variant reads, and real geospatial read/write are close to an exact list of the gaps that kept teams on v2 — while the one significant removal is a v2-only feature. 1.12 is the release that makes v3 a practical target, and that, more than any single feature, is what to take from it.
Start with what breaks
Removals go first, because they decide whether this upgrade is a version bump or a project.
| Removed in 1.12 | Affects | Do this first |
|---|---|---|
| Spark 3.4 support | Spark 3.4 | Move to 3.5, 4.0, or 4.1 before bumping Iceberg |
| Flink 2.0 support | Flink 2.0 pipelines | Move to 1.20, 2.1, 2.2, or 2.3 (2.2/2.3 are new) |
| Position deletes with row data | v2 tables from older writers | Remediate while still on 1.11 |
Deprecated Spark* conf and util methods | Custom Spark integrations | Recompile against 1.12 |
DataReader, GenericAppenderFactory, BaseFileWriterFactory | Custom readers/writers | Use PlannedDataReader |
| AWS S3 signer classes and properties | Custom S3 signing | Move to remote-signing config |
| Partition stats read path (Core, ORC) | Early adopters | Use the partition statistics scan API |
If you are jumping more than one version: Java 11 was already dropped in 1.11, so JDK 17 is the baseline, and Flink 1.19 went away there too.
The position-delete removal is the sharp edge
This one is a silent data-access failure rather than a compile error. In the v2 spec a position delete file records rows to skip as a (file_path, pos) pair, and it could carry an optional third column holding the full deleted row, for engines that wanted the content without re-reading the data file. That field was deprecated in 1.11 and is removed in 1.12 across Core, Data, Spark, and Flink.
The risk is that these files were written by older writers, they are perfectly valid v2 metadata, and nothing in your catalog flags them — so a table that reads fine on 1.11 can fail on 1.12, and maintenance jobs touching those deletes fail with it. The discussion that accepted this removal settled on two remediations, and every affected table needs one of them chosen deliberately: upgrade the table to v3 and rewrite the deletes as deletion vectors, or stay on v2 and let data compaction fold those deletes into the data files. Either way the work happens before the upgrade, not after the first failed job:
1-- 1. Narrow the search set. content = 1 is position deletes, 2 is equality2-- deletes; only position deletes could ever carry row data.3SELECT content, count(*) AS delete_files4FROM prod.analytics."orders_fact$delete_files"5GROUP BY content;6 7-- 2. List the candidates. The metadata table names the files; the Parquet/Avro8-- footer is what confirms a third top-level `row` field beside file_path9-- and pos.10SELECT file_path, file_format FROM prod.analytics."orders_fact$delete_files"11WHERE content = 1;12 13-- 3. Rewriting data files rewrites their deletes, clearing row data. On 1.11.14CALL prod.system.rewrite_data_files(15 table => 'analytics.orders_fact',16 options => map('delete-file-threshold', '1')17);18-- On v3, converting to deletion vectors sidesteps this entirely.Two behavior changes that will not throw an error
The default AWS SDK HTTP client moved to Apache HttpClient 5. If the Iceberg AWS bundle manages your dependencies, nothing changes. If you pin AWS SDK versions yourself — most Spark, Trino, and Flink deployments do — swap software.amazon.awssdk:apache-client for apache5-client. Miss it and you get a runtime classpath failure on the first S3 call, not a build error.
The REST client now retries POST requests carrying an Idempotency-Key on retriable errors (408, 500, 502, 503, 504). Commits get more resilient to transient catalog failures — and your catalog genuinely has to honor those keys, because a server treating a retried POST as a second distinct request will now see duplicate work.
The question the release notes cannot answer
Every item above is a per-table question asked across the whole lake. Which tables carry position deletes with row data? Which are accumulating equality deletes worth converting? Which have enough manifests that a rewrite should precede the upgrade? Which are Z-ordered on a float column and therefore clustered incorrectly?
No catalog answers these. Glue, Polaris, Unity, Nessie, and the REST catalogs hold pointers, swap them atomically, and authorize access — none is responsible for the health of the tables underneath. The information exists in Iceberg's own metadata tables, but reading it across hundreds of tables in several catalogs is a job nobody has written.
That gap is what a control plane fills. A control plane such as LakeOps sits above your catalogs and engines, reads table metadata and query telemetry through the same APIs those engines already use, and runs a closed loop over the estate: sense file layout, metadata health, and query patterns; plan by severity, sequencing operations in dependency order; optimize by actually executing them; then learn, so sort orders and cadence adapt on the next pass. It does not store your data, own table pointers, or sit in the query path.

Why this is not just a pile of scripts
Every team with an Iceberg lake already has some version of this: Spark procedures on a cron, an Airflow DAG per catalog, an alert when storage jumps. For twenty tables that is enough. Four things break as the estate grows, and they are the difference between automation and a control plane:
- Scripts trigger on a clock; a control plane triggers on state. A nightly compaction job rewrites healthy tables and under-serves hot ones. Health-driven execution runs when a table's measured condition calls for it — which is also how you find the twelve tables that actually block a version upgrade.
- Scripts are per-catalog; a control plane is estate-wide. A DAG written against Glue does not understand Polaris, and nothing joins them into one health picture.
- Scripts do not rank. At 2,000 tables the useful output is not "maintain this table," it is "these forty are worst, and here is why."
- Scripts have no feedback loop. A sort order hardcoded in a DAG is a guess made once. Reading which columns queries actually filter on is what makes a new strategy like Hilbert clustering an evidence-based choice instead of a coin flip.
The capability groups that matter here are observe (one health score per table plus cross-engine query telemetry), maintain (snapshot expiry, orphan cleanup, and manifest rewrite in dependency order), compact (query-aware file rewriting, with sort and clustering chosen from real filters and tested on a branch), and govern (declarative policy with inheritance and an audit trail). LakeOps attaches to Glue, Polaris, Nessie, Gravitino, Lakekeeper, S3 Tables, and REST-compatible catalogs through metadata only, so adding a catalog moves no data and changes no pipelines.
For an upgrade specifically, the useful sequence is lake-wide storage and compute first, then every table's health in one list, then the operation log showing what ran.



With that map in place, here is what 1.12 actually gives you.
Deletion vectors finally have a migration path
Deletion vectors stabilized in 1.11: instead of a separate positional delete file per operation, a v3 table keeps one Roaring bitmap per data file in a Puffin file, with a strict 1:1 relationship. Reads apply a bitmap mask instead of opening and cross-referencing a pile of small files. The delete files and merge-on-read guide covers that mechanism in depth.
The gap 1.11 left was equality deletes. CDC pipelines on Flink overwhelmingly write equality deletes, which identify rows by value rather than position. They are cheap to write and expensive to read, because the engine must evaluate the predicate against candidate files rather than applying a positional mask. There was no supported route from an equality-delete table to a deletion-vector table.
1.12 adds one: a Flink maintenance task called ConvertEqualityDeletes, integrated with IcebergSink. It converts accumulated equality deletes into deletion vectors as a maintenance operation rather than requiring a full table rewrite. Supporting fixes landed alongside it — deleted rows no longer reappear after a failed conversion cycle, and unpartitioned equality deletes now resolve correctly across all partitions.
What it means: if you run CDC into Iceberg and your read latency has been slowly degrading, this is the most valuable thing in the release. The remediation is no longer "rewrite the table"; it is a maintenance task you can schedule. Two caveats before you reach for it: it targets v3 tables with deletion vectors enabled, and it is a Flink task, so Spark-only shops are still waiting.
Three related improvements matter if you are already on deletion vectors. A new scan-based action removes dangling delete files — still referenced in metadata but no longer applicable to any live data file, which accumulate quietly after compaction. Co-located DVs in the same Puffin file are now exposed through DataFile, and commit validation was fixed where one data file has multiple DVs across snapshots. Snapshot expiration now reads delete manifests correctly, a real source of orphaned delete files.
Hilbert-curve clustering, and a Z-order bug worth knowing about
The layout change in 1.12 is a new clustering strategy, and it arrives alongside a correctness fix that may have been quietly costing you query performance for months.
Most lakehouse performance problems are pruning problems, and pruning is several mechanisms that fail independently:
- 1.Partition pruning (coarse). The partition spec with hidden transforms —
year,month,day,bucket,truncate— decides which groups of files can be skipped entirely. Over-partitioning (hourly on a moderate table) creates a metadata problem worse than the scan it was meant to fix. - 2.File pruning via sort order, or linear clustering (fine). Rewriting so rows with nearby key values land in the same files tightens per-file min/max statistics. A sort on
(customer_id, event_date)prunescustomer_idexcellently andevent_datealone poorly — the leading column gets almost all the benefit. - 3.Z-order (multi-dimensional clustering). Bit-interleaving across columns produces a space-filling curve so two to four filter columns all prune reasonably well, at the cost of none pruning as well as a leading sort column would.
- 4.Hilbert curve — new in 1.12 for Spark 4.1. Also a space-filling curve, but one with better locality preservation than Z-order.
- 5.Row-group pruning inside Parquet, which only helps if the file is clustered in the first place.
- 6.File sizing, typically 128–512 MB for analytical tables, since file count drives object-storage request cost directly.
- 7.Metadata health — manifest count, snapshot depth, delete-file burden — which sets planning cost before any data is read.
Before going further, the distinction that causes the most wasted effort: binpack compaction fixes file count, it does not cluster rows. Merging small files by size alone leaves min/max ranges just as wide as before, so file pruning still fails. That is how a table gets compacted on a schedule for months while queries never speed up. And cluster means three different things in this space: a clustering strategy is a sort order, file sizes grouping around a target is file-size distribution, and a Spark cluster is compute. Databricks Liquid Clustering, BigQuery clustering, and Snowflake micro-partitions are not Iceberg sort order either — compare them, do not equate them.
Why Hilbert beats Z-order on paper
Both Z-order and Hilbert curves map multi-dimensional points onto a single dimension so nearby points in the original space tend to land near each other in the sorted output. The difference is continuity. A Z-order curve (Morton order) interleaves bits, producing large jumps between some adjacent regions — two rows that are neighbors in your filter space can land far apart in the file layout. A Hilbert curve visits every cell with single-step moves and never jumps, so locality holds more consistently.
In practice that means tighter min/max bounds per file for the same number of clustering columns, and better file skipping on multi-column predicates. The cost profile matches Z-order: a full rewrite, more expensive than a linear sort, with the benefit diluting past roughly four columns.
| Strategy | Best for | Cost | Caveat |
|---|---|---|---|
binpack | File-count problems only | Cheapest | Does not improve pruning |
sort | One dominant filter column | Moderate | Non-leading columns prune poorly |
zorder | 2–4 equally weighted filter columns | Expensive | Discontinuities weaken locality; float encoding was buggy before 1.12 |
hilbert (Spark 4.1, new) | 2–4 equally weighted filter columns | Expensive | Newest code path; validate before estate-wide adoption |
The Z-order fix you should actually act on
1.12 includes a fix for Z-order byte encoding of floating-point values. This is the kind of entry that looks minor in a changelog and is not: if you Z-ordered a table on a float or double column on an earlier version, the encoding was incorrect, meaning the resulting clustering did not deliver the pruning you paid a full rewrite for. Spark also got a fix for a Z-order NPE on null booleans and case-insensitive column resolution.
What it means: audit which tables you have Z-ordered and on which column types. Any that include floating-point columns are candidates for re-clustering on 1.12 — and that is a good moment to evaluate whether Hilbert is the better choice for those tables anyway.
Which brings up the discipline that matters more than the strategy choice. Changing clustering rewrites every file in scope and is awkward to undo. Derive candidate keys from observed access — the columns your engines actually put in WHERE, JOIN, and GROUP BY, not the ones someone guessed at table-create time — then test on an Iceberg branch and diff against the current layout before promoting anything.


If you are choosing between these strategies for a specific table, the compaction strategies guide works through when each one wins, and query-aware compaction covers deriving the sort key from observed filters instead of declaring it up front.
Variant and geospatial cross into usable
Both of these types existed before 1.12. What changed is that engines can now actually read and write them efficiently.
Variant gets vectorized Parquet reads for unshredded columns in Spark 4.0 and 4.1 — the difference between a semi-structured column being technically supported and being fast enough for production scans. Shredding gained type uniformity, Flink 2.1 through 2.3 support variant in Avro plus shredded writes, Kafka Connect can enable Parquet variant shredding, and VariantType was added to the REST catalog spec. A cluster of metrics bugs was fixed too, including STRING bounds now using UTF-8 byte order and large decimals shredding properly.
What it means: if you shelved variant adoption because scan performance was not there, revisit it. One caution: a variant column widens a table's effective schema, and shredded sub-columns are what queries actually touch — so sort-order and compaction decisions should account for which shredded paths get filtered, not just the top-level column.
Geospatial moves from bounding-box groundwork in 1.11 to actual data movement. Geometry and geography values can now be read and written as WKB in Parquet and in Avro, Spark 4.1 maps them to Spark types, and single-value binary serialization landed in the API. One trap if you have tooling that parses type names: GeometryType and GeographyType toString() now include the resolved CRS, printing geometry(OGC:CRS84) and geography(OGC:CRS84, spherical).
V4 is being built in the open
A large share of 1.12's core work is V4 format plumbing. V4 is not the default and you should not be migrating tables to it — but the direction is now legible, and one piece of it will matter a lot.
The pieces that landed: a V4 manifest reader, the ability to write Parquet and Avro manifests in the V4 layout, TrackedFile adapters bridging data and delete files with a format_version field, a key_metadata extension for V4 deletion vectors, and relative paths — in the spec, resolved in the manifest reader, with relativization utilities in core. On the spec side, 1.12 adds an expressions spec, content stats, and finer-grained read restrictions as part of loadTable.
Relative paths are the one to understand now. Today, Iceberg metadata records absolute file locations, which is why moving a table's data between buckets, regions, or accounts means rewriting metadata rather than copying bytes. Relative paths let the metadata describe locations relative to the table root, so the same metadata stays valid when the root moves. For anyone who has run a cross-region copy, a bucket migration, or a DR drill, that removes the most tedious part of the exercise.
Also new as an interface: a RepairTable action is now defined in the API. It is a contract rather than a finished implementation, but it signals that metadata repair is becoming a first-class operation rather than something each team scripts by hand.
What it means: nothing to adopt today. The planning implication is that V4 will make table relocation and metadata repair meaningfully easier, so large migrations that are not urgent may be worth sequencing after V4 rather than before it.
Two things worth knowing are not in it
Spark 4.2 support did not land. It was on the 1.12 milestone, then pulled: the change touched too large a surface to review in time, and the community preferred to ship it once the integration matures rather than rush a baseline. If your Spark roadmap assumed 4.2 on Iceberg this quarter, re-plan that dependency.
The companion Rust release is still in flight. iceberg-rust 0.11.0 — which brings deletion-vector Puffin decoding, V3 deletion vectors applied during scans, variant support, and row-lineage metadata columns to non-JVM engines — has not shipped; the latest release remains 0.10.1, with 0.11.0 at release candidate. RC2 was rejected for an instructive reason: the Rust encryption path skipped tamper-proofing checks Java requires, so files written by Rust were unreadable from Java. If you run Iceberg through a Rust-based engine, stage exactly that test — write an encrypted table with the candidate, read it back with Java — rather than assuming cross-implementation parity.
The operational changes that touch your runbook
Several smaller items in 1.12 change day-to-day maintenance more than the headline features do.
sort_byonrewrite_manifests(Spark 3.5 and 4.0) lets you control manifest organization rather than accepting the default grouping — directly useful for scan-planning latency on large tables.max-file-group-input-filesis now a valid rewrite option, giving you a ceiling on how much a single file group pulls in. Pair it withpartial-progress.enabledto keep long rewrites recoverable.- An executor cache for delete files during rewrites reduces repeated delete-file reads on merge-on-read tables being compacted.
rest-catalog-purgedelegatesDROP TABLE PURGEto the REST catalog instead of having the engine enumerate and delete files.ignore_missing_fileson bothmigrateandsnapshotprocedures makes onboarding messy existing datasets survivable.- An
unregistertable endpoint and alabelsfield for catalog metadata were added to the REST spec, exposed throughSupportsLabels. - Encryption got more complete: manifests written by
rewrite_manifestsare encrypted, manifest-list keys commit with the snapshot using them, and DV encryption metadata survives merges. - Parquet adaptive bloom filter sizing and per-column dictionary encoding give finer control over the write path.
Three correctness fixes deserve a close read. Time-travel snapshot lookup no longer assumes snapshot-log order, which could previously resolve the wrong snapshot. Delete manifest pruning in entries metadata tables was fixed, as was pruning for negated all_manifests filters — both affect anyone building health tooling on metadata tables. And scans now fail loudly when position deletes or deletion vectors do not match the data file partition, rather than silently returning wrong results.

Turning a release into an upgrade plan
Everything above is a per-table decision, and that is the part that does not scale by hand. Reading 1.12's notes takes twenty minutes. Determining which of 800 tables carries row-data position deletes, which has equality deletes worth converting, which was Z-ordered on a float, and which needs a manifest rewrite first is a different kind of work.
What a control plane does here: it already holds the per-table signals an upgrade plan needs — file-size distribution, small-file ratio, manifest count and depth, snapshot accumulation, delete-file ratio and type, partition skew, and how well the declared sort order matches the filters engines actually issue. Those are the same inputs that drive routine health scoring, which makes upgrade readiness a query against state you already collect rather than a one-off audit script.
How it works in practice: findings are ranked by severity with the issue named per table, so the output is an ordered work queue rather than a dashboard — tables blocking your upgrade surface above ones that are merely untidy.

For the remediation itself, the sequencing is where correctness lives: expire snapshots, then compact, then clean orphans, then rewrite manifests. Run compaction before expiry and you rewrite files whose old snapshots still pin the originals, reclaiming nothing. Adaptive cadence matters for the same reason it matters generally — a streaming CDC table accumulating equality deletes needs attention hourly, a monthly dimension effectively never — and a uniform nightly job gets both wrong.

Compaction execution is where the cost shows up at estate scale: running it on a purpose-built Rust engine on Apache DataFusion rather than a general-purpose JVM cluster is the difference between remediating a few tables and remediating all of them. The platform overview covers how that loop is assembled and whether it runs on autopilot, under policy, or with manual approval; table maintenance and observability cover the operations themselves and the signals behind the health score.
The upgrade sequence
Putting it together, in order:
- 1.Move engines first. Spark 3.4 and Flink 2.0 are gone. Get to a supported engine version before touching the Iceberg dependency.
- 2.Confirm JDK 17. Already required since 1.11, but easy to miss when skipping versions.
- 3.Audit delete files while still on 1.11. Find anything carrying row data and rewrite it. This is the step that causes real incidents if skipped.
- 4.Fix your AWS dependency pin.
apache-clientbecomesapache5-clientif you manage AWS SDK deps yourself. - 5.Capture a health baseline. File counts, manifest counts, snapshot depth, and query latency per table, so a post-upgrade regression is detectable rather than debatable.
- 6.Upgrade, then adopt deliberately. Convert equality deletes to deletion vectors where Flink is in the path. Re-cluster tables that were Z-ordered on floating-point columns, evaluating Hilbert against Z-order on a branch first. Revisit variant if scan performance was the blocker.
- 7.Leave V4 alone for now, and factor it into the timing of any large migration you have not yet started.
The honest summary of 1.12 is that its most valuable contents are a correctness fix, a migration path, and a removal list rather than a headline feature — which is what a maturing format looks like. Taken together they point one direction: the gaps that justified staying on v2 are closing, and the reasons to plan a v3 move are now concrete.
What does not change is that the release only matters table by table. Whether this upgrade is a quiet afternoon or a month of firefighting comes down to whether you can answer, for every table in the lake, what shape it is in right now. Teams who would rather not build that layer themselves can run LakeOps on the catalogs they already have. For the full changelog, the official Apache Iceberg release notes are the source of truth.



