Back to blog

Cross-Catalog Sync: Iceberg on Polaris, Glue, and Unity

Three Iceberg catalogs, one lake, and no single operational view. Why bidirectional catalog sync is a correctness bug, how federation actually works across Polaris, Glue, and Unity Catalog, and how to run one operating model over catalogs that will never merge.

Cross-catalog sync between Apache Polaris, AWS Glue, and Databricks Unity Catalog over shared Apache Iceberg tables

A lake with tables in Apache Polaris, AWS Glue, and Databricks Unity Catalog does not need a sync job. It needs a decision about authority. Apache Iceberg gives every table exactly one authoritative metadata pointer, and the catalog holding that pointer is the only system that can safely commit to it. Copying that pointer into a second catalog so both sides can write is not an integration pattern — it is a corruption pattern, and it fails in the specific way Iceberg was designed to prevent.

What has genuinely changed is reachability. All three catalogs now speak the Iceberg REST protocol, and all three federate. Glue can mount a remote Iceberg REST catalog, including Unity Catalog. Polaris can register an external catalog that forwards to another REST implementation. Unity Catalog exposes its tables outbound at an Iceberg REST endpoint and crawls external catalogs inbound as foreign catalogs. Getting an engine to see a table in another platform is close to a solved problem.

What is not solved is operational authority: which system is allowed to compact, expire snapshots, remove orphan files, and rewrite manifests on each table. The three platforms are not symmetric here, and the sharpest example surprises most teams. External Iceberg clients connected to Unity Catalog can read and write data in managed Iceberg tables, but they cannot run table maintenance operations such as expiring snapshots or removing orphan files. Reads converge. Writes mostly converge. Maintenance does not converge at all.

That asymmetry is why three catalogs still produce no single operational view, and more sync will never fix it. This guide covers what makes a catalog authoritative, the three different things people call cross-catalog sync, what each platform permits an outside operator to do, and how to run one operating model over catalogs that are never going to merge.

The catalog is a compare-and-swap, not a copy

An Iceberg table is a tree. The catalog holds a pointer to the current metadata file, which holds the schema, partition spec, sort order, properties, and snapshot list. A snapshot points to a manifest list, which points to manifests, which list data files with per-column statistics. Readers walk that tree top down, and the only mutable part of it is the pointer.

Commits work by conditional replacement. A writer reads the current metadata location, writes new data files, builds new metadata, then asks the catalog to move the pointer only if it still holds the value the writer started from. In Hive-backed and JDBC-backed catalogs you can see the mechanism directly: the table stores metadata_location and previous_metadata_location, and the commit is an update guarded by the old value. That guarded update is the compare-and-swap, and it is the entire reason several engines can write one table correctly.

The consequence is strict. A compare-and-swap only serializes writers inside one catalog. If the same table is registered in Glue and in Polaris, there are two independent pointers and no shared guard. Both catalogs will happily accept commits, neither will observe the other, and the table ends up with divergent metadata lineages over a shared set of data files. That is precisely the failure optimistic concurrency exists to prevent, reintroduced one layer up.

register_table is what makes this easy to do by accident. The procedure points a catalog at an existing metadata file. No data moves, nothing is rewritten, and it completes in milliseconds, which makes it feel like a cheap read-only mirror. It is not read-only — the second registration is a fully functional table that accepts writes. Treat register_table as a migration and recovery tool, not a replication tool. Registration with an overwrite or force flag deserves extra caution: implementations may drop and recreate the entry rather than swap it atomically, opening a window where readers and writers see no table at all.

Three different things people call cross-catalog sync

Most arguments about catalog synchronization are three separate mechanisms discussed as one. They have different safety properties and different correct uses.

MechanismWhat actually movesSafe for concurrent writesRight when
Pointer registrationA metadata-file path is registered in a second catalogNo — two independent compare-and-swap domainsOne-time migration, disaster recovery, or a frozen read-only copy you accept will go stale
FederationNothing. The local catalog proxies metadata requests to the remote owner at query timeYes — authority stays with one catalogPermanent multi-platform access to tables owned elsewhere
Platform-native syncThe owning platform publishes its own tables outward on a schedule or triggerYes, by construction — the publisher stays the only writerExposing Snowflake-managed or Delta-backed tables to outside engines

Only federation is safe by construction for a lake with multiple reader platforms, because it never duplicates authority — the local catalog becomes a permissioned view over a remote pointer. That is also why it carries a latency cost: every metadata resolution is a call to a system you do not control.

Platform-native sync is one-directional on purpose. Snowflake syncs its managed Iceberg tables into an Open Catalog external catalog, where other engines can read them but not write them. Delta UniForm generates Iceberg metadata over Delta files for external readers while Databricks keeps the write path. Neither is a round trip.

The layer that sits above all three catalogs

Before the platform specifics, it is worth naming the layer this problem belongs to, because it is the one piece of the open lakehouse stack that nothing in the stack provides.

Each catalog does its own job well: it holds pointers, swaps them atomically, authorizes access, and vends credentials. None of them is responsible for the health of the tables underneath — not the ones owned by the other two, and not even its own in any continuous sense. No catalog notices that a table has accumulated 40,000 small files, that 900 snapshots are pinning storage nobody reads, that manifest count has grown until planning takes seconds, or that a declared sort order stopped matching the queries hitting the table six months ago.

That gap is what a control plane fills. A control plane such as LakeOps sits above your catalogs and engines, reads metadata and query telemetry through the same APIs those engines already use, and runs a closed loop over the estate: observe table and query signals, classify each table's health, plan the operations it needs, execute them in dependency order where the owning platform permits it, and verify the outcome. It does not store your data, own table pointers, or sit in the query path.

Modern Iceberg lakehouse with a control plane above the catalogs, object storage below, and query engines on the right
Catalogs on the left, engines on the right, Iceberg tables in object storage below, and the operational layer above. The control plane reads metadata through standard catalog APIs rather than owning table pointers.

Why this is not just a pile of scripts

Every team with an Iceberg lake already has some version of this: Spark procedures on a cron, an Airflow DAG per catalog, a Slack alert when storage jumps. That is the right starting point, and for twenty tables it is enough. Five things break as the estate grows, and they are the difference between automation and a control plane:

  • Scripts trigger on a clock; a control plane triggers on state. A nightly job rewrites healthy tables and under-serves hot ones. Health-driven execution runs when the table's measured condition calls for it.
  • Scripts are per-catalog; a control plane is estate-wide. A Glue DAG does not understand Polaris. Nothing joins the two into one health picture — which is the three-catalog problem.
  • Scripts do not rank. At 2,000 tables the question is which forty are worst right now, and why. Severity ordering is the difference between a dashboard and a work queue.
  • Scripts have no feedback loop. A sort order hardcoded in a DAG is a guess made once. A control plane reads the columns queries actually filter and join on.
  • Scripts leave no coherent record. Three platforms, three audit schemas. Proving a retention claim means reconciling them by hand.

The capabilities that matter for this problem group into five areas: observability (one health score per table, plus query telemetry from every engine), maintenance (expire, compact, clean orphans, rewrite manifests — in that order), layout (sort and clustering chosen from real filters, tested on a branch), governance (declarative policy with inheritance and an audit trail), and routing (one endpoint over several engines). LakeOps is built on that model: it attaches to Glue, Polaris, Nessie, Gravitino, Lakekeeper, S3 Tables, and REST-compatible catalogs through metadata only, so adding a catalog does not move data or change pipelines.

Watch demoYouTube ↗
Autonomous Iceberg data lake management — telemetry collection across catalogs and engines, table health observability, then compaction, snapshot management, manifest optimization, and orphan cleanup running on their own.

The view a fragmented estate cannot assemble from three native consoles is the first thing to look at: lake-wide storage and compute, then every table's health in one list, then the operation log.

Dashboard overview with lake-wide KPIs for storage, CPU, and table health across connected catalogs
Storage, compute, and health trends computed across every connected catalog. Three platform consoles cannot produce this view, which is the operational gap federation leaves behind.
Tables list showing health status, sizes, and inventory across the lake
Healthy, Warning, and Critical across every attached catalog. This is the inventory a Head of Data cannot get by opening Databricks, Snowflake, and the Glue console in three tabs.

Keep that model in mind through the rest of this guide. The sections that follow are the engineering — what each platform permits, who should own what, and where the operational layer earns its place.

What each platform lets an outsider do

Read these as permission models, not feature lists. The question is not whether a catalog supports Iceberg REST — all three do. It is what an engine or operator authenticated from outside that platform is allowed to perform.

AWS Glue

Glue exposes an Iceberg REST endpoint at https://glue.<region>.amazonaws.com/iceberg, authenticated with SigV4 using the glue signing name. REST paths carry a /catalogs/{catalog} prefix, and engines pass the catalog identifier as warehouse. One structural constraint: Glue supports single-level namespaces only.

Inbound, catalog federation connects Glue to remote Iceberg REST catalogs, with a dedicated DATABRICKSICEBERGRESTCATALOG connection type for Unity Catalog and standard REST connections for Polaris and custom implementations. Metadata is fetched at query time, and Lake Formation vends scoped credentials so Athena, Redshift, and EMR read the underlying S3 files directly with fine-grained permissions applied.

The limit to design around: for federated Iceberg catalogs, Glue enforces a maximum metadata size of 5 MB per REST API call and rejects requests for tables above it. That turns metadata hygiene into an availability requirement rather than a performance nicety — a point that returns later.

Apache Polaris and Snowflake Open Catalog

Polaris supports Iceberg REST federation behind ENABLE_CATALOG_FEDERATION. You create an EXTERNAL catalog with an iceberg-rest connection, OAuth2 (or SigV4) credentials, and an optional remote catalog name when the far side multiplexes several under one URI. Two behaviors matter before you automate it: Polaris validates connectivity eagerly — creation fails if the remote endpoint is unreachable or rejects auth — and overwrite=true registration works only for internal catalogs. Federation surfaces Iceberg tables only.

Snowflake Open Catalog is the managed Polaris service, and its internal versus external distinction drives architecture. Tables in an internal catalog are read-write for both Snowflake and third-party engines. Tables in an external catalog are synced in from elsewhere and are read-only inside Open Catalog, with Snowflake retaining read and write on its own managed tables. Setting CATALOG_SYNC on a Snowflake schema publishes its managed Iceberg tables into the designated external catalog.

One sharp edge: Open Catalog prevents overlapping table directories in internal catalogs but does not enforce that for Snowflake-managed tables in an external catalog. You must set a unique BASE_LOCATION per table yourself. Shared or nested prefixes are exactly the condition that makes orphan-file cleanup destructive.

Databricks Unity Catalog

Unity Catalog serves the Iceberg REST protocol at /api/2.1/unity-catalog/iceberg-rest, and support depends entirely on table type:

Unity Catalog table typeExternal readExternal writeCredential vending
Managed IcebergYesYesRead, write, create
Foreign Iceberg (owned by Glue, HMS, Snowflake Horizon)YesNoNot supported
Managed Delta with Iceberg reads (UniForm)YesNoRead, limited write
External Delta with Iceberg readsYesNoRead, limited write

Enabling external access takes three steps: turn on external data access for the metastore, grant EXTERNAL USE SCHEMA on the schemas in scope, and authenticate with OAuth or a personal access token. Unity Catalog then vends short-lived storage credentials inheriting the configured principal's privileges.

The limitations are the architecturally interesting part. External Iceberg clients cannot access views, and cannot create managed tables using UUID, Fixed(L), TIME, or nested required STRUCTs. Foreign Iceberg tables are not refreshed automatically over the REST path — an external reader sees a stale snapshot until REFRESH FOREIGN TABLE. And external clients cannot run maintenance on managed Iceberg tables. Snapshot expiry and orphan removal stay inside Databricks, where Predictive Optimization owns them.

Inbound, Unity Catalog federates Glue, Hive metastores, and Snowflake Horizon into foreign catalogs that are read-only in Databricks, with their own constraints: Parquet only, no partition evolution, only the main branch visible, and time travel limited to snapshots Databricks has already read.

Reading the matrix

CapabilityGluePolaris / Open CatalogUnity Catalog
Serves Iceberg REST outboundYes, SigV4Yes, OAuth2Yes, OAuth or PAT
Federates remote catalogs inboundYes, incl. a Unity connectorYes, external catalogsYes, as read-only foreign catalogs
External engines can write dataYesInternal catalogs yes, external noManaged Iceberg only
External operator can run maintenanceYesInternal catalogs yesNo
Branching and tagging visible to consumersYesYesNo on foreign tables
Hard limit to plan for5 MB federated metadata per call; single-level namespacesEager connectivity check; no overwrite on federated catalogsNo views; no external maintenance; manual refresh on foreign tables

Three catalogs, three different answers to the only question that matters operationally. Any uniform maintenance strategy that assumes it can treat all tables the same way is wrong before it runs.

The maintenance authority matrix

Four procedures keep an Iceberg table healthy, and they run in dependency order: expire snapshots, compact data files, remove orphan files, rewrite manifests. Expiry comes first because old snapshots pin data files that compaction would otherwise rewrite without reclaiming anything. Orphan removal follows compaction because compaction itself produces orphans when a rewrite loses the commit race. Manifest rewrite goes last, against the final file set.

That ordering is a property of Iceberg. What varies is who may execute each step.

OperationGlue-ownedPolaris internalUnity managed IcebergSnowflake-managed, synced out
expire_snapshotsExternal operatorExternal operatorDatabricks onlySnowflake only
rewrite_data_filesExternal operatorExternal operatorDatabricks onlySnowflake only
remove_orphan_filesExternal operatorExternal operatorDatabricks onlySnowflake only
rewrite_manifestsExternal operatorExternal operatorDatabricks onlySnowflake only
Who decides layoutYou, explicitlyYou, explicitlyPredictive OptimizationSnowflake

So a single uniform maintenance job across the estate cannot exist. You get a split operating model whether you plan for it or not: tables in Glue and Polaris internal catalogs are externally operable, and tables owned by Unity Catalog or Snowflake are operated by their platform, using that platform's signals.

The design that works accepts the split and makes it observable. For externally operable tables, run maintenance with full control over strategy and sequencing. For platform-owned tables, delegate execution but keep measurement — you still want file-count distribution, snapshot depth, manifest count, and delete-file ratio in the same health model, because that is how you discover a platform optimizer is not keeping up with a table your dashboards depend on. Delegating execution is reasonable. Delegating visibility is how tables rot unnoticed.

This is where a uniform health model stops being a nicety. Scoring every table the same way — Healthy, Warning, or Critical from file count, small-file ratio, snapshot age, manifest bloat, and write velocity — is what lets a Glue table and a Unity Catalog table appear in one ranked list even though only one of them is yours to fix (lakehouse observability covers the signals that feed the score). The ranking matters more than the dashboard: with thousands of tables, the useful output is the forty worst right now and the specific reason each one is on the list.

Insights view ranking Iceberg table health issues by severity across catalogs
Findings ranked by severity with the issue named per table — excessive manifests, small-file ratio above threshold, partition file counts, scan amplification. Ranking is what turns observability into a work queue.

Orphan file removal is the cross-catalog footgun

Of the four procedures, remove_orphan_files is the one that destroys data, and multi-catalog estates create exactly the conditions where it does.

The mechanism explains the danger. The procedure lists files under the table's storage location, compares that listing against files referenced by live metadata, and deletes the difference. The reference set comes from one catalog's view of the table. Anything that makes that view incomplete turns the procedure into a deletion of live data.

  • Shared or nested storage prefixes. If two tables live under overlapping paths, cleanup on one deletes files belonging to the other — the Open Catalog BASE_LOCATION hazard from earlier.
  • Running from a non-owning catalog. A federated or registered copy may resolve to an older metadata file. Files written since that snapshot look unreferenced and get deleted while the owner still points at them.
  • A stale federated view. Unity Catalog does not auto-refresh foreign Iceberg tables over the Iceberg REST path, so a cleanup job driven by that view may be reasoning about a snapshot hours old.
  • In-flight writes. A writer that has staged data files but not yet committed has nothing in metadata referencing them. An aggressive older_than window deletes them mid-transaction.

Here is the sequence to automate, with the parameters that make it safe. It assumes you are running against the owning catalog, which is the first precondition and the easiest to get wrong:

sql
1-- Run against the OWNING catalog only. Never a federated or registered copy.2 3-- 1. Expire first: old snapshots pin data files that compaction would rewrite4--    without reclaiming anything.5CALL prod.system.expire_snapshots(6  table       => 'analytics.orders_fact',7  older_than  => TIMESTAMP '2026-09-18 00:00:00',8  retain_last => 109);10 11-- 2. Compact. Sort by columns production queries actually filter on, not by a12--    guess made at table-create time. Partial progress keeps work after failure.13CALL prod.system.rewrite_data_files(14  table       => 'analytics.orders_fact',15  strategy    => 'sort',16  sort_order  => 'customer_id ASC NULLS LAST, event_date ASC NULLS LAST',17  options     => map(18    'target-file-size-bytes',              '268435456',19    'min-input-files',                     '5',20    'delete-file-threshold',               '2',21    'partial-progress.enabled',            'true',22    'max-concurrent-file-group-rewrites',  '4'23  )24);25 26-- 3. Clean the orphans that rewrite just produced. The safety window must exceed27--    the longest in-flight write AND the longest federated staleness window of28--    any consumer catalog. Multi-day is normal; hours is not.29CALL prod.system.remove_orphan_files(30  table      => 'analytics.orders_fact',31  older_than => TIMESTAMP '2026-09-25 00:00:00'32);33 34-- 4. Rewrite manifests last, against the final file set.35CALL prod.system.rewrite_manifests(table => 'analytics.orders_fact');

Two parameters deserve attention in a federated context. max-concurrent-file-group-rewrites is the throughput-versus-contention dial: higher finishes sooner and loses more commit races against active writers, and in a multi-platform lake you often cannot see who is writing. And the older_than value on orphan removal should be derived from your worst-case consumer staleness, not a tidy default — a seven-day window costs very little storage and removes an entire class of incident.

The other half of running this safely is being able to prove what happened. A per-operation record — trigger, duration, files before and after, bytes reclaimed, status — is what turns "cleanup ran" into evidence, and it is the only practical way to answer a GDPR erasure request spanning three platforms, where the chain is delete rows, compact to rewrite the affected files, expire the snapshots still referencing the originals, then remove the resulting orphans, in that order.

Lake-wide events audit trail showing every maintenance operation with trigger, status, and bytes reclaimed
One operation log across catalogs, with trigger, duration, files affected, and bytes reclaimed per event — rather than three native audit schemas to reconcile at audit time.

Assign write authority per table, and write it down

This is the decision that prevents most multi-catalog incidents, and it is almost always implicit. For every table, exactly one catalog owns the pointer; every other catalog reaches it by federation. Make that explicit, version it, and enforce it with credentials rather than convention.

yaml
1# catalog-authority.yaml — reviewed in PRs, enforced by IAM and RBAC2tables:3  - name: analytics.orders_fact4    owner_catalog: glue:prod-analytics      # holds the pointer, accepts commits5    writers: [spark-etl-prod]               # principals allowed to commit6    consumers:7      - polaris:open-catalog-bi             # federated, read-only8      - unity:prod_federated                # foreign catalog, read-only9    maintenance_owner: control-plane        # external operator permitted10    retention_days: 1411 12  - name: ml.feature_store_v313    owner_catalog: unity:prod               # UC managed Iceberg14    writers: [databricks-ml-jobs]15    consumers:16      - glue:prod-analytics                 # federated via Unity REST connector17    maintenance_owner: databricks           # external maintenance not permitted18    retention_days: 30

The maintenance_owner field is the one most teams are missing, and for a Unity Catalog managed table it is not a preference — the platform forces it. Recording it stops someone from scheduling a compaction job that will fail, or worse, writing a cleanup job that runs against storage without going through the owning catalog at all.

Enforcement belongs in the credential layer. Only the owning catalog's principal should hold write permission on the table's storage prefix; consumers get read-only credentials vended by their federated catalog. If a second catalog can mint write credentials for the same prefix, the registry is documentation rather than a control.

A registry in YAML tells you the intent. Keeping reality matched to it is the harder half, and it is why retention and maintenance are better expressed as policy scoped above the catalog than as parameters inside per-catalog jobs. A rule written once at namespace scope — retention window, compaction threshold, cleanup cadence — applies to tables in Glue and Polaris alike, and new tables inherit it the moment they are created instead of being discovered unmanaged months later. The per-table retention windows the earlier section argued for only stay accurate if they live somewhere versioned and reviewable — the model policy-based governance is built around.

Policy list showing compaction, orphan cleanup, snapshot expiry, and manifest rewrite schedules by scope
Operational policies by type and scope, with inheritance from organization through catalog and namespace down to table. Namespace-level policy governs a table at creation rather than being applied months later.

Pick a topology before you pick tools

TopologyHow it worksBest whenMain failure mode
Hub and spokeOne catalog is authoritative for most tables; others federate inA clear platform of record exists and most writes originate thereThe hub becomes a hard availability dependency for every consumer
Domain meshEach domain owns tables in its own platform; every platform federates the othersTeams genuinely own different platforms and write locallyFederation topology grows quadratically; credential and network config multiply
Metalake brokerA dedicated federated metadata layer fronts all catalogs under one namespaceMany catalogs, many engines, and a need for one logical hierarchyAnother system in the metadata path to operate and keep available

Apache Gravitino is the open-source option for the third row. It organizes metadata into a metalake providing a catalog.schema.table namespace and fronts multiple Iceberg backends — Hive, JDBC, REST, and Glue — behind a single Iceberg REST service, with nested multi-level namespaces, vended-credential refresh, and geo-distributed deployments where a local REST service proxies to a remote one. That last capability is the cleanest answer when cross-region latency is the real constraint.

A broker is genuinely useful for namespace unification and engine simplification. It does not change authority: the backing catalog still owns the pointer, and a broker in front of Unity Catalog still cannot run maintenance on UC managed tables. Choose a topology based on where write authority should live, not where you want metadata to appear.

Worth separating from topology is the observability question. Attaching a catalog read-only to collect metadata costs you nothing architecturally — no pointer ownership, no second compare-and-swap domain, no change to how engines resolve tables — which means the decision to get one health picture across all three catalogs is independent of, and much cheaper than, the decision about where writes live.

Multiple catalogs connected side by side — Glue, DynamoDB-backed, Iceberg REST, and S3 Tables
Different catalog types attached in parallel through the APIs engines already use. Namespaces and tables are discovered rather than hand-registered, so a table created directly in a platform does not silently stay invisible.

Retention and the federated staleness problem

Snapshot retention is a cost dial and a correctness contract at once. Every commit creates a snapshot pinning the data files it references, so ninety days of history on a high-velocity table can pin several times the live footprint. Retention generous enough for an auditor is expensive on a streaming table and free on a monthly dimension, which is why a single lake-wide number is always wrong somewhere.

Federation adds a second consequence: expiry can break a consumer. Time-travel queries from another platform reference snapshots by ID or timestamp, and expiry invalidates them. Unity Catalog makes this concrete — time travel on foreign Iceberg tables works only for snapshots Databricks has already read, so a consumer's usable history is a subset of the real history and shrinks when you expire. If downstream platforms have time-travel contracts, retention has to be negotiated rather than set per table in isolation. Tagging deserves the same check: a tag you created as an audit anchor in Glue is invisible to a Databricks consumer, which sees only the main branch.

Iceberg snapshot history with tags, branches, and rollback controls
Snapshot history with tagging, branching, and rollback. Retention sits next to the snapshots it governs, which is what makes a per-table window reviewable instead of a lake-wide guess.

There is a non-obvious interaction worth internalizing: metadata hygiene becomes an availability requirement. Glue rejects federated Iceberg requests for tables whose metadata exceeds 5 MB per REST call, and metadata grows with snapshot count, manifest count, and schema history. A table nobody has expired, or whose manifests were never consolidated, can cross that ceiling and stop being readable through Glue federation — not slowly, but as a hard rejection. Manifest rewrite and snapshot expiry are usually discussed as planning-latency optimizations. In a federated topology they keep the table reachable.

Physical layout diverges per platform, and queries pay for it

Most lakehouse performance problems are pruning problems, and the word pruning covers several mechanisms that fail independently. Across three platforms they get tuned by three different parties.

  1. 1.Partition pruning (coarse). The Iceberg partition spec — year, month, day, bucket, truncate — decides which file groups can be ignored entirely. Over-partitioning (hourly on a moderate table) creates a metadata problem worse than the scan it was meant to fix.
  2. 2.File pruning via sort order, or linear clustering (fine). Nearby key values in the same files tighten min/max stats. A sort on (customer_id, event_date) prunes customer_id excellently and event_date alone poorly.
  3. 3.Z-order, or multi-dimensional clustering. Bit-interleaving so two to four filter columns all prune reasonably; past four columns the benefit dilutes toward none.
  4. 4.Row-group pruning inside Parquet, which only helps if the file is clustered.
  5. 5.File sizing, typically 128–512 MB for analytical tables.
  6. 6.Metadata health — manifests, snapshot depth, delete files — which sets planning cost before any data is read.

Binpack fixes file count. It does not cluster rows. Merging small files by size alone leaves min/max ranges wide, so file pruning still fails — which is how a table gets compacted for months while queries never speed up. And cluster means three different things: a sort strategy, file sizes grouping around a target, and a Spark cluster. Databricks Liquid Clustering, BigQuery clustering, and Snowflake micro-partitions are not Iceberg sort order. Compare them; do not equate them.

On Unity Catalog managed tables, Predictive Optimization decides layout from Databricks workload signals. A Trino or Athena consumer reaching that table through Glue federation contributes nothing to that decision and pays full scan cost for filters the optimizer never saw. Tables in Glue and Polaris keep whatever sort order someone declared at create time, because nothing in Iceberg re-derives it. One logical dataset, three layout regimes, and query cost that depends on which platform happened to own the pointer.

The fix is to derive layout from observed access: collect WHERE, JOIN, and GROUP BY frequency from every engine reading the table, rank candidate sort columns, and test the winner on an Iceberg branch before rewriting production files. When three platforms read the table, all three should inform the choice.

Layout simulations comparing query-aware sort strategies against the current baseline before any rewrite
Candidate clustering strategies scored against observed field-access frequency and diffed against the current layout on a branch. Changing clustering rewrites files, which is why the comparison belongs here first.

If you are weighing binpack against sort and Z-order for a specific table, the compaction strategies guide works through when each wins. At estate scale the execution engine matters too: compaction on Rust and Apache DataFusion rather than a general-purpose JVM cluster is the cost difference between occasional and continuous maintenance. The platform overview covers how the observe-classify-plan-execute loop is assembled — autopilot, policy-driven, or with manual approval.

A drift and validation runbook

Federation fails quietly. A consumer reads a stale snapshot, a credential expires, a schema change does not propagate, and nothing alerts because every individual component is healthy. These are the checks worth scheduling.

  1. 1.Pointer ownership. Confirm exactly one catalog holds a writable registration per table. Unexpected duplicates are the highest-severity finding available, because they mean two compare-and-swap domains are live on one table.
  2. 2.Snapshot parity. Compare the current snapshot ID each consumer catalog sees against the owner. A persistent gap means federation is not refreshing; a growing gap means it stopped.
  3. 3.Schema drift. Verify column sets and types match across federated views. Paths that bypass the Databricks runtime never trigger a foreign-table refresh, so a consumer can hold an old schema indefinitely.
  4. 4.Metadata size headroom. Track metadata size per federated table against the 5 MB Glue ceiling and alert at 60 percent, not at failure.
  5. 5.Credential reachability. Confirm each federated connection can still vend credentials. Eager validation at creation time does not help when a secret rotates six months later.
  6. 6.Maintenance coverage. For every table, confirm something is actually maintaining it, and record what. A table whose maintenance_owner is a platform optimizer still needs a health check, because delegation is not evidence.

Iceberg metadata tables give you most of this without extra infrastructure. Run them against the owning catalog:

sql
1-- Snapshot depth and age: inputs to retention and to metadata size2SELECT count(*) AS snapshot_count, min(committed_at) AS oldest, max(committed_at) AS newest3FROM prod.analytics."orders_fact$snapshots";4 5-- Manifest count: planning overhead and the main driver of metadata size6SELECT count(*) AS manifest_count FROM prod.analytics."orders_fact$manifests";7 8-- File size distribution: is compaction keeping up with ingest?9SELECT count(*) AS file_count,10       avg(file_size_in_bytes) AS avg_bytes,11       sum(CASE WHEN file_size_in_bytes < 33554432 THEN 1 ELSE 0 END) AS small_files12FROM prod.analytics."orders_fact$files";13 14-- Delete-file burden: merge-on-read cost paid on every read15SELECT content, count(*) AS files FROM prod.analytics."orders_fact$all_files" GROUP BY content;

The queries are simple. The point is that nobody runs them across every table in three catalogs by hand, and the cadence question compounds it — a streaming table needs compaction hourly, a daily batch table daily, a slowly-changing dimension effectively never. A uniform schedule either under-maintains the hot tables or burns compute rewriting cold ones, and both failures stay invisible until someone goes looking.

Per-table adaptive maintenance settings for compaction, snapshot expiry, manifest rewrite, and orphan cleanup
Each operation taking its cadence from that table's write velocity rather than a shared nightly window — the only way a split operating model across three catalogs stays coherent.

Where this leaves you

At twenty tables across three catalogs this is tractable by hand — do not over-build. At two hundred the registry stops matching reality, and the first orphan-cleanup incident usually arrives because a job written for one catalog gets pointed at a table owned by another. At two thousand the operating model is the system: nobody can say which table is worst right now without joining metadata across every catalog, and most of the lake is unmanaged by default.

So stop trying to make the catalogs converge. In most enterprises they will not, and the Iceberg REST protocol has made that acceptable — federation gives every platform access to tables it does not own, without copying data and without a second compare-and-swap domain. Reachability is a solved problem you configure once and stop thinking about.

Converge the operating model instead, which is three concrete things. Write down which catalog owns the pointer for every table and enforce it with credentials rather than convention. Record which system is permitted to maintain each table, accepting that Unity Catalog and Snowflake reserve that for themselves. And measure every table the same way regardless of who maintains it, because delegating execution is reasonable while delegating visibility is how tables degrade for six months before anyone notices.

The decision worth making deliberately is not which catalog wins. It is whether the operational layer across them is a designed system with an owner, or a pile of per-platform jobs that gets rewritten the next time the estate changes shape. Teams who would rather not build that layer themselves can run LakeOps on the catalogs they already have — attach Glue and Polaris, observe Unity Catalog tables through federation, and get one health picture without moving data. If the adjacent question is which catalog should be authoritative, the Iceberg catalog comparison weighs the options; catalog migration covers the mechanics if the answer is consolidation after all.

Related articles

Found this useful? Share it with your team.