Governance

Iceberg governance — define once, every table follows

Declarative maintenance policies, GDPR-compliant retention, and per-operation audit trails — across every catalog and engine in your Iceberg lakehouse.

LakeOps LogoLakeOps

Last 30 days Optimization Activity

Total Operations
12,211
Last 90 days
Query Speed
12.4×
Avg. acceleration across engines
Cost Savings
$1,374,672
Saved in last 3 months
CPU & Storage
-76%
Last 90 days
Data Optimized
46.8 PB
Last 30 days

Key Metrics

Total Tables
786
Tables in all catalogs
Critical Tables
70
Require immediate attention
Warning Tables
105
Should be addressed or auto-piloted
Healthy Tables
566
Tables in optimal state
Total Data
112.4 PB
Total lake data size

Storage

-56% reclaimed
30d
112 TB75 TB38 TB
Storage used

CPU

-76% reduction
30d
100%66%33%
Compute hours

Recent Operations

Last 10 operations
OperationTableDurationImpactTimeStatus
Compact Data Files
customer_orders
orders
4s1.24 TB, 16 → 1 files57 minutes agoSUCCESS
Expire Snapshots
payment_transactions
payments
27s8.2 TB4 hours agoSUCCESS
Expire Snapshots
inventory_snapshots_20250702
warehouse
3s2.1 TB4 hours agoSUCCESS
Rewrite Manifests
raw_clickstream
analytics
1.9s3 → 1 manifests5 hours agoSUCCESS
Compact Data Files
product_catalog
products
6m 11.3s3,008 → 1,256 files6 hours agoSUCCESS

Lake Events

LiveLast 24 hours
Compact Data Files·customer_orders
ecommerce_prod·1.24 TB, 16 → 1 files
4s57 min agoOK
Expire Snapshots·payment_transactions
ecommerce_prod·12 snapshots expired
4.6s1h agoOK
Compact Data Files·raw_clickstream
marketing_events·970 → 87 files
9m 31.6s2h agoOK
Remove Orphan Files·user_sessions
marketing_events·847 MB reclaimed, 1,203 files
1m 12s3h agoOK
Rewrite Manifests·search_query_logs
ecommerce_prod·487 → 12 manifests
2.1s3h agoOK
Expire Snapshots·inventory_levels
warehouse_analytics·62 snapshots, 18.4 GB freed
27s4h agoOK
Compact Data Files·product_catalog
ecommerce_prod·3,008 → 1,256 files
6m 11.3s5h agoOK
Rewrite Manifests·shipping_events
warehouse_analytics·14 → 3 manifests
1.0s6h agoOK
Remove Orphan Files·balance_snapshots
warehouse_analytics·59,831 files, 74.8 GB
13m 6.9s7h agoOK
Compact Data Files·ad_impressions
marketing_events·42,633 → 69 files
2m 18s8h agoOK
No vendor lock-in
No code / infra changes
No data changes
REST API & MCP

Open and runs on your stack

AWS
Azure
Google Cloud
Snowflake
Databricks
Apache Flink
Apache Iceberg
Delta Lake
DuckDB
Dremio
Lakekeeper
ClickHouse
AWS
Azure
Google Cloud
Snowflake
Databricks
Apache Flink
Apache Iceberg
Delta Lake
DuckDB
Dremio
Lakekeeper
ClickHouse

How it works

From lake-wide policies to per-table compliance proof

Define governance at the organization level. Drill into each table's policy state, execution history, and retention status.

LakeOps Policies dashboard — organization-wide maintenance policies with status toggles, types, schedules, and last run timestamps

Organization-wide policies

Every policy in one dashboard — toggle, schedule, audit

The Policies dashboard lists every active and inactive policy with its type, scope, schedule, last run, and status toggle. Create, edit, enable, or disable any policy from one place — across every catalog and namespace in your lake.

  • Status toggle per policy — enable or disable in one click
  • Cron-based scheduling with event-driven override triggers
  • Full version history — rollback any policy change instantly
LakeOps policy creation wizard — six maintenance operation types: Expire Snapshots, Remove Orphan Files, Compact Data Files, Rewrite Manifests, Rewrite Position Delete Files, and Rewrite Equality Delete Files

Six policy types

Every maintenance operation as an auditable policy

Expire snapshots, remove orphan files, compact data files, rewrite manifests, rewrite position delete files, and rewrite equality delete files — each configurable with retention windows, age thresholds, target sizes, and merge parameters.

  • Expire Snapshots — time-based and count-based retention with safety windows
  • Remove Orphan Files — age thresholds coordinated after snapshot expiration
  • Compact, Rewrite Manifests, and Delete Files — target sizes and merge thresholds
LakeOps events — lake-wide audit trail for compaction, snapshot expiration, orphan cleanup, and manifest rewrite operations

Audit trail

Every operation logged — lake-wide and per table

Every compaction, snapshot expiration, orphan removal, and manifest rewrite is logged with full context: duration, files before and after, bytes reclaimed, status, and the policy that triggered it. Filter by catalog, operation type, or time range.

  • Before/after file and manifest counts on every operation
  • Policy execution history — which rule triggered each action
  • Compliance-ready trail for SOC 2, HIPAA, and regulatory evidence
Policy InheritanceHierarchy
Organization·All tables

Baseline retention, orphan cleanup

Catalog·ecommerce_prod

Daily compaction, 7-day retention

Namespace·orders

Sort compaction, hourly expiry

Table·customer_orders

GDPR 30-day, 6h compaction

Table-level overrides take precedence. New tables inherit from their namespace automatically.

Policy inheritance

Organization → catalog → namespace → table

Set baseline policies at the organization level. Override at the catalog for specific environments. Fine-tune per namespace or individual table. New tables inherit governance rules automatically — no manual setup, no gaps.

  • New tables inherit rules from their namespace and catalog automatically
  • Table-level overrides take precedence without breaking the hierarchy
  • Exclude specific tables from inherited policies when needed
Retention RulesEnforcing

Snapshot retention

Count max: 50

7 days

GDPR deletion

Full audit trail

30 days

Orphan cleanup

Safety window active

3 days

Pipeline: DELETE → compact → expire → orphan cleanup

Retention & compliance

GDPR-compliant deletion with full audit proof

Set snapshot retention periods, orphan cleanup schedules, and data deletion policies. LakeOps enforces the full GDPR pipeline — DELETE → compact → expire snapshots → remove orphan files — and logs every step for compliance evidence.

  • Time-based and count-based retention with configurable safety windows
  • Automated GDPR deletion pipeline with deterministic physical erasure
  • Per-operation audit trail for SOC 2, HIPAA, and regulatory evidence
Cross-Catalog GovernanceUnified

AWS Glue

1,240 tables

Polaris

860 tables

Nessie

430 tables

Gravitino

215 tables

One policy, all catalogs

Zero drift between governance and table state

Cross-catalog enforcement

Same governance on Glue, Polaris, Nessie, and Gravitino

Whether your tables live in AWS Glue, Apache Polaris, Project Nessie, Apache Gravitino, Lakekeeper, or S3 Tables — governance policies apply uniformly. No per-catalog scripts, no engine-specific maintenance logic.

  • Uniform policy enforcement across all connected catalogs
  • Engine-agnostic — works with Trino, Spark, Snowflake, Databricks, and more
  • Zero drift between declared governance and actual table state

Two governance layers

Access governance lives in the catalog. Operational governance lives in LakeOps.

Catalogs like Polaris and Glue control who accesses data. LakeOps controls how data is maintained, retained, and audited — the layer the catalog doesn't cover.

Catalog RBAC

Polaris, Glue IAM, and Nessie enforce role-based access at the catalog layer. LakeOps connects to all of them — governance policies apply alongside your existing access controls.

Credential vending

Short-lived, table-scoped storage tokens replace long-lived cloud keys. Catalogs issue credentials per request — LakeOps operates within the same security boundary.

Read Restrictions (Iceberg 2026)

The new REST catalog spec standardizes row filters and column masks at the catalog level. LakeOps complements access governance with operational governance — the layer the catalog doesn't cover.

Lakehouse Control Plane

Governance is one piece. See the full platform

Policies sit next to compaction, snapshot expiry, query routing, observability, and MCP access — same catalogs, same tables.

Table Health Overview786 tables
566
Healthy
105
Warning
70
Critical
92%
Optimized
TableNSSizeStatus
customer_ordersorders1.24 TBHEALTHY
payment_transactionspayments860 GBWARNING
raw_clickstreamanalytics4.6 TBCRITICAL
product_catalogproducts42 GBHEALTHY
user_sessionsanalytics1.9 TBWARNING
inventory_levelsoperations320 GBHEALTHY
shipping_eventslogistics580 GBHEALTHY
search_query_logsanalytics3.2 TBCRITICAL

Observability

Learns from query patterns and table signals to decide what to optimize, when, and how — no schedules, no manual tuning.

  • Closed-loop sense → plan → execute → learn
  • Zero manual scheduling or threshold tuning
  • Adapts to workload changes in real time

The gap

Iceberg gives you metadata — not governance

Without declarative policies and continuous enforcement, table maintenance depends on scripts that drift, tribal knowledge that doesn't scale, and compliance evidence assembled after the fact.

Maintenance scripts drift — policies don't

Cron jobs and Spark scripts break on cluster upgrades, skip tables after schema changes, and nobody notices until storage costs spike or queries time out. Tribal knowledge replaces auditable control.

Access control doesn't cover operations

Catalog RBAC controls who reads and writes tables. It says nothing about when snapshots expire, how often compaction runs, or which orphan files get cleaned up. Operational governance is a separate layer entirely.

Compliance evidence is assembled manually

Auditors need proof that GDPR deletion requests were fully resolved — data deleted, compacted, expired, orphan files removed. Without per-operation audit trails, teams spend days assembling evidence from scattered logs.

Multiple catalogs mean multiple scripts

Glue tables get one set of maintenance rules. Polaris tables get another. REST catalog tables get nothing. Governance fragments across catalogs instead of applying uniformly from one place.

Resources

Learn more

Explore all articles

Connect a catalog. Policies apply from Iceberg metadata.

LakeOps reads Iceberg metadata from every connected catalog, discovers tables, and applies governance rules automatically — no scripts, no pipeline changes.

1

Connect catalogs & engines

Only metadata is processed — never retained or stored.

Apache Iceberg
AWS
Snowflake
DuckDB
2

Get visibility & insights

Telemetry reveals table health and actions needed.

Table health scores
Optimization opportunities
Cost & performance insights
3

Choose your mode

Autopilot, manual approval, or policy-driven.

Autopilot
Manual
Policies
4

Lakehouse optimized

Queries 10x faster
Cost down 76%
Engines optimized
AIs managed
Tables healthy
Fully governed
Set up in 10 minutes · Works with your existing stack