
Managed Iceberg in 2026: Autonomous Data Lake
Iceberg tables degrade silently — small files pile up, snapshots bloat metadata, and query latency creeps higher. A breakdown of the nine components every production data lake needs to stay healthy.
Platform
Autonomously optimizes Iceberg table health, query speed, and cost — across your entire lakehouse.
Open and runs on your stack
| Table | NS | Size | Status |
|---|---|---|---|
| customer_orders | orders | 1.24 TB | HEALTHY |
| payment_transactions | payments | 860 GB | WARNING |
| raw_clickstream | analytics | 4.6 TB | CRITICAL |
| product_catalog | products | 42 GB | HEALTHY |
| user_sessions | analytics | 1.9 TB | WARNING |
| inventory_levels | operations | 320 GB | HEALTHY |
| shipping_events | logistics | 580 GB | HEALTHY |
| search_query_logs | analytics | 3.2 TB | CRITICAL |
Every table scored as Healthy, Warning, or Critical — based on file count, small-file ratio, snapshot age, manifest bloat, and write velocity. LakeOps surfaces degradation before it reaches your queries, with unified telemetry across every engine in your stack.
Compaction
38% small files — merging 970 → 87 at 512 MB target
Expire Snapshots
154 snapshots, 62 past 30-day retention
Rewrite Manifests
12 manifests — below threshold, waiting for compaction
Orphan Cleanup
847 MB unreferenced — scheduled after expiration
Query patterns
event_date, region
Top sort columns (Trino + Spark)
Improvement
12.4× faster
Avg query speed after optimization
Cycle
Self-tuning
Sort orders adapt as patterns change
No cron jobs. No Airflow DAGs. No runbooks. LakeOps scores each table's health, decides which operations to run, and sequences them in dependency order — expiry before compaction, compaction before cleanup. Each step's output is the next step's clean input.
View and track table operations and history
| Table Name | Namespace | Catalog | Operation Type | Impact | Duration | Start Time | Status |
|---|---|---|---|---|---|---|---|
| customer_orders | orders | ecommerce_prod | Compact Data Files | 1.24 TB, 16 → 1 files | 4s | 57 min ago | SUCCESS |
| customer_orders | orders | ecommerce_prod | Compact Data Files | 11 → 8 files | 10.3s | 2h ago | SUCCESS |
| customer_orders | orders | ecommerce_prod | Compact Data Files | 970 → 87 files | 6m 1.0s | 3h ago | SUCCESS |
| payment_transactions | payments | ecommerce_prod | Expire Snapshots | 1 snapshot | 4.6s | 4h ago | SUCCESS |
| payment_transactions | payments | ecommerce_prod | Rewrite Manifests | 3 → 1 manifests | 1.9s | 4h ago | SUCCESS |
| raw_clickstream | analytics | marketing_events | Compact Data Files | 970 → 87 files | 9m 31.6s | 5h ago | SUCCESS |
| user_sessions | analytics | marketing_events | Remove Orphan Files | 1,203 files, 847 MB | 1m 12s | 5h ago | SUCCESS |
| product_catalog | products | ecommerce_prod | Compact Data Files | 3,008 → 1,256 files | 6m 11.3s | 6h ago | SUCCESS |
| user_sessions | analytics | marketing_events | Rewrite Manifests | 2 → 1 manifests | 1.8s | 7h ago | SUCCESS |
| inventory_levels | operations | warehouse_analytics | Remove Orphan Files | 4,218 files, 1.9 GB | 3m 44s | 7h ago | SUCCESS |
| inventory_levels | operations | warehouse_analytics | Expire Snapshots | 12 snapshots expired | 2.0s | 8h ago | SUCCESS |
| shipping_events | logistics | warehouse_analytics | Rewrite Manifests | 3 → 1 manifests | 1.0s | 8h ago | SUCCESS |
Every compaction, snapshot expiration, orphan cleanup, and manifest rewrite is logged as an event — across all catalogs and namespaces. Filter by table, operation type, or status to trace exactly what happened, when, and what it reclaimed. The same feed powers governance audit trails and compliance evidence for SOC 2 and GDPR.
raw_clickstreamCritical312 partitions exceed file threshold
Query scan amplified 8×
search_query_logsHighExcessive manifests (487) — planning overhead
Planner latency +2.1s
payment_transactionsWarningSmall file ratio 38% — compaction recommended
S3 GET costs elevated
1.8 TB
Space freed
4.2 TB
Optimized
247
Tables
2%
Failure rate
Register your catalogs and every table gets maintained continuously. Streaming tables hourly, batch tables daily, idle tables skipped — cadence adapts to write velocity. The full metadata pipeline runs lake-wide without a single additional ops task.
customer_idevent_dateproduct_idregionBefore
After
12×
Faster queries
76%
Less CPU
95%
vs Spark
90%
Cheaper ops
200 GB compaction · same hardware
Most compaction tools merge small files blindly. LakeOps watches which columns your queries actually filter, join, and group on — then physically re-sorts data to match. Engines skip entire file groups via min/max pruning, cutting I/O by up to 90%. Built on Rust and Apache DataFusion: 95% faster and 90% cheaper than Spark-based compaction.
Run simulations on customer_orders
| Simulation | data_rel | Strategy | Avg Size | |
|---|---|---|---|---|
| meta | clusterByOrderDate | customer_id, order_status | order_date (day) | 343 MB / file |
| meta | cluster.order_type.by.status | order_status, payment_method | order_status, payment_method | 511 MB / file |
| layout | cluster.insert-time-line | customer_id, store_id | created_at (hour) | 128 MB / file |
LakeOps tests layout changes on Iceberg branches before they ever touch production — sort orders, partition strategies, and file targets against real query patterns. Compare scan reduction, file count, and estimated speedup side by side, then apply the winning change with one action.
Manage maintenance, configuration, and lifecycle policies for your data lakehouse
| On | Policy | Type | Next Run | Last Run | Updated | Actions |
|---|---|---|---|---|---|---|
orders_critical | Compact Files | Apr 25, 2026, 8:12 AM | Apr 25, 2026, 03:05 AM | Feb 01, 2025, 3:46 PM | ••• | |
payments_compact | Compact Files | Feb 15, 2026, 12:06 AM | Feb 1, 2025, 02:18 PM | Feb 5, 2025, 4:03 PM | ••• | |
Remove orphan files (e-ip...) For all tables in all catalogs every 7 days | Orphan Files | Apr 25, 2026, 8:12 AM | Apr 25, 2026, 04:07 PM | Jun 23, 2025, 04:01 PM | ••• | |
clickstream_cdc_events_p | Expire Snapshots | Apr 25, 2026, 12:03 AM | Apr 25, 2026, 03:05 AM | Jan 28, 2025, 03:25 PM | ••• | |
sessions_cdc_events_p | Expire Snapshots | Apr 25, 2026, 12:03 AM | Apr 26, 2026, 03:05 AM | Jun 26, 2025, 11:11 PM | ••• | |
global_expire_snapshots Runs snapshot expiration on all tables once a day | Expire Snapshots | Apr 26, 2026, 1:18 PM | Apr 07, 2026, 01:08 PM | Mar 14, 2026, 8:42 AM | ••• | |
manifest_rewrite_weekly Rewrite manifests for all critical tables weekly | Rewrite Manifests | Apr 28, 2026, 2:00 AM | Apr 21, 2026, 02:00 AM | Mar 10, 2026, 9:15 AM | ••• | |
staging_config | Configuration | — | — | Dec 31, 2025, 02:45 PM | ••• |
Compaction thresholds, retention windows, cleanup schedules, sort strategies — set them once as declarative policies. LakeOps enforces them across every catalog and table, continuously. Policies are auditable, versioned, and toggled with one switch. New tables inherit them automatically.
Compare engines side-by-side on cost, latency, throughput, and data scanned.
| Metric | Spark | Trino | Athena | Snowflake |
|---|---|---|---|---|
| Query success rate | 99.2% | 99.5% | 99.9% | 99.8% |
| Average runtime | 3.1s | 1.8s | 2.3s | 2.1s |
| Cost per query | $0.04 | $0.03 | $0.05 | $0.08 |
| Total queries | 3,120 | 2,456 | 1,280 | 1,876 |
| Data scanned | 4.2 TB | 2.8 TB | 1.5 TB | 3.5 TB |
Lower-left is ideal
Your applications connect to a single SQL endpoint. LakeOps routes each query to Trino, Spark, Snowflake, Athena, DuckDB, or Flink — picking the best engine by cost, latency, or workload type. Add or remove engines without touching application code.
list_catalogsreadEnumerate all registered catalogsget_table_healthreadHealth score + maintenance statusrun_queryexecuteRouted through guardrails pipelineanalyze_storage_reclaimreadSnapshot bloat + orphan insightsget_optimization_statusreadCompaction + manifest stateReadOnly
Blocks DDL and DML from agent sessions
CostEstimate
Rejects queries exceeding scan thresholds
PIIMask
Hashes sensitive columns before results reach the model
HumanApproval
Pauses high-stakes operations for review
AI agents connect via MCP with standard Postgres, MySQL, or Arrow Flight protocols — no SDK needed. Automate pipelines and CI/CD via the REST API. LakeOps routes agent queries through the same engine pool, enforces guardrails (read-only mode, cost caps, PII masking, human approval), and feeds access patterns back into compaction decisions.
Results
Benchmarks from production-grade tables across multiple engines and clouds.
Query speed
After compaction + layout optimization
CPU reduction
Compute hours across all engines
Storage saved
Orphans, snapshots & bloat removed
Table health
Autonomous maintenance keeps every table optimized
See It In Action
From catalog connection to autonomous maintenance, compaction, routing, and governance — one walkthrough.
LakeOps| Operation | Table | Duration | Impact | Time | Status |
|---|---|---|---|---|---|
| Compact Data Files | customer_orders orders | 4s | 1.24 TB, 16 → 1 files | 57 minutes ago | SUCCESS |
| Expire Snapshots | payment_transactions payments | 27s | 8.2 TB | 4 hours ago | SUCCESS |
| Expire Snapshots | inventory_snapshots_20250702 warehouse | 3s | 2.1 TB | 4 hours ago | SUCCESS |
| Rewrite Manifests | raw_clickstream analytics | 1.9s | 3 → 1 manifests | 5 hours ago | SUCCESS |
| Compact Data Files | product_catalog products | 6m 11.3s | 3,008 → 1,256 files | 6 hours ago | SUCCESS |
Only metadata is processed — never retained or stored.
Telemetry reveals table health and actions needed.
Autopilot, manual approval, or policy-driven.
Resources
Get started
Get a personalized walkthrough with your data and architecture.
Short call, no commitment.
Typically 30 min · Free