Platform

Intelligent Iceberg Lakehouse Control Plane

Autonomously optimizes Iceberg table health, query speed, and cost — across your entire lakehouse.

Sense

Monitors file layout, metadata health, and query patterns.

Plan

Prioritizes by severity, sequences operations, picks optimal engines.

Optimize

Compaction, cleanup, routing, and governance on a Rust engine.

Learn

Sort orders adapt, routing improves, cadence tunes every run.

Open and runs on your stack

AWS
Azure
Google Cloud
Snowflake
Databricks
Apache Flink
Apache Iceberg
Delta Lake
DuckDB
Dremio
Lakekeeper
ClickHouse
AWS
Azure
Google Cloud
Snowflake
Databricks
Apache Flink
Apache Iceberg
Delta Lake
DuckDB
Dremio
Lakekeeper
ClickHouse
1Observe — see what's degrading your lake
Table Health Overview786 tables
566
Healthy
105
Warning
70
Critical
92%
Optimized
TableNSSizeStatus
customer_ordersorders1.24 TBHEALTHY
payment_transactionspayments860 GBWARNING
raw_clickstreamanalytics4.6 TBCRITICAL
product_catalogproducts42 GBHEALTHY
user_sessionsanalytics1.9 TBWARNING
inventory_levelsoperations320 GBHEALTHY
shipping_eventslogistics580 GBHEALTHY
search_query_logsanalytics3.2 TBCRITICAL

Real-time health detection across every Iceberg table

Every table scored as Healthy, Warning, or Critical — based on file count, small-file ratio, snapshot age, manifest bloat, and write velocity. LakeOps surfaces degradation before it reaches your queries, with unified telemetry across every engine in your stack.

  • Per-table health scores that flag degradation before it hits query performance
  • Unified telemetry across Trino, Spark, Snowflake, Athena, DuckDB, and Flink
  • Full operation history with duration, files affected, and bytes reclaimed
2Maintain — keep every table healthy, automatically
customer_ordersMaintenance Health
Autonomous
Compaction Scope— what will be targeted this run
already compacted
38% hot zone
#4,812 baseline#7,104 watermark#8,506 now

Coordinated Operations

execution order →
In Progress

Compaction

72%

38% small files — merging 970 → 87 at 512 MB target

then →Expire Snapshots
Cooling

Expire Snapshots

45%

154 snapshots, 62 past 30-day retention

then →Rewrite Manifests
Idle

Rewrite Manifests

18%

12 manifests — below threshold, waiting for compaction

then →Orphan Cleanup
Idle

Orphan Cleanup

8%

847 MB unreferenced — scheduled after expiration

Learning from telemetry

Query patterns

event_date, region

Top sort columns (Trino + Spark)

Improvement

12.4× faster

Avg query speed after optimization

Cycle

Self-tuning

Sort orders adapt as patterns change

Autonomous Iceberg maintenance triggered by health, not schedules

No cron jobs. No Airflow DAGs. No runbooks. LakeOps scores each table's health, decides which operations to run, and sequences them in dependency order — expiry before compaction, compaction before cleanup. Each step's output is the next step's clean input.

  • Health-driven triggers — runs when tables need it, not on fixed schedules
  • Dependency-ordered: compaction → snapshot expiry → orphan cleanup → manifest rewrite
  • Learns from outcomes — each execution refines future decisions

Events

View and track table operations and history

Table NameNamespaceCatalogOperation TypeImpactDurationStart TimeStatus
customer_ordersordersecommerce_prodCompact Data Files1.24 TB, 16 → 1 files4s57 min agoSUCCESS
customer_ordersordersecommerce_prodCompact Data Files11 → 8 files10.3s2h agoSUCCESS
customer_ordersordersecommerce_prodCompact Data Files970 → 87 files6m 1.0s3h agoSUCCESS
payment_transactionspaymentsecommerce_prodExpire Snapshots1 snapshot4.6s4h agoSUCCESS
payment_transactionspaymentsecommerce_prodRewrite Manifests3 → 1 manifests1.9s4h agoSUCCESS
raw_clickstreamanalyticsmarketing_eventsCompact Data Files970 → 87 files9m 31.6s5h agoSUCCESS
user_sessionsanalyticsmarketing_eventsRemove Orphan Files1,203 files, 847 MB1m 12s5h agoSUCCESS
product_catalogproductsecommerce_prodCompact Data Files3,008 → 1,256 files6m 11.3s6h agoSUCCESS
user_sessionsanalyticsmarketing_eventsRewrite Manifests2 → 1 manifests1.8s7h agoSUCCESS
inventory_levelsoperationswarehouse_analyticsRemove Orphan Files4,218 files, 1.9 GB3m 44s7h agoSUCCESS
inventory_levelsoperationswarehouse_analyticsExpire Snapshots12 snapshots expired2.0s8h agoSUCCESS
shipping_eventslogisticswarehouse_analyticsRewrite Manifests3 → 1 manifests1.0s8h agoSUCCESS

Lake-wide event trail — every Iceberg table, every operation

Every compaction, snapshot expiration, orphan cleanup, and manifest rewrite is logged as an event — across all catalogs and namespaces. Filter by table, operation type, or status to trace exactly what happened, when, and what it reclaimed. The same feed powers governance audit trails and compliance evidence for SOC 2 and GDPR.

  • Cross-catalog visibility — events from every engine and catalog in one timeline
  • Per-operation impact: files merged, snapshots expired, bytes reclaimed, duration
  • Audit-ready — ties directly into governance policies and compliance reporting
Operations MonitoringAll engines connected

Operations Timeline (7d)

1,892 total
Mon
Tue
Wed
Thu
Fri
Sat
Sun
Compaction
Expire
Rewrite
Orphans

Proactive insights

raw_clickstreamCritical

312 partitions exceed file threshold

Query scan amplified 8×

search_query_logsHigh

Excessive manifests (487) — planning overhead

Planner latency +2.1s

payment_transactionsWarning

Small file ratio 38% — compaction recommended

S3 GET costs elevated

1.8 TB

Space freed

4.2 TB

Optimized

247

Tables

2%

Failure rate

From 50 tables to 50,000 — zero additional work

Register your catalogs and every table gets maintained continuously. Streaming tables hourly, batch tables daily, idle tables skipped — cadence adapts to write velocity. The full metadata pipeline runs lake-wide without a single additional ops task.

  • Scales from 50 to 5,000+ tables — every one scored, maintained, and verified
  • Cadence adapts to write velocity: streaming hourly, batch daily, idle skipped
  • Full metadata pipeline: snapshots → orphans → manifests → deletes → statistics
3Compact — sort files for how you query
Compaction EngineQuery-Aware
Rust + DataFusion
Detected Access Patterns— 2,847 queries · 3 engines
WHEREcustomer_id
89%
WHEREevent_date
76%
JOINproduct_id
64%
GROUPregion
51%
Sort Optimization Applied

Before

Files970
Avg size3.2 MB
Sortappend order

After

Files87
Avg size256 MB
Sortcustomer_id, date

12×

Faster queries

76%

Less CPU

95%

vs Spark

90%

Cheaper ops

200 GB compaction · same hardware

S3 Tables
6,300s
Spark
1,612s
LakeOps
221s

Iceberg compaction that sorts for how you query

Most compaction tools merge small files blindly. LakeOps watches which columns your queries actually filter, join, and group on — then physically re-sorts data to match. Engines skip entire file groups via min/max pruning, cutting I/O by up to 90%. Built on Rust and Apache DataFusion: 95% faster and 90% cheaper than Spark-based compaction.

  • 12× faster queries — sort order matches real access patterns, not guesswork
  • 76% less CPU — fewer files scanned, zero Spark clusters required
  • Self-improving — sort strategy adapts as query patterns evolve
Layout Simulations

Run simulations on customer_orders

Simulations 3metalayout
✓clusterByOrderDate
Partition: order_date → daily buckets. Cluster by customer_id, order_status. Sort: merchant_id, discount_pct
customer_id, order_status
141.6s
✓cluster.order_type.by.status
Partition: order_status = 'completed'. Rebuild partition keys: payment_method, currency.
order_status, payment_method
120.4s
✓cluster.insert-time-line
Partition: created_at → hourly. Sort by customer_id, store_id. File-size balanced partitions.
customer_id, store_id
89.2s
▸ Field access frequency by query mix
How this table's fields correlate — the foundation for choosing the right layout strategy.
order_id
customer_id
order_status
total_amount
item_ids
created_at
payment_method
discount_pct
02,000,0004,000,0006,000,0008,000,000
SELECTFILTERJOIN Rows
↕ Layout Customization Diff
Analyze layout changes across all selected simulations → highlight what sets differ from baseline.
Simulationdata_relStrategyAvg Size
metaclusterByOrderDatecustomer_id, order_statusorder_date (day)343 MB / file
metacluster.order_type.by.statusorder_status, payment_methodorder_status, payment_method511 MB / file
layoutcluster.insert-time-linecustomer_id, store_idcreated_at (hour)128 MB / file

Test layout changes on branches before production

LakeOps tests layout changes on Iceberg branches before they ever touch production — sort orders, partition strategies, and file targets against real query patterns. Compare scan reduction, file count, and estimated speedup side by side, then apply the winning change with one action.

  • Every candidate change runs on a branch — production tables stay untouched
  • Side-by-side results: scan reduction, file count, and query speedup
  • Promote the winning layout when you're ready — one action, no rewrite risk
4Govern — enforce rules across your entire lake

Policies

Manage maintenance, configuration, and lifecycle policies for your data lakehouse

OnPolicyTypeNext RunLast RunUpdatedActions
orders_critical
Compact FilesApr 25, 2026, 8:12 AMApr 25, 2026, 03:05 AMFeb 01, 2025, 3:46 PM•••
payments_compact
Compact FilesFeb 15, 2026, 12:06 AMFeb 1, 2025, 02:18 PMFeb 5, 2025, 4:03 PM•••
Remove orphan files (e-ip...)
For all tables in all catalogs every 7 days
Orphan FilesApr 25, 2026, 8:12 AMApr 25, 2026, 04:07 PMJun 23, 2025, 04:01 PM•••
clickstream_cdc_events_p
Expire SnapshotsApr 25, 2026, 12:03 AMApr 25, 2026, 03:05 AMJan 28, 2025, 03:25 PM•••
sessions_cdc_events_p
Expire SnapshotsApr 25, 2026, 12:03 AMApr 26, 2026, 03:05 AMJun 26, 2025, 11:11 PM•••
global_expire_snapshots
Runs snapshot expiration on all tables once a day
Expire SnapshotsApr 26, 2026, 1:18 PMApr 07, 2026, 01:08 PMMar 14, 2026, 8:42 AM•••
manifest_rewrite_weekly
Rewrite manifests for all critical tables weekly
Rewrite ManifestsApr 28, 2026, 2:00 AMApr 21, 2026, 02:00 AMMar 10, 2026, 9:15 AM•••
staging_config
Configuration——Dec 31, 2025, 02:45 PM•••

Declarative policies that enforce themselves

Compaction thresholds, retention windows, cleanup schedules, sort strategies — set them once as declarative policies. LakeOps enforces them across every catalog and table, continuously. Policies are auditable, versioned, and toggled with one switch. New tables inherit them automatically.

  • One-toggle policies for compaction, retention, orphan cleanup, and manifests
  • Scope inheritance — new tables comply from day one
  • Full audit trail: every execution logged with duration, impact, and outcome
5Route — send every query to the right engine

Engine comparison

Compare engines side-by-side on cost, latency, throughput, and data scanned.

Select engines

Spark
Active
Trino
Active
Athena
Active
Snowflake
Active
DuckDB
Active
Flink
Active

Performance comparison

MetricSparkTrinoAthenaSnowflake
Query success rate99.2%99.5%99.9%99.8%
Average runtime3.1s1.8s2.3s2.1s
Cost per query$0.04$0.03$0.05$0.08
Total queries3,1202,4561,2801,876
Data scanned4.2 TB2.8 TB1.5 TB3.5 TB

Cost vs latency

Lower-left is ideal

Low cost
High cost

Success rate

Spark
99.2%
Trino
99.5%
Athena
99.9%
Snowflake
99.8%

One SQL endpoint — queries route to the optimal engine

Your applications connect to a single SQL endpoint. LakeOps routes each query to Trino, Spark, Snowflake, Athena, DuckDB, or Flink — picking the best engine by cost, latency, or workload type. Add or remove engines without touching application code.

  • Route by cost, latency, or workload type — configured per routing group
  • Side-by-side engine benchmarks on identical queries
  • Swap engines behind a stable endpoint — zero application changes
6AI Agents — make your lake agent-ready
MCP InterfaceAgent-native
Connected

Wire compatibility

PostgreSQLMySQLArrow Flight
psql -h agent.lakeops.dev -U ai_agent -d ecommerce_prod

Available MCP tools

list_catalogsreadEnumerate all registered catalogs
get_table_healthreadHealth score + maintenance status
run_queryexecuteRouted through guardrails pipeline
analyze_storage_reclaimreadSnapshot bloat + orphan insights
get_optimization_statusreadCompaction + manifest state

Layered guardrails

ReadOnly

Blocks DDL and DML from agent sessions

CostEstimate

Rejects queries exceeding scan thresholds

PIIMask

Hashes sensitive columns before results reach the model

HumanApproval

Pauses high-stakes operations for review

Agent query telemetry feeds back into compaction and sort-order decisions

AI-ready lakehouse — agents query with guardrails

AI agents connect via MCP with standard Postgres, MySQL, or Arrow Flight protocols — no SDK needed. Automate pipelines and CI/CD via the REST API. LakeOps routes agent queries through the same engine pool, enforces guardrails (read-only mode, cost caps, PII masking, human approval), and feeds access patterns back into compaction decisions.

  • Standard wire protocols — agents connect with any Postgres or MySQL driver
  • Layered guardrails: ReadOnly, CostEstimate, PIIMask, HumanApproval per session
  • Closed-loop: agent query patterns inform sort order and compaction priorities
  • REST API for automation, MCP server for AI agents — full docs for both

Results

Cut costs and boost performance

Benchmarks from production-grade tables across multiple engines and clouds.

Query speed

12×faster

After compaction + layout optimization

CPU reduction

76%less compute

Compute hours across all engines

Storage saved

56%reclaimed

Orphans, snapshots & bloat removed

Table health

100%healthy

Autonomous maintenance keeps every table optimized

TPC-DS benchmark suiteProduction Iceberg tablesMulti-cloud, multi-engine

See It In Action

Watch how the platform works end to end

From catalog connection to autonomous maintenance, compaction, routing, and governance — one walkthrough.

Watch demoYouTube ↗
LakeOps platform — end-to-end walkthrough

Explore the dashboard

LakeOps LogoLakeOps

Last 30 days Optimization Activity

Total Operations
12,211
Last 90 days
Query Speed
12.4×
Avg. acceleration across engines
Cost Savings
$1,374,672
Saved in last 3 months
CPU & Storage
-76%
Last 90 days
Data Optimized
46.8 PB
Last 30 days

Key Metrics

Total Tables
786
Tables in all catalogs
Critical Tables
70
Require immediate attention
Warning Tables
105
Should be addressed or auto-piloted
Healthy Tables
566
Tables in optimal state
Total Data
112.4 PB
Total lake data size

Storage

-56% reclaimed
30d
112 TB75 TB38 TB
Storage used

CPU

-76% reduction
30d
100%66%33%
Compute hours

Recent Operations

Last 10 operations
OperationTableDurationImpactTimeStatus
Compact Data Files
customer_orders
orders
4s1.24 TB, 16 → 1 files57 minutes agoSUCCESS
Expire Snapshots
payment_transactions
payments
27s8.2 TB4 hours agoSUCCESS
Expire Snapshots
inventory_snapshots_20250702
warehouse
3s2.1 TB4 hours agoSUCCESS
Rewrite Manifests
raw_clickstream
analytics
1.9s3 → 1 manifests5 hours agoSUCCESS
Compact Data Files
product_catalog
products
6m 11.3s3,008 → 1,256 files6 hours agoSUCCESS

Lake Events

LiveLast 24 hours
Compact Data Files·customer_orders
ecommerce_prod·1.24 TB, 16 → 1 files
4s57 min agoOK
Expire Snapshots·payment_transactions
ecommerce_prod·12 snapshots expired
4.6s1h agoOK
Compact Data Files·raw_clickstream
marketing_events·970 → 87 files
9m 31.6s2h agoOK
Remove Orphan Files·user_sessions
marketing_events·847 MB reclaimed, 1,203 files
1m 12s3h agoOK
Rewrite Manifests·search_query_logs
ecommerce_prod·487 → 12 manifests
2.1s3h agoOK
Expire Snapshots·inventory_levels
warehouse_analytics·62 snapshots, 18.4 GB freed
27s4h agoOK
Compact Data Files·product_catalog
ecommerce_prod·3,008 → 1,256 files
6m 11.3s5h agoOK
Rewrite Manifests·shipping_events
warehouse_analytics·14 → 3 manifests
1.0s6h agoOK
Remove Orphan Files·balance_snapshots
warehouse_analytics·59,831 files, 74.8 GB
13m 6.9s7h agoOK
Compact Data Files·ad_impressions
marketing_events·42,633 → 69 files
2m 18s8h agoOK
No vendor lock-in
No code / infra changes
No data changes
REST API & MCP

Connect in minutes
- no vendor lock-in

1

Connect catalogs & engines

Only metadata is processed — never retained or stored.

Apache Iceberg
AWS
Snowflake
DuckDB
2

Get visibility & insights

Telemetry reveals table health and actions needed.

Table health scores
Optimization opportunities
Cost & performance insights
3

Choose your mode

Autopilot, manual approval, or policy-driven.

Autopilot
Manual
Policies
4

Lakehouse optimized

Queries 10x faster
Cost down 76%
Engines optimized
AIs managed
Tables healthy
Fully governed
Set up in 10 minutes · Works with your existing stack

Resources

Learn more

Explore all articles

Get started

See LakeOps on your stack

Get a personalized walkthrough with your data and architecture.Short call, no commitment.

Typically 30 min · Free