Migration

Migrate to Iceberg without the operational gap

Snowflake and Databricks handle maintenance invisibly. Open Iceberg does not. LakeOps fills that gap — compaction, cleanup, and governance run on both your old and new stack while you migrate workloads at your own pace.

Run on both stacks at once

Snowflake + Iceberg in one control plane

No ops gap during migration

Maintenance runs from the moment you connect

Hit the ground running

Full operations already at scale — zero ramp-up after

LakeOps LogoLakeOps

Last 30 days Optimization Activity

Total Operations
12,211
Last 90 days
Query Speed
12.4×
Avg. acceleration across engines
Cost Savings
$1,374,672
Saved in last 3 months
CPU & Storage
-76%
Last 90 days
Data Optimized
46.8 PB
Last 30 days

Key Metrics

Total Tables
786
Tables in all catalogs
Critical Tables
70
Require immediate attention
Warning Tables
105
Should be addressed or auto-piloted
Healthy Tables
566
Tables in optimal state
Total Data
112.4 PB
Total lake data size

Storage

-56% reclaimed
30d
112 TB75 TB38 TB
Storage used

CPU

-76% reduction
30d
100%66%33%
Compute hours

Recent Operations

Last 10 operations
OperationTableDurationImpactTimeStatus
Compact Data Files
customer_orders
orders
4s1.24 TB, 16 → 1 files57 minutes agoSUCCESS
Expire Snapshots
payment_transactions
payments
27s8.2 TB4 hours agoSUCCESS
Expire Snapshots
inventory_snapshots_20250702
warehouse
3s2.1 TB4 hours agoSUCCESS
Rewrite Manifests
raw_clickstream
analytics
1.9s3 → 1 manifests5 hours agoSUCCESS
Compact Data Files
product_catalog
products
6m 11.3s3,008 → 1,256 files6 hours agoSUCCESS

Lake Events

LiveLast 24 hours
Compact Data Files·customer_orders
ecommerce_prod·1.24 TB, 16 → 1 files
4s57 min agoOK
Expire Snapshots·payment_transactions
ecommerce_prod·12 snapshots expired
4.6s1h agoOK
Compact Data Files·raw_clickstream
marketing_events·970 → 87 files
9m 31.6s2h agoOK
Remove Orphan Files·user_sessions
marketing_events·847 MB reclaimed, 1,203 files
1m 12s3h agoOK
Rewrite Manifests·search_query_logs
ecommerce_prod·487 → 12 manifests
2.1s3h agoOK
Expire Snapshots·inventory_levels
warehouse_analytics·62 snapshots, 18.4 GB freed
27s4h agoOK
Compact Data Files·product_catalog
ecommerce_prod·3,008 → 1,256 files
6m 11.3s5h agoOK
Rewrite Manifests·shipping_events
warehouse_analytics·14 → 3 manifests
1.0s6h agoOK
Remove Orphan Files·balance_snapshots
warehouse_analytics·59,831 files, 74.8 GB
13m 6.9s7h agoOK
Compact Data Files·ad_impressions
marketing_events·42,633 → 69 files
2m 18s8h agoOK
No vendor lock-in
No code / infra changes
No data changes
REST API & MCP

Open and runs on your stack

AWS
Azure
Google Cloud
Snowflake
Databricks
Apache Flink
Apache Iceberg
Delta Lake
DuckDB
Dremio
Lakekeeper
ClickHouse
AWS
Azure
Google Cloud
Snowflake
Databricks
Apache Flink
Apache Iceberg
Delta Lake
DuckDB
Dremio
Lakekeeper
ClickHouse

The migration gap

Migrating the data is the easy half. Operating it is where teams get stuck.

Snowflake and Databricks handle compaction, expiry, and cleanup invisibly. Move to open Iceberg and that responsibility lands on your platform team — unless someone else picks it up.

Nobody maintains the new tables

Snowflake and Databricks auto-compact behind the scenes. Open Iceberg does not. The moment tables land in your new catalog, maintenance becomes your team’s responsibility.

Two stacks, zero unified visibility

During migration you run both the old platform and the new lake. Health, cost, and query performance split across two dashboards that don’t talk to each other.

Query regressions block the next batch

Unmaintained Iceberg tables degrade within weeks — small files pile up, manifests bloat, planning slows. Analysts blame the migration, and the next batch of workloads stalls.

The ops gap kills confidence

Leadership approved the migration for cost savings. But three months in, the platform team is firefighting table health instead of delivering features. Trust erodes.

How it works

Connect both catalogs. Migrate workloads at your own pace.

LakeOps manages table health across your old platform and your new Iceberg lake simultaneously. Move one workload at a time — tables stay healthy on both sides.

01 · Connect both catalogs

Unified visibility across your old and new stack

Connect your Snowflake, Glue, Polaris, REST, or S3 Tables catalogs — all of them, at once. LakeOps discovers every table across both your legacy platform and your new Iceberg lake, scores health uniformly, and gives you one dashboard for the entire migration.

  • Multi-catalog, multi-cloud — connect Snowflake and Glue in the same view
  • Table health scored from the first metadata read — no agents, no pipeline changes
  • 10-minute setup per catalog — not a project
Table Health Overview786 tables
566
Healthy
105
Warning
70
Critical
92%
Optimized
TableNSSizeStatus
customer_ordersorders1.24 TBHEALTHY
payment_transactionspayments860 GBWARNING
raw_clickstreamanalytics4.6 TBCRITICAL
product_catalogproducts42 GBHEALTHY
user_sessionsanalytics1.9 TBWARNING
inventory_levelsoperations320 GBHEALTHY
shipping_eventslogistics580 GBHEALTHY
search_query_logsanalytics3.2 TBCRITICAL

02 · Maintenance from day one

Tables are compacted and cleaned from the moment they land

On Snowflake or Databricks, compaction runs behind the scenes. Open Iceberg has no built-in maintenance. LakeOps fills that gap immediately: snapshot expiry, compaction, orphan cleanup, and manifest rewrite start running the moment your catalog connects — sequenced in the right order, adapted to each table’s write pattern.

  • No maintenance gap — Iceberg tables never sit unmaintained during migration
  • Health-driven cadence — high-write streaming tables run hourly, batch tables daily
  • 95% faster compaction on a Rust engine — no Spark cluster to provision
Adaptive Maintenance
Running
Compaction Scope— what will be targeted this run

Monitoring started at seq #4,812. Everything written from that point on is in scope.

already compacted
38% hot zone — will compact
#4,812 baseline#7,104 watermark#8,506 now
📦Compaction
Last: 57 minutes agoNext: in 2h
62%
42% of data files are below the 384 MB size threshold
1.24 TB compactable — will merge into ~3 files at 512 MB target
16 delete files amplify read cost — compaction will absorb them
Snapshot Expiration
Last: 4 hours agoNext: in 20h
31%
12 of 83 snapshots have passed the retention window
Retention policy: 5.0 days
Table creates 3.2 snapshots/hr — stale ones accumulate quickly
📝Rewrite Manifests
Last: 5 hours agoNext: in 1h
78%
92 total manifests: 89 data + 3 delete
Accumulating 2.1 manifests/hr — metadata overhead grows
High manifest count degrades query planning and scan performance
🧹Orphan File Cleanup
Last: 7 hours agoNext: in 17h
15%
No orphan files detected in current scan window

03 · Route queries across both stacks

Migrate workloads gradually — not all at once

One SQL endpoint routes each query to the right engine — Trino, Spark, Snowflake, DuckDB, Athena. Move workloads one at a time: start with batch ETL, then analytics, then dashboards. If a workload isn’t ready, it stays on the old engine. No big-bang cutover.

  • Gradual migration — route by workload type, team, or latency target
  • Automatic fallback if the new engine underperforms on a query shape
  • Cost comparison across engines — prove savings before you commit

Engine comparison

Compare engines side-by-side on cost, latency, throughput, and data scanned.

Select engines

Spark
Active
Trino
Active
Athena
Active
Snowflake
Active
DuckDB
Active
Flink
Active

Performance comparison

MetricSparkTrinoAthenaSnowflake
Query success rate99.2%99.5%99.9%99.8%
Average runtime3.1s1.8s2.3s2.1s
Cost per query$0.04$0.03$0.05$0.08
Total queries3,1202,4561,2801,876
Data scanned4.2 TB2.8 TB1.5 TB3.5 TB

Cost vs latency

Lower-left is ideal

Low cost
High cost

Success rate

Spark
99.2%
Trino
99.5%
Athena
99.9%
Snowflake
99.8%

04 · Policies apply automatically

Governance rules carry over to every new table

Declare compaction targets, retention windows, and cleanup schedules at the catalog or namespace level. As you migrate tables into the new catalog, they inherit the policies automatically. No manual setup per table, no governance gaps.

  • New tables inherit policies from their namespace — zero manual work
  • Same rules across Glue, Polaris, Nessie, and Gravitino
  • Full audit trail — every operation logged for compliance evidence

Policies

Manage maintenance, configuration, and lifecycle policies for your data lakehouse

OnPolicyTypeNext RunLast RunUpdatedActions
orders_critical
Compact FilesApr 25, 2026, 8:12 AMApr 25, 2026, 03:05 AMFeb 01, 2025, 3:46 PM•••
payments_compact
Compact FilesFeb 15, 2026, 12:06 AMFeb 1, 2025, 02:18 PMFeb 5, 2025, 4:03 PM•••
Remove orphan files (e-ip...)
For all tables in all catalogs every 7 days
Orphan FilesApr 25, 2026, 8:12 AMApr 25, 2026, 04:07 PMJun 23, 2025, 04:01 PM•••
clickstream_cdc_events_p
Expire SnapshotsApr 25, 2026, 12:03 AMApr 25, 2026, 03:05 AMJan 28, 2025, 03:25 PM•••
sessions_cdc_events_p
Expire SnapshotsApr 25, 2026, 12:03 AMApr 26, 2026, 03:05 AMJun 26, 2025, 11:11 PM•••
global_expire_snapshots
Runs snapshot expiration on all tables once a day
Expire SnapshotsApr 26, 2026, 1:18 PMApr 07, 2026, 01:08 PMMar 14, 2026, 8:42 AM•••
manifest_rewrite_weekly
Rewrite manifests for all critical tables weekly
Rewrite ManifestsApr 28, 2026, 2:00 AMApr 21, 2026, 02:00 AMMar 10, 2026, 9:15 AM•••
staging_config
ConfigurationDec 31, 2025, 02:45 PM•••

05 · Smooth landing

Operations are already running when migration ends

Because LakeOps managed both stacks throughout the migration, your new Iceberg lake arrives at full operational maturity. Compaction cadences are tuned, governance policies are enforced, health baselines are established. No ramp-up period — the platform team ships features instead of fighting table health.

  • Zero ramp-up — maintenance, governance, and routing already running at scale
  • Proven baselines — health scores, cost trends, and performance metrics from day one
  • No new overhead — the same control plane you used during migration runs everything after
Proprietary → Open IcebergMigration

Before — Proprietary

Snowflake auto-compacts
Maintenance is invisible
One engine, one dashboard
Vendor owns the operations
You pay the premium

After — Open Iceberg + LakeOps

Connect Iceberg catalogs
LakeOps maintains every table
Route across any engine
You own the data and ops
No vendor lock-in

From teams who migrated

“We ran LakeOps throughout the migration”

We ran LakeOps throughout our migration off Snowflake—it worked with both stacks at the same time. That single control plane took away almost all the overhead of managing our own Iceberg lake, so we could focus on the move instead of maintenance.

Itay B.

Itay B.

Data Platform Lead

LakeOps works with our Snowflake and Databricks stack—no vendor lock-in. We cut weeks of manual maintenance down to a few hours. One control plane across engines, and the team finally stopped firefighting small files and compaction.

Egor S.

Egor S.

Senior Data Engineer

Results

Cut costs and boost performance

Benchmarks from production-grade tables across multiple engines and clouds.

Query speed

12×faster

After compaction + layout optimization

CPU reduction

76%less compute

Compute hours across all engines

Storage saved

56%reclaimed

Orphans, snapshots & bloat removed

Table health

100%healthy

Autonomous maintenance keeps every table optimized

TPC-DS benchmark suiteProduction Iceberg tablesMulti-cloud, multi-engine
Watch demoYouTube ↗

Connect in minutes
- no vendor lock-in

1

Connect catalogs & engines

Only metadata is processed — never retained or stored.

Apache Iceberg
AWS
Snowflake
DuckDB
2

Get visibility & insights

Telemetry reveals table health and actions needed.

Table health scores
Optimization opportunities
Cost & performance insights
3

Choose your mode

Autopilot, manual approval, or policy-driven.

Autopilot
Manual
Policies
4

Lakehouse optimized

Queries 10x faster
Cost down 76%
Engines optimized
AIs managed
Tables healthy
Fully governed
Set up in 10 minutes · Works with your existing stack
Enterprise-grade

Built for enterprisegrade data lakes

SOC 2, SSO, RBAC, dedicated support, and the scale your largest Iceberg lakes demand.

Security & compliance

SOC 2 Type II, encryption, SSO/RBAC, and audit trails for regulated teams.

Scale & control

One control plane for your full lake. Real-time visibility, policies, and predictable performance.

Support & training

Dedicated onboarding, training, and enterprise SLAs. Deploy in VPC or on-prem.

Get started

See LakeOps on your stack

Get a personalized walkthrough with your data and architecture.Short call, no commitment.

Typically 30 min · Free