Agentic AI

Let agents query Iceberg with guardrails and the right engine

Connect over MCP — no custom SDK. Every query hits read-only, cost, and PII guards, then routes to the cheapest engine that fits.

Open and runs on your stack

AWS
Azure
Google Cloud
Snowflake
Databricks
Apache Flink
Apache Iceberg
Delta Lake
DuckDB
Dremio
Lakekeeper
ClickHouse
AWS
Azure
Google Cloud
Snowflake
Databricks
Apache Flink
Apache Iceberg
Delta Lake
DuckDB
Dremio
Lakekeeper
ClickHouse

The problem

Agents timeout on small files and scan the wrong engine

Uncompacted tables, no cost ceiling, and batch clusters for interactive SQL — that is what most agents hit today.

Agents hit slow, stale tables

Agents expect sub-second SQL. Small files and bloated manifests force full scans — a simple lookup becomes 10× slower.

No guardrails on autonomous SQL

Unsupervised agents can run DDL, scan petabytes, or leak PII. Every query is a cost and compliance risk.

Wrong engine, wrong cost

An interactive agent lookup on a batch cluster wastes minutes. Without routing, agents cannot pick the cheapest path.

No feedback into table layout

Agent workloads change constantly. Tables stay static. The lake never sorts for the columns agents actually filter on.

How it works

From MCP connect to tables that stay fast for agents

One endpoint for any MCP agent. Every query is guarded and routed. Access patterns drive compaction, snapshots, orphan cleanup, and manifest rewrites.

01 · MCP

Agents connect over MCP — no custom SDK

Claude, LangChain, or a custom agent. Schema-aware tools, async SQL with SSE, and Postgres, MySQL, or Arrow Flight wire compatibility. Point the agent at one endpoint.

  • Any MCP-compatible agent — zero integration code
  • Standard drivers: Postgres, MySQL, Arrow Flight
  • Schema-aware tools are discovered automatically
MCP endpointNo SDK
Claude
LangChain
Custom

MCP · Postgres · MySQL · Flight

lakeops MCP server

Schema tools · async SQL · SSE

Trino
DuckDB
Spark

02 · Guardrails

Layered safety before a query reaches an engine

Stack guards per agent, team, or globally. Read-only blocks DDL. Cost caps reject expensive scans. PII masking scrubs sensitive columns. Human approval pauses high-stakes SQL.

  • ReadOnly, CostEstimate, PIIMask, HumanApproval — stackable
  • Scope per session, team, or the entire lake
  • Every fired guard is logged with the query
Session guardrailsStacked

ReadOnly

Blocks DDL and DML

On

CostEstimate

Rejects scans over 10 GB

On

PIIMask

Hashes email, ssn, phone

On

HumanApproval

Pauses writes for review

Off

03 · Routing

Send agent SQL to the cheapest engine that fits

Interactive lookups go to DuckDB or Trino. Heavy scans go elsewhere. Cached decisions are 0ms. Unhealthy tables get steered off expensive engines.

  • Dedicated endpoints so agent traffic does not sit behind ETL
  • Cost vs latency strategy per routing group
  • Cached templates skip the planner — 0ms

Routing groups

Map workloads to engine pools, publish stable URLs, and tune priority without touching clients.

Groups
4
Active
3
Accepting traffic
Paused
1
Inactive groups
Engines
7
Unique in routes
All groups

Edit engines, query types, and published URLs

Analyticsactive
📡 e1fa3c3c.lakeops.dev
Engines: TrinoDuckDB
Query types: SELECTAGGREGATE
Priority: High
📅 06/02/25, 3:20:16 PM
•••
BIactive

Handles BI, DT queries on Transactional databases

📡 1d0e4f1c6.lakeops.dev
Engines: SnowflakeTrino
Query types: INSERTUPDATEDELETE
Priority: Medium
📅 06/02/25, 8:16:52 PM
•••
Data-Team ETLinactive

Batch data transformation and scheduled loads

📡 e11d1ef1.lakeops.dev
Engines: SparkFlink
Query types: BATCHSTREAM
Priority: Medium
📅 06/15/25, 5:40:10 PM
•••
Reportsactive

Routine reporting and dashboard queries

📡 c480dr51.lakeops.dev
Engines: SnowflakeClickHouse
Query types: SELECTVIEW
Priority: Medium
📅 06/02/25, 7:00:38 PM
•••

04 · Closed loop

Agent traffic shapes the entire maintenance cycle

Agent SQL feeds back into every optimization — sort-order decisions, snapshot retention, orphan cleanup cadence, and manifest rewrites. Hot tables get prioritized across the full pipeline, and routing weights update as table health improves.

  • Filter and join columns from agents inform compaction sort order
  • Snapshot expiry, orphan cleanup, and manifest rewrites adapt to access patterns
  • Routing shifts as tables get healthier — agents land on faster engines automatically
Adaptive Maintenance
Running
Compaction Scope— what will be targeted this run

Monitoring started at seq #4,812. Everything written from that point on is in scope.

already compacted
38% hot zone — will compact
#4,812 baseline#7,104 watermark#8,506 now
📦Compaction
Last: 57 minutes agoNext: in 2h
62%
42% of data files are below the 384 MB size threshold
1.24 TB compactable — will merge into ~3 files at 512 MB target
16 delete files amplify read cost — compaction will absorb them
Snapshot Expiration
Last: 4 hours agoNext: in 20h
31%
12 of 83 snapshots have passed the retention window
Retention policy: 5.0 days
Table creates 3.2 snapshots/hr — stale ones accumulate quickly
📝Rewrite Manifests
Last: 5 hours agoNext: in 1h
78%
92 total manifests: 89 data + 3 delete
Accumulating 2.1 manifests/hr — metadata overhead grows
High manifest count degrades query planning and scan performance
🧹Orphan File Cleanup
Last: 7 hours agoNext: in 17h
15%
No orphan files detected in current scan window

Connect in minutes
- no vendor lock-in

1

Connect catalogs & engines

Only metadata is processed — never retained or stored.

Apache Iceberg
AWS
Snowflake
DuckDB
2

Get visibility & insights

Telemetry reveals table health and actions needed.

Table health scores
Optimization opportunities
Cost & performance insights
3

Choose your mode

Autopilot, manual approval, or policy-driven.

Autopilot
Manual
Policies
4

Lakehouse optimized

Queries 10x faster
Cost down 76%
Engines optimized
AIs managed
Tables healthy
Fully governed
No vendor lock-in
No code / infra changes
No data changes
Set up in 10 minutes · Works with your existing stack

Get started

See LakeOps on your stack

Get a personalized walkthrough with your data and architecture.Short call, no commitment.

Typically 30 min · Free