Google Unveils Borderless Lakehouse to Connect AWS, Databricks, and Snowflake Data with AI Agents
Google Cloud has introduced the Borderless Lakehouse, an open data platform built on Apache Iceberg that enables AI agents to securely access and analyze enterprise data across AWS, Databricks, Snowflake, SaaS applications, and on-premises environments without moving data.
Xcademia Team
Xcademia Research Team

Google Introduces the Borderless Lakehouse for Enterprise AI
Enterprise data platforms are rapidly evolving as organizations embrace AI agents capable of continuously monitoring operations, analyzing business data, and executing workflows. Unlike traditional analytics that depend on scheduled reports, these intelligent agents require immediate access to trusted enterprise data wherever it resides. However, fragmented data environments, complex integration pipelines, and governance challenges often prevent organizations from fully realizing the potential of enterprise AI.
At Google Cloud Next Tokyo, Google announced major enhancements to its Borderless Lakehouse, an open data platform built on Apache Iceberg that enables enterprises to securely analyze and activate data across cloud providers, on-premises systems, and SaaS applications without moving it. The platform is designed to help AI agents securely query and reason over distributed enterprise data while reducing the costs and complexity associated with traditional data architectures.
By combining open standards with Google Cloud's analytics and AI capabilities, the Borderless Lakehouse is designed to transform enterprise data from a traditional repository into a system of action that supports real-time reasoning, analytics, and business workflows.
From a Data Repository to a System of Action
Traditional lakehouses have primarily served as centralized repositories for storing enterprise data. While effective for reporting and business intelligence, they often require organizations to build complex ETL pipelines before data becomes available for analytics or machine learning.
Google envisions the next generation of lakehouses as systems of action. In this model, always-on AI agents continuously monitor business operations, identify anomalies, analyze events as they happen, and automate workflows instead of simply generating reports.
To support this shift, AI agents need secure access to an organization's entire data estate along with the business context required to understand which data to use and when. The Borderless Lakehouse is designed to provide this foundation while allowing enterprises to keep their data in its existing location.
Built on an Open Apache Iceberg Foundation
The Borderless Lakehouse is built on Apache Iceberg, an open table format that enables consistent access to analytical data across multiple cloud environments and processing engines.
Rather than requiring organizations to migrate data into a proprietary platform, Google is expanding support for open interoperability. This allows enterprises to continue using their existing cloud platforms while enabling BigQuery, Managed Service for Apache Spark, and other Iceberg-compatible engines to securely analyze the same datasets.
By adopting open standards, organizations gain greater flexibility to modernize their analytics environments while avoiding vendor lock-in.
Federating Enterprise Data with Iceberg REST
A key enhancement announced at Next Tokyo is catalog federation, currently available in preview.
Using the Iceberg REST Catalog, the Borderless Lakehouse can discover and query remote enterprise data without requiring organizations to build complex data pipelines. Google announced support for catalog federation with:
These integrations provide secure, bidirectional connectivity through BigQuery, Managed Service for Apache Spark, and any Apache Iceberg-compatible engine. Instead of maintaining multiple copies of the same information, organizations can securely analyze distributed datasets while preserving consistent metadata and governance.
This open architecture helps reduce infrastructure complexity while accelerating access to enterprise data across multiple cloud environments.

Zero-Copy Data Integration Across Enterprise Applications
Google is extending the Borderless Lakehouse beyond cloud data platforms by introducing zero-copy integrations with leading enterprise SaaS applications, including SAP, Salesforce, and Workday.
Traditionally, organizations have relied on Extract, Transform, and Load (ETL) pipelines to move data from business applications into centralized warehouses before it could be analyzed. These processes often increase storage costs, introduce delays, and require continuous maintenance.
With the Borderless Lakehouse, BigQuery can directly query live application data without copying it into another environment. At the same time, these applications can leverage Google's AI capabilities on their own data in place, making it easier to combine finance, human resources, and customer information for enterprise-wide analytics.
This approach enables organizations to work with current business data while maintaining a single source of truth.
Three Key Benefits of the Borderless Lakehouse
Google highlighted three major advantages of its new architecture.
(1) Zero-Copy Cross-Cloud Analytics
Organizations can instantly discover and query enterprise data stored across multiple platforms without duplicating files. This allows data teams to analyze the same Apache Iceberg datasets regardless of where they are stored.
(2) Bidirectional Interoperability
The Borderless Lakehouse supports secure read and write operations across connected environments. Organizations can query external tables from BigQuery or Managed Service for Apache Spark, enrich datasets using Google's AI capabilities, and securely share those results back to partner platforms for downstream business processes.
(3) Unified Governance and Access Control
The platform provides built-in governance with trusted business context for AI agents. Table-level security and credential vending help ensure secure access regardless of which platform initiates a query, allowing organizations to maintain consistent governance across their multi-cloud environments.
Extending the Lakehouse to Operational Databases
Google is also expanding the Borderless Lakehouse across its database portfolio to better connect operational systems with analytical workloads.
Spanner Omni enables organizations to run Google's globally scalable database outside Google Cloud while remaining connected to the broader analytics ecosystem.
Meanwhile, Lakehouse Federation for AlloyDB allows transactional databases to directly query analytical warehouses without requiring data movement.
By combining live operational information with historical analytical data, organizations can perform real-time analysis while reducing the complexity and cost associated with maintaining duplicate datasets.
Together, these capabilities move the Borderless Lakehouse closer to Google's vision of a truly borderless enterprise data platform.
Bringing Google AI Directly to AWS and Azure Data
One of the biggest challenges in enterprise analytics has been the cost of working across multiple cloud providers. Running AI models or advanced analytics against data stored in AWS or Azure often requires moving large datasets, resulting in network latency, data transfer fees, and complex ETL pipelines.
The Borderless Lakehouse addresses these challenges through Cross-Cloud Interconnects, providing secure, dedicated private connections between Google Cloud and other cloud environments.
Unlike public internet connections, these interconnects deliver consistent bandwidth, lower latency, and predictable pricing through subscription-based connectivity ranging from 1 Gbps to 100 Gbps.
Google also announced support for zero variable egress costs when accessing AWS data through Partner Cross-Cloud Interconnect pricing. Instead of unpredictable transfer charges, organizations can benefit from flat-rate pricing backed by service-level agreements, making multi-cloud analytics more cost-effective.
Intelligent Cross-Cloud Caching Improves Performance
To further reduce latency and network costs, the Borderless Lakehouse introduces intelligent cross-cloud caching.
Frequently accessed remote data fragments are temporarily stored inside Google Cloud, allowing subsequent business intelligence and ad hoc analytics queries to retrieve cached information instead of repeatedly transferring the same data across cloud environments.
Combined with:
BigQuery's vectorized processing
BigQuery AI Functions for multimodal analysis
Managed Spark with Lightning Engine
Scalable metadata storage
the platform enables organizations to analyze petabyte-scale datasets while maintaining high performance across distributed cloud environments.

Bridging the Agent Trust Gap with Knowledge Catalog
As enterprises deploy more AI agents, access to data alone is no longer enough. AI systems also need business context to accurately understand and interpret enterprise information. Without that context, agents risk generating inaccurate responses or hallucinations when working with complex datasets.
To address this challenge, Google Cloud is introducing Knowledge Catalog, an always-on agentic context engine that provides AI agents with a unified view of enterprise metadata across multiple cloud environments. Rather than moving physical data, Knowledge Catalog aggregates and organizes metadata, helping AI systems understand the meaning, relationships, and governance of enterprise information.
This semantic layer enables AI agents to discover trusted data more efficiently while maintaining enterprise security and compliance requirements.
Creating a Unified Enterprise Context
The Borderless Lakehouse automatically synchronizes its runtime catalog with AWS Glue, Databricks Unity Catalog, and Snowflake Horizon, with these integrations currently available in preview.
Knowledge Catalog ingests metadata from these connected catalogs, extracts business context, and indexes enterprise information into a searchable semantic layer. Instead of relying on technical database schemas alone, it translates metadata into business-friendly terminology while providing column-level lineage across datasets.
As enterprise schemas evolve, Knowledge Catalog continuously updates these definitions, helping AI agents work with current and trusted business information without requiring manual metadata management.
Improving Trust, Governance, and Discovery
Google highlighted several advantages of using Knowledge Catalog as the intelligence layer for enterprise AI.
Organizations can reduce the need for expensive data migration projects while maintaining a consolidated view of metadata across their entire multi-cloud environment. AI agents gain access to trusted business definitions and data lineage, enabling them to interpret information more accurately and produce grounded responses.
Knowledge Catalog also embeds governance directly into the metadata layer, allowing AI agents to respect enterprise access permissions and compliance policies automatically whenever they query business data.
Building Enterprise AI Agents with Gemini Enterprise
Google is pairing the Borderless Lakehouse with Gemini Enterprise to make conversational analytics more accessible across organizations.
By combining the open source Google Cloud Data Agent Kit with the Conversational Analytics API, developers can build and deploy custom data agents capable of operating across the Borderless Lakehouse. Once published to Gemini Enterprise, these agents allow business users to interact with enterprise data using natural language rather than writing SQL queries or navigating traditional dashboards.
This approach enables employees across finance, operations, sales, and other business functions to retrieve insights more quickly while reducing dependence on technical teams.
Simplifying AI Agent Development
The Google Cloud Data Agent Kit provides developers with a complete toolkit for creating enterprise data agents using familiar development environments such as Visual Studio Code.
Built-in Model Context Protocol (MCP) tools establish secure connections to services including BigQuery, Managed Service for Apache Spark, and Cloud Storage, eliminating the need to develop custom integrations or manually include large database schemas within AI prompts.
Because these tools are integrated into the Borderless Lakehouse, developers can build AI agents that securely access governed enterprise data while significantly reducing development complexity.
Self-Service Analytics with Grounded AI
The Borderless Lakehouse allows business users to move beyond static dashboards by asking questions in natural language through Gemini Enterprise.
Instead of manually searching through reports, users can request insights from datasets distributed across multiple cloud environments. AI agents translate these requests into analytical queries while relying on Knowledge Catalog to ensure responses are grounded in approved business terminology and trusted metadata.
This combination of conversational analytics, semantic understanding, and enterprise governance helps improve the accuracy of AI-generated insights while supporting secure, organization-wide access to data.
The Economics of the Borderless Lakehouse
Google also emphasized the financial benefits of its new architecture, positioning the Borderless Lakehouse as a cost-efficient foundation for enterprise AI.
Traditional multi-cloud analytics often involve expensive data transfers, duplicated storage, and growing AI inference costs. The Borderless Lakehouse addresses these challenges through Cross-Cloud Interconnects, zero-copy data sharing, and Knowledge Catalog, reducing unnecessary data movement while supplying AI models with only the business context required for each task.
BigQuery AI further helps organizations manage AI spending through built-in token estimation, usage limits, and optimized processing modes that automatically select smaller distilled models when appropriate.
According to Google, customers using BigQuery's cost-optimized AI functions are already achieving up to a 230× reduction in token consumption, helping lower AI operating costs while enabling enterprise-scale deployments.

Next Steps
Google says the future belongs to systems of action, where AI agents can securely query, reason over, and act on enterprise data wherever it is stored. To help organizations get started, Google recommends exploring the official Google Cloud Lakehouse documentation and the Building a Borderless Lakehouse codelab to learn more about catalog federation, zero-copy analytics, and AI-powered data workflows.
Conclusion
Google Cloud's Borderless Lakehouse marks a significant step toward making enterprise AI more practical in multi-cloud environments. Built on Apache Iceberg and enhanced with catalog federation, zero-copy integrations, Cross-Cloud Interconnects, Knowledge Catalog, and Gemini Enterprise, the platform enables organizations to securely query, understand, and activate data wherever it resides.
By reducing data movement, simplifying governance, and giving AI agents access to trusted business context, the Borderless Lakehouse helps organizations build AI systems that can securely query, reason over, and act on distributed enterprise data while controlling infrastructure and AI costs. Built on open technologies such as Apache Iceberg, Google's interoperable approach provides a strong foundation for the next generation of enterprise analytics and AI-driven business operations.
Source: Google Cloud Blog
About the Author