Activate your data, regardless where it lives, with Google Cloud's borderless Lakehouse

The reality is, your data lives everywhere—on Google, across other clouds like AWS and Azure, and in your SaaS applications. Built on open Apache Iceberg and connecting to your operational systems both on-premises and across clouds, the borderless Lakehouse lets you activate your data wherever it lives, without moving it.

Bi-directional catalog interoperability

Empower analysts and conversational agents to instantly query and act on multi-cloud data wherever it physically resides, removing the overhead of building and maintaining traditional ETL pipelines. Built on the open Iceberg REST catalog, this architecture delivers secure, bi-directional catalog federation (now in preview) across AWS Glue, Databricks Unity, and Snowflake Horizon—unlocking unified, real-time insights for BigQuery, Managed Service for Apache Spark, and any Iceberg-compatible engine.

Zero copy data from SaaS applications

Break down SaaS data silos and fuel your AI strategy without the risk, lag, or cost of data movement. By extending your lakehouse directly to the application layer, you achieve secure, zero-copy integration with core enterprise systems like SAP, Salesforce, and Workday. This allows BigQuery to query live transactional data in real time—eliminating complex ETL pipelines—while enabling you to run BigQuery’s powerful AI engines directly on your source data in-place. The result is a unified, secure foundation across finance, HR, and customer operations that dramatically accelerates time-to-decision and lowers total cost of ownership.

Get the openness of Apache Iceberg with enterprise-grade storage management

The future of enterprise data isn't just about storage—it's about taking immediate action. The borderless Lakehouse gives your analyst and agents the power to safely analyze and work with your data wherever it lives. Eliminate the need for slow, expensive data-copying projects so your business can confidently make faster, automated decisions in real time while keeping your cloud budgets under control.

Bring Google AI directly to your AWS and Azure data

Eliminate the multi-cloud tax and accelerate your data strategy. The borderless Lakehouse delivers predictable, flat-rate economics and private, SLA-backed Partner Cross-Cloud Interconnect connectivity that eliminates variable egress costs for your AWS data. By pairing intelligent cross-cloud caching with BigQuery’s vectorized processing, Spark’s Lightning Engine runtime, and native multimodal AI functions, you can securely analyze petabytes of remote data with sub-second performance—without repeated, costly transfers or performance trade-offs.

Empower analysts and conversational agents to instantly query and act on multi-cloud data wherever it physically resides, removing the overhead of building and maintaining traditional ETL pipelines. Built on the open Iceberg REST catalog, this architecture delivers secure, bi-directional catalog federation (now in preview) across AWS Glue, Databricks Unity, and Snowflake Horizon—unlocking unified, real-time insights for BigQuery, Managed Service for Apache Spark, and any Iceberg-compatible engine.

Break down SaaS data silos and fuel your AI strategy without the risk, lag, or cost of data movement. By extending your lakehouse directly to the application layer, you achieve secure, zero-copy integration with core enterprise systems like SAP, Salesforce, and Workday. This allows BigQuery to query live transactional data in real time—eliminating complex ETL pipelines—while enabling you to run BigQuery’s powerful AI engines directly on your source data in-place. The result is a unified, secure foundation across finance, HR, and customer operations that dramatically accelerates time-to-decision and lowers total cost of ownership.

Get the openness of Apache Iceberg with enterprise-grade storage management

The future of enterprise data isn't just about storage—it's about taking immediate action. The borderless Lakehouse gives your analyst and agents the power to safely analyze and work with your data wherever it lives. Eliminate the need for slow, expensive data-copying projects so your business can confidently make faster, automated decisions in real time while keeping your cloud budgets under control.

Bi-directional catalog interoperability

Empower analysts and conversational agents to instantly query and act on multi-cloud data wherever it physically resides, removing the overhead of building and maintaining traditional ETL pipelines. Built on the open Iceberg REST catalog, this architecture delivers secure, bi-directional catalog federation (now in preview) across AWS Glue, Databricks Unity, and Snowflake Horizon—unlocking unified, real-time insights for BigQuery, Managed Service for Apache Spark, and any Iceberg-compatible engine.

Zero copy data from SaaS applications

Break down SaaS data silos and fuel your AI strategy without the risk, lag, or cost of data movement. By extending your lakehouse directly to the application layer, you achieve secure, zero-copy integration with core enterprise systems like SAP, Salesforce, and Workday. This allows BigQuery to query live transactional data in real time—eliminating complex ETL pipelines—while enabling you to run BigQuery’s powerful AI engines directly on your source data in-place. The result is a unified, secure foundation across finance, HR, and customer operations that dramatically accelerates time-to-decision and lowers total cost of ownership.

Bring Google AI directly to your AWS and Azure data

Eliminate the multi-cloud tax and accelerate your data strategy. The borderless Lakehouse delivers predictable, flat-rate economics and private, SLA-backed Partner Cross-Cloud Interconnect connectivity that eliminates variable egress costs for your AWS data. By pairing intelligent cross-cloud caching with BigQuery’s vectorized processing, Spark’s Lightning Engine runtime, and native multimodal AI functions, you can securely analyze petabytes of remote data with sub-second performance—without repeated, costly transfers or performance trade-offs.

Empower analysts and conversational agents to instantly query and act on multi-cloud data wherever it physically resides, removing the overhead of building and maintaining traditional ETL pipelines. Built on the open Iceberg REST catalog, this architecture delivers secure, bi-directional catalog federation (now in preview) across AWS Glue, Databricks Unity, and Snowflake Horizon—unlocking unified, real-time insights for BigQuery, Managed Service for Apache Spark, and any Iceberg-compatible engine.

Break down SaaS data silos and fuel your AI strategy without the risk, lag, or cost of data movement. By extending your lakehouse directly to the application layer, you achieve secure, zero-copy integration with core enterprise systems like SAP, Salesforce, and Workday. This allows BigQuery to query live transactional data in real time—eliminating complex ETL pipelines—while enabling you to run BigQuery’s powerful AI engines directly on your source data in-place. The result is a unified, secure foundation across finance, HR, and customer operations that dramatically accelerates time-to-decision and lowers total cost of ownership.

Differentiated engines like BigQuery and Managed Spark

Google Cloud Lakehouse delivers managed Apache Iceberg with true read/write interoperability across BigQuery and Spark. Get the price-performance and scale of BigQuery and Google Managed Service for Apache Spark with flexible storage to accelerate your multimodal analytics, data science agentic workloads.

Advanced BigQuery workloads on Iceberg

Empower your teams to make split-second business decisions using a secure, unified platform that synchronizes live data in real-time without costly processing delays. By seamlessly bringing together structured business metrics and unstructured assets using BigQuery ObjectRef, you unlock the advanced analytics and next-generation AI needed to drive immediate competitive advantage.

Supercharge your data science initiatives by running Apache Spark directly on your lakehouse with up to 4.9x faster performance. Powered by the Lightning Engine, this managed service eliminates complex infrastructure management so your data scientists can train models and analyze massive datasets in record time. By combining lightning-fast processing with direct, secure access to your lakehouse data, you dramatically accelerate your team's time-to-insight and lower machine learning development costs.

Accelerate data science with Managed Service for Apache Spark

Supercharge your data science initiatives by running Apache Spark directly on your lakehouse with up to 4.9x faster performance. Powered by the Lightning Engine, this managed service eliminates complex infrastructure management so your data scientists can train models and analyze massive datasets in record time. By combining lightning-fast processing with direct, secure access to your lakehouse data, you dramatically accelerate your team's time-to-insight and lower machine learning development costs.

Differentiated engines like BigQuery and Managed Spark

Google Cloud Lakehouse delivers managed Apache Iceberg with true read/write interoperability across BigQuery and Spark. Get the price-performance and scale of BigQuery and Google Managed Service for Apache Spark with flexible storage to accelerate your multimodal analytics, data science agentic workloads.

Accelerate data science with Managed Service for Apache Spark

Supercharge your data science initiatives by running Apache Spark directly on your lakehouse with up to 4.9x faster performance. Powered by the Lightning Engine, this managed service eliminates complex infrastructure management so your data scientists can train models and analyze massive datasets in record time. By combining lightning-fast processing with direct, secure access to your lakehouse data, you dramatically accelerate your team's time-to-insight and lower machine learning development costs.

Advanced BigQuery workloads on Iceberg

Empower your teams to make split-second business decisions using a secure, unified platform that synchronizes live data in real-time without costly processing delays. By seamlessly bringing together structured business metrics and unstructured assets using BigQuery ObjectRef, you unlock the advanced analytics and next-generation AI needed to drive immediate competitive advantage.

Supercharge your data science initiatives by running Apache Spark directly on your lakehouse with up to 4.9x faster performance. Powered by the Lightning Engine, this managed service eliminates complex infrastructure management so your data scientists can train models and analyze massive datasets in record time. By combining lightning-fast processing with direct, secure access to your lakehouse data, you dramatically accelerate your team's time-to-insight and lower machine learning development costs.

Enrich data and generate meaning through continuous learning

Accelerate AI-readiness and eliminate manual curation by turning raw, unstructured data into a self-updating map of your business. By leveraging Gemini to automate data enrichment—mining schemas, files, and logs to auto-tag assets—you establish a trusted natural-language glossary of your enterprise operations. These built-in semantic guardrails and pre-validated logic eliminate costly query errors and AI hallucinations, enabling both your analysts and autonomous agents make decisions based on trusted data.

Power search and unleash agents with secure retrieval

Give your AI assistants the exact business context they need in under a second—safely and securely. Our advanced search instantly connects your AI to the right information for accurate decision-making, while strict security permissions guarantee they only see data they are authorized to access. This allows you to easily build and deploy trustworthy AI tools using our pre-built Data Cloud Agents or flexible Data Agent Kit.

Open lakehouse governance with always on context for your agents

Knowledge Catalog is the AI-driven context engine for your lakehouse. It auto-catalogs all Google Cloud and third-party data, eliminating manual curation to deliver the grounded enterprise truth and data relationships essential for reliable agentic AI.

Aggregate context across your data estate

Knowledge Catalog eliminates data fragmentation by automatically unifying metadata from Google Cloud, partner platforms, and third-party catalogs (such as Atlan, Collibra, and Datahub) into a single, governed source of truth. By automating the discovery and indexing of your entire data footprint, you dramatically reduce compliance risk, eliminate manual data-mapping overhead, and help both your teams and AI systems can find and trust the data they need instantly.

Accelerate AI-readiness and eliminate manual curation by turning raw, unstructured data into a self-updating map of your business. By leveraging Gemini to automate data enrichment—mining schemas, files, and logs to auto-tag assets—you establish a trusted natural-language glossary of your enterprise operations. These built-in semantic guardrails and pre-validated logic eliminate costly query errors and AI hallucinations, enabling both your analysts and autonomous agents make decisions based on trusted data.

Give your AI assistants the exact business context they need in under a second—safely and securely. Our advanced search instantly connects your AI to the right information for accurate decision-making, while strict security permissions guarantee they only see data they are authorized to access. This allows you to easily build and deploy trustworthy AI tools using our pre-built Data Cloud Agents or flexible Data Agent Kit.

Open lakehouse governance with always on context for your agents

Knowledge Catalog is the AI-driven context engine for your lakehouse. It auto-catalogs all Google Cloud and third-party data, eliminating manual curation to deliver the grounded enterprise truth and data relationships essential for reliable agentic AI.

Enrich data and generate meaning through continuous learning

Accelerate AI-readiness and eliminate manual curation by turning raw, unstructured data into a self-updating map of your business. By leveraging Gemini to automate data enrichment—mining schemas, files, and logs to auto-tag assets—you establish a trusted natural-language glossary of your enterprise operations. These built-in semantic guardrails and pre-validated logic eliminate costly query errors and AI hallucinations, enabling both your analysts and autonomous agents make decisions based on trusted data.

Power search and unleash agents with secure retrieval

Give your AI assistants the exact business context they need in under a second—safely and securely. Our advanced search instantly connects your AI to the right information for accurate decision-making, while strict security permissions guarantee they only see data they are authorized to access. This allows you to easily build and deploy trustworthy AI tools using our pre-built Data Cloud Agents or flexible Data Agent Kit.

Aggregate context across your data estate

Knowledge Catalog eliminates data fragmentation by automatically unifying metadata from Google Cloud, partner platforms, and third-party catalogs (such as Atlan, Collibra, and Datahub) into a single, governed source of truth. By automating the discovery and indexing of your entire data footprint, you dramatically reduce compliance risk, eliminate manual data-mapping overhead, and help both your teams and AI systems can find and trust the data they need instantly.

Accelerate AI-readiness and eliminate manual curation by turning raw, unstructured data into a self-updating map of your business. By leveraging Gemini to automate data enrichment—mining schemas, files, and logs to auto-tag assets—you establish a trusted natural-language glossary of your enterprise operations. These built-in semantic guardrails and pre-validated logic eliminate costly query errors and AI hallucinations, enabling both your analysts and autonomous agents make decisions based on trusted data.

Give your AI assistants the exact business context they need in under a second—safely and securely. Our advanced search instantly connects your AI to the right information for accurate decision-making, while strict security permissions guarantee they only see data they are authorized to access. This allows you to easily build and deploy trustworthy AI tools using our pre-built Data Cloud Agents or flexible Data Agent Kit.

Combine analytical and operational data for real-time, agentic applications

Google Cloud's borderless Lakehouse bridges the divide between real-time transactional systems and analytical intelligence. By connecting your active databases directly to your analytical engine, you can transition from simply analyzing historical reports to triggering instant, automated operational actions

Introducing Spanner Omni

Spanner Omni brings the "unbreakable" DNA of Google Spanner—can achieve high availability with zonal and regional, strong consistency, and virtually unlimited scale—to your local hardware, your private clouds, and every edge in between.

Lakehouse Federation for AlloyDB enables real-time business insights by allowing your transactional systems to directly query data warehouses without costly data movement. This eliminates data silos and expensive infrastructure, empowering your team to make faster, data-driven decisions using both operational and historical data together.

Lakehouse federation with AlloyDB

Lakehouse Federation for AlloyDB enables real-time business insights by allowing your transactional systems to directly query data warehouses without costly data movement. This eliminates data silos and expensive infrastructure, empowering your team to make faster, data-driven decisions using both operational and historical data together.

Combine analytical and operational data for real-time, agentic applications

Google Cloud's borderless Lakehouse bridges the divide between real-time transactional systems and analytical intelligence. By connecting your active databases directly to your analytical engine, you can transition from simply analyzing historical reports to triggering instant, automated operational actions

Lakehouse federation with AlloyDB

Lakehouse Federation for AlloyDB enables real-time business insights by allowing your transactional systems to directly query data warehouses without costly data movement. This eliminates data silos and expensive infrastructure, empowering your team to make faster, data-driven decisions using both operational and historical data together.

Introducing Spanner Omni

Spanner Omni brings the "unbreakable" DNA of Google Spanner—can achieve high availability with zonal and regional, strong consistency, and virtually unlimited scale—to your local hardware, your private clouds, and every edge in between.

Lakehouse Federation for AlloyDB enables real-time business insights by allowing your transactional systems to directly query data warehouses without costly data movement. This eliminates data silos and expensive infrastructure, empowering your team to make faster, data-driven decisions using both operational and historical data together.

Customer stories


How Etsy supports search relevance and ML velocity with Google Cloud’s Lakehouse
"Our cost savings go hand-in-hand with our experimentation velocity and the ability for teams to iterate quickly. If you can move both of those things in a positive direction at the same time, that’s a massive win.":
Matthew Hall
Senior Engineering Manager, Etsy
Etys office
Roundel™ Media designed by Target: Helping retail media ads hit the bullseye with Google Cloud
"Our lakehouse architecture, with all the agentic, analytical, and developer tools on top, allows us to start building the future faster than we could have before. Google’s AI ecosystem gives us a natural runway into the next set of capabilities."
Guthrie Collin
VP of Product Management, Roundel
Target and Google Cloud
2:55
Achieve interoperability and unified context with the governed, open Lakehouse
"The unified data and context layer we’ve built is now the foundation for accelerating our agentic and AI journey... pushing toward autonomous network operations, spanning terrestrial networks and satellite capacity management."
Sumeet Singh
Chief Data and AI Officer
The governed open lakeouse

Start your borderless Lakehouse journey today

Act on data anywhere, instantly.

Cloud logo

What is Google Cloud's borderless Lakehouse, and how does it work?

Google Cloud's borderless Lakehouse is an open, decentralized data architecture built on Apache Iceberg that enables analytics engines and AI agents to query, analyze, and act on cross-cloud, on-premises, and SaaS data in place without data movement or ETL pipelines. It operates by decoupling analytical compute from underlying storage, using the open Iceberg REST catalog standard to synchronize table schemas across distributed clouds. Execution engines such as BigQuery and Managed Service for Apache Spark run vectorized queries and in-place AI functions over remote data. This transforms distributed enterprise data estates into a single, unified system of action without forcing teams to migrate physical files.

How does catalog federation work with AWS Glue, Databricks Unity, and Snowflake Horizon?

Multi-cloud catalog federation in the borderless Lakehouse leverages the open Apache Iceberg REST catalog standard to provide secure, bi-directional metadata discovery and execution across AWS Glue, Databricks Unity Catalog, and Snowflake Horizon Catalog. Instead of copying underlying datasets, query engines read table schemas and manifest files remotely over secure endpoints. This allows BigQuery, Managed Service for Apache Spark, and open-source engines to query multi-cloud tables in place while maintaining unified governance across disparate object stores.

How does zero-copy SaaS integration work for applications like SAP, Salesforce, and Workday?

Zero-copy SaaS integration extends the borderless Lakehouse directly to the enterprise application layer, enabling BigQuery to run real-time queries against live transactional systems like SAP, Salesforce, and Workday without constructing ETL pipelines. By querying data where it natively resides, organizations eliminate pipeline latency and eliminate data sync failures. This direct connectivity also allows BigQuery AI functions to execute directly on live ERP, CRM, and HR data, streamlining cross-domain enterprise analytics.



How are transactional databases like Spanner and AlloyDB integrated into the borderless Lakehouse?

Transactional databases are integrated into the borderless Lakehouse through Spanner Omni, which operates in multi-cloud environments outside Google Cloud, and Lakehouse Federation for AlloyDB. This bidirectional bridge allows operational systems to query analytical lakehouse data directly without waiting for batch syncs. As a result, developers and data engineers can join live transactional records with petabytes of historical analytical data in real time, powering low-latency operational applications and real-time dashboards.

How does the borderless Lakehouse empower AI agents and prevent hallucinations?

The borderless Lakehouse connects autonomous reasoning agents to Knowledge Catalog, an always-on context engine that provides precise business metadata, lineage, and semantic definitions without moving source data. By feeding precise, pre-filtered business context into LLM prompts, the architecture can help prevent model hallucinations and eliminates unnecessary agent reasoning loops. This semantic grounding improves answer accuracy for autonomous agents like Gemini Enterprise while reducing token consumption.

What are the economic and performance benefits of the borderless Lakehouse architecture?

The borderless Lakehouse lowers total cost of ownership (TCO) and provides predictability by using Cross-Cloud Interconnects with flat-rate pricing that eliminates variable data egress costs for AWS data. To maximize query performance across distributed clouds, the system pairs intelligent cross-cloud local caching with BigQuery’s vectorized execution engine, delivering sub-second response times. By replacing duplicate storage and fragile ETL compute, organizations significantly reduce cloud infrastructure spend and operational overhead.

Google Cloud