Activate your data, regardless where it lives, with Google Cloud's borderless Lakehouse
The reality is, your data lives everywhere—on Google, across other clouds like AWS and Azure, and in your SaaS applications. Built on open Apache Iceberg and connecting to your operational systems both on-premises and across clouds, the borderless Lakehouse lets you activate your data wherever it lives, without moving it.
News and events
Bi-directional catalog interoperability
Empower analysts and conversational agents to instantly query and act on multi-cloud data wherever it physically resides, removing the overhead of building and maintaining traditional ETL pipelines. Built on the open Iceberg REST catalog, this architecture delivers secure, bi-directional catalog federation (now in preview) across AWS Glue, Databricks Unity, and Snowflake Horizon—unlocking unified, real-time insights for BigQuery, Managed Service for Apache Spark, and any Iceberg-compatible engine.
Zero copy data from SaaS applications
Break down SaaS data silos and fuel your AI strategy without the risk, lag, or cost of data movement. By extending your lakehouse directly to the application layer, you achieve secure, zero-copy integration with core enterprise systems like SAP, Salesforce, and Workday. This allows BigQuery to query live transactional data in real time—eliminating complex ETL pipelines—while enabling you to run BigQuery’s powerful AI engines directly on your source data in-place. The result is a unified, secure foundation across finance, HR, and customer operations that dramatically accelerates time-to-decision and lowers total cost of ownership.
The future of enterprise data isn't just about storage—it's about taking immediate action. The borderless Lakehouse gives your analyst and agents the power to safely analyze and work with your data wherever it lives. Eliminate the need for slow, expensive data-copying projects so your business can confidently make faster, automated decisions in real time while keeping your cloud budgets under control.
Bring Google AI directly to your AWS and Azure data
Eliminate the multi-cloud tax and accelerate your data strategy. The borderless Lakehouse delivers predictable, flat-rate economics and private, SLA-backed Partner Cross-Cloud Interconnect connectivity that eliminates variable egress costs for your AWS data. By pairing intelligent cross-cloud caching with BigQuery’s vectorized processing, Spark’s Lightning Engine runtime, and native multimodal AI functions, you can securely analyze petabytes of remote data with sub-second performance—without repeated, costly transfers or performance trade-offs.
Bi-directional catalog interoperability
Empower analysts and conversational agents to instantly query and act on multi-cloud data wherever it physically resides, removing the overhead of building and maintaining traditional ETL pipelines. Built on the open Iceberg REST catalog, this architecture delivers secure, bi-directional catalog federation (now in preview) across AWS Glue, Databricks Unity, and Snowflake Horizon—unlocking unified, real-time insights for BigQuery, Managed Service for Apache Spark, and any Iceberg-compatible engine.
Zero copy data from SaaS applications
Break down SaaS data silos and fuel your AI strategy without the risk, lag, or cost of data movement. By extending your lakehouse directly to the application layer, you achieve secure, zero-copy integration with core enterprise systems like SAP, Salesforce, and Workday. This allows BigQuery to query live transactional data in real time—eliminating complex ETL pipelines—while enabling you to run BigQuery’s powerful AI engines directly on your source data in-place. The result is a unified, secure foundation across finance, HR, and customer operations that dramatically accelerates time-to-decision and lowers total cost of ownership.
The future of enterprise data isn't just about storage—it's about taking immediate action. The borderless Lakehouse gives your analyst and agents the power to safely analyze and work with your data wherever it lives. Eliminate the need for slow, expensive data-copying projects so your business can confidently make faster, automated decisions in real time while keeping your cloud budgets under control.
Bi-directional catalog interoperability
Empower analysts and conversational agents to instantly query and act on multi-cloud data wherever it physically resides, removing the overhead of building and maintaining traditional ETL pipelines. Built on the open Iceberg REST catalog, this architecture delivers secure, bi-directional catalog federation (now in preview) across AWS Glue, Databricks Unity, and Snowflake Horizon—unlocking unified, real-time insights for BigQuery, Managed Service for Apache Spark, and any Iceberg-compatible engine.
Zero copy data from SaaS applications
Break down SaaS data silos and fuel your AI strategy without the risk, lag, or cost of data movement. By extending your lakehouse directly to the application layer, you achieve secure, zero-copy integration with core enterprise systems like SAP, Salesforce, and Workday. This allows BigQuery to query live transactional data in real time—eliminating complex ETL pipelines—while enabling you to run BigQuery’s powerful AI engines directly on your source data in-place. The result is a unified, secure foundation across finance, HR, and customer operations that dramatically accelerates time-to-decision and lowers total cost of ownership.
Bring Google AI directly to your AWS and Azure data
Eliminate the multi-cloud tax and accelerate your data strategy. The borderless Lakehouse delivers predictable, flat-rate economics and private, SLA-backed Partner Cross-Cloud Interconnect connectivity that eliminates variable egress costs for your AWS data. By pairing intelligent cross-cloud caching with BigQuery’s vectorized processing, Spark’s Lightning Engine runtime, and native multimodal AI functions, you can securely analyze petabytes of remote data with sub-second performance—without repeated, costly transfers or performance trade-offs.
Bi-directional catalog interoperability
Empower analysts and conversational agents to instantly query and act on multi-cloud data wherever it physically resides, removing the overhead of building and maintaining traditional ETL pipelines. Built on the open Iceberg REST catalog, this architecture delivers secure, bi-directional catalog federation (now in preview) across AWS Glue, Databricks Unity, and Snowflake Horizon—unlocking unified, real-time insights for BigQuery, Managed Service for Apache Spark, and any Iceberg-compatible engine.
Zero copy data from SaaS applications
Break down SaaS data silos and fuel your AI strategy without the risk, lag, or cost of data movement. By extending your lakehouse directly to the application layer, you achieve secure, zero-copy integration with core enterprise systems like SAP, Salesforce, and Workday. This allows BigQuery to query live transactional data in real time—eliminating complex ETL pipelines—while enabling you to run BigQuery’s powerful AI engines directly on your source data in-place. The result is a unified, secure foundation across finance, HR, and customer operations that dramatically accelerates time-to-decision and lowers total cost of ownership.
Google Cloud Lakehouse delivers managed Apache Iceberg with true read/write interoperability across BigQuery and Spark. Get the price-performance and scale of BigQuery and Google Managed Service for Apache Spark with flexible storage to accelerate your multimodal analytics, data science agentic workloads.
Advanced BigQuery workloads on Iceberg
Empower your teams to make split-second business decisions using a secure, unified platform that synchronizes live data in real-time without costly processing delays. By seamlessly bringing together structured business metrics and unstructured assets using BigQuery ObjectRef, you unlock the advanced analytics and next-generation AI needed to drive immediate competitive advantage.
Accelerate data science with Managed Service for Apache Spark
Supercharge your data science initiatives by running Apache Spark directly on your lakehouse with up to 4.9x faster performance. Powered by the Lightning Engine, this managed service eliminates complex infrastructure management so your data scientists can train models and analyze massive datasets in record time. By combining lightning-fast processing with direct, secure access to your lakehouse data, you dramatically accelerate your team's time-to-insight and lower machine learning development costs.
Accelerate data science with Managed Service for Apache Spark
Supercharge your data science initiatives by running Apache Spark directly on your lakehouse with up to 4.9x faster performance. Powered by the Lightning Engine, this managed service eliminates complex infrastructure management so your data scientists can train models and analyze massive datasets in record time. By combining lightning-fast processing with direct, secure access to your lakehouse data, you dramatically accelerate your team's time-to-insight and lower machine learning development costs.
Google Cloud Lakehouse delivers managed Apache Iceberg with true read/write interoperability across BigQuery and Spark. Get the price-performance and scale of BigQuery and Google Managed Service for Apache Spark with flexible storage to accelerate your multimodal analytics, data science agentic workloads.
Accelerate data science with Managed Service for Apache Spark
Supercharge your data science initiatives by running Apache Spark directly on your lakehouse with up to 4.9x faster performance. Powered by the Lightning Engine, this managed service eliminates complex infrastructure management so your data scientists can train models and analyze massive datasets in record time. By combining lightning-fast processing with direct, secure access to your lakehouse data, you dramatically accelerate your team's time-to-insight and lower machine learning development costs.
Advanced BigQuery workloads on Iceberg
Empower your teams to make split-second business decisions using a secure, unified platform that synchronizes live data in real-time without costly processing delays. By seamlessly bringing together structured business metrics and unstructured assets using BigQuery ObjectRef, you unlock the advanced analytics and next-generation AI needed to drive immediate competitive advantage.
Accelerate data science with Managed Service for Apache Spark
Supercharge your data science initiatives by running Apache Spark directly on your lakehouse with up to 4.9x faster performance. Powered by the Lightning Engine, this managed service eliminates complex infrastructure management so your data scientists can train models and analyze massive datasets in record time. By combining lightning-fast processing with direct, secure access to your lakehouse data, you dramatically accelerate your team's time-to-insight and lower machine learning development costs.
Enrich data and generate meaning through continuous learning
Accelerate AI-readiness and eliminate manual curation by turning raw, unstructured data into a self-updating map of your business. By leveraging Gemini to automate data enrichment—mining schemas, files, and logs to auto-tag assets—you establish a trusted natural-language glossary of your enterprise operations. These built-in semantic guardrails and pre-validated logic eliminate costly query errors and AI hallucinations, enabling both your analysts and autonomous agents make decisions based on trusted data.
Power search and unleash agents with secure retrieval
Give your AI assistants the exact business context they need in under a second—safely and securely. Our advanced search instantly connects your AI to the right information for accurate decision-making, while strict security permissions guarantee they only see data they are authorized to access. This allows you to easily build and deploy trustworthy AI tools using our pre-built Data Cloud Agents or flexible Data Agent Kit.
Knowledge Catalog is the AI-driven context engine for your lakehouse. It auto-catalogs all Google Cloud and third-party data, eliminating manual curation to deliver the grounded enterprise truth and data relationships essential for reliable agentic AI.
Aggregate context across your data estate
Knowledge Catalog eliminates data fragmentation by automatically unifying metadata from Google Cloud, partner platforms, and third-party catalogs (such as Atlan, Collibra, and Datahub) into a single, governed source of truth. By automating the discovery and indexing of your entire data footprint, you dramatically reduce compliance risk, eliminate manual data-mapping overhead, and help both your teams and AI systems can find and trust the data they need instantly.
Enrich data and generate meaning through continuous learning
Accelerate AI-readiness and eliminate manual curation by turning raw, unstructured data into a self-updating map of your business. By leveraging Gemini to automate data enrichment—mining schemas, files, and logs to auto-tag assets—you establish a trusted natural-language glossary of your enterprise operations. These built-in semantic guardrails and pre-validated logic eliminate costly query errors and AI hallucinations, enabling both your analysts and autonomous agents make decisions based on trusted data.
Power search and unleash agents with secure retrieval
Give your AI assistants the exact business context they need in under a second—safely and securely. Our advanced search instantly connects your AI to the right information for accurate decision-making, while strict security permissions guarantee they only see data they are authorized to access. This allows you to easily build and deploy trustworthy AI tools using our pre-built Data Cloud Agents or flexible Data Agent Kit.
Knowledge Catalog is the AI-driven context engine for your lakehouse. It auto-catalogs all Google Cloud and third-party data, eliminating manual curation to deliver the grounded enterprise truth and data relationships essential for reliable agentic AI.
Enrich data and generate meaning through continuous learning
Accelerate AI-readiness and eliminate manual curation by turning raw, unstructured data into a self-updating map of your business. By leveraging Gemini to automate data enrichment—mining schemas, files, and logs to auto-tag assets—you establish a trusted natural-language glossary of your enterprise operations. These built-in semantic guardrails and pre-validated logic eliminate costly query errors and AI hallucinations, enabling both your analysts and autonomous agents make decisions based on trusted data.
Power search and unleash agents with secure retrieval
Give your AI assistants the exact business context they need in under a second—safely and securely. Our advanced search instantly connects your AI to the right information for accurate decision-making, while strict security permissions guarantee they only see data they are authorized to access. This allows you to easily build and deploy trustworthy AI tools using our pre-built Data Cloud Agents or flexible Data Agent Kit.
Aggregate context across your data estate
Knowledge Catalog eliminates data fragmentation by automatically unifying metadata from Google Cloud, partner platforms, and third-party catalogs (such as Atlan, Collibra, and Datahub) into a single, governed source of truth. By automating the discovery and indexing of your entire data footprint, you dramatically reduce compliance risk, eliminate manual data-mapping overhead, and help both your teams and AI systems can find and trust the data they need instantly.
Enrich data and generate meaning through continuous learning
Accelerate AI-readiness and eliminate manual curation by turning raw, unstructured data into a self-updating map of your business. By leveraging Gemini to automate data enrichment—mining schemas, files, and logs to auto-tag assets—you establish a trusted natural-language glossary of your enterprise operations. These built-in semantic guardrails and pre-validated logic eliminate costly query errors and AI hallucinations, enabling both your analysts and autonomous agents make decisions based on trusted data.
Power search and unleash agents with secure retrieval
Give your AI assistants the exact business context they need in under a second—safely and securely. Our advanced search instantly connects your AI to the right information for accurate decision-making, while strict security permissions guarantee they only see data they are authorized to access. This allows you to easily build and deploy trustworthy AI tools using our pre-built Data Cloud Agents or flexible Data Agent Kit.
Google Cloud's borderless Lakehouse bridges the divide between real-time transactional systems and analytical intelligence. By connecting your active databases directly to your analytical engine, you can transition from simply analyzing historical reports to triggering instant, automated operational actions
Introducing Spanner Omni
Spanner Omni brings the "unbreakable" DNA of Google Spanner—can achieve high availability with zonal and regional, strong consistency, and virtually unlimited scale—to your local hardware, your private clouds, and every edge in between.
Lakehouse federation with AlloyDB
Lakehouse Federation for AlloyDB enables real-time business insights by allowing your transactional systems to directly query data warehouses without costly data movement. This eliminates data silos and expensive infrastructure, empowering your team to make faster, data-driven decisions using both operational and historical data together.
Lakehouse federation with AlloyDB
Lakehouse Federation for AlloyDB enables real-time business insights by allowing your transactional systems to directly query data warehouses without costly data movement. This eliminates data silos and expensive infrastructure, empowering your team to make faster, data-driven decisions using both operational and historical data together.
Google Cloud's borderless Lakehouse bridges the divide between real-time transactional systems and analytical intelligence. By connecting your active databases directly to your analytical engine, you can transition from simply analyzing historical reports to triggering instant, automated operational actions
Lakehouse federation with AlloyDB
Lakehouse Federation for AlloyDB enables real-time business insights by allowing your transactional systems to directly query data warehouses without costly data movement. This eliminates data silos and expensive infrastructure, empowering your team to make faster, data-driven decisions using both operational and historical data together.
Introducing Spanner Omni
Spanner Omni brings the "unbreakable" DNA of Google Spanner—can achieve high availability with zonal and regional, strong consistency, and virtually unlimited scale—to your local hardware, your private clouds, and every edge in between.
Lakehouse federation with AlloyDB
Lakehouse Federation for AlloyDB enables real-time business insights by allowing your transactional systems to directly query data warehouses without costly data movement. This eliminates data silos and expensive infrastructure, empowering your team to make faster, data-driven decisions using both operational and historical data together.
Customer stories






Google Cloud's borderless Lakehouse is an open, decentralized data architecture built on Apache Iceberg that enables analytics engines and AI agents to query, analyze, and act on cross-cloud, on-premises, and SaaS data in place without data movement or ETL pipelines. It operates by decoupling analytical compute from underlying storage, using the open Iceberg REST catalog standard to synchronize table schemas across distributed clouds. Execution engines such as BigQuery and Managed Service for Apache Spark run vectorized queries and in-place AI functions over remote data. This transforms distributed enterprise data estates into a single, unified system of action without forcing teams to migrate physical files.
Multi-cloud catalog federation in the borderless Lakehouse leverages the open Apache Iceberg REST catalog standard to provide secure, bi-directional metadata discovery and execution across AWS Glue, Databricks Unity Catalog, and Snowflake Horizon Catalog. Instead of copying underlying datasets, query engines read table schemas and manifest files remotely over secure endpoints. This allows BigQuery, Managed Service for Apache Spark, and open-source engines to query multi-cloud tables in place while maintaining unified governance across disparate object stores.
Zero-copy SaaS integration extends the borderless Lakehouse directly to the enterprise application layer, enabling BigQuery to run real-time queries against live transactional systems like SAP, Salesforce, and Workday without constructing ETL pipelines. By querying data where it natively resides, organizations eliminate pipeline latency and eliminate data sync failures. This direct connectivity also allows BigQuery AI functions to execute directly on live ERP, CRM, and HR data, streamlining cross-domain enterprise analytics.
Transactional databases are integrated into the borderless Lakehouse through Spanner Omni, which operates in multi-cloud environments outside Google Cloud, and Lakehouse Federation for AlloyDB. This bidirectional bridge allows operational systems to query analytical lakehouse data directly without waiting for batch syncs. As a result, developers and data engineers can join live transactional records with petabytes of historical analytical data in real time, powering low-latency operational applications and real-time dashboards.
The borderless Lakehouse connects autonomous reasoning agents to Knowledge Catalog, an always-on context engine that provides precise business metadata, lineage, and semantic definitions without moving source data. By feeding precise, pre-filtered business context into LLM prompts, the architecture can help prevent model hallucinations and eliminates unnecessary agent reasoning loops. This semantic grounding improves answer accuracy for autonomous agents like Gemini Enterprise while reducing token consumption.
The borderless Lakehouse lowers total cost of ownership (TCO) and provides predictability by using Cross-Cloud Interconnects with flat-rate pricing that eliminates variable data egress costs for AWS data. To maximize query performance across distributed clouds, the system pairs intelligent cross-cloud local caching with BigQuery’s vectorized execution engine, delivering sub-second response times. By replacing duplicate storage and fragile ETL compute, organizations significantly reduce cloud infrastructure spend and operational overhead.