Create an Apache Spark machine learning pipeline

Accelerate your distributed machine learning journey by creating a Spark ML pipeline with zero-setup runtimes on Google Cloud.

What is distributed machine learning?

Distributed machine learning allows data scientists and ML engineers to scale machine learning (ML) model training beyond the memory and compute limitations of a single machine. By distributing data partitions and parallelizing workloads across a cluster of worker nodes, you can train complex models on massive datasets significantly faster.

Why use Apache Spark?

Apache Spark is the industry standard for distributed data processing. Its native machine learning library, MLlib, provides distributed algorithms for classification, regression, clustering, and collaborative filtering, allowing you to move from data ingestion to model deployment within a single, unified environment.

What is a Spark ML pipeline?

A Spark ML pipeline is a high-level API designed to help you combine multiple machine learning workflows into a cohesive pipeline. It allows you to chain together various data transformations (such as feature extraction and scaling) and machine learning algorithms (such as model training) within a structured, declarative workflow.


By using a pipeline, you ensure that the same sequence of data preprocessing steps is applied consistently during both model training and prediction, which is crucial for preventing data leakage and ensuring reproducible results. The key components of a Spark ML pipeline are transformers and estimators.

Why choose Google Cloud for your Spark ML journey?

Scaling ML workloads to apply hardware acceleration often introduces the challenge of managing complex Spark GPU dependencies. Provisioning VM instances, installing the correct NVIDIA drivers, and ensuring CUDA and cuDNN versions match your PyTorch or TensorFlow libraries can halt developer productivity for days.


Google Cloud solves this with Managed Service for Apache Spark, which offers pre-packaged machine learning runtimes. Built on stable, Ubuntu-based images (starting from version 2.3), these runtimes come pre-installed and pre-configured with the exact GPU drivers (CUDA, cuDNN, NCCL) and industry-standard ML frameworks (PyTorch, XGBoost, tokenizers, and transformers) your workloads require. This zero-setup environment lets your team focus on writing code rather than configuring infrastructure.

Frequently asked questions

In Spark MLlib, a transformer is an algorithm that transforms one DataFrame into another (for example, converting text into numerical vectors). An estimator is an algorithm that can be fit on a DataFrame to produce a transformer (for example, fitting a logistic regression algorithm on training data to produce a trained model).

By relying on a Spark ML pipeline, you ensure that your data preparation logic is strictly coupled with your model. This guarantees that testing or inference data is processed identically to your training data.

With Google Cloud ML runtimes on Managed Service for Apache Spark, dependency management is offloaded. The custom Ubuntu-based images are continually updated by Google and are pre-packaged with the correct, natively compatible versions of NVIDIA drivers, CUDA toolkits, and distributed ML frameworks.

Components of Spark MLlib and Google Cloud data tools

Understanding the building blocks of Spark MLlib and Google Cloud data tools is essential for executing distributed ML at scale.

Component

Description

Primary use case


Pricing considerations

Spark MLlib

Apache Spark’s scalable machine learning library containing common learning algorithms, featurization tools, and pipeline utilities.

Executing massive-scale ML model training across a distributed cluster without relying on single-node libraries.


Costs are incurred through the underlying Google Cloud services used to run Apache Spark, primarily Managed Service for Apache Spark. Pricing depends on the Compute Engine resources (vCPUs, memory, GPUs, disk) provisioned and the duration the job is running.


*See Managed Service for Apache Spark pricing.

Vector assembler

A core feature transformer in Spark MLlib that merges multiple columns (continuous numbers, one-hot encoded categories, etc.) into a single vector column.

Preparing raw tabular data into the specific dense or sparse vector format expected by Spark MLlib machine learning algorithms.

As a component within Spark MLlib, its execution contributes to the overall compute costs of running your Spark workloads on Managed Service for Apache Spark.

Spark ML pipeline

A high-level API built on top of DataFrames that helps users chain multiple feature engineering steps and models together.

Grouping Transformers and Estimators into a unified pipeline to ensure identical data processing during training and inference.

Similar to other Spark MLlib components, costs are part of the execution charges for Managed Service for Apache Spark.

BigQuery ML


A service that lets you create and execute machine learning models in BigQuery using standard SQL queries.

Training models directly where your data lives before moving highly complex, iterative distributed jobs to Spark.


BigQuery ML has its own pricing model. You are charged for

  1. Model training (based on data processed during CREATE MODEL queries),
  2. Model prediction (standard BigQuery query pricing)
  3. Model storage (billed at standard BigQuery storage rates).

Component

Description

Primary use case


Pricing considerations

Spark MLlib

Apache Spark’s scalable machine learning library containing common learning algorithms, featurization tools, and pipeline utilities.

Executing massive-scale ML model training across a distributed cluster without relying on single-node libraries.


Costs are incurred through the underlying Google Cloud services used to run Apache Spark, primarily Managed Service for Apache Spark. Pricing depends on the Compute Engine resources (vCPUs, memory, GPUs, disk) provisioned and the duration the job is running.


*See Managed Service for Apache Spark pricing.

Vector assembler

A core feature transformer in Spark MLlib that merges multiple columns (continuous numbers, one-hot encoded categories, etc.) into a single vector column.

Preparing raw tabular data into the specific dense or sparse vector format expected by Spark MLlib machine learning algorithms.

As a component within Spark MLlib, its execution contributes to the overall compute costs of running your Spark workloads on Managed Service for Apache Spark.

Spark ML pipeline

A high-level API built on top of DataFrames that helps users chain multiple feature engineering steps and models together.

Grouping Transformers and Estimators into a unified pipeline to ensure identical data processing during training and inference.

Similar to other Spark MLlib components, costs are part of the execution charges for Managed Service for Apache Spark.

BigQuery ML


A service that lets you create and execute machine learning models in BigQuery using standard SQL queries.

Training models directly where your data lives before moving highly complex, iterative distributed jobs to Spark.


BigQuery ML has its own pricing model. You are charged for

  1. Model training (based on data processed during CREATE MODEL queries),
  2. Model prediction (standard BigQuery query pricing)
  3. Model storage (billed at standard BigQuery storage rates).

How it works

Google Cloud's Managed Service for Apache Spark abstracts away infrastructure complexities, allowing you to execute your Spark ML pipeline for large-scale model training using either serverless environments or customizable managed clusters.


To eliminate difficult Spark GPU dependencies, the service utilizes pre-packaged ML runtimes fully equipped with natively compatible NVIDIA drivers, CUDA toolkits, and distributed frameworks, giving you a zero-setup path to deploy hardware-accelerated workloads.

Learn more about machine learning with Spark

Common use cases

Build a resilient feature engineering workflow

Learn how to use transformers to extract and scale data, and estimators to train algorithms.

How-to guides

  • Building a Spark ML pipeline: For general guidance on using Spark on Google Cloud, refer to the Managed Service for Apache Spark documentation.
  • Pipeline estimators in PySpark: Check the Python API samples provided within the Managed Service for Apache Spark documentation. These samples include patterns for using Spark MLlib components. You can also find Spark ML examples referenced in guides like the attach GPUs to clusters documentation.

Additional resources

Deploy zero-setup ML runtimes

Launch GPU-accelerated Spark clusters without managing underlying infrastructure. Google Cloud's ML runtimes come pre-configured so you can bypass complex Spark GPU dependencies.

How-to guides

  • Managed Service for Apache Spark ML runtimes overview: Read more about Managed Service for Apache Spark runtime image versions, which detail pre-installed libraries, including TensorFlow, PyTorch, and XGBoost, along with GPU-specific libraries.


Additional resources

  • Distributed deep learning on Google Cloud: For a broader architectural perspective on scaling ML, check out the best practices for implementing machine learning documentation. This document provides recommendations for developing custom-trained models, including distributed training and accelerator usage on Google Cloud, integrated with the Gemini platform Model Registry.

Take the next step

Start building on Google Cloud with $300 in free credits and 20+ always free products.

Google Cloud