Skip to content
Home/ Engineering Big Data Pipelines with Hadoop and Spark: Store, Process, and Analyze Large-Scale Datase
Engineering Big Data Pipelines with Hadoop and Spark: Store, Process, and Analyze Large-Scale Datase

Engineering Big Data Pipelines with Hadoop and Spark: Store, Process, and Analyze Large-Scale Datase

No customer reviews yet ISBN 9798192589724

Learn how to build practical big data pipelines with Hadoop, Apache Spark, and PySpark.

As datasets grow beyond the practical limits of a single computer, organizations need reliable ways to store, process, clean, transform, and analyze information at scale. Engineering Big Data Pipelines with Hadoop and Spark provides a practical introduction to the tools and workflows used to solve these problems.

This beginner friendly guide takes you from the foundations of big data and distributed computing to complete processing pipelines using Hadoop, Spark, and Python. You will learn not only what the technologies do, but how their individual components fit together in a repeatable data engineering workflow.

Inside this book, you will learn how to:

  • Understand big data, distributed computing, batch processing, and streaming
  • Understand the Hadoop ecosystem and how its major components work together
  • Work with HDFS for distributed storage
  • Understand MapReduce and parallel batch processing
  • Use YARN for cluster resource management
  • Set up Hadoop, Spark, and PySpark on a practical local environment
  • Create and work with PySpark applications
  • Understand RDDs, transformations, actions, and Spark execution
  • Work with Spark DataFrames for large scale data processing
  • Query large datasets using Spark SQL
  • Read and clean CSV, JSON, and Parquet datasets
  • Transform large datasets using PySpark
  • Understand partitions, shuffling, caching, persistence, and broadcast joins
  • Read Spark execution plans and use the Spark Web UI for performance analysis
  • Process continuously arriving data with Structured Streaming
  • Build introductory machine learning workflows using Spark MLlib
  • Develop classification and regression models with Spark
  • Save and reload machine learning pipelines
  • Build and monitor a complete end to end big data processing project

The book emphasizes a practical workflow rather than memorizing commands. You will learn to move data through a complete pipeline: store it, inspect it, clean it, transform it, analyze it, optimize the processing job, and save the results.

Practical examples use familiar business data including sales transactions, customer orders, product records, website logs, incoming events, and customer behavior data. The final project combines the major skills developed throughout the book by moving raw transaction data into HDFS, processing it with PySpark, analyzing it with Spark SQL, optimizing the job, and saving the results as Parquet.

You do not need previous experience with Hadoop, Spark, distributed systems, or data engineering. Basic computer skills are enough to begin, while some familiarity with Python and SQL can be helpful as you progress into PySpark and Spark SQL.

Whether you are a student learning data processing, a Python user working with larger datasets, an analyst exploring distributed computing, a software developer moving into data engineering, or a technical professional preparing to work with Hadoop and Spark, this book provides a structured path from fundamentals to practical implementation.

Build the skills to design, process, optimize, and understand large scale data pipelines with Hadoop, Spark, and PySpark.

About the author

Product details

Pub dateAug 14, 2026
ISBN-109798192589724
ISBN-139798192589724
LanguageEnglish
Last updated 2026-08-17 13:13
$16.76 $18.99 11% off
You save $2.23 · list price $18.99
In stock — ships in 24 hours with free tracking
Delivery by Monday, September 14, 2026
Qty
Sign in to Add to Saved list
Free delivery on orders over $35.
15-day returns. Any reason.
Secure checkout. We never store card details.