Project reference ↗

Apache Hadoop provides several components for storing and processing large datasets across a group of machines. HDFS supplies distributed file storage, YARN allocates processing resources, and MapReduce runs jobs using a particular distributed processing model. It is useful when applications depend on that ecosystem rather than a single standalone database. A deployment should include only the components actually required, with their metadata and persistent data managed independently of the containers running them.

Deployment and operating notes

Apache’s current documentation covers Hadoop Common, HDFS, YARN and MapReduce with separate configuration, security and upgrade guidance. The deprecated stable/hadoop chart was intended for YARN and MapReduce jobs, with HDFS transporting small artifacts and application data stored in external cloud datastores. Its README explicitly says the chart is no longer supported. Running those processes as pods does not automatically preserve Hadoop’s assumptions about host identity, disks, rack awareness or service availability.

First identify which components are actually needed. A workload using object storage and Kubernetes-native compute may not require a new HDFS and YARN installation, while an existing Hadoop application may depend on their APIs. Inventory NameNode metadata, DataNode storage, ResourceManager state, Kerberos configuration and job dependencies before changing packaging. Rehearse recovery of metadata and representative job execution with the exact Hadoop and Java versions. Verify permissions and data locality rather than relying on pod readiness. Preserve a supported rollback path for data-format upgrades, and avoid attaching production storage to an unreviewed chart. No current vendor support or production-scale benchmark is implied by this directory entry.

Historical upstream link check · 2026-10-09

The recorded upstream address responded successfully (HTTP 200) on 2026-10-09. GitHub confirms that helm/charts is archived: this is a historical chart distribution, not evidence that the application itself is retired. Link availability does not certify the historical installation instructions or current security support.

Source for this check ↗

Website availability is separate from project, chart and image support. Use the current guidance and primary sources on this page to assess the distribution.

The original record

Historical Kubedex content

Preserved for context. Commands, versions, prices and results below reflect the original research.

Hadoop is a framework for running large scale distributed applications. The Hadoop project develops open-source software for reliable, scalable, distributed computing. The Apache Hadoop software library is a framework that allows for the distributed processing of large data sets across clusters of computers using simple programming models.

It is designed to scale up from single servers to thousands of machines, each offering local computation and storage. Rather than rely on hardware to deliver high-availability, the library itself is designed to detect and handle failures at the application layer, so delivering a highly-available service on top of a cluster of computers, each of which may be prone to failures.

 

Prerequisites

Supported Platforms

  • GNU/Linux is supported as a development and production platform. Hadoop has been demonstrated on GNU/Linux clusters with 2000 nodes.
  • Windows is also a supported platform but the followings steps are for Linux only. To set up Hadoop on Windows, see wiki page.

Required software for Linux include:

  • Java™ must be installed. Recommended Java versions are described at HadoopJavaVersions.
  • ssh must be installed and sshd must be running to use the Hadoop scripts that manage remote Hadoop daemons.

 

Modules

  • Hadoop Common: The common utilities that support the other Hadoop modules.
  • Hadoop Distributed File System (HDFS™): A distributed file system that provides high-throughput access to application data.
  • Hadoop YARN: A framework for job scheduling and cluster resource management.
  • Hadoop MapReduce: A YARN-based system for parallel processing of large data sets.

 

Who Uses Hadoop?

A wide variety of companies and organizations use Hadoop for both research and production. Users are encouraged to add themselves to the Hadoop PoweredBy wiki page.

 

Components of Hadoop

The core components in the first iteration of Hadoop were MapReduce, the Hadoop Distributed File System (HDFS) and Hadoop Common, a set of shared utilities and libraries. As its name indicates, MapReduce uses map and reduce functions to split processing jobs into multiple tasks that run at the cluster nodes where data is stored and then to combine what the tasks produce into a coherent set of results. MapReduce initially functioned as both Hadoop’s processing engine and cluster resource manager, which tied HDFS directly to it and limited users to running MapReduce batch applications.

 

Hadoop applications and use cases

Hadoop is primarily geared to analytics uses, and its ability to process and store different types of data makes it a particularly good fit for big data analytics applications. Big data environments typically involve not only large amounts of data, but also various kinds, from structured transaction data to semistructured and unstructured forms of information, such as internet clickstream records, web server and mobile application logs, social media posts, customer emails and sensor data from the internet of things (IoT).

Sources & further reading

  1. Apache Hadoop current documentation
  2. Recovered historical source (Common Crawl index)
  3. Archived Hadoop chart scope and deprecation

Spotted something that needs another look?

Help improve this page →