HDFS is Hadoop's distributed filesystem. It splits stored files across data servers while a NameNode tracks the filesystem's metadata, allowing Hadoop applications to work with data spread across machines. It is useful when those applications require HDFS's storage interface and behavior. Recovering the system needs both the file blocks and the metadata that describes them; keeping a collection of worker disks is not a complete filesystem backup.
Deployment and operating notes
The recorded apache-spark-on-k8s/kubernetes-HDFS repository contains Helm charts and integration-test instructions. Its name does not make it the current Apache Hadoop deployment authority, and this review did not establish a supported matrix for modern Kubernetes and Hadoop releases. HDFS itself remains documented within Apache Hadoop rather than being a separate top-level Hadoop project.
For an existing installation, inventory NameNode namespace metadata, edit logs, DataNode identities, replication, rack mapping, security and persistent volumes. Protect metadata backups as carefully as data blocks; keeping DataNode disks alone is not a complete recovery plan. Rehearse a supported NameNode recovery and representative client reads before changing chart ownership. Determine whether workloads truly require HDFS semantics or can use another storage system, then evaluate the application changes explicitly. Do not assume object storage is a drop-in filesystem replacement or that new PVCs will discover old blocks. Keep the old cluster recoverable through a controlled data-copy and validation window.
Historical upstream link check · 2026-10-09
The recorded upstream address responded successfully (HTTP 200) on 2026-10-09. GitHub does not mark apache-spark-on-k8s/kubernetes-HDFS archived or disabled; this does not establish active maintenance, support or compatibility. Link availability does not certify the historical installation instructions or current security support.
Website availability is separate from project, chart and image support. Use the current guidance and primary sources on this page to assess the distribution.
Historical Kubedex content
Original publication: 2018-09-22T08:51:43+00:00. Preserved for context. Commands, versions, prices and results below reflect the original research.
HDFS is the storage system of the Hadoop framework. It is a distributed file system that can conveniently run on commodity hardware for processing unstructured data. It is highly fault-tolerant. The same data is stored in multiple locations and in the event that one storage location fails the same data can be easily fetched from another location. It began under the Apache Nutch project but today is a top level Apache Hadoop project. HDFS is a major constituent of Hadoop along with Hadoop YARN, Hadoop MapReduce and Hadoop Common.
HDFS key features
| HDFS key features | Description |
| Stores bulks of data | Capable of storing terabytes and petabytes of data |
| Minimum intervention | HDFS manages thousands of nodes without operator intervention |
| Computing | Benefits of distributed and parallel computing at once |
| Scaling out | It works on scaling out rather than scaling up without single downtime |
| Rollback | Allows returning to its previous version post an upgrade |
| Data integrity | Deals with corrupted data by replicating it several times |
The servers in HDFS are fully connected and communicate through TCP-based protocols. Though designed for huge databases, normal file systems (FAT, NTFS) can also be viewed.
The post HDFS appeared first on kubedex.com.
Sources & further reading
Spotted something that needs another look?
Help improve this page →