Teams running repeated Spark jobs may want Kubernetes to manage their definitions and report execution status. Kubeflow Spark Operator watches application resources and submits the corresponding Spark work, with support for scheduled applications and lifecycle handling. It adds that management layer around Spark rather than replacing Spark’s processing engine. Controller retries can repeat work, so applications need suitable output and recovery behavior, and the operator’s version compatibility must be checked separately from the Spark runtime image.
Current guidance
The recovered upstream identity corresponds to the Kubeflow Spark Operator, which provides SparkApplication and scheduled application resources and runs spark-submit on their behalf. Current upstream documentation provides its own Helm repository, describes the project as beta and identifies the v1beta2 API. This is a specific operator lineage, not every project called a Spark operator.
Review the selected controller’s version matrix and application image separately. Controller retries and automatic resubmission can repeat external writes, so applications need an appropriate output and recovery design. Namespace watching, driver service accounts, admission webhook behavior and CRD ownership also affect a shared cluster beyond a single Spark job.
Before replacing an old installation, export custom resources and inspect stored API versions. Do not casually delete CRDs while following a historical migration example: custom-resource deletion can remove the specifications needed for recovery. Rehearse conversion and controller handover in isolation, then test scheduled execution, failure retry and cleanup with an idempotent job. Retain streaming checkpoints and application dependencies explicitly; upgrading the operator alone does not prove that a different Spark major version can resume them.
Historical upstream link check · 2026-10-09
The recorded upstream address redirects to https://github.com/kubeflow/spark-operator/blob/master/README.md and returned HTTP 200 on 2026-10-09. GitHub does not mark kubeflow/spark-operator archived or disabled; this does not establish active maintenance, support or compatibility. GitHub resolves the old repository identity to kubeflow/spark-operator. Link availability does not certify the historical installation instructions or current security support.
Website availability is separate from project, chart and image support. Use the current guidance and primary sources on this page to assess the distribution.
Historical Kubedex content
Preserved for context. Commands, versions, prices and results below reflect the original research.
Spark Operator aims to make specifying and running Spark applications as easy and idiomatic as running other workloads on Kubernetes. It uses Kubernetes custom resources for specifying, running, and surfacing status of Spark applications.
Spark Operator currently supports the following list of features:
- Supports Spark 2.3 and up.
- Enables declarative application specification and management of applications through custom resources.
- Automatically runs spark-submit on behalf of users for each SparkApplication eligible for submission.
- Provides native cron support for running scheduled applications.
- Supports customization of Spark pods beyond what Spark natively is able to go through the mutating admission webhook, e.g., mounting ConfigMaps and volumes, and setting pod affinity/anti-affinity.
- Supports automatic application re-submission for updated SparkAppliation objects with the updated specification.
- Supports automatic application restart with a configurable restart policy.
- Supports automatic retries of failed submissions with optional linear back-off.
- Supports mounting local Hadoop configuration as a Kubernetes ConfigMap automatically via sparkctl.
- Supports automatically staging local application dependencies to Google Cloud Storage (GCS) via sparkctl.
- Supports collecting and exporting application-level metrics and driver/executor metrics to Prometheus.
Sources & further reading
Spotted something that needs another look?
Help improve this page →