Data exploration often benefits from keeping code, query results and explanatory text in one document. Apache Zeppelin provides browser-based notebooks with interpreters for supported languages and processing systems, including Spark. Teams can use those notebooks to investigate data and share the steps behind an analysis. Each interpreter runs code with access to particular resources, so notebook permissions and data credentials matter. On Kubernetes, the server and its interpreter pods also need controlled resource limits and cleanup.
Current guidance
This resource identifies Apache Zeppelin, the notebook and interpreter platform; the recovered incubation description is historical. Current Apache documentation describes a Kubernetes server pod that creates interpreter pods, with the Spark interpreter using Spark on Kubernetes in client mode. Choose a compatible Zeppelin, interpreter and Spark image set rather than reading old minimum-version examples as a current support matrix.
Notebook paragraphs can execute code through several interpreters. Review authentication, notebook permissions, interpreter sharing and the Kubernetes permissions used to create pods. Credentials for databases and object storage should not be broadly available to every notebook author. Plan limits and cleanup for interpreter and executor pods so abandoned sessions do not consume the cluster indefinitely.
For migration, export notebooks, interpreter settings, dependencies and required configuration, then re-run representative paragraphs in an isolated environment. Check JDBC drivers, Spark session behavior, visualization outputs and saved data paths. Preserve notebook storage and external datasets independently from ephemeral interpreter pods. A successful server upgrade does not establish that custom interpreters or notebooks using older language/runtime APIs remain compatible, so retain a controlled reference environment for important analytical workflows.
Historical upstream link check · 2026-10-09
The recorded upstream address responded successfully (HTTP 200) on 2026-10-09. Link availability does not certify the historical installation instructions or current security support.
Website availability is separate from project, chart and image support. Use the current guidance and primary sources on this page to assess the distribution.
Historical Kubedex content
Original publication: 2018-09-20T13:23:29+00:00. Preserved for context. Commands, versions, prices and results below reflect the original research.
Zeppelin, a web-based notebook that enables interactive data analytics. You can make beautiful data-driven, interactive and collaborative documents with SQL, Scala and more. A completely open web-based notebook that enables interactive data analytics
Apache Zeppelin is a new and incubating multi-purposed web-based notebook which brings data ingestion, data exploration, visualization, sharing and collaboration features to Hadoop and Spark.
Core feature:
- Web-based notebook style editor.
- Built-in Apache Spark support
WHAT IS APACHE ZEPPELIN?
Multi-purpose Notebook
The Notebook is the place for all your needs
- Data Ingestion
- Data Discovery
- Data Analytics
- Data Visualization & Collaboration
Multiple Language Backend
Apache Zeppelin interpreter concept allows any language/data-processing-backend to be plugged into Zeppelin. Currently, Apache Zeppelin supports many interpreters such as Apache Spark, Python, JDBC, Markdown and Shell.
Apache Spark integration
Especially, Apache Zeppelin provides built-in Apache Spark integration. You don’t need to build a separate module, plugin or library for it.
Apache Zeppelin with Spark integration provides:
- Automatic SparkContext and SQLContext injection
- Runtime jar dependency loading from local filesystem or maven repository. Learn more about dependency loader.
- Canceling job and displaying its progress
Data visualization
Some basic charts are already included in Apache Zeppelin. Visualizations are not limited to SparkSQL query, any output from any language backend can be recognized and visualized.
Pivot chart
Apache Zeppelin aggregates values and displays them in pivot chart with simple drag and drop. You can easily create a chart with multiple aggregated values including sum, count, average, min, max.
Dynamic forms
Apache Zeppelin can dynamically create some input forms in your notebook.
Collaborate by sharing your Notebook & Paragraph
Your notebook URL can be shared among collaborators. Then Apache Zeppelin will broadcast any changes in real-time, just like the collaboration in Google Docs. Apache Zeppelin provides an URL to display the result only, that page does not include any menus and buttons inside of notebooks. You can easily embed it as an iframe inside of your website in this way.
100% Opensource
Apache Zeppelin is Apache2 Licensed software. Please check out the source repository and how to contribute. Apache Zeppelin has a very active development community. Join to our Mailing list and report issues on Jira Issue tracker.
What Zeppelin Does
Interactive browser-based notebooks enable data engineers, data analysts and data scientists to be more productive by developing, organizing, executing, and sharing data code and visualizing results without referring to the command line or needing the cluster details. Notebooks allow these users not only allow to execute but to interactively work with long workflows. There are a number of notebooks available with Spark. iPython remains a mature choice and great example of a data science notebook. The Hortonworks Gallery provides an Ambari stack definition to help our customers quickly set up iPython on their Hadoop clusters.
Apache Zeppelin is a new and upcoming web-based notebook which brings data exploration, visualization, sharing and collaboration features to Spark. It supports Python, but also a growing list of programming languages such as Scala, Hive, SparkSQL, shell, and markdown.
Data discovery, exploration, reporting, and visualization are key components of the data science workflow. Zeppelin provides a “Modern Data Science Studio” that supports Spark and Hive out of the box. Actually, Zeppelin supports multiple language backends which has support for a growing ecosystem of data sources. Zeppelin’s notebooks provide interactive snippet-at-time experience to data scientist. You can see a collection of Zeppelin notebooks in the Hortonworks Gallery.
Also when you are done with your notebook and found some insight you want to share, you can easily create a report out of it and either print it or send it out.
The post Zeppelin appeared first on kubedex.com.
Sources & further reading
Spotted something that needs another look?
Help improve this page →