Distributed TensorFlow uses several workers to train a machine-learning model when the work benefits from more computing resources than one machine provides. The training strategy determines how those workers share data, calculations and model updates. Kubernetes can provide places to run the workers, but the training code must use an appropriate distribution model. This entry records an old TensorFlow chart, so its deployment cannot substitute for choosing a supported modern training strategy.
Deployment and operating notes
The archived distributed-tensorflow chart records TensorFlow 1.7.0. Current TensorFlow guidance describes distributed strategies such as mirrored, multi-worker and parameter-server training. These are application-level execution models, so replacing the old image with a current TensorFlow image does not automatically migrate training code or cluster configuration.
Choose a supported training strategy first, then define worker discovery, task roles, dataset sharding and checkpoint storage. Align TensorFlow, Python, accelerator drivers and communication-library versions across all workers. Kubernetes scheduling must provide the required devices and network connectivity, while the training framework controls collective progress and recovery. Rehearse a short run with known expected results, checkpoint restart and one worker failure before allocating a large cluster. Preserve model artifacts and input-data versions outside ephemeral pods. If using a training operator, verify that its job API matches the framework and failure policy. The historical discussion of algorithms is background, not a measured performance comparison or a validated modern GPU deployment.
Historical upstream link check · 2026-10-09
The recorded upstream address responded successfully (HTTP 200) on 2026-10-09. Link availability does not certify the historical installation instructions or current security support.
Website availability is separate from project, chart and image support. Use the current guidance and primary sources on this page to assess the distribution.
Historical Kubedex content
Preserved for context. Commands, versions, prices and results below reflect the original research.
TensorFlow™ is an open source software library for high-performance numerical computation. Its flexible architecture allows easy deployment of computation across a variety of platforms (CPUs, GPUs, TPUs), and from desktops to clusters of servers to mobile and edge devices. Originally developed by researchers and engineers from the Google Brain team within Google’s AI organization, it comes with strong support for machine learning and deep learning and the flexible numerical computation core is used across many other scientific domains.
Parameter server architecture
When parallel SGD uses parameter servers, the algorithm starts by broadcasting the model to the workers (devices). In each training iteration, each worker reads its own split from the mini-batch, computing its own gradients, and sending those gradients to one or more parameter servers. The parameter servers aggregate all the gradients from the workers and wait until all workers have completed before they calculate the new model for the next iteration, which is then broadcast to all workers.
Ring-allreduce architecture
In the ring-allreduce architecture, there is no central server that aggregates gradients from workers. Instead, in a training iteration, each worker reads its own split for a mini-batch, calculates its gradients, sends its gradients to its successor neighbor on the ring, and receives gradients from its predecessor neighbor on the ring. For a ring with N workers, all workers will have received the gradients necessary to calculate the updated model after N-1 gradient messages are sent and received by each worker.
Ring-allreduce is bandwidth optimal, as it ensures that the available upload and download network bandwidth at each host is fully utilized (in contrast to the parameter server model). Ring-allreduce can also overlap the computation of gradients at lower layers in a deep neural network with the transmission of gradients at higher layers, further reducing training time.
Glossary
Client
A client is typically a program that builds a TensorFlow graph and constructs a tensorflow::Session to interact with a cluster. Clients are typically written in Python or C++. A single client process can directly interact with multiple TensorFlow servers (see “Replicated training” above), and a single server can serve multiple clients.
Cluster
A TensorFlow cluster comprises a one or more “jobs”, each divided into lists of one or more “tasks”. A cluster is typically dedicated to a particular high-level objective, such as training a neural network, using many machines in parallel. A cluster is defined by a tf.train.ClusterSpec object.
Job
A job comprises a list of “tasks”, which typically serve a common purpose. For example, a job named ps (for “parameter server”) typically hosts nodes that store and update variables; while a job named worker typically hosts stateless nodes that perform compute-intensive tasks. The tasks in a job typically run on different machines. The set of job roles is flexible: for example, a worker may maintain some state.
Master service
An RPC service that provides remote access to a set of distributed devices, and acts as a session target. The master service implements the tensorflow::Session interface, and is responsible for coordinating work across one or more “worker services”. All TensorFlow servers implement the master service.
Sources & further reading
Spotted something that needs another look?
Help improve this page →