Project reference ↗

Some Python and machine-learning applications need work spread across many machines rather than one process. Ray provides that distributed execution system; KubeRay manages its clusters, jobs and services on Kubernetes. It creates and maintains the coordinating process and the workers that do the computation. Teams already using Ray can use it to bring those workloads into their cluster, while still planning the resources, startup time and recovery behavior the application requires.

Choose Ray for a workload requirement

A straightforward single-model endpoint may need fewer components than a distributed Ray service. Identify the reason for Ray, then pin compatible operator, Ray and runtime images. Separate the resources needed by the head and worker processes and account for data movement and startup time.

Test the failure model

Exercise worker loss, a head restart, unavailable model storage and partial capacity. Confirm how the service handles queued or running work during replacement. Ensure autoscaling requests can actually be satisfied by the cluster's node and GPU capacity.

The inference guide compares platform fit, while GPU scheduling covers the capacity layer. This entry describes upstream operator use; it does not report a distributed inference benchmark.

Sources & further reading

  1. Official KubeRay installation

Spotted something that needs another look?

Help improve this page →