When an application slows down or fails, you need its measurements and events to reach the tools your team uses to investigate it. In Kubernetes, that often means collecting metrics from many services, reading logs from nodes and receiving traces that follow requests across applications.

Prometheus collects and stores numerical measurements and evaluates alerting rules. OpenTelemetry provides standards and software for producing and moving telemetry; its Collector receives, processes and forwards that data. Grafana Alloy is a collection agent that combines Prometheus and OpenTelemetry components with supported log-collection features. Grafana dashboards display data from configured backends; a collector still needs somewhere to send and store it.

Keep Prometheus when its metrics, queries and alerts already serve your team. Choose the OpenTelemetry Collector when you need its supported processing and export pipeline. Evaluate Alloy when its combination of Prometheus, OpenTelemetry and log components fits the data you collect. They can work together, but two agents reading the same input can create duplicates. Start by assigning one collection path to each signal.

The guide then shows how to change that path safely and presents a recorded one-log Collector test. That small test does not certify an Alloy migration or production capacity.

What each component does

Responsibilities in a Kubernetes observability stack
ComponentUse it forDecide separately
Application instrumentationProduce request context, metrics and traces using the application's supported SDK or instrumentation.Service naming, sensitive attributes, sampling and the overhead your application can tolerate.
Prometheus metrics pathA full Prometheus server discovers and scrapes targets, stores metrics, evaluates rules and serves queries. Preserve the rules and queries your team operates.Scrape ownership, retention and whether a remote backend receives another copy.
OpenTelemetry CollectorReceive, process and export supported signals using components included in the selected distribution.Agent or gateway placement, permissions, buffering and destination support.
Grafana AlloyBuild collection pipelines with Prometheus and OpenTelemetry components, plus supported log collection.Configuration compatibility, component-specific clustering and persistent state.
Grafana and telemetry backendsQuery and present stored telemetry through configured data sources.The actual metrics, log and trace storage, access control, retention and recovery owners.

Alloy and a separate Collector do not both have to sit in every path. If the existing component handles the required processing and export, adding another hop creates another queue and failure boundary. Keep a second hop when it has a clear purpose, such as central filtering or routing across backends. See the signal comparison for choosing which data answers the operational question.

Assign one owner to each input

Write down the producer, receiver, processors, exporter and destination for every signal. A node agent can read local container files; a shared gateway can accept application OTLP traffic. Cluster-wide watchers need an explicit ownership or leader strategy. An agent cannot read another node's local files merely because both Pods share a namespace.

A practical starting layout is existing Prometheus scraping for application metrics, one node-local log reader per node, and an OTLP gateway for traces. It is a design example, not a required topology. If you move metrics scraping into the Collector or Alloy, retire the previous owner after the comparison period rather than leaving both paths active indefinitely.

The Collector scaling guidance warns that identical Prometheus receivers scrape the same targets. Split the target set or use the Operator's Target Allocator to assign it. The Target Allocator's Prometheus resource discovery still needs the appropriate CRDs and access; installing a Collector alone does not establish ownership of existing ServiceMonitors.

For Alloy, prometheus.scrape clustering requires both Alloy clustered mode and an opted-in scrape component. Peers must converge on the same target set and ownership labels. Turning on a component's clustering block without running Alloy in clustered mode does not distribute the work. Check target ownership while adding or removing replicas.

Keep useful labels without creating millions of metric series

Choose one place to enrich data with cluster, namespace, workload and service identity. Compare these attributes before and after a collector change: a renamed service can appear as a new application even when requests still succeed. Test the queries, recording rules and alerts that depend on those names, including error rate and latency calculations.

A useful review starts with the labels that change on every request. User IDs, raw URLs, session tokens and arbitrary request identifiers rarely belong in metric labels. Keep only dimensions needed for an operational question, and examine whether they belong in traces or logs with appropriate access controls instead. Measure active series and ingestion rate during the canary; a lower collector CPU reading does not show whether backend cost increased.

Migrate Promtail to Alloy in a bounded slice

Use Grafana's Promtail conversion procedure to generate and inspect a candidate configuration. Treat conversion warnings as work to resolve. Bypassing conversion errors can change behavior. The generated configuration also needs review if Kubernetes Pod discovery will run outside a DaemonSet.

Plan the log-reader state as well as the configuration. Alloy uses a different positions-file location, and its own monitoring metric names can differ. Check persistent volume paths and update collector alerts. In a test slice, verify where reading resumes after a restart, whether multiline entries retain their boundaries, and whether timestamps and labels match the queries operators use.

During comparison, send the candidate's logs to an isolated test destination or otherwise keep the two paths distinguishable. Reading the same files into the same tenant through both agents can duplicate records. Once the target slice passes, switch its collection owner and watch for gaps before moving another slice.

Diagnose delivery from the producer to the backend

  1. Producer: generate a unique synthetic event and record when and where it was emitted. Confirm the application is using the expected endpoint and protocol.
  2. Receiver: inspect listener configuration, service routing, network policy and authentication. A component declared in configuration must also be connected to a service pipeline.
  3. Processing: check filters, sampling and resource changes. Confirm the chosen Collector image includes every configured component.
  4. Export: inspect rejected requests, queue occupancy, retry and drop signals. A ready Pod can still fail to deliver data.
  5. Storage and query: check the destination tenant, timestamp range and resource filters. Verify a known count rather than assuming an empty dashboard proves an empty pipeline.

Define outage and rollback behavior

For each signal, choose the tolerable loss, maximum delay and buffer budget. In-memory queues do not survive every restart; persistence adds a storage and recovery task. Retry and batching cannot absorb an unlimited outage. Alert on sustained export failures through a path that does not depend solely on that exporter.

Before increasing replicas, identify the bottleneck and any stateful processing. Scraping requires target assignment; trace-processing decisions may require related spans to reach the same processing instance. A general CPU autoscaler does not solve those routing requirements. Scale the responsible stage and repeat the end-to-end check.

Keep the last working configuration and record how to restore input ownership. Rollback can restore future collection without recovering data already dropped. Rehearse a bounded backend rejection, a collector restart and recovery into a test destination. Capture accepted counts, missing or duplicate records, identity changes and queue recovery before calling the change successful.

A tested local OTLP log pipeline

Kubedex installed the official Collector chart 0.175.1 with core Collector 0.161.0, Helm 4.3.0 and Kubernetes 1.37.0 in a single Linux arm64 kind node on 9 October 2026. The test sent one synthetic OTLP/HTTP log and verified the resource attribute and body in the debug exporter's output.

Download the values, synthetic payload, instructions and results. The README includes the tested chart archive checksum and cleanup procedure. The values pin the observed image digest, select the core otelcol executable, disable unused receivers and pipelines, and retain the health-check extension used by the chart's probes. This is an isolated diagnostic fixture; its detailed exporter must not receive sensitive data.

With an explicit disposable kubeconfig and context, the central installation commands were:

fixture_config=/absolute/path/to/disposable-kubeconfig
fixture_context=your-disposable-context
helm4=/absolute/path/to/helm-v4.3.0
fixture_namespace=kubedex-otel-fixture
k=(kubectl --kubeconfig "$fixture_config" --context "$fixture_context")
h=(--kubeconfig "$fixture_config" --kube-context "$fixture_context" --namespace "$fixture_namespace")
"${k[@]}" get nodes
"${k[@]}" create namespace "$fixture_namespace"
"$helm4" pull opentelemetry-collector --version 0.175.1 --repo https://open-telemetry.github.io/opentelemetry-helm-charts
"$helm4" install otel-demo ./opentelemetry-collector-0.175.1.tgz "${h[@]}" --values values.yaml --wait --timeout 180s

The Helm installation returned a deployed revision. A localhost-only port-forward exposed the OTLP/HTTP endpoint for the synthetic request:

"${k[@]}" -n "$fixture_namespace" port-forward deployment/otel-demo-opentelemetry-collector 14318:4318 --address 127.0.0.1
# Run the request in a second terminal while port-forward remains open:
curl --fail --header 'Content-Type: application/json' --data-binary @log-payload.json http://127.0.0.1:14318/v1/logs

Observed HTTP status: 200. The response body was {"partialSuccess":{}}. The captured debug output included:

     -> service.name: Str(kubedex-fixture)
LogRecord #0
Body: Str(kubedex-otel-fixture-20261009)

The automated check found one occurrence of that body in the captured output. This verifies configuration and delivery for one input; it does not prove exactly-once delivery during retries, persistence across restarts, production throughput, backend authentication or outage recovery. Those require the separate failure and capacity experiments above.

Stop the port-forward and remove only the fixture namespace when finished. The recorded result includes the installed image identity and successful cleanup. The upstream chart archive is retrieved from its maintainer repository, not bundled into the download.

Sources & further reading

  1. OpenTelemetry Collector chart
  2. OpenTelemetry Kubernetes collection
  3. Collector scaling and scrape ownership
  4. OpenTelemetry Target Allocator design and Prometheus discovery
  5. Alloy Prometheus scrape clustering
  6. Promtail to Alloy migration
  7. Alloy migration paths
  8. Prometheus architecture and responsibilities

Spotted something that needs another look?

Help improve this page →