Argo CD deploys applications by comparing the configuration stored in Git with the resources running in Kubernetes. When those differ, a sync applies the desired changes. This saves repeated manual deployment work, but it also means a fix made directly in the cluster can be overwritten unless the Git configuration is corrected too.
As more teams and clusters depend on Argo CD, the practical questions become: which application will a change affect, what may the controller delete, and how can the on-call engineer stop a bad rollout? This guide explains those controls and walks through investigation and recovery for an existing installation.
An Application connects a source of configuration to a destination cluster and namespace. An ApplicationSet generates several Applications from a shared template. An AppProject limits permitted sources, destinations and resource types. The sections below build on those three objects. If you are choosing a tool first, read Argo CD vs Flux.
Commands below are diagnostic suggestions, not a deployment or recovery test executed by Kubedex.
Find the Git configuration that controls each application
For each service, record its source repository and revision, chart or manifest path, destination cluster and namespace, AppProject, and owning team. Also record whether the Application is maintained directly, generated by an ApplicationSet, or managed by a parent Application. During an incident, a change to the wrong level may disappear at the next reconciliation.
Keep the deployment artifact immutable across promotion. Record the image digest and chart version alongside the environment configuration. Argo CD uses Helm as a renderer, so a successful Helm CLI rollback elsewhere is not evidence that this Application can recover through the same operation. See the documented Helm lifecycle.
Separate synchronization, drift repair and deletion
Automatic sync applies desired changes. selfHeal repairs live drift; prune removes resources missing from the desired manifests; allowEmpty permits an empty desired application with pruning. Review each separately. Enabling one is not a reason to enable all four. The automated sync documentation also explains retry behavior. Its restriction on rollback applies to Argo CD's history rollback operation; reverting the desired manifests in Git can be delivered through normal synchronization. Choose the recovery path explicitly.
Use a disposable application to rehearse three changes before broad enablement: remove an ordinary Deployment, remove a protected storage resource, and accidentally render no resources. Write down the expected result for each. These are proposed acceptance checks, not results observed by Kubedex.
Deletion through pruning and deletion of the Application are different paths. Argo CD provides options such as Prune=confirm and Delete=false for different purposes. Scope exceptions narrowly and record why they exist. Replace=true or forceful recreation can cause an outage; neither is a general cure for an unexplained sync error. Review the exact sync-option semantics before using them.
Control the reach of ApplicationSet changes
An ApplicationSet is useful when several Applications share a template and differ by a cluster, directory or parameter. Before merging a generator or template change, compare the generated application names, destinations, projects and source revisions. Count creations and deletions, not just changed YAML lines. Test a representative cluster cohort before extending a change to the fleet.
A child's auto-sync setting can be restored by its ApplicationSet. Put the pause at the owning declaration, or use a deliberately configured exception supported by your installed version. ApplicationSet modification policies and resource preservation are separate controls: create-update does not by itself prevent owner-reference deletion when the ApplicationSet is deleted. Verify parent deletion, finalizers and child-resource behavior together. The resource modification guide describes these distinctions.
Keep the authority to change projects or fleet-wide generators narrower than the authority to change one service. Review the source and destination permissions in AppProjects, including cluster-scoped kinds and access to the Argo CD namespace. A reviewed template can still be dangerous if untrusted input controls its project or destination.
Triage the failure before requesting another sync
First capture the desired revision, the last operation result and the affected resource. Distinguish these states:
- Manifest generation failed: examine repository authentication, chart retrieval, values, rendering and plugin errors. Kubernetes has not necessarily received a new desired object.
- Apply failed: inspect the rejected resource for admission, permission, schema or immutable-field errors.
- Applied but unhealthy: inspect workload readiness, scheduling, image pulls, dependencies and application behavior.
- Repeated drift: identify another field owner, nondeterministic rendering or a mutating controller before suppressing the difference.
For an Application stored in the Argo CD namespace, these read-only commands collect a starting point. Replace the context, namespace and Application name; Applications in other namespaces require that actual namespace. Review output before sharing it because repository identifiers and error messages may contain operational details.
ops_context=your-reviewed-context
argo_namespace=argocd
application=your-application
kubectl --context "$ops_context" -n "$argo_namespace" get applications.argoproj.io "$application" -o wide
kubectl --context "$ops_context" -n "$argo_namespace" describe applications.argoproj.io "$application"
kubectl --context "$ops_context" -n "$argo_namespace" get events --field-selector "involvedObject.name=$application" --sort-by=.metadata.creationTimestamp
Inspect the workload namespace separately once you know which resource failed. Avoid exporting every Secret or every manifest merely to find one error. If an HPA legitimately owns replicas, document that ownership and scope a diff exception to that field. Ignoring a displayed difference is not automatically the same as ignoring it during sync; review RespectIgnoreDifferences in the sync options.
Use hooks and waves for order, then verify service behavior
Use sync waves to order resources within an Application, such as establishing a controller before its custom resources. Identical wave numbers in independent Applications do not create a shared dependency graph. A failing early wave can stop later work, so give that dependency an understandable health signal. Hooks need bounded execution and an explicit retention or cleanup policy. Hooks do not run during selective sync, which matters if a migration or check normally lives in a hook. See sync phases and waves.
For an app-of-apps arrangement, check how the parent observes child Application health before relying on waves to gate the next child. Argo CD removed its built-in Application health assessment in 1.8; the health documentation explains restoring a custom assessment where needed. Creating a child Application object is not proof that the child's service is ready.
Keep schema changes backward compatible where possible. A green Application does not establish that a checkout, login or other user transaction works. Run an application-level check against the routed service and compare its error rate and latency after the rollout.
Fix the controlling configuration before resuming deployment
- Contain: identify the Application's parent and stop automatic delivery at the controlling declaration. For an in-progress operation, separately assess whether to terminate it; changing future auto-sync policy does not undo writes already made.
- Preserve evidence: record the failed revision, resource error and any data-changing hooks before replacing them.
- Choose recovery: restore a known-good desired revision only if the application remains compatible with current data. Otherwise deploy the forward fix or follow the service's data-recovery plan.
- Reconcile narrowly: review the diff and deletion list for the affected Application or cohort before resuming.
- Verify: check the running image and configuration, service transaction and user-facing signals. Confirm the owning Git declaration now describes the recovered state.
Scale the component that is constrained
Measure reconciliation duration, failed syncs, Git/rendering latency, controller resource use and Kubernetes API pressure. A slow UI, a busy repo-server and a saturated application controller are different problems. Argo CD metrics help distinguish them.
Repository rendering consumes memory and temporary storage, and monorepo changes can trigger many applications at once. More concurrency can increase memory pressure. Controller sharding also depends on the installed release and configuration. Start with the measured bottleneck, change one limit or replica setting, and repeat a representative reconciliation burst. Use the high-availability and scaling guide; this article does not prescribe a universal applications-per-replica ratio.
Rehearse loss of the control plane
Git is only part of recovery. Inventory cluster access, repository credentials, projects, controller configuration and any settings held outside Git. Protect backups as credential-bearing material. Follow the version-matched export/import procedure and verify the intended namespace; upstream warns that an export from the wrong namespace may not fail.
Restore into an isolated environment and verify what would reconcile before granting access to production destinations. A recovery rehearsal should end with a known service available, its desired state restored, and a record of what was missing from the backup. That is the evidence required before calling this an exercised recovery plan.
Sources & further reading
- Argo CD Helm ownership
- Argo CD automated synchronization
- Argo CD synchronization options
- ApplicationSet modification and deletion controls
- Argo CD project permissions
- Argo CD synchronization phases and waves
- Argo CD operational metrics
- Argo CD high availability and scaling
- Argo CD disaster recovery
- Argo CD history rollback command
- Argo CD Application and custom-resource health assessment
Spotted something that needs another look?
Help improve this page →