Skip to content

Round-robin sharding: adding a cluster can leave other clusters owned by no application-controller replica (silent, until restart) #29476

Description

@SimonSQLML

Checklist:

  • I've searched in the docs and FAQ for my answer: https://bit.ly/argocd-faq.
  • I've included steps to reproduce the bug.
  • I've pasted the output of argocd version.

Describe the bug

With controller.sharding.algorithm: round-robin and dynamicClusterDistribution: true, registering a new cluster can leave other clusters owned by no application-controller replica. Applications on the orphaned clusters silently stop reconciling — with no error, no Application condition, and no change to their reported health/sync — until the application-controller is restarted.

In our case this lasted 57 hours across 313 Applications before we noticed, because the frozen Applications keep reporting their last-known Synced/Healthy status. argocd app wait --health --sync in CI therefore returns success while nothing is deployed.

Round-robin assigns shard = index-in-sorted-cluster-list % replicas (RoundRobinDistributionFunction -> createClusterIndexByClusterIdMap -> getSortedClustersList, controller/sharding/sharding.go), and the list is sorted by Cluster.ID, which is the cluster Secret's Kubernetes UID (ID: string(s.UID) in SecretToCluster, util/db/cluster.go). A newly created cluster Secret therefore lands at an effectively random position in that ordering, and every cluster sorting after it shifts by one index — changing the shard of roughly half the fleet on average.

The re-shuffle was then applied partially by one replica: it ended up simultaneously holding clusters that had moved away from it and missing clusters that had moved onto it. Two clusters were left owned by nobody.

To Reproduce

  1. Run the application-controller as a Deployment with multiple replicas (we run 10) and:
    • ARGOCD_ENABLE_DYNAMIC_CLUSTER_DISTRIBUTION=true
    • controller.sharding.algorithm: round-robin
    • a non-trivial number of clusters (we have 48) and Applications (~3,300)
  2. Let the controllers run for several days so their in-memory distribution is settled (no pod restarts).
  3. Register one new cluster. Because the ordering key is the Secret UID, the new cluster lands at an arbitrary index and every cluster after it shifts. (In our case the count went 47 → 48; the new Secret's UID placed it at index 21 of 48, so 26 of the 47 existing clusters changed shard.)
  4. Observe: one or more previously-healthy clusters are no longer processed by any replica.

Signals that identify the state:

  • argocd cluster list shows the affected clusters with empty connection state, applicationsCount: 0, and no cacheInfo.lastCacheSyncTime, while every other cluster is Successful with a fresh cache.
  • Their Applications have status.reconciledAt frozen at the moment of the re-shuffle; status.health / status.sync are frozen at their last values, and status.conditions is empty.
  • The affected cluster names appear nowhere in any controller pod's logs (we grepped ~41,000 lines over 5 minutes across all 10 pods; a cluster with 304 Applications should be continuously present).
  • One replica logs The cluster <server> has no assigned shard. continuously (~2.8 lines/second in our case).
  • Comparing dest-name= in each pod's logs against argocd-app-controller-shard-cm shows two clusters processed by two replicas simultaneously, and the two orphans processed by none.
  • All shard entries in argocd-app-controller-shard-cm were heartbeating normally the entire time, and no pod had restarted (pods were 4–5 days old).

What the affected replica actually ended up with

Reconstructing the distribution from the cluster Secret UIDs (this model reproduces the post-restart ownership we observed in the logs for 14 of 14 spot-checked clusters), the shard that owned our two orphans should have changed membership like this:

members of that shard
before the add in-cluster, stg-1, stg-2, stg-3, stg-4
after the add in-cluster, stg-4, dev-1, cluster-A, cluster-B (bold = moved in)

What the replica owning that shard was actually processing, 57 h later (derived from dest-name= in its logs):

in-cluster, stg-4 (unchanged — fine), dev-1 (moved in — correctly picked up), stg-1, stg-3 (moved away — still being processed here, and simultaneously by the replica that now owns them), and cluster-A, cluster-B were absent (moved in — never picked up).

So the update was applied per-cluster and partially: one of the three incoming clusters was picked up, two were not, and two outgoing clusters were never released. This does not look like a wholesale stale map on that replica.

Expected behavior

After a cluster is added or removed, every cluster is owned by exactly one replica. If a replica's view of the distribution becomes inconsistent, it should either self-heal (e.g. on the next resync) or surface the condition — an Application whose destination cluster is owned by no shard should not keep reporting a stale Synced/Healthy status indefinitely.

Recovery

kubectl -n argocd rollout restart deploy/argocd-application-controller fixed it within 60 seconds: both clusters reconnected, caches initialized, all 313 Applications reconciled within 5 minutes, and the has no assigned shard log spam stopped. Nothing else (cluster credentials, network, RBAC, shard ConfigMap) was touched.

Version

argocd: v3.4.5+564b949
argocd-server: v3.4.4

Deployed via the argo-cd Helm chart 9.7.1. Relevant settings:

controller:
  replicas: 10
  dynamicClusterDistribution: true
configs:
  params:
    controller.sharding.algorithm: "round-robin"
    controller.status.processors: "30"
    controller.operation.processors: "15"
    timeout.reconciliation: "360s"

Logs

From the replica holding the inconsistent view (repeating continuously, ~2.8/s):

level=warning msg="The cluster https://cluster-A.example:6443 has no assigned shard."

The two orphaned clusters produced no log lines at all on any replica.

Application status of an affected app, ~57 h after the re-shuffle (the desired image had been updated 3 h earlier, and the pipeline reported success):

status:
  reconciledAt: "2026-08-29T10:40:36Z"   # frozen at the moment of the re-shuffle
  health:   { status: Healthy }          # stale
  sync:     { status: Synced }           # stale
  conditions: []                         # no error surfaced
  summary:
    images: ["registry.example/app:OLD"] # spec requested NEW three hours earlier

Additional context

Verified against upstream source before filing:

  • controller/sharding/cache.go (which holds the cluster→shard map and emits has no assigned shard) is byte-identical at v3.4.4, v3.4.8 and v3.5.2; master differs only by a comment typo.
  • controller/cache/cache.go is the same git blob dc66e786 at v3.4.4 and v3.5.2. In handleModEvent, cluster, ok := c.clusters[newCluster.Server] is followed by if ok { ... } with no else branch, so a cluster that becomes newly owned by this replica and has no existing cache entry is not handled there. Unchanged on master today.
  • RoundRobinDistributionFunction(clusters, replicas) takes no apps argument — application counts are an input only to ConsistentHashingWithBoundedLoadsDistributionFunction.
  • getSortedClustersList sorts by Cluster.ID; SecretToCluster sets ID: string(s.UID). That is why creating a single cluster Secret re-indexes an arbitrary fraction of the fleet rather than appending at the end.

Relation to existing reports and PRs:

I'm willing to work on a fix and tests if maintainers can indicate the preferred layer (sharding cache vs. live-state cache).

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

bugSomething isn't workingbug/priority:urgentBug should be fixed immediatelybug/severity:criticialA critical bug in ArgoCD, possibly resulting in data loss or severe degraded overall functionalitycomponent:cluster-managementIssue related to multi-clusters managementcomponent:coreIssues on core functionalities such as tracking, reconciling, managing resources, etc.

Type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions