Skip to content

Agent retries forever when its tenant names a namespace the server does not have #7064

Description

@geovannewashington

Found while evaluating v0.27.0-rc.9. Not a regression: rc.8 behaves the same.

What happens

An agent holding a tenant ID the server cannot resolve retries /api/devices/auth forever. The device never appears, and nothing says why.

On rc.9 the refusal was logged on every attempt:

level=warning msg="failed to authenticate device" data="{\"message\":\"namespace not found\"}" status_code=404 tenant_id=20a1c530-f6e5-4837-ad2d-16a0cb96f117
level=warning msg="retrying request after a random time period" retry_after=19s

On master it is logged once. The retry hook reports through connectivity.Refused, which logs at level(attempt), and severity drops as attempts climb, so later refusals fall below the configured level. SetRetryCount(math.MaxInt32) is unchanged, so the loop is the same. Only the evidence is gone. A wedged agent now looks like an idle one.

Mechanism

agent/pkg/agentd/agent.go:187 adopts the tenant persisted at <private-key>.tenant when the environment supplies none. That satisfies HasNamespaceCredential(), so the pairing branch at agent/main.go:104 never runs and the agent goes straight to Authorize().

Server side, authDevice resolves the namespace and fails:

namespace, err := s.store.NamespaceResolve(ctx, store.NamespaceTenantIDResolver, req.TenantID)
if err != nil {
    return nil, NewErrNamespaceNotFound(req.TenantID, err)
}

ErrCodeNotFound maps to 404, and 404 is retried on purpose: a namespace may not exist yet. The server answers the same 404 for "not created yet" and "deleted forever", so the agent cannot distinguish a wait from a dead end.

The installer reports success for every path that can wedge

enroll_agent_interactively in install.sh has four branches:

if [ -n "$CODE" ];        then echo "...pre-authorized, accepted automatically..."; return 0; fi
if [ -n "$INSTALL_KEY" ]; then echo "...will enroll into the install key's namespace..."; return 0; fi
if [ -n "$TENANT_ID" ];   then echo "...will appear as pending, accept it there."; return 0; fi

$_AGENT_CMD login || true

Three print a prediction and return. The fourth runs the login flow in the foreground and waits on what actually happens. Since the container is started with docker run -d --restart=unless-stopped, the script exits before the agent has attempted anything at all.

So the only enrollment path that reports honestly is the one that cannot silently fail. Install with a tenant naming a namespace that does not exist and the script says the device will appear in the console. Install with an install key that does not exist and it says the device will enroll into the key's namespace. Both are already doomed when printed, and the only way to find out is docker logs on a container the operator was never told to inspect.

Scope

This issue was originally filed as a persisted-tenant problem. It is wider than that, and narrower in one respect.

An agent ends up holding an unresolvable tenant two ways:

  1. From <private-key>.tenant, written only by pairing (all four PersistTenant calls are in waitForPairing and pairingLogin).
  2. From SHELLHUB_TENANT_ID, the plain TENANT_ID= install. Wedges identically.

An install key alone does not wedge. installKeyTenant returns NewErrAuthInvalid → ErrCodeInvalid → 400, which is not in the retry set, so the agent exits fatally instead. An install key alongside TENANT_ID does wedge, because the req.TenantID == "" && req.InstallKey != "" guard skips key resolution.

So the infinite loop is a tenant-ID problem affecting every such enrollment. What is pairing-specific is only the invisibility, and that part is severe: the operator never typed the value, nothing displays it, and the fix is deleting a file they have no reason to know exists.

A case this has already caused

A mistyped tenant ID (O for 0) on a TENANT_ID= install. Permanently invalid, retried forever, no diagnostic. TenantID carries no validate tag:

PrivateKey string `env:"PRIVATE_KEY,required" validate:"required"`
TenantID   string `env:"TENANT_ID"`

while validate:"required,uuid" is the standard for tenant IDs elsewhere in the repo, and LoadConfigFromEnv already runs validator.New().StructWithFields(cfg) over this struct. A typo is catchable at startup, before any network call.

Diagnostics that actively mislead

  • enrollment_summary() in install.sh checks only CODE, INSTALL_KEY and TENANT_ID. It never reads the tenant file, so it prints Enrollment: none on exactly the affected machine. It fires on uninstall too, since the uninstall branch sits at the end of main() after the installer preamble.
  • The TENANT_ID comments at install.sh:167 and install.sh:255 claim an empty value overrides a persisted tenant. The [ -n "${TENANT_ID}" ] guard makes that impossible; absent and empty both reach agent.go:187 as "".
  • docker_uninstall and podman_uninstall say the private key was left in place and say nothing about the .tenant sibling next to it. Both are on the host bind mount, so removing the container never touches either. A reinstall reads the old tenant.
  • An invalid install key exits fatally with a generic Failed to initialize agent and the whole configuration struct in the log fields, never naming the key as the problem.

Prerequisite: provenance

agent.go:187 assigns the persisted value into the same field as the environment value, so nothing downstream can tell where a tenant came from. Any fix that clears or self-heals a tenant needs that distinction, or a wrong SHELLHUB_SERVER_ADDRESS would destroy a valid operator-supplied enrollment.

Correction to the original proposal

Item 1 of this issue is stale. 7ead6ff9f already split terminal from transient in pkg/api/client/client_public.go, and the retry set is now deliberate: 404, 402 and 403, the refusals an operator resolves without touching the device. The wedge survives that change because 404 is on the list by design.

Note on device identity

The device UID is sha256 over hostname, MAC, public key and tenant ID. A changed tenant produces a different device, so the tenant is not only where a device enrolls, it is part of who the device is. Worth keeping in view for any option that re-pairs into a different namespace.

What to do

Tier 1, no design decision required.

  • Give TenantID a uuid validate tag. Rejects the typo class at startup with a named field.
  • Make the three early-returning branches of enroll_agent_interactively observe the agent instead of predicting, or at minimum tell the operator how to check.
  • Restore a periodic log line for a refusal that keeps repeating, so a wedged agent is distinguishable from an idle one.
  • Make enrollment_summary() read the tenant file.
  • Fix the two install.sh comments.
  • Have uninstall name the .tenant file alongside the key.
  • On a rejection, log the tenant and the path it came from, and name an invalid install key as such.

Tier 2, design-neutral, prerequisite for anything else.

  • Track whether a tenant came from the environment or the file.
  • Bound the 404 retry instead of leaving it unbounded. This is the quarantine mechanism without committing to a policy.

Tier 3, the open decision.

What happens when the bound is reached, and where that policy lives. One proposal is that the namespace records it and the device receives it at acceptance time; note that the policy cannot be fetched at failure time, because the failure is not being able to reach the namespace, so it has to be written to the device up front.

The two deployment shapes pull differently here. An unattended fleet install has nobody to accept a re-pairing, and a typo'd tenant has no recovery at all, so fail-loudly is required for that shape regardless. A device with a person in front of it can genuinely re-pair, which is also the only shape where re-pairing can land the device in a namespace it never belonged to. That narrows the open question to attended devices.

Out of scope

install.sh uninstall re-detects INSTALL_METHOD rather than recording how the agent was installed, so a Docker install uninstalled on a host where Docker is no longer reachable falls through to standalone and leaves the container running. Separate issue.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Labels

No labels
No labels

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions