Found while evaluating v0.27.0-rc.9. Not a regression: rc.8 behaves the same.
What happens
An agent holding a tenant ID the server cannot resolve retries /api/devices/auth forever. The device never appears, and nothing says why.
On rc.9 the refusal was logged on every attempt:
level=warning msg="failed to authenticate device" data="{\"message\":\"namespace not found\"}" status_code=404 tenant_id=20a1c530-f6e5-4837-ad2d-16a0cb96f117
level=warning msg="retrying request after a random time period" retry_after=19s
On master it is logged once. The retry hook reports through connectivity.Refused, which logs at level(attempt), and severity drops as attempts climb, so later refusals fall below the configured level. SetRetryCount(math.MaxInt32) is unchanged, so the loop is the same. Only the evidence is gone. A wedged agent now looks like an idle one.
Mechanism
agent/pkg/agentd/agent.go:187 adopts the tenant persisted at <private-key>.tenant when the environment supplies none. That satisfies HasNamespaceCredential(), so the pairing branch at agent/main.go:104 never runs and the agent goes straight to Authorize().
Server side, authDevice resolves the namespace and fails:
namespace, err := s.store.NamespaceResolve(ctx, store.NamespaceTenantIDResolver, req.TenantID)
if err != nil {
return nil, NewErrNamespaceNotFound(req.TenantID, err)
}
ErrCodeNotFound maps to 404, and 404 is retried on purpose: a namespace may not exist yet. The server answers the same 404 for "not created yet" and "deleted forever", so the agent cannot distinguish a wait from a dead end.
The installer reports success for every path that can wedge
enroll_agent_interactively in install.sh has four branches:
if [ -n "$CODE" ]; then echo "...pre-authorized, accepted automatically..."; return 0; fi
if [ -n "$INSTALL_KEY" ]; then echo "...will enroll into the install key's namespace..."; return 0; fi
if [ -n "$TENANT_ID" ]; then echo "...will appear as pending, accept it there."; return 0; fi
$_AGENT_CMD login || true
Three print a prediction and return. The fourth runs the login flow in the foreground and waits on what actually happens. Since the container is started with docker run -d --restart=unless-stopped, the script exits before the agent has attempted anything at all.
So the only enrollment path that reports honestly is the one that cannot silently fail. Install with a tenant naming a namespace that does not exist and the script says the device will appear in the console. Install with an install key that does not exist and it says the device will enroll into the key's namespace. Both are already doomed when printed, and the only way to find out is docker logs on a container the operator was never told to inspect.
Scope
This issue was originally filed as a persisted-tenant problem. It is wider than that, and narrower in one respect.
An agent ends up holding an unresolvable tenant two ways:
- From
<private-key>.tenant, written only by pairing (all four PersistTenant calls are in waitForPairing and pairingLogin).
- From
SHELLHUB_TENANT_ID, the plain TENANT_ID= install. Wedges identically.
An install key alone does not wedge. installKeyTenant returns NewErrAuthInvalid → ErrCodeInvalid → 400, which is not in the retry set, so the agent exits fatally instead. An install key alongside TENANT_ID does wedge, because the req.TenantID == "" && req.InstallKey != "" guard skips key resolution.
So the infinite loop is a tenant-ID problem affecting every such enrollment. What is pairing-specific is only the invisibility, and that part is severe: the operator never typed the value, nothing displays it, and the fix is deleting a file they have no reason to know exists.
A case this has already caused
A mistyped tenant ID (O for 0) on a TENANT_ID= install. Permanently invalid, retried forever, no diagnostic. TenantID carries no validate tag:
PrivateKey string `env:"PRIVATE_KEY,required" validate:"required"`
TenantID string `env:"TENANT_ID"`
while validate:"required,uuid" is the standard for tenant IDs elsewhere in the repo, and LoadConfigFromEnv already runs validator.New().StructWithFields(cfg) over this struct. A typo is catchable at startup, before any network call.
Diagnostics that actively mislead
enrollment_summary() in install.sh checks only CODE, INSTALL_KEY and TENANT_ID. It never reads the tenant file, so it prints Enrollment: none on exactly the affected machine. It fires on uninstall too, since the uninstall branch sits at the end of main() after the installer preamble.
- The
TENANT_ID comments at install.sh:167 and install.sh:255 claim an empty value overrides a persisted tenant. The [ -n "${TENANT_ID}" ] guard makes that impossible; absent and empty both reach agent.go:187 as "".
docker_uninstall and podman_uninstall say the private key was left in place and say nothing about the .tenant sibling next to it. Both are on the host bind mount, so removing the container never touches either. A reinstall reads the old tenant.
- An invalid install key exits fatally with a generic
Failed to initialize agent and the whole configuration struct in the log fields, never naming the key as the problem.
Prerequisite: provenance
agent.go:187 assigns the persisted value into the same field as the environment value, so nothing downstream can tell where a tenant came from. Any fix that clears or self-heals a tenant needs that distinction, or a wrong SHELLHUB_SERVER_ADDRESS would destroy a valid operator-supplied enrollment.
Correction to the original proposal
Item 1 of this issue is stale. 7ead6ff9f already split terminal from transient in pkg/api/client/client_public.go, and the retry set is now deliberate: 404, 402 and 403, the refusals an operator resolves without touching the device. The wedge survives that change because 404 is on the list by design.
Note on device identity
The device UID is sha256 over hostname, MAC, public key and tenant ID. A changed tenant produces a different device, so the tenant is not only where a device enrolls, it is part of who the device is. Worth keeping in view for any option that re-pairs into a different namespace.
What to do
Tier 1, no design decision required.
- Give
TenantID a uuid validate tag. Rejects the typo class at startup with a named field.
- Make the three early-returning branches of
enroll_agent_interactively observe the agent instead of predicting, or at minimum tell the operator how to check.
- Restore a periodic log line for a refusal that keeps repeating, so a wedged agent is distinguishable from an idle one.
- Make
enrollment_summary() read the tenant file.
- Fix the two
install.sh comments.
- Have uninstall name the
.tenant file alongside the key.
- On a rejection, log the tenant and the path it came from, and name an invalid install key as such.
Tier 2, design-neutral, prerequisite for anything else.
- Track whether a tenant came from the environment or the file.
- Bound the 404 retry instead of leaving it unbounded. This is the quarantine mechanism without committing to a policy.
Tier 3, the open decision.
What happens when the bound is reached, and where that policy lives. One proposal is that the namespace records it and the device receives it at acceptance time; note that the policy cannot be fetched at failure time, because the failure is not being able to reach the namespace, so it has to be written to the device up front.
The two deployment shapes pull differently here. An unattended fleet install has nobody to accept a re-pairing, and a typo'd tenant has no recovery at all, so fail-loudly is required for that shape regardless. A device with a person in front of it can genuinely re-pair, which is also the only shape where re-pairing can land the device in a namespace it never belonged to. That narrows the open question to attended devices.
Out of scope
install.sh uninstall re-detects INSTALL_METHOD rather than recording how the agent was installed, so a Docker install uninstalled on a host where Docker is no longer reachable falls through to standalone and leaves the container running. Separate issue.
Found while evaluating v0.27.0-rc.9. Not a regression: rc.8 behaves the same.
What happens
An agent holding a tenant ID the server cannot resolve retries
/api/devices/authforever. The device never appears, and nothing says why.On rc.9 the refusal was logged on every attempt:
On master it is logged once. The retry hook reports through
connectivity.Refused, which logs atlevel(attempt), and severity drops as attempts climb, so later refusals fall below the configured level.SetRetryCount(math.MaxInt32)is unchanged, so the loop is the same. Only the evidence is gone. A wedged agent now looks like an idle one.Mechanism
agent/pkg/agentd/agent.go:187adopts the tenant persisted at<private-key>.tenantwhen the environment supplies none. That satisfiesHasNamespaceCredential(), so the pairing branch atagent/main.go:104never runs and the agent goes straight toAuthorize().Server side,
authDeviceresolves the namespace and fails:ErrCodeNotFoundmaps to 404, and 404 is retried on purpose: a namespace may not exist yet. The server answers the same 404 for "not created yet" and "deleted forever", so the agent cannot distinguish a wait from a dead end.The installer reports success for every path that can wedge
enroll_agent_interactivelyininstall.shhas four branches:Three print a prediction and return. The fourth runs the login flow in the foreground and waits on what actually happens. Since the container is started with
docker run -d --restart=unless-stopped, the script exits before the agent has attempted anything at all.So the only enrollment path that reports honestly is the one that cannot silently fail. Install with a tenant naming a namespace that does not exist and the script says the device will appear in the console. Install with an install key that does not exist and it says the device will enroll into the key's namespace. Both are already doomed when printed, and the only way to find out is
docker logson a container the operator was never told to inspect.Scope
This issue was originally filed as a persisted-tenant problem. It is wider than that, and narrower in one respect.
An agent ends up holding an unresolvable tenant two ways:
<private-key>.tenant, written only by pairing (all fourPersistTenantcalls are inwaitForPairingandpairingLogin).SHELLHUB_TENANT_ID, the plainTENANT_ID=install. Wedges identically.An install key alone does not wedge.
installKeyTenantreturnsNewErrAuthInvalid→ErrCodeInvalid→ 400, which is not in the retry set, so the agent exits fatally instead. An install key alongsideTENANT_IDdoes wedge, because thereq.TenantID == "" && req.InstallKey != ""guard skips key resolution.So the infinite loop is a tenant-ID problem affecting every such enrollment. What is pairing-specific is only the invisibility, and that part is severe: the operator never typed the value, nothing displays it, and the fix is deleting a file they have no reason to know exists.
A case this has already caused
A mistyped tenant ID (
Ofor0) on aTENANT_ID=install. Permanently invalid, retried forever, no diagnostic.TenantIDcarries no validate tag:while
validate:"required,uuid"is the standard for tenant IDs elsewhere in the repo, andLoadConfigFromEnvalready runsvalidator.New().StructWithFields(cfg)over this struct. A typo is catchable at startup, before any network call.Diagnostics that actively mislead
enrollment_summary()ininstall.shchecks onlyCODE,INSTALL_KEYandTENANT_ID. It never reads the tenant file, so it printsEnrollment: noneon exactly the affected machine. It fires on uninstall too, since theuninstallbranch sits at the end ofmain()after the installer preamble.TENANT_IDcomments atinstall.sh:167andinstall.sh:255claim an empty value overrides a persisted tenant. The[ -n "${TENANT_ID}" ]guard makes that impossible; absent and empty both reachagent.go:187as"".docker_uninstallandpodman_uninstallsay the private key was left in place and say nothing about the.tenantsibling next to it. Both are on the host bind mount, so removing the container never touches either. A reinstall reads the old tenant.Failed to initialize agentand the wholeconfigurationstruct in the log fields, never naming the key as the problem.Prerequisite: provenance
agent.go:187assigns the persisted value into the same field as the environment value, so nothing downstream can tell where a tenant came from. Any fix that clears or self-heals a tenant needs that distinction, or a wrongSHELLHUB_SERVER_ADDRESSwould destroy a valid operator-supplied enrollment.Correction to the original proposal
Item 1 of this issue is stale.
7ead6ff9falready split terminal from transient inpkg/api/client/client_public.go, and the retry set is now deliberate: 404, 402 and 403, the refusals an operator resolves without touching the device. The wedge survives that change because 404 is on the list by design.Note on device identity
The device UID is
sha256over hostname, MAC, public key and tenant ID. A changed tenant produces a different device, so the tenant is not only where a device enrolls, it is part of who the device is. Worth keeping in view for any option that re-pairs into a different namespace.What to do
Tier 1, no design decision required.
TenantIDauuidvalidate tag. Rejects the typo class at startup with a named field.enroll_agent_interactivelyobserve the agent instead of predicting, or at minimum tell the operator how to check.enrollment_summary()read the tenant file.install.shcomments..tenantfile alongside the key.Tier 2, design-neutral, prerequisite for anything else.
Tier 3, the open decision.
What happens when the bound is reached, and where that policy lives. One proposal is that the namespace records it and the device receives it at acceptance time; note that the policy cannot be fetched at failure time, because the failure is not being able to reach the namespace, so it has to be written to the device up front.
The two deployment shapes pull differently here. An unattended fleet install has nobody to accept a re-pairing, and a typo'd tenant has no recovery at all, so fail-loudly is required for that shape regardless. A device with a person in front of it can genuinely re-pair, which is also the only shape where re-pairing can land the device in a namespace it never belonged to. That narrows the open question to attended devices.
Out of scope
install.sh uninstallre-detectsINSTALL_METHODrather than recording how the agent was installed, so a Docker install uninstalled on a host where Docker is no longer reachable falls through tostandaloneand leaves the container running. Separate issue.