Troubleshoot CloudGrange

Diagnostic steps for the on-premises deployment. The control plane runs as a Helm release on K3s (or, on installs still running the legacy engine, as a Compose container stack) on a single management host; the runtime agent is a Windows service on each managed host.

CloudGrange is in active development and is not GA. See current product and release status.

Control plane (management host) — K3s/Helm engine

# Pod status across the release
sudo k3s kubectl get pods

# Helm release status and history
sudo helm status cloudgrange
sudo helm history cloudgrange

# Logs for a specific service
sudo k3s kubectl logs deploy/cloudgrange-api --tail 100
sudo k3s kubectl logs deploy/cloudgrange-api --tail 100 --previous   # after a crash/restart

# Describe a pod that's not Ready — shows scheduling/probe/image-pull failures
sudo k3s kubectl describe pod -l app.kubernetes.io/name=cloudgrange-api

# Certificate not issuing (cert-manager)
sudo k3s kubectl get clusterissuer,certificate
sudo k3s kubectl describe certificate cloudgrange-tls

# Relay not reachable from managed hosts — is MetalLB assigning it an address?
sudo k3s kubectl get svc cloudgrange-relay

sudo k3s kubectl requires root because the K3s installer generates a root-only kubeconfig by default; copy /etc/rancher/k3s/k3s.yaml to ~/.kube/config (and fix server: if accessing remotely) to run kubectl/helm as a non-root operator instead.

Portal or API not reachable

  • Confirm the Traefik ingress has an address: k3s kubectl get ingress — a blank ADDRESS column usually means no node has a routable IP, or (on a multi-node bare-metal install) MetalLB isn't configured yet.
  • Check the API pod is Ready (k3s kubectl get pods -l app.kubernetes.io/name=cloudgrange-api) — the API exposes /health/live, /health/ready, /health/startup.
  • The install generates a self-signed certificate via cert-manager's default ClusterIssuer; browser warnings are expected until you replace it with a CA-signed certificate or a customer-provided ClusterIssuer.

CreateContainerConfigError or a pod stuck Pending

  • CreateContainerConfigError most often means a securityContext.runAsUser mismatch — describe the pod for the exact reason (k3s kubectl describe pod <name>).
  • Pending with no node assigned: check k3s kubectl get nodes and k3s kubectl describe pod <name> for scheduling/resource-pressure reasons.

Control plane (management host) — legacy Compose engine

Still applicable to installs run with --engine compose / -Engine Compose:

# Stack service status
systemctl status cloudgrange

# Container status and logs
sudo docker compose -f /opt/cloudgrange/docker-compose.yml ps
sudo docker compose -f /opt/cloudgrange/docker-compose.yml logs --tail 100 cloudgrange-api

API health probes

Probe Meaning
GET /health/live Process is running.
GET /health/ready Database connected and migrations complete.
GET /health/startup Initialization finished.

A failing readiness probe usually means PostgreSQL is unreachable or migrations have not completed — check the postgres and api container logs.

"The updater service is not running on this appliance"

The in-app update flow depends on how CloudGrange was deployed. Platform updates are designed to run inside the cluster on every deployment method. Foundation updates (operating system and K3s) run through a host service only where CloudGrange provisioned the host. On your own Kubernetes cluster or on AKS there is no host service, and none is expected. See Updates.

On a managed foundation (appliance, Windows script or Linux script), check the host service with sudo systemctl status cloudgrange-updater-k3s. Until the in-cluster Platform updater ships, you can always apply a Platform release with Helm directly. See Updating with Helm directly.

Sign-in fails after the platform moves to a new hostname

Symptom. The platform is reachable at its new address, but signing in fails. Keycloak shows "Invalid parameter: redirect_uri", or the portal shows an unexplained error at the end of the sign-in redirect.

Cause. The identity realm records the exact addresses it will return a browser to. Those were written when the realm was first created, so they still name the OLD hostname.

What CloudGrange does about it. The platform reconciles the realm against its configured public address on every start, and again every five minutes. Open Platform administration -> Platform Health and look at the Identity Realm card:

Card says Means
In sync with https://... The realm matches the address shown. Nothing to do.
Reconciled to https://... The realm was out of date and has just been corrected. Sign in again.
Waiting for Keycloak: ... The identity service is still starting. Normal for the first few minutes after an install or upgrade.
Realm reconcile is off: ... No public address is configured for this deployment (see below).
anything else, shown red The realm could not be corrected. The card names the reason.

If the card says the reconcile is off, set the platform public origin and restart the API:

  • Helm / BYO Kubernetes: global.hostname in your values file. The chart derives Keycloak__PublicOrigin (https://<global.hostname>) from it, so this must be the address your users type in a browser — not an in-cluster Service name.
  • Azure Container Apps: set automatically from the portal application URL.
  • Docker Compose: CLOUDGRANGE_HOSTNAME in .env.

Agent and relay

Agent does not appear in the portal

  1. Confirm the service is running: Get-Service CloudGrangeAgent.
  2. Confirm the host can reach the relay: Test-NetConnection -ComputerName <relay-host> -Port 8443.
  3. Verify AGENT_RELAY_URL and AGENT_ENROLLMENT_TOKEN are set correctly for the service.
  4. If the host was previously enrolled, re-enrollment requires admin approval on the relay (see Install the Relay).

Relay is not routing jobs

  • The relay rejects jobs for hosts that are not enrolled (no-route) and jobs with no target (target-required). Check the relay container logs for these rejection reasons.
  • Confirm the relay's outbound connection to the control plane is up (RELAY_PAAS_URL) and its health thresholds (RELAY_HEALTH_*) are not tripped — the relay reports degraded status with a reason.

Where to look next

  • Logs: Loki (aggregated container logs) via Grafana.
  • Metrics: Prometheus, viewed through Grafana dashboards.
  • Traces: OpenTelemetry collector.

See current product and release status.