Prerequisites — bring your own Kubernetes

This page lists everything a Kubernetes cluster needs before you install the CloudGrange Helm chart on it. It is for advanced customers who already run Kubernetes. If you do not run Kubernetes today, use one of the installer paths instead. They build a supported cluster for you and then install the same chart. See Prerequisites.

CloudGrange is in active development and is not GA. These are preview requirements, not certified production support minimums. See current product and release status.

What you are responsible for

On a cluster you provide, CloudGrange owns the Platform: the Helm release, its images, its schema and its data. You own the Foundation: the nodes, their operating system, the Kubernetes distribution and version, the ingress controller, storage, cert-manager and the network. CloudGrange never patches or upgrades your cluster. See Support boundary.

Checklist

# Requirement Detail
1 Kubernetes version A conformant cluster. v1.36 is the version CloudGrange tests on
2 Helm Helm 3.x or 4.x (not v4.2.1). Tested: v3.22.0 and v4.3.0
3 Ingress controller Any controller. You name its IngressClass in global.ingress.className
4 LoadBalancer for the relay The relay Service is type: LoadBalancer on TCP 8443
5 Default StorageClass ReadWriteOnce. About 29.4 GiB across 7 volumes in the default profile
6 TLS cert-manager installed first, or your own certificate in a Secret
7 CPU and memory 870m CPU and 1.75 GiB memory requested. Memory limits total about 5.4 GiB
8 Pod Security The release namespace must allow the privileged level
9 RBAC and namespace cluster-admin to install. On the defaults the chart creates nothing outside the release namespace
10 In-app Platform updates Updates run in-cluster. A chart version that changes a cluster-scoped object needs one helm upgrade from you
11 Network egress or air gap ghcr.io, docker.io, quay.io and the update channel
12 DNS A name that resolves to your ingress, or an IP address

Kubernetes version

Any CNCF-conformant Kubernetes distribution can run the chart, for example K3s, RKE2, kubeadm, microk8s, kind, AKS, EKS or GKE. The chart uses only GA APIs: apps/v1, batch/v1, networking.k8s.io/v1 and rbac.authorization.k8s.io/v1.

CloudGrange tests on, and pins its own managed foundations to, K3s v1.36.4+k3s1 (Kubernetes 1.36).

The chart will declare its supported Kubernetes range, and helm install will refuse a cluster outside that range. After install, Platform → Updates shows the detected Kubernetes version and whether it is supported. See Updates.

The chart's secrets-bootstrap hook Job runs the docker.io/alpine/k8s:1.31.1 image, which provides kubectl 1.31.

Helm and kubectl

  • Helm 3.x or Helm 4.x. CloudGrange tests every release with Helm v3.22.0 (what its own installers install) and v4.3.0. Do not use Helm v4.2.1: a Helm bug (helm/helm#32214) makes helm install --wait hang until its timeout.
  • kubectl pointed at the target cluster, to verify the install and read the first-login password.

Ingress controller

The chart creates one Ingress and does not install an ingress controller. Your cluster must already run one, and you must name its IngressClass in global.ingress.className. If nothing claims the Ingress, the portal is unreachable and no error explains why.

Value Default Meaning
global.ingress.className traefik The IngressClass to bind to. Set "" to omit the field and let a default IngressClass claim it
global.ingress.annotations traefik.ingress.kubernetes.io/router.entrypoints: websecure Controller-specific annotations, applied verbatim. Set to null on any controller other than Traefik
Your controller global.ingress.className
Traefik (the K3s default) traefik (the default)
ingress-nginx nginx
microk8s built-in (microk8s enable ingress) public
AKS Application Gateway azure-application-gateway
AWS Load Balancer Controller alb

List the classes your cluster offers:

kubectl get ingressclass

The Ingress terminates TLS for the hostname in global.hostname, using Secret <release>-tls. It sends /api/v1/platform/update/upload (the release-bundle upload, which can be several GB) straight to the API Service and everything else to the portal Service on port 8080. If your controller limits request body size, raise the limit, for example with the ingress-nginx annotation nginx.ingress.kubernetes.io/proxy-body-size: "0".

If global.hostname is an IPv4 address, the chart leaves out the Ingress host rule, because Kubernetes rejects an IP address in that field. The rule then matches every host header.

LoadBalancer for the relay

The site relay, which the CloudGrange agents on your Hyper-V hosts connect to, is published by Service <release>-relay with type: LoadBalancer on TCP 8443. It does not go through the Ingress.

  • On a cloud cluster (AKS, EKS, GKE) the cloud load balancer gives the Service an address.
  • On bare metal you need a LoadBalancer implementation, for example MetalLB, or K3s's built-in ServiceLB. The chart can render MetalLB's IPAddressPool and L2Advertisement for you when you set metallb.enabled=true and metallb.addressPool. MetalLB itself must already be installed as its own Helm release.
  • Without a LoadBalancer implementation, the Service stays Pending. The platform still works inside the cluster, but agents on your network cannot reach the relay. As an alternative you can set relay.service.type=NodePort and point agents at a node address.

Storage

The cluster needs a default StorageClass that can provision ReadWriteOnce volumes. Exactly one class should be marked (default):

kubectl get storageclass

Only the database volume's class can be overridden, with postgres.persistence.storageClass. Every other volume uses the cluster default.

Volumes created by the default profile (base values.yaml, which is what a plain helm install uses):

PersistentVolumeClaim Access mode Size Holds Value
data-<release>-postgres-0 RWO 10Gi PostgreSQL database postgres.persistence.size
<release>-api-updates RWO 8Gi Uploaded release bundles api.persistence.updates.size
<release>-prometheus-data RWO 5Gi Metrics observability.prometheus.persistence.size
<release>-loki-data RWO 5Gi Logs observability.loki.persistence.size
<release>-grafana-data RWO 1Gi Dashboards observability.grafana.persistence.size
<release>-api-secrets RWO 256Mi API local configuration and secrets (/etc/cloudgrange) api.persistence.secrets.size
<release>-relay-identity RWO 128Mi Relay identity relay.persistence.identity.size
Total about 29.4 GiB 7 volumes

The multi-node profile (values-multi-node.yaml) adds a 1Gi Redis volume and replaces the single PostgreSQL volume with a three-instance CloudNativePG cluster of 10Gi each. That is about 50.4 GiB across 9 volumes. CloudNativePG must be installed first as its own Helm release.

The database, Prometheus and Loki volumes grow with the number of managed hosts and with retention. Size them for your estate. PersistentVolumeClaims survive helm uninstall by design.

TLS: cert-manager or your own certificate

The Ingress serves the certificate in Secret <release>-tls. Choose one of these:

cert-manager (default). With certManager.installOperator: true (the default), the chart creates a self-signed Issuer named <release>-selfsigned in the release namespace and a Certificate for global.hostname, valid for 90 days and renewed 15 days before expiry. The issuer is namespaced deliberately: it is used by exactly one Certificate, so it needs no cluster scope, and keeping it out of cluster scope is what lets the in-cluster updater apply later releases on its own (see In-app Platform updates). To use an issuer you already run instead — ACME/Let's Encrypt, an internal CA — set certManager.issuerRef.name (and certManager.issuerRef.kind, ClusterIssuer or Issuer); the chart then creates only the Certificate. cert-manager must already be running. Install it as its own Helm release before the chart. It ships the CRDs that the chart's templates use, and Helm validates every object in a release before it applies any of them. A combined install fails with no matches for kind "Certificate". CloudGrange's own installers use cert-manager v1.21.2 and include it in the bundle as charts/vendor/cert-manager-v1.21.2.tgz.

Your own certificate. Set certManager.installOperator=false and create the Secret yourself before you install, in the release namespace:

kubectl create secret tls cloudgrange-tls -n cloudgrange \
  --cert=cloudgrange.crt --key=cloudgrange.key

The certificate must cover global.hostname. The Secret name is <release>-tls, which is cloudgrange-tls for a release named cloudgrange.

CPU and memory

These totals come from the chart's own requests and limits. They cover CloudGrange's workloads only, not cert-manager, your ingress controller or the Kubernetes system pods.

Profile CPU requests Memory requests Memory limits
Default / single-node 870m 1792Mi (1.75 GiB) 5504Mi (about 5.4 GiB)
multi-node 1470m 2688Mi (2.6 GiB) 8704Mi (8.5 GiB)

No workload sets a CPU limit. The log shipper (Promtail) is a DaemonSet, so each extra node adds 20m CPU, 64Mi requested and 128Mi limit. Two hook Jobs run briefly at install and upgrade time, for example the realm-admin Job at 50m and 128Mi.

Per workload (default profile):

Workload Kind Replicas CPU request Memory request Memory limit
api Deployment 1 100m 256Mi 768Mi
portal Deployment 1 50m 64Mi 128Mi
relay Deployment 1 100m 128Mi 384Mi
keycloak Deployment 1 200m 512Mi 1Gi
postgres StatefulSet 1 200m 256Mi 1Gi
otel-collector, prometheus, loki, grafana Deployment 1 each 50m each 128Mi each 512Mi each
promtail DaemonSet 1 per node 20m 64Mi 128Mi

Schedule against the limits, not the requests. The sum of the limits is what the workloads can actually use.

Pod Security

The release namespace must allow the Pod Security Standards privileged level. The baseline and restricted levels reject the chart's log shipper:

  • The Promtail DaemonSet mounts the node's /var/log and /var/lib/docker/containers as hostPath volumes, read-only, to collect container logs.
  • Its init container runs privileged to raise the node's fs.inotify.max_user_instances to 8192. The Ubuntu/Debian default of 128 is exhausted on real nodes.
  • It tolerates every taint, so it also runs on control-plane nodes.

The other workloads run as non-root users (API non-root, portal uid 101, Keycloak uid 10002, Redis uid 999).

Label the namespace before you install:

kubectl create namespace cloudgrange
kubectl label namespace cloudgrange \
  pod-security.kubernetes.io/enforce=privileged \
  pod-security.kubernetes.io/warn=privileged

If your cluster also runs an admission policy engine (Kyverno, Gatekeeper, Azure Policy), exempt the release namespace or allow the Promtail DaemonSet.

RBAC and namespace

The chart is namespace-agnostic and uses the release namespace for everything. Install it into a dedicated namespace, for example cloudgrange.

The identity that runs helm install needs cluster-admin, or an equivalent. The chart creates cluster-scoped objects, and granting RBAC requires holding the rights being granted:

Object Scope Why
ClusterRole and ClusterRoleBinding <release>-promtail Cluster Promtail reads pods and nodes (get, list, watch) to label log lines. Only when observability.promtail.enabled=true, which is not the default on your own cluster
ClusterRole and ClusterRoleBinding <release>-platform-updater Cluster Read-only get on the two objects above, so that the in-cluster updater can upgrade a release that has them. Rendered only alongside them
Issuer <release>-selfsigned and Certificate <release>-tls Namespace The self-signed issuer and the platform certificate. Only when certManager.installOperator=true (the default) and cert-manager is installed
ServiceAccount, Role and RoleBinding <release>-secrets-bootstrap Namespace A pre-install hook Job. get on Secret <release>-secrets and create on secrets. It never updates or deletes
ServiceAccount <release>-api, Role and RoleBinding <release>-module-runtime Namespace Lets the API run installed modules: create/update/delete deployments and services, and read pods
ServiceAccount, Role and RoleBinding <release>-platform-updater Namespace The in-cluster Platform updater: what helm upgrade and helm rollback need on this release's own objects, plus configmaps and secrets for Helm's release storage
Role and RoleBinding <release>-api-platform-updates Namespace Lets the API create and watch the updater Job and read the update-status ConfigMap
IPAddressPool, L2Advertisement Namespace Only when metallb.enabled=true

With the defaults, the chart creates nothing outside the release namespace at all. The cluster-scoped rows above appear only when you turn on the log shipper. The platform's own service accounts get no cluster-wide write rights, and no access to pods/exec, RBAC objects or other namespaces.

In-app Platform updates on your cluster

Platform updates are applied from Platform → Updates in the portal, on your cluster as on every other. The API creates a Kubernetes Job that runs under ServiceAccount <release>-platform-updater; the Job verifies the signed release manifest, backs up the database, runs helm upgrade, and rolls both back if the health gate fails. Nothing touches your nodes and no host access is needed. See Updates.

The updater's rights stop at the release namespace, by design. It is effectively namespace-admin inside the release namespace and holds almost nothing outside it — and it cannot grant itself more, because granting RBAC requires already holding what you are granting. So there is exactly one thing an in-app update cannot do: apply a chart version that adds, changes or removes a cluster-scoped object.

The updater checks this before it changes anything — before the database backup, before helm upgrade. If a release needs cluster-scoped rights it does not have, the update is refused with no changes made and no rollback, and Platform → Updates shows you the exact command to run. It looks like this:

Platform <version> changes cluster-scoped objects the in-cluster updater is not allowed to change: create ClusterRole/<release>-platform-updater (cluster-scoped). …

Run that one upgrade yourself, with a cluster-admin kubeconfig:

helm upgrade cloudgrange oci://ghcr.io/cloudgrange/charts/cloudgrange \
  --version <new-version> --namespace cloudgrange \
  --reset-then-reuse-values \
  --set global.image.tag=<new-version> \
  --wait --timeout 10m

Use --reset-then-reuse-values, never --reuse-values. --reuse-values keeps the old release's values verbatim and therefore drops every default the new chart introduces. That is what broke the in-app updates to 2609.0.0-preview.18 and .19: the new templates read a airgap values key the old release did not have, and the upgrade failed with nil pointer evaluating interface {}.registry. --reset-then-reuse-values (Helm 3.14 and later) resets to the new chart's defaults and then re-applies only the values you supplied. It is the flag the in-cluster updater itself uses. Passing your full values file with -f instead is equally correct.

After that single command, in-app updates work again for every later release.

This used to happen on the defaults, and no longer does. Up to 2609.0.0-preview.26 the chart's self-signed issuer was a cluster-scoped ClusterIssuer and the updater's own ClusterRole/ClusterRoleBinding were rendered alongside it. A release installed with certManager.installOperator=true before cert-manager was in the cluster therefore had no ClusterIssuer yet, and the next update tried to create one — forbidden, after the database backup, so it rolled back. From the first release after 2609.0.0-preview.26 the self-signed issuer is a namespaced Issuer and the updater's cluster-scoped RBAC comes only with the log shipper, so on the defaults there is nothing cluster-scoped left for an update to trip over. A 2609.0.0-preview.12 release takes that update in-app, with no manual step (verified on a kind cluster, for BYO defaults, installOperator=true and with the log shipper on).

Network egress

The nodes pull these images. The versions are those in the chart for this preview.

Registry Images
ghcr.io cloudgrange/cloudgrange-api, cloudgrange/cloudgrange-portal, cloudgrange/cloudgrange-relay, and the chart itself (oci://ghcr.io/cloudgrange/charts/cloudgrange). With the multi-node profile also cloudnative-pg/postgresql:17.11
docker.io (Docker Hub) postgres:17.11-alpine, alpine/k8s:1.31.1, busybox:1.36, grafana/grafana:12.1.1, grafana/loki:3.7.7, grafana/promtail:3.6.3, prom/prometheus:v3.14.0, otel/opentelemetry-collector-contrib:0.160.0. With the multi-node profile also redis:7.4-alpine
quay.io keycloak/keycloak:26.6.4, and cert-manager's images if you install cert-manager from its chart

At run time the API also checks for Platform updates at the update channel, over HTTPS:

  • https://pub-ab113af532ff44ef827c176e42118f17.r2.dev/channels/preview.json (value api.updateChannelUrl)

If the update channel is unreachable, the update check reports it as unavailable. The platform itself keeps working.

If you install cert-manager from its public repository instead of the vendored chart, you also need charts.jetstack.io.

Air-gapped clusters

On a cluster with no internet access, mirror CloudGrange's pinned images into a registry your nodes can reach, then point the chart at it with one value, global.imageRegistry.

The image list: images.txt

Each release publishes images.txt with its release artifacts. It lists every container image that release can run, in every profile. Each line is:

<repository> <tag> <digest>
  • Lines starting with # are comments.
  • <repository> is fully qualified: ghcr.io/cloudgrange/cloudgrange-api, quay.io/keycloak/keycloak, and docker.io/library/postgres for Docker Hub short names.
  • Every image is pinned by <digest>. The release manifest pins images.txt itself by SHA-256, and the update channel pins the manifest.

The mirror path of an image is its repository without the registry host. For example, ghcr.io/cloudgrange/cloudgrange-api becomes cloudgrange/cloudgrange-api, and docker.io/library/postgres becomes library/postgres. With global.imageRegistry=registry.example.com/cg, the chart pulls registry.example.com/cg/cloudgrange/cloudgrange-api:<tag>@<digest>.

Mirror the images

On a machine that can reach both the internet and your registry, copy each image with crane. Copying by digest keeps the digest, so the chart's digest pins still match:

MIRROR=registry.example.com/cg          # your registry (and optional path prefix)
while read -r repo tag digest; do
  case "$repo" in ''|\#*) continue ;; esac
  crane copy "$repo@$digest" "$MIRROR/${repo#*/}:$tag"
done < images.txt

skopeo copy --all --preserve-digests docker://<repo>@<digest> docker://<mirror>/<mirror path>:<tag> works as well.

If you cannot reach both networks from one machine, copy the images to a portable format first, for example with crane pull --format=oci, and push them from inside the isolated network.

Install or upgrade from the mirror

helm upgrade --install cloudgrange ./cloudgrange-<version>.tgz \
  --namespace <namespace> --set global.imageRegistry=registry.example.com/cg

With global.imageRegistry set, every image the chart renders comes from your mirror: the CloudGrange images, the third-party images (including the digest-pinned ones such as busybox and promtail), and the image of the in-app Platform updater Job. Digests are kept, so the images are still verified by content.

Notes:

  • Vendor charts are separate. cert-manager, the CloudNativePG operator, MetalLB and Velero are installed as their own Helm releases and are not in images.txt. Mirror their images and use those charts' own image overrides.
  • Pull authentication is yours. If your mirror needs credentials, add an imagePullSecret to the namespace's default ServiceAccount, or configure credentials on the nodes.
  • Updates. The in-app Platform updater needs the release files it would otherwise download. Either host the update channel, the release manifest and the chart internally over HTTPS and set api.updateChannel.url, or upload the offline Platform bundle (cloudgrange-platform-<version>.zip) on the Platform card. On your own cluster there is no in-cluster registry, so the updater Job does not push the bundle's images: mirror that release's images.txt first. You can also run helm upgrade yourself. See Air-gapped Platform updates.

The Linux installer bundle, Install-CloudGrange-K3s-Bundled.zip from the release page, also carries the chart (charts/cloudgrange/), the vendor charts (charts/vendor/) and the images as one archive (airgap/cloudgrange-images-amd64.tar, listed in airgap/images.txt), if you prefer to load images from a file.

DNS

Choose the name users will browse to and pass it as global.hostname:

  • A DNS name (recommended). Create an A or CNAME record that resolves to your ingress controller's external address. The name is used for the Ingress host rule, the TLS certificate and Keycloak's issuer URL, so it must be the name users actually type.
  • An IPv4 address. This is supported. The Ingress host rule is omitted, so any host header matches.

Agents reach the relay by the address of the <release>-relay LoadBalancer Service. Give it a DNS name as well if agents should not use a raw IP address.

Changing global.hostname after users have signed in changes the identity issuer. Choose it before first-run setup. If you do change it later, the platform repairs the identity realm's redirect addresses itself on the next API start — watch the Identity Realm card on Platform Health, and see Sign-in fails after the platform moves to a new hostname. Signed-in sessions still need to sign in again, because the issuer in their tokens no longer matches.

Install

Once every prerequisite is met, run the following, replacing the version with the release you are installing.

1. cert-manager (skip if it already runs in the cluster, or if you bring your own certificate):

helm repo add jetstack https://charts.jetstack.io
helm install cert-manager jetstack/cert-manager \
  --version v1.21.2 \
  --namespace cert-manager --create-namespace \
  --set crds.enabled=true \
  --wait --timeout 5m

2. The namespace, with Pod Security set:

kubectl create namespace cloudgrange
kubectl label namespace cloudgrange pod-security.kubernetes.io/enforce=privileged

3. A values file (cloudgrange-values.yaml). This example is for ingress-nginx with a cloud LoadBalancer:

global:
  hostname: cloudgrange.example.com
  image:
    # Pin the image tag to the chart version you install. Never use "latest".
    tag: 2609.0.0-preview.7
  ingress:
    className: nginx
    annotations:
      nginx.ingress.kubernetes.io/proxy-body-size: "0"

certManager:
  installOperator: true   # namespaced self-signed Issuer + Certificate.
                          # false = you created Secret cloudgrange-tls yourself
  # issuerRef:            # or point the Certificate at an issuer you already run
  #   name: letsencrypt-prod
  #   kind: ClusterIssuer

postgres:
  persistence:
    size: 10Gi
    storageClass: ""      # "" = the cluster's default StorageClass

relay:
  service:
    type: LoadBalancer    # NodePort if the cluster has no LoadBalancer implementation

4. Install the chart:

helm install cloudgrange oci://ghcr.io/cloudgrange/charts/cloudgrange \
  --version 2609.0.0-preview.7 \
  --namespace cloudgrange \
  -f cloudgrange-values.yaml \
  --wait --timeout 10m

5. Verify:

kubectl get pods -n cloudgrange
kubectl get ingress,svc -n cloudgrange

Every pod should reach Running and ready. The <release>-secrets-bootstrap hook Job correctly ends as Completed. The relay Service should have an external address.

6. Read the first-login password and open the portal. The chart generates every credential. You supply none.

kubectl get secret cloudgrange-secrets -n cloudgrange \
  -o jsonpath='{.data.realm-admin-password}' | base64 -d; echo

Browse to https://<global.hostname>/ and complete the first-run setup wizard.

For the rest of the lifecycle (upgrade, uninstall, the full values reference), see Install with Helm.

See current product and release status.