Prerequisites — bring your own Kubernetes
This page lists everything a Kubernetes cluster needs before you install the CloudGrange Helm chart on it. It is for advanced customers who already run Kubernetes. If you do not run Kubernetes today, use one of the installer paths instead. They build a supported cluster for you and then install the same chart. See Prerequisites.
CloudGrange is in active development and is not GA. These are preview requirements, not certified production support minimums. See current product and release status.
What you are responsible for
On a cluster you provide, CloudGrange owns the Platform: the Helm release, its images, its schema and its data. You own the Foundation: the nodes, their operating system, the Kubernetes distribution and version, the ingress controller, storage, cert-manager and the network. CloudGrange never patches or upgrades your cluster. See Support boundary.
Checklist
| # | Requirement | Detail |
|---|---|---|
| 1 | Kubernetes version | A conformant cluster. v1.36 is the version CloudGrange tests on |
| 2 | Helm | Helm 3.x or 4.x (not v4.2.1). Tested: v3.22.0 and v4.3.0 |
| 3 | Ingress controller | Any controller. You name its IngressClass in global.ingress.className |
| 4 | LoadBalancer for the relay | The relay Service is type: LoadBalancer on TCP 8443 |
| 5 | Default StorageClass | ReadWriteOnce. About 29.4 GiB across 7 volumes in the default profile |
| 6 | TLS | cert-manager installed first, or your own certificate in a Secret |
| 7 | CPU and memory | 870m CPU and 1.75 GiB memory requested. Memory limits total about 5.4 GiB |
| 8 | Pod Security | The release namespace must allow the privileged level |
| 9 | RBAC and namespace | cluster-admin to install. On the defaults the chart creates nothing outside the release namespace |
| 10 | In-app Platform updates | Updates run in-cluster. A chart version that changes a cluster-scoped object needs one helm upgrade from you |
| 11 | Network egress or air gap | ghcr.io, docker.io, quay.io and the update channel |
| 12 | DNS | A name that resolves to your ingress, or an IP address |
Kubernetes version
Any CNCF-conformant Kubernetes distribution can run the chart, for example K3s, RKE2, kubeadm, microk8s, kind, AKS, EKS or GKE. The chart uses only GA APIs: apps/v1, batch/v1, networking.k8s.io/v1 and rbac.authorization.k8s.io/v1.
CloudGrange tests on, and pins its own managed foundations to, K3s v1.36.4+k3s1 (Kubernetes 1.36).
The chart will declare its supported Kubernetes range, and helm install will refuse a cluster outside that range. After install, Platform → Updates shows the detected Kubernetes version and whether it is supported. See Updates.
The chart's secrets-bootstrap hook Job runs the docker.io/alpine/k8s:1.31.1 image, which provides kubectl 1.31.
Helm and kubectl
- Helm 3.x or Helm 4.x. CloudGrange tests every release with Helm v3.22.0 (what its own installers install) and v4.3.0. Do not use Helm v4.2.1: a Helm bug (helm/helm#32214) makes
helm install --waithang until its timeout. - kubectl pointed at the target cluster, to verify the install and read the first-login password.
Ingress controller
The chart creates one Ingress and does not install an ingress controller. Your cluster must already run one, and you must name its IngressClass in global.ingress.className. If nothing claims the Ingress, the portal is unreachable and no error explains why.
| Value | Default | Meaning |
|---|---|---|
global.ingress.className |
traefik |
The IngressClass to bind to. Set "" to omit the field and let a default IngressClass claim it |
global.ingress.annotations |
traefik.ingress.kubernetes.io/router.entrypoints: websecure |
Controller-specific annotations, applied verbatim. Set to null on any controller other than Traefik |
| Your controller | global.ingress.className |
|---|---|
| Traefik (the K3s default) | traefik (the default) |
| ingress-nginx | nginx |
microk8s built-in (microk8s enable ingress) |
public |
| AKS Application Gateway | azure-application-gateway |
| AWS Load Balancer Controller | alb |
List the classes your cluster offers:
kubectl get ingressclass
The Ingress terminates TLS for the hostname in global.hostname, using Secret <release>-tls. It sends /api/v1/platform/update/upload (the release-bundle upload, which can be several GB) straight to the API Service and everything else to the portal Service on port 8080. If your controller limits request body size, raise the limit, for example with the ingress-nginx annotation nginx.ingress.kubernetes.io/proxy-body-size: "0".
If global.hostname is an IPv4 address, the chart leaves out the Ingress host rule, because Kubernetes rejects an IP address in that field. The rule then matches every host header.
LoadBalancer for the relay
The site relay, which the CloudGrange agents on your Hyper-V hosts connect to, is published by Service <release>-relay with type: LoadBalancer on TCP 8443. It does not go through the Ingress.
- On a cloud cluster (AKS, EKS, GKE) the cloud load balancer gives the Service an address.
- On bare metal you need a LoadBalancer implementation, for example MetalLB, or K3s's built-in ServiceLB. The chart can render MetalLB's
IPAddressPoolandL2Advertisementfor you when you setmetallb.enabled=trueandmetallb.addressPool. MetalLB itself must already be installed as its own Helm release. - Without a LoadBalancer implementation, the Service stays
Pending. The platform still works inside the cluster, but agents on your network cannot reach the relay. As an alternative you can setrelay.service.type=NodePortand point agents at a node address.
Storage
The cluster needs a default StorageClass that can provision ReadWriteOnce volumes. Exactly one class should be marked (default):
kubectl get storageclass
Only the database volume's class can be overridden, with postgres.persistence.storageClass. Every other volume uses the cluster default.
Volumes created by the default profile (base values.yaml, which is what a plain helm install uses):
| PersistentVolumeClaim | Access mode | Size | Holds | Value |
|---|---|---|---|---|
data-<release>-postgres-0 |
RWO | 10Gi | PostgreSQL database | postgres.persistence.size |
<release>-api-updates |
RWO | 8Gi | Uploaded release bundles | api.persistence.updates.size |
<release>-prometheus-data |
RWO | 5Gi | Metrics | observability.prometheus.persistence.size |
<release>-loki-data |
RWO | 5Gi | Logs | observability.loki.persistence.size |
<release>-grafana-data |
RWO | 1Gi | Dashboards | observability.grafana.persistence.size |
<release>-api-secrets |
RWO | 256Mi | API local configuration and secrets (/etc/cloudgrange) |
api.persistence.secrets.size |
<release>-relay-identity |
RWO | 128Mi | Relay identity | relay.persistence.identity.size |
| Total | about 29.4 GiB | 7 volumes |
The multi-node profile (values-multi-node.yaml) adds a 1Gi Redis volume and replaces the single PostgreSQL volume with a three-instance CloudNativePG cluster of 10Gi each. That is about 50.4 GiB across 9 volumes. CloudNativePG must be installed first as its own Helm release.
The database, Prometheus and Loki volumes grow with the number of managed hosts and with retention. Size them for your estate. PersistentVolumeClaims survive helm uninstall by design.
TLS: cert-manager or your own certificate
The Ingress serves the certificate in Secret <release>-tls. Choose one of these:
cert-manager (default). With certManager.installOperator: true (the default), the chart creates a self-signed Issuer named <release>-selfsigned in the release namespace and a Certificate for global.hostname, valid for 90 days and renewed 15 days before expiry. The issuer is namespaced deliberately: it is used by exactly one Certificate, so it needs no cluster scope, and keeping it out of cluster scope is what lets the in-cluster updater apply later releases on its own (see In-app Platform updates). To use an issuer you already run instead — ACME/Let's Encrypt, an internal CA — set certManager.issuerRef.name (and certManager.issuerRef.kind, ClusterIssuer or Issuer); the chart then creates only the Certificate. cert-manager must already be running. Install it as its own Helm release before the chart. It ships the CRDs that the chart's templates use, and Helm validates every object in a release before it applies any of them. A combined install fails with no matches for kind "Certificate". CloudGrange's own installers use cert-manager v1.21.2 and include it in the bundle as charts/vendor/cert-manager-v1.21.2.tgz.
Your own certificate. Set certManager.installOperator=false and create the Secret yourself before you install, in the release namespace:
kubectl create secret tls cloudgrange-tls -n cloudgrange \
--cert=cloudgrange.crt --key=cloudgrange.key
The certificate must cover global.hostname. The Secret name is <release>-tls, which is cloudgrange-tls for a release named cloudgrange.
CPU and memory
These totals come from the chart's own requests and limits. They cover CloudGrange's workloads only, not cert-manager, your ingress controller or the Kubernetes system pods.
| Profile | CPU requests | Memory requests | Memory limits |
|---|---|---|---|
Default / single-node |
870m | 1792Mi (1.75 GiB) | 5504Mi (about 5.4 GiB) |
multi-node |
1470m | 2688Mi (2.6 GiB) | 8704Mi (8.5 GiB) |
No workload sets a CPU limit. The log shipper (Promtail) is a DaemonSet, so each extra node adds 20m CPU, 64Mi requested and 128Mi limit. Two hook Jobs run briefly at install and upgrade time, for example the realm-admin Job at 50m and 128Mi.
Per workload (default profile):
| Workload | Kind | Replicas | CPU request | Memory request | Memory limit |
|---|---|---|---|---|---|
api |
Deployment | 1 | 100m | 256Mi | 768Mi |
portal |
Deployment | 1 | 50m | 64Mi | 128Mi |
relay |
Deployment | 1 | 100m | 128Mi | 384Mi |
keycloak |
Deployment | 1 | 200m | 512Mi | 1Gi |
postgres |
StatefulSet | 1 | 200m | 256Mi | 1Gi |
otel-collector, prometheus, loki, grafana |
Deployment | 1 each | 50m each | 128Mi each | 512Mi each |
promtail |
DaemonSet | 1 per node | 20m | 64Mi | 128Mi |
Schedule against the limits, not the requests. The sum of the limits is what the workloads can actually use.
Pod Security
The release namespace must allow the Pod Security Standards privileged level. The baseline and restricted levels reject the chart's log shipper:
- The Promtail DaemonSet mounts the node's
/var/logand/var/lib/docker/containersashostPathvolumes, read-only, to collect container logs. - Its init container runs
privilegedto raise the node'sfs.inotify.max_user_instancesto 8192. The Ubuntu/Debian default of 128 is exhausted on real nodes. - It tolerates every taint, so it also runs on control-plane nodes.
The other workloads run as non-root users (API non-root, portal uid 101, Keycloak uid 10002, Redis uid 999).
Label the namespace before you install:
kubectl create namespace cloudgrange
kubectl label namespace cloudgrange \
pod-security.kubernetes.io/enforce=privileged \
pod-security.kubernetes.io/warn=privileged
If your cluster also runs an admission policy engine (Kyverno, Gatekeeper, Azure Policy), exempt the release namespace or allow the Promtail DaemonSet.
RBAC and namespace
The chart is namespace-agnostic and uses the release namespace for everything. Install it into a dedicated namespace, for example cloudgrange.
The identity that runs helm install needs cluster-admin, or an equivalent. The chart creates cluster-scoped objects, and granting RBAC requires holding the rights being granted:
| Object | Scope | Why |
|---|---|---|
ClusterRole and ClusterRoleBinding <release>-promtail |
Cluster | Promtail reads pods and nodes (get, list, watch) to label log lines. Only when observability.promtail.enabled=true, which is not the default on your own cluster |
ClusterRole and ClusterRoleBinding <release>-platform-updater |
Cluster | Read-only get on the two objects above, so that the in-cluster updater can upgrade a release that has them. Rendered only alongside them |
Issuer <release>-selfsigned and Certificate <release>-tls |
Namespace | The self-signed issuer and the platform certificate. Only when certManager.installOperator=true (the default) and cert-manager is installed |
ServiceAccount, Role and RoleBinding <release>-secrets-bootstrap |
Namespace | A pre-install hook Job. get on Secret <release>-secrets and create on secrets. It never updates or deletes |
ServiceAccount <release>-api, Role and RoleBinding <release>-module-runtime |
Namespace | Lets the API run installed modules: create/update/delete deployments and services, and read pods |
ServiceAccount, Role and RoleBinding <release>-platform-updater |
Namespace | The in-cluster Platform updater: what helm upgrade and helm rollback need on this release's own objects, plus configmaps and secrets for Helm's release storage |
Role and RoleBinding <release>-api-platform-updates |
Namespace | Lets the API create and watch the updater Job and read the update-status ConfigMap |
IPAddressPool, L2Advertisement |
Namespace | Only when metallb.enabled=true |
With the defaults, the chart creates nothing outside the release namespace at all. The cluster-scoped rows above appear only when you turn on the log shipper. The platform's own service accounts get no cluster-wide write rights, and no access to pods/exec, RBAC objects or other namespaces.
In-app Platform updates on your cluster
Platform updates are applied from Platform → Updates in the portal, on your cluster as on every other. The API creates a Kubernetes Job that runs under ServiceAccount <release>-platform-updater; the Job verifies the signed release manifest, backs up the database, runs helm upgrade, and rolls both back if the health gate fails. Nothing touches your nodes and no host access is needed. See Updates.
The updater's rights stop at the release namespace, by design. It is effectively namespace-admin inside the release namespace and holds almost nothing outside it — and it cannot grant itself more, because granting RBAC requires already holding what you are granting. So there is exactly one thing an in-app update cannot do: apply a chart version that adds, changes or removes a cluster-scoped object.
The updater checks this before it changes anything — before the database backup, before helm upgrade. If a release needs cluster-scoped rights it does not have, the update is refused with no changes made and no rollback, and Platform → Updates shows you the exact command to run. It looks like this:
Platform
<version>changes cluster-scoped objects the in-cluster updater is not allowed to change: create ClusterRole/<release>-platform-updater (cluster-scoped). …
Run that one upgrade yourself, with a cluster-admin kubeconfig:
helm upgrade cloudgrange oci://ghcr.io/cloudgrange/charts/cloudgrange \
--version <new-version> --namespace cloudgrange \
--reset-then-reuse-values \
--set global.image.tag=<new-version> \
--wait --timeout 10m
Use
--reset-then-reuse-values, never--reuse-values.--reuse-valueskeeps the old release's values verbatim and therefore drops every default the new chart introduces. That is what broke the in-app updates to2609.0.0-preview.18and.19: the new templates read aairgapvalues key the old release did not have, and the upgrade failed withnil pointer evaluating interface {}.registry.--reset-then-reuse-values(Helm 3.14 and later) resets to the new chart's defaults and then re-applies only the values you supplied. It is the flag the in-cluster updater itself uses. Passing your full values file with-finstead is equally correct.
After that single command, in-app updates work again for every later release.
This used to happen on the defaults, and no longer does. Up to 2609.0.0-preview.26 the chart's self-signed issuer was a cluster-scoped ClusterIssuer and the updater's own ClusterRole/ClusterRoleBinding were rendered alongside it. A release installed with certManager.installOperator=true before cert-manager was in the cluster therefore had no ClusterIssuer yet, and the next update tried to create one — forbidden, after the database backup, so it rolled back. From the first release after 2609.0.0-preview.26 the self-signed issuer is a namespaced Issuer and the updater's cluster-scoped RBAC comes only with the log shipper, so on the defaults there is nothing cluster-scoped left for an update to trip over. A 2609.0.0-preview.12 release takes that update in-app, with no manual step (verified on a kind cluster, for BYO defaults, installOperator=true and with the log shipper on).
Network egress
The nodes pull these images. The versions are those in the chart for this preview.
| Registry | Images |
|---|---|
ghcr.io |
cloudgrange/cloudgrange-api, cloudgrange/cloudgrange-portal, cloudgrange/cloudgrange-relay, and the chart itself (oci://ghcr.io/cloudgrange/charts/cloudgrange). With the multi-node profile also cloudnative-pg/postgresql:17.11 |
docker.io (Docker Hub) |
postgres:17.11-alpine, alpine/k8s:1.31.1, busybox:1.36, grafana/grafana:12.1.1, grafana/loki:3.7.7, grafana/promtail:3.6.3, prom/prometheus:v3.14.0, otel/opentelemetry-collector-contrib:0.160.0. With the multi-node profile also redis:7.4-alpine |
quay.io |
keycloak/keycloak:26.6.4, and cert-manager's images if you install cert-manager from its chart |
At run time the API also checks for Platform updates at the update channel, over HTTPS:
https://pub-ab113af532ff44ef827c176e42118f17.r2.dev/channels/preview.json(valueapi.updateChannelUrl)
If the update channel is unreachable, the update check reports it as unavailable. The platform itself keeps working.
If you install cert-manager from its public repository instead of the vendored chart, you also need charts.jetstack.io.
Air-gapped clusters
On a cluster with no internet access, mirror CloudGrange's pinned images into a registry your nodes can reach, then point the chart at it with one value, global.imageRegistry.
The image list: images.txt
Each release publishes images.txt with its release artifacts. It lists every container image that release can run, in every profile. Each line is:
<repository> <tag> <digest>
- Lines starting with
#are comments. <repository>is fully qualified:ghcr.io/cloudgrange/cloudgrange-api,quay.io/keycloak/keycloak, anddocker.io/library/postgresfor Docker Hub short names.- Every image is pinned by
<digest>. The release manifest pinsimages.txtitself by SHA-256, and the update channel pins the manifest.
The mirror path of an image is its repository without the registry host. For example, ghcr.io/cloudgrange/cloudgrange-api becomes cloudgrange/cloudgrange-api, and docker.io/library/postgres becomes library/postgres. With global.imageRegistry=registry.example.com/cg, the chart pulls registry.example.com/cg/cloudgrange/cloudgrange-api:<tag>@<digest>.
Mirror the images
On a machine that can reach both the internet and your registry, copy each image with crane. Copying by digest keeps the digest, so the chart's digest pins still match:
MIRROR=registry.example.com/cg # your registry (and optional path prefix)
while read -r repo tag digest; do
case "$repo" in ''|\#*) continue ;; esac
crane copy "$repo@$digest" "$MIRROR/${repo#*/}:$tag"
done < images.txt
skopeo copy --all --preserve-digests docker://<repo>@<digest> docker://<mirror>/<mirror path>:<tag> works as well.
If you cannot reach both networks from one machine, copy the images to a portable format first, for example with crane pull --format=oci, and push them from inside the isolated network.
Install or upgrade from the mirror
helm upgrade --install cloudgrange ./cloudgrange-<version>.tgz \
--namespace <namespace> --set global.imageRegistry=registry.example.com/cg
With global.imageRegistry set, every image the chart renders comes from your mirror: the CloudGrange images, the third-party images (including the digest-pinned ones such as busybox and promtail), and the image of the in-app Platform updater Job. Digests are kept, so the images are still verified by content.
Notes:
- Vendor charts are separate. cert-manager, the CloudNativePG operator, MetalLB and Velero are installed as their own Helm releases and are not in
images.txt. Mirror their images and use those charts' own image overrides. - Pull authentication is yours. If your mirror needs credentials, add an
imagePullSecretto the namespace'sdefaultServiceAccount, or configure credentials on the nodes. - Updates. The in-app Platform updater needs the release files it would otherwise download. Either host the update channel, the release manifest and the chart internally over HTTPS and set
api.updateChannel.url, or upload the offline Platform bundle (cloudgrange-platform-<version>.zip) on the Platform card. On your own cluster there is no in-cluster registry, so the updater Job does not push the bundle's images: mirror that release'simages.txtfirst. You can also runhelm upgradeyourself. See Air-gapped Platform updates.
The Linux installer bundle, Install-CloudGrange-K3s-Bundled.zip from the release page, also carries the chart (charts/cloudgrange/), the vendor charts (charts/vendor/) and the images as one archive (airgap/cloudgrange-images-amd64.tar, listed in airgap/images.txt), if you prefer to load images from a file.
DNS
Choose the name users will browse to and pass it as global.hostname:
- A DNS name (recommended). Create an
AorCNAMErecord that resolves to your ingress controller's external address. The name is used for the Ingress host rule, the TLS certificate and Keycloak's issuer URL, so it must be the name users actually type. - An IPv4 address. This is supported. The Ingress
hostrule is omitted, so any host header matches.
Agents reach the relay by the address of the <release>-relay LoadBalancer Service. Give it a DNS name as well if agents should not use a raw IP address.
Changing global.hostname after users have signed in changes the identity issuer. Choose it before first-run setup. If you do change it later, the platform repairs the identity realm's redirect addresses itself on the next API start — watch the Identity Realm card on Platform Health, and see Sign-in fails after the platform moves to a new hostname. Signed-in sessions still need to sign in again, because the issuer in their tokens no longer matches.
Install
Once every prerequisite is met, run the following, replacing the version with the release you are installing.
1. cert-manager (skip if it already runs in the cluster, or if you bring your own certificate):
helm repo add jetstack https://charts.jetstack.io
helm install cert-manager jetstack/cert-manager \
--version v1.21.2 \
--namespace cert-manager --create-namespace \
--set crds.enabled=true \
--wait --timeout 5m
2. The namespace, with Pod Security set:
kubectl create namespace cloudgrange
kubectl label namespace cloudgrange pod-security.kubernetes.io/enforce=privileged
3. A values file (cloudgrange-values.yaml). This example is for ingress-nginx with a cloud LoadBalancer:
global:
hostname: cloudgrange.example.com
image:
# Pin the image tag to the chart version you install. Never use "latest".
tag: 2609.0.0-preview.7
ingress:
className: nginx
annotations:
nginx.ingress.kubernetes.io/proxy-body-size: "0"
certManager:
installOperator: true # namespaced self-signed Issuer + Certificate.
# false = you created Secret cloudgrange-tls yourself
# issuerRef: # or point the Certificate at an issuer you already run
# name: letsencrypt-prod
# kind: ClusterIssuer
postgres:
persistence:
size: 10Gi
storageClass: "" # "" = the cluster's default StorageClass
relay:
service:
type: LoadBalancer # NodePort if the cluster has no LoadBalancer implementation
4. Install the chart:
helm install cloudgrange oci://ghcr.io/cloudgrange/charts/cloudgrange \
--version 2609.0.0-preview.7 \
--namespace cloudgrange \
-f cloudgrange-values.yaml \
--wait --timeout 10m
5. Verify:
kubectl get pods -n cloudgrange
kubectl get ingress,svc -n cloudgrange
Every pod should reach Running and ready. The <release>-secrets-bootstrap hook Job correctly ends as Completed. The relay Service should have an external address.
6. Read the first-login password and open the portal. The chart generates every credential. You supply none.
kubectl get secret cloudgrange-secrets -n cloudgrange \
-o jsonpath='{.data.realm-admin-password}' | base64 -d; echo
Browse to https://<global.hostname>/ and complete the first-run setup wizard.
For the rest of the lifecycle (upgrade, uninstall, the full values reference), see Install with Helm.