How the platform is deployed¶
rendimiento deploys apps, but it does not deploy itself. The platform is installed with plain Kubernetes manifests in deploy/, applied with kubectl apply -k deploy. Keeping the platform outside its own control means a broken rendimiento can always be fixed with kubectl.
Your own settings¶
The manifests in deploy/ carry example values (example.com, registry.example.lan, [email protected]). A real cluster's settings belong in a private overlay, a kustomization in a private repository that builds on deploy/ and replaces what differs:
# kustomization.yaml in a private repository, next to a clone of this one
resources:
- ../../rendimiento.ai/deploy
images:
- name: registry.example.lan:5000/rendimiento
newName: registry.home.lan:5000/rendimiento # your registry
patches:
- path: configmap.yaml # the rendimiento ConfigMap with your settings
- target: { kind: Ingress, name: rendimiento }
patch: |-
- { op: replace, path: /spec/rules/0/host, value: rendimiento.your-domain.com }
- { op: replace, path: /spec/tls/0/hosts/0, value: rendimiento.your-domain.com }
Then point the Makefile at it from a local.mk (git-ignored) at the repository root:
IMAGE := registry.home.lan:5000/rendimiento
DEPLOY_DIR := $(HOME)/private-repo/rendimiento-platform
TEST_EXCLUDE_NODES := my-gpu-node # nodes remote tests must avoid
ITEST_REGISTRY := registry.home.lan:5000
make deploy renders DEPLOY_DIR (deploy/ by default) and applies it. Your hostnames, node names and email then never enter the public repository.
What gets installed¶
flowchart TB
subgraph rs[namespace rendimiento-system]
dep[Deployment rendimiento<br/>1 replica, Recreate]
svc[Service rendimiento :80]
ing[Ingress rendimiento.joserod.space<br/>TLS by cert-manager]
db[(StatefulSet rendimiento-db<br/>Postgres 17, 5Gi Longhorn)]
cm[ConfigMap rendimiento<br/>settings]
sa[ServiceAccount rendimiento]
sa2[ServiceAccount rendimiento-addons<br/>no pods, impersonated]
end
subgraph rb[namespace rendimiento-builds]
np[NetworkPolicy isolate-builds]
end
crds[CRDs apps.rendimiento.ai<br/>addons.rendimiento.ai]
ing --> svc --> dep --> db
dep -. reads .-> cm
| File | What it contains |
|---|---|
kustomization.yaml |
The list below, applied together. |
namespace.yaml |
rendimiento-system (the platform) and rendimiento-builds (CI pods). |
crds/rendimiento.ai_apps.yaml, crds/rendimiento.ai_addons.yaml |
The two custom resource definitions, generated from api/v1alpha1 by make generate. Never edit them by hand. |
rbac.yaml |
What the platform may do (see below), the rendimiento-addons identity bound to cluster-admin, and the build namespace's Role. |
postgres.yaml |
The platform database: a one-replica StatefulSet on a 5 Gi Longhorn volume. |
rendimiento.yaml |
The settings ConfigMap, the Deployment, its Service and Ingress. |
networkpolicy.yaml |
Isolation for CI pods: DNS, BuildKit and the public internet only. |
railpack/Dockerfile |
Not applied: the image with the Railpack CLI that build pods use (make railpack-image). |
deploy/rendimiento.yaml (settings, Deployment, Service, Ingress)
apiVersion: v1
kind: ConfigMap
metadata:
name: rendimiento
namespace: rendimiento-system
data:
# Example values: a real cluster's settings live outside this public
# repository (an overlay applied over these manifests; see the book).
BASE_URL: https://rendimiento.example.com
ALLOWED_USERS: your-github-login
GITHUB_APP_NAME: rendimiento
DNS_ZONE: example.com
DNS_PROXIED: "true"
REGISTRY: registry.example.lan:5000
REGISTRY_INSECURE: "true"
BUILDKIT_ADDR: tcp://buildkitd.devops-tools.svc.cluster.local:1234
# One BuildKit per worker (the "buildkit" add-on); each image builds on
# the daemon its name hashes to. BUILDKIT_ADDR is the fallback.
BUILDKIT_POOL: buildkitd-pool.devops-tools.svc.cluster.local
MAX_PARALLEL_STEPS: "4"
# Longest a single test or build step may run. First builds on the Pis
# start with a cold cache (e.g. compiling a Python wheel from source).
STEP_TIMEOUT: 45m
# Nodes that must never run builds: e.g. a node whose kernel cannot
# enforce the build NetworkPolicy, or a GPU node reserved for inference.
BUILD_EXCLUDE_NODES: gpu-node
# One replica with the Recreate strategy never overlaps, so leader
# election only adds a way to crash when the API server is slow.
LEADER_ELECTION: "false"
CLUSTER_ISSUER: letsencrypt-prod
INGRESS_CLASS: nginx
STORAGE_CLASS: longhorn
# Labels added to every app volume claim; these put it in a Longhorn backup job.
# VOLUME_LABELS: recurring-job.longhorn.io/source=enabled,recurring-job-group.longhorn.io/backup-nightly=enabled
# Services without a Dockerfile are built from source with Railpack. The
# CLI image comes from `make railpack-image`; keep both versions in step.
RAILPACK_IMAGE: registry.example.lan:5000/rendimiento-railpack:0.40.0
RAILPACK_FRONTEND: ghcr.io/railwayapp/railpack-frontend:v0.40.0
# How a service with `gpu: 1` gets a GPU (here: an NVIDIA device plugin): its
# runtime class, the host's driver libraries (read-only) and a larger /dev/shm.
GPU_RESOURCE: nvidia.com/gpu
GPU_RUNTIME_CLASS: nvidia
GPU_HOST_PATHS: /usr/lib/aarch64-linux-gnu/nvidia
GPU_ENV: NVIDIA_VISIBLE_DEVICES=all;NVIDIA_DRIVER_CAPABILITIES=all;LD_LIBRARY_PATH=/usr/lib/aarch64-linux-gnu/nvidia:/usr/local/cuda/lib64
GPU_SHARED_MEMORY: 1Gi
# Failed builds, rolled-back releases, outages and recoveries are emailed
# here, through Resend (RESEND_API_KEY in the rendimiento-notify secret).
NOTIFY_EMAIL_TO: [email protected]
NOTIFY_LANG: en # en | es (Mexican Spanish)
NOTIFY_EMAIL_FROM: rendimiento <[email protected]>
# GET /api/public/stats: aggregate numbers for a public page (a portfolio,
# a status page), served only on the internal port 8081 (the Service's
# "stats" port, which the public ingress does not route), so it is
# reachable from inside the cluster but not from the internet.
PUBLIC_STATS: "true"
STATS_LISTEN: ":8081"
PUBLIC_STATS_TZ: America/New_York
# Finished runs' step logs move to S3-compatible storage (gzip; Garage
# here), which deletes them after LOG_RETENTION_DAYS. Credentials: the
# rendimiento-logs secret (a key that can only use this bucket).
LOG_ARCHIVE_ENDPOINT: garage.garage.svc.cluster.local:3900
LOG_ARCHIVE_BUCKET: rendimiento-logs
LOG_RETENTION_DAYS: "365"
# Visitor numbers from Umami (each app's Visits tab), read inside the
# cluster with a view-only user (the rendimiento-umami secret).
# UMAMI_URL: http://umami.umami.svc.cluster.local:3000
# UMAMI_PUBLIC_URL: https://umami.example.com
---
apiVersion: apps/v1
kind: Deployment
metadata:
name: rendimiento
namespace: rendimiento-system
spec:
replicas: 1
strategy: { type: Recreate }
selector:
matchLabels: { app: rendimiento }
template:
metadata:
labels: { app: rendimiento }
spec:
serviceAccountName: rendimiento
containers:
- name: rendimiento
image: registry.example.lan:5000/rendimiento:latest
imagePullPolicy: Always
envFrom:
- configMapRef: { name: rendimiento }
# CLOUDFLARE_API_TOKEN and DNS_TARGET; optional until DNS automation is wanted.
- secretRef: { name: rendimiento-dns, optional: true }
# RESEND_API_KEY for notification emails; optional (no email without it).
- secretRef: { name: rendimiento-notify, optional: true }
# LOG_ARCHIVE_ACCESS_KEY / LOG_ARCHIVE_SECRET_KEY for the log archive.
- secretRef: { name: rendimiento-logs, optional: true }
# UMAMI_USERNAME / UMAMI_PASSWORD (a view-only Umami user) for visitor numbers.
- secretRef: { name: rendimiento-umami, optional: true }
env:
- name: DB_PASSWORD
valueFrom: { secretKeyRef: { name: rendimiento-db, key: password } }
- name: DATABASE_URL
value: postgres://rendimiento:$(DB_PASSWORD)@rendimiento-db:5432/rendimiento?sslmode=disable
- name: SETUP_TOKEN
valueFrom: { secretKeyRef: { name: rendimiento-setup, key: token } }
ports:
- { name: http, containerPort: 8080 }
- { name: metrics, containerPort: 9090 }
- { name: stats, containerPort: 8081 }
readinessProbe:
httpGet: { path: /healthz, port: http }
periodSeconds: 10
livenessProbe:
httpGet: { path: /healthz, port: http }
initialDelaySeconds: 20
periodSeconds: 20
resources:
requests: { cpu: 100m, memory: 128Mi }
limits: { memory: 512Mi }
securityContext:
allowPrivilegeEscalation: false
readOnlyRootFilesystem: true
runAsNonRoot: true
runAsUser: 65532
runAsGroup: 65532
capabilities: { drop: [ALL] }
---
apiVersion: v1
kind: Service
metadata:
name: rendimiento
namespace: rendimiento-system
spec:
selector: { app: rendimiento }
ports:
- { name: http, port: 80, targetPort: http }
# In-cluster only: the public stats (the Ingress routes only "http").
- { name: stats, port: 8081, targetPort: stats }
---
apiVersion: networking.k8s.io/v1
kind: Ingress
metadata:
name: rendimiento
namespace: rendimiento-system
annotations:
cert-manager.io/cluster-issuer: letsencrypt-prod
nginx.ingress.kubernetes.io/ssl-redirect: "true"
nginx.ingress.kubernetes.io/force-ssl-redirect: "true"
# Live logs use server-sent events.
nginx.ingress.kubernetes.io/proxy-buffering: "off"
nginx.ingress.kubernetes.io/proxy-read-timeout: "3600"
nginx.ingress.kubernetes.io/proxy-body-size: 25m
spec:
ingressClassName: nginx
tls:
- hosts: [rendimiento.example.com]
secretName: rendimiento-tls
rules:
- host: rendimiento.example.com
http:
paths:
- path: /
pathType: Prefix
backend:
service: { name: rendimiento, port: { name: http } }
deploy/rbac.yaml
apiVersion: v1
kind: ServiceAccount
metadata:
name: rendimiento
namespace: rendimiento-system
---
# Cluster-wide because each app gets its own namespace. The controller
# refuses namespaces and hostnames it does not own (see checkOwnership).
apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRole
metadata:
name: rendimiento
rules:
- apiGroups: [rendimiento.ai]
resources: [apps, apps/status, apps/finalizers, addons, addons/status, addons/finalizers]
verbs: ["*"]
# Add-ons are applied as the rendimiento-addons identity (below); the
# platform itself may only act as it, not hold its rights.
- apiGroups: [""]
resources: [serviceaccounts]
resourceNames: [rendimiento-addons]
verbs: [impersonate]
- apiGroups: [""]
resources: [namespaces, services, persistentvolumeclaims, secrets]
verbs: [get, list, watch, create, update, patch, delete]
- apiGroups: [""]
resources: [pods]
verbs: [get, list, watch]
- apiGroups: [apps]
resources: [deployments]
verbs: [get, list, watch, create, update, patch, delete]
- apiGroups: [networking.k8s.io]
resources: [ingresses]
verbs: [get, list, watch, create, update, patch, delete]
- apiGroups: [batch]
resources: [cronjobs]
verbs: [get, list, watch, create, update, patch, delete]
# Post-deploy tasks run as Jobs in the app's namespace; their logs are kept.
- apiGroups: [batch]
resources: [jobs]
verbs: [get, list, watch, create, delete]
- apiGroups: [""]
resources: [pods/log]
verbs: [get]
- apiGroups: [cert-manager.io]
resources: [certificates]
verbs: [get, list, watch]
- apiGroups: [""]
resources: [events]
verbs: [create, patch]
# Read-only, for the environment page.
- apiGroups: [""]
resources: [nodes]
verbs: [get, list]
- apiGroups: [metrics.k8s.io]
resources: [nodes]
verbs: [get, list]
- apiGroups: [networking.k8s.io]
resources: [ingressclasses, networkpolicies]
verbs: [get, list]
- apiGroups: [storage.k8s.io]
resources: [storageclasses]
verbs: [get, list]
- apiGroups: [apiextensions.k8s.io]
resources: [customresourcedefinitions]
verbs: [get, list]
- apiGroups: [cert-manager.io]
resources: [clusterissuers]
verbs: [get, list]
- apiGroups: [longhorn.io]
resources: [volumes, recurringjobs, backuptargets]
verbs: [get, list]
# Read-only, for the services catalog (who calls what).
- apiGroups: [apps]
resources: [statefulsets]
verbs: [get, list]
---
apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRoleBinding
metadata:
name: rendimiento
roleRef:
apiGroup: rbac.authorization.k8s.io
kind: ClusterRole
name: rendimiento
subjects:
- kind: ServiceAccount
name: rendimiento
namespace: rendimiento-system
---
# CI pods: create, watch, read logs, clean up.
apiVersion: rbac.authorization.k8s.io/v1
kind: Role
metadata:
name: rendimiento-builds
namespace: rendimiento-builds
rules:
- apiGroups: [""]
resources: [pods, secrets]
verbs: [get, list, watch, create, delete]
- apiGroups: [""]
resources: [pods/log]
verbs: [get]
---
apiVersion: rbac.authorization.k8s.io/v1
kind: RoleBinding
metadata:
name: rendimiento-builds
namespace: rendimiento-builds
roleRef:
apiGroup: rbac.authorization.k8s.io
kind: Role
name: rendimiento-builds
subjects:
- kind: ServiceAccount
name: rendimiento
namespace: rendimiento-system
---
# Leader election.
apiVersion: rbac.authorization.k8s.io/v1
kind: Role
metadata:
name: rendimiento-leader
namespace: rendimiento-system
rules:
- apiGroups: [coordination.k8s.io]
resources: [leases]
verbs: [get, list, watch, create, update, patch]
---
apiVersion: rbac.authorization.k8s.io/v1
kind: RoleBinding
metadata:
name: rendimiento-leader
namespace: rendimiento-system
roleRef:
apiGroup: rbac.authorization.k8s.io
kind: Role
name: rendimiento-leader
subjects:
- kind: ServiceAccount
name: rendimiento
namespace: rendimiento-system
---
# Add-ons install cluster software (Helm charts such as Longhorn: CRDs,
# ClusterRoles, DaemonSets), which needs cluster-admin, as ArgoCD has had.
# No pod runs as this account: the platform impersonates it only while
# applying add-ons, so the audit log shows exactly what add-ons changed.
apiVersion: v1
kind: ServiceAccount
metadata:
name: rendimiento-addons
namespace: rendimiento-system
automountServiceAccountToken: false
---
apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRoleBinding
metadata:
name: rendimiento-addons
roleRef:
apiGroup: rbac.authorization.k8s.io
kind: ClusterRole
name: cluster-admin
subjects:
- kind: ServiceAccount
name: rendimiento-addons
namespace: rendimiento-system
deploy/networkpolicy.yaml
# CI steps run code from the repos being built (tests, Dockerfile RUN lines).
# Keep them away from everything else: they may reach DNS, the shared
# BuildKit daemon and the public internet (git, package registries), but
# not other apps, databases, the Kubernetes API or the home network.
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
name: isolate-builds
namespace: rendimiento-builds
labels:
app.kubernetes.io/managed-by: rendimiento
spec:
podSelector: {}
policyTypes: [Ingress, Egress]
ingress: [] # nothing connects to build pods
egress:
- to:
- namespaceSelector:
matchLabels: { kubernetes.io/metadata.name: kube-system }
podSelector:
matchLabels: { k8s-app: kube-dns }
ports:
- { protocol: UDP, port: 53 }
- { protocol: TCP, port: 53 }
- to:
- namespaceSelector:
matchLabels: { kubernetes.io/metadata.name: devops-tools }
podSelector:
matchLabels: { app: buildkitd }
ports:
- { protocol: TCP, port: 1234 }
- to:
- ipBlock:
cidr: 0.0.0.0/0
except:
- 10.0.0.0/8 # cluster pods and services (10.42/16, 10.43/16)
- 172.16.0.0/12
- 192.168.0.0/16 # home network, including the router
- 169.254.0.0/16 # link-local / cloud metadata
- 100.64.0.0/10 # carrier-grade NAT
Permissions, and why they are shaped this way¶
The platform's own service account (rendimiento) can:
- manage the
AppandAddonresources; - manage namespaces, Deployments, Services, Ingresses, CronJobs, PVCs and Secrets cluster-wide, because every app gets its own namespace. The controller refuses namespaces and hostnames it does not own (ownership checks);
- read nodes, metrics, ingress classes, storage classes, CRDs, cluster issuers and StatefulSets, for the Environment and Services pages;
- create and delete pods and secrets in
rendimiento-builds(CI); - impersonate the
rendimiento-addonsservice account, and nothing more.
rendimiento-addons is bound to cluster-admin, because installing software like Longhorn means creating CRDs, ClusterRoles and DaemonSets, exactly what ArgoCD needed. No pod runs as it: only the add-on controller acts as it, so the Kubernetes audit log shows every add-on change under that name. See Security model.
Secrets you create once¶
Secret (namespace rendimiento-system) |
Keys | Used for |
|---|---|---|
rendimiento-db |
password |
Postgres password (the Deployment builds DATABASE_URL from it) |
rendimiento-setup |
token |
One-time token that protects the GitHub App setup page |
rendimiento-dns (optional) |
CLOUDFLARE_API_TOKEN, DNS_TARGET |
DNS automation. DNS_TARGET=auto follows the home network's public IP. |
rendimiento-github |
created by the setup flow | The GitHub App's ID, private key, webhook secret and OAuth client |
kubectl -n rendimiento-system create secret generic rendimiento-db --from-literal=password="$(openssl rand -hex 24)"
kubectl -n rendimiento-system create secret generic rendimiento-setup --from-literal=token="$(openssl rand -hex 16)"
# a Cloudflare token with Zone:DNS:Edit on your zones, entered without echoing it:
read -rs CF && kubectl -n rendimiento-system create secret generic rendimiento-dns \
--from-literal=CLOUDFLARE_API_TOKEN="$CF" --from-literal=DNS_TARGET=auto; unset CF
The GitHub App¶
rendimiento creates its own GitHub App with the manifest flow, so nobody fills in GitHub's forms by hand:
sequenceDiagram
actor You
participant R as rendimiento
participant G as GitHub
You->>R: open /api/setup/github?token=… (the rendimiento-setup token)
R->>You: a page that posts the App manifest to GitHub
You->>G: create the App (name, permissions, webhook URL)
G->>R: redirect to /api/setup/github/callback?code=…
R->>G: exchange the code for the App's credentials
R->>R: store them in the rendimiento-github secret
You->>G: install the App on your account (all or chosen repos)
The manifest asks for these repository permissions: Contents read & write (read code, push onboarding branches), Pull requests read & write (open onboarding PRs), Checks read & write (report CI status), Metadata read, Issues read & write and Commit statuses read (both for the Renovate add-on). It subscribes to push events. Login to the dashboard uses the same App's OAuth, and only users in ALLOWED_USERS get a session.
Changing the App's permissions later
Change them at https://github.com/settings/apps/<app-name>/permissions, then accept the new permissions on the installation (https://github.com/settings/installations → Configure). Until accepted, installations keep the old ones; the Add-ons page warns when Renovate lacks what it needs.
Shipping a new version of rendimiento¶
make test-remote # the full test suite, on a worker node
make image # build and push registry.example.lan:5000/rendimiento:latest on the BuildKit pool
make deploy # kubectl apply -k deploy (only needed when deploy/ or the CRDs changed)
kubectl -n rendimiento-system rollout restart deploy/rendimiento
kubectl -n rendimiento-system rollout status deploy/rendimiento
The Deployment pulls :latest with imagePullPolicy: Always and uses the Recreate strategy (one replica, never two at once), so a restart takes about 20 seconds during which the UI and webhooks are unavailable. Apps keep running: they do not depend on the platform being up. GitHub retries webhooks that fail.
On restart the platform:
- applies database migrations (
internal/store/migrations), each once; - requeues CI runs a previous process left running (up to two retries) and deletes leftover build pods;
- closes Renovate runs that were interrupted;
- starts the controllers, the CI worker, the Renovate scheduler, the add-on sync from git, and dynamic DNS.
:latest has no history
Rolling back the platform means rebuilding the previous commit. Tagging each build with its commit is on the roadmap; until then, git checkout <good commit> && make image is the rollback.
rendimiento builds itself¶
rendimiento's repository is onboarded like any app. Its rendimiento.yaml has the book (docs, a service) and the platform's own image (builds: platform), so every push runs:
platform:test:hack/ci-test.sh(generate, vet and the Go tests, with envtest and a throwaway Postgres) ingolang, with a kept cache;platform:build: the Dockerfile, which also checks the UI's types and translations, pushed as<registry>/rendimiento-ai-platform:<commit>.
Pull requests get the same check without a release. On the default branch the image's digest is kept with the release, next to the book's. It is not deployed by itself yet: shipping is still the steps above, which build the same Dockerfile. Letting a release update the platform (a rolling update that keeps the old version until the new one is ready, so a broken image cannot take the platform down) is the next step.
A field the running platform doesn't know
The platform reads rendimiento.yaml strictly. A change that adds a field to it (as builds: did) must be shipped before the commit that uses the field is pushed, or that push's run fails to read its own spec.
The build side¶
| What | Where it comes from |
|---|---|
| BuildKit daemons | the buildkit add-on (p0dxD/gitops/buildkit/): a StatefulSet with one daemon per worker node |
| Railpack CLI image | make railpack-image from deploy/railpack/Dockerfile, checksum-pinned |
| Railpack frontend | pulled by BuildKit from ghcr.io/railwayapp/railpack-frontend at the pinned version |
| Clone and build client images | alpine/git, moby/buildkit (the client must match the daemon's version) |