Scaling¶
Scaling here means handling more apps, more builds, more users and more clusters. For each part of the platform, this page says what limits it today, how you will notice when it hits that limit, and what to do about it.
flowchart TB
subgraph today[One Deployment today]
api[API + UI] --- worker[CI worker] --- ctrl[Controllers] --- sched[Schedulers]
end
subgraph later[Split when needed]
api2[API x N<br/>stateless] --> pg[(Postgres)]
worker2[Workers x N<br/>SKIP LOCKED] --> pg
ctrl2[Controllers<br/>leader-elected] --> k8s[(Kubernetes API)]
api2 -. events .-> bus[LISTEN/NOTIFY]
worker2 -. events .-> bus
end
today --> later
Builds¶
Today: four BuildKit daemons (buildkitd-0..3) run on the workers. Each image is sent to one daemon by rendezvous hashing, so that daemon's cache stays warm. MAX_PARALLEL_STEPS=4 caps the number of concurrent steps across the platform. The registry cache (mode=max) lets any daemon reuse layers another daemon built.
Limits: CPU and memory on the Pis, then the registry's disk and network.
Signs: runs sit in queued for a long time, steps hit STEP_TIMEOUT, and daemons are OOM-killed.
Levers, cheapest first:
- Add replicas to the
buildkitadd-on (inp0dxD/gitops). The pool finds new daemons through DNS SRV records, and only about 1/N of the images move to a different daemon. - Raise
MAX_PARALLEL_STEPSonce the daemons can take more work. - Give daemons more memory, and set
max-parallelismin buildkitd's config. - Make the cap per daemon rather than global: the runner would track steps per daemon.
- Add an amd64 or cloud builder for heavy stacks (Java, Rust), chosen by a
build.platformfield or by the size of the stack.
Queue and workers¶
Today: a Postgres table claimed with FOR UPDATE SKIP LOCKED, and one worker loop inside the platform.
Limits: thousands of runs a minute, far beyond this cluster. Postgres is not the bottleneck; build capacity is.
To scale: run the worker as its own Deployment (rendimiento worker, a subcommand that only starts the runner) with several replicas. SKIP LOCKED already makes concurrent claims safe. Add a heartbeat column so a crashed worker's runs are requeued without waiting for a restart.
Controllers¶
Today: one replica, LEADER_ELECTION=false.
To scale: turn leader election on and run two replicas. The standby takes over within about 15 seconds. Controllers scale up (more concurrent reconciles, via MaxConcurrentReconciles), not out. One leader per controller is the Kubernetes norm and handles thousands of objects.
Live updates across replicas¶
Today: the API is almost stateless: sessions live in Postgres. The exception is events.Hub, which is in memory. A browser connected to replica A would miss events published by a worker on replica B.
To scale: send events through Postgres LISTEN/NOTIFY, with no new infrastructure. Every replica listens and republishes to its own SSE clients. The hub's Publish and Subscribe API stays the same; only the transport changes. After that, run the API with several replicas behind the Service.
Database¶
Today: one Postgres pod in rendimiento-system on Longhorn storage. Build logs are rows while a run is going, then move to object storage (the log archive).
Signs: the database grows mostly from run history and uptime data.
Levers:
- Add retention: delete runs older than N days, or keep only the last N runs per app.
- Use a managed or operator-run Postgres (CloudNativePG) for backups, point-in-time recovery and replicas.
Registry¶
Today: registry.example.lan:5000 over plain HTTP, with no garbage collection.
Levers: run registry garbage collection on a schedule, keep only the last N digests per app (rollback only needs recent ones), or move to Harbor or Zot for TLS, retention policies, replication and vulnerability scanning.
Many clusters¶
Today: the platform applies objects to the cluster it runs in.
Direction: a pull-based agent per cluster:
flowchart LR
cp[rendimiento<br/>control plane] -->|desired App objects| q[(per-cluster queue)]
a1[agent on cluster A] -->|pull + report status| cp
a2[agent on cluster B, cloud] -->|pull + report status| cp
a1 --> k1[(cluster A)]
a2 --> k2[(cluster B)]
- The agent is the existing App controller with a different source: it fetches desired
Apps from the control plane's API instead of reading them locally. - Clusters behind NAT need only outbound HTTPS, with no inbound ports.
- Add
environment:ortarget:torendimiento.yamlto choose the cluster, and use promotion (dev → prod) to move a release digest between targets.
Users and teams¶
Today: a GitHub login allowlist (ALLOWED_USERS), where everyone can do everything.
Next: roles (viewer, deployer, admin) per app, taken from GitHub team membership; an audit log table recording who deployed, rolled back or changed a secret; and per-namespace RBAC for the platform itself instead of broad cluster rights.
Numbers to watch¶
| Metric | Why |
|---|---|
| queue wait (claimed − created) | whether build capacity is enough |
| step duration per service, p50/p95 | whether caching works and whether builds are getting slower |
| reconcile errors per controller | misconfiguration or cluster trouble |
DB size, and the size of the steps table's logs |
when to archive logs |
| registry disk use | when to garbage-collect |
These metrics are not exported yet. Adding them is in Refactoring.