Skip to content

Problems

When something goes wrong inside rendimiento itself, it shows up on the Problems page (in the top bar, with the number of open problems). Examples: a push whose rendimiento.yaml was rejected, an alert email that could not be sent, a DNS or GitHub call that failed, an add-on that would not sync.

Outages of your apps are not problems: they're on each app's Reliability tab.

What you see

  • Each problem has its level (warning or error), its message and the error it reported, the app it concerns (if any), and the part of the platform that reported it.
  • Repeats are grouped. The same problem happening again (even with a different commit or number in it) adds to its count: "40 times since 2:10 PM, last 3 min ago".
  • A problem is open for 24 hours after it last happened, unless it is resolved or dismissed first:
    • Resolved means rendimiento saw proof it's fixed, and says what: a rejected rendimiento.yaml is resolved by the next push accepted on the same branch ("fixed by c6eae54"), a failed email by the next email sent.
    • Dismiss hides a problem by hand.
    • Either way, it reopens if it happens again.
  • Problems are kept for 30 days after they last happened.
  • Problems about an app also appear as a banner on the app's page.

A rejected rendimiento.yaml

If a push has a rendimiento.yaml that cannot be used (a typo, a host listed twice), nothing is built or released, and the app keeps running its current release. The push still shows up:

  • as a failed run on the app's Runs tab, whose page says why;
  • as a failed check (red ✗) on the commit in GitHub;
  • in an email, for the default branch (see Notifications);
  • on the Problems page and the app's banner.

Fix the file and push again: once that push is accepted, the problem is marked resolved.

How it works

flowchart LR
    code[platform code] -- log warning / error --> h[problems handler]
    h --> out[the platform's log, as before]
    h -- queue, never blocks --> rec[Recorder]
    rec -- upsert by fingerprint --> pg[(problems table)]
    pg --> api[/api/problems/] --> ui[Problems page, nav count, app banner]
  • The platform's logger passes every record to its normal output; warnings and errors are also queued for the Problems page (internal/problems). Recording never slows the code that logs: if the queue is full, problems are counted and dropped.
  • The fingerprint of a problem is its level, component, app, message and error, with numbers and hashes blanked, so repeats land on one row.
  • Some warnings are left out on purpose: Kubernetes write conflicts that the controllers retry immediately, apps going down (the Reliability tab tracks those), and work stopped by a shutdown.
  • API: GET /api/problems (?app=, ?dismissed=1), GET /api/problems/count, POST /api/problems/{id}/dismiss.