The alert said the failure was a cache issue. It wasn’t. The job had run zero steps, in two seconds, twice in a row, and the actual reason was sitting in a GitHub check-run annotation our triage script never read.
The alert that lied by omission
ci-alert/triage fires on red CI. That morning it fired three times between 9:47 and 1:26, each time producing a plausible-sounding cause. All three were wrong in the same way: they were guesses generated from an error blob, not from anything GitHub had actually told us.
Here’s what ci_alert.py was doing. On a failed run, it called gh run view --log-failed to pull the failing step’s log and fed that text to a triage model. That works fine when a step actually ran and printed something. It does not work when the job never started, because there is no failing step and no log. gh run view --log-failed returns nothing useful in that case, in this instance an Azure BlobNotFound document, since the log artifact for a job that never ran doesn’t exist. The script had no check for that. It handed the model a storage-service error page and asked it to name a cause. The model, doing what it’s told, invented one. That’s how we got a plausible cache-related root cause for a job that had nothing to do with caching.
What actually happened, found by hand
I went and pulled the real data instead of trusting the alert:
gh api repos/OWNER/REPO/commits/SHA/check-runs \ --jq '.check_runs[] | {name,conclusion,output_title,output_summary}'That gave a check run with conclusion: failure but empty output_summary, so I went one level deeper to the annotations on that check run:
gh api repos/OWNER/REPO/check-runs/CHECK_RUN_ID/annotationsAnd there it was, annotation_level: failure, with the message GitHub actually wrote:
“The job was not started because recent account payments have failed or your spending limit needs to be increased. Please check the ‘Billing & plans’ section in your settings”
Zero steps recorded. Re-ran it as attempt 2, same result: 2 seconds, 0 steps, identical annotation. Not transient, not a code problem. GitHub Actions had been switched off at the account level because a usage allowance had been hit.
The wrong turn: chasing the billing API before the annotation
Before I found the annotation endpoint, I burned real time on the wrong lead. My first instinct was that this had to be a billing-API question, so I went straight for gh api users/OWNER/settings/billing/actions. That came back 410 This endpoint has been moved. I tried the replacement, users/OWNER/settings/billing/usage, and got a different failure: invalid API endpoint: "C:/Program Files/Git/users/OWNER/settings/billing/usage", because Git Bash on Windows was rewriting the leading slash into a filesystem path. I stripped the slash, got the endpoint working, and pulled back a usage report. It confirmed the account had hit exactly 2000.0 minutes for the month, all of it from one repo’s deploy-on-every-push workflow. That number was real and useful for the follow-up conversation about trimming that workflow, but it told me nothing about why this specific job had failed. I’d spent the better part of the diagnosis chasing account-level usage data when the actual cause was already sitting in the check-run annotation for the commit in question, one API call away, in GitHub’s own words. The billing-usage detour was a real cost: it’s the reason this ticket took over five hours instead of the twenty minutes the annotation lookup actually took once I ran it.
The fix
The triage function had a comment that already named the shape of this bug: startup_failure with no failing step was called out as a known case, but the code path that would have handled it didn’t call the check-run annotation lookup. The fix was to make the zero-step case route there before falling back to logs:
text = _gh_cli(["run", "view", str(run_id), "--repo", owner_repo, "--log-failed"])if text.strip(): return _error_window(text, cap)jobs = _failing_jobs(owner_repo, run_id)if _never_started(jobs): ann = _job_annotations(owner_repo, jobs, cap) if ann: return annreturn _whole_job_logs(owner_repo, jobs, cap)_never_started checks for the zero-steps signature and, when true, pulls annotations before ever touching a log. I also tightened the triage prompt itself: if the evidence is a GitHub annotation saying the job ran no steps, the model is now told to report that reason in GitHub’s own words as authoritative, rather than inferring a cause from repo lines that are just environment noise (DB roles, connection strings, checkout, teardown).
Cost stayed at one gh call in the common case, since both extra lookups sit past the --log-failed early return, so a run with a normal failing step still makes exactly one call.
Verified end to end against a real red run: same annotation, same message, now surfaced directly instead of being paraphrased from a blob. Root cause reported correctly: GitHub had halted Actions on the account over payments, not a code failure. We moved the offending repo to a separate billing account and left the rest waiting on the month to flip. The lesson that stuck: when a CI job fails with nothing in its log, the log isn’t missing data, it’s telling you to stop looking at logs. GitHub already wrote down why, in a different API, and the debugging time went to finding that out instead of trusting the first plausible story.