---
title: "The Signal Is in the Metadata, Not the Log"
canonical: https://dxdev.com/blog/2026-09-07_ci-annotation-not-log/
datePublished: 2026-08-29
---
The alert said the failure was a cache issue. It wasn't. The job had run zero steps, in two seconds, twice in a row, and the actual reason was sitting in a GitHub check-run annotation our triage script never read.

## The alert that lied by omission

`ci-alert/triage` fires on red CI. That morning it fired three times between 9:47 and 1:26, each time producing a plausible-sounding cause. All three were wrong in the same way: they were guesses generated from an error blob, not from anything GitHub had actually told us.

Here's what `ci_alert.py` was doing. On a failed run, it called `gh run view --log-failed` to pull the failing step's log and fed that text to a triage model. That works fine when a step actually ran and printed something. It does not work when the job never started, because there is no failing step and no log. `gh run view --log-failed` returns nothing useful in that case, in this instance an Azure `BlobNotFound` document, since the log artifact for a job that never ran doesn't exist. The script had no check for that. It handed the model a storage-service error page and asked it to name a cause. The model, doing what it's told, invented one. That's how we got a plausible cache-related root cause for a job that had nothing to do with caching.

## What actually happened, found by hand

I went and pulled the real data instead of trusting the alert:

```
gh api repos/OWNER/REPO/commits/SHA/check-runs \
  --jq '.check_runs[] | {name,conclusion,output_title,output_summary}'
```

That gave a check run with `conclusion: failure` but empty `output_summary`, so I went one level deeper to the annotations on that check run:

```
gh api repos/OWNER/REPO/check-runs/CHECK_RUN_ID/annotations
```

And there it was, `annotation_level: failure`, with the message GitHub actually wrote:

> "The job was not started because recent account payments have failed or your spending limit needs to be increased. Please check the 'Billing & plans' section in your settings"

Zero steps recorded. Re-ran it as attempt 2, same result: 2 seconds, 0 steps, identical annotation. Not transient, not a code problem. GitHub Actions had been switched off at the account level because a usage allowance had been hit.

## The wrong turn: chasing the billing API before the annotation

Before I found the annotation endpoint, I burned real time on the wrong lead. My first instinct was that this had to be a billing-API question, so I went straight for `gh api users/OWNER/settings/billing/actions`. That came back `410 This endpoint has been moved`. I tried the replacement, `users/OWNER/settings/billing/usage`, and got a different failure: `invalid API endpoint: "C:/Program Files/Git/users/OWNER/settings/billing/usage"`, because Git Bash on Windows was rewriting the leading slash into a filesystem path. I stripped the slash, got the endpoint working, and pulled back a usage report. It confirmed the account had hit exactly 2000.0 minutes for the month, all of it from one repo's deploy-on-every-push workflow. That number was real and useful for the follow-up conversation about trimming that workflow, but it told me nothing about why *this specific job* had failed. I'd spent the better part of the diagnosis chasing account-level usage data when the actual cause was already sitting in the check-run annotation for the commit in question, one API call away, in GitHub's own words. The billing-usage detour was a real cost: it's the reason this ticket took over five hours instead of the twenty minutes the annotation lookup actually took once I ran it.

## The fix

The triage function had a comment that already named the shape of this bug: `startup_failure with no failing step` was called out as a known case, but the code path that would have handled it didn't call the check-run annotation lookup. The fix was to make the zero-step case route there before falling back to logs:

```python
text = _gh_cli(["run", "view", str(run_id), "--repo", owner_repo, "--log-failed"])
if text.strip():
    return _error_window(text, cap)
jobs = _failing_jobs(owner_repo, run_id)
if _never_started(jobs):
    ann = _job_annotations(owner_repo, jobs, cap)
    if ann:
        return ann
return _whole_job_logs(owner_repo, jobs, cap)
```

`_never_started` checks for the zero-steps signature and, when true, pulls annotations before ever touching a log. I also tightened the triage prompt itself: if the evidence is a GitHub annotation saying the job ran no steps, the model is now told to report that reason in GitHub's own words as authoritative, rather than inferring a cause from repo lines that are just environment noise (DB roles, connection strings, checkout, teardown).

Cost stayed at one `gh` call in the common case, since both extra lookups sit past the `--log-failed` early return, so a run with a normal failing step still makes exactly one call.

Verified end to end against a real red run: same annotation, same message, now surfaced directly instead of being paraphrased from a blob. Root cause reported correctly: GitHub had halted Actions on the account over payments, not a code failure. We moved the offending repo to a separate billing account and left the rest waiting on the month to flip. The lesson that stuck: when a CI job fails with nothing in its log, the log isn't missing data, it's telling you to stop looking at logs. GitHub already wrote down why, in a different API, and the debugging time went to finding that out instead of trusting the first plausible story.
