---
title: "Tests That Know When to Get Out of the Way"
canonical: https://dxdev.com/blog/2026-06-23_graceful-test-degradation/
datePublished: 2026-06-23
---
At 2:32 PM, a test command that had been broken for months finally started running in CI, and it immediately found the tests we had stopped seeing.

The first problem was one character in the pnpm setup step. The setup action interpreted the resulting configuration as a duplicate version pin and rejected it before the test runner ever had a chance to start. The test job existed. The workflow looked like it had a test step. But the step that mattered was not executing.

That distinction cost us more than a green check. A CI system can report a clean-looking path while an earlier setup failure prevents the thing you think it is validating from running at all. We had a test suite in the repository, a workflow that named the suite, and stale failures waiting behind a broken setup action. None of that added up to tested code.

Once I removed the duplicate pin, the setup action accepted the workflow and the runner reached the tests for the first time in a long time. CI went red immediately. That was progress.

For months I looked at a green CI badge and read it as "the tests pass," when the honest sentence was "the tests did not run." I don't know how many commits shipped in that window on the strength of a check that was never actually checking anything. That is the part I'd rather have caught myself, before a broken setup step had to surface it for me.

## The red build was doing its job

The first run surfaced two different kinds of debt. Some tests were simply stale. They described behavior that the code no longer had, so we corrected them. The more interesting failures were environment-dependent. A database-backed suite expected a database. Another suite expected the vault. The CI runner had neither.

The diagnostic path was straightforward once the actual test step was alive. I separated the stale expectations from failures that began at resource access. The stale tests failed because their expected behavior was obsolete. The database and vault suites failed because their prerequisites did not exist in that runner. They were not proving a regression in application behavior. They were proving that an integration suite cannot integrate with infrastructure that was never provided.

Updating an obsolete assertion restores a useful test. Treating a missing database as an application failure turns an optional integration check into a permanent source of noise.

## We made the prerequisite explicit

The fix made the boundary explicit instead of loosening what every test was allowed to assume.

The database-dependent suite now checks whether its database resource is available. The vault-dependent suite does the same for the vault. If the resource is absent, the suite skips itself. The ordinary tests continue to run. In an environment where the database or vault is deliberately present, the dependent suite runs too.

The sequence matters:

1. The workflow sets up pnpm without a duplicate version pin.
2. The test runner starts.
3. Tests with no external prerequisites execute normally.
4. A database- or vault-dependent suite checks for its required resource.
5. It runs when the resource exists, or records a skip when it does not.

A skip is not a pass. That is why it works. A pass says the test executed and the expected behavior held. A skip says the test did not run because its declared environment was unavailable. Those are different facts, and the CI output needs to preserve the difference.

If a resource is meant to be present and its connection fails, that should still fail the suite. We are not swallowing operational errors. We are only avoiding the false claim that a test can validate an integration that the runner was never configured to supply.

## The alternatives that lost

We considered three other directions, all of which would have been worse for this job.

The first was to provision every external dependency inside every CI run. That can be the right choice for a dedicated integration pipeline. It was not the right first response here. A database and vault setup would add moving parts to a check that otherwise needs to be fast and reliable. It would also make the test workflow responsible for infrastructure that the code under test did not need for most of its coverage.

The second was to mock the database and vault everywhere. That would keep the runner green, but it would erase the distinction we needed. A mock tests behavior against a substitute. It cannot show that the real dependency is present or usable. For a resource-specific test, a universal fake turns an integration test into a unit test with a misleading name.

The third was to leave the tests failing and mentally classify them as environmental. That is how broken signals become background noise. Once a build contains failures that nobody expects to fix, it becomes much easier to miss the failure that matters. We had already seen the larger version of that problem, where the test step itself was not running and the workflow did not force the issue.

Self-skipping was the narrowest fix. It says exactly what the runner knows: this suite requires a resource, that resource is absent, and the suite has not made a claim about behavior. It keeps the default CI path honest without pretending that every test can run in every environment.

## The test boundary is part of the design

I used to think of graceful degradation mainly as an application concern. A feature should handle a missing service without crashing the whole request. The same rule applies to test infrastructure.

A test suite is also a program with dependencies. It needs a defined behavior when those dependencies are unavailable. Failing by default is sometimes correct, especially when the dependency is part of the environment the pipeline promises to build. Of the two failure modes, silently passing without execution does more damage than hanging, since a hang at least stops the pipeline and forces someone to look. Skipping an explicitly resource-bound suite gives a third, more precise option: it names the missing prerequisite instead of pretending the run was either a pass or a failure.

The one-character typo was the small defect. The larger problem was allowing CI to imply coverage it was not producing. Fixing the setup action exposed that gap. Giving resource-bound suites a truthful skip path preserved the line between a successful test, an unavailable environment, and a real failure.

The badge on the repository looked identical before and after this fix. What changed is underneath it: a green check now reflects the pnpm setup succeeding, the test runner actually starting, and every resource-bound suite reporting its own real status, run or skipped, rather than a setup failure nobody had connected to the word "tests."
