---
title: "Tearing Down Domains at Scale: Timing, Email Coupling, and the Owner Question"
canonical: https://dxdev.com/blog/2026-08-22_automated-domain-teardown-engine/
datePublished: 2026-06-24
---
There were **65 expired domains** sitting in the backlog when I shipped the teardown process on June 24. Clearing them was the visible part. The harder part was deciding exactly when the system should act, what could be allowed to fail, and which person should receive the message when it did.

A domain expiry is not one row changing state. In our system it touches Cloudflare, the relay, and the database. Remove one piece and leave the other two alive, and we have made a future support problem that is harder to identify because it is now only partly visible.

The first version of the idea sounded straightforward: find expired sites, remove their domain resources, notify the customer. The 65-domain backlog forced the actual questions into view. What counts as expired when a scheduled process runs around a date boundary? Does a mail problem stop the lifecycle monitor? And which contact field represents the person accountable for a site?

## Expired does not mean act immediately

My first instinct was to make teardown eligible as soon as the site crossed its expiry time. That condition looks tidy in a query. It is also wrong for the customer experience.

A domain can become expired during a day when someone is still responding to renewal notices, reconciling a card failure, or simply reading email in a different time zone. Tearing down the same day turns the expiry timestamp into a trapdoor.

I changed the rule to run **the day after** a site expires. The expiration flow warns the customer about the domain. The automated teardown then handles the expired site on the following day. The delay gives the warning a real window to be seen, and it makes the selection rule legible when we inspect it later.

I considered two alternatives. Same-day teardown lost because it made the notice and the destructive action effectively concurrent. A longer grace period lost because it would keep dead infrastructure alive without a policy reason in the actual expiry process. The next-day rule was specific enough to monitor and conservative enough not to surprise someone whose domain had just crossed the line.

The diagnostic path here was not an error message. It was the mismatch between a technically valid condition and the behavior it would cause. “Expired at or before now” looked complete until I traced it through the email timing and the scheduler. The real rule was temporal, but it was also communicative.

## Cloudflare cleanup cannot depend on a mailbox

The teardown has three concrete jobs: remove the Cloudflare configuration, remove the relay configuration, and update the database. Those are the cleanup path. Sending an email is not.

That distinction matters because the lifecycle monitor has one primary obligation: when a site is past the threshold, it must tear the site down. If the mail provider has a transient failure, a bad recipient record, or any other delivery problem, it must not leave the domain active or block the monitor from reaching the next expired site.

I made notification **best effort**. The monitor completes the teardown across Cloudflare, the relay, and the database. It then attempts to email the owner, with the mail outcome recorded as part of the run rather than as a dependency that can invalidate cleanup.

The other design was a single all-or-nothing transaction. It had an appealing shape: no notification, no completed teardown. But that puts the least controllable integration in charge of the thing we actually need to guarantee. A mail failure should be observable and repairable. It should not leave a domain active with a broken lifecycle state.

That is the coupling test I now apply to automation like this. Ask which failure is permitted to delay the system's central obligation. Here, failed cleanup must be loud because resources may still be serving. Failed email must be loud too, but it cannot reverse the cleanup decision or freeze the monitor.

## The primary contact was the wrong recipient

The last problem looked like field selection. It was an ownership decision.

The system had a primary contact and an owner. It would have been easy to notify the primary contact because that field is often present and familiar. But the person who happens to be the first contact record is not necessarily the person accountable for the site and its domain. For a destructive lifecycle event, familiarity is not enough.

The teardown email goes to the **owner**. That is the record that represents responsibility for the site. The rule is not “send to whoever we normally email.” It is “send to the person who owns the thing being removed.”

I am treating that as a data-model lesson, not an email preference. Billing, support, operations, and ownership can all exist on the same account. An automated action needs to name the role it is addressing, then use the corresponding field. Choosing the most convenient field is how a notification becomes a record that nobody was actually responsible for receiving.

## The backlog was the first production run

After the new path was in place, I cleared the 65 already-expired domains through it. That was more than housekeeping. A backlog exercises the sequence repeatedly: identify the expired site, remove Cloudflare, remove the relay, update the database, and notify the owner without making notification a gate.

I did not treat the completed backlog as proof that the scheduler was done. I left a follow-up check in place for June 27 to confirm that the automatic run fired correctly on its own. A manual clearance proves the teardown path can work. It does not prove the scheduled trigger will select the right records on the right day.

There is also a separate follow-up for the downgrade-at-checkout case. That is exactly the kind of adjacent lifecycle path that can bypass a clean implementation if we pretend every domain reaches expiry through the same sequence.

The final system is not complicated. A site expires. The customer is warned. The next day, the monitor tears down Cloudflare, relay, and database state. The owner receives a best-effort notice. Each part has a distinct job and a distinct failure mode.

The 65-domain backlog made those boundaries impossible to ignore. The real work was not deleting domains. It was refusing to let timing, email delivery, or the wrong contact field quietly decide what deletion meant.
