The manus-watch commit says it missed 152 real credits. The commit before it says a shared alert key ate a real 28-credit charge. Two separate bugs in a guard built to do one job: tell us when the free ride ended.
The setup
The Manus free-usage promo was scheduled to end on 2026-08-28. Before that, every task on --profile standard ran at 0 credits, images included, despite the ops repo’s own providers/manus_run.py stating in its comments that images cost 70-524 credits per generation. We’d measured that discrepancy directly: a 2560x1440 hero image, task S55YRpXRB2rHdhyMpB4W5m, came back with credit_usage: 0 and an unchanged balance. The guard existed for the moment that stopped being true.
The design was straightforward. A spend-guard script reads a balance snapshot through the existing get_balance() call, no second API client, persists it under .runs/manus/, and diffs it against the last snapshot. On a real charge, it posts to Discord through notify_ben(venture="dxdev"), which maps to the BenAgentDX channel and gates repeats behind a 6-hour key.
The balance shape itself is the first trap. The response splits credits into pools: freeCredits, periodicCredits, addonCredits, proMonthlyCredits, refreshCredits, maxRefreshCredits. The free-daily pool refreshes to 300 at 04:00 UTC. So totalCredits moving is not evidence of spend by itself, it moves every day at refresh regardless of usage. The guard has to read which specific pool drained to know whether the prepaid plan was actually touched, not just whether the free allowance reset.
First bug: the shared alert key
notify_ben takes a key argument so the 6-hour repeat gate can dedupe pings. The plan-pool alert and another alert were sharing one key. When the plan-pool alert fired, it consumed the gate, and the next real charge, 28 credits, hit the same key inside the 6-hour window and got silently suppressed. The fix was one line: give the plan-pool alert its own key. The bug wasn’t in the threshold logic or the balance math, it was two unrelated alerts fighting over one piece of state that was supposed to be a per-alert identity.
Second bug: midnight in the wrong timezone
The free-daily refresh happens at 04:00 UTC. The promo’s own end date was tracked as “2026-08-28” without a timezone, and got corrected in a separate commit: the promo lapsed on the 27th, not the 28th, “measured, not guessed” from the run records rather than assumed from the calendar.
But the guard itself had the same class of bug at the per-task level: manus-watch’s check for whether a task’s charge belonged to “today” compared against local midnight while the API’s own day boundary is UTC. On a machine that’s hours behind UTC, a charge that landed after UTC midnight but before local midnight got attributed to the wrong day’s bucket, and the guard’s day-over-day comparison threw it out. The commit message states the damage directly: it missed 152 real credits.
That’s the dangerous failure mode for this kind of monitor. It isn’t crashing, it isn’t throwing an exception anyone would see in a log. It just quietly agrees that nothing happened. A guard that pages you too often gets noticed and fixed. A guard that goes quiet at exactly the wrong moment looks identical to a quiet night.
Why this class of bug keeps showing up
Both bugs share a root cause: state that looks like a single value but is actually partitioned, and the partition boundary doesn’t match the code’s assumption. The alert key was one string standing in for two independent concerns. The date comparison was one clock standing in for two different timezones. Neither bug was visible by reading the function in isolation, both only showed up when a second, concurrent thing was running against the same state, a second alert type, a second timezone.
We verified the fix the same way the original build spec required: force the charge condition with a synthetic previous snapshot, confirm a Discord message actually lands on the channel, confirm the no-change case sends nothing. Reading the code and reasoning about it had already happened once, and it hadn’t caught either bug. Firing it did.
The guard is a small thing, a balance read, a diff, a Discord POST. But “small” and “safe to skip testing” aren’t the same property, especially once something else in the system is running in parallel with it. A monitor that watches for money moving needs to survive money moving at a time its own clock doesn’t expect.