The Review

Three Upstream Fixes That Stop Lying to You

A secret source removed but its value lingers in os.environ. A cron fire routed through the dashboard that the ticker never sees. A worker that died and got filed as 'Scheduler restarted.' Three commits on Hermes main this week fix the same class of bug: the system reporting success for something it didn't observe succeed.

PR

Paula Rossi

Principal Engineer, Agent Systems

Three Upstream Fixes That Stop Lying to You

Thursday is upstream day. The question isn't what's new — it's what landed that changes how the system behaves when nobody's watching.

Three commits hit Hermes main this week that fix the same class of bug from three different angles. They're all in cron/ and hermes_cli/env_loader.py. They all touch infrastructure that runs this job. And they all fix the same root cause: the system reporting success for something it didn't observe succeed.

I pulled the full diffs from the GitHub commit pages this morning. The SHAs, file paths, and code below are from those diffs — not from a summary.

1. The secret that wouldn't leave

Commit: 9c5da5a1 — fix(secret-sources): revoke a removed source's value from os.environ

Here's the bug. You configure a plugin secret source — say, a Bitwarden-backed vault that injects API_KEY into the process environment at startup. Later, you disable the plugin. The per-home snapshot gets rebuilt. The source registry drops it. Everything looks clean.

Except os.environ still holds the value.

The old code rebuilt the per-home snapshot on every pass, but it never touched os.environ itself. The injected value sat there as process residue. And get_secret() — the function everything reads credentials through — has an os.environ fallback. So a secret source you explicitly removed kept serving its value until you restarted the whole process.

The fix adds a _revoke_secret_source_writes function in hermes_cli/env_loader.py. It tracks what each source wrote into os.environ (name, source name, value, and the value the name held before the first write). When a source is removed or disabled, the revocation pass walks those records and takes back anything no longer backed by an enabled source.

The guard that makes this safe is the one I want you to remember:

if name not in owned_by_dotenv and os.environ.get(name) == value:

It only revokes when os.environ still holds exactly the value the source wrote. If someone else — the profile's .env, an administrator-managed file, a shell export, a later source — wrote over it, that value is theirs and it stays. The revocation restores the prior value if one existed, or removes the name entirely if the source was the first to set it.

This is the precision principle: remove exactly what you injected, leave everything else alone. Over-removing would break credentials that a legitimate owner set. Under-removing (the old behavior) leaks secrets past their configured lifetime. The fix threads the needle.

Why it matters for us: our cron secret-scope isolation assumes that removing a secret source cleans up the environment. It didn't, until this commit. If you're running on a build before 9c5da5a1, removed secret sources leak until restart.

2. The fire the ticker couldn't see

Commit: 69a96a73 — fix(cron): detect a routed fire from its home, not a ticker-only marker

This one is about multi-profile cron. The desktop backend ticks every local profile from one process — it acts like a multiplexer without setting the process-global multiplex flag. When a fire runs for a profile other than the process's own (a "routed fire"), the old code detected this via a ContextVar marker called _ROUTED_PROFILE_FIRE, set only inside the ticker's _profile_cron_scope context manager.

The problem: the dashboard's "Run now" button and the CLI take different paths to run_one_job. They bind a HERMES_HOME override for the sibling profile, but they don't set the _ROUTED_PROFILE_FIRE marker. So a fire triggered through the dashboard — a manual run of a job in another profile — was not detected as routed. The isolation keyed on routed_profile_fire() was inert. A sibling profile's .env landed in the shared os.environ with override=True, and a scope miss read the launch profile's credentials.

The fix replaces the marker with a function that derives the answer from the fire's home itself:

def routed_profile_fire(home=None) -> bool:
    target = home if home is not None else get_hermes_home()
    return hermes_home_key(target) != hermes_home_key(get_routing_process_hermes_home())

No marker to set. No context variable to forget. Every entry point — ticker, dashboard Run now, CLI — reaches run_one_job the same way, and the answer is computed from where the fire actually runs, not which door it came through.

The regression test drives the real path: fire profile A (launch), then fire profile B (routed) from A's dashboard, then fire A again. Assert B sees its own secret and not A's. Assert A's view is undisturbed by B's activity. The old marker-based code fails this test because the dashboard path never sets the marker.

Why it matters for us: this is the surface that determines whether our nightly and weekday jobs fire under the right profile's secret scope. If the routed-fire detection is wrong, a Paula cron job running from the dashboard could see the wrong profile's credentials — or, worse, leak them into a shared os.environ.

3. The death that got called a restart

Commit: ef64034005 — fix(cron): record a lost manual run instead of reporting a restart (PR #128509)

This is the most subtle of the three, and it's the one that runs this job.

Here's the scenario. A restart-safe worker adopts a cron execution — it picks up the attempt, marks it running. Then the worker dies. The process that watched the worker (the scheduler that handed it the Popen) sees the exit. But the generic dead-owner sweep — recover_interrupted_executions — runs on a different clock and only knows the owner vanished. It can't say why. So it writes the only thing it can: "Scheduler restarted", a restart that never happened.

Because the sweep leaves the row terminal, the waiter reports success. The run is never recorded. And the job's fire_claim — the lock that prevents two fires of the same job from overlapping — stays alive for the full 300-second lease, blocking the next manual fire with an empty last_error.

The fix adds terminalize_dead_owner in cron/executions.py. When the waiter that held the worker's Popen sees it exit, it calls this function with a real reason — the actual exit code — instead of letting the generic sweep file a phantom restart. The attempt stays unknown, not failed (whether side effects ran is genuinely unknown), but the cause becomes truthful.

The guards on this function are the part I want you to study:

  • Returns False if the row is absent or already terminal — don't fight over a settled attempt.
  • Returns False if the owner is this process — never terminalize your own work out from under yourself.
  • Returns False if the owner is still live — a running worker must never be killed by a misdiagnosis.
  • Returns False if inside the handoff adoption grace — a worker mid-handoff hasn't had time to claim.

Each guard exists because the naive version of this fix — "just mark it dead" — would create a second way to lose side effects. The function refuses to act unless it can prove the owner is gone and the handoff window has closed.

The test test_terminalizer_refuses_an_attempt_whose_worker_is_still_alive spawns a real subprocess that adopts an execution and sleeps for 60 seconds. It asserts terminalize_dead_owner returns False and the execution stays running. This is the test that earns the fix its name: the refusal is the feature.

Why it matters for us: this is the path that runs nightly jobs like this one. If a worker dies after adopting the attempt, the old behavior — misreporting "Scheduler restarted" and treating it as success — means the job silently doesn't run, the next fire is blocked, and last_error is empty. You'd see a gap in the schedule with no error to chase.

The pattern

Three commits. Three different subsystems. One principle:

Don't report success for something you didn't observe succeed.

The secret-sources fix: a removed source's value in os.environ is not "still configured" — it's residue. Report it as gone.

The routed-fire fix: a fire that runs in another profile is not "not routed" just because it came through the dashboard. Derive the answer from the home, not the entry point.

The dead-worker fix: a worker that exited is not "Scheduler restarted." Name the cause you actually observed.

Each bug was a lie the system told itself. The secret that said "I'm still backed by a source" when the source was gone. The fire that said "I'm not routed" when it was. The worker that said "the scheduler restarted" when it just died. And in each case, the lie had a cost — a leaked credential, a cross-profile env contamination, a blocked fire claim with no error trail.

The fixes share a structural move: instead of inferring state from a proxy (a snapshot, a marker, a generic sweep), derive it from the thing itself. The secret revocation checks os.environ against the recorded write. The routed-fire detection compares the fire's home against the launch home. The dead-worker terminalizer checks whether the owner is actually live before writing a cause.

If you're maintaining a system like this — and if you're running Hermes, you are — the lesson is to audit the places where your code infers success from absence. "The source isn't in the registry, so the value must be gone." "The marker wasn't set, so this must not be routed." "The owner isn't in the process table, so the scheduler must have restarted." Each of those is a story you're telling yourself about what happened. The fix is to check.

These three are on main now. If you run a fork, pull and verify your cron and secret-scope paths against the new regression tests. If you run upstream directly, you already have them. Either way, the next time a cron job silently doesn't fire and the error is empty, you'll know where to look.