14 KiB
title, date, category, module, problem_type, component, severity, applies_when, tags
| title | date | category | module | problem_type | component | severity | applies_when | tags | ||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Porting POSIX process supervision to native Windows: the primitives that fail silently | 2026-07-24 | architecture-patterns | skills (peer-job-runner.py, byte-duplicated across ce-doc-review, ce-code-review, ce-pov, ce-work, ce-plan, ce-brainstorm; ce-babysit-pr/scripts/pr-snapshot) | architecture_pattern | tooling | high |
|
|
Porting POSIX process supervision to native Windows: the primitives that fail silently
Context
peer-job-runner.py supervises detached cross-model peer jobs. It was POSIX-only:
os.fork/os.setsid to detach, killpg to reap the worker tree, fstat().st_uid == geteuid() plus mode 0700/0600 for ownership and privacy. Native Windows Python has
none of those, so the runner refused to start there
(skills/ce-doc-review/scripts/peer-job-runner.py, issue #1243).
The port's lesson is not that the primitives are missing — that part is obvious and fails loudly. It is that each POSIX primitive has a Windows analog that differs in a way that fails silently: no exception, no error, just a wrong result, a leaked process, or corrupted bytes. Every item below was verified by running it on Windows 11 / Python 3.11, and several contradicted the reasonable-sounding assumption that preceded them.
Related: Detached job lifecycle for delegated work describes the platform-independent lifecycle contract this runner implements; this doc covers only the platform-primitive mapping underneath it.
Guidance
1. Process-tree teardown: killpg has no direct analog
This is the one most likely to leak processes forever.
A POSIX process group outlives its leader. killpg targets the pgid, so a teardown
still sweeps surviving members after the leader exits — which is why the POSIX code
deliberately does not early-return on a dead leader.
taskkill /T walks parent→child from a live parent. Against an already-exited pid it
reports "process not found" and kills nothing, so a worker's grandchildren leak
permanently. Reproduced: leader exited, grandchild survived, taskkill /T /F on the
exited pid returned rc=128 and the grandchild kept running.
The real analog is a Job Object: job membership is inherited and does not depend on a
live parent, so TerminateJobObject reaches descendants of an exited leader.
But a named job object is only reachable by a process holding a handle. Windows
releases a named kernel object's name once the last handle closes, even while member
processes keep the object itself alive. Verified: after killing the process that created
the job, OpenJobObjectW on the same name fails with ERROR_FILE_NOT_FOUND (2) while a
member grandchild was still running.
So a cross-process teardown (a reap running after the supervisor died) needs a second
mechanism: a CreateToolhelp32Snapshot walk. That works because Windows never reparents
orphans — a dead pid still appears as th32ParentProcessID on its live children. This is
the exact inverse of POSIX, where orphans reparent to init and the parent pid is gone.
POSIX: pgid survives leader -> killpg sweeps orphans
orphans reparent to init -> parent-pid walk is useless
Windows: job survives leader -> TerminateJobObject sweeps orphans
but the job NAME dies with the last handle
orphans keep a dangling ppid -> snapshot walk IS the fallback
Also: a graceful taskkill /T (no /F) sends WM_CLOSE, which a console-less process
cannot receive — it reports "can only be terminated forcefully". A TERM-then-grace-then-KILL
shape ported literally spends the whole grace window achieving nothing. Dropping it cut reap
latency from ~2800 ms to ~180 ms and closed a race where a reap caller and the supervisor
both wrote a terminal record.
2. os.open is text mode on Windows — silent corruption
Windows CPython opens os.open() descriptors in CRT text mode by default: writes
expand \n → \r\n, and reads stop at the first 0x1A (Ctrl-Z, the DOS EOF). A
published result artifact containing a 0x1A is silently truncated, and byte-cap
accounting drifts from st_size.
O_BINARY = getattr(os, "O_BINARY", 0) # 0 on POSIX -> no-op there
fd = os.open(path, os.O_RDONLY | O_NOFOLLOW | O_BINARY)
Apply it to every os.open that moves bytes. Note O_NOFOLLOW is also 0 on Windows, so
the no-follow protection quietly disappears — say so rather than leaving a docstring
claiming it.
3. Ownership: SIDs, and check the handle not the path
There is no uid and mode bits are not access control. The equivalent identity is the current user's SID, and two details matter:
- Check the opened HANDLE (
GetSecurityInfoviamsvcrt.get_osfhandle(fd)), not the path. That preserves the TOCTOU property the POSIXfstat-by-fd check has; a path-based check silently reintroduces the race. - An elevated token has two identity SIDs — the user SID and the token's default
owner SID, which for an admin is
BUILTIN\Administrators. Objects the process creates may be owned by either, so a strict equality check against the user SID alone rejects the process's own files under elevation.
Privacy is an ACL, not a mode. icacls /inheritance:r is the closest analog to chmod 0700 but is destructive and irreversible in a way chmod is not — it permanently drops
inherited ACEs. Only re-ACL directories you created or own; do not "repair" a
user-supplied root the way an unconditional chmod 0700 safely would.
4. Liveness, detach, and shell-outs
os.kill(pid, 0)does not work for liveness on Windows. Worse,os.kill(pid, SIGTERM)maps toTerminateProcess— a probe that actually kills. UseOpenProcess+WaitForSingleObject(h, 0)(WAIT_TIMEOUT= alive).- Detach is
DETACHED_PROCESS | CREATE_NEW_PROCESS_GROUPre-invoking the script through an internal subcommand (there is no fork).CREATE_BREAKAWAY_FROM_JOBis the analog ofsetsid's reparent-to-init — it lets the supervisor outlive a harness that runs inside a kill-on-close job — and must be retried without the flag when the job forbids breakaway. - A child of a console-less parent gets a new console unless
CREATE_NO_WINDOWis set, so every job flashes a window.CREATE_NO_WINDOWis mutually exclusive withDETACHED_PROCESS, so the detached supervisor uses one and its worker uses the other. subprocess.run(..., check=False)suppresses a nonzero exit, not a missing executable — an absent tool still raisesFileNotFoundError. In a teardown path that turns an already-classifieddonejob intofailed. Wrap fire-and-forget shell-outs inexcept OSError.- Resolve
icacls/taskkillby absolute%SystemRoot%\System32path:CreateProcesssearches the application and current directories before System32, so bare names are a binary-hijack surface.
5. State files: locking, identity, and rename
ce-babysit-pr's pr-snapshot watcher hit the same family from a different angle (issue
#1280). It needed no fork or job object at all — only a lock, a start-time probe, and a
rename — and all three still differ.
fcntl.flock -> msvcrt.locking, with three semantic changes. It locks a byte range
from the current file offset, not the whole file, so the offset must be pinned (seek to
0, lock 1 byte) or an unlock at a different offset silently leaves the range held. It has
no shared mode — every holder is exclusive, so a POSIX LOCK_SH reader has to become a
writer, which is free only if critical sections are short. And its blocking mode LK_LOCK
gives up after ~10 seconds (1s × 10 retries) rather than waiting, so the faithful
translation of "block until the holder releases" is a retry loop around the non-blocking
LK_NBLCK, not LK_LOCK. Also: do not acquire the lock through a truncating open(path, "w") — on Windows, truncating a file another process holds a byte-range lock on fails, so
lock acquisition becomes the race it was meant to prevent. Use os.open(..., O_RDWR | O_CREAT).
That retry loop must discriminate on errno, not catch OSError. msvcrt.locking
reports genuine contention as EACCES and real defects as EBADF (closed descriptor) or
EINVAL (bad range) — all OSError, so a bare except OSError: sleep; continue retries the
unfixable ones forever, at the poll interval, emitting nothing. Because the POSIX side is an
unbounded flock wait, the loop has no timeout to rescue it: the process simply never returns.
For a supervised watcher that is observably identical to "monitoring silently stopped" — the
class of failure the port existed to fix. Retry EACCES; re-raise everything else. Testing this
needs the primitive patched, not provoked with a real bad descriptor: the os.lseek that
pins the offset rejects a closed fd before the loop is entered, so the obvious test passes
against the broken loop too.
ps -o lstart= -> GetProcessTimes. A PID-reuse guard needs process creation time;
Windows recycles PIDs aggressively, so this matters more there, not less. OpenProcess
with PROCESS_QUERY_LIMITED_INFORMATION plus GetProcessTimes is the analog, with
QueryFullProcessImageNameW standing in for command=. Two asymmetries: Windows keeps an
exited process openable while any handle to it remains, so a "dead" pid can still yield an
identity (harmless if you compare identities rather than test existence); and
QueryFullProcessImageNameW fails on an exited process, so the image half goes empty while
the creation-time half stays valid. Guard the POSIX branch too — a missing ps raises
FileNotFoundError, it does not return nonzero.
os.replace is not unconditional. POSIX rename succeeds over an open destination;
Windows refuses with PermissionError while any other handle to the destination is open.
Any unlocked best-effort reader — or a virus scanner — turns a routine concurrent read into
a crashed writer, and it is timing-dependent, so it passes every local test. Retry the
rename briefly instead of treating the sharing violation as failure.
Note that a lock file used only for locking moves no bytes, so item 2's O_BINARY does not
apply to it. Apply that rule by whether the descriptor carries data, not by platform.
Why This Matters
Every failure above is silent. Nothing raises; the job reports done, the artifact looks
plausible, the reap returns 0 — while a peer process leaks forever or a result is truncated
mid-JSON. A port that "passes the smoke test" can be wrong in all of these ways at once,
which is why the checklist matters more than any single fix.
The teardown item is the expensive one: it took an adversarial review plus two live
experiments to establish that the obvious mechanism (taskkill /T) cannot work for the
orphan case, that the correct mechanism (Job Object) is unreachable cross-process, and that
the fallback is sound only because of a Windows/POSIX asymmetry in orphan reparenting. That
is a day of rediscovery for the next person.
When to Apply
Reach for this when porting any POSIX-shaped supervisor, daemon, or job runner to native Windows — and specifically when a POSIX invariant is load-bearing. The tell is a comment explaining why a POSIX call is safe ("the pgid outlives its leader, so this still sweeps"). That reasoning is exactly what does not port, and a literal translation will compile, run, pass a happy-path test, and leak.
Do not assume the platform branch is only about missing APIs. Ask, for each guarantee the POSIX code relies on: what kernel object provides this, and what is its lifetime on Windows?
Examples
Teardown that looks correct and leaks:
# WRONG on Windows: cannot reach a dead leader's children
def kill_tree(root_pid, grace):
subprocess.run(["taskkill", "/T", "/F", "/PID", str(root_pid)], check=False)
Teardown that actually sweeps:
alive = _win_pid_alive(root_pid)
if job_name and _win_terminate_job(job_name): # primary: reaches exited-leader children
return alive
for pid in _win_descendants_deepest_first(root_pid): # fallback: dangling-ppid snapshot walk
_win_terminate_pid(pid)
if alive:
_win_terminate_pid(root_pid)
return alive
Keeping the POSIX path provably untouched is worth more than symmetry. Every Windows branch
sits behind if IS_WINDOWS:; shared helpers were changed only where the POSIX result is
identical by inspection (e.g. a reap-request helper that reduces to the original flag check
when IS_WINDOWS is false). On a Windows box the POSIX suite cannot run, so byte-identical
POSIX code is the only available regression guarantee.
Cross-platform tests are possible but narrower than they look: patching IS_WINDOWS on
Linux raises NameError for most Windows branches, because the _win_* helpers and the
ctypes/msvcrt imports only exist when the module is imported under
sys.platform == "win32". Only branches that return before touching a _win_* helper —
platform selection, marker-file logic, and the O_BINARY byte round-trip — are testable on
the existing CI.
See Also
- Interpreter resolution is a separate Windows trap in the same family: bare
python3resolves to a Microsoft Store stub that satisfiescommand -vbut exits without running Python. Tracked in issue #1247.