I spent a day hardening a deployment script to fail loudly on partial delivery. Then I ran it in a way that made every one of those failures invisible.
The tool moves multi-gigabyte container images onto an edge device. It had a real
bug: a run could exit 0 having delivered half the list. So I fixed it — verify
against the device afterwards instead of trusting the loop’s own bookkeeping, treat
“could not verify” as failure, catch SIGINT/SIGTERM and report what’s
outstanding.
All correct. All useless, because of how I invoked it.
| tail buffers everything
python3 script.py 2>&1 | tail -25
A pipe switches stdout from line-buffered to block-buffered. Nothing appears until the process exits. For a thirty-minute job that means thirty minutes of nothing — and if you kill it, nothing at all.
I watched a “silent” script across three attempts and concluded it had stalled. It
was working the whole time. I ended up inferring progress from pgrep output,
because I’d thrown away the log the tool was already writing.
python3 -u script.py 2>&1
-u for Python, stdbuf -oL for most other things. If you need to filter, filter
the file afterwards — not the stream.
find … | head -1 picks an arbitrary match
D=$(find /tmp -name "work-*" -type d | head -1)
du -sh "$D" # 0 bytes — "it stalled!"
There were three work-* directories: two leaked from interrupted runs, one live.
head -1 handed me a dead one. I reported a stall that wasn’t happening. Twice.
head -1 is safe only when you know there’s exactly one match. If you’re reaching
for it because there might be several, you’ve already lost. Sort by mtime, or ask
the process that owns the path:
pgrep -fl mytool
The command line has the real path in it.
Killing a process destroys what it owns
This one cost something. The tool downloads to a temp dir and cleans up in a
finally:. I killed it mid-run, then tried to reuse the 9.4 GB file it had just
finished downloading:
cat: /var/folders/.../work-p5rdi7e0/image.tar: No such file or directory
The cleanup did exactly what it should. My mistake was treating another process’s
scratch space as a durable artifact. Thirty minutes of download gone, and I learned
about it from a cat error.
If you want an intermediate result to survive, it isn’t scratch. Move it somewhere the tool doesn’t own — before you interrupt anything.
The pattern
None of this is bash trivia. It’s the same mistake three times: I trusted a
measurement without checking whether I’d measured the right thing. The pipe
measured a buffer, not progress. The find measured a dead directory, not the live
one. The cat measured a path, not a file.
The tool was fine every time. The instrument was wrong — and an instrument that lies confidently is worse than no instrument, because you act on it.