fix: one deploy per host at a time, under a lock the host frees on its own #4

Merged
christianmanivong merged 1 commits from fix/host-deploy-lock into main 2026-10-07 04:43:24 +00:00
Owner

Two deploy.sh runs against the same host overlapped twice in two days:

  • #1 (2026-10-05): two up -d --force-recreate runs recreated each other's containers, and the API was down for a minute.
  • #3 (2026-10-06): two runs renamed each other's *.new compose files, and one broke off. With two different tags, one tag's compose files could have started the other's images.

Several sessions deploy to .50, so a third overlap was only a matter of time.

What changes

  • A host lock. hold_host_lock opens an ssh session (a bash coproc) that takes flock on ~/netork/.deploy.lock on the host, prints LOCKED, and then waits on its stdin. release_host_lock closes that stdin.
  • The lock lives with the session. A deploy that fails (set -e ends the server's subshell), is interrupted, or loses its connection frees it without any cleanup step, so no stale lock file is left behind.
  • A second deploy waits visibly. It prints another deploy holds <host>: since <UTC time>, by <user>@<host>, deploying <tag> — waiting up to 900s... (DEPLOY_LOCK_WAIT). After that it gives up with FAILED on: <host> and never reaches the compose steps.
  • A host without flock (util-linux) is deployed as before, with a warning, instead of becoming undeployable.
  • deploy_server is now hold_host_lock + deploy_steps (the previous body, unchanged) + release_host_lock.
  • README: step 0 and DEPLOY_LOCK_WAIT.

Tests

  • New, tests/test_deploy_host_lock.py. The fake ssh runs the lock holder for real on this machine, with HOME in a temp dir, so the lock is a real flock. Covered:
    • two deploys take turns (start A, end A, start B, end B), and B names A's holder, time and tag;
    • the lock is free after a deploy, and after a failed one;
    • a deploy waiting longer than DEPLOY_LOCK_WAIT gives up and touches nothing;
    • a host without flock gets a warning and is deployed.
  • test_deploy_recreate_retry.py: its fake ssh hands the lock holder through the same way.
  • All 24 green three times in a row. shellcheck and ruff are clean, with the CI's pinned versions.
  • Checked over real ssh on 172.22.8.50, with a throwaway lock file in /tmp:
    • the lock is held while the session's stdin is open, and free right after it closes;
    • after kill -9 of the ssh client it is free at once.

Closes #1
Closes #3

Two `deploy.sh` runs against the same host overlapped twice in two days: - **#1 (2026-10-05):** two `up -d --force-recreate` runs recreated each other's containers, and the API was down for a minute. - **#3 (2026-10-06):** two runs renamed each other's `*.new` compose files, and one broke off. With two different tags, one tag's compose files could have started the other's images. Several sessions deploy to .50, so a third overlap was only a matter of time. ## What changes - **A host lock.** `hold_host_lock` opens an ssh session (a bash coproc) that takes `flock` on `~/netork/.deploy.lock` on the host, prints `LOCKED`, and then waits on its stdin. `release_host_lock` closes that stdin. - **The lock lives with the session.** A deploy that fails (`set -e` ends the server's subshell), is interrupted, or loses its connection frees it without any cleanup step, so no stale lock file is left behind. - **A second deploy waits visibly.** It prints `another deploy holds <host>: since <UTC time>, by <user>@<host>, deploying <tag> — waiting up to 900s...` (`DEPLOY_LOCK_WAIT`). After that it gives up with `FAILED on: <host>` and never reaches the compose steps. - **A host without `flock`** (util-linux) is deployed as before, with a warning, instead of becoming undeployable. - **`deploy_server`** is now `hold_host_lock` + `deploy_steps` (the previous body, unchanged) + `release_host_lock`. - README: step 0 and `DEPLOY_LOCK_WAIT`. ## Tests - **New, `tests/test_deploy_host_lock.py`.** The fake `ssh` runs the lock holder for real on this machine, with `HOME` in a temp dir, so the lock is a real `flock`. Covered: - two deploys take turns (`start A, end A, start B, end B`), and B names A's holder, time and tag; - the lock is free after a deploy, and after a failed one; - a deploy waiting longer than `DEPLOY_LOCK_WAIT` gives up and touches nothing; - a host without `flock` gets a warning and is deployed. - **`test_deploy_recreate_retry.py`:** its fake `ssh` hands the lock holder through the same way. - All 24 green three times in a row. `shellcheck` and `ruff` are clean, with the CI's pinned versions. - **Checked over real ssh on 172.22.8.50,** with a throwaway lock file in `/tmp`: - the lock is held while the session's stdin is open, and free right after it closes; - after `kill -9` of the ssh client it is free at once. Closes #1 Closes #3
christianmanivong added 1 commit 2026-10-07 04:42:25 +00:00
Several sessions deploy to the same test server, and two runs used to
overlap. On 2026-10-05 two `up -d --force-recreate` runs recreated each
other's containers and the API was down for a minute (#1). On 2026-10-06
two runs renamed each other's *.new compose files and one broke off; with
two different tags, one tag's compose files could have started the
other's images (#3).

- Each deploy first takes flock on ~/netork/.deploy.lock on the host. An
  ssh session holds it: the remote side takes the lock on fd 9, reports
  LOCKED and waits on its stdin, so ending the session frees it, whether
  the deploy finished, failed, was interrupted or lost its connection.
  Checked over real ssh on .50, including a client killed with -9.
- A second deploy prints who holds the lock, since when and with which
  tag, and waits up to DEPLOY_LOCK_WAIT seconds (default 900); then it
  gives up without touching the host.
- A host without flock is deployed without the lock, with a warning.
- deploy_server is now the lock around deploy_steps, the old body.

Closes #1
Closes #3
christianmanivong merged commit 9a93897cd2 into main 2026-10-07 04:43:24 +00:00
Sign in to join this conversation.