Commit Graph
2 Commits
Author SHA1 Message Date
Christian Manivong ec93619f8f fix: one deploy per host at a time, under a lock the host frees on its own
CI / check (pull_request) Successful in 17s
Several sessions deploy to the same test server, and two runs used to
overlap. On 2026-10-05 two `up -d --force-recreate` runs recreated each
other's containers and the API was down for a minute (#1). On 2026-10-06
two runs renamed each other's *.new compose files and one broke off; with
two different tags, one tag's compose files could have started the
other's images (#3).

- Each deploy first takes flock on ~/netork/.deploy.lock on the host. An
  ssh session holds it: the remote side takes the lock on fd 9, reports
  LOCKED and waits on its stdin, so ending the session frees it, whether
  the deploy finished, failed, was interrupted or lost its connection.
  Checked over real ssh on .50, including a client killed with -9.
- A second deploy prints who holds the lock, since when and with which
  tag, and waits up to DEPLOY_LOCK_WAIT seconds (default 900); then it
  gives up without touching the host.
- A host without flock is deployed without the lock, with a warning.
- deploy_server is now the lock around deploy_steps, the old body.

Closes #1
Closes #3
2026-10-07 06:42:13 +02:00
Christian Manivong 923ec4ce65 fix: a failed container recreate is tried once more before the deploy gives up
CI / check (pull_request) Successful in 13s
On 2026-10-05 a deploy to 172.22.8.50 failed in "Starting containers":
docker compose up -d --force-recreate recreates the services in parallel,
lost a container it had just renamed ("No such container: 02586df7…") and
stopped with every engine container Created and none running. The old API
was already gone, so netOrk was down for about two minutes, until the same
deploy was run again and went through cleanly.

The script now does that second run itself, after DEPLOY_RECREATE_RETRY_DELAY
seconds (5 by default). If the second attempt fails too, it prints the
container states (docker compose ps -a) and fails as before.

Tests run the script against a fake ssh that fails the recreate zero, one
or two times. README step 3 says what happens.

Refs NetOrk/netork#586
2026-10-05 21:58:36 +02:00