fix: a failed container recreate is tried once more before the deploy gives up #2

Merged
christianmanivong merged 1 commits from fix/recreate-retry into main 2026-10-05 20:03:55 +00:00
Owner

On 2026-10-05 a deploy to 172.22.8.50 failed in "Starting containers". docker compose up -d --force-recreate recreates the services in parallel. It lost a container it had just renamed (No such container: 02586df7…) and stopped with every engine container Created and none running. The old API was already gone, so netOrk was down for about two minutes, until the same deploy was run again and went through cleanly (NetOrk/netork#586).

Changes

  • deploy.sh now does that second run itself, after DEPLOY_RECREATE_RETRY_DELAY seconds (5 by default). If the second attempt fails too, it prints docker compose ps -a and fails as before.
  • tests/test_deploy_recreate_retry.py runs the script against a fake ssh that fails the recreate zero, one or two times.
  • README step 3 describes the retry.
  • shellcheck 0.11.0 (the CI pin) and all 19 tests pass.

Closes NetOrk/netork#586

On 2026-10-05 a deploy to 172.22.8.50 failed in "Starting containers". `docker compose up -d --force-recreate` recreates the services in parallel. It lost a container it had just renamed (`No such container: 02586df7…`) and stopped with every engine container `Created` and none running. The old API was already gone, so netOrk was down for about two minutes, until the same deploy was run again and went through cleanly (NetOrk/netork#586). ## Changes - `deploy.sh` now does that second run itself, after `DEPLOY_RECREATE_RETRY_DELAY` seconds (5 by default). If the second attempt fails too, it prints `docker compose ps -a` and fails as before. - `tests/test_deploy_recreate_retry.py` runs the script against a fake `ssh` that fails the recreate zero, one or two times. - README step 3 describes the retry. - shellcheck 0.11.0 (the CI pin) and all 19 tests pass. Closes NetOrk/netork#586
christianmanivong added 1 commit 2026-10-05 19:58:45 +00:00
On 2026-10-05 a deploy to 172.22.8.50 failed in "Starting containers":
docker compose up -d --force-recreate recreates the services in parallel,
lost a container it had just renamed ("No such container: 02586df7…") and
stopped with every engine container Created and none running. The old API
was already gone, so netOrk was down for about two minutes, until the same
deploy was run again and went through cleanly.

The script now does that second run itself, after DEPLOY_RECREATE_RETRY_DELAY
seconds (5 by default). If the second attempt fails too, it prints the
container states (docker compose ps -a) and fails as before.

Tests run the script against a fake ssh that fails the recreate zero, one
or two times. README step 3 says what happens.

Refs NetOrk/netork#586
christianmanivong merged commit 6444aa97f3 into main 2026-10-05 20:03:55 +00:00
Sign in to join this conversation.
No Reviewers
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: NetOrk/deploy#2