fix: a failed container recreate is tried once more before the deploy gives up
CI / check (pull_request) Successful in 13s

On 2026-10-05 a deploy to 172.22.8.50 failed in "Starting containers":
docker compose up -d --force-recreate recreates the services in parallel,
lost a container it had just renamed ("No such container: 02586df7…") and
stopped with every engine container Created and none running. The old API
was already gone, so netOrk was down for about two minutes, until the same
deploy was run again and went through cleanly.

The script now does that second run itself, after DEPLOY_RECREATE_RETRY_DELAY
seconds (5 by default). If the second attempt fails too, it prints the
container states (docker compose ps -a) and fails as before.

Tests run the script against a fake ssh that fails the recreate zero, one
or two times. README step 3 says what happens.

Refs NetOrk/netork#586
This commit is contained in:
Christian Manivong
2026-10-05 21:58:36 +02:00
parent a928d5eec3
commit 923ec4ce65
3 changed files with 131 additions and 3 deletions
+4 -1
View File
@@ -70,7 +70,10 @@ started earlier pulls whatever image the registry held before, which is stale.
2. Pulls the engine image and copies `docker-compose.yml` and
`docker-compose.registry.yml` out of it into `~/netork/`.
3. Pulls the netOrk images, then force-recreates the API, the workers, `netork-beat`,
`flower` and, on UI servers, `netork-ui`.
`flower` and, on UI servers, `netork-ui`. If the recreate fails, it is tried once
more after 5 s (`DEPLOY_RECREATE_RETRY_DELAY`). Compose's parallel recreate can lose
a container it just renamed and leave the rest stopped. If the second attempt fails
too, the container states are printed and the deploy fails.
4. Reconciles `registry`, `apt-cacher-ng` and `signal-api`: each is recreated only if its
definition changed. It never touches `postgres` or `redis`.
5. Verifies that the containers run exactly the image that was pulled. If they don't, the