fix: a failed container recreate is tried once more before the deploy gives up
CI / check (pull_request) Successful in 13s
CI / check (pull_request) Successful in 13s
On 2026-10-05 a deploy to 172.22.8.50 failed in "Starting containers":
docker compose up -d --force-recreate recreates the services in parallel,
lost a container it had just renamed ("No such container: 02586df7…") and
stopped with every engine container Created and none running. The old API
was already gone, so netOrk was down for about two minutes, until the same
deploy was run again and went through cleanly.
The script now does that second run itself, after DEPLOY_RECREATE_RETRY_DELAY
seconds (5 by default). If the second attempt fails too, it prints the
container states (docker compose ps -a) and fails as before.
Tests run the script against a fake ssh that fails the recreate zero, one
or two times. README step 3 says what happens.
Refs NetOrk/netork#586
This commit is contained in:
@@ -194,12 +194,31 @@ deploy_server() {
|
||||
# observed on netork-ui, which kept serving a stale image after its tag had
|
||||
# already advanced. Recreating unconditionally costs a restart per deploy,
|
||||
# which a deliberate rollout wants anyway. See NetOrk/netork#90.
|
||||
#
|
||||
# Tried twice. compose recreates the services in parallel, and on 2026-10-05 it
|
||||
# lost a container it had just renamed ("No such container") and stopped with
|
||||
# every engine container Created and none running: the old API was gone, so
|
||||
# netOrk was down until someone ran the deploy again, which then went through
|
||||
# (NetOrk/netork#586). The second attempt is that rerun. Should it fail as well,
|
||||
# the container states go to the log and the deploy fails as before.
|
||||
echo "[${SERVER}] Starting containers..."
|
||||
run_remote "$SERVER" "recreating ${SERVICES}" "" \
|
||||
"cd ~/netork && \
|
||||
local RECREATE="cd ~/netork && \
|
||||
REGISTRY_HOST='${REGISTRY_HOST}' NETORK_VERSION='${SERVER_VERSION}' \
|
||||
docker compose -f docker-compose.yml -f docker-compose.registry.yml \
|
||||
up -d --no-deps --force-recreate ${SERVICES}"
|
||||
if ! run_remote "$SERVER" "recreating ${SERVICES}" "" "$RECREATE"; then
|
||||
local delay="${DEPLOY_RECREATE_RETRY_DELAY:-5}"
|
||||
echo "[${SERVER}] Recreating failed; trying once more in ${delay}s..." >&2
|
||||
sleep "$delay"
|
||||
if ! run_remote "$SERVER" "recreating ${SERVICES} (second attempt)" "" "$RECREATE"; then
|
||||
echo "[${SERVER}] Containers after two failed attempts:" >&2
|
||||
ssh -n "$SERVER" "cd ~/netork && \
|
||||
REGISTRY_HOST='${REGISTRY_HOST}' NETORK_VERSION='${SERVER_VERSION}' \
|
||||
docker compose -f docker-compose.yml -f docker-compose.registry.yml ps -a" 2>&1 \
|
||||
| sed "s/^/[${SERVER}] /" >&2 || true
|
||||
return 1
|
||||
fi
|
||||
fi
|
||||
|
||||
# Infrastructure containers, deliberately left out of the force-recreate
|
||||
# above. Plain `up -d`: compose compares each service definition against the
|
||||
|
||||
Reference in New Issue
Block a user