CI / check (pull_request) Successful in 17s
Several sessions deploy to the same test server, and two runs used to overlap. On 2026-10-05 two `up -d --force-recreate` runs recreated each other's containers and the API was down for a minute (#1). On 2026-10-06 two runs renamed each other's *.new compose files and one broke off; with two different tags, one tag's compose files could have started the other's images (#3). - Each deploy first takes flock on ~/netork/.deploy.lock on the host. An ssh session holds it: the remote side takes the lock on fd 9, reports LOCKED and waits on its stdin, so ending the session frees it, whether the deploy finished, failed, was interrupted or lost its connection. Checked over real ssh on .50, including a client killed with -9. - A second deploy prints who holds the lock, since when and with which tag, and waits up to DEPLOY_LOCK_WAIT seconds (default 900); then it gives up without touching the host. - A host without flock is deployed without the lock, with a warning. - deploy_server is now the lock around deploy_steps, the old body. Closes #1 Closes #3
109 lines
3.5 KiB
Python
109 lines
3.5 KiB
Python
"""A failed container recreate is tried once more before the deploy gives up.
|
|
|
|
On 2026-10-05 a deploy to 172.22.8.50 failed in "Starting containers":
|
|
`docker compose up -d --force-recreate` recreates the services in parallel, lost
|
|
a container it had just renamed ("No such container: 02586df7…"), and stopped
|
|
with every engine container `Created` and none running. The old API was already
|
|
gone, so netOrk was down until somebody ran the same deploy again, which then
|
|
went through cleanly (NetOrk/netork#586).
|
|
|
|
So the script runs that step a second time on its own. Should the second attempt
|
|
fail too, it shows what the containers look like and fails loudly, as before.
|
|
|
|
The remote side is a fake `ssh` that answers each step and can fail the recreate
|
|
a given number of times.
|
|
"""
|
|
|
|
from __future__ import annotations
|
|
|
|
import os
|
|
import subprocess
|
|
from pathlib import Path
|
|
|
|
import pytest
|
|
|
|
DEPLOY = Path(__file__).resolve().parent.parent / "deploy.sh"
|
|
|
|
_FAKE_SSH = """#!/bin/sh
|
|
for arg in "$@"; do cmd="$arg"; done
|
|
# The host lock's holder runs for real, here (test_deploy_host_lock.py).
|
|
case "$cmd" in *.deploy.lock*) exec sh -c "$cmd" ;; esac
|
|
# Swallow stdin (docker login, the .env heredoc), then answer by command.
|
|
cat > /dev/null
|
|
case "$cmd" in
|
|
*--force-recreate*)
|
|
n=$(cat "{state}/recreates" 2>/dev/null || echo 0)
|
|
n=$((n + 1))
|
|
echo "$n" > "{state}/recreates"
|
|
if [ "$n" -le "{failures}" ]; then
|
|
echo "Error response from daemon: No such container: 02586df7e659"
|
|
exit 1
|
|
fi
|
|
echo " Container netork-netork-api-1 Started"
|
|
;;
|
|
*" ps -a"*)
|
|
echo "netork-netork-api-1 Created"
|
|
;;
|
|
esac
|
|
exit 0
|
|
"""
|
|
|
|
|
|
@pytest.fixture
|
|
def deploy(tmp_path: Path):
|
|
"""Run a deploy against the fake ssh; *failures* recreates fail before one works."""
|
|
|
|
def _deploy(failures: int) -> tuple[subprocess.CompletedProcess[str], int]:
|
|
bin_dir = tmp_path / "bin"
|
|
bin_dir.mkdir(exist_ok=True)
|
|
fake = bin_dir / "ssh"
|
|
fake.write_text(
|
|
_FAKE_SSH.replace("{state}", str(tmp_path)).replace("{failures}", str(failures))
|
|
)
|
|
fake.chmod(0o755)
|
|
result = subprocess.run(
|
|
["bash", str(DEPLOY), "testhost"],
|
|
env={
|
|
"PATH": f"{bin_dir}:{os.environ['PATH']}",
|
|
"HOME": str(tmp_path),
|
|
"DEPLOY_ENV_FILE": "/dev/null",
|
|
"REGISTRY_HOST": "registry.example",
|
|
"NETORK_VERSION": "latest-dev",
|
|
"DEPLOY_RECREATE_RETRY_DELAY": "0",
|
|
},
|
|
stdin=subprocess.DEVNULL,
|
|
capture_output=True,
|
|
text=True,
|
|
timeout=60,
|
|
check=False,
|
|
)
|
|
recreates = tmp_path / "recreates"
|
|
return result, int(recreates.read_text()) if recreates.exists() else 0
|
|
|
|
return _deploy
|
|
|
|
|
|
def test_a_recreate_that_works_runs_once(deploy) -> None:
|
|
result, recreates = deploy(failures=0)
|
|
|
|
assert result.returncode == 0, result.stderr
|
|
assert recreates == 1
|
|
|
|
|
|
def test_a_recreate_that_fails_once_is_tried_again(deploy) -> None:
|
|
result, recreates = deploy(failures=1)
|
|
|
|
assert result.returncode == 0, result.stderr
|
|
assert recreates == 2
|
|
assert "trying once more" in result.stderr
|
|
assert "All servers deployed successfully." in result.stdout
|
|
|
|
|
|
def test_two_failed_recreates_fail_the_deploy_and_show_the_containers(deploy) -> None:
|
|
result, recreates = deploy(failures=2)
|
|
|
|
assert result.returncode != 0
|
|
assert recreates == 2
|
|
assert "netork-netork-api-1 Created" in result.stderr
|
|
assert "FAILED on: testhost" in result.stderr
|