Author SHA1 Message Date
Christian Manivong de46ab8a0f fix: remove the bundled registry from every host, and stop starting it
CI / check (pull_request) Successful in 16s
netOrk's docker-compose.yml shipped a registry:2 for the satellite image,
and this script started it on every deploy as infrastructure. Satellites
pull from registry.netork.io; on 172.22.8.50 the registry held one old
image, nothing had pulled from it in 30 days, and it accepted anonymous
pushes on port 5000 of every instance (NetOrk/netork#763).

- registry leaves INFRA_SERVICES.
- A new step removes services netOrk no longer ships: the container
  netork-registry-1, then the volume netork_registry_data. `docker
  compose up` never removes a container whose service left the compose
  file, so without this each host would keep it until someone removed it
  by hand. The step is idempotent and works with the old compose file as
  well as the new one, so it can go out before netOrk drops the service.
2026-10-07 16:17:23 +02:00
christianmanivong 9a93897cd2 Merge pull request 'fix: one deploy per host at a time, under a lock the host frees on its own' (#4) from fix/host-deploy-lock into main
CI / check (push) Successful in 26s
2026-10-07 04:43:24 +00:00
Christian Manivong ec93619f8f fix: one deploy per host at a time, under a lock the host frees on its own
CI / check (pull_request) Successful in 17s
Several sessions deploy to the same test server, and two runs used to
overlap. On 2026-10-05 two `up -d --force-recreate` runs recreated each
other's containers and the API was down for a minute (#1). On 2026-10-06
two runs renamed each other's *.new compose files and one broke off; with
two different tags, one tag's compose files could have started the
other's images (#3).

- Each deploy first takes flock on ~/netork/.deploy.lock on the host. An
  ssh session holds it: the remote side takes the lock on fd 9, reports
  LOCKED and waits on its stdin, so ending the session frees it, whether
  the deploy finished, failed, was interrupted or lost its connection.
  Checked over real ssh on .50, including a client killed with -9.
- A second deploy prints who holds the lock, since when and with which
  tag, and waits up to DEPLOY_LOCK_WAIT seconds (default 900); then it
  gives up without touching the host.
- A host without flock is deployed without the lock, with a warning.
- deploy_server is now the lock around deploy_steps, the old body.

Closes #1
Closes #3
2026-10-07 06:42:13 +02:00
christianmanivong 6444aa97f3 Merge pull request 'fix: a failed container recreate is tried once more before the deploy gives up' (#2) from fix/recreate-retry into main
CI / check (push) Successful in 12s
2026-10-05 20:03:55 +00:00
5 changed files with 365 additions and 5 deletions
+10 -1
View File
@@ -41,6 +41,7 @@ elsewhere.
| `NETORK_VERSION` | Tag to deploy (default `latest`) | | `NETORK_VERSION` | Tag to deploy (default `latest`) |
| `REGISTRY_USER`, `REGISTRY_PASSWORD` | Registry login used on every server | | `REGISTRY_USER`, `REGISTRY_PASSWORD` | Registry login used on every server |
| `REGISTRY_USER_<server>`, `REGISTRY_PASSWORD_<server>`, `NETORK_VERSION_<server>` | Per-server overrides. `<server>` has its dots replaced by underscores, e.g. `_10_0_0_2` | | `REGISTRY_USER_<server>`, `REGISTRY_PASSWORD_<server>`, `NETORK_VERSION_<server>` | Per-server overrides. `<server>` has its dots replaced by underscores, e.g. `_10_0_0_2` |
| `DEPLOY_LOCK_WAIT` | Seconds to wait for another deploy to the same host (default 900) |
## Usage ## Usage
@@ -65,6 +66,12 @@ started earlier pulls whatever image the registry held before, which is stale.
## What a deploy does, per server ## What a deploy does, per server
0. Takes the host's deploy lock, `flock` on `~/netork/.deploy.lock`, and holds it until
the deploy ends. A second deploy to the same host waits for it and says who holds
it, since when, and which tag they are deploying. It gives up after 15 minutes
(`DEPLOY_LOCK_WAIT`, in seconds) without touching the host. The lock belongs to an
ssh session, so a deploy that fails, is interrupted or loses its connection frees it
on its own. A host without `flock` is deployed without the lock, with a warning.
1. Logs in to the registry. The password travels over ssh's stdin, never on a command 1. Logs in to the registry. The password travels over ssh's stdin, never on a command
line. line.
2. Pulls the engine image and copies `docker-compose.yml` and 2. Pulls the engine image and copies `docker-compose.yml` and
@@ -74,8 +81,10 @@ started earlier pulls whatever image the registry held before, which is stale.
more after 5 s (`DEPLOY_RECREATE_RETRY_DELAY`). Compose's parallel recreate can lose more after 5 s (`DEPLOY_RECREATE_RETRY_DELAY`). Compose's parallel recreate can lose
a container it just renamed and leave the rest stopped. If the second attempt fails a container it just renamed and leave the rest stopped. If the second attempt fails
too, the container states are printed and the deploy fails. too, the container states are printed and the deploy fails.
4. Reconciles `registry`, `apt-cacher-ng` and `signal-api`: each is recreated only if its 4. Reconciles `apt-cacher-ng` and `signal-api`: each is recreated only if its
definition changed. It never touches `postgres` or `redis`. definition changed. It never touches `postgres` or `redis`.
It then removes services netOrk no longer ships, container and volume, where a host
still has them: today the bundled `registry` (NetOrk/netork#763).
5. Verifies that the containers run exactly the image that was pulled. If they don't, the 5. Verifies that the containers run exactly the image that was pulled. If they don't, the
deploy fails. deploy fails.
6. Records `REGISTRY_HOST` and `NETORK_VERSION` in `~/netork/.env`, so that a 6. Records `REGISTRY_HOST` and `NETORK_VERSION` in `~/netork/.env`, so that a
+94 -3
View File
@@ -34,6 +34,9 @@ NETORK_VERSION="${NETORK_VERSION:-latest}"
# The compose files the engine image carries under /app/deploy/. # The compose files the engine image carries under /app/deploy/.
COMPOSE_FILES="docker-compose.yml docker-compose.registry.yml" COMPOSE_FILES="docker-compose.yml docker-compose.registry.yml"
# How long a deploy waits for another one on the same host before giving up.
DEPLOY_LOCK_WAIT="${DEPLOY_LOCK_WAIT:-900}"
read -ra SERVERS <<< "$DEPLOY_SERVERS" read -ra SERVERS <<< "$DEPLOY_SERVERS"
# Run one step on a server, and let its failure stop the deploy. # Run one step on a server, and let its failure stop the deploy.
@@ -70,6 +73,63 @@ run_remote() {
fi fi
} }
# ── One deploy per host at a time ────────────────────────────────────────────
# Several sessions deploy to the same test server, and two runs used to overlap:
# on 2026-10-05 two `up -d --force-recreate` recreated each other's containers,
# on 2026-10-06 two runs renamed each other's *.new compose files, and with two
# tags one tag's compose files could start the other's images (NetOrk/deploy#1,
# #3). So each deploy first takes flock on ~/netork/.deploy.lock on the host.
#
# An ssh session holds it: the remote side opens the lock file on fd 9, takes
# the lock, says LOCKED, and then waits on its stdin. Closing that stdin ends
# the session and frees the lock, and so does anything else that ends it: the
# deploy failing, being interrupted, or losing the network. No stale lock file
# can outlive the deploy that took it.
#
# A host without flock (util-linux) is deployed without the lock, with a
# warning, rather than becoming undeployable.
#
# hold_host_lock <server> <who and what, for the next deploy to read>
hold_host_lock() {
local server=$1 info=$2 line
# shellcheck disable=SC2029 # the wait and the info are meant to expand here
coproc HOST_LOCK {
ssh "$server" "mkdir -p ~/netork && cd ~/netork && exec 9>.deploy.lock && \
if ! command -v flock >/dev/null 2>&1; then echo NOLOCK; exec cat >/dev/null; fi; \
if ! flock -n 9; then \
echo \"BUSY \$(cat .deploy.lock.info 2>/dev/null)\"; \
flock -w ${DEPLOY_LOCK_WAIT} 9 || { echo TIMEOUT; exit 1; }; \
fi; \
echo '${info}' > .deploy.lock.info; echo LOCKED; exec cat >/dev/null" 2>&1
}
while IFS= read -r -t "$(( DEPLOY_LOCK_WAIT + 60 ))" line <&"${HOST_LOCK[0]}"; do
case "$line" in
LOCKED) return 0 ;;
NOLOCK)
echo "[${server}] WARNING: no flock on ${server}; deploying without the host lock." >&2
return 0
;;
BUSY*)
echo "[${server}] another deploy holds ${server}: ${line#BUSY } — waiting up to ${DEPLOY_LOCK_WAIT}s..." >&2
;;
TIMEOUT)
echo "[${server}] the deploy lock on ${server} is still held after ${DEPLOY_LOCK_WAIT}s; giving up." >&2
return 1
;;
*) echo "[${server}] ${line}" >&2 ;;
esac
done
echo "[${server}] could not take the deploy lock on ${server} (the ssh session ended)." >&2
return 1
}
# Free the lock: closing the session's stdin ends the remote side.
release_host_lock() {
local fd="${HOST_LOCK[1]}" pid="${HOST_LOCK_PID}"
exec {fd}>&-
wait "$pid" 2>/dev/null || true
}
VERSION_OVERRIDE="" VERSION_OVERRIDE=""
EXPLICIT_SERVERS=() EXPLICIT_SERVERS=()
i=1 i=1
@@ -128,7 +188,20 @@ if [[ ${#SERVERS[@]} -eq 0 ]]; then
exit 1 exit 1
fi fi
# One server: under its host lock, the steps below.
#
# A step that fails stops the subshell (set -e) before release_host_lock runs.
# That is fine: the subshell's end closes the lock session's stdin all the same.
deploy_server() { deploy_server() {
local SERVER="$1" KEY="${1//./_}" VER_VAR
VER_VAR="NETORK_VERSION_${KEY}"
hold_host_lock "$SERVER" \
"since $(date -u +%Y-%m-%dT%H:%M:%SZ), by ${USER:-?}@$(hostname -s), deploying ${VERSION_OVERRIDE:-${!VER_VAR:-${NETORK_VERSION}}}"
deploy_steps "$SERVER"
release_host_lock
}
deploy_steps() {
local SERVER="$1" local SERVER="$1"
# Per-server lookup — falls back to global values # Per-server lookup — falls back to global values
@@ -179,7 +252,7 @@ deploy_server() {
# Containers on upstream images. Reconciled, not force-recreated (below). # Containers on upstream images. Reconciled, not force-recreated (below).
# postgres and redis are deliberately absent: restarting a database on every # postgres and redis are deliberately absent: restarting a database on every
# deploy would be worse than any compose change one could miss. # deploy would be worse than any compose change one could miss.
local INFRA_SERVICES="registry apt-cacher-ng signal-api" local INFRA_SERVICES="apt-cacher-ng signal-api"
echo "[${SERVER}] Pulling images..." echo "[${SERVER}] Pulling images..."
run_remote "$SERVER" "pulling images for ${SERVER_VERSION}" 'Pulled|Already|Error' \ run_remote "$SERVER" "pulling images for ${SERVER_VERSION}" 'Pulled|Already|Error' \
@@ -238,6 +311,24 @@ deploy_server() {
docker compose -f docker-compose.yml -f docker-compose.registry.yml \ docker compose -f docker-compose.yml -f docker-compose.registry.yml \
up -d --no-deps ${INFRA_SERVICES}" || true up -d --no-deps ${INFRA_SERVICES}" || true
# Services netOrk no longer ships, with the data only they used. `docker
# compose up` never removes a container whose service has left the compose
# file, so each would keep running on every host until someone removed it by
# hand. Idempotent: a host that never had one, or was cleaned already, passes
# untouched. Containers first; a volume still in use cannot be removed.
# registry (netork-registry-1, netork_registry_data): the bundled registry:2
# for satellite images. Satellites pull from registry.netork.io; it held one
# old image, nothing had pulled from it in 30 days, and it accepted
# anonymous pushes on port 5000 (NetOrk/netork#763).
local RETIRED_CONTAINERS="netork-registry-1"
local RETIRED_VOLUMES="netork_registry_data"
echo "[${SERVER}] Removing retired services..."
run_remote "$SERVER" "removing ${RETIRED_CONTAINERS} ${RETIRED_VOLUMES}" "" \
"for c in ${RETIRED_CONTAINERS}; do \
docker rm -f \$c >/dev/null 2>&1 && echo \"removed container \$c\"; done; \
for v in ${RETIRED_VOLUMES}; do \
docker volume rm \$v >/dev/null 2>&1 && echo \"removed volume \$v\"; done; true" || true
# Verify the containers actually run the image we just pulled. `docker # Verify the containers actually run the image we just pulled. `docker
# compose up -d` reports "Running" (not "Started") when it decides nothing # compose up -d` reports "Running" (not "Started") when it decides nothing
# changed, and `pull` prints "Pulled" even when the tag was already local — # changed, and `pull` prints "Pulled" even when the tag was already local —
@@ -317,9 +408,9 @@ REMOTE_ENV
echo "Deploying from ${REGISTRY_HOST} to: ${SERVERS[*]}" echo "Deploying from ${REGISTRY_HOST} to: ${SERVERS[*]}"
export -f deploy_server _validate_version run_remote export -f deploy_server deploy_steps hold_host_lock release_host_lock _validate_version run_remote
export UI_SERVER REGISTRY_HOST REGISTRY_USER REGISTRY_PASSWORD NETORK_VERSION VERSION_OVERRIDE \ export UI_SERVER REGISTRY_HOST REGISTRY_USER REGISTRY_PASSWORD NETORK_VERSION VERSION_OVERRIDE \
VALID_VERSION_RE COMPOSE_FILES VALID_VERSION_RE COMPOSE_FILES DEPLOY_LOCK_WAIT
# Export all per-server credential vars so subshells can resolve them # Export all per-server credential vars so subshells can resolve them
while IFS='=' read -r key _; do while IFS='=' read -r key _; do
if [[ "$key" =~ ^(REGISTRY_(USER|PASSWORD)|NETORK_VERSION)_ ]]; then if [[ "$key" =~ ^(REGISTRY_(USER|PASSWORD)|NETORK_VERSION)_ ]]; then
+185
View File
@@ -0,0 +1,185 @@
"""One deploy per host at a time (NetOrk/deploy#1, #3).
Several sessions deploy to the same test server, and nothing stopped two runs
from overlapping:
- 2026-10-05: two `up -d --force-recreate` runs recreated each other's
containers. The API was down for a minute, with leftover `<id>_netork-…`
containers.
- 2026-10-06: two runs renamed each other's `*.new` compose files, and one broke
off at `mv: cannot stat 'docker-compose.yml.new'`. With two different tags,
one tag's compose files could have started the other tag's images.
Now each deploy first takes `flock` on `~/netork/.deploy.lock` on the host,
through an ssh session that holds it until the deploy ends. A second deploy
says who holds it, since when and with which tag, and waits. The lock lives as
long as that session does, so a deploy that dies, or a laptop that loses its
network, frees it on its own.
The "host" here is this machine with `HOME` in a temporary directory: the fake
`ssh` runs the lock holder for real, so the lock is a real `flock`.
"""
from __future__ import annotations
import os
import subprocess
import time
from pathlib import Path
import pytest
DEPLOY = Path(__file__).resolve().parent.parent / "deploy.sh"
_FAKE_SSH = """#!/bin/sh
for arg in "$@"; do cmd="$arg"; done
case "$cmd" in
*.deploy.lock*)
# The lock holder runs for real, on this "host".
if [ -n "$FAKE_NO_FLOCK" ]; then exec env PATH="$FAKE_NO_FLOCK" sh -c "$cmd"; fi
exec sh -c "$cmd"
;;
esac
cat > /dev/null
case "$cmd" in
*--force-recreate*)
echo "start $RUN_ID" >> "$FAKE_LOG"
sleep "${FAKE_RECREATE_SLEEP:-0}"
echo "end $RUN_ID" >> "$FAKE_LOG"
if [ -n "$FAKE_RECREATE_FAILS" ]; then echo "Error response from daemon"; exit 1; fi
;;
esac
exit 0
"""
@pytest.fixture
def host(tmp_path: Path) -> dict[str, Path]:
bin_dir = tmp_path / "bin"
bin_dir.mkdir()
fake = bin_dir / "ssh"
fake.write_text(_FAKE_SSH)
fake.chmod(0o755)
home = tmp_path / "home"
home.mkdir()
return {"bin": bin_dir, "home": home, "log": tmp_path / "log"}
def _start(host: dict[str, Path], run_id: str, **env: str) -> subprocess.Popen[str]:
return subprocess.Popen(
["bash", str(DEPLOY), "testhost"],
env={
"PATH": f"{host['bin']}:{os.environ['PATH']}",
"HOME": str(host["home"]),
"USER": "tester",
"DEPLOY_ENV_FILE": "/dev/null",
"REGISTRY_HOST": "registry.example",
"NETORK_VERSION": "latest-dev",
"DEPLOY_RECREATE_RETRY_DELAY": "0",
"RUN_ID": run_id,
"FAKE_LOG": str(host["log"]),
**env,
},
stdin=subprocess.DEVNULL,
stdout=subprocess.PIPE,
stderr=subprocess.PIPE,
text=True,
)
def _run(host: dict[str, Path], run_id: str, **env: str) -> subprocess.CompletedProcess[str]:
proc = _start(host, run_id, **env)
out, err = proc.communicate(timeout=60)
return subprocess.CompletedProcess(proc.args, proc.returncode, out, err)
def _lock_file(host: dict[str, Path]) -> Path:
return host["home"] / "netork" / ".deploy.lock"
def _lock_is_free(host: dict[str, Path], seconds: float = 5.0) -> bool:
deadline = time.monotonic() + seconds
while time.monotonic() < deadline:
probe = subprocess.run(["flock", "-n", str(_lock_file(host)), "true"], check=False)
if probe.returncode == 0:
return True
time.sleep(0.05)
return False
def _events(host: dict[str, Path]) -> list[str]:
return host["log"].read_text().split("\n")[:-1] if host["log"].exists() else []
def _wait_for(host: dict[str, Path], event: str, seconds: float = 20.0) -> None:
deadline = time.monotonic() + seconds
while time.monotonic() < deadline:
if event in _events(host):
return
time.sleep(0.05)
raise AssertionError(f"{event!r} never happened: {_events(host)}")
def test_two_deploys_to_one_host_take_turns(host) -> None:
first = _start(host, "A", FAKE_RECREATE_SLEEP="1.5")
try:
_wait_for(host, "start A")
second = _run(host, "B")
finally:
_out, err = first.communicate(timeout=60)
assert first.returncode == 0, err
assert second.returncode == 0, second.stderr
assert _events(host) == ["start A", "end A", "start B", "end B"]
assert "another deploy holds testhost" in second.stderr
assert "deploying latest-dev" in second.stderr
assert "by tester@" in second.stderr
def test_the_lock_is_free_once_a_deploy_is_done(host) -> None:
result = _run(host, "A")
assert result.returncode == 0, result.stderr
assert _lock_file(host).exists()
assert _lock_is_free(host)
def test_a_failed_deploy_frees_the_lock_too(host) -> None:
result = _run(host, "A", FAKE_RECREATE_FAILS="1")
assert result.returncode != 0
assert _lock_is_free(host)
def test_a_deploy_that_waits_too_long_gives_up_without_touching_the_host(host) -> None:
lock = _lock_file(host)
lock.parent.mkdir(parents=True)
holder = subprocess.Popen(["flock", str(lock), "sleep", "30"])
try:
time.sleep(0.2)
result = _run(host, "B", DEPLOY_LOCK_WAIT="1")
finally:
holder.kill()
holder.wait()
assert result.returncode != 0
assert "still held after 1s" in result.stderr
assert "FAILED on: testhost" in result.stderr
assert _events(host) == []
def test_a_host_without_flock_is_deployed_with_a_warning(host, tmp_path) -> None:
"""flock is util-linux; a host without it should not become undeployable."""
no_flock = tmp_path / "no-flock"
no_flock.mkdir()
for tool in ("sh", "mkdir", "cat", "date"):
found = subprocess.run(
["sh", "-c", f"command -v {tool}"], capture_output=True, text=True, check=True
)
(no_flock / tool).symlink_to(found.stdout.strip())
result = _run(host, "A", FAKE_NO_FLOCK=str(no_flock))
assert result.returncode == 0, result.stderr
assert "no flock on testhost" in result.stderr
assert _events(host) == ["start A", "end A"]
+3 -1
View File
@@ -25,9 +25,11 @@ import pytest
DEPLOY = Path(__file__).resolve().parent.parent / "deploy.sh" DEPLOY = Path(__file__).resolve().parent.parent / "deploy.sh"
_FAKE_SSH = """#!/bin/sh _FAKE_SSH = """#!/bin/sh
for arg in "$@"; do cmd="$arg"; done
# The host lock's holder runs for real, here (test_deploy_host_lock.py).
case "$cmd" in *.deploy.lock*) exec sh -c "$cmd" ;; esac
# Swallow stdin (docker login, the .env heredoc), then answer by command. # Swallow stdin (docker login, the .env heredoc), then answer by command.
cat > /dev/null cat > /dev/null
for arg in "$@"; do cmd="$arg"; done
case "$cmd" in case "$cmd" in
*--force-recreate*) *--force-recreate*)
n=$(cat "{state}/recreates" 2>/dev/null || echo 0) n=$(cat "{state}/recreates" 2>/dev/null || echo 0)
+73
View File
@@ -0,0 +1,73 @@
"""A service netOrk no longer ships is removed from the host (NetOrk/netork#763).
`docker compose up` never removes a container whose service has left the compose
file, so it would keep running on every host until somebody removed it by hand.
The first of these is the bundled `registry:2`: it held only an old satellite
image, nothing had pulled from it in 30 days, and it accepted anonymous pushes on
port 5000 of every instance.
The remote side is a fake `ssh` that records every command it is given.
"""
from __future__ import annotations
import os
import subprocess
from pathlib import Path
DEPLOY = Path(__file__).resolve().parent.parent / "deploy.sh"
_FAKE_SSH = """#!/bin/sh
for arg in "$@"; do cmd="$arg"; done
# The host lock's holder runs for real, here (test_deploy_host_lock.py).
case "$cmd" in *.deploy.lock*) exec sh -c "$cmd" ;; esac
cat > /dev/null
printf '%s\\n----\\n' "$cmd" >> "{state}/commands"
exit 0
"""
def _deploy(tmp_path: Path) -> tuple[subprocess.CompletedProcess[str], list[str]]:
bin_dir = tmp_path / "bin"
bin_dir.mkdir()
fake = bin_dir / "ssh"
fake.write_text(_FAKE_SSH.replace("{state}", str(tmp_path)))
fake.chmod(0o755)
result = subprocess.run(
["bash", str(DEPLOY), "testhost"],
env={
"PATH": f"{bin_dir}:{os.environ['PATH']}",
"HOME": str(tmp_path),
"DEPLOY_ENV_FILE": "/dev/null",
"REGISTRY_HOST": "registry.example",
"NETORK_VERSION": "latest-dev",
},
stdin=subprocess.DEVNULL,
capture_output=True,
text=True,
timeout=60,
check=False,
)
commands = (tmp_path / "commands").read_text().split("\n----\n")
return result, commands
def test_a_deploy_removes_the_retired_registry_and_its_volume(tmp_path: Path) -> None:
result, commands = _deploy(tmp_path)
assert result.returncode == 0, result.stderr
removal = [c for c in commands if "netork-registry-1" in c]
assert removal, "no command removes the netork-registry-1 container"
assert "docker rm -f" in removal[0]
assert "docker volume rm" in removal[0] and "netork_registry_data" in removal[0]
# The container goes first: a volume still in use cannot be removed.
assert removal[0].index("docker rm -f") < removal[0].index("docker volume rm")
def test_the_infrastructure_reconcile_no_longer_starts_the_registry(tmp_path: Path) -> None:
result, commands = _deploy(tmp_path)
assert result.returncode == 0, result.stderr
reconcile = [c for c in commands if "up -d --no-deps" in c and "--force-recreate" not in c]
assert reconcile, "the infrastructure reconcile step is missing"
assert " registry" not in reconcile[0].split("up -d --no-deps", 1)[1]