Christian Manivong ec93619f8f
CI / check (pull_request) Successful in 17s
fix: one deploy per host at a time, under a lock the host frees on its own
Several sessions deploy to the same test server, and two runs used to
overlap. On 2026-10-05 two `up -d --force-recreate` runs recreated each
other's containers and the API was down for a minute (#1). On 2026-10-06
two runs renamed each other's *.new compose files and one broke off; with
two different tags, one tag's compose files could have started the
other's images (#3).

- Each deploy first takes flock on ~/netork/.deploy.lock on the host. An
  ssh session holds it: the remote side takes the lock on fd 9, reports
  LOCKED and waits on its stdin, so ending the session frees it, whether
  the deploy finished, failed, was interrupted or lost its connection.
  Checked over real ssh on .50, including a client killed with -9.
- A second deploy prints who holds the lock, since when and with which
  tag, and waits up to DEPLOY_LOCK_WAIT seconds (default 900); then it
  gives up without touching the host.
- A host without flock is deployed without the lock, with a warning.
- deploy_server is now the lock around deploy_steps, the old body.

Closes #1
Closes #3
2026-10-07 06:42:13 +02:00

netOrk deploy

Deploys netOrk to one or more Docker hosts in parallel, from pre-built images in a container registry.

The machine running the deploy needs this repository, ssh, and a deploy.env. It does not need a netOrk checkout. Each server pulls the images itself. The compose files come out of the engine image for the tag being deployed, so they always match the images they start, and database migrations run from that same image.

Requirements

On the machine running the deploy:

  • bash
  • ssh with key-based access to every target

On every target server:

  • Docker with the compose plugin
  • a directory ~/netork/ with netOrk's .env, which holds the database and application settings the compose file reads (env_file: .env)

Images: netork/engine and netork/ui must be published in the registry under the tag you deploy. The engine image must carry its compose files under /app/deploy/. Tags built before that change are refused with a clear message.

Setup

cp deploy.env.example deploy.env
$EDITOR deploy.env

deploy.env is gitignored. Point DEPLOY_ENV_FILE at another file to keep it elsewhere.

Variable Meaning
DEPLOY_SERVERS Space-separated default targets
UI_SERVER Targets that also run netork-ui
REGISTRY_HOST Registry to pull from; required
NETORK_VERSION Tag to deploy (default latest)
REGISTRY_USER, REGISTRY_PASSWORD Registry login used on every server
REGISTRY_USER_<server>, REGISTRY_PASSWORD_<server>, NETORK_VERSION_<server> Per-server overrides. <server> has its dots replaced by underscores, e.g. _10_0_0_2
DEPLOY_LOCK_WAIT Seconds to wait for another deploy to the same host (default 900)

Usage

./deploy.sh                          # every server in DEPLOY_SERVERS
./deploy.sh 10.0.0.1                 # one server
./deploy.sh 10.0.0.1 10.0.0.2        # several, in parallel
./deploy.sh --version=main-1a2b3c4 10.0.0.1   # one-off tag, this run only

Only tags netOrk's CI publishes are accepted:

  • latest
  • latest-dev
  • main-<sha>
  • feature-<branch>
  • X.Y.Z

Anything else is refused before any server is touched.

If you deploy a branch build, wait until its CI run has published the images. A deploy started earlier pulls whatever image the registry held before, which is stale.

What a deploy does, per server

  1. Takes the host's deploy lock, flock on ~/netork/.deploy.lock, and holds it until the deploy ends. A second deploy to the same host waits for it and says who holds it, since when, and which tag they are deploying. It gives up after 15 minutes (DEPLOY_LOCK_WAIT, in seconds) without touching the host. The lock belongs to an ssh session, so a deploy that fails, is interrupted or loses its connection frees it on its own. A host without flock is deployed without the lock, with a warning.
  2. Logs in to the registry. The password travels over ssh's stdin, never on a command line.
  3. Pulls the engine image and copies docker-compose.yml and docker-compose.registry.yml out of it into ~/netork/.
  4. Pulls the netOrk images, then force-recreates the API, the workers, netork-beat, flower and, on UI servers, netork-ui. If the recreate fails, it is tried once more after 5 s (DEPLOY_RECREATE_RETRY_DELAY). Compose's parallel recreate can lose a container it just renamed and leave the rest stopped. If the second attempt fails too, the container states are printed and the deploy fails.
  5. Reconciles registry, apt-cacher-ng and signal-api: each is recreated only if its definition changed. It never touches postgres or redis.
  6. Verifies that the containers run exactly the image that was pulled. If they don't, the deploy fails.
  7. Records REGISTRY_HOST and NETORK_VERSION in ~/netork/.env, so that a hand-typed docker compose on the host uses the same images.
  8. Runs alembic upgrade head inside netork-api.
  9. Prunes unused images.

A failing step stops that server's deploy with a non-zero exit, and the other servers carry on. The script exits non-zero if any server failed and names those servers.

Operating notes

  • Never delete these Docker volumes:
    • celerybeat_schedule holds the scheduler state.
    • signal_cli_data holds netOrk's Signal device link. Losing it means pairing again by QR code.
  • netork-beat always rolls out together with the workers. The script does this for you; keep it that way if you deploy by hand.
  • Always deploy with this script. Do not point a local Docker client at a remote host (DOCKER_HOST=ssh://…): it resolves volume paths locally and breaks the remote containers.

Development

pip install pytest ruff shellcheck-py
shellcheck deploy.sh
ruff format --check . && ruff check .
pytest

The tests are static and behavioural checks of deploy.sh. None of them reach a real host.

License

MIT, see LICENSE.

S
Description
Deploys netOrk to Docker hosts from pre-built registry images
Readme MIT
81 KiB
Languages
Python 52.3%
Shell 47.7%