Files
deploy/README.md
T
Christian Manivong 923ec4ce65
CI / check (pull_request) Successful in 13s
fix: a failed container recreate is tried once more before the deploy gives up
On 2026-10-05 a deploy to 172.22.8.50 failed in "Starting containers":
docker compose up -d --force-recreate recreates the services in parallel,
lost a container it had just renamed ("No such container: 02586df7…") and
stopped with every engine container Created and none running. The old API
was already gone, so netOrk was down for about two minutes, until the same
deploy was run again and went through cleanly.

The script now does that second run itself, after DEPLOY_RECREATE_RETRY_DELAY
seconds (5 by default). If the second attempt fails too, it prints the
container states (docker compose ps -a) and fails as before.

Tests run the script against a fake ssh that fails the recreate zero, one
or two times. README step 3 says what happens.

Refs NetOrk/netork#586
2026-10-05 21:58:36 +02:00

116 lines
4.3 KiB
Markdown

# netOrk deploy
Deploys [netOrk](https://git.netork.io/NetOrk/netork) to one or more Docker hosts in
parallel, from pre-built images in a container registry.
The machine running the deploy needs this repository, `ssh`, and a `deploy.env`. It does
not need a netOrk checkout. Each server pulls the images itself. The compose files come
out of the engine image for the tag being deployed, so they always match the images they
start, and database migrations run from that same image.
## Requirements
**On the machine running the deploy:**
- bash
- `ssh` with key-based access to every target
**On every target server:**
- Docker with the compose plugin
- a directory `~/netork/` with netOrk's `.env`, which holds the database and application
settings the compose file reads (`env_file: .env`)
**Images:** `netork/engine` and `netork/ui` must be published in the registry under the
tag you deploy. The engine image must carry its compose files under `/app/deploy/`.
Tags built before that change are refused with a clear message.
## Setup
```bash
cp deploy.env.example deploy.env
$EDITOR deploy.env
```
`deploy.env` is gitignored. Point `DEPLOY_ENV_FILE` at another file to keep it
elsewhere.
| Variable | Meaning |
|---|---|
| `DEPLOY_SERVERS` | Space-separated default targets |
| `UI_SERVER` | Targets that also run `netork-ui` |
| `REGISTRY_HOST` | Registry to pull from; required |
| `NETORK_VERSION` | Tag to deploy (default `latest`) |
| `REGISTRY_USER`, `REGISTRY_PASSWORD` | Registry login used on every server |
| `REGISTRY_USER_<server>`, `REGISTRY_PASSWORD_<server>`, `NETORK_VERSION_<server>` | Per-server overrides. `<server>` has its dots replaced by underscores, e.g. `_10_0_0_2` |
## Usage
```bash
./deploy.sh # every server in DEPLOY_SERVERS
./deploy.sh 10.0.0.1 # one server
./deploy.sh 10.0.0.1 10.0.0.2 # several, in parallel
./deploy.sh --version=main-1a2b3c4 10.0.0.1 # one-off tag, this run only
```
Only tags netOrk's CI publishes are accepted:
- `latest`
- `latest-dev`
- `main-<sha>`
- `feature-<branch>`
- `X.Y.Z`
Anything else is refused before any server is touched.
If you deploy a branch build, wait until its CI run has published the images. A deploy
started earlier pulls whatever image the registry held before, which is stale.
## What a deploy does, per server
1. Logs in to the registry. The password travels over ssh's stdin, never on a command
line.
2. Pulls the engine image and copies `docker-compose.yml` and
`docker-compose.registry.yml` out of it into `~/netork/`.
3. Pulls the netOrk images, then force-recreates the API, the workers, `netork-beat`,
`flower` and, on UI servers, `netork-ui`. If the recreate fails, it is tried once
more after 5 s (`DEPLOY_RECREATE_RETRY_DELAY`). Compose's parallel recreate can lose
a container it just renamed and leave the rest stopped. If the second attempt fails
too, the container states are printed and the deploy fails.
4. Reconciles `registry`, `apt-cacher-ng` and `signal-api`: each is recreated only if its
definition changed. It never touches `postgres` or `redis`.
5. Verifies that the containers run exactly the image that was pulled. If they don't, the
deploy fails.
6. Records `REGISTRY_HOST` and `NETORK_VERSION` in `~/netork/.env`, so that a
hand-typed `docker compose` on the host uses the same images.
7. Runs `alembic upgrade head` inside `netork-api`.
8. Prunes unused images.
A failing step stops that server's deploy with a non-zero exit, and the other servers
carry on. The script exits non-zero if any server failed and names those servers.
## Operating notes
- **Never delete these Docker volumes:**
- `celerybeat_schedule` holds the scheduler state.
- `signal_cli_data` holds netOrk's Signal device link. Losing it means pairing again
by QR code.
- `netork-beat` always rolls out together with the workers. The script does this for
you; keep it that way if you deploy by hand.
- Always deploy with this script. Do not point a local Docker client at a remote host
(`DOCKER_HOST=ssh://…`): it resolves volume paths locally and breaks the remote
containers.
## Development
```bash
pip install pytest ruff shellcheck-py
shellcheck deploy.sh
ruff format --check . && ruff check .
pytest
```
The tests are static and behavioural checks of `deploy.sh`. None of them reach a real
host.
## License
MIT, see [LICENSE](LICENSE).