deploy.sh: two deploys to one host at once recreate each other's containers #1

Open
opened 2026-10-05 05:45:07 +00:00 by christianmanivong · 0 comments
Owner

What happened (2026-10-05, 07:42 CEST, 172.22.8.50)

Two sessions ran deploy.sh 172.22.8.50 (both latest-dev) within the same minute. Each run's docker compose up -d --force-recreate stopped and renamed the containers the other run was creating:

  • Error when allocating new name: Conflict. The container name "/netork-netork-worker-1" is already in use by container "4e895ef6…"
  • leftover containers named <id>_netork-netork-api-1 / <id>_netork-netork-worker-1 in state Created
  • dependency failed to start: container netork-netork-api-1 exited (137) — killed by the other run's recreate, not by OOM (OOMKilled=false)

For about a minute the API and the workers were down, and the UI container sat in Created. It recovered only because one of the two runs happened to finish last and cleanly. Nothing in the tooling made that the outcome: a slightly different interleaving would have left the host half-recreated, with both runs reporting FAILED.

Why it happens

deploy.sh holds no lock. Nothing on the target host stops a second docker compose up while one is in flight, and nothing on the workstation does either. With several sessions deploying to the same test server, overlap is a matter of time.

Options

  • A lock on the target host, around the compose and migration steps: flock on a file under ~/netork, taken over ssh. It also covers deploys from different workstations. A second run should either wait with a visible "another deploy is running on since
  • A local lock alone (flock on the workstation) would not cover another machine, and the test server is the shared resource.

The migration step matters here too: two alembic upgrade head runs at once are exactly what this would serialise.

## What happened (2026-10-05, 07:42 CEST, 172.22.8.50) Two sessions ran `deploy.sh 172.22.8.50` (both `latest-dev`) within the same minute. Each run's `docker compose up -d --force-recreate` stopped and renamed the containers the other run was creating: - `Error when allocating new name: Conflict. The container name "/netork-netork-worker-1" is already in use by container "4e895ef6…"` - leftover containers named `<id>_netork-netork-api-1` / `<id>_netork-netork-worker-1` in state `Created` - `dependency failed to start: container netork-netork-api-1 exited (137)` — killed by the other run's recreate, not by OOM (`OOMKilled=false`) For about a minute the API and the workers were down, and the UI container sat in `Created`. It recovered only because one of the two runs happened to finish last and cleanly. Nothing in the tooling made that the outcome: a slightly different interleaving would have left the host half-recreated, with both runs reporting `FAILED`. ## Why it happens `deploy.sh` holds no lock. Nothing on the target host stops a second `docker compose up` while one is in flight, and nothing on the workstation does either. With several sessions deploying to the same test server, overlap is a matter of time. ## Options - **A lock on the target host**, around the compose and migration steps: `flock` on a file under `~/netork`, taken over ssh. It also covers deploys from different workstations. A second run should either wait with a visible "another deploy is running on <host> since <time>" or fail fast with that message. - A local lock alone (`flock` on the workstation) would not cover another machine, and the test server is the shared resource. The migration step matters here too: two `alembic upgrade head` runs at once are exactly what this would serialise.
christianmanivong added the
prio
P3
label 2026-10-05 05:45:07 +00:00
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: NetOrk/deploy#1