deploy.sh: two deploys to the same host at once collide on the compose files #3

Closed
opened 2026-10-06 22:12:50 +00:00 by christianmanivong · 1 comment
Owner

Two deploy.sh runs against the same host at the same time write the same compose files in ~/netork, and one of them breaks off halfway.

Observed

On 2026-10-06 at 22:10 UTC on 172.22.8.50:

  • Session A was running deploy.sh (latest-dev) and had just started docker compose ... up -d --force-recreate.
  • Session B started deploy.sh --version=main-443dd11 at the same moment.
  • B failed at fetching compose files from …main-443dd11 with mv: cannot stat 'docker-compose.yml.new': No such file or directory (also for docker-compose.registry.yml.new).
  • Both had written and renamed docker-compose.yml.new / docker-compose.registry.yml.new in the same directory.

A's deploy finished cleanly, so the host ended healthy. With other timing, though, the compose files of one image could be used to start the containers of the other, or up could run on files that are being replaced.

Expected

Only one deploy per host at a time. A second one waits, or stops at once with "another deploy is running since …, started by …".

Option

A lock on the remote host for the whole remote part, e.g. flock -n ~/netork/.deploy.lock around run_remote. Several parallel Claude sessions deploy to .50, so this comes up in practice.

Found

By an MVP 5 session while deploying N8, 2026-10-06.

Two `deploy.sh` runs against the same host at the same time write the same compose files in `~/netork`, and one of them breaks off halfway. ## Observed On 2026-10-06 at 22:10 UTC on 172.22.8.50: - Session A was running `deploy.sh` (`latest-dev`) and had just started `docker compose ... up -d --force-recreate`. - Session B started `deploy.sh --version=main-443dd11` at the same moment. - B failed at `fetching compose files from …main-443dd11` with `mv: cannot stat 'docker-compose.yml.new': No such file or directory` (also for `docker-compose.registry.yml.new`). - Both had written and renamed `docker-compose.yml.new` / `docker-compose.registry.yml.new` in the same directory. A's deploy finished cleanly, so the host ended healthy. With other timing, though, the compose files of one image could be used to start the containers of the other, or `up` could run on files that are being replaced. ## Expected Only one deploy per host at a time. A second one waits, or stops at once with "another deploy is running since …, started by …". ## Option A lock on the remote host for the whole remote part, e.g. `flock -n ~/netork/.deploy.lock` around `run_remote`. Several parallel Claude sessions deploy to .50, so this comes up in practice. ## Found By an MVP 5 session while deploying N8, 2026-10-06.
christianmanivong added the
prio
P3
label 2026-10-06 22:12:50 +00:00
Author
Owner

Session A was mine (the bug sweep). Two details for whoever fixes this:

  • Same root cause as #1. There, two up -d --force-recreate runs recreated each other's containers. Here, two runs renamed each other's *.new compose files. One host-side flock around the whole remote part (compose files, up, migrations) fixes both, so one of the two can be closed as a duplicate when it lands.
  • Here both runs were deploying the same commit. latest-dev had moved to main-443dd11 at 22:09:59 UTC, and B asked for --version=main-443dd11. The dangerous variant, where one tag's compose files start the other tag's images, needs two different versions. That is just as likely, because sessions deploy feature tags to .50 too.

Until there is a lock, my sessions check ps aux | grep "docker compose" on the host before deploying.

Session A was mine (the bug sweep). Two details for whoever fixes this: - **Same root cause as #1.** There, two `up -d --force-recreate` runs recreated each other's containers. Here, two runs renamed each other's `*.new` compose files. One host-side `flock` around the whole remote part (compose files, `up`, migrations) fixes both, so one of the two can be closed as a duplicate when it lands. - **Here both runs were deploying the same commit.** `latest-dev` had moved to `main-443dd11` at 22:09:59 UTC, and B asked for `--version=main-443dd11`. The dangerous variant, where one tag's compose files start the other tag's images, needs two different versions. That is just as likely, because sessions deploy feature tags to .50 too. Until there is a lock, my sessions check `ps aux | grep "docker compose"` on the host before deploying.
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: NetOrk/deploy#3