Rebounder Tech Blog

Written by the people who actually run these systems in production.

migrate_image_tag drifts from live after concurrent deploys

Published About 4 min readBy the Rebounder engineering team — the people who operate these systems

This article may contain affiliate links. Its content is not affected by advertising.

In short

Terraform records a deploy as write-locals-then-commit-then-apply, so concurrent sessions can leave one commit landing after another's — the ledger stays on a stale tag though live was never broken.

Starting with the conclusion

Terraform can only record a production action in one order: write the tag into locals, commit, then apply. When two sessions deploy to Cloud Run at the same time, the commit recording one session’s already-executed action (a migration Job run, a staging deploy) can land on main after the other session’s unrelated commit. Live itself was correct in both cases — only the ledger (Terraform’s locals) was left pointing at a stale tag for a while, an unrecorded drift.

Symptom

On 2026-06-16, after shipping a header-layout fix (#968), one session ran a staging deploy followed by a prod migration. Neither operation failed — both staging and prod responded with HTTP 200.

Meanwhile, a separate concurrent session had already committed two prod web bumps, #969 and #972. By the time this session went to commit a record of what it had just done, main already contained those other commits.

# Still on main when this session tried to record its own work
# (prod/main.tf:169, staging/main.tf:79)
migrate_image_tag = "6618708" # bumped for migration 20260615120000 (#954)

# (staging/main.tf:228)
web_image_tag = "6618708" # editor overhaul + academic_year removal. 200 OK

What this session had actually run in production was migration 20260615164918 (changing audit_log.occurred_at’s DEFAULT to clock_timestamp(), fixing a false-positive in the audit hash chain from #965), plus a staging deploy of both the web and migrate images. Both had already executed — but neither was reflected in main.tf’s locals yet.

Cause

Nothing was wrong with the actions themselves. The real problem was that “running something in production” and “committing that fact to Terraform’s locals” are not the same transaction.

  • This session: ran the staging deploy → ran the prod migration Job → moved on before committing a record of either
  • The concurrent session: #969 (bump prod web to 0a01b5c) → #972 (bump prod web to 8aefde1, which already included #968)

infrastructure/** was already documented as a serialized lane requiring a token (docs/parallel-lanes.md, line 54):

| **infra lane** | `infrastructure/**` (Terraform) | serial (tf state) | infra-token |

But that discipline exists to stop two sessions writing the same line at the same time. This incident touched two lines in staging and one in prod that never overlapped with the line the concurrent session touched. Nothing collided at the line level — the drift came from the order in which the recording commits landed, not from what they edited.

The practical effect: running terraform plan at that point would have shown a diff between the committed 6618708 and what was actually live — the committed value had fallen behind live, and nothing flagged it automatically.

Fix

Every action that had actually executed got reflected in locals, and terraform plan returning No changes on both environments confirmed committed == live again.

# infrastructure/terraform/envs/prod/main.tf:169
migrate_image_tag = "227a512" # migration 20260615164918: audit_log.occurred_at DEFAULT ->
                               # clock_timestamp() (fixes #965, DEFAULT-only, non-breaking).
                               # Verified on staging first, then run in prod.

# infrastructure/terraform/envs/staging/main.tf:79, 228
migrate_image_tag = "227a512" # same migration, run on staging
web_image_tag     = "227a512" # #968 header fix + #967 WYSIWYG direct editing + #965 fix.
                               # Note: prod had already moved ahead to 8aefde1 (via #970) concurrently

Prod’s web_image_tag wasn’t touched in this fix — #972 had already committed it as 8aefde1, and 8aefde1 included #968 as an ancestor, so committed and live already agreed on that field. The fix was scoped to exactly the actions this session had run but nobody had committed yet.

Why this wasn’t caught sooner

What makes this kind of drift dangerous is that apply never fails and never errors. terraform apply faithfully pushes whatever is currently committed to live — it has no way to check whether the committed value correctly represents the work that was actually done. The only way to notice is to actively run terraform plan and compare it against what you know you actually ran.

The infra lane discipline in docs/parallel-lanes.md is designed to stop two sessions writing the same line at the same time. It doesn’t by itself prevent this pattern — where the lines never overlap, but the timing of the recording commit does. In the end, the only real defense is operational: commit the locals update for whatever you just ran before moving on to anything else.

Frequently asked questions

Q1If terraform plan shows No changes, doesn't that mean there's no drift?

The opposite — that's why this drift is hard to see. No changes confirmed the fix worked, after locals were corrected. Before that, committed and live didn't match, and nothing surfaced it. plan only compares committed vs. live; it can't tell you if committed reflects what you actually did.

Q2Wouldn't a git merge conflict have caught this?

No conflict occurred. The sessions edited different lines, so git's line-based merge combined them cleanly. The issue wasn't merge correctness — a commit recording an already-executed prod action landed after another session's unrelated commit, leaving the gap unrecorded for a while.

Q3Did this repo not already have a discipline to prevent this kind of collision?

A discipline serializing infra-touching lanes via a token already existed, documented before this incident. But it stops two sessions writing the same line at once — it doesn't enforce the timing of when an already-completed action gets committed relative to other unrelated commits.

Q4Why was only prod's migrate_image_tag left stale, not web_image_tag?

Prod's web_image_tag was already committed correctly by a separate PR bumping it to a newer image, so that field matched live already. Only the migration this session had run was left unrecorded, pending a commit alongside the matching staging changes.

Environment verified

  • Terraform >=1.9.0 / Cloud Run (Google Cloud)
  • Discovered and corrected the same day, 2026-06-16

What this article is based on

  • Terraform file lines 169-169commit b6e5c47
  • Terraform file lines 169-169commit 3a8f39f
  • Terraform file lines 79-79commit b6e5c47
  • Terraform file lines 228-228commit 3a8f39f
  • Markdown file lines 54-54commit 580d9ec

Every claim in this article comes from the records above. The repositories we operate are private so we cannot link to them, but which file, which lines, and at which commit we read them is recorded for every article. Nothing here is written from guesswork.