Overlapping Resource Ownership Across Three Working Directories

Resources duplicated across three states and connections maintained through manually copied names

In part 20 a single main.tf split into three working directories. The server, preemption recovery, and backup cleanup each held their own state and deployed through separate apply runs. Changing one service no longer required reading the plans for the other two.

The three archived states still contain targets that were not separated. Some resources sat in two states at once, and some connections used only a string a person had copied by hand.

I reconstructed which state managed each resource by reading only the type and name of each managed resource in the three states. I did not read attribute values or real identifiers, and the names below are placeholders. A state is a snapshot of the moment it was saved, so none of this means those resources exist now.

What the Three States Managed

server                 8   static IP, 5 firewall rules, instance, backup bucket
preemption recovery   15   function, log metric, alert policy and channel, 3 IAM, 5 API, topic, bucket, object
backup cleanup        15   Scheduler, function, 2 IAM, 8 API, topic, bucket, object

The server state has no API resources at all. The Compute API needed to create the server was already on, and I never brought it under code management. The other two directories each put the APIs they needed into their own state.

Two States Managed the Same API

Five google_project_service resources were managed by preemption recovery and backup cleanup at the same time.

cloudbuild  cloudfunctions  logging  monitoring  pubsub

Both directories deploy through Cloud Functions, go through a build, take triggers from Pub/Sub, and use logging and monitoring. Each wrote down the APIs it needed in its own code, and this duplication is the result. Read the code alone and each directory looks self-sufficient in the APIs it needs.

An API has one enabled state per project. Both states recorded that same API as something they managed.

flowchart TB
  R["Preemption recovery state"] --> API["5 project APIs<br/>cloudbuild·cloudfunctions<br/>logging·monitoring·pubsub"]
  C["Backup cleanup state"] --> API
  API --> REAL["One switch actually flipped on the project"]

Run destroy in one directory and that state tries to turn off the APIs it manages. Terraform does not check whether the other state is using that API. In the other direction, if one side turns an API back on, the other side’s next plan shows no change at all, because the real state already matches what it wants.

In this structure the order in which the two directories are torn down changes the result. Which one to destroy first is written nowhere in the repository.

Connections Held Together Only by a Name

The preemption recovery function restarts the server VM. The backup cleanup function empties the backup bucket the server created. Both directories received the server resource names through variable defaults.

variable "instance_name" {
  default = "<INSTANCE_NAME>"
}

The same name sits in the server directory too. The same string is written in two places and a person keeps them matched. The backup bucket name is written again the same way, in a variable in the cleanup directory.

None of the three roots has a terraform_remote_state or a module reference. There is no place where one root’s output is consumed by another. The moment the server name changes, the other two directories point at something that does not exist, and that only surfaces when the function runs.

There is also a trace of the defaults drifting apart. The default region in the server directory and the one in the backup cleanup directory differ. Whether the backup cleanup function really lived in a different region, or whether it was deployed by another path without the default ever being fixed, the code alone cannot tell. What can be confirmed is that three directories working on the same project each carried their own defaults, and nothing existed to check whether those values agreed.

Two Places Decided the Retention Policy

Backup retention was configured in two places. The server state owns the backup bucket, and the bucket itself carries a lifecycle rule.

lifecycle_rule {
  action    { type = "Delete" }
  condition { age = 1 }
}

The backup cleanup directory uses Scheduler and a function to delete old objects in that same bucket once a day. It deletes the contents of a bucket it does not own.

flowchart LR
  S["Server state"] -->|owns| B["Backup bucket"]
  S -->|lifecycle 1 day| B
  C["Backup cleanup state"] -->|Scheduler + function| B

Both settings were one day, so the effective retention was the same. But raise only one of them and the shorter one still applies, and when backups disappear sooner than expected there are two places to look.

Resource addresses overlap in one place too. google_storage_bucket_object.function_zip exists in both the preemption recovery and the backup cleanup state. They are different objects in different buckets, so this is not a collision, but in a log or a plan the address alone does not tell you which directory it belongs to.

Deployment Units and Resource Ownership

Splitting the directories let me review the server and recovery-function plans separately. I never decided which state managed each resource or how values would reach the other states. So two places managed the APIs together, a person copied the names by hand, and the retention policy split across two places.

If I split it again, I would apply these three criteria.

  1. Give each resource exactly one managing state. Manage project-wide resources such as APIs in one shared setup root.
  2. Export instance and bucket names from the owning state and pass them to consumers as explicit inputs. A name change must appear in the next plan.
  3. Define retention in either the bucket lifecycle or the cleanup function, not both.

Splitting it again, I would make the three deployment units shared project setup, the Minecraft server, and preemption recovery. With one-day retention, the backup cleanup function goes away and the bucket lifecycle handles it alone.

bootstrap/
  project APIs
  remote state bucket
  shared service account

server/
  VPC and subnet
  firewall rules and static IP
  VM
  backup bucket and lifecycle rule

spot-recovery/
  Logging sink
  Pub/Sub
  recovery function

In this layout, project API changes appear only in the bootstrap plan. Remove the other services first and clean up bootstrap last so their APIs are not disabled early. The project, zone, and instance name exported by server are passed as explicit inputs to spot-recovery. Retention is decided in one place, the bucket lifecycle.

The GCS backend uses a bucket that already exists. bootstrap creates the state bucket with local state first, then moves it with terraform init -migrate-state. server and spot-recovery use different prefixes in the same bucket.

The public example in the build track applies some of these criteria. The server root in part 04 owns the backup bucket and its lifecycle together, and that is the only place retention is decided. The recovery root in part 09 creates no server resources and takes only the target instance name as input.

In the public example, the operator still enters the instance name in two places, and the Compute and Storage APIs overlap across two roots. The google_project_service resources in both roots set disable_on_destroy = false, so destroying one root does not disable a shared API. That setting prevents an outage, but it does not remove the duplicate project service from the two states.

The deletion procedure removes the recovery root first so that a late preemption event cannot start a server while it is being deleted. That order is not a solution to the overlapping API ownership.

Before splitting the directories, I should have decided which state managed each resource and how values crossed between roots. Without those decisions, the original configuration overlapped APIs and retention policies.

References

Comments

Comments

    Image preview