Overlapping Resource Ownership Across Three Working Directories
Resources duplicated across three states and connections maintained through manually copied names
In part 20 a single
main.tf split into three working directories. The server, preemption recovery, and backup cleanup
each held their own state and deployed through separate apply runs. Changing one service no longer
required reading the plans for the other two.
The three archived states still contain targets that were not separated. Some resources sat in two states at once, and some connections used only a string a person had copied by hand.
I reconstructed which state managed each resource by reading only the type and name of each managed resource in the three states. I did not read attribute values or real identifiers, and the names below are placeholders. A state is a snapshot of the moment it was saved, so none of this means those resources exist now.
What the Three States Managed
server 8 static IP, 5 firewall rules, instance, backup bucket
preemption recovery 15 function, log metric, alert policy and channel, 3 IAM, 5 API, topic, bucket, object
backup cleanup 15 Scheduler, function, 2 IAM, 8 API, topic, bucket, object
The server state has no API resources at all. The Compute API needed to create the server was already on, and I never brought it under code management. The other two directories each put the APIs they needed into their own state.
Two States Managed the Same API
Five google_project_service resources were managed by preemption recovery and backup cleanup at
the same time.
cloudbuild cloudfunctions logging monitoring pubsub
Both directories deploy through Cloud Functions, go through a build, take triggers from Pub/Sub, and use logging and monitoring. Each wrote down the APIs it needed in its own code, and this duplication is the result. Read the code alone and each directory looks self-sufficient in the APIs it needs.
An API has one enabled state per project. Both states recorded that same API as something they managed.
flowchart TB R["Preemption recovery state"] --> API["5 project APIs<br/>cloudbuild·cloudfunctions<br/>logging·monitoring·pubsub"] C["Backup cleanup state"] --> API API --> REAL["One switch actually flipped on the project"]
Run destroy in one directory and that state tries to turn off the APIs it manages. Terraform does
not check whether the other state is using that API. In the other direction, if one side turns an
API back on, the other side’s next plan shows no change at all, because the real state already
matches what it wants.
In this structure the order in which the two directories are torn down changes the result. Which one
to destroy first is written nowhere in the repository.
Connections Held Together Only by a Name
The preemption recovery function restarts the server VM. The backup cleanup function empties the backup bucket the server created. Both directories received the server resource names through variable defaults.
variable "instance_name" {
default = "<INSTANCE_NAME>"
}
The same name sits in the server directory too. The same string is written in two places and a person keeps them matched. The backup bucket name is written again the same way, in a variable in the cleanup directory.
None of the three roots has a terraform_remote_state or a module reference. There is no place
where one root’s output is consumed by another. The moment the server name changes, the other two
directories point at something that does not exist, and that only surfaces when the function runs.
There is also a trace of the defaults drifting apart. The default region in the server directory and the one in the backup cleanup directory differ. Whether the backup cleanup function really lived in a different region, or whether it was deployed by another path without the default ever being fixed, the code alone cannot tell. What can be confirmed is that three directories working on the same project each carried their own defaults, and nothing existed to check whether those values agreed.
Two Places Decided the Retention Policy
Backup retention was configured in two places. The server state owns the backup bucket, and the bucket itself carries a lifecycle rule.
lifecycle_rule {
action { type = "Delete" }
condition { age = 1 }
}
The backup cleanup directory uses Scheduler and a function to delete old objects in that same bucket once a day. It deletes the contents of a bucket it does not own.
flowchart LR S["Server state"] -->|owns| B["Backup bucket"] S -->|lifecycle 1 day| B C["Backup cleanup state"] -->|Scheduler + function| B
Both settings were one day, so the effective retention was the same. But raise only one of them and the shorter one still applies, and when backups disappear sooner than expected there are two places to look.
Resource addresses overlap in one place too. google_storage_bucket_object.function_zip exists in
both the preemption recovery and the backup cleanup state. They are different objects in different
buckets, so this is not a collision, but in a log or a plan the address alone does not tell you
which directory it belongs to.
Deployment Units and Resource Ownership
Splitting the directories let me review the server and recovery-function plans separately. I never decided which state managed each resource or how values would reach the other states. So two places managed the APIs together, a person copied the names by hand, and the retention policy split across two places.
If I split it again, I would apply these three criteria.
- Give each resource exactly one managing state. Manage project-wide resources such as APIs in one shared setup root.
- Export instance and bucket names from the owning state and pass them to consumers as explicit
inputs. A name change must appear in the next
plan. - Define retention in either the bucket lifecycle or the cleanup function, not both.
Splitting it again, I would make the three deployment units shared project setup, the Minecraft server, and preemption recovery. With one-day retention, the backup cleanup function goes away and the bucket lifecycle handles it alone.
bootstrap/
project APIs
remote state bucket
shared service account
server/
VPC and subnet
firewall rules and static IP
VM
backup bucket and lifecycle rule
spot-recovery/
Logging sink
Pub/Sub
recovery function
In this layout, project API changes appear only in the bootstrap plan. Remove the other services
first and clean up bootstrap last so their APIs are not disabled early. The project, zone, and
instance name exported by server are passed as explicit inputs to spot-recovery. Retention is
decided in one place, the bucket lifecycle.
The GCS backend uses a bucket that already
exists. bootstrap creates the
state bucket with local state first, then moves it with terraform init -migrate-state. server
and spot-recovery use different prefixes in the same bucket.
The public example in the build track applies some of these criteria. The server root in part 04 owns the backup bucket and its lifecycle together, and that is the only place retention is decided. The recovery root in part 09 creates no server resources and takes only the target instance name as input.
In the public example, the operator still enters the instance name in two places, and the Compute
and Storage APIs overlap across two roots. The google_project_service resources in both roots set
disable_on_destroy = false, so destroying one root does not disable a shared API. That setting
prevents an outage, but it does not remove the duplicate project service from the two states.
The deletion procedure removes the recovery root first so that a late preemption event cannot start a server while it is being deleted. That order is not a solution to the overlapping API ownership.
Before splitting the directories, I should have decided which state managed each resource and how values crossed between roots. Without those decisions, the original configuration overlapped APIs and retention policies.
References
Comments
No comments yet. Be the first to leave one.
Pending review