Deploying the Server, Recovery, and Backup Separately
How the server, recovery function, and backup cleanup were deployed from three working directories with independent plan and apply cycles
The deploy procedure in start-minecraft-server/apply.bat was short. It built the Discord bot into a
ZIP and then applied Terraform straight away.
powershell -NoLogo -NoProfile -Command ^
"Compress-Archive -Path 'scripts\discord-v2\*' ^
-DestinationPath 'scripts\discord-v2.zip' -Force"
terraform apply -auto-approve
The batch file in the preemption recovery directory had the same shape. It built the Python function
ZIP and ran terraform apply -auto-approve. They sat in one repository, but the server and the
recovery function were deployed by different commands. Backup cleanup had its own working directory,
but no apply batch file of the same shape survives.
Three Directories and Three Applies
The three working directories were per-service deployment entry points. The operator moved into the directory for the service being changed and ran Terraform.
gcp-terraform-minecraft/
├── start-minecraft-server/
│ ├── main.tf
│ ├── variables.tf
│ ├── outputs.tf
│ └── apply.bat
├── restart-preempted-instance/
│ ├── main.tf
│ ├── variabels.tf
│ ├── outputs.tf
│ └── apply.bat
└── bucket-backup-cleanup/
├── main.tf
├── variables.tf
└── outputs.tf
variabels.tf is the actual filename left in the repository at the time. It is a misspelling of the
usual variables.tf, but Terraform read every .tf file in the working directory. The variable
declarations were loaded regardless of the typo in the filename.
Each directory had different resources appearing in its plan, different services changed by its apply, and different resource addresses stored in its state. The three folders were each used as one deployment entry point.
Deploying the Minecraft Server
start-minecraft-server created the server players connected to.
| Terraform resource | What it deployed |
|---|---|
google_compute_address | Static external IP |
google_compute_firewall | Game, RCON, Dynmap, and SSH access rules |
google_compute_instance | Minecraft VM and metadata |
google_storage_bucket | World backup bucket |
The Discord bot ZIP was also included in the VM resource’s metadata. When the batch file changed the ZIP contents, the next apply changed the VM metadata. Infrastructure and the application files inside the server sat in the same deployment unit.
The server configuration had no build number and no artifact repository. The bot source stayed in Git, but the generated ZIP and its SHA-256 were not tracked. Git alone cannot identify which commit produced the ZIP that was actually applied.
Deploying the Preemption Recovery Function
restart-preempted-instance deployed the Google Cloud resources that would detect a shut-down VM and
send a start request to the same instance.
| Terraform resource | What it deployed |
|---|---|
google_logging_metric | Metric aggregating VM stop logs |
google_monitoring_alert_policy | Policy watching the metric condition |
google_pubsub_topic | Topic delivering the alert to the function |
google_cloudfunctions_function | VM state check and start request |
| IAM resources | Permissions the function and Monitoring needed |
The batch file compressed function/main.py and requirements.txt into function.zip. The order
was to move the ZIP to the parent directory and then run Terraform.
cd function
powershell -NoLogo -NoProfile -Command ^
"Compress-Archive -Path 'main.py','requirements.txt' ^
-DestinationPath 'function.zip' -Force"
move function.zip ..
cd ..
terraform apply -auto-approve
After the directories were split, I could deploy only the recovery code without touching the server VM spec or the firewalls. The instance name the function was to start did not arrive automatically from the server configuration.
Deploying the Backup Cleanup Function
bucket-backup-cleanup was a configuration where Scheduler sent a Pub/Sub message and a Cloud
Function looked for old objects.
flowchart LR SCHEDULER["Cloud Scheduler"] --> TOPIC["Pub/Sub"] TOPIC --> FUNCTION["Backup cleanup function"] FUNCTION --> BACKUP["backups/ objects in the backup bucket"]
The cleanup function took the name of the backup bucket the server had created as an input value. The bucket for the function source, Scheduler, Pub/Sub, and IAM were managed by this directory’s Terraform.
The backup bucket on the server side also had a lifecycle rule with the same retention period. This overlap, where the state owning the bucket and the state deleting the objects came apart, is handled separately in part 24.
The Local State Left on the Operator’s PC
There was no backend declaration in any of the three directories. Terraform stored the local state in the directory it ran in.
| Working directory | Remote targets stored in state |
|---|---|
start-minecraft-server | IP, firewalls, VM, and backup bucket |
restart-preempted-instance | Function source, Pub/Sub, Monitoring, function and IAM |
bucket-backup-cleanup | Function source, Pub/Sub, Scheduler, function and IAM |
The three state files that were kept each carried a different lineage value, the field that
distinguishes state lineages. An apply in one directory did not read another directory’s state. To
run from a different computer after losing the state files on the operator’s PC, a procedure to
reconnect the existing remote resources was needed. The repository still holds import_all.bat,
import_vm_restart.bat, and import_existing_resources.sh. The three scripts were written to attach
existing resources to state. Whether they actually ran successfully cannot be confirmed from the
scripts alone.
Project, Instance, and Bucket Inputs
The server configuration exported the instance name and the backup bucket name as outputs.
output "instance_name" {
value = google_compute_instance.minecraft_server.name
}
output "backup_bucket" {
value = google_storage_bucket.backup_bucket.name
}
The recovery configuration took the project, region, zone, and instance name again as its own
variables. Backup cleanup also used the project, region, and target bucket name as separate inputs.
There was no shared module and no terraform_remote_state reference.
Reduced to just the variable declarations, the recovery configuration looked like this.
variable "project_id" {}
variable "region" {}
variable "zone" {}
variable "instance_name" {}
Similar names on an output and an input do not make Terraform connect the two values. The operator had to type the same value in for the recovery function to look at the real server.
An Unsaved Plan and Automatic Approval
Neither apply batch file has a terraform plan -out step. -auto-approve skips the approval
Terraform asks for before applying. The plan still appeared in the terminal, but after the ZIP was
built there was no separate input step for reading the changes and stopping or approving them.
I did not find a record of this approach causing an actual failure. With no saved plan file and no CI run history, which changes were read and approved at the time cannot be reconstructed from Git. Git shows the HCL and the scripts, and it did not keep the variables, the state, or the generated ZIP as they were just before apply.
The Deploy Procedure I Would Change Now
I would leave the per-service working directories as they are and change the order so that each directory builds a plan and approves it.
| Step | Evidence to keep | Stop condition |
|---|---|---|
| Artifact build | Source hash and ZIP SHA-256 | Untracked files or secrets included |
| Static checks | Terraform version, provider lock, and check results | fmt or validate fails |
| Plan | The saved plan and a human-readable summary | Unexpected deletions or replacements |
| Apply | Execution record of the plan that was applied | Inputs or state changed after the plan was built |
| Verify | VM, function, Scheduler, and log status | A service does not reach a ready state |
A saved plan can carry sensitive values in plaintext, so I do not commit it to Git. HashiCorp’s
terraform plan documentation also
says to treat a plan file as a sensitive artifact. The
.terraform.lock.hcl
that records the provider selection goes the other way and is version-controlled in each working
directory.
The local state moves to a GCS backend, which gives multiple runners shared central storage, remote locking, and IAM access control. The three working directories use different prefixes. The project, zone, and instance name the server outputs get passed to the recovery configuration as an explicit input file, and the file holding the real identifiers is excluded from Git.
The three directories lined up the service being deployed with the scope of the plan. Keeping the ZIP hash, the reviewed plan, and remote state with them makes the inputs and approval record available for the next deploy.
The VM read this deployment bundle from metadata and placed its contents into files and systemd units. The metadata and startup-script contract in part 22 compares those keys with the code that consumed them.
References
Comments
No comments yet. Be the first to leave one.
Pending review