Deploying the Server, Recovery, and Backup Separately

How the server, recovery function, and backup cleanup were deployed from three working directories with independent plan and apply cycles

The deploy procedure in start-minecraft-server/apply.bat was short. It built the Discord bot into a ZIP and then applied Terraform straight away.

powershell -NoLogo -NoProfile -Command ^
  "Compress-Archive -Path 'scripts\discord-v2\*' ^
  -DestinationPath 'scripts\discord-v2.zip' -Force"

terraform apply -auto-approve

The batch file in the preemption recovery directory had the same shape. It built the Python function ZIP and ran terraform apply -auto-approve. They sat in one repository, but the server and the recovery function were deployed by different commands. Backup cleanup had its own working directory, but no apply batch file of the same shape survives.

Three Directories and Three Applies

The three working directories were per-service deployment entry points. The operator moved into the directory for the service being changed and ran Terraform.

gcp-terraform-minecraft/
├── start-minecraft-server/
│   ├── main.tf
│   ├── variables.tf
│   ├── outputs.tf
│   └── apply.bat
├── restart-preempted-instance/
│   ├── main.tf
│   ├── variabels.tf
│   ├── outputs.tf
│   └── apply.bat
└── bucket-backup-cleanup/
    ├── main.tf
    ├── variables.tf
    └── outputs.tf

variabels.tf is the actual filename left in the repository at the time. It is a misspelling of the usual variables.tf, but Terraform read every .tf file in the working directory. The variable declarations were loaded regardless of the typo in the filename.

Each directory had different resources appearing in its plan, different services changed by its apply, and different resource addresses stored in its state. The three folders were each used as one deployment entry point.

Deploying the Minecraft Server

start-minecraft-server created the server players connected to.

Terraform resourceWhat it deployed
google_compute_addressStatic external IP
google_compute_firewallGame, RCON, Dynmap, and SSH access rules
google_compute_instanceMinecraft VM and metadata
google_storage_bucketWorld backup bucket

The Discord bot ZIP was also included in the VM resource’s metadata. When the batch file changed the ZIP contents, the next apply changed the VM metadata. Infrastructure and the application files inside the server sat in the same deployment unit.

The server configuration had no build number and no artifact repository. The bot source stayed in Git, but the generated ZIP and its SHA-256 were not tracked. Git alone cannot identify which commit produced the ZIP that was actually applied.

Deploying the Preemption Recovery Function

restart-preempted-instance deployed the Google Cloud resources that would detect a shut-down VM and send a start request to the same instance.

Terraform resourceWhat it deployed
google_logging_metricMetric aggregating VM stop logs
google_monitoring_alert_policyPolicy watching the metric condition
google_pubsub_topicTopic delivering the alert to the function
google_cloudfunctions_functionVM state check and start request
IAM resourcesPermissions the function and Monitoring needed

The batch file compressed function/main.py and requirements.txt into function.zip. The order was to move the ZIP to the parent directory and then run Terraform.

cd function
powershell -NoLogo -NoProfile -Command ^
  "Compress-Archive -Path 'main.py','requirements.txt' ^
  -DestinationPath 'function.zip' -Force"
move function.zip ..
cd ..

terraform apply -auto-approve

After the directories were split, I could deploy only the recovery code without touching the server VM spec or the firewalls. The instance name the function was to start did not arrive automatically from the server configuration.

Deploying the Backup Cleanup Function

bucket-backup-cleanup was a configuration where Scheduler sent a Pub/Sub message and a Cloud Function looked for old objects.

flowchart LR
  SCHEDULER["Cloud Scheduler"] --> TOPIC["Pub/Sub"]
  TOPIC --> FUNCTION["Backup cleanup function"]
  FUNCTION --> BACKUP["backups/ objects in the backup bucket"]

The cleanup function took the name of the backup bucket the server had created as an input value. The bucket for the function source, Scheduler, Pub/Sub, and IAM were managed by this directory’s Terraform.

The backup bucket on the server side also had a lifecycle rule with the same retention period. This overlap, where the state owning the bucket and the state deleting the objects came apart, is handled separately in part 24.

The Local State Left on the Operator’s PC

There was no backend declaration in any of the three directories. Terraform stored the local state in the directory it ran in.

Working directoryRemote targets stored in state
start-minecraft-serverIP, firewalls, VM, and backup bucket
restart-preempted-instanceFunction source, Pub/Sub, Monitoring, function and IAM
bucket-backup-cleanupFunction source, Pub/Sub, Scheduler, function and IAM

The three state files that were kept each carried a different lineage value, the field that distinguishes state lineages. An apply in one directory did not read another directory’s state. To run from a different computer after losing the state files on the operator’s PC, a procedure to reconnect the existing remote resources was needed. The repository still holds import_all.bat, import_vm_restart.bat, and import_existing_resources.sh. The three scripts were written to attach existing resources to state. Whether they actually ran successfully cannot be confirmed from the scripts alone.

Project, Instance, and Bucket Inputs

The server configuration exported the instance name and the backup bucket name as outputs.

output "instance_name" {
  value = google_compute_instance.minecraft_server.name
}

output "backup_bucket" {
  value = google_storage_bucket.backup_bucket.name
}

The recovery configuration took the project, region, zone, and instance name again as its own variables. Backup cleanup also used the project, region, and target bucket name as separate inputs. There was no shared module and no terraform_remote_state reference.

Reduced to just the variable declarations, the recovery configuration looked like this.

variable "project_id" {}
variable "region" {}
variable "zone" {}
variable "instance_name" {}

Similar names on an output and an input do not make Terraform connect the two values. The operator had to type the same value in for the recovery function to look at the real server.

An Unsaved Plan and Automatic Approval

Neither apply batch file has a terraform plan -out step. -auto-approve skips the approval Terraform asks for before applying. The plan still appeared in the terminal, but after the ZIP was built there was no separate input step for reading the changes and stopping or approving them.

I did not find a record of this approach causing an actual failure. With no saved plan file and no CI run history, which changes were read and approved at the time cannot be reconstructed from Git. Git shows the HCL and the scripts, and it did not keep the variables, the state, or the generated ZIP as they were just before apply.

The Deploy Procedure I Would Change Now

I would leave the per-service working directories as they are and change the order so that each directory builds a plan and approves it.

StepEvidence to keepStop condition
Artifact buildSource hash and ZIP SHA-256Untracked files or secrets included
Static checksTerraform version, provider lock, and check resultsfmt or validate fails
PlanThe saved plan and a human-readable summaryUnexpected deletions or replacements
ApplyExecution record of the plan that was appliedInputs or state changed after the plan was built
VerifyVM, function, Scheduler, and log statusA service does not reach a ready state

A saved plan can carry sensitive values in plaintext, so I do not commit it to Git. HashiCorp’s terraform plan documentation also says to treat a plan file as a sensitive artifact. The .terraform.lock.hcl that records the provider selection goes the other way and is version-controlled in each working directory.

The local state moves to a GCS backend, which gives multiple runners shared central storage, remote locking, and IAM access control. The three working directories use different prefixes. The project, zone, and instance name the server outputs get passed to the recovery configuration as an explicit input file, and the file holding the real identifiers is excluded from Git.

The three directories lined up the service being deployed with the scope of the plan. Keeping the ZIP hash, the reviewed plan, and remote state with them makes the inputs and approval record available for the next deploy.

The VM read this deployment bundle from metadata and placed its contents into files and systemd units. The metadata and startup-script contract in part 22 compares those keys with the code that consumed them.

References

Comments

Comments

    Image preview