Restarting the Minecraft Server After a Spot Preemption

A separate Terraform root that matches only Spot preemptions and starts the same VM again

The server’s compute.tf sets the Spot termination action to STOP, so a preempted VM remains TERMINATED. A Scheduler that sends a start command on a fixed schedule also turns on a server an operator stopped for cost or maintenance. The configuration below selects only the compute.instances.preempted events Google Cloud recorded and starts the same VM.

Prerequisites

The manual start and failure inspection in part 08 must work first. Do not add a recovery function to a server a person cannot start when the automation fails.

Verification scope

Terraform static checks and Python branch tests ran locally. No Cloud Run functions, Pub/Sub topic, or Logging sink was created, and no real preemption was triggered. The status of this configuration is verified locally and in a container, not verified against a real Google Cloud apply or a real preemption.

Why the Server and Recovery Root Are Separate

The recovery code has to run after the VM is off. Instead of putting it in the server root’s startup script, I run it in a separately deployed Cloud Run function.

The Logging sink exports only the logs that match the condition. The Pub/Sub topic holds that event as a message and delivers it to Cloud Run functions. The function runs briefly, only when it receives a message.

sequenceDiagram
  accTitle: Automatic VM recovery after a Spot preemption
  accDescr: A preemption system event passes through the Logging sink and Pub/Sub to the recovery function, and the function re-checks the target and the state before starting the same VM.
  participant VM as Spot VM
  participant LOG as Cloud Logging
  participant TOPIC as Pub/Sub
  participant FN as Recovery function
  participant API as Compute API
  VM->>LOG: compute.instances.preempted
  LOG->>TOPIC: LogEntry matching the filter
  TOPIC->>FN: CloudEvent
  FN->>FN: Re-check project, zone, VM, method
  FN->>API: instances.get
  alt VM is TERMINATED
    FN->>API: instances.start
    FN->>API: Check operation result
  else Starting or already running
    FN-->>FN: Record no_action
  end

The server root’s state contains the VM, disk, network, and static IP. The recovery root’s state contains only the Logging sink, Pub/Sub, function, and recovery service account. With separate backend prefixes, each destroy plan lists only the resources from its own root.

Preparing the Example and Input Values

Download the example into a directory beside the server root.

Run

cd "${HOME}"
curl --fail --location \
  "https://juntiger-assets.pages.dev/minecraft-spot-recovery.zip" \
  --output minecraft-spot-recovery.zip
unzip -q minecraft-spot-recovery.zip
cd minecraft-spot-recovery

cp backend.hcl.example backend.hcl
cp terraform.tfvars.example terraform.tfvars
nano backend.hcl

Use the state bucket created in part 03, but write a different prefix.

bucket = "my-state-bucket-name"
prefix = "minecraft/spot-recovery"

Put the server root outputs and the function source bucket name into terraform.tfvars. Before copying them over, put the server root’s real output values on screen.

Check: server root outputs

(cd "${HOME}/minecraft-one-root" && terraform output)

Run

nano terraform.tfvars
project_id                  = "my-project-id"
region                      = "asia-northeast3"
zone                        = "asia-northeast3-a"
instance_name               = "minecraft"
function_source_bucket_name = "globally-unique-function-source-bucket"

project_id, zone, and instance_name must match the server root outputs exactly. This root has no preflight check like the server root’s scripts/preflight.sh. Variable validation rejects project and bucket placeholders and invalid instance-name formats, but it cannot confirm a match with the server root outputs. A typo can still pass plan, so compare the values against the output above character by character. Give the function source bucket a name not used by the backup bucket or the state bucket. Unless a different path is given, run the commands that follow in Cloud Shell under ${HOME}/minecraft-spot-recovery.

Local Branch Test of the Function

Before waiting for a real preemption, check the input classification and the API call conditions with a local Python test.

Run

python3 -m venv .venv
. .venv/bin/activate
pip install -r function/requirements.txt
PYTHONPATH=function python -m unittest -v function/test_main.py

If the first command fails with ensurepip is not available, the runtime has no venv module. On Ubuntu, install it with sudo apt-get install python3-venv, then run the commands again.

The tests cover the following inputs.

  • A valid preemption event and a TERMINATED VM
  • A manual stop and other methods
  • A different project, zone, or VM
  • Missing Pub/Sub data and malformed JSON
  • PROVISIONING, STAGING, and RUNNING states on duplicate delivery
  • The STOPPING to TERMINATED transition and a wait timeout
  • An event older than 15 minutes
  • Compute API errors and an operation timeout

The last line of a passing run has this form.

Ran ... tests in ...

OK

If even one test says FAILED, do not move on to the Terraform apply. Only the Python package install changes the local environment; it creates no Google Cloud resources.

Double-Check the Log and VM State

The Logging sink sends to Pub/Sub only the system events that satisfy all of the following conditions.

log_id("cloudaudit.googleapis.com/system_event")
protoPayload.serviceName="compute.googleapis.com"
protoPayload.methodName="compute.instances.preempted"
protoPayload.resourceName="projects/PROJECT/zones/ZONE/instances/INSTANCE"

An operator’s v1.compute.instances.stop, a Terraform delete, and a replacement do not match this filter. The function also compares the method and the full resourceName of the delivered LogEntry against the allowed target in its environment variables.

Even when the input matches, ignore the event if the log timestamp is more than 15 minutes older than now. If the Compute API reports STOPPING, query again until the time limit. Call instances.start only after the state becomes TERMINATED, and check until the operation finishes. When a duplicate Pub/Sub delivery arrives and the VM is already PROVISIONING, STAGING, or RUNNING, the run ends as no_action.

Permissions of the Recovery Account

Put only three permissions in the custom role of the function runtime service account.

compute.instances.get
compute.instances.start
compute.zoneOperations.get

This account cannot delete the VM, change the machine type, or work on disks. Deploying the function also needs the Cloud Functions, Cloud Build, Artifact Registry, Eventarc, Pub/Sub, Logging, and Storage APIs. The Terraform execution account needs permission to create these resources and to use the build and runtime service accounts.

Read one exception as written. To build the function source into a container, the example’s iam.tf grants project-level roles/cloudbuild.builds.builder to the Compute Engine default service account. This role is broader in scope than the runtime custom role and is needed only during the build step, so on a shared project, discuss the grant with the administrator first.

The Eventarc trigger uses a separate service account. Attach only roles/eventarc.eventReceiver and roles/run.invoker to it, and name it in the function’s event_trigger.service_account_email. Only the runtime account holds permission to start the VM; the trigger account only receives events and invokes the function.

On a shared project, hand the following scope of work to the administrator.

TaskRepresentative role
Enabling APIsroles/serviceusage.serviceUsageAdmin
Creating the Logging sinkroles/logging.configWriter
Pub/Sub topic and IAMroles/pubsub.admin
Creating the functionroles/cloudfunctions.admin
Function source bucketroles/storage.admin
Service accounts and custom roleroles/iam.serviceAccountAdmin, roles/iam.roleAdmin
Account attachment and project IAMroles/iam.serviceAccountUser, permission to change project IAM

The First Plan for the Recovery Root

Run: no recovery resource is created yet

terraform init -backend-config=backend.hcl
terraform fmt
terraform fmt -check -recursive
terraform validate
terraform plan -out=recovery.tfplan
terraform show recovery.tfplan

In the plan, check the project, zone, and instance in the sink filter, the three permissions of the custom role, the separate function source bucket, and the runtime and Eventarc trigger accounts. The exclude list of the function ZIP must contain __pycache__ and test_main.py. Stop if a VM or network owned by the server root appears in the plan.

Applying the Recovery Resources

The command below enables APIs and creates resources for Pub/Sub, Storage, Logging, and Cloud Run functions. Small usage cannot be assumed to be free. Do not run it if you have not checked the budget alert and the deletion path.

Costs money: once the automatic recovery cost and permissions are approved

terraform apply recovery.tfplan

After Apply complete!, check the sink, the function, and the recovery target VM in the outputs. A failure partway through can leave some resources behind, so read the current state again with a new plan.

Run: save the target values in the recovery root

export TARGET_PROJECT_ID="$(
  terraform output -raw target_project_id
)"
export TARGET_ZONE="$(
  terraform output -raw target_zone
)"
export TARGET_INSTANCE="$(
  terraform output -raw target_instance_name
)"
printf 'Recovery target: %s / %s / %s\n' \
  "${TARGET_PROJECT_ID}" \
  "${TARGET_ZONE}" \
  "${TARGET_INSTANCE}"

If any of the three values differs from the server root outputs, do not run the manual stop test.

Manual Stop Test

Only readers who applied the real recovery root run this. A VM an operator stopped must not be recovered automatically.

Conditional run: when the recovery root is applied

gcloud compute instances stop "${TARGET_INSTANCE}" \
  --project="${TARGET_PROJECT_ID}" \
  --zone="${TARGET_ZONE}"

It must still be TERMINATED several minutes later. If it turned on by itself, check the Logging sink filter and any other automation before the function logs, and disable the recovery root.

Start it manually after the test.

gcloud compute instances start "${TARGET_INSTANCE}" \
  --project="${TARGET_PROJECT_ID}" \
  --zone="${TARGET_ZONE}"

Observing a Real Preemption

An operator cannot decide when Google Cloud preempts a Spot VM. Even where maintenance simulation is available for the product state and the VM settings, do not force an event on a server that holds a real world.

The results in the function logs are distinguished as follows.

ResultMeaning and next action
ignored_eventNot started, because the message was wrong, the method or the target differed, it was older than 15 minutes, or a permanent error such as missing permission occurred (permanent_compute_error)
no_actionThe VM is already starting or running
start_requestedA start was requested for a TERMINATED VM after a verified preemption
Only an error log, with no resultA transient error such as an operation timeout or a Spot capacity shortage propagates as an exception, and the Eventarc retry policy (RETRY_POLICY_RETRY) delivers the same event again

Even when the function requests a start, it can fail if that zone has no Spot capacity. The recovery function does not create a new VM in another zone or switch to a regular VM. That needs a separate recovery design that also moves ownership of the static IP and the world disk.

If you observed a real preemption and Minecraft’s new Done (...)! after start_requested, record the real-environment test date in the operations log. Otherwise do not overstate the completion state; keep it as verified locally and in a container, not verified against a real Google Cloud apply or a real preemption.

Automatic Recovery Failure and Manual Start

Check

gcloud compute instances describe "${TARGET_INSTANCE}" \
  --project="${TARGET_PROJECT_ID}" \
  --zone="${TARGET_ZONE}" \
  --format="value(status)"

Conditional run: the status is TERMINATED and a person is recovering it

gcloud compute instances start "${TARGET_INSTANCE}" \
  --project="${TARGET_PROJECT_ID}" \
  --zone="${TARGET_ZONE}"

Once the VM is RUNNING, check Minecraft readiness in the startup script → systemd → Done (...)! order from part 05. The design principles are covered in more detail in Automatic Recovery for a Preemptible VM on GCP and VM Restart and Minecraft Readiness.

Completion Check

  • Every Python unit test passed.
  • The server root and the recovery root use different backend prefixes.
  • A manual stop does not start automatically.
  • If you have not seen a real preemption, you recorded the unverified status.
  • The manual start procedure still works.

On a server with automatic recovery applied, Procedure for Deleting the Minecraft Server and Its Backups removes the recovery root before the server root.

References

Comments

Comments

    Image preview