Restarting the Minecraft Server After a Spot Preemption
A separate Terraform root that matches only Spot preemptions and starts the same VM again
The server’s compute.tf sets the Spot termination action to STOP, so a preempted VM remains
TERMINATED. A Scheduler that sends a start command on a fixed schedule also turns on a server an
operator stopped for cost or maintenance. The configuration below selects only the
compute.instances.preempted events Google Cloud recorded and starts the same VM.
Prerequisites
The manual start and failure inspection in part 08 must work first. Do not add a recovery function to a server a person cannot start when the automation fails.
Verification scope
Terraform static checks and Python branch tests ran locally. No Cloud Run functions, Pub/Sub topic, or Logging sink was created, and no real preemption was triggered. The status of this configuration is
verified locally and in a container, not verified against a real Google Cloud apply or a real preemption.
Why the Server and Recovery Root Are Separate
The recovery code has to run after the VM is off. Instead of putting it in the server root’s startup script, I run it in a separately deployed Cloud Run function.
The Logging sink exports only the logs that match the condition. The Pub/Sub topic holds that event as a message and delivers it to Cloud Run functions. The function runs briefly, only when it receives a message.
sequenceDiagram
accTitle: Automatic VM recovery after a Spot preemption
accDescr: A preemption system event passes through the Logging sink and Pub/Sub to the recovery function, and the function re-checks the target and the state before starting the same VM.
participant VM as Spot VM
participant LOG as Cloud Logging
participant TOPIC as Pub/Sub
participant FN as Recovery function
participant API as Compute API
VM->>LOG: compute.instances.preempted
LOG->>TOPIC: LogEntry matching the filter
TOPIC->>FN: CloudEvent
FN->>FN: Re-check project, zone, VM, method
FN->>API: instances.get
alt VM is TERMINATED
FN->>API: instances.start
FN->>API: Check operation result
else Starting or already running
FN-->>FN: Record no_action
end
The server root’s state contains the VM, disk, network, and static IP. The recovery root’s state contains only the Logging sink, Pub/Sub, function, and recovery service account. With separate backend prefixes, each destroy plan lists only the resources from its own root.
Preparing the Example and Input Values
Download the example into a directory beside the server root.
Run
cd "${HOME}"
curl --fail --location \
"https://juntiger-assets.pages.dev/minecraft-spot-recovery.zip" \
--output minecraft-spot-recovery.zip
unzip -q minecraft-spot-recovery.zip
cd minecraft-spot-recovery
cp backend.hcl.example backend.hcl
cp terraform.tfvars.example terraform.tfvars
nano backend.hcl
Use the state bucket created in part 03, but write a different prefix.
bucket = "my-state-bucket-name"
prefix = "minecraft/spot-recovery"
Put the server root outputs and the function source bucket name into terraform.tfvars.
Before copying them over, put the server root’s real output values on screen.
Check: server root outputs
(cd "${HOME}/minecraft-one-root" && terraform output)
Run
nano terraform.tfvars
project_id = "my-project-id"
region = "asia-northeast3"
zone = "asia-northeast3-a"
instance_name = "minecraft"
function_source_bucket_name = "globally-unique-function-source-bucket"
project_id, zone, and instance_name must match the server root outputs exactly.
This root has no preflight check like the server root’s scripts/preflight.sh. Variable validation
rejects project and bucket placeholders and invalid instance-name formats, but it cannot confirm a
match with the server root outputs. A typo can still pass plan, so compare the values against the
output above character by character.
Give the function source bucket a name not used by the backup bucket or the state bucket.
Unless a different path is given, run the commands that follow in Cloud Shell under
${HOME}/minecraft-spot-recovery.
Local Branch Test of the Function
Before waiting for a real preemption, check the input classification and the API call conditions with a local Python test.
Run
python3 -m venv .venv
. .venv/bin/activate
pip install -r function/requirements.txt
PYTHONPATH=function python -m unittest -v function/test_main.py
If the first command fails with ensurepip is not available, the runtime has no venv module. On
Ubuntu, install it with sudo apt-get install python3-venv, then run the commands again.
The tests cover the following inputs.
- A valid preemption event and a
TERMINATEDVM - A manual stop and other methods
- A different project, zone, or VM
- Missing Pub/Sub data and malformed JSON
PROVISIONING,STAGING, andRUNNINGstates on duplicate delivery- The
STOPPINGtoTERMINATEDtransition and a wait timeout - An event older than 15 minutes
- Compute API errors and an operation timeout
The last line of a passing run has this form.
Ran ... tests in ...
OK
If even one test says FAILED, do not move on to the Terraform apply. Only the Python
package install changes the local environment; it creates no Google Cloud resources.
Double-Check the Log and VM State
The Logging sink sends to Pub/Sub only the system events that satisfy all of the following conditions.
log_id("cloudaudit.googleapis.com/system_event")
protoPayload.serviceName="compute.googleapis.com"
protoPayload.methodName="compute.instances.preempted"
protoPayload.resourceName="projects/PROJECT/zones/ZONE/instances/INSTANCE"
An operator’s v1.compute.instances.stop, a Terraform delete, and a replacement do not
match this filter. The function also compares the method and the full resourceName of
the delivered LogEntry against the allowed target in its environment variables.
Even when the input matches, ignore the event if the log timestamp is more than 15
minutes older than now. If the Compute API reports STOPPING, query again until the time
limit. Call instances.start only after the state becomes TERMINATED, and check until
the operation finishes. When a duplicate Pub/Sub delivery arrives and the VM is already
PROVISIONING, STAGING, or RUNNING, the run ends as no_action.
Permissions of the Recovery Account
Put only three permissions in the custom role of the function runtime service account.
compute.instances.get
compute.instances.start
compute.zoneOperations.get
This account cannot delete the VM, change the machine type, or work on disks. Deploying the function also needs the Cloud Functions, Cloud Build, Artifact Registry, Eventarc, Pub/Sub, Logging, and Storage APIs. The Terraform execution account needs permission to create these resources and to use the build and runtime service accounts.
Read one exception as written. To build the function source into a container, the
example’s iam.tf grants project-level roles/cloudbuild.builds.builder to the Compute
Engine default service account. This role is broader in scope than the runtime custom
role and is needed only during the build step, so on a shared project, discuss the grant
with the administrator first.
The Eventarc trigger uses a separate service account. Attach only
roles/eventarc.eventReceiver and roles/run.invoker to it, and name it in the
function’s event_trigger.service_account_email. Only the runtime account holds
permission to start the VM; the trigger account only receives events and invokes the
function.
On a shared project, hand the following scope of work to the administrator.
| Task | Representative role |
|---|---|
| Enabling APIs | roles/serviceusage.serviceUsageAdmin |
| Creating the Logging sink | roles/logging.configWriter |
| Pub/Sub topic and IAM | roles/pubsub.admin |
| Creating the function | roles/cloudfunctions.admin |
| Function source bucket | roles/storage.admin |
| Service accounts and custom role | roles/iam.serviceAccountAdmin, roles/iam.roleAdmin |
| Account attachment and project IAM | roles/iam.serviceAccountUser, permission to change project IAM |
The First Plan for the Recovery Root
Run: no recovery resource is created yet
terraform init -backend-config=backend.hcl
terraform fmt
terraform fmt -check -recursive
terraform validate
terraform plan -out=recovery.tfplan
terraform show recovery.tfplan
In the plan, check the project, zone, and instance in the sink filter, the three
permissions of the custom role, the separate function source bucket, and the runtime and
Eventarc trigger accounts. The exclude list of the function ZIP must contain
__pycache__ and test_main.py. Stop if a VM or network owned by the server root
appears in the plan.
Applying the Recovery Resources
The command below enables APIs and creates resources for Pub/Sub, Storage, Logging, and Cloud Run functions. Small usage cannot be assumed to be free. Do not run it if you have not checked the budget alert and the deletion path.
Costs money: once the automatic recovery cost and permissions are approved
terraform apply recovery.tfplan
After Apply complete!, check the sink, the function, and the recovery target VM in the
outputs. A failure partway through can leave some resources behind, so read the current
state again with a new plan.
Run: save the target values in the recovery root
export TARGET_PROJECT_ID="$(
terraform output -raw target_project_id
)"
export TARGET_ZONE="$(
terraform output -raw target_zone
)"
export TARGET_INSTANCE="$(
terraform output -raw target_instance_name
)"
printf 'Recovery target: %s / %s / %s\n' \
"${TARGET_PROJECT_ID}" \
"${TARGET_ZONE}" \
"${TARGET_INSTANCE}"
If any of the three values differs from the server root outputs, do not run the manual stop test.
Manual Stop Test
Only readers who applied the real recovery root run this. A VM an operator stopped must not be recovered automatically.
Conditional run: when the recovery root is applied
gcloud compute instances stop "${TARGET_INSTANCE}" \
--project="${TARGET_PROJECT_ID}" \
--zone="${TARGET_ZONE}"
It must still be TERMINATED several minutes later. If it turned on by itself, check the
Logging sink filter and any other automation before the function logs, and disable the
recovery root.
Start it manually after the test.
gcloud compute instances start "${TARGET_INSTANCE}" \
--project="${TARGET_PROJECT_ID}" \
--zone="${TARGET_ZONE}"
Observing a Real Preemption
An operator cannot decide when Google Cloud preempts a Spot VM. Even where maintenance simulation is available for the product state and the VM settings, do not force an event on a server that holds a real world.
The results in the function logs are distinguished as follows.
| Result | Meaning and next action |
|---|---|
ignored_event | Not started, because the message was wrong, the method or the target differed, it was older than 15 minutes, or a permanent error such as missing permission occurred (permanent_compute_error) |
no_action | The VM is already starting or running |
start_requested | A start was requested for a TERMINATED VM after a verified preemption |
| Only an error log, with no result | A transient error such as an operation timeout or a Spot capacity shortage propagates as an exception, and the Eventarc retry policy (RETRY_POLICY_RETRY) delivers the same event again |
Even when the function requests a start, it can fail if that zone has no Spot capacity. The recovery function does not create a new VM in another zone or switch to a regular VM. That needs a separate recovery design that also moves ownership of the static IP and the world disk.
If you observed a real preemption and Minecraft’s new Done (...)! after
start_requested, record the real-environment test date in the operations log. Otherwise
do not overstate the completion state; keep it as verified locally and in a container, not verified against a real Google Cloud apply or a real preemption.
Automatic Recovery Failure and Manual Start
Check
gcloud compute instances describe "${TARGET_INSTANCE}" \
--project="${TARGET_PROJECT_ID}" \
--zone="${TARGET_ZONE}" \
--format="value(status)"
Conditional run: the status is TERMINATED and a person is recovering it
gcloud compute instances start "${TARGET_INSTANCE}" \
--project="${TARGET_PROJECT_ID}" \
--zone="${TARGET_ZONE}"
Once the VM is RUNNING, check Minecraft readiness in the startup script → systemd → Done (...)!
order from part 05. The design principles are covered in more detail in
Automatic Recovery for a Preemptible VM on GCP and
VM Restart and Minecraft Readiness.
Completion Check
- Every Python unit test passed.
- The server root and the recovery root use different backend prefixes.
- A manual stop does not start automatically.
- If you have not seen a real preemption, you recorded the unverified status.
- The manual start procedure still works.
On a server with automatic recovery applied, Procedure for Deleting the Minecraft Server and Its Backups removes the recovery root before the server root.
References
Comments
No comments yet. Be the first to leave one.
Pending review