Migrating the 2025 Terraform Configuration to the Current Baseline

Separating the product settings I would choose now from the original design defects, and the criteria for picking a migration path you can reverse

The 2025 configuration had separate defects in its metadata contract, import checks, resource ownership, and secret handling. Planning the migration requires separating the product settings I would choose now from design changes that were possible even then.

The product information and prices below were confirmed on August 28, 2026. Rates change often, so check them again with the calculator before any real decision.

Product Settings I Would Choose Now

Choose a Spot VM instead of a legacy preemptible VM. The 2025 configuration set only preemptible = true and automatic_restart = false. Those options were supported at the time, but they select a legacy preemptible VM with a 24-hour maximum runtime. The current configuration uses provisioning_model = "SPOT" and instance_termination_action. With the termination action set to STOP, the instance remains in the TERMINATED state and its attached persistent disks remain available. A Spot VM has no minimum or maximum runtime unless its runtime is limited separately.

Use the Cloud Run function deployment path for new functions. The 2025 configuration sat on the first-generation google_cloudfunctions_function. The second generation is now called Cloud Run functions. The compatible google_cloudfunctions2_function resource remains, but Google Cloud documentation now guides new Terraform deployments to build a function container and deploy it with google_cloud_run_v2_service.

The provider major version moved up. The server root was pinned at ~> 5.0, while the other two roots used >= 5.0.0 and so have 6.x in their locks. The latest today is 7.x. Each root starts from a different point, and because a major upgrade changes attribute names and defaults, you do not skip across them in one jump.

Terraform was dropped from Cloud Shell. As of the June 2026 release notes, the CLI is not in the default image. That adds one more step for preparing the execution environment.

Network egress has to be calculated separately. In the August 2026 public price table, the first 1 TiB tier of internet traffic from Seoul to users in Korea costs USD 0.19 per GiB. The rate varies by destination and usage tier, so recalculate it for the actual traffic conditions. A VM-only estimate omits this charge.

Design Defects That Did Not Depend on Product Changes

The following items are independent of product changes and could have been configured differently at the time.

Remote backends existed in 2025 too. Local state in the working directory had no remote locking or access control shared by multiple runners, and backup copies were committed to Git. The apply -auto-approve batch files had no step for saving and reading a plan, and how far the import had got was something I learned only after apply stopped.

I split the directories but never decided which state was responsible for what, so APIs overlapped in two places. Putting secrets in variable defaults was avoidable regardless of the year. The metadata key nothing read and the restore branch that never ran would both have surfaced from watching a single boot through from the start.

Two Migration Paths

There are broadly two ways to migrate while keeping the existing server alive. They divide on whether the connection address has to stay the same and whether you can move the state directly.

Path A: Migrate in Place

Keep the existing resources and state, and raise only the configuration to the current baseline.

  1. Prepare to roll back. Back up the world and finish a restore test. Skip this step and a failure at any later step leaves the world unrecoverable.
  2. Move state to a remote backend. Add the backend and run terraform init -migrate-state. Put .terraform.lock.hcl under version control.
  3. Invalidate the secrets first. Retiring and reissuing come before cleaning up code. Then remove the variable defaults and the fallback constants in the scripts.
  4. Raise the provider one major at a time. Save and read the plan at each step, and stop if there is a replacement you did not expect.
  5. Sort out ownership. Gather the overlapping APIs into one root. Use terraform state rm and import, but one at a time, checking with plan every time.
  6. Switch the preemptible settings to the current expression. Always check in the plan whether this change is an attribute update or an instance replacement. If it is a replacement, decide how to protect the boot disk and the world before going ahead.
  7. Move the functions to Cloud Run functions. Verify the new service and trigger before deleting the old first-generation function.

The rollback point differs by step. To reverse step 2, keep a verified state copy and migrate back to the local backend with terraform init -migrate-state; do not copy state files manually. Steps 5 and 6 are hard to reverse. Touch the state wrongly and you lose the link between the code and the real resource; replace the instance and the disk can disappear.

Path B: Build New and Move Only the Data

Build the server in a new project or a new namespace with a current-baseline configuration, and bring over only the world. Stop new writes to the old world before taking the final backup.

  1. Build the new server by the build track procedure.
  2. Disconnect the players, stop the old server cleanly, and take and verify the final backup.
  3. Import the world into the new server.
  4. Keep the old server stopped, confirm the new connection, and tell the players the new address.
  5. Check the new backups and play records for a few days, then tear down the old server.

Before anyone plays on the new server, rollback means stopping it and turning the old server back on. After the new world has changed, do not simply turn the old server on. Stop and back up the new server first, then decide whether to move those changes to the old world or discard them and return to the final migration point. As long as both worlds are never written at the same time, this path does not require changing the old state or walking the provider upgrade path.

In exchange, the connection address changes, and costs can overlap for the new VM and both environments’ disks, static IPs, and storage during the move.

Criteria for Choosing a Migration Path

For a single server run by one person, I recommend Path B. What the hard steps in Path A protect is a static IP and resource names. The risk of mishandling state outweighs the cost of telling a few friends a new address. Before play starts on the new server, Path B also preserves a rollback point where the old server can be turned back on. After play starts, you must explicitly decide whether to move or discard the changes made on the new world.

There are cases where Path A is right: the address is registered in several places and hard to change, the data is large enough that the transfer takes long, or organization policy does not allow creating a new project. In that case, follow the order above, and always take a fresh backup before steps 5 and 6.

The Original Then and the Example Now

Item2025 configurationBuild track example
stateLocal file in the working directoryGCS backend, versioning
Deployapply -auto-approveSaved plan reviewed, then applied
Preemptible expressionLegacy preemptibleprovisioning_model = "SPOT", STOP on termination
Administrative accessroot SSH and a key file in the working treeOS Login and IAM
Server file deliveryScripts, units, and the app ZIP in metadataJAR verified by URL and SHA-256
SecretsVariable defaultsNo secrets kept on the server
Backup retentionBoth bucket lifecycle and a cleanup functionBucket lifecycle alone
Recovery function1st-gen function, via a Monitoring alertCloud Run functions, via a log sink
Function permissionsProject roleLeast-privilege custom role

The left column records the configuration confirmed in the archived repository and state. The right column shows the public example that reduces those defects.

Why the Current Example Adds Checks

I never matched the two sides of the contract, never read how far the import got, never wrote down ownership, and never counted how far the secrets had traveled. All four happened because the procedure had no step that checks.

The build track example carries check scripts and container tests so that mismatches between the implementation and its expectations can be checked repeatedly. While preparing this series and rereading the code, I found a defect in the example’s backup verification script. The container test was passing without catching it, because the fake Java never checked the working directory. Only after I made the test closer to the real thing did the same defect reproduce in the test.

After this experience, I include verification with the implementation and review whether that verification reproduces the real failure conditions.

References

Comments

Comments

    Image preview