Splitting One Terraform Root into Three

How a single main.tf that accumulated server, recovery, and backup settings became three working directories with separate deployment units

The earliest terraform/main.tf still in the repository had a Compute Engine VM and a firewall. It was written into Terraform only as far as bringing up a VM and opening the Minecraft port.

While I ran the server, the work grew. I needed a static IP, and the world had to be backed up outside the VM. The function that restarted a preempted VM and the Discord bot deployment came into the same file. A terraform plan that once created one server eventually included backup and recovery automation as well.

flowchart LR
  FIRST["First repository<br/>VM and firewall"] --> GROWN["Grown main.tf<br/>server, backup and recovery automation"]
  GROWN --> NEW["New repository<br/>server configuration rewritten"]
  NEW --> SPLIT["Three working directories<br/>server, preemption recovery, backup cleanup"]

The main.tf That Started With a VM and a Firewall

The first HCL shows two kinds of managed resources.

terraform/main.tf
├── google_compute_firewall
└── google_compute_instance

The firewall opened the Minecraft port, and the VM started on Ubuntu 22.04 with a 30GB standard persistent disk. The small machine type later changed to a larger one, and the legacy preemptible setting went into the VM.

At that scope one file was easy to read. The VM name, the network tag, and the firewall could be reviewed in the same plan. With only resources that changed together with the server, there was little reason to split the working directory further.

Operational Resources Added Later

Every time an operational need appeared outside the Minecraft process, a Google Cloud resource and a deployment file came with it.

What operations neededWhat entered TerraformWhat the plan then showed alongside it
Keeping the connection addressStatic external IPThe VM network interface and its IP binding
Separating the admin pathGame, RCON, SSH, and Dynmap firewallsPorts and allowed CIDRs
Keeping the worldBackup bucket and IAMBucket retention settings and VM permissions
Restarting after preemptionScheduler, Pub/Sub, Cloud Functions, a Cloud Run experimentThe function, the message path, and IAM
Running DiscordBot ZIP and VM metadataInfrastructure changes and application file deployment

Adding preemption recovery brought in resources outside the VM. A VM that has been shut down cannot start itself. I needed a Google Cloud service outside the VM to read its state and send the start request. HCL from several attempts stayed in the one file, and some resources ended up declared without a finished call path. That is why the final HCL as a whole cannot be read as the operating configuration at any single point in time.

The role of metadata widened too. The startup script, the shutdown script, the systemd unit, the Minecraft settings, and the Discord bot ZIP all collected inside the VM resource. terraform apply created the VM and deployed the files to be run inside the server.

Rewriting the Server Configuration in a New Repository

The current repository did not inherit the Git history of the earlier one. There is no common commit, and the first file is not a straight copy of the old HCL. The first main.tf in the new repository shows these five resource declarations.

resource "google_compute_address" "minecraft_static_ip" {}
resource "google_storage_bucket" "backup_bucket" {}
resource "google_compute_network" "minecraft_network" {}
resource "google_compute_firewall" "minecraft_firewall" {}
resource "google_compute_instance" "minecraft_server" {}

The code above is an excerpt that keeps only the resource declarations. The actual attributes and identifiers are not published.

Using the repository that had managed everything with one Terraform as a reference, I rewrote the resource names and the file layout. The first version still started with main.tf, variables.tf, and outputs.tf at the repository root. Creating a new repository did not immediately produce three working directories.

No document survives that records the command that created the new repository or the reasoning at the time. What I can confirm reaches as far as the Git history and file contents of the two repositories, plus my present memory that I meant to configure the server, recovery, and backup cleanup separately.

Plan and Apply Split Three Ways

The server files moved to server-terraform first. Git records that change as a file move with no content edit. The same change added preempted-terraform for preemption recovery, and the backup cleanup directory came in afterward. One more round of renaming produced the three directories that exist now.

Working directoryWhat the plan reviewedWhat apply deployed
start-minecraft-serverStatic IP, firewalls, VM, backup bucketMinecraft server
restart-preempted-instanceLog metric, Monitoring, Pub/Sub, function and IAMPreemption recovery
bucket-backup-cleanupScheduler, Pub/Sub, function and IAMOld backup cleanup
flowchart TB
  subgraph BEFORE["One working directory"]
    ONE_PLAN["One plan"] --> ONE_APPLY["Apply the server and the external automation"]
  end

  subgraph AFTER["Three working directories"]
    SERVER_PLAN["Server plan"] --> SERVER_APPLY["Server apply"]
    RESTART_PLAN["Preemption recovery plan"] --> RESTART_APPLY["Recovery function apply"]
    CLEANUP_PLAN["Backup cleanup plan"] --> CLEANUP_APPLY["Cleanup function apply"]
  end

I split them to line up the service being deployed with the scope of Terraform I had to review. Changing the VM spec did not require reading the recovery function and the backup cleanup settings. When I edited only the recovery function, I did not want to check firewall and VM changes again.

I have no record and no memory of state conflicts or failure isolation being a direct goal at the time. The fact that a local state appeared in each location after the split has to be kept apart from the retrospective account of why I split them.

Operational Problems the Split Exposed

Splitting the files three ways made the plan shorter. The values passed between services and the ownership of remote resources still needed separate decisions.

Problem left overWhat the code and state showWhat the operator had to check
Duplicated inputsThe project, zone, instance, and bucket names typed again into several variablesCompare whether the three directories point at the same targets
Unused outputsNo other configuration reads the server’s instance and bucket outputsEdit the other input files whenever a value changes
Skipped deploy approvalTwo Windows batch files run apply -auto-approve right after building the ZIPAdd a manual step to read a separate plan before applying

Even after the split, the states still overlapped. Part 24 puts the three states side by side and identifies the duplicated project APIs and backup-retention settings.

The three directories produced separate plans for the server, preemption recovery, and backup cleanup. That met the goal of splitting the deployment unit. The instance and bucket names still had to be matched up by hand, though, and the project APIs were duplicated across two states.

The exact ZIP build, plan, and apply commands used by the three roots remain in part 21.

References

Comments

Comments

    Image preview