The Deployment Contract in VM Metadata and the Startup Script

How Terraform inputs became files and systemd services at boot, including values with no consuming code

The 2025 configuration passed nearly every input needed to build the server through the metadata block on google_compute_instance.

metadata = {
  startup-script      = file("${path.module}/scripts/startup.sh")
  shutdown-script     = file("${path.module}/scripts/shutdown.sh")
  server-properties   = file("${path.module}/scripts/minecraft/server.properties")
  backup-script       = file("${path.module}/scripts/backup.sh")
  minecraft-service   = file("${path.module}/scripts/systemd/minecraft-server.service")
  backup-timer        = file("${path.module}/scripts/systemd/minecraft-backup.timer")

  discord-bot-zip = filebase64("${path.module}/scripts/discord-v2.zip")
}

Shell scripts, systemd units, the server settings file, and the entire Discord bot application went down the same metadata channel. Terraform stored the text files as strings and encoded only the ZIP with filebase64. Inside the booted VM, startup.sh pulled the values out one by one and placed them on the file system.

I compared the values Terraform sent with the keys read inside the VM, using the code preserved in the repository. Every quote below also comes from that repository. Whether a deploy actually succeeded, and when each value was uploaded, cannot be confirmed because there are no plan or apply logs.

Two Kinds of Metadata Key

The metadata keys split into names handled by Compute Engine and custom names read by the startup script.

One side is names Compute Engine fixes by convention. startup-script and shutdown-script are run by the guest agent at boot and shutdown on its own. ssh-keys is also used by the platform on the login path. When the guest agent and VM configuration are in place, the platform handles these keys for their defined purposes.

The rest were names I made up. Settings like backup-bucket and rcon-port, file bodies like backup-script and minecraft-service, and application bundles like discord-bot-zip belong here. Compute Engine stores these values, and scripts in the VM read and use them through the metadata server.

get_md(){
  curl -fsS -H "Metadata-Flavor: Google" ... \
    "http://metadata.google.internal/computeMetadata/v1/$1" || true
}

This function in startup.sh read the custom keys. Terraform sent the values and the shell script read them, but nothing compared the key names in the two files. A misspelled key name still passes plan, and a value sent but never read still applies successfully.

What Happened in a Single Boot

main() in startup.sh called fourteen steps in order.

configure_sshd          Change SSH settings
fetch_metadata          Look up instance name, zone, external IP
install_packages        Install base packages
install_java            Install Temurin 21
prepare_world           Create directories, attempt automatic backup restore
download_purpur         Download Purpur 1.21.6
agree_eula              Write eula.txt
extract_metadata_files  metadata → script files
setup_discord_env       metadata → .env file
extract_and_unzip_...   metadata (base64) → bot application
extract_service_files   metadata → systemd units
set_script_permissions  Grant execute permission
install_python_...      Install Python dependencies
enable_and_start_...    Start services and send Discord notification

After the VM was created, one boot ran operating system configuration, runtime installation, application deployment, and data restore. Terraform stored the script in metadata; the shell script then placed the files and started the services.

extract_service_files writes the bodies it pulls from metadata straight to the unit paths.

declare -A SVC_MAP=(
  [minecraft-service]="/etc/systemd/system/minecraft-server.service"
  [backup-service]="/etc/systemd/system/minecraft-backup.service"
  [backup-timer]="/etc/systemd/system/minecraft-backup.timer"
  ...
)
for attr in "${!SVC_MAP[@]}"; do
  get_md "instance/attributes/$attr" > "${SVC_MAP[$attr]}" ...
done

That is the path by which a repository file read by Terraform’s file() function passes through metadata and arrives at the VM’s /etc/systemd/system. There is no validation on the way. If the value is empty, an empty unit file is created.

Five Mismatches Found in the Code

The code that sent and the code that read sat in different files, and there was no check comparing the two. Putting the two sides together identifies keys that no code reads.

flowchart LR
  TF["Terraform metadata block"] --> K1["startup-script<br/>shutdown-script"]
  TF --> K2["backup-script<br/>4 systemd units"]
  TF --> K3["discord-bot-zip"]
  TF --> K4["server-properties<br/>minecraft-port"]
  K1 --> GA["Guest agent runs it"]
  K2 --> SH["startup.sh places it as a file"]
  K3 --> SH
  K4 --> NONE["No code reads it"]

The code shows five mismatches.

Values With No Code Reading Them

server-properties rode the metadata all the way to the VM, but no script and no bot code reads that key. The same goes for project-id. Look only at the Terraform side and it appears to be a configuration that manages server settings as code, but the real server.properties was either the file Purpur generated on first run or the file inside a restored backup.

A Value Another Program Was Reading

minecraft-port is slightly different. startup.sh does not read this key; it wrote the port into the startup notification text as a constant. Instead, the Discord bot deployed over the same metadata queries this key and builds the connection information from it.

One side ignores the value and the other uses it. Which behavior is the correct one cannot be told from the code alone, and changing the port makes the bot’s announcement and the startup notification text diverge. One key had two consumers, and those two were looking at different places.

A Value That Never Arrived Because the Key Name Was Misspelled

The spelling of a name also differed. Terraform sent the key with hyphens, while the receiving script queried it with underscores.

# The key main.tf sent
discord-webhook-url-backup

# The key backup.sh looked for
get_metadata "discord_webhook_url_backup"

The lookup always returned an empty value, and the script used the fallback constant written inside itself. The shutdown notification had the same shape. In both cases it looks like the value is managed through metadata, but the value actually used was a constant baked into the script, and no error appeared to say so.

The Same Unit File Written to Two Paths

backup-timer is written twice. extract_metadata_files writes it once to /opt/minecraft/scripts/systemd/, and extract_service_files writes it again to /etc/systemd/system/. Only the latter is what systemd reads. The earlier file stays where nobody looks at it, holding the same content.

Why that copy was made at the time cannot be told from the code alone. It made the file to inspect when fixing the timer less obvious.

A Restore Path That Could Not Run

Two defects overlapped between backup and restore. The function that brings the world back when a new VM boots is prepare_world.

prepare_world(){
  ...
  mkdir -p "$MINECRAFT_DIR"/{world,scripts/discord,scripts/systemd}
  cd "$MINECRAFT_DIR"
  if [ ! -d world ]; then
    log "Attempting backup restore"
    ...
  else
    log "world directory exists—skipping restore"
  fi
}

mkdir creates the world directory, and two lines later the script asks whether that directory is missing. The condition is always false. On every boot the script logged world directory exists—skipping restore and moved on. Whatever was inside, it was code that never ran.

The decompression command inside the restore branch also differed from the backup format. backup.sh compresses with tar and pigz.

BACKUP_FILE="minecraft-backup-$(date +%Y%m%d-%H%M%S).tar.gz"
tar -I 'pigz -9' -cvf "$BACKUP_PATH" -C "$TMP_BACKUP_DIR" .

The unreachable restore branch fetches the most recent object from the same bucket and opens it with unzip.

latest=$(gsutil ls "gs://$BUCKET/backups/" | sort | tail -n1)
gsutil cp "$latest" "$tmp"
unzip -q "$tmp" -d "$MINECRAFT_DIR"

A .tar.gz does not open with unzip. The two scripts were deployed from the same repository to the same VM through the same metadata block, and they assumed different archive formats. The defect is two layers deep, so fixing one does not make restore work. Correct the condition and the format mismatch surfaces; match the formats and the condition still blocks the branch.

When a VM restarted with its disk kept, the world remained on the disk. Automatic restore was needed when the disk was created fresh. Yet the world directory created by mkdir prevented the restore branch from running on that fresh disk too.

A Backup Interval Where the Announcement and the Reality Differed

startup.sh had the backup interval as a constant.

BACKUP_FREQ_MIN=60

This value was used only to build the Discord notification text. Every time the server came up, the message “the server is automatically backed up every 60 minutes” went to the channel. The timer unit that decides the actual run was different.

[Timer]
OnCalendar=*-*-* HH:00:00
AccuracySec=1s
Persistent=true

Once a day. The announcement posted to the channel said every 60 minutes; the timer unit said once a day. The same person managed both values in the same repository, but one was text and the other was configuration, so they did not change together.

The Constraints This Contract Created

Metadata put files on the VM without a separate artifact server and let the Terraform plan show infrastructure and file changes together. This configuration had three constraints.

Fixing one line of a script becomes an instance change. When a file read by file() changes, the metadata value changes, and a VM attribute change appears in the plan. Application edits and infrastructure edits mix into the same plan.

Applying it requires another boot. Even when metadata changes, /etc/systemd/system on an already running VM stays as it was, because startup-script runs only at boot. On top of that, nothing calls systemctl daemon-reload after writing the unit files anew, so it is hard to say the first boot where the files are created behaves the same as a reboot where the contents changed.

Secrets land on disk. setup_discord_env builds an environment file from values pulled out of metadata.

cat > "$env_file" << EOF
DISCORD_BOT_TOKEN=$bot_token
RCON_PASSWORD=$rcon_password
EOF
chmod 600 "$env_file"

The permissions were restricted to 0600. That does not change the fact that a value that came through metadata is left as a file on the VM disk.

There is one more constraint in the server unit itself. The heap is a fixed value.

ExecStart=/usr/bin/java -Xms8G -Xmx8G ... -jar purpur.jar nogui
Restart=always

Growing the machine type does not raise this number with it. To increase memory I had to edit the unit file, apply the metadata again, and reboot the VM. Both services ran as User=root, and configure_sshd forcing PermitRootLogin yes was a choice from the same period.

What the Public Example Cut

The build track’s minecraft-one-root also uses metadata as a deployment channel. It reduced the amount loaded onto the channel and narrowed the contract.

It does not put in an application bundle. The server JAR is passed as a URL and a SHA-256 hash over metadata, and the VM downloads it itself and starts the service only after verifying the hash.

It keeps no keys that are never read. The game port and backup settings go over metadata, and the startup script actually uses those values to build server.properties.

Backup and restore use the same format. A backup uploads a tar.gz together with a .sha256 manifest, and the restore and verify commands work on that same format. The isolated restore test actually opens that archive and checks that the server comes up.

Properties remain even so. Editing the startup script is still a VM change, and applying it needs a reboot. As long as metadata is the deployment channel, this property does not go away. At the scale of one person running one server it is an acceptable cost, and part 08 treats that reboot procedure as part of a change operation.

The 2025 configuration loaded scripts, units, settings, an application, and secrets onto this channel all at once. apply could still succeed with unread keys, files written twice, and unreachable branches. A test that ran the first-boot path end to end could have exercised those code paths. The public example includes a container test to find the same class of omission.

References

Comments

Comments

    Image preview