The Deployment Contract in VM Metadata and the Startup Script
How Terraform inputs became files and systemd services at boot, including values with no consuming code
The 2025 configuration passed nearly every input needed to build the server through the metadata
block on google_compute_instance.
metadata = {
startup-script = file("${path.module}/scripts/startup.sh")
shutdown-script = file("${path.module}/scripts/shutdown.sh")
server-properties = file("${path.module}/scripts/minecraft/server.properties")
backup-script = file("${path.module}/scripts/backup.sh")
minecraft-service = file("${path.module}/scripts/systemd/minecraft-server.service")
backup-timer = file("${path.module}/scripts/systemd/minecraft-backup.timer")
discord-bot-zip = filebase64("${path.module}/scripts/discord-v2.zip")
}
Shell scripts, systemd units, the server settings file, and the entire Discord bot application went
down the same metadata channel. Terraform stored the text files as strings and encoded only the ZIP
with filebase64. Inside the booted VM, startup.sh pulled the values out one by one and placed
them on the file system.
I compared the values Terraform sent with the keys read inside the VM, using the code preserved in the repository. Every quote below also comes from that repository. Whether a deploy actually succeeded, and when each value was uploaded, cannot be confirmed because there are no plan or apply logs.
Two Kinds of Metadata Key
The metadata keys split into names handled by Compute Engine and custom names read by the startup script.
One side is names Compute Engine fixes by convention. startup-script and shutdown-script are run
by the guest agent at boot and shutdown on its own. ssh-keys is also used by the platform on the
login path. When the guest agent and VM configuration are in place, the platform handles these keys
for their defined purposes.
The rest were names I made up. Settings like backup-bucket and rcon-port, file bodies like
backup-script and minecraft-service, and application bundles like discord-bot-zip belong here.
Compute Engine stores these values, and scripts in the VM read and use them through the metadata
server.
get_md(){
curl -fsS -H "Metadata-Flavor: Google" ... \
"http://metadata.google.internal/computeMetadata/v1/$1" || true
}
This function in startup.sh read the custom keys. Terraform sent the values and the shell script
read them, but nothing compared the key names in the two files. A misspelled key name still
passes plan, and a value sent but never read still applies successfully.
What Happened in a Single Boot
main() in startup.sh called fourteen steps in order.
configure_sshd Change SSH settings
fetch_metadata Look up instance name, zone, external IP
install_packages Install base packages
install_java Install Temurin 21
prepare_world Create directories, attempt automatic backup restore
download_purpur Download Purpur 1.21.6
agree_eula Write eula.txt
extract_metadata_files metadata → script files
setup_discord_env metadata → .env file
extract_and_unzip_... metadata (base64) → bot application
extract_service_files metadata → systemd units
set_script_permissions Grant execute permission
install_python_... Install Python dependencies
enable_and_start_... Start services and send Discord notification
After the VM was created, one boot ran operating system configuration, runtime installation, application deployment, and data restore. Terraform stored the script in metadata; the shell script then placed the files and started the services.
extract_service_files writes the bodies it pulls from metadata straight to the unit paths.
declare -A SVC_MAP=(
[minecraft-service]="/etc/systemd/system/minecraft-server.service"
[backup-service]="/etc/systemd/system/minecraft-backup.service"
[backup-timer]="/etc/systemd/system/minecraft-backup.timer"
...
)
for attr in "${!SVC_MAP[@]}"; do
get_md "instance/attributes/$attr" > "${SVC_MAP[$attr]}" ...
done
That is the path by which a repository file read by Terraform’s file() function passes through
metadata and arrives at the VM’s /etc/systemd/system. There is no validation on the way. If the
value is empty, an empty unit file is created.
Five Mismatches Found in the Code
The code that sent and the code that read sat in different files, and there was no check comparing the two. Putting the two sides together identifies keys that no code reads.
flowchart LR TF["Terraform metadata block"] --> K1["startup-script<br/>shutdown-script"] TF --> K2["backup-script<br/>4 systemd units"] TF --> K3["discord-bot-zip"] TF --> K4["server-properties<br/>minecraft-port"] K1 --> GA["Guest agent runs it"] K2 --> SH["startup.sh places it as a file"] K3 --> SH K4 --> NONE["No code reads it"]
The code shows five mismatches.
Values With No Code Reading Them
server-properties rode the metadata all the way to the VM, but no script and no bot code reads that
key. The same goes for project-id. Look only at the Terraform side and it appears to be a
configuration that manages server settings as code, but the real server.properties was either the
file Purpur generated on first run or the file inside a restored backup.
A Value Another Program Was Reading
minecraft-port is slightly different. startup.sh does not read this key; it wrote the port into
the startup notification text as a constant. Instead, the Discord bot deployed over the same metadata
queries this key and builds the connection information from it.
One side ignores the value and the other uses it. Which behavior is the correct one cannot be told from the code alone, and changing the port makes the bot’s announcement and the startup notification text diverge. One key had two consumers, and those two were looking at different places.
A Value That Never Arrived Because the Key Name Was Misspelled
The spelling of a name also differed. Terraform sent the key with hyphens, while the receiving script queried it with underscores.
# The key main.tf sent
discord-webhook-url-backup
# The key backup.sh looked for
get_metadata "discord_webhook_url_backup"
The lookup always returned an empty value, and the script used the fallback constant written inside itself. The shutdown notification had the same shape. In both cases it looks like the value is managed through metadata, but the value actually used was a constant baked into the script, and no error appeared to say so.
The Same Unit File Written to Two Paths
backup-timer is written twice. extract_metadata_files writes it once to
/opt/minecraft/scripts/systemd/, and extract_service_files writes it again to
/etc/systemd/system/. Only the latter is what systemd reads. The earlier file stays where nobody
looks at it, holding the same content.
Why that copy was made at the time cannot be told from the code alone. It made the file to inspect when fixing the timer less obvious.
A Restore Path That Could Not Run
Two defects overlapped between backup and restore. The function that brings the world back when a
new VM boots is prepare_world.
prepare_world(){
...
mkdir -p "$MINECRAFT_DIR"/{world,scripts/discord,scripts/systemd}
cd "$MINECRAFT_DIR"
if [ ! -d world ]; then
log "Attempting backup restore"
...
else
log "world directory exists—skipping restore"
fi
}
mkdir creates the world directory, and two lines later the script asks whether that directory is
missing. The condition is always false. On every boot the script logged
world directory exists—skipping restore and moved on. Whatever was inside, it was code that never
ran.
The decompression command inside the restore branch also differed from the backup format.
backup.sh compresses with tar and pigz.
BACKUP_FILE="minecraft-backup-$(date +%Y%m%d-%H%M%S).tar.gz"
tar -I 'pigz -9' -cvf "$BACKUP_PATH" -C "$TMP_BACKUP_DIR" .
The unreachable restore branch fetches the most recent object from the same bucket and opens it with
unzip.
latest=$(gsutil ls "gs://$BUCKET/backups/" | sort | tail -n1)
gsutil cp "$latest" "$tmp"
unzip -q "$tmp" -d "$MINECRAFT_DIR"
A .tar.gz does not open with unzip. The two scripts were deployed from the same repository to the
same VM through the same metadata block, and they assumed different archive formats. The defect is
two layers deep, so fixing one does not make restore work. Correct the condition and the format
mismatch surfaces; match the formats and the condition still blocks the branch.
When a VM restarted with its disk kept, the world remained on the disk. Automatic restore was needed
when the disk was created fresh. Yet the world directory created by mkdir prevented the restore
branch from running on that fresh disk too.
A Backup Interval Where the Announcement and the Reality Differed
startup.sh had the backup interval as a constant.
BACKUP_FREQ_MIN=60
This value was used only to build the Discord notification text. Every time the server came up, the message “the server is automatically backed up every 60 minutes” went to the channel. The timer unit that decides the actual run was different.
[Timer]
OnCalendar=*-*-* HH:00:00
AccuracySec=1s
Persistent=true
Once a day. The announcement posted to the channel said every 60 minutes; the timer unit said once a day. The same person managed both values in the same repository, but one was text and the other was configuration, so they did not change together.
The Constraints This Contract Created
Metadata put files on the VM without a separate artifact server and let the Terraform plan show infrastructure and file changes together. This configuration had three constraints.
Fixing one line of a script becomes an instance change. When a file read by file() changes, the
metadata value changes, and a VM attribute change appears in the plan. Application edits and
infrastructure edits mix into the same plan.
Applying it requires another boot. Even when metadata changes, /etc/systemd/system on an
already running VM stays as it was, because startup-script runs only at boot. On top of that,
nothing calls systemctl daemon-reload after writing the unit files anew, so it is hard to say the
first boot where the files are created behaves the same as a reboot where the contents changed.
Secrets land on disk. setup_discord_env builds an environment file from values pulled out of
metadata.
cat > "$env_file" << EOF
DISCORD_BOT_TOKEN=$bot_token
RCON_PASSWORD=$rcon_password
EOF
chmod 600 "$env_file"
The permissions were restricted to 0600. That does not change the fact that a value that came through metadata is left as a file on the VM disk.
There is one more constraint in the server unit itself. The heap is a fixed value.
ExecStart=/usr/bin/java -Xms8G -Xmx8G ... -jar purpur.jar nogui
Restart=always
Growing the machine type does not raise this number with it. To increase memory I had to edit the
unit file, apply the metadata again, and reboot the VM. Both services ran as User=root, and
configure_sshd forcing PermitRootLogin yes was a choice from the same period.
What the Public Example Cut
The build track’s minecraft-one-root also uses metadata as a deployment channel. It reduced the
amount loaded onto the channel and narrowed the contract.
It does not put in an application bundle. The server JAR is passed as a URL and a SHA-256 hash over metadata, and the VM downloads it itself and starts the service only after verifying the hash.
It keeps no keys that are never read. The game port and backup settings go over metadata, and the
startup script actually uses those values to build server.properties.
Backup and restore use the same format. A backup uploads a tar.gz together with a .sha256
manifest, and the restore and verify commands work on that same format. The
isolated restore test actually opens
that archive and checks that the server comes up.
Properties remain even so. Editing the startup script is still a VM change, and applying it needs a reboot. As long as metadata is the deployment channel, this property does not go away. At the scale of one person running one server it is an acceptable cost, and part 08 treats that reboot procedure as part of a change operation.
The 2025 configuration loaded scripts, units, settings, an application, and secrets onto this channel
all at once. apply could still succeed with unread keys, files written twice, and unreachable
branches. A test that ran the first-boot path end to end could have exercised those code paths. The
public example includes a container test to find the same class of omission.
References
Comments
No comments yet. Be the first to leave one.
Pending review