Cybserve Learning Terraform series

Real failures and interview traps

Real incidents where Terraform went wrong in production, and the mix-ups that trip people up in interviews and in the Terraform Associate exam. Each one links back to the chapter that would have prevented it. Stories stick. Use them to lock in what you learned.

How these were chosen. Only cases with a public, named source: a first-person postmortem, a company incident record, a conference talk or a vendor security bulletin. Each card says which. We state only what the source states, and we say where the record is silent. We found no well-documented, named breach traced to a leaked state file, so we don't invent one. The risk is documented by HashiCorp, and our lab shows it. There's also no reliable public data on why people fail interviews. The traps below are topics the official exam tests, answered from HashiCorp's own documentation.

Production failures

Checked on 30 September 2026.

26 February 2026DataTalks.Club course platform

Terraform lost track of production, then destroyed it

  1. TriggerNew computer; the state file stayed on the old one
  2. What Terraform didSaw no state, so planned duplicates
  3. TriggerAn older state file replaced the current one
  4. What Terraform didterraform destroy removed production
  5. Blast radiusVPC, database, containers, load balancers, bastion, and the snapshots
  6. RecoveryAWS support restored a snapshot after about 24 hours
  7. What changedState in S3 · deletion protection · backups outside Terraform · restore tests · plans read by hand

What happened

Where the course would have helped

Beyond this course

Deletion protection on databases, Terraform's lifecycle { prevent_destroy = true }, backups that don't depend on Terraform, and actually testing a restore. None of these are covered in the course yet. One caution: prevent_destroy only protects resources whose blocks are still in the code being applied. Remove the block and Terraform will destroy the resource anyway (HashiCorp docs), so use it alongside provider deletion protection, not instead of it. Worth knowing: in Terraform's AWS provider, aws_db_instance deletes automated backups along with the database by default (delete_automated_backups = true) and leaves deletion_protection off. Alexey's post doesn't say which settings he used, so this is not presented as the cause.

Source

7 July 2023GitLab.com

An out-of-date Terraform run took GitLab.com down

  1. TriggerA pipeline prepared three weeks earlier was restarted
  2. Missing guardrailFresh plan, old code: the saved-plan check can't catch that
  3. What Terraform didPlanned and applied the old configuration: 617 resources to destroy
  4. Blast radiusGitLab.com down 16:25–18:42 UTC; three Gitaly nodes deleted
  5. RecoveryServices restored; under 30 minutes of data lost on each node
  6. What changedCorrective action: fail the apply if several data disks would be deleted

What happened

Where the course would have helped

Beyond this course

Deployment freshness: only the current, reviewed commit should ever reach apply, so an old pipeline can't be re-run against production. Also automatic stops on large destroy counts, and prevent_destroy on resources that hold data. How GitLab recovered services and data is in its incident review; it doesn't say Terraform state history was part of that, so we don't link it to state recovery. One caution: prevent_destroy only protects resources whose blocks are still in the code being applied. Remove the block and Terraform will destroy the resource anyway (HashiCorp docs), so use it alongside provider deletion protection, not instead of it.

Source

Told at KubeCon Europe, May 2019Spotify

Production clusters deleted by accident, twice

  1. TriggerA 50-node production cluster deleted by accident
  2. Recovery3.25 hours to restore, slowed by buggy scripts and thin docs
  3. TriggerA month later: review builds changed global state; pull requests merged out of order
  4. What Terraform didIts view of the clusters changed; second incident ran 8 PM to 5 AM
  5. Blast radiusNo end-user impact: traffic failed over outside Kubernetes
  6. What changedPlan on every pull request · up-to-date branches · fail builds on "destroy" · recovery drills

What happened

Where the course would have helped

Beyond this course

Cluster backups you have actually restored from, disaster-recovery drills, and moving gradually with a fallback in place.

Source

April 2021HashiCorp (via Codecov)

The key that signs Terraform releases was exposed

  1. TriggerAn attacker modified Codecov's tool, which HashiCorp used
  2. ExposureHashiCorp's release-signing key
  3. ResponseKey rotated and releases re-signed (bulletin, 22 April 2021)
  4. ImpactTerraform 0.12.0 to 0.12.30 couldn't verify new providers
  5. What changedUpgrade to 0.12.31 or later

What happened

Where the course would have helped

Beyond this course

Supply-chain security: verifying downloads, and keeping Terraform itself up to date, not just your code.

Source

2025 (forum post)An engineer on AWS re:Post

One command with -auto-approve deleted a production database

Forum report by the person involved. Not independently verified.

  1. Triggerapply -auto-approve with the production variables file
  2. Missing guardrailsUndeclared variables only warned · one state key for every workspace · backup retention 0 days
  3. What Terraform didDeleted the MySQL database, then failed to recreate it
  4. RecoveryNot stated in the post

What happened

Where the course would have helped

Beyond this course

Database backup retention above zero, and prevent_destroy on the database. One caution: prevent_destroy only protects resources whose blocks are still in the code being applied. Remove the block and Terraform will destroy the resource anyway (HashiCorp docs), so use it alongside provider deletion protection, not instead of it.

Source

Interview and exam traps

Things people say with confidence that aren't true. Each is tied to an objective in HashiCorp's Terraform Associate (004) exam content list, which tests Terraform 1.12. Knowing the "why" is what interviewers listen for.

1"Marking a variable sensitive encrypts it."

No. sensitive hides the value in Terraform's output. The real value is still written to state in plain text. Our lab found the password there in plain text (twice, in our OpenTofu 1.12.6 test).

2"State is a list of what Terraform created."

Not quite. State maps real-world resources to your configuration, and that includes objects you imported. Its main job is storing the link between each remote object and a resource in your code.

Exam: Objective 2d: how Terraform uses and manages stateEpisode 2, chapter 2: What state is →HashiCorp docs

3"Use terraform refresh to check for drift."

terraform refresh is deprecated. It's the same as apply -refresh-only -auto-approve, which changes state without asking. To look safely, run terraform plan -refresh-only.

Exam: Objective 7: maintain infrastructureEpisode 2, chapter 8: When reality drifts →HashiCorp docs

4"The lock file pins my module versions."

No. The dependency lock file records providers only. Terraform picks the newest module version your constraint allows. Use an exact version, such as = 5.2.1, or a commit ref when you need the same one every time.

Exam: Objective 3b: initialise a working directoryEpisode 1, chapter 8: Where modules live →HashiCorp docs

5"You can set version = on a Git module."

No. The version argument only works for modules from a registry. For Git sources, ?ref= chooses a revision: a release tag, or a commit hash for a reference that never changes. A branch can move.

6"Renaming a resource or moving it into a module is harmless."

By default Terraform treats a new address as destroy-and-create. A moved block tells it the object simply moved. In our lab: 4 to destroy without it, 0 with it.

7"terraform state rm deletes the resource."

No. It removes Terraform's record without destroying the real object. The next plan tries to create it again, which may fail if the name or ID is already taken.

8"Import creates the resource."

No. Import adds an existing object to state, so Terraform manages it from now on. You still write the matching configuration, and Terraform records that it imported the object rather than creating it.

9"S3 state locking needs a DynamoDB table."

Not any more in Terraform. Set use_lockfile = true (Terraform 1.10+). DynamoDB-based locking is deprecated there. OpenTofu currently supports both.

Exam: Objective 6: state managementEpisode 2, chapter 6: Locking →HashiCorp docs

10"A backend block can use variables."

Not in Terraform: pass values at init with -backend-config. OpenTofu 1.8+ does allow variables here, which is a good way to show you know the difference.

11"Every child module needs its own provider block."

No. Configure providers in the root module; children inherit that configuration. Each child still declares the providers it needs in required_providers.

12"Sensitive outputs can't be seen by anyone."

They're hidden in normal plan and apply output, but terraform output -json and -raw print them in plain text, and they're stored in state.

Exam: Objective 4: Terraform configurationEpisode 2, chapter 3: What's inside state →HashiCorp docs

13"Local state is fine for a team if we commit it to Git."

HashiCorp advises against storing state in version control: it can lead to data loss or exposed secrets, and Git doesn't lock state.

14"prevent_destroy means the resource can't be deleted."

No. It makes Terraform reject plans that would destroy the resource while its block is in the configuration. Remove the block and Terraform destroys it anyway. HashiCorp says to use it sparingly. Pair it with provider deletion protection, such as deletion_protection = true on a database, and reviewed plans.

Official exam objectives: Terraform Associate (004) exam content list.

← Back to the course