Real incidents where Terraform went wrong in production, and the mix-ups that trip people up in interviews and in the Terraform Associate exam. Each one links back to the chapter that would have prevented it. Stories stick. Use them to lock in what you learned.
How these were chosen. Only cases with a public, named source: a first-person postmortem, a company incident record, a conference talk or a vendor security bulletin. Each card says which. We state only what the source states, and we say where the record is silent. We found no well-documented, named breach traced to a leaked state file, so we don't invent one. The risk is documented by HashiCorp, and our lab shows it. There's also no reliable public data on why people fail interviews. The traps below are topics the official exam tests, answered from HashiCorp's own documentation.
Production failures
Checked on 30 September 2026.
26 February 2026DataTalks.Club course platform
Terraform lost track of production, then destroyed it
TriggerNew computer; the state file stayed on the old one
What Terraform didSaw no state, so planned duplicates
TriggerAn older state file replaced the current one
What Terraform didterraform destroy removed production
Blast radiusVPC, database, containers, load balancers, bastion, and the snapshots
RecoveryAWS support restored a snapshot after about 24 hours
What changedState in S3 · deletion protection · backups outside Terraform · restore tests · plans read by hand
What happened
The founder had moved to a new computer. The Terraform state file was still on the old one, so Terraform believed nothing existed and started creating duplicate infrastructure.
An old state file, describing the real production platform, was then unpacked over the current one. A terraform destroy meant to clean up the duplicates removed the production VPC, database, container cluster, load balancers and bastion host.
The database's automated snapshots were deleted with it. AWS support found a snapshot and restored the data about 24 hours later: 2.5 years of course submissions.
The plan and apply had been delegated to an AI agent, and the plan wasn't reviewed by hand before it ran. His fixes: state moved to S3, deletion protection in Terraform and in AWS, backups outside Terraform, daily restore tests, and reviewing every plan himself.
Deletion protection on databases, Terraform's lifecycle { prevent_destroy = true }, backups that don't depend on Terraform, and actually testing a restore. None of these are covered in the course yet. One caution: prevent_destroy only protects resources whose blocks are still in the code being applied. Remove the block and Terraform will destroy the resource anyway (HashiCorp docs), so use it alongside provider deletion protection, not instead of it. Worth knowing: in Terraform's AWS provider, aws_db_instance deletes automated backups along with the database by default (delete_automated_backups = true) and leaves deletion_protection off. Alexey's post doesn't say which settings he used, so this is not presented as the cause.
A ── B ── C ── D↑↑oldcurrentA pipeline restarted at B made a fresh plan of old code, while D was current. Illustration of the failure shape, not GitLab’s real history.
TriggerA pipeline prepared three weeks earlier was restarted
Missing guardrailFresh plan, old code: the saved-plan check can't catch that
What Terraform didPlanned and applied the old configuration: 617 resources to destroy
Blast radiusGitLab.com down 16:25–18:42 UTC; three Gitaly nodes deleted
RecoveryServices restored; under 30 minutes of data lost on each node
What changedCorrective action: fail the apply if several data disks would be deleted
What happened
GitLab.com was unavailable from 16:25 to 18:42 UTC, and the container registry for some users until 19:36 UTC.
GitLab's incident review gives the root cause as an out-of-sync Terraform configuration run against production. It had been prepared three weeks earlier for a database upgrade. When the pipeline was restarted, it removed and replaced production services.
The plan for that pipeline showed 617 resources to be destroyed. With GitLab.com down, the team ran the incident from Google Docs.
A follow-up records the accidental deletion of three recently created Gitaly nodes (GitLab's Git storage service), with under 30 minutes of data lost on each.
GitLab's own description, as quoted in public discussion, stresses this was not a stale saved plan. It was a fresh plan and apply of an old commit, which Terraform's saved-plan check can't catch.
Deployment freshness: only the current, reviewed commit should ever reach apply, so an old pipeline can't be re-run against production. Also automatic stops on large destroy counts, and prevent_destroy on resources that hold data. How GitLab recovered services and data is in its incident review; it doesn't say Terraform state history was part of that, so we don't link it to state recovery. One caution: prevent_destroy only protects resources whose blocks are still in the code being applied. Remove the block and Terraform will destroy the resource anyway (HashiCorp docs), so use it alongside provider deletion protection, not instead of it.
TriggerA 50-node production cluster deleted by accident
Recovery3.25 hours to restore, slowed by buggy scripts and thin docs
TriggerA month later: review builds changed global state; pull requests merged out of order
What Terraform didIts view of the clusters changed; second incident ran 8 PM to 5 AM
Blast radiusNo end-user impact: traffic failed over outside Kubernetes
What changedPlan on every pull request · up-to-date branches · fail builds on "destroy" · recovery drills
What happened
First incident: an engineer deleted a 50-node production Kubernetes cluster running dozens of workloads. Restoring it took 3.25 hours, slowed by bugs in cluster-creation scripts, incomplete documentation and an all-or-nothing creation process.
A month later, while the team was putting clusters into code to prevent accidental deletions, review builds unknowingly modified global state and two pull requests were merged out of order. Recreating the cluster needed different permissions, which changed Terraform's view of the clusters. That incident lasted from 8 PM to 5 AM.
End users weren't affected: teams had moved only part of each service to Kubernetes, and traffic failed over to instances outside it.
Afterwards the team posted the dry-run output on every pull request, required up-to-date branches and approved reviews, failed review builds when the dry run contained "destroy", and practised disaster recovery. As one slide puts it: if you've never restored from backups, you don't have backups.
TriggerAn attacker modified Codecov's tool, which HashiCorp used
ExposureHashiCorp's release-signing key
ResponseKey rotated and releases re-signed (bulletin, 22 April 2021)
ImpactTerraform 0.12.0 to 0.12.30 couldn't verify new providers
What changedUpgrade to 0.12.31 or later
What happened
Codecov, a code-coverage tool, disclosed on 15 April 2021 that an attacker had modified a component its customers download and run. HashiCorp was affected.
The exposed secret was HashiCorp's GPG private key, used to sign the checksums that verify its product downloads.
HashiCorp revoked and rotated the key, re-signed most existing releases, and published security bulletin HCSEC-2021-12 on 22 April 2021. Updated Terraform binaries followed on 26 April.
HashiCorp's notice covered Terraform 0.12.0 to 0.12.30: those versions couldn't verify new provider releases until upgraded to 0.12.31 or later. A public GitHub issue shows the warning on 0.12.6.
One command with -auto-approve deleted a production database
Forum report by the person involved. Not independently verified.
Triggerapply -auto-approve with the production variables file
Missing guardrailsUndeclared variables only warned · one state key for every workspace · backup retention 0 days
What Terraform didDeleted the MySQL database, then failed to recreate it
RecoveryNot stated in the post
What happened
The engineer reports running terraform apply -auto-approve with a production variables file. The apply deleted the existing MySQL database and then failed to recreate it.
They list the traps: undeclared variables raised only warnings, one remote state key was shared by every workspace, and automated backup retention was set to 0 days.
Database backup retention above zero, and prevent_destroy on the database. One caution: prevent_destroy only protects resources whose blocks are still in the code being applied. Remove the block and Terraform will destroy the resource anyway (HashiCorp docs), so use it alongside provider deletion protection, not instead of it.
Things people say with confidence that aren't true. Each is tied to an objective in HashiCorp's Terraform Associate (004) exam content list, which tests Terraform 1.12. Knowing the "why" is what interviewers listen for.
1"Marking a variable sensitive encrypts it."
No. sensitive hides the value in Terraform's output. The real value is still written to state in plain text. Our lab found the password there in plain text (twice, in our OpenTofu 1.12.6 test).
Not quite. State maps real-world resources to your configuration, and that includes objects you imported. Its main job is storing the link between each remote object and a resource in your code.
terraform refresh is deprecated. It's the same as apply -refresh-only -auto-approve, which changes state without asking. To look safely, run terraform plan -refresh-only.
No. The dependency lock file records providers only. Terraform picks the newest module version your constraint allows. Use an exact version, such as = 5.2.1, or a commit ref when you need the same one every time.
No. The version argument only works for modules from a registry. For Git sources, ?ref= chooses a revision: a release tag, or a commit hash for a reference that never changes. A branch can move.
6"Renaming a resource or moving it into a module is harmless."
By default Terraform treats a new address as destroy-and-create. A moved block tells it the object simply moved. In our lab: 4 to destroy without it, 0 with it.
No. It removes Terraform's record without destroying the real object. The next plan tries to create it again, which may fail if the name or ID is already taken.
No. Import adds an existing object to state, so Terraform manages it from now on. You still write the matching configuration, and Terraform records that it imported the object rather than creating it.
Not in Terraform: pass values at init with -backend-config. OpenTofu 1.8+ does allow variables here, which is a good way to show you know the difference.
11"Every child module needs its own provider block."
No. Configure providers in the root module; children inherit that configuration. Each child still declares the providers it needs in required_providers.
14"prevent_destroy means the resource can't be deleted."
No. It makes Terraform reject plans that would destroy the resource while its block is in the configuration. Remove the block and Terraform destroys it anyway. HashiCorp says to use it sparingly. Pair it with provider deletion protection, such as deletion_protection = true on a database, and reviewed plans.