PRACTICE TRACK / 115 QUESTIONS

Terraform
Think it through.

State, plan/apply, modules, and drift.

Choose a question, explain your approach, then reveal the supplied answer. Difficulty labels come from the existing question library.

115 questions

Answers stay closed until you choose to reveal them.

QUESTION 01TerraformEasy

What is Terraform state and why does it matter?

#
Reveal answer guidance

State is Terraform's record (terraform.tfstate) mapping your configuration to real-world resource IDs. It's how Terraform knows what it already created, computes diffs, and stores metadata/dependencies. Without it Terraform couldn't tell create from update from delete. It can contain secrets, so it should live in a secure remote backend, never plain in git.

QUESTION 02TerraformEasy

What is the difference between terraform plan and apply?

#
Reveal answer guidance

plan computes and shows the diff between desired config and current state — what will be created, changed, or destroyed — without making changes. apply executes that plan against the real infrastructure. In CI you typically save the plan and apply exactly that artifact so what you reviewed is what runs.

QUESTION 03TerraformMedium

Why use remote state with locking, and how does locking work?

#
Reveal answer guidance

Remote state (S3, GCS, Terraform Cloud) lets a team share one source of truth and keeps secrets off laptops. Locking (e.g. DynamoDB for the S3 backend) prevents two applies from running at once and corrupting state — the first apply takes the lock, others wait or fail fast. Always pair remote state with a lock for any shared environment.

QUESTION 04TerraformMedium

What is configuration drift and how do you handle it?

#
Reveal answer guidance

Drift is when real infrastructure no longer matches state because someone changed it out-of-band (console click, another tool). terraform plan surfaces it as unexpected changes; terraform refresh/-refresh-only reconciles state with reality. You fix it by re-applying to bring resources back to config, or by importing/updating config to match the intended new reality — and ideally prevent it with policy and least-privilege console access.

QUESTION 05TerraformHard

How do count and for_each differ, and when do you prefer for_each?

#
Reveal answer guidance

count creates resources indexed by integer (resource[0], [1]…). Removing an item in the middle shifts every later index, causing destroy/recreate churn. for_each keys resources by a stable map/set key, so adding or removing one item only affects that key. Prefer for_each whenever the set can change or items have meaningful identities; reserve count for simple "N identical copies" or conditional (count = var.enabled ? 1 : 0) cases.

QUESTION 06TerraformMedium

Your Terraform plan output shows a security group rule will be destroyed and recreated even though you only added a new rule. What causes this, and how do you prevent unnecessary churn?

#
Reveal answer guidance

This happens when you use a resource "aws_security_group_rule" with count or inline ingress/egress blocks that are order-sensitive. Terraform sees the entire block as an unordered set, but if you use TypeSet with ingress blocks inline, adding a rule in the middle shifts indices. Fix: use separate aws_security_group_rule resources (one per rule) with for_each keyed by a stable value (port+protocol+CIDR), not count. Or use aws_vpc_security_group_ingress_rule / aws_vpc_security_group_egress_rule (AWS provider v4+) which use a separate resource per rule and never cause churn on unrelated rules. Always prefer for_each over count for rule-like resources.

QUESTION 07TerraformHard

You need to split a monolithic Terraform state into two separate states (e.g., network and application). Walk through the complete migration with exact commands.

#
Reveal answer guidance

Step 1: Create new root modules (terraform/network, terraform/app) and new empty state backends (e.g., separate S3 keys). Step 2: In the monolith directory, use terraform state list to identify resources to move. Step 3: terraform state mv -state-out=../network/terraform.tfstate aws_vpc.main aws_vpc.main — this moves the VPC along with implicit dependencies (subnets, route tables) because Terraform tracks them via references. However, parent-child moves must be explicit; Terraform does not automatically move dependents. You must move each resource individually: terraform state mv aws_subnet.main[0] aws_subnet.main[0]. Step 4: Import moved state into the new backend: in the new module dir, terraform init then terraform state push ../monolith/terraform.tfstate. Step 5: Remove moved resources from the monolith config and state: terraform state rm aws_vpc.main. Step 6: terraform plan in both directories — they should show no changes. Risk: if dependent resources reference moved resources across state files, you need terraform_remote_state data sources, which add latency and potential for stale reads. Test in a clone of prod state first.

QUESTION 08TerraformHard

A terraform apply fails halfway — 10 resources created, then an API rate limit hits and the process is killed. What state is the infrastructure in? Walk through recovery.

#
Reveal answer guidance

The infrastructure has 10 resources created, but Terraform's state file was never updated (or was partially updated if the process was killed mid-write). Recovery: (1) Run terraform plan — Terraform shows 10 resources it thinks need creating (it doesn't know they exist) and the failed resource as needing creation. (2) Import the 10 orphaned resources: terraform import <resource_type>.<name> <id> for each. (3) Alternatively, use terraform refresh -target=<resource> to pull the 10 resources into state without touching the failed one. (4) Delete the orphaned resources if they're safe to recreate: either via cloud console/CLI or terraform destroy -target=<resource> after importing. (5) Once state is reconciled, terraform apply succeeds. Prevention: always use -auto-approve with caution, set higher max_retries and retry_max_attempts in the provider config, and use a remote backend with state locking so partial state writes don't corrupt the file.

QUESTION 09TerraformHard

How does Terraform resolve provider versions, and what happens when a provider upgrade introduces breaking changes to your resources? Walk through a safe provider upgrade strategy.

#
Reveal answer guidance

Terraform resolves providers via required_providers blocks, consulting version constraints (e.g. >= 4.0, < 5.0), the lock file (.terraform.lock.hcl), and available provider registries. When upgrading, the lock file pins exact versions by hash. Safe upgrade strategy: (1) Run terraform init -upgrade to update the lock file to latest matching constraint. (2) Review the provider changelog for breaking attributes. (3) Run terraform plan — any ForceNew attributes will show as destroy/recreate; note them. (4) Use -refresh-only first to check for unexpected drift. (5) Apply in a non-production environment first. (6) If the provider changed the internal type of an attribute (e.g. from string to list(string)), use terraform state replace-provider hashicorp/aws hashicorp/aws to migrate state, then use terraform plan to verify. (7) For major version bumps, run both providers side-by-side during migration using provider aliases. Key insight: never run terraform init -upgrade in production without first running the exact same upgrade in a lower environment and verifying zero unexpected changes.

QUESTION 10TerraformHard

Explain the moved block (Terraform 1.1+). When would you use it instead of terraform state mv?

#
Reveal answer guidance

The moved block lets you declare in HCL that a resource was renamed or moved to a new module path, and Terraform handles the state migration automatically during plan/apply. Example: moved { from = aws_instance.web_old; to = aws_instance.web }. Use it over terraform state mv when: (1) the refactor is permanent and committed to version control — the moved block lives in config so all team members and CI get the migration automatically; (2) you're refactoring a module — moved blocks in the module automatically migrate consumers of that module; (3) you're changing count to for_each — moved { from = aws_subnet.private[0]; to = aws_subnet.private["az1a"] } handles key changes. Keep moved blocks for one release cycle, then remove them after everyone has applied. terraform state mv is better for ad-hoc operations, emergency state fixes, or cross-state moves where the target lives in a different backend.

QUESTION 11TerraformHard

What is the removed block in Terraform 1.7+, and how does it differ from deleting a resource block from config?

#
Reveal answer guidance

The removed block explicitly tells Terraform to destroy a resource and remove it from state when you apply, but keeps the config visible for documentation. Example: removed { from = aws_s3_bucket.old_bucket; lifecycle { destroy = true } }. Difference from just deleting the config block: if you delete the resource block and run apply, Terraform prompts you to destroy the resource (or you confirm with -target). With removed, you declare intent explicitly — Terraform will destroy it on next apply without needing -target. The lifecycle { destroy = false } variant creates a terraform state rm effect (forget the resource without destroying it). This is useful for: (1) migrating resources to a new state file — remove with destroy = false and import into the new state; (2) explicitly documenting that a resource is being decommissioned; (3) suppressing the destructive plan until you're ready. Unlike simply deleting the block, removed gives you a safer, tracked process visible to your team in the same config.

QUESTION 12TerraformHard

How do Terraform dynamic blocks work, and what are their limitations compared to generating resources with for_each?

#
Reveal answer guidance

dynamic blocks generate nested configuration blocks inside a single resource, not separate resources. They iterate over a collection and produce block instances. Example: dynamic "tag" { for_each = var.tags; content { key = tag.key; value = tag.value } }. Limitations: (1) Only supports nested blocks that a resource schema defines as TypeList or TypeSet — you cannot use dynamic to create multiple resources. (2) The generated block arguments must all be optional in the schema; required arguments inside a dynamic block will fail if the collection is empty. (3) Debugging is harder — a complex dynamic block produces opaque errors. (4) Cannot use count or for_each inside a dynamic block. (5) for_each on resource creation (resource "x" { for_each = var.items }) is fundamentally different — it creates N separate resources, each with its own state, lifecycle, and ability to be targeted independently. Use dynamic for optional inline rule blocks within a single resource (e.g., security group rules inside one aws_security_group). Use for_each on the resource itself when you need independent lifecycle management for each instance.

QUESTION 13TerraformHard

What are provider-defined functions (Terraform 1.8+), and how do they differ from templatefile, local values, or terraform_data?

#
Reveal answer guidance

Provider-defined functions are functions exposed by providers (e.g., AWS, GCP) that can be called directly in HCL as provider::namespace::function_name(args). For example, provider::aws::arn_build("s3", var.bucket_name, var.region) to construct an ARN without string interpolation. They differ from: (1) templatefile — renders a file template with variables; it's a built-in function, provider-defined functions go beyond what built-ins can offer, like parsing cloud-specific resource IDs. (2) local values — computed expressions used to reduce repetition; provider functions can be called inside locals but are not a replacement — they add new computational capabilities that previously required external data sources or null_resource with local-exec. (3) terraform_data — a resource used as a container for arbitrary data; provider functions replace many use cases where you needed a terraform_data to call an external script just to compute something provider-specific. The key difference: provider functions are pure computation — they don't create, read, or manage infrastructure state, so they can be used in locals, variables, and outputs without creating additional resources in the state file.

QUESTION 14TerraformHard

What is the Terraform test framework (introduced in 1.6+)? Walk through the structure of a test file and how you would test a module that creates an S3 bucket with versioning.

#
Reveal answer guidance

Terraform tests use .tftest.hcl files that define run blocks, each running terraform plan and apply against the module with different inputs, then asserting conditions. Structure: a mock_provider or real provider block at the top, then run blocks. Example: run "setup_network" { command = apply; variables = { cidr_block = "10.0.0.0/16" } } followed by assertions: assert { condition = aws_vpc.main.cidr_block == "10.0.0.0/16"; error_message = "CIDR mismatch" }. For a versioned S3 bucket: (1) Create tests/basic.tftest.hcl with a run "create_bucket" block that passes bucket name and enables versioning. (2) Assert the bucket exists and versioning[0].enabled == true. (3) Add a run "update_bucket" block that changes a tag and re-applies, asserting the change applied without recreating. (4) Run terraform test — it creates real resources (unless using mock_provider), runs all assertions, and destroys everything in the correct order. Best practices: use command = plan for validation-only tests (cheap), command = apply for state-dependent assertions (slower). Condition functions like can() and try() help write robust assertions. The framework runs in isolation — state is separate from your working state file.

QUESTION 15TerraformMedium

What is the difference between terraform init, terraform plan, and terraform apply in a CI/CD pipeline?

#
Reveal answer guidance

In CI/CD: (1) terraform init is run once per job to initialize the working directory, download providers, and configure the backend. It should use the lock file (.terraform.lock.hcl) committed to version control for deterministic provider versions. (2) terraform plan is run on pull requests/merge requests to produce a diff artifact (plan file). The plan should be saved and reviewed as part of the PR. (3) terraform apply runs only on the main branch after merge, consuming the exact plan artifact from the plan step — never re-planning. This ensures what was reviewed is exactly what runs. The apply step also requires access to the remote backend (S3/GCS/TFC) for state locking. Additional CI patterns: use -input=false everywhere to prevent blocking on prompts; use terraform validate and terraform fmt -check as pre-merge checks; use terraform show -json on the plan file for programmatic review (e.g., checking for unintended resource destruction).

QUESTION 16TerraformHard

A production terraform apply wants to replace an RDS database because identifier changed. How do you stop data loss and still complete the rename requirement?

#
Reveal answer guidance

First stop the apply and confirm whether the rename is a provider ForceNew attribute. For an RDS identifier, create a migration plan: snapshot the DB, add lifecycle { prevent_destroy = true }, use a blue/green or replica cutover if downtime matters, update DNS or application endpoints, then intentionally remove the old instance after validation. Terraform should model the safe migration steps, not blindly accept a destroy/create plan.

QUESTION 17TerraformHard

Two engineers run Terraform against the same AWS account and the second run fails with a state lock error. What do you check before using force-unlock?

#
Reveal answer guidance

Check the backend lock record and CI history to confirm no apply is still running. For S3 backends, inspect the DynamoDB lock item or native lock file if configured; for Terraform Cloud, check the active run. Use force-unlock only when the original process is definitely dead, because unlocking during an active apply can corrupt state and cause duplicate infrastructure changes.

QUESTION 18TerraformHard

A module used count for subnets and removing one subnet from the middle of a list causes Terraform to recreate several others. How do you fix the design without downtime?

#
Reveal answer guidance

Move from list indexes to stable keys with for_each, such as subnet names or AZ IDs. Use moved blocks or terraform state mv to map aws_subnet.private[0] to aws_subnet.private["app-a"] and so on before changing the resource. Run a plan that shows only the intended deletion; if replacements appear, stop and repair the state mapping.

QUESTION 19TerraformMedium

A team stores Terraform state in S3 but sees occasional concurrent applies. What backend controls should be enabled?

#
Reveal answer guidance

Use a remote backend with locking and versioning. For S3, enable bucket versioning, encryption, strict bucket policy, and state locking using the backend-supported lock mechanism for the Terraform version in use. CI should serialize environment applies and require reviewed saved plans so state locking is the last line of defense, not the only one.

QUESTION 20TerraformHard

Terraform shows drift because an autoscaling group desired capacity changed during a production incident. Should you add ignore_changes?

#
Reveal answer guidance

Only ignore fields that are intentionally controlled outside Terraform. If the ASG desired capacity is managed by autoscaling policies or incident runbooks, ignore_changes = [desired_capacity] is reasonable while keeping min and max capacity managed. Do not use broad ignore rules to hide real drift such as launch template, subnet, security group, or tag changes.

QUESTION 21TerraformHard

A provider upgrade changes default behavior and the plan shows hundreds of tag diffs. How do you roll out the upgrade safely?

#
Reveal answer guidance

Pin the current provider, read the upgrade guide, test in a non-production workspace, and separate the provider version bump from functional infrastructure changes. If defaults changed, set explicit arguments or provider default_tags to preserve behavior. Review JSON plan output in CI and merge only when the diff is understood and intentionally accepted.

QUESTION 22TerraformMedium

A resource was created manually in AWS and now must be managed by Terraform. What is the safest import workflow?

#
Reveal answer guidance

Write the matching resource block first, import the real ID with an import block or terraform import, then run terraform plan and adjust arguments until the plan is empty or contains only intentional changes. Never import into a guessed config and immediately apply; that often overwrites production settings.

QUESTION 23TerraformHard

A Terraform module outputs a password generated by random_password. Why is marking the output sensitive = true not enough?

#
Reveal answer guidance

sensitive = true hides the value in CLI output, but the value still exists in Terraform state. The backend must be encrypted and access-controlled, state access should be limited, and secrets should preferably be written to a secrets manager with only references exposed to downstream modules.

QUESTION 24TerraformHard

A CI pipeline runs terraform plan on pull requests from forks. What security risk does that create?

#
Reveal answer guidance

Terraform can execute provider code and data sources that may use credentials, and PR code can alter providers, modules, or external data commands. Do not expose cloud credentials to untrusted fork plans. Use read-only speculative plans in isolated accounts, restrict module sources, avoid external data sources for untrusted changes, and require trusted maintainers to run privileged plans.

QUESTION 25TerraformMedium

Your plan file was saved in CI but apply fails because variables changed between plan and apply. What should the pipeline do?

#
Reveal answer guidance

Generate a binary plan with all variables and provider versions fixed, store it as an immutable artifact, and apply exactly that file with terraform apply tfplan. The apply job should not recompute a fresh plan with different environment variables, because the reviewed diff would no longer be the executed diff.

QUESTION 26TerraformHard

A VPC module is used by 20 services and a breaking output rename is needed. How do you release it?

#
Reveal answer guidance

Use semantic versioning and publish a backwards-compatible version first that exposes both old and new outputs. Migrate consumers gradually, then release a major version removing the old output. Pin module versions in every consumer so one module release does not unexpectedly break all environments.

QUESTION 27TerraformHard

Terraform wants to delete and recreate an IAM role used by live workloads because the role name changed. How do you avoid outage?

#
Reveal answer guidance

Treat IAM names as stable API contracts. If a rename is mandatory, create the new role in parallel, attach equivalent policies, update workloads to assume or attach the new role, validate, then remove the old role. In Terraform, use separate resources or staged applies; do not let a single plan replace an in-use identity unexpectedly.

QUESTION 28TerraformMedium

When should a value be a module input variable versus discovered through a data source?

#
Reveal answer guidance

Use variables for explicit dependencies and environment decisions that callers own. Use data sources for stable existing infrastructure that Terraform does not manage or that is managed in a separate state. Avoid hidden data-source lookups for critical dependencies because they make plans less predictable and can accidentally bind to the wrong resource.

QUESTION 29TerraformHard

A data.aws_ami lookup always picks the latest image and causes unexpected EC2 replacement. How do you control this?

#
Reveal answer guidance

Use a tested AMI promotion process instead of resolving latest directly in production. Pin the AMI ID through a variable, SSM parameter, or environment release artifact. Let Terraform consume the promoted value, so instance replacement happens only when the platform team intentionally promotes a new image.

QUESTION 30TerraformHard

A root module has 2,000 resources and plans take 35 minutes. What practical changes improve Terraform performance and operability?

#
Reveal answer guidance

Split state by ownership and blast radius, such as network, shared data, cluster, and application layers. Reduce expensive data sources, avoid giant module graphs, use provider aliases only where needed, and run plans only for affected stacks in CI. Do not split purely by resource type if teams still need cross-stack coordination for every change.

QUESTION 31TerraformMedium

A developer used depends_on between two modules to fix ordering. What is the downside?

#
Reveal answer guidance

Module-level depends_on forces all resources in one module to depend on all resources in another, making plans conservative and slower and causing unnecessary replacement or unknown values. Prefer passing the exact output needed by the dependent module; that creates a precise dependency edge.

QUESTION 32TerraformHard

An S3 bucket policy references a CloudFront distribution ARN from another state. How should the dependency be wired?

#
Reveal answer guidance

Prefer explicit remote state outputs or a shared configuration registry such as SSM Parameter Store. The producer state should output only the distribution ID or ARN needed, and the consumer should read that value by environment. Avoid duplicating naming logic in both stacks because it can silently point the policy at the wrong distribution.

QUESTION 33TerraformHard

A terraform destroy in a shared sandbox would delete a shared VPC used by other stacks. How do you design for safer teardown?

#
Reveal answer guidance

Separate shared infrastructure into its own state and restrict destroy permissions for that state. Add prevent_destroy on critical shared resources, use ownership tags and policy checks, and make application stacks consume shared IDs as inputs rather than owning the shared VPC resources.

QUESTION 34TerraformMedium

How do you structure Terraform for dev, staging, and production without copy-pasting entire directories?

#
Reveal answer guidance

Keep reusable modules for common infrastructure and use thin environment root modules that pass different inputs. Pin module versions, keep backend state separate per environment, and keep production changes gated by reviews. Avoid workspaces as the only isolation boundary for complex environments with different blast radii or approval rules.

QUESTION 35TerraformHard

A Terraform workspace points to the wrong AWS account and the plan wants to create duplicate resources. What controls prevent this?

#
Reveal answer guidance

Use explicit provider account checks with data.aws_caller_identity and preconditions or validation against the expected account ID. Configure CI credentials per environment, name backend keys clearly, and fail fast when workspace, region, or account does not match the expected values.

QUESTION 36TerraformHard

A resource argument is unknown during plan and causes a module to fail validation. How do you redesign it?

#
Reveal answer guidance

Avoid requiring apply-time values for plan-time decisions such as for_each keys or provider configuration. Use stable input keys, split the deployment into layers, or output the generated value from one state and consume it in a later run. Terraform needs collection keys known during planning.

QUESTION 37TerraformMedium

What is the difference between terraform taint, -replace, and deleting from state?

#
Reveal answer guidance

-replace=addr asks Terraform to replace a resource in the next plan and is the modern explicit option. taint marks state for replacement but is less preferred for CI workflows. Removing from state only makes Terraform forget the object; it does not destroy it and can lead to duplicate resource creation if the config still exists.

QUESTION 38TerraformHard

An apply failed halfway after creating a load balancer but before updating DNS. What is your recovery flow?

#
Reveal answer guidance

Do not rerun blindly. Inspect state and cloud resources, run terraform plan to see what Terraform believes is done, and complete or roll back using Terraform where possible. If a resource exists but is missing from state, import it or remove the duplicate before applying again. Then add tests or staged changes for the failure point.

QUESTION 39TerraformHard

A module accepts a list of security group rules. Reordering the list causes rule replacements. How do you avoid noisy and risky diffs?

#
Reveal answer guidance

Model rules as a map keyed by a stable semantic ID, then use for_each. Each rule key should represent intent, such as ingress_https_from_alb, not its list position. This keeps Terraform from treating a reorder as deletion and recreation of unrelated rules.

QUESTION 40TerraformMedium

When should you use lifecycle { create_before_destroy = true }?

#
Reveal answer guidance

Use it for replaceable resources where the new object can exist alongside the old one, such as launch templates or some target groups. It is not a universal zero-downtime switch; names, quotas, uniqueness constraints, and dependencies may prevent parallel creation. Always verify the plan and naming strategy.

QUESTION 41TerraformHard

A security group rule resource uses create_before_destroy, but replacement fails due to duplicate rule errors. Why?

#
Reveal answer guidance

Some cloud APIs treat equivalent rules as unique even if Terraform sees them as separate resources. Terraform tries to create the replacement first, but AWS rejects a duplicate. The fix is to avoid unnecessary replacement, use stable keyed rules, or allow destroy-before-create for that specific rule when the brief change is acceptable.

QUESTION 42TerraformHard

A team wants to use Terraform to manage Kubernetes resources immediately after creating an EKS cluster in the same apply. What can go wrong?

#
Reveal answer guidance

The Kubernetes provider needs cluster endpoint, auth, and API readiness during planning and applying. In one graph, provider configuration may be unknown or the cluster API may not be ready when Kubernetes resources apply. A safer pattern is two stages: create the cluster first, then configure Kubernetes and Helm providers from the stable cluster outputs.

QUESTION 43TerraformMedium

How do you prevent accidental public S3 buckets with Terraform?

#
Reveal answer guidance

Set bucket public access block resources, restrictive bucket policies, encryption, and ownership controls in the module. Add policy-as-code checks in CI for public ACLs and wildcard principals. Use account-level S3 Block Public Access so Terraform mistakes are blocked by AWS as well.

QUESTION 44TerraformHard

A module uses random_id for resource names and every environment gets different names. When is that good and when is it harmful?

#
Reveal answer guidance

Random suffixes are useful for globally unique names such as S3 buckets or test environments. They are harmful when names are operational contracts used by monitoring, IAM, DNS, or humans. For production, prefer deterministic names with controlled suffixes and only use randomness where the cloud API requires uniqueness.

QUESTION 45TerraformHard

Terraform plan shows an IAM policy JSON diff on every run even though permissions did not change. How do you fix it?

#
Reveal answer guidance

Generate IAM JSON with aws_iam_policy_document or jsonencode from structured values so key ordering is stable. Avoid heredocs assembled by string concatenation. If AWS normalizes the policy differently, compare the effective document and update the config to match provider normalization.

QUESTION 46TerraformMedium

How do you handle provider aliases for deploying resources across multiple AWS regions?

#
Reveal answer guidance

Define one provider block per region with an alias, then pass the correct provider explicitly into modules using the providers map. Keep region-specific resources in keyed maps so the module knows which provider to use. Avoid relying on the default provider for multi-region stacks because accidental region placement is hard to detect later.

QUESTION 47TerraformHard

An ACM certificate in us-east-1 is needed for CloudFront, but the rest of the stack is in ap-south-1. How should Terraform model this?

#
Reveal answer guidance

Use an aliased AWS provider for us-east-1 for the ACM certificate and DNS validation resources as needed, while keeping regional infrastructure on the primary provider. Pass the certificate ARN to the CloudFront resource. This avoids creating the certificate in the wrong region, which CloudFront cannot use.

QUESTION 48TerraformHard

A module uses null_resource with local-exec to configure production servers. What risks does this create?

#
Reveal answer guidance

local-exec runs outside Terraform stateful resource semantics and depends on the runner environment, network access, and command idempotency. It can leak secrets, fail partially, and be hard to roll back. Prefer cloud-init, image baking, SSM documents, Kubernetes jobs, or provider-native resources that expose declarative state.

QUESTION 49TerraformMedium

What is a good use case for terraform_data compared with null_resource?

#
Reveal answer guidance

terraform_data is useful for storing computed values or creating replacement triggers inside Terraform without the external provisioner behavior of null_resource. It can help coordinate lifecycle triggers, but it should not become a replacement for real provider resources or configuration management.

QUESTION 50TerraformHard

A Terraform module must enforce that production RDS has deletion protection enabled. Where do you enforce it?

#
Reveal answer guidance

Enforce it in multiple layers: module defaults and validation, resource preconditions, CI policy checks, and cloud IAM or SCP controls where possible. Terraform validation gives fast feedback, but cloud-side controls protect against manual changes and misconfigured modules outside the repo.

QUESTION 51TerraformHard

A plan includes known after apply for a value used in a DNS record name. Why is this dangerous?

#
Reveal answer guidance

DNS record names are public and often part of dependencies. If the name is unknown until apply, reviewers cannot verify the exact endpoint being created and downstream references may fail. Prefer deterministic names from inputs and use apply-time values only for record values such as generated load balancer hostnames.

QUESTION 52TerraformMedium

How do you migrate Terraform state from local files to an S3 backend?

#
Reveal answer guidance

Commit backend configuration without the local state file, initialize with terraform init -migrate-state, and let Terraform copy the state to S3. Then verify versioning, encryption, and locking, remove local state from developer machines, and ensure .tfstate files are ignored by git.

QUESTION 53TerraformHard

A state file accidentally committed to Git contains secrets. What is the incident response?

#
Reveal answer guidance

Assume compromise. Rotate all exposed credentials, remove the file from the repository history, invalidate any derived tokens, and audit access logs. Moving the state to a secure backend is necessary but not sufficient because Git history may have already exposed the values.

QUESTION 54TerraformHard

A module output exposes a full database connection string. What is a better design?

#
Reveal answer guidance

Expose non-secret metadata such as endpoint, port, and secret ARN separately. Store credentials in a secrets manager and let applications retrieve them through IAM. If a connection string must be assembled, do it at deploy time in the consuming platform, not as a Terraform output stored in state.

QUESTION 55TerraformMedium

How do variable validation blocks improve production safety?

#
Reveal answer guidance

They fail fast before planning or applying invalid inputs, such as unsupported instance classes, CIDR ranges, or environment names. They are not a substitute for policy-as-code, but they make module contracts explicit and prevent common caller mistakes.

QUESTION 56TerraformHard

A caller passes 0.0.0.0/0 to a database security group module. How should a mature module handle it?

#
Reveal answer guidance

The module should reject that input with validation or a precondition unless an explicit emergency override is provided and audited. Security-sensitive modules should encode safe defaults and narrow allowed sources. CI policy should also block public database ingress regardless of module behavior.

QUESTION 57TerraformHard

Terraform Cloud speculative plans pass, but applies fail due to missing runtime permissions. How do you diagnose this?

#
Reveal answer guidance

Compare credentials used for speculative plan and apply, including workspace variables, run tasks, and cloud role assumptions. Plans may need only read permissions while applies need writes. Ensure the apply role has the intended least privilege and that policy checks are not masking provider authentication failures.

QUESTION 58TerraformMedium

Why should production Terraform runs avoid developer laptops?

#
Reveal answer guidance

Laptop runs vary by Terraform version, provider cache, credentials, local environment variables, and network access. CI or Terraform Cloud gives auditable runs, consistent versions, saved plan artifacts, controlled credentials, and review gates. Emergency manual applies should be rare and documented.

QUESTION 59TerraformHard

A provider data source calls an API that rate-limits heavily and makes every plan flaky. What do you change?

#
Reveal answer guidance

Cache or promote the discovered value into a controlled input such as SSM Parameter Store, remote state output, or release metadata. Reduce repeated data source calls by centralizing lookups in locals. If the value changes independently, create a scheduled update process rather than making every Terraform plan depend on a flaky API.

QUESTION 60TerraformHard

You need to rotate a KMS key used by many encrypted resources. How should Terraform represent the rotation?

#
Reveal answer guidance

For automatic annual rotation, enable rotation on the existing key if supported. For a new key, create it in parallel, update aliases and resource encryption settings in stages, re-encrypt or recreate dependent resources according to service behavior, then retire the old key only after all consumers and backups no longer need it.

QUESTION 61TerraformMedium

What should be included in a Terraform module README for interview-quality production work?

#
Reveal answer guidance

Inputs, outputs, provider requirements, examples, security assumptions, resource ownership, upgrade notes, known limitations, and migration guidance. The README should state what the module manages and what it intentionally leaves to callers, because unclear ownership creates drift and unsafe changes.

QUESTION 62TerraformHard

A module version upgrade changes a default from private to public subnets. How do you prevent accidental exposure?

#
Reveal answer guidance

Do not rely on risky defaults. Make sensitive placement explicit with required variables, validate allowed subnet types, and add CI policies that detect public route table associations or public IP assignment. Module upgrades should be reviewed with a plan from every affected environment.

QUESTION 63TerraformHard

Terraform destroys and recreates an EKS node group when changing labels. How do you reduce disruption?

#
Reveal answer guidance

Check whether the provider marks the field ForceNew. If replacement is required, create a new node group with the desired labels, let workloads drain through PDB-aware rolling migration, then remove the old node group. For mutable settings, prefer managed update mechanisms, but verify real provider behavior before applying.

QUESTION 64TerraformMedium

How do you model optional resources in Terraform without creating index errors?

#
Reveal answer guidance

Use for_each with an empty or single-item map, or expose outputs with try and clear null handling. Avoid scattered resource[0] references that fail when count = 0. Optional modules should have well-defined null outputs and callers should handle the disabled case explicitly.

QUESTION 65TerraformHard

A module has a boolean create_vpc and many conditional resources. When does this become a bad design?

#
Reveal answer guidance

If the module has many modes, complex conditional outputs, and hidden dependencies, it becomes hard to reason about and test. Split into smaller modules: one to create the VPC and another to consume an existing VPC. Simple feature toggles are fine; whole ownership modes usually deserve separate modules.

QUESTION 66TerraformHard

A Terraform apply updates Route 53 records before the new load balancer is healthy. How do you design a safer cutover?

#
Reveal answer guidance

Separate provisioning from traffic cutover. Create the new load balancer and target group, validate health checks, then update weighted or failover DNS records gradually. Terraform can manage both stages, but the pipeline should require an explicit approval before shifting production traffic.

QUESTION 67TerraformMedium

How does Terraform decide resource order?

#
Reveal answer guidance

Terraform builds a dependency graph from references, provider dependencies, lifecycle rules, and explicit depends_on. Resources without dependencies can run in parallel. Correctly passing outputs between resources is better than adding broad manual dependencies because it gives Terraform an accurate graph.

QUESTION 68TerraformHard

A resource must be replaced whenever a rendered user-data template changes. What is a robust approach?

#
Reveal answer guidance

For EC2 instances, put user data in a launch template and let an autoscaling group instance refresh or deployment process replace instances safely. Use hash-based triggers only where the provider lacks native change detection. Avoid directly replacing pets in production without drain and rollback steps.

QUESTION 69TerraformHard

A Terraform run fails because the provider plugin version differs between developers. How do you fix version consistency?

#
Reveal answer guidance

Commit .terraform.lock.hcl, pin provider constraints deliberately, and run terraform init -upgrade only during planned upgrades. CI should use a fixed Terraform version and the lock file. This ensures everyone gets the same provider checksums and avoids surprise behavior changes.

QUESTION 70TerraformMedium

What is the difference between required_version and provider version constraints?

#
Reveal answer guidance

required_version constrains the Terraform CLI version. required_providers constrains provider plugins such as AWS, AzureRM, or Kubernetes. Both matter because language features, state behavior, and provider schemas can change independently.

QUESTION 71TerraformHard

A new Terraform CLI version changes plan behavior in CI. What rollout process should be used?

#
Reveal answer guidance

Pin the old version, test the new version in a branch against representative stacks, review release notes, and roll it out environment by environment. Keep provider upgrades separate from CLI upgrades where possible so failures have a clear cause.

QUESTION 72TerraformHard

A module references remote state from production while planning staging. What is the risk?

#
Reveal answer guidance

It can accidentally couple staging to production IDs, secrets, or endpoints, causing tests to hit real production infrastructure or apply policies to the wrong resources. Remote state keys should be parameterized and validated by environment, and sensitive production outputs should not be broadly readable.

QUESTION 73TerraformMedium

How do you handle secrets needed by providers, such as cloud credentials?

#
Reveal answer guidance

Use short-lived credentials from OIDC or a workload identity system in CI. Avoid static keys in variables or tfvars files. Provider credentials should come from the runner identity or secure workspace variables, and state access should be separated from infrastructure write permissions where possible.

QUESTION 74TerraformHard

A Terraform module manages IAM policies for many teams. How do you prevent privilege escalation through module inputs?

#
Reveal answer guidance

Do not accept raw arbitrary policy JSON for high-trust roles unless the caller is equally trusted. Offer constrained inputs, validate allowed actions and resources, run policy-as-code checks, and use permission boundaries or SCPs so even a bad module call cannot grant unrestricted admin access.

QUESTION 75TerraformHard

A plan wants to remove a prevent_destroy resource after someone deleted the lifecycle block. How can this be caught?

#
Reveal answer guidance

Use code review and CI policies that require prevent_destroy for critical resource types such as production databases, KMS keys, and DNS zones. Lifecycle blocks are config, so deleting them removes the guard. Cloud-side deletion protection and IAM deny rules provide stronger independent protection.

QUESTION 76TerraformMedium

What is an appropriate use of terraform refresh-only?

#
Reveal answer guidance

Use it to update state to match real infrastructure without changing resources, typically during drift investigation or after emergency manual changes. Review the refresh-only plan carefully because it can record drift into state, making later plans treat the changed real-world values as the new baseline unless config disagrees.

QUESTION 77TerraformHard

An emergency console change fixed production. How do you bring Terraform back under control?

#
Reveal answer guidance

Document the change, run a plan or refresh-only plan to understand drift, then decide whether Terraform config should adopt the change or revert it. Update code and tests, import any new objects if needed, and apply from CI. The goal is to make the emergency change intentional in version control.

QUESTION 78TerraformHard

Terraform is used to manage DNS records across many microservices. How do you avoid teams overwriting each other?

#
Reveal answer guidance

Give each service ownership of only its record names, or centralize DNS in a platform module with clear inputs. Use separate states by zone or service, enforce naming conventions, and avoid broad aws_route53_record resources generated from shared mutable maps unless ownership and review are strict.

QUESTION 79TerraformMedium

How should .tfvars files be handled in a repository?

#
Reveal answer guidance

Non-secret environment values can be committed if they are reviewed configuration. Secret tfvars should not be committed; use secure workspace variables or a secrets manager. Keep examples such as example.tfvars to document required inputs without exposing real credentials.

QUESTION 80TerraformHard

A module depends on an external service account existing before apply. Should Terraform create it or read it?

#
Reveal answer guidance

That depends on ownership. If the platform team owns the service account lifecycle, Terraform should create and output it. If another team owns it, read it through a controlled input or data source and validate it exists. Avoid a module that sometimes owns and sometimes borrows the same identity without clear mode separation.

QUESTION 81TerraformHard

A Terraform apply changes an ALB listener rule priority and causes routing conflicts. How do you manage listener rules safely?

#
Reveal answer guidance

Use deterministic priority assignment, validate uniqueness, and avoid deriving priorities from unstable list indexes. For large shared ALBs, centralize listener rule ownership or allocate priority ranges per service. Plans should clearly show old and new routing behavior before apply.

QUESTION 82TerraformMedium

What is the problem with using timestamps in resource names or triggers?

#
Reveal answer guidance

Timestamps change on every plan or apply, causing perpetual diffs and unnecessary replacement. Use content hashes for real change detection or stable release versions for intentional rollouts. Terraform configuration should be deterministic wherever possible.

QUESTION 83TerraformHard

A Docker image tag latest is passed into Terraform for ECS task definitions. Why is this a production problem?

#
Reveal answer guidance

Terraform cannot know when the image behind latest changes, and rollbacks are ambiguous. Use immutable image digests or versioned tags produced by CI. Terraform should deploy a specific artifact, so the plan represents a reproducible release.

QUESTION 84TerraformHard

An ECS service update through Terraform causes all tasks to restart at once. What controls should be checked?

#
Reveal answer guidance

Review deployment configuration such as minimum healthy percent and maximum percent, health check grace period, load balancer health checks, capacity, and task definition changes. Terraform submits the desired service update, but ECS controls rollout behavior. Set safe deployment parameters and verify enough capacity exists before applying.

QUESTION 85TerraformMedium

How do you avoid circular dependencies between security groups?

#
Reveal answer guidance

Create security groups first, then create rules as separate resources that reference both group IDs. Inline rules can create cycles or force replacement in complex graphs. Separate rule resources also make ownership and diffs clearer.

QUESTION 86TerraformHard

A state split is needed because one stack became too large. How do you migrate safely?

#
Reveal answer guidance

Create the new root module and backend, move selected resources with terraform state mv or equivalent moved/import workflow, and verify both old and new plans are clean. Do not recreate resources. Migrate in small batches, back up state first, and coordinate locks so no one applies during the split.

QUESTION 87TerraformHard

A state backend bucket is accidentally deleted. What recovery options exist?

#
Reveal answer guidance

If bucket versioning and backups exist, restore the latest valid state object and lock metadata. If not, rebuild state by importing existing infrastructure into matching config, starting with critical shared resources. This is why remote state needs versioning, retention, backups, and restricted delete permissions.

QUESTION 88TerraformMedium

What are moved blocks and when are they better than terraform state mv?

#
Reveal answer guidance

moved blocks declare address changes in code so Terraform can migrate state during plan for everyone consistently. They are better for reviewed refactors that should travel with the module. terraform state mv is still useful for one-off emergency or cross-state migrations.

QUESTION 89TerraformHard

A resource address changes because a module was renamed. What happens without a moved block?

#
Reveal answer guidance

Terraform treats the old address as removed and the new address as a separate desired resource, so the plan may destroy and recreate infrastructure. A moved block maps the old module address to the new one and preserves the real object in state.

QUESTION 90TerraformHard

How do you test Terraform modules before production?

#
Reveal answer guidance

Use static checks (terraform fmt, validate, linting), policy checks, example plans, and ephemeral integration tests in a sandbox account for critical modules. Validate outputs and important resource attributes. For modules managing risky resources, test upgrade paths and replacement behavior, not just initial creation.

QUESTION 91TerraformMedium

What should a good Terraform pull request include?

#
Reveal answer guidance

It should include the code change, the relevant plan output or link, explanation of replacements and destroys, blast radius, rollback notes, and any required manual steps. Reviewers should be able to tell what real infrastructure will change without running Terraform themselves.

QUESTION 92TerraformHard

A plan shows -/+ replacement for a production load balancer. What questions should you answer before approval?

#
Reveal answer guidance

Why is replacement required, can the new load balancer exist in parallel, how will DNS or clients cut over, are certificates and security groups ready, what is rollback, and what downtime is expected? If those answers are not clear, split the change into provision, validate, cutover, and cleanup stages.

QUESTION 93TerraformHard

A Terraform module provisions both network and application resources. Why might that be bad?

#
Reveal answer guidance

Network and application resources usually have different owners, lifecycles, blast radius, and approval requirements. Combining them means a routine app change can accidentally affect shared networking. Separate stacks let network changes be rare and controlled while application deployments stay fast.

QUESTION 94TerraformMedium

When is a monorepo for Terraform helpful, and when is it painful?

#
Reveal answer guidance

A monorepo helps standardize modules, policies, and review workflows across environments. It becomes painful if every change triggers plans for every stack, ownership is unclear, or state boundaries are too broad. Use path-based CI and clear CODEOWNERS to keep it manageable.

QUESTION 95TerraformHard

A provider bug creates bad state for one resource. How do you proceed?

#
Reveal answer guidance

Back up state, read provider issues and upgrade notes, test a provider upgrade, and use state surgery only as a last resort. If editing state is unavoidable, do it with locks held, minimal changes, peer review, and a clean plan afterward. Prefer import or provider-supported fixes where possible.

QUESTION 96TerraformHard

A cloud resource was deleted manually but still exists in Terraform state. What will the next plan show?

#
Reveal answer guidance

After refresh, Terraform should detect the object is gone and plan to recreate it if the config still declares it. If recreation is unsafe, remove or modify the config first. If the deletion was intentional, remove it from config and state through a reviewed change.

QUESTION 97TerraformMedium

How should Terraform handle generated files such as rendered Helm values?

#
Reveal answer guidance

Prefer generating values in memory with templatefile, yamlencode, or structured inputs rather than committing generated artifacts. If generated files are needed for review, ensure they are deterministic and produced by CI. Avoid local files that differ between developer machines.

QUESTION 98TerraformHard

A Helm release managed by Terraform fails and leaves Kubernetes resources partially updated. What is the recovery approach?

#
Reveal answer guidance

Inspect Helm release status, Kubernetes events, and Terraform state. Fix the chart values or cluster condition, then rerun the plan. If the release is stuck, use Helm-native rollback or uninstall carefully, then reconcile Terraform state. For critical apps, consider GitOps controllers for app rollout and Terraform only for cluster prerequisites.

QUESTION 99TerraformHard

Terraform manages both a database and the application schema migration job. Why is this risky?

#
Reveal answer guidance

Terraform is infrastructure orchestration, not an application deployment transaction manager. Schema migrations need ordering, rollback, retries, and app version coordination. Terraform can provision the database and permissions, but migrations are usually safer in a deployment pipeline or migration tool with explicit operational controls.

QUESTION 100TerraformMedium

What is the difference between module composition and deeply nested modules?

#
Reveal answer guidance

Composition keeps root modules wiring small focused modules together with visible inputs and outputs. Deep nesting hides dependencies and makes changes hard to reason about. Prefer shallow composition so environment owners can see the infrastructure graph and override important settings deliberately.

QUESTION 101TerraformHard

A cost spike happens after Terraform creates NAT gateways in every AZ. How should this be reviewed?

#
Reveal answer guidance

Check whether high availability required one NAT gateway per AZ or whether a lower-cost non-production design is acceptable. Terraform modules should expose cost-impacting choices clearly and use environment-specific defaults. Add cost estimation or policy checks in CI for expensive resources.

QUESTION 102TerraformHard

A module creates CloudWatch alarms but alert thresholds differ by service. How should inputs be designed?

#
Reveal answer guidance

Provide sensible defaults and allow explicit per-service threshold maps with validation. Avoid hardcoding thresholds in the module or accepting an unstructured blob that bypasses validation. Outputs should make alarm names and ARNs visible to incident tooling.

QUESTION 103TerraformMedium

How do you use tags effectively in Terraform-managed infrastructure?

#
Reveal answer guidance

Set common tags at the provider or module level and require ownership, environment, cost center, and data classification tags. Avoid tag drift by centralizing defaults, but allow resource-specific tags where needed. Use policy checks to block missing mandatory tags.

QUESTION 104TerraformHard

A provider-level default_tags change updates thousands of resources. How do you roll it out?

#
Reveal answer guidance

Treat it as a broad infrastructure change. Plan by environment, check which resources support tag updates in place versus replacement, and avoid mixing tag rollout with functional changes. If many APIs are touched, apply during a low-risk window and monitor rate limits and service-specific behavior.

QUESTION 105TerraformHard

Terraform wants to recreate a KMS key because its description changed. What should you suspect?

#
Reveal answer guidance

A description alone normally should not require key replacement, so suspect an address change, provider schema issue, changed key spec, multi-region setting, or imported state mismatch. Inspect the exact ForceNew attribute in the plan JSON and provider docs before approving any KMS replacement.

QUESTION 106TerraformMedium

Why should backend configuration usually not depend on Terraform variables?

#
Reveal answer guidance

Backend initialization happens before normal Terraform evaluation, so standard variables are not available in the usual way. Use partial backend configuration with CI-provided backend config files or flags. Keep backend keys explicit and predictable to avoid writing state to the wrong location.

QUESTION 107TerraformHard

A CI job initializes the wrong backend key for production. How can this be prevented?

#
Reveal answer guidance

Make backend config derived from trusted pipeline metadata, not user-edited PR files, and validate account, region, and workspace before planning. Use separate credentials and protected branches for production. A plan should fail if the backend key or caller identity does not match the environment.

QUESTION 108TerraformHard

A service team wants direct write access to the production Terraform state bucket. What is the concern?

#
Reveal answer guidance

State write access is equivalent to control over Terraform ownership and may expose secrets. Teams should trigger controlled runs through CI or Terraform Cloud, not edit state directly. Grant least-privilege read only when needed, and separate state admin permissions from infrastructure authoring permissions.

QUESTION 109TerraformMedium

How do you handle Terraform outputs consumed by non-Terraform systems?

#
Reveal answer guidance

Publish stable outputs to a controlled registry such as SSM Parameter Store, Secrets Manager, DNS, or a service catalog. Directly reading state from scripts couples external systems to Terraform internals and may expose secrets. Outputs should have clear compatibility expectations.

QUESTION 110TerraformHard

An apply creates a resource but times out waiting for readiness. The next plan wants to create it again. What do you do?

#
Reveal answer guidance

Check whether the resource exists in the provider and whether it was recorded in state. If it exists but is missing from state, import it before applying again. Then tune timeouts or readiness dependencies if the API is slow. Reapplying blindly may create duplicates or fail on name conflicts.

QUESTION 111TerraformHard

A Terraform run is blocked by API throttling in a large AWS account. What mitigations are available?

#
Reveal answer guidance

Reduce unnecessary parallelism, split state by ownership, remove repeated data sources, use provider retry settings where available, and avoid broad account-wide lookups. Schedule large applies away from other automation and consider service quota increases if the workload is legitimate.

QUESTION 112TerraformMedium

When should you lower Terraform -parallelism?

#
Reveal answer guidance

Lower it when provider APIs throttle, dependencies are fragile, or a service cannot handle many concurrent changes. It can make applies slower but safer. Do not use it to hide bad dependency modeling; fix missing dependencies where order actually matters.

QUESTION 113TerraformHard

A module configures backup retention for databases. What production safeguards should it include?

#
Reveal answer guidance

Require minimum retention for production, enable deletion protection where supported, configure final snapshots, tag backups, and expose restore-related outputs. Use validation and policy checks so a caller cannot accidentally set retention to zero in a protected environment.

QUESTION 114TerraformHard

A resource must be adopted from another Terraform state. How do you avoid duplicate ownership?

#
Reveal answer guidance

Coordinate a state handoff: remove or move the resource from the old state and import or move it into the new state while both states are locked and no applies are running. Verify the old state no longer plans changes for it and the new state has a clean plan. Two states must never manage the same object.

QUESTION 115TerraformMedium

What is the danger of using the same module source branch for production?

#
Reveal answer guidance

A branch is mutable, so production can get different module code without an explicit version change. Use immutable tags or registry versions. Branch sources are acceptable for development testing but not for stable production dependencies.

CONTINUE PRACTICING

Try another perspective.