Skip to main content

Terraform Infrastructure

Terraform lives in orphaned-wells-ui-server/deployment/terraform.

It manages the current GKE deployment infrastructure, upload buckets, DNS, namespace-level Kubernetes RBAC, and deployment target outputs. Legacy VM definitions remain in Terraform for future re-enablement, but Terraform manages legacy VM modules only for names listed in enabled_legacy_backend_vms.

What Terraform manages​

  • Shared GKE Autopilot cluster.
  • One Cloud Storage upload bucket per unique GKE backend bucket name.
  • One global static IP per backend collaborator for Kubernetes Ingress.
  • Optional <collaborator>-k8s-server.uow-carbon.org test DNS records.
  • Primary backend DNS records, such as staging-server.uow-carbon.org, pointing to the GKE static IP.
  • One Kubernetes namespace plus the backend-api/processing-worker ServiceAccounts and worker-dispatch Role/RoleBinding per backend.
  • kubernetes_deploy_targets, the output consumed by backend GitHub Actions.
  • Legacy Compute Engine VM resources only for names listed in enabled_legacy_backend_vms.
caution

Add a collaborator to gke_backends or gke_backend_overrides for GKE-only infrastructure. Add a collaborator to legacy_backend_vms only when you need to preserve a reusable VM definition, and add it to enabled_legacy_backend_vms only when Terraform should actively manage that VM.

Set enable_kubernetes_workloads = false in a GKE backend definition to retain its cloud resources but exclude it from Kubernetes namespaces, runtime RBAC, and kubernetes_deploy_targets. CA currently uses this setting. Change it to true when CA is ready for a complete GKE deployment.

Prerequisites​

  • Terraform at the backend's pinned .terraform-version (1.13.5).
  • Google Cloud SDK.
  • gke-gcloud-auth-plugin, installed with gcloud components install gke-gcloud-auth-plugin.
  • jq for output export and state verification commands.
  • GitHub CLI gh for updating GitHub Actions secrets from the command line.
  • Access to the target GCP project.
  • Access to the shared Terraform state bucket: gs://tidy-outlet-412020-ogrre-terraform-state.
  • A Terraform platform identity with the deployment and Terraform service-account permissions, including permission to create Kubernetes Roles and RoleBindings. Use a privileged human identity or dedicated infrastructure account; the GitHub deployment account's roles/container.developer is intentionally insufficient for RBAC creation.
  • A local terraform.tfvars file only when you need uncommitted local overrides.

Shared non-secret defaults live in variables.tf. Add long-lived public collaborators to the gke_backends map there so the repo remains the source of truth.

Example terraform.tfvars override for local testing:

gke_backend_overrides = {
"boots" = {}
}

*.tfvars and .env files are local operational files and should not be committed. Terraform state is stored remotely in GCS; do not commit local state backups or state migration files.

Authenticate​

Use your human Google account when it has the required infrastructure roles:

cd orphaned-wells-ui-server/deployment/terraform

gcloud auth login
gcloud config set project <PROJECT_ID>
unset GOOGLE_APPLICATION_CREDENTIALS GOOGLE_AUTHORIZED_USER_CREDENTIALS CLOUDSDK_AUTH_CREDENTIAL_FILE_OVERRIDE
gcloud auth application-default login

Alternatively, use a separately authorized Terraform platform account for this shell only. The GitHub deployment account is not the Terraform apply identity:

cd orphaned-wells-ui-server/deployment/terraform

export GOOGLE_APPLICATION_CREDENTIALS=/secure/path/ogrre-terraform-platform-service-key.json
gcloud auth activate-service-account --key-file "$GOOGLE_APPLICATION_CREDENTIALS"
gcloud config set project <PROJECT_ID>

Do not use the storage runtime or Document AI runtime service-account keys for Terraform. Those accounts are intentionally limited to backend runtime tasks.

Plan and apply​

Terraform PRs targeting upstream main, including fork PRs, run Deployment checks: formatting and terraform init -backend=false / terraform validate. These checks receive no cloud credentials, OIDC permission, or remote state. They do not run a live plan or request an Environment approval. Normal PR reviews remain required. GitHub may separately require permission to start an ordinary fork CI run, according to the repository's Actions policy.

With backend repository variable ENABLE_TERRAFORM_CI=true, the upstream staging workflow reconciles shared infrastructure after merge or a manual run on main:

ResultNext action
Tracked Terraform inputs and state match the last successful reconciliationSkip planning and applying; deployment reads live targets.
A fresh plan has no changesAutomatically verify the saved plan, current inputs/state, and outputs; record readiness without applying.
A fresh plan has changes, including output-only changesWait for terraform-apply approval, then apply the exact saved plan.
Plan, verification, or apply fails, or approval is rejectedBlock staging deployment; fix the issue and start a new full staging run.

No-change detection uses Terraform's detailed exit code, not a search for text or a count of changed resources. A change to kubernetes_deploy_targets can require approval even when no cloud resources change. Automatic completion, apply, and deployments share a concurrency group per workspace. A superseded Terraform revision or changed state prevents stale no-change completion.

Syncing a fork's main skips staging jobs, including infrastructure readiness. Merges into collaborator branches do not plan or apply Terraform; their deploy workflows require current upstream infrastructure readiness. Application promotion to those environments still follows their own branch workflows.

Set up or upgrade CI​

Bootstrap is an operator-run script, outside GitHub Actions. From the backend repository root, set the existing project, workspace, deployment account, and dedicated CI bucket, then run bash deployment/ci/bootstrap_terraform_ci.sh. It creates/configures WIF, separate plan/apply accounts, and private plan storage; it does not apply application Terraform or create a workspace. Reuse the existing bucket on reruns (ogrre-terraform-ci for the current installation).

Configure only the terraform-apply GitHub Environment with required reviewers, deployment branch main, and administrator bypass disabled. Its Prevent self-review setting is independent of normal PR approval rules. The former terraform-plan workflow and Environment are retired.

Existing installations must rerun the updated bootstrap before merging the automatic no-change update. The plan account needs permission to replace exactly status/<workspace>.json in the CI bucket, in addition to its existing reads and state-lock access. It receives no Terraform state or infrastructure writes. No new variables, keys, or Environments are required for this upgrade.

Follow the backend Terraform CI setup and rollout guide for initial setup, the no-change upgrade, and retiring the old fork-plan workflow. Bootstrap prints repository variables, not secrets. Cloud CI is disabled unless ENABLE_TERRAFORM_CI=true; static PR checks run independently. Phase one keeps DEPLOYMENT_SERVICE_KEY_JSON for deployment; Terraform CI uses WIF.

Review a live plan​

In upstream Actions → Deploy Staging Server to GKE, open the run. The Summary contains the Terraform plan section with its commit and workspace. For full output, open infrastructure / plan → Generate Terraform plan. Review resource and output changes for the shared ogrre workspace. If apply is waiting, use Review deployments → terraform-apply → Approve and deploy or Reject. A verified no-change plan has no approval request.

To run a fresh reconciliation, including after a failed or stale attempt:

gh workflow run deploy-k8s-staging.yml \
--repo CATALOG-Historic-Records/orphaned-wells-ui-server --ref main \
-f force_terraform_plan=true

Forcing a plan does not force apply or approval. This dispatch can deploy the backend image after reconciliation. A readiness record with status: verified means no-change completion; status: applied means a successful apply.

Manual plan​

Manual planning and applying remain supported alongside CI. Both use the same shared remote workspace. A staging workflow applies this shared Terraform configuration; it does not isolate infrastructure changes to staging.

Continue in orphaned-wells-ui-server/deployment/terraform after authenticating as described above. Use the pinned Terraform CLI and the Kubernetes authentication plugin:

gcloud components install gke-gcloud-auth-plugin
terraform version
unset TF_WORKSPACE
terraform init -input=false -lockfile=readonly
terraform workspace select ogrre
terraform workspace show
terraform fmt -check -recursive
terraform validate
terraform plan -input=false -lock-timeout=5m

Confirm the CLI matches .terraform-version and the workspace is ogrre. A speculative plan can run with CI enabled; it may wait for the state lock. Keep locking enabled. Local terraform.tfvars files load automatically, while CI uses committed defaults. Incorporate required production overrides into reviewed shared configuration so local and CI plans agree.

Manual apply​

Use the reviewed infrastructure configuration from current main and coordinate with other operators. If CI is enabled, pause the deployment entry points and wait for all active or approval-waiting runs to finish. Leave ENABLE_TERRAFORM_CI=true while pausing workflows, so deployment cannot fall back to stale secret-based targets. The backend guide provides the exact pause and resume commands.

Save and review the plan outside the repository:

OGRRE_PLAN_DIR="$(mktemp -d)"
chmod 700 "$OGRRE_PLAN_DIR"
terraform plan -input=false -lock-timeout=5m -out="$OGRRE_PLAN_DIR/manual.tfplan"
terraform show -no-color "$OGRRE_PLAN_DIR/manual.tfplan"

After reviewing the changes, apply that exact plan:

terraform apply -input=false -lock-timeout=5m "$OGRRE_PLAN_DIR/manual.tfplan"
terraform output -json kubernetes_deploy_targets | jq .
rm -f "$OGRRE_PLAN_DIR/manual.tfplan"
rmdir "$OGRRE_PLAN_DIR"

A saved-plan apply executes without another approval prompt. Keep the plan private; it can contain sensitive values. Generate a new plan if the state or desired configuration changed during review.

If CI is enabled, resume the staging workflow first and dispatch it with force_terraform_plan=true. Review the new plan; no-change completion restores readiness automatically when local inputs match main. Approve apply only if the plan contains changes. Resume other deployment workflows after that run succeeds; never edit the readiness record by hand. If CI is disabled, refresh the fallback target secret before backend deployment using the next section's commands.

Existing namespace migration​

Terraform now owns OGRRE namespaces and runtime Kubernetes RBAC. Before the first apply, import every namespace that already exists in the cluster. For example:

terraform import 'kubernetes_namespace_v1.backend["staging"]' uow-staging

Repeat for each existing collaborator namespace, then review the plan. It should create the ServiceAccounts, Job-dispatch Role, and RoleBinding without requiring a GitHub Actions RBAC bootstrap. New collaborators are created by a normal Terraform apply after being added to gke_backends or gke_backend_overrides.

To exclude GKE resources from a plan:

terraform plan -var='enable_gke=false'

Export GKE deployment targets​

Enabled CI selects the existing workspace and reads kubernetes_deploy_targets live. It does not use K8S_DEPLOY_TARGETS. Deployments stop when current main infrastructure has not been successfully reconciled; later backend-only commits cannot bypass that check. Collaborator application promotion still follows its existing branch workflow.

Only while ENABLE_TERRAFORM_CI is disabled, export the target map and store it as the fallback backend repository secret:

env -u TF_WORKSPACE terraform workspace select ogrre
gh auth login
gh secret set K8S_DEPLOY_TARGETS \
--repo CATALOG-Historic-Records/orphaned-wells-ui-server \
--body "$(terraform output -json kubernetes_deploy_targets | jq -c .)"

Refresh this fallback after target changes while automation is disabled. Once enabled and verified across all active collaborator branches, the secret can be removed. Keep the deployment credential for phase one.

Keep the backend repository deployment docs and these frontend docs in sync when changing Kubernetes or Terraform behavior. The source of truth lives in orphaned-wells-ui-server, but this docs site is the operator-facing deployment guide.

Upload buckets​

Each GKE backend gets a Cloud Storage upload bucket. The default bucket name is <collaborator>_uploads; use upload_bucket_name only for existing exceptions or explicit custom names. Staging uses uploaded_documents_v0.

Terraform keys bucket resources by bucket name, so collaborators can intentionally share a bucket without creating duplicate resources. Buckets use force_destroy=false and prevent_destroy=true to protect uploaded files.

GKE resource sizing​

Collaborator APIs target 2 replicas with 1 CPU and 4 GiB memory per pod. Staging targets 1 API replica with a 500m CPU / 1 GiB request and a 1 CPU / 2 GiB limit, using two Uvicorn workers. Processing-worker resources are configured separately.

Changing these values in Terraform does not update running Kubernetes pods by itself. Approve the Terraform apply, then deploy the affected backend environments. Refresh K8S_DEPLOY_TARGETS only for the disabled rollout fallback.

For an otherwise up-to-date workspace, the reduction changes only the kubernetes_deploy_targets output. Staging API cpu_request goes from 1 to 500m, memory_request goes from 4Gi to 1Gi, and memory_limit goes from 4Gi to 2Gi; cpu_limit stays at 1, production targets stay unchanged, and worker settings stay unchanged. It requires no infrastructure resources to be added, changed, or destroyed. Review any other plan changes separately. Explicit target values take precedence over workflow defaults. Applying these outputs does not resize live pods; a backend deployment activates the target for that environment.

Follow the staging validation and rollback steps before considering production changes. Validate production behavior before considering further API reductions. Use gke_backend_overrides for environment-specific adjustments or rollback.

Batch-worker sizing​

GCS-backed batch processing now runs in a short-lived Kubernetes Job rather than in a FastAPI background task. Terraform exports, but does not directly create, the Job settings below through kubernetes_deploy_targets:

  • api_uvicorn_workers
  • processing_job_cpu_request and processing_job_memory_request
  • processing_job_cpu_limit and processing_job_memory_limit
  • processing_job_ephemeral_storage
  • processing_job_active_deadline_seconds
  • processing_job_ttl_seconds_after_finished
  • processing_job_max_active

The worker defaults to 1850m CPU and 12 GiB memory for collaborator environments; staging uses 1 CPU and 6 GiB. One batch worker may be active per environment initially, and failed Jobs are not automatically retried. These settings are generated by backend GitHub Actions and should not be copied into the collaborator runtime-env secrets.

Single-file/ZIP processing, record-image uploads, imports, rotation, and exports still use API resources. Include these workflows in staging load checks before choosing a smaller production API allocation.

Legacy Compute Engine VMs​

Legacy VM definitions are kept in legacy_backend_vms so they can be re-enabled later without reconstructing their machine, disk, image, or zone settings. They are disabled by default because enabled_legacy_backend_vms defaults to an empty set.

To re-enable a legacy VM, add its name to enabled_legacy_backend_vms:

enabled_legacy_backend_vms = ["isgs"]

To disable it again without deleting the definition, remove the name from enabled_legacy_backend_vms. If Terraform state still contains old module resources and you want Terraform to stop managing them rather than destroy them, remove only those bindings from state with terraform state rm.

Stopped legacy VMs, detached disks, static IPs, or automatic snapshots that are no longer in Terraform state require explicit Google Cloud cleanup. Confirm rollback is no longer needed before deleting them.

Workspaces and state​

This Terraform setup uses a shared GCS backend configured in backend.tf:

gs://tidy-outlet-412020-ogrre-terraform-state/orphaned-wells-ui-server

The backend bucket is bootstrap infrastructure. It is configured by backend.tf and created or updated out-of-band with scripts/bootstrap_terraform_state_bucket.sh; it is not itself a Terraform-managed resource in this root module. The bucket should have uniform bucket-level access, public access prevention, and object versioning enabled.

Use the shared ogrre workspace for normal infrastructure work:

terraform init
env -u TF_WORKSPACE terraform workspace select ogrre
terraform workspace show
terraform plan

If this is your first time using the backend, confirm the remote workspace is visible:

terraform workspace list

You can also verify that Terraform is reading remote state:

terraform state pull > /private/tmp/ogrre-remote.tfstate
jq -r '.resources | length' /private/tmp/ogrre-remote.tfstate
gcloud storage ls -r gs://tidy-outlet-412020-ogrre-terraform-state/orphaned-wells-ui-server
caution

Treat the shared remote state as production infrastructure metadata. Do not run terraform state push, terraform state rm, terraform state mv, terraform import, terraform workspace new, or terraform workspace delete against the shared backend unless you are intentionally doing state maintenance.

warning

If terraform init asks whether to migrate all local workspaces to gcs, do not answer yes unless you intentionally want to overwrite or copy every local workspace into the shared backend. For selective migration, back up the local state and push only the known-good state file.

Workspaces are still available with the GCS backend, but use a separate workspace only when you intentionally want an isolated remote state:

terraform workspace new <workspace_name>
terraform workspace select <workspace_name>

Do not use ad hoc workspaces for shared collaborator infrastructure. A plan from an empty or incorrect workspace may propose recreating existing cloud resources.

State recovery and imports​

Imports are normally unnecessary now that the shared ogrre state is remote. If you are recovering or bootstrapping a state file, use the backend repository import scripts and review the summary before applying:

bash scripts/import_existing_infrastructure.sh --target-workspace ogrre --dry-run
bash scripts/import_existing_infrastructure.sh --target-workspace ogrre

The comprehensive importer reads this repo's Terraform naming conventions and attempts to import the configured firewalls, GKE resources, upload buckets, DNS records, static IPs, and enabled legacy VM resources. Missing resources are reported and do not stop the script unless --strict is passed.

Primary DNS ownership​

Primary backend DNS records are managed by root-level GKE resources:

google_dns_record_set.gke_backend_primary["<collaborator>"]

Legacy VM modules can still manage VMs and VM IPs, but they do not create the primary backend DNS record for a collaborator whose GKE backend has create_primary_dns_record=true. Existing DNS records were moved into the GKE resource addresses with Terraform moved blocks, so a normal plan should not propose duplicate DNS records for existing <collaborator>-server.uow-carbon.org names.

Targeted plans​

Use -target only for isolated Terraform work:

terraform plan -target='module.backend_vms["staging"]'
terraform apply -target='module.backend_vms["staging"]'

Targeting bypasses part of Terraform's normal dependency planning, so use a full plan afterward when possible.