Terraform Infrastructure
Terraform lives in orphaned-wells-ui-server/deployment/terraform.
It manages the current GKE deployment infrastructure, upload buckets, DNS, namespace-level Kubernetes RBAC, and deployment target outputs. Legacy VM definitions remain in Terraform for future re-enablement, but Terraform manages legacy VM modules only for names listed in enabled_legacy_backend_vms.
What Terraform manages
- Shared GKE Autopilot cluster.
- One Cloud Storage upload bucket per unique GKE backend bucket name.
- One global static IP per backend collaborator for Kubernetes Ingress.
- Optional
<collaborator>-k8s-server.uow-carbon.orgtest DNS records. - Primary backend DNS records, such as
staging-server.uow-carbon.org, pointing to the GKE static IP. - One Kubernetes namespace plus the
backend-api/processing-workerServiceAccounts and worker-dispatch Role/RoleBinding per backend. kubernetes_deploy_targets, the output consumed by backend GitHub Actions.- Legacy Compute Engine VM resources only for names listed in
enabled_legacy_backend_vms.
Add a collaborator to gke_backends or gke_backend_overrides for GKE-only infrastructure. Add a collaborator to legacy_backend_vms only when you need to preserve a reusable VM definition, and add it to enabled_legacy_backend_vms only when Terraform should actively manage that VM.
Set enable_kubernetes_workloads = false in a GKE backend definition to retain
its cloud resources but exclude it from Kubernetes namespaces, runtime RBAC,
and kubernetes_deploy_targets. CA currently uses this setting. Change it to
true when CA is ready for a complete GKE deployment.
Prerequisites
- Terraform at the backend's pinned
.terraform-version(1.13.5). - Google Cloud SDK.
gke-gcloud-auth-plugin, installed withgcloud components install gke-gcloud-auth-plugin.jqfor output export and state verification commands.- GitHub CLI
ghfor updating GitHub Actions secrets from the command line. - Access to the target GCP project.
- Access to the shared Terraform state bucket:
gs://tidy-outlet-412020-ogrre-terraform-state. - A Terraform platform identity with the deployment and Terraform service-account permissions, including permission to create Kubernetes Roles and RoleBindings. Use a privileged human identity or dedicated infrastructure account; the GitHub deployment account's
roles/container.developeris intentionally insufficient for RBAC creation. - A local
terraform.tfvarsfile only when you need uncommitted local overrides.
Shared non-secret defaults live in variables.tf. Add long-lived public collaborators to the gke_backends map there so the repo remains the source of truth.
Example terraform.tfvars override for local testing:
gke_backend_overrides = {
"boots" = {}
}
*.tfvars and .env files are local operational files and should not be committed. Terraform state is stored remotely in GCS; do not commit local state backups or state migration files.
Authenticate
Use your human Google account when it has the required infrastructure roles:
cd orphaned-wells-ui-server/deployment/terraform
gcloud auth login
gcloud config set project <PROJECT_ID>
unset GOOGLE_APPLICATION_CREDENTIALS GOOGLE_AUTHORIZED_USER_CREDENTIALS CLOUDSDK_AUTH_CREDENTIAL_FILE_OVERRIDE
gcloud auth application-default login
Alternatively, use a separately authorized Terraform platform account for this shell only. The GitHub deployment account is not the Terraform apply identity:
cd orphaned-wells-ui-server/deployment/terraform
export GOOGLE_APPLICATION_CREDENTIALS=/secure/path/ogrre-terraform-platform-service-key.json
gcloud auth activate-service-account --key-file "$GOOGLE_APPLICATION_CREDENTIALS"
gcloud config set project <PROJECT_ID>
Do not use the storage runtime or Document AI runtime service-account keys for Terraform. Those accounts are intentionally limited to backend runtime tasks.
Plan and apply
Terraform PRs targeting upstream main, including fork PRs, run Deployment
checks: formatting and terraform init -backend=false / terraform validate.
These checks receive no cloud credentials, OIDC permission, or remote state.
They do not run a live plan or request an Environment approval. Normal PR
reviews remain required. GitHub may separately require permission to start an
ordinary fork CI run, according to the repository's Actions policy.
With backend repository variable ENABLE_TERRAFORM_CI=true, the upstream staging
workflow reconciles shared infrastructure after merge or a manual run on main:
| Result | Next action |
|---|---|
| Tracked Terraform inputs and state match the last successful reconciliation | Skip planning and applying; deployment reads live targets. |
| A fresh plan has no changes | Automatically verify the saved plan, current inputs/state, and outputs; record readiness without applying. |
| A fresh plan has changes, including output-only changes | Wait for terraform-apply approval, then apply the exact saved plan. |
| Plan, verification, or apply fails, or approval is rejected | Block staging deployment; fix the issue and start a new full staging run. |
No-change detection uses Terraform's detailed exit code, not a search for text
or a count of changed resources. A change to kubernetes_deploy_targets can
require approval even when no cloud resources change. Automatic completion,
apply, and deployments share a concurrency group per workspace. A superseded
Terraform revision or changed state prevents stale no-change completion.
Syncing a fork's main skips staging jobs, including infrastructure readiness.
Merges into collaborator branches do not plan or apply Terraform; their deploy
workflows require current upstream infrastructure readiness. Application
promotion to those environments still follows their own branch workflows.
Set up or upgrade CI
Bootstrap is an operator-run script, outside GitHub Actions. From the backend
repository root, set the existing project, workspace, deployment account, and
dedicated CI bucket, then run bash deployment/ci/bootstrap_terraform_ci.sh.
It creates/configures WIF, separate plan/apply accounts, and private plan storage;
it does not apply application Terraform or create a workspace. Reuse the existing
bucket on reruns (ogrre-terraform-ci for the current installation).
Configure only the terraform-apply GitHub Environment with required
reviewers, deployment branch main, and administrator bypass disabled. Its
Prevent self-review setting is independent of normal PR approval rules.
The former terraform-plan workflow and Environment are retired.
Existing installations must rerun the updated bootstrap before merging the
automatic no-change update. The plan account needs permission to replace exactly
status/<workspace>.json in the CI bucket, in addition to its existing reads and
state-lock access. It receives no Terraform state or infrastructure writes. No
new variables, keys, or Environments are required for this upgrade.
Follow the backend Terraform CI setup and rollout guide
for initial setup, the no-change upgrade, and retiring the old fork-plan workflow.
Bootstrap prints repository variables, not secrets. Cloud CI is disabled
unless ENABLE_TERRAFORM_CI=true; static PR checks run independently. Phase one
keeps DEPLOYMENT_SERVICE_KEY_JSON for deployment; Terraform CI uses WIF.
Review a live plan
In upstream Actions → Deploy Staging Server to GKE, open the run. The
Summary contains the Terraform plan section with its commit and workspace.
For full output, open infrastructure / plan → Generate Terraform plan.
Review resource and output changes for the shared ogrre workspace. If apply
is waiting, use Review deployments → terraform-apply → Approve and deploy
or Reject. A verified no-change plan has no approval request.
To run a fresh reconciliation, including after a failed or stale attempt:
gh workflow run deploy-k8s-staging.yml \
--repo CATALOG-Historic-Records/orphaned-wells-ui-server --ref main \
-f force_terraform_plan=true
Forcing a plan does not force apply or approval. This dispatch can deploy the
backend image after reconciliation. A readiness record with status: verified
means no-change completion; status: applied means a successful apply.
Manual plan
Manual planning and applying remain supported alongside CI. Both use the same shared remote workspace. A staging workflow applies this shared Terraform configuration; it does not isolate infrastructure changes to staging.
Continue in orphaned-wells-ui-server/deployment/terraform after authenticating
as described above. Use the pinned Terraform CLI and the Kubernetes
authentication plugin:
gcloud components install gke-gcloud-auth-plugin
terraform version
unset TF_WORKSPACE
terraform init -input=false -lockfile=readonly
terraform workspace select ogrre
terraform workspace show
terraform fmt -check -recursive
terraform validate
terraform plan -input=false -lock-timeout=5m
Confirm the CLI matches .terraform-version and the workspace is ogrre.
A speculative plan can run with CI enabled; it may wait for the state lock.
Keep locking enabled. Local terraform.tfvars files load automatically, while
CI uses committed defaults. Incorporate required production overrides into
reviewed shared configuration so local and CI plans agree.
Manual apply
Use the reviewed infrastructure configuration from current main and coordinate
with other operators. If CI is enabled, pause the deployment entry points and
wait for all active or approval-waiting runs to finish. Leave
ENABLE_TERRAFORM_CI=true while pausing workflows, so deployment cannot fall
back to stale secret-based targets. The backend guide provides the exact
pause and resume commands.
Save and review the plan outside the repository:
OGRRE_PLAN_DIR="$(mktemp -d)"
chmod 700 "$OGRRE_PLAN_DIR"
terraform plan -input=false -lock-timeout=5m -out="$OGRRE_PLAN_DIR/manual.tfplan"
terraform show -no-color "$OGRRE_PLAN_DIR/manual.tfplan"
After reviewing the changes, apply that exact plan:
terraform apply -input=false -lock-timeout=5m "$OGRRE_PLAN_DIR/manual.tfplan"
terraform output -json kubernetes_deploy_targets | jq .
rm -f "$OGRRE_PLAN_DIR/manual.tfplan"
rmdir "$OGRRE_PLAN_DIR"
A saved-plan apply executes without another approval prompt. Keep the plan private; it can contain sensitive values. Generate a new plan if the state or desired configuration changed during review.
If CI is enabled, resume the staging workflow first and dispatch it with
force_terraform_plan=true. Review the new plan; no-change completion restores
readiness automatically when local inputs match main. Approve apply only if
the plan contains changes. Resume other deployment
workflows after that run succeeds; never edit the readiness record by hand.
If CI is disabled, refresh the fallback target secret before backend deployment
using the next section's commands.
Existing namespace migration
Terraform now owns OGRRE namespaces and runtime Kubernetes RBAC. Before the first apply, import every namespace that already exists in the cluster. For example:
terraform import 'kubernetes_namespace_v1.backend["staging"]' uow-staging
Repeat for each existing collaborator namespace, then review the plan. It
should create the ServiceAccounts, Job-dispatch Role, and RoleBinding without
requiring a GitHub Actions RBAC bootstrap. New collaborators are created by a
normal Terraform apply after being added to gke_backends or
gke_backend_overrides.
To exclude GKE resources from a plan:
terraform plan -var='enable_gke=false'
Export GKE deployment targets
Enabled CI selects the existing workspace and reads kubernetes_deploy_targets
live. It does not use K8S_DEPLOY_TARGETS. Deployments stop when current main
infrastructure has not been successfully reconciled; later backend-only commits
cannot bypass that check. Collaborator application promotion still follows its
existing branch workflow.
Only while ENABLE_TERRAFORM_CI is disabled, export the target map and store it
as the fallback backend repository secret:
env -u TF_WORKSPACE terraform workspace select ogrre
gh auth login
gh secret set K8S_DEPLOY_TARGETS \
--repo CATALOG-Historic-Records/orphaned-wells-ui-server \
--body "$(terraform output -json kubernetes_deploy_targets | jq -c .)"
Refresh this fallback after target changes while automation is disabled. Once enabled and verified across all active collaborator branches, the secret can be removed. Keep the deployment credential for phase one.
Keep the backend repository deployment docs and these frontend docs in sync when changing Kubernetes or Terraform behavior. The source of truth lives in orphaned-wells-ui-server, but this docs site is the operator-facing deployment guide.
Upload buckets
Each GKE backend gets a Cloud Storage upload bucket. The default bucket name is <collaborator>_uploads; use upload_bucket_name only for existing exceptions or explicit custom names. Staging uses uploaded_documents_v0.
Terraform keys bucket resources by bucket name, so collaborators can intentionally share a bucket without creating duplicate resources. Buckets use force_destroy=false and prevent_destroy=true to protect uploaded files.
GKE resource sizing
Collaborator APIs target 2 replicas with 1 CPU and 4 GiB memory per pod. Staging targets 1 API replica with a 500m CPU / 1 GiB request and a 1 CPU / 2 GiB limit, using two Uvicorn workers. Processing-worker resources are configured separately.
Changing these values in Terraform does not update running Kubernetes pods by itself. Approve the Terraform apply, then deploy the affected backend environments. Refresh K8S_DEPLOY_TARGETS only for the disabled rollout fallback.
For an otherwise up-to-date workspace, the reduction changes only the
kubernetes_deploy_targets output. Staging API cpu_request goes from 1 to
500m, memory_request goes from 4Gi to 1Gi, and memory_limit goes
from 4Gi to 2Gi; cpu_limit stays at 1, production targets stay
unchanged, and worker settings stay unchanged. It requires no infrastructure
resources to be added, changed, or destroyed. Review any other plan changes separately.
Explicit target values take precedence over workflow defaults. Applying
these outputs does not resize live pods; a backend deployment
activates the target for that environment.
Follow the staging validation and rollback steps
before considering production changes. Validate production behavior before
considering further API reductions. Use gke_backend_overrides for
environment-specific adjustments or rollback.
Batch-worker sizing
GCS-backed batch processing now runs in a short-lived Kubernetes Job rather
than in a FastAPI background task. Terraform exports, but does not directly
create, the Job settings below through kubernetes_deploy_targets:
api_uvicorn_workersprocessing_job_cpu_requestandprocessing_job_memory_requestprocessing_job_cpu_limitandprocessing_job_memory_limitprocessing_job_ephemeral_storageprocessing_job_active_deadline_secondsprocessing_job_ttl_seconds_after_finishedprocessing_job_max_active
The worker defaults to 1850m CPU and 12 GiB memory for collaborator environments; staging uses 1 CPU and 6 GiB. One batch worker may be active per environment initially, and failed Jobs are not automatically retried. These settings are generated by backend GitHub Actions and should not be copied into the collaborator runtime-env secrets.
Single-file/ZIP processing, record-image uploads, imports, rotation, and exports still use API resources. Include these workflows in staging load checks before choosing a smaller production API allocation.
Legacy Compute Engine VMs
Legacy VM definitions are kept in legacy_backend_vms so they can be re-enabled later without reconstructing their machine, disk, image, or zone settings. They are disabled by default because enabled_legacy_backend_vms defaults to an empty set.
To re-enable a legacy VM, add its name to enabled_legacy_backend_vms:
enabled_legacy_backend_vms = ["isgs"]
To disable it again without deleting the definition, remove the name from enabled_legacy_backend_vms. If Terraform state still contains old module resources and you want Terraform to stop managing them rather than destroy them, remove only those bindings from state with terraform state rm.
Stopped legacy VMs, detached disks, static IPs, or automatic snapshots that are no longer in Terraform state require explicit Google Cloud cleanup. Confirm rollback is no longer needed before deleting them.
Workspaces and state
This Terraform setup uses a shared GCS backend configured in backend.tf:
gs://tidy-outlet-412020-ogrre-terraform-state/orphaned-wells-ui-server
The backend bucket is bootstrap infrastructure. It is configured by backend.tf and created or updated out-of-band with scripts/bootstrap_terraform_state_bucket.sh; it is not itself a Terraform-managed resource in this root module. The bucket should have uniform bucket-level access, public access prevention, and object versioning enabled.
Use the shared ogrre workspace for normal infrastructure work:
terraform init
env -u TF_WORKSPACE terraform workspace select ogrre
terraform workspace show
terraform plan
If this is your first time using the backend, confirm the remote workspace is visible:
terraform workspace list
You can also verify that Terraform is reading remote state:
terraform state pull > /private/tmp/ogrre-remote.tfstate
jq -r '.resources | length' /private/tmp/ogrre-remote.tfstate
gcloud storage ls -r gs://tidy-outlet-412020-ogrre-terraform-state/orphaned-wells-ui-server
Treat the shared remote state as production infrastructure metadata. Do not run terraform state push, terraform state rm, terraform state mv, terraform import, terraform workspace new, or terraform workspace delete against the shared backend unless you are intentionally doing state maintenance.
If terraform init asks whether to migrate all local workspaces to gcs, do not answer yes unless you intentionally want to overwrite or copy every local workspace into the shared backend. For selective migration, back up the local state and push only the known-good state file.
Workspaces are still available with the GCS backend, but use a separate workspace only when you intentionally want an isolated remote state:
terraform workspace new <workspace_name>
terraform workspace select <workspace_name>
Do not use ad hoc workspaces for shared collaborator infrastructure. A plan from an empty or incorrect workspace may propose recreating existing cloud resources.
State recovery and imports
Imports are normally unnecessary now that the shared ogrre state is remote. If you are recovering or bootstrapping a state file, use the backend repository import scripts and review the summary before applying:
bash scripts/import_existing_infrastructure.sh --target-workspace ogrre --dry-run
bash scripts/import_existing_infrastructure.sh --target-workspace ogrre
The comprehensive importer reads this repo's Terraform naming conventions and attempts to import the configured firewalls, GKE resources, upload buckets, DNS records, static IPs, and enabled legacy VM resources. Missing resources are reported and do not stop the script unless --strict is passed.
Primary DNS ownership
Primary backend DNS records are managed by root-level GKE resources:
google_dns_record_set.gke_backend_primary["<collaborator>"]
Legacy VM modules can still manage VMs and VM IPs, but they do not create the primary backend DNS record for a collaborator whose GKE backend has create_primary_dns_record=true. Existing DNS records were moved into the GKE resource addresses with Terraform moved blocks, so a normal plan should not propose duplicate DNS records for existing <collaborator>-server.uow-carbon.org names.
Targeted plans
Use -target only for isolated Terraform work:
terraform plan -target='module.backend_vms["staging"]'
terraform apply -target='module.backend_vms["staging"]'
Targeting bypasses part of Terraform's normal dependency planning, so use a full plan afterward when possible.