Skip to main content

Backend on GKE

The current backend deployment path runs each collaborator as a Kubernetes workload in the shared GKE cluster.

Terraform owns the cloud infrastructure, namespaces, and long-lived Kubernetes RBAC. GitHub Actions renders orphaned-wells-ui-server/deployment/kubernetes/backend.yaml with collaborator-specific values and applies the workload-only manifest to GKE.

Architecture​

Terraform creates:

  • Shared GKE Autopilot cluster.
  • One Cloud Storage upload bucket per unique backend bucket name.
  • One global static IP per collaborator backend.
  • Optional test DNS records at <collaborator>-k8s-server.uow-carbon.org.
  • Primary backend DNS records at <collaborator>-server.uow-carbon.org.
  • Each backend namespace, backend-api and processing-worker ServiceAccounts, and the API Job-dispatch Role/RoleBinding.
  • kubernetes_deploy_targets, the JSON map consumed by GitHub Actions.

Terraform creates one namespace per Kubernetes-workload-enabled collaborator, using the uow-<collaborator> naming pattern. Terraform owns these resources:

  • ServiceAccount/backend-api and namespace-scoped RBAC for dispatching document-processing Jobs
  • ServiceAccount/processing-worker for short-lived document-processing Pods
  • namespace labels

GitHub Actions creates these workload resources in that namespace:

  • Deployment/backend
  • Service/backend
  • BackendConfig/backend-config
  • ManagedCertificate/backend-cert
  • FrontendConfig/backend-frontend-config
  • Ingress/backend
  • Secret/backend-runtime-env
  • Secret/backend-runtime-files
  • Secret/dockerhub-pull

An entry can retain its cloud infrastructure while remaining unavailable for Kubernetes deployment by setting enable_kubernetes_workloads = false in its Terraform backend definition. Terraform then omits the namespace, runtime ServiceAccounts/RBAC, and kubernetes_deploy_targets entry. CA currently uses this setting; change it to true when CA is ready for a full GKE deployment.

Batch processing workers​

The API Deployment remains the only always-running backend workload. When a user finalizes a local directory upload or submits a GCS batch, the API stores durable job state in MongoDB and creates one Kubernetes Job. Kubernetes schedules a high-memory worker Pod from the same immutable backend image; it processes the batch, records progress and failures in MongoDB, then exits. The worker is not a public service and does not run FastAPI.

Local directory files now transfer directly from the browser to GCS and use the same worker path. Single-file and ZIP uploads still use API Pods. Keep API resource requests until staging measurements cover these remaining workflows. See Uploads and processing for diagrams, progress states, interrupted transfers, and retry behavior.

Directory upload storage configuration​

Apply Terraform's upload bucket CORS and lifecycle settings before deploying the paired frontend and backend changes. Review existing rules in the plan, since Terraform manages these bucket settings. CORS defaults to the custom frontend origins: https://uow-carbon.org for staging and https://<collaborator>.uow-carbon.org for collaborators. The staging bucket also allows http://localhost:3000 for native/Docker development and http://localhost:3001 for the isolated Docker E2E stack.

For other hosts or ports, override the full origin list per bucket, retaining the default origins you still use. For example, to also use port 3002:

upload_bucket_cors_origins = {
uploaded_documents_v0 = [
"https://uow-carbon.org",
"http://localhost:3000",
"http://localhost:3001",
"http://localhost:3002",
]
}

The backend serving the frontend must allow its browser origin in ALLOWED_ORIGINS. Configure this on the local backend for a local stack; only add localhost to a deployed backend if developers call that backend directly. Bucket CORS does not update backend settings. Use the browser's published host port, not the container's internal port; 127.0.0.1 and localhost are distinct origins. See Docker development.

The bucket rules govern GCS XML API requests. Directory uploads use JSON API resumable sessions, which receive the frontend origin at creation and have their own CORS handling. The browser uses GCS resumable PUT requests with Content-Range and reads Range responses. The existing storage runtime identity authorizes uploads and verifies object metadata. Document AI needs the same GCS input/output permissions as an existing batch workflow. Users do not need their own GCS credentials.

Staged originals under directory_uploads/ and temporary output under directory_upload_outputs/ expire through a 14-day bucket lifecycle rule. Sessions last seven days. Keep retention longer than the session window plus worker runtime. Record display images under uploads/, deleted-record audit files, and caller-owned GCS batch sources are outside these rules.

DIRECTORY_UPLOAD_MAX_BYTES controls the total original bytes per directory (default 20 GiB). The manifest is limited to 1,000 files. Existing Document AI file-size and processor limits still apply.

Capacity, failure recovery, and resource sizing​

PROCESSING_JOB_MAX_ACTIVE reserves worker capacity atomically in MongoDB. Additional submissions remain queued. The API runs job maintenance every 30 seconds to dispatch queued work and reconcile failed pods without browser polling. No additional Kubernetes role or deployment is needed.

At the smaller staging API size, validate a representative 500-file directory with large multipage PDFs while normal API traffic continues. Test transfer interruption, repeated finalization, simultaneous submissions, worker termination, and reopening the upload dialog after a refresh. Compare API and worker memory, CPU, throttling, temporary storage, OOM/restarts, duration, and API latency separately using their app.kubernetes.io/component=api and processor labels.

Staging now targets a 500m CPU / 1 GiB request and a 1 CPU / 2 GiB limit for the API, with one replica and two Uvicorn workers. Its processing workers retain 1 CPU / 6 GiB. Production targets two API replicas at 1 CPU / 4 GiB each, while its processing workers retain 1850m CPU / 12 GiB. Include single-file/ZIP uploads, record-image uploads, imports, rotation, and exports in load checks because those workflows still execute in the API. Use peak measurements and response times under load to validate production sizing; an idle usage snapshot is insufficient.

Approve Terraform apply and deploy the affected environments with the updated workflow. The sizing change updates staging API output from a 1 CPU / 4 GiB request and limit to a 500m CPU / 1 GiB request and 1 CPU / 2 GiB limit. Applying and exporting the outputs does not resize running pods. Verify admitted pod resources after deployment because Autopilot can adjust requests:

kubectl -n uow-staging get pods -l app.kubernetes.io/component=api \
-o 'custom-columns=NAME:.metadata.name,CPU_REQUEST:.spec.containers[*].resources.requests.cpu,MEMORY_REQUEST:.spec.containers[*].resources.requests.memory,CPU_LIMIT:.spec.containers[*].resources.limits.cpu,MEMORY_LIMIT:.spec.containers[*].resources.limits.memory'
kubectl -n uow-staging top pods -l app.kubernetes.io/component=api

If reducing production later, deploy production environments one at a time at 1 CPU / 4 GiB per API pod, preserving two replicas and all processing-worker settings. Change API requests and limits together and compare peak usage, throttling, and latency after each deployment.

To restore staging to its previous API allocation, merge this entry into the existing gke_backend_overrides map in terraform.tfvars, apply, refresh the deploy-target secret, and redeploy staging:

gke_backend_overrides = {
staging = {
cpu_request = "1"
memory_request = "4Gi"
cpu_limit = "1"
memory_limit = "4Gi"
}
}

Remove the override when resuming validation at the smaller staging target. To restore a production API to its previous memory allocation, set memory_request and memory_limit to 6Gi under that environment in gke_backend_overrides, then apply, refresh the secret, and redeploy it. The backend Kubernetes README contains the full rollout checks.

GitHub Actions does not create namespaces or RBAC. A Terraform apply creates them before the first worker-capable deployment. Existing namespaces must be imported into Terraform state once; the exact migration commands are in the backend Terraform README. Do not run the former GitHub Actions RBAC bootstrap.

Prerequisites​

Required local tools for manual operations:

  • gcloud
  • kubectl
  • terraform
  • jq
  • gh

Required Google APIs:

  • Kubernetes Engine API: container.googleapis.com
  • Compute Engine API: compute.googleapis.com
  • Cloud DNS API: dns.googleapis.com
  • Cloud Storage API: storage.googleapis.com

Other required access:

  • Ability to edit GitHub Secrets and Actions for the OGRRE GitHub project CATALOG-HISTORIC-RECORDS.
  • Permissions to manage GKE, Compute addresses, Cloud DNS, Cloud Storage buckets, OAuth, and project services in Google Cloud.
  • The OGRRE service accounts for storage runtime, Document AI runtime, GitHub deployment, and the Terraform platform identity. The last two may share an identity only when that broader privilege is intentional.
  • A Terraform platform identity with Kubernetes permission to create namespaces, ServiceAccounts, Roles, and RoleBindings. Keep this separate from the GitHub deployment identity when practical.

Prepare deployment targets​

With ENABLE_TERRAFORM_CI=true, merge Terraform changes to upstream main. Plans with changes require approval; verified no-change plans complete automatically without applying. Staging waits for successful reconciliation; updated deploy workflows read targets live from the existing remote workspace. See Terraform CI setup for WIF, Environment protection, and collaborator-branch rollout requirements. The existing deployment JSON key remains in use for phase one.

The commands below are for the disabled rollout fallback. Apply Terraform first:

cd orphaned-wells-ui-server/deployment/terraform
terraform init
terraform workspace select ogrre
terraform plan
terraform apply

Then export and store the target map. The gh command requires a GitHub account with permission to edit repository Actions secrets:

terraform workspace select ogrre
gh auth login
gh secret set K8S_DEPLOY_TARGETS \
--repo CATALOG-Historic-Records/orphaned-wells-ui-server \
--body "$(terraform output -json kubernetes_deploy_targets | jq -c .)"

The host field controls the Kubernetes Ingress host and the Google-managed certificate domain. The storage_bucket_name field controls the runtime STORAGE_BUCKET_NAME written to Secret/backend-runtime-env. After Terraform changes these values, redeploy the backend. Refresh K8S_DEPLOY_TARGETS first only when automation is disabled.

Resource target changes work the same way: Terraform updates kubernetes_deploy_targets, and the next backend deployment reads, renders, and applies those values to Kubernetes.

Required backend secrets​

The GKE deployment workflows require:

  • PROJECT_ID
  • DOCKERHUB_USERNAME
  • DOCKERHUB_ACCESS_TOKEN
  • DEPLOYMENT_SERVICE_KEY_JSON
  • STORAGE_SERVICE_KEY_JSON
  • DOCUMENT_AI_SERVICE_KEY_JSON
  • K8S_DEPLOY_TARGETS only for the disabled Terraform CI rollout fallback
  • One runtime env secret per currently enabled collaborator: STAGING_ENV, ISGS_ENV, NEWTS_ENV, OSAGE_ENV, and RRC_ENV. CA_ENV may remain configured for future use, but CA cannot deploy while enable_kubernetes_workloads = false.

DEPLOYMENT_SERVICE_KEY_JSON is used only by GitHub Actions to authenticate to Google Cloud and apply the deployment. STORAGE_SERVICE_KEY_JSON and DOCUMENT_AI_SERVICE_KEY_JSON are written into the Kubernetes backend-runtime-files secret and mounted into backend pods as runtime keys.

The environment-file secret should contain the same key/value pairs used by the VM .env file. The workflow overrides these Kubernetes-owned values:

  • ENVIRONMENT
  • BACKEND_URL
  • LOG_DIR
  • LOCAL_STORAGE_ROOT
  • LOCAL_STORAGE_URL_BASE
  • STORAGE_BUCKET_NAME
  • STORAGE_SERVICE_KEY
  • DOCUMENT_AI_SERVICE_KEY
  • GOOGLE_APPLICATION_CREDENTIALS

Keep COLLABORATOR in the runtime secret when the backend needs collaborator-specific processor metadata.

When adding a collaborator, update deploy-k8s-dispatch.yml with the new DEPLOY_ENV, <COLLABORATOR>_ENV secret, and runtime-env case branch before relying on a dedicated workflow. A dedicated deploy-k8s-<collaborator>.yml workflow calls the dispatch workflow and will fail until dispatch supports that collaborator.

Deploy with GitHub Actions​

The staging workflow builds and pushes the backend image, then deploys staging:

gh workflow run deploy-k8s-staging.yml \
--repo CATALOG-Historic-Records/orphaned-wells-ui-server \
--ref main

Collaborator workflows deploy an existing image tag through IMAGE_TAG: auto. On collaborator branch merge commits, auto resolves to the second parent commit, which is the main commit that staging built and tested. On fast-forward or non-merge commits, auto resolves to the current commit:

gh workflow run deploy-k8s-isgs.yml --repo CATALOG-Historic-Records/orphaned-wells-ui-server --ref isgs
gh workflow run deploy-k8s-newts.yml --repo CATALOG-Historic-Records/orphaned-wells-ui-server --ref newts
gh workflow run deploy-k8s-osage.yml --repo CATALOG-Historic-Records/orphaned-wells-ui-server --ref osage
gh workflow run deploy-k8s-rrc.yml --repo CATALOG-Historic-Records/orphaned-wells-ui-server --ref rrc

The exact set of deployable collaborators is controlled by both kubernetes_deploy_targets and .github/workflows/deploy-k8s-dispatch.yml. Confirm that the collaborator has an output entry, is present in workflow_dispatch.inputs.DEPLOY_ENV.options, and has a runtime-env secret mapping before dispatching it.

Merging source changes into a collaborator branch only changes the live backend when that branch's deployment workflow runs. If automatic deployments are enabled for the branch, the merge can trigger the rollout. Otherwise, manually dispatch the collaborator workflow after the merge.

Automatic deployments are controlled by repository variables:

ENABLE_GKE_DEPLOYMENTS=true

Or one collaborator at a time:

ENABLE_GKE_STAGING_DEPLOY=true
ENABLE_GKE_ISGS_DEPLOY=true
ENABLE_GKE_NEWTS_DEPLOY=true
ENABLE_GKE_OSAGE_DEPLOY=true

The current RRC workflow checks ENABLE_GKE_DEPLOYMENTS; add an ENABLE_GKE_RRC_DEPLOY workflow check before relying on an RRC-specific deploy variable.

Check rollout status​

Authenticate to the cluster:

gcloud container clusters get-credentials uow-backend-gke \
--region us-central1 \
--project <PROJECT_ID>

Check one backend:

kubectl -n uow-staging get deployment backend
kubectl -n uow-staging get pods -l app.kubernetes.io/name=orphaned-wells-ui-server -o wide
kubectl -n uow-staging get ingress backend
kubectl -n uow-staging get managedcertificate backend-cert
kubectl -n uow-staging get jobs -l app.kubernetes.io/component=processor

Read logs:

kubectl -n uow-staging logs deployment/backend --tail=200
kubectl -n uow-staging logs deployment/backend --tail=200 -f
kubectl -n uow-staging logs job/<worker-job-name> --tail=200

Restart a backend deployment:

kubectl -n uow-staging rollout restart deployment/backend
kubectl -n uow-staging rollout status deployment/backend --timeout=10m

Check resource usage:

kubectl top pods -n uow-staging
kubectl top nodes

Notes​

  • GKE replaces the legacy VM nginx/certbot path with GKE Ingress, ManagedCertificate, FrontendConfig, and BackendConfig.
  • Google-managed certificates require DNS to point at the GKE load balancer before they become active.
  • The rendered manifest in deployment/kubernetes/rendered/backend.yaml is an ephemeral artifact. Do not commit it.
  • The Kubernetes Deployment uses pod-local emptyDir volumes for /logs and /data. Production document storage should continue using the Terraform-managed Google Cloud Storage upload bucket.
  • Collaborator APIs target two replicas at 1 CPU and 4 GiB memory each; their processing workers retain 1850m CPU and 12 GiB. Staging targets one API replica with a 500m CPU / 1 GiB request and 1 CPU / 2 GiB limit; its workers retain 1 CPU and 6 GiB.
  • Worker Jobs have independent CPU, memory, scratch-storage, deadline, retention, and maximum-active-job settings. GitHub Actions supplies these generated settings from K8S_DEPLOY_TARGETS; do not place PROCESSING_JOB_* values in collaborator runtime-env secrets.