Backend on GKE
The current backend deployment path runs each collaborator as a Kubernetes workload in the shared GKE cluster.
Terraform owns the cloud infrastructure, namespaces, and long-lived Kubernetes RBAC. GitHub Actions renders orphaned-wells-ui-server/deployment/kubernetes/backend.yaml with collaborator-specific values and applies the workload-only manifest to GKE.
Architecture
Terraform creates:
- Shared GKE Autopilot cluster.
- One Cloud Storage upload bucket per unique backend bucket name.
- One global static IP per collaborator backend.
- Optional test DNS records at
<collaborator>-k8s-server.uow-carbon.org. - Primary backend DNS records at
<collaborator>-server.uow-carbon.org. - Each backend namespace,
backend-apiandprocessing-workerServiceAccounts, and the API Job-dispatch Role/RoleBinding. kubernetes_deploy_targets, the JSON map consumed by GitHub Actions.
Terraform creates one namespace per Kubernetes-workload-enabled collaborator,
using the uow-<collaborator> naming pattern. Terraform owns these resources:
ServiceAccount/backend-apiand namespace-scoped RBAC for dispatching document-processing JobsServiceAccount/processing-workerfor short-lived document-processing Pods- namespace labels
GitHub Actions creates these workload resources in that namespace:
Deployment/backendService/backendBackendConfig/backend-configManagedCertificate/backend-certFrontendConfig/backend-frontend-configIngress/backendSecret/backend-runtime-envSecret/backend-runtime-filesSecret/dockerhub-pull
An entry can retain its cloud infrastructure while remaining unavailable for
Kubernetes deployment by setting enable_kubernetes_workloads = false in its
Terraform backend definition. Terraform then omits the namespace, runtime
ServiceAccounts/RBAC, and kubernetes_deploy_targets entry. CA currently uses
this setting; change it to true when CA is ready for a full GKE deployment.
Batch processing workers
The API Deployment remains the only always-running backend workload. When a user finalizes a local directory upload or submits a GCS batch, the API stores durable job state in MongoDB and creates one Kubernetes Job. Kubernetes schedules a high-memory worker Pod from the same immutable backend image; it processes the batch, records progress and failures in MongoDB, then exits. The worker is not a public service and does not run FastAPI.
Local directory files now transfer directly from the browser to GCS and use the same worker path. Single-file and ZIP uploads still use API Pods. Keep API resource requests until staging measurements cover these remaining workflows. See Uploads and processing for diagrams, progress states, interrupted transfers, and retry behavior.
Directory upload storage configuration
Apply Terraform's upload bucket CORS and lifecycle settings before deploying the
paired frontend and backend changes. Review existing rules in the plan, since
Terraform manages these bucket settings. CORS defaults to the custom frontend
origins: https://uow-carbon.org for staging and
https://<collaborator>.uow-carbon.org for collaborators. The staging bucket
also allows http://localhost:3000 for native/Docker development and
http://localhost:3001 for the isolated Docker E2E stack.
For other hosts or ports, override the full origin list per bucket, retaining the default origins you still use. For example, to also use port 3002:
upload_bucket_cors_origins = {
uploaded_documents_v0 = [
"https://uow-carbon.org",
"http://localhost:3000",
"http://localhost:3001",
"http://localhost:3002",
]
}
The backend serving the frontend must allow its browser origin in
ALLOWED_ORIGINS. Configure this on the local backend for a local stack; only
add localhost to a deployed backend if developers call that backend directly.
Bucket CORS does not update backend settings. Use the browser's published host
port, not the container's internal port; 127.0.0.1 and localhost are distinct
origins. See Docker development.
The bucket rules govern GCS XML API requests. Directory uploads use JSON API
resumable sessions, which receive the frontend origin at creation and have
their own CORS handling.
The browser
uses GCS resumable PUT requests with Content-Range and reads Range responses.
The existing storage runtime identity authorizes uploads and verifies object
metadata. Document AI needs the same GCS input/output permissions as an existing
batch workflow. Users do not need their own GCS credentials.
Staged originals under directory_uploads/ and temporary output under
directory_upload_outputs/ expire through a 14-day bucket lifecycle rule.
Sessions last seven days. Keep retention longer than the session window plus
worker runtime. Record display images under uploads/, deleted-record audit
files, and caller-owned GCS batch sources are outside these rules.
DIRECTORY_UPLOAD_MAX_BYTES controls the total original bytes per directory
(default 20 GiB). The manifest is limited to 1,000 files. Existing Document AI
file-size and processor limits still apply.
Capacity, failure recovery, and resource sizing
PROCESSING_JOB_MAX_ACTIVE reserves worker capacity atomically in MongoDB.
Additional submissions remain queued. The API runs job maintenance every 30
seconds to dispatch queued work and reconcile failed pods without browser polling.
No additional Kubernetes role or deployment is needed.
At the smaller staging API size, validate a representative 500-file directory with
large multipage PDFs while normal API traffic continues. Test transfer interruption,
repeated finalization, simultaneous submissions, worker termination, and reopening
the upload dialog after a refresh. Compare API and worker memory, CPU, throttling,
temporary storage, OOM/restarts, duration, and API latency separately using their
app.kubernetes.io/component=api and processor labels.
Staging now targets a 500m CPU / 1 GiB request and a 1 CPU / 2 GiB limit for the API, with one replica and two Uvicorn workers. Its processing workers retain 1 CPU / 6 GiB. Production targets two API replicas at 1 CPU / 4 GiB each, while its processing workers retain 1850m CPU / 12 GiB. Include single-file/ZIP uploads, record-image uploads, imports, rotation, and exports in load checks because those workflows still execute in the API. Use peak measurements and response times under load to validate production sizing; an idle usage snapshot is insufficient.
Approve Terraform apply and deploy the affected environments with the updated workflow. The sizing change updates staging API output from a 1 CPU / 4 GiB request and limit to a 500m CPU / 1 GiB request and 1 CPU / 2 GiB limit. Applying and exporting the outputs does not resize running pods. Verify admitted pod resources after deployment because Autopilot can adjust requests:
kubectl -n uow-staging get pods -l app.kubernetes.io/component=api \
-o 'custom-columns=NAME:.metadata.name,CPU_REQUEST:.spec.containers[*].resources.requests.cpu,MEMORY_REQUEST:.spec.containers[*].resources.requests.memory,CPU_LIMIT:.spec.containers[*].resources.limits.cpu,MEMORY_LIMIT:.spec.containers[*].resources.limits.memory'
kubectl -n uow-staging top pods -l app.kubernetes.io/component=api
If reducing production later, deploy production environments one at a time at 1 CPU / 4 GiB per API pod, preserving two replicas and all processing-worker settings. Change API requests and limits together and compare peak usage, throttling, and latency after each deployment.
To restore staging to its previous API allocation, merge this entry into the existing
gke_backend_overrides map in terraform.tfvars, apply, refresh the deploy-target
secret, and redeploy staging:
gke_backend_overrides = {
staging = {
cpu_request = "1"
memory_request = "4Gi"
cpu_limit = "1"
memory_limit = "4Gi"
}
}
Remove the override when resuming validation at the smaller staging target. To restore a production
API to its previous memory allocation, set memory_request and memory_limit
to 6Gi under that environment in gke_backend_overrides, then apply, refresh
the secret, and redeploy it. The backend
Kubernetes README
contains the full rollout checks.
GitHub Actions does not create namespaces or RBAC. A Terraform apply creates them before the first worker-capable deployment. Existing namespaces must be imported into Terraform state once; the exact migration commands are in the backend Terraform README. Do not run the former GitHub Actions RBAC bootstrap.
Prerequisites
Required local tools for manual operations:
gcloudkubectlterraformjqgh
Required Google APIs:
- Kubernetes Engine API:
container.googleapis.com - Compute Engine API:
compute.googleapis.com - Cloud DNS API:
dns.googleapis.com - Cloud Storage API:
storage.googleapis.com
Other required access:
- Ability to edit GitHub Secrets and Actions for the OGRRE GitHub project
CATALOG-HISTORIC-RECORDS. - Permissions to manage GKE, Compute addresses, Cloud DNS, Cloud Storage buckets, OAuth, and project services in Google Cloud.
- The OGRRE service accounts for storage runtime, Document AI runtime, GitHub deployment, and the Terraform platform identity. The last two may share an identity only when that broader privilege is intentional.
- A Terraform platform identity with Kubernetes permission to create namespaces, ServiceAccounts, Roles, and RoleBindings. Keep this separate from the GitHub deployment identity when practical.
Prepare deployment targets
With ENABLE_TERRAFORM_CI=true, merge Terraform changes to upstream main.
Plans with changes require approval; verified no-change plans complete
automatically without applying. Staging waits for successful reconciliation; updated deploy workflows read
targets live from the existing remote workspace. See Terraform CI setup
for WIF, Environment protection, and collaborator-branch rollout requirements.
The existing deployment JSON key remains in use for phase one.
The commands below are for the disabled rollout fallback. Apply Terraform first:
cd orphaned-wells-ui-server/deployment/terraform
terraform init
terraform workspace select ogrre
terraform plan
terraform apply
Then export and store the target map. The gh command requires a GitHub account with permission to edit repository Actions secrets:
terraform workspace select ogrre
gh auth login
gh secret set K8S_DEPLOY_TARGETS \
--repo CATALOG-Historic-Records/orphaned-wells-ui-server \
--body "$(terraform output -json kubernetes_deploy_targets | jq -c .)"
The host field controls the Kubernetes Ingress host and the Google-managed certificate domain. The storage_bucket_name field controls the runtime STORAGE_BUCKET_NAME written to Secret/backend-runtime-env. After Terraform changes these values, redeploy the backend. Refresh K8S_DEPLOY_TARGETS first only when automation is disabled.
Resource target changes work the same way: Terraform updates kubernetes_deploy_targets, and the next backend deployment reads, renders, and applies those values to Kubernetes.
Required backend secrets
The GKE deployment workflows require:
PROJECT_IDDOCKERHUB_USERNAMEDOCKERHUB_ACCESS_TOKENDEPLOYMENT_SERVICE_KEY_JSONSTORAGE_SERVICE_KEY_JSONDOCUMENT_AI_SERVICE_KEY_JSONK8S_DEPLOY_TARGETSonly for the disabled Terraform CI rollout fallback- One runtime env secret per currently enabled collaborator:
STAGING_ENV,ISGS_ENV,NEWTS_ENV,OSAGE_ENV, andRRC_ENV.CA_ENVmay remain configured for future use, but CA cannot deploy whileenable_kubernetes_workloads = false.
DEPLOYMENT_SERVICE_KEY_JSON is used only by GitHub Actions to authenticate to
Google Cloud and apply the deployment. STORAGE_SERVICE_KEY_JSON and
DOCUMENT_AI_SERVICE_KEY_JSON are written into the Kubernetes
backend-runtime-files secret and mounted into backend pods as runtime keys.
The environment-file secret should contain the same key/value pairs used by the VM .env file. The workflow overrides these Kubernetes-owned values:
ENVIRONMENTBACKEND_URLLOG_DIRLOCAL_STORAGE_ROOTLOCAL_STORAGE_URL_BASESTORAGE_BUCKET_NAMESTORAGE_SERVICE_KEYDOCUMENT_AI_SERVICE_KEYGOOGLE_APPLICATION_CREDENTIALS
Keep COLLABORATOR in the runtime secret when the backend needs collaborator-specific processor metadata.
When adding a collaborator, update deploy-k8s-dispatch.yml with the new DEPLOY_ENV, <COLLABORATOR>_ENV secret, and runtime-env case branch before relying on a dedicated workflow. A dedicated deploy-k8s-<collaborator>.yml workflow calls the dispatch workflow and will fail until dispatch supports that collaborator.
Deploy with GitHub Actions
The staging workflow builds and pushes the backend image, then deploys staging:
gh workflow run deploy-k8s-staging.yml \
--repo CATALOG-Historic-Records/orphaned-wells-ui-server \
--ref main
Collaborator workflows deploy an existing image tag through IMAGE_TAG: auto. On collaborator branch merge commits, auto resolves to the second parent commit, which is the main commit that staging built and tested. On fast-forward or non-merge commits, auto resolves to the current commit:
gh workflow run deploy-k8s-isgs.yml --repo CATALOG-Historic-Records/orphaned-wells-ui-server --ref isgs
gh workflow run deploy-k8s-newts.yml --repo CATALOG-Historic-Records/orphaned-wells-ui-server --ref newts
gh workflow run deploy-k8s-osage.yml --repo CATALOG-Historic-Records/orphaned-wells-ui-server --ref osage
gh workflow run deploy-k8s-rrc.yml --repo CATALOG-Historic-Records/orphaned-wells-ui-server --ref rrc
The exact set of deployable collaborators is controlled by both
kubernetes_deploy_targets and .github/workflows/deploy-k8s-dispatch.yml.
Confirm that the collaborator has an output entry, is present in
workflow_dispatch.inputs.DEPLOY_ENV.options, and has a runtime-env secret
mapping before dispatching it.
Merging source changes into a collaborator branch only changes the live backend when that branch's deployment workflow runs. If automatic deployments are enabled for the branch, the merge can trigger the rollout. Otherwise, manually dispatch the collaborator workflow after the merge.
Automatic deployments are controlled by repository variables:
ENABLE_GKE_DEPLOYMENTS=true
Or one collaborator at a time:
ENABLE_GKE_STAGING_DEPLOY=true
ENABLE_GKE_ISGS_DEPLOY=true
ENABLE_GKE_NEWTS_DEPLOY=true
ENABLE_GKE_OSAGE_DEPLOY=true
The current RRC workflow checks ENABLE_GKE_DEPLOYMENTS; add an ENABLE_GKE_RRC_DEPLOY workflow check before relying on an RRC-specific deploy variable.
Check rollout status
Authenticate to the cluster:
gcloud container clusters get-credentials uow-backend-gke \
--region us-central1 \
--project <PROJECT_ID>
Check one backend:
kubectl -n uow-staging get deployment backend
kubectl -n uow-staging get pods -l app.kubernetes.io/name=orphaned-wells-ui-server -o wide
kubectl -n uow-staging get ingress backend
kubectl -n uow-staging get managedcertificate backend-cert
kubectl -n uow-staging get jobs -l app.kubernetes.io/component=processor
Read logs:
kubectl -n uow-staging logs deployment/backend --tail=200
kubectl -n uow-staging logs deployment/backend --tail=200 -f
kubectl -n uow-staging logs job/<worker-job-name> --tail=200
Restart a backend deployment:
kubectl -n uow-staging rollout restart deployment/backend
kubectl -n uow-staging rollout status deployment/backend --timeout=10m
Check resource usage:
kubectl top pods -n uow-staging
kubectl top nodes
Notes
- GKE replaces the legacy VM nginx/certbot path with GKE Ingress,
ManagedCertificate,FrontendConfig, andBackendConfig. - Google-managed certificates require DNS to point at the GKE load balancer before they become active.
- The rendered manifest in
deployment/kubernetes/rendered/backend.yamlis an ephemeral artifact. Do not commit it. - The Kubernetes Deployment uses pod-local
emptyDirvolumes for/logsand/data. Production document storage should continue using the Terraform-managed Google Cloud Storage upload bucket. - Collaborator APIs target two replicas at 1 CPU and 4 GiB memory each; their processing workers retain 1850m CPU and 12 GiB. Staging targets one API replica with a 500m CPU / 1 GiB request and 1 CPU / 2 GiB limit; its workers retain 1 CPU and 6 GiB.
- Worker Jobs have independent CPU, memory, scratch-storage, deadline, retention, and maximum-active-job settings. GitHub Actions supplies these generated settings from
K8S_DEPLOY_TARGETS; do not placePROCESSING_JOB_*values in collaborator runtime-env secrets.