Skip to main content

Create and connect Google project

OGRRE relies on Google Cloud Platform (GCP) for multiple services. To process documents, OGRRE uses Document AI. To store document files, OGRRE uses Cloud Storage. To authenticate users, OGRRE uses Google OAuth Platform. It is configured to work with any Google Cloud project, as long as the proper credentials are provided. To learn how to configure OGRRE with GCP, read the following tutorials.

Setting up a GCP project​

  1. Create an account:
  2. To create a GCP account, you must have a Google account. Then, simply sign up on the GCP console.
  3. Create a project:
  4. To create a project and grant access to it, see the Google documentation on creating and managing projects.
  5. Create a bucket:
  6. OGRRE uses cloud storage buckets to store and retrieve well document image files. To create a bucket, see the documentation on creating a storage bucket.
  7. Set up OAuth:
  8. OGRRE relies on Google OAuth to authenticate users. To create an OAuth client, see the documentation on using Google OAuth.
  9. Get access credentials:
  10. OGRRE uses separate service accounts for storage runtime, Document AI runtime, and deployment/Terraform operations. See Create Google service-account keys for the OGRRE-specific setup steps, or see Google's documentation on creating a service account and creating service account keys.

Create Google service-account keys​

OGRRE separates runtime, deployment, and Terraform CI identities. The three runtime/deployment accounts below use JSON keys in phase one; the two Terraform CI accounts use Workload Identity Federation (WIF), without JSON keys.

Service accountSuggested nameUsed byCredential setting
Storage runtimeogrre-storage-runtimeBackend runtime in production GKE, local Docker, and local non-Docker runs when STORAGE_BACKEND=google. Reads, writes, lists, moves, signs, and deletes uploaded/generated files in Cloud Storage.Local backend env: STORAGE_SERVICE_KEY. Backend GitHub secret: STORAGE_SERVICE_KEY_JSON.
Document AI runtimeogrre-document-aiBackend runtime in production GKE, local Docker, and local non-Docker runs when DOCUMENT_AI_BACKEND=google. Runs online and batch processing and deploys, undeploys, and checks processor versions.Local backend env: DOCUMENT_AI_SERVICE_KEY. Backend GitHub secret: DOCUMENT_AI_SERVICE_KEY_JSON.
Deploymentogrre-deployment-ciGitHub Actions deployment for backend GKE and frontend App Engine.Backend and frontend GitHub secret: DEPLOYMENT_SERVICE_KEY_JSON.
Terraform plan/no-change completiongithub-terraform-planReads infrastructure and state, manages the workspace lock, and writes only the workspace's CI readiness record after no-change verification.WIF; backend repository variable TF_PLAN_SERVICE_ACCOUNT.
Terraform applygithub-terraform-ciApplies infrastructure changes after terraform-apply approval.WIF; backend repository variable TF_APPLY_SERVICE_ACCOUNT.
note

The backend's operator-run deployment/ci/bootstrap_terraform_ci.sh creates and configures both Terraform CI accounts and their WIF grants. Do not create JSON keys for them. Manual Terraform uses an authorized human ADC login or a separate platform account; keep infrastructure administration off the deployment identity. See Terraform setup and manual operations.

Create the accounts​

Repeat these steps for the storage, Document AI, and deployment accounts. Use the bootstrap script for the Terraform CI accounts:

  1. Open the Google Cloud Console.
  2. Select the project that owns the OGRRE infrastructure.
  3. Open IAM & Admin -> Service Accounts.
  4. Click Create service account.
  5. Enter the service account name and description.
  6. Click Create and continue.
  7. Grant only the roles required for that account's task, described below.
  8. Click Done.

Grant storage runtime access​

Grant the storage runtime service account Storage Object User (roles/storage.objectUser) on each OGRRE upload bucket used by STORAGE_BUCKET_NAME or by Terraform's kubernetes_deploy_targets output.

Bucket-level access is the preferred scope. Project-level access also works, but it applies to every bucket in the project. If an older project policy blocks roles/storage.objectUser, Storage Object Admin (roles/storage.objectAdmin) is the broader fallback.

See Google's Cloud Storage IAM roles and bucket IAM policy documentation for the current role definitions and console flow.

Grant Document AI runtime access​

Grant the Document AI runtime service account Document AI API User (roles/documentai.apiUser) on the project or on the specific processors it needs to use.

OGRRE also deploys, undeploys, and checks processor versions from the backend, so grant a custom role containing:

  • documentai.processorVersions.get
  • documentai.processorVersions.list
  • documentai.processorVersions.update

If you do not want to maintain a custom role, use Document AI Editor (roles/documentai.editor) as the broader predefined fallback. See Google's Document AI IAM roles documentation for the current role definitions.

For batch processing, OGRRE sends gs://... input and output locations to the Document AI Batch API. Grant this same Document AI service account:

  • Storage Object Viewer (roles/storage.objectViewer) on each batch input bucket.
  • Storage Object Creator (roles/storage.objectCreator) on each batch output bucket.

If input and output use the same OGRRE bucket and you prefer simpler administration, Storage Object User (roles/storage.objectUser) on that bucket is the broader fallback.

Grant deployment and Terraform access​

Grant each identity the roles needed for its own operations.

For Terraform infrastructure updates, bootstrap grants the apply account these roles. A human or separate platform account performing manual applies needs equivalent access:

  • Kubernetes Engine Admin (roles/container.admin)
  • Compute Network Admin (roles/compute.networkAdmin)
  • Cloud DNS Administrator (roles/dns.admin)
  • Storage Admin (roles/storage.admin)
  • Service Usage Admin (roles/serviceusage.serviceUsageAdmin) when Terraform manages project services

The apply identity also needs read/write access to the shared Terraform state bucket, gs://tidy-outlet-412020-ogrre-terraform-state. The project-level roles/storage.admin grant covers this when the state bucket is in the same project. If it is scoped more tightly, grant the account access on the state bucket directly.

For backend GKE deployments, give the deployment identity Kubernetes Engine Developer (roles/container.developer). Terraform owns namespaces and runtime RBAC; deployment applies application workloads. Bootstrap also grants the deployment account reads on the state and CI buckets so deployment can verify readiness and read live targets. The plan account's exact read, lock, and readiness grants are documented in the backend CI identity guide.

For frontend App Engine deployments from GitHub Actions:

  • App Engine Deployer (roles/appengine.deployer)
  • App Engine Service Admin (roles/appengine.serviceAdmin) so gcloud app deploy can promote the new version by updating service traffic
  • Cloud Build Editor (roles/cloudbuild.builds.editor)
  • Storage Object Admin (roles/storage.objectAdmin) on the App Engine staging/build buckets or the project
  • Service Account User (roles/iam.serviceAccountUser) on the App Engine runtime service account

roles/appengine.deployer can create a new App Engine version, but it cannot update the service traffic split. The default frontend workflow promotes the new version, which requires appengine.services.update; roles/appengine.serviceAdmin is the least-broad predefined App Engine role in this list that grants that promotion permission. If a workflow is changed to deploy with --no-promote, this role is not needed for the deploy step, but a separate operator still needs traffic-update permission to make the version live.

Project-level roles/iam.serviceAccountUser works, but it lets this deployment account act as service accounts across the project. Grant it on the App Engine runtime service account when practical.

Do not use the storage runtime or Document AI runtime keys for Terraform, gcloud app deploy, or GKE deployment workflows.

Project-level versus resource-level roles​

Google Cloud roles granted on a project are inherited by resources in that project. Granting roles/storage.objectUser at the project level therefore works for OGRRE buckets in that project, but it also grants the same access to other project buckets. Resource-level grants are narrower and are preferred for the runtime accounts.

Use project-level roles for the Terraform apply identity where it must create and update infrastructure across the project. Keep deployment and runtime grants scoped to the operations and resources those accounts use.

Generate the JSON key files​

Repeat these steps for each service account that needs a JSON key:

  1. Open IAM & Admin -> Service Accounts.
  2. Click the email address for the service account.
  3. Open the Keys tab.
  4. Click Add key -> Create new key.
  5. Select JSON.
  6. Click Create.
  7. Save the downloaded file in a secure local location. Google only downloads the private key file once.
  8. Rename the file with a -service-key.json suffix.

Suggested local filenames:

ogrre-storage-runtime-service-key.json
ogrre-document-ai-service-key.json
ogrre-deployment-ci-service-key.json

Configure local development​

For Docker-based local development, the simplest location for runtime keys is next to orphaned-wells-ui/deployment/.env:

orphaned-wells-ui/deployment/ogrre-storage-runtime-service-key.json
orphaned-wells-ui/deployment/ogrre-document-ai-service-key.json

Then set these values in orphaned-wells-ui/deployment/.env:

PROJECT_ID=<project-id>
LOCATION=us
STORAGE_BACKEND=google
STORAGE_BUCKET_NAME=<bucket-name>
STORAGE_SERVICE_KEY=ogrre-storage-runtime-service-key.json
DOCUMENT_AI_BACKEND=google
DOCUMENT_AI_SERVICE_KEY=ogrre-document-ai-service-key.json

STORAGE_SERVICE_KEY and DOCUMENT_AI_SERVICE_KEY can also be absolute paths. The Docker start script mounts these key files into the backend container automatically when the matching Google backend is configured.

For local non-Docker backend runs, put the same runtime key files in orphaned-wells-ui-server/ogrre/ or use absolute paths in orphaned-wells-ui-server/ogrre/.env.

Do not set GOOGLE_APPLICATION_CREDENTIALS to the storage or Document AI runtime key for normal backend development. The backend reads STORAGE_SERVICE_KEY and DOCUMENT_AI_SERVICE_KEY directly.

For manual Terraform, prefer an authorized human login. From the Terraform directory, clear inherited credential overrides before configuring ADC:

cd orphaned-wells-ui-server/deployment/terraform
unset GOOGLE_APPLICATION_CREDENTIALS GOOGLE_AUTHORIZED_USER_CREDENTIALS CLOUDSDK_AUTH_CREDENTIAL_FILE_OVERRIDE
gcloud auth login
gcloud config set project <PROJECT_ID>
gcloud auth application-default login
gcloud components install gke-gcloud-auth-plugin
unset TF_WORKSPACE
terraform init -input=false -lockfile=readonly
terraform workspace select ogrre
terraform plan -input=false -lock-timeout=5m

Use the backend's pinned Terraform version. A separately authorized platform account is also supported; see the manual plan/apply guide. Do not leave a platform credential override set when returning to runtime work.

Configure production GitHub secrets​

Store the key JSON contents as GitHub Actions secrets:

gh secret set DEPLOYMENT_SERVICE_KEY_JSON \
--repo CATALOG-Historic-Records/orphaned-wells-ui-server \
< /secure/path/ogrre-deployment-ci-service-key.json

gh secret set STORAGE_SERVICE_KEY_JSON \
--repo CATALOG-Historic-Records/orphaned-wells-ui-server \
< /secure/path/ogrre-storage-runtime-service-key.json

gh secret set DOCUMENT_AI_SERVICE_KEY_JSON \
--repo CATALOG-Historic-Records/orphaned-wells-ui-server \
< /secure/path/ogrre-document-ai-service-key.json

gh secret set DEPLOYMENT_SERVICE_KEY_JSON \
--repo CATALOG-Historic-Records/orphaned-wells-ui \
< /secure/path/ogrre-deployment-ci-service-key.json

Backend GKE deployments create Kubernetes secrets from STORAGE_SERVICE_KEY_JSON and DOCUMENT_AI_SERVICE_KEY_JSON and mount them into the backend pods. The frontend App Engine deployment uses DEPLOYMENT_SERVICE_KEY_JSON only at deploy time; the static frontend does not need runtime access to service-account keys.

Clean up old credentials​

After validating storage uploads/downloads, Document AI processing, processor deploy/undeploy, backend deployment, and frontend deployment with the new accounts, remove old GitHub secrets and local files:

gh secret delete SERVICE_KEY_JSON \
--repo CATALOG-Historic-Records/orphaned-wells-ui-server

gh secret delete CREDS_JSON \
--repo CATALOG-Historic-Records/orphaned-wells-ui-server

gh secret delete GCLOUD_SERVICE_ACCOUNT_JSON \
--repo CATALOG-Historic-Records/orphaned-wells-ui

Update ignored local .env files so they no longer reference michael2-service-key.json, then delete local copies of the old key. In Google Cloud, disable the old service account or key first, verify nothing breaks, then delete the old key and remove the old account's IAM bindings.

Keep service-account keys out of git

Service-account JSON files are long-lived secrets. Do not commit them to either repository. Use the -service-key.json naming convention so the frontend and backend .gitignore rules catch the file before commit.

Before committing setup changes, verify that git is ignoring the key file:

git check-ignore -v deployment/ogrre-storage-runtime-service-key.json
git check-ignore -v deployment/ogrre-document-ai-service-key.json
git check-ignore -v deployment/ogrre-deployment-ci-service-key.json

If the command prints no output, the file is not ignored and should not be committed.

Some Google Cloud organizations disable service-account key creation by policy. If the key creation button is unavailable, ask your Google Cloud administrator whether the project can be exempted or whether a different authentication method should be used.

Environment variables for OGRRE​

When using Google Cloud for storage, Document AI, or authentication, configure the relevant GCP environment variables in your backend .env file. See the backend environment variables section for when each variable is required.
  • PROJECT_ID: The project ID of your GCP project.
  • LOCATION: The location you chose for document AI processors, likely to be "us".
  • STORAGE_BUCKET_NAME: The name you chose for your storage bucket.
  • STORAGE_SERVICE_KEY: The name of the file storing your Google Cloud Storage service-account key. This file should be stored in the same directory as the .env file and named with a -service-key.json suffix so it is ignored by git.
  • DOCUMENT_AI_SERVICE_KEY: The name of the file storing your Google Document AI service-account key. This file should be stored in the same directory as the .env file and named with a -service-key.json suffix so it is ignored by git.
  • token_uri: The endpoint URL where an application requests and receives access tokens. This is necessary for user authentication. For more information, see the Google's documentation on OAuth2.
  • client_id: The client ID of your OAuth client.
  • client_secret: The client secret of your OAuth client.