Production deployment

Use the Helm chart with managed PostgreSQL and Redis, an ingress with TLS, and your organization’s sign-in provider. Start from the checked-in production values and adapt the hostname, storage classes, resources, and network policy to your cluster.

Warning

Optio currently requires one combined API/web replica. The API owns in-memory Local relays, session broadcasts, and cancellation state. Keep api.replicas at 1 and API autoscaling disabled. Recreate upgrades include a brief control-plane outage; this is not a highly available API. Agent pods scale independently.

Managed services and existing Secrets

Provision Secrets in the Optio namespace through your secret manager, then reference their keys. Helm does not need to read their contents. Redis must retain queue state without evicting queue keys; configure database and Redis TLS according to your managed services. Redis Cluster and Amazon ElastiCache Serverless run with externalRedis.mode: cluster, which keeps every queue under one hash-tagged key prefix. The Redis guide covers TLS, authentication, eviction behavior, capacity alarms, and a smoke test to run before switching.

values.production.yaml
publicUrl: https://optio.example.com
api:
  replicas: 1
  strategy:
    type: Recreate
  autoscaling:
    enabled: false
postgresql:
  enabled: false
redis:
  enabled: false
auth:
  disabled: false
existingSecrets:
  DATABASE_URL: { name: optio-runtime, key: database-url }
  REDIS_URL: { name: optio-runtime, key: redis-url }
  OPTIO_ENCRYPTION_KEY: { name: optio-runtime, key: encryption-key }
  GOOGLE_OAUTH_CLIENT_ID: { name: optio-sign-in, key: client-id }
  GOOGLE_OAUTH_CLIENT_SECRET: { name: optio-sign-in, key: client-secret }

The same mechanism supports GitHub/GitLab OAuth, generic OIDC, and other API environment settings. Secret references override inline settings. Missing keys prevent startup. Restart the deployment after rotating an external Secret so containers receive the new environment values.

Generate the encryption key once
openssl rand -hex 32

Warning

Back up the encryption key separately and retain the same key across upgrades and restores. Replacing it makes saved credentials unreadable; changing a Secret is not an encryption-key migration.

Sign-in and ingress

Configure OAuth or OIDC through existing Secrets, or use the bootstrap wizard for organization Google sign-in. Keep authentication enabled. Set publicUrl to the browser-facing HTTPS origin and use the production values’ ingress paths for /, /api, and /ws, with a valid TLS certificate and WebSocket timeouts.

The combined pod contains separate API and web containers. There is no independent web replica setting. With a shared ingress origin, the web app’s API and WebSocket connections use that origin.

Storage and worker identity

On EKS, install the EBS CSI driver and choose a StorageClass available in your cluster. Repo StatefulSets keep home and workspace claims; retained claims need their own capacity planning, backups, and cleanup policy.

Worker settings
agent:
  pvc:
    storageClass: gp3 # example; must exist in your cluster
    size: 10Gi
  cache:
    storageClass: gp3
  serviceAccount:
    create: true
    annotations: {} # configure an agent IAM role only if needed

Agents have a separate service account, no API control-plane RBAC, and no automounted Kubernetes token. Use scoped Connections for per-owner credentials. Pool isolation includes workspace, owner, purpose, and explicit access settings; work within one pool is mutually trusted. Pods share a kernel and are not a sufficient boundary for hostile tenants.

Install and upgrade

Use the chart from the release you intend to deploy and pin matching API, web, and agent image versions. Back up PostgreSQL, the encryption key, and needed PVC data first.

From the matching release checkout
helm upgrade --install optio helm/optio \
  --namespace optio --create-namespace \
  -f values.production.yaml

kubectl rollout status deployment/optio-api -n optio

Database migrations run at API startup. Local daemons reconnect with their existing PTYs. Pod terminals use tmux to reconnect to an existing shell, but a recreated pod loses its processes even when its files survive on a persistent volume.

Understand recovery

Session status distinguishes live, reconnecting, resumable, lost, and ended. An uncertain execution outcome requires inspection before retrying; it is not treated as successful completion or blindly replayed. Durable chat receipts prevent the same accepted request ID from starting another turn after restart. External side effects still need application-level idempotency or human review.

Read the complete EKS, recovery, and collaboration guide for migration handling, legacy pods, claim retention, and sharing boundaries. See the configuration reference for other settings.