Self-hosting
Reference

Configure a self-hosted platform

Configure Agent Barn’s Product, Ingest, and Communications services, identity, databases, Agent runtimes, networking, storage, integrations, workers, and monitoring.

For
Platform engineers, self-hosted operators, and Kubernetes administrators
On this page
  1. Overview
  2. Choose values for your deployment
  3. Configuration paths
  4. Configuration boundaries
  5. Before you begin
  6. 1. Create the configuration file
  7. 2. Configure identity and stable keys
  8. 3. Configure data services
  9. 4. Configure models and Agent runtimes
  10. 5. Configure URLs, DNS, and TLS
  11. 6. Configure Kubernetes access and storage
  12. 7. Configure optional integrations
  13. 8. Configure workers, telemetry, and Communications
  14. Communication credential ownership
  15. 9. Configure monitoring
  16. Separate production and staging
  17. Configuration variable reference
  18. Apply configuration changes
  19. Validate the configuration
  20. Security and rotation
  21. Current configuration constraints
  22. Troubleshooting
  23. Next steps
  24. Staging and custom namespaces
  • Self-hosting
  • Configuration
  • 18 minutes

Configure the services, credentials, networking, storage, and runtime settings a self-hosted Agent Barn installation requires.

Overview

This reference covers both supported configuration paths: a local Docker and k3d environment, and a persistent Kubernetes deployment using Helmfile. It prepares the configuration the deployment tools consume; Deploy Agent Barn to Kubernetes applies it to a cluster.

Agent Barn has three separately served HTTP application boundaries:

ServicePrefixDefault portResponsibility
Product API/api/v18000Human-facing product operations
Ingest API/ingest/v18001Runtime Tool Call telemetry
Communications/communications/v18002Provider ingress, Runtime communication, durable Deliveries, and outbound replies

Communications is a separately served process. It owns canonical Conversation writes; Ingest owns Tool Call and Tool Result writes. The paths use separate Runtime credentials and neither is a substitute for the other.

Activity write paths
Provider → Communications :8002 → Conversation Messages
Runtime replies → Communications :8002 → Conversation Messages
Runtime Tool Calls → Ingest :8001 → Tool Call records

Configuration moves through five layers:

  1. Environment file A local, secret-bearing .env for Docker and k3d, or .env.deploy for Kubernetes. Neither file is committed.
  2. Deployment tooling run.sh and compose.yml locally, or deploy.sh and helmfile.yaml.gotmpl for Kubernetes.
  3. Kubernetes Secrets and Helm values The tooling maps environment variables into chart values and Secrets. A value only takes effect when something actually maps it.
  4. Platform services The Product API, Ingest API, Communications service, UI, PostgreSQL releases, LiteLLM, Redis, Firecrawl, workers, and monitoring.
  5. Generated Agent workloads The API renders Deployments, Secrets, Services, ConfigMaps, and PVCs for each Agent when it starts.

The diagram reads top to bottom: an environment file feeds the deployment tooling, which produces Kubernetes Secrets and Helm values, which configure the three application processes and supporting services, which in turn generate each Agent’s workload.

Configuration has two different lifetimes:

Lifetime Examples
Platform configuration Database credentials, signing key, ingress hosts, email provider, registry, monitoring
Generated Agent configuration Runtime image, model proxy address, Ingest and Communications addresses, generated credentials, rendered Templates and Skills

Changing platform configuration does not necessarily rebuild an already running Agent. Generated Agent settings normally take effect the next time the Agent is stopped and started.

Choose values for your deployment

The deployment configuration describes your installation. Use addresses, credentials, and infrastructure names supplied by your own providers or cluster administrator.

For the repository's deploy.sh workflow, settings are read from .env.deploy. AAI Labs GitHub Actions workflows use additional names such as PUBLIC_* and STAGING_* to select values for their environments. Those workflow names are not interchangeable with the deployment input names below. See AAI Labs hosted-service operations for those workflows.

Input What to supply
KUBECONFIGThe kubeconfig path on the machine running deployment tools.
NAMESPACEThe target namespace. The default is agent-farm; bootstrap resources and permissions must agree with the selected namespace.
REGISTRY_PREFIXThe prefix used to construct your image repository paths.
REGISTRY_SERVERThe registry host used for authentication. It may differ from a prefix that includes a repository path.
API_IMAGE_REPOSITORY, UI_IMAGE_REPOSITORY, HERMES_IMAGE_REPOSITORY, OPENCLAW_IMAGE_REPOSITORYThe repository names used by your chosen image distribution.
API_IMAGE_TAG, UI_IMAGE_TAG, HERMES_IMAGE_TAG, OPENCLAW_IMAGE_TAGThe corresponding image versions for your deployment. Keep component version choices explicit.
API_HOST, UI_HOSTHostnames served by your ingress, without a URL scheme.
WEB_APP_URLThe web application's full external URL, including its scheme.
INGRESS_CLUSTER_ISSUERThe name of the cert-manager ClusterIssuer in your cluster.
STORAGE_CLASSA suitable StorageClass from your cluster. A company-specific class name is not portable to another cluster.
POD_KUBECONFIG_B64The encoded kubeconfig used by the API to manage Agent workloads. In the current deploy path, omitting it falls back to the deployment kubeconfig.

Required and optional status, defaults, encoding requirements, and provider restrictions for every setting are documented in the detailed sections below.

Configuration paths

Choose the file that matches how you run Agent Barn.

Local: .env

Copied from .env.spec at the repository root. Applied by ./run.sh and compose.yml, and read directly by the API, UI, worker, and Make targets during native development.

run.sh validates the variables the complete local stack requires, and exits listing any that are missing.

Kubernetes: .env.deploy

Copied from .env.deploy.spec. Applied by ./deploy.sh, which sources the file, derives POD_KUBECONFIG_B64 when it is omitted, applies the bootstrap RBAC manifest, and runs Helmfile.

Variable names and network addresses differ from the local file. Do not use .env as a production deployment file.

Environment Template Working file Applied by
Local Docker and k3d .env.spec .env ./run.sh and compose.yml
Local native development .env.spec .env API, UI, worker, and Make targets
Kubernetes .env.deploy.spec .env.deploy ./deploy.sh and helmfile.yaml.gotmpl
Staging or custom namespace .env.deploy.spec Complete environment-specific Helmfile inputs Provision matching namespace/RBAC/kubeconfig prerequisites first; use the environment deployment workflow or Helmfile directly. deploy.sh is not a staging bootstrap.

Configuration boundaries

Configuration is owned at several layers.

Layer Owns
Deployment environment Secrets, image versions, URLs, storage, registry access, and external-provider credentials
Helmfile Release order, and mapping environment variables to chart values
Helm charts Kubernetes Secrets, Deployments, Services, hooks, probes, ingress, worker, and reconciliation
API configuration Authentication, encryption, the model catalogue, runtime images, and integration clients
Organization settings Allowed models, Members, Templates, Skills, and Shared Credentials
Agent settings Runtime, model, Communication Connections and their access policies, Skills, Integration material, and the Template pin

Environment configuration should establish a secure platform baseline. Organization and Agent configuration belongs in the product’s authenticated workflows.

Before you begin

For local configuration

  • The public agent-barn repository
  • Docker with Linux containers
  • kubectl
  • An OpenRouter API key
  • Enough memory for the application, k3d, LiteLLM, and Agent runtimes
  • Compatible Hermes and OpenClaw image references
  • A GitHub token only when the local base-image build needs authenticated repository access

For Kubernetes configuration

  • The Agent Barn deployment files or release bundle
  • An existing Kubernetes cluster and a kubeconfig for it
  • Helm, Helmfile, and the Helm diff plugin
  • A registry holding compatible API, UI, Hermes, and OpenClaw images
  • Three DNS names: UI, API, and Grafana
  • Traefik and cert-manager
  • A usable StorageClass
  • OpenRouter, registry, monitoring, and database credentials

Create the configuration file

Local environment

From the repository root:

Copy the template
cp .env.spec .env

Open .env and replace the placeholder values. The full local launcher checks its required values before it creates the cluster or starts containers:

Start the stack
./run.sh

If a required value is missing, the script exits and lists the missing variable names.

For native local development, start Product API, Ingest, and Communications together with make dev-api, or run Communications alone with make dev-communications. The complete ./run.sh Docker stack includes Communications as a separate process. Where local tooling supports it, its default port is:

Communications port
COMMUNICATIONS_PORT=8002

Kubernetes environment

Open the configuration for your deployment path:

Starting point Configuration step
Extracted Source code ZIP or tar.gzOpen the extracted repository directory and copy .env.deploy.spec to .env.deploy. Supply the deployment's infrastructure values, image locations, tags, and credentials.
Git checkout using deploy.shCopy .env.deploy.spec to .env.deploy and configure the selected checkout's deployment inputs.
Separate deployment bundle, when providedEdit the bundle's included .env.deploy. Preserve its intended image selections while supplying your installation settings.

The bundle and repository spec can contain different image repository defaults. Use the values for your chosen image distribution rather than mixing the two configurations.

For bundle discovery and image-access prerequisites, see Deploy to Kubernetes.

For a repository checkout, copy the template from the deployment directory:

Copy the template
cp .env.deploy.spec .env.deploy

Restrict access to the file:

Restrict permissions
chmod 600 .env.deploy

The file is sourced by a shell, so use plain assignments:

Assignment shape
KEY=value

Avoid spaces around =, shell command substitutions, values copied from examples, unescaped shell metacharacters, multiline secret values, and comments placed after a secret on the same line.

Use a secret manager as the authoritative backup. The deployment file is an input, not an adequate recovery system.

Configure identity and stable keys

Agent Barn requires three stable cryptographic values and one bootstrap administrator credential.

Generate the stable values

Generate the signing key:

Signing key
openssl rand -hex 32

Generate the Fernet encryption key:

Encryption key
python -c "from cryptography.fernet import Fernet; print(Fernet.generate_key().decode())"

Generate the LiteLLM master key:

LiteLLM master key
echo "sk-$(openssl rand -hex 24)"

Add them to the configuration:

Stable keys
SECRET_SIGNING_KEY=REPLACE_WITH_GENERATED_SIGNING_KEY
AGENT_TOKEN_ENCRYPTION_KEY=REPLACE_WITH_GENERATED_FERNET_KEY
LITELLM_MASTER_KEY=sk-REPLACE_WITH_GENERATED_VALUE

Stable-key responsibilities

Variable Requirement Used for Effect of an unplanned change
SECRET_SIGNING_KEY Stable Signing Agent Barn access tokens Existing signed sessions can become invalid
AGENT_TOKEN_ENCRYPTION_KEY Stable Encrypting Agent Secrets, Connection credentials, Shared Credentials, LiteLLM keys, generated Runtime credentials, and stored OAuth tokens Existing encrypted data can no longer be decrypted
LITELLM_MASTER_KEY Stable LiteLLM administration, and encryption of its virtual-key data Existing Agent LiteLLM keys and LiteLLM database data can become unusable

Configure the bootstrap administrator

Bootstrap administrator
PLATFORM_ADMIN_CREDENTIALS=admin@example.com:REPLACE_WITH_STRONG_PASSWORD

The password must contain at least eight characters, an uppercase letter, a lowercase letter, and a digit. Do not use a colon in the password, because the value uses a colon to separate the email from the password.

The bootstrap credential is used only when Agent Barn finds no existing Platform Administrator. Changing the environment variable after the initial bootstrap does not change the existing administrator’s password.

Set the environment label

Environment label
ENVIRONMENT=production

Use a distinct value such as staging for another stack. The label appears in operational metrics and alerts.

Configure data services

The Kubernetes deployment creates three independent PostgreSQL releases.

Database Purpose
Application PostgreSQL Users, Organizations, Agents, Templates, Skills, Activity, and product state
LiteLLM PostgreSQL LiteLLM virtual keys and spend data
Firecrawl PostgreSQL Self-hosted Firecrawl state

Configure each database independently:

Database credentials
POSTGRES_APP_USER=agentfarm
POSTGRES_APP_PASSWORD=REPLACE_WITH_UNIQUE_PASSWORD
POSTGRES_APP_DB=agentfarm

POSTGRES_LITELLM_USER=litellm
POSTGRES_LITELLM_PASSWORD=REPLACE_WITH_DIFFERENT_PASSWORD
POSTGRES_LITELLM_DB=litellm

POSTGRES_FIRECRAWL_USER=firecrawl
POSTGRES_FIRECRAWL_PASSWORD=REPLACE_WITH_ANOTHER_PASSWORD
POSTGRES_FIRECRAWL_DB=firecrawl

Generate URL-safe passwords with:

Password
openssl rand -hex 24

Hex values avoid connection-string escaping problems. Do not reuse one password across all three databases.

Local database URL

Native host commands require an explicit application database URL:

Local database URL
DB_CONNECTION_URL=postgresql://${POSTGRES_USER}:${POSTGRES_PASSWORD}@localhost:${POSTGRES_PORT}/${POSTGRES_DB}

Compose overrides this with the internal db service hostname, and Kubernetes derives the application connection URL from its POSTGRES_APP_* values.

Redis

Redis carries the general background worker and the Event Delivery pipeline:

Redis URL
REDIS_URL=redis://localhost:6379/0

Local Compose replaces the hostname with redis. Kubernetes uses:

In-cluster Redis
redis://redis:6379/0

Redis unavailability does not prevent the Product API from committing supported Domain Events, but low-latency Event Delivery processing stops until the worker and transport recover.

Firecrawl

Generate a platform Firecrawl key with openssl rand -hex 24, then configure it for Kubernetes:

Firecrawl key
FIRECRAWL_API_KEY=REPLACE_WITH_GENERATED_VALUE

The current deployment uses the same platform key for the self-hosted Firecrawl service and the default Agent integration. Individual Agents can later receive their own credential overrides; see Connect Firecrawl.

Configure models and Agent runtimes

Agent Barn sends model requests through LiteLLM, which uses OpenRouter as the configured provider.

Model request path
Agent runtime
      │
      ▼
LiteLLM proxy
      │
      ▼
OpenRouter

OpenRouter

Set a real OpenRouter inference key:

Provider key
OPENROUTER_API_KEY=sk-or-REPLACE_WITH_PROVIDER_KEY

OPENROUTER_API_KEY and LITELLM_MASTER_KEY are different credentials:

Variable Authority
OPENROUTER_API_KEY Makes provider model requests
LITELLM_MASTER_KEY Administers the self-hosted LiteLLM proxy and its virtual keys

Default model

Configure the default using the full Agent Barn model path:

Default model
AGENT_DEFAULT_MODEL=litellm/openrouter/z-ai/glm-5.2

The segment following litellm/openrouter/ must be a valid OpenRouter model ID.

Model catalogue allowlist

Limit the model picker with comma-separated fnmatch globs:

Allowlist
AGENT_MODEL_ALLOWLIST=z-ai/glm-5.2,openai/gpt-5*

Patterns match OpenRouter model IDs without the litellm/openrouter/ prefix. An empty value exposes the full available catalogue:

Unrestricted catalogue
AGENT_MODEL_ALLOWLIST=

An Organization can narrow its own allowed models further. The environment allowlist controls what the platform catalogue offers; it does not replace Organization model governance.

Runtime images

Kubernetes deployments use four compatible images:

Image tags
API_IMAGE_TAG=RELEASE_API_TAG
UI_IMAGE_TAG=RELEASE_UI_TAG
OPENCLAW_IMAGE_TAG=RELEASE_OPENCLAW_TAG
HERMES_IMAGE_TAG=RELEASE_HERMES_TAG

Local development uses full Agent image references:

Local Agent images
OPENCLAW_IMAGE=registry.example.com/agentbarn-openclaw-base:RELEASE_TAG
HERMES_IMAGE=registry.example.com/agentbarn-hermes-base:RELEASE_TAG

Kubernetes Agent workloads use OPENCLAW_IMAGE and HERMES_IMAGE. The older AGENT_IMAGE setting is not used by the current API.

Use the runtime image versions supplied with the selected Agent Barn release. Hermes and OpenClaw versions must stay compatible with that release; do not combine an arbitrary API image with unrelated runtime versions.

LiteLLM addresses

In Kubernetes, both the API and the Agents use the in-cluster service:

In-cluster LiteLLM
http://litellm:4000

In local k3d, the API container reaches LiteLLM through the Compose network, while Agent pods reach the host-published proxy through host.docker.internal on default port 7070:

Local Agent-facing LiteLLM
AGENT_LITELLM_BASE_URL=http://host.docker.internal:7070

Running Agents keep their current pod image and generated configuration until they are stopped and started again.

Configure URLs, DNS, and TLS

A Kubernetes deployment uses separate UI, API, and Grafana hostnames:

Hostnames
UI_HOST=agentbarn.example.com
API_HOST=api.agentbarn.example.com
GRAFANA_HOST=grafana.agentbarn.example.com
WEB_APP_URL=https://agentbarn.example.com

Use hostnames without a scheme for UI_HOST, API_HOST, and GRAFANA_HOST, and the complete public origin for WEB_APP_URL. The API chart derives API_EXTERNAL_URL from the first configured API hostname.

DNS

Point all three hostnames at the public address of the Traefik ingress controller, then verify resolution before deploying:

Resolve hostnames
dig +short agentbarn.example.com
dig +short api.agentbarn.example.com
dig +short grafana.agentbarn.example.com

Ingress exposure

The UI ingress exposes /. The API ingress exposes only:

API ingress path
/api

The API ingress deliberately does not expose /metrics, the Ingest API port, or internal worker endpoints. Prometheus scrapes metrics through Kubernetes Services, and Agent Runtimes send telemetry to the in-cluster Ingest service rather than a public ingress.

API_EXTERNAL_URL is the public Agent Barn origin used to construct provider webhook URLs. A saved webhook-based Connection receives a URL shaped like:

Provider webhook URL
https://<public-host>/communications/v1/webhooks/<connection-id>

The origin must use valid public HTTPS for providers such as Microsoft Teams, and public DNS must resolve to the deployment ingress. Route only /communications/v1/webhooks to Communications; this is not the internal Runtime Communications URL, and the webhook URL contains no provider credential.

SourceDestinationExposure
Browser or API clientProduct API :8000/api/v1Public or trusted application network
Agent RuntimeIngest :8001/ingest/v1Internal
Agent RuntimeCommunications :8002/communications/v1Internal
Platform DriverCommunications Connection event routesInternal
Provider webhookCommunications :8002/communications/v1/webhooks/{connection_id}Public HTTPS
CommunicationsPostgreSQLInternal
CommunicationsSlack, Teams, Telegram, or Discord APIsOutbound internet access
PrometheusCommunications :8002/metricsInternal
Kubernetes probesCommunications :8002/healthInternal

TLS

The API and UI charts currently use this cert-manager ClusterIssuer:

ClusterIssuer
letsencrypt-http01

Confirm that it exists:

Check the issuer
kubectl get clusterissuer letsencrypt-http01

HTTP-01 requires publicly reachable DNS and ingress. A private installation needs a certificate strategy compatible with its network.

Configure Kubernetes access and storage

The hook ServiceAccount and the API's mounted kubeconfig serve different purposes. Creating agent-farm-user does not create or export the API kubeconfig. See Deploy to Kubernetes for the access preparation and encoding steps.

Deployment kubeconfig

Set the kubeconfig used by kubectl, Helm, and Helmfile, using an absolute path:

Deployment kubeconfig
KUBECONFIG=/absolute/path/to/agent-barn-production.yaml

KUBECONFIG is the path to the kubeconfig used by deployment tools on the deployment machine. With the supplied deploy.sh, its identity must also be able to perform the script's bootstrap apply.

Namespace

Production uses:

Production namespace
NAMESPACE=agent-farm

Staging uses:

Staging namespace
NAMESPACE=agent-farm-staging

NAMESPACE is the namespace for the Helmfile releases, defaulting to agent-farm. Bootstrap resources and workload permissions must match; changing this value alone does not adapt the shipped bootstrap manifest.

The API chart sets K8S_NAMESPACE from the Helm release namespace, which prevents a staging API from creating Agent workloads in production.

API-facing kubeconfig

POD_KUBECONFIG_B64 is the single-line base64 encoding of the kubeconfig the API uses to manage Agent workloads:

API kubeconfig
POD_KUBECONFIG_B64=

Supply a separately authorized workload identity when required by your deployment policy. A missing or empty value falls back to the deployment kubeconfig in deploy.sh; that implementation default is not a statement about your installation's security requirements.

That identity needs to manage Deployments, Services, ConfigMaps, Secrets, PersistentVolumeClaims, Pods, Pod logs, and the exec and port-forward operations the health and log workflows use.

For the deployed API, use a kubeconfig for the installation's cluster. When running inside Kubernetes, the client can rewrite the configured server address to the in-cluster API endpoint; supplying another cluster's kubeconfig is not a supported way to redirect this deployment path.

When rotating this credential, update the deployment configuration and roll the API through the deployment process. The client caches its Kubernetes configuration, so replacing a mounted file alone should not be treated as a live credential reload. A token copied into a static kubeconfig does not automatically renew itself.

Container registry

Registry and repositories
REGISTRY_PREFIX=registry.example.com/agent-barn
REGISTRY_SERVER=registry.example.com
REGISTRY_USERNAME=REPLACE_WITH_USERNAME
REGISTRY_PASSWORD=REPLACE_WITH_ACCESS_TOKEN

API_IMAGE_REPOSITORY=api
UI_IMAGE_REPOSITORY=ui
HERMES_IMAGE_REPOSITORY=hermes-base
OPENCLAW_IMAGE_REPOSITORY=openclaw-base

REGISTRY_PREFIX is prepended to image repository names, and REGISTRY_SERVER identifies the registry for authentication.

Storage

Select a StorageClass, or leave it empty to use the cluster default:

StorageClass
STORAGE_CLASS=REPLACE_WITH_STORAGE_CLASS

The StorageClass is used by the application, LiteLLM, and Firecrawl PostgreSQL releases, by Prometheus, and by generated Agent PVCs. Use network-replicated storage when node-loss durability is required.

Configure optional integrations

Transactional email

Agent Barn uses Cloudflare Email Sending for invitations, password recovery, and Agent lifecycle notifications. Email is enabled only when all three values are set:

Email configuration
CLOUDFLARE_ACCOUNT_ID=REPLACE_WITH_ACCOUNT_ID
CLOUDFLARE_API_TOKEN=REPLACE_WITH_EMAIL_SENDING_TOKEN
SENDER_EMAIL=noreply@mail.agentbarn.example.com

The Cloudflare token needs Email Sending: Edit permission, and the SENDER_EMAIL domain must be verified for Email Sending. Use a separate mail.-style subdomain per environment:

Sender domains
Production: noreply@mail.agentbarn.example.com
Staging:    noreply@mail-staging.agentbarn.example.com

To disable delivery, leave all three empty:

Email disabled
CLOUDFLARE_ACCOUNT_ID=
CLOUDFLARE_API_TOKEN=
SENDER_EMAIL=

With delivery disabled, send attempts are logged and treated as no-ops. Avoid configuring only one or two values, because the platform still considers email disabled.

Google Workspace OAuth

Create a Google OAuth 2.0 Web application client and register this redirect URI:

Callback URI
https://agentbarn.example.com/api/v1/integrations/google/callback

Then configure the client values:

Google OAuth client
GOOGLE_CLOUD_CLIENT_ID=REPLACE_WITH_CLIENT_ID
GOOGLE_CLOUD_CLIENT_SECRET=REPLACE_WITH_CLIENT_SECRET

Google OAuth is enabled only when both values are set and the callback matches <WEB_APP_URL>/api/v1/integrations/google/callback exactly. Leave both empty to disable it.

These are application-owned OAuth client credentials. The refresh token created when a user connects Google Workspace remains Agent-specific and encrypted by Agent Barn; see Connect Google Workspace.

Configure workers, telemetry, and Communications

Event Delivery worker

The Kubernetes API chart deploys a general Dramatiq worker and a scheduled Event Delivery reconciliation CronJob. Both use Redis. The default reconciliation schedule is:

Reconciliation schedule
*/5 * * * *

Local development uses:

Local worker
make redis-up
make dev-worker

A one-shot local reconciliation pass is available through:

Reconcile once
make reconcile

Ingest API

Ingest is a separate FastAPI application on port 8001. INGEST_BASE_URL is the internal endpoint used only for Runtime Tool Call telemetry, authenticated with a per-start Ingest key. Its value must end in /ingest/v1.

The default Kubernetes address is:

In-cluster Ingest
INGEST_BASE_URL=http://agentbarn-api:8001/ingest/v1

Local Agent pods use:

Local Ingest
INGEST_BASE_URL=http://host.docker.internal:8001/ingest/v1

Do not use Ingest for Conversation Messages, provider events, Communication Deliveries, Runtime replies, or platform routing.

Communications service

COMMUNICATIONS_BASE_URL is the internal base URL supplied to Agent Runtimes for the versioned Communications protocol. It must end in /communications/v1; prefer the internal Service whenever it is available.

Local Communications
COMMUNICATIONS_BASE_URL=http://host.docker.internal:8002/communications/v1
Kubernetes Communications
COMMUNICATIONS_BASE_URL=http://agentbarn-api-communications:8002/communications/v1

The exact Kubernetes hostname depends on the Helm release name. Runtime communication uses a separate per-start Communications credential and protocol version; it is not interchangeable with the Ingest credential.

Make sure Agent pods can reach both internal services in Kubernetes:

Agent reachability
LiteLLM: http://litellm:4000
Ingest:  http://agentbarn-api:8001/ingest/v1
Communications: http://agentbarn-api-communications:8002/communications/v1

Communications uses the shared application database and these process environment variables:

VariableDefault or requirementPurpose
DB_CONNECTION_URLRequiredShared PostgreSQL database
AGENT_TOKEN_ENCRYPTION_KEYRequired and shared with Product APIDecrypt Connection credentials and validate Runtime communication identity
COMMUNICATION_JOURNAL_RETENTION_DAYS31Content-free journal retention
SLACK_DIRECTORY_CACHE_TTL_SECONDS600Slack directory discovery cache
SLACK_REQUEST_TIMEOUT_SECONDS30Slack provider API timeout
TEAMS_PUBLISHER_NAMEAgent BarnPublisher name in generated Teams app packages
TEAMS_PUBLISHER_WEBSITE_URLPublic HTTPS URLTeams package publisher website
TEAMS_PRIVACY_URLPublic HTTPS URLTeams package privacy URL
TEAMS_TERMS_URLPublic HTTPS URLTeams package terms URL

COMMUNICATION_JOURNAL_RETENTION_DAYS accepts values from 1 through 3650. If the Helm chart does not expose a dedicated value for an optional setting, configure it as a process environment variable instead of inventing chart behavior.

The development and test controls SKIP_SLACK_TOKEN_VALIDATION, SKIP_TELEGRAM_TOKEN_VALIDATION, SKIP_DISCORD_TOKEN_VALIDATION, and SKIP_TEAMS_TOKEN_VALIDATION default to false and bypass external credential validation. Use them only in controlled development or test environments; they must remain disabled in production and do not make invalid credentials usable.

Communications in Kubernetes

Communications chart values
communications:
  enabled: true
  replicaCount: 1
  service:
    port: 8002

The API chart creates a separate Communications Deployment and ClusterIP Service. Product API receives the internal Communications Service URL, while the Communications pod reads shared application configuration from the API Secret. Readiness and liveness use /health on port 8002; /metrics remains internal. Disabling Communications does not leave Platform messaging functional.

Increasing communications.replicaCount is supported by PostgreSQL-backed Connection ingress leases. One replica supervises a given provider Connection at a time, durable Deliveries remain in PostgreSQL, and Connection revision changes trigger provider-session reconciliation. Every replica shares the database and encryption configuration; webhook requests may be load-balanced across replicas. Do not run independent Communications replicas against separate databases.

Communication Deliveries are durable PostgreSQL records. Provider ingress and outbound processing do not use the Domain Event Dramatiq queue as their source of truth; Redis remains for the general worker and Domain Event delivery.

Communication credential ownership

CredentialOwner
Slack xoxb- and xapp- tokensSlack Communication Connection
Teams App ID, client secret, and Tenant IDMicrosoft Teams Communication Connection
Telegram bot tokenTelegram Communication Connection
Discord bot tokenDiscord Communication Connection
Tool Integration credentialsAgent Secrets or Shared Credentials
Runtime Ingest keyGenerated per Agent start
Runtime Communications credentialGenerated per Agent start

Communication Connection credentials are submitted through Connection management, encrypted at rest, and never returned through read APIs. They are not materialized into Hermes or OpenClaw and are independent from Agent Secrets and Shared Credentials. One Agent can have multiple Communication Connections.

An Agent Runtime receives its internal Ingest and Communications base URLs, per-start Ingest credential, per-start Communications protocol credential, Communications protocol version, Runtime configuration, and tool Integration material. It does not receive provider tokens, Teams client secrets, platform channel policies, provider webhook credentials, or provider-session configuration.

See Communication Connections for ownership and administration, and Self-hosting Communications for service operation.

Configure monitoring

The current Helmfile deploys Prometheus, Grafana, Alertmanager, kube-state-metrics, and the Agent Barn dashboards and alert rules.

Add the following variables to .env.deploy:

Monitoring configuration
SLACK_ALERTS_WEBHOOK_URL=REPLACE_WITH_ALERT_WEBHOOK
GRAFANA_ADMIN_PASSWORD=REPLACE_WITH_STRONG_PASSWORD
GRAFANA_HOST=grafana.agentbarn.example.com

Build the chart dependencies before deployment:

Chart dependencies
helm dependency build helm/monitoring

The monitoring stack is namespace-scoped and does not create its own cluster-scoped RBAC.

Prometheus reads Product API, Ingest API, Communications, LiteLLM, and Agent health metrics, Kubernetes state for the release namespace, and the remaining OpenRouter credits when the provider key has a configured credit limit.

Grafana is the only monitoring component exposed through ingress. Use a dedicated Slack webhook, and restrict access to the Grafana administrator password.

Communications exposes these internal endpoints:

Communications endpoints
GET http://<communications-host>:8002/health
GET http://<communications-host>:8002/metrics

/health is a process-level liveness and readiness response; it does not summarize every Connection. /metrics exports HTTP and low-cardinality Communications metrics. Use Communication diagnostics for per-Connection investigation. Neither endpoint should be public.

Separate production and staging

Use separate configuration and infrastructure identities for every environment.

Setting Production Staging
Namespace agent-farm agent-farm-staging
Environment label production staging
UI hostname Production hostname Staging hostname
API hostname Production hostname Staging hostname
Grafana hostname Production hostname Staging hostname
Sender domain mail. subdomain Separate mail-staging. subdomain
Image tags Production release tags Compatible staging tags
Databases and PVCs Production namespace Staging namespace
Platform stable keys Production values Separate staging values
Kubeconfig Production-scoped Staging-scoped

Do not reuse the production signing key, encryption key, database passwords, or LiteLLM master key in staging.

The OpenRouter inference key and the Cloudflare account can technically be shared, but doing so shares provider quota and failure impact. Use independent credentials when operational isolation matters.

To use an environment-specific file:

Deploy staging
ENV_FILE=.env.deploy.staging ./deploy.sh

Configuration variable reference

Required Kubernetes deployment inputs

Group Requirement Variables
Cluster Required KUBECONFIG, NAMESPACE
Registry Required REGISTRY_PREFIX, REGISTRY_SERVER, REGISTRY_USERNAME, REGISTRY_PASSWORD
Images Required API_IMAGE_REPOSITORY, UI_IMAGE_REPOSITORY, HERMES_IMAGE_REPOSITORY, OPENCLAW_IMAGE_REPOSITORY, and all four image tags
Application database Required POSTGRES_APP_USER, POSTGRES_APP_PASSWORD, POSTGRES_APP_DB
LiteLLM database Required POSTGRES_LITELLM_USER, POSTGRES_LITELLM_PASSWORD, POSTGRES_LITELLM_DB
Firecrawl database Required POSTGRES_FIRECRAWL_USER, POSTGRES_FIRECRAWL_PASSWORD, POSTGRES_FIRECRAWL_DB
Models Required LITELLM_MASTER_KEY, OPENROUTER_API_KEY
Application identity Required SECRET_SIGNING_KEY, AGENT_TOKEN_ENCRYPTION_KEY, PLATFORM_ADMIN_CREDENTIALS, ENVIRONMENT
URLs Required UI_HOST, API_HOST, WEB_APP_URL, GRAFANA_HOST
Firecrawl Required FIRECRAWL_API_KEY
Monitoring Required SLACK_ALERTS_WEBHOOK_URL, GRAFANA_ADMIN_PASSWORD, MONITORING_WEB_PASSWORD, GRAFANA_HOST

Conditional and optional inputs

Variable Requirement Purpose
POD_KUBECONFIG_B64 Optional Recommended for production A dedicated, namespace-scoped Kubernetes identity for the API
STORAGE_CLASS Conditional When the cluster default is unsuitable Select non-default persistent storage
AGENT_DEFAULT_MODEL Optional Override the built-in default model
AGENT_MODEL_ALLOWLIST Optional Narrow the platform model catalogue
CLOUDFLARE_ACCOUNT_ID Conditional Required with the other two email values Enable transactional email
CLOUDFLARE_API_TOKEN Conditional Required with the other two email values Cloudflare Email Sending authorization
SENDER_EMAIL Conditional Required with the other two email values The verified From address
GOOGLE_CLOUD_CLIENT_ID Conditional Required with the client secret Enable Google Workspace OAuth
GOOGLE_CLOUD_CLIENT_SECRET Conditional Required with the client ID The Google OAuth client secret
FIRECRAWL_BULL_AUTH_KEY Optional Protect the internal Firecrawl BullMQ administration interface
INGRESS_CLUSTER_ISSUER Optional Currently ineffective Present in the template, but not wired into the API or UI chart by Helmfile

Generated values

Value Requirement Generated by
Application database URL Generated Helmfile, from the POSTGRES_APP_* values
LiteLLM database URL Generated Helmfile, from the POSTGRES_LITELLM_* values
API external URL Generated Helmfile, from the first API_HOST
API pod kubeconfig Generated deploy.sh, from KUBECONFIG unless POD_KUBECONFIG_B64 is supplied
Agent Barn LiteLLM API key Generated The API chart pre-install and pre-upgrade hook
Per-Agent LiteLLM keys Generated The Agent creation workflow
Per-start Ingest key and Communications credential Generated The Agent start workflow
Registry pull Secret Generated The API chart
TLS Secrets Generated cert-manager

Do not manually create the generated per-Agent keys.

Apply configuration changes

Local changes require restarting the affected processes or containers:

Restart locally
./stop.sh
./run.sh --detach

Kubernetes changes require another deployment:

Redeploy
./deploy.sh

The API chart hashes Secret content into its pod template, so Secret changes roll the API and worker pods.

Changes that can require an Agent restart include runtime image changes, LiteLLM, Ingest, or Communications addresses, generated Runtime credential material, and Template or Skill changes applied to the Agent. Connection policy or provider-session changes are reconciled by Communications and do not require restarting the Agent.

Validate the configuration

Before deployment, verify the Kubernetes context:

Context
kubectl config current-context
kubectl cluster-info

Confirm storage:

Storage
kubectl get storageclass

Confirm the ingress and certificate prerequisites:

Ingress and issuer
kubectl get ingressclass traefik
kubectl get clusterissuer letsencrypt-http01

Confirm namespace permissions:

Permissions
kubectl auth can-i create deployments --namespace agent-farm
kubectl auth can-i create services --namespace agent-farm
kubectl auth can-i create secrets --namespace agent-farm
kubectl auth can-i create persistentvolumeclaims --namespace agent-farm

Confirm that DNS resolves:

DNS
dig +short agentbarn.example.com
dig +short api.agentbarn.example.com
dig +short grafana.agentbarn.example.com

Confirm the monitoring chart is prepared:

Chart dependencies
helm dependency build helm/monitoring

Review the deployment file without printing its secret values:

List variable names
grep -E '^[A-Z0-9_]+=' .env.deploy | cut -d= -f1 | sort

After deployment, verify the release:

Release state
helm list --namespace agent-farm
kubectl get deployments,statefulsets,pods --namespace agent-farm
kubectl get pvc,ingress,certificate --namespace agent-farm

Verify the API:

Health check
curl --fail https://api.agentbarn.example.com/api/v1/health

A healthy response resembles:

Response
{
  "status": "ok",
  "db": "connected"
}

Security and rotation

Apply these rules to the configuration:

  • Never commit .env or .env.deploy
  • Store the stable keys in an encrypted secret manager with tested recovery access
  • Use a namespace-scoped kubeconfig for the API
  • Use registry access tokens instead of personal passwords
  • Give Cloudflare tokens only the required Email Sending permission
  • Use unique passwords for each of the three databases
  • Keep secrets out of ENVIRONMENT, hostnames, model allowlists, and other visible values
  • Do not render Helm Secrets into shared CI logs
  • Restrict access to deployment artifacts and backups
  • Back up PostgreSQL and the stable keys together
  • Test recovery using the restored keys and database data
  • Rotate provider and registry tokens independently from the stable encryption keys
  • Treat stable-key rotation as a migration with rollback and verification steps
  • DB_CONNECTION_URL points to the shared application database
  • INGEST_BASE_URL reaches port 8001 and COMMUNICATIONS_BASE_URL reaches port 8002, each with its versioned path
  • API_EXTERNAL_URL is the correct public HTTPS origin, while Runtime and Driver routes remain private
  • Public ingress exposes only the Communications provider-webhook prefix
  • Provider-validation bypasses are disabled and provider credentials remain on Connections rather than in environment files
  • Journal retention is intentional and Prometheus reaches the internal Communications metrics endpoint

A database backup without AGENT_TOKEN_ENCRYPTION_KEY cannot recover encrypted credentials. A LiteLLM database backup without its original LITELLM_MASTER_KEY can leave its virtual-key data unusable.

Current configuration constraints

Account for these current limitations:

  1. .env.deploy.spec does not list SLACK_ALERTS_WEBHOOK_URL, GRAFANA_ADMIN_PASSWORD, MONITORING_WEB_PASSWORD, or GRAFANA_HOST, but Helmfile requires them.
  2. INGRESS_CLUSTER_ISSUER is present in .env.deploy.spec but is not passed into the API or UI chart.
  3. deploy.sh does not build the monitoring chart dependency.
  4. deploy.sh applies the production k8s/agent-farm-user.yaml bootstrap manifest regardless of a changed NAMESPACE.
  5. The API-facing kubeconfig defaults to the deployment kubeconfig, which may be more privileged than the API requires.
  6. The deployment uses in-cluster LiteLLM, Ingest, Firecrawl, and Redis addresses that are not all exposed as .env.deploy overrides.
  7. Some api/core/config.py settings are not wired through the Helm deployment and are not supported .env.deploy inputs.
  8. The namespace and associated ServiceAccount names intentionally retain agent-farm naming after the product and repository rebrand.

Troubleshooting

Symptom Likely cause Resolution
run.sh reports missing variables A required local value is blank Fill the named values in .env and run the launcher again.
API startup fails while initializing data The bootstrap password does not satisfy the policy Use at least eight characters with an uppercase letter, a lowercase letter, and a digit.
Changing PLATFORM_ADMIN_CREDENTIALS does not change the login The bootstrap administrator already exists Change the password through the supported account workflow instead.
Existing credentials can no longer be decrypted AGENT_TOKEN_ENCRYPTION_KEY changed Restore the original key, or perform a planned credential migration.
Existing Agent model keys stop working LITELLM_MASTER_KEY or the LiteLLM database changed Restore the original master key together with the LiteLLM database.
An Agent starts but never answers LiteLLM is unreachable, or the OpenRouter key is invalid Verify the Agent-facing LiteLLM address, proxy health, and the provider key.
An Agent answers but its Activity stays empty The Agent cannot reach the Ingest API Verify the Ingest base URL, port 8001, versioned path, DNS, and any NetworkPolicy.
Provider messages do not become Conversations Communications cannot receive, admit, or deliver a Connection event Verify Communications on port 8002, the Connection, internal Runtime reachability, and the webhook ingress prefix where applicable.
An Agent pod uses the wrong image Image tags are incompatible or stale Configure the release’s matching runtime tags, then stop and start the Agent.
Kubernetes reports ImagePullBackOff The registry credentials, prefix, repository, or tag is wrong Inspect the pod events and the generated pull Secret.
PersistentVolumeClaims remain Pending STORAGE_CLASS is missing or incompatible Select a dynamically provisioned ReadWriteOnce StorageClass.
PostgreSQL rejects its configured password after an edit The Kubernetes Secret changed but the initialized database user did not Perform a coordinated database credential rotation.
Email sends are logged but never delivered One or more of the three email values is missing Set all three Cloudflare values and verify the sender domain.
Cloudflare rejects the email request The token lacks permission, or the sender domain is unverified Grant Email Sending permission and complete domain verification.
Google OAuth reports a redirect mismatch The registered URI does not exactly match the WEB_APP_URL origin Register the exact callback URI and deploy again.
Event Deliveries remain Pending Redis, the worker, or reconciliation is unavailable Check Redis, the worker Deployment, and the reconciliation CronJob.
Certificates remain not ready DNS, ingress, or the ClusterIssuer is incorrect Inspect the Ingress, Certificate, Challenge, and ClusterIssuer resources.
Changing INGRESS_CLUSTER_ISSUER has no effect The variable is not wired into the charts Use the currently supported issuer, or update the deployment wiring.
Staging Agents appear in the production namespace The namespace or API kubeconfig wiring is wrong Verify the release namespace, K8S_NAMESPACE, and the staging bootstrap manifest.
Platform changes apply but running Agents are unchanged Agent workloads keep the configuration generated at their last start Stop and start the affected Agents.

Next steps

  1. Confirm every required value is present and every stable key is backed up.
  2. Validate the cluster, storage, ingress, and DNS prerequisites.
  3. Apply the configuration to a cluster.
  4. Verify the API health endpoint and the generated Agent workloads.
  5. Continue to Deploy Agent Barn to Kubernetes.

Staging and custom namespaces

Provision the target namespace and its deployment/runtime identities separately. Ensure the selected ServiceAccount, its permissions, the LiteLLM key bootstrap Job, monitoring, and the pod kubeconfig agree on that namespace. Then apply Helmfile with the target environment's complete configuration. An existing environment file alone does not supply these prerequisites.

For AAI Labs' k3s staging environment, use the configured deploy.yml staging workflow after provisioning its prerequisites. It is separate from the public-tag deployment. The staging manifest currently defines agent-farm-user, while the workflow selects agent-farm-staging-user for the LiteLLM key Job; those names must be aligned before treating the manifest as complete workflow bootstrap.

Documentation