Agents
How-to

Review Agent health and logs

Interpret Agent lifecycle and Runtime health, inspect live and retained Runtime logs, and determine when a failure belongs to a Communication Connection instead.

For
Agent operators, Agent viewers, Agent editors, Agent owners, Organization administrators, and support engineers
On this page
  1. Overview
  2. Runtime health model
  3. Diagnostic workflow
  4. 1. Open the Agent and Connection
  5. 2. Interpret Runtime health
  6. 3. Review live Runtime logs
  7. 4. Review retained Runtime logs
  8. 5. Correlate Runtime and Communications activity
  9. When the problem is a Communication Connection
  10. Diagnostic decision table
  11. Recover the failing resource
  12. Runtime log retention
  13. Runtime health, logs, and Connection diagnostics API
  14. Log security
  15. Troubleshooting
  16. Next steps

Interpret Agent lifecycle and Runtime health, inspect live and retained Runtime logs, and determine when a failure belongs to a Communication Connection instead.

Runtime health and Connection health are independent. A running, healthy Runtime can have an errored Connection, and a connected provider session does not prove that the Runtime can process a Delivery.

Overview

The Agent detail page and each Communication Connection expose separate diagnostic planes. Start with the signal that owns the symptom.

PlaneSignals
Agent RuntimePersisted lifecycle (STOPPED, RUNNING, or ERROR), displayed Runtime condition, Runtime error reason, live logs, retained snapshots, and Tool Calls written through Ingest.
Communication ConnectionEnabled state; observed provider state (pending, connected, degraded, or error); end-to-end Delivery health; counts; queue depth and oldest pending age; safe failures; incidents; provider-state history; and Delivery transitions, retries, reconnects, and dead letters.
Communications recordCanonical Conversation messages for the Connection, location, and thread.

Runtime health and logs require access to the Agent and activity.read. Connection summary and journal reads require visibility of the Agent through agent.read. Reconnect and Delivery retry actions require agent.update.

Runtime health model

Agent Barn combines persisted lifecycle with the Runtime health probe to produce the condition shown for the Agent. It does not include Communication Connection provider state.

Persisted lifecycleRuntime healthDisplayed conditionMeaning
STOPPEDNot queriedIdleThe Agent is intentionally stopped.
RUNNINGStarting or not readyInitializingRuntime resources or its health endpoint are not ready yet.
RUNNINGHealthyWorkingThe Runtime health endpoint reports healthy.
RUNNINGUnreachable or unhealthyDisconnectedAgent Barn cannot currently confirm Runtime health.
RUNNING or ERRORCrashed or stored failureNeeds attentionThe Runtime workload crashed or a lifecycle operation failed.
  • Idle means the Agent is intentionally STOPPED; no active Runtime is expected.
  • Initializing means Runtime resources or the Runtime health endpoint are not ready yet.
  • Working means the Runtime health endpoint reports healthy.
  • Disconnected means the Agent is persisted as RUNNING, but Runtime health cannot be confirmed.
  • Needs attention means the Runtime workload crashed or the Agent is persisted in ERROR.

Runtime causes include image pulls, missing Runtime configuration, model or LiteLLM failures, pod readiness, crashes, out-of-memory termination, Kubernetes access, and unexpected process exits.

Diagnostic workflow

  1. Determine whether the symptom is Runtime execution or message delivery.
  2. Check the persisted Agent lifecycle and displayed Runtime condition.
  3. Inspect the Runtime error reason and logs when the Runtime is not healthy.
  4. Open the affected Communication Connection when the Runtime is healthy but messages are missing, delayed, rejected, or failing.
  5. Use Conversations for canonical message history and Tool Calls for Runtime execution telemetry.
  6. Apply recovery only to the failing resource.

Open the Agent and Connection

  1. Select the Organization that owns the Agent.
  2. Open the Agent from Home.
  3. Review its lifecycle, displayed Runtime condition, and error reason.
  4. Open Logs for Runtime evidence.
  5. Open the affected Communication Connection for provider and Delivery evidence.

The Agent detail page exposes Conversations, Tool calls, Logs, and Work to users with activity.read. The Connection detail surface exposes its own summary diagnostics and Delivery transitions.

Interpret Runtime health

ConditionFirst action
IdleReview the latest snapshot; start only if the Agent should run.
InitializingInspect Runtime deployment, readiness, configuration, and health.
WorkingInvestigate the specific Conversation, Tool Call, or Connection symptom.
DisconnectedRead the Runtime error reason and inspect live logs.
Needs attentionInspect Runtime logs and the lifecycle failure before changing the Agent.

For an extended Initializing or Disconnected condition, inspect image pulls, Runtime configuration, Template and Skill loading, model or LiteLLM access, Kubernetes readiness, and unexpected Runtime exits. A Working condition confirms Runtime availability only; it is not a provider connectivity or Delivery guarantee.

Review live Runtime logs

When an Agent is RUNNING, the Logs tab initially loads recent Runtime output and then opens a server-sent event stream for new lines. While connected, the header displays Streaming.

Live logsStreaming
[Runtime startup output]
[Template and Skill loading]
[Runtime operation output]
[model, tool, or Runtime error output]
Jump to latest

This is a visual example of the interface layout. The lines are placeholders, not product output.

Runtime logs are appropriate for startup and shutdown, Template and Skill loading, LiteLLM or model errors, Tool Integration failures, Runtime exceptions, Kubernetes workload failures, and the shared Communications adapter's interaction with the local Runtime endpoint. They are not the authoritative record of Connection health or Communication Delivery state.

Follow new output

The viewer follows the newest line while you remain at the bottom. Scrolling upward pauses automatic following; select Jump to latest to resume.

Find the first Runtime failure

Read from the start of the affected session and use focused search terms:

Runtime log search terms
error
failed
invalid
unauthorized
timeout
CrashLoopBackOff
OOMKilled
ImagePullBackOff

The first error is often more useful than later effects of the same failure. The browser retains up to 10,000 displayed lines; this browser buffer is separate from snapshot retention.

Review retained Runtime logs

When an Agent is not RUNNING, Logs shows the latest retained Runtime session. Agent Barn attempts a best-effort snapshot before stopping an Agent, including a stop performed by Apply & Restart.

Scroll toward the beginning to load older sessions. The viewer paginates retained snapshots and separates sessions with an end marker, so you can compare a working start with the first failing start or configuration before and after a Runtime change.

Correlate Runtime and Communications activity

Conversations

Answer whether Communications persisted the inbound or outbound message for the correct Connection, location, and thread.

Tool calls

Answer whether Runtime tool execution began, succeeded, or failed. Ingest stores Tool Calls, not Conversation messages.

Runtime logs

Answer what Hermes or OpenClaw reported during execution.

Connection diagnostics

Answer whether the provider session connected and how a durable Delivery moved through Communications.

If a message is missing, open the affected Connection's details and use Communication diagnostics. Do not infer provider admission or Delivery state from Runtime logs.

When the problem is a Communication Connection

Each Connection has independent diagnostics for provider connectivity, end-to-end health, last successful provider connection, current error age and consecutive failures, Delivery success rate, queue depth and oldest pending Delivery, a Connection-state timeline and incidents, recent structured content-free failures, and Delivery transitions and dead-letter state.

Safe failure details can include category, operation, HTTP status, provider code, retryability, bounded retry-after, and provider request ID. They never expose provider URLs, headers, bodies, credentials, message content, or raw exception text.

Use Communication diagnostics for the complete workflow.

Diagnostic decision table

CaseNext action
Agent STOPPEDReview the latest Runtime snapshot; start only if it should run.
Agent ERROR or Runtime Needs attentionInspect Runtime logs and the lifecycle failure.
Runtime Initializing or DisconnectedInspect Runtime deployment and health.
Runtime Working, Connection pending, degraded, or errorInspect Connection health and incidents.
Runtime Working, Connection connected, missing messageInspect policy admission and Delivery transitions.
Dead-lettered outbound DeliveryInspect safe failure details and retry only when eligible.
Conversation exists but Tool Call failedInspect Tool Calls and Runtime logs.

Recover the failing resource

Runtime failures

  • Inspect logs before stopping the Agent.
  • Correct Runtime configuration, Template, Skills, model, or Tool Integration credentials.
  • Use the Agent lifecycle workflow to stop and start when required.
  • Stopping attempts a best-effort Runtime log snapshot.

Agent Secrets belong to Runtime and Tool Integration configuration. Communication Connection credentials are encrypted Connection credentials managed independently.

Connection failures

  • Refresh status to read the most recently observed supervisor state; it does not actively probe the provider.
  • Request a reconnect for an enabled Connection when its provider session must be recreated. Reconnecting preserves queued Deliveries and does not duplicate them.
  • Retry only an active, dead-lettered outbound Delivery. Retry preserves its message and stable idempotency identity.
  • Do not restart the Agent merely to reconnect a provider session. Connection changes do not require a Runtime restart.

Runtime log retention

Retention per Agent
Snapshot timing
Captured before the Agent stops, on a best-effort basis
Retained sessions
The newest five snapshots
Lines requested during capture
Up to 50,000 recent lines
Maximum stored snapshot size
One MiB; newest content is retained when truncated
Live browser buffer
Up to 10,000 displayed lines

Older snapshots are deleted as newer snapshots are saved. Use external observability for permanent retention.

Runtime health, logs, and Connection diagnostics API

Agent Runtime endpoints

These endpoints require Agent access and activity.read.

Method and endpointPurpose
GET /agents/{agent_id}/healthzReturn current Runtime health for a running Agent, or the stored reason for an Agent in error.
GET /agents/{agent_id}/logsReturn recent live Runtime lines, or the latest snapshot when not running.
GET /agents/{agent_id}/logs/streamStream live Runtime lines through server-sent events.
GET /agents/{agent_id}/logs/historyReturn a retained snapshot and the next older snapshot identifier.

Connection diagnostics endpoints

Method and endpointPurpose
GET /agents/{agent_id}/connections/{connection_id}/summaryRead Connection health, provider state, Delivery signals, and incidents.
GET /agents/{agent_id}/connections/{connection_id}/journalRead Connection journal and Delivery transitions.
POST /agents/{agent_id}/connections/{connection_id}/reconnectRequest Connection supervisor reconnect.
POST /agents/{agent_id}/connections/{connection_id}/deliveries/{delivery_id}/retryRetry an eligible dead-lettered outbound Delivery.

These routes sit beneath the active Organization API base path.

Health response

Runtime health response
{
  "status": "ok",
  "reason": null
}

Recent Runtime logs

GET /agents/{agent_id}/logs accepts tail_lines from 1 through 10000, defaulting to 100.

Recent Runtime logs response
{
  "lines": ["runtime output"],
  "source": "live",
  "has_snapshots": true,
  "snapshot_id": null,
  "session_started_at": null,
  "session_ended_at": null
}

GET /agents/{agent_id}/logs/stream returns text/event-stream; its optional tail_lines accepts 0 through 1000. Snapshot history uses next_snapshot_id until has_more is false.

Log security

Troubleshooting

Runtime remains Initializing or Disconnected

Inspect deployment readiness, image pulls, Runtime configuration, model or LiteLLM access, Kubernetes access, and the Runtime health endpoint. Correct the Runtime cause before using the lifecycle workflow.

Runtime is Working but a message is missing

Open the affected Connection. Inspect policy admission, canonical Conversations, and Delivery transitions before changing or restarting the Agent Runtime.

A Tool Call failed

Inspect the Tool Call and Runtime logs. Correct the Tool Integration or Runtime credential/configuration issue; do not confuse it with encrypted Communication Connection credentials.

No Runtime logs are available after stopping

Snapshot capture is best effort. Use Tool Calls, Connection diagnostics where relevant, deployment observability, and external log aggregation.

Next steps

Documentation