Interpret Agent lifecycle and Runtime health, inspect live and retained Runtime logs, and determine when a failure belongs to a Communication Connection instead.
Runtime health and Connection health are independent. A running, healthy Runtime can have an errored Connection, and a connected provider session does not prove that the Runtime can process a Delivery.
Overview
The Agent detail page and each Communication Connection expose separate diagnostic planes. Start with the signal that owns the symptom.
| Plane | Signals |
|---|---|
| Agent Runtime | Persisted lifecycle (STOPPED, RUNNING, or ERROR), displayed Runtime condition, Runtime error reason, live logs, retained snapshots, and Tool Calls written through Ingest. |
| Communication Connection | Enabled state; observed provider state (pending, connected, degraded, or error); end-to-end Delivery health; counts; queue depth and oldest pending age; safe failures; incidents; provider-state history; and Delivery transitions, retries, reconnects, and dead letters. |
| Communications record | Canonical Conversation messages for the Connection, location, and thread. |
Runtime health and logs require access to the Agent and activity.read. Connection summary and journal reads require visibility of the Agent through agent.read. Reconnect and Delivery retry actions require agent.update.
Runtime health model
Agent Barn combines persisted lifecycle with the Runtime health probe to produce the condition shown for the Agent. It does not include Communication Connection provider state.
| Persisted lifecycle | Runtime health | Displayed condition | Meaning |
|---|---|---|---|
STOPPED | Not queried | Idle | The Agent is intentionally stopped. |
RUNNING | Starting or not ready | Initializing | Runtime resources or its health endpoint are not ready yet. |
RUNNING | Healthy | Working | The Runtime health endpoint reports healthy. |
RUNNING | Unreachable or unhealthy | Disconnected | Agent Barn cannot currently confirm Runtime health. |
RUNNING or ERROR | Crashed or stored failure | Needs attention | The Runtime workload crashed or a lifecycle operation failed. |
- Idle means the Agent is intentionally
STOPPED; no active Runtime is expected. - Initializing means Runtime resources or the Runtime health endpoint are not ready yet.
- Working means the Runtime health endpoint reports healthy.
- Disconnected means the Agent is persisted as
RUNNING, but Runtime health cannot be confirmed. - Needs attention means the Runtime workload crashed or the Agent is persisted in
ERROR.
Runtime causes include image pulls, missing Runtime configuration, model or LiteLLM failures, pod readiness, crashes, out-of-memory termination, Kubernetes access, and unexpected process exits.
Diagnostic workflow
- Determine whether the symptom is Runtime execution or message delivery.
- Check the persisted Agent lifecycle and displayed Runtime condition.
- Inspect the Runtime error reason and logs when the Runtime is not healthy.
- Open the affected Communication Connection when the Runtime is healthy but messages are missing, delayed, rejected, or failing.
- Use Conversations for canonical message history and Tool Calls for Runtime execution telemetry.
- Apply recovery only to the failing resource.
Open the Agent and Connection
- Select the Organization that owns the Agent.
- Open the Agent from Home.
- Review its lifecycle, displayed Runtime condition, and error reason.
- Open Logs for Runtime evidence.
- Open the affected Communication Connection for provider and Delivery evidence.
The Agent detail page exposes Conversations, Tool calls, Logs, and Work to users with activity.read. The Connection detail surface exposes its own summary diagnostics and Delivery transitions.
Interpret Runtime health
| Condition | First action |
|---|---|
| Idle | Review the latest snapshot; start only if the Agent should run. |
| Initializing | Inspect Runtime deployment, readiness, configuration, and health. |
| Working | Investigate the specific Conversation, Tool Call, or Connection symptom. |
| Disconnected | Read the Runtime error reason and inspect live logs. |
| Needs attention | Inspect Runtime logs and the lifecycle failure before changing the Agent. |
For an extended Initializing or Disconnected condition, inspect image pulls, Runtime configuration, Template and Skill loading, model or LiteLLM access, Kubernetes readiness, and unexpected Runtime exits. A Working condition confirms Runtime availability only; it is not a provider connectivity or Delivery guarantee.
Review live Runtime logs
When an Agent is RUNNING, the Logs tab initially loads recent Runtime output and then opens a server-sent event stream for new lines. While connected, the header displays Streaming.
[Runtime startup output] [Template and Skill loading] [Runtime operation output] [model, tool, or Runtime error output]
This is a visual example of the interface layout. The lines are placeholders, not product output.
Runtime logs are appropriate for startup and shutdown, Template and Skill loading, LiteLLM or model errors, Tool Integration failures, Runtime exceptions, Kubernetes workload failures, and the shared Communications adapter's interaction with the local Runtime endpoint. They are not the authoritative record of Connection health or Communication Delivery state.
Follow new output
The viewer follows the newest line while you remain at the bottom. Scrolling upward pauses automatic following; select Jump to latest to resume.
Find the first Runtime failure
Read from the start of the affected session and use focused search terms:
error
failed
invalid
unauthorized
timeout
CrashLoopBackOff
OOMKilled
ImagePullBackOffThe first error is often more useful than later effects of the same failure. The browser retains up to 10,000 displayed lines; this browser buffer is separate from snapshot retention.
Review retained Runtime logs
When an Agent is not RUNNING, Logs shows the latest retained Runtime session. Agent Barn attempts a best-effort snapshot before stopping an Agent, including a stop performed by Apply & Restart.
Scroll toward the beginning to load older sessions. The viewer paginates retained snapshots and separates sessions with an end marker, so you can compare a working start with the first failing start or configuration before and after a Runtime change.
Correlate Runtime and Communications activity
Conversations
Answer whether Communications persisted the inbound or outbound message for the correct Connection, location, and thread.
Tool calls
Answer whether Runtime tool execution began, succeeded, or failed. Ingest stores Tool Calls, not Conversation messages.
Runtime logs
Answer what Hermes or OpenClaw reported during execution.
Connection diagnostics
Answer whether the provider session connected and how a durable Delivery moved through Communications.
If a message is missing, open the affected Connection's details and use Communication diagnostics. Do not infer provider admission or Delivery state from Runtime logs.
When the problem is a Communication Connection
Each Connection has independent diagnostics for provider connectivity, end-to-end health, last successful provider connection, current error age and consecutive failures, Delivery success rate, queue depth and oldest pending Delivery, a Connection-state timeline and incidents, recent structured content-free failures, and Delivery transitions and dead-letter state.
Safe failure details can include category, operation, HTTP status, provider code, retryability, bounded retry-after, and provider request ID. They never expose provider URLs, headers, bodies, credentials, message content, or raw exception text.
Use Communication diagnostics for the complete workflow.
Diagnostic decision table
| Case | Next action |
|---|---|
Agent STOPPED | Review the latest Runtime snapshot; start only if it should run. |
Agent ERROR or Runtime Needs attention | Inspect Runtime logs and the lifecycle failure. |
| Runtime Initializing or Disconnected | Inspect Runtime deployment and health. |
| Runtime Working, Connection pending, degraded, or error | Inspect Connection health and incidents. |
| Runtime Working, Connection connected, missing message | Inspect policy admission and Delivery transitions. |
| Dead-lettered outbound Delivery | Inspect safe failure details and retry only when eligible. |
| Conversation exists but Tool Call failed | Inspect Tool Calls and Runtime logs. |
Recover the failing resource
Runtime failures
- Inspect logs before stopping the Agent.
- Correct Runtime configuration, Template, Skills, model, or Tool Integration credentials.
- Use the Agent lifecycle workflow to stop and start when required.
- Stopping attempts a best-effort Runtime log snapshot.
Agent Secrets belong to Runtime and Tool Integration configuration. Communication Connection credentials are encrypted Connection credentials managed independently.
Connection failures
- Refresh status to read the most recently observed supervisor state; it does not actively probe the provider.
- Request a reconnect for an enabled Connection when its provider session must be recreated. Reconnecting preserves queued Deliveries and does not duplicate them.
- Retry only an active, dead-lettered outbound Delivery. Retry preserves its message and stable idempotency identity.
- Do not restart the Agent merely to reconnect a provider session. Connection changes do not require a Runtime restart.
Runtime log retention
- Snapshot timing
- Captured before the Agent stops, on a best-effort basis
- Retained sessions
- The newest five snapshots
- Lines requested during capture
- Up to 50,000 recent lines
- Maximum stored snapshot size
- One MiB; newest content is retained when truncated
- Live browser buffer
- Up to 10,000 displayed lines
Older snapshots are deleted as newer snapshots are saved. Use external observability for permanent retention.
Runtime health, logs, and Connection diagnostics API
Agent Runtime endpoints
These endpoints require Agent access and activity.read.
| Method and endpoint | Purpose |
|---|---|
GET /agents/{agent_id}/healthz | Return current Runtime health for a running Agent, or the stored reason for an Agent in error. |
GET /agents/{agent_id}/logs | Return recent live Runtime lines, or the latest snapshot when not running. |
GET /agents/{agent_id}/logs/stream | Stream live Runtime lines through server-sent events. |
GET /agents/{agent_id}/logs/history | Return a retained snapshot and the next older snapshot identifier. |
Connection diagnostics endpoints
Method and endpoint Purpose GET /agents/{agent_id}/connections/{connection_id}/summaryRead Connection health, provider state, Delivery signals, and incidents. GET /agents/{agent_id}/connections/{connection_id}/journalRead Connection journal and Delivery transitions. POST /agents/{agent_id}/connections/{connection_id}/reconnectRequest Connection supervisor reconnect. POST /agents/{agent_id}/connections/{connection_id}/deliveries/{delivery_id}/retryRetry an eligible dead-lettered outbound Delivery.
These routes sit beneath the active Organization API base path.
Health response
{
"status": "ok",
"reason": null
}
Recent Runtime logs
GET /agents/{agent_id}/logs accepts tail_lines from 1 through 10000, defaulting to 100.
{
"lines": ["runtime output"],
"source": "live",
"has_snapshots": true,
"snapshot_id": null,
"session_started_at": null,
"session_ended_at": null
}
GET /agents/{agent_id}/logs/stream returns text/event-stream; its optional tail_lines accepts 0 through 1000. Snapshot history uses next_snapshot_id until has_more is false.
Log security
Troubleshooting
Runtime remains Initializing or Disconnected
Inspect deployment readiness, image pulls, Runtime configuration, model or LiteLLM access, Kubernetes access, and the Runtime health endpoint. Correct the Runtime cause before using the lifecycle workflow.
Runtime is Working but a message is missing
Open the affected Connection. Inspect policy admission, canonical Conversations, and Delivery transitions before changing or restarting the Agent Runtime.
A Tool Call failed
Inspect the Tool Call and Runtime logs. Correct the Tool Integration or Runtime credential/configuration issue; do not confuse it with encrypted Communication Connection credentials.
No Runtime logs are available after stopping
Snapshot capture is best effort. Use Tool Calls, Connection diagnostics where relevant, deployment observability, and external log aggregation.