Example setup · Technical operations

How agents and people can run one incident room

Logs, repos, dashboards, server agents, and responders stayed together without one crowded chat.

A human operator supervising several active agents in a shared technical workshop
Find why the jobs failed. Protect production, test the fix, and keep a clear timeline.

Why this gets messy

Incident work happens in parallel. One person checks logs, another checks the last deploy, another prepares a rollback, and someone else shares updates.

A group chat records the conversation, but it does not show which agent is on which machine or whether two workers are about to make conflicting changes.

Put the whole job in one place

October put the agents, repos, dashboards, timeline, and human responders in one shared room. Each worker stayed separate but saw the updates it needed.

Give each agent one clear job

Log analysis

Goose · Production observer. Read limited logs and reported patterns without changing production.

Regression search

Codex · Application repository. Found likely changes and prepared a small fix in its own worktree.

Verification

Claude Code · Staging environment. Reproduced the failure and tested the proposed fix.

Incident record

Hermes · Shared timeline. Kept the decisions, evidence, status changes, and follow-up tasks.

Keep every handoff visible

Protect production first

Incident lead. The lead named the affected systems, banned risky actions, and set the proof needed before any change.

Investigate in parallel

Specialist agents. Each worker took one clear question. Everyone could see the findings and failed guesses.

Test the smallest fix

Repository and verification agents. The proposed fix went to staging with its evidence. The operator could stop duplicate or risky work.

Recover and save the timeline

Human lead. A person approved recovery. The shared room already held the timeline, open risks, and follow-up tasks.

You still make the important calls

Reading production and changing production required different access.

Every active task had a visible owner.

The human lead approved production changes and public updates.

What changed

Responders could see the whole incident, not just the chat. New evidence reached the right agents while the lead kept control of production.

When the incident ended, the timeline, decisions, evidence, and follow-up work were already in one place.

All case studies · Download October