会持续学习的 Incident CopilotThe Incident Copilot That Teaches Back
Platform Reliability Guild
一套带有安全护栏的 AI 故障排查流程,减少了在 Runbook 之间反复搜索的时间,并把每次已解决事件沉淀成下一支团队更清晰的诊断路径。A guarded AI troubleshooting workflow reduced time spent searching across runbooks and turned every resolved incident into a better diagnostic path for the next team.
值班工程师使用监控面板、日志和现有 Runbook 诊断事件,并在服务恢复目标内完成处置。On-call engineers diagnose incidents using dashboards, logs, and existing runbooks while meeting the service recovery target.
团队没有直接上线无约束聊天机器人,而是构建了经过评估的检索流程,加入引用核验、置信度提示和人工决策节点。The guild built an evaluated retrieval workflow with citation checks, confidence prompts, and a human decision point instead of deploying an unguarded chatbot.
在试点样本中,获得有效诊断的中位时间从 42 分钟降至 18 分钟,同时减少 Runbook 搜索成本,且没有引入任何自动生产操作。Median time to a useful diagnosis fell from 42 to 18 minutes across the pilot set, while runbook search effort dropped and no autonomous production action was introduced.
评估集、提示词模式、安全检查表和事件后学习闭环,现已成为其他运营知识场景可复用的起点。The evaluation set, prompt pattern, safety checklist, and post-incident learning loop are now a reusable starter for other operational knowledge domains.
- 24 起事件的复盘样本24-incident retrospective sample
- 诊断中位时间由 42 分钟降至 18 分钟42 → 18 min median diagnosis
- 带引用回答接受率 87%87% cited-answer acceptance
- 自动生产操作为 00 autonomous production actions






