Larry Carvalho, Anca Sailer, et al.
KubeCon EU 2025
When a multi-step LLM agent fails, the terminal accuracy score saysthatsomething1 went wrong but is silent onwhereandwhy. We address this diagnostic problem with2 a structural causal framework that localizes faults to specific nodes in the workflow3 DAG and distinguishes failures of insufficient context from failures of inadequate4 decision-making, mapping each diagnosis onto a well-defined intervention. We5 validate on IT-Bench, a Kubernetes incident diagnosis benchmark, across three6 models. For GPT-OSS-120B, the framework traces failures to a long-context7 bottleneck mid-workflow; addressing it lifts end-to-end task accuracy from 59% to8 74% — +15pp absolute / 27% relative improvement.
Larry Carvalho, Anca Sailer, et al.
KubeCon EU 2025
Gaetano Rossiello, Shankar Subramaniam
ACM CAIS 2026
Haoran Qiu, Weichao Mao, et al.
ASPLOS 2024
Matthew Arnold, Jeffrey Boston, et al.
MLSys 2020