Jihun Yun, Aurelie Lozano, et al.
NeurIPS 2021
Agentic systems operating over heterogeneous data ecosystems face data management challenges distinct from single-turn LLM inference. When a pipeline ingests large artifacts---log streams, database query results, document corpora---into a monolithic context window, the result is unbounded, non-queryable state that is expensive to recover on failure.
We present \emph{structured state management} (SSM), a data-centric architectural pattern in which each pipeline stage reads only its declared upstream keys, every stage output is schema-validated before commit to a persistent state store, and the execution engine performs selective recovery---retrying only the failed stage over its bounded context.
We evaluate SSM on two data-intensive workloads: Kubernetes root cause analysis (84,000-record input artifact) and multi-file code debugging. Across three configurations (monolithic, static-decomposed, SSM), we find that static decomposition \emph{increases} retry cost by 80.5% over monolithic ( vs.\ tokens) due to cascading re-reads of upstream state, while SSM reduces retry cost by 73.2% over static decomposition and 51.7% over monolithic ( tokens). Schema contracts also create structured human-in-the-loop intervention points at every stage boundary.
Jihun Yun, Aurelie Lozano, et al.
NeurIPS 2021
Ge Gao, Xi Yang, et al.
AAAI 2024
Imran Nasim, Michael E. Henderson
Mathematics
Daniele Lotito
Dynamical Systems in Lecce 2025