Thinking in DAGs: How Humans and Agents Collaborate in Data Science

Chen Li

As LLM agents increasingly participate in data science, a critical question emerges: what abstraction best supports humans and agents in jointly constructing, understanding, and revising analyses? While scripts and notebooks are expressive and detailed, they often obscure data dependencies within mutable state, overwhelm agents with excessive low-level code, and hinder human inspection of intermediate steps. However, natural language, despite its ease of use, frequently suffers from inherent ambiguity and imprecision.

This talk proposes the directed acyclic graph (DAG) as a superior abstraction for human-agent collaborative data science. By modeling analyses as operators connected by explicit data dependencies, DAGs provide a compact structure for agent reasoning, a visual interface for human oversight, and a foundation for execution, optimization, and provenance. We illustrate this approach with our work called BobFlow, demonstrating how agents can iteratively construct dataflows to solve data science tasks and manage the context using dataflow-level metadata. We show how an analyst easily understands the agent’s behaviors through an intuitive dataflow-based interface, how the analyst efficiently inspects the agent’s past actions to provide feedback, and how the agent finishes a task with high accuracy and low cost. We also address key implementation challenges, including seamless migration between scripts and workflows, efficient caching of intermediate results, managing reproducible environments, and control blocks such as IfElse and ForLoop. We showcase our progress within Apache Texera (Incubating), demonstrating how these techniques foster transparent, efficient, and trustworthy human-agent collaboration on data science.

PDF