ipynb notebooks

How AI agents are reshaping the data scientist's toolset

Jupyter Notebooks gave data scientists a linear home for a linear workflow.

3 min readMachine Learning

The notebook cell was a good container for a bygone workflow, but it is no longer the right one for the work that data scientists actually do. A data scientist on Reddit recently asked whether Jupyter notebooks are still good enough or if we are just used to them, and the answer is increasingly clear: the abstraction of "code → output" is holding us back when AI agents can already write most of that code for us.

The classical pipeline, EDA, data prep, fit, evaluate, tune, save, mapped neatly onto notebook cells because each step was a discrete, human-written block of logic. Today, tools like Claude and Codex can generate those blocks on demand, which should free the data scientist to focus on hypothesis formation and decision-making. Instead, we are still trapped in a linear cell structure designed for manual execution. The real question is not whether notebooks are obsolete; it is whether we are willing to replace the cell with a better abstraction. The user who posed the question suggested moving to "prompt → result" as the new unit of work, and that idea deserves serious attention. When the model can write the code, the container should be the question, not the syntax. This shift has practical consequences for how we Build Guardrails for Safer AI Data Access, because an agent that responds to prompts must be constrained by architectural guardrails that a static notebook never required.

What makes this moment different is that the data scientist's toolset is no longer just code, it is interaction. A notebook presumes you will read the output, write the next cell, and repeat. An agentic workflow presumes you will state an intent, review a result, and refine the prompt. That is a fundamentally different rhythm, and it demands tools that are built for conversation, not for linear execution. Data scientists who adopt this approach can Build a complete data science stack for free with these open-source AI tools, because the shift to prompt-driven work does not require expensive infrastructure, it requires thinking about the cell as a dialogue unit rather than a code block. The open-source ecosystem already supports this model, and the barrier to trying it is lower than most realize.

The specific consequence to watch is how this changes the role of the data scientist in classical ML applications. If the agent writes the code and the human writes the prompts, then the bottleneck shifts from implementation to experimentation design. The data scientist who masters prompt-driven workflows will run more hypotheses per day, not because they type faster, but because they stop organizing work around code cells that the machine can fill. That is the real opportunity, and the real challenge, because it means the notebook's greatest strength, its explicitness, becomes its greatest limitation. The tools that emerge next will not look like notebooks at all.

From Machine Learning

I am a data scientist who started working in the industry before the LLM revolution. Back then, Jupyter Notebooks were a perfect fit for the classical DS pipeline: EDA -> data prep -> fit -> eval -> tune -> save model artefact and notebook.

Lately, I have been thinking a lot about how agentic development and LLMs are changing the way data scientists work. Especially in classical ML applications, where you still need to explore data, run experiments, check different hypotheses and decide what to do next based on the results.

Read the original at Machine Learning