The Pit of Success, and What a Data Science Language Would Need to Make It

[This article was first published on Econometrics and Free Software, and kindly contributed to R-bloggers]. (You can report issue about the content on this page here)
Want to share your content on R-bloggers? click here if you have a blog, or here if you don't.

Rico Mariani, a long-time performance engineer at Microsoft, coined the idea of the pit of success. A system has a pit of success when the natural, lazy way of using it leads to good outcomes. You don’t need heroics or perfect memory. You fall into the right result, and climbing out to do something wrong takes deliberate effort.

I came across this framing in a talk that I was recently recommended and I further recommend it to anyone interested in the topic: Functional architecture – The pits of success – Mark Seemann. It put a name to something I had been circling around for years.

I had been arguing a version of the same point long before I knew the term. In my 2022 post titled Functional programming explains why containerization is needed for reproducibility, I argued that a pipeline is only as reproducible as the functions it is built from, and that even a pure function hides inputs such as the R version, the package versions, and the operating system. Containers are one way to make those hidden inputs explicit. In another 2022 post on learnings from functional programming, I made a related case for writing functions that fail early and never reach into the global environment. Both posts were, in a sense, attempts to build a pit of success out of habits and tooling. Neither was framed that way, and I think the framing is more useful.

Pits of failure are everywhere in data science

Most of the tools we use for analysis are pits of failure. You can be meticulous and still end up with an analysis that cannot be rerun, because the meticulousness is entirely up to you.

Think about how a typical script behaves. It reads a file by path, and nothing says that the path is an input. It uses a package, and nothing records which version. It reassigns a variable halfway through, so the value at the bottom of the file depends on what happened at the top. It converts a column from text to numbers without complaint, and quietly turns the bad entries into missing values. Each of these is convenient in the moment. Together, they make the result depend on things nobody wrote down.

Good practice is then a set of instructions: pin your dependencies, avoid globals, check your inputs, document your steps. These are road signs. They tell you where to go, and they leave the driving to you. Under deadline pressure, road signs get ignored, and not out of malice. Only a speed bump would make you drive more responsibly.

What makes a pit of success

If the goal is to make the disciplined path the easy one, we need to ask what the language itself would have to do. If reproducibility depends on remembering a checklist, then reproducibility is a convention. If the language makes the checklist unnecessary, reproducibility becomes a property of the program. From my own experience of building and using R packages, and from thinking about where analyses usually go wrong, I think a language designed for this purpose would need at least the points raised in this blog post from earlier this year, but let me rephrase them here again.

The program should be a flat text file. Notebooks have become popular in data science, but I think they are a serious mistake. They mix code, output, and state in a format that is hard to diff, review, and reuse, and they encourage the hidden-state problems that make analyses irreproducible. Notebooks suffer from exactly the same problem that gets blamed on Excel. The main argument for notebooks is that they make exploration easy. But that says more about the shortcomings of the alternatives than about any real strength of notebooks. Environments like RStudio and Positron (or Emacs) already make interactive work with data natural, and they do it while keeping the program in a plain text file that can be run, tested, and versioned like any other code.

Reproducibility should be the default, not an option. If pinning an environment requires a separate file that someone has to remember to update, it will drift. A language that declares its runtime and dependencies as part of the program, and builds them the same way every time, removes a decision from the analyst. You cannot opt out of it by forgetting.

Inputs should be declared, not discovered. A program should not be able to read a file or an environment variable that it does not name. When every dependency is visible in the code, the tooling can see the whole picture: which step depends on which, and what changes when an input changes. Undeclared inputs become a build error rather than a silent difference between two machines. This is the same point as the hidden-inputs argument from my containerization post, moved from the operating system into the language.

Hidden state should be hard to create. Immutable bindings, explicit reassignment when you really need it, and no implicit global state make a result depend only on its inputs. This is also what makes caching safe. If a step is a pure function of its declared inputs, the system can reuse its previous output with confidence, and nobody has to wonder whether a cached result is stale.

Types and schemas should fail loudly. Silent coercion is one of the most common ways data goes wrong without anyone noticing. A language built for analysis should refuse to turn text into numbers without being asked, should make missing values explicit rather than inventing them, and should check that a column a downstream step needs actually exists before anything runs.

An analysis should be laid out as a pipeline. Traditional scripts tend to morph into spaghetti very quickly: steps depend on each other in ways that are hard to see, and a change in one place can silently affect results three sections later. Writing the analysis as an explicit pipeline, with each step a named stage that takes declared inputs and produces declared outputs, makes it much easier to debug a failing step, reuse results that have not changed, and hand the work over to someone (or something) else.

The whole graph should be checkable before it runs. If the pipeline is a first-class object that the language can inspect, you can catch missing columns, broken references, and cycles in milliseconds. Waiting three hours for a simulation to discover a typo in a column name is a failure of design, not of the analyst.

Errors should be values. When a step fails, the rest of the pipeline should not necessarily stop. An error that is captured as data can be inspected, reported, and handled, and independent branches can keep going. This turns debugging from reading a stack trace from code you didn’t write into reading a result that tells you where the problem is.

Interoperability should not require glue code. Real analysis mixes R, Python, Julia, SQL, and whatever comes next. Every hand-written CSV export between two of them is a place where types get lost and assumptions go unrecorded. A shared, typed interchange format between runtimes removes a whole category of errors and a lot of boilerplate.

Feedback should be structured. Error messages that are only meant for a human reading a terminal are hard for tools to act on. If the language reports problems in a machine-readable form, the same checks serve analysts, continuous integration, and increasingly, AI coding assistants that write much of the code.

What makes this set of properties attractive is that the same features that make an analysis reproducible also make it easier for an LLM to work with. An LLM is much better at modifying a program when its inputs are explicit, its dependencies are named, its state is visible, and its errors are structured. In that sense, good language design doesn’t just put guardrails around the human analyst. It gives the machine a better representation of what the analyst is trying to do.

This suggests an interesting convergence: the language we would want to make data science safer for humans is also close to the language we would want to make data science easier for AI to operate on. A language built around these principles could be a pit of success for both the human analyst and the AI assistant helping them. Instead of pulling in different directions, the two would be working with the same constraints.

Defaults, constraints, and choice architecture

There is also a useful way to think about this from economics. We often distinguish between telling people what they ought to do and changing the rules and incentives under which they make decisions. A recommendation such as “pin your dependencies” or “don’t rely on global state” is a rule in the first sense: it asks the analyst to remember and follow a good practice. Defaults, constraints, and institutions work differently. They change what is easy to do, what is costly to do, and what is possible to do at all.

This is closely related to the idea of choice architecture in behavioural economics, and to the broader observation that the design of an environment can shape behaviour without requiring people to consciously follow a set of instructions. In software, the architecture is unusually literal: a language can make some actions easy, others difficult, and some impossible to express without making the programmer explicitly acknowledge what they are doing.

That suggests a useful question for data science. Instead of asking how we can get analysts to consistently follow reproducibility best practices, what if we designed the language so that following those practices was the easiest way to write a program?

Design tools that make falling into a pit of success easy

But have you tried asking someone who does none of this to adopt these practices? People may agree with every point, but applying them consistently takes effort and frankly brings very little immediate reward. Investing early in good practice makes life much easier later, but “later” is hard to sell to someone with a deadline yesterday.

So what is the alternative? Rather than persuading people to jump into the pit, we can build tools that make falling into it the default. I’ve been working on a domain-specific language called T that aims at exactly this. The idea is that by simply using it, you end up in a pit of success. I won’t go into the details here, since you can read more about it here and here. What I want to discuss is why LLMs make this kind of tool practical to build.

Historically, building a language or a tool was a large project that only a few people could justify. LLMs lower that cost dramatically. Someone with experience in a domain can now build a small tool that solves one specific problem well, without a team or a long development cycle. Taste, knowing which problems matter and which habits are worth enforcing, is exactly what a domain expert brings and what an LLM does not supply on its own.

I don’t know yet how widely T will be adopted. But I think it is a good example of the pit-of-success idea applied to data science, and the broader point is that the tools themselves can carry the discipline, so that individual analysts don’t have to. The future will tell whether that bet pays off.

Where the pit needs walls and an exit

A pit of success can become a trap if it is the only way to work. Exploration is messy by nature. You try a model, plot something odd, throw away three approaches before finding the one that works. A language that forbids all of that makes people go back to their old tools, usually a notebook on the side, and then the discipline lives in two places at once.

So the language needs a deliberate place for exploration that does not contaminate production. Exploration should be easy and interactive, with the results promoted into the strict pipeline once they stabilise. The pit should catch you when you go into production, not when you’re trying to think. This is the part I find hardest to get right, and I suspect any language that ignores it will be abandoned quickly. For now, the way this works in T is that you have much more freedom to try things from its REPL. But I suspect that much of this exploratory work will be LLM-driven in the future. Interactive data exploration may move from “playing around” with the data to asking an LLM to do things for you and show you the results.

If that happens, interactive exploration may become less important as a feature of the language itself. The important requirement would instead be that the language provides a good interface for an LLM to explore the data and then turn what it discovers into a reproducible pipeline.

I think much of the answer lies in the tools we already use. Editors such as Emacs and Positron or VS Code already give analysts a place to explore interactively while writing plain text files that can later be run as a pipeline. What a new language needs is an editor integration that makes this transition smooth: completions, diagnostics, and navigation working in the same window as the exploratory code. That integration is a solvable engineering task, and LLMs are now quite capable of producing the plugins, language servers, or tree-sitter grammars it requires. One question is whether the language makes the move from exploration to production feel like tidying up rather than rewriting from scratch, which is, again, the typical life-cycle of an analysis. Quick and dirty exploration in a Notebook, clean rewrite in a script (sometimes even in another language). T solves this by forcing the analysis to be a pipeline that runs in a reproducible environment by design.

What the pit cannot do

A language, even one supported by an LLM, can make reproducibility cheap. It cannot make the analysis correct. A misspecified regression run reproducibly is still misspecified, and catching that remains the work of a statistician and a reviewer. Nor can a language change the culture around it. People need to believe that rebuilding last year’s figures is worth the time, and someone needs to ask whether they still can.

The pit of success makes the right thing easy. It does not decide what the right thing is, and it does not make anyone want to jump in. Those are still human jobs.

This is roughly the bet behind T, though I think the general ideas matter more than any one implementation. If you have thought about what a data science language should refuse to let you do, I would be interested to hear what you would put on that list.

To leave a comment for the author, please follow the link and comment on their blog: Econometrics and Free Software.

R-bloggers.com offers daily e-mail updates about R news and tutorials about learning R and many other topics. Click here if you're looking to post or find an R/data-science job.
Want to share your content on R-bloggers? click here if you have a blog, or here if you don't.

Never miss an update!
Subscribe to R-bloggers to receive
e-mails with the latest R posts.
(You will not see this message again.)

Click here to close (This popup will not appear again)