Escaping the extrapolator's circle
A review and discussion of Daniel Steel's "Across the Boundaries".
Research on large language models (LLMs) faces a number of epistemological challenges. I am increasingly convinced that many of these challenges can be fruitfully described as questions about the boundaries or limits of extrapolation. To better understand these problems, I’ve been reading some of the relevant literature, including philosophy of science. Philosophy, in my experience, is not necessarily helpful for “solving” methodological or theoretical questions in science, but it can be extraordinarily helpful in charting out the space of concepts and arguments, ultimately in the service of greater clarity about the problem you’re actually facing. One of the texts I’ve found particularly useful in this regard is philosopher Daniel Steel’s Across the Boundaries. The following essay is a review of that book, and was originally published on the Leaky Margin.
Researchers often wish to draw generalizations about some “target” population, but for various reasons—ethical concerns, feasibility, etc.—the population of interest cannot be studied directly or exhaustively. Instead, researchers study some other population, with the goal of extrapolating causal relationships from the sample to the target.
Daniel Steel’s Across the Boundaries begins with several examples of this nature, such as investigating whether a substance is carcinogenic in humans by studying whether it is carcinogenic in rats (a “model organism”). As Steel writes of this example and others:
If the populations in question were perfectly homogenous, extrapolation would be easy: the result from the first population could be directly carried over to the second. But it is not reasonable to assume that the populations in the foregoing examples are homogeneous: they almost certainly differ with respect to characteristics that affect the causal relationship in question.
The challenge, then, is whether and when we can “transfer” causal generalizations from one context to another in the absence of homogeneity.
Steel calls this challenge the problem of extrapolation. The book, in large part, consists of detailed elaboration on this problem, proposed solutions to the problem, and epistemological or practical problems with those solutions (such as the extrapolator’s circle, as discussed below). Along the way, Steel introduces his preferred notion (and notation) of causality, drawing heavily from Judea Pearl’s do-calculus. Various concepts from philosophy of science such as “mechanism” and “capacity” are also examined in the context of the problem of extrapolation.
His focus is on two broad academic disciplines: biology and social science.
In biology, extrapolation arises most clearly when researchers wish to draw generalizations from the study of some model organism (e.g., a rat) to some target organism or species (e.g., humans). His running example is the carcinogenic properties of aflatoxins: having established that these molecules cause cancer in rats, can scientists draw the inference that they cause cancer in humans—and what epistemic principles can we learn about extrapolation in the efforts of scientists to draw such inferences?
The challenge, though, is not restricted to questions about medicine or health: much neurophysiological research, for example, is conducted on non-human organisms (such as monkeys, rats, or even lobster stomatogastric systems). And the motivation for addressing this challenge, moreover, is not merely epistemological. If studies of non-human animals do not reliably tell us anything about humans—and if knowledge about humans is taken as the ostensible goal of those studies—one might wonder whether it is ethically appropriate to conduct such studies in the first place.
(There might, of course, be many other reasons to study non-human animals! Though notably, the problem of extrapolation in biology is not restricted to cases where researchers want to extrapolate to humans. Given that one has studied a sample of rats, what has one learned about rats in general, or rodents in general? What can we reliably conclude?)
Steel’s focus in social science is primarily on work in economics, e.g., the study of whether a “welfare-to-work program” will improve the lives of its recipients. Even if a pilot study yields interesting and promising results, on what grounds can researchers conclude that those results will reliably generalize to a different population? The original population and the target population are, presumably, different along numerous dimensions, some of which may plausibly interact with the assumed causal mechanisms at play. This is, in a sense, the problem underlying the design of many public policies or assistance programs: something that worked in one context may not work in another.
As with biology, however, the challenges of extrapolation extend to many areas of social science, perhaps most notably research in cognitive science, in which experimental studies are often conducted on so-called “WEIRD” populations. Cognitive scientists are increasingly aware of these problems, in part due to extensive writing by researchers like Joe Henrich or Asifa Majid, both of whom have made compelling cases for the importance of cultural and linguistic diversity in studying human cognition.
When it comes to public policy, these challenges are further complicated by the observation that intervening in a system may, in many cases, change the causal structure of that system in ways that are relevant to the intervention. We might usefully gloss these cases as one kind of “unintended consequence”. The simpler kind of unintended consequence occurs when one doesn’t fully understand the system in question, turns some knob (literally or metaphorically) anticipating some target effect, and instead (or in addition) observes a different effect, which is possibly undesired. Structure-altering interventions, however, can yield unintended consequences even if one has the correct understanding of a system at time T; because the intervention changes the structure of the system, one’s understanding at time T may not be helpful for predicting or understanding its behavior at time T + 1. Steel writes:
Even if the generalization accurately described the causal relationship under the original set of background conditions, it might be an inaccurate representation of that relationship in the new circumstances brought about by the intervention. (pg. 155)
As an example of such a structure-altering intervention, Steel cites anthropologist James Scott’s account of a government-sponsored irrigation project in a Malaysian village, which made it possible to grow two rice crops, rather than one, per year (bolding mine):
Initially, the project significantly improved the economic situation of villagers, from the land-poor who relied on wage labor to the larger landowners…However, landowners soon discovered that renting combine-harvester machines was better than hiring field labor under the new system, since double-cropping required quickly harvesting one crop so that the next could be planted. Not only did this turn of events adversely affect the wages of land-poor villagers, it also undermined traditional demonstrations of generosity through which wealthy farmers attempted to ensure reliable sources of labor in the future. These practices included sumptuous feasts to which all in the village were invited, bonuses for laborers at the end of the harvest, and tenancy agreements that made allowances for poor harvests. In short, the innovation of double-cropping fundamentally altered the economic structure of mutual dependence between poor and wealthier villagers and all of the practices that went along with it. (pg. 156)
More abstractly: the intervention changed the constraints and incentives at play in the causal system, producing unintended and unanticipated effects.
Yet clearly, researchers in biology and social science would like to learn generalizable things about their target populations of interest, even when they don’t have direct access to those populations. What are they to do?
The simplest and most naive strategy is what Steel calls simple induction. Here, a researcher simply assumes that causal relationships identified in a model organism can be extrapolated “by default” to their target of inquiry, as long as there is no explicit reason to assume otherwise.
This approach may have its place, but it’s not difficult to identify problems with it. There are, after all, many cases in which models and their targets are sufficiently different such that a finding in the model does not generalize to the target. It seems strange—and perhaps foolhardy—to stipulate that we can extrapolate as long as we aren’t aware of such differences. In a new domain, we will know very little about either the model or the target, and thus we will know very little about the ways in which model and target are relevantly different with respect to the causal relationship of interest. As a consequence, extrapolation will always be more justifiable in cases where we have little knowledge about the model or target.
Perhaps there is a pragmatic justification for this assumption: we have to start somewhere, so let’s proceed as if the model and target are sufficiently similar. But in the absence of knowledge either way, why should the assumption be that they’re sufficiently similar rather than sufficiently different? Moreover, as Steel writes, simple induction doesn’t really offer any guidance about what to do when there’s reason to suspect that there are differences between model and target: can anything be learned about the target in such situations?
Simple induction, then, is somewhat unsatisfying as an epistemic principle. The question, however, is whether a more satisfying and rigorous approach can be identified. A number of such proposals exist, but as Steel points out, many of them face some version of what he calls the extrapolator’s circle (bolding mine):
One challenge, which I call extrapolator’s circle, arises from the fact that extrapolation is worthwhile only when there are important limitations on what one can learn about the target by studying it directly. The challenge, then, is to explain how the suitability of the model as a basis for extrapolation can be established given only limited, partial information about the target. Critics of animal extrapolation sometimes present this challenge in the form of a vicious circle: establishing the suitability of the model would require already possessing detailed knowledge of the causal relationship in the target, in which case extrapolation would be unnecessary. (pg. 4)
A similar definition is given later in the book, again emphasizing the circularity at play (bolding mine):
Simple induction relies on some criterion of relatedness, such as phylogeny or type of economic system. The shortcomings of simple induction stem from the fact that satisfying such criteria is often not sufficient for being a reliable basis for extrapolation. Consequently, additional information about the similarity between the model and the target—for instance, that the relevant mechanisms are the same in both—is needed to justify the extrapolation. The extrapolator’s circle is the challenge of explaining how we could acquire this additional information, given the limitations on what we can know about the target. In other words, it needs to be explained how we could know that the model and the target are similar in causally relevant respects without already knowing the causal relationship in the target. (pg. 78)
Put another way: many proposals for solving the problem of extrapolation rely on the assumption that we know enough about the target to determine whether some particular property of the model can be extrapolated.
For instance, suppose we stipulate that extrapolation is only licensed when we are confident the “same mechanisms” are at play in both the model and the target. In biology, this might be a useful heuristic for evolved, highly conserved mechanisms. But it is not clear how this helps for questions involving mechanisms that are (presumably) not explicitly conserved through evolution, such as responses to a new drug or synthetic compound.
Perhaps, then, we can stipulate that the question of whether any given finding can be extrapolated is an empirical problem. In this approach, we might treat findings obtained in a model as hypotheses that can be subsequently tested (and either confirmed or disconfirmed in the target). This is all well and good for cases where such hypotheses are indeed directly testable, but there are many situations where we can’t empirically test something in the target. Indeed, our inability to directly study the target (e.g., for ethical or practical reasons) is typically a major motivation for using a model organism in the first place!
The circle is expressible in the form of two questions. If we haven’t done the right kinds of empirical confirmation in the target, how can extrapolation be licensed from the model? And if we have done the right kinds of empirical confirmation in the target, what good is the model?
Steel’s preferred solution is something he calls comparative process tracing (CPT).
To understand CPT, we must first understand the notion of “process tracing”. The basic logic here is present in the name: we can study a process in terms of its component parts and mechanisms, literally “tracing” the causal chain from some initial point to whatever outcome we’re interested in. For example, suppose we are interested in whether exposure to aflatoxins causes cancer. We might characterize that causal process in terms of the component mechanisms involved in the metabolization of aflatoxins, perhaps focusing on how the compounds are altered at the molecular level at each stage of this process, as well as which enzymes catalyze these processes.
The diagram below represents a simplified (and abstracted) view of a potential causal chain, where each component stage (e.g., “A”) exerts a causal effect on the subsequent stage (e.g., “B”).
Comparative process tracing amounts to comparing, as best one can, the relevant mechanisms thought to be involved in some process in both the model and the target.
Now, we’ve already established above that we can’t simply enumerate the entire causal chain in the target: if we could, we wouldn’t need to study the model in the first place. The key insight, however, is that we may not need the entire causal chain in the target. Instead, we can figure out which stages we can compare, focusing on stages that might provide particularly rich sources of information about whether the overall chains are functionally equivalent in the right ways. Steel writes:
The reliability of comparative process tracing depends on correctly identifying the points at which significant differences between the model and the target are likely to arise. Significant differences are those that would make a difference to whether the causal generalization to be extrapolated is true in the target. For instance, metabolism is a source of potentially significant difference in carcinogenesis, since how a compound is metabolized often matters to whether it is carcinogenic or not. (pg. 89)
As Steel points out, researchers can focus on particularly informative “downstream” stages, i.e., those in which observing a similarity or difference between model and target provides high amounts of information about the extent to which other aspects of the (upstream and unobserved) causal chain are analogous. For example, if differences at points “B” and “C” must result in differences at point “D”, but only “D” is observable, then researchers can simply compare “D”. If “D” is analogous across model and target, researchers can infer that “B” and “C” are sufficiently similar in the right ways; if “D” is different across model and target, then there must be some difference upstream (perhaps “B” or “C”, or perhaps some other unobserved and unknown stage).
For me, this idea was much clearer with an example.
Returning to the case of aflatoxins: for decades, scientists have known both that exposure to aflatoxins cause cancer in rats, and that exposure to aflatoxins is a risk factor for cancer in humans. The latter is a correlational relationship of the kind typically uncovered by population-level studies in epidemiology. A correlation is consistent with a causal relationship but could, of course, be produced by some other set of variables. The gold standard for confirming causality in humans would be a randomized controlled trial in which people were exposed to aflatoxins, but this would (obviously) be unethical to run. Moreover, it is unclear whether rats are the most appropriate model for humans with respect to this particular set of mechanisms. This is not an idle question, as aflatoxins were not found to have carcinogenic effects in mice. Are humans more like mice or rats in this case?
The comparative process tracing approach to this conundrum is to compare, where possible, the relevant stages and mechanisms in each organism. As Steel describes, researchers compared the metabolism of aflatoxins in humans, rats, and mice:
It was found that although the phase I metabolism of [aflatoxins] proceeded similarly among mice, rats, and humans (and in fact at a higher rate in mice), the phase II metabolism among mice was extremely effective in detoxifying [aflatoxins] but not among rats or humans (Hengstler et al., 1999, 928-31). Furthermore, this metabolite bound to DNA in rat liver cells in vivo at sites at which the nucleotide base guanine was present to form complexes called DNA adducts (ibid., 927). It was further found that such cells suffered unusually frequent mutations in which guanine-cytosine base pairs were replaced with adenine-thymine pairs, a mutagenic effect found in vivo among rats and in vitro among cells of a variety of origins, including bacteria and human (ibid., 923, 927). In addition, guanine-cytosine to adenine-thymine mutations were found in activated oncogenes present in rats exposed to [aflatoxins] but were absent in the controls (ibid., 130-33). Thus, comparative process tracing yielded the conclusion that the rat was a better model than the mouse. (pg. 91)
Abstracting away from some of the molecular details, the takeaway here is that researchers didn’t have to analyze each process end-to-end; instead, they compared the component stages of carcinogenesis in each organism where possible. In rats and mice, cells could be analyzed in vivo (i.e., in living organisms); in humans, cells were analyzed in vitro (i.e., outside of a living organism). Crucially, however, in vitro human cells looked more like in vivo rat cells than in vivo mice cells with respect to how they responded to aflatoxin exposure.
Further evidence comes from the comparison of other biomarkers relevant to assessing the causal effect of aflatoxins (aflatoxin DNA adducts). Here, rat cells again looked the most similar to in vitro human cells in terms of the quantity of these DNA adducts. Interestingly, as Steel notes, even in rats, the quantity of aflatoxin DNA adducts was substantially less than in humans. This suggests that the effect in rats might be seen as a lower-bound for the strength of the effect in humans. If the effect in rats is (to choose an arbitrary number) something like 0.5, we might say that the effect in humans is probably at least 0.5, but might be larger. This conclusion would in turn be based on the fact that human cells in vitro were “more affected” (in terms of the quantity of aflatoxin DNA adducts) by aflatoxin exposure than rat cells in vivo. (A skeptic might appropriately note, here, that one might also wish to compare to rat cells in vitro, in case the effect here is driven by in vitro vs. in vivo comparisons, rather than human vs. rat cells.)
A coarse way to think about all this is that extrapolation is a matter of building up a body of convergent evidence.
But people mean many different things when they use the phrase “convergent evidence”, and I think Steel’s formulation here gets at a more precise notion. In this framework, model organisms are used both to discover causal relationships (“X —> Y”) and to identify the component mechanisms and stages involved in those relationships (“X —> A —> B —> Y”). Researchers can’t, of course, study the causal relationship directly (“X —> Y”) end-to-end in their target of interest—that’s the point of the extrapolator’s circle—but they might decide whether the relationship can be extrapolated from model to target by comparing the relevant mechanisms and stages where possible.
Steel covers a number of other topics in the book, including a much more detailed account of causal dependencies, the role of reductionism in developing explanatory accounts, and perhaps most importantly, the question of whether “social mechanisms” can play an analogous role in licensing extrapolations in the social sciences. He also provides a compelling example of how process tracing can be used to identify causal relationships even when causal inference might be virtually impossible on the basis of statistical analysis.
In my view, the book is most effective in characterizing the problem of extrapolation and in articulating the challenges associated with the extrapolator’s circle. I found Steel’s deflationary responses to various proposals quite compelling, and I think many people (including practicing scientists) would benefit from reading it. For example, in Cognitive Science, there is much discussion about the use of large language models (LLMs) as models or “model organisms” for answering questions about human cognition, but much less discussion (let alone consensus) about the actual inferences that can be reliably drawn from such use cases. Trying to draw out the details of particular cases—with the extrapolator’s circle in mind—might reveal problems that don’t show up when speaking in abstractions.
The book is also effective at conveying the optimistic case for how something like comparative process tracing (CPT) might work. It’s possible I’m misreading Steel here, but I don’t see his goal as primarily prescriptive; rather, the contribution (as I see it) is more akin to articulating the processes involved in actual scientific practice, much as Hasok Chang characterized the historical processes involved in the development of thermometry. Clearly, scientists do regularly make extrapolations about populations other than the populations they’re studying—and while some of these extrapolations might be in error, presumably some of them are successful. The question is thus akin to asking what scientists are doing when they extrapolate successfully. Based on conversations I’ve had with researchers in biology (specifically, neurophysiology), I think Steel’s account is representative of the kinds of epistemic heuristics and reasoning processes researchers seem to draw on.
One question, of course, is whether something like CPT reliably works for causal processes throughout biology, beyond the aflatoxin example that Steel highlights. An even more challenging question is whether it works for the social sciences, in which the concept of a “mechanism” is harder to pin down, and in which the system under investigation might be less stable.
Yet I don’t think Steel can really be accused of being overly optimistic about the prospects of something like CPT. Indeed, the lesson I drew from the book—and, I think, the lesson any honest reader should draw—is that principled extrapolation is extremely difficult, and that many (perhaps most) examples of extrapolation in biology and the social sciences are probably premature. Whether or not this is an acceptable error to make presumably depends on the relative costs and benefits at play, our values, and our sense of how science actually works in practice.



Does the extrapolation problem get nastiest when the target population starts adapting to the model? With LLMs, deployment seems to change the prompts, workflows, and expectations you’re trying to generalize across.