The file drawer

Ideally, social science should be an institutionalised reflective equilibrium between theory and empirical work—through which arguments guide empirical inquiry and findings prompt revisions to our assumptions, explanations, and questions (Ashworth et al., 2021)—with each scholar contributing according to their comparative advantage. Realising this ideal requires clear theorising and, as the credibility revolution has taught us, attention to the identification challenges associated with testing the empirical implications of theories. It also requires care in thinking about the scope conditions and external validity of findings.1

Contributing to this collective endeavour requires individual researchers to invest time and effort under considerable uncertainty about which ideas will prove fruitful. Indeed, research is often a non-linear process of going down dead ends, trying other avenues, failing, and trying again until one makes progress in answering the research question of interest. This process requires a willingness to be usefully wrong (though being right never hurts, of course)—an aspiration reflected in the title of my Substack, Often wrong, but sometimes useful. Even projects that fall short of their original aims can advance understanding by clarifying the limits of an argument or revealing why a design cannot answer the intended research question.2

These ideals are (all too) often in tension with the incentives of academic publishing and career progression, which can encourage researchers to retrofit questions or theoretical intuitions to familiar methods, explore the ‘garden of forking paths’ somewhat too assiduously, and consign ‘unsuccessful’ projects to the file drawer (Gelman and Loken, 2013; Smaldino and McElreath, 2016; McElreath, 2020). Academic websites tend to reproduce this emphasis on polished, successful projects, leaving much of the research process unreported. Unclear theoretical ideas, plausible but ultimately uninformative empirical designs,3 and inconclusive or contradictory findings consequently receive less attention.

This selective reporting creates a form of survivorship bias,4 which has at least three negative effects. The first is that students see only polished projects that “work out” while regularly “failing” with their own projects. This can be demoralising, with frustration—that is normal in research and often also productive—being mistaken for a lack of ability. It can also discourage them from exploring new avenues, leading to excessive risk aversion in the form of hewing closely to what has already been done.

Second, it can distort the allocation of research effort. Without access to earlier attempts and the reasons they proved unproductive, researchers may repeat avoidable mistakes or pursue approaches that others have already found unsuitable. Third, it inhibits the accumulation of knowledge. When credible findings that provide little support for an argument remain unpublished, researchers may overestimate both the strength of its empirical support and the range of circumstances in which it applies (see footnote 1).

While there is a severe limit to what any single individual can do, this part of my website is meant to list ‘failed’ theoretical ideas and empirical projects that can potentially inform subsequent work. Sharing these attempts is one way I try to resist the temptation to let unsuccessful work disappear in my file drawer, as it were. Doing so at scale, however, requires institutional changes rather than merely individual resolve. Aletheia, which makes both submitted research and reviewers’ feedback publicly available, illustrates what better institutions might look like and how difficult institutional change is.

Critical scrutiny of the incentive structures that social scientists are subject to and suggestions for reform are much needed, particularly in the age of generative artificial intelligence (Munger, 2026). But a single-minded focus on what is wrong risks giving rise to the impression that the entire system is broken. My own experience gives me reasons for a more hopeful view. In many ways, this website is an homage to the generous senior academics who have mentored me, from my supervisors to my co-authors and beyond, and to the ideals of intellectual curiosity, rigour, and generosity that they embody. Finally, the website also serves as a commitment device through which I hope to hold myself accountable to these ideals in my own research, knowing that I will sometimes fall short.

Green transition, brave new feminine world?

2025–26 · Survey experiment

With Henri Gruhl (RWI – Leibniz Institute for Economic Research and Vrije Universiteit Amsterdam), Johannes Brehm (RWI – Leibniz Institute for Economic Research and Hertie School), and Lara Hankeln (University of Oxford)

Motivation

Women consistently express greater support for climate policy than men, yet the mechanisms behind this difference remain contested. We examine how people who are not personally exposed to job loss respond when climate policy displaces workers from traditionally masculine occupations. Even when relatively few workers lose their jobs, the broader public’s judgements about their losses and the compensation they deserve may influence support for the policy.

The theoretical intuition

We treat occupational status as the degree of prestige an occupation commands, independent of income. We focus on its masculinity dimension—the extent to which an occupation’s prestige derives from its association with masculinity. These associations involve traits such as physical risk, toughness, or manual skill. By focusing on a single dimension, we avoid the interpretative ambiguity of composite measures, in which it can be unclear which component produces an observed political response. Occupations are also central sites of preference formation. Two jobs can have similar overall prestige, wages, and employment security while differing in the masculinity component of occupational status.

When occupational transitions preserve the masculinity component of status, disagreement among the unaffected chiefly reflects the balance between other-regarding motivations and efficiency concerns. When transitions reduce this component, respondents also disagree over whether the loss is normatively deserving of compensation. We expect men without tertiary education, who tend to have stronger affinities with masculine occupations, to be more supportive of compensation than women without tertiary education. Among the tertiary-educated, we expect lower support and a narrower gender gap.

The “status-substitution problem” creates a further obstacle. While material harms can be offset with sufficient resources, money is less effective at compensating losses in the masculinity component of occupational status. Compensation may therefore fail to sustain support for climate policy among those who regard these losses as normatively significant, while others may oppose transfers that they construe as sustaining patriarchal structures.

The research design

The pre-analysis plan specifies a survey experiment in Germany’s Social-Ecological Panel, with respondents randomly assigned in equal proportions to two vignettes. Both describe a 45-year-old male coal worker who loses his job through the coal phase-out, receives temporary financial compensation equal to his previous income, and undergoes retraining and placement in a new occupation. In the condition intended to preserve the masculinity component of occupational status, the new occupation involves physically demanding technical and manual tasks and regular use of tools and machinery. In the condition intended to reduce this component, it is a social occupation involving emotionally demanding tasks and direct interaction with people. Both descriptions specify the same income and social standing as before, a secure position in the same region, and a working environment, hours, and team composition similar to the previous job.

We describe broad task domains because specific job titles, such as construction worker or nurse, could introduce stereotypes about migration, wages, prestige, and working conditions. We also avoid explicit references to masculinity or gender norms to reduce concerns about social desirability and respondents inferring the intended answer. Immediately after the vignette, respondents estimate the new job’s income and describe its occupation in their own words. The vignette remains visible during these checks, then disappears before the outcome questions. These measure support for temporary financial compensation, retraining and placement, and climate policy. They also assess the perceived fairness of compensation. Later questions check recall of the new occupation and whether respondents inferred the study’s purpose. The analysis examines differences by respondents’ gender and education.

The result

The preregistered predictions were largely not borne out. Respondents initially noticed the occupational contrast, but their later answers often failed to retain it. Counting missing answers as failures, 74% later recalled an occupation consistent with the technical and manual vignette, compared with 21% for the social vignette. Some respondents who initially recognised the social occupation subsequently reverted to technical or manual work. Respondents given the social vignette also expected lower earnings, despite the statement that income was unchanged. These problems leave the proposed mechanism inadequately tested.

Lessons for subsequent research

The central lesson is that noticing a vignette’s wording does not establish that respondents believe its premises or retain the intended distinction. I suspect that the transition into social and interpersonal work, with unchanged income, social standing, and working conditions, lacked credibility; the evidence cannot distinguish all possible explanations. I would therefore consider other experimental approaches, including images, while testing whether respondents accept the stipulated conditions and retain the occupational contrast. The theory was also too complex for the design, with low statistical power for its predictions. A subsequent study should examine a simpler argument with sufficient power for comparisons by gender and education.

Materials

OSF project · Thread on the pre-analysis plan · Thread on the results

Footnotes

  1. Questions of external validity may seem abstract, yet they have immediate practical relevance in my work on the IPCC chapter, where we have to synthesise the literature and assess how far its findings can be generalised. This requires specifying both the quantities that individual studies estimate and the questions that the synthesis seeks to answer. The MIDA framework of Blair et al. (2023)—model, inquiry, data strategy, and answer strategy—distinguishes the question of interest, including the units to which it refers, from the procedures used to collect and analyse evidence.

    Because a synthesis typically concerns settings other than those studied, external validity requires that the mechanism producing an effect in the studied setting also operates in the settings of interest. Gailmard (2021) examines what could justify such a claim. A mechanism credibly identified in one case can be inferred to operate in a second case only if similarity in the observable characteristics of the two cases is taken to imply that the same mechanism operates in both. That step requires the same assumption as identification via selection on observables in a single case, namely that the relationship between observable characteristics and outcomes is known well enough to infer an outcome that has not been observed. Researchers who reject this assumption for causal inference in a single case cannot consistently accept it for generalisation across cases. Put simply, external validity requires a theory that specifies which characteristics of a case determine whether the mechanism operates in it. Such theory is rarely made explicit in the synthesis of evidence and in what is often called evidence-based policy advice, where the results of well-identified studies are applied to settings that resemble the original study in observable respects.

    There has been growing interest in how to think carefully about external validity, though the literature remains nascent and its insights have not yet percolated through the profession. Three strands of this literature strike me as particularly useful. The first examines how much a single study teaches us about a population beyond its sample. The second examines when estimates from different studies can be combined. The third examines why treatment effects vary across studies even when their designs are harmonised.

    Little and Pepinsky (2021) provide a Bayesian framework for the first question. Any estimate equals the true effect plus the bias of the design. Once researchers specify prior beliefs about both, Bayes’s rule determines how much an estimate teaches them about the effect (see Gelman et al., 2013, for a general treatment of Bayesian inference); the more of the estimate’s variation they attribute to bias, the less they learn. Restricting a study to a subpopulation in which the design is credible, such as compliers or units near a discontinuity threshold, removes the bias but introduces a different unknown, namely the difference between the effect in the subpopulation and the effect in the population of interest. Researchers therefore face a choice between two sources of error. The estimate from the whole population is biased, and the estimate from the subpopulation answers a different question.

    Which estimate teaches researchers more about the population effect depends on which source of error they believe to be smaller. The subpopulation estimate is more informative if the effect in the subpopulation is likely to track the effect in the population closely. The whole-population estimate is more informative if its bias is likely to be small. The more uncertain researchers are about that bias, the less closely the subpopulation effect needs to track the population effect for the subpopulation estimate to be preferable. In either case, the answer depends on prior beliefs about quantities that the study itself does not estimate.

    Slough and Tyson (2023, 2024) address the second question. Studies estimate the same target if and only if they draw the same contrast, measure outcomes in the same way, and examine a mechanism that has external validity across their settings. Without such harmonisation, a meta-analysis may fail to find consistent evidence for a mechanism that operates in every setting, or find apparent consistency where the mechanism is absent. Meta-analysis presupposes these assumptions and cannot establish them.

    Two further contributions address the third question. Izzo, Dewan, and Wolton (2025) show that even harmonised studies in the same context may estimate different quantities because the average treatment effect depends on the information available to individuals, which varies with realised circumstances even when the structural characteristics of the environment remain constant. Only effects conditional on a sufficient statistic for that information are invariant across circumstances. Tosh et al. (2025) show that many variables can have large effects on the same outcome only if they are strongly dependent on one another or interact strongly, and interactions make the effect of any one factor depend on the levels of the other factors in a particular study. Both results imply that a synthesis has to specify the conditions under which an effect is expected to hold (see also Findley, Denly, and Kikuta, 2026).

  2. Spirling and Stewart (2025) provide an epistemological rationale for this claim. They argue that almost all empirical social science proceeds by inference to the best explanation, with researchers generating alternative explanations, collecting facts relevant to their implications, and discriminating among the explanations in light of those facts. Evidence that fails to support an explanation contributes to this endeavour because it eliminates candidate explanations and prevents researchers from mistaking an absence of evidence for support. Spirling and Stewart therefore recommend reporting null and negative results even for preferred explanations. On this view, a study that fails to reject the null hypothesis has not identified a best explanation, yet it can still show that a particular explanation is inconsistent with the available facts.

  3. Examples include a regression discontinuity design with insufficient statistical power near the threshold to detect effects of substantive interest (Cattaneo et al., 2019) or a survey experiment in which respondents do not perceive the intended distinction between treatments.

  4. The related problem of file drawer bias arises when the likelihood of reporting and publishing findings depends on their statistical significance. Franco, Malhotra, and Simonovits (2014) find that statistically insignificant results are less likely to be published. Using a more recent sample of papers, Moniz, Druckman, and Freese (2025) confirm these findings, though the difference is substantially smaller than in previous decades. Both studies point to researchers’ decisions about writing up and submitting results as an important source of selection.

    Huntington-Klein et al. (2025) find substantial variation among 146 research teams answering the same question, with decisions about data wrangling accounting for a substantial amount of this variation. Stevenson and Fischman (2026) show how favouring statistically significant results can inflate reported effects, particularly when statistical power is low. They also argue that conventional standard errors omit uncertainty arising from reasonable alternative analytical choices. These findings reinforce the value of documenting unsuccessful projects and the decisions behind them, helping others assess the evidence and inform subsequent research.