1. Estimating the Cost of Informal Care with a Novel Two-Stage Approach to Individual Synthetic Control

With Maria Petrillo, Daniel Valdenegro Ibarra, Yanan Zhang, Gwilym Pryce, Matthew R. Bennett. Preprint available here, and code available here.

Abstract: Informal carers provide the majority of care for people living with challenges related to older age, long-term illness, or disability. However, the care they provide often results in a significant income penalty for carers, a factor largely overlooked in the economics literature and policy discourse. Leveraging data from the UK Household Longitudinal Study, this paper provides the first robust causal estimates of the caring income penalty using a novel individual synthetic control based method that accounts for unit-level heterogeneity in post-treatment trajectories over time. Our baseline estimates identify an average relative income gap of up to 45%, with an average decrease of £162 in monthly income, peaking at £192 per month after 4 years, based on the difference between informal carers providing the highest-intensity of care and their synthetic counterparts. We find that the income penalty is more pronounced for women than for men, and varies by ethnicity and age.

2. From a Seed of Doubt Grows a Forest of Uncertainty

With Jiani Yan and Mark Verhagen. Code is available here. Some slides from recent talks available here. Preprint coming soon!

Abstract: The current best practice in ‘Open’, ‘Reproducible’ and ‘Responsible’ research is to hide the variation caused by pseudo-random number generators (PRNGs) through the arbitrary use of a ‘seed’ or ‘random state’ in algorithmic pipelines. However, in the process of doing so, we argue that researchers index into the scientific record an insurmountable number of outcomes with seeming certainty, when they are in fact anything from certain: eliminating this variation is the opposite of what responsible researchers and practitioners should be doing. PNRGs are almost ubiquitous in some research areas – occurring in a large proportion of quantitative and computational research designs – and the potential variation in the estimand or outcome of interest is hitherto significantly under-appreciated. We undertake a series of substantial replication projects of highly published work, primarily in the form of Monte Carlo simulations, Machine Learning, and more traditional inferential designs. We show just how large the variation caused by the instantiation of PRNGs can be. We conclude with recommendations on how to embrace this variatiaon for the betterment of scientific society, and how to responsibly conduct research designs which legitimately reduce it where possible.

3. The Legacy of Longevity: Persistent inequalities in UK life expectancy

With Aaron Reeves, Felix Tropf and Darryl Lundy. Working paper coming soon!

Abstract: That global life expectancy has more than doubled within the previous two centuries is–by any objective standard–something miraculous to behold, and the academic literature across the fields of economics, demography, public health and evolutionary biology have all contributed to our understanding of the mechanisms behind the regional variations in the demographic transitions in mortality. We focus on the effect of the income differential on health gradients through the life expectancies of the tertiary universe of descendants of the British aristocracy and the general population. We use a dataset of 127,523 offspring up to three generations deep, meticulously curated from 7,161 individual sources including 6,756 instances of direct correspondence with aristocratic families. Using this unstructured free-text data on date of birth and death and information on the general population, we develop lifetable based methodologies to provide five distinct findings. We first fail to replicate and generally rally against the so called `peerage paradox’: that lifespans between aristocrats (and their families) was equivalent to the general population until the turn of the 19th century. Secondly, the mortality transition of elites occurred around 100 years earlier than for the general public (with considerable relative improvements of approximately 30\% during the industrial revolution(s)). Thirdly, male aristocratic offspring fared less well than the general population during both the Great War and the Second World War, consistent with the existing evidence base. Fourthly, life expectancies equalized at the same time as the introduction of the National Health Service Act 1946. Finally, tentative evidence suggests that this gap has, however, begun to re-emerge since the 1980s.

A slide deck related to this workstream can be found here.

4. Gendered Impact in Academic Research

With Sander Wagner and Melinda C. Mills. Code is available here. Preprint coming soon!

Abstract: Evaluating the impact of scientific research beyond academia—on policy, health, the economy, and cultural life—has become a cornerstone of science policy and research-funding allocation worldwide. Yet which researchers produce the research underpinning this impact, and how this production is shaped by gender, remains poorly understood. We combine structured and unstructured records from the United Kingdom’s latest Research Excellence Framework, the largest national research assessment currently in operation, with large-scale bibliometric data to quantify gender differences among the researchers underpinning documented impact. Women account for 38.16% of these contributors: underrepresented overall, but with a consistently higher share than in research-output authorships (33.63%), both overall and across all four REF panels. Impact production is also strongly gendered across domains: women are better represented in case studies concerning education, health, cultural, and civil-society impact, whereas those concerning commercialisation pathways such as patenting and manufacturing remain dominated by men. These findings reveal critical inequalities across the pathways that connect research to impact beyond academia, offering crucial evidence for policymakers and academic institutions aiming to build more equitable and representative systems for evaluating scientific contributions.

5. A Grid Based Approach to Analysing Spatial Weighting Matrix Specification

Code library available here, with a link to the working paper version here (I hope to finish this paper one day, but it’s going to involve rewriting a lot of MatLab code into Python).

Abstract: We outline a grid-based approach to provide further evidence against the misconception that the results of spatial econometric models are sensitive to the exact specification of the exogenously set weighting matrix (otherwise known as the ‘biggest myth in spatial econometrics’). Our application estimates three large sets of specifications using an original dataset which contains information on the Prime Central London housing market. We show that while posterior model probabilities may indicate a strong preference for an extremely small number of models, and while the spatial autocorrelation parameter varies substantially, median direct effects remain stable across the entire permissible spatial weighting matrix space. We argue that spatial econometric models should be estimated across this entire space, as opposed to the current convention of merely estimating a cursory number of points for robustness.