All DPhil projects
Entry
2027 / 2026/2027 applications
Supervisors
Jakub Bijak, Charles Rahal
Research unit
Demographic Science Unit
Oxford Population Health, University of Oxford

How can incomplete, overlapping, and selectively reported records support defensible mortality estimates? This project will connect machine-assisted evidence extraction with demographic estimation, carrying uncertainty through the full computational pipeline.

Background

Conflict mortality is difficult to estimate precisely because the available evidence is not a representative census of deaths. Sources omit events, duplicate reports, and capture different populations with different probabilities. Combining more records is therefore insufficient unless the model also addresses how those records were generated.

This project will integrate demographic methods and machine learning to identify, reconcile, and model casualty records and contextual evidence. Primary sources will include ACLED, GDELT, UCDP, and selected digital archives, with empirical applications across contrasting conflict settings.

Computational Approach

  1. Extract and harmonise evidence. Investigate NLP and large language models for information extraction, with possible approaches including retrieval augmentation, weak supervision, prompt optimisation, and domain-adaptive fine-tuning.
  2. Resolve overlapping reports probabilistically. Combine record linkage, learned representations, and graph-based data integration to investigate inconsistent and duplicated evidence without treating uncertain matches as established facts.
  3. Estimate mortality under incomplete observation. Develop multiple-systems estimation in a Bayesian hierarchical framework, accounting for dependence between sources and variation in ascertainment across places and periods.
  4. Propagate and validate uncertainty. Carry extraction and linkage uncertainty into the final mortality estimates, using scalable approximate inference where necessary. Evaluate performance through simulation, sensitivity analysis, and out-of-sample validation.

The central contribution is an integrated inferential framework: uncertainty introduced during data construction must remain visible in the conclusions, rather than disappearing between pipeline stages.

Training and Research Environment

The student will join the Demographic Science Unit within Oxford Population Health. Training will cover reproducible data engineering, NLP, probabilistic linkage, Bayesian hierarchical modelling, scalable inference, and model validation, supported by relevant advanced courses in Statistics and Computer Science. The work is computational and uses public conflict-event databases and selected archival sources; no fieldwork is planned.

Prospective Student

This project suits applicants with a strong quantitative background in applied statistics, econometrics, machine learning, or data science, and proficient programming in R, Python, or a comparable language. An interest in demographic modelling and conflict mortality is important. Experience with Bayesian analysis, probabilistic programming, NLP, or record linkage is desirable; familiarity with Stan or JAGS would be advantageous.

Enquiries and Applications

For an informal discussion, contact Charles Rahal, quoting DSU003 and outlining your research interests and relevant computational experience.

This is a project for 2027 entry to the DPhil in Population Health. Please consult the official Oxford project advert and course page for current application requirements, deadlines, and funding information.