All DPhil projects
Entry
2027 / 2026/2027 applications
Supervisors
Charles Rahal, Jiani Yan
Research unit
Demographic Science Unit
Oxford Population Health, University of Oxford

A good risk predictor is not necessarily a good measure of disease burden. This project will use population-scale longitudinal records to develop interpretable measures of multimorbidity that remain meaningful across health systems and social groups.

Background

Counting diagnoses treats very different conditions as equivalent; legacy indices often prioritise inpatient mortality rather than patients’ daily lives. Rich diagnostic and prescribing histories offer an alternative, but increasingly accurate predictions do not automatically yield valid measurements of multimorbidity.

This project will develop transparent, multidimensional measures that distinguish shortened life, day-to-day health impact, and the burden of managing multiple medicines. Development in UK Biobank and external evaluation in China Kadoorie Biobank will make transportability a central methodological question rather than an afterthought.

Computational Approach

  1. Construct longitudinal representations. Harmonise diagnostic and prescribing histories using established phenotypes, then apply representation learning to identify structure and interactions among conditions.
  2. Distil interpretable measures. Translate that structure into ordered measures of disease burden. Deep predictive models will act as benchmarks for quantifying the cost of interpretability, rather than becoming the final measurement instrument.
  3. Validate across outcomes and populations. Assess survival, lived health burden, and polypharmacy alongside calibration, discrimination, reliability, and temporal stability. Use psychometric measurement-invariance methods to test whether scores retain their meaning across countries and social strata.
  4. Stress-test measurement assumptions. Examine sensitivity to disease-recording differences, clinical code mappings, condition definitions, and missing-data assumptions.

The methodological aim is not merely to maximise predictive performance, but to establish what a score measures, how reliably it does so, and where its interpretation travels.

Training and Research Environment

The project brings together electronic health record phenotyping, longitudinal and survival analysis, representation learning, interpretable machine learning, and psychometrics. Analyses will use secure biobank environments and high-performance computing, with training in data governance and reproducible research software. The student will participate in the Leverhulme Centre for Demographic Science and the Metrics and Models lab. No fieldwork or industry placement is envisaged.

Prospective Student

Applicants should have computational training in mathematics, statistics, data science, computer science, formal mathematical epidemiology, or a related discipline, and be confident programming in Python. Experience with machine learning or electronic health records is helpful, but not essential; an interest in rigorous validation and meaningful health measurement is central.

Enquiries and Applications

For an informal discussion, contact Charles Rahal, quoting DSU002 and outlining your research interests and relevant computational experience.

This is a project for 2027 entry to the DPhil in Population Health. Please consult the official Oxford project advert and course page for current application requirements, deadlines, and funding information.