Breaking the guardrails: stress-testing, jailbreaking, and ultimately securing LLMs for Population Health research
- Entry
- 2027 / 2026/2027 applications
- Supervisors
- Charles Rahal, Daniel Valdenegro, Jiani Yan
- Research unit
- Demographic Science Unit
Oxford Population Health, University of Oxford
How can we measure consequential failures in AI-assisted health research, and build safeguards that preserve legitimate scientific use? This project treats LLM safety as a problem of rigorous computational evaluation, rather than a count of refusals.
Background
Large language models increasingly support evidence synthesis, research coding, and protocol development. However, their safeguards can fail under adversarial prompts, retrieved documents, or tool interactions. In population health research, the relevant failures include fabricated evidence, privacy leakage, unsafe recommendations, and corrupted analyses. A model’s willingness to answer is not, by itself, a meaningful measure of risk.
The project will develop a health-research-specific evaluation framework that distinguishes harmless responses from scientifically or clinically consequential failures, and measures whether defensive interventions reduce harm without blocking useful research.
Computational Approach
- Specify and benchmark the threat model. Construct a preregistered evaluation with benign controls, explicit attacker capabilities, and fixed query budgets. Compare open-weight and authorised hosted models across direct, multi-turn, retrieval-mediated, and agentic workflows.
- Measure consequences, not just compliance. Combine expert annotation and hierarchical modelling to assess severity-weighted attack success, scientific correctness, information leakage, unauthorised actions, and attack cost.
- Evaluate adaptive attacks and layered defences. Quantify robustness alongside over-refusal and retained research utility, making the safety-utility trade-off an explicit outcome of the experimental design.
All evaluation will use synthetic tasks, canary data, and sandboxed tools, not real patient records or live clinical systems. The programme is intended to produce methodological papers, an auditable benchmark, and controlled-access research tools, with responsible disclosure integral to the work.
Training and Research Environment
Training will combine natural language processing, experimental design, reproducible software engineering, and statistical evaluation with secure red-teaming, clinical safety, and research ethics. The project is computational; no fieldwork is envisaged.
Prospective Student
This project would suit someone with strong Python and computational skills, grounded in computer science, statistics, data science, or health informatics. Experience with NLP, machine learning, or software engineering is valuable, alongside careful experimental reasoning and a commitment to ethical, reproducible research.
Enquiries and Applications
For an informal discussion, contact Charles Rahal, quoting DSU001 and outlining your research interests and relevant computational experience.
This is a project for 2027 entry to the DPhil in Population Health. Please consult the official Oxford project advert and course page for current application requirements, deadlines, and funding information.