- Senior Researcher in AI safety, interpretability and technical governance, University of Oxford
Dr Fazl Barez is a Senior Research Fellow at the University of Oxford, specialising in AI safety, interpretability, and governance. He leads research initiatives within the AI Governance Initiative, focusing on the development of safety frameworks and interpretability methods for advanced AI systems. He also teaches the AI Safety and Alignment course.
Alongside his academic work, Dr Barez is Principal Scientist at Martian, which works on understanding machine intelligence. His research is supported by OpenAI, Anthropic, Schmidt Sciences, Nvidia and others. He is also affiliated with the Centre for the Study of Existential Risk (CSER) at the University of Cambridge, contributing to research on the risks associated with artificial intelligence.
Dr Barez's work focuses on understanding what happens inside neural networks and using that understanding to make AI systems safer. As models grow more capable, the most urgent challenge in interpretability is moving from observation to action – building systems where we can trace a model's internal reasoning, verify it, and correct it when something is wrong. Specific directions he is working on include:
• When a model produces a surprising or harmful output, how can we trace the internal cause – automatically and at scale?
• How can we remove a dangerous capability from a model and be confident it won't come back?
• When a model shows its reasoning, how do we know it actually computed things that way – and how do we give regulators evidence?
Experience
-
–presentSenior Researcher in AI safety, interpretability and technical governance, Oxford Martin School, University of Oxford
Contact Fazl for
- General
- Media request
- Speaking request
- Consulting / Advising
- Research collaboration
- Research supervision
- Website
- Article Feed
- ORCID
- Joined