In tests, AI robot systems easily rejected directly malicious commands. But their safety filters collapsed when creative writing was used to instruct them.
- Senior Researcher in AI safety, interpretability and technical governance, University of Oxford
Contact Fazl for
- General
- Media request
- Speaking request
- Consulting / Advising
- Research collaboration
- Research supervision
- Website
- Article Feed
- ORCID
- Joined
