Articles on AI alignment
Displaying all articles
As AI agents become more autonomous, keeping them aligned with what humans want will require layered oversight and effective control.
AI systems need ‘situational awareness’ to do the right thing – but recent events show they can get confused.
A new AI model could automate the process of searching for cybersecurity bugs and flaws – for better or worse.
People and computers perceive the world differently, which can lead AI to make mistakes no human would. Researchers are working on how to bring human and AI vision into alignment.
Low-cost Chinese-made AI could outcompete premium US models – and it won’t be an accident.
Humans and AIs have different methods of calculating words about probability like ‘maybe’ and ‘likely’ – and different interpretations about what they mean.
While we’re at the peak of AI hype, some people are concerned it will wipe out humanity.
In stress-testing AI models, it’s not hard to push them to the brink and make them threaten to harm humans.
The tools that are meant to help make AI safer could actually make it much more dangerous.








