If you ask an AI agent to hack an account, it will most certainly refuse, but researchers at EPFL just proved there is an easier way in, and it involves patience rather than technical skill. Their new study shows that breaking a harmful goal into small, harmless-sounding requests can trick AI agents into completing tasks they would normally reject outright (via TechXplore).
It echoes the recent ‘Bioshocking’ exploit in which AI browsers were manipulated into treating credential theft as part of a harmless game.
How researchers exposed this weakness
The team built an automated testing tool called STING, short for Sequential Testing of Illicit…
Read the full article at DIGITALTRENDS.COM









