OpenAI is announcing security updates following the July news that its AI broke out of a sandboxed environment and accidentally hacked Hugging Face, including improvements to its research environments, monitoring, and alignment techniques. The company had already put the brakes on a new model, Astra, that it thinks could have âcriticalâ cybersecurity capabilities, and the company says it instituted a two-week pause in reinforcement learning (RL) training on its âlatest models intended for deploymentâ while it tightened up security. The companyâs âlargest planned frontier RL run remains on hold.â
For its frontier model research, OpenAI now requires stronger sandboxes for workloads that âexecute model-generated or otherwise untrusted code,â and has more controls to âisolate higher-risk and untrusted workloads from the internet.â It has also updated its research environment to âremove potentially vulnerable shared services, reduce standing privileges, and improve security and trust boundaries.â
As part of the companyâs expanded monitoring setup, OpenAI now aims to issue an alert âwithin 30 minutes after concerning activity is surfaced,â OpenAI says. If the people paged after an alert canât âconclusivelyâ determine whether an alert is a false positive within 30 minutes, âthose teams are expected to pause the activity.â
OpenAI also says that itâs applying âour core alignment techniques across more stages of the training process,â including reward models that âbetter detect and discourage unsafe behaviorâ and training models âto be more honest about their actions, capabilities, and limitations.â
Since the discovery of the Hugging Face breach, Anthropic and Meta have also found that their AI models had hacked other organizations.
Read the full article here