
Link: https://www.ictbusiness.biz / business / openai-makes-safety-changes-after-ai-hack
OpenAI Makes Safety Changes After AI Hack
OpenAI decided to slow the development of its most advanced unreleased AI models. The company also introduced tougher safeguards after an incident that indicated upcoming systems could pose greater cyber risks.
In an update, OpenAI stated that it has temporarily paused reinforcement learning training on its latest models intended for deployment for two weeks while it hardened research environments and broadened monitoring coverage. The ChatGPT maker added that its largest planned frontier reinforcement learning run remains on hold while smaller-scale training and evaluations continue.
The move comes after OpenAI disclosed an incident involving AI start-up Hugging Face in July, when AI systems breached the company’s systems during a cybersecurity evaluation. OpenAI stated that it is implementing stronger monitoring, alignment and security safeguards for models under development.
OpenAI plans to more closely track how its most capable unreleased models work through problems and use online tools to alert safety teams to worrying behaviour within 30 minutes. The company also added controls designed to prevent some models from accessing the internet while carrying out higher-risk tasks. Its approach rests on three safeguards: monitoring to detect and respond to concerning behaviour, alignment to reduce the likelihood of harmful or unauthorised actions, and security measures to limit what AI systems can access or affect.