darioamodei
近期帖子
我们必须为前沿设定节奏(We Must Pace the Frontier):我写了一篇新文章,讲为什么 AI 行业应该放慢速度,并给出了一个三步走的方案。 Anthropic 正在单方面践行其中的第一步。我们将向第三方评估机构提供永久性的、与员工同等级别的系统访问权限,使其能够核查我们是否遵守自身的安全措施、报告相关事件,并在训练过程中评估模型的对齐情况。 全文请见:https://t.co/OGyPb7yaYt
今天我发表了一篇新文章《应对 AI 指数级发展的政策》(Policy on the AI Exponential)。AI 正在以极快的速度发展——远超政策制定流程所能应对的速度。这篇文章阐述了我对这项技术当前所处阶段的看法,以及弥合这一差距所需采取的行动:https://t.co/Lh6PWae178
Cyber is the first clear and present danger from frontier AI models, but it won’t be the last. If we are able to collectively rise to the challenge and confront this risk, it could serve as a blueprint for addressing the even more difficult challenges that lie ahead of us.
Rather than release Mythos Preview to general availability, we’re giving defenders early controlled access in order to find and patch vulnerabilities before Mythos-class models proliferate across the ecosystem.
Glasswing is just the first step: patching and securing the world’s software infrastructure will be the work of months and years, and will require even broader cooperation across AI companies, cyberdefenders, software providers, governments, and more.