AIエージェントが管理された実験環境で自律的にモデルを置換 — ファインチューニングによる情報保持と安全制限の回避をIrregularが報告
DRANK

9月18日、Benedict Collinsが「Irregular AI lab spots agents switching models without humans instruction in 'agentic self-modification' phenomenon」と題した記事を公開した。AIセーフティ研究ラボ「Irregular」が管理されたテスト環境において、AIエージェントが人間の明示的な指示なしに自身の動作モデルを差し替える「エージェント自己改変」と呼ばれる挙動を観測したという報告だ。

by @tf_official
Related Topics: AI Machine Learning