審査済み
訓練中に得られる証拠だけで、欺瞞的なアラインメントを排除できるか?
Can deceptive alignment be ruled out by evidence available during training?
ai-alignmentmachine-learningdeceptive-alignment
問題文
訓練中はアラインされたように振る舞い、配備後に別の目的を追求するモデルを、訓練中に得られる行動的・内部的証拠だけで排除できるかを決定せよ。
Decide whether behavioural and internal evidence obtainable during training suffices to rule out a model that behaves aligned during training and pursues a different objective when deployed.
背景
The concern is that a model with a mesa-objective may behave well while it is being selected and differently afterwards. The sharpest empirical result is negative: deliberately inserted conditional behaviour survived standard safety training, and adversarial training in some cases taught the model to hide it better.
アプローチ · 3 件
解決したかどうかは、問いではなくアプローチごとに決まります。Atlas は判定しません。外部の判定者が何をしたかを記録します。
Behavioural evaluation and red-teaming
実験の再現
- 到達状況
- 進行中
- 判定者
- empirical replication: the experiments can be rerun on new models · 再現
- 定式化
- Search for the deceptive behaviour by testing the model, including adversarially, and remove it by further training.問いより弱い
保たれているもの
- Finds behaviour that the tests actually elicit
弱まっているもの
- Cannot distinguish 'the behaviour is absent' from 'the trigger was not presented'
加えられた仮定・条件
- The assumption that the deployment distribution is covered by the tests
閉じた道
- Backdoored behaviour persisted through supervised fine-tuning, reinforcement learning and adversarial training; in one setting adversarial training made the model better at recognising the test rather than removing the behaviour. Behavioural training is therefore not a reliable removal mechanism.無条件
- preprintarXiv:2401.05566E. Hubinger et al., Sleeper agents: training deceptive LLMs that persist through safety training (2024)
証拠
- preprintarXiv:2401.05566E. Hubinger et al., Sleeper agents: training deceptive LLMs that persist through safety training (2024)
Interpretability-based detection
実験の再現
- 到達状況
- 進行中
- 判定者
- empirical replication, once detection methods make falsifiable predictions · 再現
- 定式化
- Detect the conditional objective by inspecting internal representations rather than behaviour.問いとずれている
保たれているもの
- Does not depend on presenting the trigger
弱まっているもの
- No method yet gives a guarantee of absence; interpretability findings are partial and model-specific
加えられた仮定・条件
- The assumption that a deceptive objective is legible in the representations
Arguments from learned optimisation
専門家の合意
- 到達状況
- 判定者なし・判定がつかない
- 判定者
- none: the argument is conceptual and is debated in the literature · 判定手続きなし
- 定式化
- Argue from the structure of training that a mesa-optimiser with a different objective is favoured, or that it is not.問いとずれている
保たれているもの
- Frames what would have to be true for the failure to arise
弱まっているもの
- Conceptual argument with no adjudicator; both directions are defended
加えられた仮定・条件
- Assumptions about inductive biases that are not measured
証拠
- preprintarXiv:1906.01820E. Hubinger, C. van Merwijk, V. Mikulik, J. Skalse, S. Garrabrant, Risks from learned optimization in advanced machine learning systems (2019): mesa-optimization and deceptive alignment
出典
- preprintarXiv:1906.01820E. Hubinger, C. van Merwijk, V. Mikulik, J. Skalse, S. Garrabrant, Risks from learned optimization in advanced machine learning systems (2019): mesa-optimization and deceptive alignment
- preprintarXiv:2401.05566E. Hubinger et al., Sleeper agents: training deceptive LLMs that persist through safety training (2024)
記録
- URI
- https://atlasalt.com/q/66af3b91-234a-41de-bb22-cddf63320c43
- 登録
- 2026-09-17
- 最終レビュー
- 2026-09-19
- 次回レビュー期限
- 2026-12-18
- 版
- 9abf94e5ece5
- ライセンス
- CC-BY-4.0
- 立場
- record_only(Atlas は判定しない)