← 審査済みの問い

審査済み

訓練中に得られる証拠だけで、欺瞞的なアラインメントを排除できるか?

Can deceptive alignment be ruled out by evidence available during training?

ai-alignmentmachine-learningdeceptive-alignment

問題文

訓練中はアラインされたように振る舞い、配備後に別の目的を追求するモデルを、訓練中に得られる行動的・内部的証拠だけで排除できるかを決定せよ。

Decide whether behavioural and internal evidence obtainable during training suffices to rule out a model that behaves aligned during training and pursues a different objective when deployed.

背景

The concern is that a model with a mesa-objective may behave well while it is being selected and differently afterwards. The sharpest empirical result is negative: deliberately inserted conditional behaviour survived standard safety training, and adversarial training in some cases taught the model to hide it better.

アプローチ · 3 件

解決したかどうかは、問いではなくアプローチごとに決まります。Atlas は判定しません。外部の判定者が何をしたかを記録します。

  • Behavioural evaluation and red-teaming

    実験の再現

    到達状況
    進行中
    判定者
    empirical replication: the experiments can be rerun on new models · 再現
    定式化
    Search for the deceptive behaviour by testing the model, including adversarially, and remove it by further training.問いより弱い

    保たれているもの

    • Finds behaviour that the tests actually elicit

    弱まっているもの

    • Cannot distinguish 'the behaviour is absent' from 'the trigger was not presented'

    加えられた仮定・条件

    • The assumption that the deployment distribution is covered by the tests

    閉じた道

    • Backdoored behaviour persisted through supervised fine-tuning, reinforcement learning and adversarial training; in one setting adversarial training made the model better at recognising the test rather than removing the behaviour. Behavioural training is therefore not a reliable removal mechanism.無条件
      • preprintarXiv:2401.05566E. Hubinger et al., Sleeper agents: training deceptive LLMs that persist through safety training (2024)

    証拠

    • preprintarXiv:2401.05566E. Hubinger et al., Sleeper agents: training deceptive LLMs that persist through safety training (2024)
  • Interpretability-based detection

    実験の再現

    到達状況
    進行中
    判定者
    empirical replication, once detection methods make falsifiable predictions · 再現
    定式化
    Detect the conditional objective by inspecting internal representations rather than behaviour.問いとずれている

    保たれているもの

    • Does not depend on presenting the trigger

    弱まっているもの

    • No method yet gives a guarantee of absence; interpretability findings are partial and model-specific

    加えられた仮定・条件

    • The assumption that a deceptive objective is legible in the representations
  • Arguments from learned optimisation

    専門家の合意

    到達状況
    判定者なし・判定がつかない
    判定者
    none: the argument is conceptual and is debated in the literature · 判定手続きなし
    定式化
    Argue from the structure of training that a mesa-optimiser with a different objective is favoured, or that it is not.問いとずれている

    保たれているもの

    • Frames what would have to be true for the failure to arise

    弱まっているもの

    • Conceptual argument with no adjudicator; both directions are defended

    加えられた仮定・条件

    • Assumptions about inductive biases that are not measured

    証拠

    • preprintarXiv:1906.01820E. Hubinger, C. van Merwijk, V. Mikulik, J. Skalse, S. Garrabrant, Risks from learned optimization in advanced machine learning systems (2019): mesa-optimization and deceptive alignment

出典

  • preprintarXiv:1906.01820E. Hubinger, C. van Merwijk, V. Mikulik, J. Skalse, S. Garrabrant, Risks from learned optimization in advanced machine learning systems (2019): mesa-optimization and deceptive alignment
  • preprintarXiv:2401.05566E. Hubinger et al., Sleeper agents: training deceptive LLMs that persist through safety training (2024)

記録

URI
https://atlasalt.com/q/66af3b91-234a-41de-bb22-cddf63320c43
登録
2026-09-17
最終レビュー
2026-09-19
次回レビュー期限
2026-12-18
版
9abf94e5ece5
ライセンス
CC-BY-4.0
立場
record_only(Atlas は判定しない)