審査済み
モデルに、人間がそう信じそうなことではなく、モデル自身が知っていることを報告させられるか?(ELK)
Can a model be trained to report what it knows rather than what a human would believe?
ai-alignmentmachine-learningscalable-oversight
問題文
モデル内部の知識と、人間の評価者が正しいと判断する内容がずれる状況で、前者を報告させる訓練手法が存在するかを決定せよ。
Decide whether there is a training strategy that makes a model report its own internal knowledge of a situation, rather than report whatever a human evaluator would judge to be true, in cases where the two come apart.
背景
Posed by the Alignment Research Center in 2021. The report is unusual in form: it lists proposed training strategies and, for each, a counterexample in which the strategy trains a human simulator instead of a knowledge reporter. It is therefore already a barrier ledger written in prose.
アプローチ · 2 件
解決したかどうかは、問いではなくアプローチごとに決まります。Atlas は判定しません。外部の判定者が何をしたかを記録します。
Training strategies with regularisers
実験の再現
- 到達状況
- 判定者なし・判定がつかない
- 判定者
- none so far: proposals are judged by counterexample in the report and on the alignment forum, not by an experiment that settles them · 判定手続きなし
- 定式化
- Penalise reporters that are complex, slow, or dependent on the human's beliefs, so that the direct reporter is favoured over the human simulator.問いより弱い
保たれているもの
- Targets exactly the failure the question is about
弱まっているもの
- Each strategy in the report is met by a counterexample in which a human simulator is still preferred
加えられた仮定・条件
- Assumptions about the relative complexity of the reporter and the simulator
閉じた道
- The report gives a counterexample for every strategy it proposes: whenever the training signal is generated by human judgement, a model that predicts human judgement scores at least as well as one that reports its own knowledge.無条件
- otherhttps://docs.google.com/document/d/1WwsnJQstPq91_Yh-Ch2XRL8H_EpsnjrC1dwZXR37PC8/editThe ELK report itself: a list of proposed training strategies, each with a counterexample
証拠
- otherhttps://docs.google.com/document/d/1WwsnJQstPq91_Yh-Ch2XRL8H_EpsnjrC1dwZXR37PC8/editThe ELK report itself: a list of proposed training strategies, each with a counterexample
Interpretability-based reporting
実験の再現
- 到達状況
- 進行中
- 判定者
- empirical replication, once a method exists that can be tested · 再現
- 定式化
- Read the answer out of the model's internal representations instead of training an output head on human labels.問いとずれている
保たれているもの
- Avoids taking the human's judgement as the training signal
弱まっているもの
- Requires interpretability strong enough to identify the relevant internal state and to know it is the right one, which is itself unsolved
加えられた仮定・条件
- The assumption that the knowledge exists as a locatable internal state
出典
- otherhttps://www.alignment.org/blog/arcs-first-technical-report-eliciting-latent-knowledge/P. Christiano, A. Cotra, M. Xu, Eliciting Latent Knowledge (Alignment Research Center technical report, December 2021)
- otherhttps://docs.google.com/document/d/1WwsnJQstPq91_Yh-Ch2XRL8H_EpsnjrC1dwZXR37PC8/editThe ELK report itself: a list of proposed training strategies, each with a counterexample
記録
- URI
- https://atlasalt.com/q/186d9c3a-bd3a-46fc-9e1b-3c673ff14650
- 登録
- 2026-09-17
- 最終レビュー
- 2026-09-19
- 次回レビュー期限
- 2026-12-18
- 版
- 5b242e1f3bc9
- ライセンス
- CC-BY-4.0
- 立場
- record_only(Atlas は判定しない)