← 審査済みの問い

審査済み

モデルに、人間がそう信じそうなことではなく、モデル自身が知っていることを報告させられるか?(ELK)

Can a model be trained to report what it knows rather than what a human would believe?

ai-alignmentmachine-learningscalable-oversight

問題文

モデル内部の知識と、人間の評価者が正しいと判断する内容がずれる状況で、前者を報告させる訓練手法が存在するかを決定せよ。

Decide whether there is a training strategy that makes a model report its own internal knowledge of a situation, rather than report whatever a human evaluator would judge to be true, in cases where the two come apart.

背景

Posed by the Alignment Research Center in 2021. The report is unusual in form: it lists proposed training strategies and, for each, a counterexample in which the strategy trains a human simulator instead of a knowledge reporter. It is therefore already a barrier ledger written in prose.

アプローチ · 2 件

解決したかどうかは、問いではなくアプローチごとに決まります。Atlas は判定しません。外部の判定者が何をしたかを記録します。

  • Training strategies with regularisers

    実験の再現

    到達状況
    判定者なし・判定がつかない
    判定者
    none so far: proposals are judged by counterexample in the report and on the alignment forum, not by an experiment that settles them · 判定手続きなし
    定式化
    Penalise reporters that are complex, slow, or dependent on the human's beliefs, so that the direct reporter is favoured over the human simulator.問いより弱い

    保たれているもの

    • Targets exactly the failure the question is about

    弱まっているもの

    • Each strategy in the report is met by a counterexample in which a human simulator is still preferred

    加えられた仮定・条件

    • Assumptions about the relative complexity of the reporter and the simulator

    閉じた道

    証拠

  • Interpretability-based reporting

    実験の再現

    到達状況
    進行中
    判定者
    empirical replication, once a method exists that can be tested · 再現
    定式化
    Read the answer out of the model's internal representations instead of training an output head on human labels.問いとずれている

    保たれているもの

    • Avoids taking the human's judgement as the training signal

    弱まっているもの

    • Requires interpretability strong enough to identify the relevant internal state and to know it is the right one, which is itself unsolved

    加えられた仮定・条件

    • The assumption that the knowledge exists as a locatable internal state

出典

記録

URI
https://atlasalt.com/q/186d9c3a-bd3a-46fc-9e1b-3c673ff14650
登録
2026-09-17
最終レビュー
2026-09-19
次回レビュー期限
2026-12-18
版
5b242e1f3bc9
ライセンス
CC-BY-4.0
立場
record_only(Atlas は判定しない)