審査済み
AI 同士の討論によって、弱い審判が正しい結論に到達できるか?
Can debate between AI systems let a weaker judge reach correct answers?
ai-alignmentmachine-learningscalable-oversight
問題文
答えを直接検証できない審判が、AI 同士を討論させることで、真の答えを確実に選べるかを決定せよ。
Decide whether a debate protocol between competing AI systems lets a judge who cannot verify the answer directly reliably select the true answer.
背景
Debate was proposed as a scalable oversight mechanism: an honest debater should be able to expose a dishonest one, so the judge only has to evaluate the final step. Experiments found a failure mode in which neither debater can locate the flaw in an argument that both know to be flawed.
アプローチ · 2 件
解決したかどうかは、問いではなくアプローチごとに決まります。Atlas は判定しません。外部の判定者が何をしたかを記録します。
Debate protocols with honesty as an equilibrium
実験の再現
- 到達状況
- 進行中
- 判定者
- empirical replication with human and model judges; no decision procedure settles the general claim · 再現
- 定式化
- Design a debate protocol in which the honest strategy wins, and show experimentally that judges track the truth.問いと同値
保たれているもの
- Targets the claim as stated: a weaker judge reaching correct answers
弱まっているもの
- Results are protocol-specific and depend on debater capability; no general guarantee has been established empirically
加えられた仮定・条件
- Assumptions about debater optimality and judge behaviour
閉じた道
- Obfuscated arguments: a debater can present an argument that is flawed somewhere, where neither side can identify the flawed step within the debate's budget. The honest debater then cannot win by pointing at the error, which breaks the mechanism's central assumption.無条件
- otherhttps://www.alignmentforum.org/posts/PJLABqQ962hZEqhdB/debate-update-obfuscated-arguments-problemB. Barnes, P. Christiano, Debate update: obfuscated arguments problem (AI Alignment Forum, December 2020)
証拠
- preprintarXiv:1805.00899G. Irving, P. Christiano, D. Amodei, AI safety via debate (2018)
Complexity-theoretic guarantees
査読付きの証明
- 到達状況
- 進行中
- 判定者
- peer review of the theoretical claims, under explicitly stated assumptions · 査読
- 定式化
- Prove that the protocol lets a bounded judge decide a class of problems, given stated assumptions about the debaters.問いより弱い
保たれているもの
- Gives a guarantee rather than an experiment
弱まっているもの
- The guarantee holds under assumptions (debater optimality, access to the same computation) that real systems do not satisfy
加えられた仮定・条件
- Formal assumptions standing in for empirical facts about the models and the judge
閉じた道
- A theorem about idealised debaters says nothing about whether the assumptions hold of trained systems; the obfuscated-arguments failure appears exactly where the idealisation breaks.無条件
- otherhttps://www.alignmentforum.org/posts/PJLABqQ962hZEqhdB/debate-update-obfuscated-arguments-problemB. Barnes, P. Christiano, Debate update: obfuscated arguments problem (AI Alignment Forum, December 2020)
証拠
- preprintarXiv:2311.14125Scalable AI safety via doubly-efficient debate (2023): debate protocols with complexity-theoretic guarantees under stated assumptions
出典
- preprintarXiv:1805.00899G. Irving, P. Christiano, D. Amodei, AI safety via debate (2018)
- otherhttps://www.alignmentforum.org/posts/PJLABqQ962hZEqhdB/debate-update-obfuscated-arguments-problemB. Barnes, P. Christiano, Debate update: obfuscated arguments problem (AI Alignment Forum, December 2020)
記録
- URI
- https://atlasalt.com/q/651bfc14-cefc-4ba6-96a7-9b33e2aa149e
- 登録
- 2026-09-17
- 最終レビュー
- 2026-09-19
- 次回レビュー期限
- 2026-12-18
- 版
- 8a1a19f20d9c
- ライセンス
- CC-BY-4.0
- 立場
- record_only(Atlas は判定しない)