← 審査済みの問い

審査済み

AI 同士の討論によって、弱い審判が正しい結論に到達できるか?

Can debate between AI systems let a weaker judge reach correct answers?

ai-alignmentmachine-learningscalable-oversight

問題文

答えを直接検証できない審判が、AI 同士を討論させることで、真の答えを確実に選べるかを決定せよ。

Decide whether a debate protocol between competing AI systems lets a judge who cannot verify the answer directly reliably select the true answer.

背景

Debate was proposed as a scalable oversight mechanism: an honest debater should be able to expose a dishonest one, so the judge only has to evaluate the final step. Experiments found a failure mode in which neither debater can locate the flaw in an argument that both know to be flawed.

アプローチ · 2 件

解決したかどうかは、問いではなくアプローチごとに決まります。Atlas は判定しません。外部の判定者が何をしたかを記録します。

  • Debate protocols with honesty as an equilibrium

    実験の再現

    到達状況
    進行中
    判定者
    empirical replication with human and model judges; no decision procedure settles the general claim · 再現
    定式化
    Design a debate protocol in which the honest strategy wins, and show experimentally that judges track the truth.問いと同値

    保たれているもの

    • Targets the claim as stated: a weaker judge reaching correct answers

    弱まっているもの

    • Results are protocol-specific and depend on debater capability; no general guarantee has been established empirically

    加えられた仮定・条件

    • Assumptions about debater optimality and judge behaviour

    閉じた道

    証拠

    • preprintarXiv:1805.00899G. Irving, P. Christiano, D. Amodei, AI safety via debate (2018)
  • Complexity-theoretic guarantees

    査読付きの証明

    到達状況
    進行中
    判定者
    peer review of the theoretical claims, under explicitly stated assumptions · 査読
    定式化
    Prove that the protocol lets a bounded judge decide a class of problems, given stated assumptions about the debaters.問いより弱い

    保たれているもの

    • Gives a guarantee rather than an experiment

    弱まっているもの

    • The guarantee holds under assumptions (debater optimality, access to the same computation) that real systems do not satisfy

    加えられた仮定・条件

    • Formal assumptions standing in for empirical facts about the models and the judge

    閉じた道

    証拠

    • preprintarXiv:2311.14125Scalable AI safety via doubly-efficient debate (2023): debate protocols with complexity-theoretic guarantees under stated assumptions

出典

記録

URI
https://atlasalt.com/q/651bfc14-cefc-4ba6-96a7-9b33e2aa149e
登録
2026-09-17
最終レビュー
2026-09-19
次回レビュー期限
2026-12-18
版
8a1a19f20d9c
ライセンス
CC-BY-4.0
立場
record_only(Atlas は判定しない)