← 審査済みの問い

審査済み

行動だけから、その主体の選好を推定できるか?

Can an agent's preferences be inferred from its behaviour alone?

ai-alignmentmachine-learningvalue-learning

問題文

合理的とは限らない主体の報酬関数や選好を、追加の仮定なしに、その行動の観察だけから同定できるかを決定せよ。

Decide whether the reward function or preferences of a possibly irrational agent can be identified from observations of its behaviour without further assumptions.

背景

Value learning from human behaviour is a central alignment proposal. The decomposition of behaviour into a planner and a reward is not unique: Armstrong and Mindermann show simplicity priors do not break the tie, which makes the identification question, not the optimisation question, the hard part.

アプローチ · 2 件

解決したかどうかは、問いではなくアプローチごとに決まります。Atlas は判定しません。外部の判定者が何をしたかを記録します。

  • Inverse reinforcement learning

    実験の再現

    到達状況
    進行中
    判定者
    empirical replication on benchmarks; no decision procedure settles identifiability in general · 再現
    定式化
    Recover the reward function that best explains observed behaviour, assuming the agent is (approximately) optimising it.問いより弱い

    保たれているもの

    • Recovers a reward that reproduces the observed behaviour

    弱まっているもの

    • Reward functions are not identifiable from behaviour: many pairs of planner and reward explain the same actions equally well

    加えられた仮定・条件

    • A rationality assumption that stands in for the missing information

    閉じた道

    • Armstrong and Mindermann show that for any behaviour there are planner-reward decompositions that are equally simple and radically different, including one whose reward is the negation of another. Simplicity priors therefore cannot pick out the intended preferences; the gap must be closed by assumptions from outside the data.無条件
      • preprintarXiv:1712.05812S. Armstrong, S. Mindermann, Occam's razor is insufficient to infer the preferences of irrational agents (NeurIPS 2018)

    証拠

    • otherA. Ng, S. Russell, Algorithms for inverse reinforcement learning, ICML 2000, 663-670Inverse reinforcement learning; the reward function is not identifiable from behaviour alone
  • Interactive and assistive formulations

    実験の再現

    到達状況
    進行中
    判定者
    empirical replication in simulated and human studies · 再現
    定式化
    Treat value learning as a two-player game in which the human's actions are informative and the machine can ask, rather than as passive observation.問いとずれている

    保たれているもの

    • Adds information the passive setting lacks, which is where the identifiability failure comes from

    弱まっているもの

    • Still requires a model of how the human's answers relate to their preferences; the unidentifiability reappears in that model

    加えられた仮定・条件

    • Assumptions about human rationality in the interaction protocol

    閉じた道

    • Goodhart-style failures: optimising a learned proxy diverges from the intended objective precisely where the proxy is least constrained by the data, which is the region the interaction did not cover.無条件
      • preprintarXiv:1803.04585D. Manheim, S. Garrabrant, Categorizing variants of Goodhart's Law (2018)

    証拠

    • preprintarXiv:1606.03137D. Hadfield-Menell, A. Dragan, P. Abbeel, S. Russell, Cooperative inverse reinforcement learning (2016)

出典

  • preprintarXiv:1712.05812S. Armstrong, S. Mindermann, Occam's razor is insufficient to infer the preferences of irrational agents (NeurIPS 2018)
  • otherA. Ng, S. Russell, Algorithms for inverse reinforcement learning, ICML 2000, 663-670Inverse reinforcement learning; the reward function is not identifiable from behaviour alone

記録

URI
https://atlasalt.com/q/3e7d19e2-45c6-4eba-8506-4ccda02bd6f5
登録
2026-09-17
最終レビュー
2026-09-19
次回レビュー期限
2026-12-18
版
7cd5a48cc7b4
ライセンス
CC-BY-4.0
立場
record_only(Atlas は判定しない)