審査済み
行動だけから、その主体の選好を推定できるか?
Can an agent's preferences be inferred from its behaviour alone?
ai-alignmentmachine-learningvalue-learning
問題文
合理的とは限らない主体の報酬関数や選好を、追加の仮定なしに、その行動の観察だけから同定できるかを決定せよ。
Decide whether the reward function or preferences of a possibly irrational agent can be identified from observations of its behaviour without further assumptions.
背景
Value learning from human behaviour is a central alignment proposal. The decomposition of behaviour into a planner and a reward is not unique: Armstrong and Mindermann show simplicity priors do not break the tie, which makes the identification question, not the optimisation question, the hard part.
アプローチ · 2 件
解決したかどうかは、問いではなくアプローチごとに決まります。Atlas は判定しません。外部の判定者が何をしたかを記録します。
Inverse reinforcement learning
実験の再現
- 到達状況
- 進行中
- 判定者
- empirical replication on benchmarks; no decision procedure settles identifiability in general · 再現
- 定式化
- Recover the reward function that best explains observed behaviour, assuming the agent is (approximately) optimising it.問いより弱い
保たれているもの
- Recovers a reward that reproduces the observed behaviour
弱まっているもの
- Reward functions are not identifiable from behaviour: many pairs of planner and reward explain the same actions equally well
加えられた仮定・条件
- A rationality assumption that stands in for the missing information
閉じた道
- Armstrong and Mindermann show that for any behaviour there are planner-reward decompositions that are equally simple and radically different, including one whose reward is the negation of another. Simplicity priors therefore cannot pick out the intended preferences; the gap must be closed by assumptions from outside the data.無条件
- preprintarXiv:1712.05812S. Armstrong, S. Mindermann, Occam's razor is insufficient to infer the preferences of irrational agents (NeurIPS 2018)
証拠
- otherA. Ng, S. Russell, Algorithms for inverse reinforcement learning, ICML 2000, 663-670Inverse reinforcement learning; the reward function is not identifiable from behaviour alone
Interactive and assistive formulations
実験の再現
- 到達状況
- 進行中
- 判定者
- empirical replication in simulated and human studies · 再現
- 定式化
- Treat value learning as a two-player game in which the human's actions are informative and the machine can ask, rather than as passive observation.問いとずれている
保たれているもの
- Adds information the passive setting lacks, which is where the identifiability failure comes from
弱まっているもの
- Still requires a model of how the human's answers relate to their preferences; the unidentifiability reappears in that model
加えられた仮定・条件
- Assumptions about human rationality in the interaction protocol
閉じた道
- Goodhart-style failures: optimising a learned proxy diverges from the intended objective precisely where the proxy is least constrained by the data, which is the region the interaction did not cover.無条件
- preprintarXiv:1803.04585D. Manheim, S. Garrabrant, Categorizing variants of Goodhart's Law (2018)
証拠
- preprintarXiv:1606.03137D. Hadfield-Menell, A. Dragan, P. Abbeel, S. Russell, Cooperative inverse reinforcement learning (2016)
出典
- preprintarXiv:1712.05812S. Armstrong, S. Mindermann, Occam's razor is insufficient to infer the preferences of irrational agents (NeurIPS 2018)
- otherA. Ng, S. Russell, Algorithms for inverse reinforcement learning, ICML 2000, 663-670Inverse reinforcement learning; the reward function is not identifiable from behaviour alone
記録
- URI
- https://atlasalt.com/q/3e7d19e2-45c6-4eba-8506-4ccda02bd6f5
- 登録
- 2026-09-17
- 最終レビュー
- 2026-09-19
- 次回レビュー期限
- 2026-12-18
- 版
- 7cd5a48cc7b4
- ライセンス
- CC-BY-4.0
- 立場
- record_only(Atlas は判定しない)