| 2 OctFriday |
What are we aligning? Computation, agency, and incomplete delegation
readings
|
Daniele CondorelliWarwick |
| 9 OctFriday |
Training is not incentive provision: Rewards, learning rules, and the systems they produce
readings
|
Alkis Georgiadis-HarrisWarwick |
| 16 OctFriday |
Learning intentions and identifying preferences: Common interests do not eliminate inference problems
readings
|
Tanay VenkataLBS |
| 30 OctFriday |
Instrumental convergence: Resources, survival, and control
readings
|
Massimiliano FurlanWarwick |
| 6 NovFriday |
Corrigibility and effective authority: Preserving the principal’s ability to intervene
readings
|
Joe BasfordLSE |
| 13 NovFriday |
Measurement, incentives, and reward hacking: A useful diagnostic need not be a good target
readings
|
Joe BasfordLSE |
| 20 NovFriday |
Strategic evaluation and the limits of inspection: Strategic compliance, avoiding modification, and verifiability
readings
|
Hersh ChopraLSE |
| 27 NovFriday |
Oversight beyond the supervisor’s competence: Investigation, debate, and weak-to-strong generalisation
readings
|
Cecilia WoodsUK AISI |
| 4 DecFriday |
Interpretability as identification: Representation, intervention, and misleading explanations
readings
|
Matthew LevyLSE |
| 11 DecFriday |
Control, monitoring, and safe composition: Useful systems built from imperfectly trusted components
readings
|
Carlos VicheWarwick |