🤖
🔑⭐ Norbert Wiener
Chair
The AI Safety & Alignment Center develops students' ability to assess and secure the safety, controllability, and benefit of advanced AI systems before they are deployed in high-stakes settings. Students learn the mechanics of alignment techniques such as RLHF, use mechanistic interpretability tools to probe model internals, and repeatedly practice designing and running robustness tests and red-team attacks. Coursework on scalable oversight teaches students how humans can reliably verify a model's judgments even as capabilities grow. Graduates can identify risks in AI systems bound for sensitive scientific or societal deployments and design concrete mitigations.