β AI & Computing Sciences Division
AI Safety & Alignment Center
Researches AI alignment, mechanistic interpretability, robustness, and evaluation methods - including RLHF, red-teaming, and scalable oversight - to ensure advanced AI systems remain safe, controllable, and beneficial when deployed in high-stakes scientific and societal settings.
Research Fields 15















