🎓GeoAcademyGeoVerse Lab
AI & Computing Sciences College

AI Safety & Alignment Center

The AI Safety & Alignment Center develops students' ability to assess and secure the safety, controllability, and benefit of advanced AI systems before they are deployed in high-stakes settings. Students learn the mechanics of alignment techniques such as RLHF, use mechanistic interpretability tools to probe model internals, and repeatedly practice designing and running robustness tests and red-team attacks. Coursework on scalable oversight teaches students how humans can reliably verify a model's judgments even as capabilities grow. Graduates can identify risks in AI systems bound for sensitive scientific or societal deployments and design concrete mitigations.

⚙️ GeoAcademy System Administration College🖥️ AI & Computing Sciences College🔬 Basic Sciences College⚡ Intelligent Geophysical Exploration College🛢️ Resource & Energy Engineering College🌍 Applied Geoscience Solutions College💼 Economics, Policy & Strategy for the Future College🎓 Education & Training Development College🚀 Innovation & International Collaboration College
Machine Learning DepartmentComputer Vision DepartmentNatural Language Processing DepartmentHigh-Performance Computing DepartmentQuantum Computing DepartmentGenerative AI & Foundation Models CenterReinforcement Learning & Agents CenterAI Safety & Alignment CenterAI for Science CenterRobotics & Embodied AI Center
🤖
🔑
Norbert Wiener
Chair

Researchers (real GeoVerse Lab members) 15

W. Ross Ashby🔑
1903–1972
W. Ross Ashby
Researcher
💡 Founded cybernetic control theory (Law of Requisite Variety, the Homeostat), establishing how a regulator must model a system's variety to control it — the systems-theoretic root of AI alignment
Frank P. Ramsey🔑
1903–1930
Frank P. Ramsey
Researcher
💡 Founded subjective expected-utility theory, providing the axiomatic basis for representing an agent's preferences as a utility function derivable from choices — the mathematical foundation of reward modeling from human feedback
🧑‍🔬
🔑
1904–1985
Donald O. Hebb
Researcher
💡 Proposed the theory that neurons wire together via correlated activity ('cells that fire together wire together'), founding the study of how distributed representations encode concepts — the theoretical root of interpreting learned features in neural networks
Auguste Kerckhoffs🔑
1835–1903
Auguste Kerckhoffs
Researcher
💡 Formulated Kerckhoffs's Principle — a system must remain secure even if everything about it except the key is public knowledge — the foundational security-engineering assumption underlying adversarial robustness and red-teaming evaluation of AI systems
John Forbes Nash Jr.🔑
1928–2015
John Forbes Nash Jr.
Researcher
💡 Founded non-cooperative game theory and the Nash equilibrium, the formal framework for analyzing strategic interaction between agents — the game-theoretic foundation of AI safety debate and multi-agent oversight protocols
Hans Jonas🔑
1903–1993
Hans Jonas
Researcher
💡 Formulated the 'imperative of responsibility' — that the unprecedented power of modern technology demands a new ethics of long-term, precautionary responsibility to future generations, foundational to AI governance frameworks
🧑‍🔬
🔑
1936–2001
Robert W. Floyd
Researcher
💡 Founded formal program verification (Floyd-Hoare logic, assigning preconditions/postconditions to prove program correctness), the mathematical foundation for formally verifying that safety-critical AI systems satisfy specified properties
Herman Kahn🔑
1922–1983
Herman Kahn
Researcher
💡 Founded systematic future/scenario analysis for low-probability, catastrophic risks ('thinking about the unthinkable'), pioneering the methodology of existential-risk analysis later applied to advanced AI
Leonid Hurwicz🔑
1917–2008
Leonid Hurwicz
Researcher
💡 Founded mechanism design theory, showing how to design incentive structures so that self-interested agents' rational behavior reveals their true preferences and produces desired outcomes — directly foundational to inverse reinforcement learning and incentive alignment
John Rawls🔑
1921–2002
John Rawls
Researcher
💡 Founded modern theory of distributive justice (justice as fairness, the veil of ignorance, the difference principle), the philosophical foundation most cited in algorithmic fairness research for defining what a 'fair' AI decision procedure should achieve
J. C. R. Licklider🔑
1915–1990
J. C. R. Licklider
Researcher
💡 Founded the vision of 'man-computer symbiosis' — humans and machines cooperating with complementary strengths under meaningful human oversight — the foundational framework for trustworthy, human-compatible AI system design
🧑‍🔬
🔑
1925–2019
Charles Perrow
Researcher
💡 Founded Normal Accident Theory — that catastrophic failures are inevitable in tightly coupled, complex systems regardless of safeguards — a foundational framework for analyzing systemic risk in complex, opaque AI systems
Immanuel Kant🔑
1724–1804
Immanuel Kant
Researcher
💡 Founded deontological ethics (the Categorical Imperative — act only on principles universalizable and treating persons as ends, not merely means), one of the two dominant ethical frameworks (alongside utilitarianism) debated for encoding constraints in AI systems
🧑‍🔬
🔑
1906–1985
Bruno de Finetti
Researcher
💡 Founded the subjectivist (Bayesian) theory of probability and the exchangeability theorem, and defined calibration as the criterion for a good subjective probability forecaster — foundational to evaluating whether an AI system's confidence estimates are trustworthy
Lev Pontryagin🔑
1908–1988
Lev Pontryagin
Researcher
💡 Founded optimal control theory (Pontryagin's Maximum Principle), the mathematical framework for steering a dynamical system toward a goal under constraints — directly relevant to the theory of maintaining control/corrigibility over autonomous AI agents