GeoVerse Lab
โ† AI & Computing Sciences Division

AI Safety & Alignment Center

Researches AI alignment, mechanistic interpretability, robustness, and evaluation methods - including RLHF, red-teaming, and scalable oversight - to ensure advanced AI systems remain safe, controllable, and beneficial when deployed in high-stakes scientific and societal settings.

โš™๏ธ GeoVerse System Administration Division๐Ÿ–ฅ๏ธ AI & Computing Sciences Division๐Ÿ”ฌ Basic Sciences Divisionโšก Intelligent Geophysical Exploration Division๐Ÿ›ข๏ธ Resource & Energy Engineering Division๐ŸŒ Applied Geoscience Solutions Division๐Ÿ’ผ Economics, Policy & Strategy Division๐ŸŽ“ Education & Training Development Division๐Ÿš€ Innovation & International Collaboration Division
Machine Learning CenterNatural Language Processing CenterComputer Vision CenterHigh-Performance Computing CenterQuantum Computing CenterGenerative AI & Foundation Models CenterReinforcement Learning & Agents CenterAI Safety & Alignment CenterAI for Science CenterRobotics & Embodied AI Center
Norbert Wiener๐Ÿ”‘
โญ Center Chief
Norbert Wiener
Center Head
๐Ÿ’ก Cybernetics founder, early warnings on AI risk
Auguste Kerckhoffs๐Ÿ”‘
1835โ€“1903
Auguste Kerckhoffs
Researcher
๐Ÿ’ก Formulated Kerckhoffs's Principle โ€” a system must remain secure even if everything about it except the key is public knowledge โ€” the foundational security-engineering assumption underlying adversarial robustness and red-teaming evaluation of AI systems
๐Ÿ–ฅ๏ธ AI & Computing Sciences DivisionAI Safety & Alignment Center
W. Ross Ashby๐Ÿ”‘
1903โ€“1972
W. Ross Ashby
Researcher
๐Ÿ’ก Founded cybernetic control theory (Law of Requisite Variety, the Homeostat), establishing how a regulator must model a system's variety to control it โ€” the systems-theoretic root of AI alignment
๐Ÿ–ฅ๏ธ AI & Computing Sciences DivisionAI Safety & Alignment Center
๐Ÿง‘โ€๐Ÿ”ฌ
๐Ÿ”‘
1906โ€“1985
Bruno de Finetti
Researcher
๐Ÿ’ก Founded the subjectivist (Bayesian) theory of probability and the exchangeability theorem, and defined calibration as the criterion for a good subjective probability forecaster โ€” foundational to evaluating whether an AI system's confidence estimates are trustworthy
๐Ÿ–ฅ๏ธ AI & Computing Sciences DivisionAI Safety & Alignment Center
Lev Pontryagin๐Ÿ”‘
1908โ€“1988
Lev Pontryagin
Researcher
๐Ÿ’ก Founded optimal control theory (Pontryagin's Maximum Principle), the mathematical framework for steering a dynamical system toward a goal under constraints โ€” directly relevant to the theory of maintaining control/corrigibility over autonomous AI agents
๐Ÿ–ฅ๏ธ AI & Computing Sciences DivisionAI Safety & Alignment Center
John Forbes Nash Jr.๐Ÿ”‘
1928โ€“2015
John Forbes Nash Jr.
Researcher
๐Ÿ’ก Founded non-cooperative game theory and the Nash equilibrium, the formal framework for analyzing strategic interaction between agents โ€” the game-theoretic foundation of AI safety debate and multi-agent oversight protocols
๐Ÿ–ฅ๏ธ AI & Computing Sciences DivisionAI Safety & Alignment Center
Immanuel Kant๐Ÿ”‘
1724โ€“1804
Immanuel Kant
Researcher
๐Ÿ’ก Founded deontological ethics (the Categorical Imperative โ€” act only on principles universalizable and treating persons as ends, not merely means), one of the two dominant ethical frameworks (alongside utilitarianism) debated for encoding constraints in AI systems
๐Ÿ–ฅ๏ธ AI & Computing Sciences DivisionAI Safety & Alignment Center
John Rawls๐Ÿ”‘
1921โ€“2002
John Rawls
Researcher
๐Ÿ’ก Founded modern theory of distributive justice (justice as fairness, the veil of ignorance, the difference principle), the philosophical foundation most cited in algorithmic fairness research for defining what a 'fair' AI decision procedure should achieve
๐Ÿ–ฅ๏ธ AI & Computing Sciences DivisionAI Safety & Alignment Center
Hans Jonas๐Ÿ”‘
1903โ€“1993
Hans Jonas
Researcher
๐Ÿ’ก Formulated the 'imperative of responsibility' โ€” that the unprecedented power of modern technology demands a new ethics of long-term, precautionary responsibility to future generations, foundational to AI governance frameworks
๐Ÿ–ฅ๏ธ AI & Computing Sciences DivisionAI Safety & Alignment Center
๐Ÿง‘โ€๐Ÿ”ฌ
๐Ÿ”‘
1904โ€“1985
Donald O. Hebb
Researcher
๐Ÿ’ก Proposed the theory that neurons wire together via correlated activity ('cells that fire together wire together'), founding the study of how distributed representations encode concepts โ€” the theoretical root of interpreting learned features in neural networks
๐Ÿ–ฅ๏ธ AI & Computing Sciences DivisionAI Safety & Alignment Center
Leonid Hurwicz๐Ÿ”‘
1917โ€“2008
Leonid Hurwicz
Researcher
๐Ÿ’ก Founded mechanism design theory, showing how to design incentive structures so that self-interested agents' rational behavior reveals their true preferences and produces desired outcomes โ€” directly foundational to inverse reinforcement learning and incentive alignment
๐Ÿ–ฅ๏ธ AI & Computing Sciences DivisionAI Safety & Alignment Center
Frank P. Ramsey๐Ÿ”‘
1903โ€“1930
Frank P. Ramsey
Researcher
๐Ÿ’ก Founded subjective expected-utility theory, providing the axiomatic basis for representing an agent's preferences as a utility function derivable from choices โ€” the mathematical foundation of reward modeling from human feedback
๐Ÿ–ฅ๏ธ AI & Computing Sciences DivisionAI Safety & Alignment Center
๐Ÿง‘โ€๐Ÿ”ฌ
๐Ÿ”‘
1925โ€“2019
Charles Perrow
Researcher
๐Ÿ’ก Founded Normal Accident Theory โ€” that catastrophic failures are inevitable in tightly coupled, complex systems regardless of safeguards โ€” a foundational framework for analyzing systemic risk in complex, opaque AI systems
๐Ÿ–ฅ๏ธ AI & Computing Sciences DivisionAI Safety & Alignment Center
J. C. R. Licklider๐Ÿ”‘
1915โ€“1990
J. C. R. Licklider
Researcher
๐Ÿ’ก Founded the vision of 'man-computer symbiosis' โ€” humans and machines cooperating with complementary strengths under meaningful human oversight โ€” the foundational framework for trustworthy, human-compatible AI system design
๐Ÿ–ฅ๏ธ AI & Computing Sciences DivisionAI Safety & Alignment Center
๐Ÿง‘โ€๐Ÿ”ฌ
๐Ÿ”‘
1936โ€“2001
Robert W. Floyd
Researcher
๐Ÿ’ก Founded formal program verification (Floyd-Hoare logic, assigning preconditions/postconditions to prove program correctness), the mathematical foundation for formally verifying that safety-critical AI systems satisfy specified properties
๐Ÿ–ฅ๏ธ AI & Computing Sciences DivisionAI Safety & Alignment Center
Herman Kahn๐Ÿ”‘
1922โ€“1983
Herman Kahn
Researcher
๐Ÿ’ก Founded systematic future/scenario analysis for low-probability, catastrophic risks ('thinking about the unthinkable'), pioneering the methodology of existential-risk analysis later applied to advanced AI
๐Ÿ–ฅ๏ธ AI & Computing Sciences DivisionAI Safety & Alignment Center