🎓GeoAcademyGeoVerse Lab
AI & Computing Sciences College

Reinforcement Learning & Agents Center

The Reinforcement Learning & Agents Center teaches sequential decision-making theory and autonomous agent architecture spanning deep RL, model-based world models, multi-agent systems, and LLM-based agents. Students formulate scientific exploration, experiment design, and adaptive control problems as reinforcement learning problems and train agents to solve them, progressing from single-agent settings to multi-agent collaboration. Through this progression, graduates reach the point where they can design and implement autonomous systems that make their own decisions in uncertain, changing environments.

⚙️ GeoAcademy System Administration College🖥️ AI & Computing Sciences College🔬 Basic Sciences College⚡ Intelligent Geophysical Exploration College🛢️ Resource & Energy Engineering College🌍 Applied Geoscience Solutions College💼 Economics, Policy & Strategy for the Future College🎓 Education & Training Development College🚀 Innovation & International Collaboration College
Machine Learning DepartmentComputer Vision DepartmentNatural Language Processing DepartmentHigh-Performance Computing DepartmentQuantum Computing DepartmentGenerative AI & Foundation Models CenterReinforcement Learning & Agents CenterAI Safety & Alignment CenterAI for Science CenterRobotics & Embodied AI Center
🤖
🔑
Richard Bellman
Chair

Researchers (real GeoVerse Lab members) 15

David Blackwell🔑
1919–2010
David Blackwell
Researcher
💡 Proved foundational optimality and convergence results for Markov decision processes (Blackwell optimality), establishing the rigorous mathematical conditions under which dynamic-programming-based sequential decision policies are provably optimal — extending and rigorizing Bellman's framework
Ivan Pavlov🔑
1849–1936
Ivan Pavlov
Researcher
💡 Discovered classical conditioning — that an organism learns to associate a predictive stimulus with a delayed reward through repeated temporal pairing — the foundational associative-learning phenomenon that temporal-difference learning algorithms in RL directly formalize mathematically
Magnus Hestenes🔑
1906–1991
Magnus Hestenes
Researcher
💡 Founded modern numerical methods for the calculus of variations and constrained optimization (conjugate gradient method, augmented Lagrangian methods), providing the gradient-based optimization machinery that policy gradient reinforcement learning algorithms directly apply to optimize policy parameters
🧑‍🔬
🔑
1941–1997
Harry Klopf
Researcher
💡 Proposed the 'hedonistic neuron' hypothesis — that individual neurons act to maximize local reward-like signals — directly inspiring Sutton and Barto's actor-critic architecture, which separates a policy ('actor') from a value-estimating critic
🧑‍🔬
🔑
1918–2016
Jay Wright Forrester
Researcher
💡 Founded System Dynamics, the discipline of building explicit feedback-loop simulation models of complex systems to predict and reason about their behavior — the direct conceptual and terminological ancestor of the 'world models' that model-based reinforcement learning agents learn and plan within
Lloyd Shapley🔑
1923–2016
Lloyd Shapley
Researcher
💡 Founded the theory of stochastic games, generalizing Markov decision processes to multiple interacting agents with competing or cooperative objectives — the exact mathematical framework underlying multi-agent reinforcement learning
🧑‍🔬
🔑
1901–1990
Arthur Samuel
Researcher
💡 Built the first self-improving game-playing program (checkers) that learned by playing against itself and evaluating board positions via minimax search, coining the term 'machine learning' and founding the self-play tree-search paradigm that MCTS and AlphaZero-class systems directly descend from
Arturo Rosenblueth🔑
1900–1970
Arturo Rosenblueth
Researcher
💡 Co-founded cybernetics with Wiener and Bigelow, formalizing the theory of goal-directed, feedback-driven behavior in machines — the founding theoretical framework for agents that pursue goals via perception-action-feedback loops, directly relevant to LLM-based autonomous agents using tools and feedback
Gustav Elfving🔑
1908–1984
Gustav Elfving
Researcher
💡 Founded the theory of optimal experimental design, determining how to choose experimental conditions (the 'Elfving set') to maximize the statistical information gained per experiment — foundational to Bayesian optimal experimental design for autonomous scientific exploration
🧑‍🔬
🔑
1927–2021
Peter Whittle
Researcher
💡 Developed the theory of restless bandits and Whittle indices, generalizing multi-armed bandit theory to more realistic settings where unobserved arms continue to evolve — foundational to modern exploration-exploitation algorithms balancing information gathering against exploitation of known-good options
🧑‍🔬
🔑
1927–1992
Allen Newell
Researcher
💡 Co-developed the General Problem Solver and means-ends analysis, the foundational method of decomposing a complex goal into a hierarchy of subgoals — directly foundational to hierarchical reinforcement learning's options and subgoal frameworks
Albert Bandura🔑
1925–2021
Albert Bandura
Researcher
💡 Founded Social Learning Theory, demonstrating that organisms learn complex behaviors by observing and imitating others (the famous Bobo doll experiments) without direct trial-and-error reward — the founding psychological phenomenon that imitation learning and behavior cloning in RL directly formalize
🧑‍🔬
🔑
1924–1976
Daniel Berlyne
Researcher
💡 Founded the experimental psychology of curiosity, showing that novelty, uncertainty, and complexity themselves drive exploratory behavior independent of external reward — the foundational psychological theory that intrinsic-motivation and curiosity-driven exploration bonuses in RL directly formalize
🧑‍🔬
🔑
1925–2014
Harold W. Kuhn
Researcher
💡 Developed the theory of extensive-form games (game trees with sequential moves and information sets), providing the exact formal representation of sequential, multi-step decision-making under uncertainty that both game-tree search and sequential multi-agent RL directly operate on
🧑‍🔬
🔑
1919–1997
Yakov Tsypkin
Researcher
💡 Founded the theory of adaptive and learning control systems, showing how a controller can adjust its own parameters online based on observed system behavior — directly foundational to adaptive control-theoretic reinforcement learning agents that tune their policies in response to a changing environment