🎓GeoAcademyGeoVerse Lab
AI & Computing Sciences College

Generative AI & Foundation Models Center

The Generative AI & Foundation Models Center gives students a rigorous grasp of how large-scale foundation models - LLMs, diffusion models, and multimodal systems - are pretrained, aligned via RLHF, and governed by scaling laws, paired with hands-on training and fine-tuning of such models. Students run projects adapting general-purpose models to geoscience-specific applications, including scientific reasoning and data synthesis, using domain adaptation and retrieval-augmented generation. Graduates emerge as practicing AI engineers capable of customizing and deploying foundation models for a specific scientific domain.

⚙️ GeoAcademy System Administration College🖥️ AI & Computing Sciences College🔬 Basic Sciences College⚡ Intelligent Geophysical Exploration College🛢️ Resource & Energy Engineering College🌍 Applied Geoscience Solutions College💼 Economics, Policy & Strategy for the Future College🎓 Education & Training Development College🚀 Innovation & International Collaboration College
Machine Learning DepartmentComputer Vision DepartmentNatural Language Processing DepartmentHigh-Performance Computing DepartmentQuantum Computing DepartmentGenerative AI & Foundation Models CenterReinforcement Learning & Agents CenterAI Safety & Alignment CenterAI for Science CenterRobotics & Embodied AI Center
🤖
🔑
Ludwig Boltzmann
Chair

Researchers (real GeoVerse Lab members) 15

David Rumelhart🔑
1942–2011
David Rumelhart
Researcher
💡 Co-developed the backpropagation algorithm for training multilayer neural networks and founded parallel distributed processing (PDP) theory, the direct architectural ancestor of the deep sequence and transformer networks underlying modern foundation models
Paul Langevin🔑
1872–1946
Paul Langevin
Researcher
💡 Formulated the Langevin equation describing stochastic (Brownian) motion under random forcing, the exact stochastic differential equation whose discretization is the sampling procedure in modern diffusion/score-based generative models
Andrey Markov🔑
1856–1922
Andrey Markov
Researcher
💡 Founded the theory of Markov chains — stochastic processes where the next state depends only on the current state — first applied by Markov himself to model letter sequences in text, the direct mathematical ancestor of all autoregressive language modeling including LLMs
Pierre-Simon Laplace🔑
1749–1827
Pierre-Simon Laplace
Researcher
💡 Founded Bayesian probabilistic inference and the systematic theory of inferring unobserved causes (latent variables) from observed data, the direct philosophical and mathematical ancestor of latent-variable generative models (VAEs, diffusion models)
Emile Borel🔑
1871–1956
Emile Borel
Researcher
💡 Founded early game theory and the minimax concept for zero-sum games, the exact mathematical framework (a two-player minimax game between generator and discriminator) that defines Generative Adversarial Networks
Karl Pearson🔑
1857–1936
Karl Pearson
Researcher
💡 Invented Principal Component Analysis, the first method for learning a compressed low-dimensional latent representation that best reconstructs high-dimensional data — the linear ancestor of autoencoder-based representation learning
George Kingsley Zipf🔑
1902–1950
George Kingsley Zipf
Researcher
💡 Discovered Zipf's Law — the power-law distribution of word frequency in natural language — the founding empirical power-law regularity that motivates and parallels the power-law scaling laws governing large language model performance versus size, data, and compute
Ferdinand de Saussure🔑
1857–1913
Ferdinand de Saussure
Researcher
💡 Founded structural linguistics and semiotics, establishing the theory of the arbitrary sign (signifier/signified relationship) that underlies how multimodal models must learn shared, grounded representations across language and vision
🧑‍🔬
🔑
1927–1995
Gerard Salton
Researcher
💡 Founded modern information retrieval, inventing the vector space model representing documents and queries as vectors ranked by similarity — the exact retrieval mechanism (dense vector search) that retrieval-augmented generation systems use to fetch context for language models
🧑‍🔬
🔑
1890–1960
John Rupert Firth
Researcher
💡 Formulated the distributional hypothesis of meaning ('you shall know a word by the company it keeps'), the foundational linguistic principle that word embedding methods (word2vec, GloVe, and transformer embeddings) directly operationalize
Abraham Wald🔑
1902–1950
Abraham Wald
Researcher
💡 Founded statistical decision theory and sequential analysis, formalizing how an agent should update inference adaptively as new observations arrive within a single decision process — the statistical-decision-theoretic root of in-context learning within a model's forward pass
Edward Thorndike🔑
1874–1949
Edward Thorndike
Researcher
💡 Founded the psychological theory of transfer of learning (the 'identical elements' theory of why training on one task improves performance on another), the conceptual ancestor of fine-tuning and transfer learning in foundation models
🧑‍🔬
🔑
1925–1999
David A. Huffman
Researcher
💡 Invented Huffman coding, the optimal variable-length prefix-code compression algorithm based on symbol frequency, the same frequency-driven merging principle that Byte-Pair Encoding (BPE) tokenization in modern LLMs directly generalizes
🧑‍🔬
🔑
1926–2009
Ray Solomonoff
Researcher
💡 Founded algorithmic information theory and universal induction, defining the shortest program (most compressed description) that reproduces a given data pattern as its ideal predictor — the theoretical ideal that knowledge distillation approximates when compressing a large model into a smaller one
Melvil Dewey🔑
1851–1931
Melvil Dewey
Researcher
💡 Invented the Dewey Decimal Classification system for systematically organizing knowledge into hierarchical subject domains, the founding methodology for the domain classification and curation that shapes how domain-specific corpora (e.g. geoscience text) are organized for adapting foundation models