GeoVerse Lab
← AI & Computing Sciences Division

Generative AI & Foundation Models Center

Advances large-scale foundation models - LLMs, diffusion models, and multimodal systems, spanning pretraining, RLHF, and scaling laws - and adapts them for scientific reasoning, data synthesis, and geoscience-specific applications through domain adaptation and retrieval-augmented generation.

Machine Learning CenterNatural Language Processing CenterComputer Vision CenterHigh-Performance Computing CenterQuantum Computing CenterGenerative AI & Foundation Models CenterReinforcement Learning & Agents CenterAI Safety & Alignment CenterAI for Science CenterRobotics & Embodied AI Center
Research Fields 15
Sequence Modeling & Connectionist Architectures
πŸ”„ Sequence Modeling & Connectionist Architectures
Sequence modeling trains a neural network to process or generate ordered data, such as tex…
Diffusion Models & Score-Based Generative Modeling
🌫️ Diffusion Models & Score-Based Generative Modeling
Diffusion models generate new data by learning to reverse a fixed process that gradually a…
Statistical & Autoregressive Language Modeling
πŸ”€ Statistical & Autoregressive Language Modeling
A language model assigns a probability to a sequence of symbols, most commonly words or su…
Generative Probabilistic Modeling & Latent Variable Inference
🎲 Generative Probabilistic Modeling & Latent Variable Inference
A latent variable generative model posits that observed data is produced by first sampling…
Generative Adversarial Networks & Minimax Game-Theoretic Generation
βš”οΈ Generative Adversarial Networks & Minimax Game-Theoretic Generation
A generative adversarial network trains two neural networks in opposition: a generator tha…
Autoencoders & Latent Representation Learning
πŸ—œοΈ Autoencoders & Latent Representation Learning
An autoencoder is a neural network trained to reconstruct its own input after passing it t…
Scaling Laws & Emergent Behavior in Large Models
πŸ“ˆ Scaling Laws & Emergent Behavior in Large Models
A neural scaling law is an empirically observed power-law relationship between a language …
Multimodal & Vision-Language Foundation Models
πŸ”— Multimodal & Vision-Language Foundation Models
A multimodal foundation model learns a joint representation space in which corresponding p…
Retrieval-Augmented Generation & Information Retrieval
πŸ“š Retrieval-Augmented Generation & Information Retrieval
Retrieval-augmented generation retrieves a small set of relevant documents from an externa…
Word/Vector Embeddings & Distributional Semantics
🧡 Word/Vector Embeddings & Distributional Semantics
A word embedding represents each word or subword token as a dense numerical vector, learne…
In-Context Learning & Sequential/Adaptive Statistical Inference
πŸ’‘ In-Context Learning & Sequential/Adaptive Statistical Inference
In-context learning is the ability of a large language model to adapt its behaviour to a n…
Fine-Tuning & Transfer of Learning
πŸ”§ Fine-Tuning & Transfer of Learning
Fine-tuning adapts a foundation model already pretrained on a large, general-purpose datas…
Subword Tokenization
πŸ”€ Subword Tokenization
Subword tokenization splits raw text into a fixed vocabulary of frequently recurring chara…
Knowledge Distillation & Model Compression
πŸ§ͺ Knowledge Distillation & Model Compression
Knowledge distillation trains a smaller student network to reproduce the output behaviour …
Corpus Taxonomy & Domain-Adapted Data Curation
πŸ“š Corpus Taxonomy & Domain-Adapted Data Curation
Domain-adapted data curation organises a training corpus into a hierarchical taxonomy of s…