The R2L Lab at the University of Waterloo's Cheriton Department of Computer Science builds intelligent generalist agents that read, reason, and act across digital and physical environments. We organize our work around four long-running research threads: building and evaluating generalist agents, generalization, adaptation, and reasoning, language as supervision, and information access for autonomous reasoners. We are recruiting students and postdocs across all four.
Please read this before applying.
We rarely hire MS students or remote undergraduate students.
The University of Waterloo admits students each semester. We admit for the Fall cycle, with an option to start earlier in the Summer.
Our lab suits students who are curious about graduate school and open-ended research. Students who want to move into an industry engineering job are a poor fit.
Read our papers before you apply. An LLM summary of them is insufficient preparation.
Core threads:
Post-deployment adaptation and continual learning. Every deployed model meets situations it never saw during training. What algorithms let a model keep learning and adapt to those situations? Our active projects include test-time adaptation through environment interactions, eliciting and learning from automatic language feedback, and self-improvement through unsupervised bootstrapping.
Interfacing with complex data. Current ML relies on intensive data cleaning and hand-built data pipelines. How do we automate that work, so that a system dropped into a heterogeneous data environment deduces for itself which data to learn from? Our active projects include natural-language querying over heterogeneous data lakes, hybrid vector search, and adaptive indexing for vector databases.
Beyond the core threads, we are looking for graduate students and postdocs in two further areas:
AI for natural-science discovery (chemistry, materials). Building agents that read scientific literature, plan experiments, and reason over multi-modal scientific data. Active projects include unified multimodal models for chemistry. Suited to candidates with backgrounds in ML/NLP, scientific computing, or computational chemistry.
Systems and algorithms for scaling ML deployments. Infrastructure-level research to make large-scale agent training and inference tractable, including data and compute scheduling for long-horizon agentic learning. Suited to candidates with strong systems backgrounds who want to apply that lens to modern ML workloads.
Apply even if you lack traditional ML/NLP training and want to move into it.
As a PhD student at R2L, you will own a research project within one of our threads, from framing through publication. You will publish at top AI/ML/NLP venues (NeurIPS, ICML, ICLR, ACL, EMNLP) and contribute to the intellectual life of the lab. We admit 1-2 PhD students every year.
Required:
Strong background in computer science, mathematics, or a related field.
Demonstrated experience in machine learning and deep learning.
Proficiency in Python and ML frameworks (PyTorch, JAX).
Prior research experience, including publications at relevant venues.
Experience training and systematically running inference with large language models.
Foundations in theoretical machine learning, NLP on real-world text, and reinforcement learning in realistic environments.
Agents
We are hiring a postdoc to lead research within our agents threads: building and evaluating generalist agents, generalization and adaptation, language as supervision, or information access. You will define and lead ambitious projects, co-mentor PhD students, and help shape the lab's research direction. Apply on AcademicJobsOnline.
AI for Chemistry
We are also hiring a postdoc focused on agents for natural-science discovery, in collaboration with industrial and academic partners. This position suits candidates trained in either ML/NLP or computational chemistry who want to work across both. Apply on AcademicJobsOnline.
Required for either postdoc:
PhD in Computer Science or a related field.
Strong publication record at top AI/ML/NLP venues (or, for the chemistry posting, equivalent record in computational chemistry).
Demonstrated ability to conduct research independently.
Each semester we have 1–2 slots for undergraduate researchers. We prioritize students through the full-time URF program, followed by the part-time URA program. Students outside the University of Waterloo can apply through URF and through the Vector Internship program. At R2L, undergrads lead their own investigation: identify a compelling research problem, formulate a hypothesis, design experiments, and drive the project toward a meaningful outcome such as a publication at a top-tier conference. This is not a task-execution role.
Required:
Outstanding academic record in Computer Science.
Strong programming skills in Python and experience with a deep learning framework (PyTorch).
Genuine curiosity about AI and a proactive, self-motivated mindset.
Prior research experience is a plus, and a strong demonstration of initiative matters more.
Strong low-level systems experience (e.g. CUDA, operating systems, database systems) is a strong bonus.
Four R2L-led projects show the range of the work:
GTTA: Test-Time Adaptation for LLM Agents via Environment Interaction (ICLR 2026). Agents fail at deployment because of two distinct mismatches: a syntactic gap (unfamiliar observation formats) and a semantic gap (unknown state-transition dynamics). GTTA closes both with lightweight online adaptation and a persona-driven exploration phase that probes the environment before task execution.
AgentIR: Reasoning-Aware Retrieval for Deep Research Agents (COLM 2026). Deep research agents emit explicit natural-language reasoning before each search call, and existing retrievers ignore that signal. AgentIR embeds the reasoning trace together with the query. On BrowseComp-Plus, its 4B retriever reaches 68% accuracy with an open-weight agent, against 50% from conventional embeddings twice its size.
SynQuE: Synthetic Dataset Quality Estimation Without Annotations (TMLR 2026). Synthetic data is abundant, and its quality varies. SynQuE ranks synthetic datasets by their expected real-world task performance using only limited unannotated real data. On text-to-SQL parsing, training on the top-3 synthetic datasets that SynQuE selects raises accuracy from 30.4% to 38.4% on average.
ASH: Agents that Self-Hone via Embodied Learning (Preprint). Long-horizon embodied learning usually requires hand-engineered rewards or action-labeled demonstrations. ASH instead learns from unlabeled, noisy internet video through a self-improvement loop. When it gets stuck, it trains its own inverse dynamics model and uses that model to extract supervision from relevant video. On Pokemon Emerald and Zelda: Minish Cap, ASH sustains progression across 8-hour evaluations where VPT plateaus.
Questions after reading this page? Reach out to Victor Zhong.