Join Us: Openings at the Reading to Learn (R2L) Lab

The R2L Lab at the University of Waterloo's Cheriton Department of Computer Science builds intelligent generalist agents that read, reason, and act across digital and physical environments. We organize our work around four long-running research threads: building and evaluating generalist agents, generalization, adaptation, and reasoning, language as supervision, and information access for autonomous reasoners. We are recruiting students and postdocs across all four.

Please read this before applying.

  • We rarely hire MS students or remote undergraduate students.

  • The University of Waterloo admits students each semester. We admit for the Fall cycle, with an option to start earlier in the Summer.

  • Our lab suits students who are curious about graduate school and open-ended research. Students who want to move into an industry engineering job are a poor fit.

  • Read our papers before you apply. An LLM summary of them is insufficient preparation.


Areas We Are Recruiting In

Core threads:

Beyond the core threads, we are looking for graduate students and postdocs in two further areas:

Apply even if you lack traditional ML/NLP training and want to move into it.


How to Apply

  1. Review our work. Read our recent publications and identify the thread your interests align with.
  2. Complete our Lab Application. Submit this form so we can track your application. We track applications through the form, and an email alone will likely be missed. Undergraduate students should also attach a transcript.
  3. For prospective PhD students: also submit an official application for PhD to Waterloo's Computer Science department, and name Victor Zhong as your prospective supervisor. The deadline is December 1st in most years.
  4. For postdocs: apply through the AcademicJobsOnline link on the relevant posting below.

Open Positions

PhD Student

As a PhD student at R2L, you will own a research project within one of our threads, from framing through publication. You will publish at top AI/ML/NLP venues (NeurIPS, ICML, ICLR, ACL, EMNLP) and contribute to the intellectual life of the lab. We admit 1-2 PhD students every year.

Required:

Postdoctoral Fellow

Agents

We are hiring a postdoc to lead research within our agents threads: building and evaluating generalist agents, generalization and adaptation, language as supervision, or information access. You will define and lead ambitious projects, co-mentor PhD students, and help shape the lab's research direction. Apply on AcademicJobsOnline.

AI for Chemistry

We are also hiring a postdoc focused on agents for natural-science discovery, in collaboration with industrial and academic partners. This position suits candidates trained in either ML/NLP or computational chemistry who want to work across both. Apply on AcademicJobsOnline.

Required for either postdoc:

Undergraduate Research Assistants

Each semester we have 1–2 slots for undergraduate researchers. We prioritize students through the full-time URF program, followed by the part-time URA program. Students outside the University of Waterloo can apply through URF and through the Vector Internship program. At R2L, undergrads lead their own investigation: identify a compelling research problem, formulate a hypothesis, design experiments, and drive the project toward a meaningful outcome such as a publication at a top-tier conference. This is not a task-execution role.

Required:


Example Research Projects

Four R2L-led projects show the range of the work:

GTTA: Test-Time Adaptation for LLM Agents via Environment Interaction (ICLR 2026). Agents fail at deployment because of two distinct mismatches: a syntactic gap (unfamiliar observation formats) and a semantic gap (unknown state-transition dynamics). GTTA closes both with lightweight online adaptation and a persona-driven exploration phase that probes the environment before task execution.

AgentIR: Reasoning-Aware Retrieval for Deep Research Agents (COLM 2026). Deep research agents emit explicit natural-language reasoning before each search call, and existing retrievers ignore that signal. AgentIR embeds the reasoning trace together with the query. On BrowseComp-Plus, its 4B retriever reaches 68% accuracy with an open-weight agent, against 50% from conventional embeddings twice its size.

SynQuE: Synthetic Dataset Quality Estimation Without Annotations (TMLR 2026). Synthetic data is abundant, and its quality varies. SynQuE ranks synthetic datasets by their expected real-world task performance using only limited unannotated real data. On text-to-SQL parsing, training on the top-3 synthetic datasets that SynQuE selects raises accuracy from 30.4% to 38.4% on average.

ASH: Agents that Self-Hone via Embodied Learning (Preprint). Long-horizon embodied learning usually requires hand-engineered rewards or action-labeled demonstrations. ASH instead learns from unlabeled, noisy internet video through a self-improvement loop. When it gets stuck, it trains its own inverse dynamics model and uses that model to extract supervision from relevant video. On Pokemon Emerald and Zelda: Minish Cap, ASH sustains progression across 8-hour evaluations where VPT plateaus.


Contact

Questions after reading this page? Reach out to Victor Zhong.