All Events

Sparks of New Pre-Training

Yoav Artzi · Cornell University

Friday, September 18, 2026 · 11:00 AM

SSW Room 903

This talk covers new pre-training techniques. First, we introduce the hypothesis that state and prediction representations, which are entangled in transformers, are better separated. We design a simple architectural modification that effectively separates them, and provides 2.6x token efficiency during pre-training. The second technique is focused on externalizing knowledge by pre-training an LLM to rely on an external knowledge base, while inducing this KB from the pre-training data. We re-model the Limited Memory Language Model (LMLM) paradigm we introduced in prior work, with a new expressive continuous query mechanism. This dramatically increases the expressivity of the LMLM paradigm, and allows scaling to general web text. Our Co-LMLM model gives lower perplexity than a model trained on 40x the amount of tokens, and provides SimpleQA performance on par with Claude Sonnet 4.5. All that while presenting all the advantages of the LMLM class -- knowledge control, provenance, editing, and factuality. Together, these techniques demonstrate two of many possible avenues to bring about fundamental change in LLMs through new pre-training paradigms.

About the speaker

Yoav Artzi is an Associate Professor in the Department of Computer Science and Cornell Tech at Cornell University, and a visiting faculty research at Google DeepMind. His research focuses on language modeling and learning in interactive and situated scenarios. His work was acknowledged by awards and honorable mentions at ACL, EMNLP, NAACL, and IROS, as well as a TACL test-of-time award. Yoav holds a B.Sc. from Tel Aviv University and a Ph.D. from the University of Washington. He was previously arXiv's associate faculty director, and co-founded COLM.