BEGIN:VCALENDAR
VERSION:2.0
PRODID:-//CAIL//Events//EN
CALSCALE:GREGORIAN
METHOD:PUBLISH
X-WR-CALNAME:CAIL Events
BEGIN:VEVENT
UID:2026-09-18-yoav-artzi@cail.columbia.edu
DTSTAMP:20260921T212555Z
DTSTART:20260918T150000Z
DTEND:20260918T160000Z
SUMMARY:ML Seminar: Yoav Artzi - Sparks of New Pre-Training
LOCATION:SSW Room 903
DESCRIPTION:This talk covers new pre-training techniques. First\, we intro
 duce the hypothesis that state and prediction representations\, which are 
 entangled in transformers\, are better separated. We design a simple archi
 tectural modification that effectively separates them\, and provides 2.6x 
 token efficiency during pre-training. The second technique is focused on e
 xternalizing knowledge by pre-training an LLM to rely on an external knowl
 edge base\, while inducing this KB from the pre-training data. We re-model
  the Limited Memory Language Model (LMLM) paradigm we introduced in prior 
 work\, with a new expressive continuous query mechanism. This dramatically
  increases the expressivity of the LMLM paradigm\, and allows scaling to g
 eneral web text. Our Co-LMLM model gives lower perplexity than a model tra
 ined on 40x the amount of tokens\, and provides SimpleQA performance on pa
 r with Claude Sonnet 4.5. All that while presenting all the advantages of 
 the LMLM class -- knowledge control\, provenance\, editing\, and factualit
 y. Together\, these techniques demonstrate two of many possible avenues to
  bring about fundamental change in LLMs through new pre-training paradigms
 .\n\nhttps://cail.columbia.edu/events/2026-09-18-yoav-artzi
URL:https://cail.columbia.edu/events/2026-09-18-yoav-artzi
END:VEVENT
END:VCALENDAR
