BEGIN:VCALENDAR
VERSION:2.0
PRODID:-//CAIL//Events//EN
CALSCALE:GREGORIAN
METHOD:PUBLISH
X-WR-CALNAME:CAIL Events
BEGIN:VEVENT
UID:2026-03-27-eric-wong@cail.columbia.edu
DTSTAMP:20260823T060646Z
DTSTART:20260327T150000Z
DTEND:20260327T160000Z
SUMMARY:ML Seminar: Eric Wong - A Mechanistic Theory of Safety: How Jailbr
 eaking 1-Layer Transformers Taught us how to Steer LLMs
LOCATION:School of Social Work\, Room C03
DESCRIPTION:Why are LLM guardrails fundamentally so easily broken\, and ho
 w can we enforce them? This talk formalizes a mechanistic theory for study
 ing safety problems. We begin with one-layer transformers\, identifying ru
 le-breaking as an inherent architectural vulnerability in the model's atte
 ntion mechanism. This mechanistic theory framework (LogicBreaks) taught us
  a critical lesson: if attention is the key to breaking rules\, it may als
 o be the key to enforcing them.\n\nBuilding upon this insight\, we expand 
 the mechanistic theory to analyze attention-based interventions\, arriving
  at InstaBoost: an incredibly simple yet highly effective steering method 
 that boosts the model's attention on user-provided instructions during gen
 eration. This technique\, developed from analysis on one-layer transformer
 s\, provides state-of-the-art control over large-scale LLMs with just five
  lines of code.\n\nhttps://cail.columbia.edu/events/2026-03-27-eric-wong
URL:https://cail.columbia.edu/events/2026-03-27-eric-wong
END:VEVENT
END:VCALENDAR
