Sharon Levy
Rutgers University
Beyond Surface-Level Safety: Uncovering Implicit Biases
Sharon Levy · Rutgers University
Current safety alignment techniques effectively mitigate explicit biases but can fail on more subtle implicit biases. First, I will discuss our methodology to discover implicit biases using logic puzzles. This provides automatic generation and evaluation and can be easily tailored to various demographic attributes. Next, I will highlight the issue of selective refusal bias, where models may refuse to generate harmful content targeting some demographic groups and not others. Together, this research highlights areas of improvement for model alignment.
About the speaker
Sharon is an Assistant Professor in Computer Science at Rutgers University. Her research focuses on natural language processing, with an emphasis on Responsible AI. She works on problems relating to fairness, trustworthiness, and safety. Previously, Sharon was a postdoctoral fellow at the Center for Language and Speech Processing at Johns Hopkins University and obtained her Ph.D. from the University of California Santa Barbara.