Research

Human-In-The-Loop AI Evaluation

Recently, my research has focused on how to integrate human feedback into training and evaluating large language models, with a focus on improving performance and ensuring alignment with user preferences in real-world contexts. This work demonstrates the limitations of automated metrics in capturing nuance, especially in sensitive tasks like summarization or toxicity detection, and proposes human-centered methodologies that prioritize meaningful evaluation and data curation.

A few recent papers and defensive publications can be found here:

Pragmatics & Prosody

I also work on exploring the connections between formal semantics, pragmatics, and prosody, with a focus on how intonation and discourse particles influence speech acts in empirically grounded ways. A central theme is understanding how these elements, especially in expressions of surprise, function to modify the speech act itself across different languages.

You can find some of this work here: