Recently, my research has focused on how to integrate human feedback into training and evaluating large language models, with a focus on improving performance and ensuring alignment with user preferences in real-world contexts. This work demonstrates the limitations of automated metrics in capturing nuance, especially in sensitive tasks like summarization or toxicity detection, and proposes human-centered methodologies that prioritize meaningful evaluation and data curation.
A few recent papers and defensive publications can be found here:
I also work on exploring the connections between formal semantics, pragmatics, and prosody, with a focus on how intonation and discourse particles influence speech acts in empirically grounded ways. A central theme is understanding how these elements, especially in expressions of surprise, function to modify the speech act itself across different languages.
You can find some of this work here: