Sankaran Vaidyanathan
shun-ka-run • /ʃʌŋkəˈɹʌn/ • சங்கரன்
I am a PhD candidate at the College of Information and Computer Sciences, UMass Amherst, where I am advised by David Jensen. My research focuses on developing principled tools grounded in causal reasoning for explaining and evaluating complex AI systems, including large language models (LLMs) and reinforcement learning agents.
In particular, I focus on problems where subjective human judgments play a central role, including mechanistic interpretability in neural networks, evaluating LLM outputs, blame and responsibility attribution, and alignment with human social norms. These domains are often difficult to model using conventional statistical approaches in causality and machine learning: human judgments are shaped by implicit expectations, context-sensitive reasoning, and the tendency to highlight some causes over others based on agreed-upon social norms.
By developing methods grounded in scientific rigor and the human values that guide real-world decision-making, I aim to enable reliable evaluation and responsible governance of AI systems.
news
| Sep 26, 2026 | The Curse of Multiple Mediators: Hidden Interaction Effects in Activation Patching was accepted as a spotlight paper at NeurIPS 2026. See you in Sydney! |
|---|---|
| Jun 25, 2026 | Released Interphyre, a 2D physics puzzle environment with level editing, full simulator state access, and intervention capabilities to support causal analysis and agentic experimentation. |
| Jun 01, 2026 | Started an internship at Basis Research Institute working with Rafal Urbaniak, Emily Bunnapradist, and Michelangelo Naim on probabilistic actual causality and its applications to LLM interpretability. |
| Apr 10, 2026 | Guest lecture on mechanistic interpretability for COMPSCI 690S: AI Alignment. Slides here. |
| Dec 07, 2025 | Presented our work on Detecting and Characterizing Planning in Language Models at the NeurIPS 2025 Mechanistic Interpretability Workshop. |