The ICINAi community

Building the science of interpretable AI, together.

Our steering group brings together researchers across institutions, disciplines, and sectors to help shape ICINAi’s scientific direction and shared infrastructure.

Steering group

Researchers helping guide the consortium’s priorities, programs, and standards.

David Bau

David Bau

Northeastern University

David Bau is an assistant professor at Northeastern University and director of the National Deep Inference Fabric. His research studies emergent internal mechanisms in large neural networks across language and vision, including causal tracing, model editing, and shared interpretability infrastructure.

Mor Geva

Mor Geva

Tel Aviv University

Mor Geva is an assistant professor in the School of Computer Science and AI at Tel Aviv University. Her research opens up the inner workings of large language models to improve their transparency, control, and reasoning.

Yonatan Belinkov

Yonatan Belinkov

Technion – Israel Institute of Technology

Yonatan Belinkov is an associate professor at the Technion. He develops probing, causal-analysis, and evaluation methods for understanding what language models represent and how their internal computations produce behavior.

Ellie Pavlick

Ellie Pavlick

Brown University · Google DeepMind

Ellie Pavlick is an associate professor at Brown University and a research scientist at Google DeepMind. She studies conceptual representations, language, reasoning, and learning, often by comparing the behavior of humans and language models.

Christopher Potts

Christopher Potts

Stanford University

Christopher Potts is a professor of linguistics at Stanford University, with a courtesy appointment in computer science. His research connects compositionality, causal abstraction, interpretability, information retrieval, and foundation-model programming.

Roger Grosse

Roger Grosse

University of Toronto · Vector Institute · Anthropic

Roger Grosse is a professor at the University of Toronto, a Vector Institute faculty member, and a researcher at Anthropic. His work spans mechanistic interpretability, alignment science, neural-network optimization, and representation learning.

Aaron Mueller

Aaron Mueller

Boston University

Aaron Mueller is an assistant professor of computer science at Boston University. His lab studies how concepts are learned and represented in language models, developing tools for circuit discovery, interpretability, and model control.

Naomi Saphra

Naomi Saphra

Boston University

Naomi Saphra is an assistant professor in Computing & Data Sciences at Boston University. Her research examines language-model training dynamics, how structure emerges during learning, and what those processes reveal about generalization and interpretability.

Tamar Rott Shaham

Tamar Rott Shaham

Weizmann Institute of Science

Tamar Rott Shaham is an assistant professor of computer science and applied mathematics at the Weizmann Institute. She develops automated tools to discover and explain model operations, then uses those insights to control model behavior.

Ana Marasović

Ana Marasović

University of Utah

Ana Marasović is an assistant professor in the Kahlert School of Computing at the University of Utah. Her work in natural language processing and human-centered AI develops more faithful explanations and evaluations of language models.

Ivan Titov

Ivan Titov

University of Edinburgh · University of Amsterdam

Ivan Titov is a professor at the University of Edinburgh and the University of Amsterdam. He studies trustworthy, robust, interpretable, and controllable language models and leads major European training and research programs in natural language processing.

Tal Linzen

Tal Linzen

New York University · Google

Tal Linzen is an associate professor at New York University and a staff research scientist at Google. His work connects language-model evaluation and interpretability with linguistics and cognitive science, especially the study of syntactic structure.

Byron Wallace

Byron Wallace

Northeastern University

Byron Wallace is a professor at Northeastern University and director of its BS in Artificial Intelligence program. His research develops machine-learning and natural-language-processing methods for health informatics, with particular attention to faithful explanations, biomedical evidence synthesis, and human–AI systems.

Yanai Elazar

Yanai Elazar

Bar-Ilan University

Yanai Elazar is an assistant professor of computer science and AI at Bar-Ilan University. He builds tools for a science of generative models, including methods for interpretability, data attribution, and rigorous evaluation.

Mario Giulianelli

Mario Giulianelli

University College London · Parallax

Mario Giulianelli is an associate professor of computational linguistics at University College London and a cofounder of Parallax. He combines behavioral evidence with analysis of internal mechanisms to make AI evaluation more rigorous and explanatory.

Gabriele Sarti

Gabriele Sarti

Parallax

Gabriele Sarti is a cofounder and technical director at Parallax. He develops white-box methods for context attribution, user modeling, deception detection, and model monitoring, alongside open-source libraries for language-model interpretability.

Eric J. Michaud

Eric J. Michaud

Simplex · Astera Institute

Eric J. Michaud is a research scientist at Simplex and the Astera Institute. He pursues empirical and conceptual accounts of learned computation, including work on sparse feature circuits, representation geometry, and mechanistic program synthesis.

Noah D. Goodman

Noah D. Goodman

Stanford University

Noah D. Goodman is a professor of psychology and computer science at Stanford University. His research uses probabilistic models to study language, cognition, and reasoning, including what language-model behavior reveals about internal representations and thought.

Zachary Lipton

Zachary Lipton

Carnegie Mellon University

Zachary Lipton is the Raj Reddy Associate Professor of Machine Learning at Carnegie Mellon University. His research examines the conceptual foundations of interpretability, robust and adaptive machine learning, natural language processing, and real-world decision systems.

Hinrich Schuetze

Hinrich Schuetze

LMU Munich

Hinrich Schuetze is professor and chair of computational linguistics at LMU Munich. His research spans statistical natural language processing, representation learning, lexical semantics, and the analysis of multilingual language models.

Amir Globerson

Amir Globerson

Tel Aviv University

Amir Globerson is a professor of computer science and AI at Tel Aviv University. He studies machine learning and natural language processing, including the theory of learned representations, attention, and interpretable reasoning.

Fazl Barez

Fazl Barez

University of Oxford

Fazl Barez is a senior researcher at the University of Oxford and leads the TSG Lab. His work connects interpretability, evaluation, model control, and machine unlearning with the technical foundations needed for effective AI governance.

Robert West

Robert West

EPFL

Robert West is an associate professor at EPFL, where he leads the Data Science and AI Lab. His work spans AI, natural language processing, and computational social science, including safe AI and structured human–AI collaboration.

Emre Yavuz

Emre Yavuz

Cambridge Boston Alignment Initiative

Emre Yavuz directs programs at the Cambridge Boston Alignment Initiative. He builds research communities and talent pathways for AI safety and biosecurity, connecting researchers through fellowships, training, and cross-institutional collaboration.