Catherine Chen
Postdoctoral Researcher

I am a postdoc at Leiden University working with Suzan Verberne in the Text Mining and Retreival Group. My research focuses on how explanations shape trust in AI systems, and is funded by the Hybrid Intelligence Centre.

I completed my PhD at Brown University, where I worked with Carsten Eickhoff in the Health NLP Lab. My PhD research centered on explainable information retrieval (XIR), particularly developing methods for interpretable and explainable search. I was also involved in projects relating to the broader interpretability of LLMs and the applications of ML/NLP/IR to the biomedical domain.

Before my PhD, I worked as a full-stack software engineer for FreeWheel in NYC to develop advertising technology solutions. I received my BA in Computer Science and French from Wellesley College, where I was also a member of the Track & Field and Swimming & Diving Teams.

In my spare time, I enjoy running, ice skating, cooking, and exploring the local food scene.

c.s.chen[at]liacs.leiden.nl

Education
  • Brown University
    Brown University
    • Ph.D. in Computer Science
      Sep. 2021 - July 2026
    • Sc.M. in Computer Science
      Sep 2021 - May 2023
  • Wellesley College
    Wellesley College
    • B.A. in Computer Science and French
      Sep 2015 - May 2019
Employment
  • Centre for the Governance of AI
    Centre for the Governance of AI
    Summer Fellow
    June - Aug 2025
  • Blue Cross Blue Shield of MN
    Blue Cross Blue Shield of MN
    Data Science PhD Intern
    June - Sept 2024
  • FreeWheel, a Comcast Company
    FreeWheel, a Comcast Company
    Software Engineer
    Aug 2019 - Aug 2021
News
2026
Moved to the Netherlands and started my postdoc at Leiden University!
Oct 01
Our paper, RankSteer, was accepted to Findings at EMNLP
Aug 21
Attended SIGIR in Melbourne, Australia to co-organize the second edition of WExIR
Jul 20
Successfully defended my PhD dissertation!
Jun 15
Attended ECIR in Delft, Netherlands and presented a tutorial on Mechanistic Interpretability in IR
Mar 25
2025
Had a full paper and a tutorial accepted to ECIR 2026
Dec 16
Succesfully proposed my PhD thesis!
Dec 03
Our paper, Pathway to Relevance: How Cross-Encoders Implement a Semantic Variant of BM25, was accepted to EMNLP 2025
Aug 20
Co-organized the 1st edition of WExIR at SIGIR in Padua, Italy
Jul 19
Started as a Summer Fellow at GovAI in London
Jun 08
Presented my work at the University of Amsterdam
May 13
Presented my work at Radboud University
May 08
Presented my work at Leiden University
Apr 30
Attended ECIR in Lucca, Italy and presented our demo paper, MechIR
Apr 06
Our short paper, Towards Best Practices of Axiomatic Activation Patching in Information Retrieval, was accepted at SIGIR 2025
Apr 04
Our workshop on Explainability in Information Retrieval (WExIR) was accepted at SIGIR 2025 (check us out!)
Feb 10
Started my 3.5 month visit in the IRLab at the University of Amsterdam!
Feb 01
2024
Our demo paper, MechIR: A Mechanistic Interpretability Framework for Information Retrieval, was accepted at ECIR 2025
Dec 16
Gave a talk at the Glasgow IR Seminar
Sep 30
Attended SIGIR in DC and presented my first (two) first-author paper(s)!
Jul 14
2/2 of our long papers were accepted to SIGIR 2024!
Mar 25
2023
Attended EMNLP in Singapore and presented our paper Outlier Dimensions Encode Task Specific Knowledge
Dec 06
Attended SIGIR in Taiwan and workshopped my dissertation in the Doctoral Consortium
Jul 23
Received my Master’s degree and advanced to candidacy
May 28
2021
Started my PhD at Brown University!
Sep 06
Selected Publications (view all )
RankSteer: Can Pointwise LLM Rankers Be Calibrated at the Representation Level?

Yumeng Wang, Catherine Chen, Suzan Verberne

EMNLP 2026

Large language models (LLMs) are strong zero-shot pointwise rankers, but lag behind pairwise and listwise methods. Beyond missing comparative signals, we identify a calibration gap: ranking-relevant information encoded in hidden states is not fully captured by the scalar output head. We propose RankSteer, a post-hoc activation-steering framework that calibrates ranking via projection-based interventions along multiple directions at inference time: decision, evidence, and, optionally, role. This is achieved without updating model weights or introducing cross-document comparisons. We instantiate RankSteer on two structurally distinct pointwise variants %4. What did we find and observe improvements over their respective baselines on most TREC DL and BEIR datasets across three backbones. This suggests that the calibration gap is a general property of pointwise rankers. Our additional geometric analysis shows that steering improves ranking by concentrating each query's document representations along an existing ranking geometry, offering new insight into how LLMs internally represent and calibrate relevance judgments.

RankSteer: Can Pointwise LLM Rankers Be Calibrated at the Representation Level?

Yumeng Wang, Catherine Chen, Suzan Verberne

EMNLP 2026

Large language models (LLMs) are strong zero-shot pointwise rankers, but lag behind pairwise and listwise methods. Beyond missing comparative signals, we identify a calibration gap: ranking-relevant information encoded in hidden states is not fully captured by the scalar output head. We propose RankSteer, a post-hoc activation-steering framework that calibrates ranking via projection-based interventions along multiple directions at inference time: decision, evidence, and, optionally, role. This is achieved without updating model weights or introducing cross-document comparisons. We instantiate RankSteer on two structurally distinct pointwise variants %4. What did we find and observe improvements over their respective baselines on most TREC DL and BEIR datasets across three backbones. This suggests that the calibration gap is a general property of pointwise rankers. Our additional geometric analysis shows that steering improves ranking by concentrating each query's document representations along an existing ranking geometry, offering new insight into how LLMs internally represent and calibrate relevance judgments.

How Role-Play Shapes Relevance Judgment in Zero-Shot LLM Rankers

Yumeng Wang, Jirui Qi, Catherine Chen, Panagiotis Eustratiadis, Suzan Verberne

ECIR 2026

Large Language Models (LLMs) have emerged as promising zero-shot rankers, but their performance is highly sensitive to prompt formulation. In particular, role-play prompts, where the model is assigned a functional role or identity, often give more robust and accurate relevance rankings. However, the mechanisms and diversity of role-play effects remain underexplored, limiting both effective use and interpretability. In this work, we systematically examine how role-play variations influence zero-shot LLM rankers. We employ causal intervention techniques from mechanistic interpretability to trace how role-play information shapes relevance judgments in LLMs. Our analysis reveals that (1) careful formulation of role descriptions have a large effect on the ranking quality of the LLM; (2) role-play signals are predominantly encoded in early layers and communicate with task instructions in middle layers, while receiving limited interaction with query or document representations. Specifically, we identify a group of attention heads that encode information critical for role-conditioned relevance. These findings not only shed light on the inner workings of role-play in LLM ranking but also offer guidance for designing more effective prompts in IR and beyond, pointing toward broader opportunities for leveraging role-play in zero-shot applications.

How Role-Play Shapes Relevance Judgment in Zero-Shot LLM Rankers

Yumeng Wang, Jirui Qi, Catherine Chen, Panagiotis Eustratiadis, Suzan Verberne

ECIR 2026

Large Language Models (LLMs) have emerged as promising zero-shot rankers, but their performance is highly sensitive to prompt formulation. In particular, role-play prompts, where the model is assigned a functional role or identity, often give more robust and accurate relevance rankings. However, the mechanisms and diversity of role-play effects remain underexplored, limiting both effective use and interpretability. In this work, we systematically examine how role-play variations influence zero-shot LLM rankers. We employ causal intervention techniques from mechanistic interpretability to trace how role-play information shapes relevance judgments in LLMs. Our analysis reveals that (1) careful formulation of role descriptions have a large effect on the ranking quality of the LLM; (2) role-play signals are predominantly encoded in early layers and communicate with task instructions in middle layers, while receiving limited interaction with query or document representations. Specifically, we identify a group of attention heads that encode information critical for role-conditioned relevance. These findings not only shed light on the inner workings of role-play in LLM ranking but also offer guidance for designing more effective prompts in IR and beyond, pointing toward broader opportunities for leveraging role-play in zero-shot applications.

Pathway to Relevance: How Cross-Encoders Implement a Semantic Variant of BM25

Meng Lu, Catherine Chen, Carsten Eickhoff

EMNLP 2025

Mechanistic interpretation has greatly contributed to a more detailed understanding of generative language models, enabling significant progress in identifying structures that implement key behaviors through interactions between internal components. In contrast, interpretability in information retrieval (IR) remains relatively coarse-grained, and much is still unknown as to how IR models determine whether a document is relevant to a query. In this work, we address this gap by mechanistically analyzing how one commonly used model, a cross-encoder, estimates relevance. We find that the model extracts traditional relevance signals, such as term frequency and inverse document frequency, in early-to-middle layers. These concepts are then combined in later layers, similar to the well-known probabilistic ranking function, BM25. Overall, our analysis offers a more nuanced understanding of how IR models compute relevance. Isolating these components lays the groundwork for future interventions that could enhance transparency, mitigate safety risks, and improve scalability.

Pathway to Relevance: How Cross-Encoders Implement a Semantic Variant of BM25

Meng Lu, Catherine Chen, Carsten Eickhoff

EMNLP 2025

Mechanistic interpretation has greatly contributed to a more detailed understanding of generative language models, enabling significant progress in identifying structures that implement key behaviors through interactions between internal components. In contrast, interpretability in information retrieval (IR) remains relatively coarse-grained, and much is still unknown as to how IR models determine whether a document is relevant to a query. In this work, we address this gap by mechanistically analyzing how one commonly used model, a cross-encoder, estimates relevance. We find that the model extracts traditional relevance signals, such as term frequency and inverse document frequency, in early-to-middle layers. These concepts are then combined in later layers, similar to the well-known probabilistic ranking function, BM25. Overall, our analysis offers a more nuanced understanding of how IR models compute relevance. Isolating these components lays the groundwork for future interventions that could enhance transparency, mitigate safety risks, and improve scalability.

MechIR: A Mechanistic Interpretability Framework for Information Retrieval

Andrew Parry, Catherine Chen, Carsten Eickhoff, Sean MacAvaney

ECIR 2025

Mechanistic interpretability is an emerging diagnostic approach for neural models that has gained traction in broader natural language processing domains. This paradigm aims to provide attribution to components of neural systems where causal relationships between hidden layers and output were previously uninterpretable. As the use of neural models in IR for retrieval and evaluation becomes ubiquitous, we need to ensure that we can interpret why a model produces a given output for both transparency and the betterment of systems. This work comprises a flexible framework for diagnostic analysis and intervention within these highly parametric neural systems specifically tailored for IR tasks and architectures. In providing such a framework, we look to facilitate further research in interpretable IR with a broader scope for practical interventions derived from mechanistic interpretability. We provide preliminary analysis and look to demonstrate our framework through an axiomatic lens to show its applications and ease of use for those IR practitioners inexperienced in this emerging paradigm.

MechIR: A Mechanistic Interpretability Framework for Information Retrieval

Andrew Parry, Catherine Chen, Carsten Eickhoff, Sean MacAvaney

ECIR 2025

Mechanistic interpretability is an emerging diagnostic approach for neural models that has gained traction in broader natural language processing domains. This paradigm aims to provide attribution to components of neural systems where causal relationships between hidden layers and output were previously uninterpretable. As the use of neural models in IR for retrieval and evaluation becomes ubiquitous, we need to ensure that we can interpret why a model produces a given output for both transparency and the betterment of systems. This work comprises a flexible framework for diagnostic analysis and intervention within these highly parametric neural systems specifically tailored for IR tasks and architectures. In providing such a framework, we look to facilitate further research in interpretable IR with a broader scope for practical interventions derived from mechanistic interpretability. We provide preliminary analysis and look to demonstrate our framework through an axiomatic lens to show its applications and ease of use for those IR practitioners inexperienced in this emerging paradigm.

Axiomatic Causal Interventions for Reverse Engineering Relevance Computation in Neural Retrieval Models

Catherine Chen, Jack Merullo, Carsten Eickhoff

SIGIR 2024

Neural models have demonstrated remarkable performance across diverse ranking tasks. However, the processes and internal mechanisms along which they determine relevance are still largely unknown. Existing approaches for analyzing neural ranker behavior with respect to IR properties rely either on assessing overall model behavior or employing probing methods that may offer an incomplete understanding of causal mechanisms. To provide a more granular understanding of internal model decision-making processes, we propose the use of causal interventions to reverse engineer neural rankers, and demonstrate how mechanistic interpretability methods can be used to isolate components satisfying term-frequency axioms within a ranking model. We identify a group of attention heads that detect duplicate tokens in earlier layers of the model, then communicate with downstream heads to compute overall document relevance. More generally, we propose that this style of mechanistic analysis opens up avenues for reverse engineering the processes neural retrieval models use to compute relevance. This work aims to initiate granular interpretability efforts that will not only benefit retrieval model development and training, but ultimately ensure safer deployment of these models.

Axiomatic Causal Interventions for Reverse Engineering Relevance Computation in Neural Retrieval Models

Catherine Chen, Jack Merullo, Carsten Eickhoff

SIGIR 2024

Neural models have demonstrated remarkable performance across diverse ranking tasks. However, the processes and internal mechanisms along which they determine relevance are still largely unknown. Existing approaches for analyzing neural ranker behavior with respect to IR properties rely either on assessing overall model behavior or employing probing methods that may offer an incomplete understanding of causal mechanisms. To provide a more granular understanding of internal model decision-making processes, we propose the use of causal interventions to reverse engineer neural rankers, and demonstrate how mechanistic interpretability methods can be used to isolate components satisfying term-frequency axioms within a ranking model. We identify a group of attention heads that detect duplicate tokens in earlier layers of the model, then communicate with downstream heads to compute overall document relevance. More generally, we propose that this style of mechanistic analysis opens up avenues for reverse engineering the processes neural retrieval models use to compute relevance. This work aims to initiate granular interpretability efforts that will not only benefit retrieval model development and training, but ultimately ensure safer deployment of these models.

Evaluating Search System Explainability with Psychometrics and Crowdsourcing

Catherine Chen, Carsten Eickhoff

SIGIR 2024

Information retrieval (IR) systems have become an integral part of our everyday lives. As search engines, recommender systems, and conversational agents are employed across various domains from recreational search to clinical decision support, there is an increasing need for transparent and explainable systems to guarantee accountable, fair, and unbiased results. Despite many recent advances towards explainable AI and IR techniques, there is no consensus on what it means for a system to be explainable. Although a growing body of literature suggests that explainability is comprised of multiple subfactors, virtually all existing approaches treat it as a singular notion. In this paper, we examine explainability in Web search systems, leveraging psychometrics and crowdsourcing to identify human-centered factors of explainability.

Evaluating Search System Explainability with Psychometrics and Crowdsourcing

Catherine Chen, Carsten Eickhoff

SIGIR 2024

Information retrieval (IR) systems have become an integral part of our everyday lives. As search engines, recommender systems, and conversational agents are employed across various domains from recreational search to clinical decision support, there is an increasing need for transparent and explainable systems to guarantee accountable, fair, and unbiased results. Despite many recent advances towards explainable AI and IR techniques, there is no consensus on what it means for a system to be explainable. Although a growing body of literature suggests that explainability is comprised of multiple subfactors, virtually all existing approaches treat it as a singular notion. In this paper, we examine explainability in Web search systems, leveraging psychometrics and crowdsourcing to identify human-centered factors of explainability.

Outlier Dimensions Encode Task Specific Knowledge
Outlier Dimensions Encode Task Specific Knowledge

William Rudman, Catherine Chen, Carsten Eickhoff

EMNLP 2023

Representations from large language models (LLMs) are known to be dominated by a small subset of dimensions with exceedingly high variance. Previous works have argued that although ablating these outlier dimensions in LLM representations hurts downstream performance, outlier dimensions are detrimental to the representational quality of embeddings. In this study, we investigate how fine-tuning impacts outlier dimensions and show that 1) outlier dimensions that occur in pre-training persist in fine-tuned models and 2) a single outlier dimension can complete downstream tasks with a minimal error rate. Our results suggest that outlier dimensions can encode crucial task-specific knowledge and that the value of a representation in a single outlier dimension drives downstream model decisions.

Outlier Dimensions Encode Task Specific Knowledge
Outlier Dimensions Encode Task Specific Knowledge

William Rudman, Catherine Chen, Carsten Eickhoff

EMNLP 2023

Representations from large language models (LLMs) are known to be dominated by a small subset of dimensions with exceedingly high variance. Previous works have argued that although ablating these outlier dimensions in LLM representations hurts downstream performance, outlier dimensions are detrimental to the representational quality of embeddings. In this study, we investigate how fine-tuning impacts outlier dimensions and show that 1) outlier dimensions that occur in pre-training persist in fine-tuned models and 2) a single outlier dimension can complete downstream tasks with a minimal error rate. Our results suggest that outlier dimensions can encode crucial task-specific knowledge and that the value of a representation in a single outlier dimension drives downstream model decisions.

All publications