SU
S. Udagawa
info
Please Note
<p>This page displays the records of the person named above and is not linked to a unique person identifier. This record may need to be merged to a profile.</p>
1 records found
1
This study examines how effectively widely used offline information retrieval (IR) metrics reflect changes in online performance. As offline evaluation plays a central role in model development, understanding its alignment with user‑oriented signals is essential. Using 52 diverse ranking pipelines and approximately 2,000 queries from the MS MARCO DL19 and DL20 benchmarks, we analyze the sensitivity of five offline metrics: Precision@10, Recall@10, MAP, MRR, and NDCG@10, to five simulated online metrics: CTR, SSR, ZRR, ADT, and SAR. Sensitivity is quantified through slope-based analysis, and alignment is assessed using the Pearson correlation coefficient. Our results show that NDCG@10 and Recall@10 are the most sensitive offline metrics across multiple online behaviors, while Precision@10 consistently exhibits low sensitivity. Furthermore, we demonstrate that sensitivity and alignment capture complementary aspects of offline–online relationships: some metric pairs show strong responsiveness but weak linear consistency. Overall, this study provides a detailed and reproducible evaluation of how offline metrics behave in relation to simulated online performance, offering practical guidance for selecting offline metrics that better reflect user-centric outcomes.
https://github.com/AinzOoalGown123/Metric-Sensitivity-Analysis ...
https://github.com/AinzOoalGown123/Metric-Sensitivity-Analysis ...
This study examines how effectively widely used offline information retrieval (IR) metrics reflect changes in online performance. As offline evaluation plays a central role in model development, understanding its alignment with user‑oriented signals is essential. Using 52 diverse ranking pipelines and approximately 2,000 queries from the MS MARCO DL19 and DL20 benchmarks, we analyze the sensitivity of five offline metrics: Precision@10, Recall@10, MAP, MRR, and NDCG@10, to five simulated online metrics: CTR, SSR, ZRR, ADT, and SAR. Sensitivity is quantified through slope-based analysis, and alignment is assessed using the Pearson correlation coefficient. Our results show that NDCG@10 and Recall@10 are the most sensitive offline metrics across multiple online behaviors, while Precision@10 consistently exhibits low sensitivity. Furthermore, we demonstrate that sensitivity and alignment capture complementary aspects of offline–online relationships: some metric pairs show strong responsiveness but weak linear consistency. Overall, this study provides a detailed and reproducible evaluation of how offline metrics behave in relation to simulated online performance, offering practical guidance for selecting offline metrics that better reflect user-centric outcomes.
https://github.com/AinzOoalGown123/Metric-Sensitivity-Analysis
https://github.com/AinzOoalGown123/Metric-Sensitivity-Analysis