Model-based Feedback in the Language Modeling Approach to Information Retrieval

Abstract

The language modeling approach to retrieval has been shown to perform well empirically. One advantage of this new approach is its statistical foundations. However, feedback, as one important component in a retrieval system, has only been dealt with heuristically in this new retrieval approach: the original query is usually literally expanded by adding additional terms to it. Such <i>expansion-based</i> feedback creates an inconsistent interpretation of the original and the expanded query. In this paper, we present a more principled approach to feedback in the language modeling approach. Specifically, we treat feedback as updating the query language model based on the extra evidence carried by the feedback documents. Such a <i>model-based</i> feedback strategy easily fits into an extension of the language modeling approach. We propose and evaluate two different approaches to updating a query language model based on feedback documents, one based on a generative probabilistic model of feedback documents and one based on minimization of the KL-divergence over feedback documents. Experiment results show that both approaches are effective and outperform the Rocchio feedback approach.

DOI: 10.1145/502585.502654

Extracted Key Phrases

Unfortunately, ACM prohibits us from displaying non-influential references for this paper.

To see the full reference list, please visit http://dl.acm.org/citation.cfm?id=502654.

Showing 1-10 of 430 extracted citations
050100'02'04'06'08'10'12'14'16
Citations per Year

786 Citations

Semantic Scholar estimates that this publication has received between 662 and 933 citations based on the available data.

See our FAQ for additional information.