Exploiting concept clusters for content-based information retrieval

Citations

WEB OF SCIENCE

29
Citations

SCOPUS

36

초록

Current approaches to index weighting for information retrieval from texts are based on statistical analysis of the texts' contents. A key shortcoming of these indexing schemes, which consider only the terms in a document, is that they cannot extract semantically exact indexes that represent the semantic content of a document. To address this issue, we proposed a new indexing formalism that considers not only the terms in a document, but also the concepts. In the proposed method, concepts are extracted by exploiting clusters of terms that are semantically related, referred to as concept clusters. Through experiments on the TREC-2 collection of Wall Street Journal documents, we show that the proposed method outperforms an indexing method based on term frequency (TF), especially in regard to the highest-ranked documents. Moreover, the index term dimension was 53.3% lower for the proposed method than for the TF-based method, which is expected to significantly reduce the document search time in a real environment. (C) 2004 Elsevier Inc. All rights reserved.

키워드

information retrieval; indexing; term frequency; weighting function
제목
Exploiting concept clusters for content-based information retrieval
저자
Kang, B.Y.; Kim, Dae-Won; Lee, S.J.
DOI
10.1016/j.ins.2004.03.013
발행일
2005-02
유형
Article
저널명
Information Sciences
권
170
호
2-4
페이지
443 ~ 462