상세 보기
Exploiting concept clusters for content-based information retrieval
- Kang, B.Y.;
- Kim, Dae-Won;
- Lee, S.J.
WEB OF SCIENCE
29SCOPUS
36초록
Current approaches to index weighting for information retrieval from texts are based on statistical analysis of the texts' contents. A key shortcoming of these indexing schemes, which consider only the terms in a document, is that they cannot extract semantically exact indexes that represent the semantic content of a document. To address this issue, we proposed a new indexing formalism that considers not only the terms in a document, but also the concepts. In the proposed method, concepts are extracted by exploiting clusters of terms that are semantically related, referred to as concept clusters. Through experiments on the TREC-2 collection of Wall Street Journal documents, we show that the proposed method outperforms an indexing method based on term frequency (TF), especially in regard to the highest-ranked documents. Moreover, the index term dimension was 53.3% lower for the proposed method than for the TF-based method, which is expected to significantly reduce the document search time in a real environment. (C) 2004 Elsevier Inc. All rights reserved.
키워드
- 제목
- Exploiting concept clusters for content-based information retrieval
- 저자
- Kang, B.Y.; Kim, Dae-Won; Lee, S.J.
- 발행일
- 2005-02
- 유형
- Article
- 권
- 170
- 호
- 2-4
- 페이지
- 443 ~ 462
- 언어
- ENG
- 출판사
- ELSEVIER SCIENCE INC
- 발행국가
- 미국
- 분량
- 20 페이지
- ISSN
- E 1872-6291
P 0020-0255