Spelling suggestions: "subject:"eeb document indexing"" "subject:"beb document indexing""
1 |
WebDoc an Automated Web Document Indexing SystemTang, Bo 13 December 2002 (has links)
This thesis describes WebDoc, an automated system that classifies Web documents according to the Library of Congress classification system. This work is an extension of an early version of the system that successfully generated indexes for journal articles. The unique features of Web documents, as well as how they will affect the design of a classification system, are discussed. We argue that full-text analysis of Web documents is inevitable, and contextual information must be used to assist the classification. The architecture of the WebDoc system is presented. We performed experiments on it with and without the assistance of contextual information. The results show that contextual information improved the system?s performance significantly.
|
Page generated in 0.0777 seconds