Global ETD Search

1	adXtractor – Automated and Adaptive Generation of Wrappers for Information Retrieval Ademi, Muhamet January 2017 (has links) The aim of this project is to investigate the feasibility of retrieving unstructured automotive listings from structured web pages on the Internet. The research has two major purposes: (1) to investigate whether it is feasible to pair information extraction algorithms and compute wrappers (2) demonstrate the results of pairing these techniques and evaluate the measurements. We merge two training sets available on the web to construct reference sets which is the basis for the information extraction. The wrappers are computed by using information extraction techniques to identify data properties with a variety of techniques such as fuzzy string matching, regular expressions and document tree analysis. The results demonstrate that it is possible to pair these techniques successfully and retrieve the majority of the listings. Additionally, the findings also suggest that many platforms utilise lazy loading to populate image resources which the algorithm is unable to capture. In conclusion, the study demonstrated that it is possible to use information extraction to compute wrappers dynamically by identifying data properties. Furthermore, the study demonstrates the ability to open non-queryable domain data through a unified service. wrapper generation information extraction content of interest identification wrapper rules text extraction key value pair wrapper generate main content identification web scraping information extraction algorithms web extraction dom tree analysis dom analysis Engineering and Technology Teknik och teknologier

Search results

adXtractor – Automated and Adaptive Generation of Wrappers for Information Retrieval