Global ETD Search

1	A Domain Based Approach to Crawl the Hidden Web Pandya, Milan 04 December 2006 (has links) There is a lot of research work being performed on indexing the Web. More and more sophisticated Web crawlers are been designed to search and index the Web faster. But all these traditional crawlers crawl only the part of Web we call “Surface Web”. They are unable to crawl the hidden portion of the Web. These traditional crawlers retrieve contents only from surface Web pages which are just a set of Web pages linked by some hyperlinks and ignoring the hidden information. Hence, they ignore tremendous amount of information hidden behind these search forms in Web pages. Most of the published research has been done to detect such searchable forms and make a systematic search over these forms. Our approach here will be based on a Web crawler that analyzes search forms and fills tem with appropriate content to retrieve maximum relevant information from the database. web crawler search spider web bot best first crawler focused web crawler web page domain based Computer Sciences
2	A Distributed Approach to Crawl Domain Specific Hidden Web Desai, Lovekeshkumar 03 August 2007 (has links) A large amount of on-line information resides on the invisible web - web pages generated dynamically from databases and other data sources hidden from current crawlers which retrieve content only from the publicly indexable Web. Specially, they ignore the tremendous amount of high quality content "hidden" behind search forms, and pages that require authorization or prior registration in large searchable electronic databases. To extracting data from the hidden web, it is necessary to find the search forms and fill them with appropriate information to retrieve maximum relevant information. To fulfill the complex challenges that arise when attempting to search hidden web i.e. lots of analysis of search forms as well as retrieved information also, it becomes eminent to design and implement a distributed web crawler that runs on a network of workstations to extract data from hidden web. We describe the software architecture of the distributed and scalable system and also present a number of novel techniques that went into its design and implementation to extract maximum relevant data from hidden web for achieving high performance. Deep Web Breadth-first crawler Search spider Distributed Web crawler task-specific and Domain Specific Hidden Web Content Extraction Computer Sciences

Search results

A Domain Based Approach to Crawl the Hidden Web

A Distributed Approach to Crawl Domain Specific Hidden Web