Global ETD Search

Return to search

Identifying Search Engine Spam Using DNS

Web crawlers encounter both finite and infinite elements during crawl. Pages and hosts can be infinitely generated using automated scripts and DNS wildcard entries. It is a challenge to rank such resources as an entire web of pages and hosts could be created to manipulate the rank of a target resource. It is crucial to be able to differentiate genuine content from spam in real-time to allocate crawl budgets. In this study, ranking algorithms to rank hosts are designed which use the finite Pay Level Domains(PLD) and IPv4 addresses. Heterogenous graphs derived from the webgraph of IRLbot are used to achieve this. PLD Supporters (PSUPP) which is the number of level-2 PLD supporters for each host on the host-host-PLD graph is the first algorithm that is studied. This is further improved by True PLD Supporters(TSUPP) which uses true egalitarian level-2 PLD supporters on the host-IP-PLD graph and DNS blacklists. It was found that support from content farms and stolen links could be eliminated by finding TSUPP. When TSUPP was applied on the host graph of IRLbot, there was less than 1% spam in the top 100,000 hosts.

http://hdl.handle.net/1969.1/ETD-TAMU-2011-12-10235

search engines

web crawling

spam

Identifer	oai:union.ndltd.org:tamu.edu/oai:repository.tamu.edu:1969.1/ETD-TAMU-2011-12-10235
Date	2011 December 1900
Creators	Mathiharan, Siddhartha Sankaran
Contributors	Loguinov, Dmitri, Caverlee, James, Reddy, A. L. Narasimha
Source Sets	Texas A and M University
Language	en_US
Detected Language	English
Type	Thesis, thesis, text
Format	application/pdf

Page generated in 0.0025 seconds

Identifying Search Engine Spam Using DNS

Description

Links & Downloads

Tags

Additional Fields