A fast distributed focused-web crawling

Harry T.Yani Achsan, Wahyu Catur Wibowo

Research output: Contribution to journalConference articlepeer-review

22 Citations (Scopus)

Abstract

Mining data from a web database becomes more challenging in recent years due to the exploding size of data, the rising of dynamic web, and the increasing performance of web security. Mining data from a web database differs from mining data from web sites because it is intended to collect specific data from a single web site. Collecting a very large data in a limited time tends to be detected as a cyber attack and will be banned from connecting into the web server. To avoid the problem, this paper proposes a crawling method to mine web database faster and cheaper than conventional web crawlers. The method used is to run hundreds of threads from a single web crawler in a single computer and to distribute the threads into hundreds or thousands publicly available proxy servers. This web crawler strategy highly increases the speed of mining and is more secure than using single thread of web crawler.

Original languageEnglish
Pages (from-to)492-499
Number of pages8
JournalProcedia Engineering
Volume69
DOIs
Publication statusPublished - 2014
Event2013 24th DAAAM International Symposium on Intelligent Manufacturing and Automation - Zadar, Croatia
Duration: 23 Oct 201326 Oct 2013

Keywords

  • Focused web crawler
  • Multi thread
  • Proxy
  • Web database

Fingerprint

Dive into the research topics of 'A fast distributed focused-web crawling'. Together they form a unique fingerprint.

Cite this