Implementing web data extraction and making mashup with Xtractorz

Rudy A.G. Gultom, Riri Fitri Sari, Bagio Budiardjo

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

4 Citations (Scopus)

Abstract

Implementing web data extraction means we can directly extract data from various web pages, where they mostly formed in an unstructured HTML format, into a new structured format such as XML or XHTML. In this paper we review the implementation of web data extraction and stages in making a Mashup. We implement web data extraction by visually extract targeted data from data sources (web pages). Afterward, we combined web data extraction with the stages of making a Mashup, e.g. data retrieval, data source modeling, data cleaning/filtering, data integration and data visualization. Problems arise in querying data sources due to unstructured contents of web pages (HTML), we cannot directly extract data into a new structured form. To address this problem, we propose a system, called Xtractorz, that can perform web data extraction in a Mashup format. We provide a fully visual and interactive user interface with new technique and approach using PHP and AJAX as the programming languages, and MySQL as the Data Repository. Furthermore, Xtractorz enables the user to conduct their job without the need to write a script or program or even without any knowledge of computer programming. The test results shows that Xtractorz requires less number of steps in making a Mashup compared with RoboMaker and Karma.

Original languageEnglish
Title of host publication2010 IEEE 2nd International Advance Computing Conference, IACC 2010
Pages385-393
Number of pages9
DOIs
Publication statusPublished - 27 Apr 2010
Event2010 IEEE 2nd International Advance Computing Conference, IACC 2010 - Patiala, India
Duration: 19 Feb 201020 Feb 2010

Publication series

Name2010 IEEE 2nd International Advance Computing Conference, IACC 2010

Conference

Conference2010 IEEE 2nd International Advance Computing Conference, IACC 2010
CountryIndia
CityPatiala
Period19/02/1020/02/10

Keywords

  • AJAX
  • DOM tree
  • HTML
  • Making mashup
  • Mashup stages
  • MySQL
  • PHP
  • Web data extraction
  • XML

Fingerprint Dive into the research topics of 'Implementing web data extraction and making mashup with Xtractorz'. Together they form a unique fingerprint.

Cite this