Abstract
The paper presents a workflow application for efficient parallel processing of data downloaded from an Internet portal. The workflow partitions input files into subdirectories which are further split for parallel processing by services installed on distinct computer nodes. This way, analysis of the first ready subdirectories can start fast and is handled by services implemented as parallel multithreaded applications using multiple cores of modern CPUs. The goal is to assess achievable speed-ups and determine which factors influence scalability and to what degree. Data processing services were implemented for assessment of context (positive or negative) in which the given keyword appears in a document. The testbed application used these services to determine how a particular brand was recognized by either authors of articles or readers in comments in a specific Internet portal focused on new technologies. Obtained execution times as well as speed-ups are presented for data sets of various sizes along with discussion on how factors such as load imbalance and memory/disk bottlenecks limit performance
Citations
-
3
CrossRef
-
0
Web of Science
-
4
Scopus
Author (1)
Cite as
Full text
full text is not available in portal
Keywords
Details
- Category:
- Conference activity
- Type:
- materiały konferencyjne indeksowane w Web of Science
- Title of issue:
- Procedia Computer Science, vol. 29 strony 499 - 508
- ISSN:
- 1877-0509
- Language:
- English
- Publication year:
- 2014
- Bibliographic description:
- Czarnul P..: A Workflow Application for Parallel Processing of Big Data from an Internet Portal, W: Procedia Computer Science, vol. 29, 2014, Elsevier,.
- DOI:
- Digital Object Identifier (open in new tab) 10.1016/j.procs.2014.05.045
- Verified by:
- Gdańsk University of Technology
seen 113 times