Smart Data Placement Using Storage-as-a-Service Model for Big Data Pipelines

Khan, Akif Quddus; Nikolov, Nikolay; Matskin, Minhail; Prodan, Radu; Roman, Dumitru; Sahin, Bekir; Bussler, Christoph; Soylu, Ahmet

dc.contributor.author	Khan, Akif Quddus
dc.contributor.author	Nikolov, Nikolay
dc.contributor.author	Matskin, Minhail
dc.contributor.author	Prodan, Radu
dc.contributor.author	Roman, Dumitru
dc.contributor.author	Sahin, Bekir
dc.contributor.author	Bussler, Christoph
dc.contributor.author	Soylu, Ahmet
dc.date.accessioned	2023-10-06T11:40:09Z
dc.date.available	2023-10-06T11:40:09Z
dc.date.created	2023-01-08T15:20:33Z
dc.date.issued	2023
dc.identifier.citation	Sensors. 2023, 23 (2), 564.	en_US
dc.identifier.issn	1424-8220
dc.identifier.uri	https://hdl.handle.net/11250/3094958
dc.description.abstract	Big data pipelines are developed to process data characterized by one or more of the three big data features, commonly known as the three Vs (volume, velocity, and variety), through a series of steps (e.g., extract, transform, and move), making the ground work for the use of advanced analytics and ML/AI techniques. Computing continuum (i.e., cloud/fog/edge) allows access to virtually infinite amount of resources, where data pipelines could be executed at scale; however, the implementation of data pipelines on the continuum is a complex task that needs to take computing resources, data transmission channels, triggers, data transfer methods, integration of message queues, etc., into account. The task becomes even more challenging when data storage is considered as part of the data pipelines. Local storage is expensive, hard to maintain, and comes with several challenges (e.g., data availability, data security, and backup). The use of cloud storage, i.e., storage-as-a-service (StaaS), instead of local storage has the potential of providing more flexibility in terms of scalability, fault tolerance, and availability. In this article, we propose a generic approach to integrate StaaS with data pipelines, i.e., computation on an on-premise server or on a specific cloud, but integration with StaaS, and develop a ranking method for available storage options based on five key parameters: cost, proximity, network performance, server-side encryption, and user weights/preferences. The evaluation carried out demonstrates the effectiveness of the proposed approach in terms of data transfer performance, utility of the individual parameters, and feasibility of dynamic selection of a storage option based on four primary user scenarios.	en_US
dc.language.iso	eng	en_US
dc.publisher	MDPI	en_US
dc.relation.uri	https://www.mdpi.com/1424-8220/23/2/564
dc.rights	Navngivelse 4.0 Internasjonal	*
dc.rights.uri	http://creativecommons.org/licenses/by/4.0/deed.no	*
dc.title	Smart Data Placement Using Storage-as-a-Service Model for Big Data Pipelines	en_US
dc.title.alternative	Smart Data Placement Using Storage-as-a-Service Model for Big Data Pipelines	en_US
dc.type	Peer reviewed	en_US
dc.type	Journal article	en_US
dc.description.version	publishedVersion	en_US
dc.rights.holder	© 2023 by the authors. Licensee MDPI, Basel, Switzerland.	en_US
dc.source.volume	23	en_US
dc.source.journal	Sensors	en_US
dc.source.issue	2	en_US
dc.identifier.doi	10.3390/s23020564
dc.identifier.cristin	2102762
dc.relation.project	Norges forskningsråd: 309691	en_US
dc.source.articlenumber	564	en_US
cristin.ispublished	true
cristin.fulltext	original
cristin.qualitycode	1

Tilhørende fil(er)

Filnavn:: sensors-23-00564.pdf
Størrelse:: 1.074Mb
Format:: PDF

Åpne

Denne innførselen finnes i følgende samling(er)

Publikasjoner fra CRIStin - SINTEF AS [5802]
SINTEF Digital [2501]

Vis enkel innførsel

Med mindre annet er angitt, så er denne innførselen lisensiert som Navngivelse 4.0 Internasjonal