Development of an ETL Process Based on Open Source Technologies to Solve the Problem of Data Delivery to Consumers

 
Audio is AI-generated
352

Abstract

The article discusses the issues of developing an ETL process for a data warehouse based on open source technologies, instead of private software supplied by the vendor. The process allows you to deliver data from the source to the consumer, focusing on the speed of delivery, the resources spent and the convenience of development. The architecture for solving the problem with a description of the processes being replaced is presented, data transmission over a new process is implemented. Modern tools used to work with data are involved, methods of interaction with them and selection of technical characteristics for the process are described.

General Information

Keywords: database, open source, software, ETL process, data delivery

Journal rubric: Software

OpenAlex citations: 1

OpenAlex trends: Economic and Technological Systems Analysis, Economic and Technological Developments in Russia, Engineering Education and Technology

Information about the work in OpenAlex

Number of citations: 1

Topics

Economic and Technological Systems Analysis

This cluster of papers covers a wide range of topics related to digital transformation, innovation management, and the integration of information technologies in various domains. It includes research on quality management, cyber-physical systems, big data, sustainability, neural networks, environmental safety, and project management.

Number of works: 36814  |  Total number of citations: 64075

Topic detailsв OpenAlex

Economic and Technological Developments in Russia

This cluster of papers covers a wide range of topics related to economic development, innovation, climate change, global economy, digital transformation, sustainable development, knowledge clusters, regional development, and macroeconomic assessment. The focus is on understanding and addressing the challenges and opportunities in the global economic landscape, with specific emphasis on the socio-economic development and policies in Russia.

Number of works: 86421  |  Total number of citations: 127399

Topic detailsв OpenAlex

Engineering Education and Technology

This cluster of papers explores the concept of Smart University within the context of Digital Ecosystems, focusing on topics such as Internet of Things, Artificial Intelligence, cognitive modeling, educational technology, innovation management, predictive maintenance, knowledge management, and industrial control systems. The papers discuss the principles and semantics of digital ecosystems, machine learning algorithms for smart data analysis, ontology of smart classrooms, mobile social networking for smart campuses, and the development and evaluation models for smart universities.

Number of works: 12917  |  Total number of citations: 36014

Topic detailsв OpenAlex

Work details in OpenAlex

Article type: scientific article

DOI: https://doi.org/10.17759/mda.2023130210

Received 12.04.2023

Published

For citation: Starkov, V.V., Gorbatova, S.S., Vodolaga, V.I. (2023). Development of an ETL Process Based on Open Source Technologies to Solve the Problem of Data Delivery to Consumers. Modelling and Data Analysis, 13(2), 180–193. (In Russ.). https://doi.org/10.17759/mda.2023130210

© Starkov V.V., Gorbatova S.S., Vodolaga V.I., 2023

License: CC BY-NC 4.0

References

  1. David Loshin. ETL (Extract, Transform, Load) . Business Intelligence. — 2nd. — Morgan Kaufmann, 2012. — 400 p
  2. Ralph Kimball, Joe Caserta. The Data Warehouse ETL Toolkit: Practical Techniques for Extracting, Cleaning, Conforming, and Delivering Data. — John Wiley & Sons, 2004. — 528 p.
  3. David Haertzen. ETL Tools . The Analytical Puzzle: Profitable Data Warehousing, Business Intelligence and Analytics. — Technics Publications, 2012. — 346 p.
  4. S. Riza, U. Lezerson, Sh. Ouen, D. Uills. Spark dlya professionalov: sovremennye patterny obrabotki bol'shikh dannykh = Advanced Analytics with Spark. Patterns for Learning from Data at Scale (O’Reilly, 2015). 2017. — 272 p.
  5. Uorren R., Karau Kh. Effektivnyi Spark. Masshtabirovanie i optimizatsiya = High Performance Spark. Best Practices for Scaling and Optimizing Apache Spark. 2018. — 352 s.
  6. Kh. Karau, E. Konvinski, P. Vendell, M. Zakhariya. Izuchaem Spark. Molnienosnyi analiz dannykh = Learning Spark: Lightning-Fast Big Data Analytics (O’Reilly, 2015). 2015. — 304 s.
  7. Narkhid Niya, Shapira Gven, Palino Todd. Apache Kafka. Potokovaya obrabotka i analiz dannykh. — SPb., 2019 p = 320.
  8. Vohra, Deepak (October 2016). Practical Hadoop Ecosystem: A Definitive Guide to Hadoop-Related Frameworks and Tools (1st ed.). Apress. p. 429.

Information About the Authors

Viacheslav V. Starkov, e-mail: starkov.viatcheslav@yandex.ru

Svetlana S. Gorbatova, Senior Lecturer, Moscow Institute of Steel and Alloys (National Research Technological University) (NUST MISIS), Moscow, Russian Federation, ORCID: https://orcid.org/0009-0005-5213-6780, e-mail: ssgorbatova@misis.ru

Victoria I. Vodolaga, Master's Degree, Lomonosov Moscow State University (MSU), Moscow, Russian Federation, ORCID: https://orcid.org/0009-0003-1816-0088, e-mail: vikavodolaga1@gmail.com

Metrics

 Web Views

Whole time: 716
Previous month: 14
Current month: 15

 PDF Downloads

Whole time: 352
Previous month: 16
Current month: 11

 Total

Whole time: 1068
Previous month: 30
Current month: 26