Analisis Performa Hadoop Multi-Node Cluster pada Skenario Pemrosesan Small Files

Authors

  • I Putu Ngurah Rio Aditya Pratama Universitas Pendidikan Nasional
  • Ngakan Kutha Krisnawijaya Universitas Pendidikan Nasional

DOI:

https://doi.org/10.55606/jutiti.v6i2.7642

Keywords:

Big Data, Hadoop, Multi-Node Cluster, Resource Distribution, Small Files

Abstract

The rapid growth of Big Data demands computing systems that are capable of processing data efficiently and in a distributed manner. Apache Hadoop, as a distributed computing framework, is widely used to address the limitations of centralized systems; however, it still faces challenges in handling small files, which can lead to increased file management overhead and imbalanced resource utilization. This study aims to analyze the performance of a Hadoop Multi-Node Cluster in small file processing scenarios, particularly in terms of workload distribution and CPU and memory resource utilization across nodes. The research method involves implementing a Hadoop Multi-Node Cluster in a local environment with a configuration of one master node and five worker nodes, followed by concurrent small file processing tests using the MapReduce mechanism. Performance evaluation is conducted by monitoring resource utilization through Hadoop ResourceManager, as well as Prometheus and Grafana monitoring systems to obtain real-time performance visualizations. The results of this study are expected to provide insights into the effectiveness of resource distribution in Hadoop Multi-Node Clusters for handling small files and to serve as a reference for the development and evaluation of distributed computing systems in academic and research environments.

Downloads

Download data is not yet available.

References

Aggarwal, R., Verma, J., & Siwach, M. (2022a). Small files’ problem in Hadoop: A systematic literature review. Journal of King Saud University - Computer and Information Sciences, 34(10), 8658–8674. https://doi.org/10.1016/j.jksuci.2021.09.007

Aggarwal, R., Verma, J., & Siwach, M. (2022b). Small files’ problem in Hadoop: A systematic literature review. Journal of King Saud University - Computer and Information Sciences, 34(10), 8658–8674. https://doi.org/10.1016/j.jksuci.2021.09.007

Bakni, N.-E., & Assayad, I. (2024). Analysis of the MapReduce Performance in Hadoop. Revue d’Intelligence Artificielle, 38(6), 1391–1397. https://doi.org/10.18280/ria.380601

Bakni, N.-E., & Assayad, I. (2025). Smart Data Placement Strategy in Heterogeneous Hadoop. HighTech and Innovation Journal, 6(1), 42–53. https://doi.org/10.28991/HIJ-2025-06-01-03

Buhl, H. U., Röglinger, M., Moser, F., & Heidemann, J. (2013). Big Data: A Fashionable Topic with(out) Sustainable Relevance for Research and Practice? Business & Information Systems Engineering, 5(2), 65–69. https://doi.org/10.1007/s12599-013-0249-5

El-Sayed, T., Badawy, M., & El-Sayed, A. (2019). Impact of Small Files on Hadoop Performance: Literature Survey and Open Points. Menoufia Journal of Electronic Engineering Research, 28(1), 109–120. https://doi.org/10.21608/mjeer.2019.62728

Gaykar, R. S., Khanaa, V., & Joshi, S. D. (2022). Faulty Node Detection in HDFS Using Machine Learning Techniques. Revue d’Intelligence Artificielle, 36(4), 553–560. https://doi.org/10.18280/ria.360406

Giri, P. R., & Sharma, G. (2022). Apache Hadoop Architecture, Applications, and Hadoop Distributed File System. Semiconductor Science and Information Devices, 4(1), 14–20. https://doi.org/10.30564/ssid.v4i1.4619

Gupta, A., Santhiya, P., Thiyagarajan, C., Gupta, A., Gupta, M., & Dwivedi, R. Kr. (2025). The Cutting-Edge Hadoop Distributed File System: Un-leashing Optimal Performance. ICST Transactions on Scalable Information Systems, 12(5). https://doi.org/10.4108/eetsis.9027

Helwan University, Sawiris, B., El-Gaber, S., Helwan University, Abdel-Fattah, M., & Helwan University. (2021). Centralization of Big Data Using Distributed Computing Approach in IoT. International Journal of Intelligent Engineering and Systems, 14(4), 393–409. https://doi.org/10.22266/ijies2021.0831.35

Hong, Z., Xiao-ming, W., Jie, C., Yan-hong, M., Yi-rong, G., & Min, W. (2016). An Optimized Model for MapReduce Based on Hadoop. TELKOMNIKA (Telecommunication Computing Electronics and Control), 14(4), 1552. https://doi.org/10.12928/telkomnika.v14i4.3606

Julia, Mutahari, M. I., Renaldi, & Saepullah. (2024). Analisis Kinerja Basis Data Terdistribusi dalam Linkungan Cloud Computing. Karimah Tauhid, 3(2), 1771–1782. https://doi.org/10.30997/karimahtauhid.v3i2.11907

Kalia, K., & Gupta, N. (2021). Analysis of hadoop MapReduce scheduling in heterogeneous environment. Ain Shams Engineering Journal, 12(1), 1101–1110. https://doi.org/10.1016/j.asej.2020.06.009

Kretzer, A. R., Barreto Vavassori Benitti, F., & Siqueira, F. (2025). Challenges and Opportunities in Big Data Analytics for Industry 4.0: A Systematic Evaluation of Current Architectures. IEEE Access, 13, 183419–183447. https://doi.org/10.1109/ACCESS.2025.3624558

Mandala Putra, H., Akbar, T., Ahmadi, A., & Iman Darmawan, M. (2021). Analisa Performa Klastering Data Besar pada Hadoop. Infotek: Jurnal Informatika dan Teknologi, 4(2), 174–183. https://doi.org/10.29408/jit.v4i2.3565

Perwiratama, R. (2024). Desain Infrastruktur: Kubernetes dan Hadoop sebagai Penyimpanan Data Terdistribusi. KONSTELASI: Konvergensi Teknologi dan Sistem Informasi, 4(2). https://doi.org/10.24002/konstelasi.v4i2.10446

Rahul Dev Singh, Vikram Kumar Gupta, & Priya Anjali Patel. (2024). Optimization Of Big Data Processing Using Distributed Computing In Cloud Environments. International Journal of Computer Technology and Science, 1(2), 01–07. https://doi.org/10.62951/ijcts.v1i2.58

Saadoon, M., Hamid, S. H. A., Sofian, H., Altarturi, H., Nasuha, N., Azizul, Z. H., Sani, A. A., & Asemi, A. (2021). Experimental Analysis in Hadoop MapReduce: A Closer Look at Fault Detection and Recovery Techniques. Sensors, 21(11), 3799. https://doi.org/10.3390/s21113799

Setyowati, L., & Nasir Ahmad, D. (2021). Pemanfaatan Big Data Dalam Era Teknologi 5.0. ABDINE: Jurnal Pengabdian Masyarakat, 1(2), 117–122. https://doi.org/10.52072/abdine.v1i2.205

Shafiyah, S., Ahsan, A. S., & Asmara, R. (2022). Big Data Infrastructure Design Optimizes Using Hadoop Technologies Based on Application Performance Analysis. SISTEMASI, 11(1), 55. https://doi.org/10.32520/stmsi.v11i1.1510

Sudirman, A., Irawan, & Saharuna, Z. (2024). Pengembangan Sistem Big Data: Rancang Bangun Infrastruktur dengan Framework Hadoop. Journal of Informatics and Computer Engineering Research, 1(1), 25–32. https://doi.org/10.31963/jicer.v1i1.4919

Veri Ferdiansyah & Muhammad Irwan Padli Nasution. (2023). Penerapan Teknologi Big Data Dalam Pengembangan Database Pendidikan. Jurnal Riset Manajemen, 1(3), 22–29. https://doi.org/10.54066/jurma.v1i3.591

Zagan, E., & Danubianu, M. (2021). HADOOP: A Comparative Study between Single-Node and Multi-Node Cluster. International Journal of Advanced Computer Science and Applications, 12(2). https://doi.org/10.14569/IJACSA.2021.0120207

Zhang, D., Dai, Z.-Y., Sun, X.-P., Wu, X.-T., Li, H., Tang, L., & He, J.-H. (2024). A distributed data processing scheme based on Hadoop for synchrotron radiation experiments. Journal of Synchrotron Radiation, 31(3), 635–645. https://doi.org/10.1107/S1600577524002637

Downloads

Published

2026-07-29

How to Cite

I Putu Ngurah Rio Aditya Pratama, & Ngakan Kutha Krisnawijaya. (2026). Analisis Performa Hadoop Multi-Node Cluster pada Skenario Pemrosesan Small Files. Jurnal Teknik Informatika Dan Teknologi Informasi, 6(2), 498–516. https://doi.org/10.55606/jutiti.v6i2.7642

Similar Articles

<< < 4 5 6 7 8 9 10 11 12 13 > >> 

You may also start an advanced similarity search for this article.