Analisis Performa Hadoop Multi-Node Cluster pada Skenario Pemrosesan Small Files
DOI:
https://doi.org/10.55606/jutiti.v6i2.7642Keywords:
Big Data, Hadoop, Multi-Node Cluster, Resource Distribution, Small FilesAbstract
The rapid growth of Big Data demands computing systems that are capable of processing data efficiently and in a distributed manner. Apache Hadoop, as a distributed computing framework, is widely used to address the limitations of centralized systems; however, it still faces challenges in handling small files, which can lead to increased file management overhead and imbalanced resource utilization. This study aims to analyze the performance of a Hadoop Multi-Node Cluster in small file processing scenarios, particularly in terms of workload distribution and CPU and memory resource utilization across nodes. The research method involves implementing a Hadoop Multi-Node Cluster in a local environment with a configuration of one master node and five worker nodes, followed by concurrent small file processing tests using the MapReduce mechanism. Performance evaluation is conducted by monitoring resource utilization through Hadoop ResourceManager, as well as Prometheus and Grafana monitoring systems to obtain real-time performance visualizations. The results of this study are expected to provide insights into the effectiveness of resource distribution in Hadoop Multi-Node Clusters for handling small files and to serve as a reference for the development and evaluation of distributed computing systems in academic and research environments.
Downloads
References
Aggarwal, R., Verma, J., & Siwach, M. (2022a). Small files’ problem in Hadoop: A systematic literature review. Journal of King Saud University - Computer and Information Sciences, 34(10), 8658–8674. https://doi.org/10.1016/j.jksuci.2021.09.007
Aggarwal, R., Verma, J., & Siwach, M. (2022b). Small files’ problem in Hadoop: A systematic literature review. Journal of King Saud University - Computer and Information Sciences, 34(10), 8658–8674. https://doi.org/10.1016/j.jksuci.2021.09.007
Bakni, N.-E., & Assayad, I. (2024). Analysis of the MapReduce Performance in Hadoop. Revue d’Intelligence Artificielle, 38(6), 1391–1397. https://doi.org/10.18280/ria.380601
Bakni, N.-E., & Assayad, I. (2025). Smart Data Placement Strategy in Heterogeneous Hadoop. HighTech and Innovation Journal, 6(1), 42–53. https://doi.org/10.28991/HIJ-2025-06-01-03
Buhl, H. U., Röglinger, M., Moser, F., & Heidemann, J. (2013). Big Data: A Fashionable Topic with(out) Sustainable Relevance for Research and Practice? Business & Information Systems Engineering, 5(2), 65–69. https://doi.org/10.1007/s12599-013-0249-5
El-Sayed, T., Badawy, M., & El-Sayed, A. (2019). Impact of Small Files on Hadoop Performance: Literature Survey and Open Points. Menoufia Journal of Electronic Engineering Research, 28(1), 109–120. https://doi.org/10.21608/mjeer.2019.62728
Gaykar, R. S., Khanaa, V., & Joshi, S. D. (2022). Faulty Node Detection in HDFS Using Machine Learning Techniques. Revue d’Intelligence Artificielle, 36(4), 553–560. https://doi.org/10.18280/ria.360406
Giri, P. R., & Sharma, G. (2022). Apache Hadoop Architecture, Applications, and Hadoop Distributed File System. Semiconductor Science and Information Devices, 4(1), 14–20. https://doi.org/10.30564/ssid.v4i1.4619
Gupta, A., Santhiya, P., Thiyagarajan, C., Gupta, A., Gupta, M., & Dwivedi, R. Kr. (2025). The Cutting-Edge Hadoop Distributed File System: Un-leashing Optimal Performance. ICST Transactions on Scalable Information Systems, 12(5). https://doi.org/10.4108/eetsis.9027
Helwan University, Sawiris, B., El-Gaber, S., Helwan University, Abdel-Fattah, M., & Helwan University. (2021). Centralization of Big Data Using Distributed Computing Approach in IoT. International Journal of Intelligent Engineering and Systems, 14(4), 393–409. https://doi.org/10.22266/ijies2021.0831.35
Hong, Z., Xiao-ming, W., Jie, C., Yan-hong, M., Yi-rong, G., & Min, W. (2016). An Optimized Model for MapReduce Based on Hadoop. TELKOMNIKA (Telecommunication Computing Electronics and Control), 14(4), 1552. https://doi.org/10.12928/telkomnika.v14i4.3606
Julia, Mutahari, M. I., Renaldi, & Saepullah. (2024). Analisis Kinerja Basis Data Terdistribusi dalam Linkungan Cloud Computing. Karimah Tauhid, 3(2), 1771–1782. https://doi.org/10.30997/karimahtauhid.v3i2.11907
Kalia, K., & Gupta, N. (2021). Analysis of hadoop MapReduce scheduling in heterogeneous environment. Ain Shams Engineering Journal, 12(1), 1101–1110. https://doi.org/10.1016/j.asej.2020.06.009
Kretzer, A. R., Barreto Vavassori Benitti, F., & Siqueira, F. (2025). Challenges and Opportunities in Big Data Analytics for Industry 4.0: A Systematic Evaluation of Current Architectures. IEEE Access, 13, 183419–183447. https://doi.org/10.1109/ACCESS.2025.3624558
Mandala Putra, H., Akbar, T., Ahmadi, A., & Iman Darmawan, M. (2021). Analisa Performa Klastering Data Besar pada Hadoop. Infotek: Jurnal Informatika dan Teknologi, 4(2), 174–183. https://doi.org/10.29408/jit.v4i2.3565
Perwiratama, R. (2024). Desain Infrastruktur: Kubernetes dan Hadoop sebagai Penyimpanan Data Terdistribusi. KONSTELASI: Konvergensi Teknologi dan Sistem Informasi, 4(2). https://doi.org/10.24002/konstelasi.v4i2.10446
Rahul Dev Singh, Vikram Kumar Gupta, & Priya Anjali Patel. (2024). Optimization Of Big Data Processing Using Distributed Computing In Cloud Environments. International Journal of Computer Technology and Science, 1(2), 01–07. https://doi.org/10.62951/ijcts.v1i2.58
Saadoon, M., Hamid, S. H. A., Sofian, H., Altarturi, H., Nasuha, N., Azizul, Z. H., Sani, A. A., & Asemi, A. (2021). Experimental Analysis in Hadoop MapReduce: A Closer Look at Fault Detection and Recovery Techniques. Sensors, 21(11), 3799. https://doi.org/10.3390/s21113799
Setyowati, L., & Nasir Ahmad, D. (2021). Pemanfaatan Big Data Dalam Era Teknologi 5.0. ABDINE: Jurnal Pengabdian Masyarakat, 1(2), 117–122. https://doi.org/10.52072/abdine.v1i2.205
Shafiyah, S., Ahsan, A. S., & Asmara, R. (2022). Big Data Infrastructure Design Optimizes Using Hadoop Technologies Based on Application Performance Analysis. SISTEMASI, 11(1), 55. https://doi.org/10.32520/stmsi.v11i1.1510
Sudirman, A., Irawan, & Saharuna, Z. (2024). Pengembangan Sistem Big Data: Rancang Bangun Infrastruktur dengan Framework Hadoop. Journal of Informatics and Computer Engineering Research, 1(1), 25–32. https://doi.org/10.31963/jicer.v1i1.4919
Veri Ferdiansyah & Muhammad Irwan Padli Nasution. (2023). Penerapan Teknologi Big Data Dalam Pengembangan Database Pendidikan. Jurnal Riset Manajemen, 1(3), 22–29. https://doi.org/10.54066/jurma.v1i3.591
Zagan, E., & Danubianu, M. (2021). HADOOP: A Comparative Study between Single-Node and Multi-Node Cluster. International Journal of Advanced Computer Science and Applications, 12(2). https://doi.org/10.14569/IJACSA.2021.0120207
Zhang, D., Dai, Z.-Y., Sun, X.-P., Wu, X.-T., Li, H., Tang, L., & He, J.-H. (2024). A distributed data processing scheme based on Hadoop for synchrotron radiation experiments. Journal of Synchrotron Radiation, 31(3), 635–645. https://doi.org/10.1107/S1600577524002637
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Jurnal Teknik Informatika dan Teknologi Informasi

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.






