Analisis Perbandingan Algoritma TF-IDF dan Word2Vec dalam Rekomendasi Film Berbasis Konten Letterboxd

Systematic Literature Review

Authors

  • Diva Bulan Universitas Muhammadiyah Riau
  • Yulia Fatma Universitas Muhammadiyah Riau

DOI:

https://doi.org/10.55606/jutiti.v6i2.7244

Keywords:

Content-Based Filtering, Letterboxd, Natural Language Processing, Recomendation Film System, TF-IDF, Word2Vec

Abstract

The exponential growth of film streaming platforms has created significant challenges for users in discovering content that aligns with their preferences. Recommendation systems have proven to be a strategic solution. This study aims to compare the performance of Term Frequency-Inverse Document Frequency (TF-IDF) and Word2Vec algorithms in the context of content-based filtering recommendation systems, utilizing data from the Letterboxd platform. The research was conducted through a Systematic Literature Review (SLR) of 15 prior studies, with a selection process comprising the definition of inclusion and exclusion criteria, literature searches across academic databases, and comparative synthesis of findings. The review findings indicate that TF-IDF excels in weighting explicit terms within structured metadata such as genre, director, and cast, and can be implemented without a complex training process. Word2Vec, on the other hand, demonstrates superiority in capturing semantic relationships between words through dense vector representations, yet requires a large corpus and careful hyperparameter tuning. Based on the literature synthesis, TF-IDF combined with Cosine Similarity is deemed more optimal for a content-based film recommendation system using Letterboxd metadata, given the structured and explicit nature of the data. Recall@5 of 73% and Recall@10 of 80% reported in prior studies further support this conclusion

Downloads

Download data is not yet available.

References

Agyemang, Edmund, Lawrence Agbota, Vincent Agbenyeavu, Peggy Akabuah, Bismark Bimpong, and Christopher Attafuah. 2025. “Prediction of Coffee Ratings Based On Influential Attributes Using SelectKBest and Optimal Hyperparameters.” 1–13. http://arxiv.org/abs/2509.18124.

Ahmad, Fiaz, Nisar Hussain, Amna Qasim, Momina Hafeez, Muhammad Usman Grigori, and Alexander Gelbukh. n.d. “Irony Detection in Urdu Text : A Comparative Study Using Machine Learning Models and Large Language Models.” 1–5.

Amalia, Junita, Juanda Pakpahan, Melani Pakpahan, and Yeni Panjaitan. 2022. “Model Klasifikasi Berita Palsu Menggunakan Bidirectional LSTM Dan Word2Vec Sebagai Vektorisasi.” 9(4):3319–31.

Analyzing the Evolution of Graphs and Texts. 2023. (May).

Gifari, Okta Ihza, Muh. Adha, Fernandito Freddy, and Fernandito Freddy Setlight Durrand. 2022. “Analisis Sentimen Review Film Menggunakan TF-IDF Dan Support Vector Machine.” Journal of Information Technology 2(1):36–40. doi:10.46229/jifotech.v2i1.330.

Huda, Arif Akbarul, Rohmad Fajarudin, and Arifiyanto Hadinegoro. 2022. “Sistem Rekomendasi Content-Based Filtering Menggunakan TF-IDF Vector Similarity Untuk Rekomendasi Artikel Berita.” Building of Informatics, Technology and Science (BITS) 4(3):1679–86. doi:10.47065/bits.v4i3.2511.

Khomsah, Siti, Rima Dias Ramadhani, and Sena Wijayanto. 2022. “The Accuracy Comparison Between Word2Vec and FastText On.” Rekayasa Sistem Dan Teknologi Informasi 5(158):352–58.

Kokot, Robin, and Wessel Poelman. 2025. “Type and Complexity Signals in Multilingual Question Representations.” 411–25. doi:10.18653/v1/2025.mrl-main.28.

Majidi, Mohammad Zana, Sajjad Karimi, Teng Wang, Robert Kluger, and Reginald Souleyrette. 2025. “Predicting Person-Level Injury Severity Using Crash Narratives: A Balanced Approach with Roadway Classification and Natural Language Process Techniques.”

Mamatov, Nursultan, and Philipp Kellmeyer. 2025. “Enhancing Mortality Prediction in Cardiac Arrest ICU Patients through Meta-Modeling of Structured Clinical Data from MIMIC-IV.” http://arxiv.org/abs/2510.18103.

Rahman, Shadikur, Hasibul Karim Shanto, Umme Ayman Koana, and Syed Muhammad Danish. 2025. “Automated Research Article Classification and Recommendation Using NLP and ML.” http://arxiv.org/abs/2510.05495.

Royyan, Alvi Rahmy, and Erwin Budi Setiawan. 2022. “Feature Expansion Word2Vec for Sentiment Analysis of Public Policy in Twitter.” Jurnal RESTI 6(1):78–84. doi:10.29207/resti.v6i1.3525.

Safira, Stevani Dean, Nanda Rosma Anwar, and Medistra Aldrin. 2025. “Implementasi Metode Term Frequency-Inverse Document Frequency Dan Cosine Similarity Pada Bot Telegram Untuk Mendukung Rekomendasi Laptop Berdasarkan Preferensi Pengguna Implementation of the TF-IDF and Cosine Similarity Methods in Telegram Bots.” 4(1):63–73.

Yerramsetty, Surya Tejaswi, and Almas Fathimah. 2025. “Multi-Label Clinical Text Eligibility Classification and Summarization System.” https://arxiv.org/pdf/2510.13115.

Zhang, Zizhao, Tianxiang Zhao, Yu Sun, Liping Sun, and Jichuan Kang. 2024. “Graph-Structured Data Analysis of Component Failure in Autonomous Cargo Ships Based on Feature Fusion.” 1–22.

Downloads

Published

2026-06-13

How to Cite

Diva Bulan, & Yulia Fatma. (2026). Analisis Perbandingan Algoritma TF-IDF dan Word2Vec dalam Rekomendasi Film Berbasis Konten Letterboxd: Systematic Literature Review. Jurnal Teknik Informatika Dan Teknologi Informasi, 6(2), 23–33. https://doi.org/10.55606/jutiti.v6i2.7244

Similar Articles

<< < 2 3 4 5 6 7 8 9 10 11 > >> 

You may also start an advanced similarity search for this article.