Comparative Study of BERT Variants for Sentiment Analysis with Error Analysis

Year : 2026 | Volume : 04 | Issue : 01 | Page : 39 48
By

Deepshikha Prajapati,

Shruti Prajapati,

Thara Chakkingal,

  1. Research Scholar, Master of Computer Application, Thakur Institute of Management Studies, Career Development & Research, Mumbai, Maharashtra, India
  2. Research Scholar, Master of Computer Application, Thakur Institute of Management Studies, Career Development & Research, Mumbai, Maharashtra, India
  3. Assistant Professor, Master of Computer Application, Thakur Institute of Management Studies, Career Development & Research, Mumbai, Maharashtra, India

Abstract

The use of media is going up fast in India, and this has led to the rise of Hinglish. Hinglish is an informal blend of Hindi and English that people commonly use in everyday conversations, especially across social media platforms such as Twitter, Facebook, and WhatsApp. People use Hinglish to talk to each other in a way that is not very formal. Hinglish blends English vocabulary with informal usage, often ignoring standard grammatical rules, which makes it challenging for computers to accurately interpret and process it. It is especially hard for computers to figure out how people are feeling when they use Hinglish in the media. Hinglish poses a significant challenge for natural language processing, particularly when it comes to accurately interpreting people’s emotions and opinions. The problem with Hinglish is that it often has words from languages mixed together, and the spelling and sentence structure can be weird. This makes it hard for regular language models to understand. Even though models like BERT are really good at understanding text in languages, they need a lot of computer power to work. This means they are not good for situations where we need to get answers, and we do not have a lot of computer power. So, we looked at how three smaller models work: DistilBERT, Multilingual Representations for Indian Languages (MuRIL), and XLM-RoBERTa. We used the Kaggle Hinglish Sentiment Dataset to test these models. When we look at how these models work and where they make mistakes, the research helps us understand how they can handle the complexities of Hinglish language. This is important because Hinglish is a mix of Hindi and English. The research shows that these models can work with Hinglish while still being good at classifying things and not using much computer power. Studying is helpful because it adds to what we know about working with languages that do not have a lot of resources. It also helps us make systems that can figure out how people feel about things. We can use these systems in real life. The research on Hinglish models is useful for sentiment analysis systems.

Keywords: Hinglish, code-mixed text, sentiment analysis, natural language processing (NLP), transformer models

[This article belongs to International Journal of Computer Science Languages ]

How to cite this article: Deepshikha Prajapati, Shruti Prajapati, Thara Chakkingal. Comparative Study of BERT Variants for Sentiment Analysis with Error Analysis. International Journal of Computer Science Languages. 2026; 04(01):39-48.
How to cite this URL: Deepshikha Prajapati, Shruti Prajapati, Thara Chakkingal. Comparative Study of BERT Variants for Sentiment Analysis with Error Analysis. International Journal of Computer Science Languages. 2026; 04(01):39-48. Available from: https://journals.stmjournals.com/ijcsl/article=2026/view=247731

References

  1. Devlin J, Chang MW, Lee K, Toutanova K. BERT: pre-training of deep bidirectional transformers for language understanding. In: Burstein J, Doran C, Solorio T, editors. Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. Vol. 1, Long and Short Papers; 2019 Jun; Minneapolis, MN, USA. Stroudsburg (PA): Association for Computational Linguistics; 2019. p. 4171–4186. doi:10.18653/v1/N19-1423.
  2. Çetinoğlu Ö, Schulz S, Vu NT. Challenges of computational processing of code-switching. In: Diab M, Fung P, Ghoneim M, Hirschberg J, Solorio T, editors. Proceedings of the Second Workshop on Computational Approaches to Code Switching; 2016 Nov; Austin, TX, USA. Stroudsburg (PA): Association for Computational Linguistics; 2016. p. 1–11. doi:10.18653/v1/W16-5801.
  3. Patil A, Patwardhan V, Phaltankar A, Takawane G, Joshi R. Comparative study of pre-trained BERT models for code-mixed Hindi-English data. 2023 IEEE 8th International Conference for Convergence in Technology (I2CT), Lonavla, India. 2023. p. 1–7. doi:10.1109/I2CT57861.2023.10126273.
  4. Hashmi E, Yayilgan SY, Shaikh S. Augmenting sentiment prediction capabilities for code-mixed tweets with multilingual transformers. Soc Netw Anal Min. 2024;14:86. doi:10.1007/s13278-024-01245-6.
  5. Astuti LW, Sari Y, S. Code-mixed sentiment analysis using transformer for Twitter social media data. Int J Adv Comput Sci Appl. 2023;14(10). doi:10.14569/IJACSA.2023.0141053.
  6. Sampath KK, Supriya M. Transformer based sentiment analysis on code mixed data. Procedia Comput Sci. 2024;233:682–691. doi:10.1016/j.procs.2024.03.257.
  7. Tewari P, Gumber M, Tyagi S, Seekhwal P. Hinglish text analysis: challenges and opportunities in multilingual natural language processing. In: Proceedings of the International Conference on Innovative Computing & Communication (ICICC 2024); 2024. SSRN Electron J. 2025. doi:10.2139/ssrn.5193470.
  8. Jadon AS, Parmar M, Agrawal R. Hinglish sentiment analysis: deep learning models for nuanced sentiment classification in multilingual digital communication. 2024 2nd International Conference on Device Intelligence, Computing and Communication Technologies (DICCT), Dehradun, India. IEEE; 2024. p. 318–323. doi:10.1109/DICCT61038.2024.10533057.
  9. Singh SK, Sharma A, Sahil D, Singh D, Pandit S, Saghir U. Sentiment analysis of English-Hindi code-mixed text using mBERT model. 2025 3rd International Conference on Inventive Computing and Informatics (ICICI), Bangalore, India. 2025. p. 552–556. doi:10.1109/ICICI65870.2025.11069692.
  10. Mohana Priya KT, Shrinithi G, Nithish P, Pranesh AC. Comparative analysis of transformer models for sentiment classification in code-mixed Indic languages. Int J Eng Res Sustain Technol. 2025;3(1):1–9. doi:10.63458/ijerst.v3i1.101.
  11. Almalki SS. Sentiment analysis and emotion detection using transformer models in multilingual social media data. Int J Adv Comput Sci Appl. 2025;16(3). doi:10.14569/IJACSA.2025.0160332.
  12. Aliyu Y, Sarlan A, Danyaro KU, Sani Abd Rahman AS, Muazu AA, Abubakar MY. Deep learning techniques for sentiment analysis in code-switched Hausa-English tweets. Int J Inf Manag Data Insights. 2025;5(1):100330. doi:10.1016/j.jjimei.2025.100330.
  13. Ramesh G, Doddapaneni S, Bheemaraj A, Jobanputra M, Ak R, Sharma A, et al. Samanantar: the largest publicly available parallel corpora collection for 11 Indic languages. Trans Assoc Comput Linguist. 2022;10:145–162. doi:10.1162/tacl_a_00452.
  14. Nazir MK, Faisal CN, Habib MA, Ahmad H. Leveraging multilingual transformer for multiclass sentiment analysis in code-mixed data of low-resource languages. IEEE Access. 2025;13:7538–7554. doi:10.1109/ACCESS.2025.3527710.
  15. Mamta EA, Ekbal A. Transformer based multilingual joint learning framework for code-mixed and English sentiment analysis. J Intell Inf Syst. 2024;62:231–253. doi:10.1007/s10844-023-00808-x.
  16. Veeramani H, Thapa S, Naseem U. MLInitiative@WILDRE7: hybrid approaches with large language models for enhanced sentiment analysis in code-switched and code-mixed texts. In: Jha GN, Sobha L, Bali K, Ojha AK, editors. Proceedings of the 7th Workshop on Indian Language Data: Resources and Evaluation; 2024 May; Torino, Italia. Paris: ELRA and ICCL; 2024. p.
    66–72.
  17. Thakur V, Sahu R, Omer S. Current state of Hinglish text sentiment analysis [Preprint]. SSRN Electron J. 2020. doi:10.2139/ssrn.3614442.
  18. Yuan LS, Ming LT. Sentiment prediction using multilingual bidirectional encoder representations and cross-lingual language model robustly optimized BERT approach from transformers on code-mixed text. 2025 IEEE International Conference on Computation, Big-Data and Engineering (ICCBE), Penang, Malaysia. 2025. p. 901–905. doi:10.1109/ICCBE65177.2025.11256168.
  19. Kumar A, Susan S. Supervised sentiment analysis of movie reviews with SHAP-based interpretability analysis. In: Swaroop A, Virdee B, Correia SD, Polkowski Z, editors. Proceedings of Data Analytics and Management. ICDAM 2024. Lecture Notes in Networks and Systems. Vol. 1300. Singapore: Springer; 2025. p. 381–388. doi:10.1007/978-981-96-3361-6_28.

Regular Issue Subscription Original Research
Volume 04
Issue 01
Received 10/03/2026
Accepted 11/04/2026
Published 27/04/2026
Publication Time 48 Days


Login

My IP

PlumX Metrics