Toward Integrated CNN-based Sentiment Analysis of Tweets for Scarce-resource Language—Hindi

Vedika Gupta, Nikita Jain, Shubham Shubham, Agam Madan, Ankit Chaudhary, Qin Xin

Research output: Contribution to journalArticlepeer-review

31 Citations (Scopus)

Abstract

Linguistic resources for commonly used languages such as English and Mandarin Chinese are available in abundance, hence the existing research in these languages. However, there are languages for which linguistic resources are scarcely available. One of these languages is the Hindi language. Hindi, being the fourth-most popular language, still lacks in richly populated linguistic resources, owing to the challenges involved in dealing with the Hindi language. This article first explores the machine learning-based approaches—Naïve Bayes, Support Vector Machine, Decision Tree, and Logistic Regression—to analyze the sentiment contained in Hindi language text derived from Twitter.

Further, the article presents lexicon-based approaches (Hindi Senti-WordNet, NRC Emotion Lexicon) for sentiment analysis in Hindi while also proposing a Domain-specific Sentiment Dictionary. Finally, an integrated convolutional neural network (CNN)—Recurrent Neural Network and Long Short-term Memory—is proposed to analyze sentiment from Hindi language tweets, a total of 23,767 tweets classified into positive, negative, and neutral. The proposed CNN approach gives an accuracy of 85%.
Original languageEnglish
Article number80
Pages (from-to)1-23
Number of pages23
JournalACM Transactions on Asian and Low-Resource Language Information Processing
Volume20
Issue number5
DOIs
Publication statusPublished - 23 Jun 2021

Fingerprint

Dive into the research topics of 'Toward Integrated CNN-based Sentiment Analysis of Tweets for Scarce-resource Language—Hindi'. Together they form a unique fingerprint.

Cite this