Machine Learning based Depression Detection with NLP Applicationon Twitter Text Data

Main Article Content

Priyanka Srivastava, Mohammad Suaib

Abstract

Depression detection through social media analysis has gained increasing attention due to the accessibility of large-scale user-generated content. This study proposes a hybrid sentiment analysis framework that integrates Latent Dirichlet Allocation (LDA), Valence Aware Dictionary and Sentiment Reasoner (VADER), and Term Frequency–Inverse Document Frequency (TF-IDF) features to capture thematic, emotional, and statistical representations of text. The dataset, sourced from Twitter, was preprocessed with advanced filtering techniques to remove noise and irrelevant tokens, producing a robust feature space. ML classifiers are evaluated like Support Vector Machine (SVM), K-Nearest Neighbor (KNN), logistic regression, and random forest, and SVM achieved the best accuracy (95.1%), precision (98.8%), recall (89.4%), and F1-score (93.9%). The results show that the combination of probabilistic, lexicon-based and statistical approaches improves classification results over the individual approaches and provides an interpretable, scalable framework for computational depression detection.

Article Details

Section
Articles