An Efficient Extractive Summarization Framework Using TF-IDF, Centroid Scoring, and Dynamic Semantic Redundancy Control

Main Article Content

Himanshu Kumar, V K Jain, Vivek Kumar

Abstract

Automatic text summarization has become an essential Natural Language Processing tasks for dealing with the escalating amount of information available in textual formats digitally. Extractive summarization algorithms have been extensively used due to their efficiency and explainability, yet the existing models suffer from issues like semantic redundancy, lack of context-awareness, and dependence on static sentence selection processes. In order to solve these issues, this paper presents NewsCos-R a lightweight extractive summarization system combining the use of TF-IDF weighting, centroid-based cosine scoring, and the innovative approach of Dynamic Semantic Redundancy Control (DSRC). The proposed DSRC technique is capable of dynamically controlling redundancy elimination based on the degree of sentence relevance and context-aware semantic overlap. As opposed to typical threshold-based sentence deletion in static settings, the proposed DSRC allows for more flexible and informative sentence filtering. First, the input text is preprocessed. Then, TF-IDF sentence vectors are calculated, centroids are constructed and assigned scores of semantic relevance, which are used for the DSRC sentence selection process. Empirical validation performed on CNN/DailyMail data shows promising results of the proposed summarizer, which provides comparable performance with respect to ROUGE, BLEU, and METEOR scores while being computationally efficient. The results indicate that lightweight statistical summarization systems can achieve enhanced semantic quality when combined with adaptive redundancy-aware sentence selection strategies.

Article Details

Section
Articles