Fasttext Deep Recurrent Network And Optimized Densenet Based Aggregated Max Fusion For Multimodal Sarcasm Detection In Local Self-Government Social Media Platforms

Authors

  • Mr. K. Sathyaseelan, Dr. N. Yuvaraj

DOI:

https://doi.org/10.52152/1emrrd36

Keywords:

Multimodal sarcasm detection, multimodal fusion, Deep Recurrent Neural Network, word embedding, DenseNet.

Abstract

Multimodal sarcasm detection is a complex but promising area of research and application, leveraging diverse data types to improve understanding and interpretation of sarcastic expressions. Traditional sarcasm detection often relies solely on textual analysis, which can miss the nuances conveyed through tone of voice or facial expressions. Deep learning models, particularly those leveraging architectures like convolutional neural networks (CNNs) and recurrent neural networks (RNNs), can effectively process and integrate these multiple data modalities. By analyzing patterns across different inputs, these models can improve the accuracy of sarcasm detection, capturing the subtleties that indicate sarcasm, such as intonation or context-specific gestures. In this paper, Pearson Correlated FastText and Bray Curtis Deep Recurrent Neural Network (PCFT-BDRNN) model is proposed for detecting sarcasm using text data on social media. Initially, number of text data from the given dataset is collected as input. After that, text preprocessing using Gensim tokenizer, stop-word filtering, word stemming and Lemmatization is carried out. Subsequently, FastText word embedding is performed to determine the semantic information from the preprocessed results. Later, classification of text is achieved by designing Bray Curtis Deep Recurrent Perceptive Neural Learning with better accuracy. Optimized Canonical DenseNet (OP-DenseNet) Model is proposed for sarcasm detection using images on social media. The designed OP-DenseNet model includes multiple layers such as input layer, convolutional layers, global average pooling layer, fully connected layer and output layer. The input layer gets number of images collected from the given dataset.  The multiple convolutional layers perform texture, color, intensity, standard deviation and edge feature extraction. Followed by this, spatial dimensions of the images are minimized using global average pooling layer. In addition, the classification of images is made in fully connected layer using Canonical variate analysis. ReLU activation function provides classification results at the output layer for sarcasm detection. Lastly, multimodal fusion is developed using aggregated max fusion model for improving the sarcasm detection where it combines both the results of text and image classifiers. Experimental evaluation is carried out using Multi-modal sarcasm detection Dataset using different performance metrics. The results show that the proposed models significantly improve the accuracy, precision, recall with minimum time than the state-of-the-art methods.

Downloads

Published

2026-06-15

Issue

Section

Article

How to Cite

Fasttext Deep Recurrent Network And Optimized Densenet Based Aggregated Max Fusion For Multimodal Sarcasm Detection In Local Self-Government Social Media Platforms. (2026). Lex Localis - Journal of Local Self-Government, 23-48. https://doi.org/10.52152/1emrrd36