Sie befinden Sich nicht im Netzwerk der Universität Paderborn. Der Zugriff auf elektronische Ressourcen ist gegebenenfalls nur via VPN oder Shibboleth (DFN-AAI) möglich. mehr Informationen...
Ergebnis 15 von 171
PeerJ. Computer science, 2022-04, Vol.8, p.e896-e896, Article e896
2022
Volltextzugriff (PDF)

Details

Autor(en) / Beteiligte
Titel
Multi-label emotion classification of Urdu tweets
Ist Teil von
  • PeerJ. Computer science, 2022-04, Vol.8, p.e896-e896, Article e896
Ort / Verlag
United States: PeerJ. Ltd
Erscheinungsjahr
2022
Quelle
EZB Free E-Journals
Beschreibungen/Notizen
  • Urdu is a widely used language in South Asia and worldwide. While there are similar datasets available in English, we created the first multi-label emotion dataset consisting of 6,043 tweets and six basic emotions in the Urdu Nastalíq script. A multi-label (ML) classification approach was adopted to detect emotions from Urdu. The morphological and syntactic structure of Urdu makes it a challenging problem for multi-label emotion detection. In this paper, we build a set of baseline classifiers such as machine learning algorithms (Random forest (RF), Decision tree (J48), Sequential minimal optimization (SMO), AdaBoostM1, and Bagging), deep-learning algorithms (Convolutional Neural Networks (1D-CNN), Long short-term memory (LSTM), and LSTM with CNN features) and transformer-based baseline (BERT). We used a combination of text representations: stylometric-based features, pre-trained word embedding, word-based n-grams, and character-based n-grams. The paper highlights the annotation guidelines, dataset characteristics and insights into different methodologies used for Urdu based emotion classification. We present our best results using micro-averaged F1, macro-averaged F1, accuracy, Hamming loss (HL) and exact match (EM) for all tested methods.

Weiterführende Literatur

Empfehlungen zum selben Thema automatisch vorgeschlagen von bX