Publications des agents du Cirad


Mining Tweet Data - Statistic and semantic information for political tweet classification

Tisserant G., Roche M., Prince V.. 2014. In : Proceedings of the 6th International Conference on Knowledge Discovery and Information Retrieval : KDIR 2014, Rome, Italy, 21-24 October, 2014. s.l. : s.n., p. 523-529. International Conference on Knowledge Discovery and Information Retrieval. 6, 2014-10-21/2014-10-24, Rome (Italie).

This paper deals with the quality of textual features in messages in order to classify tweets. The aim of our study is to show how improving the representation of textual data affects the performance of learning algorithms. We will first introduce our method GENDESC. It generalizes less relevant words for tweet classification. Secondly, we compare and discuss the types of textual features given by different approaches. More precisely we discuss the semantic specificity of textual features, e.g. Named Entities, HashTags.

Documents associés

Communication de congrès

Agents Cirad, auteurs de cette publication :