-
Automatic Labelling of Topics with Neural Embeddings
Abstract: Topics generated by topic models are typically represented as list of terms. To reduce the cognitive overhead of interpreting these topics for end-users, we propose labelling a topic with a succinct phrase that summarises its theme or idea. Using Wikipedia document titles as label candidates, we compute neural embeddings for documents and words to select the most relevant labels for topics. Compar… ▽ More
Submitted 22 December, 2016; v1 submitted 15 December, 2016; originally announced December 2016.
Comments: 11 pages, 3 figures, published in COLING2016
Journal ref: Proceedings of the 26th International Conference on Computational Linguistics (COLING 2016), pp 953--963
-
An Empirical Evaluation of doc2vec with Practical Insights into Document Embedding Generation
Abstract: Recently, Le and Mikolov (2014) proposed doc2vec as an extension to word2vec (Mikolov et al., 2013a) to learn document-level embeddings. Despite promising results in the original paper, others have struggled to reproduce those results. This paper presents a rigorous empirical evaluation of doc2vec over two tasks. We compare doc2vec to two baselines and two state-of-the-art document embedding metho… ▽ More
Submitted 18 July, 2016; originally announced July 2016.
Comments: 1st Workshop on Representation Learning for NLP
Journal ref: Proceedings of the 1st Workshop on Representation Learning for NLP, Berlin, Germany, pp. 78--86