Discovering underlying topics in newsgroups
A topic model is a type of statistical model for discovering the probability distributions of words linked to the topic. The topic in topic modeling does not exactly match the dictionary definition but corresponds to a nebulous statistical concept, which is an abstraction that occurs in a collection of documents.
When we read a document, we expect certain words appearing in the title or the body of the text to capture the semantic context of the document. An article about Python programming might have words such as class and function, while a story about snakes might have words such as eggs and afraid. Documents usually have multiple topics; for instance, this section is about three things: topic modeling, non-negative matrix factorization, and latent Dirichlet allocation, which we will discuss shortly. We can therefore define an additive model for topics by assigning different weights to topics.
Topic modeling is widely used for...