You're reading from Artificial Intelligence with Python Your complete guide to building intelligent apps using Python 3.x

Product type Paperback

Published in Jan 2020

Publisher Packt

ISBN-13 9781839219535

Length 618 pages

Edition 2nd Edition

Languages

Python

Tools

TensorFlow

Concepts

Artificial Intelligence

Authors (2):

Prateek Joshi

Alberto Artasanchez

View More author details

Table of Contents (26) Chapters

Preface

1. Introduction to Artificial Intelligence

2. Fundamental Use Cases for Artificial Intelligence FREE CHAPTER

3. Machine Learning Pipelines

4. Feature Selection and Feature Engineering

5. Classification and Regression Using Supervised Learning

6. Predictive Analytics with Ensemble Learning

7. Detecting Patterns with Unsupervised Learning

8. Building Recommender Systems

9. Logic Programming

10. Heuristic Search Techniques

11. Genetic Algorithms and Genetic Programming

12. Artificial Intelligence on the Cloud

13. Building Games with Artificial Intelligence

14. Building a Speech Recognizer

15. Natural Language Processing

16. Chatbots

17. Sequential Data and Time Series Analysis

18. Image Recognition

19. Neural Networks

20. Deep Learning with Convolutional Neural Networks

21. Recurrent Neural Networks and Other Deep Learning Models

22. Creating Intelligent Agents with Reinforcement Learning

23. Artificial Intelligence and Big Data

24. Other Books You May Enjoy

25. Index

Extracting the frequency of terms using the Bag of Words model

One of the main goals of text analysis with the Bag of Words model is to convert text into a numerical form so that we can use machine learning on it. Let's consider text documents that contain many millions of words. In order to analyze these documents, we need to extract the text and convert it into a form of numerical representation.

Machine learning algorithms need numerical data to work with so that they can analyze the data and extract meaningful information. This is where the Bag of Words model comes in. This model extracts vocabulary from all the words in the documents and builds a model using a document-term matrix. This allows us to represent every document as a bag of words. We just keep track of word counts and disregard the grammatical details and the word order.

Let's see what a document-term matrix is all about. A document-term matrix is basically a table that gives us counts...