0

Explore Products

Best Sellers

New Releases

Books

Videos

Audiobooks

Free Learning

Natural Language Processing and Computational Linguistics

You're reading from Natural Language Processing and Computational Linguistics A practical guide to text analysis with Python, Gensim, spaCy, and Keras

Product type Paperback

Published in Jun 2018

Publisher Packt

ISBN-13 9781788838535

Length 306 pages

Edition 1st Edition

Languages

Processing

Tools

Keras

Concepts

Mobile Application Development

Author (1):

Bhargav Srinivasa-Desikan

View More author details

Table of Contents (17) Chapters

Preface

1. What is Text Analysis? FREE CHAPTER

2. Python Tips for Text Analysis

3. spaCy's Language Models

4. Gensim – Vectorizing Text and Transformations and n-grams

5. POS-Tagging and Its Applications

6. NER-Tagging and Its Applications

7. Dependency Parsing

8. Topic Models

9. Advanced Topic Modeling

10. Clustering and Classifying Text

11. Similarity Queries and Summarization

12. Word2Vec, Doc2Vec, and Gensim

13. Deep Learning for Text

14. Keras and spaCy for Deep Learning

15. Sentiment Analysis and ChatBots

16. Other Books You May Enjoy

Leave a review - let other readers know what you think

Tokenizing text

You can see that the first step in this pipeline is tokenizing – what exactly is this?

Tokenization is the task of splitting a text into meaningful segments, called tokens. These segments could be words, punctuation, numbers, or other special characters that are the building blocks of a sentence. In spaCy, the input to the tokenizer is a Unicode text, and the output is a Doc object [19].

Different languages will have different tokenization rules. Let's look at an example of how tokenization might work in English. For the sentence – Let us go to the park., it's quite straightforward, and would be broken up as follows, with the appropriate numerical indices:

0	1	2	3	4	5	6
Let	us	go	to	the	park	.

This looks awfully like the result when we just run text.split(' ') – when does tokenizing...

The rest of the chapter is locked

Register for a free Packt account to unlock a world of extra content!

A free Packt account unlocks extra newsletters, articles, discounted offers, and much more. Start advancing your knowledge today.

Unlock this book and the full library FREE for 7 days

Get unlimited access to 7000+ expert-authored eBooks and videos courses covering every tech area you can think of

Start free trial

Renews at €18.99/month. Cancel anytime

Authors (1)

Bhargav Srinivasa-Desikan

Bhargav Srinivasa-Desikan

Bhargav Srinivasa-Desikan is a research engineer working for INRIA in Lille, France. He is a part of the MODAL (Models of Data Analysis and Learning) team, and he works on metric learning, predictor aggregation, and data visualization. He is a regular contributor to the Python open source community, and completed Google Summer of Code in 2016 with Gensim where he implemented Dynamic Topic Models. He is a regular speaker at PyCons and PyDatas across Europe and Asia, and conducts tutorials on text analysis using Python.

See other products by Bhargav Srinivasa-Desikan

Other recommended products

Related to this chapter

Mastering spaCy

Mastering spaCy

Using machine learning-based NLP models, you can speed up business processes, make more accurate predictions, and uncover new insights from your existing data, where spaCy, an advanced industrial-grade natural language processing library, can help. With this book, you'll learn how to use it and create high-impact ML solutions for NLP.

Jul 2021 11h 52m

Python Natural Language Processing Cookbook

Python Natural Language Processing Cookbook

Leverage your natural language processing skills to make sense of text. With this book, you'll learn fundamental and advanced NLP techniques in Python that will help you to make your data fit for application in a wide variety of industries. You'll also find recipes for overcoming common challenges in implementing NLP pipelines.

Mar 2021 9h 28m

Natural Language Processing Fundamentals

Natural Language Processing Fundamentals

Natural Language Processing Fundamentals starts with basics and goes on to explain various NLP tools and techniques that equip you with all that you need to solve common business problems for processing text.

Mar 2019 12h 28m

Natural Language Processing with Python Quick Start Guide

Natural Language Processing with Python Quick Start Guide

NLP in Python is among the most sought-after skills among data scientists. With code and relevant case studies, this book will show how you can use industry grade tools to implement NLP programs capable of learning from relevant data. We will explore many modern methods ranging from spaCy to word vectors that have reinvented NLP.

fastText Quick Start Guide

fastText Quick Start Guide

Facebook's fastText library handles text representation and classification, used for Natural Language Processing (NLP). Most organizations have to deal with enormous amounts of text data on a daily basis, and efficient data insights requires powerful NLP tools like fastText. This book is your ideal introduction to fastText.

Jul 2018 6h 28m

Hands-On Natural Language Processing with PyTorch 1.x

Hands-On Natural Language Processing with PyTorch 1.x

Developers working with NLP will be able to put their knowledge to work with this practical guide to PyTorch. You will learn to use PyTorch offerings and how to understand and analyze text using Python. You will learn to extract the underlying meaning in the text using deep neural networks and modern deep learning algorithms.

Jul 2020 9h 12m

Natural Language Processing with Java

Natural Language Processing with Java

Natural Language Processing with Java will explore how to automatically organize text using approaches such as full-text search, proper name recognition, clustering, tagging, information extraction, and summarization. You will leverage the power of Java to extract relationships within different elements of text and documents.

Jul 2018 10h 36m

Hands-On Python Natural Language Processing

Hands-On Python Natural Language Processing

This book provides a blend of both the theoretical and practical aspects of Natural Language Processing (NLP). It covers the concepts essential to develop a thorough understanding of NLP and also delves into a detailed discussion on NLP based use-cases such as language translation, sentiment analysis, etc. Every module covers real-world examples

Jun 2020 10h 32m

The Natural Language Processing Workshop

The Natural Language Processing Workshop

The Natural Language Processing Workshop takes you through fundamental NLP techniques, such as preparing datasets, collecting text, extracting text, and sentiment analysis. As you progress, you'll get to grips with creating your own chatbots and dynamic models.

Aug 2020 15h 4m

Python Artificial Intelligence Projects for Beginners

Python Artificial Intelligence Projects for Beginners

This book demonstrates AI projects in Python covering modern techniques that make up the world of Artificial Intelligence. You will come across a variety of real-world projects on classifying data, text processing techniques, deep learning and neural networks

Jul 2018 5h 24m

Hands-On Natural Language Processing with Python

Hands-On Natural Language Processing with Python

This book teaches you to leverage deep learning models in performing various NLP tasks along with showcasing the best practices in dealing with the NLP challenges. The book equips you with practical knowledge to implement deep learning in your linguistic applications using NLTk and Python's popular deep learning library, TensorFlow.

Jul 2018 10h 24m

Python Natural Language Processing

Python Natural Language Processing

Natural Language Processing is a field of computational linguistics and artificial intelligence that deals with human-computer interaction. The numbers of human-computer interaction instances are increasing so it's becoming imperative that computers comprehend all major natural languages. Python's powerful tools and libraries are evolved so much that natural language processing becomes much simpler and accurate with it. This book will get you up and running with Python's library for Natural Language Processing-- NLTK-- in no time.

Jul 2017 16h 12m