Subscription

Explore Products

Best Sellers

New Releases

Books

Videos

Audiobooks

Learning Hub

Conferences

Free Learning

You're reading from Python Natural Language Processing Advanced machine learning and deep learning techniques for natural language processing

Product type Paperback

Published in Jul 2017

Publisher Packt

ISBN-13 9781787121423

Length 486 pages

Edition 1st Edition

Languages

Processing

Tools

Processing

Concepts

Artificial Intelligence

Author (1):

Jalaj Thanaki

View More author details

Table of Contents (13) Chapters

Preface

1. Introduction FREE CHAPTER

2. Practical Understanding of a Corpus and Dataset

3. Understanding the Structure of a Sentences

4. Preprocessing

5. Feature Engineering and NLP Algorithms

6. Advanced Feature Engineering and NLP Algorithms

7. Rule-Based System for NLP

8. Machine Learning for NLP Problems

9. Deep Learning for NLU and NLG Problems

10. Advanced Tools

11. How to Improve Your NLP Skills

12. Installation Guide

Apache Spark as a processing framework

Apache Spark is a large-scale data processing framework. It is a fast and general-purpose engine. It is one of the fastest processing frameworks. Spark can perform in-memory data processing, as well as on-disk data processing.

Spark's important features are as follows:

Speed: Apache Spark can run programs up to 100 times faster than Hadoop MapReduce in-memory or 10 times faster on-disk
Ease of use: There are various APIs available for Scala, Java, Spark, and R to develop your application
Generality: Spark provides features of Combine SQL, streaming, and complex analytics
Run everywhere: Spark can run on Hadoop, Mesos, standalone, or in the cloud. You can access diverse data sources by including HDFS, Cassandra, HBase, and S3

I have used Spark to train my models using MLlib. I have used Spark Java as well as PySpark API. The result...

The rest of the chapter is locked

A free Packt account unlocks extra newsletters, articles, discounted offers, and much more. Start advancing your knowledge today.

Unlock this book and the full library FREE for 7 days

Get unlimited access to 7000+ expert-authored eBooks and videos courses covering every tech area you can think of

Start free trial

Renews at $19.99/month. Cancel anytime

Authors (1)

Jalaj Thanaki

Jalaj Thanaki is an experienced data scientist with a demonstrated history of working in the information technology, publishing, and finance industries. She is author of the book Python Natural Language Processing, Packt publishing. Her research interest lies in Natural Language Processing, Machine Learning, Deep Learning, and Big Data Analytics. Besides being a data scientist, Jalaj is also a social activist, traveler, and nature-lover.

See other products by Jalaj Thanaki