Search icon CANCEL
Subscription
0
Cart icon
Cart
Close icon
You have no products in your basket yet
Save more on your purchases!
Savings automatically calculated. No voucher code required
Arrow left icon
All Products
Best Sellers
New Releases
Books
Videos
Audiobooks
Learning Hub
Newsletters
Free Learning
Arrow right icon
Arrow up icon
GO TO TOP
Artificial Intelligence for Big Data

You're reading from  Artificial Intelligence for Big Data

Product type Book
Published in May 2018
Publisher Packt
ISBN-13 9781788472173
Pages 384 pages
Edition 1st Edition
Languages
Authors (2):
Anand Deshpande Anand Deshpande
Profile icon Anand Deshpande
Manish Kumar Manish Kumar
Profile icon Manish Kumar
View More author details
Toc

Table of Contents (19) Chapters close

Title Page
Copyright and Credits
Packt Upsell
Contributors
Preface
1. Big Data and Artificial Intelligence Systems 2. Ontology for Big Data 3. Learning from Big Data 4. Neural Network for Big Data 5. Deep Big Data Analytics 6. Natural Language Processing 7. Fuzzy Systems 8. Genetic Programming 9. Swarm Intelligence 10. Reinforcement Learning 11. Cyber Security 12. Cognitive Computing 1. Other Books You May Enjoy Index

Text preprocessing


Preprocessing the data is the process of cleaning and preparing the text for classification and derivation of meaning. Since our data may have a lot of noise, uninformative parts, such as HTML tags, need to be eliminated or re-aligned. At the word level, there might be many words that do not make much impact on the overall semantic of the textual context. Text preprocessing involves a few steps, such as extraction, tokenization, stop words removal, text enrichment, and normalization with stemming and lemmatization. In addition to these, some of the basic and generic techniques that improve accuracy involve converting the text to lower case, removing numbers (based on the context), removing punctuation, stripping white spaces (sometimes these add to noise in the input signal), and eliminating the sparse terms that are infrequent terms in the document. In the subsequent sections, we'll analyze some of these techniques in detail.

Removing stop words

Stop words are words that...

lock icon The rest of the chapter is locked
Register for a free Packt account to unlock a world of extra content!
A free Packt account unlocks extra newsletters, articles, discounted offers, and much more. Start advancing your knowledge today.
Unlock this book and the full library FREE for 7 days
Get unlimited access to 7000+ expert-authored eBooks and videos courses covering every tech area you can think of
Renews at $15.99/month. Cancel anytime