Packt+ | Advance your knowledge in tech

You're reading from Practical Data Analysis For small businesses, analyzing the information contained in their data using open source technology could be game-changing. All you need is some basic programming and mathematical skills to do just that.

Product type Paperback

Published in Oct 2013

Publisher Packt

ISBN-13 9781783280995

Length 360 pages

Edition 1st Edition

Languages

Python

Tools

NLTK

Concepts

Data Analysis

Author (1):

Hector Cuesta

View More author details

Table of Contents (17) Chapters

Preface

1. Getting Started FREE CHAPTER

2. Working with Data

3. Data Visualization

4. Text Classification

5. Similarity-based Image Retrieval

6. Simulation of Stock Prices

7. Predicting Gold Prices

8. Working with Support Vector Machines

9. Modeling Infectious Disease with Cellular Automata

10. Working with Social Graphs

11. Sentiment Analysis of Twitter Data

12. Data Processing and Aggregation with MongoDB

13. Working with MapReduce

14. Online Data Analysis with IPython and Wakari

A. Setting Up the Infrastructure

Index

Learning and classification

When we want to automatically identify to which category a specific value (categorical value) belongs, we need to implement an algorithm that can predict the most likely category for the value, based on the previous data. This is called Classification. In the words of Tom Mitchell:

"How can we build computer systems that automatically improve with experience, and what are the fundamental laws that govern all learning processes?"

The keyword here is learning (supervised learning in this case), and also how to train an algorithm to identify categorical elements. The common examples are spam classification , speech recognition , search engines , computer vision , and language detection ; but there are a large number of applications for a classifier. We can find two kinds of problems in classification. The binary classification is where we have only two categories (spam or not spam) and multiclass classification is where many categories are involved (for example...

The rest of the chapter is locked

Tech Concepts

Programming languages

Tech Tools

Unlimited access to the largest independent learning library in tech of over 8,000 expert-authored tech books and videos.

Innovative learning tools, including AI book assistants, code context explainers, and text-to-speech.

50+ new titles added per month and exclusive early access to books as they are being written.

A free Packt account unlocks extra newsletters, articles, discounted offers, and much more. Start advancing your knowledge today.

Unlock this book and the full library FREE for 7 days

Get unlimited access to 7000+ expert-authored eBooks and videos courses covering every tech area you can think of

Start free trial

Renews at €18.99/month. Cancel anytime

Authors (1)

Hector Cuesta

Hector Cuesta is founder and Chief Data Scientist at Dataxios, a machine intelligence research company. Holds a BA in Informatics and a M.Sc. in Computer Science. He provides consulting services for data-driven product design with experience in a variety of industries including financial services, retail, fintech, e-learning and Human Resources. He is an enthusiast of Robotics in his spare time. You can follow him on Twitter at https://twitter.com/hmCuesta.

See other products by Hector Cuesta