Packt+ | Advance your knowledge in tech

0

Explore Products

Best Sellers

New Releases

Books

Videos

Audiobooks

Free Learning

Programming MapReduce with Scalding

You're reading from Programming MapReduce with Scalding A practical guide to designing, testing, and implementing complex MapReduce applications in Scala

Product type Paperback

Published in Jun 2014

Publisher

ISBN-13 9781783287017

Length 148 pages

Edition 1st Edition

Languages

Scala

Tools

Hadoop

Concepts

Front End Web Development

Author (1):

Antonios Chalkiopoulos

View More author details

Table of Contents (11) Chapters

Preface

1. Introduction to MapReduce

2. Get Ready for Scalding FREE CHAPTER

3. Scalding by Example

4. Intermediate Examples

5. Scalding Design Patterns

6. Testing and TDD

7. Running Scalding in Production

8. Using External Data Stores

9. Matrix Calculations and Machine Learning

Index

Setting a similarity using the Jaccard index

Quite often, we have to work with sets of data in machine learning. Users like posts, buy products, listen to music, or watch movies. In this case, data is structured in the two columns: 'user and 'item.

In order to calculate correlations, we need to work with sets. The Jaccard similarity coefficient is a statistic that measures the similarity between sets. The level of similarity is the calculation of the size of the intersection divided by the size of the union of the sample sets, as shown.

For example, if two users in the dataset are related to the same two items, and each user is also related to a distinct item, the Jaccard similarity indicates the following:

The similarity between item1 and item2 is 100 percent
The similarity between the common and distinct items is 50 percent
The similarity between two distinct items is 0 percent

To begin the implementation, we first need to calculate the item popularity, and then add the popularity back to the...

The rest of the chapter is locked

Register for a free Packt account to unlock a world of extra content!

A free Packt account unlocks extra newsletters, articles, discounted offers, and much more. Start advancing your knowledge today.

Unlock this book and the full library FREE for 7 days

Get unlimited access to 7000+ expert-authored eBooks and videos courses covering every tech area you can think of

Start free trial

Renews at €18.99/month. Cancel anytime

Authors (1)

Antonios Chalkiopoulos

Antonios Chalkiopoulos

Antonios Chalkiopoulos is a developer living in London and a professional working with Hadoop and Big Data technologies. He completed a number of complex MapReduce applications in Scalding into 40-plus production nodes HDFS Cluster. He is a contributor to Scalding and other open source projects, and he is interested in cloud technologies, NoSQL databases, distributed real-time computation systems, and machine learning. He was involved in a number of Big Data projects before discovering Scala and Scalding. Most of the content of this book comes from his experience and knowledge accumulated while working with a great team of engineers.

See other products by Antonios Chalkiopoulos

Personalised recommendations for you

Based on your interests and search pattern

Modern Full-Stack React Projects

Modern Full-Stack React Projects

Full-Stack React Projects is a complete guide to learning full-stack web development, understanding the creation and integration of backend systems, and advancing your career as a frontend developer.

Jun 2024 16h 52m

Mastering Node.js Web Development

Mastering Node.js Web Development

Explore Node.js with practical examples that will teach you how to utilize open-source packages for real-world solutions. Gain the skills to develop and deploy server-side applications that enhance your client-side projects.

Jun 2024 25h 56m

Reactive Patterns with RxJS and Angular Signals

Reactive Patterns with RxJS and Angular Signals

This RxJS book will help you understand the core concepts of RxJS and provide practical patterns to make your code more reactive and declarative. You'll also understand Angular Signals, which provide another way to improve code reactivity.

Jul 2024 8h 28m

API Testing and Development with Postman

API Testing and Development with Postman

Whether you are a tester or a developer working with APIs, you'll be able to put your knowledge to work with this practical guide to using Postman. The book provides a hands-on approach to implementing and learning the associated methodologies that will have you up-and-running and productive in no time.

Jun 2024 11h 56m

API Testing and Development with Postman

API Testing and Development with Postman

Whether you are a tester or a developer working with APIs, you'll be able to put your knowledge to work with this practical guide to using Postman. The book provides a hands-on approach to implementing and learning the associated methodologies that will have you up-and-running and productive in no time.

Jun 2024 11h 56m

FastAPI Cookbook

FastAPI Cookbook

This book helps you unlock the power of FastAPI to build high-performing web apps and APIs by taking you through the basics like routing and data validation through to advanced topics, such as custom middleware and WebSockets.

Aug 2024 11h 56m

Mastering Spring Boot 3.0

Mastering Spring Boot 3.0

This hands-on guide empowers you to develop scalable and efficient applications. You'll also learn microservices patterns, reactive programming, and security measures for building robust backend systems.

Jun 2024 8h 32m

Nuxt 3 Projects

Nuxt 3 Projects

This book is a comprehensive guide to Nuxt.js, which takes you from the basics to advanced topics. Uniquely, this book emphasizes practical, project-based learning, tackling real-world problems.

Jun 2024 7h 40m

Vue.js 3 for Beginners

Vue.js 3 for Beginners

Learning a new language by following video tutorials, blog posts, and documentation is a tiresome activity. This book will take you on an exciting journey of becoming a proficient Vue.js developer through a practical, step-by-step approach.

Sep 2024 10h 4m

Full-Stack Web Development with TypeScript 5

Full-Stack Web Development with TypeScript 5

The book emphasizes best practices, debugging, performance optimization, and scalable code structure, helping you develop practical skills in frontend and backend development, database integration, and AI integration.

Mastering Flask Web and API Development

Mastering Flask Web and API Development

The book is an introduction to Flask that will showcase its baseline, core, and advanced integration features to enable you to solve enterprise-related problems and issues in both web and API development.

Aug 2024 16h 28m