Search icon CANCEL
Arrow left icon
Explore Products
Best Sellers
New Releases
Books
Videos
Audiobooks
Learning Hub
Conferences
Free Learning
Arrow right icon
Arrow up icon
GO TO TOP
Mastering Java for Data Science

You're reading from   Mastering Java for Data Science Analytics and more for production-ready applications

Arrow left icon
Product type Paperback
Published in Apr 2017
Publisher Packt
ISBN-13 9781782174271
Length 364 pages
Edition 1st Edition
Languages
Tools
Arrow right icon
Author (1):
Arrow left icon
Alexey Grigorev Alexey Grigorev
Author Profile Icon Alexey Grigorev
Alexey Grigorev
Arrow right icon
View More author details
Toc

Search engine - preparing data

In the first chapter, we introduced the running example, building a search engine. A search engine is a program that, given a query from the user, returns results ordered by relevance with respect to the query. In this chapter, we will perform the first steps--obtaining and processing data.

Suppose we are working on a web portal where users generate a lot of content, but they have trouble finding what other people have created. To overcome this problem, we propose to build a search engine, and product management has identified the typical queries that the users will put in.

For example, "Chinese food", "homemade pizza", and "how to learn programming" are typical queries from this list.

Now we need to collect the data. Luckily for us, there are already search engines on the Internet that can take in a query and return a list of URLs they consider relevant...

You have been reading a chapter from
Mastering Java for Data Science
Published in: Apr 2017
Publisher: Packt
ISBN-13: 9781782174271
Register for a free Packt account to unlock a world of extra content!
A free Packt account unlocks extra newsletters, articles, discounted offers, and much more. Start advancing your knowledge today.
Unlock this book and the full library FREE for 7 days
Get unlimited access to 7000+ expert-authored eBooks and videos courses covering every tech area you can think of
Renews at $19.99/month. Cancel anytime