ebook img

Data science algorithms in a week: data analysis, machine learning, and more PDF

205 Pages·2017·7.71 MB·English
Save to my drive
Quick download
Download
Most books are stored in the elastic cloud where traffic is expensive. For this reason, we have a limit on daily download.

Preview Data science algorithms in a week: data analysis, machine learning, and more

Data Science Algorithms in a Week Data analysis, machine learning, and more Dávid Natingga BIRMINGHAM - MUMBAI Data Science Algorithms in a Week Copyright © 2017 Packt Publishing All rights reserved. No part of this book may be reproduced, stored in a retrieval system, or transmitted in any form or by any means, without the prior written permission of the publisher, except in the case of brief quotations embedded in critical articles or reviews. Every effort has been made in the preparation of this book to ensure the accuracy of the information presented. However, the information contained in this book is sold without warranty, either express or implied. Neither the author, nor Packt Publishing, and its dealers and distributors will be held liable for any damages caused or alleged to be caused directly or indirectly by this book. Packt Publishing has endeavored to provide trademark information about all of the companies and products mentioned in this book by the appropriate use of capitals. However, Packt Publishing cannot guarantee the accuracy of this information. First published: August 2017 Production reference: 1080817 Published by Packt Publishing Ltd. Livery Place 35 Livery Street Birmingham B3 2PB, UK. ISBN 978-1-78728-458-6 www.packtpub.com Credits Author Copy Editor Dávid Natingga Safis Editing Reviewer Project Coordinator Surendra Pepakayala Kinjal Bari Commissioning Editor Proofreader Veena Pagare Safis Editing Acquisition Editor Indexer Chandan Kumar Pratik Shirodkar Content Development Editor Production Coordinator Mamata Walkar Shantanu Zagade Technical Editor Naveenkumar Jain About the Author Dávid Natingga graduated in 2014 from Imperial College London in MEng Computing with a specialization in Artificial Intelligence. In 2011, he worked at Infosys Labs in Bangalore, India, researching the optimization of machine learning algorithms. In 2012 and 2013, at Palantir Technologies in Palo Alto, USA, he developed algorithms for big data. In 2014, as a data scientist at Pact Coffee, London, UK, he created an algorithm suggesting products based on the taste preferences of customers and the structure of coffees. In 2017, he work at TomTom in Amsterdam, Netherlands, processing map data for navigation platforms. As a part of his journey to use pure mathematics to advance the field of AI, he is a PhD candidate in Computability Theory at, University of Leeds, UK. In 2016, he spent 8 months at Japan, Advanced Institute of Science and Technology, Japan, as a research visitor. Dávid Natingga married his wife Rheslyn and their first child will soon behold the outer world. I would like to thank Packt Publishing for providing me with this opportunity to share my knowledge and experience in data science through this book. My gratitude belongs to my wife Rheslyn who has been patient, loving, and supportive through out the whole process of writing this book. About the Reviewer Surendra Pepakayala is a seasoned technology professional and entrepreneur with over 19 years of experience in the US and India. He has broad experience in building enterprise/web software products as a developer, architect, software engineering manager, and product manager at both start-ups and multinational companies in India and the US. He is a hands- on technologist/hacker with deep interest and expertise in Enterprise/Web Applications Development, Cloud Computing, Big Data, Data Science, Deep Learning, and Artificial Intelligence. A technologist turned entrepreneur, after 11 years in corporate US, Surendra has founded an enterprise BI / DSS product for school districts in the US. He subsequently sold the company and started a Cloud Computing, Big Data, and Data Science consulting practice to help start-ups and IT organizations streamline their development efforts and reduce time to market of their products/solutions. Also, Surendra takes pride in using his considerable IT experience for reviving / turning-around distressed products / projects. He serves as an advisor to eTeki, an on-demand interviewing platform, where he leads the effort to recruit and retain world-class IT professionals into eTeki’s interviewer panel. He has reviewed drafts, recommended changes and formulated questions for various IT certifications such as CGEIT, CRISC, MSP, and TOGAF. His current focus is on applying Deep Learning to various stages of the recruiting process to help HR (staffing and corporate recruiters) find the best talent and reduce friction involved in the hiring process. www.PacktPub.com For support files and downloads related to your book, please visit www.PacktPub.com. Did you know that Packt offers eBook versions of every book published, with PDF and ePub files available? You can upgrade to the eBook version at www.PacktPub.com and as a print book customer, you are entitled to a discount on the eBook copy. Get in touch with us at [email protected] for more details. At www.PacktPub.com, you can also read a collection of free technical articles, sign up for a range of free newsletters and receive exclusive discounts and offers on Packt books and eBooks. https:/​/​www.​packtpub.​com/​mapt Get the most in-demand software skills with Mapt. Mapt gives you full access to all Packt books and video courses, as well as industry-leading tools to help you plan your personal development and advance your career. Why subscribe? Fully searchable across every book published by Packt Copy and paste, print, and bookmark content On demand and accessible via a web browser Customer Feedback Thanks for purchasing this Packt book. At Packt, quality is at the heart of our editorial process. To help us improve, please leave us an honest review on this book's Amazon page at link. If you'd like to join our team of regular reviewers, you can e-mail us at [email protected]. We award our regular reviewers with free eBooks and videos in exchange for their valuable feedback. Help us be relentless in improving our products! Table of Contents Preface 1 Chapter 1: Classification Using K Nearest Neighbors 6 Mary and her temperature preferences 6 Implementation of k-nearest neighbors algorithm 10 Map of Italy example - choosing the value of k 15 House ownership - data rescaling 18 Text classification - using non-Euclidean distances 20 Text classification - k-NN in higher-dimensions 23 Summary 25 Problems 25 Chapter 2: Naive Bayes 29 Medical test - basic application of Bayes' theorem 30 Proof of Bayes' theorem and its extension 31 Extended Bayes' theorem 32 Playing chess - independent events 33 Implementation of naive Bayes classifier 34 Playing chess - dependent events 37 Gender classification - Bayes for continuous random variables 40 Summary 42 Problems 43 Chapter 3: Decision Trees 51 Swim preference - representing data with decision tree 52 Information theory 53 Information entropy 53 Coin flipping 54 Definition of information entropy 54 Information gain 55 Swim preference - information gain calculation 55 ID3 algorithm - decision tree construction 57 Swim preference - decision tree construction by ID3 algorithm 57 Implementation 58 Classifying with a decision tree 64 Classifying a data sample with the swimming preference decision tree 65 Playing chess - analysis with decision tree 65 Going shopping - dealing with data inconsistency 69 Summary 70 Problems 71 Chapter 4: Random Forest 75 Overview of random forest algorithm 76 Overview of random forest construction 76 Swim preference - analysis with random forest 77 Random forest construction 78 Construction of random decision tree number 0 78 Construction of random decision tree number 1 80 Classification with random forest 83 Implementation of random forest algorithm 83 Playing chess example 86 Random forest construction 88 Construction of a random decision tree number 0: 88 Construction of a random decision tree number 1, 2, 3 92 Going shopping - overcoming data inconsistency with randomness and measuring the level of confidence 94 Summary 96 Problems 97 Chapter 5: Clustering into K Clusters 102 Household incomes - clustering into k clusters 102 K-means clustering algorithm 103 Picking the initial k-centroids 104 Computing a centroid of a given cluster 104 k-means clustering algorithm on household income example 104 Gender classification - clustering to classify 105 Implementation of the k-means clustering algorithm 109 Input data from gender classification 112 Program output for gender classification data 112 House ownership – choosing the number of clusters 113 Document clustering – understanding the number of clusters k in a semantic context 119 Summary 126 Problems 126 Chapter 6: Regression 135 Fahrenheit and Celsius conversion - linear regression on perfect data 136 Weight prediction from height - linear regression on real-world data 139 [ ]

See more

The list of books you might like

Most books are stored in the elastic cloud where traffic is expensive. For this reason, we have a limit on daily download.