Cloudera Administration Handbook A complete, hands-on guide to building and maintaining large Apache Hadoop clusters using Cloudera Manager and CDH5 Rohit Menon BIRMINGHAM - MUMBAI Cloudera Administration Handbook Copyright © 2014 Packt Publishing All rights reserved. No part of this book may be reproduced, stored in a retrieval system, or transmitted in any form or by any means, without the prior written permission of the publisher, except in the case of brief quotations embedded in critical articles or reviews. Every effort has been made in the preparation of this book to ensure the accuracy of the information presented. However, the information contained in this book is sold without warranty, either express or implied. Neither the author, nor Packt Publishing, and its dealers and distributors will be held liable for any damages caused or alleged to be caused directly or indirectly by this book. Packt Publishing has endeavored to provide trademark information about all of the companies and products mentioned in this book by the appropriate use of capitals. However, Packt Publishing cannot guarantee the accuracy of this information. First published: July 2014 Production reference: 1110714 Published by Packt Publishing Ltd. Livery Place 35 Livery Street Birmingham B3 2PB, UK. ISBN 978-1-78355-896-4 www.packtpub.com Cover image by John Michael Harkness ([email protected]) Credits Author Project Coordinators Rohit Menon Swati Kumari Amey Sawant Reviewers Skanda Bhargav Proofreaders Brandon Forehand Simran Bhogal Mike Hordila Ameesha Green Maria Gould Commissioning Editor Akram Hussain Indexer Rekha Nair Acquisition Editor Gregory Wild Graphics Disha Haria Content Development Editor Priya Singh Production Coordinator Nitesh Thakur Technical Editors Kunal Anil Gaikwad Cover Work Edwin Moses Nitesh Thakur Siddhi Rane Copy Editors Janbal Dharmaraj Deepa Nambiar Alfida Paiva Laxmi Subramanian Notice Packt Publishing has endeavored to provide trademark information about all of the companies and products mentioned in this book by the appropriate use of capitals. However, Packt Publishing cannot guarantee the accuracy of this information. CLOUDERA® is a registered trademark of Cloudera, Inc. Except where otherwise indicated, all screenshots are the copyrighted material of Cloudera, Inc. For the latest documentation on use of Cloudera software, visit http://www.cloudera.com/. About the Author Rohit Menon is a senior system analyst living in Denver, Colorado. He has over 7 years of experience in the field of Information Technology, which started with the role of a real-time applications developer back in 2006. He now works for a product-based company specializing in software for large telecom operators. He graduated with a master's degree in Computer Applications from University of Pune, where he built an autonomous maze-solving robot as his final year project. He later joined a software consulting company in India where he worked on C#, SQL Server, C++, and RTOS to provide software solutions to reputable organizations in USA and Japan. After this, he started working for a product-based company where most of his time was dedicated to programming the finer details of products using C++, Oracle, Linux, and Java. He is a person who always likes to learn new technologies and this got him interested in web application development. He picked up Ruby, Ruby on Rails, HTML, JavaScript, CSS, and built www.flicksery.com, a Netflix search engine that makes searching for titles on Netflix much easier. On the Hadoop front, he is a Cloudera Certified Apache Hadoop Developer. He blogs at www.rohitmenon.com, mainly on topics related to Apache Hadoop and its components. To share his learning, he has also started www.hadoopscreencasts. com, a website that teaches Apache Hadoop using simple, short, and easy-to-follow screencasts. He is well versed with wide variety of tools and techniques such as MapReduce, Hive, Pig, Sqoop, Oozie, and Talend Open Studio. I would like to thank my parents for instilling the qualities of perseverance and hard work. I would also like to thank my wife, Madhuri, and my daughter, Anushka, for being patient and allowing me to spend most of my time studying and researching. About the Reviewers Skanda Bhargav is an engineering graduate from Visvesvaraya Technological University (VTU), Belgaum in Karnataka, India. He did his majors in Computer Science Engineering. He is currently employed with Happiest Minds Technologies, a MNC based out of Bangalore. He is a Cloudera Certified Developer for Apache Hadoop. His interests are Big Data and Hadoop. He has been a reviewer for the following books: • Instant MapReduce Patterns – Hadoop Essentials How-to, Srinath Perera, Packt Publishing • Hadoop Cluster Deployment, Danil Zburivsky, Packt Publishing He has also reviewed Building Hadoop Clusters [Video], Sean Mikha, Packt Publishing. I would like to thank my family for their immense support and faith in me throughout my learning stage. My friends have brought the confidence in me to a level that makes me bring out the best out of myself. I am happy that God has blessed me with such wonderful people around me, without which this work might not have been the success that it is today. Brandon Forehand started programming at an early age and loves solving problems. He is a Cloudera Certified Apache Hadoop Developer and currently works at Moz as a principal software engineer on the Big Data team, developing systems to index links on the web and providing data to help online marketers improve their websites' visibility. Previously, he worked at Amazon on Kindle and developed software to convert physical books to e-books. He has also worked at a research laboratory, developing sonar systems for the Navy. He earned a BSc in Computer Science from the University of Texas, Austin. I would like to thank my wife for putting up with me all of these years and the countless people who have helped me along the way in my career. Mike Hordila has worked with very large databases and high availability systems for more than 20 years. He consults for major organizations, always looking for new ways and technologies. He has shared some of his experience in a number of articles in major Oracle magazines and also in a couple of books. www.PacktPub.com Support files, eBooks, discount offers, and more You might want to visit www.PacktPub.com for support files and downloads related to your book. Did you know that Packt offers eBook versions of every book published, with PDF and ePub files available? You can upgrade to the eBook version at www.PacktPub.com and as a print book customer, you are entitled to a discount on the eBook copy. Get in touch with us at [email protected] for more details. At www.PacktPub.com, you can also read a collection of free technical articles, sign up for a range of free newsletters, and receive exclusive discounts and offers on Packt books and eBooks. TM http://PacktLib.PacktPub.com Do you need instant solutions to your IT questions? PacktLib is Packt's online digital book library. Here, you can access, read, and search across Packt's entire library of books. Why subscribe? • Fully searchable across every book published by Packt • Copy and paste, print, and bookmark content • On demand and accessible via web browser Free access for Packt account holders If you have an account with Packt at www.PacktPub.com, you can use this to access PacktLib today and view nine entirely free books. Simply use your login credentials for immediate access. Table of Contents Preface 1 Chapter 1: Getting Started with Apache Hadoop 7 History of Apache Hadoop and its trends 7 Components of Apache Hadoop 10 Understanding the Apache Hadoop daemons 10 Namenode 11 Secondary namenode 12 Jobtracker 14 Tasktracker 14 ResourceManager 17 NodeManager 17 Job submission in YARN 17 Introducing Cloudera 19 Introducing CDH 19 Responsibilities of a Hadoop administrator 20 Summary 22 Chapter 2: HDFS and MapReduce 23 Essentials of HDFS 23 Configuring HDFS 24 The read/write operational flow in HDFS 26 Writing files in HDFS 26 Reading files in HDFS 28 Understanding the namenode UI 29 Understanding the secondary namenode UI 33 Exploring HDFS commands 34 Commonly used HDFS commands 34 Commands to administer HDFS 37