Social Network Analysis for Startups Maksim Tsvetovat and Alexander Kouznetsov Beijing • Cambridge • Farnham • Köln • Sebastopol • Tokyo Social Network Analysis for Startups by Maksim Tsvetovat and Alexander Kouznetsov Copyright © 2011 Maksim Tsvetovat and Alexander Kouznetsov. All rights reserved. Printed in the United States of America. Published by O’Reilly Media, Inc., 1005 Gravenstein Highway North, Sebastopol, CA 95472. O’Reilly books may be purchased for educational, business, or sales promotional use. Online editions are also available for most titles (http://my.safaribooksonline.com). For more information, contact our corporate/institutional sales department: (800) 998-9938 or [email protected]. Editors: Shawn Wallace and Mike Hendrickson Cover Designer: Karen Montgomery Production Editor: Kristen Borg Interior Designer: David Futato Proofreader: O’Reilly Production Services Illustrator: Robert Romano Revision History for the First Edition: See http://oreilly.com/catalog/errata.csp?isbn=9781449306465 for release details. Nutshell Handbook, the Nutshell Handbook logo, and the O’Reilly logo are registered trademarks of O’Reilly Media, Inc. Social Network Analysis for Startups, the image of a hawfinch, and related trade dress are trademarks of O’Reilly Media, Inc. Many of the designations used by manufacturers and sellers to distinguish their products are claimed as trademarks. Where those designations appear in this book, and O’Reilly Media, Inc., was aware of a trademark claim, the designations have been printed in caps or initial caps. While every precaution has been taken in the preparation of this book, the publisher and authors assume no responsibility for errors or omissions, or for damages resulting from the use of the information con- tained herein. ISBN: 978-1-449-30646-5 [LSI] 1316789838 Table of Contents Preface ..................................................................... vii 1. Introduction ........................................................... 1 Analyzing Relationships to Understand People and Groups 2 Binary and Valued Relationships 2 Symmetric and Asymmetric Relationships 3 Multimode Relationships 3 From Relationships to Networks—More Than Meets the Eye 3 Social Networks vs. Link Analysis 4 The Power of Informal Networks 6 Terrorists and Revolutionaries: The Power of Social Networks 10 Social Networks in Prison 10 Informal Networks in Terrorist Cells 11 The Revolution Will Be Tweeted 14 2. Graph Theory—A Quick Introduction ..................................... 19 What Is a Graph? 19 Adjacency Matrices 21 Edge-Lists and Adjacency Lists 22 7 Bridges of Königsberg 23 Graph Traversals and Distances 25 Depth-First Traversal 27 Breadth-First Traversal 30 Paths and Walks 31 Dijkstra’s Algorithm 33 Graph Distance 35 Graph Diameter 36 Why This Matters 36 6 Degrees of Separation is a Myth! 37 Small World Networks 37 iii 3. Centrality, Power, and Bottlenecks ....................................... 39 Sample Data: The Russians are Coming! 39 Get Oriented in Python and NetworkX 39 Read Nodes and Edges from LiveJournal 41 Snowball Sampling 43 Saving and Loading a Sample Dataset from a File 44 Centrality 45 Who Is More Important in this Network? 45 Find the “Celebrities” 45 Find the Gossipmongers 49 Find the Communication Bottlenecks and/or Community Bridges 51 Putting It Together 54 Who Is a “Gray Cardinal?” 55 Klout Score 57 PageRank—How Google Measures Centrality 58 What Can’t Centrality Metrics Tell Us? 60 4. Cliques, Clusters and Components ........................................ 61 Components and Subgraphs 61 Analyzing Components with Python 62 Islands in the Net 62 Subgraphs—Ego Networks 65 Extracting and Visualizing Ego Networks with Python 65 Triads 67 Fraternity Study—Tie Stability and Triads 68 Triads and Terrorists 68 The “Forbidden Triad” and Structural Holes 72 Structural Holes and Boundary Spanning 73 Triads in Politics 74 Directed Triads 76 Analyzing Triads in Real Networks 77 Real Data 79 Cliques 79 Detecting Cliques 79 Hierarchical Clustering 81 The Algorithm 82 Clustering Cities 83 Preparing Data and Clustering 84 Block Models 86 Triads, Network Density, and Conflict 88 5. 2-Mode Networks ...................................................... 93 Does Campaign Finance Influence Elections? 93 iv | Table of Contents Theory of 2-Mode Networks 96 Affiliation Networks 96 Attribute Networks 98 A Little Math 98 2-Mode Networks in Practice 100 PAC Networks 102 Candidate Networks 102 Expanding Multimode Networks 105 Exercise 107 6. Going Viral! Information Diffusion ....................................... 109 Anatomy of a Viral Video 109 What Did Facebook Do Right? 110 How Do You Estimate Critical Mass? 111 Wikinomics of Critical Mass 112 Content is (Still) King 113 How Does Information Shape Networks (and Vice Versa)? 116 Birds of a Feather? 117 Homophily vs. Curiosity 117 Weak Ties 119 Dunbar Number and Weak Ties 119 A Simple Dynamic Model in Python 121 Influencers in the Midst 125 Exercises for the Reader 127 Coevolution of Networks and Information 127 Exercises for the Reader 133 Why Model Networks? 134 7. Graph Data in the Real World ........................................... 137 Medium Data: The Tradition 138 Big Data: The Future, Starting Today 138 “Small Data”—Flat File Representations 139 EdgeList Files 139 .net Format 140 GML, GraphML, and other XML Formats 141 Ancient Binary Format—##h Files 142 “Medium Data”: Database Representation 142 What are Cursors? 143 What are Transactions? 144 Names 144 Nodes as Data, Attributes as ? 145 The Class 145 Functions and Decorators 146 Table of Contents | v The Adaptor 148 Working with 2-Mode Data 150 Exercises for the Reader 151 Social Networks and Big Data 151 NoSQL 152 Structural Realities 153 Computational Complexities 156 Big Data is Big 156 Big Data at Work 156 What Are We Distributing? 157 Hadoop, S3, and MapReduce 157 Hive 158 SQL is Still Our Friend 160 A. Data Collection ....................................................... 161 B. Installing Software .................................................... 171 vi | Table of Contents Preface Almost every startup company in 2011 uses the word “social” in their business plans —although few actually know how to analyze and understand the social processes that can result in their firms’ success or failure. If you are working in social media, social CRM, social marketing, organizational consulting, etc., you should read this book for insights on how social systems evolve and change, and how to detect what is going on. Despite the title, the book is not just for startups. In fact, it is a “course-in-a-book”, encapsulating nearly a semester’s worth of theoretical and practical material—read it and you will know enough about social network analysis to be “dangerous”. If you are a student in the field, we strongly encourage you to seek out and read every paper and book referred to in the footnotes. This will make you conversant in the classic literature of the field and enable you to confidently start your own research project. If you are of a technical or computer science background, this book will introduce you to major sociological concepts and tie them back into things that can be programmed, and data that can be analyzed. If your background is social sciences or marketing, you may find some of the background material familiar, but at the same time will learn the quantitative and programmatic approaches to understanding humans in a social setting. Prerequisites This books is written to be accessible by a wide audience. We keep jargon to a mini- mum, and explain terms as they come along. However, there is a serious amount of technical content (as one expects from an O’Reilly book). We expect you to be at least marginally conversant in Python—i.e., able to write your own scripts, understand the basic control structures and data structures of the lan- guage. If you are not yet there, we suggest starting with an online Python tutorial or Head First Python by Paul Barry (O’Reilly). vii We do not cover in detail the process of harvesting data from Twitter, Facebook, and other data sources. Other books in O’Reilly’s “Animal Guide” series provide ample coverage, including Twitter API: Up and Running by Kevin Makice, and Mining the Social Web by Matthew Russell. Open-Source Tools This book utilizes open-source Python libraries including NetworkX, NumPy, and MatPlotLib. • NetworkX packages and documentation can be found at http://networkx.lanl.gov/. Aric A. Hagberg, Daniel A. Schult and Pieter J. Swart, “Exploring network struc- ture, dynamics, and function using NetworkX”, in Proceedings of the 7th Python in Science Conference (SciPy2008), eds. Gäel Varoquaux, Travis Vaught, and Jarrod Millman, pp. 11–15, Aug 2008. • NumPy packages and documentation are at http://numpy.scipy.org/. Oliphant, Travis E. “Python for Scientific Computing”. Computing in Science & Engineering 9(2007), 10-20 Ascher, D. et al. Numerical Python, tech. report UCRL- MA-128569, Lawrence Livermore National Laboratory, 2001. • MatplotLib can be found at http://matplotlib.sourceforge.net. Hunter, J.D. “Matplotlib: A 2D Graphics Environment”. Computing in Science & Engineering 9(2007), 90-95 Conventions Used in This Book The following typographical conventions are used in this book: Italic Indicates new terms, URLs, email addresses, filenames, and file extensions. Constant width Used for program listings, as well as within paragraphs to refer to program elements such as variable or function names, databases, data types, environment variables, statements, and keywords. Constant width bold Shows commands or other text that should be typed literally by the user. Constant width italic Shows text that should be replaced with user-supplied values or by values deter- mined by context. viii | Preface
Description: