ebook img

Automating Data Quality Monitoring at Scale: Going Deeper than Data Observability (Third Early Release) PDF

02023·3.48 MB·English
Save to my drive
Quick download
Download
Most books are stored in the elastic cloud where traffic is expensive. For this reason, we have a limit on daily download.

Preview Automating Data Quality Monitoring at Scale: Going Deeper than Data Observability (Third Early Release)

Description:
The world's businesses ingest a combined 2.5 quintillion bytes of data every day. But how much of this vast amount of data--used to build products, power AI systems, and drive business decisions--is poor quality or just plain bad? This practical book shows you how to ensure that the data your organization relies on contains only high-quality records.Most data engineers, data analysts, and data scientists genuinely care about data quality, but they often don't have the time, resources, or understanding to create a data quality monitoring solution that succeeds at scale. In this book, Jeremy Stanley and Paige Schwartz from Anomalo explain how you can use automated data quality monitoring to cover all your tables efficiently, proactively alert on every category of issue, and resolve problems immediately.We’ll wrap up by introducing the data quality monitoring strategy we advocate for in this book: a three-pillar approach combining rules, metrics monitoring, and Unsupervised Machine Learning. As we’ll show, this approach has multiple benefits. It allows subject matter experts to enforce essential constraints and track KPIs for important tables, while providing a base level of automated monitoring for a large volume of diverse data. This approach doesn’t require massive computer power or legions of analysts to maintain rules and thresholds. With machine learning, it will detect “unknown unknowns” in the data and reduce alert fatigue by understanding correlations and trends in the data values across columns and even across tables, alerting only when changes are new and significantThis book will help you:Learn why data quality is a business imperativeUnderstand and assess unsupervised learning models for detecting data issuesImplement notifications that reduce alert fatigue and let you triage and resolve issues quicklyIntegrate automated data quality monitoring with data catalogs, orchestration layers, and BI and ML systemsUnderstand the limits of automated data quality monitoring and how to overcome themLearn how to deploy and manage your monitoring solution at scaleMaintain automated data quality monitoring for the long term
See more

The list of books you might like

Most books are stored in the elastic cloud where traffic is expensive. For this reason, we have a limit on daily download.