Data Profiling
Home > Computing and Information Technology > Computer hardware > Network hardware > Data Profiling: (Synthesis Lectures on Data Management)
Data Profiling: (Synthesis Lectures on Data Management)

Data Profiling: (Synthesis Lectures on Data Management)


     0     
5
4
3
2
1



International Edition


X
About the Book

Data profiling refers to the activity of collecting data about data, {i.e.}, metadata. Most IT professionals and researchers who work with data have engaged in data profiling, at least informally, to understand and explore an unfamiliar dataset or to determine whether a new dataset is appropriate for a particular task at hand. Data profiling results are also important in a variety of other situations, including query optimization, data integration, and data cleaning. Simple metadata are statistics, such as the number of rows and columns, schema and datatype information, the number of distinct values, statistical value distributions, and the number of null or empty values in each column. More complex types of metadata are statements about multiple columns and their correlation, such as candidate keys, functional dependencies, and other types of dependencies. This book provides a classification of the various types of profilable metadata, discusses popular data profiling tasks,and surveys state-of-the-art profiling algorithms. While most of the book focuses on tasks and algorithms for relational data profiling, we also briefly discuss systems and techniques for profiling non-relational data such as graphs and text. We conclude with a discussion of data profiling challenges and directions for future work in this area.

Table of Contents:
Preface.- Acknowledgments.- Discovering Metadata.- Data Profiling Tasks.- Single-Column Analysis.- Dependency Discovery.- Relaxed and Other Dependencies.- Use Cases.- Profiling Non-Relational Data.- Data Profiling Tools.- Data Profiling Challenges.- Conclusions.- Bibliography.- Authors' Biographies .

About the Author :
Ziawasch Abedjan is Assistant Professor and Head of the ""Big Data Management"" (BigDaMa) Group at the Technische Universitat Berlin. Before Ziawasch was a postdoc at the ""Computer Science and Artificial Intelligence Laboratory"" at MIT working on various data integration topics. Ziawasch received his Ph.D. from the Hasso Plattner Institute in Potsdam, Germany. His research interests include, data mining, data integration, and data profiling.Lukasz Golab is an Associate Professor at the University of Waterloo and a Canada Research Chair. Prior to joining Waterloo, he was a Senior Member of Research Staff at AT&T Labs in Florham Park, NJ, USA. He holds a B.Sc. in Computer Science (with High Distinction) from the University of Toronto and a Ph.D. in Computer Science (with Alumni Gold Medal) from the University of Waterloo. His publications span several research areas within data management and data analytics, including data stream management, data profiling, data quality, data science for social good, and educational data mining.Felix Naumann studied mathematics, economy, and computer sciences at the University of Technology in Berlin. After receiving his diploma in 1997 he joined the graduate school ""Distributed Information Systems"" at Humboldt University of Berlin. He completed his Ph.D. thesis on ""Quality-driven Query Answering"" in 2000. In 2001 and 2002 he worked at the IBM Almaden Research Center on topics around data integration. From 2003-2006 he was an assistant professor of information integration at the Humboldt University of Berlin. Since 2006 he has held the chair for information systems at the Hasso Plattner Institute at the University of Potsdam in Germany. He is Editor-in-Chief of the Information Systems journal. His research interests are in the areas of information integration, data quality, data cleansing, text extraction, and-of course-data profiling. He has given numerous invited talks and tutorials on the topic of the book.Thorsten Papenbrock is a researcher and lecturer at the Hasso Plattner Institute at the University of Potsdam in Germany. He received his M.Sc. in IT-Systems Engineering in 2014 and his Ph.D. in Computer Science in 2017. His thesis on ""Data Profiling-Efficient Discovery of Dependencies"" inspired many sections of this book. In research, his main interests are data profiling, data cleaning, distributed and parallel computing, database systems, and data analytics.


Best Sellers


Product Details
  • ISBN-13: 9783031007378
  • Publisher: Springer International Publishing AG
  • Publisher Imprint: Springer International Publishing AG
  • Height: 235 mm
  • No of Pages: 136
  • Returnable: Y
  • Width: 191 mm
  • ISBN-10: 3031007379
  • Publisher Date: 08 Nov 2018
  • Binding: Paperback
  • Language: English
  • Returnable: Y
  • Series Title: Synthesis Lectures on Data Management


Similar Products

Add Photo
Add Photo

Customer Reviews

REVIEWS      0     
Click Here To Be The First to Review this Product
Data Profiling: (Synthesis Lectures on Data Management)
Springer International Publishing AG -
Data Profiling: (Synthesis Lectures on Data Management)
Writing guidlines
We want to publish your review, so please:
  • keep your review on the product. Review's that defame author's character will be rejected.
  • Keep your review focused on the product.
  • Avoid writing about customer service. contact us instead if you have issue requiring immediate attention.
  • Refrain from mentioning competitors or the specific price you paid for the product.
  • Do not include any personally identifiable information, such as full names.

Data Profiling: (Synthesis Lectures on Data Management)

Required fields are marked with *

Review Title*
Review
    Add Photo Add up to 6 photos
    Would you recommend this product to a friend?
    Tag this Book Read more
    Does your review contain spoilers?
    What type of reader best describes you?
    I agree to the terms & conditions
    You may receive emails regarding this submission. Any emails will include the ability to opt-out of future communications.

    CUSTOMER RATINGS AND REVIEWS AND QUESTIONS AND ANSWERS TERMS OF USE

    These Terms of Use govern your conduct associated with the Customer Ratings and Reviews and/or Questions and Answers service offered by Bookswagon (the "CRR Service").


    By submitting any content to Bookswagon, you guarantee that:
    • You are the sole author and owner of the intellectual property rights in the content;
    • All "moral rights" that you may have in such content have been voluntarily waived by you;
    • All content that you post is accurate;
    • You are at least 13 years old;
    • Use of the content you supply does not violate these Terms of Use and will not cause injury to any person or entity.
    You further agree that you may not submit any content:
    • That is known by you to be false, inaccurate or misleading;
    • That infringes any third party's copyright, patent, trademark, trade secret or other proprietary rights or rights of publicity or privacy;
    • That violates any law, statute, ordinance or regulation (including, but not limited to, those governing, consumer protection, unfair competition, anti-discrimination or false advertising);
    • That is, or may reasonably be considered to be, defamatory, libelous, hateful, racially or religiously biased or offensive, unlawfully threatening or unlawfully harassing to any individual, partnership or corporation;
    • For which you were compensated or granted any consideration by any unapproved third party;
    • That includes any information that references other websites, addresses, email addresses, contact information or phone numbers;
    • That contains any computer viruses, worms or other potentially damaging computer programs or files.
    You agree to indemnify and hold Bookswagon (and its officers, directors, agents, subsidiaries, joint ventures, employees and third-party service providers, including but not limited to Bazaarvoice, Inc.), harmless from all claims, demands, and damages (actual and consequential) of every kind and nature, known and unknown including reasonable attorneys' fees, arising out of a breach of your representations and warranties set forth above, or your violation of any law or the rights of a third party.


    For any content that you submit, you grant Bookswagon a perpetual, irrevocable, royalty-free, transferable right and license to use, copy, modify, delete in its entirety, adapt, publish, translate, create derivative works from and/or sell, transfer, and/or distribute such content and/or incorporate such content into any form, medium or technology throughout the world without compensation to you. Additionally,  Bookswagon may transfer or share any personal information that you submit with its third-party service providers, including but not limited to Bazaarvoice, Inc. in accordance with  Privacy Policy


    All content that you submit may be used at Bookswagon's sole discretion. Bookswagon reserves the right to change, condense, withhold publication, remove or delete any content on Bookswagon's website that Bookswagon deems, in its sole discretion, to violate the content guidelines or any other provision of these Terms of Use.  Bookswagon does not guarantee that you will have any recourse through Bookswagon to edit or delete any content you have submitted. Ratings and written comments are generally posted within two to four business days. However, Bookswagon reserves the right to remove or to refuse to post any submission to the extent authorized by law. You acknowledge that you, not Bookswagon, are responsible for the contents of your submission. None of the content that you submit shall be subject to any obligation of confidence on the part of Bookswagon, its agents, subsidiaries, affiliates, partners or third party service providers (including but not limited to Bazaarvoice, Inc.)and their respective directors, officers and employees.

    Accept

    Fresh on the Shelf


    Inspired by your browsing history


    Your review has been submitted!

    You've already reviewed this product!