Data cleaning is the largest part of any data analysis project: we discuss our thoughts on all the decisions we make about data and the implications of those decisions.

What are the filtering and manipulation decisions we make when cleaning data that have larger effects on final datasets? Who handles that data cleaning and how much latitude do they have to make decisions? How does a manager understand what was done to the data without reading the code? Beyond that, how do the huge masses of data we deal with change the landscape for modeling and finding signals in the noise?

Download MP3 | Subscribe: RSS

Links:
Fired by Bot at Amazon: ‘It’s You Against the Machine’ [Bloomberg]