Datacast: Recent Episodes

James Le

Datacast follows the narrative journey of data practitioners and researchers to unpack the career lessons they learned along the way. James Le hosts the show.

View Details

Show Notes* (01:41) Salma reflected on her upbringing in Paris (France) and her love for math since a young age. * (04:05) Salma talked about her decision to dive deeper into statistics. * (07:39) Salma recalled her 5 years in investment banking in Hong Kong. * (10:44) Salma doubled down on the cultivation of her resilience during this period. * (13:45) Salma shared the founding story Sifflet. * (15:53) Salma touched on her co-founder dynamics with Wajdi Fathallah and Wissem Fathallah. * (18:50) Salma explained the concept of Full Data Stack Observability to the uninitiated. * (22:14) Salma extrapolated on Sifflet's approach to data observability. * (24:37) Salma discussed Sifflet's data quality focus with automated monitoring coverage and over 50 data quality templates. * (27:28) Salma discussed Sifflet's data lineage solution with field-level lineage, root cause analysis, and incident management/business impact assessment. * (31:47) Salma discussed Sifflet's data catalog with a powerful metadata search engine and centralized documentation for all data assets. * (33:46) Salma expanded on the full-stack mindset as a data vendor. * (37:03) Salma emphasized the importance of integrations with other data tools. * (38:52) Salma touched on product features such as Flow Stopper to stop vulnerable pipelines from running at the orchestration layer and Metrics Observability to extend the observability framework to the semantic layer. * (43:40) Salma unpacked her article on building a modern data team. * (47:54) Salma shared some hiring lessons to attract the right people to Sifflet. * (54:02) Salma emphasized the importance of company branding. * (55:01) Salma shared her thoughts on building a startup culture. * (57:26) Salma briefly mentioned her process of working with design partners in the early stage. * (01:00:17) Salma emphasized her enthusiasm for the broader data community. * (01:02:05) Conclusion.

Salma's Contact Info* LinkedIn * Twitter * Medium

Sifflet's Resources* Website | LinkedIn | Twitter | Docs * Data Catalog | Data Quality Monitoring | Data Lineage | Integrations

Mentioned ContentPeople* Zhamak Dehghani (Creator of Data Mesh and Founder of NextData) * Benoit Dageville, Thierry Cruanes, and Marcin Zukowski (Founders of Snowflake)

Books* "The Hard Thing About Hard Things" (by Ben Horowitz) * "The Boys In The Boat" (by Daniel James Brown)

NotesMy conversation with Salma was recorded back in late 2022. Since then, I recommend checking out the launch of Sifflet AI Assistant and this blog post on 2024 data trends.

About the showDatacast features long-form, in-depth conversations with practitioners and researchers in the data community to walk through their professional journeys and unpack the lessons learned along the way. I invite guests coming from a wide range of career paths — from scientists and analysts to founders and investors — to analyze the case for using data in the real world and extract their mental models (“the WHY and the HOW”) behind their pursuits. Hopefully, these conversations can serve as valuable tools for early-stage data professionals as they navigate their own careers in the exciting data universe.

Datacast is produced and edited by James Le. For inquiries about sponsoring the podcast, email khanhle.1013@gmail.com.

Subscribe by searching for Datacast wherever you get podcasts, or click one of the links below:

  • Listen on Spotify
  • Listen on Apple Podcasts
  • Listen on Google Podcasts

If you’re new, see the podcast homepage for the most recent episodes to listen to, or browse the full guest list.

View Details

Show Notes* (01:53) Suresh went over his college experience studying Electronics Engineering at the National Institute of Technology Karnataka. * (04:35) Suresh recalled his 9-year engineering career at Sylantro Systems. * (08:47) Suresh talked about the origin of Apache Hadoop at Yahoo. * (11:05) Suresh dissected the high-level design architecture of the Hadoop Distributed File System (HDFS). * (15:36) Suresh reflected on his decision to become a co-founder of Hortonworks, which focused on bringing Hadoop training and support to enterprise customers. * (17:36) Suresh unpacked the evolution of the Hortonworks Data Platform - which includes Hadoop technology such as HDFS, MapReduce, Pig, Hive, HBase, ZooKeeper, and additional components. * (20:30) Suresh shared his lessons from developing and supporting open-source software designed to manage big data processing. * (23:43) Suresh walked through the evolution of Uber’s Data Platform. * (28:03) Suresh described Uber's journey toward better data culture from first principles. * (34:00) Suresh explained his motivation to start the OpenMetadata Project. * (37:21) Suresh elaborated on OpenMetadata's five design principles: schema-first, extensibility, API-centric, vendor-neural, and open-source. * (40:17) Suresh highlighted OpenMetadata's built-in features to power multiple applications, such as data collaboration, metadata versioning, and data lineage. * (44:38) Suresh emphasized his priority for the open-source roadmap to adapt to the community's needs. * (47:05) Suresh explained the architecture of OpenMetadata - which goes deep into the push-based and pull-based characteristics of metadata ingestion and consumption. * (51:47) Suresh shared the long-term vision of his new company Collate, which powers the OpenMetadata initiative. * (53:36) Suresh shared valuable hiring lessons as a startup founder. * (56:30) Suresh shared fundraising advice to founders who want to seek the right investors for their startups. * (57:50) Closing segment.

Suresh's Contact Info* LinkedIn * Twitter * GitHub

OpenMetadata's Resources* Website | Twitter * Slack | GitHub | Community * Documentation * Collate

Mentioned ContentPeople* Joe Littlejohn (jsonschema2pojo) * Sriharsha Chintalapani

Book* The Innovator's Dilemma (by Clayton Christensen)

About the showDatacast features long-form, in-depth conversations with practitioners and researchers in the data community to walk through their professional journeys and unpack the lessons learned along the way. I invite guests coming from a wide range of career paths — from scientists and analysts to founders and investors — to analyze the case for using data in the real world and extract their mental models (“the WHY and the HOW”) behind their pursuits. Hopefully, these conversations can serve as valuable tools for early-stage data professionals as they navigate their own careers in the exciting data universe.

Datacast is produced and edited by James Le. For inquiries about sponsoring the podcast, email khanhle.1013@gmail.com.

Subscribe by searching for Datacast wherever you get podcasts, or click one of the links below:

  • Listen on Spotify
  • Listen on Apple Podcasts
  • Listen on Google Podcasts

If you’re new, see the podcast homepage for the most recent episodes to listen to, or browse the full guest list.

View Details

Show Notes (02:20) Krishna described his academic experience getting an MS in Computer Science from the University of Minnesota - where he developed efficient tools for text document clustering and pattern discovery. * (05:36) Krishna recalled his 5.5 years at Microsoft working on Bing's search engine. * (08:32) Krishna talked about the challenges of competing against Google Search. * (10:22) Krishna shared the high-level technical and operational challenges encountered during the development and scaling phase of Twitter Search. * (14:55) Krishna revealed vital lessons from building critical data infrastructure at Twitter. * (17:54) Krishna touched on his time at Pinterest as the head of data engineering - leading a team working on all things data from analytics, experimentation, logging, and infrastructure. * (20:05) Krishna reviewed the design and implementation of real-time analytics, ETL-as-a-Service, and an A/B testing platform at Pinterest. * (24:40) Krishna unpacked the major ML model performance issues while running Facebook's feed ranking platform. * (28:18) Krishna distilled lessons learned about algorithmic governance from Facebook. * (31:38) Krishna provided leadership lessons from building teams that create scalable platforms and delightful consumer products on Twitter, Pinterest, and Facebook. * (33:19) Krishna shared the founding story of Fiddler AI, whose mission is to build trust into AI. * (37:56) Krishna unpacked the key challenges and tools in his 2019 article "AI needs a new developer stack." * (40:49) Krishna discussed the evolution of MLOps over the past 4 years. * (42:48) Krishna explained the benefits of using the Model Performance Management (MPM) framework to address enterprise MLOps challenges. * (47:01) Krishna gave a brief overview of capabilities within Fiddler's MPM platform, such as model monitoring, explainable AI, analytics, and fairness. * (50:28) Krishna highlighted research efforts inside Fiddler concerning explainability, drift metric calculation, and fairness. * (53:17) Krishna discussed the challenges with monitoring for NLP and Computer Vision models. * (57:18) Krishna zoomed in on Fiddler's approach to model governance for the modern enterprise. * (01:02:24) Krishna distilled v*aluable lessons learned to attract the right people who are excited about Fiddler's mission and aligned with Fiddler's culture. * (01:06:08) Krishna reflected on the evolution of Fiddler's company culture. * (01:09:19) Krishna shared the challenges of finding the early design partners and defining a new category of Responsible AI. * (01:12:23) Krishna gave fundraising advice to founders who are seeking the right investors for their startups. * (01:14:45) Closing segment.

Krishna's Contact Info* LinkedIn * Twitter * Medium

Fiddler's Resources* Website | LinkedIn | Twitter | YouTube * About | Customers | Careers * AI Observability | Model Monitoring | Explainable AI | Fairness | Analytics * Blog | Docs | Resources

Mentioned ContentPeople1. Goku Mohamandas (Made With ML and Anyscale) 2. Krishnaram Kenthapadi (Chief AI Officer & Chief Scientist at Fiddler)

Books1. "The Hard Thing About Hard Things" (Ben Horowitz) 2. "The Five Dysfunctions of A Team" (Patrick Lencioni)

NotesMy conversation with Krishna was recorded more than a year ago. Since then, I'd recommend checking out these Fiddler's resources:

  1. Strategic investments in Fiddler by Alteryx Ventures, Mozilla Ventures, Dentsu Ventures, and Scale Asia Ventures.
  2. Fiddler introduces an end-to-end workflow for robust Generative AI back in May 2023.
  3. Krishna's thought leadership on LLMOps and the missing link in Generative AI.

About the showDatacast features long-form, in-depth conversations with practitioners and researchers in the data community to walk through their professional journeys and unpack the lessons learned along the way. I invite guests coming from a wide range of career paths — from scientists and analysts to founders and investors — to analyze the case for using data in the real world and extract their mental models (“the WHY and the HOW”) behind their pursuits. Hopefully, these conversations can serve as valuable tools for early-stage data professionals as they navigate their own careers in the exciting data universe.

Datacast is produced and edited by James Le. For inquiries about sponsoring the podcast, email khanhle.1013@gmail.com.

Subscribe by searching for Datacast wherever you get podcasts, or click one of the links below:

  • Listen on Spotify
  • Listen on Apple Podcasts
  • Listen on Google Podcasts

If you’re new, see the podcast homepage for the most recent episodes to listen to, or browse the full guest list.

View Details

Show Notes* (01:59) Emanuel reflected on his upbringing in Switzerland and his 4-year apprenticeship in Software Engineering and Business at Credit Suisse. * (06:09) Emanuel recalled his 4-year program in Computer Science at HSR (University of Applied Sciences Rapperswil). * (08:43) Emanuel touched on his decision to pursue a Master’s degree at Brown University in the US. * (11:58) Emanuel explained his decision to continue with a Ph.D. degree at Brown under the advisement of Professor Andy van Dam and summarized the arc of his Ph.D. research focus. * (14:50) Emanuel highlighted his first research paper called PanoramicData on Interactive Data Exploration. * (16:42) Emanuel shared his thoughts on common traits of a successful researcher. * (18:03) Emanuel emphasized the focus of his research at the intersection of Human-Computer Interaction, Information Visualization, and Data Analysis. * (20:38) Emanuel shared valuable lessons from interning twice at Microsoft Research in Redmond. * (22:34) Emanuel talked about his time as a postdoc in Professor Tim Kraska’s group at the CSAIL at MIT. * (24:59) Emanuel shared the founding story of Einblick - a visual computing platform that enables data teams to answer tougher, more meaningful questions by making advanced analytics and model building more streamlined and accessible. * (29:06) Emanuel touched on the responsibilities of Einblick's 5 co-founders. * (30:23) Emanuel highlighted technical challenges of building Einblick's integrated environment for descriptive, predictive, and prescriptive analytics. * (32:09) Emanuel mentioned the collaboration challenge in data and brought up Einblick's real-time remote collaboration through video-enabled data whiteboards. * (35:39) Emanuel highlighted the challenges of working with computational notebooks and brought up the benefits of using Einblick's collaborative visual canvas. * (38:54) Emanuel unpacked the challenges of commercializing an academic research project. * (40:27) Emanuel gave a broad overview of Einblick's go-to-market strategy. * (43:55) Emanuel shared valuable hiring lessons to attract the right people who are aligned with Einblick’s cultural values. * (47:05) Emanuel shared fundraising advice to founders who are seeking the right investors for their startups. * (49:20) Emanuel shared the similarities and differences between being a researcher and being a founder. * (50:43) Closing segment.

Emanuel's Contact Info* Website * Google Scholar * LinkedIn * Twitter

Einblick's Resources* Website | Twitter | LinkedIn * Docs | Blog * ChartGen AI * Notebook Feature Release (2022) * Video-Based Collaboration Release (2021)

Mentioned ContentPapers and Projects* PanoramicData is a hybrid pen and touch system for visual data exploration (Infovis 2014 Paper | Video) * (s|qu)eries (pronounced “Squeries”) is a visual query interface for creating queries on sequences (series) of data based on regular expressions (CHI 2015 Paper | Summary Video) * Vizdom is an interactive visual analytics system that scales to large datasets through progressive computation (VLDB Demo 2015 Paper | Health Video | Election Video) * Tableur is a spreadsheet-like pen- and touch-based system that revolves around handwriting recognition - all data is represented as digital ink (CHI 2016 LBW Paper | Video) * Towards Accessible Data Analysis (Emanuel's Ph.D. Dissertation at Brown, 2018) * Northstar is an interactive data science platform that combines data exploration with automated machine learning (SIGMOD DEEM Paper | Video)

People* Wes McKinney * Fei-Fei Li

Books* "The Book of Why" (by Judea Pearl) * "The Signal and The Noise" (by Nate Silver)

NotesMy conversation with Emanuel was recorded back in late 2022. Since then, I recommend checking out the launch of Einblick Prompt and ChartGenAI.

About the showDatacast features long-form, in-depth conversations with practitioners and researchers in the data community to walk through their professional journeys and unpack the lessons learned along the way. I invite guests coming from a wide range of career paths — from scientists and analysts to founders and investors — to analyze the case for using data in the real world and extract their mental models (“the WHY and the HOW”) behind their pursuits. Hopefully, these conversations can serve as valuable tools for early-stage data professionals as they navigate their own careers in the exciting data universe.

Datacast is produced and edited by James Le. For inquiries about sponsoring the podcast, email khanhle.1013@gmail.com.

Subscribe by searching for Datacast wherever you get podcasts, or click one of the links below:

  • Listen on Spotify
  • Listen on Apple Podcasts
  • Listen on Google Podcasts

If you’re new, see the podcast homepage for the most recent episodes to listen to, or browse the full guest list.

About the showDatacast features long-form, in-depth conversations with practitioners and researchers in the data community to walk through their professional journeys and unpack the lessons learned along the way. I invite guests coming from a wide range of career paths — from scientists and analysts to founders and investors — to analyze the case for using data in the real world and extract their mental models (“the WHY and the HOW”) behind their pursuits. Hopefully, these conversations can serve as valuable tools for early-stage data professionals as they navigate their own careers in the exciting data universe.

Datacast is produced and edited by James Le. For inquiries about sponsoring the podcast, email khanhle.1013@gmail.com.

Subscribe by searching for Datacast wherever you get podcasts, or click one of the links below:

  • Listen on Spotify
  • Listen on Apple Podcasts
  • Listen on Google Podcasts

If you’re new, see the podcast homepage for the most recent episodes to listen to, or browse the full guest list.

View Details

Timestamps* (01:41) Diana shared her upbringing in Orlando and her undergraduate experience studying Economics at MIT. * (03:27) Diana reflected on valuable lessons from her internships during college. * (05:58) Diana brought up her learning during the year working as an investment banking analyst in the technology group at Morgan Stanley. * (09:10) Diana recalled her transition from investment banking into venture capital at Norwest Venture Partners. * (11:52) Diana discussed the takeaways from her time as a venture associate meeting entrepreneurs on a regular cadence. * (14:23) Diana recalled her decision to leave venture capital and become the first hire at the product organization at Cockroach Labs. * (19:00) Diana went over the challenges and learning curves as a non-technical first product hire at Cockroach. * (21:26) Diana extrapolated on the idea of determining the best-fit product strategy rather than blindly following frameworks. * (23:55) Diana described her experience as the first hire into the product organization at TimescaleDB. * (26:40) Diana highlighted the challenges of open-source GTM. * (27:56) Diana reflected on her 3-part blog series on building a product for the most dissatisfied customers first, the majority next, and the full need in the long run. * (30:40) Diana shared 2 tactical lessons to cultivate focus as a product manager. * (32:44) Diana share the founding story of Correlated alongside her co-founders, Tim Geisenheimer and John Pena. * (35:38) Diana briefly touched on her 2 entrepreneurial attempts during COVID-19. * (38:30) Diana unpacked the notion of Product-Led Revenue and described how Correlated works at a high level. * (40:49) Diana highlighted the role of integrations within Correlated's product strategy. * (43:04) Diana mentioned Correlated's product-led playbooks to help users manage their product-led strategy from start to finish. * (45:40) Diana explained how she leveraged customer feedback to ship the feature called PQL Scoring that leverages machine learning to identify the best leads. * (48:51) Diana shared the consistent principles that have remained the same for successful communication in Product Management. * (52:30) Diana discussed her learnings on customer discovery at early-stage startups. * (56:08) Diana reflected on the early signs of product-market fit that carry through all of her startups. * (58:58) Conclusion

Diana's Contact Info* LinkedIn * Twitter * Medium * Substack

Correlated's Resources* Website | LinkedIn | Twitter * Product Overview | How Correlated Works * Blog | Podcast | Docs * PLG Playbook Library * Correlated Launches to Bring Product-Led Revenue to Market with $8.3M in Funding * What Is Product-Led Revenue? * Correlated launches PQL Scoring to accelerate your product-led strategy

Mentioned ContentBlog Posts* "The Standard Due Diligence Process" (Jan 2016) * "Mistakes to Avoid when Pitching to a VC" (Jan 2016) * "My Startup Litmus Test" (Feb 2016) * "Why I left VC to join Cockroach Labs" (April 2017) * "My First 90 Days as the First Product Hire" (May 2017) * "Coding != Technical: What It Means to be Technical as a PM" (Aug 2017) * "How learning to sell makes for a better product manager" (Nov 2017) * "Roadmap Planning: Users First, Features Second" (March 2018) * "Build something people will use more than once" (May 2019) * "Focus on the unhappiest, most dissatisfied customers first" (May 2019) * "Build for the majority" (May 2019) * "Why user interviews can fail you when starting a startup" (Sep 2021) * "Tackling the challenges of communicating effectively in product management" (Jan 2022) * "Some Learnings on Customer Discovery at Early-Stage Startups" (May 2022) * "4 early signs of product-market fit" (Sep 2022) * "Give customers what they want, but not what they ask for" (Sep 2022)

People1. Lenny Rachitsky 2. Julie Zhuo 3. Nate Stewart 4. Jeff Sposetti

Book* "Crossing The Chasm" (by Geoffrey Moore)

About the showDatacast features long-form, in-depth conversations with practitioners and researchers in the data community to walk through their professional journeys and unpack the lessons learned along the way. I invite guests coming from a wide range of career paths — from scientists and analysts to founders and investors — to analyze the case for using data in the real world and extract their mental models (“the WHY and the HOW”) behind their pursuits. Hopefully, these conversations can serve as valuable tools for early-stage data professionals as they navigate their own careers in the exciting data universe.

Datacast is produced and edited by James Le. For inquiries about sponsoring the podcast, email khanhle.1013@gmail.com.

Subscribe by searching for Datacast wherever you get podcasts, or click one of the links below:

  • Listen on Spotify
  • Listen on Apple Podcasts
  • Listen on Google Podcasts

If you’re new, see the podcast homepage for the most recent episodes to listen to, or browse the full guest list.

About the showDatacast features long-form, in-depth conversations with practitioners and researchers in the data community to walk through their professional journeys and unpack the lessons learned along the way. I invite guests coming from a wide range of career paths — from scientists and analysts to founders and investors — to analyze the case for using data in the real world and extract their mental models (“the WHY and the HOW”) behind their pursuits. Hopefully, these conversations can serve as valuable tools for early-stage data professionals as they navigate their own careers in the exciting data universe.

Datacast is produced and edited by James Le. For inquiries about sponsoring the podcast, email khanhle.1013@gmail.com.

Subscribe by searching for Datacast wherever you get podcasts, or click one of the links below:

  • Listen on Spotify
  • Listen on Apple Podcasts
  • Listen on Google Podcasts

If you’re new, see the podcast homepage for the most recent episodes to listen to, or browse the full guest list.

View Details

Timestamps* (01:45) Jason shared the formative experiences of his upbringing in the Bay Area and coming of age in the “Moneyball” era of baseball. * (05:03) Jason described his overall academic experience at Stanford - where he studied Mathematical and Computational Science with a minor in Classical Studies. * (09:15) Jason reflected on his experience participating in the Mayfield Fellowship at Stanford. * (12:03) Jason recalled his time being a part of the business operations team during a high-growth period at Opendoor. * (14:25) Jason talked about lessons learned working as a management consultant at McKinsey’s Bay Area practice. * (15:59) Jason reminisced about his time at the AI Fund startup studio - where he launched AI-enabled SaaS startups by iterating on prototypes, signing design partners, and recruiting the founding team. * (19:25) Jason explained his decision to join the investment team at Greylock Partners. * (22:24) Jason walked through his journey proving value as a new investor. * (24:41) Jason unpacked his checklist for evaluating early-stage enterprise investment opportunities. * (27:09) Jason explained his seed investment in Onehouse - a cloud-native managed lakehouse service that makes data lakes easier, faster, and cheaper. * (30:31 ) Jason explained his Series A investment in Baseten - which builds a powerful software toolkit that empowers technical data science teams to serve, integrate, design, and ship their custom ML models efficiently. * (33:23) Jason touched on advice for his portfolio companies in hiring decisions and navigating product/GTM strategy. * (37:00) Jason unpacked key takeaways from Greylock’s Castles in the Cloud project. * (39:58) Jason dissected key trends in the markets of security, AI/ML, management and governance, and edge computing (as shown in "VC Funding for the Cloud"). * (46:24) Jason elaborated on his vision of "The Next Cloud Data Platform" - which examines how the data warehouse, lakehouse, and semantic layer could combine to create a platform for data applications. * (50:55) Jason shared a few books that have greatly influenced his life. * (52:22) Closing segment.

Jason's Contact Info* LinkedIn * Twitter * Greylock

Mentioned ContentBooks1. "Moneyball" (by Michael Lewis) 2. "Why The West Rules For Now" (by Ian Morris) 3. "Snow Crash" (by Neil Stephenson) 4. "Cryptonomicon" (by Neil Stephenson) 5. "Termination Shock" (by Neil Stephenson) 6. "Principles for Dealing with the Changing World Order" (by Ray Dalio)

People1. David Luan (Founder and CEO of Adept) 2. Alex Ratner (Co-Founder and CEO of Snorkel AI) 3. Frank Slootman (CEO of Snowflake) 4. Clement Delangue (Co-Founder and CEO of HuggingFace)

NotesMy conversation with Jason was recorded back in late 2022. Since then, I recommend checking out these resources:

  1. His blog post on the next platform opportunity in cybersecurity
  2. Greylock's investment in LlamaIndex

About the showDatacast features long-form, in-depth conversations with practitioners and researchers in the data community to walk through their professional journeys and unpack the lessons learned along the way. I invite guests coming from a wide range of career paths — from scientists and analysts to founders and investors — to analyze the case for using data in the real world and extract their mental models (“the WHY and the HOW”) behind their pursuits. Hopefully, these conversations can serve as valuable tools for early-stage data professionals as they navigate their own careers in the exciting data universe.

Datacast is produced and edited by James Le. For inquiries about sponsoring the podcast, email khanhle.1013@gmail.com.

Subscribe by searching for Datacast wherever you get podcasts, or click one of the links below:

  • Listen on Spotify
  • Listen on Apple Podcasts
  • Listen on Google Podcasts

If you’re new, see the podcast homepage for the most recent episodes to listen to, or browse the full guest list.

About the showDatacast features long-form, in-depth conversations with practitioners and researchers in the data community to walk through their professional journeys and unpack the lessons learned along the way. I invite guests coming from a wide range of career paths — from scientists and analysts to founders and investors — to analyze the case for using data in the real world and extract their mental models (“the WHY and the HOW”) behind their pursuits. Hopefully, these conversations can serve as valuable tools for early-stage data professionals as they navigate their own careers in the exciting data universe.

Datacast is produced and edited by James Le. For inquiries about sponsoring the podcast, email khanhle.1013@gmail.com.

Subscribe by searching for Datacast wherever you get podcasts, or click one of the links below:

  • Listen on Spotify
  • Listen on Apple Podcasts
  • Listen on Google Podcasts

If you’re new, see the podcast homepage for the most recent episodes to listen to, or browse the full guest list.

View Details

Show Notes* (01:44) Heather talked about her upbringing, her education, and her 14-year career at Liberty Mutual Insurance. * (06:50) Heather emphasized the benefits of her education in organizational design and management. * (08:56) Heather walked through her decision to shift from underwriting to technology and data at Liberty. * (13:55) Heather commented on her 5 years as a product manager at Crum & Forster. * (20:14) Heather described her 2-year experience as the Global Head of Technology at Innovisk (a Wills Towers Watson ). * (23:54) Heather distinguished the work environments in a startup and a large company. * (26:17) Heather recalled different data strategy initiatives she led at Brown and Brown Insurance. * (28:33) Heather explained issues in the insurance value chain and the role of data to help tackle them. * (34:17) Heather discussed her role as the founding Chief Data Officer at Accelerant Holdings. * (40:18) Heather brought up the data quality issues that Accelerant risk exchange helped solve. * (45:52) Heather gave advice to organizations to move from data governance to Data Intelligence. * (48:45) Heather provided her perspective on hiring data talent. * (50:40) Heather looked at the insurance transformation from a technology perspective. * (52:52) Heather talked about engaging women in technical fields. * (54:22) Closing segment.

Heather's Contact Info* LinkedIn * Accelerant | About

Relevant Reading* Bloomberg | Boehly'sHeather's Eldridge Bets on Accelerant at $2.2 Billion Valuation * Insurance Business Mag | Accelerant: An insurtech that defies categories * Business Insurance | Accelerant establishes $175 million sidecar reinsurer * AI Times Journal | Data Intelligence is Key to Understanding our Customers – Chief Data Officer, Accelerant Holdings * Lightco | Insurance Innovators Top 100 * Business Wire | Accelerant Launches the Accelerant Risk Exchange to Reimagine Insurance

Mentioned ResourcesPeople1. Zhamak Dehgani 2. Cassie Kozyrkov 3. Allie Miller

Book* Data Mesh: Delivering Data-Driven Value at Scale (by Zhamak Dehghani)

About the showDatacast features long-form, in-depth conversations with practitioners and researchers in the data community to walk through their professional journeys and unpack the lessons learned along the way. I invite guests coming from a wide range of career paths — from scientists and analysts to founders and investors — to analyze the case for using data in the real world and extract their mental models (“the WHY and the HOW”) behind their pursuits. Hopefully, these conversations can serve as valuable tools for early-stage data professionals as they navigate their own careers in the exciting data universe.

Datacast is produced and edited by James Le. For inquiries about sponsoring the podcast, email khanhle.1013@gmail.com.

Subscribe by searching for Datacast wherever you get podcasts, or click one of the links below:

  • Listen on Spotify
  • Listen on Apple Podcasts
  • Listen on Google Podcasts

If you’re new, see the podcast homepage for the most recent episodes to listen to, or browse the full guest list.

About the showDatacast features long-form, in-depth conversations with practitioners and researchers in the data community to walk through their professional journeys and unpack the lessons learned along the way. I invite guests coming from a wide range of career paths — from scientists and analysts to founders and investors — to analyze the case for using data in the real world and extract their mental models (“the WHY and the HOW”) behind their pursuits. Hopefully, these conversations can serve as valuable tools for early-stage data professionals as they navigate their own careers in the exciting data universe.

Datacast is produced and edited by James Le. For inquiries about sponsoring the podcast, email khanhle.1013@gmail.com.

Subscribe by searching for Datacast wherever you get podcasts, or click one of the links below:

  • Listen on Spotify
  • Listen on Apple Podcasts
  • Listen on Google Podcasts

If you’re new, see the podcast homepage for the most recent episodes to listen to, or browse the full guest list.

View Details

Show Notes* (01:36) Bob shared formative experiences of his upbringing with exposure to technology. * (05:08) Bob discussed his time building software for fun as a teenager. * (07:08) Bob reflected on his education in music and his decision to transition to a career in software development. * (10:52) Bob explained his project control(human, data, sound) while working on his consultancy agency Kubrickology. * (14:25) Bob discussed how using software can enhance our creativity. * (17:28) Bob talked about his fascination with merging physical and digital realms. * (19:41) Bob recalled his TEDx talk that introduced three high-level ideas about why software works well based on its ability to adapt to our language. * (24:57) Bob shared the founding story of Weaviate. * (29:40) Bob talked about his process of choosing his co-founders. * (31:29) Bob unpacked his high-level thinking around creating a business model around the open-source project. * (38:06) Bob defined a vector search engine for the uninitiated. * (40:49) Bob gave a brief overview of the high-level design of Weaviate. * (43:15) Bob talked about Weaviate's production-ready features, such as horizontal scalability and graph-like connections between objects. * (45:45) Bob reviewed the use cases for Weaviate that he is most proud of. * (49:31) Bob emphasized the importance of engaging open-source contributors to generate valuable product feedback. * (55:03) Bob talked about the pricing model for Weaviate Cloud Service. * (57:59) Bob anticipated the evolution of the tooling landscape within the AI-first database ecosystem to support the increasing adoption of unstructured data. * (01:02:31) Bob shared valuable hiring lessons to attract the right people to join Weaviate. * (01:04:36) Bob explained his process of identifying people who align with the cultural values of Weaviate. * (01:08:27) Bob gave fundraising advice to founders who are seeking the right investors for their startups. * (01:12:17) Bob highlighted his thinking around being a remote-first company and building an open-source brand. * (01:16:28) Closing segment.

Bob's Contact Info* Wikipedia * LinkedIn * Twitter * GitHub * WTF Medium Blog * YouTube

Weaviate's Resources* Website | Twitter | Slack | Forum | GitHub * Blog | Podcast | Playbook

Mentioned ContentPeople1. Sam Ramji (DataStax) 2. Paul Graham (Y Combinator)

Book* "Hackers and Painters" (by Paul Graham)

NotesMy conversation with Bob was recorded back in late 2022. Since then, I recommend checking out these resources:

  • SeMI Tech becomes Weaviate
  • The $50M Series B funding led by Index with participation from Battery
  • The public beta of Weaviate Cloud Service
  • Bob's posts on Weaviate's organic growth and 4th birthday

About the showDatacast features long-form, in-depth conversations with practitioners and researchers in the data community to walk through their professional journeys and unpack the lessons learned along the way. I invite guests coming from a wide range of career paths — from scientists and analysts to founders and investors — to analyze the case for using data in the real world and extract their mental models (“the WHY and the HOW”) behind their pursuits. Hopefully, these conversations can serve as valuable tools for early-stage data professionals as they navigate their own careers in the exciting data universe.

Datacast is produced and edited by James Le. For inquiries about sponsoring the podcast, email khanhle.1013@gmail.com.

Subscribe by searching for Datacast wherever you get podcasts, or click one of the links below:

  • Listen on Spotify
  • Listen on Apple Podcasts
  • Listen on Google Podcasts

If you’re new, see the podcast homepage for the most recent episodes to listen to, or browse the full guest list.

View Details

Show Notes* (01:40) Sakib shared formative experiences of his upbringing in SoCal and his undergraduate experience at the University of Pennsylvania. * (05:24) Sakib recalled his favorite classes at Penn. * (07:55) Sakib reflected on his internship experience at Innova Dynamics and Morgan Stanley. * (11:04) Sakib reflected on his decision to pursue a career in venture capital at Bessemer Venture Partners. * (14:02) Sakib walked through his process of proving value as a new investor. * (16:21) Sakib explained his process of forming clear investment theses. * (18:35) Sakib talked about his brief year as a product manager at Viagogo before going back to Bessemer. * (22:04) Sakib dissected his investments in LaunchDarkly and PagerDuty (in the domain of developer-centric platforms). * (24:16) Sakib explained his investments in Coiled, Prefect, and Arcion Labs (in the domain of data infrastructure). * (25:59) Sakib walked through his investment in Guild Education and Tribe (in the domain of education and community management). * (29:06) Sakib shared advice to portfolio companies in terms of navigating hard decisions and growth strategy. * (34:18) Sakib outlined Bessemer's roadmap on data infrastructure - which looks at the wave of startups enabling the next generation of data-driven businesses. * (39:57) Sakib brought up the products that help abstract away complexity from data engineering problems. * (41:34) Sakib highlighted the tools that power the next generation of data scientists. * (43:27) Sakib emphasized the emergence and evolution of metadata management * (46:48) Sakib unpacked the evolution of ML infrastructure. * (49:28) Sakib examined the key trends and opportunities that will define the next wave of BI and data analytics software. * (52:56) Sakib shared his investment perspectives on climate change and student builders. * (55:29) Closing segment.

Sakib's Contact Info* Profile Page * LinkedIn * Twitter

Mentioned ContentPeople1. Sarah Catanzaro (General Partner of Amplify Partners) 2. Ed Sim (Founder of Boldstart Ventures) 3. Mike Speiser (Managing Partner of Sutter Hill Ventures)

Book* "The Idea Factory" (by Jon Gertner)

NotesMy conversation with Sakib was recorded back in late 2022. Since then, I recommend checking out these resources:

  • This blog post on the era of intelligent search
  • Bessermer's AI Roadmap and the ChatBVP bot
  • Bessemer's 2023 Cloud 100 Benchmarks Report

View Details

Show Notes* (01:56) Casber reflected on his experience growing up in China and moving to the US to pursue an undergraduate degree in business at UC Berkeley. * (04:45) Casber recalled working on his startup Etch.ai and interning at Wish during his time at Berkeley. * (10:03) Casber differentiated investing patterns for B2B and B2C startups. * (12:10) Casber reflected on his investment banking experience at Bank of America Merrill Lynch and transitioning to venture capital at Sapphire Ventures. * (17:06) Casber gave some advice for analysts who want to transition into the tech and venture industry. * (20:06) Casber provided a high-level overview of Sapphire Ventures and its investment focus. * (21:54) Casber recalled his early days as a new investor and his process of adding value to portfolio companies. * (24:58) Casber dissected his investments in the Series F round of JumpCloud and the Series B round of Uptycs (in the domain of security). * (33:03) Casber explained his investments in the Series B round of Tetrate and the Series A of Zesty (in the domain of enterprise infrastructure). * (35:59) Casber walked through his investment in the Series D round of Dremio (in the domain of data and analytics). * (39:09) Casber shared his advice to his portfolio companies in terms of navigating hiring decisions and growth strategy. * (43:06) Casber unpacked the three strategies software companies can borrow from the open-source cloud playbook. * (46:42) Casber highlighted the key trends he is most bullish on in the Open Data Ecosystem. * (51:49) Casber emphasized the importance of interoperability in the modern data stack tooling landscape. * (54:15) Casber painted the modular future of AI infrastructure. * (59:44) Casber highlighted the key trends propelling the dynamic evolution of the software development lifecycle. * (01:03:06) Casber reflected on his learning process for any new industry as an investor. * (01:05:53) Closing segment.

Casber's Contact Info* Sapphire Ventures Profile * Twitter * LinkedIn

Mentioned ContentArticles1. 3 Strategies Software Companies Can Borrow from the Open-Source Cloud Playbook (Aug 2020) 2. What is the Open Data Ecosystem and Why It's Here to Stay (April 2021) 3. The Future of AI Infrastructure is Becoming Modular: Why Best-of-Breed MLOps Solutions are Taking Off and Top Players to Watch (March 2022) 4. Evolution of the Software Development Lifecycle and the Future of DevOps (June 2022)

Books1. "The Power Law" (by Sebastian Mallaby) 2. "Engines That Move Markets" (by Alasdair Naim)

NotesMy conversation with Casber was recorded back in late 2022. Since then, I recommend checking out these resources:

  • Casber's appearance on Bloomberg News
  • Casber's analysis of the next wave of cybersecurity
  • Sapphire's investments in Huntress and Weights & Biases
  • Sapphire's new $1B fund to invest in AI-powered enterprise tech startups.

About the showDatacast features long-form, in-depth conversations with practitioners and researchers in the data community to walk through their professional journeys and unpack the lessons learned along the way. I invite guests coming from a wide range of career paths — from scientists and analysts to founders and investors — to analyze the case for using data in the real world and extract their mental models (“the WHY and the HOW”) behind their pursuits. Hopefully, these conversations can serve as valuable tools for early-stage data professionals as they navigate their own careers in the exciting data universe.

Datacast is produced and edited by James Le. For inquiries about sponsoring the podcast, email khanhle.1013@gmail.com.

Subscribe by searching for Datacast wherever you get podcasts, or click one of the links below:

  • Listen on Spotify
  • Listen on Apple Podcasts
  • Listen on Google Podcasts

If you’re new, see the podcast homepage for the most recent episodes to listen to, or browse the full guest list.

View Details

Show Notes* (01:46) Itai reflected on his education at The Hebrew University of Jerusalem, studying Math and Computer Science. * (04:18) Itai walked through his time as a software engineer at Google working in Google Trends. * (06:56) Itai emphasized the importance of a software checklist within Google's engineering culture. * (08:55) Itai explained how he became fascinated with AI/ML engineering. * (10:31) Itai touched on his period working as an AI consultant. * (13:28) Itai talked about his side hustle as a co-owner of Lia's Kitchen, a 100% vegan restaurant in Berlin. * (16:13) Itai shared the founding story of Mona Labs, whose mission is to make AI and machine learning impactful, effective, reliable, and safe for fast-growth teams and businesses. * (21:25) Itai unpacked the architecture overview of the Mona monitoring platform. * (24:50) Itai talked about the early days of Mona finding design partners. * (27:15) Itai dissected his perspective on a comprehensive monitoring strategy. * (31:42) Itai explained why the secret to comprehensive monitoring lies in granular tracking and avoiding noise. * (38:35) Itai explained how Mona can support real-time monitoring across the layers of the platform. * (43:18) Itai mentioned the integration with New Relic to display the variability of use cases for Mona. * (46:08) Itai discussed the shift for data science teams from being research-oriented to product-oriented. * (51:03) Itai provided four tactics for data science teams to become "product-oriented." * (58:04) Itai shared valuable hiring lessons to attract the right people who are excited about the mission of Mona Labs. * (01:01:46) Itai provided his mental model for finding exceptional engineering talent. * (01:03:38) Itai brought up again the importance of finding lighthouse customers. * (01:05:46) Itai gave his thoughts on building the product to satisfy different customer needs. * (01:07:55) Itai described the thriving ML engineering community in Israel. * (01:09:42) Closing thoughts

Itai's Contact Info* LinkedIn

Mona Labs' Resources* Website | LinkedIn | Twitter | YouTube * About | Customers | Careers * Platform * Blog | Case Studies | Docs

Mentioned ContentBlog Posts and Talks* We are building Mona to bring ML observability to production AI * The definitive guide to AI/ML monitoring * The secret to successful AI monitoring: Get granular, but avoid noise * Taking AI from good to great by understanding it in the real world (June 2022) * Data drift, concept drift, and how to monitor for them * The issues ML model retraining won't solve * Common pitfalls to avoid when evaluating an ML monitoring solution * Introducing automated exploratory data analysis powered by Mona * Best practices for setting up monitoring operations for your AI team * The challenges of specificity in monitoring AI * Is your LLM application ready for the public? * Overcoming cultural shifts from data science to prompt engineering

People1. Goku Mohandas (Creator of Made With ML) 2. Ville Tuulos (CEO and Co-Founder of Outerbounds) 3. Nimrod Tamir (CTO and Co-Founder of Mona Labs)

NotesMy conversation with Itai was recorded back in October 2022. Since then, Mona Labs has introduced a new self-service monitoring solution for GPT! Read Itai's blog post for the technical details.

View Details

Show Notes* (01:33) Gabi shared her professional interests growing up - from painting and drawing to graphic design. * (04:44) Gabi touched on her entrance to the field of data visualization. * (06:30) Gabi described her graduate school experience studying Data Visualization at Parsons School of Design. * (08:50) Gabi talked about the benefits of teaching data visualization classes later in her career. * (12:30) Gabi recalled working as a Data Visualization specialist at The Washington Post. * (14:50) Gabi gave her perspective on the evolution of data journalism. * (18:29) Gabi talked about her experience co-founding Raw Haus, a creative community bringing together emerging talent in design, technology, and entrepreneurship. * (21:20) Gabi emphasized the magic of community gatherings. * (23:51) Gabi reflected on her time at WeWork as a senior data visualization engineer to design and build graphics, dashboards, and tools that tell stories using data. * (28:02) Gabi walked through the evolution of the Data Cult initiative - which she co-created with Leah Weiss. * (31:44) Gabi recalled her decision to leave WeWork in early 2020 and start Data Culture - a data engineering and visualization consultancy focused on helping organizations build data capabilities, implement modern infrastructure and create lasting data culture. * (36:31) Gabi unpacked Data Culture's well-defined blueprint for each client engagement. * (39:21) Gabi explained how Data Culture leveraged tools in the modern data stack for its consulting services. * (40:42) Gabi brought up Data Culture's Studio - which offers data storytelling and visualization services to mission-aligned organizations. * (42:50) Gabi reviewed her experience working with Kode with Klossy to empower young scholars to solve important issues using data science. * (46:13) Gabi shared her perspective on how companies can scale their respective data cultures. * (49:52) Gabi shared the story behind the founding of Preql. * (52:59) Gabi touched on the process of working with design partners for Preql. * (56:42) Gabi shared valuable hiring lessons to attract the right people at Data Culture and Preql. * (58:10) Gabi provided her perspective on building a diverse team. * (01:00:57) Gabi shared fundraising advice to data founders who are seeking the right investors for their startups. * (01:03:34) Closing segment.

Gabi's Contact Info* LinkedIn * Twitter

Preql's Resources* Website | Twitter | LinkedIn * Introducing Preql: The Future of Data Transformation (April 2022)

Mentioned ResourcesPeople1. Umi Syam (Graphics and Multimedia Editor at the New York Times) 2. Giorgia Lupi (Information Designer and Partner at Pentagram) 3. Susie Lu (Senior Data Visualization Engineer at Netflix)

Book* Invisible Women: Exploring Data Bias in a World Designed for Men (by Caroline Criado Perez)

NotesMy conversation with Gabi was recorded back in August 2022. Since then, Preql has officially launched and currently supports strategy and operations teams at B2B and vertical SaaS companies!

About the showDatacast features long-form, in-depth conversations with practitioners and researchers in the data community to walk through their professional journeys and unpack the lessons learned along the way. I invite guests coming from a wide range of career paths — from scientists and analysts to founders and investors — to analyze the case for using data in the real world and extract their mental models (“the WHY and the HOW”) behind their pursuits. Hopefully, these conversations can serve as valuable tools for early-stage data professionals as they navigate their own careers in the exciting data universe.

Datacast is produced and edited by James Le. For inquiries about sponsoring the podcast, email khanhle.1013@gmail.com.

Subscribe by searching for Datacast wherever you get podcasts, or click one of the links below:

  • Listen on Spotify
  • Listen on Apple Podcasts
  • Listen on Google Podcasts

If you’re new, see the podcast homepage for the most recent episodes to listen to, or browse the full guest list.

View Details

Show Notes* (01:55) Alex reflected on his upbringing as an immigrant moving from Colombia to the US at 14. * (07:06) Alex recalled his undergraduate experience at NYU’s Polytechnic School of Engineering. where he study Computer Science and do research in cryptography. * (16:40) Alex went over his first job working as a software engineer at FactSet Research System. * (20:13) Alex walked through his time as the first employee and the first engineer at YieldMo. * (24:30) Alex talked about his hiring philosophy for engineers who care about their craft. * (28:03) Alex touched on the backstory behind the creation of Concord, with Shinji Kim and Robert Blafford, while working at YieldMo. * (32:26) Alex shared lessons learned from his first-time founder experience with Concord. * (35:22) Alex went over his two years at Akamai as a Platform Infrastructure Engineer after the Concord acquisition. * (40:01) Alex introduced his work on SMF, an RPC framework designed for microsecond tail latency. * (43:41) Alex shared the story behind the founding of Redpanda Data, which builds a high-performance, Apache Kafka-compatible data streaming platform for mission-critical workloads. * (47:19) Alex walked through the major benefits of choosing Redpanda over Kafka. * (51:03) Alex explained his decision to open-source Redpanda in November 2020 under the Source Available License BSL. * (56:08) Alex mentioned successful tactics his team employed in order to raise the adoption and contribution to the open-source library. * (01:01:13) Alex unpacked the design of Redpanda's Intelligent Data API. * (01:08:55) Alex provided his perspective on the modern streaming data architecture. * (01:13:24) Alex shared valuable hiring lessons to attract the right people who are excited about Redpanda’s mission. * (01:18:30) Alex talked about his experience choosing customers for Redpanda. * (01:20:33) Alex shared fundraising advice to founders who are seeking the right investors for their startups. * (01:23:23) Alex gave advice to a smart, driven minority who aspires to work on ambitious, technically deep, and challenging problems. * (01:28:18) Closing segment.

Alex's Contact Info* LinkedIn * Twitter * Website * GitHub

Redpanda's Resources* Website | Twitter | LinkedIn | Slack | GitHub | Contributing Doc * About Redpanda | Platform Capabilities | Customers * Docs | Redpanda University * Reports and Guides | Benchmarks * Hack The Planet Scholarship

Mentioned ContentBlog Posts* Redpanda raison d'etre (Feb 2019) * Thread-per-core buffer management for a modern Kafka-API storage system (Sep 2020) * Redpanda is now free and Source Available (Nov 2020) * Redpanda creates Redpanda, the Intelligent Data API Platform, backed by $15.5M initial funding from Lightspeed Venture Partners and GV (Jan 2021) * The Intelligent Data API (Jan 2021) * Redpanda Wasm engine architecture (June 2021) * We raised an additional $50M to drive the future of streaming data. Join us! (Feb 2022) * Redpanda gives Kafka a Run for Its Money (InfoWorld, May 2022) * Alex Gallego Builds Redpanda To Simplify And Unify Real-Time Streaming Data (Forbes, June 2022)

Talks* Distributed Stream Processing over thousands of Datacenters (GeeCON, Aug 2017) * How to Build the Fastest RPC (Nov 2017) * Co-designing Raft + thread-per-core execution model for the Kafka-API (Dec 2021)

People1. Andy Pavlo 2. Leslie Lamport 3. Kyle Kingsbury

NotesMy conversation with Alex was recorded back in August 2022. Since then, I recommend checking out these resources:

  1. The $100M Series C funding announcement
  2. This guide for developers on streaming data
  3. Customer case studies with Lacework, Exein, and SmartLunch
  4. Resources on the advantage of Redpanda over Apache Kafka (cost of ownership comparison, data sovereignty, and this holistic comparison)

View Details

Show Notes* (01:44) Chetan reflected on his undergraduate experience at Stanford studying Electrical Engineering and Statistics back in the late 2000s. * (06:10) Chetan recalled his experience interning at IBM and Quantcast and doing research at Stanford Center for Minds, Brain, and Computation. * (08:41) Chetan talked about his first job working as a research analyst focused on healthcare policy at Acumen. * (11:15) Chetan walked through his decision to join Airbnb as their 4th data scientist and work on building Airbnb's original ETL framework for online risk mitigation. * (15:12) Chetan recalled the early state of data science at Airbnb. * (18:17) Chetan touched on the development of Airbnb's knowledge management and sharing platform called Knowledge Repo. * (23:10) Chetan explained why an experimentation program is the most impactful thing a data team can do. * (26:06) Chetan walked through the evolution of Airbnb's experimentation platform since its inception in 2014. * (31:24) Chetan recalled fond memories from taking a year off from work to travel. * (35:16) Chetan touched on his transition back to work by way of living in Atlanta and co-founding a logistics software startup called Saltbox. * (39:28) Chetan described his time as a data scientist at Webflow, building their experimentation system from scratch. * (42:48) Chetan shared the story behind the founding of Eppo. * (46:12) Chetan dissected the key capabilities that are baked into the Eppo product. * (48:32) Chetan dived deeper into the problems caused by long experiment durations and the benefits of using CUPED to bend time in experiments. * (52:04) Chetan talked about the role of a statistics engineer. * (54:45) Chetan shared his perspective on the role of experimentation in the Modern Data Stack and the Modern Growth Stack. * (01:00:47) Chetan discussed the core elements of the modern experimentation stack. * (01:04:54) Chetan talked about the experiment overhead. * (01:06:24) Chetan emphasized the designer gap in experimentation tools * (01:08:50) Chetan shared his thoughts about metric strategy. * (01:10:51) Chetan shared valuable hiring lessons to attract the right people who are excited about Eppo's mission. * (01:14:10) Chetan provided his perspectives on finding design partners for an early-stage startup. * (01:17:04) Chetan shared fundraising advice to founders who are seeking the right investors for their startups. * (01:19:40) Closing segment.

Chetan's Contact Info* LinkedIn * Twitter * GitHub * AngelList

Eppo's Resources* Website | Twitter | LinkedIn * Blog | Updates | Doc * About | Careers * Experimentation Product * Feature Flagging Product

Mentioned ContentArticles* Travel Year Facts and Superlatives (Dec 2019) * Why I Started Eppo (Feb 2021) * Reducing Experiment Durations (June 2021) * The Designer Gap in Experimentation Tools (June 2021) * We're Hiring A Statistics Engineer! (Aug 2021) * Should You Always Run An Experiment? (Aug 2021) * Stop Micromanaging Product Strategy (Sep 2021) * The most impactful thing a data team can do is establish an experimentation program (Dec 2021) * Bending Time in Experimentation (June 2022) * We Raised $19.5M! (June 2022) * Experimentation for the Modern Growth Stack: Our Investment in Eppo (June 2022)

People1. Mike Kaminsky 2. Sean Taylor 3. Jeremy Howard

Book* The Mom's Test (by Rob Fitzpatrick)

NotesMy conversation with Chetan was recorded back in August 2022. Since then, Eppo has launched feature flagging, and now offers the first "flags on top of your warehouse" experimentation platform. They also have Miro, Twitch, DraftKings, and Zapier as customers.

About the showDatacast features long-form, in-depth conversations with practitioners and researchers in the data community to walk through their professional journeys and unpack the lessons learned along the way. I invite guests coming from a wide range of career paths — from scientists and analysts to founders and investors — to analyze the case for using data in the real world and extract their mental models (“the WHY and the HOW”) behind their pursuits. Hopefully, these conversations can serve as valuable tools for early-stage data professionals as they navigate their own careers in the exciting data universe.

Datacast is produced and edited by James Le. For inquiries about sponsoring the podcast, email khanhle.1013@gmail.com.

Subscribe by searching for Datacast wherever you get podcasts, or click one of the links below:

  • Listen on Spotify
  • Listen on Apple Podcasts
  • Listen on Google Podcasts

If you’re new, see the podcast homepage for the most recent episodes to listen to, or browse the full guest list.

View Details

Show Notes* (01:38) Chad reflected on his early career as a freelance journalist working in Southeast Asia. * (04:27) Chad explained the benefits of writing for anyone working in a technical field. * (06:15) Chad touched on his entrance to data analytics through Conversion Rate Optimization. * (09:06) Chad walked through his decision to dive deep into the field of experimentation. * (13:28) Chad recalled how he spent time learning the basics of statistics. * (15:32) Chad discussed the differences in experimentation cultures at Subway, SEPHORA, and Microsoft. * (19:11) Chad shared the technical details behind the evolution of Convoy's data platform since he joined in 2019. * (23:05) Chad emphasized the role of data in Convoy's digital freight business. * (26:17) Chad brought up the importance of solving data discovery at Convoy and their decision to choose Amundsen. * (29:02) Chad shared lessons learned setting up a flexible experimentation platform at Convoy. * (32:46) Chad unpacked the problems with Change Data Capture and how his team built an internal change management platform called Chassis (a source of truth for definitions of events, entities, and relationships.). * (41:33) Chad discussed the existential threat of data quality. * (44:38) Chad explainedwhy the modern data warehouse is broken and why the "Immutable Data Warehouse" can be a solution. * (51:39) Chad zoomed in on the death of data modeling. * (57:42) Chad is bullish on the rise of the knowledge layer and data contracts in the upcoming years. * (01:03:37) Chad talked at length about the data collaboration problem. * (01:09:28) Chad gave advice for data organizations to be more customer-centric. * (01:11:55) Chad shared components of a high-quality Data UX function that any centralized data team should consider when developing data experiences. * (01:14:46) Chad touched on his mental framework for evaluating potential investments in the data space. * (01:18:57) Chad brought up the valuable skills he acquired as an internal product manager. * (01:20:05) Closing segment.

Chad's Contact Info* LinkedIn * Data Products Substack * Data Quality Camp

Mentioned ContentTalks* Aligning Experimentation Across Product Development and Marketing (CXL Live 2019) * Chassis: Entities, Events, and Change Management (Data Quality Meetup, 2021) * 1,000 Experiments Club with AB Tasty (July 2021) * Data Discovery at Lyft and Convoy (July 2021) (with Mark Grover) * The growth of the data platform product manager role (The Tech Trek, Dec 2021) * Implementing Amundsen at Convoy (Building the Backend, Jan 2022) * Getting ROI from Experimentation: How AB Experimentation plays out in Organizations (Data Council, March 2022) * Why are we so bad at this modern data stack? (Catalog and Cocktails, April 2022)

Articles* Experimentation not only protects your KPIs but your job as well (Dec 2019) * Is The Modern Data Warehouse Broken? (April 2022) (with Barr Moses) * The Existential Threat of Data Quality (May 2022) * The Death of Data Modeling (June 2022) * Data Collaboration Problem (June 2022) * The Rise of Data Contracts (Aug 2022)

People1. Barr Moses (Monte Carlo Data) 2. Juan Sequeda (data.world) 3. Adrian Kreuziger (Convoy)

Book* Agile Data Warehouse Design (by Lawrence Corr)

NotesMy conversation with Chad was recorded back in July 2022. Since then, I'd recommend looking at:

  • His two-part series on engineering guide to data contracts (Part 1 and Part 2)
  • The Data Quality Camp community
  • His most recent post on how scale kills data teams
  • Data Facade (of which he is an angel investor)

View Details

Show Notes* (01:56) Curtis reflected on his upbringing in rural Kentucky and his gift of education. * (07:20) Curtis explained how he cultivated mental focus and intellectual fortitude while growing up in Kentucky. * (10:30) Curtis shared his view regarding online misinformation on social media. * (14:27) Curtis recalled his undergraduate experience at Vanderbilt University in the early 2010s. * (22:39) Curtis explained how he learned best via teaching and mentoring. * (24:04) Curtis walked through the research and industry experiences he obtained throughout college. * (32:45) Curtis recalled his decision to embark on a Ph.D. in Computer Science at MIT. * (38:53) Curtis told the story of how he ended up finding his advisor - Professor Isaac Chuang (the inventor of the first working quantum computer). * (40:36) Curtis mentioned how he invented the CAMEO Detection Algorithm to detect “multiple-account” cheating in massive open online courses. * (44:47) Curtis unpacked his Ph.D. research on dataset uncertainty estimation. * (50:08) Curtis dissected confident learning, a family of theories and algorithms for supervised ML with label errors. * (53:22) Curtis encapsulated how he strategically iterated cleanlab at his various graduate internships. * (01:00:22) Curtis recalled his time founding his first startup ChipBrain, before founding Cleanlab. * (01:06:42) Curtis brought up the creation of the labelerrors.com project. * (01:12:12) Curtis provided lessons learned as a second-time founder. * (01:14:25) Curtis elaborated on the open-source roadmap of cleanlab. * (01:17:08) Curtis highlighted the key capabilities of Cleanlab Studio - the no-code, automatic data correction solution for data and engineering teams with robust enterprise features. * (01:18:50) Curtis touched on Cleanlab Vizzy - an interactive visualization of confident learning. * (01:20:29) Curtis shared valuable hiring lessons to attract the right people who are excited about Cleanlab’s mission. * (01:23:23) Curtis gave his thoughts on shaping Cleanlab’s culture. * (01:26:06) Curtis explained the similarity and differences between being a founder and a researcher. * (01:29:09) Curtis mentioned how he had helped researchers build affordable state-of-the-art deep learning machines. * (01:31:46) Curtis brought up his alter ego PomDP the Ph.D. rapper, and how rapping has been an outlet for him to express emotions and creativity. * (01:40:12) Curtis emphasized how his success had been due to a function of grit, resourcefulness, and friends made along the way. * (01:44:04) Closing segment.

Curtis' Contact Info* Academic Website * LinkedIn | Twitter | Facebook | Instagram * Google Scholar | GitHub * PhD Rapper (YouTube | Spotify | SoundCloud | Facebook | Twitter | Instagram) * L7 Machine Learning Blog

Cleanlab's Resources* Website | GitHub | Slack | Twitter | LinkedIn * Blog | Research | Doc * About | Careers * Cleanlab Studio * Cleanlab Vizzy * The Cleanlab Culture

Mentioned ContentPapers* Detecting and preventing “multiple-account” cheating in massive open online courses, Curtis G. Northcutt, Andrew Ho, & Isaac L. Chuang, Computers & Education, 2016. [paper | code | arXiv] * Comment Ranking Diversification in Forum Discussions, Curtis G. Northcutt, Kimberly Leon, & Naichun Chen, Learning at Scale, 2017. [paper | code | free-access] * Learning with Confident Examples: Rank Pruning for Robust Classification with Noisy Labels, Curtis G. Northcutt, Tailin Wu, & Isaac L. Chuang, 33rd Conference on Uncertainty in Artificial Intelligence (UAI 2017). [paper | code] * Confident Learning: Estimating Uncertainty for Dataset Labels, Curtis G. Northcutt, Lu Jiang, & Isaac L. Chuang, Journal of Artificial Intelligence Research (JAIR), Vol. 70 (2021). [paper | code | blog] * Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks, Curtis Northcutt, Anish Athalye, and Jonas Mueller, 35th Conference on Neural Information Processing Systems (NeurIPS 2021) Track on Datasets and Benchmarks [paper| demo | code | blog]

Blog Posts* Founder’s Medal recipient chooses MIT over Microsoft (May 2013) * Build a Pro Deep Learning Workstation... for Half the Price (Feb 2019) * An Introduction to Confident Learning: Finding and Learning with Label Errors in Datasets (Nov 2019) * Announcing cleanlab: a Python Package for ML and Deep Learning on Datasets with Label Errors (Nov 2019) * Double Deep Learning Speed by Changing the Position of your GPUs (Dec 2019) * Benchmarking: Which GPU for Deep Learning? (Dec 2019) * The Best 4-GPU Deep Learning Rig only costs $7000 not $11,000 (April 2020) * Pervasive Label Errors in ML Datasets Destabilize Benchmarks (March 2021) * Cleanlab: The History, Present, and Future (April 2022) * cleanlab 2.0: Automatically Find Errors in ML Datasets (April 2022) * How We Built Cleanlab Vizzy (August 2022)

Talks and Podcasts* Tedx Talk: The MIT Rap Challenge (July 2020) * Talk at NLP Summit (March 2022) * Talk at Data + AI Summit (June 2022) * MLOps Coffee Chat (July 2022) * Talk at Snorkel's Future of Data-Centric AI Conference (July 2022) * Open-Source Startup Podcast (March 2023)

People1. Leslie Kaelbling 2. Geoff Hinton 3. Jeff Dean

Book* Play Bigger: How Pirates, Dreamers, and Innovators Create and Dominate Markets (by Al Ramadan, Dave Peterson, Chris Lockhead, and Kevin Maney)

NotesMy conversation with Curtis was recorded back in August 2022. The Cleanlab team has had some important announcements in 2023 that I recommend looking at:

  1. The launches of CleanVision, Datalab, and ActiveLab
  2. This blog post on using Cleanlab to improve LLMs
  3. His new single "Clarity In My Vision"
  4. Cleanlab's partnership with Databricks (Video)

Cleanlab is about to announce its Series A announcement soon. Stay on the look for it!

About the showDatacast features long-form, in-depth conversations with practitioners and researchers in the data community to walk through their professional journeys and unpack the lessons learned along the way. I invite guests coming from a wide range of career paths — from scientists and analysts to founders and investors — to analyze the case for using data in the real world and extract their mental models (“the WHY and the HOW”) behind their pursuits. Hopefully, these conversations can serve as valuable tools for early-stage data professionals as they navigate their own careers in the exciting data universe.

Datacast is produced and edited by James Le. For inquiries about sponsoring the podcast, email khanhle.1013@gmail.com.

Subscribe by searching for Datacast wherever you get podcasts, or click one of the links below:

  • Listen on Spotify
  • Listen on Apple Podcasts
  • Listen on Google Podcasts

If you’re new, see the podcast homepage for the most recent episodes to listen to, or browse the full guest list.

View Details

Show Notes* (01:41) Frank shared formative experiences of his upbringing moving from China to the US. * (04:45) Frank described his overall academic experience at Stanford, studying Electrical Engineering with a minor in Computer Science. * (08:41) Frank talked about his research and industry experience while at Stanford. * (11:34) Frank shared his proudest accomplishments working at Yahoo as a research engineer in the Vision and Machine Learning group. * (16:37) Frank went over his experience co-founding a company that developed indoor localization and navigation solutions called Orion. * (23:06) Frank walked through his decision to leave Silicon Valley for China. * (26:02) Frank talked about his experience living and doing business in China (check out his two-part blog series that has covered normal life and the pandemic story in China). * (32:44) Frank elaborated on the work culture differences between the East and the West. * (37:58) Frank reflected on his decision to join Zilliz back in August 2021. * (42:55) Frank unpacked the notion of vector databases for the un-initiated. * (47:44) Frank provided a brief overview on the high-level design of Milvus, Zilliz's advanced open-source vector database solution. * (51:38) Frank highlighted three unique use cases of Milvus - malware detection, reverse image search, and drug discovery. * (56:51) Frank introduced Towhee, an open-source project that helps software engineers develop and deploy applications that utilize embeddings in just a few lines of code. * (01:01:59) Frank anticipated the evolution of the embedding tooling landscape to support the increasing adoption of unstructured data. * (01:04:21) Frank gave a primer on Zilliz Cloud, Zilliz's enterprise vector database solution. * (01:06:30) Closing segment.

Frank's Contact Info* LinkedIn * Twitter * GitHub * Website

Zilliz's Resources* Website | Twitter | LinkedIn | GitHub | YouTube * Zilliz Cloud Database * Milvus (Docs | GitHub) * Towhee (Docs | GitHub)

Mentioned ContentArticles and Presentations* A Gentle Introduction to Vector Databases (Dec 2021) * My Experience Living and Working in China, Part I (Feb 2022) * My Experience Living and Working in China, Part II (March 2022) * Making ML More Accessible for Application Developers (April 2022) * Understanding Neural Network Embeddings (April 2022) * Building An Open-Source Platform for Generating Embedding Vectors (Berlin Buzzwords, 2022)

People1. Yann LeCun (Chief AI Scientist at Meta, Professor at NYU) 2. Yangqing Jia (Creator of the Caffe deep learning framework) 3. Soumith Chintala (Creator of the PyTorch deep learning framework)

Book* A Short History of Nearly Everything (by Bill Bryson)

NotesMy conversation with Frank was recorded back in August 2022. The Zilliz team has had some important announcements in 2023 that I recommend looking at:

  1. The landing page of Zilliz Cloud
  2. The beta launch of Milvus 2.3
  3. The development of GPTCache
  4. The OSS Chat demo application

About the showDatacast features long-form, in-depth conversations with practitioners and researchers in the data community to walk through their professional journeys and unpack the lessons learned along the way. I invite guests coming from a wide range of career paths — from scientists and analysts to founders and investors — to analyze the case for using data in the real world and extract their mental models (“the WHY and the HOW”) behind their pursuits. Hopefully, these conversations can serve as valuable tools for early-stage data professionals as they navigate their own careers in the exciting data universe.

Datacast is produced and edited by James Le. For inquiries about sponsoring the podcast, email khanhle.1013@gmail.com.

Subscribe by searching for Datacast wherever you get podcasts, or click one of the links below:

  • Listen on Spotify
  • Listen on Apple Podcasts
  • Listen on Google Podcasts

If you’re new, see the podcast homepage for the most recent episodes to listen to, or browse the full guest list.

View Details

Show Notes (01:58) Vinoth shared his college experience studying IT at the Madras Institute of Technology in Chennai, India. * (07:09) Vinoth reflected on his time at UT Austin, getting a Master's degree in Computer Science - where he did research on high-bandwidth content distribution and large-scale parallel processing with shell pipes. * (11:20) Vinoth recalled his two years as a software engineer at Oracle, working on their database replication engine, HPC, and stream processing. * (15:30) Vinoth walked over his transition to LinkedIn as a senior software engineer, working primarily on Voldemort - a key-value store that handles a big chunk of traffic on Linkedin and serves thousands of requests per second over terabytes of data. * (24:41) Vinoth talked about his career transition to Uber in late 2014 as a founding engineer on Uber's data team and architect of Uber's data architecture. * (28:39) Vinoth reflected on the state of Uber's data infrastructure when he joined. * (34:31) Vinoth elaborated on Uber's case for incremental processing on Hadoop. * (38:53) Vinoth reviewed the initial design and implementation of Hudi across the Hadoop ecosystem at Uber in 2016. * (41:33) Vinoth shared t*he evolution of Hudi after it was initially open-sourced by Uber in 2017 and eventually incubated into the Apache Software Foundation in 2019. * (46:49) Vinoth explained how to keep the development of Apache Hudi vendor-neutral. * (49:36) Vinoth provided lessons learned about establishing standards for open-source data projects. * (53:45) Vinoth went over the valuable leadership lessons that he absorbed throughout his 4.5 years at Uber. * (57:17) Vinoth reflected on his 1.5 years as a principal engineer at Confluent working on ksqlDB, which makes it easy to create event streaming applications. * (01:02:16) Vinoth articulated the vision for Apache Hudi as a Streaming Data Lake platform. * (01:08:00) Vinoth highlighted the challenges with databases around indexing and concurrency control. * (01:11:37) Vinoth shared the unique challenges around prioritizing the Hudi roadmap and engaging an open-source community. * (01:16:32) Vinoth shared the founding story of Onehouse, a cloud-native, fully-managed lakehouse service built on Apache Hudi. * (01:22:02 ) Vinoth emphasized Onehouse's commitment towards openness. * (01:24:36) Vinoth shared valuable hiring lessons to attract the right people who are excited about Onehouse's mission. * (01:26:40) Vinoth shared fundraising advice to founders who are seeking the right investors for their startups. * (01:28:24) Closing segment.

Vinoth's Contact Info* LinkedIn * Twitter

Onehouse's Resources* Website | Twitter | LinkedIn * About | Product | Blog | Careers

Apache Hudi's Resources* User Docs | Technical Wiki | Roadmap * GitHub | Twitter | Slack

Mentioned ContentArticles and Presentations* Voldemort : Prototype to Production (May 2014) * Uber's Case for Incremental Processing on Hadoop (Aug 2016) * Hoodie: An Open Source Incremental Processing Framework From Uber (2017) * The Past, Present, and Future of Efficient Data Lake Architectures (2021) * Highly Available, Fault-Tolerant Pull Queries in ksqlDB (May 2020) * Apache Hudi - The Data Lake Platform (July 2021) * Introducing Onehouse (Feb 2022) * Automagic Data Lake Infrastructure (Feb 2022) * Onehouse Commitment to Openness (Feb 2022)

People* Leslie Lamport * Jeff Dean * Michael Stonebreaker

Book* Zero To One (by Peter Thiel)

NotesMy conversation with Vinoth was recorded back in August 2022. The Onehouse team has had some announcements in 2023 that I recommend looking at:

  • The Launch Announcement of Onetable
  • The $25M Series A Funding Announcement
  • Onehouse Availability in AWS Marketplace
  • Onehouse Product Demo on building a data lake for GitHub analytics at scale
  • Walmart's recent study on different open-source data lakehouse formats
  • This discussion around the Hudi 1.x vision

About the showDatacast features long-form, in-depth conversations with practitioners and researchers in the data community to walk through their professional journeys and unpack the lessons learned along the way. I invite guests coming from a wide range of career paths — from scientists and analysts to founders and investors — to analyze the case for using data in the real world and extract their mental models (“the WHY and the HOW”) behind their pursuits. Hopefully, these conversations can serve as valuable tools for early-stage data professionals as they navigate their own careers in the exciting data universe.

Datacast is produced and edited by James Le. For inquiries about sponsoring the podcast, email khanhle.1013@gmail.com.

Subscribe by searching for Datacast wherever you get podcasts, or click one of the links below:

  • Listen on Spotify
  • Listen on Apple Podcasts
  • Listen on Google Podcasts

If you’re new, see the podcast homepage for the most recent episodes to listen to, or browse the full guest list.

View Details

Show Notes* (01:55) Alexa shared formative experiences of her upbringing in Philadelphia. * (03:47) Alexa reflected on her undergraduate experience at Vanderbilt studying Engineering Science. * (05:49) Alexa recalled her first job out of college in management consulting at KPMG. * (08:20) Alexa walked over her transition from consulting to technology when she joined the Sales Operations team at Dataminr. * (12:35) Alexa talked about her proudest accomplishments at Dataminr - seeding the initial idea for Pocus and building a community for women in the workplace. * (20:23) Alexa reflected on her MBA experience at the Stanford Graduate School of Business. * (24:05) Alexa elaborated on the mindset difference between investing and operating. * (25:27) Alexa briefly touched on her internship at Monte Carlo. * (27:58) Alexa shared the founding story of Pocus. * (32:27) Alexa unpacked the concept of Product-Led Sales as a GTM approach. * (35:40) Alexa provided two example use cases of Pocus. * (39:35) Alexa explained the concepts of Product-Qualified Leads and Sales-Assist. * (42:20) Alexa discussed the long-term vision of Pocus' product roadmap. * (45:33) Alexa shared valuable hiring lessons to attract the right people who are aligned to Pocus' values. * (51:15) Alexa went over the journey of building the Product-Led Sales community. * (54:54) Alexa shared the unique opportunities of evolving a category, a community, and a product all at once. * (57:56) Alexa shared fundraising advice to founders who are seeking the right investors for their startups. * (01:01:15 ) Alexa provided advice to a smart, driven female operator who wants to take the leap of founding her company. * (01:03:09) Closing segment.

Alexa' Contact Info* LinkedIn * Twitter

Pocus' Resources* Website | Twitter | LinkedIn | YouTube * About | Product | Blog | Careers * Community | Newsletter

Mentioned ContentBlog Posts* What is Product-Led Sales? (July 2022) * The Myth of "No Sales" at PLG Companies (July 2021) * When To Add A Sales Team to Your PLG Company (Sep 2021) * The Definitive PQL Guide: Part 1, Part 2, Part 3 (Nov 2021) * What Is The Sales-Assist Role? (Nov 2021) * Introducing Pocus' PLS Platform (Nov 2021) * Product-Led Sales Community Wisdom Highlights 2021 (Dec 2021) * Notes on Community-Led Category Creation with Pocus' Co-Founder, Alexa Grabell (Feb 2022) * Sneak Peek at Pocus' PLS Platform (March 2022) * Announcing $23M to Transform How GTM Teams Use Data to Drive Revenue (June 2022) * Year One: The Product-Led Sales Platform is Here to Stay (July 2022)

People* Kyle Poyar (OpenView Ventures) * Melissa Ross (Clockwise) * Aaron Geller (QuickNode)

NotesMy conversation with Alexa was recorded back in July 2022. The Pocus team has had some announcements in 2023 that I recommend looking at:

  1. The launch announcement of Pocus' Revenue Data Platform
  2. The Product-Led Sales Playbook Volume 2
  3. The Unlocking Revenue podcast
  4. The Playbook Library for product-led go-to-market

About the showDatacast features long-form, in-depth conversations with practitioners and researchers in the data community to walk through their professional journeys and unpack the lessons learned along the way. I invite guests coming from a wide range of career paths — from scientists and analysts to founders and investors — to analyze the case for using data in the real world and extract their mental models (“the WHY and the HOW”) behind their pursuits. Hopefully, these conversations can serve as valuable tools for early-stage data professionals as they navigate their own careers in the exciting data universe.

Datacast is produced and edited by James Le. For inquiries about sponsoring the podcast, email khanhle.1013@gmail.com.

Subscribe by searching for Datacast wherever you get podcasts, or click one of the links below:

  • Listen on Spotify
  • Listen on Apple Podcasts
  • Listen on Google Podcasts

If you’re new, see the podcast homepage for the most recent episodes to listen to, or browse the full guest list.

View Details

Show Notes* (02:06) Carlos shared formative experiences of his upbringing tinkering with robots and websites. * (04:03) Carlos reflected on his education, studying Mechanical and Aerospace Engineering at Cornell University. * (05:34) Carlos discussed the technical details of his research on machine learning applications in robotics and art. * (10:11) Carlos explained his work as a robotic system analyst at Kiva Systems. * (15:41) Carlos discussed building his first data product at Kiva. * (20:24) Carlos recalled his stint working on warehouse-automating distributed robots at Amazon Robotics (after the Kiva acquisition). * (24:31) Carlos revealed his decision in 2013 to join an early-stage healthcare startup called Flatiron Health as the first data hire. * (28:43) Carlos shared his experience building Flatiron's Data Insights team from scratch. * (31:51) Carlos reviewed different data products built and deployed at Flatiron Health. * (38:41) Carlos shared the key learnings from hiring for his data team at Flatiron. * (44:08) Carlos shared the founding story of Glean, which is building a new way to make data exploration and visualization accessible to everyone. * (50:52) Carlos explained the pain points in data visualization/exploration and the product features of Glean that address them. * (55:03) Carlos dissected Glean DataOps, which brings modern developer workflow to the business intelligence layer and prevents broken dashboards. * (59:28) Carlos outlined the long-term product vision for Glean. * (01:03:11) Carlos shared valuable hiring lessons to attract the right people who are excited about Glean's mission. * (01:07:15) Carlos discussed his team's challenges in finding the early design partners. * (01:10:13) Carlos shared fundraising advice to founders who are seeking the right investors for their startups. * (01:11:57) Closing segment.

Carlos' Contact Info* Twitter * LinkedIn * GitHub * Website * Medium

Glean's Resources* Website | Twitter | LinkedIn * About | Docs | Blog * Interactive Public Demo | DataOps

Mentioned ContentBlog Posts* How the Data Insights team helps Flatiron build useful data products (May 2018) * The biggest mistake making your first data hire: not interviewing for product (July 2020) * How to interview your first data hire (Aug 2020) * My hack for getting started with data as a product (May 2021) * Introducing Glean (March 2022) * Your dashboard is probably broken (April 2022)

People1. Vicki Boykis 2. Anthony Goldbloom 3. Wes McKinney

Book* The Toyota Way: 14 Management Principles from the World's Greatest Manufacturer (by Jeffrey Liker)

NotesMy conversation with Carlos was recorded back in June 2022. The Glean team has had some announcements in 2023 that I recommend looking at:

  1. The recently launched, interactive public demo site
  2. This recent integration with DuckDB
  3. This post about Version Control for BI
  4. Their Public Roadmap

About the showDatacast features long-form, in-depth conversations with practitioners and researchers in the data community to walk through their professional journeys and unpack the lessons learned along the way. I invite guests coming from a wide range of career paths — from scientists and analysts to founders and investors — to analyze the case for using data in the real world and extract their mental models (“the WHY and the HOW”) behind their pursuits. Hopefully, these conversations can serve as valuable tools for early-stage data professionals as they navigate their own careers in the exciting data universe.

Datacast is produced and edited by James Le. For inquiries about sponsoring the podcast, email khanhle.1013@gmail.com.

Subscribe by searching for Datacast wherever you get podcasts, or click one of the links below:

  • Listen on Spotify
  • Listen on Apple Podcasts
  • Listen on Google Podcasts

If you’re new, see the podcast homepage for the most recent episodes to listen to, or browse the full guest list.

View Details

Show Notes* (01:49) Shruti shared her upbringing in India - where she studied Engineering and Computer Science in the early 2000s. * (03:11) Shruti reflected on her early career as a software engineer at Hewlett-Packard and IBM. * (07:29) Shruti recalled the early days of cloud computing. * (09:01) Shruti reflected on her time pursuing an MBA at UCLA Anderson School of Management. * (11:55) Shruti explained her shift from software engineering to product management. * (14:19) Shruti revisited her years at VMware as a product line manager for cloud infrastructure - owning all aspects of go-to-market strategy and execution for VMware's entire software-defined storage portfolio. * (18:30) Shruti talked about her time as the VP of Marketing at Ravello Systems - growing the business from zero customers to a successful multi-million dollar acquisition by Oracle. * (23:07) Shruti went over her time as a senior director of product management for Oracle's cloud portfolio. * (27:20) Shruti recalled the founding story of Rockset - where she is a co-founder and Chief Product Officer. * (30:40) Shruti explained the concepts of real-time analytics and data applications for the uninitiated. * (37:31) Shruti unpacked the high-level design of Rockset architecture - which brings together cloud-native architecture, schemaless ingestion, converged indexing, and full-featured SQL. * (40:23) Shruti elaborated on the concept of converged indexing. * (42:43) Shruti dissected the technology requirements and the key layers of "the modern real-time data stack." * (46:17) Shruti talked about the role of partnerships in Rockset's product strategy. * (51:29) Shruti highlighted some of Rockset's customer use cases. * (56:06) Shruti shared valuable hiring lessons to attract high-integrity and diverse people for Rockset. * (58:51) Shruti shared her take on interviewing on strengths over weaknesses. * (01:01:32) Shruti shared the strategy Rockset used to find design partners in the early days. * (01:05:12) Shruti shared the tactics to combine the power of product-led adoption with sales-driven growth for rapidly scaling Rockset's business. * (01:08:45) Shruti shared fundraising advice to founders who are seeking the right investors for their startups. * (01:10:27) Shruti described the evolution of enterprise marketing and GTM strategy in the past decade. * (01:12:59) Closing segment.

Shruti's Contact Info* LinkedIn * Twitter * Forbes

Rockset's Resources* Website | Twitter | LinkedIn | Facebook * Docs | Blog | Community * Product | Architecture | Customers * Real-Time Analytics Explained * What Is A Data Application?

Mentioned ContentArticles* "Building Data Applications Powered by Real-Time Analytics" (May 2021) * "How startups can create a culture where women can win" (May 2021) * "Streaming Data and the Modern Real-Time Data Stack" (Nov 2021)

People* Barr Moses (Monte Carlo Data) * Jay Kreps (Confluent) * Alex DeBrie (DynamoDB Expert)

Book* Competing Against Luck (by Clayton Christensen)

NotesMy conversation with Shruti was recorded back in June 2022. Since then, a lot has happened. I recommend looking at the resources below:

  • The launch of compute-compute separation for real-time analytics (March 2023)
  • This benchmark on top real-time analytics databases in 2023 (Feb 2023)
  • This talk on emerging architectures for real-time CDC (Dec 2022)

About the showDatacast features long-form, in-depth conversations with practitioners and researchers in the data community to walk through their professional journeys and unpack the lessons learned along the way. I invite guests coming from a wide range of career paths — from scientists and analysts to founders and investors — to analyze the case for using data in the real world and extract their mental models (“the WHY and the HOW”) behind their pursuits. Hopefully, these conversations can serve as valuable tools for early-stage data professionals as they navigate their own careers in the exciting data universe.

Datacast is produced and edited by James Le. For inquiries about sponsoring the podcast, email khanhle.1013@gmail.com.

Subscribe by searching for Datacast wherever you get podcasts, or click one of the links below:

  • Listen on Spotify
  • Listen on Apple Podcasts
  • Listen on Google Podcasts

If you’re new, see the podcast homepage for the most recent episodes to listen to, or browse the full guest list.

View Details

Show Notes* (01:18) Arjun shared formative experiences of his upbringing - growing up in Bangalore, India; going to UWC Mahindra College for high school; and pursuing a liberal arts education in the US. * (04:45) Arjun described his overall academic experience at Willams College - where he studied Computer Science and Economics and did a one-year stint at the Computer Lab at the University of Cambridge. * (11:19) Arjun talked about his specialization within academic computer science: distributed systems. * (14:17) Arjun unpacked the arc of his Ph.D. experience at the University of Pennsylvania, advised by Professor Andreas Haeberlen. * (19:25) Arjun dissected the technical challenges and novelty of his Ph.D. dissertation on distributed systems that computed differentially private things. * (23:20) Arjun shared his love for teaching which benefits his industry career. * (25:55) Arjun walked through his decision to join Cockroach Labs as a software engineer. * (32:25) Arjun unpacked the CockroachDB Performance Guide and a RocksDB deep-dive on the Cockroach Labs blog. * (37:24) Arjun shared valuable lessons learned from his scaling journey with Cockroach. * (41:36) Arjun mentioned how his writing practice benefited his day-to-day work designing database systems in a production setting (Check out his posts on database transaction isolation semantics and the history of log-structured merge trees). * (45:46) Arjun unpacked his 2019 blog post titled "The Philosophy of Computational Complexity." * (52:52) Arjun emphasized the importance of writing evergreen and authoritative long-form content that attracts a small amount of audience. * (55:54) Arjun shared the story behind the founding of Materialize, which builds a SQL streaming database on top of Timely Dataflow and Differential Dataflow, two research projects created by his co-founder Frank McSherry. * (01:00:04) Arjun unpacked the architecture design of Materialize at a high level. * (01:04:36) Arjun explained a core capability of Materialize called Streaming SQL. * (01:07:37) Arjun discussed successful tactics to raise the adoption and contribution to Materialize's open-source project. * (01:11:23) Arjun walked through the major enterprise-grade features baked into Materialize Cloud. * (01:15:54) Arjun dissected a blog post about Materialize’s unbundled cloud architecture detailing the shift from the Materialize single binary to Materialize Cloud. * (01:21:13) Arjun envisioned how Materialize fits into the quickly evolving modern data stack. * (01:25:07) Arjun shared valuable hiring lessons to attract the right people who are excited about Materialize's mission. * (01:27:59) Arjun shared his brief take on building a high-performance company culture. * (01:29:19) Arjun discussed the challenges for his team to find the early design partners. * (01:31:17) Arjun walked through notable use cases of Materialize. * (01:34:24) Arjun shared fundraising advice with founders who are seeking the right investors for their startups. * (01:41:21) Arjun highlighted the similarities and differences between being a researcher and a founder. * (01:42:46) Closing segment.

Arjun's Contact Info* LinkedIn * Twitter * GitHub * Google Scholar

Materialize's Resources* Website | Twitter | LinkedIn | Slack * Docs | GitHub * Blog | Events | Guides * Careers

Mentioned ContentResearch + Articles* Distributed Differential Privacy and Applications (2015) * Performance Report: Benchmarking CockroachDB's TPC-C Performance * Why We Built CockroachDB on top of RocksDB (2019) * A History of Transaction Histories (2018) * A Brief History of Log Structured Merge Trees (2018)

People* Kyle Kingsbury * Bob Muglia * Frank McSherry

Book* Zero To One (by Peter Thiel)

NotesMy conversation with Arjun was recorded back in May 2022. Since then, a lot has happened. I recommend looking at the resources below:

  • About Materialize webpage (which shows the team building Materialize as well as the pedigree)
  • Guide: What is a Streaming Database (which walks through why Materialize is important and different from a normal database)
  • Case Study: Real-time Delivery Tracking UI in a Single Sprint at Onward
  • Tech Demo: CI/CD Workflows for dbt+Materialize (March 2023)
  • Announcing The Next Generation of Materialize (Oct 2022)

About the showDatacast features long-form, in-depth conversations with practitioners and researchers in the data community to walk through their professional journeys and unpack the lessons learned along the way. I invite guests coming from a wide range of career paths — from scientists and analysts to founders and investors — to analyze the case for using data in the real world and extract their mental models (“the WHY and the HOW”) behind their pursuits. Hopefully, these conversations can serve as valuable tools for early-stage data professionals as they navigate their own careers in the exciting data universe.

Datacast is produced and edited by James Le. For inquiries about sponsoring the podcast, email khanhle.1013@gmail.com.

Subscribe by searching for Datacast wherever you get podcasts, or click one of the links below:

  • Listen on Spotify
  • Listen on Apple Podcasts
  • Listen on Google Podcasts

If you’re new, see the podcast homepage for the most recent episodes to listen to, or browse the full guest list.

View Details

Show Notes* (01:30) Doris walked through her time doing research in physics and astrophysics at UC Berkeley and getting involved with data science. * (04:11) Doris reflected on her decision to pursue the Ph.D. program in computer science at the University of Illinois, Urbana-Champaign. * (05:53) Doris discussed her development of no-code, interactive visualization interfaces accelerating users toward data insight discovery. * (10:37) Doris explained how the RISE Lab and I School at UC Berkeley helped shape her thinking around working with end-users and building something to serve the data science community. * (16:05) Doris unpacked the focus of her Ph.D. dissertation - which is to make data exploration and visualization easier and more accessible through automation. * (17:27) Doris shared the motivation and high-level design behind the development of Lux, a general-purpose visual exploration assistant situated within a computational notebook. * (21:25) Doris revealed the recipe for open-source community engagement and roadmap prioritization with Lux. * (26:17) Doris shared the founding story of Ponder, whose mission is to improve data science productivity by empowering users to do data science at all scales. * (31:02) Doris explained how Ponder helps solve the fragmentation challenges across the data stack. * (34:27) Doris provided a brief overview of Modin, which improves the scalability of data frames. * (38:41) Doris discussed Ponder's go-to-market strategy to drive more enterprise interest toward the product. * (41:23) Doris discussed her team's challenges in finding early design partners across various industries. * (44:16) Doris shared valuable hiring lessons to attract the right people who are excited about Ponder's mission. * (47:42) Doris shared fundraising advice to founders who are seeking the right investors for their startups. * (49:33) Doris highlighted the difference between being a researcher and a founder. * (51:06) Closing segment.

Doris' Contact Info* Website * Twitter * LinkedIn * GitHub

Ponder's Resources* Website | Twitter | LinkedIn | Slack * Modin | Lux * Events

Mentioned ContentPublications* The Case for a Visual Discovery Assistant:A Holistic Solution for Accelerating Visual Data Exploration (IEEE Data Bulletin 2018) * Understanding Sense-making in Visual Query Systems (IEEE Visual Analytics Science and Tech 2019) * Deconstructing Categorization in Visualization Recommendation: A Taxonomy and Comparative Study (IEEE Transactions on Visualization and Computer Graphics 2021) * Lux: Always-On Visualization Recommendation for Exploratory Data Science (Dec 2021)

Blog Posts* Insight Machines: The Past, Present, and Future of Visualization Recommendation (Multiple Views, Feb 2020) * Announcing Ponder (March 2022) * How we parallelized 600+ pandas functions with Modin (March 2022) * Using Lux to visualize your pandas dataframes with zero effort (March 2022) * Ph.D. Alum Doris Lee Wants to Democratize Data Science Tools (March 2022)

People* Chip Huyen * Shreyar Shankar * Parul Pandey

NotesMy conversation with Doris was recorded back in May 2022. Earlier this year, Ponder developed the first-of-its-kind technology that allows anyone to run their pandas code directly in your data warehouse, be it Snowflake, BigQuery, or Redshift. With Ponder, you get the same pandas-native experience that you love, but with the power and scalability of cloud-native data warehouses. More details are in this blog post.

Additionally, you can run NumPy commands on your data warehouse as well. This means you can work with the NumPy API to build data and ML pipelines, and let Snowflake / BigQuery / Redshift take care of scaling, security, and compliance. More details are in this blog post.

If you are interested in trying these new capabilities out, sign up here!

About the showDatacast features long-form, in-depth conversations with practitioners and researchers in the data community to walk through their professional journeys and unpack the lessons learned along the way. I invite guests coming from a wide range of career paths — from scientists and analysts to founders and investors — to analyze the case for using data in the real world and extract their mental models (“the WHY and the HOW”) behind their pursuits. Hopefully, these conversations can serve as valuable tools for early-stage data professionals as they navigate their own careers in the exciting data universe.

Datacast is produced and edited by James Le. For inquiries about sponsoring the podcast, email khanhle.1013@gmail.com.

Subscribe by searching for Datacast wherever you get podcasts, or click one of the links below:

  • Listen on Spotify
  • Listen on Apple Podcasts
  • Listen on Google Podcasts

If you’re new, see the podcast homepage for the most recent episodes to listen to, or browse the full guest list.

View Details

Show Notes* (01:47) Chris reflected on his educational experience at Santa Clara University in the mid-2000s, where he also interned at NeoMagic and Intacct Corporation. * (07:31) Chris recalled valuable lessons from his first job as a software engineer at PayPal, researching new fraud prevention techniques. * (11:28) Chris shared the technical and operational challenges associated with his work at LinkedIn as a data scientist - scaling LinkedIn's Hadoop cluster, improving LinkedIn's "People You May Know" algorithm, and delivering the next generation of LinkedIn's "Who's Viewed My Profile" product. * (22:00) Chris provided criteria that his team relied on when choosing their big data solutions (which include Aster Data, Greenplum, and Hadoop). * (25:22) Chris gave advice to early-stage startups that want to start adopting best practices in observability and deployment. * (28:02) Chris expanded on his concept that models and microservices should be running on the same continuous delivery stack. * (30:52) Chris discussed his strategy to become a better interviewer - as he performed ~1,500 interviews at LinkedIn and WePay. * (37:39) Chris explained the motivation behind the creation of Apache Samza (LinkedIn's streaming system infrastructure built on top of Apache Kafka) and discussed its high-level design philosophy. * (46:19) Chris shared lessons learned from evangelizing Samza to the broader open-source community outside of LinkedIn. * (52:44) Chris talked about his decision to join the Data Infrastructure team at WePay as a principal software engineer after 7 years at LinkedIn. * (01:00:53) Chris shared the technical details behind the evolution of WePay's data infrastructure throughout his time there. * (01:12:40) Chris shared an insider perspective on the adoption of Apache Airflow from his experience as a Project Committee Member. * (01:20:15) Chris discussed the fundamental design principles that make Apache Kafka such a powerful technology. * (01:25:40) Chris reflected on his experience building out WePay's engineering team. * (01:27:14) Chris shared the story behind the writing journey of the "Missing README" - which he co-authored with Dmitriy Ryaboy. * (01:38:16) Chris revisited his predictions in a 2019 post called "The Future of Data Engineering" and discussed key trends such as real-time data warehouses, data mesh, and headless BI. * (01:44:27) Chris gave advice to a smart, driven engineer who wants to explore angel investing - given his experience as a strategic investor and advisor for startups in the data space since 2015. * (01:48:17) Chris shared advice on hiring engineers and navigating open-source product strategies for companies he invested in. * (01:53:57) Chris reflected on his consistency in adding value to the relationships he has formed over the years. * (01:58:00) Closing segment.

Chris's Contact Info* Website * Twitter * LinkedIn * Github * AngelList

Mentioned ContentBlog Posts* Joel Spolsky's Blog * Models and microservices should be running on the same continuous delivery stack (Oct 2018) * Using checksums to verify syncing 100M database records (Napkin Math, Jan 2021) * Datacast episode with Jeremiah Lowin, CEO of Prefect (March 2022) * Kafka CDC breaks database encapsulation (Nov 2018) * Kafka provides data portability and infrastructure agility (Jan 2019) * The Future of Data Engineering (July 2019) * Work For Two Companies (Nov 2021)

People* Will Larson * Maxime Beauchemin * Julia Evans * Gunnar Morling * Coda Hale

Books* Google's Site Reliability Engineering Books * "On Writing Well" * "The Missing README" * "Empire of Light: Tesla, Edison, Westinghouse, and the Race to Electrify the World"

NotesMy conversation with Chris was recorded back in May 2022. Earlier this year, Chris released Recap, a dead simple data catalog for engineers, written in Python. Recap makes it easy for engineers to build infrastructure and tools that need metadata. Check out his blog post and get started with Recap's documentation!

About the showDatacast features long-form, in-depth conversations with practitioners and researchers in the data community to walk through their professional journeys and unpack the lessons learned along the way. I invite guests coming from a wide range of career paths — from scientists and analysts to founders and investors — to analyze the case for using data in the real world and extract their mental models (“the WHY and the HOW”) behind their pursuits. Hopefully, these conversations can serve as valuable tools for early-stage data professionals as they navigate their own careers in the exciting data universe.

Datacast is produced and edited by James Le. For inquiries about sponsoring the podcast, email khanhle.1013@gmail.com.

Subscribe by searching for Datacast wherever you get podcasts, or click one of the links below:

  • Listen on Spotify
  • Listen on Apple Podcasts
  • Listen on Google Podcasts

If you’re new, see the podcast homepage for the most recent episodes to listen to, or browse the full guest list.

View Details

Show Notes* (01:32) Nnamdi shared formative experiences of his upbringing, where he spent countless hours building computers, coding up websites, and finding ways to game Google search. * (04:54) Nnamdi described his undergraduate experience studying Economics at Yale University and interning at McKinsey and J.P. Morgan. * (08:10) Nnamdi reflected on the decline of the investment banking industry - given his one year working for the technology, media, and telecommunications group at J.P. Morgan in New York. * (12:52) Nnamdi discussed his career transition into venture investing at ICONIQ Capital, where he deployed over $500 million into high-growth technology companies. * (15:00) Nnamdi reflected on his proudest accomplishments during his four formative years at ICONIQ. * (17:35) Nnamdi talked about his excitement for GitLab, one of his investments. * (21:27) Nnamdi touched on his time getting an MBA from the Stanford Graduate School of Business. * (24:21) Nnamdi also completed coursework in Stanford's Computer Science department (such as CS 231N and CS 224N) during his MBA. * (26:37) Nnamdi explained the venture ecosystem at Stanford, given his experience serving as the Co-President and Vice President of the Venture Capital and Tech Clubs, respectively. * (28:57) Nnamdi unpacked his experience working at Confluent as a product manager and conducting independent research on trends in developer productivity. * (32:23) Nnamdi reflected on his decision to join Lightspeed Venture in mid-2020, investing in early-stage software startups to enhance the productivity of technical knowledge workers. * (34:17) Nnamdi shared how he proved his value upfront in potential deals and started forming his investment theses as a new investor at Lightspeed. * (36:24) Nnamdi dissected the key factors that triggered him to make investments in the seed rounds of Ponder and Voltron Data (in the domain of developer tools). * (40:36) Nnamdi explained his Series A investment in Redpanda and Materialize (in the domain of real-time data infrastructure). * (45:45) Nnamdi shared advice he had been giving his portfolio companies in hiring decisions and navigating growth strategy. * (49:07) Nnamdi walked through his 3-part series on major industry trends, top strategic priorities, and biggest challenges for software and infrastructure startups pushing the developer productivity frontier. * (52:37) Nnamdi shared advice to startups thinking about scaling their developer relations, given the challenge of hiring developer advocates for dev-focused startups. * (56:27) Nnamdi unpacked his 3-part series on the developer productivity manifesto that introduces the developer productivity flywheel, explains how more developers lead to lower productivity, and argues that we are leaving on the table $670B of software by not maximizing developer employment and developer productivity. * (01:01:26) Nnamdi examined his obsession with the fat-tailed nature of high-growth startups, such as why VCs don't index-invest, why Saas monetization is concentrated on the tails, and why product-market fit gets harder to achieve the longer you search for it. * (01:04:26) Nnamdi explained his new and improved SaaS metric called Weighted ACV, which is the weight of the revenue that a customer represents and tells founders where to look if they want to best understand the revenue of their businesses. * (01:07:53) Nnamdi thought about his recognition as equal to his credibility as an investor on a mission to increase total software output by investing in technical tools for technical people. * (01:11:03) Closing segment.

Nnamdi's Contact Info* Website * Lightspeed Profile * LinkedIn * Twitter * GitHub * Medium

Lightspeed's Resources* Website | Twitter | LinkedIn * Global Presence * Medium Blog

Mentioned ContentArticles* Six Trends Shaping Developer Productivity * Top Three Strategic Priorities of Developer Productivity Startups * Four Challenges Facing Developer Productivity Startups * Awesome Developer Advocates Are Hiding in Plain Sight * The Developer Productivity Manifesto Part 1 — The Flywheel * The Developer Productivity Manifesto Part 2 — More (Developers) Isn’t Always More * The Developer Productivity Manifesto Part 3 — Leaving Software on the Table * You Don't Understand Compound Growth * Funding Simply Shifts the Bottleneck * Why Don't VCs Index Invest? * Enterprise Software Monetization is Fat-Tailed * Product-Market Fit is Lindy * Introducing a New and Improved SaaS Metric: Weighted ACV

People1. Mike Volpi (Index Ventures) 2. Keith Rabois (Founders Fund)

BooksNassim Taleb's Incerto Series:

  • Fooled By Randomness
  • The Black Swan
  • The Bed of Procrustes
  • Antifragile
  • Skin In The Game

NotesMy conversation with Nnamdi was recorded in May 2022. Since then, many things have happened. I'd recommend checking out:

  1. Lightspeed's announcement of the new three funds last year
  2. Nnamdi’s new series on software valuations (1, 2, 3)
  3. Nnamdi’s recent posts on the reality of tech layoffs and the need for more startups
  4. Nnamdi’s recent investment in Select Star

About the showDatacast features long-form, in-depth conversations with practitioners and researchers in the data community to walk through their professional journeys and unpack the lessons learned along the way. I invite guests coming from a wide range of career paths — from scientists and analysts to founders and investors — to analyze the case for using data in the real world and extract their mental models (“the WHY and the HOW”) behind their pursuits. Hopefully, these conversations can serve as valuable tools for early-stage data professionals as they navigate their own careers in the exciting data universe.

Datacast is produced and edited by James Le. For inquiries about sponsoring the podcast, email khanhle.1013@gmail.com.

Subscribe by searching for Datacast wherever you get podcasts, or click one of the links below:

  • Listen on Spotify
  • Listen on Apple Podcasts
  • Listen on Google Podcasts

If you’re new, see the podcast homepage for the most recent episodes to listen to, or browse the full guest list.

View Details

Show Notes* (02:00) Tom shared formative experiences of his upbringing. * (04:17) Tom described his educational experience at MIT and his research thesis in Computer Vision. * (07:08) Tom talked about his interest in computer vision and computational neuroscience. * (08:38) Tom recalled lessons from his first job out of school as a software engineer at Silicon Graphics, building high-performance visualization systems. * (11:32) Tom reflected on his time as a product lead at Autodesk, launching a location services platform for global wireless carriers and a developer ecosystem for GIS applications. * (13:48) Tom reflected on his MBA experience at Harvard Business School. * (16:54) Tom reflected on his first stint at Google - leading a sales operations team in AdWords and building YouTube's monetization systems. * (20:06) Tom recalled lessons learned as a first-time founder of a social commerce startup called Renown Labs. * (23:11) Tom walked over his time as the Director of Product at the SaaS social media marketing startup Wildfire (which was acquired by Google in August 2012). * (27:00) Tom explained his career transition from tech operating into venture investing - after joining the enterprise investment team at a16z as a partner in 2012. * (29:55) Tom revisited his thesis, discussing the rise of Enterprise Hackers back in 2013. * (32:25) Tom talked about his decision to join NextWorld Capital as a partner in 2014, leading investments across enterprise applications, the Internet of Things, and AI. * (34:50) Tom unpacked his investment thesis on enterprise technology that helps the blue-collar working class. * (36:50) Tom shared his mental checklist used to evaluate investment opportunities in enterprise AI at NextWorld. * (40:39) Tom shared the founding story of Masterful AI - where he has been a co-founder and CEO since 2019. * (43:08) Tom expanded upon the 2-year incubation period from the inception to the announcement of the Masterful AI platform. * (45:17) Tom unpacked major inefficiencies of ML development and explained how the Masterful platform works at a high level. * (48:18) Tom shared exciting initiatives in Masterful's product roadmap. * (50:09) Tom highlighted the principles that stood the test of time in computer vision over the past two decades. * (51:41) Tom shared valuable hiring lessons to attract the right people who are excited about Masterful AI's mission. * (55:07) Tom discussed his team's challenges in finding early design partners across various industries. * (58:24) Tom shared fundraising advice to founders who are seeking the right investors for their startups. * (01:01:01) Tom reflected on his career traversing across product management, venture capital, and startup founder. * (01:05:20) Closing segment.

Tom's Contact Info* LinkedIn * Twitter * Medium * Website

Masterful AI's Resources* Website | Twitter | LinkedIn * Docs | Slack Community * "Building Things with Machine Learning" Podcast

Mentioned ContentArticles* "The Enterprise Hacker Rises" (a16z Blog, Dec 2013) * "Joining NextWorld Capital" (Personal Blog, Nov 2014) * "The Next Big Opportunity In Enterprise Starts In The Field" (TechCrunch, July 2015) * "My visit to the Obama White House: AI, the future of jobs, and a VC’s Letter to the next administration" (NextWorld Insights, Jan 2017) * "AI hype has peaked so what’s next?" (TechCrunch, Sep 2017) * "AI is bringing superpowers to the specialist" (LinkedIn, Oct 2018) * "Introducing Masterful AI" (Masterful Blog, Nov 2021)

People* Andrew Ng (Founder of DeepLearning.AI, Founder and CEO of Landing AI, Co-Founder of Coursera) * Chris Dixon (General Partner at a16z)

Book* AI Superpowers (by Kai-Fu Lee)

NoteMy conversation with Tom was recorded back in May 2022. Here is the note from Tom regarding updates with Masterful:

The latest at Masterful AI is that we’re launching a new generative AI product. We saw a need to make generative models more customizable and more reliable, so companies can trust them for real business applications. We’re starting by enabling companies to tell a more vivid and personalized story about their products at scale.

About the showDatacast features long-form, in-depth conversations with practitioners and researchers in the data community to walk through their professional journeys and unpack the lessons learned along the way. I invite guests coming from a wide range of career paths — from scientists and analysts to founders and investors — to analyze the case for using data in the real world and extract their mental models (“the WHY and the HOW”) behind their pursuits. Hopefully, these conversations can serve as valuable tools for early-stage data professionals as they navigate their own careers in the exciting data universe.

Datacast is produced and edited by James Le. For inquiries about sponsoring the podcast, email khanhle.1013@gmail.com.

Subscribe by searching for Datacast wherever you get podcasts, or click one of the links below:

  • Listen on Spotify
  • Listen on Apple Podcasts
  • Listen on Google Podcasts

If you’re new, see the podcast homepage for the most recent episodes to listen to, or browse the full guest list.

View Details

Show Notes* (02:02) Grace shared formative experiences of her upbringing - being heavily influenced by the financial sector from growing up near New York and getting an appreciation for diverse perspectives from studying abroad in Tokyo. * (03:47) Grace described her college experience at Stanford studying Management Science and Engineering. * (08:04) Grace talked about her participation in the Mayfield Fellowship and her service as a Co-President of Stanford Women in Business. * (12:34) Grace walked through her internship experiences as an investor at the Stanford Management Company, in product at ed-tech startup Handshake, and in growth equity at Stripes Group. * (16:20) Grace reflected on her time at Canvas Ventures - where she joined as a campus scout while still a student. * (19:11) Grace shared three approaches to prove her value upfront in potential deals and to form her investment thesis as a new VC associate. * (23:07) Grace dissected her Series A investment in Vendia, a blockchain-powered, real-time data-sharing platform that solves the growing inter-organization data collaboration problem. * (25:17) Grace examined her Series A investment in Robocorp, which offers the first cloud-native, open-source automation stack and orchestration platform to power any automation process. * (26:43) Grace shared three pieces of advice in hiring decisions, navigating go-to-market strategy, and growing product offerings that she had given her portfolio companies. * (29:49) Grace shared trends in the API-first economy that she is most excited about in the upcoming years. * (31:56) Grace unpacked key takeaways from her article "The Mindset of a Data Leader." * (34:02) Grace discussed under-hyped and over-hyped trends in Web3 - taken from her incredibly detailed deck on the Web3 World. * (36:31) Grace dissected the major categories of the Web3 infrastructure, including Decentralized Finance, Decentralized Apps, DAOs, NFTs, and Guild Education/Reskilling. * (41:43) Grace walked through her decision to join Lux Capital, a firm that invests in emerging science and technology ventures at the outermost edges of what is possible, as a principal investor in early 2022. * (45:12) Grace shared her mental checklist to evaluate entrepreneurs and make investment decisions at the nexus of web3, data infrastructure, and applications of AI/ML. * (47:10) Grace talked about her community-building work to promote women's voices in tech. * (48:59) Grace reflected on her consistency in adding value to every conversation with people in her community. * (51:26) Closing segment.

Grace's Contact Info* Website * Lux Profile * LinkedIn * Twitter

Lux Capital* Website | Twitter | LinkedIn * Securities (Podcast & Newsletter)

Mentioned ResourcesArticles* "The Third-Party API Economy: Part I" (Sep 2020) * "The Third-Party API Economy: Part II" (Feb 2021) * "The Mindset of a Data Leader" (Nov 2020) * "The Web3 World" (Jan 2022) * "Welcoming our newest investor Grace Isford to Lux Capital" (Feb 2022)

People* Fred Wilson (Union Square Ventures) * Matt Huang and Fred Ehrsam (Paradigm Ventures) * Katie Haun (Haun Ventures)

Book* "Wanting" (by Luke Burgis)

NotesMy conversation with Grace was recorded back in April 2022. Since then, many things have happened. I'd recommend:

  • Listening to her chats with Tina Seelig and Christian Catalini on the Securities podcast
  • Reading her thoughts on building the next AI/ML infrastructure stack
  • Check out her reflections on the intersection of AI and creativity

Additionally, Grace invested in RunwayML's Series C, a pioneer in the Generative AI space. If you are in NYC, be sure to stop by the upcoming first annual AI film festival powered by Runway!

About the showDatacast features long-form, in-depth conversations with practitioners and researchers in the data community to walk through their professional journeys and unpack the lessons learned along the way. I invite guests coming from a wide range of career paths — from scientists and analysts to founders and investors — to analyze the case for using data in the real world and extract their mental models (“the WHY and the HOW”) behind their pursuits. Hopefully, these conversations can serve as valuable tools for early-stage data professionals as they navigate their own careers in the exciting data universe.

Datacast is produced and edited by James Le. For inquiries about sponsoring the podcast, email khanhle.1013@gmail.com.

Subscribe by searching for Datacast wherever you get podcasts, or click one of the links below:

  • Listen on Spotify
  • Listen on Apple Podcasts
  • Listen on Google Podcasts

If you’re new, see the podcast homepage for the most recent episodes to listen to, or browse the full guest list.

View Details

Show Notes* (01:32) Dânia shared her upbringing in Brazil and her college experience studying Applied Mathematics at the University of Campinas. * (05:58) Dânia touched on her early career working in marketing intelligence in Brazil. * (10:38) Dânia described her thesis on scalable implementations of the Alternating Least Squares algorithm for Collaborative Filtering recommendation, conducted during her Master's degree in Computer Science from the University of Fluminense. * (16:10) Dânia recalled her hustling phase working and getting a Master's degree simultaneously. * (24:19) Dânia reflected on her move to Berlin to work as a data scientist in several startups. * (31:00) Dânia looked back at her time working at MYTOYS GROUP's Analytics team, responsible for Predictive Analytics and Machine Learning Modeling. * (34:12) Dânia compared doing data science to practicing mixed martial arts. * (38:35) Dânia reflected on her involvement with Data Science for Social Good Berlin as a data ambassador and Data Science Retreat as a SQL Masterclass Teacher. * (43:14) Dânia shared the founding story of AI Guild - the go-to community for data and business professionals advancing AI adoption - where she is a founding member. * (47:36) Dânia gave her thoughts on barriers preventing more women from entering the data field. * (51:21) Dânia discussed the #datalift initiative, which pushes to productionize more data analytics and machine learning solutions. * (58:27) Dânia explained her work supporting the advancement of #datacareer talents and experts. * (01:01:22) Dânia gave her take on the evolution of the data field over the past decade. * (01:03:16) Closing segment.

Dânia's Contact Info* LinkedIn * Twitter * Website * GitHub * Medium

AI Guild's Resources* Website | LinkedIn | YouTube * Join As A Member * #datalift * #datacareer

Mentioned ContentPeople1. Andrew Ng: Founder of deeplearning.ai, co-founder of Coursera 2. Alessandra Sala: President of Women in AI, Sr. Director of Artificial Intelligence and Data Science at Shutterstock 3. Joy Buolamwini: Founder and Executive director of The Algorithmic Justice League and maker of the "Coded Bias" documentary, available on Netflix

Book* Weapons of Math Destruction: How Big Data Increases Inequality and Threatens Democracy by Cathy O'Neil

About the showDatacast features long-form, in-depth conversations with practitioners and researchers in the data community to walk through their professional journeys and unpack the lessons learned along the way. I invite guests coming from a wide range of career paths — from scientists and analysts to founders and investors — to analyze the case for using data in the real world and extract their mental models (“the WHY and the HOW”) behind their pursuits. Hopefully, these conversations can serve as valuable tools for early-stage data professionals as they navigate their own careers in the exciting data universe.

Datacast is produced and edited by James Le. For inquiries about sponsoring the podcast, email khanhle.1013@gmail.com.

Subscribe by searching for Datacast wherever you get podcasts, or click one of the links below:

  • Listen on Spotify
  • Listen on Apple Podcasts
  • Listen on Google Podcasts

If you’re new, see the podcast homepage for the most recent episodes to listen to, or browse the full guest list.

View Details

Show Notes* (01:33) Bobby shared his upbringing in DC and high-school experience at St. Albans School. * (04:10) Bobby described his academic experience at Stanford studying Management Science and Engineering. * (07:39) Bobby recalled valuable career lessons learned working as a Finance Analyst at IBM and Inflection. * (09:56) Bobby reflected on his rationale for joining Intercom as one of the company's early employees right after its Series A financing in 2013. * (14:16) Bobby unpacked his 2016 talk "Scaling Analytics at Intercom," which explained the analytics journey at Intercom. * (18:46) Bobby shared a few metrics that are fundamental to the health of a startup across its growth stages (read his Intercom blog about the data points that startups should measure). * (22:50) Bobby shared the founding story of Equals. * (27:33) Bobby explained his decision to choose Ben McRedmond as his co-founder. * (29:35) Bobby expanded on the appealing traits of using spreadsheets. * (31:54) Bobby described the evolution of spreadsheet-like products and how the Equals product works at a high level. * (34:35) Bobby gave his take on how the concept of a next-generation spreadsheet fits into the quickly evolving modern data stack. * (38:31) Bobby shared valuable hiring lessons to attract the right people who are excited about Equals' mission. * (44:34) Bobby shared the challenges of finding Equals' early design partners and lighthouse customers. * (47:17) Bobby recapped key lessons about hiring financial analysts at Intercom. * (51:45) Bobby shared advice to a smart, driven finance operator looking to get more influence within a startup environment. * (56:26) Bobby emphasized the valuable skills acquired from his analyst career for his current founder journey. * (58:45) Closing segment.

Bobby's Contact Info* LinkedIn * Twitter

Equals Resources* Website | Twitter | LinkedIn * Spreadsheet Templates * Insights In Action interview series * Introducing Pivot Tables for Equals (Aug 2022) * Equals raises $16M Series A from a16z to replace Excel (Nov 2022)

Equals is hiring across Engineering, Design, Growth, and an Executive Assistant. Reach out to Bobby if you are interested!

Mentioned ContentArticles + Talk* 23 SaaS Metrics for Fundraising + Optimization (March 2015) * Scaling Analytics at Intercom (Intercom Analytics Meetup, April 2016) * Data Points: What Should Your Startup Measure? (Oct 2017) * Every analyst is a finance analyst (May 2021) * The only question that matters when interviewing analysts (May 2021) * When to make your first finance hire (May 2021) * The hardest leap to make as a scaling finance leader (June 2021) * Finance and describing product-market fit (Sep 2021) * The curious analyst (Sep 2021) * The less lonely finance leader (Sep 2021) * Why every scaling finance team is understaffed (Nov 2021) * Revenue is the best North Star metric (March 2022)

People* Karen Church (VP of Research and Data Science at Intercom, Founder of HER+Data) * Noah Goodman (President at DataCRT) * Peter Fishman (Co-Founder of Mozart Data)

About the showDatacast features long-form, in-depth conversations with practitioners and researchers in the data community to walk through their professional journeys and unpack the lessons learned along the way. I invite guests coming from a wide range of career paths — from scientists and analysts to founders and investors — to analyze the case for using data in the real world and extract their mental models (“the WHY and the HOW”) behind their pursuits. Hopefully, these conversations can serve as valuable tools for early-stage data professionals as they navigate their own careers in the exciting data universe.

Datacast is produced and edited by James Le. For inquiries about sponsoring the podcast, email khanhle.1013@gmail.com.

Subscribe by searching for Datacast wherever you get podcasts, or click one of the links below:

  • Listen on Spotify
  • Listen on Apple Podcasts
  • Listen on Google Podcasts

If you’re new, see the podcast homepage for the most recent episodes to listen to, or browse the full guest list.

View Details

Show Notes* (01:34) Ran reflected on his time working as a Technical Product Manager at the Israeli Intelligence army. * (04:07) Ran recalled his favorite classes on Machine Learning and Computer Graphics during his education in Computer Science at Reichman University. * (05:24) Ran talked about a valuable lesson learned as a Software Engineer at VMware's Cloud Provider Software Business Unit. * (08:07) Ran shared his thoughts on how engineers could be more impactful in startup organizations. * (09:52) Ran talked about his decision to join Wix.com to work as a software engineer focusing on data infrastructure. * (12:48) Ran explained the motivation for building Wix's internal ML platform, designed to address the end-to-end ML workflow. * (16:48) Ran discussed the main components of Wix's ML platform: feature store, CI/CD mechanism, UI management console, and API prediction service. * (18:51) Ran unpacked the virtual feature store and the CI/CD components of Wix's ML platform. * (24:41) Ran expanded on the distinction between virtual and materialized feature stores. * (27:01) Ran provided three key lessons for organizations looking to build an internal ML platform (as brought upon his 2020 talk discussing Wix's ML Platform). * (31:43) Ran shared the essential attributes of exceptional data and ML engineering talent. * (33:54) Ran shared the founding story of Qwak, which aims to build an end-to-end ML engineering platform to automate the MLOps processes. * (37:07) Ran talked about his responsibilities as the VP of Engineering at Qwak. * (38:45) Ran dissected the key capabilities that are baked into the Qwak platform - a Build System, a Serving layer, a Data Lake, a Feature Store, and Automations capabilities. * (44:05) Ran explained the big engineering challenges for teams to build an in-house feature store and envisioned the future of the feature store ecosystem in the upcoming years. * (47:45) Ran shared valuable hiring lessons to attract the right people who are excited about Qwak's mission. * (50:22) Ran reflected on the challenges for Qwak to find the early design partners. * (52:43) Ran described the state of the ML Engineering community in Israel. * (54:53) Closing segment.

Ran's Contact Info* LinkedIn

Qwak's Resources* Website | Twitter | LinkedIn * Why Qwak * Blog

Mentioned ContentTalks* "Overview of Wix's Machine Learning Platform" (2020) * "Feature Stores - Unified Data Pipelines for ML" (2022)

People* Andrew Ng * Matei Zaharia * Barr Moses

Book* "Principles" (by Ray Dalio)

About the showDatacast features long-form, in-depth conversations with practitioners and researchers in the data community to walk through their professional journeys and unpack the lessons learned along the way. I invite guests coming from a wide range of career paths — from scientists and analysts to founders and investors — to analyze the case for using data in the real world and extract their mental models (“the WHY and the HOW”) behind their pursuits. Hopefully, these conversations can serve as valuable tools for early-stage data professionals as they navigate their own careers in the exciting data universe.

Datacast is produced and edited by James Le. For inquiries about sponsoring the podcast, email khanhle.1013@gmail.com.

Subscribe by searching for Datacast wherever you get podcasts, or click one of the links below:

  • Listen on Spotify
  • Listen on Apple Podcasts
  • Listen on Google Podcasts

If you’re new, see the podcast homepage for the most recent episodes to listen to, or browse the full guest list.

View Details

Show Notes* (02:15) Eric reflected on his early interest in computer science and his decision to study at Carnegie Mellon University in the early 90s. * (05:40) Eric recalled his academic and overall college experience, emphasizing the importance of the people he was surrounded with. * (08:22) Eric talked about his time working as a quant analyst early in his career, the moment he encountered the birth of the Mosaic browser, and his decision to join the tech industry. * (13:01) Eric imparted wisdom learned from venture investing during the dot-com boom. * (18:02) Eric talked about the next phase of his academic career - earning a Ph.D. in Computer Science from Carnegie Mellon and dropping out of a Ph.D. program at Stanford. * (21:06) Eric discussed his academic research on Computational Economics for corporate malfeasance during his time as a Ph.D. student. * (27:39) Eric shared different initiatives he worked on with Carnegie Mellon University - serving as the Assistant Dean and Assistant Professor of Software Engineering, launching CMU's Silicon Valley Campus, and founding CMU's Entrepreneurial Management program. * (31:54) Eric described his journey in founding Hg Analytics, a hedge fund focused on statistical arbitrage, alongside other CMU's Computer Science PhDs. * (37:36) Eric revisited his passion for AI and robotics, which eventually led to serving as a Presidential Innovation Fellow during the Obama Administration with the White House Office of Science and Technology Policy. * (42:54) Eric shared his perspective on the role of AI in geopolitics and highlighted the challenges with data integration. * (47:29) Eric explained his company Conexus, which develops a technology spin-off from MIT's Mathematics department using a branch of math called Category Theory. * (50:55) Eric went over a customer case study that uses Conexus's solution to guarantee the semantics of data integrity during data transformation. * (54:20) Eric showed his enthusiasm for the concept of data relationships. * (56:59) Eric provided a sneak peek of his forthcoming book, "The Coming Composability: The roadmap for using technology to solve society's biggest problems." * (58:38) Closing segment.

Eric's Contact Info* Twitter * LinkedIn

Conexus' Resources* Website | Resources

Mentioned ContentPeople* Kai-Fu Lee * Andrew Ng * Eric Xing

Book* "ReCulturing: Design Your Company Culture to Connect with Strategy and Purpose for Lasting Success" (by Melissa Daimler)

About the showDatacast features long-form, in-depth conversations with practitioners and researchers in the data community to walk through their professional journeys and unpack the lessons learned along the way. I invite guests coming from a wide range of career paths — from scientists and analysts to founders and investors — to analyze the case for using data in the real world and extract their mental models (“the WHY and the HOW”) behind their pursuits. Hopefully, these conversations can serve as valuable tools for early-stage data professionals as they navigate their own careers in the exciting data universe.

Datacast is produced and edited by James Le. Get in touch with feedback or guest suggestions by emailing khanhle.1013@gmail.com.

Subscribe by searching for Datacast wherever you get podcasts, or click one of the links below:

  • Listen on Spotify
  • Listen on Apple Podcasts
  • Listen on Google Podcasts

If you’re new, see the podcast homepage for the most recent episodes to listen to or browse the full guest list.

View Details

Show Notes* (01:56) Astasia shared her childhood growing up in Silicon Valley. * (05:12) Astasia reflected on her undergraduate education at Stanford - studying Political Science and International Relations. * (06:35) Astasia discussed her research at the Graduate Business School with Professor Condoleezza Rice on a case study called "San Leon Energy: Hydraulic Fracturing in Poland" - which explores how to manage the political risks of using a controversial energy extraction technology in the European Union. * (09:26) Astasia talked about her year in the UK getting a Master's in Technology Policy at the University of Cambridge's Judge Business School. * (12:52) Astasia recalled her experience as an Equity Research Analyst at Baird and Co. * (17:49) Astasia mentioned her work at Cisco Investments, driving their cloud-infrastructure M&A and venture investments. * (20:58) Astasia shared her thoughts on different M&A frameworks she learned from Cisco. * (23:27) Astasia reflected on her decision to join Redpoint Ventures in early 2017, leading investments across developer tools, cloud infrastructure, data/ML infrastructure, AI applications, and cybersecurity. * (25:44) Astasia debunked misconceptions about the venture industry. * (29:30) Astasia discussed ways to prove her value upfront in potential deals and start forming her investment theses as a new investor. * (33:01) Astasia dissected the key factors that triggered her to invest in the Series A of Solo.io and the Series B of LaunchDarkly (in the domain of cloud infrastructure). * (38:48) Astasia explained her Series A investment in Hex and Series B investment in Preset (in the domain of data infrastructure). * (44:12) Astasia shared advice she had given her portfolio companies in hiring decisions, pricing products, and navigating go-to-market strategy while at Redpoint. * (47:36) Astasia walked through her process of writing comprehensive research primers in her Medium blog Memory Leak on wide-ranging topics - from data science notebooks and data orchestration to data pipelining and ML data management. * (51:19) Astasia shared the typical challenges she has seen in companies looking to incorporate Product-Led Growth into their go-to-market motion. * (54:10) Astasia discussed building a community as a fuel for product-led growth and shared advice to startups thinking about starting their community initiatives. * (56:40) Astasia shared advice for hiring good DevRel practitioners. * (01:00:15) Astasia shared advice for a smart, driven operator who wants to explore angel investing. * (01:03:26) Astasia talked about her current journey as the Founding Partner at Quiet Capital, sitting on its early-stage enterprise team and leading opportunities across pre-seed, seed, Series A, and Series B. * (01:05:13) Astasia expanded upon her typical mental checklist to evaluate entrepreneurs and make investment decisions. * (01:07:36) Astasia briefly touched on LP fundraising for Quiet Capital to become a "modern venture firm." * (01:09:59) Astasia emphasized her enthusiasm for the Data-Centric ML movement. * (01:13:41) Closing segment.

Astasia's Contact Info* LinkedIn * Medium * Twitter

Quiet Capital* Website * LinkedIn * Twitter

Mentioned ResourcesContent* John Gannon Blog

People* Satish Dharmaraj (Redpoint Ventures) * Scott Raney (Redpoint Ventures) * Amanda Robson (Cowboy Ventures)

NotesMy conversation with Astasia was recorded back in April 2022. Since then, many things have happened. I'd recommend:

  1. Signing up for her Memory Leak newsletter
  2. Browsing through Quiet Capital's new portfolio careers page
  3. Listening to Astasia's appearance on the Data Stack Show
  4. Checking out Quiet Capital's investments in Edge Delta, Diagrid, and Omni
  5. Looking at her real-time infrastructure landscape

View Details

Show Notes* (02:24) Tarush shared his upbringing in India and his decision to study abroad in the US. * (03:51) Tarush walked through his college experience studying Computer Engineering at Carnegie Mellon University. * (06:24) Tarush described the non-existent state of data infrastructure at Salesforce when he joined as the first data engineer in 2012. * (11:21) Tarush went over his contribution to the automation and benchmarking frameworks over his tenure at Salesforce. * (15:50) Tarush recalled lessons learned from building and managing a data team as a Data Manager at Wyng. * (19:54) Tarush explained how a data team can serve other functional units more efficiently. * (22:37) Tarush elaborated on his decision to adopt Looker for Wyng's Business Intelligence needs. * (26:30) Tarush talked about his decision to join WeWork as their Director of Data Engineering in 2016. * (30:39) Tarush went over the origin and evolution of Marquez - WeWork’s first open-source project around data lineage - during his time as the director of WeWork’s Data Platform team. * (33:49) Tarush highlighted the main challenges of building an internal data platform. * (35:43) Tarush recalled his move to China to help establish WeWork’s Asia operations and focus on the hyper-growing Chinese market. * (39:01) Tarush shared the founding story of 5x during his sabbatical in 2020. * (42:39) Tarush explained the industry's need for a managed data stack. * (45:20) Tarush went over 5x’s process of sourcing, interviewing, and onboarding data engineers who are pre-trained on the modern data stack. * (48:37) Tarush talked about finding the right vendors that make up the modern data stack to partner with. * (50:06) Tarush walked through his production process to put together a lot of good videos to explain what 5x does and raise awareness about the company. * (51:52) Closing segment.

Tarush's Contact Info* LinkedIn * Twitter * Medium

5x Resources* Website | LinkedIn | Twitter | YouTube | Instagram * 5x Explained in 2 Minutes * Managed Data Platform * On-Demand Data Engineering Services * Integrations

Mentioned ContentPeople* George Fraser and Taylor Brown (Founders of Fivetran) * Prukalpa Sankar (Co-Founder and CEO of Atlan) * Frank Slootman (CEO and Chairman of Snowflake)

Books* Stealing Fire (by Steven Kotler and Jamie Wheal) * The 5 AM Club (by Robin Sharma)

About the showDatacast features long-form, in-depth conversations with practitioners and researchers in the data community to walk through their professional journeys and unpack the lessons learned along the way. I invite guests coming from a wide range of career paths — from scientists and analysts to founders and investors — to analyze the case for using data in the real world and extract their mental models (“the WHY and the HOW”) behind their pursuits. Hopefully, these conversations can serve as valuable tools for early-stage data professionals as they navigate their own careers in the exciting data universe.

Datacast is produced and edited by James Le. Get in touch with feedback or guest suggestions by emailing khanhle.1013@gmail.com.

Subscribe by searching for Datacast wherever you get podcasts, or click one of the links below:

  • Listen on Spotify
  • Listen on Apple Podcasts
  • Listen on Google Podcasts

If you’re new, see the podcast homepage for the most recent episodes to listen to or browse the full guest list.

View Details

Show Notes* (01:59) Hyun shared his upbringing and experience living in Korea, Singapore, and the US. * (04:18) Hyun described his undergraduate experience at Duke University. * (08:21) Hyun shared how he got a real taste of the game-changing potential of deep learning from the experience of bringing ML to diagnose Parkinson’s disease with brain MRI scans. * (10:54) Hyun talked about his journey of leveling up coding and ML knowledge. * (12:13) Hyun reflected on his motivation to pursue a Ph.D. program in computer science at Duke. * (15:22) Hyun talked about his participation in the 2016 Amazon Robotics Challenge as the “Team Duke” leader and its Motion Planning function. * (17:25) Hyun reflected on his decision to take a leave of absence from his Ph.D. program and return to Korea to work as an ML Research Engineer at the AI Research Lab of SK Telecom, a major Korean conglomerate. * (19:46) Hyun discussed his research on game AI and synthetic image generation during his time with SK Telecom. * (22:57) Hyun shared the founding story of Superb AI. * (27:11) Hyun described going through the Y Combinator Winter 2019 batch. * (32:25) Hyun unpacked the evolution of Superb AI’s Labeling platform since its inception. * (34:47) Hyun walked through the process of prioritizing the product roadmap. * (36:54) Hyun zoomed in on Superb AI’s automated labeling feature, Custom Auto-Label, which automatically detects and labels common or niche objects in images and videos. * (40:21) Hyun touched on challenges with manually reviewing and auditing labels. * (42:25) Hyun dissected the data-centric problems in computer vision that the newly released Superb DataOps platform is built to solve. * (46:46) Hyun hinted at Superb AI’s product roadmap, judging from current industry-wide pain points. * (48:53) Hyun highlighted a customer use case of Superb AI product offerings. * (51:42) Hyun shared his vision of where Superb AI fits into the quickly evolving AI Infrastructure ecosystem. * (54:15) Hyun shared valuable hiring lessons to attract people who are excited about Superb AI’s mission. * (58:01) Hyun expanded his perspectives on defining and scaling a global company culture. * (01:00:06) Hyun reflected on the challenges of running a remote-first company. * (01:01:54) Hyun shared fundraising advice for founders seeking the right investors for their startups. * (01:03:35) Hyun highlighted the difference between being a researcher and a founder. * (01:05:08) Closing segment.

Hyun’s Contact Info* LinkedIn * Twitter

Superb AI Resources* Website | LinkedIn | Twitter | YouTube | GitHub | Docs * Superb AI Suite Labeling Platform * Superb AI DataOps Platform * The Ground Truth Newsletter * Superb AI Academy

Mentioned ContentPeople1. Andrew Ng 2. Andrej Karpathy 3. Ian Goodfellow

Book1. Zero To One (by Peter Thiel)

About the showDatacast features long-form, in-depth conversations with practitioners and researchers in the data community to walk through their professional journeys and unpack the lessons learned along the way. I invite guests coming from a wide range of career paths — from scientists and analysts to founders and investors — to analyze the case for using data in the real world and extract their mental models (“the WHY and the HOW”) behind their pursuits. Hopefully, these conversations can serve as valuable tools for early-stage data professionals as they navigate their own careers in the exciting data universe.

Datacast is produced and edited by James Le. Get in touch with feedback or guest suggestions by emailing khanhle.1013@gmail.com.

Subscribe by searching for Datacast wherever you get podcasts, or click one of the links below:

  • Listen on Spotify
  • Listen on Apple Podcasts
  • Listen on Google Podcasts

If you’re new, see the podcast homepage for the most recent episodes to listen to or browse the full guest list.

View Details

Show Notes* (01:45) Gary walked through his academic experience getting a Bachelor’s degree in Business Administration at Arizona State University and an MBA in Finance at USC — Marshall School of Business. * (04:52) Gary recalled the most valuable lesson from leading a business development team in the enterprise offerings group at Verizon. * (07:45) Gary recalled the challenges of bringing a company public during his time as the Director of Corporate Development at NorthPoint Communications. * (12:18) Gary shared his learnings while holding a COO role at Vinfolio — an innovator in the wine Industry. * (15:19) Gary talked about his responsibilities in the Chief Financial Officer roles at KnowNow and Zuora. * (19:06) Gary gave advice to founders seeking the right investors for their startups. * (23:51) Gary walked through the learning curves while serving as the CFO, CRO, and COO of enterprise AI pioneer Ayasdi. * (31:06) Gary shared his playbook on building a well-oiled sales operations machine. * (33:46) Gary shared his journey as a first-time CEO at CLARA Analytics. * (36:37) Gary talked about his proudest accomplishments while driving significant growth for CLARA. * (37:52) Gary discussed the go-to-market motions implemented at CLARA. * (41:07) Gary walked through his brief stint as an Entrepreneur-In-Residence at Redpoint Ventures, a top-tier VC firm focused on early-stage investing. * (44:14) Gary rationalized his decision to become the CEO of Arcion Labs in December 2021. * (49:39) Gary explained the high-level architectural design of Arcion’s data mobility platform. * (54:19) Gary discussed strategies for finding the right technology partners to collaborate with. * (57:42) Gary highlighted a few customer use cases of Arcion. * (01:01:48) Gary shared valuable hiring lessons to attract the right people who are excited about Arcion’s mission. * (01:04:28) Gary distilled lessons learned while building a high-performance team at Arcion. * (01:09:14) Gary described the benefits of adopting usage-based pricing in enterprise technology. * (01:11:41) Closing segment.

Gary’s Contact Info* LinkedIn * Twitter * Crunchbase

Arcion’s Resources* Website | LinkedIn | Twitter | YouTube | Docs | Slack * “Dawn of the Data Mobility Era” (Feb 2022) * “Arcion lands $13M to help companies replicate data across platforms” (Venture Beat, Feb 2022)

Mentioned ContentContent* The Network Effects Bible (by James Currier of NFX) * Blog by Tomasz Tunguz of Redpoint Ventures

People1. Gurjeet Singh (Co-Founder and CEO of Oma Robotics, Ex-CEO/Co-Founder of Ayasdi) 2. Satish Dharmaraj (Managing Director at Redpoint Ventures)

NotesMy conversation with Gary was recorded back in March 2022. Since then, many things have happened at Arcion. I’d recommend checking out:

  1. The introduction of Arcion Cloud.
  2. This article about data mobility on The New Stack.
  3. This article about change data capture on Venture Beat.
  4. This big product launch on Oracle log reader availability featured by VentureBeat
  5. The article about the missing piece for the Modern Data Stack featured by Crunchbase
  6. Arcion is launched with Databricks Partner Connect, featured by Datanami
  7. Arcion is a proud sponsor of the Oracle Cloud World 2022 in Las Vegas, Oct 17–20. If any data professionals are attending the conference, they should stop by the Arcion booth to say hi!

About the showDatacast features long-form, in-depth conversations with practitioners and researchers in the data community to walk through their professional journeys and unpack the lessons learned along the way. I invite guests coming from a wide range of career paths — from scientists and analysts to founders and investors — to analyze the case for using data in the real world and extract their mental models (“the WHY and the HOW”) behind their pursuits. Hopefully, these conversations can serve as valuable tools for early-stage data professionals as they navigate their own careers in the exciting data universe.

Datacast is produced and edited by James Le. Get in touch with feedback or guest suggestions by emailing khanhle.1013@gmail.com.

Subscribe by searching for Datacast wherever you get podcasts or click one of the links below:

  • Listen on Spotify
  • Listen on Apple Podcasts
  • Listen on Google Podcasts

If you’re new, see the podcast homepage for the most recent episodes to listen to, or browse the full guest list.

View Details

Show Notes* (01:43) DeVaris reflected on his upbringing on the south side of Chicago and college experience at UIUC, studying Mathematics and Computer Science in the early 2000s. * (06:46) DeVaris shared his journey of learning how to program, make computers, and dive into the Internet. * (09:35) DeVaris recalled valuable lessons from interning at Intel and Cisco Systems. * (15:49) DeVaris shared his proudest accomplishments during his five years at Microsoft — first as a system engineer and then as an academic developer evangelist. * (22:06) DeVaris recalled his experience working in the gaming and music space as the Chief Developer Evangelist at Marmalade and the Chief Product Officer at Klick Push, respectively. * (27:49) DeVaris provided his perspective on the startup acquisition process. * (29:13) DeVaris unpacked his two years as a platform product manager at Zendesk, where he drove the adoption of the Zendesk Developer Platform for developers to create unique customer experiences. * (35:43) DeVaris revealed the challenges of building a technical community, given his experience at Zendesk. * (38:25) DeVaris recalled his time working for a year as the Lead Product Manager at VSCO — a startup that builds digital tools for the modern creative. * (45:12) DeVaris went over the challenges of building software for brand ambassadors and children’s playtime, given his time as the Head of Product Management at Slyce.io and the CTO at Super Heroic. * (49:39) DeVaris reflected on his desire to scratch his entrepreneurial itch. * (52:00) DeVaris gave advice for early-career technologists on evaluating startup opportunities. * (55:51) DeVaris unpacked the product challenges he encountered while building tools for developers as the Director of Product Management at Heroku. * (58:57) DeVaris touched on his one year as the first platform engineering PM hire at Twitter. * (01:02:18) DeVaris shared the founding story of Meroxa. * (01:04:28) DeVaris dissected how Meroxa’s platform architecture is designed at a high level — including a change data capture service, schema registry, event streaming service, API proxy, and incident automation framework. * (01:06:06) DeVaris explained the technical challenges associated with creating connections between data sources and destinations in real time. * (01:08:37) DeVaris zoomed into Conduit — Meroxa’s open-source, single-binary data integration tool written in Golang that provides developer-friendly streaming data orchestration. * (01:12:32) DeVaris highlighted a few customer use cases of Meroxa. * (01:16:16) DeVaris shared valuable hiring lessons to attract the right people who are excited about Meroxa’s mission and fit with Meroxa’s cultural values. * (01:18:37) DeVaris shared challenges to finding the early design partners & lighthouse customers for Meroxa. * (01:20:24) DeVaris gave advice to founders seeking the right investors for their startups. * (01:22:58) DeVaris gave advice to smart, driven operators looking to explore angel investing. * (01:25:17) DeVaris discussed the remaining barriers that prevent minorities from pursuing a technology career. * (01:30:42) DeVaris imparted lessons from photography and DJ that benefited his career in product. * (01:32:26) Closing segment.

DeVaris’ Contact Info* LinkedIn * Twitter * Website * GitHub

Meroxa’s Resources* Website | LinkedIn | Twitter | YouTube * Careers | Medium Blog * Documentation * Conduit (GitHub | Discord | Twitter | Docs)

Mentioned ContentArticles* “Hello World, Meroxa Style” (April 2021) * “Streaming Your Database Changes with Change Data Capture” (Part 1 + Part 2) * “Conduit: Streaming Data Integration for Developers” (Jan 2022) * “Why Conduit? An Evolutionary Leap Forward for Real-Time Data Integration” (Feb 2022) * “Hello Meroxa 2.0” (April 2022)

Resources for minorities* Kura Labs (A free training and job placement academy for Infrastructure Computing, DevOps, and SRE for students from underserved communities) * Free Code Camp (Learn to code — for free)

Books1. Zero To One (by Peter Thiel) 2. The Hard Thing About Hard Things (by Ben Horowitz)

People1. Tristan Handy (Co-Founder and CEO of dbt Labs) 2. Arjun Narayan (Co-Founder and CEO of Materialize) 3. Benn Stancil (Chief Analytics Officer at Mode Analytics) 4. Chad Sanderson (Head of Data Platform at Convoy)

NotesMy conversation with DeVaris was recorded back in April 2022. Since then, many things have happened at Meroxa. I’d recommend checking out:

  1. The introduction of Meroxa 2.0 and Turbine.
  2. This interview on data-driven work culture.
  3. New CDC Connectors built into Conduit.
  4. Meroxa is a recipient of DoD funding to help the US Space Force monitor aircraft health in real-time.

About the showDatacast features long-form, in-depth conversations with practitioners and researchers in the data community to walk through their professional journeys and unpack the lessons learned along the way. I invite guests coming from a wide range of career paths — from scientists and analysts to founders and investors — to analyze the case for using data in the real world and extract their mental models (“the WHY and the HOW”) behind their pursuits. Hopefully, these conversations can serve as valuable tools for early-stage data professionals as they navigate their own careers in the exciting data universe.

Datacast is produced and edited by James Le. Get in touch with feedback or guest suggestions by emailing khanhle.1013@gmail.com.

Subscribe by searching for Datacast wherever you get podcasts or click one of the links below:

  • Listen on Spotify
  • Listen on Apple Podcasts
  • Listen on Google Podcasts

If you’re new, see the podcast homepage for the most recent episodes to listen to, or browse the full guest list.

View Details

Show Notes* (01:49) Cody shared his upbringing in New Jersey, his childhood interest in science and technology, and the few people who have made big differences in his story. * (09:35) Cody went over his academic experience studying Electrical Engineering and Computer Science at MIT. * (17:51) Cody recalled his favorite classes taken at MIT. * (22:43) Cody talked about his engagement in serving as the president of MIT’s chapter of Eta Kappa Nu Honor Society and advancing online education at the MIT Office of Digital Learning. * (31:25) Cody is bullish on the future of digital learning. * (35:43) Cody expanded on his internships with Google throughout his time at MIT — doing local search quality and YouTube analytics. * (42:31) Cody described the challenges of dealing with high-frequency trading data from his one year working as a junior data scientist at the Vendor Data Group of Jump Trading in Chicago. * (46:50) Cody reflected on his decision to embark on a Ph.D. journey in Computer Science at Stanford University. * (51:54) Cody mentioned his participation in the DAWN project, specifically DAWNBench, an end-to-end deep learning benchmark and competition. * (54:21) Cody unpacked the evolution of MLPerf, an industry-standard benchmark for the training and inference performance of ML models. * (56:52) Cody walked through the motivation and empirical work in his paper “Selection via Proxy: Efficient Data Selection for Deep Learning.” * (59:34) Cody discussed his paper “Similarity Search for Efficient Active Learning and Search of Rare Concepts.” * (01:06:32) Cody shared his learnings about bringing ML from research to industry from his advisors, Matei Zaharia and Peter Bailis — who were both academics and startup founders simultaneously. * (01:09:19) Cody went over key trends in the emerging Data-Centric AI community — given his involvement with the Data-Centric AI workshop at NeurIPS 2021 and the DataPerf benchmark suite. * (01:12:19) Cody shared lessons learned about finding product-market fit as the founder of Coactive AI — which brings unstructured data into the world of SQL and the big data tools that teams already love. * (01:15:34) Cody emphasized the importance of focusing on the HR function and defining cultural guiding principles for any early-stage startup founder. * (01:21:05) Cody provided his perspective on the differences and similarities between being a researcher and a founder. * (01:23:47) Closing segment.

Cody’s Contact Info* Website * Twitter * LinkedIn * Google Scholar

Coactive AI’s Resources* Website * Twitter * LinkedIn * Culture Values

Mentioned ContentTalk* “Digging Deeper: How a Few Extra Moments Can Change Lives” (TEDxStanford 2017) * “Data Selection for Data-Centric AI” (Stanford MLSys 2022)

Research* “Probabilistic Use Cases: Discovering Behavioral Patterns for Predicting Certification” (2015) * DAWNBench: An End-to-End Deep Learning Benchmark and Competition (Dec 2017) * “MLPerf: An Industry Standard Benchmark Suite for Machine Learning Performance” (Feb 2020) * “Selection via Proxy: Efficient Data Selection for Deep Learning” (Oct 2020) * “Similarity Search for Efficient Active Learning and Search of Rare Concepts” (July 2021) * DataPerf, a new benchmark suite for machine learning datasets and data-centric algorithms (Dec 2021)

People* Matei Zaharia (Cody’s Ph.D. Advisor, Co-Creator of Apache Spark, Co-Founder of Databricks) * Fei-Fei Li (Professor of Computer Science at Stanford, Creator of ImageNet Dataset) * Michael Bernstein (Professor of Computer Science at Stanford with a focus on Human-Computer Interaction)

Books1. “No Rule Rules: Netflix and the Culture of Reinvention” (by Reed Hastings) 2. “What You Do Is Who You Are: How to Create Your Work Business Culture” (by Ben Horowitz) 3. “The Inner Game of Tennis: The Classical Guide to Peak Performance” (by Timothy Gallwey)

NotesMy conversation with Cody was recorded back in January 2022. Since then, many things have happened at Coactive AI. I’d recommend:

  1. Attending Cody’s upcoming talk at Snorkel’s The Future of Data-Centric AI.
  2. Reviewing the DataPerf workshop at ICML 2022.
  3. Reading the CoactiveAI blog post on bringing UI props to MLOps.
  4. Watching Cody’s CBS News interview back in February 2022.

About the showDatacast features long-form, in-depth conversations with practitioners and researchers in the data community to walk through their professional journeys and unpack the lessons learned along the way. I invite guests coming from a wide range of career paths — from scientists and analysts to founders and investors — to analyze the case for using data in the real world and extract their mental models (“the WHY and the HOW”) behind their pursuits. Hopefully, these conversations can serve as valuable tools for early-stage data professionals as they navigate their own careers in the exciting data universe.

Datacast is produced and edited by James Le. Get in touch with feedback or guest suggestions by emailing khanhle.1013@gmail.com.

Subscribe by searching for Datacast wherever you get podcasts or click one of the links below:

  • Listen on Spotify
  • Listen on Apple Podcasts
  • Listen on Google Podcasts

If you’re new, see the podcast homepage for the most recent episodes to listen to, or browse the full guest list.

View Details

Show Notes* (02:18) Merav talked about her undergraduate experience at McGill University studying Psychology and Sociology. * (04:33) Merav discussed important attributes of an exceptional teacher given her two years teaching elementary special education in NYC public schools through the Teach For America program. * (08:19) Merav commented on her time working at the International Baccalaureate Organization and working as a Kaplan GRE instructor. * (10:57) Merav shared the backstory behind the founding of Data Society, a predictive analytics training and consulting company (co-founded with Dmitri Adler and John Nader). * (14:15) Merav reflected on her journey into programming. * (17:16) Merav explained why data science training should be industry-tailored for maximum success. * (20:57) Merav talked about how Data Society creates and evaluates its training curriculum. * (23:59) Merav provided an example of how Data Society provides customized AI solutions to inform decisions, automate time-consuming manual processes, and solve complex data challenges for its clients. * (27:38) Merav brought up challenges that hinder the adoption of data science in the government sector. * (29:49) Merav unpacked the six different steps for organizations to start moving up the data analytics maturity model. * (33:07) Merav dissected meldR, Data Society’s internal product built for Learning and Development teams in healthcare. * (36:24) Merav reflected on bootstrapping Data Society in the early days (look at this 2016 Kickstarter campaign). * (39:48) Merav discussed the shift from a B2C to a B2B model for Data Society and scoring partnerships with Fortune 500 companies and federal agencies. * (42:47) Merav shared valuable hiring lessons to attract the right people who are excited about the mission of Data Society. * (45:22) Merav shared her experience shaping the remote work culture. * (49:05) Merav touched on initiatives at Data Society to bring more goodness to the world. * (50:28) Merav provided different ways to engage more women in data science (via the Women Data Scientists DC Meetup and DCFemTech). * (53:17) Merav predicted the evolution of education in the next 3 to 5 years. * (55:29) Closing segment.

Merav’s Contact Info* LinkedIn * Twitter

Data Society’s Resources* Website * Twitter * LinkedIn

Mentioned ContentArticles* “Is Your Enterprise Data-Driven?” (May 2021) * “Why Data Science Training Should Be Industry-Tailored for Maximum Success” (August 2021) * “Female Founders: Merav Yuravlivker of Data Society On The Five Things You Need To Thrive and Succeed as a Woman Founder” (Sep 2021)

People* DJ Patil (The first Chief Data Scientist of the US) * Hilary Mason (Co-Founder of Hidden Door) * Avriel Epps-Darling (Ph.D. candidate, Ford fellow, and Presidential Scholar at Harvard University)

Book* Weapons of Math Destruction (by Cathy O’Neil)

NotesMy conversation with Merav was recorded back in December 2021. Since then, many things have happened at Data Society. I’d recommend:

  1. Reading Merav’s articles on Forbes about creating a culture of data sharing, assessing data literacy, and communication in the learning process.
  2. Reading Data Society’s white papers about data science in research and data science in healthcare.
  3. Checking out the Camelsback product for risk assessment in financial services.
  4. Trying out the Data DNA assessment tool for organizations’ data maturity.

Finally, Merav was also just recognized as one of the DC region's 40 Under 40. The awards are given annually to recognize the outstanding achievements of young leaders in the Washington, DC, area who lead the community forward through hard work, philanthropy, and community engagement.

View Details

Show Notes* (01:46) Douwe went over formative experiences catching the programming virus at the age of 9, combining high school with freelance web development, and studying Computer Science at Utrecht University in college. * (03:55) Douwe shared the story behind founding a startup called Stinngo, which led him to join GitLab in 2015 as employee number 10. * (05:29) Douwe provided insights on attributes of exceptional engineering talent, given his time hiring developers and eventually becoming GitLab's first Development Lead. * (08:28) Douwe unpacked the evolution of his engineering career at GitLab. * (11:11) Douwe discussed the motivation behind the creation of the Meltano project in August 2018 to help GitLab's internal data team address the gaps that prevent them from understanding the effectiveness of business operations. * (14:38) Douwe reflected on his decision in 2019 to leave GitLab’s engineering organization and join the then 5-people Meltano team full-time. * (20:24) Douwe shared the details about Meltano's product development journey from its Version 1 to its pivot. * (26:18) Douwe reflected on the mental aspect of being the sole person whom Meltano depended on for a while. * (29:20) Douwe explained the positioning of Meltano as an open-source self-hosted platform for running data integration and transformation pipelines. * (34:54) Douwe shared details of Meltano's ideal customer profiles. * (37:45) Douwe provided a quick tour of the Meltano project, which represents the single source of truth regarding one's ELT pipelines: how data should be integrated and transformed, how the pipelines should be orchestrated, and how the various plugins that make up the pipelines should be configured. * (40:39) Douwe unpacked different components of Meltano's product strategy, including Meltano SDK, Meltano Hub, and Meltano Labs. * (45:05) Douwe discussed prioritizing Meltano's product roadmap in order to bring DataOps functionality to every step of the entire data lifecycle. * (48:53) Douwe shared the story behind spinning Meltano out of GitLab in June 2021 and raising a $4.2M Seed funding round led by GV to bring the benefits of open source data integration and DataOps to a wider audience. * (52:19) Douwe provided his thoughts behind open-source contributors in a way that can generate valuable product feedback for Meltano. * (55:43) Douwe shared valuable hiring lessons to attract the right people who align with Meltano's values. * (59:04) Douwe shared advice to startup CEOs who are experimenting with the remote work culture in our “new-normal” virtual working environments. * (01:04:10) Douwe unpacked Meltano's mission and vision as outlined in this blog post. * (01:06:40) Closing segment.

Douwe's Contact Info* GitLab * LinkedIn * Twitter * GitHub * Website

Meltano's Resources* Website | Twitter | LinkedIn | GitHub | YouTube * Meltano Documentation | Product | DataOps * Meltano SDK | Meltano Hub | Meltano Labs * Company Handbook | Community | Values | Careers

Mentioned ContentArticles* Hey, data teams - We're working on a tool just for you (Aug 2018) * To-do zero, inbox zero, calendar zero: I think that means I'm done (Sep 2019) * Meltano graduates to Version 1.0 (Oct 2019) * Revisiting the Meltano strategy: a return to our roots (May 2020) * Why we are building an open-source platform for ELT pipelines (May 2020) * Meltano spins out of GitLab, raises seed funding to bring data integration into the DataOps era (June 2021) * Meltano: The strategic foundation of the ideal data stack (Oct 2021) * Introducing your DataOps platform infrastructure: Our strategy for the future of data (Nov 2021) * Our next step for building the infrastructure for your Modern Data Stack (Dec 2021)

People* Maxime Beauchemin (Founder and CEO of Preset, Creator of Apache Airflow and Apache Superset, Angel Investor in Meltano) * Benn Stancil (Chief Analytics Officer at Mode Analytics, Well-Known Substack Writer) * The entire team at dbt Labs

NotesMy conversation with Douwe was recorded back in November 2021. Since then, many things have happened at Meltano. I'd recommend:

  • Checking out their updated company values
  • Reading Douwe's article about the DataOps Operating System on The New Stack
  • Examining Douwe's blog post about moving Meltano to GitHub
  • Looking over the announcement of Meltano 2.0 and the additional seed funding

About the showDatacast features long-form, in-depth conversations with practitioners and researchers in the data community to walk through their professional journeys and unpack the lessons learned along the way. I invite guests coming from a wide range of career paths — from scientists and analysts to founders and investors — to analyze the case for using data in the real world and extract their mental models (“the WHY and the HOW”) behind their pursuits. Hopefully, these conversations can serve as valuable tools for early-stage data professionals as they navigate their own careers in the exciting data universe.

Datacast is produced and edited by James Le. Get in touch with feedback or guest suggestions by emailing khanhle.1013@gmail.com.

Subscribe by searching for Datacast wherever you get podcasts or click one of the links below:

  • Listen on Spotify
  • Listen on Apple Podcasts
  • Listen on Google Podcasts

If you’re new, see the podcast homepage for the most recent episodes to listen to, or browse the full guest list.

View Details

Show Notes* (01:41) Mars walked through his education studying Computer Systems Engineering at The University of Auckland in New Zealand. * (03:16) Mars reflected on his overall Ph.D. experience in Computer Science at UCLA. * (05:55) Mars discussed his early research paper on a robust and scalable lane departure warning system for smartphones. * (07:13) Mars described his work on SmartFall, an automatic fall detection system to help prevent the elderly from falling. * (08:34) Mars explained his project WANDA, an end-to-end remote health monitoring and analytics system designed for heart failure patients. * (10:06) Mars recalled learnings from interning as a software engineer at Google during his Ph.D. * (14:54) Mars discussed engineering challenges while working on PHP for Google App Engine and Gboard personalization during his subsequent four years at Google. * (19:05) Mars rationalized his decision to join LinkedIn to lead an engineering team that builds the core metadata infrastructure for the entire organization. * (21:15) Mars discussed the motivation behind the creation of LinkedIn’s generalized metadata search and discovery tool, DataHub, later open-sourced in 2020. * (25:21) Mars dissected the key architecture of DataHub, which is designed to address the key scalability challenges coming in four different forms: modeling, ingestion, serving, and indexing. * (28:50) Mars expressed the challenges of finding DataHub’s early adopters internally at LinkedIn and externally later on at other companies. * (35:22) Mars shared the story behind the founding of Metaphor Data, which he co-founded with Pardhu Gunnam and Seyi Adebajo and currently serves as the CTO. * (41:55) Mars unpacked how Metaphor’s modern metadata platform serves as a system of record for any organization’s data ecosystem. * (48:07) Mars described new challenges with metadata management since the introduction of the modern data stack and key features of a great modern metadata platform (as brought up in his in-depth blog post with Ben Lorica). * (53:55) Mars explained how a modern metadata platform fits within the broader data ecosystem. * (58:30) Mars shared the hurdles to finding Metaphor Data’s early design partners and lighthouse customers. * (01:04:33) Mars shared valuable hiring lessons to attract the right people who are excited about Metaphor’s mission. * (01:07:28) Mars shared important culture-building lessons to build out a high-performing team at Metaphor. * (01:10:45) Mars shared fundraising advice for founders currently seeking the right investors for their startups. * (01:13:22) Closing segment.

Mars’ Contact Info* Twitter * LinkedIn * Google Scholar * GitHub

Metaphor Data* Website | Twitter | LinkedIn * Careers | About Page * Data Documentation | Data Collaboration

Mentioned ContentArticles* DataHub: A generalized metadata search and discovery tool (Aug 2019) * Open-sourcing DataHub: LinkedIn’s metadata search and discovery platform (Feb 2020) * Founding Metaphor Data (Dec 2020) * Metaphor and Soda partner to unify the modern data stack with trusted data (Dec 2021) * Introducing Metaphor: The Modern Metadata Platform (Nov 2021) * The Modern Metadata Platform: What, Why, and How? (Jan 2022)

Papers* SmartLDWS: A robust and scalable lane departure warning system for the smartphones (Oct 2009) * SmartFall: An automatic fall detection system based on subsequence matching for the SmartCane (April 2009) * WANDA: An end-to-end remote health monitoring and analytics system for heart failure patients (Oct 2012)

People* Benn Stancil (Chief Analytics Officer at Mode Analytics, Well-Known Substack Writer) * Tristan Handy (Co-Founder and CEO of dbt Labs, Writer of The Analytics Engineering Roundup) * Andy Pavlo (Associate Professor of Database at Carnegie Mellon University)

Books* “Working In Public” (by Nadia Eghbal) * “The Mom Test” (by Rob Fitzpatrick) * “A Thousand Brains” (by Jeff Hawkins) * “The Scout Mindset” (by Julia Galef)

NotesMy conversation with Mars was recorded back in January 2022. Since then, many things have happened at Metaphor Data. I’d recommend:

  • Visiting their brand new website
  • Reading the 3-part “Data Documentation” series on their blog (part 1, part 2, and part 3)
  • Looking over the Trusted Data landing page

About the showDatacast features long-form, in-depth conversations with practitioners and researchers in the data community to walk through their professional journeys and unpack the lessons learned along the way. I invite guests coming from a wide range of career paths — from scientists and analysts to founders and investors — to analyze the case for using data in the real world and extract their mental models (“the WHY and the HOW”) behind their pursuits. Hopefully, these conversations can serve as valuable tools for early-stage data professionals as they navigate their own careers in the exciting data universe.

Datacast is produced and edited by James Le. Get in touch with feedback or guest suggestions by emailing khanhle.1013@gmail.com.

Subscribe by searching for Datacast wherever you get podcasts or click one of the links below:

  • Listen on Spotify
  • Listen on Apple Podcasts
  • Listen on Google Podcasts

If you’re new, see the podcast homepage for the most recent episodes to listen to, or browse the full guest list.

View Details

Show Notes* (01:35) Ville recalled his education getting degrees in Computer Science from the University of Helsinki in Finland. * (04:35) Ville walked over his time working at a startup called Gurusoft that planned to commercialize self-organizing maps, a peculiar artificial neural network. * (07:17) Ville reflected on his four years as a researcher at Nokia — working on big data infrastructure, analytics, and ML open-source projects (such as Disco and Ringo). * (11:56) Ville shared the story of co-founding a startup that built a novel scriptable data platform called Bitdeli with his brother and not finding a product-market fit. * (13:58) Ville walked through AdRoll’s acquisition of Bitdeli in June 2013. * (15:49) Ville discussed the engineering challenges associated with his work at AdRoll — AdRoll Prospecting and traildb.io. * (19:33) Ville mentioned the product and leadership/management lessons during his time being AdRoll’s Head of Data and leading various data/ML efforts. * (24:43) Ville rationalized his decision to join the ML Infrastructure team at Netflix in 2017. * (27:26) Ville discussed the motivation behind the creation of Netflix’s human-centric ML infrastructure, Metaflow, later open-sourced in 2019. * (30:21) Ville unpacked the key design principles that summarize the philosophy of Metaflow, which is influenced by the unique culture at Netflix. * (35:00) Ville talked about his well-known diagram on the data infrastructure’s hierarchy of needs. * (37:33) Ville examined the technical details behind Metaflow’s integration with AWS to make it easy for users to move back and forth between their local and remote modes of development and execution. * (40:58) Ville expressed the challenges of finding Metaflow’s early adopters internally at Netflix and externally later on at other companies. * (45:13) Ville went over the strategy around prioritizing features for Metaflow’s future roadmap. * (52:22) Ville shared the story behind the founding of Outerbounds, which he co-founded with Savin Goyal and Oleg Avdeev. * (55:03) Ville provided his thoughts behind Metaflow’s contributors in a way that can generate valuable product feedback for Outerbounds. * (58:30) Ville shared valuable hiring lessons to attract the right people who are excited about Outerbounds’ mission. * (01:01:28) Ville shared upcoming initiatives that he is most excited about for Outerbounds. * (01:04:05) Ville walked through his writing process for an upcoming technical book with Manning called “Effective Data Science Infrastructure,” a hands-on guide to assembling infrastructure for data science and machine learning applications. * (01:06:34) Ville unpacked his great O’Reilly article that digs deep into the fundamentals of ML as an engineering discipline. * (01:11:03) Closing segment.

Ville’s Contact Info* LinkedIn * Twitter * GitHub

Outerbounds* Website | Twitter | LinkedIn | GitHub | YouTube * Metaflow GitHub | Metaflow Docs * Slack Community * Careers * Metaflow Resources for Data Science * Metaflow Resources for Engineering

Mentioned ContentTalks* SF Data Mining Meetup: TrailDB — Processing Trillions of Events at AdRoll (July 2016) * QConSF 2018: Human-Centric Machine Learning Infrastructure @Netflix (Feb 2019) * AWS re:Invent 2019: More Data Science with Less Engineering — ML Infrastructure at Netflix (Dec 2019) * Scale By The Bay 2019: Human-Centric ML Infrastructure at Netflix (Jan 2020) * AICamp: Metaflow — The ML Infrastructure at Netflix (Aug 2021)

Articles* Open-Sourcing Metaflow, a Human-Centric Framework for Data Science (Netflix Tech Blog, Dec 2019) * Unbundling Data Science Workflows with Metaflow and AWS Step Functions (Netflix Tech Blog, July 2020) * MLOps and DevOps: Why Data Makes It Different (O’Reilly, Oct 2021)

People* Michael Jordan (Distinguished Professor in EECS and Statistics at UC Berkeley) * Matthew Honnibal and Ines Montani (Creators of open-source NLP library spaCy) * Hadley Wickham (Chief Scientist at RStudio and Adjunct Professor of Statistics at Rice University)

Book* “The Mom Test” (by Rob Fitzpatrick)

NotesMy conversation with Ville was recorded back in October 2021. Since then, many things have happened at Outerbounds. I’d recommend:

  • Visiting Outerbounds’ new website with Metaflow resources for Data Science and Engineering
  • Watching Ville’s recent talk at Data Council Austin about the Modern Stack for ML Infrastructure
  • Buying Ville’s newly released book “Effective Data Science Infrastructure”

About the showDatacast features long-form, in-depth conversations with practitioners and researchers in the data community to walk through their professional journeys and unpack the lessons learned along the way. I invite guests coming from a wide range of career paths — from scientists and analysts to founders and investors — to analyze the case for using data in the real world and extract their mental models (“the WHY and the HOW”) behind their pursuits. Hopefully, these conversations can serve as valuable tools for early-stage data professionals as they navigate their own careers in the exciting data universe.

Datacast is produced and edited by James Le. Get in touch with feedback or guest suggestions by emailing khanhle.1013@gmail.com.

Subscribe by searching for Datacast wherever you get podcasts or click one of the links below:

  • Listen on Spotify
  • Listen on Apple Podcasts
  • Listen on Google Podcasts

If you’re new, see the podcast homepage for the most recent episodes to listen to, or browse the full guest list.

View Details

Show Notes* (01:48) Mike recalled his undergraduate experience studying Economics at Arizona State University and doing research on statistics/econometrics. * (04:59) Mike reflected on his three years working as an analyst in the Boston office of the Analysis Group. * (09:08) Mike discussed how he leveled up his programming skills at work. * (11:05) Mike shared his learnings about building effective data-driven products while working as a data scientist at Case Commons. * (17:20) Mike revisited his transition to a new role as the Director of Analytics at Harry’s, the men’s grooming brand — starting a new data team from scratch. * (23:04) Mike unpacked analytics and infrastructure challenges during his time at Harry’s — developing the data warehouse, an internal marketing attribution tool, and a fleet of systems for automated decision-making to improve efficiency. * (27:21) Mike reasoned his move to Mexico City — spending time practicing Spanish, among other things. * (32:22) Mike talked about his journey of starting a new consulting practice to help companies get more value out of their data, which was primarily shaped by his network. * (36:30) Mike shared the founding story behind Recast, whose mission is to help modern brands improve the effectiveness of their marketing dollars. * (42:09) Mike dissected the core technical problem that Recast is addressing: performing media mix modeling in the context of “programmatic” channels. * (46:14) Mike shared the story behind the inception and evolution of Locally Optimistic, a community for current and aspiring data analytics leaders. * (49:29) Mike walked through his 3-part blog series on Agile Analytics — discussing the good aspects, the bad aspects, and the adjustments needed for analytics teams to adopt the Scrum methodology. * (53:25) Mike unpacked his post “A Culture of Partnership,” — which discusses the three key activities that can help an analytics team identify the most important opportunities in the business and work effectively with key stakeholders and partner teams to drive value. * (57:25) Mike examined his seminal piece called “The Analytics Engineer,” which generated much attention from the analytics community — which argues that the analytics engineer can provide a multiplier effect on the output of an analytics team. * (01:03:24) Mike shared the motivation and pedagogical philosophy behind the Analytics Engineers Club (co-founded with Claire Carroll), which provides a training course for data analysts looking to improve their engineering skills. * (01:07:57) Mike anticipated the evolution of the quickly evolving modern data stack (read his Fivetran article “The Modern Data Science Stack”). * (01:09:22) Mike unpacked how organizations can build, start, and maintain the data quality flywheel (read his Datafold article “The Data Quality Flywheel”). * (01:11:40) Mike shared his thoughts regarding the challenge of sharing complex analyses. * (01:13:15) Closing segment.

Mike’s Contact Info* Twitter * Website * LinkedIn * GitHub

Further Resources* Recast * Locally Optimistic * Analytics Engineers Club

Mentioned ContentArticles* “Learning a language is hard” (Personal Blog, Jan 2020) * “Modern Media Mix Modeling” (Recast Blog) * “Agile Analytics, Part 1: The Good Stuff” (Locally Optimistic Blog, May 2018) * “Agile Analytics, Part 2: The Bad Stuff” (Locally Optimistic Blog, June 2018) * “Agile Analytics, Part 3: The Adjustments” (Locally Optimistic Blog, July 2018) * “A Culture of Partnership” (Locally Optimistic Blog, March 2019) * “The Analytics Engineer” (Locally Optimistic Blog, Jan 2019) * “Data Education Is Broken” (Analytics Engineering Club, June 2021) * “Teaching The Real Tools” (Analytics Engineering Club, Aug 2021) * “The Modern Data Science Stack” (Fivetran Blog, Oct 2020) * “The Data Quality Flywheel” (Datafold Blog, Nov 2020) * “Knowledge Sharing” (Personal Blog, Sep 2020) * “TDD for ELT” (Personal Blog, Sep 2020) * “Are Data Catalogs Curing the Symptom or the Disease?” (Personal Blog, Dec 2020)

People* Claire Carroll (Co-Instructor of Analytics Engineering Club, Product Manager of Hex, previous Community Manager of dbt Labs) * Drew Banin (Head of Product at dbt Labs) * Barry McCardel (Co-Founder and CEO of Hex)

NotesMy conversation with Michael was recorded back in October 2021. Since then, Michael has been active in his work projects. I’d recommend:

  • Following the Analytics Engineering Club for upcoming sessions (They are currently teaching their second summer cohort)
  • Reading his collaboration blog post with Reforge on the attribution stack
  • Consuming his Recast content explaining why marketing-mix modeling is hard and laying out the checklist for evaluating an MMM vendor

About the showDatacast features long-form, in-depth conversations with practitioners and researchers in the data community to walk through their professional journeys and unpack the lessons learned along the way. I invite guests coming from a wide range of career paths — from scientists and analysts to founders and investors — to analyze the case for using data in the real world and extract their mental models (“the WHY and the HOW”) behind their pursuits. Hopefully, these conversations can serve as valuable tools for early-stage data professionals as they navigate their own careers in the exciting data universe.

Datacast is produced and edited by James Le. Get in touch with feedback or guest suggestions by emailing khanhle.1013@gmail.com.

Subscribe by searching for Datacast wherever you get podcasts or click one of the links below:

  • Listen on Spotify
  • Listen on Apple Podcasts
  • Listen on Google Podcasts

If you’re new, see the podcast homepage for the most recent episodes to listen to, or browse the full guest list.

View Details

Show Notes* (01:37) Caitlin went over her college experience studying Computer Science at Stanford University in the early 2010s. * (03:55) Caitlin talked about her teaching experience for CS 106A and CS 103. * (07:09) Caitlin shared valuable lessons from completing software engineering internships at Harvard University, Facebook, and Palantir. * (10:06) Caitlin walked over technical and organizational challenges during her time at Palantir — building products for both government/commercial customers and working with designers/infrastructure engineers to deliver full-stack applications to the field. * (12:01) Caitlin explained why Palantir is composed of “loosely individual startups.” * (14:56) Caitlin recalled learning curves during her transition to a tech lead role at Palantir — becoming responsible for the technical architecture and code quality of the product, mentorship and growth of the engineers, and the product direction and prioritization of features. * (18:31) Caitlin discussed her time as a Data Engineering Manager at Remix Technologies — leading the team that builds geospatial data pipelines on top of AWS, Postgres/PostGIS, and Apache Airflow. * (24:45) Caitlin reflected on valuable leadership and people management lessons absorbed during her transition to growing and developing diverse and inclusive engineering teams. * (29:05) Caitlin shared the founding story of Hex, the modern data workspace for teams, alongside her co-founders Barry and Glen. * (32:58) Caitlin talked about Hex’s ideal users (the “analytically technical” who need better tools to access and manage more sophisticated workflows) and introduced Hex’s Logic View. * (35:22) Caitlin examined the collaboration challenges in data teams and revealed Hex’s Library to address some of the shortcomings. * (39:59) Caitlin shared her thoughts on the evolution of data science notebooks. * (42:14) Caitlin unpacked the nuanced problem of justifying data ROI to functional stakeholders and described Hex’s interactive App Builder. * (45:17) Caitlin shared exciting development in the horizon of Hex’s product roadmap. * (46:37) Caitlin shared valuable hiring lessons to attract the right people who are excited about Hex’s mission. * (52:10) Caitlin shared the hurdles to find the early design partners and lighthouse customers of Hex. * (56:01) Caitlin shared upcoming go-to-market initiatives that she’s most excited about for Hex. * (58:24) Caitlin shared fundraising advice for founders currently seeking the right investors for their startups. * (01:01:42) Closing segment.

Caitlin’s Contact Info* LinkedIn * Twitter

Hex’s Resources* Website | Twitter | LinkedIn * Logic View | App Builder | Knowledge Library * Docs | Blog | Gallery * Customers | Careers | Integrations | Pricing

Mentioned ContentArticles* “Long Live Code” (June 2020) * “Don’t Tell Your Data Team’s ROI Story” (Aug 2020) * “The Sharing Gap” (Oct 2020)

People* Tristan Handy (Founder and CEO of dbt Labs) * Claire Carroll (Product Manager of Hex, previous Community Manager of dbt Labs) * Wes McKinney (Creator of Pandas and Arrow, Co-Founder and CTO of Voltron Data) * DeVaris Brown (Co-Founder and CEO of Meroxa)

Book* “Mindset: The New Psychology of Success” (by Carol Dweck)

NotesMy conversation with Caitlin was recorded back in Fall 2021. Since then, many things have happened at Hex. I’d recommend looking at:

  • Caitlin’s piece announcing Hex’s SOC 2 Type II report to reflect Hex’s commitment to security
  • Caitlin’s recent talk at Data Council Austin about implementing reactive notebooks with iPython
  • The release of Hex Knowledge Library, a new way to publish and discover data work
  • Hex’s $16M Series A (led by Redpoint Ventures) and $52M Series B (led by a16z along with Snowflake, Databricks, and existing investors)
  • Hex’s increasing list of customers such as AngelList, Fivetran, Hightouch, Loom, Mixpanel, Notion, Ramp, Replicated, SeatGeek, etc.

View Details

Show Notes* (00:43) Kashish shared briefly about his upbringing in Atlanta and his early interest in STEM subjects. * (02:38) Kashish described his overall academic experience studying Economics, Management, and Computer Science at the University of Pennsylvania. * (05:53) Kashish walked over the Machine Learning classes and projects throughout his MSE degree in Robotics. * (09:02) Kashish shared valuable lessons learned from multiple internships throughout his undergraduate: data science at Implantable Provider Group, investment analysis at Tree Line, and product management at LYNK. * (13:14) Kashish told the anecdotes that enabled him to realize his passion for building startups. * (17:14) Kashish recapped his learning about venture capital from spending a summer as an analyst in early-stage deep-tech companies at Bessemer Venture Partners in New York. * (22:09) Kashish shared learnings from his entrepreneurial stints at an early age. * (26:12) Kashish talked through his decision to move to San Francisco after college (Read his blog post explaining how he moved here without a job and a home). * (29:04) Kashish recalled his experience working on a project called Carry (an executive assistant for travel on Slack) with his friend Tejas Manohar and going through Y Combinator. * (36:40) Kashish shared the founding story of Hightouch, a data platform that syncs customer data from the data warehouse to CRM, marketing, and support tools. * (44:15) Kashish emphasized the importance of speed and execution around different pivots that led to Hightouch. * (46:35) Kashish unpacked the notion of Operational Analytics, an approach to analytics that shifts the focus from simply understanding data to putting that data to work in the tools that run your business. * (49:46) Kashish dissected Hightouch’s market-leading Reverse ETL, which is the process of copying data from a data warehouse to operational systems of record. * (54:51) Kashish discussed Hightouch Audiences, used primarily by larger B2C customers, that allows marketing teams to build audiences and filters on top of existing data models. * (58:09) Kashish explained how the “Reverse ETL” concept fits into the quickly evolving modern data stack. * (01:00:26) Kashish shared how the Hightouch team prioritizes their product roadmap, given the high number of customer requests. * (01:02:47) Kashish shared valuable hiring lessons to attract the right people who are excited about Hightouch’s mission. * (01:05:13) Kashish shared the hurdles to find the early design partners and lighthouse customers of Hightouch. * (01:08:06) Kashish explained how Hightouch prices by destinations, reflecting the value customers get from using the product and helping them predict costs over time. * (01:10:32) Kashish shared upcoming go-to-market initiatives that he is most excited about for Hightouch. * (01:14:36) Kashish shared fundraising advice for founders currently seeking the right investors for their startups. * (01:17:47) Kashish emphasized the industry recognition of the Reverse ETL market. * (01:19:47) Closing segment.

Kashish’s Contact Info* LinkedIn * Twitter * GitHub * Website * Medium

Hightouch’s Resources* Website | Twitter | LinkedIn * Data Features | Hightouch Audiences | Hightouch Notify * Docs | Blog * Customers | Careers | Pricing

Mentioned ContentArticles* “On Moving to SF Jobless and Homeless” (Aug 2018) * “Hightouch Ushers In The Era of Operational Analytics” (March 2021) * “The State of Reverse ETL” (May 2021) * “What is Operational Analytics?” (July 2021) * “Hightouch Has Raised a Series A!” (July 2021) * “Hightouch Raises $12M to Empower Business Teams With Operational Analytics” (July 2021) * “The Cloud 100 Rising Stars 2021” (Aug 2021) * “What is Reverse ETL?” (Nov 2021)

Companies* dbt Labs * Shipyard * Big Time Data

Book* “The Hard Things About Hard Things” (by Ben Horowitz)

NotesMy conversation with Kashish was recorded back in August 2021. Since then, many things have happened at Hightouch. I’d recommend looking at:

  • Kashish’s piece about Hightouch’s transition from Reverse ETL to becoming a Data Activation company
  • Kashish’s recent talk at Data Council Austin about the current state of Data Apps built on top of the warehouse and the future as warehouses become even faster.
  • The release of Hightouch Notify that sends notifications on top of the data warehouse
  • Hightouch’s Series B funding of $40M back in November 2021

Finally, Kashish lets me know that back in August, Hightouch were only 25 people. Now, the company is 70-person strong!

About the showDatacast features long-form, in-depth conversations with practitioners and researchers in the data community to walk through their professional journeys and unpack the lessons learned along the way. I invite guests coming from a wide range of career paths — from scientists and analysts to founders and investors — to analyze the case for using data in the real world and extract their mental models (“the WHY and the HOW”) behind their pursuits. Hopefully, these conversations can serve as valuable tools for early-stage data professionals as they navigate their own careers in the exciting data universe.

Datacast is produced and edited by James Le. Get in touch with feedback or guest suggestions by emailing khanhle.1013@gmail.com.

Subscribe by searching for Datacast wherever you get podcasts or click one of the links below:

  • Listen on Spotify
  • Listen on Apple Podcasts
  • Listen on Google Podcasts

If you’re new, see the podcast homepage for the most recent episodes to listen to, or browse the full guest list.

View Details

Show Notes (01:53) Alessya shared her formative experiences growing up in Kazakhstan, coming to Washington during high school, and discovering a passion and extreme aptitude for mathematics. * (04:20) Alessya described her undergraduate experience studying Applied Mathematics at the University of Washington.* * (08:00) Alessya talked about impactful projects she contributed to while working as a software developer at Amazon’s quality assurance and DevOps organizations. * (12:29) Alessya went over critical responsibilities during her time as Amazon’s Technical Program Manager. * (17:06) Alessya talked about the process of building and getting adoption for an internal Machine Learning platform at Amazon. * (20:42) Alessya shared her biggest takeaways from Amazon’s culture of customer obsession and operational excellence. * (23:26) Alessya revisited her period enrolling in UW’s Master of Science in Entrepreneurship Program and highlighted two core entrepreneurial muscles developed: networking and negotiation. * (28:58) Alessya provided insights on the startup ecosystem and ML community in Seattle. * (34:47) Alessya walked through her period serving as the CTO in Residence at Allen Institute for AI and evaluating a range of AI technologies for viability and product readiness. * (37:12) Alessya shared the backstory behind the founding of WhyLabs, an AI observability platform built to enable every enterprise to run AI with certainty (read her blog post about early misadventures with AI at Amazon that inspired the incubation of WhyLabs at AI2). * (42:23) Alessya examined what makes an AI solution robust and responsible. * (46:09) Alessya dissected the anatomy of an enterprise AI Observability platform. * (49:58) Alessya explained why data logging is a critical missing component in the production ML stack and described whylogs, an open-source ML data logging library from WhyLabs. * (54:12) Alessya shared valuable hiring lessons to attract the right people who are excited about WhyLabs’ mission. * (57:03) Alessya shared tactics to find and engage contributors to whylogs. * (58:10) Alessya shared the hurdles to find the early design partners and lighthouse customers of WhyLabs. * (01:02:28) Alessya shared upcoming go-to-market initiatives that she is most excited about for WhyLabs. * (01:03:54) Alessya explained what it felt to be recognized as the CEO of the year for the Pacific Northwest startup community last year and shared her perspective on work-life balance. * (01:07:43) Closing segment.

Alessya’s Contact Info* LinkedIn * Twitter

WhyLabs’s Resources* Website * whylogs * Slack Community * Blog * LinkedIn | Twitter | Facebook | YouTube | GitHub * What is AI Observability?

Mentioned ContentArticles + Talks* “Introducing WhyLabs, a Leap Forward in AI Reliability” (Sep 2020) * “WhyLabs: The AI Observability Platform” (Sep 2020) * “whylogs: Embrace Data Logging Across Your ML Systems” (Sep 2020) * “Who Said Moms Can’t CEO?” (May 2021) * “The Critical Missing Component in the Production ML Stack” (May 2021)

People* Cassie Kozyrkov (Chief Decision Scientist at Google) * Dan Jeffries (Chief Evangelist at Pachyderm and Founder of AI Infrastructure Alliance) * Michael Petrochuk (Founder and CTO of WellSaid Labs)

Book* “The Hard Things About Hard Things” (by Ben Horowitz)

NotesMy conversation with Alessya was recorded back in August 2021. Since then, many things have happened at WhyLabs.I'd recommend looking at:

  • The self-service release of AI Observatory
  • Series A funding
  • Exploring their new integrations with Teachable Hub, UbiOps, Valohai, and Superb AI
  • Launch of their listing on AWS Marketplace
  • Their article on How Observability Uncovers the Effects of ML Technical Debt
  • Their achievement of SOC 2 Type 2 certification

whylogs is evolving to a new iteration that will be even more usable and more useful than it was before. With the launch of whylogs v1 in May, users will be able to create data profiles in a fraction of the time and with a much simpler API. Additionally, WhyLabs built-in handy features such as the profile visualizer (which allows users to visualize one or multiple profiles for exploration and comparison) and constraints (which allow users to validate the quality of their data as it flows through their data pipelines).

About the showDatacast features long-form, in-depth conversations with practitioners and researchers in the data community to walk through their professional journeys and unpack the lessons learned along the way. I invite guests coming from a wide range of career paths — from scientists and analysts to founders and investors — to analyze the case for using data in the real world and extract their mental models (“the WHY and the HOW”) behind their pursuits. Hopefully, these conversations can serve as valuable tools for early-stage data professionals as they navigate their own careers in the exciting data universe.

Datacast is produced and edited by James Le. Get in touch with feedback or guest suggestions by emailing khanhle.1013@gmail.com.

Subscribe by searching for Datacast wherever you get podcasts or click one of the links below:

  • Listen on Spotify
  • Listen on Apple Podcasts
  • Listen on Google Podcasts

If you’re new, see the podcast homepage for the most recent episodes to listen to, or browse the full guest list.

View Details

Show Notes* (02:00) Evan shared his upbringing, born and raised in a small coastal town on New Zealand’s North Island and later studied Software Engineering and Business. * (03:55) Evan recalled working as a software solution architect at NEC Corporation back in New Zealand. * (06:17) Evan talked about his decision to join Twilio in 2011 as one of the company’s early employees right after its Series B financing. * (08:40) Evan shared his perspectives on joining startups and big companies as a new grad. * (13:01) Evan provided insights on attributes of exceptional sales engineers, given his time building the first iteration of Twilio’s global pre-sales team. * (17:30) Evan unpacked the evolution of his career at Twilio — working as a product manager, a director of product & engineering, and a general manager of IoT & wireless. * (22:51) Evan dissected Twilio’s unique “middle-out” sales strategy, which has hugely impacted the company’s incredible growth from Series B through to IPO and beyond. * (29:03) Evan went over the untapped opportunity being enabled by new cellular IoT technologies. * (33:25) Evan explained his decision to embark on a new journey as the CEO of Fin.com after a decade at Twilio. * (37:26) Evan talked about the need for workflow automation and how Fin’s product features are built to address that. * (40:35) Evan went over Fin’s remote performance optimization capabilities that help teams thrive in a remote-first environment. * (42:56) Evan shared valuable hiring lessons to attract the right leaders who are excited about Fin’s mission. * (45:38) Evan shared the hurdles his team has to go through while finding early customers for Fin (as it pivoted to building a SaaS product). * (48:02) Evan talked about the qualities of Jeff Lawson that made him such a great CEO. * (50:41) Closing segment.

Evan’s Contact Info* Twitter * LinkedIn

Fin’s Resources* Website * LinkedIn * Twitter * “Fin.com Raises $20M from Coatue” (Sep 2021) * “Customers Operations Benchmarks for 2022” (Nov 2021) * “Fin’s new Experiments Product Enables CX teams to Confidently Deliver Business Process Changes that Maximize Business Impact” (Dec 2021)

Mentioned ContentPeople* Jack Dorsey * Bret Taylor * Paul Buchheit

Book* “Startup CXO: A Field Guide to Scaling Up Your Company’s Critical Functions and Teams” (by Matt Blumberg)

About the showDatacast features long-form, in-depth conversations with practitioners and researchers in the data community to walk through their professional journeys and unpack the lessons learned along the way. I invite guests coming from a wide range of career paths — from scientists and analysts to founders and investors — to analyze the case for using data in the real world and extract their mental models (“the WHY and the HOW”) behind their pursuits. Hopefully, these conversations can serve as valuable tools for early-stage data professionals as they navigate their own careers in the exciting data universe.

Datacast is produced and edited by James Le. Get in touch with feedback or guest suggestions by emailing khanhle.1013@gmail.com.

Subscribe by searching for Datacast wherever you get podcasts or click one of the links below:

  • Listen on Spotify
  • Listen on Apple Podcasts
  • Listen on Google Podcasts

If you’re new, see the podcast homepage for the most recent episodes to listen to, or browse the full guest list.

View Details

Show Notes* (01:51) Nick shared his formative experiences of her childhood — moving between different schools, becoming interested in Math, and graduating from UCLA at the age of 19. * (05:45) Nick recalled working as a quant analyst focused on emerging market debt at BlackRock. * (09:57) Nick went over his decision to join Airbnb as a data scientist on their growth team in 2014. * (12:17) Nick discussed how data science could be used to drive community growth on the Airbnb platform. * (16:35) Nick led the data architecture design and experimentation platform for Airbnb Trips, one of Airbnb’s biggest product launches in 2016. * (20:40) Nick provided insights on attributes of exceptional data science talent, given his time interviewing hundreds of candidates to build a data science team from 20 to 85+. * (23:50) Nick went over his process of leveling up his product management skillset — leading Airbnb’s Machine Learning teams and growing the data organization significantly. * (26:56) Nick emphasized the importance of flexibility in his work routine. * (29:27) Nick unpacked the technical and organizational challenges of designing and fostering the adoption of Bighead, Airbnb’s internal framework-agnostic, end-to-end platform for machine learning. * (34:54) Nick recalled his decision to leave Airbnb and become the Head of Data at Branch, which delivers world-class financial services to the mobile generation. * (37:24) Nick unpacked key takeaways from his Bay Area AI meetup in 2019 called “ML Infrastructure at an Early Stage Startup” related to his work at Branch. * (40:55) Nick discussed his decision to pursue a startup idea in the analytics space rather than the ML space. * (43:36) Nick shared the founding story of Transform, whose mission is to make data accessible by way of a metrics store. * (49:54) Nick walked through the four key capabilities of a metrics store: semantics, performance, governance, and interfaces + introduced Metrics Framework (Transform’s capability to create company-wide alignment around key metrics that scale with an organization through a unified framework). * (55:58) Nick unpacked Metrics Catalog — Transform’s capability to eliminate repetitive tasks by giving everyone a single place to collaborate, annotate data charts, and view personalized data feeds. * (59:57) Nick dissected Metrics API — Transform’s capability to generate a set of APIs to integrate metrics into any other enterprise tools for enriched data, dimensional modeling, and increased flexibility. * (01:02:41) Nick explained how metrics store fit into a modern data analytics stack * (01:05:57) Nick shared valuable hiring lessons finding talents who fit with Transform’s cultural values. * (01:12:27) Nick shared the hurdles his team has to go through while finding early design partners for Transform. * (01:15:38) Nick shared upcoming go-to-market initiatives that he’s most excited about for Transform. * (01:17:46) Nick shared fundraising advice for founders currently seeking the right investors for their startups. * (01:20:45) Closing segment.

Nick’s Contact Info* LinkedIn * Twitter * Medium

Transform's Resources* Website * Blog * LinkedIn | Twitter

Mentioned ContentArticles + Talks* “ML Infrastructure at an Early Stage” (March 2019) * “Why We Founded Transform” (June 2021) * “My Experience with Airbnb’s Early Metrics Store” (June 2021) * “The 4 Pillars of Our Workplace Culture” (Aug 2021)

People* Airbnb’s Metrics Repo Team (Paul Yang, James Mayfield, Will Moss, Jonathan Parks, and Aaron Keys) * Maxime Beauchemin (Founder and CEO of Preset, Creator of Apache Airflow and Apache Superset) * Emilie Schario (Data Strategist In Residence at Amplify Partners, Previously Head of Data at Netlify)

Book* “High-Output Management” (by Andy Grove)

NotesMy conversation with Nick was recorded back in July 2021. Since then, many things have happened at Transform. I’d recommend:

  • Registering for the Metrics Store Summit that will happen at the end of April 2022
  • Reviewing the piece about 4 Pillars of Transform’s Workplace Culture
  • Reading Nick’s post on the brief history of the metrics store
  • Exploring Transform’s integrations with Mode, Hex, and Google Sheets

About the showDatacast features long-form, in-depth conversations with practitioners and researchers in the data community to walk through their professional journeys and unpack the lessons learned along the way. I invite guests coming from a wide range of career paths — from scientists and analysts to founders and investors — to analyze the case for using data in the real world and extract their mental models (“the WHY and the HOW”) behind their pursuits. Hopefully, these conversations can serve as valuable tools for early-stage data professionals as they navigate their own careers in the exciting data universe.

Datacast is produced and edited by James Le. Get in touch with feedback or guest suggestions by emailing khanhle.1013@gmail.com.

Subscribe by searching for Datacast wherever you get podcasts or click one of the links below:

  • Listen on Spotify
  • Listen on Apple Podcasts
  • Listen on Google Podcasts

If you’re new, see the podcast homepage for the most recent episodes to listen to, or browse the full guest list.

View Details

Show Notes* (01:29) Jeremiah reflected on his academic interest studying Statistics and Economics at Harvard. * (05:33) Jeremiah recalled his four years as a market risk manager at King Street Capital Management. * (07:18) Jeremiah explained how the training in risk management has made a huge impact in his career as a startup founder. * (09:48) Jeremiah then founded his own consultancy Lowin Data Company that designed and built ML systems for time series data. * (12:38) Jeremiah mentioned his fascination with the rapid growth of machine learning in the past decade. * (15:54) Jeremiah talked about his contribution to the Apache Airflow project and lessons learned about open-source development/governance. * (21:48) Jeremiah unpacked the notion of negative engineering and shared the story behind the inception of Prefect. * (27:24) Jeremiah dissected Prefect Core, the open-source framework that is stocked with all the necessary components for designing, building, testing, and running powerful data applications. * (32:45) Jeremiah went over the advanced enterprise features of Prefect Cloud that complement users of Prefect Core. * (36:04) Jeremiah discussed Prefect's product strategy (read his blog post "Toward Dataflow Automation," which distinguishes the difference between what a company makes and what a company sells). * (40:44) Jeremiah explained how Prefect users can take advantage of the hybrid execution model. * (47:08) Jeremiah walked through Prefect Server and Prefect UI that enable users to run parts of Prefect Cloud locally. * (50:27) Jeremiah talked about how his team has gradually open-sourced the Prefect platform. * (51:38) Jeremiah explained how Prefect settles into a "success-based pricing" model, where the cost is based entirely on the number of tasks users run successfully each month. * (54:15) Jeremiah shared how to nurture a highly active community of open-source contributors to Prefect Core. * (58:23) Jeremiah unpacked Prefect's hiring strategy, which emphasizes the importance of hiring a team diverse in thoughts, backgrounds, makeups, and experiences (read this fantastic guide to building a high-performance team on Prefect's website). * (01:07:02) Jeremiah shared fundraising advice for founders currently seeking the right investors for their startups. * (01:11:53) Jeremiah unpacked the two key pillars central to Prefect’s hyper-adoption within the data world: expansion and product. * (01:14:09) Closing segment.

Jeremiah's Contact Info* LinkedIn * Twitter * Medium * GitHub

Prefect's Resources* Website * GitHub | Slack | Documentation | Twitter | Meetup * Community Updates * The Prefect Guide to Building A High-Performance Team (April 2021) * Prefect Cloud * Prefect Core * Prefect's Hybrid Model

Mentioned ContentArticles* "Positive and Negative Engineering" (Oct 2018) * "The Golden Spike" (Jan 2019) * "Prefect is Open-Source!" (March 2019) * "Towards Dataflow Automation" (June 2019) * "The Prefect Hybrid Model" (Feb 2020) * "Project Earth" (March 2020) * "Open-Sourcing The Prefect Platform" (March 2020) * "Your Code Will Fail (But That's Okay)" (May 2020) * "Liftoff: Prefect's Series A" (Feb 2021) * "Escape Velocity: Prefect's Series B" (June 2021)

Talks and Podcasts* "Invest Like The Best" (Jan 2017) * "Task Failed Successfully" (PyData DC 2018) * "Software Engineering Daily" (April 2020) * "The OSS Startup Podcast" (Nov 2021) * "The Sequel Show" (Jan 2022)

People* Vicki Boykis (ML Engineer at Tumblr, Newsletter Writer of Normcore Tech) * Chris Riccomini (Software Engineer at WePay, Contributor of Airflow, Investor/Advisor at Prefect) * Justin Gage (Newsletter Writer of Technically)

Books* "Creativity Inc." (by Ed Cadmull) * "The Hitchhiker's Guide to the Galaxy" (by Douglas Adams, Eoin Colfer, and Thomas Tidholm) * "Shoe Dog" (by Phil Knight)

NotesMy conversation with Jeremiah was recorded back in July 2021. Since then, many things have happened at Prefect:

  • The 2021 Growth Report
  • The releases of Prefect Orion and Prefect Radar as part of the product roadmap
  • The announcement of Prefect's Premier Partnership Program for trusted partners
  • The introduction of Prefect Discourse for data engineers
  • The latest drop of Prefect 2.0!

About the showDatacast features long-form, in-depth conversations with practitioners and researchers in the data community to walk through their professional journeys and unpack the lessons learned along the way. I invite guests coming from a wide range of career paths — from scientists and analysts to founders and investors — to analyze the case for using data in the real world and extract their mental models (“the WHY and the HOW”) behind their pursuits. Hopefully, these conversations can serve as valuable tools for early-stage data professionals as they navigate their own careers in the exciting data universe.

Datacast is produced and edited by James Le. Get in touch with feedback or guest suggestions by emailing khanhle.1013@gmail.com.

Subscribe by searching for Datacast wherever you get podcasts or click one of the links below:

  • Listen on Spotify
  • Listen on Apple Podcasts
  • Listen on Google Podcasts

If you’re new, see the podcast homepage for the most recent episodes to listen to, or browse the full guest list.

View Details

Show Notes* (02:00) Shinji reflected on her academic experience studying Software Engineering at the University of Waterloo in the late 2000s. * (04:19) Shinji shared valuable lessons learned from her undergraduate co-op experience with statistical analysis at Sun Microsystems, software engineering at Barclays Capital, and growth marketing at Facebook. * (08:52) Shinji shared lessons learned from being a Management Consultant at Deloitte. * (14:01) Shinji revisited her decision to quit the job at Deloitte and create a social puzzle game called Shufflepix. * (17:42) Shinji went over her time working as a Product Manager at the mobile ad exchange network YieldMo. * (22:25) Shinji discussed the problem of stream processing at YieldMo, which sparked the creation of Concord. * (26:17) Shinji unpacked the pain points with existing stream processing frameworks and the competitive advantage of using Concord. * (33:19) Shinji recalled her time at Akamai — initially as a data engineer in the Platform Engineering unit and later as a product manager for the IoT Edge Connect platform. * (37:26) Shinji explained why sharing context knowledge around data remains a largely unsolved problem. * (42:07) Shinji unpacked the three capabilities of an ideal data discovery platform: (1) exposing up-to-date operational metadata along with the documentation, (2) tracking the provenance of data back to its source, and (3) guiding data usage. * (46:59) Shinji unpacked the benefits of plugging BI tools into data discovery platforms and collecting metadata, which facilitates better visibility and understanding. * (52:36) Shinji discussed the role of a data discovery platform within the modern data stack. * (53:59) Shinji shared the hurdles that her team has to go through while finding early adopters of Select Star. * (55:48) Shinji shared valuable hiring lessons learned at Select Star. * (01:00:00) Shinji shared fundraising advice for founders currently seeking the right investors for their startups. * (01:04:41) Closing segment.

Shinji’s Contact Info* LinkedIn * Twitter * Medium

Select Star’s Resources* Website * Blog * LinkedIn | Twitter | Medium

Mentioned ContentArticles* “The Next Evolution of Data Catalogs: Data Discovery Platforms” (Feb 2021) * “Data Discovery for Business Intelligence” (May 2021)

People* Martin Kleppmann (Author of Designing Data-Intensive Applications) * Emily Riederer (Senior Analytics Manager at Capital One) * Anya Prosvetova (Tableau DataDev Ambassador)

Book* “Managing Oneself” (by Peter Drucker)

NotesMy conversation with Shinji was recorded back in July 2021. Since then, many things have happened at Select Star:

  • General Availability launch on Product Hunt: https://www.producthunt.com/posts/selectstar
  • Snowflake partnership on data governance: https://blog.selectstar.com/selectstar-and-snowflake-partner-to-take-data-governance-to-a-new-level-a9d274e1d4c6
  • Case studies with Pitney Bowes and Handshake

About the showDatacast features long-form, in-depth conversations with practitioners and researchers in the data community to walk through their professional journeys and unpack the lessons learned along the way. I invite guests coming from a wide range of career paths — from scientists and analysts to founders and investors — to analyze the case for using data in the real world and extract their mental models (“the WHY and the HOW”) behind their pursuits. Hopefully, these conversations can serve as valuable tools for early-stage data professionals as they navigate their own careers in the exciting data universe.

Datacast is produced and edited by James Le. Get in touch with feedback or guest suggestions by emailing khanhle.1013@gmail.com.

Subscribe by searching for Datacast wherever you get podcasts or click one of the links below:

  • Listen on Spotify
  • Listen on Apple Podcasts
  • Listen on Google Podcasts

If you’re new, see the podcast homepage for the most recent episodes to listen to, or browse the full guest list.

View Details

Show Notes* (02:27) Taimur reflected on his education studying Computer Science at UT Austin in the early 2000s. * (06:26) Taimur recalled his first job working as a quality assurance engineer at Vignette. * (08:47) Taimur went through his time with Oracle / Siebel, where he transitioned from a purely technical engineering-focused role to more customer-facing functions. * (13:44) Taimur reflected on his proudest accomplishments at Oracle. * (18:23) Taimur recalled dropping out of studying at the Stanford Center of Professional Development and moving to Seattle to work for Amazon Web Services. * (20:35) Taimur provided insights on attributes of exceptional sales talent, given his time as an enterprise sales manager in his first two years at AWS. * (23:55) Taimur shared anecdotes of successful product launches and their market expansion strategies while leading business development for AWS's database and compute services. * (28:33) Taimur discussed instituting the culture of customer obsession and operational excellence into his teams - while leading the incubation, market development, and technical go-to-market strategy and execution for the AWS Platform across infrastructure, data, developer services, and emerging technologies. * (33:14) Taimur talked about his decision to join Microsoft to lead the Worldwide Customer Success function for their Azure Data Platform, Analytics, and AI business. * (36:24) Taimur unpacked his talk called “Enabling Customer Success through Evolutionary Architectures.” * (43:07) Taimur compared the BizOps culture between Azure and AWS. * (46:29) Taimur discussed his decision to onboard Redis as their Chief Business Development Officer. * (50:07) Taimur went over the data challenges with operational ML, the emerging data architecture of feature stores, and the powerful capabilities of Redis as a solution. * (55:58) Taimur unpacked key ideas in his talk "First Principles in Building A Real-Time AI Platform." * (01:01:52) Taimur hinted at Redis' product vision of "caching for ML data." * (01:05:21) Taimur gave advice for a smart, driven operator who wants to explore angel investing. * (01:10:17) Taimur described the evolution of tech leadership, strategic business development, and customer success strategies in the past two decades. * (01:15:29) Taimur shared three books that have greatly influenced his life. * (01:16:48) Closing segment.

Taimur's Contact Info* LinkedIn * Twitter * Redis Profile

Redis' Resources* Website * Redis Open Source | Redis Enterprise Software | Redis Enterprise Cloud * Redis AI * LinkedIn | Twitter | Facebook | YouTube * "Redis Labs Becomes Redis" (Aug 2021)

Mentioned ContentPeople* Andy Jassy (CEO of Amazon) * Melanie Perkins (CEO of Canva) * Jeff Lawson (CEO of Twilio)

Books* "Man's Search For Meaning" (by Viktor Frankl) * "Thinking In Systems" (by Donella Meadows) * "A Treasury of Rumi" (by Muhammad Isa Waley and Rumi) * "Start With Why" (by Simon Sinek)

Talks* "First Principles in Building A Real-Time AI Platform" (March 2021) * "Redis as an Online Feature Store" (April 2021) * "Redis as an online feature store, Redis Labs" (May 2021)

About the showDatacast features long-form, in-depth conversations with practitioners and researchers in the data community to walk through their professional journeys and unpack the lessons learned along the way. I invite guests coming from a wide range of career paths — from scientists and analysts to founders and investors — to analyze the case for using data in the real world and extract their mental models (“the WHY and the HOW”) behind their pursuits. Hopefully, these conversations can serve as valuable tools for early-stage data professionals as they navigate their own careers in the exciting data universe.

Datacast is produced and edited by James Le. Get in touch with feedback or guest suggestions by emailing khanhle.1013@gmail.com.

Subscribe by searching for Datacast wherever you get podcasts or click one of the links below:

  • Listen on Spotify
  • Listen on Apple Podcasts
  • Listen on Google Podcasts

If you’re new, see the podcast homepage for the most recent episodes to listen to, or browse the full guest list.

View Details

Show Notes* (01:43) Leigh-Marie shared her formative experiences of her childhood — growing up in Alabama, solving math problems competitively, and going to Phillips Exeter Academy. * (04:21) Leigh-Marie discussed her undergraduate experience at MIT studying Math with Computer Science. * (06:41) Leigh-Marie went through her internship experience at Jane Street and Blend. * (10:07) Leigh-Marie recalled lessons learned from interning at Google — as an ML engineer for the Research and Machine Intelligence Team and an Associate Product Manager for the Chrome Web Platform team. * (13:39) Leigh-Marie talked about her decision to join the early founding team of Scale API (now known as Scale AI) after finishing MIT. * (17:30) Leigh-Marie explained why labeled data is the key bottleneck to the growth of the ML industry. * (20:02) Leigh-Marie discussed the engineering and product challenges of dealing with 3D sensor data. * (22:33) Leigh-Marie unpacked her experience building Scale’s Sensor Fusion Annotation product from scratch, from gathering customer interests to building the initial MVP. * (26:45) Leigh-Marie talked about learning curves during Scale’s scaling phase, as the product had more advanced features and the customer list grew. * (32:21) Leigh-Marie dived into Scale’s credo emphasizing a relentless speed of execution. * (35:00) Leigh-Marie shared valuable hiring lessons at Scale’s early days (Read Alex’s blog post about Scale’s hiring philosophy). * (38:05) Leigh-Marie went over the importance of developing uncompressed understandings of how everything works together as Scale grows. * (41:39) Leigh-Marie shared her advice for folks who want to get into angel investing. * (44:02) Leigh-Marie shared her motivation behind joining Founders Fund (Read Founders Fund’s investment manifesto). * (46:56) Leigh-Marie went over how she has been proving value upfront and forming investment theses as a new investor. * (49:10) Leigh-Marie shared advice she has been giving to companies regarding their product-market fit and go-to-market fit strategies. * (50:38) Leigh-Marie reflected on her transitions from software engineering to product management to venture capital. * (52:31) Leigh-Marie shared the lesson learned from playing poker that benefits her careers in startup and venture. * (54:19) Closing segment.

Leigh-Marie’s Contact Info* Substack * Twitter * LinkedIn * GitHub * Quora * Founders Fund

People* Peter Thiel * Ali Partovi * Trae Stephens

Books* “Angels” (by Jason Calanacis) * “Zero To One” (by Blake Masters and Peter Thiel) * “7 Powers: The Foundations of Business Strategy” (by Hamilton Helmer)

Blog Posts* “The One Data Platform To Rule Them All” (July 2021) * “Startup Opportunities in Machine Learning Infrastructure” (Sep 2021)

About the showDatacast features long-form, in-depth conversations with practitioners and researchers in the data community to walk through their professional journeys and unpack the lessons learned along the way. I invite guests coming from a wide range of career paths — from scientists and analysts to founders and investors — to analyze the case for using data in the real world and extract their mental models (“the WHY and the HOW”) behind their pursuits. Hopefully, these conversations can serve as valuable tools for early-stage data professionals as they navigate their own careers in the exciting data universe.

Datacast is produced and edited by James Le. Get in touch with feedback or guest suggestions by emailing khanhle.1013@gmail.com.

Subscribe by searching for Datacast wherever you get podcasts or click one of the links below:

  • Listen on Spotify
  • Listen on Apple Podcasts
  • Listen on Google Podcasts

If you’re new, see the podcast homepage for the most recent episodes to listen to, or browse the full guest list.

View Details

Show Notes* (01:47) Chen-Ping shared his upbringing growing up in Taiwan and going to boarding school in the US at the age of 14. * (04:42) Chen-Ping got his Bachelor’s and Master’s degrees in Computer Science from RIT back in the early-to-mid 2000, in which he did academic research in computational neuroscience. * (08:18) Chen-Ping walked through his MS thesis at RIT, designing and implementing a computational model of neurons from the visual cortex’s medial superior temporal area. * (10:18) Chen-Ping talked about the academic culture shock of pursuing his Master’s degree in Computer Science and Engineering at Penn State. * (13:47) Chen-Ping walked through his MS thesis at Penn State, proposing a statistical asymmetry-based automatic brain tumor detection from 3D MR images. * (18:35) Chen-Ping discussed the thread of his research as a Ph.D. student at Stony Brook, where he worked at the Computer Vision Lab and the Eye Cog Lab. * (23:19) Chen-Ping unpacked his Ph.D. dissertation at Stony Brook called computational models of visual features: from proto-objects to object categories. * (28:54) Chen-Ping went through his internship experience at Riverbed Technology and Shutterstock. * (30:20) Chen-Ping dissected the development of a neuro-inspired deep convolutional neural network called Map-CNN for modeling human early visual information processing during his time as a Postdoc at Harvard’s Cognitive and Neural Organization Lab. * (32:14) Chen-Ping mentioned research areas at the intersection of computer vision and cognitive vision that he is excited about. * (33:33) Chen-Ping shared the story behind the founding of Phiar with James Briscoe, an ex-classmate from RIT, and Ivy Lee, an ex-colleague from Shutterstock. * (36:33) Chen-Ping discussed technical challenges with developing an ultra-lightweight Spatial AI engine that allows any vehicle to perceive its surroundings using a camera that can run in real-time at the edge on a commodity automotive computing platform. * (39:36) Chen-Ping unpacked the key features of a complete Visual Mobility platform, including automobile integration, AR navigation, digitized environment, smart parking, 3rd-party integration, and reality-as-a-service. * (41:16) Chen-Ping shared details around Phiar’s ultra-efficient monocular depth estimation AI that runs efficiently on a mobile phone and achieves SOTA accuracies on the benchmark KITTI dataset. * (43:16) Chen-Ping revisited his experience going through the Y-Combinator incubator in the summer of 2018. * (44:27) Chen-Ping shared high-level fundraising advice for first-time founders. * (46:30) Chen-Ping talked about strategies he found useful to identify the right client partnerships for Phiar. * (48:10) Chen-Ping shared valuable hiring lessons learned at Phiar. * (51:37) Chen-Ping reflected on the difference between being a researcher and a founder. * (53:43) Closing segment.

Chen-Ping’s Contact Info* LinkedIn * Twitter * Google Scholar

Phiar’s Resources* Website * LinkedIn | Twitter | Facebook | YouTube * “Phiar Secures $12M Series A and Names Google Head of Android Automotive Platforms as CEO” (Sep 2021)

Mentioned ContentPeople* Fei-Fei Li * Yann LeCun * Yoshua Bengio

Books and Papers* “Zero To One” (by Blake Masters and Peter Thiel) * “Modeling Clutter Perception using Parametric Proto-object Partitioning” (NIPS 2013) * “Modeling visual clutter perception using proto-object segmentation” (June 2014) * “Searching for Category-Consistent Features: A Computational Approach to Understanding Visual Category Representation” (May 2016) * “Generating the features for category representation using a deep convolutional neural network” (Sep 2016) * “Map-CNN: A Convolutional Neural Network with Map-like Organizations” (Aug 2017) * “Mid-level visual features underlie the high-level categorical organization of the ventral stream” (Sep 2018)

About the showDatacast features long-form, in-depth conversations with practitioners and researchers in the data community to walk through their professional journeys and unpack the lessons learned along the way. I invite guests coming from a wide range of career paths — from scientists and analysts to founders and investors — to analyze the case for using data in the real world and extract their mental models (“the WHY and the HOW”) behind their pursuits. Hopefully, these conversations can serve as valuable tools for early-stage data professionals as they navigate their own careers in the exciting data universe.

Datacast is produced and edited by James Le. Get in touch with feedback or guest suggestions by emailing khanhle.1013@gmail.com.

Subscribe by searching for Datacast wherever you get podcasts or click one of the links below:

  • Listen on Spotify
  • Listen on Apple Podcasts
  • Listen on Google Podcasts

If you’re new, see the podcast homepage for the most recent episodes to listen to, or browse the full guest list.

View Details

Timestamps* (02:00) Aarti shared her upbringing growing up in India and going to New York for undergraduate. * (04:47) Aarti recalled her academic experience getting dual degrees in Computer Science and Computer Engineering at New York University. * (07:17) Aarti shared details about her involvement with the ACM chapter and the Women in Computing club at NYU. * (10:46) Aarti shared valuable lessons from her research internships. * (14:16) Aarti discussed her decision to pursue an MS degree in Computer Science at Stanford University. * (20:27) Aarti reflected on her learnings being the Head Teaching Assistant for CS 230, one of Stanford’s most popular Deep Learning courses. * (23:59) Aarti shared her thoughts on ML applications in both clinical and administrative healthcare settings. * (26:47) Aarti unpacked the motivation and empirical work behind CheXNet, an algorithm that can detect pneumonia from chest X-rays at a level exceeding practicing radiologists. * (29:39) Aarti went over the implications of MURA, a large dataset of musculoskeletal radiographs containing over 40,000 images from close to 15,000 studies, for ML applications in radiology. * (32:50) Aarti went over her experience working briefly as an ML engineer at Andrew Ng’s startup Landing AI and applying ML to visual inspection tasks in manufacturing. * (36:56) Aarti talked about her participation in external entrepreneurial initiatives such as Threshold Venture Fellowship and Greylock X Fellowship. * (43:41) Aarti reminisced her time in a hybrid ML engineer/product manager/VC associate role at AI Fund, which works intensively with entrepreneurs during their startups’ most critical and risky phase from 0 to 1. * (48:43) Aarti shared advice that AI fund companies tended to receive regarding product-market fit and go-to-market fit strategy. * (54:04) Aarti walked through her decision to onboard Snorkel AI, the startup behind the popular Snorkel open-source project capable of quickly generating training data with weak supervision. * (56:36) Aarti reflected on the difference between being an ML researcher and an ML engineer. * (01:00:18) Closing segment.

Aarti’s Contact Info* LinkedIn * Twitter * Google Scholar

People* Andrew Ng * John Langford * David Sontag

Books and Papers* “The Art of Doing Science & Engineering” (by Richard Hamming) * “Deep Medicine: How AI Can Make Healthcare Human Again” (by Eric Topol) * “CheXNet: Radiologist-Level Pneumonia Detection on Chest X-Rays with Deep Learning” (Dec 2017) * “MURA: Large Dataset for Abnormality Detection in Musculoskeletal Radiographs” (May 2018)

About the showDatacast features long-form, in-depth conversations with practitioners and researchers in the data community to walk through their professional journeys and unpack the lessons learned along the way. I invite guests coming from a wide range of career paths — from scientists and analysts to founders and investors — to analyze the case for using data in the real world and extract their mental models (“the WHY and the HOW”) behind their pursuits. Hopefully, these conversations can serve as valuable tools for early-stage data professionals as they navigate their own careers in the exciting data universe.

Datacast is produced and edited by James Le. Get in touch with feedback or guest suggestions by emailing khanhle.1013@gmail.com.

Subscribe by searching for Datacast wherever you get podcasts or click one of the links below:

  • Listen on Spotify
  • Listen on Apple Podcasts
  • Listen on Google Podcasts

If you’re new, see the podcast homepage for the most recent episodes to listen to, or browse the full guest list.

View Details

Timestamps* (02:14) Alberto briefly shared his upbringing and education at the Bayes Business School in London. * (04:01) Alberto shared key learnings from his first entrepreneurial stint at 19 by developing a 3D printing product for ed-tech. * (07:48) Alberto described his overall experience participating in Singularity University’s Graduate Studies Program at the NASA Ames Research Park under a Google-funded scholarship in 2015. * (12:52) Alberto helped develop the Aipoly product to aid the blind and visually impaired. * (17:38) Alberto showed his enthusiasm for federated learning applications within mobile devices. * (19:53) Alberto talked about the dichotomy between capitalism and social good in entrepreneurship. * (22:29) Alberto shared the backstory behind the founding of V7 Labs. * (26:40) Alberto discussed the comparison between biological and artificial neural networks. * (28:02) Alberto emphasized the importance of having a good co-founder. * (30:27) Alberto dissected the notable features developed within V7’s Annotation capability. * (33:37) Alberto went over things to look for in a video labeling tool, citing his blog post. * (37:21) Alberto unpacked key principles behind V7’s robust Dataset Management tool. * (40:53) Alberto walked through the powerful capabilities of V7 Neurons that power its Model Automation tool. * (43:33) Alberto shared fundraising advice for founders seeking the right investors for their startups. * (46:07) Alberto shared valuable hiring and culture-setting lessons learned at V7. * (50:12) Alberto emphasized the importance of not losing sight of the ‘ideal customer’ for young founders in the AI space. * (53:01) Alberto shared the hurdles his team has to go through while finding new customers in new industries. * (55:10) Alberto walked through labeling challenges dealing with medical imaging datasets. * (57:35) Alberto discussed outreach initiatives that helped drive V7’s organic growth. * (59:49) Alberto mentioned the importance of collaboration between companies within the MLOps ecosystem. * (01:02:01) Alberto touched on the scientific hunger of Europe regarding the adoption of AI technologies. * (01:03:49) Alberto briefly mentioned what public recognition means to him in the pursuit of democratizing AI for the world. * (01:06:07) Closing segment.

Alberto’s Contact Info* Website * LinkedIn * Twitter * Medium

V7’s Resources* Website * Software 2.0 Blog * Academy Tutorials * Documentation * LinkedIn | Twitter

Mentioned ContentArticles* “7 Things We Looked for in a Video Labeling Tool” (Aug 2020) * “The Biggest Mistake I’ve Ever Made: Losing Sight of the Ideal Customer” (March 2021)

Talks* “An AI Narrator for the Blind” (TEDx Geneva 2016) * “If The Blind Could See” (TEDx Melbourne 2018)

People* Geoff Hinton (for rethinking the ML field fundamentally) * Chelsea Finn (for her work on meta-learning) * Jeff Clune (for making agents that work at scale in the real world)

Book* “Start With Why” (by Simon Sinek)

NotesV7 is hiring across all departments. Take a look at their careers page for the openings!

About the showDatacast features long-form, in-depth conversations with practitioners and researchers in the data community to walk through their professional journeys and unpack the lessons learned along the way. I invite guests coming from a wide range of career paths — from scientists and analysts to founders and investors — to analyze the case for using data in the real world and extract their mental models (“the WHY and the HOW”) behind their pursuits. Hopefully, these conversations can serve as valuable tools for early-stage data professionals as they navigate their own careers in the exciting data universe.

Datacast is produced and edited by James Le. Get in touch with feedback or guest suggestions by emailing khanhle.1013@gmail.com.

Subscribe by searching for Datacast wherever you get podcasts or click one of the links below:

  • Listen on Spotify
  • Listen on Apple Podcasts
  • Listen on Google Podcasts

If you’re new, see the podcast homepage for the most recent episodes to listen to, or browse the full guest list.

View Details

Timestamps* (01:55) Jessica shared the formative experiences of her upbringing — being born in a triplet with two other sisters and growing up in an immigrant family from Russia. * (05:45) Jessica shared her experience being part of UC Berkeley’s first cohort of Data Science majors. * (09:56) Jessica talked about her campus involvements with student-run organizations such as the Mobile Developers of Berkeley and the Data Science Society at Berkeley. * (13:02) Jessica walked through her participation in initiatives like researching the CITRIS and the Banatao Institute, sitting on the leadership board of the TAMID Group, and being an Accel Scholar. * (15:31) Jessica shared valuable lessons from her summer internships. * (19:08) Jessica discussed her decision to join Ironclad, a Series D digital contracting startup building software to take legal teams to the next level. * (22:39) Jessica provided a brief explanation of digital contracting for the uninitiated. * (24:59) Jessica talked about challenges that in-house legal teams typically face and how Ironclad helps address them. * (27:04) Jessica gave a tour of Ironclad’s Contract Lifecycle Management software offerings. * (30:00) Jessica walked through her journey of building the analytics function from scratch and providing data insights to inform business decisions cross-functionally. * (33:40) Jessica shared tidbits about her time management and goal-setting systems. * (34:45) Jessica walked through the end-to-end data analysis process for Ironclad’s first legal analytics benchmark report analyzing economic trends caused by COVID-19. * (38:38) Jessica discussed the learning curves as she took on bigger analytical responsibilities at Ironclad. * (43:05) Jessica unpacked her 3-level framework for building a data analytics culture from the ground up. * (48:07) Jessica shared concrete advice on positively influencing a company’s culture to be data-driven. * (50:27) Jessica unfolded the drive behind creating the Data Angels Community, a Slack group connecting women interested in data to resources for support, education, and opportunities. * (52:25) Jessica revealed her community playbook to engage the members of Data Angels. * (57:01) Jessica shared a bit of her guilty pleasure in using data for beauty and fashion. * (01:00:44) Closing segment.

Jessica’s Contact Info* LinkedIn * Twitter * Data Angels

Mentioned ContentResources* "How to use contract data during COVID-19" (Ironclad Report) * "Building data analytics culture from the ground up" (Women In Product 2020 Talk) * "Building a data-centered culture at Ironclad" (Ironclad Article)

People* Emily Robinson and Jacqueline Nolis (Co-Authors and Co-Hosts of “How To Build A Career in Data Science” the book and the podcast) * Cassie Kozyrkov (Chief Decision Scientist at Google) * Shreya Shankar (Ph.D. Student at UC Berkeley and Entrepreneur-In-Residence at Amplify Partners) (Check out my interview with Shreya as well!)

Book* “Everybody Lies: Big Data, New Data, and What The Internet Can Tell Us About Who We Really Are” (by Seth Stephens-Davidowitz)

NotesMy conversation with Jessica was recorded back in May 2021. Jessica is now a Senior Data Analyst and Ironclad's Data Analytics team has grown to 4 so she is no longer a 1-woman show! Also, the Data Angels Slack community has over 500 members in it now!

About the showDatacast features long-form, in-depth conversations with practitioners and researchers in the data community to walk through their professional journeys and unpack the lessons learned along the way. I invite guests coming from a wide range of career paths — from scientists and analysts to founders and investors — to analyze the case for using data in the real world and extract their mental models (“the WHY and the HOW”) behind their pursuits. Hopefully, these conversations can serve as valuable tools for early-stage data professionals as they navigate their own careers in the exciting data universe.

Datacast is produced and edited by James Le. Get in touch with feedback or guest suggestions by emailing khanhle.1013@gmail.com.

Subscribe by searching for Datacast wherever you get podcasts or click one of the links below:

  • Listen on Spotify
  • Listen on Apple Podcasts
  • Listen on Google Podcasts

If you’re new, see the podcast homepage for the most recent episodes to listen to, or browse the full guest list.

View Details

Timestamps* (01:40) Julia shared the differences growing up in New York and moving to San Francisco. * (03:05) Julia discussed her overall undergraduate experience at Stanford — getting dual degrees in Computer Science and Management Science & Engineering_._ * (05:40) Julia went over her time as an Investment Banker at Qatalyst Partners — notably working on Microsoft’s acquisition of LinkedIn. * (09:11) Julia talked about her career transition to venture capital — working as an associate investor at New Enterprise Associates. * (10:46) Julia emphasized the importance of getting up-to-speed and forming an investment thesis as a new investor. * (15:05) Julia discussed her Series A investment in Metabase, an open-source business intelligence software project. * (18:36) Julia unpacked her investment(s) in Sentry, an application monitoring platform that helps developers monitor apps in real-time to catch bugs early. * (20:14) Julia explained her investment in the Series B round for Anyscale, an end-to-end computing platform that makes building and managing a scaled application across clouds as easy as developing an app on a single computer. * (23:03) Julia contextualized her investments in the seed round for Datafold, a data observability platform that equips analytics engineers with the tools to address data quality issues. * (24:24) Julia shared typical hiring and go-to-market decisions that companies need to make (depending upon their growth stages and product strategies). * (27:05) Julia mentioned her Metabase application to help investors pick winning open-source startups. * (29:05) Julia rationalized her switch to becoming a product manager at dbt Labs. * (30:34) Julia peeked into the roadmap of dbt Cloud, a hosted service that helps data analysts and engineers productionize dbt deployments. * (33:34) Julia went over an under-invested area and the role of interoperability within the broader data tooling ecosystem. * (37:56) Julia reflected on the difference between being a venture investor and a product manager. * (41:05) Closing segment.

Julia’s Contact Info* LinkedIn * Twitter

dbt’s Resources* Slack Community * Coalesce 2021 Replays * dbt Learn * GitHub * Events and Meetups

Mentioned ContentPeople* Tristan Handy (Founder and CEO of dbt Labs) * Ali Ghodsi (Co-Creator of Apache Spark, Co-Founder and CEO of Databricks) * Dan Levine (General Partner at Accel Partners)

Book* “Working Backwards: Insights, Stories, and Secrets from Inside Amazon” (by Bill Carr and Colin Bryar)

NotesMy conversation with Julia was recorded back in May 2021. Since the podcast was recorded, a lot has happened at dbt Labs! I’d recommend:

  • Reading Julia’s recent blog posts on adopting CI/CD and introducing Environment Variables in dbt Cloud.
  • Watching the talk replays from Coalesce, dbt’s 2nd annual analytics engineering conference
  • Listening to Season 1 of the Analytics Engineering Podcast, where Julia co-hosts with Tristan Handy to go deep into the hopes, dreams, motivations, and failures of leading data and analytics practitioners.

About the showDatacast features long-form, in-depth conversations with practitioners and researchers in the data community to walk through their professional journeys and unpack the lessons learned along the way. I invite guests coming from a wide range of career paths — from scientists and analysts to founders and investors — to analyze the case for using data in the real world and extract their mental models (“the WHY and the HOW”) behind their pursuits. Hopefully, these conversations can serve as valuable tools for early-stage data professionals as they navigate their own careers in the exciting data universe.

Datacast is produced and edited by James Le. Get in touch with feedback or guest suggestions by emailing khanhle.1013@gmail.com.

Subscribe by searching for Datacast wherever you get podcasts or click one of the links below:

  • Listen on Spotify
  • Listen on Apple Podcasts
  • Listen on Google Podcasts

If you’re new, see the podcast homepage for the most recent episodes to listen to, or browse the full guest list.

View Details

Timestamps* (1:33) Einat described her experience getting Bachelor’s, Master’s, and Ph.D. degrees in Mathematics from Tel Aviv University in the 90s and early 2000s. * (4:01) Einat went over her Ph.D. thesis on approximation algorithms for clustering problems. * (6:17) Einat discussed working as an algorithm developer for Compugen while being a Ph.D. student. * (8:43) Einat went over projects she contributed to as a senior algorithm developer at Flash Networks back in 2005. * (11:50) Einat mentioned achievements and lessons learned from her time as the VP of R&D at Correlix. * (17:51) Einat recalled lessons from hiring engineering talent at Correlix. * (19:24) Einat unpacked the engineering challenges of building SimilarWeb, a platform that gives a true 360-degree view of all digital activity across customers, prospects, partners, and competition. * (24:29) Einat discussed the responsibilities of her role as the CTO of SimilarWeb. * (27:40) Einat shared the founding story of Treeverse, whose mission is to simplify the lives of data engineers, data scientists, and data analysts who are transforming the world with data. * (29:52) Einat explained the pain points of working with the data lake architecture and the vision that lakeFS is built upon. * (34:31) Einat emphasized the importance of asking good questions to extract insights about customers’ pain points. * (37:57) Einat explained why data versioning-as-an-Infrastructure matters. * (42:28) Einat shared the challenges of incorporating data mesh to develop a data-intensive application. * (46:33) Einat provided her take on how to ensure data quality in a data lake environment. * (51:02) Einat discussed roadmap prioritization for an open-source project. * (52:08) Einat went over the opportunities with the metadata store, data quality, compute, and data discovery components within the data engineering ecosystem. * (55:03) Einat captured the three trends on how the data engineering landscape might look in the near future. * (01:00:59) Einat emphasized the role of open-source development in the data tooling ecosystem. * (01:04:14) Einat fleshed out the recommended pricing strategy for open-source developers. * (01:06:09) Einat revisited how lakeFS got started thanks to the Go community and evolved. * (01:08:01) Einat shared valuable hiring lessons learned at Treeverse. * (01:10:05) Einat described the state of the data community in Israel. * (01:11:49) Closing segment.

Einat’s Contact Info* LinkedIn * Twitter * Email

Mentioned ContentlakeFS* Website * GitHub * @lakeFS * Treeverse * Slack

Blog Posts* Why We Built lakeFS: Atomic and Versioned Data Lake Operations (Aug 2020) * Data Versioning — Does It Mean What You Think It Means? (Aug 2020) * How To Manage Your Data The Way You Manage Your Code (Oct 2020) * Data Mesh Applied: How to Move Beyond The Data Lake with lakeFS (Dec 2020) * Why Data Versioning as an Infrastructure Matters? (Dec 2020) * Ensuring Data Quality in a Data Lake Environment (Jan 2021) * The State of Data Engineering in 2021 (May 2021)

People* Ali Ghodsi (Co-Creator of Apache Spark, Co-Founder and CEO of Databricks) * Shay Banon (Co-Founder and CEO of Elastic) * Gwen Shapira (Engineering Leader at Confluent)

Book* “Designing For Data-Intensive Applications” (by Martin Kleppmann)

NotesMy conversation with Einat was recorded back in April 2021. Since the podcast was recorded, a lot has happened at Treeverse! I’d recommend:

  • Looking at their Series A announcement back in July.
  • Reading Einat’s recent articles on measuring data engineering teams, mapping data versioning projects, and finding a role model for lakeFS.
  • Reviewing lakeFS’s ongoing roadmap.
  • Connecting with the lakeFS community by attending their upcoming events.

About the showDatacast features long-form, in-depth conversations with practitioners and researchers in the data community to walk through their professional journeys and unpack the lessons learned along the way. I invite guests coming from a wide range of career paths — from scientists and analysts to founders and investors — to analyze the case for using data in the real world and extract their mental models (“the WHY and the HOW”) behind their pursuits. Hopefully, these conversations can serve as valuable tools for early-stage data professionals as they navigate their own careers in the exciting data universe.

Datacast is produced and edited by James Le. Get in touch with feedback or guest suggestions by emailing khanhle.1013@gmail.com.

Subscribe by searching for Datacast wherever you get podcasts or click one of the links below:

  • Listen on Spotify
  • Listen on Apple Podcasts
  • Listen on Google Podcasts

If you’re new, see the podcast homepage for the most recent episodes to listen to, or browse the full guest list.

View Details

Timestamps* (02:13) Prukalpa discussed her upbringing in India and studying Engineering at the Nanyang Tech University in Singapore. * (03:52) Prukalpa shared the key learnings from her summer internship as an Investment Banking Analyst at Goldman Sachs. * (05:37) Prukalpa went over the seed idea for SocialCops (Read her Quora answer on the fundraising story). * (11:27) Prukalpa gave a brief overview of the business model at SocialCops. * (12:45) Prukalpa unpacked her talk called “How Big Data Can Influence Decisions That Actually Matter” at TEDxGateway 2017 related to the data-for-good initiatives that SocialCops facilitated. * (15:23) Prukalpa shared her thoughts on the future of the Data-for-Good movement. * (17:49) Prukalpa discussed the challenges that SocialCops’ data teams faced and the founding story behind Atlan. * (21:38) Prukalpa went over the trust-based culture that enabled SocialCops’ 8-member data team to build out India’s National Data Platform. * (27:00) Prukalpa dissected the six principles of Atlan’s DataOps Culture Code. * (31:37) Prukalpa unpacked the notion of Data Catalog 3.0, which is a key value prop of the Atlan platform. * (36:01) Prukalpa provided the 3-level framework to ensure data quality (detect -> prevent -> cure) and strong practices to maintain high-quality data. * (40:19) Prukalpa revealed the challenges that organizations face when starting their data governance initiatives. * (45:35) Prukalpa talked about the under-invested building blocks of modern data platforms. * (49:24) Prukalpa raised the importance of integration for Atlan to work well with the rest of the modern data stack. * (50:39) Prukalpa recapped the trends that Chief Data Officers needed to watch out for in 2021. * (54:01) Prukalpa gave fundraising advice for founders currently seeking the right investors for their startups. * (58:42) Prukalpa discussed Atlan’s outreach initiatives to engage with the broader data community actively. * (01:01:03) Prukalpa went over Atlan’s hiring philosophy based on the concept of People-as-a-Moat to attract, engage, and grow top talent — as inspired by the McKinsey advantage. * (01:05:22) Prukalpa shared Atlan’s Go-To-Market initiatives in the US this year and emphasized the importance of building an execution machine. * (01:08:53) Prukalpa described the state of the data community in India. * (01:10:25) Prukalpa shared entrepreneurship books that have deeply impacted her startup journey. * (01:12:16) Prukalpa briefly mentioned what public recognition means to her in the pursuit of democratizing data for the world. * (01:14:23) Closing segment.

Prukalpa’s Contact Info* LinkedIn * Twitter

Mentioned ContentAtlan (Twitter | LinkedIn | Facebook | Instagram | YouTube | Documentation)* “Empowering Organizations to Become Masters of Their Data” (Video) * Atlan Labs (Open-Source Projects) * Humans of Data Interviews (Interviews) * The DataOps Culture Code (Document) * Building a Business Case for DataOps (EBook) * The Data Catalog Primer (EBook) * The Ultimate Guide to Evaluating a Data Catalog (EBook)

Blog Posts* Voices In The Head of a Middle-Class Aspiring Startup Founder (July 2013) * SocialCops: What We Actually Do (Oct 2016) * People-as-a-Moat: What Startups Can Learn From McKinsey About Building A Strong Company (Aug 2018) * Going from Great People to Greater Teams: How We Think About Growth at Atlan (August 2018) * Onwards and Upwards: Chapter 2 for SocialCops (July 2019) * What is data quality? (Jan 2021) * Top 5 Data Trends For CDOs to Watch Out For In 2021 (Feb 2021) * Data Catalog 3.0: Modern Metadata for the Modern Data Stack (Feb 2021) * We Failed To Setup a Data Catalog 3x. Here’s Why (March 2021) * The Building Blocks of a Modern Data Platform (March 2021) * Data Governance Has a Serious Branding Problem (Nov 2021)

Books* “The Hard Things About The Hard Things” (by Ben Horowitz) * “Hatching Twitter” (by Nick Bilton) * “The McKinsey Way” (by Ethan Rasiel) * “How Google Works” (by Eric Schmidt and Jonathan Rosenberg) * “The Mom Test” (by Rob Fitzpatrick) * “Disciplined Entrepreneurship” (by Bill Aulet) * “Big Data” (by Mayer-Schnonberger and Cukier)

Talks* Game of Life (TEDxIIMShilong — March 2014) * How Big Data Can Influence Decisions That Actually Matter (TEDx Gateway — April 2017) * Better Villages Through Big Data (TED Talks India — December 2017) * The power of data science to measure unmeasured parameters in Emerging Markets (PyData Dehli — Oct 2019) * The Girl Who Thinks In Numbers: Data Warrior Prukalpa Sankar (Feb 2020)

NotesMy conversation with Prukalpa was recorded back in April 2021. Since the podcast was recorded, a lot has happened at Atlan!

  • They raised a $16M Series A led by Insight Partners, with participation from Sequoia Capital, Waterbridge Ventures, and amazing angels such as the founding teams of Snowflake and Looker.
  • They got mentioned in Gartner’s Inaugural Market Guide for Active Metadata Management.
  • They announced a partnership with Snowflake.

Prukalpa has written more content. I’d recommend checking out:

  • The series on metadata.
  • The list of resources on the modern data stack.
  • The behind-the-scenes look at how Postman’s data team uses Atlan.
  • The new way to think about data strategy using the Data Advantage Matrix.

About the showDatacast features long-form, in-depth conversations with practitioners and researchers in the data community to walk through their professional journeys and unpack the lessons learned along the way. I invite guests coming from a wide range of career paths — from scientists and analysts to founders and investors — to analyze the case for using data in the real world and extract their mental models (“the WHY and the HOW”) behind their pursuits. Hopefully, these conversations can serve as valuable tools for early-stage data professionals as they navigate their own careers in the exciting data universe.

Datacast is produced and edited by James Le. Get in touch with feedback or guest suggestions by emailing khanhle.1013@gmail.com.

Subscribe by searching for Datacast wherever you get podcasts or click one of the links below:

  • Listen on Spotify
  • Listen on Apple Podcasts
  • Listen on Google Podcasts

If you’re new, see the podcast homepage for the most recent episodes to listen to, or browse the full guest list.

View Details

Timestamps* (01:58) Michel went over his education studying at EPITA — School of Engineering and Computer Science in France. * (03:50) Michel mentioned his first US internship at Siemens Corporate Research as an R&D engineer. * (05:48) Michel discussed the unique challenges of building systems to handle financial data through his engineering experience at FactSet Research Systems and Murex. * (07:48) Michel talked about his move to San Francisco to work as a Senior Software Engineer at Rapleaf, focusing on scaling up data integration and data management pipelines. * (10:40) Michel unpacked his work building the modern data stack at LiveRamp. * (16:18) Michel shared valuable leadership and hiring lessons absorbed during his time as Liveramp’s Head of Data Integrations — fostering a strong culture of innovation and expanding the engineering organization significantly. * (19:03) Michel dived into how to interview engineering talent for independence, autonomy, and communication. * (20:56) Michel dissected the engineering architecture of the rideOS ride-hail platform (where he was a founding member and director of engineering). * (26:03) Michel told the founding story of Airbyte, whose mission is to make data integration pipelines a commodity (+ the pivot that happened during Airbyte’s time at Y Combinator). * (32:10) Michel explained the paint points with existing data integration practices and the vision that Airbyte is moving towards. * (35:07) Michel unpacked the analogy of Airbyte’s approach to building a connector manufacturing plant, which is to think in onion layers. * (39:13) Michel shared the challenges that are still hard for an open-source solution to address (Read his list of challenges that open-source and commercial software face to solve the data integration problem). * (40:28) Michel discussed how to prioritize product roadmap while developing an open-source project. * (41:59) Michel discussed pricing strategies for open-source projects (Airbyte’s business models entail both self-hosted and hosted solutions). * (44:17) Michel revealed the hurdles that Airbyte has overcome to find the early committers for their open-source project. * (47:53) Michel shared valuable hiring lessons learned at Airbyte. * (50:16) Michel shared fundraising advice for founders seeking the right investors for their startups. * (52:41) Closing segment.

Michel’s Contact Info* Twitter * LinkedIn * GitHub

Mentioned ContentAirbyte (Docs | Community | GitHub | Twitter | LinkedIn)* Handbook * Recipes * Community Call * Office Hours * Connector Contest

Blog Posts* “The Hard Things About Pivoting” (July 2020) * “How Can We Commoditize Data Integration Pipelines” (Sep 2020) * “How to Build Thousands of Connectors” (Oct 2020) * “The Deck We Used to Raise Our Seed with Accel in 13 Days” (March 2021)

People* Jeremy Litz (Former CTO and Co-Founder of LiveRamp) * Tristan Handy (CEO of dbtLabs and Editor of the Analytics Engineering Newsletter)

Book* “High-Growth Handbook” (by Elad Gil)

NotesMy conversation with Michel was recorded back in April 2021. Since the podcast was recorded, a lot has happened at Airbyte! I'd recommend:

  • Looking over the deck that they used to raise a $26M Series A led by Benchmark.
  • Reading Michel's take on Airbyte's new OSS model and strategy to commoditize all data integration.
  • Diving into Airbyte Cloud, a hosted service that takes all of the features of the open-source version and adds hosting and management, on top of a number of additional support options and enterprise features.
  • Subscribing to Airbyte's newsletter called Weekly Bytes and exploring Airbyte Recipes.

About the showDatacast features long-form, in-depth conversations with practitioners and researchers in the data community to walk through their professional journeys and unpack the lessons learned along the way. I invite guests coming from a wide range of career paths — from scientists and analysts to founders and investors — to analyze the case for using data in the real world and extract their mental models (“the WHY and the HOW”) behind their pursuits. Hopefully, these conversations can serve as valuable tools for early-stage data professionals as they navigate their own careers in the exciting data universe.

Datacast is produced and edited by James Le. Get in touch with feedback or guest suggestions by emailing khanhle.1013@gmail.com.

Subscribe by searching for Datacast wherever you get podcasts or click one of the links below:

  • Listen on Spotify
  • Listen on Apple Podcasts
  • Listen on Google Podcasts

If you’re new, see the podcast homepage for the most recent episodes to listen to, or browse the full guest list.

View Details

Timestamps* (02:03) Cindi briefly shared her early interest in writing and her decision to major in English at the University of Maryland in the mid-80s. * (05:22) Cindi talked about her move to Zurich for a Business Systems Specialist role at Dow Chemical. * (07:35) Cindi recalled the state of Business Intelligence tools and their adoption level in the enterprises during the mid-90s. * (10:53) Cindi went over her decision to pursue an MBA from the Jones Business School at Rice University, in which her MBA Thesis was about how the Internet would reshape the first-generation BI tools. * (16:30) Cindi discussed how she balanced academic study and parenthood during her MBA. * (20:57) Cindi talked about her proudest accomplishments as a manager at Deloitte — building BI and analytics practice in Houston. * (22:48) Cindi went over her time running her independent analyst firm BI Scorecard, which advised clients on BI and analytics tool selections via rigorous evaluation criteria. * (26:14) Cindi brought up her time teaching classes at The Data Warehousing Institute, which educates business leaders on the proper deployment of data warehousing strategies and technologies. * (27:49) Cindi mentioned her move to become the Vice President in data and analytics at Gartner. * (30:33) Cindi walked through the end-to-end process of creating Gartner’s Magic Quadrant for Analytics and BI Platforms and Critical Capabilities. * (33:47) Cindi explained the culture of “Selfless Excellence” at ThoughtSpot — where she currently serves as a Chief Strategy Officer. * (36:11) Cindi explained the concept of “What’s In It For Me” (WIIFM) that helps bring a data-driven culture to organizations. * (39:34) Cindi gave a tour of ThoughtSpot’s core capabilities, ranging from SearchIQ and SpotIQ to ThoughtSpot One and ThoughtSpot Embrace. * (43:04) Cindi broke down her responsibilities as a Chief Data Strategy Officer working with internal and external stakeholders. * (44:40) Cindi emphasized the role of partnerships between startup vendors to empower the future of BI analytics (Read her article A New Era in Analytics and BI”). * (49:58) Cindi recapped takeaways from ThoughtSpot’s ebook that presents 6 Top Trends and Predictions for Data, Analytics, and AI in 2021. * (53:07) Cindi gave advice to companies that want to bring consumerization to enterprise analytics. * (56:35) Cindi gave her two cents on the movement of Data For Good in the progress of analytics and AI in the near future. * (58:58) Cindi recapped insights that she has observed from hosting The Data Chief Podcast (which features interviews with some of the most successful data leaders). * (01:03:29) Cindi gave advice to female data practitioners in the early phase of their careers (Read her article on the challenges that keep women out of tech). * (01:05:30) Closing segment.

Cindi’s Contact* LinkedIn * Twitter * ThoughtSpot * The Data Chief Podcast

Mentioned ContentBlog Posts* “Why I Joined ThoughtSpot” (April 2019) * “A New Era in Analytics and BI” (August 2019) * “Perfect Storm or Transformative Triumvirate: Data for Good, Data for Evil, and AI Ethics” (Nov 2019) * “We Can Put a Man on the Moon, But We Can’t Keep Women in Tech” (Sep 2019) * 6 Top Trends and Predictions for Data, Analytics, and AI in 2021 (2021 E-Book)

Published Books* “Successful Business Intelligence” (Nov 2013) * “SAP BusinessObjects BI 4.0” (Nov 2012)

Data for Good Resources* Datakind * Mastercard Center for Inclusive Growth * Carnegie Mellon’s Data Science for Social Good * Viz for Social Good

Women in Data Resources* Women in Data * Women in Big Data * Women in Analytics

People* Joy Buolamwini (Computer Scientist and Digital Activist at MIT Media Lab, Founder of the Algorithmic Justice League) * Cathy O’Neil (Author of “Weapons of Math Destruction”) * Kate Strachnyi (Founder of DATAcated) * Ralph Kimball (Original Architect of Data Warehousing) * Ajeet Singh and Amit Prakash (Co-Founders of ThoughtSpot)

Recommended Books* “Moneyball” (by Michael Lewis) * “Freakonomics” (by Steven Levitt and Stephen Dubner)

NotesMy conversation with Cindi was recorded back in April 2021. Since the podcast was recorded, a lot has happened at ThoughtSpot:

  • They unveiled their new vision for the Modern Analytics Cloud — a simple, actionable, and open approach to cloud analytics that’s redefining how companies deliver value from across the entire modern data stack.
  • They acquired Diyotta & Seekwell. With Diyotta, they’re expanding the number of integrations with other cloud companies, while Seekwell gives customers the ability to operationalize insights by connecting analytics to other systems to trigger action.
  • ThoughtSpot Everywhere launched as the first development platform to build interactive data apps with search and AI-driven analytics.
  • Cloud growth. They announced major growth in their SaaS and cloud offerings, including their first 100 SaaS customers, 250% ARR growth from cloud products, and planned headcount growth.

About the showDatacast features long-form, in-depth conversations with practitioners and researchers in the data community to walk through their professional journeys and unpack the lessons learned along the way. I invite guests coming from a wide range of career paths — from scientists and analysts to founders and investors — to analyze the case for using data in the real world and extract their mental models (“the WHY and the HOW”) behind their pursuits. Hopefully, these conversations can serve as valuable tools for early-stage data professionals as they navigate their own careers in the exciting data universe.

Datacast is produced and edited by James Le. Get in touch with feedback or guest suggestions by emailing khanhle.1013@gmail.com.

Subscribe by searching for Datacast wherever you get podcasts or click one of the links below:

  • Listen on Spotify
  • Listen on Apple Podcasts
  • Listen on Google Podcasts

If you’re new, see the podcast homepage for the most recent episodes to listen to, or browse the full guest list.

View Details

Timestamps* (01:30) Taivo shared briefly about his experience going through the Estonian K-12 system, as argued in his blog post written in Estonian. * (05:34) Taivo described his undergraduate experience studying Computer Science at the University of Tartu and exposing to Machine Learning. * (08:15) Taivo discussed his time interning at Skype and TransferWise. * (10:01) Taivo went over his Master's Degree in Computer Science at ETH Zurich, where he worked on a thesis called "Uncertainty-based active imitation learning" at the Learning and Adaptive Systems Group. * (17:17) Taivo talked about his time working at Starship Technologies as a Perception Engineer. * (21:26) Taivo unpacked the Data Specification Manifesto that entails 3 principles for iteratively solving complex problems. * (27:21) Taivo unpacked "The Two Loops Of Building Algorithmic Products" from his experience at Veriff - an Estonian startup that develops an identity verification platform. * (32:11) Taivo discussed how his team at Veriff developed automation-heavy products. * (36:45) Taivo shared lessons learned as a Product Manager at Veriff: leading the go-to-market strategy, establishing communication between the product and sales division, and building a unique DataOps team that creates good datasets. * (44:31) Taivo described the key characteristics and properties of a tool that can address the whole data annotation workflow (Read his article "Data Loops Are The Bottleneck In Applied AI"). * (49:33) Taivo predicted the evolution of the DataOps discipline for AI teams in the upcoming years (Read his article "Your AI Team Needs DataOps"). * (54:01) Taivo untangled the relationship between sampling and labeling, and their importance in the AI development process (Read his article "Datasets Carve The Terrain of AI"). * (56:36) Taivo talked about the tools that he's most excited about during the transition to Software 2.0. * (01:00:04) Taivo shared his journey thus far as the founder of a stealth startup. * (01:06:21) Taivo revealed insider insights about the #EstonianMafia startup ecosystem. * (01:09:36) Taivo shared the productivity tips that have been most useful to his personal/professional growth. * (01:14:10) Closing segment.

Taivo's Contact* Website * Twitter * LinkedIn * Medium * Google Scholar

Mentioned ContentBlog Posts* Data Specification Manifesto * "Building Automation-Heavy Products" (April 2019) * "Data Loops Are The Bottleneck In Applied AI" (June 2019) * "Your AI Team Needs DataOps" (July 2020) * "Datasets Carve The Terrain of AI" (Nov 2020)

Talks* "The Two Loops Of Building Algorithmic Products" (April 2019) * "How To Build Your AI Startup" (June 2020) * "Datasets: The Source Code of Software 2.0" (Nov 2020)

People* Andrej Karpathy (The Senior Director of AI at Tesla, who coined the term Software 2.0) * Mike Bostock (The Creator of D3.js)

Book* "Surely You're Joking, Mr. Feynman" (by Richard Feynman)

About the showDatacast features long-form, in-depth conversations with practitioners and researchers in the data community to walk through their professional journeys and unpack the lessons learned along the way. I invite guests coming from a wide range of career paths — from scientists and analysts to founders and investors — to analyze the case for using data in the real world and extract their mental models (“the WHY and the HOW”) behind their pursuits. Hopefully, these conversations can serve as valuable tools for early-stage data professionals as they navigate their own careers in the exciting data universe.

Datacast is produced and edited by James Le. Get in touch with feedback or guest suggestions by emailing khanhle.1013@gmail.com.

Subscribe by searching for Datacast wherever you get podcasts or click one of the links below:

  • Listen on Spotify
  • Listen on Apple Podcasts
  • Listen on Google Podcasts

If you’re new, see the podcast homepage for the most recent episodes to listen to, or browse the full guest list.

View Details

Timestamps* (01:42) Gleb shared briefly about his upbringing and studying Economics in university in Russia. * (04:15) Gleb discussed his move to the US to pursue a Master of Information Systems Management at Carnegie Mellon University. * (07:07) Gleb went over his summer internship as a Business Analyst at Autodesk. * (08:40) Gleb shared the details of his project architecting data model/ETL pipelines as a PM at Autodesk. * (11:34) Gleb unpacked the evolution of his career at Lyft — from an individual data analyst to a PM on data tooling and a high-impact project that he worked on. * (16:54) Gleb shared valuable lessons from the experience of leading multiple cross-functional teams of engineers and growing the data organization significantly. * (19:48) Gleb mentioned his time as a Product Manager at Phantom Auto, leading the development of a teleoperation product for autonomous vehicles over cellular networks. * (25:28) Gleb emphasized the critical factors to consider when choosing a working environment: trusted managers/colleagues, maturity of tools/processes, and the function of data teams within the organization. * (29:10) Gleb shared the story behind the founding of Datafold, whose mission is to help companies effectively leverage their data assets while making Data Engineering & Analytics a creative and enjoyable experience. * (33:04) Gleb dissected the pain points with regression testing and the benefits of using Data Diff (Datafold’s first product) for data engineers. * (36:54) Gleb unpacked the data monitoring feature within Datafold’s data observability platform. * (39:45) Gleb discussed how to choose data warehousing solutions for your use cases (and made the distinction between data warehouse and data lake). * (47:03) Gleb gave insights on the need for BI and data observability/quality management tools within the modern analytics stack. * (50:40) Gleb emphasized the importance of tooling integration for Datafold’s roadmap. * (52:07) Gleb has been hosting Data Quality meetups to discuss the under-explored area of data quality. * (54:02) Gleb shared his learnings from going through the YC incubator in summer 2020. * (55:45) Gleb discussed the hurdles he had to jump through to find early customers of Datafold. * (57:47) Gleb emphasized valuable lessons he has learned to attract the right people who are excited about Datafold’s mission. * (59:17) Gleb shared his advice for founders who are in the process of finding the right investors for their companies. * (01:02:11) Closing segment.

Gleb’s Contact Info* LinkedIn * Datafold (Twitter and LinkedIn) * Data Quality Meetups

Mentioned ContentCourse* Harvard’s CS50: Introduction to Computer Science

Blog Posts* Modern Analytics Stack (June 2020) * Choosing Data Warehouse for Analytics (June 2020) * 3 Ways To Be Wrong About Open-Source Data Warehousing Software (June 2020) * Buy Not Build (Aug 2020) * Datafold Raises a $2.1M Seed Round Led by NEA (Nov 2020) * Datafold + dbt: The Perfect Stack for Reliable Data Pipelines (Feb 2021)

People* Maxime Beauchemin (Founder and CEO at Preset, creator of Apache Superset and Apache Airflow) * Tobias Macey (Host of the Data Engineering Podcast)

Books* “How To Measure Anything” (by Douglas Hubbard) * “Lean Analytics” (by Benjamin Yoskovitz and Alistair Croll)

NotesMy conversation with Gleb was recorded back in March 2021. Since the podcast was recorded, a lot has happened at Datafold! I’d recommend:

  • Reading Gleb’s open-source edition of the modern data stack.
  • Listening to Gleb’s appearance on the Data Engineering podcast.
  • Watching the lightning talks and panel discussions from recent Data Quality meetups number 4 and number 5.

About the showDatacast features long-form, in-depth conversations with practitioners and researchers in the data community to walk through their professional journeys and unpack the lessons learned along the way. I invite guests coming from a wide range of career paths — from scientists and analysts to founders and investors — to analyze the case for using data in the real world and extract their mental models (“the WHY and the HOW”) behind their pursuits. Hopefully, these conversations can serve as valuable tools for early-stage data professionals as they navigate their own careers in the exciting data universe.

Datacast is produced and edited by James Le. Get in touch with feedback or guest suggestions by emailing khanhle.1013@gmail.com.

Subscribe by searching for Datacast wherever you get podcasts or click one of the links below:

  • Listen on Spotify
  • Listen on Apple Podcasts
  • Listen on Google Podcasts

If you’re new, see the podcast homepage for the most recent episodes to listen to, or browse the full guest list.

View Details

Timestamps* (01:59) Saishruthi talked about her upbringing, growing up in a rural town in India with no Internet connection and no computers. * (05:50) Saishruthi discussed her undergraduate studying Electrical Engineering at Sri Sairam Engineering College in the early 2010s. * (11:56) Saishruthi mentioned the projects and learnings during her two years working at Tata Consultancy Services as an instrumentation engineer. * (15:57) Saishruthi went over her MS degree in Electrical Engineering at San Jose State University and her journey into data science. * (22:20) Saishruthi shared the initial hurdles she faced transitioning back to school and assimilating to the US culture. * (26:10) Saishruthi touched on her work with San Jose City on disaster management. * (28:20) Saishruthi went over her job search process, eventually landing a data science position at IBM. * (32:16) Saishruthi unpacked lessons learned from public speaking. * (35:20) Saishruthi summarized IBM’s data science and machine learning initiatives. * (37:02) Saishruthi brought up various projects happening at IBM’s Center for Open Source Data and AI Technologies, whose mission is to make open-source AI models dramatically easier to create, deploy, and manage in the enterprise. * (39:40) Saishruthi unpacked the qualities needed to contribute to open-source projects and their role in shaping the development of ML technologies. * (44:50) Saishruthi dissected examples of bias in ML, identified solutions to combat unwanted bias, and presented tools for that (as delivered in her talk titled “Digital Discrimination: Cognitive Bias in Machine Learning”). * (49:12) Saishruthi shared her thoughts on the evolution of research and applications within the Trusted AI landscape. * (54:07) Saishruthi discussed the core value propositions of IBM’s Elyra, a set of AI-centric extensions to JupyterLab that aims to help data practitioners deal with the complexities of the model development lifecycle. * (56:11) Saishruthi briefly shared the challenges with developing Coursera courses on data visualization with Python and with R. * (01:00:47) Saishruthi went over her passion for movements such as Women In Tech and Girls Who Code. * (01:03:27) Saishruthi shared details about her initiative to bring education to rural children. * (01:06:36) Closing segment.

Saishruthi’s Contact Info* Twitter * LinkedIn * Medium * GitHub * Coursera

Mentioned ContentTalks* “Digital Discrimination: Cognitive Bias in Machine Learning” (All Things Open 2020)

Projects* AI Fairness 360 * AI Explainability 360 * Adversarial Robustness Toolkit * Model Asset Exchange * Data Asset Exchange * Elyra

Courses* Data Visualization with Python * Data Visualization with R

About the showDatacast features long-form, in-depth conversations with practitioners and researchers in the data community to walk through their professional journeys and unpack the lessons learned along the way. I invite guests coming from a wide range of career paths — from scientists and analysts to founders and investors — to analyze the case for using data in the real world and extract their mental models (“the WHY and the HOW”) behind their pursuits. Hopefully, these conversations can serve as valuable tools for early-stage data professionals as they navigate their own careers in the exciting data universe.

Datacast is produced and edited by James Le. Get in touch with feedback or guest suggestions by emailing khanhle.1013@gmail.com.

Subscribe by searching for Datacast wherever you get podcasts or click one of the links below:

  • Listen on Spotify
  • Listen on Apple Podcasts
  • Listen on Google Podcasts

If you’re new, see the podcast homepage for the most recent episodes to listen to, or browse the full guest list.

View Details

Timestamps* (01:44) Mohamed described his interest growing up in Egypt and studying Biomedical Engineering at Cairo University in the early 2000s. * (04:22) Mohamed commented on his experience moving to the US to pursue an MBA degree and working in various software engineering roles. * (07:35) Mohamed shared his experience authoring two books: (1) 3D Business Analyst: The Ultimate Hands-On Guide to Mastering Business Analysis and (2) Business Analysis for Beginners: Jump-Start Your BA Career in 4 Weeks. * (13:19) Mohamed discussed his move to the Bay Area for a Senior Engineering Manager role at Twilio, managing and shipping a series of communication API products using Machine and Deep Learning. * (17:39) Mohamed dissected engineering challenges building ML systems at Amazon, alongside key leadership lessons he acquired from managing Amazon’s Kindle mobile and ML engineering teams. * (20:50) Mohamed shared his insider perspective on Amazon’s practices of customer obsession, working backward, and disagree-to-commit. * (24:52) Mohamed mentioned the benefits of teaching a computer vision course for engineers at Amazon’s internal Machine Learning university. * (28:33) Mohamed went over the engineering (hardware + software) and ML challenges associated with building a proprietary threat detection platform at Synapse Tech Corporation (where he was the Head of Engineering). * (32:03) Mohamed shared concrete technical challenges with building an ML system that performs inference on edge devices. * (37:03) Mohamed revealed specific data labeling challenges while building the ML system at Synapse. * (39:57) Mohamed went over his one year as the VP of Engineering for the AI Platform at Rakuten, when he incubated the idea for Kolena. * (42:52) Mohamed explained the current state of ML testing infrastructure and unpacked his current project Kolena, a rigorous ML QA platform that lets users take control of their ML testing. * (49:07) Mohamed has been collaborating with a few institutions, podcasters, and ML influencers to raise awareness of the importance of ML testing and different approaches to tackle the problem. * (50:12) Mohamed touched on his side hustles working with Intel in autonomous drones and teaching content with Udacity’s AI Nanodegree programs. * (53:07) Mohamed dissected his project Mowgly, an educational platform with tracks curated by industry experts to guide users to master specific topics. * (54:58) Mohamed described his experience authoring a book with Manning in 2020 called “Deep Learning For Vision Systems.” * (58:51) Closing segment.

Mohamed’s Contact Info* LinkedIn * Twitter * Website * YouTube * GitHub * Kolena

Mentioned ContentPeople* Andrew Trask (Leader at OpenMined, Senior Research Scientist at DeepMind, Ph.D. Student at the University of Oxford) * Francois Chollet (Senior Software Engineer at Google, Creator of Keras) * Lex Fridman (Host of the popular Lex Fridman Podcast, AI Researcher working on autonomous vehicles and human-robot interaction at MIT)

Books* “Mindset” (by Carol Dweck) * “Outliers” (by Malcolm Gladwell)

NotesMy conversation with Mohamed was recorded back in March 2021. Here are some updates that Mohamed shared with me since then:

  • Kolena is an ML testing and validation platform that enables teams to implement testing best practices to rigorously test their models’ behavior and ship high-quality ML products much faster.
  • Mohamed and his team have signed a couple of big enterprise customers and raised a large seed round from top-tier investors and almost every industry leader in the AI space. These were strong signals that Kolena is solving a very important problem!
  • Mohamed’s first impression on the market is: the ML market is hungry for a reliable testing platform for models. Kolena has quite of a waitlist and plans to launch early next year.

About the showDatacast features long-form, in-depth conversations with practitioners and researchers in the data community to walk through their professional journeys and unpack the lessons learned along the way. I invite guests coming from a wide range of career paths — from scientists and analysts to founders and investors — to analyze the case for using data in the real world and extract their mental models (“the WHY and the HOW”) behind their pursuits. Hopefully, these conversations can serve as valuable tools for early-stage data professionals as they navigate their own careers in the exciting data universe.

Datacast is produced and edited by James Le. Get in touch with feedback or guest suggestions by emailing khanhle.1013@gmail.com.

Subscribe by searching for Datacast wherever you get podcasts or click one of the links below:

  • Listen on Spotify
  • Listen on Apple Podcasts
  • Listen on Google Podcasts

If you’re new, see the podcast homepage for the most recent episodes to listen to, or browse the full guest list.

View Details

Show Notes* (01:46) Jennifer shared her formative experiences growing up in France and wanting to be a physicist. * (03:04) Jennifer unpacked the evolution of her academic journey in France — getting Physics degrees at Louis Pasteur University, Paris-Sud University, and Sorbonne University. * (06:44) Jennifer mentioned her time as a Postdoctoral Researcher in Neutrino Physics at Duke University, where her research group lacked the funding to carry on scientific projects. * (09:35) Jennifer discussed her transition from academia to industry, working as a Quantitative Research Scientist at Quantlab Financial in Houston. * (13:31) Jennifer went over her move to the Bay Area, working for YuMe and Ayasdi — growing and managing early-stage data science teams at both places. * (19:19) Jennifer recalled her foray into becoming a Senior Data Science Manager of the Search team at Walmart Labs. She managed the Metrics-Measurements-Insights team and the Store-Search team. * (23:59) Jennifer shared the business anecdote that made her obsessed with measuring the ROI of data science. * (28:46) Jennifer reflected on the opportunity to give conference talks and become a thought leader in the data science community (watch her first industry talk, “Review Analysis: An Approach to Leveraging User-Generated Content in the Context of Retail” at MLconf 2016). * (31:10) Jennifer unpacked her interest in active learning and outlined existing challenges of making active learning performant in real-world ML systems. * (36:58) After 1.5 years with Walmart Labs, Jennifer became the Chief Data Scientist at Atlassian. She shared the tactics to grow the Search & Smarts team of scientists and engineers from 3 to 17 people in less than 6 months across 3 locations. * (40:31) Jennifer discussed the organizational and operational challenges with making ML useful in enterprises and the importance of data preparation in the modern ML stack. * (47:24) Jennifer elaborated on the topic of “Agile for Data Science Teams,” which discusses that organizations that invest in ML but do not get the organizational side of things right will fail. * (53:09) Jennifer went over her decision to accept a VP of Machine Learning role at Figure Eight, then a frontier startup that offers enterprise-grade labeling solutions to ML teams. * (57:56) Jennifer went over the inception of her startup Alectio, whose mission is to help companies do ML more efficiently with fewer data and help the world do ML more sustainably by reducing the industry’s carbon footprint. * (01:04:32) Jennifer unpacked her 4-part blog series about responsible AI that calls out the need to fight bias, increase accessibility, and create more opportunities in AI. * (01:09:06) Jennifer discussed the hurdles she had to jump through to find early adopters of Alectio. * (01:11:03) Jennifer emphasized the valuable lessons learned to attract the right people who are excited about Alectio’s mission. * (01:14:38) Jennifer cautioned the danger of taking advice without thinking through how it can be applied to one’s career. * (01:17:09) Jennifer condensed her decade of experience navigating the tech industry as a woman into concrete advice. * (01:19:19) Closing segment.

Jennifer’s Contact Info* LinkedIn * Twitter * Medium

Alectio’s Resources* Website * Twitter * LinkedIn * What Is Alectio? (Video) * Is Big Data Dragging Us Towards Another AI Winter? (Article)

Mentioned ContentTalks* The Day Big Data Died (Oct 2020 @ Interop Digital) * The Importance of Ethics in Data Science (Keynote @ Women in Analytics Conference 2019) * Introduction to Active Learning (ODSC London 2018) * Agile for Data Science Teams (Strata Data Conf — New York 2018) * Big Data and the Advent of Data Mixology (Interop ITX — The Future of Data Summit 2017) * The Limitations of Big Data In Predictive Analytics (DataEngConf SF 2017) * Review Analysis: An Approach to Leveraging User-Generated Content in the Context of Retail (MLconf 2016)

Articles1 — Women vs. The Workplace Series

  • Gender Discrimination (Oct 2015)
  • Why Leading By Example Matters (Jan 2017)
  • Data Scientist: the SexISTiest Job of the 21st Century? (Feb 2017)
  • The Role of Motherhood in Gender Discrimination (March 2017)
  • The Biggest Challenges of the Female Manager (May 2017)
  • Parity in the Workplace: Why We Are Not There Yet (Dec 2017)
  • The Pyramid of Needs of Professional Women (Dec 2017)

2 — Management Series

  • The Secrets to Successfully Managing an Underperformer (June 2017)
  • The Top Secrets to Managing a Rockstar (July 2017)
  • The Real Cost of Hiring Over-Qualified Candidates in Technology (March 2018)
  • Team Culture (May 2018)

3 — Responsible AI Series

  • How We Got Responsible AI All Wrong (Part 1)
  • Impact, Bias, and Sustainability in AI (Part 2)
  • Increasing Accessibility to AI (Part 3)
  • Creating More Opportunities in AI (Part 4)

Book* “Managing Up” (by Rosanne Badowski and Roger Gittines)

NotesJennifer told me that Alectio is about to launch a community version that people will be able to compete to get the best model with the minimum amount of data this fall. Be sure to check out their blog and follow them on LinkedIn!

About the showDatacast features long-form, in-depth conversations with practitioners and researchers in the data community to walk through their professional journeys and unpack the lessons learned along the way. I invite guests coming from a wide range of career paths — from scientists and analysts to founders and investors — to analyze the case for using data in the real world and extract their mental models (“the WHY and the HOW”) behind their pursuits. Hopefully, these conversations can serve as valuable tools for early-stage data professionals as they navigate their own careers in the exciting data universe.

Datacast is produced and edited by James Le. Get in touch with feedback or guest suggestions by emailing khanhle.1013@gmail.com.

Subscribe by searching for Datacast wherever you get podcasts or click one of the links below:

  • Listen on Spotify
  • Listen on Apple Podcasts
  • Listen on Google Podcasts

If you’re new, see the podcast homepage for the most recent episodes to listen to, or browse the full guest list.

View Details

Show Notes (01:48) Sarah talked about the formative experiences of her upbringing: growing up interested in the natural sciences and switching focus on terrorism analysis after experiencing the 9/11 tragedy with her own eyes. * (04:07) Sarah discussed her experience studying International Security Studies at Stanford and working at the Center for International Security and Cooperation. * (07:15) Sarah recalled her first job out of college as a Program Director at the Center for Advanced Defense Studies — collaborating with academic researchers to develop computational approaches that counter terrorism and piracy. * (09:48) Sarah went over her time as a cyber-intelligence analyst at Cyveillance, which provided threat intelligence services to enterprises worldwide. * (12:22) Sarah walked over her time at Palantir as an embedded analyst, where she observed the struggles that many agencies had with data integration and modeling challenges. * (15:26) Sarah unpacked the challenges of building out the data team and applying the data work at Mattermark. * (20:15) Sarah shared her opinion on the career trajectory for data analysts and data scientists, given her experience as a manager for these roles. * (23:43) Sarah shared the power of having a peer group and building a team culture that she was proud of at Mattermark. * (26:41) Sarah joined Canvas Ventures as a Data Partner in 2016 and shared her motivation for getting into venture capital. * (29:47) Sarah revealed the secret sauce to succeed in venture — stamina*. * (32:00) Sarah has been an investor at Amplify Partners since 2017 and shared what attracted her about the firm’s investment thesis and the team. * (35:28) Sarah walked through the framework she used to prove her value upfront as the new investor at Amplify. * (38:35) Sarah shared the details behind her investment on the Series A round for OctoML, a Seattle-based startup that leverages Apache TVM to enable their clients to simply, securely, and efficiently deploy any model on any hardware backend. * (44:39) Sarah dissected her investment on the seed round for Einblick, a Boston-based startup that builds a visual computing platform for BI and analytics use cases. * (48:45) Sarah mentioned the key factors inspiring her investment in the seed round for Metaphor Data, a meta-data platform that grew out of the DataHub open-source project developed at LinkedIn. * (53:57) Sarah discussed what triggered her investment in the Series A round for Runway, a New York-based team building the next-generation creative toolkit powered by machine learning. * (58:36) Sarah unpacked the advice she has been giving her portfolio companies in hiring decisions and expanding their founding team (and advice they should ignore). * (01:01:29) Sarah went over the process of curating her weekly newsletter called Projects To Know (active since 2019). * (01:05:00) Sarah predicted the 3 trends in the data ecosystem that will have a disproportionately huge impact in the future. * (01:11:15) Closing segment.

Sarah’s Contact Info

  • Amplify Page
  • Twitter
  • LinkedIn
  • Medium

Amplify Partners’ Resources

  • Website
  • Team
  • Portfolio
  • Blog

Mentioned ContentBlog Posts

  • Our Investment in OctoML
  • Announcing Our Investment in Einblick
  • Our Investment in Metaphor Data
  • Our Series A Investment in Runway

People

  • Sunil Dhaliwal (General Partner at Amplify Partners)
  • Mike Dauber (General Partner at Amplify Partners)
  • Lenny Pruss (General Partner at Amplify Partners)
  • Mike Volpi (Co-Founder and Partner at Index Ventures)
  • Gary Little (Co-Founder and General Partner at Canvas Ventures)

Book

  • “Zen and the Art of Motorcycle Maintenance” (by Robert Pirsig)

New UpdatesSince the podcast was recorded, Sarah has been keeping her stamina high!

  • Her investments in Hex (data workspace for teams) and Meroxa (real-time data platform) have been made public.
  • She has also spoken at various panels, including SIGMOD, REWORK, University of Chicago, and Utah Nerd Nights.

Be sure to follow @sarahcat21 on Twitter to subscribe to her brain on the intersection of data, VC, and startups!

View Details

Show Notes* (01:39) Aparna talked about her Bachelor’s degree in Electrical Engineering and Computer Science at UC Berkeley. * (02:50) Aparna shared her undergraduate research experience at the Energy and Sustainable Technologies lab. * (04:34) Aparna discussed valuable lessons learned from her industry internships at TubeMogul and compared the objective with that of a research environment. * (08:26) Aparna then joined Uber as a software engineer on the Marketplace Forecasting team, where she led the development of Uber’s first model lifecycle management system for running ML model computations at scale to power Uber’s dynamic pricing algorithms. * (12:40) Aparna talked about how she became interested in model monitoring while Uber’s model store. * (17:29) Aparna discussed her decision to join the Ph.D. program in Computer Vision at Cornell University, specifically about bias in model, after spending 3 years at Uber. * (23:40) Aparna shared the backstory behind co-founding MonitorML with her brother Eswar and going through the 2019 summer batch of Y-Combinator. * (26:47) Aparna discussed the acquisition of MonitorML by Arize AI, where she’s currently the Chief Product Officer. * (28:41) Aparna unpacked the key insights in her ongoing ML Observability blog series, which argues that model observability is the foundational platform that empowers teams to continually deliver and improve results from the lab to production. * (33:17) Aparna shared her verdict for the ML tooling ecosystem in the upcoming years from her in-depth exploration of ML infrastructure tools covering data preparation, model building, model validation, and model serving. * (37:01) Aparna briefly shared the challenges encountered to get the first cohort of customers for Arize. * (39:23) Aparna went over valuable lessons to attract the right people who are excited about Arize’s mission. * (41:04) Aparna shared her advice for founders who are in the process of finding the right investors for their companies. * (42:24) Aparna reasoned how participating in The Amazing Race was similar to running a startup. * (44:59) Closing segment.

Aparna’s Contact Info

  • Twitter
  • LinkedIn
  • Medium
  • Forbes Column
  • Website
  • Github
  • Google Scholar

Arize’s Resources

  • Website
  • Medium
  • LinkedIn
  • Twitter

Mentioned ContentBlog Posts

  • ML Infrastructure Tools for Data Preparation (May 2020)
  • ML Infrastructure Tools for Model Building (May 2020)
  • ML Infrastructure Tools for Production (Part 1) (May 2020)
  • ML Infrastructure Tools for Production (Part 2) (Sep 2020)
  • ML Infrastructure Tools — ML Observability (Feb 2021)
  • The Model’s Shipped — What Could Possibly Go Wrong? (Feb 2021)

People

  • Rediet Abebe (Assistant Professor of Computer Science at UC Berkeley and Junior Fellow at the Harvard Society of Fellows)
  • Timnit Gebru (Founder of Black in AI, Ex-Research Scientist at Google)
  • Serge Belongie (Professor of Computer Science at Cornell and Aparna’s past Ph.D. advisor)
  • Solon Barocas (Principal Researcher at Microsoft Research and Adjunct Assistant Professor of Information Science at Cornell)
  • Manish Raghavan (Ph.D. candidate in the Computer Science department at Cornell)
  • Kate Crawford (Principal Researcher at Microsoft Research and Co-founder/Director of research at NYU’s AI Now Institute)

Book

  • “The Hard Thing About The Hard Things” (by Ben Horowitz)

New UpdatesSince the podcast was recorded, a lot has happened at Arize AI!

  • Aparna has continued writing the ML observability series: The Playbook to Monitor Your Model’s Performance in Production (March 2021) and Beyond Monitoring: The Rise of Observability (May 2021).
  • Arize has been recognized in Forbes’s AI 50 2021: Most Promising AI Companies.
  • Aparna has also contributed to Forbes various articles: from the Chronicles of AI Ethics and Q&A with Ethics researchers, to a list of Women in AI to watch and emerging ML tooling categories.

About The ShowDatacast features long-form, in-depth conversations with practitioners and researchers in the data community to walk through their professional journeys and unpack the lessons learned along the way. I invite guests coming from a wide range of career paths — from scientists and analysts to founders and investors — to analyze the case for using data in the real world and extract their mental models (“the WHY”) behind their pursuits. Hopefully, these conversations can serve as valuable tools for early-stage data professionals as they navigate their own careers in the exciting data universe.

Datacast is produced and edited by James Le. Get in touch with feedback or guest suggestions by emailing khanhle.1013@gmail.com.

Subscribe by searching for Datacast wherever you get podcasts or click one of the links below:

  • Listen on Spotify
  • Listen on Apple Podcasts
  • Listen on Google Podcasts

If you’re new, see the podcast homepage for the most recent episodes to listen to, or browse the full guest list.

View Details

Show Notes* (02:07) Emeli shared her educational background getting degrees in Applied Mathematics and Informatics from the Peoples’ Friendship University of Russia in the early 2010s. * (04:24) Emeli went over her experience getting a Master’s Degree at Yandex School of Data Analysis. * (07:06) Emeli reflected on lessons learned from her first job out of university working as a Software Developer at Rambler, one of the biggest Russian web portals. * (09:33) Emeli walked over her first year as a Data Scientist developing e-commerce recommendation systems at Yandex. * (13:38) Emeli discussed core projects accomplished as the Chief Data Scientist at Yandex Data Factory, Yandex’s end-to-end data platform. * (17:52) Emeli shared her learnings transitioning from an IC to a manager role. * (19:21) Emeli mentioned key components of success for industrial AI, given her time as the co-founder and Chief Data Scientist at Mechanica AI. * (22:40) Emeli dissected the makings of her Coursera specializations — “Machine Learning and Data Analysis” and “Big Data Essentials.” * (26:14) Emeli discussed her teaching activities at Moscow Institute of Physics and Technology, Yandex School of Data Analysis, Harbour.Space, and Graduate School of Management — St. Petersburg State University. * (30:12) Emeli shared the story behind the founding of Evidently AI, which is building a human interface to machine learning, so that companies can trust, monitor, and improve the performance of their AI solutions. * (32:32) Emeli explained the concept of model monitoring and exposed the monitoring gap in the enterprise (read Part 1 and Part 2 of the Monitoring series). * (34:13) Emeli looked at possible data quality and integrity issues while proposing how to track them (read Part 3, Part 4, and Part 5 of the Monitoring series). * (36:47) Emeli revealed the pros and cons of building an open-source product. * (39:13) Emeli talked about prioritizing product roadmap for Evidently AI. * (41:24) Emeli described the data community in Moscow. * (42:03) Closing segment.

Emeli’s Contact Info

  • LinkedIn
  • Twitter
  • Coursera
  • GitHub
  • Medium

Evidently AI’s Resources

  • Website
  • Twitter
  • LinkedIn
  • GitHub
  • Documentation

Mentioned ContentBlog Posts

  • ML Monitoring, Part 1: What Is It and How It Differs? (Aug 2020)
  • ML Monitoring, Part 2: Who Should Care and What We Are Missing? (Aug 2020)
  • ML Monitoring, Part 3: What Can Go Wrong With Your Data? (Sep 2020)
  • ML Monitoring, Part 4: How To Track Data Quality and Data Integrity? (Oct 2020)
  • ML Monitoring, Part 5: Why Should You Care About Data And Concept Drift? (Nov 2020)
  • ML Monitoring, Part 6: Can You Build a Machine Learning Model to Monitor Another Model? (April 2021)

Courses

  • “Machine Learning and Data Analysis”
  • “Big Data Essentials”

People

  • Yann LeCun (Professor at NYU, Chief AI Scientist at Facebook)
  • Tomas Mikolov (the creator of Word2Vec, ex-scientist at Google and Facebook)
  • Andrew Ng (Professor at Stanford, Co-Founder of Google Brain, Coursera, and Landing AI, Ex-Chief Scientist at Baidu)

Book

  • “The Elements of Statistical Learning” (by Trevor Hastie, Robert Tibshirani, and Jerome Friedman)

New UpdatesSince the podcast was recorded, a lot has happened at Evidently! You can use this open-source tool (https://github.com/evidentlyai/evidently) to generate a variety of interactive reports on the ML model performance and integrate it into your pipelines using JSON profiles.

This monitoring tutorial is a great showcase of what can go wrong with your models in production and how to keep an eye on them: https://evidentlyai.com/blog/tutorial-1-model-analytics-in-production.

About The ShowDatacast features long-form conversations with practitioners and researchers in the data community to walk through their professional journey and unpack the lessons learned along the way. I invite guests coming from a wide range of career paths - from scientists and analysts to founders and investors — to analyze the case for using data in the real world and extract their mental models (“the WHY”) behind their pursuits. Hopefully, these conversations can serve as valuable tools for early-stage data professionals as they navigate their own careers in the exciting data universe.

Datacast is produced and edited by James Le. Get in touch with feedback or guest suggestions by emailing khanhle.1013@gmail.com.

Subscribe by searching for Datacast wherever you get podcasts or click one of the links below:

  • Listen on Spotify
  • Listen on Apple Podcasts
  • Listen on Google Podcasts

If you're new, see the podcast homepage for the most recent episodes to listen to, or browse the full guest list.

View Details

Show Notes* (01:59) David recalled his undergraduate experience studying Physics and Mathematics at Duke University back in the early 90s. * (05:55) David reflected on his decision to pursue a Ph.D. in Physics at the University of Maryland, College Park, specializing in Nonlinear Dynamics and Chaos Theory. * (10:18) David unpacked his Nature paper called “Topology in Chaotic Scattering.” * (14:43) David went over his two papers on fractal dimensions in higher-dimensional chaotic scattering following his Nature publication. * (21:42) David talked about his project K Desktop Environment, which provides a free, user-friendly desktop for Linux/UNIX systems (later turned into a print book with MacMillan Publishing in 2000). * (24:20) David explained the premise behind his work on Andamooka, a site that supports open content. * (27:24) David walked over his time as a quantitative analyst at Thales Fund Management after finishing his Ph.D. * (30:50) David discussed his 4-year stint at Lehman Brothers — moving up the ladder into a Vice President role, up until Barclay’s Capital acquired it. * (33:24) David talked about his proudest accomplishment during the 5-year stint as a headdesk in equities trader at KCG/GETCO. * (35:37) David shared war stories while working at an investment firm called Teza Technologies and co-founding Galaxy Digital Trading (specializing in cryptocurrency trading). * (41:34) David unpacked key concepts covered in his guest lectures on optimization of high-frequency trading systems at NYU Stern School of Business. * (44:26) David explained his career change to work as a Machine Learning Engineer at Instagram in the summer of 2019. * (47:17) David briefly mentioned his transition back to a quant trader role at 3Red Partners. * (48:05) David is writing a technical book with Manning called “Tuning Up,” which provides a toolbox of experimental methods that will boost the effectiveness of machine learning systems, trading strategies, infrastructure, and more. * (50:48) David reflected on the benefits of his physics academic background for his quant analyst career. * (52:27) Closing segment.

David’s Contact Info* Website * LinkedIn * Twitter

Mentioned ContentPublications

  • "Topology In Chaotic Scattering" (Nature, May 1999)
  • "Fractal Dimension of Higher-Dimensional Chaotic Repellors" (June 1999)
  • "Fractal Basin Boundaries in Higher-Dimensional Chaotic Scattering"

Book

  • “The Elements of Statistical Learning” (by Trevor Hastie, Robert Tibshirani, and Jerome Friedman)

People

  • Jim Simons (Founder of Renaissance Technologies)
  • Michael Kearns (Professor at the University of Pennsylvania, previously leading Morgan Stanley’s AI Center of Excellence)
  • Vasant Dhar (Professor at NYU Stern School of Business, Founder of SCT Capital)

Tuning Up — From A/B testing to Bayesian optimizationManning’s permanent 40% discount code (good for all Manning products in all formats) for Datacast listeners: poddcast19.

You can refer to this link: http://mng.bz/4MAR.

Here are two free eBook codes to get copies of Tuning Up for two lucky Datacast listeners: tngdtcr-AB2C and tngdtcr-6D43

You can refer to this link: http://mng.bz/G6Bq.

View Details

Show Notes* (02:06) Fabiana talked about her Bachelor’s degree in Applied Mathematics from the University of Lisbon in the early 2010s. * (04:18) Fabiana shared lessons learned from her first job out of college as a Siebel and BI Developer at Novabase. * (05:13) Fabiana discussed unique challenges while working as an IoT Solutions Architect at Vodafone. * (09:56) Fabiana mentioned projects she contributed to as a Data Scientist at startups such as ODYSAI and Habit Analytics. * (12:44) Fabiana talked about the two Master’s degrees she got while working in the industry (Applied Econometrics from Lisbon School of Economics and Management and Business Intelligence from NOVA IMS Information Management School). * (14:41) Fabiana distinguished the difference between data science and business intelligence. * (18:01) Fabiana shared the founding story of YData, the first data-centric platform with synthetic data, whose she is currently the Chief Data Officer. * (21:32) Fabiana discussed different techniques to generate synthetic data, including oversampling, Bayesian Networks, and generative models. * (24:01) Fabiana unpacked the key insights in her blog series on generating synthetic tabular data. * (29:40) Fabiana summarized novel design and optimization techniques to cope with the challenges of training GAN models. * (33:44) Fabiana brought up the benefits of using Differential Privacy as a complement to synthetic data generation. * (38:07) Fabiana unpacked her post “The Cost of Poor Data Quality,” — where she defined data quality as data measures based on factors such as accuracy, completeness, consistency, reliability, and above all, whether it is up to date. * (42:11) Fabiana explained the important role that data quality plays in ensuring model explainability. * (44:57) Fabiana reasoned about YData’s decision to pursue the open-source strategy. * (47:47) Fabiana discussed her podcast called “When Machine Learning Meets Privacy” in collaboration with the MLOps Slack community. * (49:14) Fabiana briefly shared the challenges encountered to get the first cohort of customers for YData. * (50:12) Fabiana went over valuable lessons to attract the right people who are excited about YData’s mission. * (51:52) Fabiana shared her take on the data community in Lisbon and her effort to inspire more women to join the tech industry. * (53:47) Closing segment.

Fabiana’s Contact Info* LinkedIn * Medium * Twitter

YData’s Resources* Website * Github * LinkedIn * Twitter * AngelList * Synthetic Data Community

Mentioned ContentBlog Posts

  • Synthetic Data: The Future Standard for Data Science Development (April 2020)
  • Generating Synthetic Tabular Data with GANs — Part 1 (May 2020)
  • Generating Synthetic Tabular Data with GANs — Part 2 (May 2020)
  • What Is Differential Privacy? (May 2020)
  • What Is Going On With My GAN? (July 2020)
  • How To Generate Synthetic Tabular Data? Wasserstein Loss for GANs (Sep 2020)
  • The Cost of Poor Data Quality (Sep 2020)
  • How Can I Explain My ML Models To The Business? (Oct 2020)
  • Synthetic Time-Series Data: A GAN Approach (Jan 2021)

Podcast

  • “When Machine Learning Meets Privacy”

People

  • Jean-Francois Rajotte (Resident Data Scientist at the University of British Columbia)
  • Sumit Mukherjee (Associate Professor of Statistics at Columbia University)
  • Andrew Trask (Leader at OpenMined, Research Scientist at DeepMind, Ph.D. Student at the University of Oxford)
  • Théo Ryffel (Co-Founder of Arkhn, Ph.D. Student at ENS and INRIA, Leader at OpenMined)

Recent Announcements/Articles

  • Partnerships with UbiOps and Algorithmia
  • The rise of DataPrepOps (March 2021)
  • From model-centric to data-centric (March 2021)

View Details

Show Notes (02:06) Azin described her childhood growing up in Iran and going to a girls-only high school in Tehran designed specifically for extraordinary talents. * (05:08) Azin went over her undergraduate experience studying Computer Science at the University of Tehran. * (10:41) Azin shared her academic experience getting a Computer Science MS degree at the University of Toronto, supervised by Babak Taati and David Fleet. * (14:07) Azin talked about her teaching assistant experience for a variety of CS courses at Toronto. * (15:54) Azin briefly discussed her 2017 report titled “Barriers to Adoption of Information Technology in Healthcare,” which takes a system thinking perspective to identify barriers to the application of IT in healthcare and outline the solutions. * (19:35) Azin unpacked her MS thesis called “Subspace Selection to Suppress Confounding Source Domain Information in AAM Transfer Learning,” which explores transfer learning in the context of facial analysis. * (28:48) Azin discussed her work as a research assistant at the Toronto Rehabilitation Institute, working on a research project that addressed algorithmic biases in facial detection technology for older adults with dementia. * (33:02) Azin has been an Applied Research Scientist at Georgian since 2018, a venture capital firm in Canada that focuses on investing in companies operating in the IT sectors. * (38:20) Azin shared the details of her initial Georgian project to develop a robust and accurate injury prediction model using a hybrid instance-based transfer learning method.* * (42:12) Azin unpacked her Medium blog post discussing transfer learning in-depth (problems, approaches, and applications). * (48:18) Azin explained how transfer learning could address the widespread “cold-start” problem in the industry. * (49:50) Azin shared the challenges of working on a fintech platform with a team of engineers at Georgian on various areas such as supervised learning, explainability, and representation learning. * (51:46) Azin went over her project with Tractable AI, a UK-based company that develops AI applications for accident and disaster recovery. * (55:26) Azin shared her excitement for ML applications using data-efficient methods to enhance life quality. * (57:46) Closing segment.

Azin’s Contact Info* Website * Twitter * LinkedIn * Google Scholar * GitHub

Mentioned ContentPublications

  • “Barriers to Adoption of Information Technology in Healthcare” (2017)
  • “Subspace Selection to Suppress Confounding Source Domain Information in AAM TransferLearning” (2017)
  • “A Hybrid Instance-based Transfer Learning Method” (2018)
  • “Prediction of Workplace Injuries” (2019)
  • “Algorithmic Bias in Clinical Populations — Evaluating and Improving Facial Analysis Technology in Older Adults with Dementia” (2019)
  • “Limitations and Biases in Facial Landmark Detection” (2019)

Blog Posts

  • “An Introduction to Transfer Learning” (Dec 2018)
  • “Overcoming The Cold-Start Problem: How We Make Intractable Tasks Tractable” (April 2021)

People

  • Yoshua Bengio (Professor of Computer Science and Operations Research at University of Montreal)
  • Geoffrey Hinton (Professor of Computer Science at University of Toronto)
  • Louis-Philippe Morency (Associate Professor of Computer Science at Carnegie Mellon University)

Book

  • “Machine Learning: A Probabilistic Approach” (by Kevin Murphy)

Note: Azin and her collaborator are going to give a talk at ODSC Europe 2021 in June about a Georgian’s project with a portfolio company, Tractable. They have written a short blog post about it too which you can find HERE.

View Details

Show Notes* (02:09) Gordon briefly talked about his undergraduate studying Psychology and Philosophy at Rutgers University in the early 90s. * (03:24) Gordon reflected on the first decade of his career getting into database technologies. * (05:34) Gordon discussed his predilection towards consulting, specifically his role in the professional services team at AB Initio Software in the early 2000s. * (08:02) Gordon recalled the challenges of leading data warehousing initiatives at Smarter Travel Media and ClickSquared in the 2000s. * (13:14) Gordon emphasized the advantage of a multi-tenant database over a traditional relational database. * (18:30) Gordon recalled his one-year stint at Cervello, leading business intelligence implementations for their clients. * (21:59) Gordon elaborated on his projects during his 3 years as the director of business intelligence infrastructure at Fitbit. * (26:09) Gordon dived into his framework of choosing data tooling vendors while at Fitbit (and how he settled with a tiny startup called Snowflake back then). * (30:02) Gordon provided recommendations for startups to be data-driven. * (33:24) Gordon recalled practices to foster effective collaboration while managing the 3 teams of data engineering, data warehousing, and data analytics at Fitbit. * (36:44) Gordon went over his proudest accomplishment as the director of data engineering at ezCater, making substantial improvements to their data warehouse platform. * (38:59) Gordon shared his framework for interviewing data engineers. * (41:39) Gordon walked through his consulting engagement in analytics engineering for Zipcar and data warehousing for edX. * (46:17) Gordon reflected on his time as the Vice President of business intelligence at HubSpot. * (50:50) Gordon unpacked his notion of “Data Hierarchy of Needs,” which entails the five pillars — data security, data quality, system reliability, user experience, and data coverage. * (56:55) Gordon discussed current opportunities for driving better social outcomes and empowering democracy through data. * (59:48) Gordon shared the key criteria that enable healthy team dynamics from his hands-on experience building data teams. * (01:02:13) Gordon unpacked the central features and benefits of Snowflake for the un-initiated. * (01:06:25) Gordon gave his verdict for the ETL tooling landscape in the next few years. * (01:08:33) Gordon described the data community in Boston. * (01:09:52) Closing segment.

Gordon’s Contact Info* LinkedIn

Mentioned ContentPeople

  • Tristan Handy (co-founder of Fishtown Analytics and co-creator of dbt)
  • Michael Kaminsky (who coined the term “Analytics Engineering”)
  • Barr Moses (co-founder and CEO of Monte Carlo, who coined the term “Data Observability”)

Book

  • “Start With Why” (By Simon Sinek)

View Details

Show Notes* (2:05) Louis went over his childhood as a self-taught programmer and his early days in school as a freelance developer. * (4:22) Louis described his overall undergraduate experience getting a Bachelor’s degree in IT Systems Engineering from Hasso Plattner Institute, a highly-ranked computer science university in Germany. * (6:10) Louis dissected his Bachelor thesis at HPI called “Differentiable Convolutional Neural Network Architectures for Time Series Classification,” — which addresses the problem of automatically designing architectures for time series classification efficiently, using a regularization technique for ConvNet that enables joint training of network weights and architecture through back-propagation. * (7:40) Louis provided a brief overview of his publication “Transfer Learning for Speech Recognition on a Budget,” — which explores Automatic Speech Recognition training by model adaptation under constrained GPU memory, throughput, and training data. * (10:31) Louis described his one-year Master of Research degree in Computational Statistics and Machine Learning at the University College London supervised by David Barber. * (12:13) Louis unpacked his paper “Modular Networks: Learning to Decompose Neural Computation,” published at NeurIPS 2018 — which proposes a training algorithm that flexibly chooses neural modules based on the processed data. * (15:13) Louis briefly reviewed his technical report, “Scaling Neural Networks Through Sparsity,” which discusses near-term and long-term solutions to handle sparsity between neural layers. * (18:30) Louis mentioned his report, “Characteristics of Machine Learning Research with Impact,” which explores questions such as how to measure research impact and what questions the machine learning community should focus on to maximize impact. * (21:16) Louis explained his report, “Contemporary Challenges in Artificial Intelligence,” which covers lifelong learning, scalability, generalization, self-referential algorithms, and benchmarks. * (23:16) Louis talked about his motivation to start a blog and discussed his two-part blog series on intelligence theories (part 1 on universal AI and part 2 on active inference). * (27:46) Louis described his decision to pursue a Ph.D. at the Swiss AI Lab IDSIA in Lugano, Switzerland, where he has been working on Meta Reinforcement Learning agents with Jürgen Schmidhuber. * (30:06) Louis created a very extensive map of reinforcement learning in 2019 that outlines the goal, methods, and challenges associated with the RL domain. * (33:50) Louis unpacked his blog post reflecting on his experience at NeurIPS 2018 and providing updates on the AGI roadmap regarding topics such as scalability, continual learning, meta-learning, and benchmarks. * (37:04) Louis dissected his ICLR 2020 paper “Improving Generalization in Meta Reinforcement Learning using Learned Objectives,” which introduces a novel algorithm called MetaGenRL, inspired by biological evolution. * (44:03) Louis elaborated on his publication “Meta-Learning Backpropagation And Improving It,” which introduces the Variable Shared Meta-Learning framework that unifies existing meta-learning approaches and demonstrates that simple weight-sharing and sparsity in a network are sufficient to express powerful learning algorithms. * (51:14) Louis expands on his idea to bootstrap AI that entails how the task, the general meta learner, and the unsupervised objective should interact (proposed at the end of his invited talk at NeurIPS 2020). * (54:14) Louis shared his advice for individuals who want to make a dent in AI research. * (56:05) Louis shared his three most useful productivity tips. * (58:36) Closing segment.

Louis’s Contact Info* Website * Twitter * LinkedIn * Google Scholar * GitHub

Mentioned ContentPapers and Reports

  • Differentiable Convolutional Neural Network Architectures for Time Series Classification (2017)
  • Transfer Learning for Speech Recognition on a Budget (2017)
  • Modular Networks: Learning to Decompose Neural Computation (2018)
  • Contemporary Challenges in Artificial Intelligence (2018)
  • Characteristics of Machine Learning Research with Impact (2018)
  • Scaling Neural Networks Through Sparsity (2018)
  • Improving Generalization in Meta Reinforcement Learning using Learned Objectives (2019)
  • Meta-Learning Backpropagation And Improving It (2020)

Blog Posts

  • Theories of Intelligence — Part 1 and Part 2 (July 2018)
  • Modular Networks: Learning to Decompose Neural Computation (May 2018)
  • How to Make Your ML Research More Impactful (Dec 2018)
  • A Map of Reinforcement Learning (Jan 2019)
  • NeurIPS 2018, Updates on the AI Roadmap (Jan 2019)
  • MetaGenRL: Improving Generalization in Meta Reinforcement Learning (Oct 2019)
  • General Meta-Learning and Variable Sharing (Nov 2020)

People

  • Jeff Clune (for his push on meta-learning research)
  • Kenneth Stanley (for his deep thoughts on open-ended learning)
  • Jürgen Schmidhuber (for being a visionary scientist)

Book

  • “Grit” (by Angela Duckworth)

View Details

Show Notes* (01:58) Dzejla described her undergraduate experience studying Computer Science at the Sarajevo School of Science and Technology back in the mid-2000s. * (07:59) Dzejla recapped her overall experience getting a Ph.D. in Computer Science at Stony Brook University. * (14:38) Dzejla unpacked the key research problem in her Ph.D. thesis titled “Upper and Lower Bounds on Sorting and Searching in External Memory.” * (19:13) Dzejla went over the details of her paper “Don’t Thrash: How to Cache Your Hash on Flash,” — which describes the Cascade Filter, an approximate-membership-query data structure that scales beyond main memory, that is an alternative to the well-known Bloom-filter data structure. * (24:41) Dzejla elaborated on her work “The batched predecessor problem in external memory,” — which studies the lower bounds in three external memory models: the I/O comparison model, the I/O pointer-machine model, and the index-ability model. * (29:56) Dzejla shared her learnings from being a Teaching Assistant for the Introduction to Algorithms course at Stony Brook (both at the undergraduate and graduate level). * (35:08) Dzejla went over her summer internships at Microsoft’s Server and Tools Division during her Ph.D. * (41:06) Dzejla reasoned about her decision to return to Sarajevo School of Science and Technology as an Assistant Professor of Computer Science. * (47:22) Dzejla dissected the essential concepts and methods covered in her Data Structures, Introductory Algorithms, Advanced Algorithms, and Algorithms for Big Data courses taught at SSIT. * (48:42) Dzejla provided a brief overview of the Computer Science/Software Engineering department at the International University of Sarajevo (where she has been a professor since 2017. * (50:57) Dzejla briefly talked about the courses that she taught at IUS, including Intro to Programming, Human-Computer Interaction, and Algorithms/Data Structures. * (52:49) Dzejla shared the challenges of writing Algorithms and Data Structures for Massive Datasets, which introduces data processing and analytics techniques specifically designed for large distributed datasets. * (56:14) Dzejla explained concepts in Part 1 of the book — including Hash Tables, Approximate Membership, Bloom Filters, Frequency/Cardinality Estimation, Count-Min Sketch, and Hyperloglog. * (58:38) Dzejla provided a brief overview of techniques to handle streaming data in Part 2 of the book. * (01:00:14) Dzejla mentioned the data structures for large databases and external-memory algorithms in Part 3 of the book. * (01:02:15) Dzejla shared her thoughts about the tech community in Sarajevo. * (01:04:16) Closing segment.

Dzejla’s Contact Info* LinkedIn * Twitter * Google Scholar

Mentioned ContentPapers

  • “Upper and Lower Bounds on Sorting and Searching in External Memory” (Dzejla’s Ph.D. Thesis, 2014)
  • “Don’t Thrash: How to Cache Your Hash on Flash” (2012)
  • “The batched predecessor problem in external memory” (2014)

People

  • Erik Demaine (Computer Science Professor at MIT)
  • Michael Bender (Computer Science Professor at Stony Brook, Dzejla’s Ph.D. Advisor)
  • Joseph Mitchell (Computational Geometry Professor at Stony Brook)
  • Steven Skiena (Computer Science Professor at Stony Brook)
  • Jeff Erickson (Computer Science Professor at UIUC)

Books

  • “Algorithms and Data Structures for Massive Datasets” (by Dzejla Medjedovic, Emin Tahirovic, and Ines Dedovic)
  • “The Algorithm Design Manual” (by Steven Skiena)

Here is a permanent 40% discount code (good for all Manning products in all formats) for Datacast listeners: poddcast19. Link at http://mng.bz/4MAR.

Here is one free eBook code good for a copy of Algorithms and Data Structures for Massive Datasets for a lucky listener: algdcsr-7135. Link at http://mng.bz/Q2y6

View Details

Show Notes* (1:45) Willem discussed his undergraduate degree in Mechatronic Engineering at Stellenbosch University in the early 2010s. * (2:34) Willem recalled his entrepreneurial journey founding and selling a networking startup that provides internet access to private residents on campus. * (5:37) Willem worked for two years as a Software Engineer focusing on data systems at Systems Anywhere in Capetown after college. * (6:49) Willem talked about his move to Bangkok working as a Senior Software Engineer at INDEFF, a company in industrial control systems. * (9:52) Willem went over his decision to join Gojek, a leading Indonesian on-demand multi-service platform and digital payment technology group. * (12:16) Willem mentioned the engineering challenges associated with building complex data systems for super-apps. * (14:50) Willem dissected Gojek’s ML platform, including these four solutions for various stages of the ML life cycle: Clockwork, Merlin, Feast, and Turing. * (19:24) Willem recapped the lessons from designing the ML platform to meet Gojek’s scaling requirements — as delivered at Cloud Next 2018. * (23:09) Willem briefly went through the key design components to incorporate Kubeflow pipelines into Gojek’s existing ML platform — as delivered at KubeCon 2019. * (26:21) Willem explained the inception of Feast, an open-source feature store that bridges the gap between data and models. * (32:20) Willem talked about prioritizing the product roadmap and engaging the community for an open-source project. * (35:07) Willem recapped the key lessons learned and envisioned Feast's future to be a lightweight modular feature store. * (37:29) Willem explained the differences between commercial and open-source feature stores (given Tecton’s recent backing of Feast). * (41:36) Willem reflected on his experience living and working in Southeast Asia. * (44:33) Closing segment.

Willem’s Contact Info* Twitter * LinkedIn * GitHub

Mentioned ContentFeast

  • Feast Project website: feast.dev
  • Feast Slack community: #Feast
  • Feast Documentation: docs.feast.dev
  • Feast GitHub repository: feast-dev/feast
  • Feast on StackOverflow: stackoverflow.com/questions/tagged/feast
  • Feast Wiki: wiki.lfaidata.foundation/display/FEAST/Feast+Home
  • Feast Twitter: @feast_dev

Article

  • An Introduction to Gojek’s Machine Learning Platform (2019)
  • Introducing Feast: An Open-Source Feature Store For Machine Learning (2019)
  • A State of Feast (2020)
  • Why Tecton is Backing The Feast Open-Source Feature Store (2020)

Talks

  • Lessons Learned Scaling Machine Learning at GoJek on Google Cloud (Cloud Next 2018)
  • Accelerating Machine Learning App Development with Kubeflow Pipelines (Cloud Next 2019)
  • Moving People and Products with Machine Learning on Kubeflow (KubeCon 2019)

People

  • David Aronchick (Open-Source ML Strategy at Azure, Ex-PM for Kubernetes at Google, Co-Founder of Kubeflow, Advisor to Tecton)
  • Jeremy Lewi (Principal Engineer at Primer.ai, Co-Founder of Kubeflow)
  • Felipe Hoffa (Developer Advocate for BigQuery, Data Cloud Advocate for Snowflake)

Book

  • Cal Newport’s “Deep Work”

Willem will be a speaker at Tecton’s apply() virtual conference (April 21-22, 2021) for data and ML teams to discuss the practical data engineering challenges faced when building ML for the real world. Participants will share best practice development patterns, tools of choice, and emerging architectures they use to successfully build and manage production ML applications. Everything is on the table from managing labeling pipelines, to transforming features in real-time, and serving at scale. Register for free now: https://www.applyconf.com/!

View Details

Show Notes* (1:56) Jim went over his education at Trinity College Dublin in the late 90s/early 2000s, where he got early exposure to academic research in distributed systems. * (4:26) Jim discussed his research focused on dynamic software architecture, particularly the K-Component model that enables individual components to adapt to a changing environment. * (5:37) Jim explained his research on collaborative reinforcement learning that enables groups of reinforcement learning agents to solve online optimization problems in dynamic systems. * (9:03) Jim recalled his time as a Senior Consultant for MySQL. * (9:52) Jim shared the initiatives at the RISE Research Institute of Sweden, in which he has been a researcher since 2007. * (13:16) Jim dissected his peer-to-peer systems research at RISE, including theoretical results for search algorithm and walk topology. * (15:30) Jim went over challenges building peer-to-peer live streaming systems at RISE, such as GradientTV and Glive. * (18:18) Jim provided an overview of research activities at the Division of Software and Computer Systems at the School of Electrical Engineering and Computer Science at KTH Royal Institute of Technology. * (19:04) Jim has taught courses on Distributed Systems and Deep Learning on Big Data at KTH Royal Institute of Technology. * (22:20) Jim unpacked his O’Reilly article in 2017 called “Distributed TensorFlow,” which includes the deep learning hierarchy of scale. * (29:47) Jim discussed the development of HopsFS, a next-generation distribution of the Hadoop Distributed File System (HDFS) that replaces its single-node in-memory metadata service with a distributed metadata service built on a NewSQL database. * (34:17) Jim rationalized the intention to commercialize HopsFS and built Hopsworks, an user-friendly data science platform for Hops. * (36:56) Jim explored the relative benefits of public research money and VC-funded money. * (41:48) Jim unpacked the key ideas in his post “Feature Store: The Missing Data Layer in ML Pipelines.” * (47:31) Jim dissected the critical design that enables the Hopsworks feature store to refactor a monolithic end-to-end ML pipeline into separate feature engineering and model training pipelines. * (52:49) Jim explained why data warehouses are insufficient for machine learning pipelines and why a feature store is needed instead. * (57:59) Jim discussed prioritizing the product roadmap for the Hopswork platform. * (01:00:25) Jim hinted at what’s on the 2021 roadmap for Hopswork. * (01:03:22) Jim recalled the challenges of getting early customers for Hopsworks. * (01:04:30) Jim intuited the differences and similarities between being a professor and being a founder. * (01:07:00) Jim discussed worrying trends in the European Tech ecosystem and the role that Logical Clocks will play in the long run. * (01:13:37) Closing segment.

Jim’s Contact Info* Logical Clocks * Twitter * LinkedIn * Google Scholar * Medium * ACM Profile * GitHub

Mentioned ContentResearch Papers

  • “The K-Component Architecture Meta-Model for Self-Adaptive Software” (2001)
  • “Dynamic Software Evolution and The K-Component Model” (2001)
  • “Using feedback in collaborative reinforcement learning to adaptively optimize MANET routing” (2005)
  • “Building Autonomic Systems Using Collaborative Reinforcement Learning” (2006)
  • “Improving ICE Service Selection in a P2P System using the Gradient Topology” (2007)
  • “gradienTv: Market-Based P2P Live Media Streaming on the Gradient Overlay” (2010)
  • “GLive: The Gradient Overlay as a Market Maker for Mesh-Based P2P Live Streaming” (2011)
  • “HopsFS: Scaling Hierarchical File System Metadata Using NewSQL Databases” (2016)
  • “Scaling HDFS to More Than 1 Million Operations Per Second with HopsFS” (2017)
  • “Hopsworks: Improving User Experience and Development on Hadoop with Scalable, Strongly Consistent Metadata” (2017)
  • “Implicit Provenance for Machine Learning Artifacts” (2020)
  • “Time Travel and Provenance for Machine Learning Pipelines” (2020)
  • “Maggy: Scalable Asynchronous Parallel Hyperparameter Search” (2020)

Articles

  • “Distributed TensorFlow” (2017)
  • “Reflections on AWS’s S3 Architectural Flaws” (2017)
  • “Meet Michelangelo: Uber’s Machine Learning Platform” (2017)
  • “Feature Store: The Missing Data Layer in ML Pipelines” (2018)
  • “What Is Wrong With European Tech Companies?” (2019)
  • “ROI of Feature Stores” (2020)
  • “MLOps With A Feature Store” (2020)
  • “ML Engineer Guide: Feature Store vs. Data Warehouse” (2020)
  • “Unifying Single-Host and Distributed Machine Learning with Maggy” (2020)
  • “How We Secure Your Data With Hopsworks” (2020)
  • “One Function Is All You Need For ML Experiments” (2020)
  • “Hopsworks: World’s Only Cloud-Native Feature Store, now available on AWS and Azure” (2020)
  • “Hopsworks 2.0: The Next Generation Platform for Data-Intensive AI with a Feature Store” (2020)
  • “Hopsworks Feature Store API 2.0, a new paradigm” (2020)
  • “Swedish startup Logical Clocks takes a crack at scaling MySQL backend for live recommendations” (2021)

Projects

  • Apache Hudi (by Uber)
  • Delta Lake (by Databricks)
  • Apache Iceberg (by Netflix)
  • MLflow (by Databricks)
  • Apache Flink (by The Apache Foundation)

People

  • Leslie Lamport (The Father of Distributed Computing)
  • Jeff Dean (Creator of MapReduce and TensorFlow, Lead of Google AI)
  • Richard Sutton (The Father of Reinforcement Learning — who wrote “The Bitter Lesson”)

Programming Books

  • C++ Programming Languages books (by Scott Meyers)
  • “Effective Java” (by Joshua Bloch)
  • “Programming Erlang” (by Joe Armstrong)
  • “Concepts, Techniques, and Models of Computer Programming” (by Peter Van Roy and Seif Haridi)

View Details

Show Notes (2:20) Pier shared his college experience at the University of Southampton studying Electronic Engineering. * (3:46) For his final undergraduate project, Pier developed a suite of games and used machine learning to analyze brainwaves data that can classify whether a child is affected or not by autism. * (11:26) Pier went over his favorite courses and involvement with the AI Society during his additional year at the University of Southampton to get a Master’s in Artificial Intelligence. * (13:40) For his Master’s thesis called “Causal Reasoning in Machine Learning,” Pier created and deployed a suite of Agent-Based and Compartmental Models to simulate epidemic disease developments in different types of communities. * (26:51) Pier went over his stints as a developer intern at Fidessa and a freelance data scientist at Digital-Dandelion. * (29:21) Pier reflected on his time (so far) as a data scientist at SAS Institute, where he helps their customers solve various data-driven challenges using cloud-based technologies and DevOps processes. * (33:37) Pier discussed the key benefits that writing and editing technical content for Towards Data Science to his professional development. * (36:31) Pier covered the threads that he kept pulling with his blog posts. * (38:50) Pier talked about his Augmented Reality Personal Business Card created in HTML using the AR.js library. * (41:12) Pier brought up data structures in two other impressive JavaScript projects using TensorFlow.js and ml5.js. * (44:19) Pier went over his experience working with data visualization tools such as Plotly, R Shiny, and Streamlit. * (47:27) Pier talked about his work on a chapter for a book called “Applied Data Science in Tourism” that is going to be published with Springer this year. * (48:37) Pier shared his thoughts r*egarding the tech community in London. * (49:19) Closing segment.

Pier’s Contact Info* Website * LinkedIn * Twitter * GitHub * Medium * Patreon * Kaggle

Mentioned Content* “Alleviate Children’s Health Issues Through Games and Machine Learning” * “Causal Reasoning in Machine Learning” * Andrej Karpathy (Director of AI and Autopilot at Tesla) * Cassie Kozyrkov (Chief Decision Scientist at Google) * Iain Brown (Head of Data Science at SAS) * “The Book Of Why” (By Judea Pearl) * “Pattern Recognition and Machine Learning” (by Christopher Bishop)

View Details

Timestamps

  • (1:55) Alba shared her background growing up interested in studying Physics and pivoting into quantum mechanics.
  • (3:33) Alba went over her Bachelor’s in Fundamental Physics at The University of Barcelona.
  • (4:54) Alba continued her education with an M.S. degree that specialized in Particle Physics and Gravitation.
  • (6:40) Alba started her Ph.D. in Physics in 2015 and discussed her first publication, “Operational Approach to Bell Inequalities: Application to Qutrits.”
  • (9:48) Alba also spent time as a visiting scholar at the University of Oxford and the University of Madrid during her Ph.D.
  • (11:25) Alba explained her second paper to understand the connection between maximal entanglement and the fundamental symmetries of high-energy physics.
  • (13:27) Alba dissected her next work titled “Multipartite Entanglement in Spin Chains and The Hyperdeterminant.”
  • (18:56) Alba shared the origin of Quantic, a quantum computation joint effort between the University of Barcelona and the Barcelona Supercomputing Center.
  • (22:27) Alba unpacked her article “Quantum Computation: Playing The Quantum Symphony,” making a metaphor between quantum computing and musical symphony.
  • (27:47) Alba discussed the motivation and contribution of her paper “Exact Ising Model Simulation On A Quantum Computer.”
  • (32:51) Alba recalled creating a tutorial that ended up winning the Teach Me QISKit challenge from IBM back in 2018.
  • (35:01) Alba elaborated on her paper “Quantum Circuits For the Maximally Entangled States,” which designs a series of quantum circuits that generate absolute maximally entangled states to benchmark a quantum computer.
  • (38:54) Alba dissected key ideas in her paper “Data Re-Uploading For a Universal Quantum Classifier.”
  • (43:51) Alba explained how she leveled up her knowledge of classical neural networks.
  • (47:40) Alba shared her experience as a Postdoctoral Fellow at The Matter Lab at the University of Toronto — working on quantum machine learning and variational quantum algorithms (checked out the Quantum Research Seminars Toronto that she has been organizing).
  • (52:18) Alba explained her work on the Meta-Variational Quantum Eigensolver algorithm capable of learning the ground state energy profile of a parametrized Hamiltonian.
  • (59:23) Alba went over Tequila, a development package for quantum algorithms in Python that her group created.
  • (01:04:49) Alba presented a quantum calling for new algorithms, applications, architectures, quantum-classical interface, and more (as presented here).
  • (01:08:59) Alba has been active in education and public outreach activities about encouraging scientific vocations for young minds, especially in Catalonia.
  • (01:12:07) Closing segment.

Her Contact Info

  • Website
  • Twitter
  • LinkedIn
  • Google Scholar
  • GitHub

Her Recommended Resources

  • Ewin Tang (Ph.D. Student in Theoretical Computer Science at the University of Washington)
  • Alán Aspuru-Guzik (Professor of Chemistry and Computer Science at the University of Toronto, Alba’s current supervisor)
  • José Ignacio Latorre (Professor of Theoretical Physics at the University of Barcelona, Alba’s former supervisor)
  • Quantum Computation and Quantum Information (by Michael Nielsen and Isaac Chuang)
  • Quantum Field Theory and The Standard Model (by Matthew Schwarz)
  • The Structure of Scientific Revolutions (by Thomas Kuhn)
  • Against Method (by Paul Feyerabend)
  • Quantum Computing Since Democritus (by Scott Aaronson)

View Details

Timestamps* (2:07) JY discussed his college time studying Computer Science and Applied Math at Ecole Polytechnique — a leading French institute in science and technology. * (3:04) JY reflected on time at Stanford getting a Master’s in Management Science and Engineering, where he served as a Teaching Assistant for CS 229 (Machine Learning) and CS 246 (Mining Massive Datasets). * (6:14) JY walked over his ML engineering internship at LiveRamp — a data connectivity platform for the safe and effective use of data. * (7:54) JY reflected on his next three years at Databricks, first as a software engineer and then as a tech lead for the Spark Infrastructure team. * (10:00) JY unpacked the challenges of packaging/managing/monitoring Spark clusters and automating the launch of hundreds of thousands of nodes in the cloud every day. * (14:48) JY shared the founding story behind Data Mechanics, whose mission is to give superpowers to the world's data engineers so they can make sense of their data and build applications at scale on top of it. * (18:09) JY explained the three tenets of Data Mechanics: (1) managed and serverless, (2) integrated into clients’ workflows, and (3) built on top of open-source software (read the launch blog post). * (22:06) JY unpacked the core concepts of Spark-On-Kubernetes and evaluated the benefits/drawbacks of this new deployment mode — as presented in “Pros and Cons of Running Apache Spark on Kubernetes.” * (26:00) JY discussed Data Mechanics’ main improvements on the open-source version of Spark-On-Kubernetes — including an intuitive user interface, dynamic optimizations, integrations, and security — as explained in “Spark on Kubernetes Made Easy.” * (28:35) JY went over Data Mechanics Delight, a customized Spark UI which was recently open-sourced. * (35:40) JY shared the key ideas in his thought-leading piece on how to be successful with Apache Spark in 2021. * (38:42) JY went over his experience going through the Y Combinator program in summer 2019. * (40:56) JY reflected on the key decisions to get the first cohort of customers for Data Mechanics. * (42:26) JY shared valuable hiring lessons for early-stage startup founders. * (44:34) JY described the data and tech community in France. * (47:19) Closing segment.

His Contact Info

  • Twitter
  • LinkedIn
  • Data Mechanics

His Recommended Resources

  • Jure Leskovec (Associate Professor of Computer Science at Stanford / Chief Scientist at Pinterest)
  • Jeff Bezos (Founder of Amazon)
  • Matei Zaharia (CTO of Databricks and creator of Apache Spark)
  • “Designing For Data-Intensive Applications” (by Martin Kleppmann)

View Details

Timestamps* (2:55) Chris went over his experience studying Computer Science at the University of Southern California for undergraduate in the late 90s. * (5:26) Chris recalled working as a Software Engineer at NASA Jet Propulsion Lab in his sophomore year at USC. * (9:54) Chris continued his education at USC with an M.S. and then a Ph.D. in Computer Science. Under the guidance of Dr. Nenad Medvidović, his Ph.D. thesis is called “Software Connectors For Highly-Distributed And Voluminous Data-Intensive Systems.” He proposed DISCO, a software architecture-based systematic framework for selecting software connectors based on eight key dimensions of data distribution. * (16:28) Towards the end of his Ph.D., Chris started getting involved with the Apache Software Foundation. More specifically, he developed the original proposal and plan for Apache Tika (a content detection and analysis toolkit) in collaboration with Jérôme Charron to extract data in the Panama Papers, exposing how wealthy individuals exploited offshore tax regimes. * (24:58) Chris discussed his process of writing “Tika In Action,” which he co-authored with Jukka Zitting in 2011. * (27:01) Since 2007, Chris has been a professor in the Department of Computer Science at USC Viterbi School of Engineering. He went over the principles covered in his course titled “Software Architectures.” * (29:49) Chris touched on the core concepts and practical exercises that students could gain from his course “Information Retrieval and Web Search Engines.” * (32:10) Chris continued with his advanced course called “Content Detection and Analysis for Big Data” in recent years (check out this USC article). * (36:31) Chris also served as the Director of the USC’s Information Retrieval and Data Science group, whose mission is to research and develop new methodology and open source software to analyze, ingest, process, and manage Big Data and turn it into information. * (41:07) Chris unpacked the evolution of his career at NASA JPL: Member of Technical Staff -> Senior Software Architect -> Principal Data Scientist -> Deputy Chief Technology and Innovation Officer -> Division Manager for the AI, Analytics, and Innovation team. * (44:32) Chris dove deep into MEMEX — a JPL’s project that aims to develop software that advances online search capabilities to the deep web, the dark web, and nontraditional content. * (48:03) Chris briefly touched on XDATA — a JPL’s research effort to develop new computational techniques and open-source software tools to process and analyze big data. * (52:23) Chris described his work on the Object-Oriented Data Technology platform, an open-source data management system originally developed by NASA JPL and then donated to the Apache Software Foundation. * (55:22) Chris shared the scientific challenges and engineering requirements associated with developing the next generation of reusable science data processing systems for NASA’s Orbiting Carbon Observatory space mission and the Soil Moisture Active Passive earth science mission. * (01:01:05) Chris talked about his work on NASA’s Machine Learning-based Analytics for Autonomous Rover Systems — which consists of two novel capabilities for future Mars rovers (Drive-By Science and Energy-Optimal Autonomous Navigation). * (01:04:24) Chris quantified the Apache Software Foundation's impact on the software industry in the past decade and discussed trends in open-source software development. * (01:07:15) Chris unpacked his 2013 Nature article called “A vision for data science” — in which he argued that four advancements are necessary to get the best out of big data: algorithm integration, development and stewardship, diverse data formats, and people power. * (01:11:54) Chris revealed the challenges of writing the second edition of “Machine Learning with TensorFlow,” a technical book with Manning that teaches the foundational concepts of machine learning and the TensorFlow library's usage to build powerful models rapidly. * (01:15:04) Chris mentioned the differences between working in academia and industry. * (01:16:20) Chris described the tech and data community in the greater Los Angeles area. * (01:18:30) Closing segment.

His Contact Info* Wikipedia * NASA Page * Google Scholar * USC Page * Twitter * LinkedIn * GitHub

His Recommended Resources* Doug Cutting (Founder of Lucene and Hadoop) * Hilary Mason (Ex Data Scientist at bit.ly and Cloudera) * Jukka Zitting (Staff Software Engineer at Google) * "The One Minute Manager" (by Ken Blanchard and Spencer Johnson)

View Details

Show Notes (2:09) Marcello described his academic experience getting a Master’s Degree in Computer Science from the Universita di Catania in the early 2000s, where his thesis is called Evolutionary Randomized Graph Embedder. * (6:14) Marcello commented on his career phase working as a web developer across various places in Europe. * (9:18) Marcello discussed his time working as a software engineer at INPS, a government-owned company that now handles most Italian citizens' pubic-related data. * (10:42) Marcello talked about his time as a data visualization engineer at SwiftIQ. He created a data visualization library that allows the inclusion of dynamic charts in HTML pages with just a few JavaScript lines. * (13:40) Marcello went over his projects while working as a full-stack software engineer for Twitter’s User Services Engineering team in Dublin. * (17:19) Marcello reflected on his time at Microsoft Zurich’s Social and Engagement team, contributing to machine learning infrastructure and tools. * (21:28) Marcello briefly touched on his one-year stint at Apple Zurich as a Senior Applied Research Engineer. * (23:49) Marcello talked about the challenges while writing “Algorithms and Data Structures in Action,” which introduces a diverse range of algorithms used in web apps, systems programming, and data manipulation. * (27:11) Marcello expanded upon part 1 of the book, including advanced data structures such as D-ary Heaps, Randomized Treaps, Bloom Filters, Disjoint Sets, Tries/Radix Trees, and Cache. * (34:51) Marcello brought up data structures to perform efficient multi-dimensional queries, including various nearest neighbor searches and clustering techniques, in part 2 of the book. * (39:21) Marcello briefly described the algorithms in part 3 of the book — graph embeddings, gradient descent, simulated annealing, and genetic algorithms. * (48:28) Marcello talked about his work on jsgraph — a lightweight library to model graphs, run graphs algorithms, and display them on screen. * (52:06) Marcello compared Python, Java, and JavaScript programming languages. * (54:13) Marcello discussed his current interest in quantum computing. * (56:18) Marcello shared his thoughts r*egarding Dublin, Zurich, and Rome's tech communities. * (57:37) Closing segment.

His Contact Info* Twitter * LinkedIn * GitHub * Blog

His Recommended Resources* "Algorithms and Data Structures in Action" (Marcello's book with Manning) * Andrew Ng * Geoffrey Hinton * Francois Chollet * "Scalability Rules" (by Martin Abbott and Michael Fischer)

This is the 40% discount code that is good for all Manning’s products in all formats: poddcast19.

These are 5 free eBook codes, each good for one copy of “Algorithms and Data Structures in Action”:

  • adsdcr-5E76
  • adsdcr-EE51
  • adsdcr-DD47
  • adsdcr-B1BF
  • adsdcr-A61F

View Details

Show Notes* (2:10) Dave talked briefly about his Electrical Engineering study at Rensselaer Polytechnic Institute back in the late 90s. * (4:03) Dave commented on his career phase working as a software engineer across various companies in Bozeman, Montana. * (7:38) Dave discussed his work as a senior architect and tech lead at Expero, a Houston-based startup that develops custom software exclusively for domain-expert users. * (11:26) Dave briefly defined common big data frameworks (Hadoop, Apache Spark) and databases (Apache Cassandra, Apache Kafka). * (13:37) Dave went over the challenges during his time as a chief software architect at Gene by Gene, a biotech company focusing on DNA-based ancestry and genealogy. * (20:00) Dave shared the common patterns and anti-patterns of using graph databases (in reference to his talk “A Practical Guide to Graph Databases”). * (26:16) Dave walked through the three categories of graph technologies: Graph Computing Engine, RDF TripleStore, and Labeled Property Graph (in reference to his talk “A Skeptics Guide to Graph Databases”). * (33:03) Dave discussed his move to DataStax’s Global Graph Practice team as a solutions architect and graph database subject matter expert. * (36:00) Dave explained the design of DataStax’s enterprise solution called Customer 360, which collapses data silos to drive business value. * (41:16) Dave talked about his current experience as a Senior Graph Architect at AWS. * (43:51) Dave mentioned the challenges while writing "Graph Databases In Action" (published last October). * (47:25) Dave explained the open-source Apache TinkerPop framework and the Gremlin language used in the book for the uninitiated. * (51:04) Dave discussed trends in big data and distributed systems that he is most excited about. * (55:06) Closing segment.

His Contact Info* Website * Twitter * LinkedIn * GitHub

His Recommended Resources* "Graph Databases In Action" (Associated Code Repository) * Martin Fowler (Founder of ThoughtWorks) * Martin Kleppmann (Author of "Designing Data-Intensive Applications") * Andrew Ng (Professor at Stanford, Co-Founder of Google Brain and Coursera, Ex-Chief Scientist at Baidu) * "Pragmatic Programmer" (by Andy Hunt and Dave Thomas) * "The Five Dysfunctions Of A Team" (by Patrick Lencioni) * "How To Observe Scientific Advice for Common Real-World Problems" (by Randall Munroe)

This is the 40% discount code that is good for all Manning's products in all formats: poddcast19.

These are 5 free eBook codes, each good for one copy of "Graph Databases In Action":

  • gdadcr-E55F
  • gdadcr-B896
  • gdadcr-8C53
  • gdadcr-AAE1
  • gdadcr-39F0

View Details

Show Notes* (2:13) Jason went over his experience studying Computer Science at Loyola College in Baltimore for undergraduate, where he got an early exposure to academic research in image registration. * (4:31) Jason described his graduate school experience at John Hopkins University, where he completed his Ph.D. on “Techniques for Vision-Based Human-Computer Interaction” that proposed the Visual Interaction Cues paradigm. * (9:31) During his time as a Post-Doc Fellow at UCLA, Jason helped develop automatic segmentation and recognition techniques for brain tumors to improve the accuracy of diagnosis and treatment accuracy * (14:27) From 2007 to 2014, Jason was a professor in the Computer Science and Engineering department at SUNY-Buffalo. He covered the content of two graduate-level courses on Bayesian Vision and Intro to Pattern Recognition that he taught. * (18:20) On the topic of metric learning, Jason proposed an approach to data analysis and modeling for computer vision called "Active Clustering." * (21:35) On the topic of image understanding, Jason created Generalized Image Understanding - a project that examined a unified methodology that integrates low-, mid-, and high-level elements for visual inference (equivalent to image captioning today). * (24:51) On the topic of video understanding, Jason worked on ISTARE: Intelligent Spatio-Temporal Activity Reasoning Engine, whose objective is to represent, learn, recognize, and reason over activities in persistent surveillance videos. * (27:46) Jason dissected Action Bank - a high-level representation of activity in video, which comprises of many individual action detectors sampled broadly in semantic space and viewpoint space. * (35:30) Jason unpacked LIBSVX - a library of super voxel and video segmentation methods coupled with a principled evaluation benchmark based on quantitative 3D criteria for good super voxels. * (40:06) Jason gave an overview of AI research activities at the University of Michigan, where he was a professor of Electrical Engineering and Computer Science from 2014 to 2020. * (41:09) Jason covered the problems and projects in his graduate-level courses on Foundations of Computer Vision and Advanced Topics in Computer Vision at Michigan. * (44:56) Jason went over his recent research on video captioning and video description. * (47:03) Jason described his exciting software called BubbleNets, which chooses the best video frame for a human to annotate. * (51:44) Jason shared anecdotes of Voxel51's inception and key takeaways that he has learned. * (01:05:25) Jason talked about Voxel51's Physical Distancing Index that tracks the coronavirus global pandemic's impact on social behavior. * (01:07:47) Jason discussed his exciting new chapter as the new director of the Stevens Institute for Artificial Intelligence. * (01:11:28) Jason identified the differences and similarities between being a professor and being a founder. * (01:14:55) Jason gave his advice to individuals who want to make a dent in AI research. * (01:16:14) Jason mentioned the trends in computer vision research that he is most excited about at the moment. * (01:17:23) Closing segment.

His Contact Info* Wikipedia * Google Scholar * Website * Twitter * LinkedIn

His Recommended Resources* Bubblenets: Video Object Segmentation for Computer Vision * Voxel51's FiftyOne Open-Sourced Library * Jeff Siskind (Professor at Purdue University) * CJ Taylor (Professor at the University of Pennsylvania) * Kristen Grauman (Professor at the University of Austin) * "An Introduction to Mathematical Statistics"

View Details

Show Notes* (2:23) Barr discussed growing up in Israel and serving as a commander of the Data Analytics unit at the Israeli Air Force. * (4:10) Barr reflected on her college experience at Stanford studying Math and Computational Science. * (7:24) Barr walked over the two career lessons learned from being a Management Consultant at Bain and Company. * (9:51) Barr reflected on her time as VP of Customer Operations at Gainsight, which offers enterprise solutions for Customer Success and Product teams. She helped build and scale a global team covering various functions such as business operations, customer success, professional services. * (12:32) Barr unpacked the notion of data downtime, introduced in her blog post “The Rise of Data Downtime.” * (17:25) Barr unveiled the four main steps in the data reliability maturity curve: reactive, proactive, automated, and scalable - as indicated in “Closing The Data Downtime Gap." * (21:09) Barr shared the founding story behind Monte Carlo, whose mission is to accelerate the world’s adoption of data by reducing data downtime. * (24:29) Barr explained the five pillars of data observability. * (27:45) Barr unpacked the rise of data catalogs as a powerful tool for data governance, along with the three categories of data catalog solutions that data teams are adopting - as presented in “What We Got Wrong About Data Governance.” * (31:32) Barr discussed the benefits of using Data Mesh - a type of data platform architecture that embraces data ubiquity in the enterprise by leveraging a domain-oriented, self-serve design. * (37:28) Barr went over a framework that looks at the business functions and the nature of the work to score the impact and allocate the ROI of the data team - as proposed in "Measuring the ROI of Your Data Organization." * (40:39) Barr shared five practices for designing a platform that maximizes data's value and impact inside an organization. * (43:27) Barr reflected on the key decisions to get the first cohort of customers for Monte Carlo. * (46:31) Barr shared valuable hiring lessons. * (48:48) Barr went over helpful resources throughout her journey as a founder. * (50:13) Barr dropped the final advice for founders on seeking the right investors. * (51:18) Closing segment.

Her Contact Info* Twitter * LinkedIn * Medium * Monte Carlo

Her Recommended Resources* Stanford's "Mathematics and Magic Tricks" course (taught by Persi Diaconis) * "The Biggest Bluff" by Maria Konnikova * Snowflake * DJ Patil (Former U.S. Chief Data Scientist) * Monte Carlo's Customers

View Details

Show Notes* (1:57) Carl recalled his undergraduate experience studying Electrical Engineering at Stanford back in the early 90s. * (3:58) Carl recalled his graduate experience pursuing Master’s degrees in Computer Science at NYU and King’s College in the late 90s. For his Master's Thesis, he investigated Support Vector Machines with a Bayesian algorithm programmed in C. * (6:45) Carl walked over his Ph.D. work in Computation and Neural Systems at CalTech, where he did a thesis on Biophysics of Extracellular Action Potentials. * (13:11) Carl provided brief thoughts about his experience working as a business analyst and consultant for HBO during his Ph.D. period. * (14:55) Carl went over his rationale behind his decision to move from academic neuroscience to quantitative finance. * (19:19) Carl discussed his proudest accomplishments and valuable lessons learned from spending seven years at Morgan Stanley Capital International and rising to a leadership role as Vice President of Risk Modeling. * (23:17) Carl uncovered his move to San Francisco to work as a lead data scientist at Sparked back in 2014, which builds a customer success SaaS solution. * (27:10) Adding to his move to Zuora in 2015, Carl explained how the subscription business model works in layman terms. * (31:44) Carl unpacked the common patterns that he saw from analyzing subscriber churn for companies across industries due to his work on Zuora Analytics. * (33:30) Carl shared the process of creating the Subscription Economy Index, Zuora’s landmark index tracking the rapid ascent of the Subscription Economy, and distilled the key trends of the 2020 edition. * (39:59) Carl unpacked the three reasons that make churn hard to fight: (1) Churn is hard to predict, (2) Churn is harder to prevent, and (3) Churn requires a multi-team effort (Watch his talks at the 2019 Data Council San Francisco and the 2020 Subscribed Online Conference). * (44:46) Carl shared advice for data scientists who want to collaborate more effectively with other functional departments. * (46:30) Carl emphasized the importance of creating great customer metrics, which are ratios of basic behavioral metrics to fight churn effectively. * (53:49) Carl went over the challenges of writing “Fighting Churn With Data,” which provides a clear overview of churn concepts, along with hands-on tricks and tips developed through years of experience analyzing customer behavior. * (55:53) Carl reflected on how his academic background in computational neuroscience contributes to his success as a quant analyst and a data scientist. * (59:37) Carl compared his experience living and working across Los Angeles, New York, and San Francisco. * (01:02:04) Closing segment.

His Contact Info* LinkedIn * Twitter * GitHub * Google Scholar * Medium

His Recommended Resources* “Fighting Churn With Data” by Carl Gold * Konrad Kording (Professor of Computational Neuroscience at the University of Pennsylvania) * Kate Crawford (Distinguished Research Professor in Tech, Culture, and Society at New York University) * Cassie Kozyrkov (Chief Decision Scientist at Google) * "Freakonomics" by Stephen Dubner and Stephen Levitt * Carl's other podcast appearances

Here are the discount codes that you can use to purchase "Fighting Churn with Data" with 40% off:

  • fcddcr-6D84
  • fcddcr-4AE7
  • fcddcr-6D9C
  • fcddcr-30BF
  • fcddcr-9705

View Details

Show Notes

  • (2:08) Jess discussed her foray into studying Software Engineering at California Polytechnic State University during college and revealed her favorite course on Computer Science Ethics taken there.
  • (4:31) Jess unpacked her argument that it is important to shift the engineering mindset away from only asking how to ask why - referring to his blog post “Changing The Engineer’s Mindset.”
  • (7:27) Jess went over her summer internship experience at GoDaddy as a software engineer.
  • (11:39) Jess talked about her time working as a research assistant for the Ethics and Emerging Sciences Group at Cal Poly, where she examined the ethical implications of AI “predictive policing” systems and survey the current role of fairness metrics for battling algorithmic bias.
  • (16:27) Jess revealed her experience being involved with the open data movement in Colombia (read her articles “The Truth About Open Data” and “How To Use Data Science For Social Impact”).
  • (24:22) Jess emphasized the importance of education to spread data literacy in developing nations.
  • (26:35) Jess discussed her experience as a current Ph.D. student in the Department of Information Science at the University of Colorado, Boulder, where you focus on value tradeoffs in technology and machine learning ethics.
  • (32:01) Jess unpacked the ETHItechniCAL framework to assist with ethical decision-making that she proposes in “The Trolley Problem Isn’t Theoretical Anymore.”
  • (35:39) Jess unpacked her argument, saying that computer scientists must be educated to code with social responsibility and equipped with the correct tools to do so - as indicated in “How Tech Shapes Society.”
  • (39:00) Jess discussed the work “Investigating Potential Factors Associated with Gender Discrimination in Collaborative Recommender Systems” with Masoud Mansoury and Himan Abdollahpouri.
  • (42:54) Jess discussed the work “Exploring User Opinions of Fairness in Recommender Systems” with Nasim Sonboli.
  • (47:12) Via her podcast The Radical AI, Jess unpacked the underrated AI and social issues that she came across.
  • (49:17) Via her YouTube show Sci-Fi in Real Life, Jess shared her 3 favorite videos: "Dying To Be Alive," "Living On The Edge," and "Black Mirror Meta Episode."
  • (52:25) Jess dug deep into her mission of cultivating positive social impacts for the world.
  • (54:32) Closing segment.

Her Contact Info

  • Website
  • Twitter
  • Medium
  • LinkedIn
  • GitHub
  • Radical AI Podcast
  • Sci-Fi In Real Life YouTube Show

Her Recommended Resources

  • UC Boulder's Internet Rules Lab
  • UC Boulder's That Recommender Systems Lab
  • Safiya Noble
  • Cathy O'Neil
  • Ruha Benjamin
  • "The Courage To Be Disliked" by Ichiro Kishimi and Fumitake Koga

View Details

Show Notes

  • (2:12) Luis shared how he got excited about learning mathematics and specialized in combinatorics.
  • (4:26) Luis discussed his experience studying Math for his Bachelor’s and Master’s degrees at the University of Waterloo - where he took many courses in combinatorics and engaged in undergraduate research.
  • (5:59) Luis pursued his Ph.D. in Mathematics at the University of Michigan - where he worked on Schubert Calculus that intersects combinatorics and geometry (check out his Ph.D. dissertation).
  • (8:45) Luis distinguished the differences between doing research in mathematics and machine learning.
  • (11:33) Luis went over his time as a Postdoc Fellow and Lecturer at the University of Quebec at Montreal - where he was a member of the LaCIM lab (whose areas of research originating in Combinatorics and its relationships to Algebra and Computer Science) and taught classes in French.
  • (13:47) Luis explained why he left academia and got his job as a Machine Learning Engineer at Google in 2014.
  • (16:33) Luis discussed the engineering and analytical challenges he encountered as part of the video recommendations team at YouTube.
  • (19:58) Luis shared lessons he learned to transition from academia to industry.
  • (22:25) Luis went over his move to become the Head of Content for AI and Data Science at Udacity, alongside his online education passion.
  • (26:08) Luis explained Udacity's educational approach to course content design in various nano degree programs, including Machine Learning, Deep Learning, and Data Science.
  • (28:46) Luis unpacked his end-to-end process of making YouTube, where he teaches concepts in Machine Learning and Math in layman terms.
  • (31:01) Luis unpacked his statement, "Humans are bad at abstraction, but great at math," from his video “You Are Much Better At Math Than You Think.”
  • (34:46) Luis shared his 3 favorite Machine Learning videos: Restricted Boltzmann Machines, A Friendly Introduction to Machine Learning, and My Story with the Thue-Morse Sequence.
  • (37:18) Luis discussed the data science culture at Apple, where he spent one-year teaching machine learning to the employees and doing internal consulting in AI-related projects.
  • (39:06) Luis revealed his interest in quantum computing. He works as a Quantum AI Research Scientist at Zapata Computing, a quantum software company that offers computing solutions for industrial and commercial use.
  • (43:19) Luis mentioned the challenges of writing “Grokking Machine Learning” - a technical book with Manning planned to be published next year - like a mystery novel.
  • (46:12) Luis shared the differences between working in Silicon Valley and Canada.
  • (47:50) Closing segment.

His Contact Info

  • Website
  • Twitter
  • LinkedIn
  • YouTube
  • GitHub
  • Google Scholar
  • Medium

His Recommended Resources

  • Sebastian Thrun
  • Andrew Ng
  • Rana el Kaliouby
  • "How Not To Be Wrong: The Power of Mathematical Thinking" by Jordan Ellenberg
  • "Weapons of Math Destruction" by Cathy O'Neil

Here are the codes for free eBook copies of Luis' book "Grokking Machine Learning": gmldcr-D659, gmldcr-2512, gmldcr-0752, gmldcr-30A2, gmldcr-01E8. Additionally, use the code poddcast19 to receive a 40% discount of all Manning products!

View Details

Show Notes

  • (2:05) Frank reflected on his undergraduate experience studying Electrical Engineering at the University of Massachusetts - Dartmouth.
  • (3:33) Frank commented on his experience working in the game industry after school.
  • (6:28) Frank went over the opportunity to work as a software engineer at Amazon, where he contributed to the personalization system that recommends products to customers at a scale of tens of thousands of requests per second.
  • (8:44) Frank brought up the challenges of building Amazon’s recommendation systems back in the early days.
  • (10:14) Frank discussed how Amazon’s recommendations and content optimization technology evolved incrementally during his time as a Senior Manager.
  • (12:05) Frank touched on the core engineering challenges during his time as a Senior Manager of Technology at IMDB.
  • (14:19) Frank spoke about his proudest accomplishments at Amazon, both from the technical and the management perspectives.
  • (18:19) Frank shared the story behind his professional transition into self-employment (check out his book “Self-Employment: Building an Internet Business of One”).
  • (24:07) Frank shared a brief overview of his business (Sundog Software)'s virtual reality products.
  • (25:15) Frank shared how he came to be an online instructor, discussed the pros/cons, and gave advice for aspiring ones.
  • (29:38) Frank has created various courses that focus on Apache Spark, ranging from Python and Scala support to Spark Streaming capability.
  • (31:34) Frank discussed how the Hadoop ecosystem has fallen out of favor (check out his popular Udemy courses titled “The Ultimate Hands-On Hadoop”).
  • (33:20) Frank touched on ElasticSearch - an industry-standard open-source search engine (check out his Manning live videos on ElasticSearch 6 and ElasticSearch 7).
  • (37:08) Frank provided his perspectives on the current landscape of recommendation systems research and applications.
  • (42:17) Frank advised scientists and engineers on how to communicate with non-technical colleagues effectively.
  • (43:25) Closing segment.

His Contact Info

  • Website
  • LinkedIn
  • Twitter
  • Facebook
  • YouTube

His Recommended Resources

  • Amazon Leadership Principles
  • Sundog Software
  • “Self-Employment: Building an Internet Business of One”
  • “Building Recommender Systems with Machine Learning and AI"
  • Jose Portilla
  • Kirill Eremenko
  • Andrew Ng
  • "Lean Startup" by Eric Ries
  • "Architecting Modern Data Platforms" by Jan Kunigk, Ian Buss, Paul Wilkinson, Lars George

Use the codes below to get a discount from Frank's live video course on Manning called "Machine Learning, Data Science and Deep Learning with Python":

  • mldldcr-4DB2
  • mldldcr-9FE8
  • mldldcr-EA35

View Details

Show Notes

  • (2:02) Amita described her educational background, studying Electronics at universities back in the 90s. She also professed her love for Asimov’s writings.
  • (3:33) Amita talked about her reason to pursue a path of an academic professor.
  • (5:13) Amita discussed her Ph.D. titled "Modeling, Design, and Applications of Optical Amplifiers and Long Period Gatings” at the University of Dehli and Karlsruhe Institute of Technology
  • (8:38) Amita shared her opinions on how the education of neural networks has evolved in the last 20 years of her teaching career - including the programming language shift from using Fortran and C++ to Python and the importance of learning computer networking and operating systems.
  • (14:29) Amita discussed her research that combines the concepts of social network analysis and neural networks to model user behavior in society (read the full paper here).
  • (17:57) Amita talked about the process of writing the TensorFlow 1.x Deep Learning Cookbook with Antonio Gulli.
  • (21:08) Amita went over TensorFlow Machine Learning Projects, co-authored alongside Ankit Jain and Armando Fandango.
  • (23:19) Amita dived into Hands-On Artificial Intelligence for IoT - which discusses different AI techniques to build smart IoT systems, covering practical case studies in personal & home devices, industrial applications, and smart cities.
  • (27:50) Amita explained the improvements in TensorFlow 2.0 from its previous version, referring to her book Deep Learning with TensorFlow 2 and Keras - 2nd Edition in collaboration with Antonio Gulli and Sujit Pal.
  • (31:33) Amita went over her experience participating in the NASA Centennial Space Robotics Challenge in 2017, in which her team finished in the top 20 out of more than 100 teams worldwide.
  • (34:49) Amita reflected on her volunteering experience with a group of friends to build an Acute Myeloid Leukemia detection system that won an award for the Intel showcase in 2019.
  • (38:21) Amita unpacked her blog post looking at COVID19 from a data science perspective.
  • (40:26) Amita described her mentoring work at Neuromatch Academy, a non-profit online course in computational neuroscience.
  • (44:19) Amita shared her opinion on the benefits of an online classroom versus an in-person classroom.
  • (51:19) Amita expressed her thoughts on the tech and data community in New Dehli.
  • (52:24) Amita shared her hobby of writing science fiction stories.
  • (53:41) Closing segment.

Her Contact Info

  • Website
  • Google Scholar
  • LinkedIn
  • Twitter
  • Github

Her Recommended Resources

  • "TensorFlow 1.x Deep Learning Cookbook."
  • "TensorFlow Machine Learning Projects"
  • "Hands-On Artificial Intelligence for IoT"
  • "Deep Learning with TensorFlow 2 and Keras - 2nd Edition."
  • TensorFlow Strategy for Distributed Training
  • Neuromatch Academy Course Content
  • Computational Neuroscience Coursera Course
  • Alan Turing
  • JJ Hopfield
  • Geoffrey Hinton
  • “The Theory of Everything” by Stephen Hawking

View Details

Show Notes

  • (2:02) Shreya discussed her initial exposure to Computer Science and her favorite CS course on Advanced Topics in Operating Systems at Stanford.
  • (4:07) Shreya emphasized the importance of distilling technical concepts to a non-technical audience, thanks to her experience as a section leader and teaching assistant for CS198.
  • (6:26) Shreya shared the lack of representation in technical roles that keep women away from considering technology as a career path, and the initiative she was involved with at SHE++.
  • (9:40) Shreya reflected on her software engineering internship experience at Facebook, working on Civic Engagement tools to help representatives connect with their constituents.
  • (12:33) Shreya went over the anecdote of how she worked on Machine Learning Security research at Google Brain.
  • (15:36) Shreya unpacked the paper “Adversarial Examples That Fool Both Computer Vision and Time-Limited Humans,” - where her team constructs adversarial examples that transfer computer vision models to the human visual system.
  • (20:08) Shreya reflected on the lessons learned from her experience working with seasoned researchers at Google Brain.
  • (23:31) Shreya gave her advice for engineers who are interested in multiple specializations.
  • (25:34) Shreya provided resources on the fundamentals of computer systems.
  • (27:15) Shreya explained her reason to work at an early-stage startup right after college (check out the blog post on her decision-making process).
  • (28:41) Shreya was the first ML Engineer at Viaduct, a startup that develops end-to-end machine learning and data analytics platform to empower OEMs to manage, analyze, and utilize their connected vehicle data.
  • (32:27) Shreya discussed two common misconceptions people have about the differences between machine learning in research and practice (read her reflection on one-year of making ML actually useful).
  • (35:24) Shreya expanded on the organizational silo challenge that hinders collaboration between data scientists and software engineers while designing a machine learning product.
  • (40:48) Shreya has been quite open about the challenge of recruiting female engineers, explaining that it is hard to sell women candidates when their alternatives are “conventionally sexy."
  • (47:24) Shreya and a few others have developed and open-sourced GPT-3 Sandbox, a library that helps users get started with the GPT-3 API.
  • (51:52) Shreya explained her prediction on why OpenAI can be the AWS of modeling.
  • (54:24) Shreya shared the benefits of going to therapy to cope with mental illness challenges.
  • (58:36) Closing segment.

Her Contact Info

  • Website
  • Twitter
  • LinkedIn
  • GitHub
  • Google Scholar
  • Medium

Her Recommended Resources

  • Martin Kleppmann’s “Designing Data-Intensive Applications”
  • Stanford’s CS110 - “Principles of Computer Systems"
  • Ada Lovelace
  • Women in AI
  • Black in AI
  • Quoc Le
  • Uber Engineering Blog
  • Steve Krug’s “Don’t Make Me Think"

View Details

Show Notes

  • (2:37) Francesca discussed her educational background in Italy, studying Economics and Institutional Studies at LUISS Guido Carli University for her Master’s and then Economics and Technology Innovation at Sant’Anna University for her Ph.D. She also mentioned her transition to studying in the US at Harvard Business School.
  • (7:43) Francesca shared the anecdote behind going to HBS to pursue a Postdoc Research Fellowship in Economics. She also revealed the differences in the educational approaches between Italy and the United States.
  • (15:15) During her Postdoc, Francesca worked on multiple patent data-driven projects to investigate and measure the impact of external knowledge networks on companies’ competitiveness and innovation. She discussed a specific project that analyzed biotech innovation in Boston, San Diego, and San Francisco clusters using social media and citation data.
  • (24:26) Francesca talked about her decision to join Microsoft as a data scientist in its Cloud and Enterprise division back in 2014, where she first worked on projects for clients from the energy and finance sectors.
  • (30:00) Francesca discussed the two types of customers who seek Microsoft’s cloud solutions to solve their data problems and explained the learning curves she went through while interacting with them.
  • (36:11) Francesca unpacked the Healthy Data Science Organization Framework - which is a portfolio of methodologies, technologies, resources that will assist organizations in becoming more data-driven (Read her InfoQ article “The Data Science Mindset: 6 Principles to Build Healthy Data-Driven Organizations”).
  • (45:31) Francesca shared the challenges of building end-to-end machine learning applications that she has observed from Microsoft Azure AI’s clients.
  • (49:56) Francesca walked through a typical day in her current leadership role at Microsoft’s Cloud AI Advocates team.
  • (53:44) Francesca discussed the different components in a typical Azure deployment workflow (Read her post “Azure Machine Learning Deployment Workflow”).
  • (58:44) Francesca explained Automated Machine Learning, a breakthrough from Microsoft Research division that is essentially a recommender system for machine learning pipelines.
  • (01:03:50) Francesca went over model interpretability features within Azure AI (as part of the InterpretML package) and touched on Microsoft’s Responsible AI principles.
  • (01:08:01) Francesca explained the differences between model fairness and model interpretability at both the training time and inference time (Check out the Fairlearn package).
  • (01:12:11) Francesca is currently writing a book with Wiley called “Machine Learning for Time Series Forecasting with Python.”
  • (01:14:39) Francesca shared her advice for undergraduate students looking to get into the field, judging from her experience being a mentor for Ph.D. and Postdoc students at institutions such as Harvard, MIT, and Columbia.
  • (01:17:27) Francesca reasoned how her educational backgrounds in economics and operations management contribute to her success in a data science career
  • (01:20:09) Closing segment.

Her Contact Info

  • Twitter
  • Medium
  • LinkedIn

Her Recommended Resources

People To Follow

  • Hilary Mason
  • Andrew Ng
  • Hannah Wallach

Book To Read

  • An Introduction to Probability Theory and Its Applications (by William Feller)

A Developer’s Introduction to Data Science

  • Video series on Data Science and Machine Learning on Azure
  • Video series on Data Science and Machine Learning on Azure GitHub repo

Azure Machine Learning

  • Azure Machine Learning Documentation
  • Azure Machine Learning Service
  • The Data Science Lifecycle
  • Algorithm Cheat Sheet
  • How to Select Machine Learning Algorithms
  • Azure Machine Learning Designer

Responsible Machine Learning

  • Responsible Machine Learning
  • Model Interpretability
  • InterpretML Repo
  • InterpretML Toolkit
  • InterpretML Documentation
  • Fairlearn Service
  • Fairlearn Documentation

Automated Machine Learning

  • Automated Machine Learning
  • Auto ML Featurization
  • AutoML Config Class

View Details

Show Notes

  • (2:55) Patricia talked about his interest in learning languages and living in different cultures.
  • (4:05) Patricia talked about her experience volunteering as a translator at the International Network of Street Papers.
  • (5:00) Patricia studied Liberal Arts at John Abbott College, English Literature at Concordia University, and Computer Science and Linguistics at McGill University during her undergraduate years.
  • (8:06) Patricia worked at McGill Language Development Lab as a Research Assistant, which studied how children learn different types of words and sentences.
  • (9:15) Patricia described her graduate school experience at the University of Toronto, where she researched lost language decipherment and writing systems.
  • (11:19) Patricia talked about MedStory, which is a text-oriented visual prototype built to support the complexity of medical narratives (spearheaded by Nicole Sultanum).
  • (12:35) Patricia explained her research paper, “Vowel and Consonant Classification through Spectral Decomposition.”
  • (15:29) Patricia unpacked her blog post, “Why is Privacy-Preserving NLP Important?”
  • (19:02) Patricia dissected her paper “Privacy-Preserving Character Language Modelling” that proposes a method for calculating character bigram and trigram probabilities over sensitive data using homomorphic encryption.
  • (21:13) Patricia wrote a two-part series called “Homomorphic Encryption for Beginners.”
  • (22:21) Patricia unwrapped her paper “Efficient Evaluation of Activation Functions over Encrypted Data” that shows how to represent the value of any function over a defined and bounded interval, given encrypted input data, without needing to decrypt any intermediate values before obtaining the function’s output.
  • (25:33) Patricia elaborated on her paper “Extracting Bark-Frequency Cepstral Coefficients from Encrypted Signals,” which claims that extracting spectral features from encrypted signals is the first step towards achieving secure end-to-end automatic speech recognition over encrypted data.
  • (27:38) Patricia explained why privacy is an essential attribute for speech recognition applications.
  • (29:53) Patricia discussed her comprehensive guide on “Perfectly Privacy-Preserving AI” which dives into the four pillars of perfectly privacy-preserving AI and outlines potential combinatorial solutions to satisfy all four pillars.
  • (37:53) Patricia shared her take on the differences working in academic and commercial settings (she is the founder and CEO of Private AI).
  • (40:50) Patricia talked about Private AI’s GALATEA Anonymization Suite, which anonymizes data at the source and encrypts them using quantum-safe cryptography.
  • (45:05) Patricia emphasized the importance of talking to customers when building a commercial product.
  • (46:58) Patricia shared her experience as a Postgraduate Affiliate at Vector Institute, which works with institutions, industry, startups, incubators, and accelerators to advance AI research and drive its application, adoption, and commercialization across Canada.
  • (49:09) Patricia shared her advice for young researchers by going deep into at least two domains and combining the knowledge.
  • (50:30) Patricia shared her excitement for privacy and NLP research in the upcoming years.
  • (52:36) Closing segment.

Her Contact Info

  • Website
  • Twitter
  • LinkedIn
  • Google Scholar
  • Medium
  • GitHub

Her Recommended Resources

  • Homomorphic Encryption
  • Secure Multiparty Computation
  • Federated Learning
  • Differential Privacy
  • Vector Institute
  • MILA Montreal Institute
  • Alberta Machine Intelligence Institute
  • Reza Shokri (Assistant Professor at National University of Singapore)
  • Parinaz Sobhani (Director of Machine Learning at Georgian Partners)
  • Doina Precup (Associate Professor at McGill University)

View Details

Show Notes

  • (2:19) Eugene got his Bachelor’s degree in Psychology and Organizational Behavior from Singapore Management University, in which he did a senior thesis titled “Competition Improves Performance.”
  • (3:29) Eugene’s first role out of school is an Investment Analyst position at Singapore’s Ministry of Trade & Industry.
  • (4:18) Eugene then moved to a Data Analyst role at IBM, working on projects such as supply-chain dashboards, social media analytics, and anti-money laundering detection.
  • (5:55) Eugene transitioned to an internal Data Scientist role at IBM, working on job forecasting and job recommendations.
  • (9:03) Eugene shared the story of how he became a Data Scientist at Lazada Group, which was a small e-commerce startup back in 2015.
  • (12:08) Eugene explained his decision to go back to school and pursued an online Master’s degree in Computer Science at Georgia Tech.
  • (19:14) Eugene shared his career milestones, as displayed in his blog post reflecting on his journey from getting a degree in Psychology to leading data science at Lazada.
  • (22:17) Eugene discussed the unique data science challenges while working at uCare.ai - a startup that aims to make healthcare more efficient and reduce costs.
  • (25:29) Eugene revealed three useful tips to deliver great data science talks (read his blog post “How to Give a Kick-Ass Data Science Talk” for the details).
  • (28:29) Eugene talked about his transition to become an Applied Scientist at Amazon - working on Amazon Kindle.
  • (30:43) Eugene unpacked his post “Commando, Soldier, Police, and Your Career Choices” that provides an interesting metaphor to help guide career decisions.
  • (33:43) Eugene went meta onto his writing process (read here) and note-taking strategy (read here).
  • (39:01) Eugene shared the lessons learned from taking on responsibilities in hiring, mentoring, and stakeholder engagement in his second year at Lazada (read his blog post on the first 100 days as a Data Science Lead).
  • (44:20) Eugene went in-depth into the engineering and cultural challenges throughout Alibaba Group’s acquisition of Lazada Group.
  • (47:51) Eugene explained Alibaba’s playbook for the technical integration of their acquisitions and the super-apps phenomenon in Asia (check out a summary of his talk on Asia’s Tech Giants).
  • (53:52) Eugene unpacked the values and essential aspects of Lazada’s data science team culture, as detailed in his post “Building a Strong Data Science Team Culture.”
  • (57:44) Eugene summarized his thoughts on the topic of data science and agile/scrum development (Read his 3-part blog series: Part 1, Part 2, and Part 3).
  • (01:03:18) Eugene was heavily involved with the development of product ranking, product recommendations, and product classification models in his first year at Lazada (check out slides to his talk “How Lazada Ranks Products”).
  • (01:09:08) Eugene helped mentor and empower teams on multiple machine learning systems while acting as VP of Data Science at Lazada (check out slides to his talk “Data Science Challenges at Lazada”).
  • (01:12:07) Eugene shared the case study of how uCare.ai developed a machine learning system for Southeast Asia’s largest healthcare group that estimates a patient’s total bill at the point of pre-admission.
  • (01:14:06) Eugene summarized his 2-part series that exposes the challenges after model deployment and yields a practical guide to maintaining models in production.
  • (01:19:04) Eugene discussed his early-career Product Classification project that uses a public Amazon dataset and builds two APIs for image classification & image search.
  • (01:22:29) Eugene discussed his 2-part series that implements seven different models on the same Amazon dataset, from matrix factorization to graphs and NLP.
  • (01:24:42) Closing segment.

His Contact Info

  • Website
  • Twitter
  • LinkedIn
  • GitHub

His Recommended Resources

  • Niklas Luhmann (well-known German sociologist)
  • Roam Research (note-taking application)
  • MLflow (A platform for ML lifecycle management)
  • Amazon Product Review Dataset (big data in JSON format)
  • Andrej Karpathy (Read “The Unreasonable Effectiveness of RNNs” and “A Recipe For Training Neural Networks”)
  • Jeremy Howard (Read the “Universal Language Model Fine-tuning for Text Classification paper)
  • Hamel Hussain (Check out GitHub Actions and fastpages)
  • “Introduction to Statistical Learning” (by Trevor Hastie and Rob Tibshirani)
  • “The Pragmatic Programmer” (by Andy Hunt and Dave Thomas)
  • applied-ml repository
  • ml-survey repository

View Details

Show Notes

  • (2:22) Matthew shared his childhood growing up interested in the field of biology.
  • (5:29) Matthew described his undergraduate experience studying Cellular and Molecular Biology at Brown University. He dropped out for a year and a half to work at MIT and test out a few company ideas in the biotech space.
  • (8:13) Matthew spent a decent amount of time in biological aging research after that, working at the Karp Lab at MIT and the Backsai Lab in Massachusetts General Hospital.
  • (13:28) Matthew recalled the story of how he switched his pursuit to a career in Machine Learning.
  • (17:14) Matthew commented on his experience as a Machine Learning Engineer freelancer on various projects in privacy and security, music analysis, and secure communications.
  • (20:36) Matthew discussed the opportunity to work with Google as a contract software developer and shared valuable lessons from contributing to the TensorFlow Probability library for probabilistic reasoning and statistical analysis.
  • (23:48) Matthew gave a quick overview of Bayesian Neural Networks (read his blog post for more details).
  • (27:18) Matthew went over his contribution to the open-source community OpenMined, whose goal is to make the world more privacy-preserving by lowering the barrier-to-entry to private AI technologies.
  • (32:29) Matthew worked on De-Moloch in late 2018, described to be "software that lets anyone easily run AI algorithms on sensitive data without it being personally identifiable" (read his blog post "Private ML Explained in 5 Levels of Complexity" for a complete description).
  • (36:17) Matthew unpacked his post "Private ML Marketplaces," - which summarizes and discusses various approaches previously proposed in this space, such as smart contracts, data encryption/transformation/approximation, and federated learning.
  • (39:45) Matthew shared his experience competing in the Pioneer Tournament.
  • (42:19) Matthew shared brief advice on how to become a Machine Learning Engineer. For the full details, read his mega-post "Lessons from becoming an ML engineer in 12 months, without a CS or Math degree."
  • (45:16) Matthew described his experience working as a Machine Learning Engineer at UnifyID, a startup that is building a revolutionary identity platform based on implicit passwordless authentication.
  • (47:52) Matthew unpacked his research paper "Model Weight Theft with Just Noise Inputs: The Curious Case of the Petulant Attacker" at UnifyID. The paper explores the scenarios under which an attacker can steal the weights of a convolutional neural network whose architecture is already known.
  • (51:55) Matthew is currently doing research with FOR.ai, a multi-disciplinary team of scientists and engineers who like researching for fun.
  • (54:14) Matthew unpacked his research at FOR.ai, namely "Optimal Brain Damage" and "BitTensor: An Intermodel Intelligence Measure."
  • (01:00:52) Matthew shared key takeaways from attending academic conferences such as ICML 2019 and NeurIPS 2019.
  • (01:03:45) Matthew unpacked his 4-part series on ML Research interview that targets aspiring ML engineers, hiring managers/senior ML engineers, and people navigating ML research that don't want to lose sight of first principles.
  • (01:07:09) Matthew unpacked his fantastic post called "Nitpicking ML Technical Debt" that breaks down relevant points of Google's famous paper on Hidden Technical Debt.
  • (01:10:49) Matthew unpacked his well-researched list that examines the under-investigated fields in 10 academic domains ranging from computer science and biology to economics and philosophy.
  • (01:14:41) Closing segment.

His Contact Info

  • Website
  • Twitter
  • LinkedIn
  • GitHub

His Recommended Resources

  • Vijay Pande
  • David Ha
  • Chip Huyen
  • John Brockman's "This Idea Must Die"

View Details

Show Notes

  • (2:22) Carl talked about his early exposure to programming and his Bachelor’s degree in Computer Science at the University of Rochester in the late 90s.
  • (5:12) Carl implemented his first fully connected, two-hidden-layer artificial neural network using the C programming language back in 2000 when using neural networks wasn’t nearly as cool as it is today.
  • (8:00) Carl started his career as a software engineer at IBM, writing software for large-scale distributed systems and voice-dialog management system.
  • (13:31) The first production machine learning system that Carl worked on is called Conversational Interaction Manager, which is a dialog management system for conversational mixed-initiative natural language applications. He brought up the challenges in DevOps and data quality.
  • (20:05) The second production machine learning system that you worked on is called Smarter Campus, which is a project that enables staffing recommendations based on social networking, optimization, and text analytics.
  • (27:16) Carl unpacked the evolution of his career at IBM, working on various leadership roles. In particular, he worked on IBM Bluemix, IBM’s cloud platform-as-a-service, with over 1 million registered users. He emphasized the importance of talking to customers and finding product-market fit.
  • (33:01) Carl discussed his decision to pursue a Master’s degree in Computer Science at the University of Florida in the mid of his career.
  • (35:24) Carl explained his research paper, which combines game theory and machine learning called “AmalgaCloud: Social Network Adaptation for Human and Computational Agent Team Formation.” The paper focuses on the relationship between network adaptation for candidate group participants and the performance of problem-solving groups.
  • (40:50) Carl discussed his patent on learning ontologies for machine learning - which maps ontologies from data warehouses to computer systems.
  • (47:00) Carl unpacked his 4-part blog series dated in 2016 that discusses server-less computing via tools such as Docker and Apache OpenWhisk.
  • (52:58) Carl emphasized the importance of learning Docker to be productive as a Machine Learning practitioner.
  • (55:02) Carl became a program manager at Google Cloud and helped manage the company’s efforts to democratize machine learning via the Advanced Solutions Lab in 2017.
  • (59:07) Carl recalled his experience as an instructor at various machine learning boot camps.
  • (01:01:44) Carl went over the growing popularity of semi-structured data, referring to his talk at Google’s 2018 Data Cloud Next event.
  • (01:06:29) Currently, Carl is the CTO of CounterFactual AI, which works with various clients using tools such as PyTorch and AWS. He brought up an example of a food delivery application.
  • (01:09:13) Carl went over his experience leading a workshop on Server-less Machine Learning with TensorFlow at the Reinforce AI Conference in Budapest last year.
  • (01:10:52) Carl is writing a book with Manning called Server-less Machine Learning In Action. He explained that server-less tools help minimize the efforts to do MLOps.
  • (01:13:47) Carl talked about the rise of PyTorch as a production-ready deep learning framework, as well as his preference for the PyTorch’s language design philosophy.
  • (01:17:10) Carl shared his opinions on choosing different cloud platforms to host and run the server-less ML pipeline.
  • (01:19:37) Carl described the data and tech community in Orlando, Florida.
  • (01:21:53) Closing segment.

His Contact Information

  • Website
  • LinkedIn
  • Twitter
  • GitHub
  • Google Scholar

His Recommended Resources

  • “An Introduction to Natural Computation” by Dana Ballard
  • IBM’s semiconductor facility FAB
  • IBM Bluemix
  • “Pattern Recognition and Machine Learning” by Christopher Bishop
  • Docker
  • PyTorch
  • PyTorch Lightning
  • Jurgen Schmidhuber (the father of LSTM)
  • Solomon Hykes (Founder, CTO, and chief architect of Docker)
  • Dana Ballard (professor of Computer Science at UT-Austin)
  • Gang of Four Design Patterns (engineering book with object-oriented design theory and practice)

Serverless Machine Learning In Action

  • Check out the book at this link: https://www.manning.com/books/serverless-machine-learning-in-action?a_aid=khanhnamle1994&a_bid=fa913283
  • Here are the 5 free Ebook codes: smldcr-0BE0, smldcr-9F05, smldcr-F807, smldcr-CBB2, smldcr-52D1
  • Here is a 40% discount code: poddcast19

View Details

Show Notes

  • (2:40) Brian discussed his career as a musician and his mission as a consultant to bring design principles into the analytics world.
  • (5:25) Brian talked about his career inception as a UX designer and his interest in human-centered design.
  • (9:48) Brian shared the backstory behind starting Designing for Analytics and his advice for anyone interested in becoming a consultant (hint: finding the minimum viable audience for your craft!).
  • (20:44) Brian shared the common problems that his clients ask him to solve - citing that many of the solutions in engineering-driven organizations are “Technically Right, Effectively Wrong” (listen to Brian’s podcast with David Stephenson).
  • (27:20) Brian explained why data product design goes well beyond user interfaces and helps define what is required to enable the desired user and business outcomes, referring to his post “Does your data product enable surgery or healing?”
  • (33:14) Brian revealed the tactical tips for designing an effective prototype for data products, as shared in his post “Designing MVPs for Data Products and Decision Support Tools.”
  • (40:31) Brian talked about the importance of using human-centered design to measure meaningful engagement in the context of data products, as shared in his post “Why Low Engagement May Not be the Problem with Your Data Product or Analytics Service” (Hint: Think about the last mile and use design to make deliberate choices to improve user engagement).
  • (47:46) Brian unpacked the design framework CED (which stands for Conclusion, Evidence, and Data), which helps build customer trust, engagement, and indispensability around advanced analytics.
  • (54:55) Brian shared his take on how to structure a quad team, including software engineers, UX designers, data scientists, and product managers to build machine learning-powered products.
  • (01:03:25) Brian emphasized the importance of trust in modern data products, after countless conversations with leaders in his podcast Experiencing Data.
  • (01:06:34) Brian unveiled his seminar called Designing Human-Centered Data Products for data scientists, technical product managers, and analytics practitioners.
  • (01:09:47) Closing segment.

His Contact Info

  • Designing For Analytics
  • Twitter
  • LinkedIn
  • Experiencing Data Podcast
  • Insights Newsletter

His Recommended Resources

  • Seth Godin’s podcast Akimbo
  • Minimum Viable Product
  • Wizard of Oz Testing
  • CED Framework
  • Chris Do
  • Scott Berkun (his book “How Design Makes The World”)
  • Juhan Sonin (Involution Studios)
  • Amanda Cox (NYT’s The Upshot)
  • “Good Charts” (by Scott Berinato)
  • “Change By Design” (by Tim Brown)
  • “Infonomics” (by Douglas Laney)
  • “Competing In The Age of AI” (by Marco Iansiti and Karim Lakhani)

View Details

Show Notes

  • (2:19) Luigi got his Bachelor’s in Mathematics and Master’s in Computer Science from Fordham University, with a break working as a Data Analyst in between.
  • (5:41) Luigi worked as a Research Engineer at Fordham's Wireless Sensor Data Mining Lab for a year during his Master’s program and got exposed to Machine Learning.
  • (9:13) Luigi’s first role out of graduate school is a Data Engineering position at Namely, a Human Resources platform for thousands of mid-sized companies.
  • (14:33) Luigi then worked as a Machine Learning Engineer at CTRL-Labs - a startup (acquired by Facebook) pioneering the development of non-invasive neural interfaces that reimagine how humans and machines collaborate.
  • (20:45) Luigi discusses the skills he picked up during his transition from Data Engineering to Machine Learning Engineering, such as data analysis, data visualization, dimensionality reduction, and domain expertise.
  • (25:38) Luigi went over his time teaching graduate courses in Applied Statistics & Probability and Big Data Programming at Fordham’s Department of Computer Science.
  • (28:37) Luigi talked about his next role working as a Data Scientist at 2U - an edTech SaaS platform providing schools with the comprehensive operating infrastructure they need to attract, enroll, educate, support, and graduate students globally.
  • (31:12) Luigi emphasized the importance of being good at data science and picking up skills from other functional domains for anyone looking into management roles.
  • (33:47) Luigi shared brief thoughts on the role of ed-tech in the current environment with remote education.
  • (35:47) Luigi unpacked his blog post called How I Hire Data Scientists that shares advice for both the hiring managers and the job applicants.
  • (42:05) Luigi shared his anecdotal journey of starting ML In Production - which provides content on the best practices of doing machine learning in production. Check out this article for more detail!
  • (47:16) Luigi discussed the nuts and bolts of setting up the weekly newsletter for his website.
  • (51:21) Luigi unpacked the 4-part series "Docker for Machine Learning” that discusses the benefits of using Docker with machine learning, how to build custom Docker images, how to perform batch inference using Docker containers, and how to perform online inference using Docker and Flask REST API.
  • (54:26) Luigi’s next post, "Batch Inference vs. Online Inference," discusses the differences between using batch inference or online inference for model serving.
  • (56:51) Luigi’s next post, "Storing Metadata from ML Experiments," reveals the importance of storing metadata during the machine learning process as well as the types of metadata to capture.
  • (01:00:49) Luigi’s following post "How Data Leakage Impacts ML Models" goes over the issues of data leakage, which occurs when data used at training time is unavailable at inference time.
  • (01:04:14) Luigi unpacked his 6-part series that first introduces Kubernetes and then goes deeper into its components, including Pods, Jobs, CronJobs, Deployment, and Services.
  • (01:07:06) Luigi reflected on his talk “Productionizing ML Models at scale with Kubernetes” at the TWIML conference last year.
  • (01:10:50) Luigi dug into his popular post "The Ultimate Guide to Model Retraining," which covers the problem of model drift as well as the necessary steps to retrain models already in production.
  • (01:14:36) Luigi laid out the benefits of using AWS SageMaker for model deployment. Check out his concise description of SageMaker’s architecture as well as his video tutorial on how to train scikit-learn models on SageMaker.
  • (01:17:50) Luigi unpacked his multi-part series on model deployment. So far, he has covered deployment in the machine learning context, software interfaces, batch inference, online inference, model registries, test-driven development, and A/B testing.
  • (01:22:12) Luigi encouraged every software engineer to learn about running ML systems in production, given the gradual shift to Software 2.0, as indicated in his post "Machine Learning is Forcing Software Development to Evolve."
  • (01:26:48) Luigi reveals what differentiates successful industry ML projects from unsuccessful ones, based on his interview series with other ML practitioners. Hint: (1) don’t focus on the hype, instead focus on the business outcomes + (2) start small.
  • (01:30:10) Luigi distinguished the skills required for the three roles: data engineers, data scientists, and machine learning engineers.
  • (01:33:46) Luigi shared his opinions on the data science community in New York City.
  • (01:35:41) Closing segment.

His Contact Info

  • Website
  • Twitter
  • LinkedIn

His Recommended Resources

  • MLinProduction.com
  • Monitor! Stop Being a Blind Data Scientist (by Ori Cohen)
  • Andrej Karpathy
  • Lex Fridman
  • Xavier Amatriain
  • “Hands-On Machine Learning with Scikit-Learn, Keras, and TensorFlow” by Aurelien Geron

A New Course From Luigi

Luigi just launched his first online course, Build, Deploy, and Monitor Machine Learning Models with Amazon SageMaker! I had a look at the course content, and I’m convinced that the course will be super valuable to any ML engineer or data scientist who wants to level up and learn how to productionize their machine learning models.

You can take the course on your own, but Luigi is also teaming up with TWiML to offer a version with virtual Study Group sessions for people who want a more interactive experience. Right now, Luigi is offering an early bird discount on the course until August 1st!

I know the course will be precious for a lot of you within my community, so Luigi created a coupon code DATACAST to save an additional 10% off the course!

Head over to the Teachable course page to learn more about AmazonSageMaker and take advantage of the discount!

View Details

Show Notes

  • (2:00) Alexey studied Information Systems and Technologies from a local university in his hometown in eastern Russia.
  • (4:54) Alexey commented on his experience working as a Java developer in the first three years after college in Russia and Poland, along with his initial exposure to Machine Learning thanks to Coursera.
  • (7:55) Alexey talked about his decision to pursue the IT4BI Master Program specializing in Large-Scale Business Intelligence in 2013.
  • (9:42) Alexey discussed his time working as a Research Assistant on Apache Flink at the DIMA Group at TU Berlin.
  • (12:28) Alexey’s Master Thesis is called Semantification of Identifiers in Mathematics for Better Math Information Retrieval, which was later presented at the SIGIR conference on R&D in Information Retrieval in 2016.
  • (14:35) Alexey discussed his first job as a Data Scientist at Searchmetrics - working on projects to help content marketers improve SEO ranking for their articles.
  • (18:54) Alexey’s next role was with the ad-tech company Simplaex. There, he designed, developed, and maintained the ML infrastructure for processing 3+ billion events per day with 100+ million unique daily users - working with tools like Spark for data engineering tasks.
  • (22:17) Alexey reflected on his journey participating in Kaggle competitions.
  • (25:35) Alexey also participated in other competitions at academic conferences: winning 2nd place at the Web Search and Data Mining 2017 challenge on Vandalism Detection and winning 1st place at the NIPS 2017 challenge on Ad Placement.
  • (29:59) Alexey authored his first book called Mastering Java for Data Science, which teaches readers how to create data science applications with Java.
  • (31:40) Alexey then transitioned to a Data Scientist role at OLX Group, a global marketplace for online classified advertisements.
  • (33:23) Alexey explained the ML system that detects duplicates of images submitted to the OLX marketplace, which he presented at PyData Berlin 2019. Read his two-part blog series: The first post presents a two-step framework for duplicate detection, and the second post explains how his team served and deployed this framework at scale.
  • (38:12) Alexey was recently involved in building an infrastructure for serving image models at OLX. Read his two-part blog series on this evolution of image model serving at OLX, including the transition from AWS SageMaker to Kubernetes for model deployment, as well as the utilization of AWS Athena and MXNet for design simplification.
  • (42:39) Alexey is in the process of writing a technical book called Machine Learning Bookcamp - which encourages readers to learn machine learning by doing projects.
  • (46:17) Alexey discussed common struggles during data science interviews, referring to his talk on Getting a Data Science Job.
  • (48:32) Alexey has put together a neat GitHub page that includes both theoretical and technical questions for people who are preparing for interviews.
  • (52:19) Alexey extrapolated on the steps needed to become a better data scientist, in conjunction to his LinkedIn post a while back.
  • (56:40) Alexey gave his advice for software engineers looking to transition into data science.
  • (58:32) Alexey shared his opinion on the data science community in Berlin.
  • (01:01:53) Closing segment.

His Contact Info

  • Website
  • Twitter
  • LinkedIn
  • GitHub
  • Kaggle
  • Quora
  • Google Scholar
  • Medium

His Recommended Resources

  • Apache Flink
  • Kubeflow
  • Data Science Interviews GitHub Repo
  • PyData Berlin
  • Berlin Buzzwords
  • Andrew Ng
  • Designing Data-Intensive Applications by Martin Kleppmann

Machine Learning Bookcamp

  • Permanent 40$ discount code: poddcast19
  • 5 free eBook codes (each good for one sample of the book): mlbdrt-D452, mlbdrt-5922, mlbdrt-2C4D, mlbdrt-3034, mlbdrt-1DD1

View Details

Show Notes

  • (2:27) Ankit studied Electrical Engineering with a focus on Communication and Signal Processing at the Indian Institute of Technology, Bombay.
  • (3:27) Ankit then worked for three years as a Senior Field Engineer at Schlumberger, an international oilfield services company.
  • (4:23) Ankit then went to the US to pursue a Masters in Financial Engineering from the Walter Hass School of Business at UC Berkeley.
  • (6:13) Ankit had an opportunity to intern as a data scientist at Facebook during his Masters and worked on detecting spam for Facebook pages.
  • (8:27) Ankit worked full-time as a Quantitative Finance Analyst at Bank of America after finishing his degree, with projects such as building models to identify risk in bank portfolio and analyzing relevance opportunities for strategic investment.
  • (9:46) Ankit discussed his transition to a Data Scientist role at ClearSlide, a B2B platform for Sales Enablement + Engagement.
  • (11:32) Ankit discussed his work on sales forecasting algorithms at ClearSlide.
  • (15:06) In 2015, Ankit moved to Bangalore to become the Head of Data Science and Analytics at Ruunr, a B2B platform that offers hyper-local logistics services that partners with merchants in India.
  • (16:18) Ankit unpacked his thorough post “How Food Delivery Can Be a Sustainable Business” that reflects his experience at Ruunr.
  • (18:35) Ankit talked about the similarities and differences of tech culture in Bangalore and San Francisco.
  • (19:39) Ankit came back to the US and started working as a Data Scientist at Uber in early 2017.
  • (20:34) Ankit discussed his work at Uber on user-level forecasting.
  • (23:12) Ankit talked about the different types of problems that researchers at Uber AI Labs work on.
  • (24:49) Ankit unpacked his in-depth technical post on Uber’s Engineering blog “Food Discovery with Uber Eats: Using Graph Learning to Power Recommendations” — including graph neural networks for food recommendations, the design of the data and training pipeline, and ways to incorporate more data for further improvement.
  • (28:55) Ankit discussed the challenges with building the Uber Eats recommendation system in production.
  • (32:15) Ankit has written a technical book called TensorFlow Machine Learning Projects — which teaches how to exploit the benefits (simplicity, efficiency, and flexibility) of using TensorFlow in various real-world projects.
  • (34:43) Ankit gave his two cents on the battle of frameworks between TensorFlow and PyTorch.
  • (36:26) Ankit shared his advice for academics looking to work in the industry: building end-to-end projects, learning how to build scalable pipelines, and keeping up with important research topics.
  • (38:33) Ankit reflected on the benefits of his electrical engineering and financial analysis education towards his career in data science.
  • (40:11) Closing segment.

His Contact Info

  • LinkedIn
  • Twitter
  • GitHub
  • Quora

His Recommended Resources

  • GraphSAGE
  • Meta-Graph: Few-Shot Link Prediction via Meta-Learning
  • Ankit’s book "TensorFlow Machine Learning Projects” published with Packt
  • Andrew Ng
  • Geoffrey Hinton
  • Jeff Dean
  • "Elements of Statistical Learning" by Trevor Hastie, Robert Tibshirani, and Jerome Friedman

View Details

Show Notes

  • (2:32) Ari discussed his undergraduate studying Physiology and Neuroscience at UC San Diego, while doing neuroscience research on adult neurogenesis at the Gage Lab.
  • (4:39) Ari discussed his decision to pursue a Ph.D. in Neurobiology at Harvard after college and extracted the importance of communication in research, thanks to his advisor Chris Harvey.
  • (7:16) Ari explained his Ph.D. thesis titled “Population dynamics in parietal cortex during evidence accumulation for decision-making” - in which he developed methods to understand how neuronal circuits perform the computations necessary for complex behavior.
  • (12:59) Ari talked about his process of learning machine learning and using that to analyze massive neuroscience datasets in his research.
  • (15:22) Ari recounted attending NIPS 2015 and serendipitously meeting people from DeepMind, which he lated joined as a Research Scientist in their London office.
  • (18:59) Ari’s research focuses on the generalization of neural networks, and shared his work called "On the Importance of Single Directions for Generalization” presented at ICLR 2018 (inspired by Chiyuan Zhang’s paper and Quoc Le’s paper previously).
  • (28:51) Ari explained the differences between generalizing networks and memorizing networks, citing the results from his work "Insights on Representational Similarity in Neural Networks with Canonical Correlation” with Maithra Raghu and Samy Bengio presented at NeurIPS 2018 (Read Maithra’s paper on SVCCA that inspired it).
  • (35:16) Another topic that Ari focuses on is representation learning and abstraction for intelligent systems. His team at DeepMind proposes a dataset and a challenge designed to probe abstract reasoning, as explained in “Measuring Abstract Reasoning in Neural Networks" presented at ICML 2018 (learn more about the IQ test Raven’s Progressive Matrices and take the challenge here).
  • (42:21) An extension from the work above is "Learning to Make Analogies by Contrasting Abstract Relational Structure" - presented at ICLR 2019. With the same authors (led by Felix Hill along with David Barrett, Adam Santoro, Tim Lillicrap), Ari showed that while architecture choice can influence generalization performance, the choice of data and the manner in which it is presented to the model is even more critical.
  • (48:18) Ari discussed "Neural Scene Representation and Rendering” (led by Ali Eslami and Danilo Rezende) that introduces Generative Query Network (GQN), a framework within which machines learn to represent scenes using only their own sensors (watch the video and check out the data).
  • (55:09) Ari explained the findings in "Analyzing Biological and Artificial Neural Networks: Challenges with Opportunities for Synergy?” published at the Current Opinion in Neurobiology (joint work with David Barrett and Jakob Macke).
  • (57:04) Ari shared the properties of pruning algorithms that influence stability and generalization, as claimed in “The Generalization-Stability Tradeoff in Neural Network Pruning” led by Brian Bartoldson.
  • (01:00:56) Ari went over the generalization of lottery tickets in neural networks, which is inspired by the lottery ticket hypothesis from Jonathan Frankle and Michael Carbin at MIT. The two papers mentioned are collaboration with Haonan Yu, Yuandong Tian, Michela Paganini, and Sergey Edunov (Check out his talk at REWORK Deep Learning Summit in Montreal 2019).
  • (01:09:00) Ari investigated "Training BatchNorm and Only BatchNorm” which looks at the performance of neural networks when trained only with the Batch Normalization parameters (joint work with Jonathan Frankle and David Schwab).
  • (01:12:12) Ari mentioned "The Early Phase of Neural Network Training” (presented at ICML 2020) that uses the lottery ticket framework to rigorously examine the early part of the training (joint work with Jonathan Frankle and David Schwab).
  • (01:16:25) Ari discussed at length “Representation Learning Through Latent Canonicalizations" (presented at ICLR 2020). This work seeks to learn representations in which semantically meaningful factors of variation (like color or shape) can be independently manipulated by learned linear transformations in latent space, termed “latent canonicalizes” (joint work with Or Litany, Srinath Sridhar, Leonidas Guibas, and Judy Hoffman).
  • (01:22:15) Ari summarized "Selectivity Considered Harmful: Evaluating the Causal Impact of Class Selectivity in DNNs" - which investigates the causal impact of class selectivity on network function (led by Matthew Leavitt).
  • (01:25:26) Ari reflected on his career and shared advice for individuals who want to make a dent in AI research.
  • (01:28:10) Ari shared his excitement on self-supervised learning, which addresses the need of neural networks to require expensive labeled data.
  • (01:29:47) Closing segment.

His Contact Information

  • Website
  • Google Scholar
  • LinkedIn
  • Twitter
  • GitHub

His Recommended Resources

  • “Understanding Deep Learning Requires Rethinking Generalization” by Chiyuan Zhang
  • “Singular Vector Canonical Correlation Analysis for Deep Learning Dynamics and Interpretability” by Maithra Raghu
  • Raven’s Progressive Matrices IQ test
  • "The Lottery Ticket Hypothesis” by Jonathan Frankle and Michael Carbin (Open-Source Framework)
  • “Random Features for Large-Scale Kernel Machines” by Ali Rahimi and Ben Recht (NIPS 2017 Test Of Time Award)
  • “beta-VAE: Learning Basic Visual Concepts with a Constrained Variational Framework” by DeepMind
  • Samy Bengio (Research Scientist at Google AI)
  • Aleksander Madry (Professor of Computer Science at MIT)
  • Jason Yosinski (Founding Member of Uber AI Labs)
  • “The Idea Factory: Bell Labs and The Great Age of American Innovation" by Jon Gertner

View Details

Show Notes:

  • (2:02) Josh studied Mathematics at Columbia University during his undergraduate and explained why he was not set out for a career as a mathematician.
  • (3:55) Josh then worked for two years as a Management Consultant at McKinsey.
  • (6:05) Josh explained his decision to go back to graduate school and pursue a Ph.D. in Mathematics at UC Berkeley.
  • (7:23) Josh shared the anecdote of taking a robotics class with professor Pieter Abbeel and switching to a Ph.D. in the Computer Science department at UC Berkeley.
  • (8:50) Josh described the period where he learned programming to make the transition from Math to Computer Science.
  • (10:46) Josh talked about the opportunity to collaborate and then work full-time as a Research Scientist at OpenAI - all during his Ph.D.
  • (12:40) Josh discussed the sim2real problem, as well as the experiments conducted in his first major work "Transfer from Simulation to Real World through Learning Deep Inverse Dynamics Model".
  • (17:43) Josh discussed his paper "Domain Randomization for Transferring Deep Neural Networks from Simulation to the Real World", which has been cited more than 600 times up until now.
  • (20:51) Josh unpacked the OpenAI’s robotics system that was trained entirely in simulation and deployed on a physical robot, which can learn a new task after seeing it done once (Read the blog post “Robots That Learn” and watch the corresponding video).
  • (24:01) Josh went over his work on Hindsight Experience Replay - a novel technique that can deal with sparse and binary rewards in Reinforcement Learning (Read the blog post “Generalizing From Simulation").
  • (28:41) Josh talked about the paper "Domain Randomization and Generative Models for Robotic Grasping”, which (1) explores a novel data generation pipeline for training a deep neural network to perform grasp planning that applies the idea of domain randomization to object synthesis; and (2) proposes an autoregressive grasp planning model that maps sensor inputs of a scene to a probability distribution over possible grasps.
  • (32:27) Josh unpacked the design of OpenAI's Dactyl - a reinforcement learning system that can manipulate objects using a Shadow Dexterous Hand (Read the paper “Learning Dexterous In-Hand Manipulation” and watch the corresponding video).
  • (35:31) Josh reflected on his time at OpenAI.
  • (36:05) Josh investigated his most recent work called “Geometry-Aware Neural Rendering” - which tackles the neural rendering problem of understanding the 3D structure of the world implicitly.
  • (28:21) Check out Josh's talk "Synthetic Data Will Help Computer Vision Make the Jump to the Real World" at the 2018 LDV Vision Summit in New York.
  • (28:55) Josh summarized the mental decision tree to debug and improve the performance of neural networks, as a reference to his talk "Troubleshooting Deep Neural Networks” at Reinforce Conf 2019 in Budapest.
  • (41:25) Josh discussed the limitations of domain randomization and what the solutions could look like, as a reference to his talk "Beyond Domain Randomization” at the 2019 Sim2Real workshop in Freiburg.
  • (44:52) Josh emphasized the importance of working on the right problems and focusing on the core principles in machine learning for junior researchers who want to make a dent in the AI research community.
  • (48:30) Josh is a co-organizer of Full-Stack Deep Learning, a training program for engineers to learn about production-ready deep learning.
  • (50:40) Closing segment.

His Contact Information:

  • Website
  • LinkedIn
  • Twitter
  • GitHub
  • Google Scholar

His Recommended Resources:

  • Full-Stack Deep Learning
  • Pieter Abbeel
  • Ilya Sutskever
  • Lukas Biewald
  • “Thinking Fast and Slow” by Daniel Kahneman

View Details

Show Notes:

  • (2:20) Sara shared her childhood growing up in Africa.
  • (4:05) Sara talked about her undergraduate experience at Carleton College studying Economics and International Relations.
  • (9:07) Sara discussed her first job working as an Economics Analyst at Compass Lexecon in the Bay Area.
  • (12:20) Sara then joined Udemy as a data analyst, then transitioned to the engineering team to work on spam detection and recommendation algorithms.
  • (14:58) Sara dig deep into the “hustling period” of her career and how she brute-forced her way to grow as an engineer.
  • (17:24) Sara founded Delta Analytics - a local Bay Area non-profit community of data scientists, engineers, and economists in 2014 that believes in using data for good.
  • (20:53) Sara shared Delta’s collaboration with Eneza Education to empower students to access quizzes by mobile texting in Kenya (check out her presentation at the ODSC West 2016).
  • (25:16) Sara shared Delta’s partnership with Rainforest Connection to identify illegal de-forestation using steamed audio from the rainforest (check out her presentation at MLconf Seattle 2017).
  • (28:22) Sara unpacked her blog post Why “data for good” lacks precision, in which she described 4 key criteria frequently used to qualify an initiative as “data for good” and discussed some open challenges associated with each.
  • (36:34) Sara unpacked her blog post Slow learning, in which she revealed her journey to get accepted into the AI Residency program at Google AI.
  • (41:03) Sara discussed her initial research interest on model interpretability for deep neural networks and her work done at Google called The (Un)reliability of Saliency Methods - which argues that saliency methods are not reliable enough to explain model prediction.
  • (45:55) Sara pushed the research above further with A Benchmark for Interpretability Methods in Deep Neural Networks, which proposes an empirical measure of the approximate accuracy of feature importance estimates in deep neural networks called RemOve And Retrain.
  • (48:46) Sara explained why model interpretability is not always required (check out her talks at PyBay 2018, REWORK Toronto 2018, and REWORK San Francisco 2019).
  • (52:10) Sara explained the typical measurements of model reliability and the limitations of them, such as localization methods and points of failure.
  • (59:04) Sara explained why model compression is an interesting research direction and her work The State of Sparsity in Deep Neural Networks - which highlights the need for large-scale benchmarks in the field of model compression.
  • (01:02:49) Sara discussed her paper Selective Brain Damage: Measuring the Disparate Impact of Model Pruning - which explores the impact of pruning techniques for neural networks trained for computer vision tasks. Check out the paper website!
  • (01:05:08) Sara shared her future research directions on efficient pruning, sparse network training, and local gradient updates.
  • (01:06:56) Sara explained the premise behind her talk Gradual Learning at the Future of Finance Summit in 2019, in which she shared the three fundamental approaches to machine learning impact.
  • (01:12:20) Sara described the AI community in Africa as well as the issues the community is currently facing: both from the investment landscape and the infrastructure ecosystem.
  • (01:18:00) Sara and her brother recently started a podcast called Underrated ML which pitches the underrated ideas in machine learning.
  • (01:20:15) Sara reflected how her background in economics influences her career outlook in machine learning.
  • (01:25:42) Sara reflected on the differences between applied ML and research ML, and shared her advice for people contemplating between these career paths.
  • (01:29:49) Closing segment.

Her Contact Information:

  • Website
  • LinkedIn
  • GitHub
  • Google Scholar
  • Twitter
  • Medium

Her Recommended Resources :

  • Deep Learning Indaba
  • Southeast Asia Machine Learning School
  • MILA - AI For Humanity
  • Why “data for good” lacks precision (Sara's take on "Data for Good" initiatives)
  • Slow learning (Sara's journey to Google AI)
  • fast.ai
  • Sanity Check for Saliency Maps by Julius Adebayo et al.
  • Focal Loss for Dense Object Detection by Tsung-Yi Lin et al.
  • MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications by Andrew Howard et al.
  • Underrated ML (Sara’s new podcast)
  • Dumitru Erhan (Research Scientist at Google AI)
  • Samy Bengio (Research Scientist at Google AI)
  • Andrea Frome (Ex-Research Engineer at Google AI)
  • Elements of Statistical Learning by Trevor Hastie, Robert Tibshirani, and Jerome Friedman

View Details

Show Notes:

  • (1:58) Colleen gave a brief overview of her professional background and her path to data science.
  • (3:27) Colleen explained how her background in medicine and social science contributes to her success as a data scientist working in different domains.
  • (5:20) Colleen share her thoughts on how data science varies by sector.
  • (7:02) Referring to her consulting company Staticlysm LLC, Colleen discussed a current medical technology project that she is working on.
  • (8:19) Colleen shared applications of quantum machine learning in the wild, referring to her work at Quantopo - where she is the cofounder and chief mathematician.
  • (12:33) Colleen discussed a new project leveraging quantum and quantum-inspired algorithms for nuclear reactor optimization at Quantopo.
  • (15:17) Colleen gave advice for data scientists who want to start a business and get into consulting.
  • (16:58) Colleen discussed topological data analysis in machine learning.
  • (23:02) Colleen discussed epidemic modeling and engaging foreign aid organizations, given her experience with the Ebola outbreak.
  • (26:01) Colleen discussed the different approaches to model the spread of diseases in epidemics, which follow a differential equation framework.
  • (30:22) Colleen explained why buy-in and corporation from those in power are critical to combat epidemics.
  • (32:20) Colleen explained the difference between the current Coronavirus pandemic and the previous Ebola epidemic (note that this conversation is recorded in mid-March 2020, so information about COVID19 may have dated).
  • (37:06) Colleen shared resources to get up-skilled in data science.
  • (39:40) Colleen talked about the benefits of writing on Quora, where she has written more than 13,000 answers.
  • (41:31) Colleen shared the traits of an excellent technical communicator and/or data translator.
  • (43:26) Colleen shared her process of writing a technical book that focuses on the use-cases of topology, geometry, and graph theory in machine learning and data science.
  • (50:32) Colleen talked about the growth of the data science community in Miami.
  • (54:30) Colleen discussed her involvement in data science within the African sectors.
  • (57:55) Colleen shared her thoughts on how the data science field will evolve in the next few years: the access to big data platforms, the dominance of Python, and the rise of quantum computing.
  • (01:01:12) Closing segment.

Her Contact Info:

  • LinkedIn
  • Quora
  • ResearchGate
  • KDNuggets

Her Recommended Resources: 

  • IBM Quantum Computing
  • Xanadu (Quantum Computing Hardware)
  • D-Wave Systems (Quantum Computing Hardware)
  • Ayasdi (Topological Data Analysis-Focused Startup)
  • The SIR Model for Spread of Disease
  • Google Scholar
  • arXiv
  • LinkedIn Learning
  • Coursera
  • Genetic Algorithms and Adaptation paper by John Holland
  • Random Forests paper by Leo Breiman
  • Andrew Ng (Founder of Coursera, DeepLearning.AI, and Landing.AI)
  • Dover Series on Mathematics

View Details

Show Notes:

  • (2:12) Parul talked about her educational background, studying Electrical Engineering at the National Institute of Technology, Hamirpur.
  • (3:18) Parul worked as a Business Analyst at Tata Power India for 7 years.
  • (4:29) Parul talked about her initial interests in writing about data science and machine learning on Medium.
  • (6:30) Parul discussed her first blog series “A Guide to Machine Learning in R for Beginners” - which covers the Fundamentals of ML, Intro to R, Distributions and EDA in R, Linear Regression, Logistic Regression, and Decision Trees.
  • (8:02) Reference to her articles on data visualization, Parul talked about matplotlib, seaborn, and plotly as the main visualization libraries she practices, in addition to Tableau for building dashboard.
  • (10:11) Parul shared her thoughts on the state of Machine Learning interpretability, in reference to her articles on this topic.
  • (13:54) Parul discussed the advantages of using Jupyter Lab over Jupyter Notebook.
  • (17:30) Parul discussed the common challenges of bringing recommendation systems from prototype into production (Read her two articles about recommendation systems: (1) an overview of different approaches and (2) an overview of the process of designing and building a recommendation system pipeline)
  • (21:00) Parul went in depth into her NLP project called "Building a Simple Chatbot from Scratch in Python (using NLTK).”
  • (23:26) Parul continued this chatbot project with a 2-part series on building a conversational chatbot with Rasa stack and Python and deploying it on Slack.
  • (28:15) Parul went over her Satellite Imagery Analysis with Python piece, which examines the vegetation cover of a region with the help of satellite data.
  • (32:22) Parul talked about the process of Recreating Gapminder in Tableau: A Humble Tribute to Hans Rosling.
  • (35:17) Parul discussed her project Music Genre Classification, which shows how to analyze an audio/music signal in Python.
  • (39:20) Parul went over her tutorials on Computer Vision: (1) Face Detection with Python using OpenCV and (2) Image Segmentation with Python’s scikit-image module.
  • (42:01) Parul unpacked her tutorial "Predicting the Future with Facebook’s Prophet” - a forecasting model to predict the number of views for her Medium articles.
  • (44:58) Parul have been working as a Data Science Evangelist at H2O.AI since July 2019.
  • (47:04) Parul described Flow - H2O's web-based interface (Read her tutorial here).
  • (49:23) Parul described Driverless AI - H2O’s product that automates the challenging and repetitive tasks in applied data science (Read her tutorial here).
  • (52:39) Parul described AutoML - H2O's automation of the end-to-end process of applying ML to real-world problems (Read her tutorial here).
  • (57:07) Parul shared her secret sauce for effective data visualization and storytelling, as illustrated in her analysis of the 2019 Kaggle Survey to figure out women’s representation in machine learning and data science.
  • (01:02:02) Parul described the data science community in Hyderabad, from her lens as an organizer for the Hyderabad Chapter of the Women in Machine Learning and Data Science.
  • (01:05:45) Parul was recognized as a LinkedIn’s Top Voices 2019 in the Software Development category.
  • (01:10:30) Closing segment.

Her Contact Info:

  • Medium
  • GitHub
  • Twitter
  • LinkedIn
  • Website
  • Kaggle

Her Recommended Resources:

  • Interpretable Machine Learning post
  • "Interpretable Machine Learning: A Guide for Making Black Box Models Explainable" by Chris Molnar
  • “Towards A Rigorous Science of Interpretable Machine Learning” by Finale Doshi-Velez and Been Kim
  • Parul’s Compilation of Data Visualization articles
  • Parul’s Programming with Python articles
  • Women in Machine Learning and Data Science
  • Rachel Thomas
  • Andreas Mueller
  • HuggingFace
  • “Factfulness” by Hans Rosling

View Details

Show Notes:

  • (2:18) Leonard discussed his undergraduate experience at Carnegie Mellon - where he studied Biology and Computer Science.
  • (5:10) Leonard decided to pursue a Ph.D. in Bioinformatics at the University of California - San Francisco.
  • (6:27) Leonard described his Ph.D. research that focused on finding hidden patterns in genetically-linked diseases.
  • (9:42) Leonard went deep into clustering algorithms (Markov Clustering and Louvain) and their applications such as protein and news article similarity.
  • (13:21) Leonard shared his story of starting a data science consultancy with various client startups.
  • (17:58) Leonard discussed the interesting consulting projects that he worked on: from detecting plagiarism to predicting bill insurance.
  • (22:04) Leonard shared practical tips to learn technical concepts.
  • (23:23) Leonard reflected on his experience working with a string of startups including Accretive Health, Quid, and Stride Health.
  • (26:06) Leonard is the founding team member of Primer AI, a startup that applies state-of-the-art NLP techniques to build machines that read and write, back in early 2015.
  • (30:31) Leonard discussed the technical challenges to develop algorithms that power Primer’s products to scale across languages other than English.
  • (34:28) Leonard unpacked his technical post "Russian NLP” on Primer’s blog.
  • (38:17) Leonard talked about the advances in the NLP research domain that he is most excited about in 2020 (XLNet >>> BERT).
  • (41:10) Leonard discussed the challenges of scaling the data-driven culture across Primer AI as the company grows.
  • (46:20) Leonard mentioned different use cases of Primer for clients in finance, government, and corporate.
  • (51:41) Leonard talked about his decision to leave Primer and become a Data Science Health Innovation Fellow at the Berkeley Institute for Data Science.
  • (54:30) Leonard went over applications of data science in healthcare that will be adopted widely in the next few years.
  • (1:02:45) Leonard discussed his process of writing a book called “Data Science Bookcamp.”
  • (1:07:21) Leonard revealed how he chose the case studies to be included in the book.
  • (1:10:27) Closing segment.

His Contact Info:

  • LinkedIn
  • Google Scholar
  • Berkeley Institute For Data Science

His Recommended Resources:

  • Semi-Supervised Learning
  • Association Rule Learning
  • spaCy (Open-Source Library for Advanced NLP)
  • fastText (NLP library from Facebook)
  • XLNet: Generalized Autoregressive Pretraining for Language Understanding
  • BERT: Pretraining of Deep Bidirectional Transformers for Language Understanding
  • Federated Learning with Differential Privacy: Algorithms and Performance Analysis
  • Differential Privacy- Enabled Federated Learning for Sensitive Health Data
  • Oasis Labs and Dr. Dawn Song
  • Fitbit and Apple Watch
  • Walter Pitts who invented neural networks
  • Paul Werbos who invented back-propagation
  • Fei-Fei Li who constructed the ImageNet dataset
  • “The Signal and The Noise” by Nate Silver

You can read the completed chapters of "Data Science Bookcamp" using the codes below:

  • Permanent discount code: poddcast19
  • 5 free eBook codes: dcdsprf-B373, dcdsprf-CA3B, dcdsprf-299E, dcdsprf-6E5, and dcdsprf-9660 (activated and will last for 2 months)

View Details

Show Notes:

  • (2:25) Vincent talked about his educational background, in which he studied Information Systems at the Singapore Management University.
  • (4:20) Vincent talked about the capstone project on Data Analytics Practicuum that he completed for his degree.
  • (5:47) Vincent went over his Software Development and Business Intelligence internship experience with VISA.
  • (9:30) Vincent landed a Data Science internship at Lazada Group, an international e-commerce company based in Singapore.
  • (11:09) Vincent decided to come back to VISA for a full-time software engineering role after college.
  • (12:32) Vicent worked on a variety of projects from designing micro-services to developing data dictionary management system during his 2-year stint at VISA
  • (15:30) Vincent shared Nigel Poulton's resources on Kubernetes and Docker.
  • (16:38) Vincent discussed his career move to become a Data Analyst at Google Singapore, focusing on Trust and Safety.
  • (20:16) Vincent went over the unique challenges of fighting abuse at Google’s scale.
  • (22:55) In the project "Stock Analysis with Pandas and Scikit-Learn," Vincent walked through an application that retrieves and displays the right financial insights quickly about a certain company stocks price.
  • (25:11) In the project “Build Your Own Data Dashboard,” Vincent showed a tutorial on how to work with Dash, an open-source Python library to build web apps which are optimized for data visualization.
  • (29:13) In the project “Deploy Your First Analytics Project,” Vincent showed a tutorial on how to deploy a dashboard web app with Heroku.
  • (31:48) Vincent shared tips on how to ace the data analysis and data science interviews from big companies in his article “Ace Your Data Analytics Interviews.”
  • (36:15) Vincent unpacked his article “Data Analytics Is Hard… Here’s How To Excel."
  • (43:52) Vincent unpacked his article "How I Overcome Imposter Syndrome in Data Analytics.”
  • (51:11) Vincent shared his thoughts regarding the tech and data community in Singapore.
  • (53:42) Closing Segment.

His Contact Info:

  • LinkedIn
  • Twitter
  • GitHub
  • Medium
  • YouTube

His Recommended Resources:

  • Nigel Poulton on Kubernetes
  • Sebastian Thrun
  • Hans Rosling
  • DJ Patil
  • “Deep Work” by Cal Newport

View Details

Show Notes:

  • (2:17) Ben talked about his past career working in the golf industry - working at the National Golf Foundation and the PGA of America.
  • (4:12) Ben discussed about his first exposure to machine learning and data science.
  • (5:06) Ben talked about his motivation for pursuing an online Master’s degree in Data Science at Southern Methodist University.
  • (6:02) Ben emphasized the importance of a Data Mining course that he took.
  • (8:12) Ben discussed his job as a Senior Data Scientist at CarePredict, an AI elder care platform that helps senior live independently, economically, and longer.
  • (8:56) Ben shared his thought about data security, the biggest challenge of adopting machine learning in healthcare.
  • (10:38) Ben talked about his next employer JM Family Enterprises, one of the largest companies in the automotive industry.
  • (12:44) Ben walked through the end-to-end model development process to solve various problems of interests in his Data Scientist work at JM Family Enterprises.
  • (14:15) Ben discussed the challenges around feature engineering and model experiments in this process.
  • (18:09) Ben shared information about his current role as Machine Learning Technical Lead at Southeast Toyota Finance.
  • (19:29) Ben talked about his passion to do IC data science work.
  • (22:37) Ben went over different conferences he has been / will be at.
  • (26:03) Ben shared the best practices/techniques/libraries to do efficient feature engineering and feature selection, as presented at Palm Beach Data Science Meetup in September 2018 and PyData Miami in January 2019.
  • (29:27) Ben talked about the importance of doing exploratory data analysis and logging experiments before engaging in any feature engineering / selection work.
  • (32:50) Ben shared his experiments performing data science for Fantasy Football - specifically using machine learning to predict the future performance of players, from his talk at the Palm Beach Data Science Meetup last year.
  • (37:25) Ben talked about his experience using H2O AutoML.
  • (40:07) Ben gave a glimpse of his talks about evaluating traditional and novel feature selection approaches at PyData LA and Strata Data Conf.
  • (51:25) Ben gave his advice for people who are interested in speaking at conferences.
  • (52:29) Ben shared his thoughts about the tech and data community in the greater Miami area.
  • (53:16) Closing Segment.

His Contact Info:

  • LinkedIn

His Recommended Resources:

  • MLflow from Databricks
  • Streamlit Library
  • PyData Conference
  • H2O World Conference
  • O’Reilly Strata Data and AI Conference
  • REWORK Summit Conference
  • Pandas Library
  • XGBFir Library
  • tsfresh Library
  • Lending Club Dataset
  • SHAP library from Scott Lundberg
  • "Interpretable Machine Learning with XGBoost" by Scott Lundberg
  • Amazon SageMaker
  • Google Cloud AutoML
  • H2O AutoML
  • Wes McKinney’s "Python for Data Analysis"

View Details

Show Notes:

  • (2:00) Arthur talked about his undergraduate studying Psychology at North Carolina State University.
  • (3:28) Arthur mentioned his time working as a research assistant at the LACElab in NCSU that does human factor and cognition research.
  • (5:08) Arthur discussed his decision to pursue a graduate degree in Cognitive Neuroscience at the University of Oregon right after college.
  • (6:35) Arthur went over his Master's thesis (Navigation performance in virtual environments varies with fractal dimension of landscape) in more detail
  • (10:30) Arthur unpacked his popular blog series called “Simple Reinforcement Learning in TensorFlow” on Medium.
  • (12:56) Arthur recalled his decision to join Unity to work on its reinforcement learning problems.
  • (14:31) Arthur recalled his choice to do the Ph.D. part-time while working full-time.
  • (16:24) Arthur discussed problems with existing reinforcement learning simulation platforms and how the Unity Machine Learning Agents Toolkit addresses those.
  • (18:30) Arthur went over the challenges of maintaining and continuously iterating the Unity ML Agents toolkit.
  • (20:36) Arthur emphasized the benefit of training the agents with an additional curiosity-based intrinsic reward, which is inspired from a paper from UC Berkeley researchers (check out the Unity blog post).
  • (22:33) Arthur talked about the challenges of implementing such curiosity-based techniques.
  • (25:15) Arthur unpacked the introduction of the Obstacle Tower - a high fidelity, 3D, third person, procedurally generated environment - released in the latest version of the toolkit (read his blog post “On “solving” Montezuma’s Revenge”).
  • (29:15) Arthur discussed the Obstacle Tower Challenge, a contest that offers researchers and developers the chance to compete to train the best-performing agents on the Obstacle Tower Environment.
  • (32:49) Referring to his fun tutorial called “GANs explained with a classic sponge bob squarepants episode,” Arthur walked through the theory behind the Generative Adversarial Network algorithm via an explanation using an episode of Spongebob Squarepants.
  • (34:30) Arthur extrapolated on his post “RL or Evolutionary Strategies? Nature has a solution: Both.”
  • (38:36) Arthur shared a couple of approaches to balance the bias and variance tradeoff in reinforcement learning models, referring to his article “Making sense of the bias/variance tradeoff in Deep RL.”
  • (41:19) Arthur talked about successor representations and their applications in deep learning, psychology, and neuroscience (read his post "The present in terms of the future: Successor representations in RL”).
  • (42:38) Arthur reflected on the benefits of his Psychology and Neuroscience background for his research career.
  • (44:21) Arthur shared his advice for graduate students who want to make a dent in the AI / ML research community.
  • (45:30) Closing segment.

His Contact Info:

  • Twitter
  • GitHub
  • Medium
  • LinkedIn
  • Google Scholar
  • Unity Blog

His Recommended Resources:

  • DeepMind
  • Google Brain
  • Being and Time (by Martin Heidegger)

View Details

Show Notes:

  • (2:27) Alex talked about his undergraduate experience studying Applied Mathematics at the Kiev Polytechnic Institute.
  • (3:15) Alex quickly went over his time working remotely as a Machine Learning Engineer for a US-company called Inma AI during university.
  • (4:24) Alex mentioned his decision to pursue a Master’s degree in Mathematics at the University of Verona.
  • (6:21) Alex went over the Math graduate classes that he took for his degree, including differential geometry and optimization theory.
  • (7:58) Alex talked about his experience working at Mlvch to create the best solutions on the market related to visual style transfer and image enhancement.
  • (10:01) Alex shed some light on his Master’s thesis work called “Boosting financial models calibration with deep neural networks” (Read his post "Meta-learning in finance: boosting models calibration with deep learning”).
  • (11:39) Alex talked about the applications of meta-learning, a powerful technique to train deep neural networks, in physics and biology.
  • (13:15) Alex shared his experience working as a partner and solutions architect at Mawi Solutions, a Ukraine-based hardware-and-software company that is disrupting the market of wearable preventive healthcare.
  • (16:29) In reference to blog post “Deep learning: the final frontier for signal processing and time series analysis?,” Alex discussed ways to apply deep learning to model time series (Watch his talk at PyCon Italia 2019 as well).
  • (18:52) Alex is currently the Co-Founder and CTO of Neurons Lab, an innovative European AI boutique based out of London that serves clients in financial tech, marketing tech, and medical tech areas.
  • (22:41) Alex is well-known for a series of blog posts that experiment neural networks for algorithmic trading time series forecasting (Check out his Deep Trading code repo).
  • (26:51) Alex emphasized the benefits of using multi-task learning, in reference to his post “Multitask learning: teach your AI more to make it better” (check out his experiments for 4 different use cases).
  • (29:35) Alex advocated for the need of using generative models in applied AI (Read his 2 blog posts “GANs beyond generation: 7 alternative use cases” and “Generative AI: A key to machine intelligence?”).
  • (34:02) Alex talked about a new approach called disentangled representation learning, which combines the best of classical math and machine learning modeling, in reference to his article “GANs” vs “ODEs”: the end of mathematical modeling? (See his Code and Talk).
  • (36:36) In reference to his post “Fantastic data scientists: where to find them and how to become one,” Alex described the type of data scientist he identifies with as well as the skills that he is looking to improve upon.
  • (42:08) Alex went over the evolution of algorithms for asset portfolio management, referring to his piece “AI for portfolio management: from Markowitz to Reinforcement Learning” (See his Code Repo).
  • (45:09) Alex gave his advice for people who want to get into blogging and public speaking.
  • (48:00) Alex gave his AI predictions for the year 2020 (Read his 2018 predictions for researchers and for developers).
  • (49:54) Alex shared his thoughts regarding the data science community in East Europe vs West Europe.
  • (52:10) Closing Segment.

His Contact Info:

  • LinkedIn
  • Medium
  • Twitter
  • GitHub

His Recommended Resources:

  • Quantopian Lectures on Quantitative Investment
  • Successful Algorithmic Trading Ebook from QuantStart
  • Advanced Algorithmic Trading Ebook from QuantStart
  • "Python for Finance” by Yves Hilpisch
  • "Derivatives Pricing in Python” by Yves Hilpisch
  • "Advances in Financial Machine Learning” by Marcos Lopez de Prado
  • “Pattern Recognition and Machine Learning” by Christopher Bishop
  • Machine Learning and Reinforcement Learning in Finance in Coursera
  • OpenAI
  • DeepMind
  • Salesforce Einstein
  • Facebook AI
  • Google AI

View Details

Show Notes:

  • (2:08) Mael recalled his experience getting a Bachelor of Science Degree in Economics from HEC Lausanne in Switzerland.
  • (4:47) Mael discussed his experience co-founding Wanago, which is the world’s first van acquisition and conversion crowdfunding platform.
  • (9:48) Mael talked about his decision to pursue a Master’s degree in Actuarial Science, also at HEC Lausanne.
  • (11:51) Mael talked about his teaching assistantships experience for courses in Corporate and Public Finance.
  • (13:30) Mael talked about his 6-month internship at Vaudoise Assurances, in which he focused on an individual non-life product pricing.
  • (16:26) Mael gave his insights on the state of adopting new tools in the actuarial science space.
  • (18:12) Mael briefly went over his decision to do a Post Master’s program in Big Data at Telecom Paris, which focuses on statistics, machine learning, deep learning, reinforcement learning, and programming.
  • (20:51) Mael explained the end-to-end process of a deep learning research project for the French employment center on multi-modal emotion recognition, where his team delivered state-of-the-art models in text, sound, and video processing for sentiment analysis (check out the GitHub repo).
  • (26:12) Mael talked about his 6-month part-time internship doing Natural Language Processing for Veamly, a productivity app for engineers.
  • (28:58) Mael talked about his involvement with VIVADATA, a specialized AI programming school in Paris, as a machine learning instructor.
  • (34:18) Mael discussed his current responsibilities at Anasen, a Paris-based startup backed by Y Combinator back in 2017.
  • (38:12) Mael talked about his interest in machine learning for healthcare, and his goal to pursue a Ph.D. degree.
  • (40:00) Mael provided a neat summary on current state of data engineering technologies, referring to his list of in-depth Data Engineering Articles.
  • (42:36) Mael discussed his NoSQL Big Data Project, in which he built a Cassandra architecture for the GDELT database.
  • (47:38) Mael talked about his generic process of writing technical content (check out his Machine Learning Tutorials GitHub Repo).
  • (52:50) Mael discussed 2 machine learning projects that I personally found to be very interesting: (1) a Language Recognition App built using Markov Chains and likelihood decoding algorithms, and (2) the Data Visualization of French traffic accidents database built with D3, Python, Flask, and Altair.
  • (56:13) Mael discussed his resources to learn deep learning (check out his Deep Learning articles on the theory of deep learning, different architectures of deep neural networks, and the applications in Natural Language Processing / Computer Vision).
  • (57:33) Mael mentioned 2 impressive computer vision projects that he did: (1) a series of face classification algorithms using deep learning architectures, and (2) face detection algorithms using OpenCV.
  • (59:47) Mael moved on to talk about his NLP project fsText, a few-shot learning text classification library on GitHub, using pre-trained embeddings and Siamese networks.
  • (01:03:09) Mael went over applications of Reinforcement Learning that he is excited about (check out his recent Reinforcement Learning Articles).
  • (01:05:14) Mael shared his advice for people who want to get into freelance technical writing.
  • (01:06:47) Mael shared his thoughts on the tech and data community in Paris.
  • (01:07:49) Closing segment.

His Contact Info:

  • Twitter
  • Website
  • LinkedIn
  • GitHub
  • Medium

His Recommended Resources:

  • Deep Learning by Ian Goodfellow, Yoshua Bengio, and Aaron Courville
  • PyImageSearch by Adrian Rosebrock
  • Station F Incubator in Paris
  • BenevolentAI
  • Econometrics Data Science: A Predictive Modeling Approach by Francis Diebold

View Details

Show Notes:

  • (1:57) Jannes discussed his undergraduate experience studying International Business Administration at Rotterdam School of Management - Erasmus University (one of Europe’s top 10 research-based business schools).
  • (3:06) Jannes talked about his work on the IHS Global Green City Index, a set of measurements that can give majors actionable insight in where cities stand in sustainability.
  • (4:05) Jannes talked about his involvement with the Turing Society in Rotterdam, where he developed a new kind of machine learning class called “Bletchley Bootcamp for Machine Learning in Financial Context".
  • (5:20) Jannes discussed how he built out the materials and inviting guest lectures for his machine learning for finance class.
  • (6:57) Jannes talked about his decision to pursue a Master’s degree in Financial Economics at Said Business School, a part of Oxford University.
  • (8:04) Jannes went over the most useful graduate course that he took for the Master’s degree, called “Information and Communication in Finance."
  • (11:23) Jannes discussed his role with the Oxford Artificial Intelligence Society, which provides a platform to educate, build, connect, and employ an AI community that constantly drives innovation for the university and the world.
  • (14:23) Jannes shared his thoughts regarding challenges of bringing different perspectives into the conversations around AI.
  • (15:40) Jannes shared a brief overview of his current employer, QuantumBlack, and an example project with Formula 1 to optimize pitch stops.
  • (18:12) Moving on to discuss his book “Machine Learning For Finance” (which introduces the study of machine learning and deep learning algorithms for financial practitioners), Jannes went over his motivation as well as the ideal audience.
  • (19:27) Jannes went over in details the process of writing this book.
  • (22:39) Jannes recommended a couple of resources for people who are new to Time Series forecasting.
  • (23:57) Jannes revealed how he uses Twitter to keep up-to-date with NLP research.
  • (25:54) Jannes explored 2 powerful financial applications of generative models: (1) perform synthetic data generation that can help with data labeling efforts and (2) generate realistic time series.
  • (27:59) Jannes explained why private equity is an exciting playground for Reinforcement Learning models.
  • (29:05) Jannes shared his opinions on common approaches to do Reinforcement Learning in the finance industry.
  • (32:13) Jannes recommended the best practices to deploy machine learning models into production.
  • (34:26) Jannes shared some interesting research projects of ethics and fairness in machine learning that attracts his attention.
  • (36:56) Jannes shared how his financial economics background contributes to him being a good data scientist.
  • (38:11) Jannes shared his thoughts on the tech and data community in London.
  • (38:56) Closing segment.

His Contact Info:

  • LinkedIn
  • Twitter
  • GitHub
  • Medium

His Recommended Resources:

  • Machine Learning For Finance
  • Oxford Internet Institute
  • Quant GANs: Deep Generation of Financial Time Series paper
  • MLFlow Open-Source Platform for End-to-End ML Lifecycle
  • Professor Sandra Wachter’s paper: “Affinity Profiling and Discrimination by Association in Online Behavioral Advertising"
  • Two Sigma
  • BenevolentAI
  • Computer Age Statistical Inference by Bradley Efron and Trevor Hastie

View Details

Show Notes:

  • (2:15) Jan discussed his undergraduate experience studying Business Administration and Economics from Goshen College in Indiana.
  • (3:48) Jan went over his first job out of college: working as a Strategy and Enterprise Intelligence consultant at the EY office in Berlin, a Big 4 consulting firm.
  • (6:09) Jan talked about his decision to pursue a part-time Master’s degree in Computer Science at the Trier University of Applied Sciences while working at EY.
  • (7:21) Jan covered the most useful graduate courses during his Master’s degree, including Advanced Programming and Distributed Systems.
  • (9:05) As part of his program, Jan did his thesis with the Scout24. In fact, he even wrote a blog post offering a glimpse of what it’s like to be a data scientist at Scout24.
  • (12:31) Jan discussed the benefits of taking deep learning online classes from Andrew Ng’s deeplearning.ai platform.
  • (15:52) Jan is also a mentor and ambassador with deeplearning.ai, in which he gives feedback on the educational content, discusses new product ideas, and writes forum entries.
  • (19:30) Jan discussed his current projects at Carmeq GmbH, the Berlin-based innovation vehicle of Volkswagen AG.
  • (21:22) Related to his work at Carmeq, in the blog post “The State of Self-Driving Cars for Everybody,” Jan outlined the 6 main infrastructural problems of self-driving cars for the masses. We discussed these problems in finer detail.
  • (30:52) Jan gave a curated list of 5 mindset-chasing books that helps him become a better data scientist, referring to his article “Top 5 Business-Related Books Every Data Scientist Should Read.”
  • (38:32) Jan shared the 5 pitfalls that young data scientists can stumble upon in their first job, referring to his article “The Power of Goal-Setting in Data Science.”
  • (44:38) Jan emphasized the importance of using Google’s goal-setting method OKRs (Objects and Key Results) to set a data science project up for success, referring to his article “The Power of Goal-Setting in Data Science.”
  • (48:32) Jan explained the AI Project Canvas, which answers the most pressing questions about the outcome and resources needed for an AI project.
  • (51:50) Jan went over the importance of learning business basics for data scientists.
  • (56:30) Referring to his post called “Becoming a Level 3.0 Data Scientist,” Jan discussed his current career trajectory as well as skills that he is looking to develop.
  • (58:20) Jan gave his advice for data scientists to make a leap from an individual contributor to a manager.
  • (01:02:04) Referring his post called “The Secrets to a Successful AI Strategy,” Jan gave his advice for data scientists to collaborate productively with their counterparts in product management and business operations.
  • (01:02:55) Jan shared his opinions on the technology and data community in Berlin.
  • (01:05:00) Closing segment.

His Contact Info:

  • Medium
  • LinkedIn
  • Twitter
  • GitHub

His Recommended Resources:

  • Deep Learning Specialization from deeplearning.ai
  • Federated Learning
  • Lex Friedman’s “Deep Learning for Self-Driving Cars” class taught at MIT
  • Nassim Taleb’s “Skin In The Game"
  • Nassim Taleb’s “Black Swan"
  • Peter Thiel’s “Zero To One"
  • Eric Ries’ “The Lean Startup"
  • Daniel Kahneman’s “Thinking, Fast and Slow"
  • Richard Rumelt’s “Good Strategy, Bad Strategy"
  • John Doerr’s “Measure What Matters"
  • AI Project Canvas
  • William Oncken’s “Managing Management Time"
  • Google AI and DeepMind
  • DeepMind: The Podcast hosted by Hannah Fry
  • Hans Rosling’s “Factfulness"

View Details

Show Notes:

  • (2:17) Admond described his undergraduate experience studying Applied Physics from Nanyang Technological University in Singapore.
  • (3:02) Admond talked about the importance of learning multivariate calculus and linear algebra in college.
  • (3:38) Admond shared his research internship experience at CERN, the European Organization for Nuclear Research, in Geneva after his junior year.
  • (4:54) Admond went over his part-time data analytics internship at SMRT Corporation while still at school.
  • (6:10) Admond shared his decision to graduate one semester early to pursue a Data Science internship at Quantum Inventions.
  • (7:42) Admond wrote a Medium blog post titled “My first data scientist internship” detailing his experience at Quantum Inventions.
  • (9:07) Admond gave his advice for job seekers to filter out the signal and the noise in data science job postings, with reference to his article “Why did I reject a data scientist job?”
  • (11:12) Admond discussed his work as a research engineer at the online gaming platform Titansoft, where he focused on Artificial Intelligence in human behavior imitation to enhance the current automation system.
  • (13:53) Admond shared the lessons he wrote about in the blog post called “5 Lessons I Have Learned from Data Science in real working experience."
  • (17:34) Admond talked about his current employer Micron Technology.
  • (19:06) Admond walked through the 4 stages to define a problem statement, as evidence in his article “How to ask the right question as a data scientist.”
  • (24:22) Admond argued for the importance of resourcefulness as a data scientist, regarding his piece “Be resourceful - one of the most important skills to succeed in data science.”
  • (28:34) Admond went over the 5 core principles that are useful to his writing from his article “How do I write about Data Science on Medium.”
  • (33:50) Admond considered specializing in Natural Language Processing after having been a generalist for 2 years (“Why You Should Be a Generalist first, Specialist later as a Data Scientist”).
  • (36:11) Admond went over the speech recognition side project that he has been working on.
  • (39:34) Admond shared his thoughts regarding ethics and privacy within speech recognition.
  • (42:42) Admond shared other projects he’s been involved with, including public speaking and consulting gigs.
  • (45:53) Admond talked about his involvement with the AI Time Journal as a committee member for the AI for Education 2019 Initiative, which identifies and showcases the most impactful and beneficial applications of Artificial Intelligence in the field of Education.
  • (50:24) Admond talked about the tech and data community in Singapore.
  • (57:38) Admond discussed his lowest point during his journey into data science and how he handled that from a psychological standpoint.
  • (01:03:18) Admond shared the 6 lessons for recent college grads, in reference to his article “One Year After Graduation, Here is What I’ve Learned.”
  • (01:11:24) Closing segment.

His Contact Info:

  • LinkedIn
  • Medium
  • GitHub
  • Twitter

His Recommended Resources:

  • Grab
  • Bitcurate
  • "The Lean Startup” by Eric Lies
  • AI Singapore
  • TensorFlow and Deep Learning Meetup at Singapore

View Details

Show Notes:

  • (02:05) Sunanda discussed her Master’s degree in Physics from the Indian Institute of Technology, Madras.
  • (03:05) Sunanda talked about her decision to go to the US to pursue a Ph.D. in Condensed Matter and Materials Physics at Purdue University.
  • (04:22) Sunanda went in-depth into her Ph.D. thesis that focused on spin transport.
  • (06:51) Sunanda accepted a Materials Science Postdoc Fellowship at Princeton University and worked there for 3 years on quantum insulators.
  • (10:05) Sunanda recalled the most valuable lessons she learned from her Postdoc time.
  • (11:49) Sunanda talked about her transition from physics to data science by getting a certificate of Professional Achievement in Data Science from Columbia University.
  • (18:42) Sunanda went over the best courses she took at Columbia that prepare her well for a career as a Data Scientist.
  • (21:13) Sunanda talked about the ease of learning Python and R from scratch given her programming experience in C++.
  • (22:29) Sunanda quickly mentioned her job search experience.
  • (24:32) Sunanda discussed her work as a Data Scientist for 3 years at DataXu (a Boston-based advertising startup), including advertising attributions measurements and marketing-mix modeling.
  • (28:35) Sunanda revealed the workflow of algorithmic conceptualization and experimental design in her work at DataXu.
  • (31:05) Sunanda talked about her career progress at DataXu.
  • (32:01) Sunanda discussed her move to Wayfair, a big e-commerce company that sells home goods.
  • (34:08) Sunanda went over her responsibilities as an Associate Director of Data Science at Wayfair.
  • (35:58) Sunanda gave her advice for people who want to make a transition from an individual contributor to a data science leader within an organization.
  • (38:17) Sunanda gave a quick overview of R&D work on recommendation systems at Wayfair.
  • (40:31) Sunanda commented on the differences between the two work environments at startup company versus at medium-sized company.
  • (41:48) Sunanda shared her thoughts on tech and data community in the Boston area.
  • (43:14) Closing segment.

Her Contact Info:

  • Twitter
  • LinkedIn

Her Recommended Resources:

  • Stitch Fix Tech Blog
  • Pinterest Engineering Blog
  • Carol Dweck’s “Mindset”

View Details

Show Notes:

  • (2:35) Tristan talked about his undergraduate experience studying Aeronautical Engineering from the University of Witwatersrand in South Africa.
  • (4:10) Tristan discussed his Technical Director experience at a company called Y2K Tec, where he managed developers working on tools using Visual C++, SQL Server, and Windows Server.
  • (5:11) Tristan talked about his 3 years working as a solutions and enterprise application architect at IKT CC.
  • (6:45) Tristan went over his 2-year experience as a software architect at Unisys, a global IT company that builds high-performance, security-centric solutions for the most demanding businesses and governments around the world.
  • (8:55) Tristan talked about his transition to become a manager at JumpCo - a young-minded IT products and service company that has serviced Enterprise Application Design, Development and Implementation within the South African market in the next 5 years.
  • (10:43) Tristan discussed his work as the Director and Cloud Architect at Ixio Analytics, which is a startup that specializes in data and management information, advanced analytics and machine learning.
  • (14:03) Tristan went into more detail about his consulting service Led By Data, which delivers whole-loop process design for business operations.
  • (17:45) Referring to his talk “Predictive Analytics in Action - Real Business Results in South Africa” back in 2015, Tristan gave an example to illustrate his proposed framework to do predictive analytics efficiently called Question, Think, Test, Analyze.
  • (23:00) Tristan went over his transition from a Director title to a Data Scientist title at Led By Data in late 2016.
  • (24:58) Tristan shared his thoughts regarding the transferrable skills from cloud architect to being data science.
  • (31:21) Tristan shared his insights on the process of deploying machine learning models into production.
  • (32:43) Tristan shared the best online resources that he used to prepare for his Data Science role.
  • (39:11) Tristan discussed various ongoing projects that he’s been dabbling on.
  • (43:24) Tristan talked about his job as a Senior Data Scientist at DirectAxis, which is a big financial services organization in South Africa.
  • (45:47) Tristan went into the details about his responsibility for delivering end-to-end predictive analytics solutions serving financial services products at DirectAxis.
  • (51:22) Tristan shared his thoughts on how machine learning and data science can help improve the personal lending and insurance processes (especially for the low risk mass middle market).
  • (57:30) Tristan discussed his data science work at his current employer aYo Holdings, an organization that serves over 2.5 million customers in South Africa who purchase medical and life insurance on MTN cellular networks.
  • (01:01:32) Tristan reflected on the benefits that his Aeronautical Engineering background has had on his Data Science career.
  • (01:05:15) Tristan shared his insights on the tech and data community in South Africa.
  • (01:09:50) Closing segment.

His Contact Info:

  • Twitter
  • LinkedIn

His Recommended Resources:

  • gbm (Generalized Boosted Regression Models) R package
  • Eric Siegel’s “Predictive Analytics”
  • Jason Brownlee’s “Machine Learning Mastery”
  • Analytics Vidha
  • Data Science Central
  • Andrew Ng
  • scikit-learn tutorials
  • David Beazley and Brian Jones' “Python Cookbok”
  • Microsoft Azure ML
  • Python Dask library
  • Discovery Health

View Details

Show Notes:

  • (2:04) Ankit discussed his college experience at the Amity School of Engineering and Tech in which he studied Computer Science and Engineering.
  • (3:34) Ankit talked about his first role out of school as a Senior Solution Integrator role at Ericsson.
  • (5:20) Ankit talked about a unique aspect of his role at Ericsson, which was global collaboration and communication with clients all over the world, including the US, UK, and Nigeria.
  • (7:50) Ankit went over his 2 years at Core Services working as an Oracle Applications Technical Consultant.
  • (9:25) Ankit gave his story about pursuing a Master’s Degree in Information Systems Management at Carnegie Mellon.
  • (10:56) Ankit talked about his Teaching Assistant experience for a couple of Python courses at CMU.
  • (12:13) Ankit talked about this capstone work for his Master’s degree.
  • (13:27) Ankit gave a brief overview about Crisis Text Line.
  • (14:48) Ankit provided the context for his article on LinkedIn called "Triage for Crisis Counseling: How our algorithm prioritizes suicidal texts" that he wrote during his summer internship at Crisis Text Line.
  • (16:40) Ankit discussed a couple of interesting projects he has been involved with since accepting a full-time Data Scientist role at Crisis Text Line in January 2017.
  • (21:19) Ankit talked about his post “Embracing AI to save lives” on Crisis Text Line’s blog, which uses machine learning to identify previously missed imminent risk conversations and reduce false alarms.
  • (28:01) Referring to another piece called "Detecting crisis - An AI solution", Ankit discussed the impacts of AI at Crisis Text Line.
  • (30:34) Ankit talked about the vision of Crisis Text Line’s CEO, Nancy Lublin.
  • (33:21) Ankit gave his advice on becoming good at data science.
  • (37:37) Ankit went over his participation in Toastmasters New York, an organization that operates worldwide for the purpose of promoting communication and public speaking skills.
  • (40:10) Ankit talked about the data science community in New York City.
  • (42:47 ) Closing segment.

His Contact Info:

  • LinkedIn
  • Twitter

His Recommended Resources:

  • Practical Deep Learning For Coders
  • DataKind
  • Elon Musk’s Tesla and SpaceX
  • Ray Dalio’s Principles
  • Cole Nussbaumer Knaflic's Storytelling with Data

View Details

Show Notes:

  • (2:09) Genevieve discussed her undergraduate experience studying Electrical Engineering and Mathematics at the University of Arizona.
  • (3:18) Genevieve talked about her Master’s work in Electrical Machines from the University of Tokyo.
  • (6:59) Genevieve went in-depth about her research work on transverse-flux motor design during her Master’s, in which she won the Outstanding Paper award at ICEMS 2009.
  • (11:39) Genevieve talked about her motivation to pursue a Ph.D. degree in Computer Science at Brown University after coming back from Japan.
  • (14:17) Genevieve shared her story of finding her research advisor (Dr. James Hays) as a graduate student.
  • (18:44) Genevieve discussed her work building and maintaining the SUN Attributes dataset, a widely used resource for scene understanding, during her first year of her Ph.D. degree.
  • (21:52) Genevieve talked about the paper Basic Level Scene Understanding (2013), her collaboration with researchers from MIT, Princeton, and University of Washington to build a system that can automatically understand 3D scenes from a single image.
  • (24:32) Genevieve talked about the paper Bootstrapping Fine-grained Classifiers: Active Learning with a Crowd in the Loop presented at the NIPS conference in 2013, her collaboration with researchers from UCSD and Cal-Tech to propose an iterative crowd-enabled active learning algorithm for building high-precision visual classifiers from unlabeled images.
  • (28:25) Genevieve discussed her Ph.D. thesis titled “Collective Insight: Crowd-Driven Image Understanding.”
  • (34:02) Genevieve mentioned her next career move - becoming a Postdoctoral Researcher at Microsoft Research New England.
  • (36:40) Genevieve talked about her teaching experience for 2 graduate-level courses: Data-Driven Computer Vision at Brown University in Spring 2016 and Deep Learning For Computer Vision at Tufts University in Spring 2017.
  • (38:04) Genevieve shared her 2 advice for graduate students who want to make a dent in the AI/Machine Learning research community.
  • (41:45) Genevieve went over her startup TRASH, which develops computational filmmaking tools for mobile iphono-graphers.
  • (43:45) Genevieve mentioned the benefit of having TRASH as part of the NYU Tandon Future Labs, which is a network of business incubator and accelerators that support early stage ventures in NYC.
  • (45:00) Genevieve talked about the research trends in computer vision, augmented reality, and scene understanding that she’s most interested in at the moment.
  • (45:59) Closing segment.

Her Contact Info:

  • Website
  • GitHub
  • LinkedIn
  • Twitter
  • CV

Her Recommended Resources:

  • The Trouble with Trusting AI to Interpret Police Body-Cam Video
  • Microsoft Research Podcast
  • Stanford’s CS231n: Convolutional Neural Networks for Visual Recognition
  • Yann LeCun's letter to CVPR chair after bad reviews on a Vision System that "learnt" features & reviews
  • Paperspace
  • CVPR 2019
  • Michael Black’s Perceiving Systems Lab at the Max Planck Institute for Intelligent Systems
  • Nassim Taleb’s “The Black Swan”

View Details

Show Notes:

  • (2:02) Peadar discussed his undergraduate experience studying Physics and Philosophy at the University of Bristol.
  • (3:05) Peadar then pursued a Master’s degree in Mathematics from the University of Luxembourg, where he did a thesis on machine learning for time series forecasting.
  • (4:16) Peadar commented on his varied work experience with various companies, particularly on data maturity and the difference of established companies and startups.
  • (7:11) Peadar talked about his latest startup called aflorithmic Labs, which develops tech platform that powers and enables the creation of a new generation hyper-personalized / super-relevant podcasts.
  • (8:13) In the series “Interviews with Data Scientists,” Peadar interviewed with 24 of the world’s most influential and innovative data scientists from across the spectrum. He talked about the common traits in the best data scientists.
  • (10:05) Peadar mentioned his contribution to PyMC3, a Python package for Bayesian statistical modeling and Probabilistic Machine Learning focusing on advanced Markov chain Monte Carlo (MCMC) and variational inference (VI) algorithms.
  • (11:32) Peadar talked about the probabilistic programming survey he conducted recently, in which A/B testing is a big use case.
  • (13:37) In his talk “Lies damned lies and statistics in Python” at PyData London 2016, Peadar compared and debugged models in Statsmodels, scikit-learn and PyMC3. He recalled the differences here.
  • (15:27) Peadar went over “Probabilistic Programming Primer” - an online course he designed to teach people to learn how to enhance modeling abilities and better communicate risk.
  • (18:32) Peadar talked about the recent development in the PyData ecosystem, in reference to his talk “A Map of the PyData Stack” at PyData Amsterdam 2016.
  • (20:18) Discussing his blog post “How to successfully deliver Data Science in the Enterprise,” Peadar went over the people, processes, and things that are required to make data science a successful component in enterprise businesses.
  • (23:25) Discussing his blog post “Building Full-Stack Vertical Data Products,” Peadar emphasized the importance of providing end-to-end value with lean metrics as a data scientist.
  • (29:50) Discussing his blog post “One weird tips to improve the success of DS projects,” Peadar shared his small practice of writing down the risks before embarking on a project.
  • (32:58) Discussing his blog post “3 pitfalls for non-technical managers managing DS teams,” Peadar described the things that non-technical managers will get wrong in managing a technical project.
  • (35:31) Discussing his blog post “What does it mean to be a Senior DS?,” Peadar explained why senior data scientists should understand the soft side of technical decision making and should care about ethics.
  • (38:57) Peadar gave a brief overview of machine learning interpretability.
  • (40:21) Closing segments.

His Contact Info:

  • LinkedIn
  • Twitter
  • GitHub
  • Medium
  • Quora
  • Website

His Recommended Resources:

  • LIME
  • SHAP
  • Stitch Fix Tech Blog
  • Ravelin Blog
  • Stripe Engineering Blog
  • Spotify Discover Weekly
  • Dale Carnegie’s How to Win Friends and Influence People

View Details

Show Notes:

  • (2:06) Nick talked about earning his Ph.D. degree in Psycho-Linguistics from the University of Texas at Austin.
  • (3:58) Nick discussed his Ph.D. dissertation, in which he looked at the role of domain-general decision making processes in human language comprehension.
  • (8:41) Nick talked about his teaching experience in graduate school, as well as why UT-Austin is a perfect place to study linguistics.
  • (12:29) Nick delved into his first job out of school as a Linguistic Associate at Lexicon Branding, a world’s premier naming agency with over 30 years of experience in the business
  • (14:50) Nick talked about his next role as a Data Scientist at Idibon, a dated AI-startup based in Silicon Valley that helped companies understand their language data.
  • (18:55) Nick mentioned his next job as a Senior Data Scientist at CrowdFlower (now known as Figure Eight), a human-in-the-loop ML and AI company based in San Francisco.
  • (23:19) Nick talked about his next Senior Data Scientist role - working on Engagement Lead at Change Healthcare, a healthcare technology company that offers software, analytics, network solutions, and technology-enabled services to help create a stronger, more collaborative healthcare system.
  • (26:28) Nick discussed his transition to become a Senior Data Scientist at Womply, a SaaS-based software that is powered by transaction and online review data for millions of small businesses.
  • (30:08) Nick talked about his current role as a Senior Manager in the Data Science department at Johnson & Johnson’s Health Technology group and shared the challenges of putting AI/ML algorithms into production in the healthcare domain.
  • (34:11) Nick shared his narrative of being a person who can facilitate communication between technical and nontechnical teams in the blog post “Drinks with a businessman” (with an anecdote including his grandfather).
  • (39:27) In reference to “A few thoughts on ML from the perspective of a behavioral scientist,” Nick shared his advice for data scientists who want to develop the core skills including being able to generate and test hypotheses, to think critically about unfamiliar data, and to gauge how much you trust your result.
  • (41:37) Nick shared his findings in his 3-part blog series title “Being a grad student is a lot like being a startup” (Part 1, Part 2, and Part 3).
  • (46:21) In reference to “Data science - the way I see it,” Nick talked about the methodological side of data science.
  • (49:05) Nick discussed how graduate students can leverage their skills to qualify for an industry role.
  • (52:17) Nick reflected on his efforts to stay active in academia while holding an industry job.
  • (54:09) Closing segment.

His Contact Info:

  • LinkedIn
  • Twitter
  • Website

His Recommended Resources:

  • Stitch Fix Technology
  • Netflix Research
  • Fast Company’s Most Innovative Data Science Companies in 2019
  • Going Pro in Data Science by Jerry Overton
  • Smart Thinking, Smart Change, and Brain Briefs by Art Markman
  • Data science is different now by Vicki Boykis

View Details

Show Notes:

  • (2:04) Conor recalled his undergraduate experience studying Computational Modeling and Data Analytics at Virginia Tech.
  • (4:07) Conor emphasized the importance of learning data structures and algorithms for interview prep.
  • (5:57) Conor mentioned his Data Analytics internship at MasterCard after his sophomore year.
  • (7:35) Conor talked about his involvement with the Bio-Complexity Institute at Virginia Tech as a Computational Research Intern during his junior year.
  • (8:58) Conor talked about his process of compiling a big list of data science internship job postings that he shared last year.
  • (12:09) Conor went over his Medium post called “Insights from Analyzing 80+ Job Rejections in Python” in which he talked about his coping mechanism towards failure.
  • (15:20) Conor discussed his Data Science internship at Unity Technologies out in the Bay Area after his junior year.
  • (17:22) Conor went over key lessons presented in his blog post “5 Lessons from a Data Science Intern at a Tech Unicorn.”
  • (23:20) Conor emphasized the importance of having conversations with colleagues to challenge your assumptions.
  • (24:54) Conor addressed the premise of his article “Minimum Viable Analysis.”
  • (27:30) Conor discussed his job search process for a full-time data science position.
  • (29:40) Conor talked about the projects he has been working on at his current employer - Squarespace.
  • (32:28) Conor shared his opinions on the difference in the data science communities in New York and San Francisco.
  • (34:32) Conor shared 3 key insights on acing data science interviews.
  • (38:52) In reference to his article “10 Reads for DS Getting Started with Business Models”, Conor shared his advice for data scientists to get better at communicating with the business folks.
  • (41:17) Conor talked about his part-time involvement with DataCamp to design an interactive online course for data scientists.
  • (43:30) Conor talked about his weekly newsletter with more than 1,000 subscribers.
  • (48:07) Conor shared his curiosity for idea sharing and development.
  • (49:51) Closing segments.

His Contact Info:

  • Website
  • LinkedIn
  • Twitter
  • Medium
  • Github

His Recommended Resources:

  • Ted Talk: What I learned from 100 days of rejection
  • 13 Essential Newsletters for Data Scientists
  • Data Skeptic podcast
  • Data Scientists are Thinkers
  • Netflix Tech Blog
  • Uber Engineering Blog
  • Stitch Fix Tech Blog
  • Facebook’s Prophet forecasting procedure
  • Spotify’s Luigi Python module
  • Thinking Fast and Slow

View Details

Show Notes:

  • (2:20) Martina recalled her experience getting Bachelor, Master’s, and Ph.D. degrees in Physics from Sapienza Università’ di Roma.
  • (3:35) Martina discussed her Ph.D. thesis, in which she looked at the study of Complex Systems related to Linguistics and studied how natural language evolves in time.
  • (6:04) Martina talked about her experience doing the S2DS bootcamp in London after finishing her Ph.D.
  • (7:10) Martina gave her reasons to move to the Greater UK while looking for a tech job.
  • (8:05) Martina talked about the importance of software engineering from her time working for the education company Twig World.
  • (10:07) Martina discussed her current job as a Data Scientist at Mallzee, also known as “Tinder for Fashion.”
  • (11:46) Martina briefly went over her work in recommendation systems, data analytics, and statistical modeling in the first two years at Mallzee.
  • (13:50) Martina explained the unique features of fashion that make it a fertile field to do data science work.
  • (17:30) Martina emphasized the importance of communication for a data scientist.
  • (22:20) Martina talked about her transition to the Data Science Lead role at Mallzee since 2017.
  • (24:26) Martina gave her opinion about the relationship between the technical side and the scientific side of data science. (“Data Science Down the Line”)
  • (26:54) Martina talked about the importance of learning statistics for people coming from an engineering background who want to get into data science.
  • (30:30) Martina explained the analogy in her post “Don’t make recipes out of them” where she compared doing data science to cooking.
  • (35:00) Martina discussed the fundamental problem in academia which makes it losing appeal to young and talented individuals. (“Scientific publishing”)
  • (38:48) Martina talked about her experience organizing the PyData Edinburgh Meetup.
  • (42:04) Martina advocated for contributing to conversations to raise awareness about women in the scientific field. (“Women amaze”)
  • (46:45) Martina discussed her project using data from the corpora present in the NLTK library and analyzed the growth of types with respect to text size. (“The growth of vocabulary in different languages”)
  • (50:52) Martina mentioned her attempt to learn D3 to visualize data. (“Rallying into D3”)
  • (53:20) Martina moved on to her project analyzing tags on Stackoverflow. (“Stackoverflow Tags”)
  • (59:30) Martina talked about using TensorFlow for her deep learning project (“TensorFlow: Create the training set for the object detection”)
  • (01:03:23) Martina gave her prediction on how Data Science and Machine Learning will evolve in the next couple years.
  • (01:07:08) Closing segments.

Her Contact Info:

  • Twitter
  • Website
  • LinkedIn
  • GitHub
  • DataLab Podcast Interview

Her Recommended Resources:

  • Natural Language Toolkit NLTK
  • Scott Murray’s “Interactive Data Visualization For The Web”
  • Dashing D3.js
  • Elijah Meeks’s “D3.js In Action”
  • Malcolm Maclean’s “D3 Tips and Tricks”
  • StackOverflow Developer Survey 2019
  • Trevor Hastie, Robert Tibshirani, Jerome Friedman’s “The Elements of Statistical Learning”
  • Francis Chollet’s “Deep Learning with Python”

View Details

Show Notes:

  • (2:16) Jim recalled his experience getting a Bachelor Degree in Chemistry at the University College of London.
  • (4:01) Jim talked about the analytical skills that he got out from his Chemistry degree.
  • (5:09) Jim gave a brief background overview of his employer KPMG, one of the big 4 consulting firms.
  • (6:56) Jim shared the major challenges of applying scientific rigor to identify and quantify business opportunities using data.
  • (9:42) Jim reflected on his professional growth working after 3 years working as a data analyst at KPMG.
  • (12:23) Jim explained his motivation behind his decision to pursue a Masters in Business Analytics at Imperial College of London.
  • (16:27) Jim recalled the most useful courses he took during his Master degree (Graph Analysis on the technical side and Marketing on the business side).
  • (18:40) Jim talked about the importance of learning econometrics for a data scientist.
  • (21:22) Jim talked about the benefit of teaching materials to other people that contribute significantly to his career, which he wrote about his post “Lessons learned teaching R.”
  • (23:51) Jim recently wrote a blog post about his experience attending the RStudio Conference at Austin in January, in which he shared several principles for teaching.
  • (28:35) Jim started working at the KPMG office in Atlanta starting January 2018.
  • (31:10) Jim talked about his blog post called “Do the simple things first,” in which he argued that “a complex method is never justified until a simple one has been tried first.”
  • (35:21) Jim talked about the use of machine learning for his projects at KPMG.
  • (37:03) Jim shared some resources to learn data engineering, including learning SQL and reading “R For Data Science.”
  • (40:40) Jim shared the key developments in the R ecosystem in 2019 that he’s most excited about, including caret and tidyverse.
  • (44:46) Jim gave his prediction on how data science will evolve in the next 5 years.
  • (49:04) Jim anticipated his career trajectory.
  • (49:49) Closing segments.

His Contact Info:

  • LinkedIn
  • Twitter
  • Website
  • GitHub

His Recommended Resources:

  • R For Data Science Learning Community
  • R For Data Science Slack Channel
  • "What’s in a name” from Lyft Engineering Blog
  • Spotify
  • RStudio
  • DataCamp
  • “Thinking, Fast and Slow” by Daniel Kahneman

View Details

Show Notes:

  • (2:10) Francisco recalled his experience getting a Bachelor of Science Degree in Psychology from Nova Southeastern University.
  • (3:42) Francisco gave a brief overview of the field of Behavioral Neuroscience.
  • (4:19) Francisco gave some recommendation for people who are interested in learning more about behavioral psychology and neuroscience.
  • (5:39) Francisco discussed his research thesis looking at “The effects of video gaming with a brain-computer interface on physiological stress and mood”.
  • (10:45) Francisco shared his opinion on the development of virtual reality technologies.
  • (12:37) Francisco talked about his experience working as a Research Assistant at the Clinical Systems Biology Lab at the Institute for Neuro-Immune Medicine while at college.
  • (16:00) Francisco mentioned potential applications of his research project towards clinical treatments.
  • (16:39) In addition to getting his degree in Psychology, Francisco also minored in Applied Statistics. He talked about some of the most useful courses that he took, including applied regression analysis.
  • (18:10) Francisco talked about his Hockey Analytics internship for the Florida Panthers during his senior year.
  • (21:01) Francisco reflected on his job search process post-graduation.
  • (22:58) Francisco shared a background overview about MotionPoint, his current employer, as well as the reasons that made him excited about working there.
  • (24:25) Francisco talked about a sentiment analysis project that he worked on, in which he extracted conversation themes and key words during sales calls, then used NLP techniques to find which sales reps produce more positive sentiment in prospective clients.
  • (26:06) Francisco went in depth on his modeling approach for that sentiment analysis project.
  • (28:50) Francisco mentioned his promotion from an intern role to a full time data analyst/scientist position.
  • (29:34) Francisco shared another data science challenge he’s tackling to score client accounts.
  • (32:17) Francisco discussed his involvement with DataCamp. In particular, he wrote two tutorials on Markov Chain Analysis in R and Fuzzy String Matching in Python.
  • (35:47) Francisco talked about the value of having a degree in psychology and neuroscience for a data science role, in particular, by having a certain perspective.
  • (39:51) Francisco gave advice for people with Math & Computer Science background to learn the business aspect of a data science role.
  • (41:35) Francisco talked about his enrollment to the Udacity Deep Learning Nanodegree program.
  • (42:45) Closing segment.

His Contact Info:

  • LinkedIn
  • DataCamp
  • GitHub

His Recommended Resources:

  • Facebook AI Research Group
  • PyTorch
  • Waymo
  • DeepMind
  • MIT’s Self-Driving Car Lectures
  • Doing Data Science: Straight Talk From the Front Line by Cathy O’Neil and Rachel Schutt

View Details

Show Notes:

  • (2:22) Matthias recalled his experience getting a Master’s Degree in English, History, and Political Science from the University of Regensburg in Germany.
  • (3:19) Matthias discussed his Master’s Thesis looking at the formation and transformation of 9/11-memory in American news weeklies as well as the importance of project management skills.
  • (6:00) Matthias talked about his 2 years with the Professional Teacher Training Program working for the Bavarian Government.
  • (7:48) Matthias extracted the most useful skills obtained from his experience as an academic high school teacher.
  • (9:33) Matthias went over his decision to pursue a Ph.D. in Linguistics at Ball State University in 2013.
  • (12:04) Matthias explained why he chose to study Linguistics.
  • (13:01) For his Ph.D. dissertation, Matthias did an in-depth analysis of German tweets, where he investigated the interaction between language, demographics, and personality of German users on Twitter.
  • (15:52) Matthias discussed in detail his end-to-end data science approach towards his Ph.D. thesis project.
  • (22:38) Matthias provided insights on the challenges in dealing with linguistic data.
  • (25:35) Matthias talked about the benefits of getting a certificate in statistical modeling, including learning about Generalized Linear Models.
  • (29:15) Matthias mentioned his work as a data science consultant at Ball State’s Digital Scholarship Lab.
  • (31:20) Matthias gave out resources to learn data science using the R statistical language, including DataCamp, Udemy, Coursera, and Lynda.
  • (35:26) Matthias reflected on his job search experience for a data science role after finishing his Ph.D. degree coming from a non-traditional background.
  • (42:30) Matthias went over interesting projects that he has involved with at LiveIntent, his current employer.
  • (45:18) Matthias talked about his recent move from a Data Analyst role to a Data Analytics Manager role.
  • (47:10) Matthias emphasized the value of communication skills as a data science leader.
  • (49:55) Matthias identified the single biggest difference between doing data science in the industry versus in academia.
  • (53:09) Matthias briefly talked about the data science community in New York (check out New York R User Group Meetup).
  • (54:13) Closing segments.

His Contact Info:

  • LinkedIn
  • Website
  • GitHub

His Recommended Resources:

  • Setting up an AWS EC2 instance with RStudio Server and TensorFlow GPU
  • Airbnb’s One Data Science Job Doesn’t Fit All
  • Hadley Wickham and Garrett Grolemund’s "R for Data Science"
  • Leonard Mlodinow’s “The Drunkard’s Walk: How Randomness Rules Our Lives"

View Details

Show Notes:

  • (2:05) Mark revisited his brief stint during college back in the 90s.
  • (3:38) Mark’s first job as a Business Consultant at Dixon Stores Group was a humbling experience.
  • (5:41) Mark stressed the importance of customer empathy that lasts with him throughout his career.
  • (6:45) Mark talked about his next long-term job as a Configuration & Support Engineer at Orange PCS, where he built solid technical skills in IT and data work.
  • (12:51) Mark discussed his next gig working as a Principal Specialist at T-Systems and then doing freelance work in IT and data analytics.
  • (16:23) Mark reflected on the sabbatical years he took a break from working.
  • (19:10) Mark shared an overview about his current employer, Mango Solutions.
  • (20:34) Mark discussed his first big project working at Mango as a senior IT consultant.
  • (25:57) Mark then transitioned into a technical architect role, where he became the bridge between the data science and the IT worlds.
  • (28:33) In reference to his talk “An operating model for R”, Mark stressed the importance of policy, procedure, people and policing to build an effective operating model that connects data scientists and IT specialists.
  • (34:32) Mark gave a client use case that his Data Engineering team has been involved with at Mango.
  • (36:48) Mark talked about the cultural challenges of deploying code into production within an organization.
  • (39:35) Mark talked about his key accomplishments in his leadership role as Head of Data Engineering.
  • (43:27) In reference to his talk “R is production safe”, Mark discussed the existing challenges in the R-community to write production code.
  • (50:39) In reference to his book “Field Guide to the R Ecosystem”, Mark shared the key developments in the R ecosystem in 2019 that he’s most excited about.
  • (53:27) Mark gave thoughts on the rise of cloud-based processing technologies for data engineering’s best practices.
  • (56:42) Mark mentioned the Google Cloud Certified Data Engineer Exam that people can take to learn about data engineering.
  • (59:35) Mark emphasized the importance of communication skills to become an organizational leader.
  • (01:01:13) Mark shared his view on the data science ecosystem in the UK.
  • (01:02:18) Closing segments.

His Contact Info:

  • Website
  • Twitter
  • GitHub
  • LinkedIn

His Recommended Resources:

  • rayshader R package
  • TensorFlow for R
  • sparklyr R interface for Apache Spark
  • Amazon S3 for Data Storage
  • Google Cloud Platform Podcast
  • Hudl
  • RStudio
  • R for Data Science by Hadley Wickham and Garrett Grolemund

View Details

Show Notes:

  • (2:07) Chintan talked about his undergraduate education in India at the Dharmsinh Desai University, studying Electronics & Communication.
  • (2:53) Chintan recalled the most useful knowledge from his undergraduate experience.
  • (3:42) Chintan explained his decision to pursue a Master degree in Communication and Signal Processing at Newcastle University in the UK.
  • (5:09) Chintan discussed advanced communication systems knowledge he learned during his Master.
  • (6:25) Chintan talked about his decision to continue his academic journey with a Ph.D. in Underwater Acoustic Communication at Newcastle University.
  • (8:15) Chintan went in depth to explain the study of Underwater Acoustic Communication and how it is different from normal communication systems.
  • (11:45) Chintan discussed his Ph.D. thesis on applying Bayesian theorem for Turbo coding to solve signal processing problems under the water.
  • (15:00) Chintan elaborated on the benefits of doing a Ph.D. degree.
  • (19:01) Chintan mentioned the importance of collaboration in the academic community.
  • (20:18) Chintan talked about his Post-Doc experience at University of Illinois Urbana Champaign.
  • (22:00) Chintan discussed his next career step moving back to the UK working as a digital communication engineer at Schlumberger, a big company in the oil and gas industry.
  • (27:30) Chintan characterize datasets that are particular to oil and gas.
  • (28:58) Chintan examined the technique of vibration analysis in his work.
  • (31:23) Chintan next landed a role as a Software Engineer at a company called SmartFocus in London.
  • (34:52) Chintan mentioned the best resources he used to learn machine learning and data science.
  • (36:51) Chintan recalled his experience searching for a data science job.
  • (38:02) Chintan briefly went over the interesting projects that he has involved with at Avanade, his current employer.
  • (40:34) Chintan talked about his recent transition from a Senior Data Scientist role to a Manager in Advanced Analytics role.
  • (43:35) Chintan shared his view about the data science community in London.
  • (44:44) Closing segments.

His Contact Info:

  • Twitter
  • GitHub
  • LinkedIn

His Recommended Resources:

  • Andrew Ng’s Machine Learning Coursera Course
  • MIT’s The Analytics Edge edX Course
  • London’s Data Science Festival
  • Nate Silver’s “The Signal and The Noise”
  • Charles Wheelan’s “Naked Statistics”
  • Philip Tetlock’s “Superforecasting”

View Details

Show Notes:

  • (2:16) Thomas talked about the study of Food Science and Technology in which he focused on microbiology.
  • (3:15) Thomas stressed the importance of user empathy, something useful he gained from his degree.
  • (4:39) Thomas discussed the reason to pursue a Ph.D. in Bioinformatics at the Technical University of Denmark.
  • (6:10) Thomas talked in-depth about the tools he developed for his Ph.D. thesis, which are able to handle large-scale pangenome analyses using sequential data.
  • (9:11) Thomas talked about using the ggplot2 package for his R package “Find My Friends.”
  • (11:11) Thomas worked on the ggforce package, which aims at providing missing functionalities to ggplot2 during his internship at RStudio.
  • (13:34) Thomas recalled the best learning he got from his internship with RStudio.
  • (15:08) Thomas gave advice for people who want to contribute to open-source projects.
  • (18:57) Thomas shared the experience working on the package ggraph, also known as the grammar of graphics for relational data.
  • (22:02) Thomas discussed 2 other packages, tidygraph and particles, that he built to bring graph and network data into the tidyverse, the very popular collection of R packages designed for data science.
  • (25:37) Thomas provided resources for R users who want to learn more about network analysis and network visualization.
  • (27:05) Thomas went over his job as a data scientist at SKAT, where he handled all the advanced analytics going on in the Danish Tax Authorities.
  • (32:15) Thomas talked about the intuition behind working on patchwork, a package that can combine multiple ggplots in the same graphics.
  • (35:50) Thomas summarized his most recent projects, gganimate (a package that extends ggplot2 to include the description of animation) and tweenr (a package for interpolating data mainly for animations).
  • (40:47) Thomas discussed his current job as a software engineer at RStudio.
  • (43:04) Thomas gave his two cent on the Python and R comparison.
  • (45:53) Thomas talked about using Twitter to share his work, where he has more than 10,000 followers.
  • (47:07) Thomas went over something he works on during his spare time, generative art visualization (check his Instagram account!).
  • (49:09) Thomas gave some thoughts on the tech community in Copenhagen.
  • (51:09) Closing segments.

His Contact Info:

  • Website
  • Twitter
  • GitHub
  • LinkedIn

His Recommended Resources:

  • Hilary Parker from Stitch Fix
  • David Robinson from DataCamp
  • Unflattening Comic Book
  • DataCamp's Network Analysis in R
  • Katya Ognyanova's Network Analysis and Visualization with R and igraph

View Details

Show Notes:

  • (2:00) Ewan recalled his undergraduate days at the University of Edinburgh studying Physics.
  • (2:43) Ewan talked about his first job out of school as a Seismic Acquisition Specialist at ION Geophysical.
  • (5:28) Ewan briefly skimmed through his next job as a Senior Research Executive in Marketing Science at BBC.
  • (7:40) Ewan gave some key insights on the intersection between marketing and data.
  • (10:15) Ewan introduced Skyscanner and the interesting data problem that he is helping to solve.
  • (12:39) Ewan went over a couple of major data science use cases at Skyscanner.
  • (18:05) Ewan discussed how Skyscanner builds its travel recommendation system.
  • (23:40) Ewan emphasized the importance of doing A/B testing experiments at Skyscanner.
  • (27:58) Ewan talked about the differences in doing experimentation on both Skyscanner's web app and the mobile app.
  • (29:31) Ewan gave a long answer on how a data scientist can effectively communicate the results to others in an organization.
  • (33:07) Ewan gave his two cents on possible ways a data science team can be situated in a company.
  • (36:40) Ewan shared the 3 most important software engineering skills that a data scientist can develop.
  • (38:25) Ewan stressed the importance of learning SQL.
  • (40:25) Ewan mentioned the important signals to look for in a data-centric company.
  • (43:01) Ewan shared more information about the data science and tech ecosystem in Scotland.
  • (45:40) Ewan gave advice for aspiring data scientists who want to succeed during a data science interview.
  • (49:45) Ewan gave advice for business executives who want to bring data talents into their organizations.
  • (51:00) Closing segments.

His Contact Info:

  • Website
  • Twitter
  • LinkedIn

His recommended resources:

  • Airbnb Data Science
  • Stitch Fix Data Science
  • Ben Goldacre’s Bad Science

View Details

Show Notes:

  • (2:05) Chris recalled his Econometrics and Quantitative Economics study during his undergraduate days.
  • (4:14) Chris talked about his research work focus on oil and gas at the Center for Energy Studies at Louisiana State University.
  • (5:53) Chris gave some insights on how data science and econometrics can solve problems in the energy industry.
  • (7:51) Chris talked about his decision to pursue a Masters in Applied Statistics.
  • (9:02) Chris emphasized the importance of learning statistical theory. In particular, survival analysis is very similar to conversion rate analysis.
  • (10:28) Chris discussed his Master’s thesis work in analyzing the churn rate for Treehouse.
  • (12:05) Chris maintained his own website called Statwonk.
  • (13:18) Chris mentioned the popularity of count data on consumer products.
  • (14:40) Chris gave some recommendations for those who want to learn fundamental statistics: Think Stats, Think Bayes, Statistical Inference, and Statistical Intervals.
  • (16:20) Chris recalled how he got a job as the first data scientist at Treehouse.
  • (16:52) Chris described Treehouse CEO Ryan Carson as a “marketing genius.”
  • (18:34) Chris spotted the number one most challenging aspect for a data scientist within a business - translating technical jargon to comprehensible concepts.
  • (20:11) Chris recalled building a company dashboard from scratch for Treehouse.
  • (22:46) Chris shared the background overview of his next employer, Zapier.
  • (24:25) Chris shared his favorite things working as a data scientist as Zapier.
  • (32:07) Chris recalled building an AnamolyBot to keep track of Slack communication between Zapier team members.
  • (36:33) Chris talked about the open-source project he has been working on during his spare time.
  • (39:10) Chris gave advice for people who want to seek remote positions.
  • (42:20) Closing segments.

His Contact Info:

  • Website
  • Twitter
  • LinkedIn
  • GitHub

His recommended resources:

  • Shopify’s Cameron-Davidson Pilon
  • Stitch Fix’s Kim Larsen
  • R for Data Science
  • Hadley Wickham
  • John Tukey’s “Exploratory Data Analysis”
  • William Cleveland’s “The Elements of Graphing Data”
  • rstats

  • PyData

View Details

Show Notes:

  • (3:20) Saurabh recalled his college experience.
  • (4:25) Saurabh talked about his first role out of school as a software engineer specializing in database at CA Technologies.
  • (8:06) Saurabh landed database consulting roles with different companies.
  • (9:14) Saurabh gave insights on the differences between database systems now and a decade ago.
  • (11:05) Saurabh shared his experience landing a senior data scientist job with Barnes & Noble.
  • (13:15) Saurabh explained major challenges in hiring good data scientists.
  • (16:10) Saurabh discussed his decision to go work for Rent The Runway.
  • (17:57) Saurabh gave insights on the data problems he had worked with at Rent The Runway.
  • (19:43) In reference to his blog post on scaling machine learning at RTR, Saurabh shared knowledge on structuring a data science team.
  • (21:36) Saurabh gave advice for data scientists to incorporate feedback loops into their workflow.
  • (26:00) Saurabh talked about how to give better pitches to business stakeholders.
  • (29:03) Saurabh showed great appreciation for Rent The Runway’s CEO and Founder, Jennifer Hyman.
  • (30:39) In his post Human in AI loop, Saurabh shared the details of building Deep Dress AI, which can show the Instagram photos on Rent The Runway’s product pages and help the customers find fashion inspiration through these high-quality shots.
  • (34:10) Saurabh talked about unique challenges of doing data science in the retail space.
  • (36:18) Saurabh gave a brief glimpse about his stealth-mode startup Virevol.
  • (36:58) Saurabh gave advice for data scientists who want to pursue the entrepreneurial route and start their own ventures.
  • (38:36) In a fun blog post, Saurabh used Generative Adversarial Network to generate dresses photos and use them as training data for the model. He discussed the potential usage of GAN models for fashion.
  • (41:02) Closing segments.

His Contact Info:

  • Website
  • Twitter
  • LinkedIn

His recommended resources:

  • Salesforce AI
  • Elucd
  • Nate Silver's "The Signal and The Noise"
  • Deep Learning Textbook by Ian Goodfellow, Yoshua Bengio and Aaron Courville

View Details

Show Notes:

  • (3:08) Leni explains how studying Theatre Critics in college help her develop analytical thinking.
  • (4:00) Leni briefly her first job out of college in brand management at a lifestyle brand called Life is Porno.
  • (5:54) Leni goes over the most interesting classes she took during her Master of Arts in New Media Studies at Charles University.
  • (7:52) Leni talks about her 1-year residency in the social media team of New Media department at Czech TV.
  • (12:02) Leni reflects on the career lesson she got out from her time at Czech TV.
  • (13:47) Leni discusses her first professional role working with data at Socialbakers, a social media marketing company in Prague.
  • (16:35) Leni talks about choosing R as her main programming language to learn data science.
  • (19:43) Leni explains the study of social network analysis.
  • (21:50) Leni discusses the technical aspects of her Master’s thesis about Czech journalists on Twitter.
  • (25:00) Leni goes over her brief stint as a marketing specialist at Zeleznakoule, a gym in Prague and a community of health enthusiasts.
  • (26:57) Leni talks about her next career move to Seznam, the most visited web portal and search engine in the Czech Republic.
  • (29:11) Leni also volunteers as a mentor for Czechitas, a non-profit organization focused on educating girls about programming, data analysis and graphics.
  • (30:17) Leni describes the data science community in Prague.
  • (32:51) Leni talks about her passion for the tech scene in Stockholm.
  • (35:22) Leni talks about her popular Facebook page Dataholka.
  • (36:45) Leni shares some tips for aspiring data scientists who want to improve their social media presence.
  • (39:00) Leni gives recommendations for social media users regarding concerns about data privacy.
  • (41:30) Leni shares her thoughts on the current state of digital journalism.
  • (43:48) Leni discusses her current plan to pursue a Ph.D. degree in the US.
  • (47:52) Closing segment.

Her Contact Info:

  • Website
  • Twitter
  • Facebook
  • Instagram

Her recommended resources:

  • Netflix Tech Blog
  • Spotify Labs

View Details

Show Notes:

  • (1:58) Deep gives a brief background of his career.
  • (3:42) Deep explains the discipline of civil engineering.
  • (4:24) Deep talks about his first job out of college as a programmer analyst at Cognizant in Pune, India.
  • (5:26) Deep then became a senior software engineer at the investment bank Citigroup in Mumbai, India.
  • (7:30) Deep discusses his next role as a senior associate for the e-commerce company Sapient in Bangalore, India.
  • (8:57) Deep then went to Barclays Investment Bank back in Pune and worked as an analyst, in which he learned a lot about big data technologies.
  • (10:07) Deep talks about the diverse Indian cultures as he had experience working in many different regions in India.
  • (12:22) Deep recounts the story that propelled him to explore study opportunities in the United States.
  • (16:10) Deep gives practical advice to international students who want to pursue a graduate degree in the US.
  • (17:53) Deep talks about the Machine Learning and Deep Learning classes he took while at Galvanize.
  • (19:27) As a graduate student researcher, Deep built an image captioning system for a startup named FoxType, which helps non-native English speaker write better emails. He goes over the challenges in building this system.
  • (24:55) As a deep learning research intern for Netomi.com, Deep built a system with memory using neural network to solve the shortcoming of understanding long-term conversation and context for conversational agent. He discusses the process he went through.
  • (27:59) Deep talks about his machine learning project in which he built a sentiment analysis engine using the 300k reviews from Trip Advisor.
  • (31:33) Deep talks about another data engineering project in which he built an end-to-end real-time mapping of geo-locations in San Francisco for Uber Price Surge.
  • (37:32) Deep recommend the big data tools and technologies that all data scientists should know about.
  • (39:27) Deep shares his job search experience and how he landed a role at Nvidia.
  • (42:02) Deep provides helpful advice to crack the machine learning interview (aka, learning the fundamentals + networking).
  • (45:14) Deep talks about the autonomous vehicle system he’s working on at Nvidia.
  • (47:20) Deep quickly goes over Nvidia’s company culture.
  • (49:32) Deep digs deep into computer vision problems.
  • (51:34) Deep predicts the future applications of deep learning.
  • (53:40) Closing segments.

His contact info:

  • Twitter
  • LinkedIn
  • GitHub

His recommended resources:

  • Google Brain
  • Facebook AI Research
  • OpenAI
  • Salesforce Research
  • Andrej Karpathy's Blog
  • Off the Convex Path
  • Nvidia AI Research Blog
  • Google AI Blog
  • Facebook Research Blog
  • Arxiv Sanity Presenter

View Details

Show Notes:

  • (2:10) Jon reflects on his academic career and explained why he pursued an advanced degree in Biology. He also went over his post-doc fellowship at Cancer Research UK.
  • (4:45) Jon discusses the zebra-fish project he worked on during his post-doc research at King’s College London.
  • (8:13) Jon talks about the developmental biology work he did at the embryology lab at University College London.
  • (10:00) Jon explains why biological data are noisy.
  • (10:58) Jon explains his motivation behind his transition from academia to industry.
  • (15:04) Jon’s advice for Ph.D. students and post-doc researchers who want to make a similar transition.
  • (16:20) Jon talks about his freelance data science project to build a recommendation systems aiming at women struggling with Type 2 diabetes.
  • (21:32) Jon talks about his experience running his data science consultancy business.
  • (25:32) Jon’s advice to improve data storytelling skill.
  • (26:59) Jon presents a high-level overview of his company Pivigo.
  • (31:11) Jon talks about his role as Pivigo’s Head of Data Science.
  • (34:50) Jon goes over the industries that are Pivigo’s clients.
  • (36:30) Jon goes over the skills that data science talents need to develop.
  • (37:53) Jon’s advice to transition from individual contributor to manager.
  • (41:51) Jon mentions the talks he gave recently at O’Reilly Strata Data Conference in London.
  • (46:48) Closing segment.

His contact info:

  • Twitter
  • LinkedIn
  • GitHub

His recommended resources:

  • Datacamp
  • David Robinson's "Understanding empirical Bayes estimation"
  • "R for Data Science" by Hadley Wickham and Garrett Grolemund
  • rstats