Datacast follows the narrative journey of data practitioners and researchers to unpack the career lessons they learned along the way. James Le hosts the show.
Show Notes* (01:41) Salma reflected on her upbringing in Paris (France) and her love for math since a young age. * (04:05) Salma talked about her decision to dive deeper into statistics. * (07:39) Salma recalled her 5 years in investment banking in Hong Kong. * (10:44) Salma doubled down on the cultivation of her resilience during this period. * (13:45) Salma shared the founding story Sifflet. * (15:53) Salma touched on her co-founder dynamics with Wajdi Fathallah and Wissem Fathallah. * (18:50) Salma explained the concept of Full Data Stack Observability to the uninitiated. * (22:14) Salma extrapolated on Sifflet's approach to data observability. * (24:37) Salma discussed Sifflet's data quality focus with automated monitoring coverage and over 50 data quality templates. * (27:28) Salma discussed Sifflet's data lineage solution with field-level lineage, root cause analysis, and incident management/business impact assessment. * (31:47) Salma discussed Sifflet's data catalog with a powerful metadata search engine and centralized documentation for all data assets. * (33:46) Salma expanded on the full-stack mindset as a data vendor. * (37:03) Salma emphasized the importance of integrations with other data tools. * (38:52) Salma touched on product features such as Flow Stopper to stop vulnerable pipelines from running at the orchestration layer and Metrics Observability to extend the observability framework to the semantic layer. * (43:40) Salma unpacked her article on building a modern data team. * (47:54) Salma shared some hiring lessons to attract the right people to Sifflet. * (54:02) Salma emphasized the importance of company branding. * (55:01) Salma shared her thoughts on building a startup culture. * (57:26) Salma briefly mentioned her process of working with design partners in the early stage. * (01:00:17) Salma emphasized her enthusiasm for the broader data community. * (01:02:05) Conclusion.
Salma's Contact Info* LinkedIn * Twitter * Medium
Sifflet's Resources* Website | LinkedIn | Twitter | Docs * Data Catalog | Data Quality Monitoring | Data Lineage | Integrations
Mentioned ContentPeople* Zhamak Dehghani (Creator of Data Mesh and Founder of NextData) * Benoit Dageville, Thierry Cruanes, and Marcin Zukowski (Founders of Snowflake)
Books* "The Hard Thing About Hard Things" (by Ben Horowitz) * "The Boys In The Boat" (by Daniel James Brown)
NotesMy conversation with Salma was recorded back in late 2022. Since then, I recommend checking out the launch of Sifflet AI Assistant and this blog post on 2024 data trends.
About the showDatacast features long-form, in-depth conversations with practitioners and researchers in the data community to walk through their professional journeys and unpack the lessons learned along the way. I invite guests coming from a wide range of career paths — from scientists and analysts to founders and investors — to analyze the case for using data in the real world and extract their mental models (“the WHY and the HOW”) behind their pursuits. Hopefully, these conversations can serve as valuable tools for early-stage data professionals as they navigate their own careers in the exciting data universe.
Datacast is produced and edited by James Le. For inquiries about sponsoring the podcast, email khanhle.1013@gmail.com.
Subscribe by searching for Datacast wherever you get podcasts, or click one of the links below:
If you’re new, see the podcast homepage for the most recent episodes to listen to, or browse the full guest list.
Show Notes* (01:53) Suresh went over his college experience studying Electronics Engineering at the National Institute of Technology Karnataka. * (04:35) Suresh recalled his 9-year engineering career at Sylantro Systems. * (08:47) Suresh talked about the origin of Apache Hadoop at Yahoo. * (11:05) Suresh dissected the high-level design architecture of the Hadoop Distributed File System (HDFS). * (15:36) Suresh reflected on his decision to become a co-founder of Hortonworks, which focused on bringing Hadoop training and support to enterprise customers. * (17:36) Suresh unpacked the evolution of the Hortonworks Data Platform - which includes Hadoop technology such as HDFS, MapReduce, Pig, Hive, HBase, ZooKeeper, and additional components. * (20:30) Suresh shared his lessons from developing and supporting open-source software designed to manage big data processing. * (23:43) Suresh walked through the evolution of Uber’s Data Platform. * (28:03) Suresh described Uber's journey toward better data culture from first principles. * (34:00) Suresh explained his motivation to start the OpenMetadata Project. * (37:21) Suresh elaborated on OpenMetadata's five design principles: schema-first, extensibility, API-centric, vendor-neural, and open-source. * (40:17) Suresh highlighted OpenMetadata's built-in features to power multiple applications, such as data collaboration, metadata versioning, and data lineage. * (44:38) Suresh emphasized his priority for the open-source roadmap to adapt to the community's needs. * (47:05) Suresh explained the architecture of OpenMetadata - which goes deep into the push-based and pull-based characteristics of metadata ingestion and consumption. * (51:47) Suresh shared the long-term vision of his new company Collate, which powers the OpenMetadata initiative. * (53:36) Suresh shared valuable hiring lessons as a startup founder. * (56:30) Suresh shared fundraising advice to founders who want to seek the right investors for their startups. * (57:50) Closing segment.
Suresh's Contact Info* LinkedIn * Twitter * GitHub
OpenMetadata's Resources* Website | Twitter * Slack | GitHub | Community * Documentation * Collate
Mentioned ContentPeople* Joe Littlejohn (jsonschema2pojo) * Sriharsha Chintalapani
Book* The Innovator's Dilemma (by Clayton Christensen)
About the showDatacast features long-form, in-depth conversations with practitioners and researchers in the data community to walk through their professional journeys and unpack the lessons learned along the way. I invite guests coming from a wide range of career paths — from scientists and analysts to founders and investors — to analyze the case for using data in the real world and extract their mental models (“the WHY and the HOW”) behind their pursuits. Hopefully, these conversations can serve as valuable tools for early-stage data professionals as they navigate their own careers in the exciting data universe.
Datacast is produced and edited by James Le. For inquiries about sponsoring the podcast, email khanhle.1013@gmail.com.
Subscribe by searching for Datacast wherever you get podcasts, or click one of the links below:
If you’re new, see the podcast homepage for the most recent episodes to listen to, or browse the full guest list.
Show Notes (02:20) Krishna described his academic experience getting an MS in Computer Science from the University of Minnesota - where he developed efficient tools for text document clustering and pattern discovery. * (05:36) Krishna recalled his 5.5 years at Microsoft working on Bing's search engine. * (08:32) Krishna talked about the challenges of competing against Google Search. * (10:22) Krishna shared the high-level technical and operational challenges encountered during the development and scaling phase of Twitter Search. * (14:55) Krishna revealed vital lessons from building critical data infrastructure at Twitter. * (17:54) Krishna touched on his time at Pinterest as the head of data engineering - leading a team working on all things data from analytics, experimentation, logging, and infrastructure. * (20:05) Krishna reviewed the design and implementation of real-time analytics, ETL-as-a-Service, and an A/B testing platform at Pinterest. * (24:40) Krishna unpacked the major ML model performance issues while running Facebook's feed ranking platform. * (28:18) Krishna distilled lessons learned about algorithmic governance from Facebook. * (31:38) Krishna provided leadership lessons from building teams that create scalable platforms and delightful consumer products on Twitter, Pinterest, and Facebook. * (33:19) Krishna shared the founding story of Fiddler AI, whose mission is to build trust into AI. * (37:56) Krishna unpacked the key challenges and tools in his 2019 article "AI needs a new developer stack." * (40:49) Krishna discussed the evolution of MLOps over the past 4 years. * (42:48) Krishna explained the benefits of using the Model Performance Management (MPM) framework to address enterprise MLOps challenges. * (47:01) Krishna gave a brief overview of capabilities within Fiddler's MPM platform, such as model monitoring, explainable AI, analytics, and fairness. * (50:28) Krishna highlighted research efforts inside Fiddler concerning explainability, drift metric calculation, and fairness. * (53:17) Krishna discussed the challenges with monitoring for NLP and Computer Vision models. * (57:18) Krishna zoomed in on Fiddler's approach to model governance for the modern enterprise. * (01:02:24) Krishna distilled v*aluable lessons learned to attract the right people who are excited about Fiddler's mission and aligned with Fiddler's culture. * (01:06:08) Krishna reflected on the evolution of Fiddler's company culture. * (01:09:19) Krishna shared the challenges of finding the early design partners and defining a new category of Responsible AI. * (01:12:23) Krishna gave fundraising advice to founders who are seeking the right investors for their startups. * (01:14:45) Closing segment.
Krishna's Contact Info* LinkedIn * Twitter * Medium
Fiddler's Resources* Website | LinkedIn | Twitter | YouTube * About | Customers | Careers * AI Observability | Model Monitoring | Explainable AI | Fairness | Analytics * Blog | Docs | Resources
Mentioned ContentPeople1. Goku Mohamandas (Made With ML and Anyscale) 2. Krishnaram Kenthapadi (Chief AI Officer & Chief Scientist at Fiddler)
Books1. "The Hard Thing About Hard Things" (Ben Horowitz) 2. "The Five Dysfunctions of A Team" (Patrick Lencioni)
NotesMy conversation with Krishna was recorded more than a year ago. Since then, I'd recommend checking out these Fiddler's resources:
About the showDatacast features long-form, in-depth conversations with practitioners and researchers in the data community to walk through their professional journeys and unpack the lessons learned along the way. I invite guests coming from a wide range of career paths — from scientists and analysts to founders and investors — to analyze the case for using data in the real world and extract their mental models (“the WHY and the HOW”) behind their pursuits. Hopefully, these conversations can serve as valuable tools for early-stage data professionals as they navigate their own careers in the exciting data universe.
Datacast is produced and edited by James Le. For inquiries about sponsoring the podcast, email khanhle.1013@gmail.com.
Subscribe by searching for Datacast wherever you get podcasts, or click one of the links below:
If you’re new, see the podcast homepage for the most recent episodes to listen to, or browse the full guest list.
Show Notes* (01:59) Emanuel reflected on his upbringing in Switzerland and his 4-year apprenticeship in Software Engineering and Business at Credit Suisse. * (06:09) Emanuel recalled his 4-year program in Computer Science at HSR (University of Applied Sciences Rapperswil). * (08:43) Emanuel touched on his decision to pursue a Master’s degree at Brown University in the US. * (11:58) Emanuel explained his decision to continue with a Ph.D. degree at Brown under the advisement of Professor Andy van Dam and summarized the arc of his Ph.D. research focus. * (14:50) Emanuel highlighted his first research paper called PanoramicData on Interactive Data Exploration. * (16:42) Emanuel shared his thoughts on common traits of a successful researcher. * (18:03) Emanuel emphasized the focus of his research at the intersection of Human-Computer Interaction, Information Visualization, and Data Analysis. * (20:38) Emanuel shared valuable lessons from interning twice at Microsoft Research in Redmond. * (22:34) Emanuel talked about his time as a postdoc in Professor Tim Kraska’s group at the CSAIL at MIT. * (24:59) Emanuel shared the founding story of Einblick - a visual computing platform that enables data teams to answer tougher, more meaningful questions by making advanced analytics and model building more streamlined and accessible. * (29:06) Emanuel touched on the responsibilities of Einblick's 5 co-founders. * (30:23) Emanuel highlighted technical challenges of building Einblick's integrated environment for descriptive, predictive, and prescriptive analytics. * (32:09) Emanuel mentioned the collaboration challenge in data and brought up Einblick's real-time remote collaboration through video-enabled data whiteboards. * (35:39) Emanuel highlighted the challenges of working with computational notebooks and brought up the benefits of using Einblick's collaborative visual canvas. * (38:54) Emanuel unpacked the challenges of commercializing an academic research project. * (40:27) Emanuel gave a broad overview of Einblick's go-to-market strategy. * (43:55) Emanuel shared valuable hiring lessons to attract the right people who are aligned with Einblick’s cultural values. * (47:05) Emanuel shared fundraising advice to founders who are seeking the right investors for their startups. * (49:20) Emanuel shared the similarities and differences between being a researcher and being a founder. * (50:43) Closing segment.
Emanuel's Contact Info* Website * Google Scholar * LinkedIn * Twitter
Einblick's Resources* Website | Twitter | LinkedIn * Docs | Blog * ChartGen AI * Notebook Feature Release (2022) * Video-Based Collaboration Release (2021)
Mentioned ContentPapers and Projects* PanoramicData is a hybrid pen and touch system for visual data exploration (Infovis 2014 Paper | Video) * (s|qu)eries (pronounced “Squeries”) is a visual query interface for creating queries on sequences (series) of data based on regular expressions (CHI 2015 Paper | Summary Video) * Vizdom is an interactive visual analytics system that scales to large datasets through progressive computation (VLDB Demo 2015 Paper | Health Video | Election Video) * Tableur is a spreadsheet-like pen- and touch-based system that revolves around handwriting recognition - all data is represented as digital ink (CHI 2016 LBW Paper | Video) * Towards Accessible Data Analysis (Emanuel's Ph.D. Dissertation at Brown, 2018) * Northstar is an interactive data science platform that combines data exploration with automated machine learning (SIGMOD DEEM Paper | Video)
People* Wes McKinney * Fei-Fei Li
Books* "The Book of Why" (by Judea Pearl) * "The Signal and The Noise" (by Nate Silver)
NotesMy conversation with Emanuel was recorded back in late 2022. Since then, I recommend checking out the launch of Einblick Prompt and ChartGenAI.
About the showDatacast features long-form, in-depth conversations with practitioners and researchers in the data community to walk through their professional journeys and unpack the lessons learned along the way. I invite guests coming from a wide range of career paths — from scientists and analysts to founders and investors — to analyze the case for using data in the real world and extract their mental models (“the WHY and the HOW”) behind their pursuits. Hopefully, these conversations can serve as valuable tools for early-stage data professionals as they navigate their own careers in the exciting data universe.
Datacast is produced and edited by James Le. For inquiries about sponsoring the podcast, email khanhle.1013@gmail.com.
Subscribe by searching for Datacast wherever you get podcasts, or click one of the links below:
If you’re new, see the podcast homepage for the most recent episodes to listen to, or browse the full guest list.
About the showDatacast features long-form, in-depth conversations with practitioners and researchers in the data community to walk through their professional journeys and unpack the lessons learned along the way. I invite guests coming from a wide range of career paths — from scientists and analysts to founders and investors — to analyze the case for using data in the real world and extract their mental models (“the WHY and the HOW”) behind their pursuits. Hopefully, these conversations can serve as valuable tools for early-stage data professionals as they navigate their own careers in the exciting data universe.
Datacast is produced and edited by James Le. For inquiries about sponsoring the podcast, email khanhle.1013@gmail.com.
Subscribe by searching for Datacast wherever you get podcasts, or click one of the links below:
If you’re new, see the podcast homepage for the most recent episodes to listen to, or browse the full guest list.
Timestamps* (01:41) Diana shared her upbringing in Orlando and her undergraduate experience studying Economics at MIT. * (03:27) Diana reflected on valuable lessons from her internships during college. * (05:58) Diana brought up her learning during the year working as an investment banking analyst in the technology group at Morgan Stanley. * (09:10) Diana recalled her transition from investment banking into venture capital at Norwest Venture Partners. * (11:52) Diana discussed the takeaways from her time as a venture associate meeting entrepreneurs on a regular cadence. * (14:23) Diana recalled her decision to leave venture capital and become the first hire at the product organization at Cockroach Labs. * (19:00) Diana went over the challenges and learning curves as a non-technical first product hire at Cockroach. * (21:26) Diana extrapolated on the idea of determining the best-fit product strategy rather than blindly following frameworks. * (23:55) Diana described her experience as the first hire into the product organization at TimescaleDB. * (26:40) Diana highlighted the challenges of open-source GTM. * (27:56) Diana reflected on her 3-part blog series on building a product for the most dissatisfied customers first, the majority next, and the full need in the long run. * (30:40) Diana shared 2 tactical lessons to cultivate focus as a product manager. * (32:44) Diana share the founding story of Correlated alongside her co-founders, Tim Geisenheimer and John Pena. * (35:38) Diana briefly touched on her 2 entrepreneurial attempts during COVID-19. * (38:30) Diana unpacked the notion of Product-Led Revenue and described how Correlated works at a high level. * (40:49) Diana highlighted the role of integrations within Correlated's product strategy. * (43:04) Diana mentioned Correlated's product-led playbooks to help users manage their product-led strategy from start to finish. * (45:40) Diana explained how she leveraged customer feedback to ship the feature called PQL Scoring that leverages machine learning to identify the best leads. * (48:51) Diana shared the consistent principles that have remained the same for successful communication in Product Management. * (52:30) Diana discussed her learnings on customer discovery at early-stage startups. * (56:08) Diana reflected on the early signs of product-market fit that carry through all of her startups. * (58:58) Conclusion
Diana's Contact Info* LinkedIn * Twitter * Medium * Substack
Correlated's Resources* Website | LinkedIn | Twitter * Product Overview | How Correlated Works * Blog | Podcast | Docs * PLG Playbook Library * Correlated Launches to Bring Product-Led Revenue to Market with $8.3M in Funding * What Is Product-Led Revenue? * Correlated launches PQL Scoring to accelerate your product-led strategy
Mentioned ContentBlog Posts* "The Standard Due Diligence Process" (Jan 2016) * "Mistakes to Avoid when Pitching to a VC" (Jan 2016) * "My Startup Litmus Test" (Feb 2016) * "Why I left VC to join Cockroach Labs" (April 2017) * "My First 90 Days as the First Product Hire" (May 2017) * "Coding != Technical: What It Means to be Technical as a PM" (Aug 2017) * "How learning to sell makes for a better product manager" (Nov 2017) * "Roadmap Planning: Users First, Features Second" (March 2018) * "Build something people will use more than once" (May 2019) * "Focus on the unhappiest, most dissatisfied customers first" (May 2019) * "Build for the majority" (May 2019) * "Why user interviews can fail you when starting a startup" (Sep 2021) * "Tackling the challenges of communicating effectively in product management" (Jan 2022) * "Some Learnings on Customer Discovery at Early-Stage Startups" (May 2022) * "4 early signs of product-market fit" (Sep 2022) * "Give customers what they want, but not what they ask for" (Sep 2022)
People1. Lenny Rachitsky 2. Julie Zhuo 3. Nate Stewart 4. Jeff Sposetti
Book* "Crossing The Chasm" (by Geoffrey Moore)
About the showDatacast features long-form, in-depth conversations with practitioners and researchers in the data community to walk through their professional journeys and unpack the lessons learned along the way. I invite guests coming from a wide range of career paths — from scientists and analysts to founders and investors — to analyze the case for using data in the real world and extract their mental models (“the WHY and the HOW”) behind their pursuits. Hopefully, these conversations can serve as valuable tools for early-stage data professionals as they navigate their own careers in the exciting data universe.
Datacast is produced and edited by James Le. For inquiries about sponsoring the podcast, email khanhle.1013@gmail.com.
Subscribe by searching for Datacast wherever you get podcasts, or click one of the links below:
If you’re new, see the podcast homepage for the most recent episodes to listen to, or browse the full guest list.
About the showDatacast features long-form, in-depth conversations with practitioners and researchers in the data community to walk through their professional journeys and unpack the lessons learned along the way. I invite guests coming from a wide range of career paths — from scientists and analysts to founders and investors — to analyze the case for using data in the real world and extract their mental models (“the WHY and the HOW”) behind their pursuits. Hopefully, these conversations can serve as valuable tools for early-stage data professionals as they navigate their own careers in the exciting data universe.
Datacast is produced and edited by James Le. For inquiries about sponsoring the podcast, email khanhle.1013@gmail.com.
Subscribe by searching for Datacast wherever you get podcasts, or click one of the links below:
If you’re new, see the podcast homepage for the most recent episodes to listen to, or browse the full guest list.
Timestamps* (01:45) Jason shared the formative experiences of his upbringing in the Bay Area and coming of age in the “Moneyball” era of baseball. * (05:03) Jason described his overall academic experience at Stanford - where he studied Mathematical and Computational Science with a minor in Classical Studies. * (09:15) Jason reflected on his experience participating in the Mayfield Fellowship at Stanford. * (12:03) Jason recalled his time being a part of the business operations team during a high-growth period at Opendoor. * (14:25) Jason talked about lessons learned working as a management consultant at McKinsey’s Bay Area practice. * (15:59) Jason reminisced about his time at the AI Fund startup studio - where he launched AI-enabled SaaS startups by iterating on prototypes, signing design partners, and recruiting the founding team. * (19:25) Jason explained his decision to join the investment team at Greylock Partners. * (22:24) Jason walked through his journey proving value as a new investor. * (24:41) Jason unpacked his checklist for evaluating early-stage enterprise investment opportunities. * (27:09) Jason explained his seed investment in Onehouse - a cloud-native managed lakehouse service that makes data lakes easier, faster, and cheaper. * (30:31 ) Jason explained his Series A investment in Baseten - which builds a powerful software toolkit that empowers technical data science teams to serve, integrate, design, and ship their custom ML models efficiently. * (33:23) Jason touched on advice for his portfolio companies in hiring decisions and navigating product/GTM strategy. * (37:00) Jason unpacked key takeaways from Greylock’s Castles in the Cloud project. * (39:58) Jason dissected key trends in the markets of security, AI/ML, management and governance, and edge computing (as shown in "VC Funding for the Cloud"). * (46:24) Jason elaborated on his vision of "The Next Cloud Data Platform" - which examines how the data warehouse, lakehouse, and semantic layer could combine to create a platform for data applications. * (50:55) Jason shared a few books that have greatly influenced his life. * (52:22) Closing segment.
Jason's Contact Info* LinkedIn * Twitter * Greylock
Mentioned ContentBooks1. "Moneyball" (by Michael Lewis) 2. "Why The West Rules For Now" (by Ian Morris) 3. "Snow Crash" (by Neil Stephenson) 4. "Cryptonomicon" (by Neil Stephenson) 5. "Termination Shock" (by Neil Stephenson) 6. "Principles for Dealing with the Changing World Order" (by Ray Dalio)
People1. David Luan (Founder and CEO of Adept) 2. Alex Ratner (Co-Founder and CEO of Snorkel AI) 3. Frank Slootman (CEO of Snowflake) 4. Clement Delangue (Co-Founder and CEO of HuggingFace)
NotesMy conversation with Jason was recorded back in late 2022. Since then, I recommend checking out these resources:
About the showDatacast features long-form, in-depth conversations with practitioners and researchers in the data community to walk through their professional journeys and unpack the lessons learned along the way. I invite guests coming from a wide range of career paths — from scientists and analysts to founders and investors — to analyze the case for using data in the real world and extract their mental models (“the WHY and the HOW”) behind their pursuits. Hopefully, these conversations can serve as valuable tools for early-stage data professionals as they navigate their own careers in the exciting data universe.
Datacast is produced and edited by James Le. For inquiries about sponsoring the podcast, email khanhle.1013@gmail.com.
Subscribe by searching for Datacast wherever you get podcasts, or click one of the links below:
If you’re new, see the podcast homepage for the most recent episodes to listen to, or browse the full guest list.
About the showDatacast features long-form, in-depth conversations with practitioners and researchers in the data community to walk through their professional journeys and unpack the lessons learned along the way. I invite guests coming from a wide range of career paths — from scientists and analysts to founders and investors — to analyze the case for using data in the real world and extract their mental models (“the WHY and the HOW”) behind their pursuits. Hopefully, these conversations can serve as valuable tools for early-stage data professionals as they navigate their own careers in the exciting data universe.
Datacast is produced and edited by James Le. For inquiries about sponsoring the podcast, email khanhle.1013@gmail.com.
Subscribe by searching for Datacast wherever you get podcasts, or click one of the links below:
If you’re new, see the podcast homepage for the most recent episodes to listen to, or browse the full guest list.
Show Notes* (01:44) Heather talked about her upbringing, her education, and her 14-year career at Liberty Mutual Insurance. * (06:50) Heather emphasized the benefits of her education in organizational design and management. * (08:56) Heather walked through her decision to shift from underwriting to technology and data at Liberty. * (13:55) Heather commented on her 5 years as a product manager at Crum & Forster. * (20:14) Heather described her 2-year experience as the Global Head of Technology at Innovisk (a Wills Towers Watson ). * (23:54) Heather distinguished the work environments in a startup and a large company. * (26:17) Heather recalled different data strategy initiatives she led at Brown and Brown Insurance. * (28:33) Heather explained issues in the insurance value chain and the role of data to help tackle them. * (34:17) Heather discussed her role as the founding Chief Data Officer at Accelerant Holdings. * (40:18) Heather brought up the data quality issues that Accelerant risk exchange helped solve. * (45:52) Heather gave advice to organizations to move from data governance to Data Intelligence. * (48:45) Heather provided her perspective on hiring data talent. * (50:40) Heather looked at the insurance transformation from a technology perspective. * (52:52) Heather talked about engaging women in technical fields. * (54:22) Closing segment.
Heather's Contact Info* LinkedIn * Accelerant | About
Relevant Reading* Bloomberg | Boehly'sHeather's Eldridge Bets on Accelerant at $2.2 Billion Valuation * Insurance Business Mag | Accelerant: An insurtech that defies categories * Business Insurance | Accelerant establishes $175 million sidecar reinsurer * AI Times Journal | Data Intelligence is Key to Understanding our Customers – Chief Data Officer, Accelerant Holdings * Lightco | Insurance Innovators Top 100 * Business Wire | Accelerant Launches the Accelerant Risk Exchange to Reimagine Insurance
Mentioned ResourcesPeople1. Zhamak Dehgani 2. Cassie Kozyrkov 3. Allie Miller
Book* Data Mesh: Delivering Data-Driven Value at Scale (by Zhamak Dehghani)
About the showDatacast features long-form, in-depth conversations with practitioners and researchers in the data community to walk through their professional journeys and unpack the lessons learned along the way. I invite guests coming from a wide range of career paths — from scientists and analysts to founders and investors — to analyze the case for using data in the real world and extract their mental models (“the WHY and the HOW”) behind their pursuits. Hopefully, these conversations can serve as valuable tools for early-stage data professionals as they navigate their own careers in the exciting data universe.
Datacast is produced and edited by James Le. For inquiries about sponsoring the podcast, email khanhle.1013@gmail.com.
Subscribe by searching for Datacast wherever you get podcasts, or click one of the links below:
If you’re new, see the podcast homepage for the most recent episodes to listen to, or browse the full guest list.
About the showDatacast features long-form, in-depth conversations with practitioners and researchers in the data community to walk through their professional journeys and unpack the lessons learned along the way. I invite guests coming from a wide range of career paths — from scientists and analysts to founders and investors — to analyze the case for using data in the real world and extract their mental models (“the WHY and the HOW”) behind their pursuits. Hopefully, these conversations can serve as valuable tools for early-stage data professionals as they navigate their own careers in the exciting data universe.
Datacast is produced and edited by James Le. For inquiries about sponsoring the podcast, email khanhle.1013@gmail.com.
Subscribe by searching for Datacast wherever you get podcasts, or click one of the links below:
If you’re new, see the podcast homepage for the most recent episodes to listen to, or browse the full guest list.
Show Notes* (01:36) Bob shared formative experiences of his upbringing with exposure to technology. * (05:08) Bob discussed his time building software for fun as a teenager. * (07:08) Bob reflected on his education in music and his decision to transition to a career in software development. * (10:52) Bob explained his project control(human, data, sound) while working on his consultancy agency Kubrickology. * (14:25) Bob discussed how using software can enhance our creativity. * (17:28) Bob talked about his fascination with merging physical and digital realms. * (19:41) Bob recalled his TEDx talk that introduced three high-level ideas about why software works well based on its ability to adapt to our language. * (24:57) Bob shared the founding story of Weaviate. * (29:40) Bob talked about his process of choosing his co-founders. * (31:29) Bob unpacked his high-level thinking around creating a business model around the open-source project. * (38:06) Bob defined a vector search engine for the uninitiated. * (40:49) Bob gave a brief overview of the high-level design of Weaviate. * (43:15) Bob talked about Weaviate's production-ready features, such as horizontal scalability and graph-like connections between objects. * (45:45) Bob reviewed the use cases for Weaviate that he is most proud of. * (49:31) Bob emphasized the importance of engaging open-source contributors to generate valuable product feedback. * (55:03) Bob talked about the pricing model for Weaviate Cloud Service. * (57:59) Bob anticipated the evolution of the tooling landscape within the AI-first database ecosystem to support the increasing adoption of unstructured data. * (01:02:31) Bob shared valuable hiring lessons to attract the right people to join Weaviate. * (01:04:36) Bob explained his process of identifying people who align with the cultural values of Weaviate. * (01:08:27) Bob gave fundraising advice to founders who are seeking the right investors for their startups. * (01:12:17) Bob highlighted his thinking around being a remote-first company and building an open-source brand. * (01:16:28) Closing segment.
Bob's Contact Info* Wikipedia * LinkedIn * Twitter * GitHub * WTF Medium Blog * YouTube
Weaviate's Resources* Website | Twitter | Slack | Forum | GitHub * Blog | Podcast | Playbook
Mentioned ContentPeople1. Sam Ramji (DataStax) 2. Paul Graham (Y Combinator)
Book* "Hackers and Painters" (by Paul Graham)
NotesMy conversation with Bob was recorded back in late 2022. Since then, I recommend checking out these resources:
About the showDatacast features long-form, in-depth conversations with practitioners and researchers in the data community to walk through their professional journeys and unpack the lessons learned along the way. I invite guests coming from a wide range of career paths — from scientists and analysts to founders and investors — to analyze the case for using data in the real world and extract their mental models (“the WHY and the HOW”) behind their pursuits. Hopefully, these conversations can serve as valuable tools for early-stage data professionals as they navigate their own careers in the exciting data universe.
Datacast is produced and edited by James Le. For inquiries about sponsoring the podcast, email khanhle.1013@gmail.com.
Subscribe by searching for Datacast wherever you get podcasts, or click one of the links below:
If you’re new, see the podcast homepage for the most recent episodes to listen to, or browse the full guest list.
Show Notes* (01:40) Sakib shared formative experiences of his upbringing in SoCal and his undergraduate experience at the University of Pennsylvania. * (05:24) Sakib recalled his favorite classes at Penn. * (07:55) Sakib reflected on his internship experience at Innova Dynamics and Morgan Stanley. * (11:04) Sakib reflected on his decision to pursue a career in venture capital at Bessemer Venture Partners. * (14:02) Sakib walked through his process of proving value as a new investor. * (16:21) Sakib explained his process of forming clear investment theses. * (18:35) Sakib talked about his brief year as a product manager at Viagogo before going back to Bessemer. * (22:04) Sakib dissected his investments in LaunchDarkly and PagerDuty (in the domain of developer-centric platforms). * (24:16) Sakib explained his investments in Coiled, Prefect, and Arcion Labs (in the domain of data infrastructure). * (25:59) Sakib walked through his investment in Guild Education and Tribe (in the domain of education and community management). * (29:06) Sakib shared advice to portfolio companies in terms of navigating hard decisions and growth strategy. * (34:18) Sakib outlined Bessemer's roadmap on data infrastructure - which looks at the wave of startups enabling the next generation of data-driven businesses. * (39:57) Sakib brought up the products that help abstract away complexity from data engineering problems. * (41:34) Sakib highlighted the tools that power the next generation of data scientists. * (43:27) Sakib emphasized the emergence and evolution of metadata management * (46:48) Sakib unpacked the evolution of ML infrastructure. * (49:28) Sakib examined the key trends and opportunities that will define the next wave of BI and data analytics software. * (52:56) Sakib shared his investment perspectives on climate change and student builders. * (55:29) Closing segment.
Sakib's Contact Info* Profile Page * LinkedIn * Twitter
Mentioned ContentPeople1. Sarah Catanzaro (General Partner of Amplify Partners) 2. Ed Sim (Founder of Boldstart Ventures) 3. Mike Speiser (Managing Partner of Sutter Hill Ventures)
Book* "The Idea Factory" (by Jon Gertner)
NotesMy conversation with Sakib was recorded back in late 2022. Since then, I recommend checking out these resources:
Show Notes* (01:56) Casber reflected on his experience growing up in China and moving to the US to pursue an undergraduate degree in business at UC Berkeley. * (04:45) Casber recalled working on his startup Etch.ai and interning at Wish during his time at Berkeley. * (10:03) Casber differentiated investing patterns for B2B and B2C startups. * (12:10) Casber reflected on his investment banking experience at Bank of America Merrill Lynch and transitioning to venture capital at Sapphire Ventures. * (17:06) Casber gave some advice for analysts who want to transition into the tech and venture industry. * (20:06) Casber provided a high-level overview of Sapphire Ventures and its investment focus. * (21:54) Casber recalled his early days as a new investor and his process of adding value to portfolio companies. * (24:58) Casber dissected his investments in the Series F round of JumpCloud and the Series B round of Uptycs (in the domain of security). * (33:03) Casber explained his investments in the Series B round of Tetrate and the Series A of Zesty (in the domain of enterprise infrastructure). * (35:59) Casber walked through his investment in the Series D round of Dremio (in the domain of data and analytics). * (39:09) Casber shared his advice to his portfolio companies in terms of navigating hiring decisions and growth strategy. * (43:06) Casber unpacked the three strategies software companies can borrow from the open-source cloud playbook. * (46:42) Casber highlighted the key trends he is most bullish on in the Open Data Ecosystem. * (51:49) Casber emphasized the importance of interoperability in the modern data stack tooling landscape. * (54:15) Casber painted the modular future of AI infrastructure. * (59:44) Casber highlighted the key trends propelling the dynamic evolution of the software development lifecycle. * (01:03:06) Casber reflected on his learning process for any new industry as an investor. * (01:05:53) Closing segment.
Casber's Contact Info* Sapphire Ventures Profile * Twitter * LinkedIn
Mentioned ContentArticles1. 3 Strategies Software Companies Can Borrow from the Open-Source Cloud Playbook (Aug 2020) 2. What is the Open Data Ecosystem and Why It's Here to Stay (April 2021) 3. The Future of AI Infrastructure is Becoming Modular: Why Best-of-Breed MLOps Solutions are Taking Off and Top Players to Watch (March 2022) 4. Evolution of the Software Development Lifecycle and the Future of DevOps (June 2022)
Books1. "The Power Law" (by Sebastian Mallaby) 2. "Engines That Move Markets" (by Alasdair Naim)
NotesMy conversation with Casber was recorded back in late 2022. Since then, I recommend checking out these resources:
About the showDatacast features long-form, in-depth conversations with practitioners and researchers in the data community to walk through their professional journeys and unpack the lessons learned along the way. I invite guests coming from a wide range of career paths — from scientists and analysts to founders and investors — to analyze the case for using data in the real world and extract their mental models (“the WHY and the HOW”) behind their pursuits. Hopefully, these conversations can serve as valuable tools for early-stage data professionals as they navigate their own careers in the exciting data universe.
Datacast is produced and edited by James Le. For inquiries about sponsoring the podcast, email khanhle.1013@gmail.com.
Subscribe by searching for Datacast wherever you get podcasts, or click one of the links below:
If you’re new, see the podcast homepage for the most recent episodes to listen to, or browse the full guest list.
Show Notes* (01:46) Itai reflected on his education at The Hebrew University of Jerusalem, studying Math and Computer Science. * (04:18) Itai walked through his time as a software engineer at Google working in Google Trends. * (06:56) Itai emphasized the importance of a software checklist within Google's engineering culture. * (08:55) Itai explained how he became fascinated with AI/ML engineering. * (10:31) Itai touched on his period working as an AI consultant. * (13:28) Itai talked about his side hustle as a co-owner of Lia's Kitchen, a 100% vegan restaurant in Berlin. * (16:13) Itai shared the founding story of Mona Labs, whose mission is to make AI and machine learning impactful, effective, reliable, and safe for fast-growth teams and businesses. * (21:25) Itai unpacked the architecture overview of the Mona monitoring platform. * (24:50) Itai talked about the early days of Mona finding design partners. * (27:15) Itai dissected his perspective on a comprehensive monitoring strategy. * (31:42) Itai explained why the secret to comprehensive monitoring lies in granular tracking and avoiding noise. * (38:35) Itai explained how Mona can support real-time monitoring across the layers of the platform. * (43:18) Itai mentioned the integration with New Relic to display the variability of use cases for Mona. * (46:08) Itai discussed the shift for data science teams from being research-oriented to product-oriented. * (51:03) Itai provided four tactics for data science teams to become "product-oriented." * (58:04) Itai shared valuable hiring lessons to attract the right people who are excited about the mission of Mona Labs. * (01:01:46) Itai provided his mental model for finding exceptional engineering talent. * (01:03:38) Itai brought up again the importance of finding lighthouse customers. * (01:05:46) Itai gave his thoughts on building the product to satisfy different customer needs. * (01:07:55) Itai described the thriving ML engineering community in Israel. * (01:09:42) Closing thoughts
Itai's Contact Info* LinkedIn
Mona Labs' Resources* Website | LinkedIn | Twitter | YouTube * About | Customers | Careers * Platform * Blog | Case Studies | Docs
Mentioned ContentBlog Posts and Talks* We are building Mona to bring ML observability to production AI * The definitive guide to AI/ML monitoring * The secret to successful AI monitoring: Get granular, but avoid noise * Taking AI from good to great by understanding it in the real world (June 2022) * Data drift, concept drift, and how to monitor for them * The issues ML model retraining won't solve * Common pitfalls to avoid when evaluating an ML monitoring solution * Introducing automated exploratory data analysis powered by Mona * Best practices for setting up monitoring operations for your AI team * The challenges of specificity in monitoring AI * Is your LLM application ready for the public? * Overcoming cultural shifts from data science to prompt engineering
People1. Goku Mohandas (Creator of Made With ML) 2. Ville Tuulos (CEO and Co-Founder of Outerbounds) 3. Nimrod Tamir (CTO and Co-Founder of Mona Labs)
NotesMy conversation with Itai was recorded back in October 2022. Since then, Mona Labs has introduced a new self-service monitoring solution for GPT! Read Itai's blog post for the technical details.
Show Notes* (01:33) Gabi shared her professional interests growing up - from painting and drawing to graphic design. * (04:44) Gabi touched on her entrance to the field of data visualization. * (06:30) Gabi described her graduate school experience studying Data Visualization at Parsons School of Design. * (08:50) Gabi talked about the benefits of teaching data visualization classes later in her career. * (12:30) Gabi recalled working as a Data Visualization specialist at The Washington Post. * (14:50) Gabi gave her perspective on the evolution of data journalism. * (18:29) Gabi talked about her experience co-founding Raw Haus, a creative community bringing together emerging talent in design, technology, and entrepreneurship. * (21:20) Gabi emphasized the magic of community gatherings. * (23:51) Gabi reflected on her time at WeWork as a senior data visualization engineer to design and build graphics, dashboards, and tools that tell stories using data. * (28:02) Gabi walked through the evolution of the Data Cult initiative - which she co-created with Leah Weiss. * (31:44) Gabi recalled her decision to leave WeWork in early 2020 and start Data Culture - a data engineering and visualization consultancy focused on helping organizations build data capabilities, implement modern infrastructure and create lasting data culture. * (36:31) Gabi unpacked Data Culture's well-defined blueprint for each client engagement. * (39:21) Gabi explained how Data Culture leveraged tools in the modern data stack for its consulting services. * (40:42) Gabi brought up Data Culture's Studio - which offers data storytelling and visualization services to mission-aligned organizations. * (42:50) Gabi reviewed her experience working with Kode with Klossy to empower young scholars to solve important issues using data science. * (46:13) Gabi shared her perspective on how companies can scale their respective data cultures. * (49:52) Gabi shared the story behind the founding of Preql. * (52:59) Gabi touched on the process of working with design partners for Preql. * (56:42) Gabi shared valuable hiring lessons to attract the right people at Data Culture and Preql. * (58:10) Gabi provided her perspective on building a diverse team. * (01:00:57) Gabi shared fundraising advice to data founders who are seeking the right investors for their startups. * (01:03:34) Closing segment.
Gabi's Contact Info* LinkedIn * Twitter
Preql's Resources* Website | Twitter | LinkedIn * Introducing Preql: The Future of Data Transformation (April 2022)
Mentioned ResourcesPeople1. Umi Syam (Graphics and Multimedia Editor at the New York Times) 2. Giorgia Lupi (Information Designer and Partner at Pentagram) 3. Susie Lu (Senior Data Visualization Engineer at Netflix)
Book* Invisible Women: Exploring Data Bias in a World Designed for Men (by Caroline Criado Perez)
NotesMy conversation with Gabi was recorded back in August 2022. Since then, Preql has officially launched and currently supports strategy and operations teams at B2B and vertical SaaS companies!
About the showDatacast features long-form, in-depth conversations with practitioners and researchers in the data community to walk through their professional journeys and unpack the lessons learned along the way. I invite guests coming from a wide range of career paths — from scientists and analysts to founders and investors — to analyze the case for using data in the real world and extract their mental models (“the WHY and the HOW”) behind their pursuits. Hopefully, these conversations can serve as valuable tools for early-stage data professionals as they navigate their own careers in the exciting data universe.
Datacast is produced and edited by James Le. For inquiries about sponsoring the podcast, email khanhle.1013@gmail.com.
Subscribe by searching for Datacast wherever you get podcasts, or click one of the links below:
If you’re new, see the podcast homepage for the most recent episodes to listen to, or browse the full guest list.
Show Notes* (01:55) Alex reflected on his upbringing as an immigrant moving from Colombia to the US at 14. * (07:06) Alex recalled his undergraduate experience at NYU’s Polytechnic School of Engineering. where he study Computer Science and do research in cryptography. * (16:40) Alex went over his first job working as a software engineer at FactSet Research System. * (20:13) Alex walked through his time as the first employee and the first engineer at YieldMo. * (24:30) Alex talked about his hiring philosophy for engineers who care about their craft. * (28:03) Alex touched on the backstory behind the creation of Concord, with Shinji Kim and Robert Blafford, while working at YieldMo. * (32:26) Alex shared lessons learned from his first-time founder experience with Concord. * (35:22) Alex went over his two years at Akamai as a Platform Infrastructure Engineer after the Concord acquisition. * (40:01) Alex introduced his work on SMF, an RPC framework designed for microsecond tail latency. * (43:41) Alex shared the story behind the founding of Redpanda Data, which builds a high-performance, Apache Kafka-compatible data streaming platform for mission-critical workloads. * (47:19) Alex walked through the major benefits of choosing Redpanda over Kafka. * (51:03) Alex explained his decision to open-source Redpanda in November 2020 under the Source Available License BSL. * (56:08) Alex mentioned successful tactics his team employed in order to raise the adoption and contribution to the open-source library. * (01:01:13) Alex unpacked the design of Redpanda's Intelligent Data API. * (01:08:55) Alex provided his perspective on the modern streaming data architecture. * (01:13:24) Alex shared valuable hiring lessons to attract the right people who are excited about Redpanda’s mission. * (01:18:30) Alex talked about his experience choosing customers for Redpanda. * (01:20:33) Alex shared fundraising advice to founders who are seeking the right investors for their startups. * (01:23:23) Alex gave advice to a smart, driven minority who aspires to work on ambitious, technically deep, and challenging problems. * (01:28:18) Closing segment.
Alex's Contact Info* LinkedIn * Twitter * Website * GitHub
Redpanda's Resources* Website | Twitter | LinkedIn | Slack | GitHub | Contributing Doc * About Redpanda | Platform Capabilities | Customers * Docs | Redpanda University * Reports and Guides | Benchmarks * Hack The Planet Scholarship
Mentioned ContentBlog Posts* Redpanda raison d'etre (Feb 2019) * Thread-per-core buffer management for a modern Kafka-API storage system (Sep 2020) * Redpanda is now free and Source Available (Nov 2020) * Redpanda creates Redpanda, the Intelligent Data API Platform, backed by $15.5M initial funding from Lightspeed Venture Partners and GV (Jan 2021) * The Intelligent Data API (Jan 2021) * Redpanda Wasm engine architecture (June 2021) * We raised an additional $50M to drive the future of streaming data. Join us! (Feb 2022) * Redpanda gives Kafka a Run for Its Money (InfoWorld, May 2022) * Alex Gallego Builds Redpanda To Simplify And Unify Real-Time Streaming Data (Forbes, June 2022)
Talks* Distributed Stream Processing over thousands of Datacenters (GeeCON, Aug 2017) * How to Build the Fastest RPC (Nov 2017) * Co-designing Raft + thread-per-core execution model for the Kafka-API (Dec 2021)
People1. Andy Pavlo 2. Leslie Lamport 3. Kyle Kingsbury
NotesMy conversation with Alex was recorded back in August 2022. Since then, I recommend checking out these resources:
Show Notes* (01:44) Chetan reflected on his undergraduate experience at Stanford studying Electrical Engineering and Statistics back in the late 2000s. * (06:10) Chetan recalled his experience interning at IBM and Quantcast and doing research at Stanford Center for Minds, Brain, and Computation. * (08:41) Chetan talked about his first job working as a research analyst focused on healthcare policy at Acumen. * (11:15) Chetan walked through his decision to join Airbnb as their 4th data scientist and work on building Airbnb's original ETL framework for online risk mitigation. * (15:12) Chetan recalled the early state of data science at Airbnb. * (18:17) Chetan touched on the development of Airbnb's knowledge management and sharing platform called Knowledge Repo. * (23:10) Chetan explained why an experimentation program is the most impactful thing a data team can do. * (26:06) Chetan walked through the evolution of Airbnb's experimentation platform since its inception in 2014. * (31:24) Chetan recalled fond memories from taking a year off from work to travel. * (35:16) Chetan touched on his transition back to work by way of living in Atlanta and co-founding a logistics software startup called Saltbox. * (39:28) Chetan described his time as a data scientist at Webflow, building their experimentation system from scratch. * (42:48) Chetan shared the story behind the founding of Eppo. * (46:12) Chetan dissected the key capabilities that are baked into the Eppo product. * (48:32) Chetan dived deeper into the problems caused by long experiment durations and the benefits of using CUPED to bend time in experiments. * (52:04) Chetan talked about the role of a statistics engineer. * (54:45) Chetan shared his perspective on the role of experimentation in the Modern Data Stack and the Modern Growth Stack. * (01:00:47) Chetan discussed the core elements of the modern experimentation stack. * (01:04:54) Chetan talked about the experiment overhead. * (01:06:24) Chetan emphasized the designer gap in experimentation tools * (01:08:50) Chetan shared his thoughts about metric strategy. * (01:10:51) Chetan shared valuable hiring lessons to attract the right people who are excited about Eppo's mission. * (01:14:10) Chetan provided his perspectives on finding design partners for an early-stage startup. * (01:17:04) Chetan shared fundraising advice to founders who are seeking the right investors for their startups. * (01:19:40) Closing segment.
Chetan's Contact Info* LinkedIn * Twitter * GitHub * AngelList
Eppo's Resources* Website | Twitter | LinkedIn * Blog | Updates | Doc * About | Careers * Experimentation Product * Feature Flagging Product
Mentioned ContentArticles* Travel Year Facts and Superlatives (Dec 2019) * Why I Started Eppo (Feb 2021) * Reducing Experiment Durations (June 2021) * The Designer Gap in Experimentation Tools (June 2021) * We're Hiring A Statistics Engineer! (Aug 2021) * Should You Always Run An Experiment? (Aug 2021) * Stop Micromanaging Product Strategy (Sep 2021) * The most impactful thing a data team can do is establish an experimentation program (Dec 2021) * Bending Time in Experimentation (June 2022) * We Raised $19.5M! (June 2022) * Experimentation for the Modern Growth Stack: Our Investment in Eppo (June 2022)
People1. Mike Kaminsky 2. Sean Taylor 3. Jeremy Howard
Book* The Mom's Test (by Rob Fitzpatrick)
NotesMy conversation with Chetan was recorded back in August 2022. Since then, Eppo has launched feature flagging, and now offers the first "flags on top of your warehouse" experimentation platform. They also have Miro, Twitch, DraftKings, and Zapier as customers.
About the showDatacast features long-form, in-depth conversations with practitioners and researchers in the data community to walk through their professional journeys and unpack the lessons learned along the way. I invite guests coming from a wide range of career paths — from scientists and analysts to founders and investors — to analyze the case for using data in the real world and extract their mental models (“the WHY and the HOW”) behind their pursuits. Hopefully, these conversations can serve as valuable tools for early-stage data professionals as they navigate their own careers in the exciting data universe.
Datacast is produced and edited by James Le. For inquiries about sponsoring the podcast, email khanhle.1013@gmail.com.
Subscribe by searching for Datacast wherever you get podcasts, or click one of the links below:
If you’re new, see the podcast homepage for the most recent episodes to listen to, or browse the full guest list.
Show Notes* (01:38) Chad reflected on his early career as a freelance journalist working in Southeast Asia. * (04:27) Chad explained the benefits of writing for anyone working in a technical field. * (06:15) Chad touched on his entrance to data analytics through Conversion Rate Optimization. * (09:06) Chad walked through his decision to dive deep into the field of experimentation. * (13:28) Chad recalled how he spent time learning the basics of statistics. * (15:32) Chad discussed the differences in experimentation cultures at Subway, SEPHORA, and Microsoft. * (19:11) Chad shared the technical details behind the evolution of Convoy's data platform since he joined in 2019. * (23:05) Chad emphasized the role of data in Convoy's digital freight business. * (26:17) Chad brought up the importance of solving data discovery at Convoy and their decision to choose Amundsen. * (29:02) Chad shared lessons learned setting up a flexible experimentation platform at Convoy. * (32:46) Chad unpacked the problems with Change Data Capture and how his team built an internal change management platform called Chassis (a source of truth for definitions of events, entities, and relationships.). * (41:33) Chad discussed the existential threat of data quality. * (44:38) Chad explainedwhy the modern data warehouse is broken and why the "Immutable Data Warehouse" can be a solution. * (51:39) Chad zoomed in on the death of data modeling. * (57:42) Chad is bullish on the rise of the knowledge layer and data contracts in the upcoming years. * (01:03:37) Chad talked at length about the data collaboration problem. * (01:09:28) Chad gave advice for data organizations to be more customer-centric. * (01:11:55) Chad shared components of a high-quality Data UX function that any centralized data team should consider when developing data experiences. * (01:14:46) Chad touched on his mental framework for evaluating potential investments in the data space. * (01:18:57) Chad brought up the valuable skills he acquired as an internal product manager. * (01:20:05) Closing segment.
Chad's Contact Info* LinkedIn * Data Products Substack * Data Quality Camp
Mentioned ContentTalks* Aligning Experimentation Across Product Development and Marketing (CXL Live 2019) * Chassis: Entities, Events, and Change Management (Data Quality Meetup, 2021) * 1,000 Experiments Club with AB Tasty (July 2021) * Data Discovery at Lyft and Convoy (July 2021) (with Mark Grover) * The growth of the data platform product manager role (The Tech Trek, Dec 2021) * Implementing Amundsen at Convoy (Building the Backend, Jan 2022) * Getting ROI from Experimentation: How AB Experimentation plays out in Organizations (Data Council, March 2022) * Why are we so bad at this modern data stack? (Catalog and Cocktails, April 2022)
Articles* Experimentation not only protects your KPIs but your job as well (Dec 2019) * Is The Modern Data Warehouse Broken? (April 2022) (with Barr Moses) * The Existential Threat of Data Quality (May 2022) * The Death of Data Modeling (June 2022) * Data Collaboration Problem (June 2022) * The Rise of Data Contracts (Aug 2022)
People1. Barr Moses (Monte Carlo Data) 2. Juan Sequeda (data.world) 3. Adrian Kreuziger (Convoy)
Book* Agile Data Warehouse Design (by Lawrence Corr)
NotesMy conversation with Chad was recorded back in July 2022. Since then, I'd recommend looking at:
Show Notes* (01:56) Curtis reflected on his upbringing in rural Kentucky and his gift of education. * (07:20) Curtis explained how he cultivated mental focus and intellectual fortitude while growing up in Kentucky. * (10:30) Curtis shared his view regarding online misinformation on social media. * (14:27) Curtis recalled his undergraduate experience at Vanderbilt University in the early 2010s. * (22:39) Curtis explained how he learned best via teaching and mentoring. * (24:04) Curtis walked through the research and industry experiences he obtained throughout college. * (32:45) Curtis recalled his decision to embark on a Ph.D. in Computer Science at MIT. * (38:53) Curtis told the story of how he ended up finding his advisor - Professor Isaac Chuang (the inventor of the first working quantum computer). * (40:36) Curtis mentioned how he invented the CAMEO Detection Algorithm to detect “multiple-account” cheating in massive open online courses. * (44:47) Curtis unpacked his Ph.D. research on dataset uncertainty estimation. * (50:08) Curtis dissected confident learning, a family of theories and algorithms for supervised ML with label errors. * (53:22) Curtis encapsulated how he strategically iterated cleanlab at his various graduate internships. * (01:00:22) Curtis recalled his time founding his first startup ChipBrain, before founding Cleanlab. * (01:06:42) Curtis brought up the creation of the labelerrors.com project. * (01:12:12) Curtis provided lessons learned as a second-time founder. * (01:14:25) Curtis elaborated on the open-source roadmap of cleanlab. * (01:17:08) Curtis highlighted the key capabilities of Cleanlab Studio - the no-code, automatic data correction solution for data and engineering teams with robust enterprise features. * (01:18:50) Curtis touched on Cleanlab Vizzy - an interactive visualization of confident learning. * (01:20:29) Curtis shared valuable hiring lessons to attract the right people who are excited about Cleanlab’s mission. * (01:23:23) Curtis gave his thoughts on shaping Cleanlab’s culture. * (01:26:06) Curtis explained the similarity and differences between being a founder and a researcher. * (01:29:09) Curtis mentioned how he had helped researchers build affordable state-of-the-art deep learning machines. * (01:31:46) Curtis brought up his alter ego PomDP the Ph.D. rapper, and how rapping has been an outlet for him to express emotions and creativity. * (01:40:12) Curtis emphasized how his success had been due to a function of grit, resourcefulness, and friends made along the way. * (01:44:04) Closing segment.
Curtis' Contact Info* Academic Website * LinkedIn | Twitter | Facebook | Instagram * Google Scholar | GitHub * PhD Rapper (YouTube | Spotify | SoundCloud | Facebook | Twitter | Instagram) * L7 Machine Learning Blog
Cleanlab's Resources* Website | GitHub | Slack | Twitter | LinkedIn * Blog | Research | Doc * About | Careers * Cleanlab Studio * Cleanlab Vizzy * The Cleanlab Culture
Mentioned ContentPapers* Detecting and preventing “multiple-account” cheating in massive open online courses, Curtis G. Northcutt, Andrew Ho, & Isaac L. Chuang, Computers & Education, 2016. [paper | code | arXiv] * Comment Ranking Diversification in Forum Discussions, Curtis G. Northcutt, Kimberly Leon, & Naichun Chen, Learning at Scale, 2017. [paper | code | free-access] * Learning with Confident Examples: Rank Pruning for Robust Classification with Noisy Labels, Curtis G. Northcutt, Tailin Wu, & Isaac L. Chuang, 33rd Conference on Uncertainty in Artificial Intelligence (UAI 2017). [paper | code] * Confident Learning: Estimating Uncertainty for Dataset Labels, Curtis G. Northcutt, Lu Jiang, & Isaac L. Chuang, Journal of Artificial Intelligence Research (JAIR), Vol. 70 (2021). [paper | code | blog] * Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks, Curtis Northcutt, Anish Athalye, and Jonas Mueller, 35th Conference on Neural Information Processing Systems (NeurIPS 2021) Track on Datasets and Benchmarks [paper| demo | code | blog]
Blog Posts* Founder’s Medal recipient chooses MIT over Microsoft (May 2013) * Build a Pro Deep Learning Workstation... for Half the Price (Feb 2019) * An Introduction to Confident Learning: Finding and Learning with Label Errors in Datasets (Nov 2019) * Announcing cleanlab: a Python Package for ML and Deep Learning on Datasets with Label Errors (Nov 2019) * Double Deep Learning Speed by Changing the Position of your GPUs (Dec 2019) * Benchmarking: Which GPU for Deep Learning? (Dec 2019) * The Best 4-GPU Deep Learning Rig only costs $7000 not $11,000 (April 2020) * Pervasive Label Errors in ML Datasets Destabilize Benchmarks (March 2021) * Cleanlab: The History, Present, and Future (April 2022) * cleanlab 2.0: Automatically Find Errors in ML Datasets (April 2022) * How We Built Cleanlab Vizzy (August 2022)
Talks and Podcasts* Tedx Talk: The MIT Rap Challenge (July 2020) * Talk at NLP Summit (March 2022) * Talk at Data + AI Summit (June 2022) * MLOps Coffee Chat (July 2022) * Talk at Snorkel's Future of Data-Centric AI Conference (July 2022) * Open-Source Startup Podcast (March 2023)
People1. Leslie Kaelbling 2. Geoff Hinton 3. Jeff Dean
Book* Play Bigger: How Pirates, Dreamers, and Innovators Create and Dominate Markets (by Al Ramadan, Dave Peterson, Chris Lockhead, and Kevin Maney)
NotesMy conversation with Curtis was recorded back in August 2022. The Cleanlab team has had some important announcements in 2023 that I recommend looking at:
Cleanlab is about to announce its Series A announcement soon. Stay on the look for it!
About the showDatacast features long-form, in-depth conversations with practitioners and researchers in the data community to walk through their professional journeys and unpack the lessons learned along the way. I invite guests coming from a wide range of career paths — from scientists and analysts to founders and investors — to analyze the case for using data in the real world and extract their mental models (“the WHY and the HOW”) behind their pursuits. Hopefully, these conversations can serve as valuable tools for early-stage data professionals as they navigate their own careers in the exciting data universe.
Datacast is produced and edited by James Le. For inquiries about sponsoring the podcast, email khanhle.1013@gmail.com.
Subscribe by searching for Datacast wherever you get podcasts, or click one of the links below:
If you’re new, see the podcast homepage for the most recent episodes to listen to, or browse the full guest list.
Show Notes* (01:41) Frank shared formative experiences of his upbringing moving from China to the US. * (04:45) Frank described his overall academic experience at Stanford, studying Electrical Engineering with a minor in Computer Science. * (08:41) Frank talked about his research and industry experience while at Stanford. * (11:34) Frank shared his proudest accomplishments working at Yahoo as a research engineer in the Vision and Machine Learning group. * (16:37) Frank went over his experience co-founding a company that developed indoor localization and navigation solutions called Orion. * (23:06) Frank walked through his decision to leave Silicon Valley for China. * (26:02) Frank talked about his experience living and doing business in China (check out his two-part blog series that has covered normal life and the pandemic story in China). * (32:44) Frank elaborated on the work culture differences between the East and the West. * (37:58) Frank reflected on his decision to join Zilliz back in August 2021. * (42:55) Frank unpacked the notion of vector databases for the un-initiated. * (47:44) Frank provided a brief overview on the high-level design of Milvus, Zilliz's advanced open-source vector database solution. * (51:38) Frank highlighted three unique use cases of Milvus - malware detection, reverse image search, and drug discovery. * (56:51) Frank introduced Towhee, an open-source project that helps software engineers develop and deploy applications that utilize embeddings in just a few lines of code. * (01:01:59) Frank anticipated the evolution of the embedding tooling landscape to support the increasing adoption of unstructured data. * (01:04:21) Frank gave a primer on Zilliz Cloud, Zilliz's enterprise vector database solution. * (01:06:30) Closing segment.
Frank's Contact Info* LinkedIn * Twitter * GitHub * Website
Zilliz's Resources* Website | Twitter | LinkedIn | GitHub | YouTube * Zilliz Cloud Database * Milvus (Docs | GitHub) * Towhee (Docs | GitHub)
Mentioned ContentArticles and Presentations* A Gentle Introduction to Vector Databases (Dec 2021) * My Experience Living and Working in China, Part I (Feb 2022) * My Experience Living and Working in China, Part II (March 2022) * Making ML More Accessible for Application Developers (April 2022) * Understanding Neural Network Embeddings (April 2022) * Building An Open-Source Platform for Generating Embedding Vectors (Berlin Buzzwords, 2022)
People1. Yann LeCun (Chief AI Scientist at Meta, Professor at NYU) 2. Yangqing Jia (Creator of the Caffe deep learning framework) 3. Soumith Chintala (Creator of the PyTorch deep learning framework)
Book* A Short History of Nearly Everything (by Bill Bryson)
NotesMy conversation with Frank was recorded back in August 2022. The Zilliz team has had some important announcements in 2023 that I recommend looking at:
About the showDatacast features long-form, in-depth conversations with practitioners and researchers in the data community to walk through their professional journeys and unpack the lessons learned along the way. I invite guests coming from a wide range of career paths — from scientists and analysts to founders and investors — to analyze the case for using data in the real world and extract their mental models (“the WHY and the HOW”) behind their pursuits. Hopefully, these conversations can serve as valuable tools for early-stage data professionals as they navigate their own careers in the exciting data universe.
Datacast is produced and edited by James Le. For inquiries about sponsoring the podcast, email khanhle.1013@gmail.com.
Subscribe by searching for Datacast wherever you get podcasts, or click one of the links below:
If you’re new, see the podcast homepage for the most recent episodes to listen to, or browse the full guest list.
Show Notes (01:58) Vinoth shared his college experience studying IT at the Madras Institute of Technology in Chennai, India. * (07:09) Vinoth reflected on his time at UT Austin, getting a Master's degree in Computer Science - where he did research on high-bandwidth content distribution and large-scale parallel processing with shell pipes. * (11:20) Vinoth recalled his two years as a software engineer at Oracle, working on their database replication engine, HPC, and stream processing. * (15:30) Vinoth walked over his transition to LinkedIn as a senior software engineer, working primarily on Voldemort - a key-value store that handles a big chunk of traffic on Linkedin and serves thousands of requests per second over terabytes of data. * (24:41) Vinoth talked about his career transition to Uber in late 2014 as a founding engineer on Uber's data team and architect of Uber's data architecture. * (28:39) Vinoth reflected on the state of Uber's data infrastructure when he joined. * (34:31) Vinoth elaborated on Uber's case for incremental processing on Hadoop. * (38:53) Vinoth reviewed the initial design and implementation of Hudi across the Hadoop ecosystem at Uber in 2016. * (41:33) Vinoth shared t*he evolution of Hudi after it was initially open-sourced by Uber in 2017 and eventually incubated into the Apache Software Foundation in 2019. * (46:49) Vinoth explained how to keep the development of Apache Hudi vendor-neutral. * (49:36) Vinoth provided lessons learned about establishing standards for open-source data projects. * (53:45) Vinoth went over the valuable leadership lessons that he absorbed throughout his 4.5 years at Uber. * (57:17) Vinoth reflected on his 1.5 years as a principal engineer at Confluent working on ksqlDB, which makes it easy to create event streaming applications. * (01:02:16) Vinoth articulated the vision for Apache Hudi as a Streaming Data Lake platform. * (01:08:00) Vinoth highlighted the challenges with databases around indexing and concurrency control. * (01:11:37) Vinoth shared the unique challenges around prioritizing the Hudi roadmap and engaging an open-source community. * (01:16:32) Vinoth shared the founding story of Onehouse, a cloud-native, fully-managed lakehouse service built on Apache Hudi. * (01:22:02 ) Vinoth emphasized Onehouse's commitment towards openness. * (01:24:36) Vinoth shared valuable hiring lessons to attract the right people who are excited about Onehouse's mission. * (01:26:40) Vinoth shared fundraising advice to founders who are seeking the right investors for their startups. * (01:28:24) Closing segment.
Vinoth's Contact Info* LinkedIn * Twitter
Onehouse's Resources* Website | Twitter | LinkedIn * About | Product | Blog | Careers
Apache Hudi's Resources* User Docs | Technical Wiki | Roadmap * GitHub | Twitter | Slack
Mentioned ContentArticles and Presentations* Voldemort : Prototype to Production (May 2014) * Uber's Case for Incremental Processing on Hadoop (Aug 2016) * Hoodie: An Open Source Incremental Processing Framework From Uber (2017) * The Past, Present, and Future of Efficient Data Lake Architectures (2021) * Highly Available, Fault-Tolerant Pull Queries in ksqlDB (May 2020) * Apache Hudi - The Data Lake Platform (July 2021) * Introducing Onehouse (Feb 2022) * Automagic Data Lake Infrastructure (Feb 2022) * Onehouse Commitment to Openness (Feb 2022)
People* Leslie Lamport * Jeff Dean * Michael Stonebreaker
Book* Zero To One (by Peter Thiel)
NotesMy conversation with Vinoth was recorded back in August 2022. The Onehouse team has had some announcements in 2023 that I recommend looking at:
About the showDatacast features long-form, in-depth conversations with practitioners and researchers in the data community to walk through their professional journeys and unpack the lessons learned along the way. I invite guests coming from a wide range of career paths — from scientists and analysts to founders and investors — to analyze the case for using data in the real world and extract their mental models (“the WHY and the HOW”) behind their pursuits. Hopefully, these conversations can serve as valuable tools for early-stage data professionals as they navigate their own careers in the exciting data universe.
Datacast is produced and edited by James Le. For inquiries about sponsoring the podcast, email khanhle.1013@gmail.com.
Subscribe by searching for Datacast wherever you get podcasts, or click one of the links below:
If you’re new, see the podcast homepage for the most recent episodes to listen to, or browse the full guest list.
Show Notes* (01:55) Alexa shared formative experiences of her upbringing in Philadelphia. * (03:47) Alexa reflected on her undergraduate experience at Vanderbilt studying Engineering Science. * (05:49) Alexa recalled her first job out of college in management consulting at KPMG. * (08:20) Alexa walked over her transition from consulting to technology when she joined the Sales Operations team at Dataminr. * (12:35) Alexa talked about her proudest accomplishments at Dataminr - seeding the initial idea for Pocus and building a community for women in the workplace. * (20:23) Alexa reflected on her MBA experience at the Stanford Graduate School of Business. * (24:05) Alexa elaborated on the mindset difference between investing and operating. * (25:27) Alexa briefly touched on her internship at Monte Carlo. * (27:58) Alexa shared the founding story of Pocus. * (32:27) Alexa unpacked the concept of Product-Led Sales as a GTM approach. * (35:40) Alexa provided two example use cases of Pocus. * (39:35) Alexa explained the concepts of Product-Qualified Leads and Sales-Assist. * (42:20) Alexa discussed the long-term vision of Pocus' product roadmap. * (45:33) Alexa shared valuable hiring lessons to attract the right people who are aligned to Pocus' values. * (51:15) Alexa went over the journey of building the Product-Led Sales community. * (54:54) Alexa shared the unique opportunities of evolving a category, a community, and a product all at once. * (57:56) Alexa shared fundraising advice to founders who are seeking the right investors for their startups. * (01:01:15 ) Alexa provided advice to a smart, driven female operator who wants to take the leap of founding her company. * (01:03:09) Closing segment.
Alexa' Contact Info* LinkedIn * Twitter
Pocus' Resources* Website | Twitter | LinkedIn | YouTube * About | Product | Blog | Careers * Community | Newsletter
Mentioned ContentBlog Posts* What is Product-Led Sales? (July 2022) * The Myth of "No Sales" at PLG Companies (July 2021) * When To Add A Sales Team to Your PLG Company (Sep 2021) * The Definitive PQL Guide: Part 1, Part 2, Part 3 (Nov 2021) * What Is The Sales-Assist Role? (Nov 2021) * Introducing Pocus' PLS Platform (Nov 2021) * Product-Led Sales Community Wisdom Highlights 2021 (Dec 2021) * Notes on Community-Led Category Creation with Pocus' Co-Founder, Alexa Grabell (Feb 2022) * Sneak Peek at Pocus' PLS Platform (March 2022) * Announcing $23M to Transform How GTM Teams Use Data to Drive Revenue (June 2022) * Year One: The Product-Led Sales Platform is Here to Stay (July 2022)
People* Kyle Poyar (OpenView Ventures) * Melissa Ross (Clockwise) * Aaron Geller (QuickNode)
NotesMy conversation with Alexa was recorded back in July 2022. The Pocus team has had some announcements in 2023 that I recommend looking at:
About the showDatacast features long-form, in-depth conversations with practitioners and researchers in the data community to walk through their professional journeys and unpack the lessons learned along the way. I invite guests coming from a wide range of career paths — from scientists and analysts to founders and investors — to analyze the case for using data in the real world and extract their mental models (“the WHY and the HOW”) behind their pursuits. Hopefully, these conversations can serve as valuable tools for early-stage data professionals as they navigate their own careers in the exciting data universe.
Datacast is produced and edited by James Le. For inquiries about sponsoring the podcast, email khanhle.1013@gmail.com.
Subscribe by searching for Datacast wherever you get podcasts, or click one of the links below:
If you’re new, see the podcast homepage for the most recent episodes to listen to, or browse the full guest list.
Show Notes* (02:06) Carlos shared formative experiences of his upbringing tinkering with robots and websites. * (04:03) Carlos reflected on his education, studying Mechanical and Aerospace Engineering at Cornell University. * (05:34) Carlos discussed the technical details of his research on machine learning applications in robotics and art. * (10:11) Carlos explained his work as a robotic system analyst at Kiva Systems. * (15:41) Carlos discussed building his first data product at Kiva. * (20:24) Carlos recalled his stint working on warehouse-automating distributed robots at Amazon Robotics (after the Kiva acquisition). * (24:31) Carlos revealed his decision in 2013 to join an early-stage healthcare startup called Flatiron Health as the first data hire. * (28:43) Carlos shared his experience building Flatiron's Data Insights team from scratch. * (31:51) Carlos reviewed different data products built and deployed at Flatiron Health. * (38:41) Carlos shared the key learnings from hiring for his data team at Flatiron. * (44:08) Carlos shared the founding story of Glean, which is building a new way to make data exploration and visualization accessible to everyone. * (50:52) Carlos explained the pain points in data visualization/exploration and the product features of Glean that address them. * (55:03) Carlos dissected Glean DataOps, which brings modern developer workflow to the business intelligence layer and prevents broken dashboards. * (59:28) Carlos outlined the long-term product vision for Glean. * (01:03:11) Carlos shared valuable hiring lessons to attract the right people who are excited about Glean's mission. * (01:07:15) Carlos discussed his team's challenges in finding the early design partners. * (01:10:13) Carlos shared fundraising advice to founders who are seeking the right investors for their startups. * (01:11:57) Closing segment.
Carlos' Contact Info* Twitter * LinkedIn * GitHub * Website * Medium
Glean's Resources* Website | Twitter | LinkedIn * About | Docs | Blog * Interactive Public Demo | DataOps
Mentioned ContentBlog Posts* How the Data Insights team helps Flatiron build useful data products (May 2018) * The biggest mistake making your first data hire: not interviewing for product (July 2020) * How to interview your first data hire (Aug 2020) * My hack for getting started with data as a product (May 2021) * Introducing Glean (March 2022) * Your dashboard is probably broken (April 2022)
People1. Vicki Boykis 2. Anthony Goldbloom 3. Wes McKinney
Book* The Toyota Way: 14 Management Principles from the World's Greatest Manufacturer (by Jeffrey Liker)
NotesMy conversation with Carlos was recorded back in June 2022. The Glean team has had some announcements in 2023 that I recommend looking at:
About the showDatacast features long-form, in-depth conversations with practitioners and researchers in the data community to walk through their professional journeys and unpack the lessons learned along the way. I invite guests coming from a wide range of career paths — from scientists and analysts to founders and investors — to analyze the case for using data in the real world and extract their mental models (“the WHY and the HOW”) behind their pursuits. Hopefully, these conversations can serve as valuable tools for early-stage data professionals as they navigate their own careers in the exciting data universe.
Datacast is produced and edited by James Le. For inquiries about sponsoring the podcast, email khanhle.1013@gmail.com.
Subscribe by searching for Datacast wherever you get podcasts, or click one of the links below:
If you’re new, see the podcast homepage for the most recent episodes to listen to, or browse the full guest list.
Show Notes* (01:49) Shruti shared her upbringing in India - where she studied Engineering and Computer Science in the early 2000s. * (03:11) Shruti reflected on her early career as a software engineer at Hewlett-Packard and IBM. * (07:29) Shruti recalled the early days of cloud computing. * (09:01) Shruti reflected on her time pursuing an MBA at UCLA Anderson School of Management. * (11:55) Shruti explained her shift from software engineering to product management. * (14:19) Shruti revisited her years at VMware as a product line manager for cloud infrastructure - owning all aspects of go-to-market strategy and execution for VMware's entire software-defined storage portfolio. * (18:30) Shruti talked about her time as the VP of Marketing at Ravello Systems - growing the business from zero customers to a successful multi-million dollar acquisition by Oracle. * (23:07) Shruti went over her time as a senior director of product management for Oracle's cloud portfolio. * (27:20) Shruti recalled the founding story of Rockset - where she is a co-founder and Chief Product Officer. * (30:40) Shruti explained the concepts of real-time analytics and data applications for the uninitiated. * (37:31) Shruti unpacked the high-level design of Rockset architecture - which brings together cloud-native architecture, schemaless ingestion, converged indexing, and full-featured SQL. * (40:23) Shruti elaborated on the concept of converged indexing. * (42:43) Shruti dissected the technology requirements and the key layers of "the modern real-time data stack." * (46:17) Shruti talked about the role of partnerships in Rockset's product strategy. * (51:29) Shruti highlighted some of Rockset's customer use cases. * (56:06) Shruti shared valuable hiring lessons to attract high-integrity and diverse people for Rockset. * (58:51) Shruti shared her take on interviewing on strengths over weaknesses. * (01:01:32) Shruti shared the strategy Rockset used to find design partners in the early days. * (01:05:12) Shruti shared the tactics to combine the power of product-led adoption with sales-driven growth for rapidly scaling Rockset's business. * (01:08:45) Shruti shared fundraising advice to founders who are seeking the right investors for their startups. * (01:10:27) Shruti described the evolution of enterprise marketing and GTM strategy in the past decade. * (01:12:59) Closing segment.
Shruti's Contact Info* LinkedIn * Twitter * Forbes
Rockset's Resources* Website | Twitter | LinkedIn | Facebook * Docs | Blog | Community * Product | Architecture | Customers * Real-Time Analytics Explained * What Is A Data Application?
Mentioned ContentArticles* "Building Data Applications Powered by Real-Time Analytics" (May 2021) * "How startups can create a culture where women can win" (May 2021) * "Streaming Data and the Modern Real-Time Data Stack" (Nov 2021)
People* Barr Moses (Monte Carlo Data) * Jay Kreps (Confluent) * Alex DeBrie (DynamoDB Expert)
Book* Competing Against Luck (by Clayton Christensen)
NotesMy conversation with Shruti was recorded back in June 2022. Since then, a lot has happened. I recommend looking at the resources below:
About the showDatacast features long-form, in-depth conversations with practitioners and researchers in the data community to walk through their professional journeys and unpack the lessons learned along the way. I invite guests coming from a wide range of career paths — from scientists and analysts to founders and investors — to analyze the case for using data in the real world and extract their mental models (“the WHY and the HOW”) behind their pursuits. Hopefully, these conversations can serve as valuable tools for early-stage data professionals as they navigate their own careers in the exciting data universe.
Datacast is produced and edited by James Le. For inquiries about sponsoring the podcast, email khanhle.1013@gmail.com.
Subscribe by searching for Datacast wherever you get podcasts, or click one of the links below:
If you’re new, see the podcast homepage for the most recent episodes to listen to, or browse the full guest list.
Show Notes* (01:18) Arjun shared formative experiences of his upbringing - growing up in Bangalore, India; going to UWC Mahindra College for high school; and pursuing a liberal arts education in the US. * (04:45) Arjun described his overall academic experience at Willams College - where he studied Computer Science and Economics and did a one-year stint at the Computer Lab at the University of Cambridge. * (11:19) Arjun talked about his specialization within academic computer science: distributed systems. * (14:17) Arjun unpacked the arc of his Ph.D. experience at the University of Pennsylvania, advised by Professor Andreas Haeberlen. * (19:25) Arjun dissected the technical challenges and novelty of his Ph.D. dissertation on distributed systems that computed differentially private things. * (23:20) Arjun shared his love for teaching which benefits his industry career. * (25:55) Arjun walked through his decision to join Cockroach Labs as a software engineer. * (32:25) Arjun unpacked the CockroachDB Performance Guide and a RocksDB deep-dive on the Cockroach Labs blog. * (37:24) Arjun shared valuable lessons learned from his scaling journey with Cockroach. * (41:36) Arjun mentioned how his writing practice benefited his day-to-day work designing database systems in a production setting (Check out his posts on database transaction isolation semantics and the history of log-structured merge trees). * (45:46) Arjun unpacked his 2019 blog post titled "The Philosophy of Computational Complexity." * (52:52) Arjun emphasized the importance of writing evergreen and authoritative long-form content that attracts a small amount of audience. * (55:54) Arjun shared the story behind the founding of Materialize, which builds a SQL streaming database on top of Timely Dataflow and Differential Dataflow, two research projects created by his co-founder Frank McSherry. * (01:00:04) Arjun unpacked the architecture design of Materialize at a high level. * (01:04:36) Arjun explained a core capability of Materialize called Streaming SQL. * (01:07:37) Arjun discussed successful tactics to raise the adoption and contribution to Materialize's open-source project. * (01:11:23) Arjun walked through the major enterprise-grade features baked into Materialize Cloud. * (01:15:54) Arjun dissected a blog post about Materialize’s unbundled cloud architecture detailing the shift from the Materialize single binary to Materialize Cloud. * (01:21:13) Arjun envisioned how Materialize fits into the quickly evolving modern data stack. * (01:25:07) Arjun shared valuable hiring lessons to attract the right people who are excited about Materialize's mission. * (01:27:59) Arjun shared his brief take on building a high-performance company culture. * (01:29:19) Arjun discussed the challenges for his team to find the early design partners. * (01:31:17) Arjun walked through notable use cases of Materialize. * (01:34:24) Arjun shared fundraising advice with founders who are seeking the right investors for their startups. * (01:41:21) Arjun highlighted the similarities and differences between being a researcher and a founder. * (01:42:46) Closing segment.
Arjun's Contact Info* LinkedIn * Twitter * GitHub * Google Scholar
Materialize's Resources* Website | Twitter | LinkedIn | Slack * Docs | GitHub * Blog | Events | Guides * Careers
Mentioned ContentResearch + Articles* Distributed Differential Privacy and Applications (2015) * Performance Report: Benchmarking CockroachDB's TPC-C Performance * Why We Built CockroachDB on top of RocksDB (2019) * A History of Transaction Histories (2018) * A Brief History of Log Structured Merge Trees (2018)
People* Kyle Kingsbury * Bob Muglia * Frank McSherry
Book* Zero To One (by Peter Thiel)
NotesMy conversation with Arjun was recorded back in May 2022. Since then, a lot has happened. I recommend looking at the resources below:
About the showDatacast features long-form, in-depth conversations with practitioners and researchers in the data community to walk through their professional journeys and unpack the lessons learned along the way. I invite guests coming from a wide range of career paths — from scientists and analysts to founders and investors — to analyze the case for using data in the real world and extract their mental models (“the WHY and the HOW”) behind their pursuits. Hopefully, these conversations can serve as valuable tools for early-stage data professionals as they navigate their own careers in the exciting data universe.
Datacast is produced and edited by James Le. For inquiries about sponsoring the podcast, email khanhle.1013@gmail.com.
Subscribe by searching for Datacast wherever you get podcasts, or click one of the links below:
If you’re new, see the podcast homepage for the most recent episodes to listen to, or browse the full guest list.
Show Notes* (01:30) Doris walked through her time doing research in physics and astrophysics at UC Berkeley and getting involved with data science. * (04:11) Doris reflected on her decision to pursue the Ph.D. program in computer science at the University of Illinois, Urbana-Champaign. * (05:53) Doris discussed her development of no-code, interactive visualization interfaces accelerating users toward data insight discovery. * (10:37) Doris explained how the RISE Lab and I School at UC Berkeley helped shape her thinking around working with end-users and building something to serve the data science community. * (16:05) Doris unpacked the focus of her Ph.D. dissertation - which is to make data exploration and visualization easier and more accessible through automation. * (17:27) Doris shared the motivation and high-level design behind the development of Lux, a general-purpose visual exploration assistant situated within a computational notebook. * (21:25) Doris revealed the recipe for open-source community engagement and roadmap prioritization with Lux. * (26:17) Doris shared the founding story of Ponder, whose mission is to improve data science productivity by empowering users to do data science at all scales. * (31:02) Doris explained how Ponder helps solve the fragmentation challenges across the data stack. * (34:27) Doris provided a brief overview of Modin, which improves the scalability of data frames. * (38:41) Doris discussed Ponder's go-to-market strategy to drive more enterprise interest toward the product. * (41:23) Doris discussed her team's challenges in finding early design partners across various industries. * (44:16) Doris shared valuable hiring lessons to attract the right people who are excited about Ponder's mission. * (47:42) Doris shared fundraising advice to founders who are seeking the right investors for their startups. * (49:33) Doris highlighted the difference between being a researcher and a founder. * (51:06) Closing segment.
Doris' Contact Info* Website * Twitter * LinkedIn * GitHub
Ponder's Resources* Website | Twitter | LinkedIn | Slack * Modin | Lux * Events
Mentioned ContentPublications* The Case for a Visual Discovery Assistant:A Holistic Solution for Accelerating Visual Data Exploration (IEEE Data Bulletin 2018) * Understanding Sense-making in Visual Query Systems (IEEE Visual Analytics Science and Tech 2019) * Deconstructing Categorization in Visualization Recommendation: A Taxonomy and Comparative Study (IEEE Transactions on Visualization and Computer Graphics 2021) * Lux: Always-On Visualization Recommendation for Exploratory Data Science (Dec 2021)
Blog Posts* Insight Machines: The Past, Present, and Future of Visualization Recommendation (Multiple Views, Feb 2020) * Announcing Ponder (March 2022) * How we parallelized 600+ pandas functions with Modin (March 2022) * Using Lux to visualize your pandas dataframes with zero effort (March 2022) * Ph.D. Alum Doris Lee Wants to Democratize Data Science Tools (March 2022)
People* Chip Huyen * Shreyar Shankar * Parul Pandey
NotesMy conversation with Doris was recorded back in May 2022. Earlier this year, Ponder developed the first-of-its-kind technology that allows anyone to run their pandas code directly in your data warehouse, be it Snowflake, BigQuery, or Redshift. With Ponder, you get the same pandas-native experience that you love, but with the power and scalability of cloud-native data warehouses. More details are in this blog post.
Additionally, you can run NumPy commands on your data warehouse as well. This means you can work with the NumPy API to build data and ML pipelines, and let Snowflake / BigQuery / Redshift take care of scaling, security, and compliance. More details are in this blog post.
If you are interested in trying these new capabilities out, sign up here!
About the showDatacast features long-form, in-depth conversations with practitioners and researchers in the data community to walk through their professional journeys and unpack the lessons learned along the way. I invite guests coming from a wide range of career paths — from scientists and analysts to founders and investors — to analyze the case for using data in the real world and extract their mental models (“the WHY and the HOW”) behind their pursuits. Hopefully, these conversations can serve as valuable tools for early-stage data professionals as they navigate their own careers in the exciting data universe.
Datacast is produced and edited by James Le. For inquiries about sponsoring the podcast, email khanhle.1013@gmail.com.
Subscribe by searching for Datacast wherever you get podcasts, or click one of the links below:
If you’re new, see the podcast homepage for the most recent episodes to listen to, or browse the full guest list.
Show Notes* (01:47) Chris reflected on his educational experience at Santa Clara University in the mid-2000s, where he also interned at NeoMagic and Intacct Corporation. * (07:31) Chris recalled valuable lessons from his first job as a software engineer at PayPal, researching new fraud prevention techniques. * (11:28) Chris shared the technical and operational challenges associated with his work at LinkedIn as a data scientist - scaling LinkedIn's Hadoop cluster, improving LinkedIn's "People You May Know" algorithm, and delivering the next generation of LinkedIn's "Who's Viewed My Profile" product. * (22:00) Chris provided criteria that his team relied on when choosing their big data solutions (which include Aster Data, Greenplum, and Hadoop). * (25:22) Chris gave advice to early-stage startups that want to start adopting best practices in observability and deployment. * (28:02) Chris expanded on his concept that models and microservices should be running on the same continuous delivery stack. * (30:52) Chris discussed his strategy to become a better interviewer - as he performed ~1,500 interviews at LinkedIn and WePay. * (37:39) Chris explained the motivation behind the creation of Apache Samza (LinkedIn's streaming system infrastructure built on top of Apache Kafka) and discussed its high-level design philosophy. * (46:19) Chris shared lessons learned from evangelizing Samza to the broader open-source community outside of LinkedIn. * (52:44) Chris talked about his decision to join the Data Infrastructure team at WePay as a principal software engineer after 7 years at LinkedIn. * (01:00:53) Chris shared the technical details behind the evolution of WePay's data infrastructure throughout his time there. * (01:12:40) Chris shared an insider perspective on the adoption of Apache Airflow from his experience as a Project Committee Member. * (01:20:15) Chris discussed the fundamental design principles that make Apache Kafka such a powerful technology. * (01:25:40) Chris reflected on his experience building out WePay's engineering team. * (01:27:14) Chris shared the story behind the writing journey of the "Missing README" - which he co-authored with Dmitriy Ryaboy. * (01:38:16) Chris revisited his predictions in a 2019 post called "The Future of Data Engineering" and discussed key trends such as real-time data warehouses, data mesh, and headless BI. * (01:44:27) Chris gave advice to a smart, driven engineer who wants to explore angel investing - given his experience as a strategic investor and advisor for startups in the data space since 2015. * (01:48:17) Chris shared advice on hiring engineers and navigating open-source product strategies for companies he invested in. * (01:53:57) Chris reflected on his consistency in adding value to the relationships he has formed over the years. * (01:58:00) Closing segment.
Chris's Contact Info* Website * Twitter * LinkedIn * Github * AngelList
Mentioned ContentBlog Posts* Joel Spolsky's Blog * Models and microservices should be running on the same continuous delivery stack (Oct 2018) * Using checksums to verify syncing 100M database records (Napkin Math, Jan 2021) * Datacast episode with Jeremiah Lowin, CEO of Prefect (March 2022) * Kafka CDC breaks database encapsulation (Nov 2018) * Kafka provides data portability and infrastructure agility (Jan 2019) * The Future of Data Engineering (July 2019) * Work For Two Companies (Nov 2021)
People* Will Larson * Maxime Beauchemin * Julia Evans * Gunnar Morling * Coda Hale
Books* Google's Site Reliability Engineering Books * "On Writing Well" * "The Missing README" * "Empire of Light: Tesla, Edison, Westinghouse, and the Race to Electrify the World"
NotesMy conversation with Chris was recorded back in May 2022. Earlier this year, Chris released Recap, a dead simple data catalog for engineers, written in Python. Recap makes it easy for engineers to build infrastructure and tools that need metadata. Check out his blog post and get started with Recap's documentation!
About the showDatacast features long-form, in-depth conversations with practitioners and researchers in the data community to walk through their professional journeys and unpack the lessons learned along the way. I invite guests coming from a wide range of career paths — from scientists and analysts to founders and investors — to analyze the case for using data in the real world and extract their mental models (“the WHY and the HOW”) behind their pursuits. Hopefully, these conversations can serve as valuable tools for early-stage data professionals as they navigate their own careers in the exciting data universe.
Datacast is produced and edited by James Le. For inquiries about sponsoring the podcast, email khanhle.1013@gmail.com.
Subscribe by searching for Datacast wherever you get podcasts, or click one of the links below:
If you’re new, see the podcast homepage for the most recent episodes to listen to, or browse the full guest list.
Show Notes* (01:32) Nnamdi shared formative experiences of his upbringing, where he spent countless hours building computers, coding up websites, and finding ways to game Google search. * (04:54) Nnamdi described his undergraduate experience studying Economics at Yale University and interning at McKinsey and J.P. Morgan. * (08:10) Nnamdi reflected on the decline of the investment banking industry - given his one year working for the technology, media, and telecommunications group at J.P. Morgan in New York. * (12:52) Nnamdi discussed his career transition into venture investing at ICONIQ Capital, where he deployed over $500 million into high-growth technology companies. * (15:00) Nnamdi reflected on his proudest accomplishments during his four formative years at ICONIQ. * (17:35) Nnamdi talked about his excitement for GitLab, one of his investments. * (21:27) Nnamdi touched on his time getting an MBA from the Stanford Graduate School of Business. * (24:21) Nnamdi also completed coursework in Stanford's Computer Science department (such as CS 231N and CS 224N) during his MBA. * (26:37) Nnamdi explained the venture ecosystem at Stanford, given his experience serving as the Co-President and Vice President of the Venture Capital and Tech Clubs, respectively. * (28:57) Nnamdi unpacked his experience working at Confluent as a product manager and conducting independent research on trends in developer productivity. * (32:23) Nnamdi reflected on his decision to join Lightspeed Venture in mid-2020, investing in early-stage software startups to enhance the productivity of technical knowledge workers. * (34:17) Nnamdi shared how he proved his value upfront in potential deals and started forming his investment theses as a new investor at Lightspeed. * (36:24) Nnamdi dissected the key factors that triggered him to make investments in the seed rounds of Ponder and Voltron Data (in the domain of developer tools). * (40:36) Nnamdi explained his Series A investment in Redpanda and Materialize (in the domain of real-time data infrastructure). * (45:45) Nnamdi shared advice he had been giving his portfolio companies in hiring decisions and navigating growth strategy. * (49:07) Nnamdi walked through his 3-part series on major industry trends, top strategic priorities, and biggest challenges for software and infrastructure startups pushing the developer productivity frontier. * (52:37) Nnamdi shared advice to startups thinking about scaling their developer relations, given the challenge of hiring developer advocates for dev-focused startups. * (56:27) Nnamdi unpacked his 3-part series on the developer productivity manifesto that introduces the developer productivity flywheel, explains how more developers lead to lower productivity, and argues that we are leaving on the table $670B of software by not maximizing developer employment and developer productivity. * (01:01:26) Nnamdi examined his obsession with the fat-tailed nature of high-growth startups, such as why VCs don't index-invest, why Saas monetization is concentrated on the tails, and why product-market fit gets harder to achieve the longer you search for it. * (01:04:26) Nnamdi explained his new and improved SaaS metric called Weighted ACV, which is the weight of the revenue that a customer represents and tells founders where to look if they want to best understand the revenue of their businesses. * (01:07:53) Nnamdi thought about his recognition as equal to his credibility as an investor on a mission to increase total software output by investing in technical tools for technical people. * (01:11:03) Closing segment.
Nnamdi's Contact Info* Website * Lightspeed Profile * LinkedIn * Twitter * GitHub * Medium
Lightspeed's Resources* Website | Twitter | LinkedIn * Global Presence * Medium Blog
Mentioned ContentArticles* Six Trends Shaping Developer Productivity * Top Three Strategic Priorities of Developer Productivity Startups * Four Challenges Facing Developer Productivity Startups * Awesome Developer Advocates Are Hiding in Plain Sight * The Developer Productivity Manifesto Part 1 — The Flywheel * The Developer Productivity Manifesto Part 2 — More (Developers) Isn’t Always More * The Developer Productivity Manifesto Part 3 — Leaving Software on the Table * You Don't Understand Compound Growth * Funding Simply Shifts the Bottleneck * Why Don't VCs Index Invest? * Enterprise Software Monetization is Fat-Tailed * Product-Market Fit is Lindy * Introducing a New and Improved SaaS Metric: Weighted ACV
People1. Mike Volpi (Index Ventures) 2. Keith Rabois (Founders Fund)
BooksNassim Taleb's Incerto Series:
NotesMy conversation with Nnamdi was recorded in May 2022. Since then, many things have happened. I'd recommend checking out:
About the showDatacast features long-form, in-depth conversations with practitioners and researchers in the data community to walk through their professional journeys and unpack the lessons learned along the way. I invite guests coming from a wide range of career paths — from scientists and analysts to founders and investors — to analyze the case for using data in the real world and extract their mental models (“the WHY and the HOW”) behind their pursuits. Hopefully, these conversations can serve as valuable tools for early-stage data professionals as they navigate their own careers in the exciting data universe.
Datacast is produced and edited by James Le. For inquiries about sponsoring the podcast, email khanhle.1013@gmail.com.
Subscribe by searching for Datacast wherever you get podcasts, or click one of the links below:
If you’re new, see the podcast homepage for the most recent episodes to listen to, or browse the full guest list.
Show Notes* (02:00) Tom shared formative experiences of his upbringing. * (04:17) Tom described his educational experience at MIT and his research thesis in Computer Vision. * (07:08) Tom talked about his interest in computer vision and computational neuroscience. * (08:38) Tom recalled lessons from his first job out of school as a software engineer at Silicon Graphics, building high-performance visualization systems. * (11:32) Tom reflected on his time as a product lead at Autodesk, launching a location services platform for global wireless carriers and a developer ecosystem for GIS applications. * (13:48) Tom reflected on his MBA experience at Harvard Business School. * (16:54) Tom reflected on his first stint at Google - leading a sales operations team in AdWords and building YouTube's monetization systems. * (20:06) Tom recalled lessons learned as a first-time founder of a social commerce startup called Renown Labs. * (23:11) Tom walked over his time as the Director of Product at the SaaS social media marketing startup Wildfire (which was acquired by Google in August 2012). * (27:00) Tom explained his career transition from tech operating into venture investing - after joining the enterprise investment team at a16z as a partner in 2012. * (29:55) Tom revisited his thesis, discussing the rise of Enterprise Hackers back in 2013. * (32:25) Tom talked about his decision to join NextWorld Capital as a partner in 2014, leading investments across enterprise applications, the Internet of Things, and AI. * (34:50) Tom unpacked his investment thesis on enterprise technology that helps the blue-collar working class. * (36:50) Tom shared his mental checklist used to evaluate investment opportunities in enterprise AI at NextWorld. * (40:39) Tom shared the founding story of Masterful AI - where he has been a co-founder and CEO since 2019. * (43:08) Tom expanded upon the 2-year incubation period from the inception to the announcement of the Masterful AI platform. * (45:17) Tom unpacked major inefficiencies of ML development and explained how the Masterful platform works at a high level. * (48:18) Tom shared exciting initiatives in Masterful's product roadmap. * (50:09) Tom highlighted the principles that stood the test of time in computer vision over the past two decades. * (51:41) Tom shared valuable hiring lessons to attract the right people who are excited about Masterful AI's mission. * (55:07) Tom discussed his team's challenges in finding early design partners across various industries. * (58:24) Tom shared fundraising advice to founders who are seeking the right investors for their startups. * (01:01:01) Tom reflected on his career traversing across product management, venture capital, and startup founder. * (01:05:20) Closing segment.
Tom's Contact Info* LinkedIn * Twitter * Medium * Website
Masterful AI's Resources* Website | Twitter | LinkedIn * Docs | Slack Community * "Building Things with Machine Learning" Podcast
Mentioned ContentArticles* "The Enterprise Hacker Rises" (a16z Blog, Dec 2013) * "Joining NextWorld Capital" (Personal Blog, Nov 2014) * "The Next Big Opportunity In Enterprise Starts In The Field" (TechCrunch, July 2015) * "My visit to the Obama White House: AI, the future of jobs, and a VC’s Letter to the next administration" (NextWorld Insights, Jan 2017) * "AI hype has peaked so what’s next?" (TechCrunch, Sep 2017) * "AI is bringing superpowers to the specialist" (LinkedIn, Oct 2018) * "Introducing Masterful AI" (Masterful Blog, Nov 2021)
People* Andrew Ng (Founder of DeepLearning.AI, Founder and CEO of Landing AI, Co-Founder of Coursera) * Chris Dixon (General Partner at a16z)
Book* AI Superpowers (by Kai-Fu Lee)
NoteMy conversation with Tom was recorded back in May 2022. Here is the note from Tom regarding updates with Masterful:
The latest at Masterful AI is that we’re launching a new generative AI product. We saw a need to make generative models more customizable and more reliable, so companies can trust them for real business applications. We’re starting by enabling companies to tell a more vivid and personalized story about their products at scale.
About the showDatacast features long-form, in-depth conversations with practitioners and researchers in the data community to walk through their professional journeys and unpack the lessons learned along the way. I invite guests coming from a wide range of career paths — from scientists and analysts to founders and investors — to analyze the case for using data in the real world and extract their mental models (“the WHY and the HOW”) behind their pursuits. Hopefully, these conversations can serve as valuable tools for early-stage data professionals as they navigate their own careers in the exciting data universe.
Datacast is produced and edited by James Le. For inquiries about sponsoring the podcast, email khanhle.1013@gmail.com.
Subscribe by searching for Datacast wherever you get podcasts, or click one of the links below:
If you’re new, see the podcast homepage for the most recent episodes to listen to, or browse the full guest list.
Show Notes* (02:02) Grace shared formative experiences of her upbringing - being heavily influenced by the financial sector from growing up near New York and getting an appreciation for diverse perspectives from studying abroad in Tokyo. * (03:47) Grace described her college experience at Stanford studying Management Science and Engineering. * (08:04) Grace talked about her participation in the Mayfield Fellowship and her service as a Co-President of Stanford Women in Business. * (12:34) Grace walked through her internship experiences as an investor at the Stanford Management Company, in product at ed-tech startup Handshake, and in growth equity at Stripes Group. * (16:20) Grace reflected on her time at Canvas Ventures - where she joined as a campus scout while still a student. * (19:11) Grace shared three approaches to prove her value upfront in potential deals and to form her investment thesis as a new VC associate. * (23:07) Grace dissected her Series A investment in Vendia, a blockchain-powered, real-time data-sharing platform that solves the growing inter-organization data collaboration problem. * (25:17) Grace examined her Series A investment in Robocorp, which offers the first cloud-native, open-source automation stack and orchestration platform to power any automation process. * (26:43) Grace shared three pieces of advice in hiring decisions, navigating go-to-market strategy, and growing product offerings that she had given her portfolio companies. * (29:49) Grace shared trends in the API-first economy that she is most excited about in the upcoming years. * (31:56) Grace unpacked key takeaways from her article "The Mindset of a Data Leader." * (34:02) Grace discussed under-hyped and over-hyped trends in Web3 - taken from her incredibly detailed deck on the Web3 World. * (36:31) Grace dissected the major categories of the Web3 infrastructure, including Decentralized Finance, Decentralized Apps, DAOs, NFTs, and Guild Education/Reskilling. * (41:43) Grace walked through her decision to join Lux Capital, a firm that invests in emerging science and technology ventures at the outermost edges of what is possible, as a principal investor in early 2022. * (45:12) Grace shared her mental checklist to evaluate entrepreneurs and make investment decisions at the nexus of web3, data infrastructure, and applications of AI/ML. * (47:10) Grace talked about her community-building work to promote women's voices in tech. * (48:59) Grace reflected on her consistency in adding value to every conversation with people in her community. * (51:26) Closing segment.
Grace's Contact Info* Website * Lux Profile * LinkedIn * Twitter
Lux Capital* Website | Twitter | LinkedIn * Securities (Podcast & Newsletter)
Mentioned ResourcesArticles* "The Third-Party API Economy: Part I" (Sep 2020) * "The Third-Party API Economy: Part II" (Feb 2021) * "The Mindset of a Data Leader" (Nov 2020) * "The Web3 World" (Jan 2022) * "Welcoming our newest investor Grace Isford to Lux Capital" (Feb 2022)
People* Fred Wilson (Union Square Ventures) * Matt Huang and Fred Ehrsam (Paradigm Ventures) * Katie Haun (Haun Ventures)
Book* "Wanting" (by Luke Burgis)
NotesMy conversation with Grace was recorded back in April 2022. Since then, many things have happened. I'd recommend:
Additionally, Grace invested in RunwayML's Series C, a pioneer in the Generative AI space. If you are in NYC, be sure to stop by the upcoming first annual AI film festival powered by Runway!
About the showDatacast features long-form, in-depth conversations with practitioners and researchers in the data community to walk through their professional journeys and unpack the lessons learned along the way. I invite guests coming from a wide range of career paths — from scientists and analysts to founders and investors — to analyze the case for using data in the real world and extract their mental models (“the WHY and the HOW”) behind their pursuits. Hopefully, these conversations can serve as valuable tools for early-stage data professionals as they navigate their own careers in the exciting data universe.
Datacast is produced and edited by James Le. For inquiries about sponsoring the podcast, email khanhle.1013@gmail.com.
Subscribe by searching for Datacast wherever you get podcasts, or click one of the links below:
If you’re new, see the podcast homepage for the most recent episodes to listen to, or browse the full guest list.
Show Notes* (01:32) Dânia shared her upbringing in Brazil and her college experience studying Applied Mathematics at the University of Campinas. * (05:58) Dânia touched on her early career working in marketing intelligence in Brazil. * (10:38) Dânia described her thesis on scalable implementations of the Alternating Least Squares algorithm for Collaborative Filtering recommendation, conducted during her Master's degree in Computer Science from the University of Fluminense. * (16:10) Dânia recalled her hustling phase working and getting a Master's degree simultaneously. * (24:19) Dânia reflected on her move to Berlin to work as a data scientist in several startups. * (31:00) Dânia looked back at her time working at MYTOYS GROUP's Analytics team, responsible for Predictive Analytics and Machine Learning Modeling. * (34:12) Dânia compared doing data science to practicing mixed martial arts. * (38:35) Dânia reflected on her involvement with Data Science for Social Good Berlin as a data ambassador and Data Science Retreat as a SQL Masterclass Teacher. * (43:14) Dânia shared the founding story of AI Guild - the go-to community for data and business professionals advancing AI adoption - where she is a founding member. * (47:36) Dânia gave her thoughts on barriers preventing more women from entering the data field. * (51:21) Dânia discussed the #datalift initiative, which pushes to productionize more data analytics and machine learning solutions. * (58:27) Dânia explained her work supporting the advancement of #datacareer talents and experts. * (01:01:22) Dânia gave her take on the evolution of the data field over the past decade. * (01:03:16) Closing segment.
Dânia's Contact Info* LinkedIn * Twitter * Website * GitHub * Medium
AI Guild's Resources* Website | LinkedIn | YouTube * Join As A Member * #datalift * #datacareer
Mentioned ContentPeople1. Andrew Ng: Founder of deeplearning.ai, co-founder of Coursera 2. Alessandra Sala: President of Women in AI, Sr. Director of Artificial Intelligence and Data Science at Shutterstock 3. Joy Buolamwini: Founder and Executive director of The Algorithmic Justice League and maker of the "Coded Bias" documentary, available on Netflix
Book* Weapons of Math Destruction: How Big Data Increases Inequality and Threatens Democracy by Cathy O'Neil
About the showDatacast features long-form, in-depth conversations with practitioners and researchers in the data community to walk through their professional journeys and unpack the lessons learned along the way. I invite guests coming from a wide range of career paths — from scientists and analysts to founders and investors — to analyze the case for using data in the real world and extract their mental models (“the WHY and the HOW”) behind their pursuits. Hopefully, these conversations can serve as valuable tools for early-stage data professionals as they navigate their own careers in the exciting data universe.
Datacast is produced and edited by James Le. For inquiries about sponsoring the podcast, email khanhle.1013@gmail.com.
Subscribe by searching for Datacast wherever you get podcasts, or click one of the links below:
If you’re new, see the podcast homepage for the most recent episodes to listen to, or browse the full guest list.
Show Notes* (01:33) Bobby shared his upbringing in DC and high-school experience at St. Albans School. * (04:10) Bobby described his academic experience at Stanford studying Management Science and Engineering. * (07:39) Bobby recalled valuable career lessons learned working as a Finance Analyst at IBM and Inflection. * (09:56) Bobby reflected on his rationale for joining Intercom as one of the company's early employees right after its Series A financing in 2013. * (14:16) Bobby unpacked his 2016 talk "Scaling Analytics at Intercom," which explained the analytics journey at Intercom. * (18:46) Bobby shared a few metrics that are fundamental to the health of a startup across its growth stages (read his Intercom blog about the data points that startups should measure). * (22:50) Bobby shared the founding story of Equals. * (27:33) Bobby explained his decision to choose Ben McRedmond as his co-founder. * (29:35) Bobby expanded on the appealing traits of using spreadsheets. * (31:54) Bobby described the evolution of spreadsheet-like products and how the Equals product works at a high level. * (34:35) Bobby gave his take on how the concept of a next-generation spreadsheet fits into the quickly evolving modern data stack. * (38:31) Bobby shared valuable hiring lessons to attract the right people who are excited about Equals' mission. * (44:34) Bobby shared the challenges of finding Equals' early design partners and lighthouse customers. * (47:17) Bobby recapped key lessons about hiring financial analysts at Intercom. * (51:45) Bobby shared advice to a smart, driven finance operator looking to get more influence within a startup environment. * (56:26) Bobby emphasized the valuable skills acquired from his analyst career for his current founder journey. * (58:45) Closing segment.
Bobby's Contact Info* LinkedIn * Twitter
Equals Resources* Website | Twitter | LinkedIn * Spreadsheet Templates * Insights In Action interview series * Introducing Pivot Tables for Equals (Aug 2022) * Equals raises $16M Series A from a16z to replace Excel (Nov 2022)
Equals is hiring across Engineering, Design, Growth, and an Executive Assistant. Reach out to Bobby if you are interested!
Mentioned ContentArticles + Talk* 23 SaaS Metrics for Fundraising + Optimization (March 2015) * Scaling Analytics at Intercom (Intercom Analytics Meetup, April 2016) * Data Points: What Should Your Startup Measure? (Oct 2017) * Every analyst is a finance analyst (May 2021) * The only question that matters when interviewing analysts (May 2021) * When to make your first finance hire (May 2021) * The hardest leap to make as a scaling finance leader (June 2021) * Finance and describing product-market fit (Sep 2021) * The curious analyst (Sep 2021) * The less lonely finance leader (Sep 2021) * Why every scaling finance team is understaffed (Nov 2021) * Revenue is the best North Star metric (March 2022)
People* Karen Church (VP of Research and Data Science at Intercom, Founder of HER+Data) * Noah Goodman (President at DataCRT) * Peter Fishman (Co-Founder of Mozart Data)
About the showDatacast features long-form, in-depth conversations with practitioners and researchers in the data community to walk through their professional journeys and unpack the lessons learned along the way. I invite guests coming from a wide range of career paths — from scientists and analysts to founders and investors — to analyze the case for using data in the real world and extract their mental models (“the WHY and the HOW”) behind their pursuits. Hopefully, these conversations can serve as valuable tools for early-stage data professionals as they navigate their own careers in the exciting data universe.
Datacast is produced and edited by James Le. For inquiries about sponsoring the podcast, email khanhle.1013@gmail.com.
Subscribe by searching for Datacast wherever you get podcasts, or click one of the links below:
If you’re new, see the podcast homepage for the most recent episodes to listen to, or browse the full guest list.
Show Notes* (01:34) Ran reflected on his time working as a Technical Product Manager at the Israeli Intelligence army. * (04:07) Ran recalled his favorite classes on Machine Learning and Computer Graphics during his education in Computer Science at Reichman University. * (05:24) Ran talked about a valuable lesson learned as a Software Engineer at VMware's Cloud Provider Software Business Unit. * (08:07) Ran shared his thoughts on how engineers could be more impactful in startup organizations. * (09:52) Ran talked about his decision to join Wix.com to work as a software engineer focusing on data infrastructure. * (12:48) Ran explained the motivation for building Wix's internal ML platform, designed to address the end-to-end ML workflow. * (16:48) Ran discussed the main components of Wix's ML platform: feature store, CI/CD mechanism, UI management console, and API prediction service. * (18:51) Ran unpacked the virtual feature store and the CI/CD components of Wix's ML platform. * (24:41) Ran expanded on the distinction between virtual and materialized feature stores. * (27:01) Ran provided three key lessons for organizations looking to build an internal ML platform (as brought upon his 2020 talk discussing Wix's ML Platform). * (31:43) Ran shared the essential attributes of exceptional data and ML engineering talent. * (33:54) Ran shared the founding story of Qwak, which aims to build an end-to-end ML engineering platform to automate the MLOps processes. * (37:07) Ran talked about his responsibilities as the VP of Engineering at Qwak. * (38:45) Ran dissected the key capabilities that are baked into the Qwak platform - a Build System, a Serving layer, a Data Lake, a Feature Store, and Automations capabilities. * (44:05) Ran explained the big engineering challenges for teams to build an in-house feature store and envisioned the future of the feature store ecosystem in the upcoming years. * (47:45) Ran shared valuable hiring lessons to attract the right people who are excited about Qwak's mission. * (50:22) Ran reflected on the challenges for Qwak to find the early design partners. * (52:43) Ran described the state of the ML Engineering community in Israel. * (54:53) Closing segment.
Ran's Contact Info* LinkedIn
Qwak's Resources* Website | Twitter | LinkedIn * Why Qwak * Blog
Mentioned ContentTalks* "Overview of Wix's Machine Learning Platform" (2020) * "Feature Stores - Unified Data Pipelines for ML" (2022)
People* Andrew Ng * Matei Zaharia * Barr Moses
Book* "Principles" (by Ray Dalio)
About the showDatacast features long-form, in-depth conversations with practitioners and researchers in the data community to walk through their professional journeys and unpack the lessons learned along the way. I invite guests coming from a wide range of career paths — from scientists and analysts to founders and investors — to analyze the case for using data in the real world and extract their mental models (“the WHY and the HOW”) behind their pursuits. Hopefully, these conversations can serve as valuable tools for early-stage data professionals as they navigate their own careers in the exciting data universe.
Datacast is produced and edited by James Le. For inquiries about sponsoring the podcast, email khanhle.1013@gmail.com.
Subscribe by searching for Datacast wherever you get podcasts, or click one of the links below:
If you’re new, see the podcast homepage for the most recent episodes to listen to, or browse the full guest list.
Show Notes* (02:15) Eric reflected on his early interest in computer science and his decision to study at Carnegie Mellon University in the early 90s. * (05:40) Eric recalled his academic and overall college experience, emphasizing the importance of the people he was surrounded with. * (08:22) Eric talked about his time working as a quant analyst early in his career, the moment he encountered the birth of the Mosaic browser, and his decision to join the tech industry. * (13:01) Eric imparted wisdom learned from venture investing during the dot-com boom. * (18:02) Eric talked about the next phase of his academic career - earning a Ph.D. in Computer Science from Carnegie Mellon and dropping out of a Ph.D. program at Stanford. * (21:06) Eric discussed his academic research on Computational Economics for corporate malfeasance during his time as a Ph.D. student. * (27:39) Eric shared different initiatives he worked on with Carnegie Mellon University - serving as the Assistant Dean and Assistant Professor of Software Engineering, launching CMU's Silicon Valley Campus, and founding CMU's Entrepreneurial Management program. * (31:54) Eric described his journey in founding Hg Analytics, a hedge fund focused on statistical arbitrage, alongside other CMU's Computer Science PhDs. * (37:36) Eric revisited his passion for AI and robotics, which eventually led to serving as a Presidential Innovation Fellow during the Obama Administration with the White House Office of Science and Technology Policy. * (42:54) Eric shared his perspective on the role of AI in geopolitics and highlighted the challenges with data integration. * (47:29) Eric explained his company Conexus, which develops a technology spin-off from MIT's Mathematics department using a branch of math called Category Theory. * (50:55) Eric went over a customer case study that uses Conexus's solution to guarantee the semantics of data integrity during data transformation. * (54:20) Eric showed his enthusiasm for the concept of data relationships. * (56:59) Eric provided a sneak peek of his forthcoming book, "The Coming Composability: The roadmap for using technology to solve society's biggest problems." * (58:38) Closing segment.
Eric's Contact Info* Twitter * LinkedIn
Conexus' Resources* Website | Resources
Mentioned ContentPeople* Kai-Fu Lee * Andrew Ng * Eric Xing
Book* "ReCulturing: Design Your Company Culture to Connect with Strategy and Purpose for Lasting Success" (by Melissa Daimler)
About the showDatacast features long-form, in-depth conversations with practitioners and researchers in the data community to walk through their professional journeys and unpack the lessons learned along the way. I invite guests coming from a wide range of career paths — from scientists and analysts to founders and investors — to analyze the case for using data in the real world and extract their mental models (“the WHY and the HOW”) behind their pursuits. Hopefully, these conversations can serve as valuable tools for early-stage data professionals as they navigate their own careers in the exciting data universe.
Datacast is produced and edited by James Le. Get in touch with feedback or guest suggestions by emailing khanhle.1013@gmail.com.
Subscribe by searching for Datacast wherever you get podcasts, or click one of the links below:
If you’re new, see the podcast homepage for the most recent episodes to listen to or browse the full guest list.
Show Notes* (01:56) Astasia shared her childhood growing up in Silicon Valley. * (05:12) Astasia reflected on her undergraduate education at Stanford - studying Political Science and International Relations. * (06:35) Astasia discussed her research at the Graduate Business School with Professor Condoleezza Rice on a case study called "San Leon Energy: Hydraulic Fracturing in Poland" - which explores how to manage the political risks of using a controversial energy extraction technology in the European Union. * (09:26) Astasia talked about her year in the UK getting a Master's in Technology Policy at the University of Cambridge's Judge Business School. * (12:52) Astasia recalled her experience as an Equity Research Analyst at Baird and Co. * (17:49) Astasia mentioned her work at Cisco Investments, driving their cloud-infrastructure M&A and venture investments. * (20:58) Astasia shared her thoughts on different M&A frameworks she learned from Cisco. * (23:27) Astasia reflected on her decision to join Redpoint Ventures in early 2017, leading investments across developer tools, cloud infrastructure, data/ML infrastructure, AI applications, and cybersecurity. * (25:44) Astasia debunked misconceptions about the venture industry. * (29:30) Astasia discussed ways to prove her value upfront in potential deals and start forming her investment theses as a new investor. * (33:01) Astasia dissected the key factors that triggered her to invest in the Series A of Solo.io and the Series B of LaunchDarkly (in the domain of cloud infrastructure). * (38:48) Astasia explained her Series A investment in Hex and Series B investment in Preset (in the domain of data infrastructure). * (44:12) Astasia shared advice she had given her portfolio companies in hiring decisions, pricing products, and navigating go-to-market strategy while at Redpoint. * (47:36) Astasia walked through her process of writing comprehensive research primers in her Medium blog Memory Leak on wide-ranging topics - from data science notebooks and data orchestration to data pipelining and ML data management. * (51:19) Astasia shared the typical challenges she has seen in companies looking to incorporate Product-Led Growth into their go-to-market motion. * (54:10) Astasia discussed building a community as a fuel for product-led growth and shared advice to startups thinking about starting their community initiatives. * (56:40) Astasia shared advice for hiring good DevRel practitioners. * (01:00:15) Astasia shared advice for a smart, driven operator who wants to explore angel investing. * (01:03:26) Astasia talked about her current journey as the Founding Partner at Quiet Capital, sitting on its early-stage enterprise team and leading opportunities across pre-seed, seed, Series A, and Series B. * (01:05:13) Astasia expanded upon her typical mental checklist to evaluate entrepreneurs and make investment decisions. * (01:07:36) Astasia briefly touched on LP fundraising for Quiet Capital to become a "modern venture firm." * (01:09:59) Astasia emphasized her enthusiasm for the Data-Centric ML movement. * (01:13:41) Closing segment.
Astasia's Contact Info* LinkedIn * Medium * Twitter
Quiet Capital* Website * LinkedIn * Twitter
Mentioned ResourcesContent* John Gannon Blog
People* Satish Dharmaraj (Redpoint Ventures) * Scott Raney (Redpoint Ventures) * Amanda Robson (Cowboy Ventures)
NotesMy conversation with Astasia was recorded back in April 2022. Since then, many things have happened. I'd recommend:
Show Notes* (02:24) Tarush shared his upbringing in India and his decision to study abroad in the US. * (03:51) Tarush walked through his college experience studying Computer Engineering at Carnegie Mellon University. * (06:24) Tarush described the non-existent state of data infrastructure at Salesforce when he joined as the first data engineer in 2012. * (11:21) Tarush went over his contribution to the automation and benchmarking frameworks over his tenure at Salesforce. * (15:50) Tarush recalled lessons learned from building and managing a data team as a Data Manager at Wyng. * (19:54) Tarush explained how a data team can serve other functional units more efficiently. * (22:37) Tarush elaborated on his decision to adopt Looker for Wyng's Business Intelligence needs. * (26:30) Tarush talked about his decision to join WeWork as their Director of Data Engineering in 2016. * (30:39) Tarush went over the origin and evolution of Marquez - WeWork’s first open-source project around data lineage - during his time as the director of WeWork’s Data Platform team. * (33:49) Tarush highlighted the main challenges of building an internal data platform. * (35:43) Tarush recalled his move to China to help establish WeWork’s Asia operations and focus on the hyper-growing Chinese market. * (39:01) Tarush shared the founding story of 5x during his sabbatical in 2020. * (42:39) Tarush explained the industry's need for a managed data stack. * (45:20) Tarush went over 5x’s process of sourcing, interviewing, and onboarding data engineers who are pre-trained on the modern data stack. * (48:37) Tarush talked about finding the right vendors that make up the modern data stack to partner with. * (50:06) Tarush walked through his production process to put together a lot of good videos to explain what 5x does and raise awareness about the company. * (51:52) Closing segment.
Tarush's Contact Info* LinkedIn * Twitter * Medium
5x Resources* Website | LinkedIn | Twitter | YouTube | Instagram * 5x Explained in 2 Minutes * Managed Data Platform * On-Demand Data Engineering Services * Integrations
Mentioned ContentPeople* George Fraser and Taylor Brown (Founders of Fivetran) * Prukalpa Sankar (Co-Founder and CEO of Atlan) * Frank Slootman (CEO and Chairman of Snowflake)
Books* Stealing Fire (by Steven Kotler and Jamie Wheal) * The 5 AM Club (by Robin Sharma)
About the showDatacast features long-form, in-depth conversations with practitioners and researchers in the data community to walk through their professional journeys and unpack the lessons learned along the way. I invite guests coming from a wide range of career paths — from scientists and analysts to founders and investors — to analyze the case for using data in the real world and extract their mental models (“the WHY and the HOW”) behind their pursuits. Hopefully, these conversations can serve as valuable tools for early-stage data professionals as they navigate their own careers in the exciting data universe.
Datacast is produced and edited by James Le. Get in touch with feedback or guest suggestions by emailing khanhle.1013@gmail.com.
Subscribe by searching for Datacast wherever you get podcasts, or click one of the links below:
If you’re new, see the podcast homepage for the most recent episodes to listen to or browse the full guest list.
Show Notes* (01:59) Hyun shared his upbringing and experience living in Korea, Singapore, and the US. * (04:18) Hyun described his undergraduate experience at Duke University. * (08:21) Hyun shared how he got a real taste of the game-changing potential of deep learning from the experience of bringing ML to diagnose Parkinson’s disease with brain MRI scans. * (10:54) Hyun talked about his journey of leveling up coding and ML knowledge. * (12:13) Hyun reflected on his motivation to pursue a Ph.D. program in computer science at Duke. * (15:22) Hyun talked about his participation in the 2016 Amazon Robotics Challenge as the “Team Duke” leader and its Motion Planning function. * (17:25) Hyun reflected on his decision to take a leave of absence from his Ph.D. program and return to Korea to work as an ML Research Engineer at the AI Research Lab of SK Telecom, a major Korean conglomerate. * (19:46) Hyun discussed his research on game AI and synthetic image generation during his time with SK Telecom. * (22:57) Hyun shared the founding story of Superb AI. * (27:11) Hyun described going through the Y Combinator Winter 2019 batch. * (32:25) Hyun unpacked the evolution of Superb AI’s Labeling platform since its inception. * (34:47) Hyun walked through the process of prioritizing the product roadmap. * (36:54) Hyun zoomed in on Superb AI’s automated labeling feature, Custom Auto-Label, which automatically detects and labels common or niche objects in images and videos. * (40:21) Hyun touched on challenges with manually reviewing and auditing labels. * (42:25) Hyun dissected the data-centric problems in computer vision that the newly released Superb DataOps platform is built to solve. * (46:46) Hyun hinted at Superb AI’s product roadmap, judging from current industry-wide pain points. * (48:53) Hyun highlighted a customer use case of Superb AI product offerings. * (51:42) Hyun shared his vision of where Superb AI fits into the quickly evolving AI Infrastructure ecosystem. * (54:15) Hyun shared valuable hiring lessons to attract people who are excited about Superb AI’s mission. * (58:01) Hyun expanded his perspectives on defining and scaling a global company culture. * (01:00:06) Hyun reflected on the challenges of running a remote-first company. * (01:01:54) Hyun shared fundraising advice for founders seeking the right investors for their startups. * (01:03:35) Hyun highlighted the difference between being a researcher and a founder. * (01:05:08) Closing segment.
Hyun’s Contact Info* LinkedIn * Twitter
Superb AI Resources* Website | LinkedIn | Twitter | YouTube | GitHub | Docs * Superb AI Suite Labeling Platform * Superb AI DataOps Platform * The Ground Truth Newsletter * Superb AI Academy
Mentioned ContentPeople1. Andrew Ng 2. Andrej Karpathy 3. Ian Goodfellow
Book1. Zero To One (by Peter Thiel)
About the showDatacast features long-form, in-depth conversations with practitioners and researchers in the data community to walk through their professional journeys and unpack the lessons learned along the way. I invite guests coming from a wide range of career paths — from scientists and analysts to founders and investors — to analyze the case for using data in the real world and extract their mental models (“the WHY and the HOW”) behind their pursuits. Hopefully, these conversations can serve as valuable tools for early-stage data professionals as they navigate their own careers in the exciting data universe.
Datacast is produced and edited by James Le. Get in touch with feedback or guest suggestions by emailing khanhle.1013@gmail.com.
Subscribe by searching for Datacast wherever you get podcasts, or click one of the links below:
If you’re new, see the podcast homepage for the most recent episodes to listen to or browse the full guest list.
Show Notes* (01:45) Gary walked through his academic experience getting a Bachelor’s degree in Business Administration at Arizona State University and an MBA in Finance at USC — Marshall School of Business. * (04:52) Gary recalled the most valuable lesson from leading a business development team in the enterprise offerings group at Verizon. * (07:45) Gary recalled the challenges of bringing a company public during his time as the Director of Corporate Development at NorthPoint Communications. * (12:18) Gary shared his learnings while holding a COO role at Vinfolio — an innovator in the wine Industry. * (15:19) Gary talked about his responsibilities in the Chief Financial Officer roles at KnowNow and Zuora. * (19:06) Gary gave advice to founders seeking the right investors for their startups. * (23:51) Gary walked through the learning curves while serving as the CFO, CRO, and COO of enterprise AI pioneer Ayasdi. * (31:06) Gary shared his playbook on building a well-oiled sales operations machine. * (33:46) Gary shared his journey as a first-time CEO at CLARA Analytics. * (36:37) Gary talked about his proudest accomplishments while driving significant growth for CLARA. * (37:52) Gary discussed the go-to-market motions implemented at CLARA. * (41:07) Gary walked through his brief stint as an Entrepreneur-In-Residence at Redpoint Ventures, a top-tier VC firm focused on early-stage investing. * (44:14) Gary rationalized his decision to become the CEO of Arcion Labs in December 2021. * (49:39) Gary explained the high-level architectural design of Arcion’s data mobility platform. * (54:19) Gary discussed strategies for finding the right technology partners to collaborate with. * (57:42) Gary highlighted a few customer use cases of Arcion. * (01:01:48) Gary shared valuable hiring lessons to attract the right people who are excited about Arcion’s mission. * (01:04:28) Gary distilled lessons learned while building a high-performance team at Arcion. * (01:09:14) Gary described the benefits of adopting usage-based pricing in enterprise technology. * (01:11:41) Closing segment.
Gary’s Contact Info* LinkedIn * Twitter * Crunchbase
Arcion’s Resources* Website | LinkedIn | Twitter | YouTube | Docs | Slack * “Dawn of the Data Mobility Era” (Feb 2022) * “Arcion lands $13M to help companies replicate data across platforms” (Venture Beat, Feb 2022)
Mentioned ContentContent* The Network Effects Bible (by James Currier of NFX) * Blog by Tomasz Tunguz of Redpoint Ventures
People1. Gurjeet Singh (Co-Founder and CEO of Oma Robotics, Ex-CEO/Co-Founder of Ayasdi) 2. Satish Dharmaraj (Managing Director at Redpoint Ventures)
NotesMy conversation with Gary was recorded back in March 2022. Since then, many things have happened at Arcion. I’d recommend checking out:
About the showDatacast features long-form, in-depth conversations with practitioners and researchers in the data community to walk through their professional journeys and unpack the lessons learned along the way. I invite guests coming from a wide range of career paths — from scientists and analysts to founders and investors — to analyze the case for using data in the real world and extract their mental models (“the WHY and the HOW”) behind their pursuits. Hopefully, these conversations can serve as valuable tools for early-stage data professionals as they navigate their own careers in the exciting data universe.
Datacast is produced and edited by James Le. Get in touch with feedback or guest suggestions by emailing khanhle.1013@gmail.com.
Subscribe by searching for Datacast wherever you get podcasts or click one of the links below:
If you’re new, see the podcast homepage for the most recent episodes to listen to, or browse the full guest list.
Show Notes* (01:43) DeVaris reflected on his upbringing on the south side of Chicago and college experience at UIUC, studying Mathematics and Computer Science in the early 2000s. * (06:46) DeVaris shared his journey of learning how to program, make computers, and dive into the Internet. * (09:35) DeVaris recalled valuable lessons from interning at Intel and Cisco Systems. * (15:49) DeVaris shared his proudest accomplishments during his five years at Microsoft — first as a system engineer and then as an academic developer evangelist. * (22:06) DeVaris recalled his experience working in the gaming and music space as the Chief Developer Evangelist at Marmalade and the Chief Product Officer at Klick Push, respectively. * (27:49) DeVaris provided his perspective on the startup acquisition process. * (29:13) DeVaris unpacked his two years as a platform product manager at Zendesk, where he drove the adoption of the Zendesk Developer Platform for developers to create unique customer experiences. * (35:43) DeVaris revealed the challenges of building a technical community, given his experience at Zendesk. * (38:25) DeVaris recalled his time working for a year as the Lead Product Manager at VSCO — a startup that builds digital tools for the modern creative. * (45:12) DeVaris went over the challenges of building software for brand ambassadors and children’s playtime, given his time as the Head of Product Management at Slyce.io and the CTO at Super Heroic. * (49:39) DeVaris reflected on his desire to scratch his entrepreneurial itch. * (52:00) DeVaris gave advice for early-career technologists on evaluating startup opportunities. * (55:51) DeVaris unpacked the product challenges he encountered while building tools for developers as the Director of Product Management at Heroku. * (58:57) DeVaris touched on his one year as the first platform engineering PM hire at Twitter. * (01:02:18) DeVaris shared the founding story of Meroxa. * (01:04:28) DeVaris dissected how Meroxa’s platform architecture is designed at a high level — including a change data capture service, schema registry, event streaming service, API proxy, and incident automation framework. * (01:06:06) DeVaris explained the technical challenges associated with creating connections between data sources and destinations in real time. * (01:08:37) DeVaris zoomed into Conduit — Meroxa’s open-source, single-binary data integration tool written in Golang that provides developer-friendly streaming data orchestration. * (01:12:32) DeVaris highlighted a few customer use cases of Meroxa. * (01:16:16) DeVaris shared valuable hiring lessons to attract the right people who are excited about Meroxa’s mission and fit with Meroxa’s cultural values. * (01:18:37) DeVaris shared challenges to finding the early design partners & lighthouse customers for Meroxa. * (01:20:24) DeVaris gave advice to founders seeking the right investors for their startups. * (01:22:58) DeVaris gave advice to smart, driven operators looking to explore angel investing. * (01:25:17) DeVaris discussed the remaining barriers that prevent minorities from pursuing a technology career. * (01:30:42) DeVaris imparted lessons from photography and DJ that benefited his career in product. * (01:32:26) Closing segment.
DeVaris’ Contact Info* LinkedIn * Twitter * Website * GitHub
Meroxa’s Resources* Website | LinkedIn | Twitter | YouTube * Careers | Medium Blog * Documentation * Conduit (GitHub | Discord | Twitter | Docs)
Mentioned ContentArticles* “Hello World, Meroxa Style” (April 2021) * “Streaming Your Database Changes with Change Data Capture” (Part 1 + Part 2) * “Conduit: Streaming Data Integration for Developers” (Jan 2022) * “Why Conduit? An Evolutionary Leap Forward for Real-Time Data Integration” (Feb 2022) * “Hello Meroxa 2.0” (April 2022)
Resources for minorities* Kura Labs (A free training and job placement academy for Infrastructure Computing, DevOps, and SRE for students from underserved communities) * Free Code Camp (Learn to code — for free)
Books1. Zero To One (by Peter Thiel) 2. The Hard Thing About Hard Things (by Ben Horowitz)
People1. Tristan Handy (Co-Founder and CEO of dbt Labs) 2. Arjun Narayan (Co-Founder and CEO of Materialize) 3. Benn Stancil (Chief Analytics Officer at Mode Analytics) 4. Chad Sanderson (Head of Data Platform at Convoy)
NotesMy conversation with DeVaris was recorded back in April 2022. Since then, many things have happened at Meroxa. I’d recommend checking out:
About the showDatacast features long-form, in-depth conversations with practitioners and researchers in the data community to walk through their professional journeys and unpack the lessons learned along the way. I invite guests coming from a wide range of career paths — from scientists and analysts to founders and investors — to analyze the case for using data in the real world and extract their mental models (“the WHY and the HOW”) behind their pursuits. Hopefully, these conversations can serve as valuable tools for early-stage data professionals as they navigate their own careers in the exciting data universe.
Datacast is produced and edited by James Le. Get in touch with feedback or guest suggestions by emailing khanhle.1013@gmail.com.
Subscribe by searching for Datacast wherever you get podcasts or click one of the links below:
If you’re new, see the podcast homepage for the most recent episodes to listen to, or browse the full guest list.
Show Notes* (01:49) Cody shared his upbringing in New Jersey, his childhood interest in science and technology, and the few people who have made big differences in his story. * (09:35) Cody went over his academic experience studying Electrical Engineering and Computer Science at MIT. * (17:51) Cody recalled his favorite classes taken at MIT. * (22:43) Cody talked about his engagement in serving as the president of MIT’s chapter of Eta Kappa Nu Honor Society and advancing online education at the MIT Office of Digital Learning. * (31:25) Cody is bullish on the future of digital learning. * (35:43) Cody expanded on his internships with Google throughout his time at MIT — doing local search quality and YouTube analytics. * (42:31) Cody described the challenges of dealing with high-frequency trading data from his one year working as a junior data scientist at the Vendor Data Group of Jump Trading in Chicago. * (46:50) Cody reflected on his decision to embark on a Ph.D. journey in Computer Science at Stanford University. * (51:54) Cody mentioned his participation in the DAWN project, specifically DAWNBench, an end-to-end deep learning benchmark and competition. * (54:21) Cody unpacked the evolution of MLPerf, an industry-standard benchmark for the training and inference performance of ML models. * (56:52) Cody walked through the motivation and empirical work in his paper “Selection via Proxy: Efficient Data Selection for Deep Learning.” * (59:34) Cody discussed his paper “Similarity Search for Efficient Active Learning and Search of Rare Concepts.” * (01:06:32) Cody shared his learnings about bringing ML from research to industry from his advisors, Matei Zaharia and Peter Bailis — who were both academics and startup founders simultaneously. * (01:09:19) Cody went over key trends in the emerging Data-Centric AI community — given his involvement with the Data-Centric AI workshop at NeurIPS 2021 and the DataPerf benchmark suite. * (01:12:19) Cody shared lessons learned about finding product-market fit as the founder of Coactive AI — which brings unstructured data into the world of SQL and the big data tools that teams already love. * (01:15:34) Cody emphasized the importance of focusing on the HR function and defining cultural guiding principles for any early-stage startup founder. * (01:21:05) Cody provided his perspective on the differences and similarities between being a researcher and a founder. * (01:23:47) Closing segment.
Cody’s Contact Info* Website * Twitter * LinkedIn * Google Scholar
Coactive AI’s Resources* Website * Twitter * LinkedIn * Culture Values
Mentioned ContentTalk* “Digging Deeper: How a Few Extra Moments Can Change Lives” (TEDxStanford 2017) * “Data Selection for Data-Centric AI” (Stanford MLSys 2022)
Research* “Probabilistic Use Cases: Discovering Behavioral Patterns for Predicting Certification” (2015) * DAWNBench: An End-to-End Deep Learning Benchmark and Competition (Dec 2017) * “MLPerf: An Industry Standard Benchmark Suite for Machine Learning Performance” (Feb 2020) * “Selection via Proxy: Efficient Data Selection for Deep Learning” (Oct 2020) * “Similarity Search for Efficient Active Learning and Search of Rare Concepts” (July 2021) * DataPerf, a new benchmark suite for machine learning datasets and data-centric algorithms (Dec 2021)
People* Matei Zaharia (Cody’s Ph.D. Advisor, Co-Creator of Apache Spark, Co-Founder of Databricks) * Fei-Fei Li (Professor of Computer Science at Stanford, Creator of ImageNet Dataset) * Michael Bernstein (Professor of Computer Science at Stanford with a focus on Human-Computer Interaction)
Books1. “No Rule Rules: Netflix and the Culture of Reinvention” (by Reed Hastings) 2. “What You Do Is Who You Are: How to Create Your Work Business Culture” (by Ben Horowitz) 3. “The Inner Game of Tennis: The Classical Guide to Peak Performance” (by Timothy Gallwey)
NotesMy conversation with Cody was recorded back in January 2022. Since then, many things have happened at Coactive AI. I’d recommend:
About the showDatacast features long-form, in-depth conversations with practitioners and researchers in the data community to walk through their professional journeys and unpack the lessons learned along the way. I invite guests coming from a wide range of career paths — from scientists and analysts to founders and investors — to analyze the case for using data in the real world and extract their mental models (“the WHY and the HOW”) behind their pursuits. Hopefully, these conversations can serve as valuable tools for early-stage data professionals as they navigate their own careers in the exciting data universe.
Datacast is produced and edited by James Le. Get in touch with feedback or guest suggestions by emailing khanhle.1013@gmail.com.
Subscribe by searching for Datacast wherever you get podcasts or click one of the links below:
If you’re new, see the podcast homepage for the most recent episodes to listen to, or browse the full guest list.
Show Notes* (02:18) Merav talked about her undergraduate experience at McGill University studying Psychology and Sociology. * (04:33) Merav discussed important attributes of an exceptional teacher given her two years teaching elementary special education in NYC public schools through the Teach For America program. * (08:19) Merav commented on her time working at the International Baccalaureate Organization and working as a Kaplan GRE instructor. * (10:57) Merav shared the backstory behind the founding of Data Society, a predictive analytics training and consulting company (co-founded with Dmitri Adler and John Nader). * (14:15) Merav reflected on her journey into programming. * (17:16) Merav explained why data science training should be industry-tailored for maximum success. * (20:57) Merav talked about how Data Society creates and evaluates its training curriculum. * (23:59) Merav provided an example of how Data Society provides customized AI solutions to inform decisions, automate time-consuming manual processes, and solve complex data challenges for its clients. * (27:38) Merav brought up challenges that hinder the adoption of data science in the government sector. * (29:49) Merav unpacked the six different steps for organizations to start moving up the data analytics maturity model. * (33:07) Merav dissected meldR, Data Society’s internal product built for Learning and Development teams in healthcare. * (36:24) Merav reflected on bootstrapping Data Society in the early days (look at this 2016 Kickstarter campaign). * (39:48) Merav discussed the shift from a B2C to a B2B model for Data Society and scoring partnerships with Fortune 500 companies and federal agencies. * (42:47) Merav shared valuable hiring lessons to attract the right people who are excited about the mission of Data Society. * (45:22) Merav shared her experience shaping the remote work culture. * (49:05) Merav touched on initiatives at Data Society to bring more goodness to the world. * (50:28) Merav provided different ways to engage more women in data science (via the Women Data Scientists DC Meetup and DCFemTech). * (53:17) Merav predicted the evolution of education in the next 3 to 5 years. * (55:29) Closing segment.
Merav’s Contact Info* LinkedIn * Twitter
Data Society’s Resources* Website * Twitter * LinkedIn
Mentioned ContentArticles* “Is Your Enterprise Data-Driven?” (May 2021) * “Why Data Science Training Should Be Industry-Tailored for Maximum Success” (August 2021) * “Female Founders: Merav Yuravlivker of Data Society On The Five Things You Need To Thrive and Succeed as a Woman Founder” (Sep 2021)
People* DJ Patil (The first Chief Data Scientist of the US) * Hilary Mason (Co-Founder of Hidden Door) * Avriel Epps-Darling (Ph.D. candidate, Ford fellow, and Presidential Scholar at Harvard University)
Book* Weapons of Math Destruction (by Cathy O’Neil)
NotesMy conversation with Merav was recorded back in December 2021. Since then, many things have happened at Data Society. I’d recommend:
Finally, Merav was also just recognized as one of the DC region's 40 Under 40. The awards are given annually to recognize the outstanding achievements of young leaders in the Washington, DC, area who lead the community forward through hard work, philanthropy, and community engagement.
Show Notes* (01:46) Douwe went over formative experiences catching the programming virus at the age of 9, combining high school with freelance web development, and studying Computer Science at Utrecht University in college. * (03:55) Douwe shared the story behind founding a startup called Stinngo, which led him to join GitLab in 2015 as employee number 10. * (05:29) Douwe provided insights on attributes of exceptional engineering talent, given his time hiring developers and eventually becoming GitLab's first Development Lead. * (08:28) Douwe unpacked the evolution of his engineering career at GitLab. * (11:11) Douwe discussed the motivation behind the creation of the Meltano project in August 2018 to help GitLab's internal data team address the gaps that prevent them from understanding the effectiveness of business operations. * (14:38) Douwe reflected on his decision in 2019 to leave GitLab’s engineering organization and join the then 5-people Meltano team full-time. * (20:24) Douwe shared the details about Meltano's product development journey from its Version 1 to its pivot. * (26:18) Douwe reflected on the mental aspect of being the sole person whom Meltano depended on for a while. * (29:20) Douwe explained the positioning of Meltano as an open-source self-hosted platform for running data integration and transformation pipelines. * (34:54) Douwe shared details of Meltano's ideal customer profiles. * (37:45) Douwe provided a quick tour of the Meltano project, which represents the single source of truth regarding one's ELT pipelines: how data should be integrated and transformed, how the pipelines should be orchestrated, and how the various plugins that make up the pipelines should be configured. * (40:39) Douwe unpacked different components of Meltano's product strategy, including Meltano SDK, Meltano Hub, and Meltano Labs. * (45:05) Douwe discussed prioritizing Meltano's product roadmap in order to bring DataOps functionality to every step of the entire data lifecycle. * (48:53) Douwe shared the story behind spinning Meltano out of GitLab in June 2021 and raising a $4.2M Seed funding round led by GV to bring the benefits of open source data integration and DataOps to a wider audience. * (52:19) Douwe provided his thoughts behind open-source contributors in a way that can generate valuable product feedback for Meltano. * (55:43) Douwe shared valuable hiring lessons to attract the right people who align with Meltano's values. * (59:04) Douwe shared advice to startup CEOs who are experimenting with the remote work culture in our “new-normal” virtual working environments. * (01:04:10) Douwe unpacked Meltano's mission and vision as outlined in this blog post. * (01:06:40) Closing segment.
Douwe's Contact Info* GitLab * LinkedIn * Twitter * GitHub * Website
Meltano's Resources* Website | Twitter | LinkedIn | GitHub | YouTube * Meltano Documentation | Product | DataOps * Meltano SDK | Meltano Hub | Meltano Labs * Company Handbook | Community | Values | Careers
Mentioned ContentArticles* Hey, data teams - We're working on a tool just for you (Aug 2018) * To-do zero, inbox zero, calendar zero: I think that means I'm done (Sep 2019) * Meltano graduates to Version 1.0 (Oct 2019) * Revisiting the Meltano strategy: a return to our roots (May 2020) * Why we are building an open-source platform for ELT pipelines (May 2020) * Meltano spins out of GitLab, raises seed funding to bring data integration into the DataOps era (June 2021) * Meltano: The strategic foundation of the ideal data stack (Oct 2021) * Introducing your DataOps platform infrastructure: Our strategy for the future of data (Nov 2021) * Our next step for building the infrastructure for your Modern Data Stack (Dec 2021)
People* Maxime Beauchemin (Founder and CEO of Preset, Creator of Apache Airflow and Apache Superset, Angel Investor in Meltano) * Benn Stancil (Chief Analytics Officer at Mode Analytics, Well-Known Substack Writer) * The entire team at dbt Labs
NotesMy conversation with Douwe was recorded back in November 2021. Since then, many things have happened at Meltano. I'd recommend:
About the showDatacast features long-form, in-depth conversations with practitioners and researchers in the data community to walk through their professional journeys and unpack the lessons learned along the way. I invite guests coming from a wide range of career paths — from scientists and analysts to founders and investors — to analyze the case for using data in the real world and extract their mental models (“the WHY and the HOW”) behind their pursuits. Hopefully, these conversations can serve as valuable tools for early-stage data professionals as they navigate their own careers in the exciting data universe.
Datacast is produced and edited by James Le. Get in touch with feedback or guest suggestions by emailing khanhle.1013@gmail.com.
Subscribe by searching for Datacast wherever you get podcasts or click one of the links below:
If you’re new, see the podcast homepage for the most recent episodes to listen to, or browse the full guest list.
Show Notes* (01:41) Mars walked through his education studying Computer Systems Engineering at The University of Auckland in New Zealand. * (03:16) Mars reflected on his overall Ph.D. experience in Computer Science at UCLA. * (05:55) Mars discussed his early research paper on a robust and scalable lane departure warning system for smartphones. * (07:13) Mars described his work on SmartFall, an automatic fall detection system to help prevent the elderly from falling. * (08:34) Mars explained his project WANDA, an end-to-end remote health monitoring and analytics system designed for heart failure patients. * (10:06) Mars recalled learnings from interning as a software engineer at Google during his Ph.D. * (14:54) Mars discussed engineering challenges while working on PHP for Google App Engine and Gboard personalization during his subsequent four years at Google. * (19:05) Mars rationalized his decision to join LinkedIn to lead an engineering team that builds the core metadata infrastructure for the entire organization. * (21:15) Mars discussed the motivation behind the creation of LinkedIn’s generalized metadata search and discovery tool, DataHub, later open-sourced in 2020. * (25:21) Mars dissected the key architecture of DataHub, which is designed to address the key scalability challenges coming in four different forms: modeling, ingestion, serving, and indexing. * (28:50) Mars expressed the challenges of finding DataHub’s early adopters internally at LinkedIn and externally later on at other companies. * (35:22) Mars shared the story behind the founding of Metaphor Data, which he co-founded with Pardhu Gunnam and Seyi Adebajo and currently serves as the CTO. * (41:55) Mars unpacked how Metaphor’s modern metadata platform serves as a system of record for any organization’s data ecosystem. * (48:07) Mars described new challenges with metadata management since the introduction of the modern data stack and key features of a great modern metadata platform (as brought up in his in-depth blog post with Ben Lorica). * (53:55) Mars explained how a modern metadata platform fits within the broader data ecosystem. * (58:30) Mars shared the hurdles to finding Metaphor Data’s early design partners and lighthouse customers. * (01:04:33) Mars shared valuable hiring lessons to attract the right people who are excited about Metaphor’s mission. * (01:07:28) Mars shared important culture-building lessons to build out a high-performing team at Metaphor. * (01:10:45) Mars shared fundraising advice for founders currently seeking the right investors for their startups. * (01:13:22) Closing segment.
Mars’ Contact Info* Twitter * LinkedIn * Google Scholar * GitHub
Metaphor Data* Website | Twitter | LinkedIn * Careers | About Page * Data Documentation | Data Collaboration
Mentioned ContentArticles* DataHub: A generalized metadata search and discovery tool (Aug 2019) * Open-sourcing DataHub: LinkedIn’s metadata search and discovery platform (Feb 2020) * Founding Metaphor Data (Dec 2020) * Metaphor and Soda partner to unify the modern data stack with trusted data (Dec 2021) * Introducing Metaphor: The Modern Metadata Platform (Nov 2021) * The Modern Metadata Platform: What, Why, and How? (Jan 2022)
Papers* SmartLDWS: A robust and scalable lane departure warning system for the smartphones (Oct 2009) * SmartFall: An automatic fall detection system based on subsequence matching for the SmartCane (April 2009) * WANDA: An end-to-end remote health monitoring and analytics system for heart failure patients (Oct 2012)
People* Benn Stancil (Chief Analytics Officer at Mode Analytics, Well-Known Substack Writer) * Tristan Handy (Co-Founder and CEO of dbt Labs, Writer of The Analytics Engineering Roundup) * Andy Pavlo (Associate Professor of Database at Carnegie Mellon University)
Books* “Working In Public” (by Nadia Eghbal) * “The Mom Test” (by Rob Fitzpatrick) * “A Thousand Brains” (by Jeff Hawkins) * “The Scout Mindset” (by Julia Galef)
NotesMy conversation with Mars was recorded back in January 2022. Since then, many things have happened at Metaphor Data. I’d recommend:
About the showDatacast features long-form, in-depth conversations with practitioners and researchers in the data community to walk through their professional journeys and unpack the lessons learned along the way. I invite guests coming from a wide range of career paths — from scientists and analysts to founders and investors — to analyze the case for using data in the real world and extract their mental models (“the WHY and the HOW”) behind their pursuits. Hopefully, these conversations can serve as valuable tools for early-stage data professionals as they navigate their own careers in the exciting data universe.
Datacast is produced and edited by James Le. Get in touch with feedback or guest suggestions by emailing khanhle.1013@gmail.com.
Subscribe by searching for Datacast wherever you get podcasts or click one of the links below:
If you’re new, see the podcast homepage for the most recent episodes to listen to, or browse the full guest list.
Show Notes* (01:35) Ville recalled his education getting degrees in Computer Science from the University of Helsinki in Finland. * (04:35) Ville walked over his time working at a startup called Gurusoft that planned to commercialize self-organizing maps, a peculiar artificial neural network. * (07:17) Ville reflected on his four years as a researcher at Nokia — working on big data infrastructure, analytics, and ML open-source projects (such as Disco and Ringo). * (11:56) Ville shared the story of co-founding a startup that built a novel scriptable data platform called Bitdeli with his brother and not finding a product-market fit. * (13:58) Ville walked through AdRoll’s acquisition of Bitdeli in June 2013. * (15:49) Ville discussed the engineering challenges associated with his work at AdRoll — AdRoll Prospecting and traildb.io. * (19:33) Ville mentioned the product and leadership/management lessons during his time being AdRoll’s Head of Data and leading various data/ML efforts. * (24:43) Ville rationalized his decision to join the ML Infrastructure team at Netflix in 2017. * (27:26) Ville discussed the motivation behind the creation of Netflix’s human-centric ML infrastructure, Metaflow, later open-sourced in 2019. * (30:21) Ville unpacked the key design principles that summarize the philosophy of Metaflow, which is influenced by the unique culture at Netflix. * (35:00) Ville talked about his well-known diagram on the data infrastructure’s hierarchy of needs. * (37:33) Ville examined the technical details behind Metaflow’s integration with AWS to make it easy for users to move back and forth between their local and remote modes of development and execution. * (40:58) Ville expressed the challenges of finding Metaflow’s early adopters internally at Netflix and externally later on at other companies. * (45:13) Ville went over the strategy around prioritizing features for Metaflow’s future roadmap. * (52:22) Ville shared the story behind the founding of Outerbounds, which he co-founded with Savin Goyal and Oleg Avdeev. * (55:03) Ville provided his thoughts behind Metaflow’s contributors in a way that can generate valuable product feedback for Outerbounds. * (58:30) Ville shared valuable hiring lessons to attract the right people who are excited about Outerbounds’ mission. * (01:01:28) Ville shared upcoming initiatives that he is most excited about for Outerbounds. * (01:04:05) Ville walked through his writing process for an upcoming technical book with Manning called “Effective Data Science Infrastructure,” a hands-on guide to assembling infrastructure for data science and machine learning applications. * (01:06:34) Ville unpacked his great O’Reilly article that digs deep into the fundamentals of ML as an engineering discipline. * (01:11:03) Closing segment.
Ville’s Contact Info* LinkedIn * Twitter * GitHub
Outerbounds* Website | Twitter | LinkedIn | GitHub | YouTube * Metaflow GitHub | Metaflow Docs * Slack Community * Careers * Metaflow Resources for Data Science * Metaflow Resources for Engineering
Mentioned ContentTalks* SF Data Mining Meetup: TrailDB — Processing Trillions of Events at AdRoll (July 2016) * QConSF 2018: Human-Centric Machine Learning Infrastructure @Netflix (Feb 2019) * AWS re:Invent 2019: More Data Science with Less Engineering — ML Infrastructure at Netflix (Dec 2019) * Scale By The Bay 2019: Human-Centric ML Infrastructure at Netflix (Jan 2020) * AICamp: Metaflow — The ML Infrastructure at Netflix (Aug 2021)
Articles* Open-Sourcing Metaflow, a Human-Centric Framework for Data Science (Netflix Tech Blog, Dec 2019) * Unbundling Data Science Workflows with Metaflow and AWS Step Functions (Netflix Tech Blog, July 2020) * MLOps and DevOps: Why Data Makes It Different (O’Reilly, Oct 2021)
People* Michael Jordan (Distinguished Professor in EECS and Statistics at UC Berkeley) * Matthew Honnibal and Ines Montani (Creators of open-source NLP library spaCy) * Hadley Wickham (Chief Scientist at RStudio and Adjunct Professor of Statistics at Rice University)
Book* “The Mom Test” (by Rob Fitzpatrick)
NotesMy conversation with Ville was recorded back in October 2021. Since then, many things have happened at Outerbounds. I’d recommend:
About the showDatacast features long-form, in-depth conversations with practitioners and researchers in the data community to walk through their professional journeys and unpack the lessons learned along the way. I invite guests coming from a wide range of career paths — from scientists and analysts to founders and investors — to analyze the case for using data in the real world and extract their mental models (“the WHY and the HOW”) behind their pursuits. Hopefully, these conversations can serve as valuable tools for early-stage data professionals as they navigate their own careers in the exciting data universe.
Datacast is produced and edited by James Le. Get in touch with feedback or guest suggestions by emailing khanhle.1013@gmail.com.
Subscribe by searching for Datacast wherever you get podcasts or click one of the links below:
If you’re new, see the podcast homepage for the most recent episodes to listen to, or browse the full guest list.
Show Notes* (01:48) Mike recalled his undergraduate experience studying Economics at Arizona State University and doing research on statistics/econometrics. * (04:59) Mike reflected on his three years working as an analyst in the Boston office of the Analysis Group. * (09:08) Mike discussed how he leveled up his programming skills at work. * (11:05) Mike shared his learnings about building effective data-driven products while working as a data scientist at Case Commons. * (17:20) Mike revisited his transition to a new role as the Director of Analytics at Harry’s, the men’s grooming brand — starting a new data team from scratch. * (23:04) Mike unpacked analytics and infrastructure challenges during his time at Harry’s — developing the data warehouse, an internal marketing attribution tool, and a fleet of systems for automated decision-making to improve efficiency. * (27:21) Mike reasoned his move to Mexico City — spending time practicing Spanish, among other things. * (32:22) Mike talked about his journey of starting a new consulting practice to help companies get more value out of their data, which was primarily shaped by his network. * (36:30) Mike shared the founding story behind Recast, whose mission is to help modern brands improve the effectiveness of their marketing dollars. * (42:09) Mike dissected the core technical problem that Recast is addressing: performing media mix modeling in the context of “programmatic” channels. * (46:14) Mike shared the story behind the inception and evolution of Locally Optimistic, a community for current and aspiring data analytics leaders. * (49:29) Mike walked through his 3-part blog series on Agile Analytics — discussing the good aspects, the bad aspects, and the adjustments needed for analytics teams to adopt the Scrum methodology. * (53:25) Mike unpacked his post “A Culture of Partnership,” — which discusses the three key activities that can help an analytics team identify the most important opportunities in the business and work effectively with key stakeholders and partner teams to drive value. * (57:25) Mike examined his seminal piece called “The Analytics Engineer,” which generated much attention from the analytics community — which argues that the analytics engineer can provide a multiplier effect on the output of an analytics team. * (01:03:24) Mike shared the motivation and pedagogical philosophy behind the Analytics Engineers Club (co-founded with Claire Carroll), which provides a training course for data analysts looking to improve their engineering skills. * (01:07:57) Mike anticipated the evolution of the quickly evolving modern data stack (read his Fivetran article “The Modern Data Science Stack”). * (01:09:22) Mike unpacked how organizations can build, start, and maintain the data quality flywheel (read his Datafold article “The Data Quality Flywheel”). * (01:11:40) Mike shared his thoughts regarding the challenge of sharing complex analyses. * (01:13:15) Closing segment.
Mike’s Contact Info* Twitter * Website * LinkedIn * GitHub
Further Resources* Recast * Locally Optimistic * Analytics Engineers Club
Mentioned ContentArticles* “Learning a language is hard” (Personal Blog, Jan 2020) * “Modern Media Mix Modeling” (Recast Blog) * “Agile Analytics, Part 1: The Good Stuff” (Locally Optimistic Blog, May 2018) * “Agile Analytics, Part 2: The Bad Stuff” (Locally Optimistic Blog, June 2018) * “Agile Analytics, Part 3: The Adjustments” (Locally Optimistic Blog, July 2018) * “A Culture of Partnership” (Locally Optimistic Blog, March 2019) * “The Analytics Engineer” (Locally Optimistic Blog, Jan 2019) * “Data Education Is Broken” (Analytics Engineering Club, June 2021) * “Teaching The Real Tools” (Analytics Engineering Club, Aug 2021) * “The Modern Data Science Stack” (Fivetran Blog, Oct 2020) * “The Data Quality Flywheel” (Datafold Blog, Nov 2020) * “Knowledge Sharing” (Personal Blog, Sep 2020) * “TDD for ELT” (Personal Blog, Sep 2020) * “Are Data Catalogs Curing the Symptom or the Disease?” (Personal Blog, Dec 2020)
People* Claire Carroll (Co-Instructor of Analytics Engineering Club, Product Manager of Hex, previous Community Manager of dbt Labs) * Drew Banin (Head of Product at dbt Labs) * Barry McCardel (Co-Founder and CEO of Hex)
NotesMy conversation with Michael was recorded back in October 2021. Since then, Michael has been active in his work projects. I’d recommend:
About the showDatacast features long-form, in-depth conversations with practitioners and researchers in the data community to walk through their professional journeys and unpack the lessons learned along the way. I invite guests coming from a wide range of career paths — from scientists and analysts to founders and investors — to analyze the case for using data in the real world and extract their mental models (“the WHY and the HOW”) behind their pursuits. Hopefully, these conversations can serve as valuable tools for early-stage data professionals as they navigate their own careers in the exciting data universe.
Datacast is produced and edited by James Le. Get in touch with feedback or guest suggestions by emailing khanhle.1013@gmail.com.
Subscribe by searching for Datacast wherever you get podcasts or click one of the links below:
If you’re new, see the podcast homepage for the most recent episodes to listen to, or browse the full guest list.
Show Notes* (01:37) Caitlin went over her college experience studying Computer Science at Stanford University in the early 2010s. * (03:55) Caitlin talked about her teaching experience for CS 106A and CS 103. * (07:09) Caitlin shared valuable lessons from completing software engineering internships at Harvard University, Facebook, and Palantir. * (10:06) Caitlin walked over technical and organizational challenges during her time at Palantir — building products for both government/commercial customers and working with designers/infrastructure engineers to deliver full-stack applications to the field. * (12:01) Caitlin explained why Palantir is composed of “loosely individual startups.” * (14:56) Caitlin recalled learning curves during her transition to a tech lead role at Palantir — becoming responsible for the technical architecture and code quality of the product, mentorship and growth of the engineers, and the product direction and prioritization of features. * (18:31) Caitlin discussed her time as a Data Engineering Manager at Remix Technologies — leading the team that builds geospatial data pipelines on top of AWS, Postgres/PostGIS, and Apache Airflow. * (24:45) Caitlin reflected on valuable leadership and people management lessons absorbed during her transition to growing and developing diverse and inclusive engineering teams. * (29:05) Caitlin shared the founding story of Hex, the modern data workspace for teams, alongside her co-founders Barry and Glen. * (32:58) Caitlin talked about Hex’s ideal users (the “analytically technical” who need better tools to access and manage more sophisticated workflows) and introduced Hex’s Logic View. * (35:22) Caitlin examined the collaboration challenges in data teams and revealed Hex’s Library to address some of the shortcomings. * (39:59) Caitlin shared her thoughts on the evolution of data science notebooks. * (42:14) Caitlin unpacked the nuanced problem of justifying data ROI to functional stakeholders and described Hex’s interactive App Builder. * (45:17) Caitlin shared exciting development in the horizon of Hex’s product roadmap. * (46:37) Caitlin shared valuable hiring lessons to attract the right people who are excited about Hex’s mission. * (52:10) Caitlin shared the hurdles to find the early design partners and lighthouse customers of Hex. * (56:01) Caitlin shared upcoming go-to-market initiatives that she’s most excited about for Hex. * (58:24) Caitlin shared fundraising advice for founders currently seeking the right investors for their startups. * (01:01:42) Closing segment.
Caitlin’s Contact Info* LinkedIn * Twitter
Hex’s Resources* Website | Twitter | LinkedIn * Logic View | App Builder | Knowledge Library * Docs | Blog | Gallery * Customers | Careers | Integrations | Pricing
Mentioned ContentArticles* “Long Live Code” (June 2020) * “Don’t Tell Your Data Team’s ROI Story” (Aug 2020) * “The Sharing Gap” (Oct 2020)
People* Tristan Handy (Founder and CEO of dbt Labs) * Claire Carroll (Product Manager of Hex, previous Community Manager of dbt Labs) * Wes McKinney (Creator of Pandas and Arrow, Co-Founder and CTO of Voltron Data) * DeVaris Brown (Co-Founder and CEO of Meroxa)
Book* “Mindset: The New Psychology of Success” (by Carol Dweck)
NotesMy conversation with Caitlin was recorded back in Fall 2021. Since then, many things have happened at Hex. I’d recommend looking at:
Show Notes* (00:43) Kashish shared briefly about his upbringing in Atlanta and his early interest in STEM subjects. * (02:38) Kashish described his overall academic experience studying Economics, Management, and Computer Science at the University of Pennsylvania. * (05:53) Kashish walked over the Machine Learning classes and projects throughout his MSE degree in Robotics. * (09:02) Kashish shared valuable lessons learned from multiple internships throughout his undergraduate: data science at Implantable Provider Group, investment analysis at Tree Line, and product management at LYNK. * (13:14) Kashish told the anecdotes that enabled him to realize his passion for building startups. * (17:14) Kashish recapped his learning about venture capital from spending a summer as an analyst in early-stage deep-tech companies at Bessemer Venture Partners in New York. * (22:09) Kashish shared learnings from his entrepreneurial stints at an early age. * (26:12) Kashish talked through his decision to move to San Francisco after college (Read his blog post explaining how he moved here without a job and a home). * (29:04) Kashish recalled his experience working on a project called Carry (an executive assistant for travel on Slack) with his friend Tejas Manohar and going through Y Combinator. * (36:40) Kashish shared the founding story of Hightouch, a data platform that syncs customer data from the data warehouse to CRM, marketing, and support tools. * (44:15) Kashish emphasized the importance of speed and execution around different pivots that led to Hightouch. * (46:35) Kashish unpacked the notion of Operational Analytics, an approach to analytics that shifts the focus from simply understanding data to putting that data to work in the tools that run your business. * (49:46) Kashish dissected Hightouch’s market-leading Reverse ETL, which is the process of copying data from a data warehouse to operational systems of record. * (54:51) Kashish discussed Hightouch Audiences, used primarily by larger B2C customers, that allows marketing teams to build audiences and filters on top of existing data models. * (58:09) Kashish explained how the “Reverse ETL” concept fits into the quickly evolving modern data stack. * (01:00:26) Kashish shared how the Hightouch team prioritizes their product roadmap, given the high number of customer requests. * (01:02:47) Kashish shared valuable hiring lessons to attract the right people who are excited about Hightouch’s mission. * (01:05:13) Kashish shared the hurdles to find the early design partners and lighthouse customers of Hightouch. * (01:08:06) Kashish explained how Hightouch prices by destinations, reflecting the value customers get from using the product and helping them predict costs over time. * (01:10:32) Kashish shared upcoming go-to-market initiatives that he is most excited about for Hightouch. * (01:14:36) Kashish shared fundraising advice for founders currently seeking the right investors for their startups. * (01:17:47) Kashish emphasized the industry recognition of the Reverse ETL market. * (01:19:47) Closing segment.
Kashish’s Contact Info* LinkedIn * Twitter * GitHub * Website * Medium
Hightouch’s Resources* Website | Twitter | LinkedIn * Data Features | Hightouch Audiences | Hightouch Notify * Docs | Blog * Customers | Careers | Pricing
Mentioned ContentArticles* “On Moving to SF Jobless and Homeless” (Aug 2018) * “Hightouch Ushers In The Era of Operational Analytics” (March 2021) * “The State of Reverse ETL” (May 2021) * “What is Operational Analytics?” (July 2021) * “Hightouch Has Raised a Series A!” (July 2021) * “Hightouch Raises $12M to Empower Business Teams With Operational Analytics” (July 2021) * “The Cloud 100 Rising Stars 2021” (Aug 2021) * “What is Reverse ETL?” (Nov 2021)
Companies* dbt Labs * Shipyard * Big Time Data
Book* “The Hard Things About Hard Things” (by Ben Horowitz)
NotesMy conversation with Kashish was recorded back in August 2021. Since then, many things have happened at Hightouch. I’d recommend looking at:
Finally, Kashish lets me know that back in August, Hightouch were only 25 people. Now, the company is 70-person strong!
About the showDatacast features long-form, in-depth conversations with practitioners and researchers in the data community to walk through their professional journeys and unpack the lessons learned along the way. I invite guests coming from a wide range of career paths — from scientists and analysts to founders and investors — to analyze the case for using data in the real world and extract their mental models (“the WHY and the HOW”) behind their pursuits. Hopefully, these conversations can serve as valuable tools for early-stage data professionals as they navigate their own careers in the exciting data universe.
Datacast is produced and edited by James Le. Get in touch with feedback or guest suggestions by emailing khanhle.1013@gmail.com.
Subscribe by searching for Datacast wherever you get podcasts or click one of the links below:
If you’re new, see the podcast homepage for the most recent episodes to listen to, or browse the full guest list.
Show Notes (01:53) Alessya shared her formative experiences growing up in Kazakhstan, coming to Washington during high school, and discovering a passion and extreme aptitude for mathematics. * (04:20) Alessya described her undergraduate experience studying Applied Mathematics at the University of Washington.* * (08:00) Alessya talked about impactful projects she contributed to while working as a software developer at Amazon’s quality assurance and DevOps organizations. * (12:29) Alessya went over critical responsibilities during her time as Amazon’s Technical Program Manager. * (17:06) Alessya talked about the process of building and getting adoption for an internal Machine Learning platform at Amazon. * (20:42) Alessya shared her biggest takeaways from Amazon’s culture of customer obsession and operational excellence. * (23:26) Alessya revisited her period enrolling in UW’s Master of Science in Entrepreneurship Program and highlighted two core entrepreneurial muscles developed: networking and negotiation. * (28:58) Alessya provided insights on the startup ecosystem and ML community in Seattle. * (34:47) Alessya walked through her period serving as the CTO in Residence at Allen Institute for AI and evaluating a range of AI technologies for viability and product readiness. * (37:12) Alessya shared the backstory behind the founding of WhyLabs, an AI observability platform built to enable every enterprise to run AI with certainty (read her blog post about early misadventures with AI at Amazon that inspired the incubation of WhyLabs at AI2). * (42:23) Alessya examined what makes an AI solution robust and responsible. * (46:09) Alessya dissected the anatomy of an enterprise AI Observability platform. * (49:58) Alessya explained why data logging is a critical missing component in the production ML stack and described whylogs, an open-source ML data logging library from WhyLabs. * (54:12) Alessya shared valuable hiring lessons to attract the right people who are excited about WhyLabs’ mission. * (57:03) Alessya shared tactics to find and engage contributors to whylogs. * (58:10) Alessya shared the hurdles to find the early design partners and lighthouse customers of WhyLabs. * (01:02:28) Alessya shared upcoming go-to-market initiatives that she is most excited about for WhyLabs. * (01:03:54) Alessya explained what it felt to be recognized as the CEO of the year for the Pacific Northwest startup community last year and shared her perspective on work-life balance. * (01:07:43) Closing segment.
Alessya’s Contact Info* LinkedIn * Twitter
WhyLabs’s Resources* Website * whylogs * Slack Community * Blog * LinkedIn | Twitter | Facebook | YouTube | GitHub * What is AI Observability?
Mentioned ContentArticles + Talks* “Introducing WhyLabs, a Leap Forward in AI Reliability” (Sep 2020) * “WhyLabs: The AI Observability Platform” (Sep 2020) * “whylogs: Embrace Data Logging Across Your ML Systems” (Sep 2020) * “Who Said Moms Can’t CEO?” (May 2021) * “The Critical Missing Component in the Production ML Stack” (May 2021)
People* Cassie Kozyrkov (Chief Decision Scientist at Google) * Dan Jeffries (Chief Evangelist at Pachyderm and Founder of AI Infrastructure Alliance) * Michael Petrochuk (Founder and CTO of WellSaid Labs)
Book* “The Hard Things About Hard Things” (by Ben Horowitz)
NotesMy conversation with Alessya was recorded back in August 2021. Since then, many things have happened at WhyLabs.I'd recommend looking at:
whylogs is evolving to a new iteration that will be even more usable and more useful than it was before. With the launch of whylogs v1 in May, users will be able to create data profiles in a fraction of the time and with a much simpler API. Additionally, WhyLabs built-in handy features such as the profile visualizer (which allows users to visualize one or multiple profiles for exploration and comparison) and constraints (which allow users to validate the quality of their data as it flows through their data pipelines).
About the showDatacast features long-form, in-depth conversations with practitioners and researchers in the data community to walk through their professional journeys and unpack the lessons learned along the way. I invite guests coming from a wide range of career paths — from scientists and analysts to founders and investors — to analyze the case for using data in the real world and extract their mental models (“the WHY and the HOW”) behind their pursuits. Hopefully, these conversations can serve as valuable tools for early-stage data professionals as they navigate their own careers in the exciting data universe.
Datacast is produced and edited by James Le. Get in touch with feedback or guest suggestions by emailing khanhle.1013@gmail.com.
Subscribe by searching for Datacast wherever you get podcasts or click one of the links below:
If you’re new, see the podcast homepage for the most recent episodes to listen to, or browse the full guest list.
Show Notes* (02:00) Evan shared his upbringing, born and raised in a small coastal town on New Zealand’s North Island and later studied Software Engineering and Business. * (03:55) Evan recalled working as a software solution architect at NEC Corporation back in New Zealand. * (06:17) Evan talked about his decision to join Twilio in 2011 as one of the company’s early employees right after its Series B financing. * (08:40) Evan shared his perspectives on joining startups and big companies as a new grad. * (13:01) Evan provided insights on attributes of exceptional sales engineers, given his time building the first iteration of Twilio’s global pre-sales team. * (17:30) Evan unpacked the evolution of his career at Twilio — working as a product manager, a director of product & engineering, and a general manager of IoT & wireless. * (22:51) Evan dissected Twilio’s unique “middle-out” sales strategy, which has hugely impacted the company’s incredible growth from Series B through to IPO and beyond. * (29:03) Evan went over the untapped opportunity being enabled by new cellular IoT technologies. * (33:25) Evan explained his decision to embark on a new journey as the CEO of Fin.com after a decade at Twilio. * (37:26) Evan talked about the need for workflow automation and how Fin’s product features are built to address that. * (40:35) Evan went over Fin’s remote performance optimization capabilities that help teams thrive in a remote-first environment. * (42:56) Evan shared valuable hiring lessons to attract the right leaders who are excited about Fin’s mission. * (45:38) Evan shared the hurdles his team has to go through while finding early customers for Fin (as it pivoted to building a SaaS product). * (48:02) Evan talked about the qualities of Jeff Lawson that made him such a great CEO. * (50:41) Closing segment.
Evan’s Contact Info* Twitter * LinkedIn
Fin’s Resources* Website * LinkedIn * Twitter * “Fin.com Raises $20M from Coatue” (Sep 2021) * “Customers Operations Benchmarks for 2022” (Nov 2021) * “Fin’s new Experiments Product Enables CX teams to Confidently Deliver Business Process Changes that Maximize Business Impact” (Dec 2021)
Mentioned ContentPeople* Jack Dorsey * Bret Taylor * Paul Buchheit
Book* “Startup CXO: A Field Guide to Scaling Up Your Company’s Critical Functions and Teams” (by Matt Blumberg)
About the showDatacast features long-form, in-depth conversations with practitioners and researchers in the data community to walk through their professional journeys and unpack the lessons learned along the way. I invite guests coming from a wide range of career paths — from scientists and analysts to founders and investors — to analyze the case for using data in the real world and extract their mental models (“the WHY and the HOW”) behind their pursuits. Hopefully, these conversations can serve as valuable tools for early-stage data professionals as they navigate their own careers in the exciting data universe.
Datacast is produced and edited by James Le. Get in touch with feedback or guest suggestions by emailing khanhle.1013@gmail.com.
Subscribe by searching for Datacast wherever you get podcasts or click one of the links below:
If you’re new, see the podcast homepage for the most recent episodes to listen to, or browse the full guest list.
Show Notes* (01:51) Nick shared his formative experiences of her childhood — moving between different schools, becoming interested in Math, and graduating from UCLA at the age of 19. * (05:45) Nick recalled working as a quant analyst focused on emerging market debt at BlackRock. * (09:57) Nick went over his decision to join Airbnb as a data scientist on their growth team in 2014. * (12:17) Nick discussed how data science could be used to drive community growth on the Airbnb platform. * (16:35) Nick led the data architecture design and experimentation platform for Airbnb Trips, one of Airbnb’s biggest product launches in 2016. * (20:40) Nick provided insights on attributes of exceptional data science talent, given his time interviewing hundreds of candidates to build a data science team from 20 to 85+. * (23:50) Nick went over his process of leveling up his product management skillset — leading Airbnb’s Machine Learning teams and growing the data organization significantly. * (26:56) Nick emphasized the importance of flexibility in his work routine. * (29:27) Nick unpacked the technical and organizational challenges of designing and fostering the adoption of Bighead, Airbnb’s internal framework-agnostic, end-to-end platform for machine learning. * (34:54) Nick recalled his decision to leave Airbnb and become the Head of Data at Branch, which delivers world-class financial services to the mobile generation. * (37:24) Nick unpacked key takeaways from his Bay Area AI meetup in 2019 called “ML Infrastructure at an Early Stage Startup” related to his work at Branch. * (40:55) Nick discussed his decision to pursue a startup idea in the analytics space rather than the ML space. * (43:36) Nick shared the founding story of Transform, whose mission is to make data accessible by way of a metrics store. * (49:54) Nick walked through the four key capabilities of a metrics store: semantics, performance, governance, and interfaces + introduced Metrics Framework (Transform’s capability to create company-wide alignment around key metrics that scale with an organization through a unified framework). * (55:58) Nick unpacked Metrics Catalog — Transform’s capability to eliminate repetitive tasks by giving everyone a single place to collaborate, annotate data charts, and view personalized data feeds. * (59:57) Nick dissected Metrics API — Transform’s capability to generate a set of APIs to integrate metrics into any other enterprise tools for enriched data, dimensional modeling, and increased flexibility. * (01:02:41) Nick explained how metrics store fit into a modern data analytics stack * (01:05:57) Nick shared valuable hiring lessons finding talents who fit with Transform’s cultural values. * (01:12:27) Nick shared the hurdles his team has to go through while finding early design partners for Transform. * (01:15:38) Nick shared upcoming go-to-market initiatives that he’s most excited about for Transform. * (01:17:46) Nick shared fundraising advice for founders currently seeking the right investors for their startups. * (01:20:45) Closing segment.
Nick’s Contact Info* LinkedIn * Twitter * Medium
Transform's Resources* Website * Blog * LinkedIn | Twitter
Mentioned ContentArticles + Talks* “ML Infrastructure at an Early Stage” (March 2019) * “Why We Founded Transform” (June 2021) * “My Experience with Airbnb’s Early Metrics Store” (June 2021) * “The 4 Pillars of Our Workplace Culture” (Aug 2021)
People* Airbnb’s Metrics Repo Team (Paul Yang, James Mayfield, Will Moss, Jonathan Parks, and Aaron Keys) * Maxime Beauchemin (Founder and CEO of Preset, Creator of Apache Airflow and Apache Superset) * Emilie Schario (Data Strategist In Residence at Amplify Partners, Previously Head of Data at Netlify)
Book* “High-Output Management” (by Andy Grove)
NotesMy conversation with Nick was recorded back in July 2021. Since then, many things have happened at Transform. I’d recommend:
About the showDatacast features long-form, in-depth conversations with practitioners and researchers in the data community to walk through their professional journeys and unpack the lessons learned along the way. I invite guests coming from a wide range of career paths — from scientists and analysts to founders and investors — to analyze the case for using data in the real world and extract their mental models (“the WHY and the HOW”) behind their pursuits. Hopefully, these conversations can serve as valuable tools for early-stage data professionals as they navigate their own careers in the exciting data universe.
Datacast is produced and edited by James Le. Get in touch with feedback or guest suggestions by emailing khanhle.1013@gmail.com.
Subscribe by searching for Datacast wherever you get podcasts or click one of the links below:
If you’re new, see the podcast homepage for the most recent episodes to listen to, or browse the full guest list.
Show Notes* (01:29) Jeremiah reflected on his academic interest studying Statistics and Economics at Harvard. * (05:33) Jeremiah recalled his four years as a market risk manager at King Street Capital Management. * (07:18) Jeremiah explained how the training in risk management has made a huge impact in his career as a startup founder. * (09:48) Jeremiah then founded his own consultancy Lowin Data Company that designed and built ML systems for time series data. * (12:38) Jeremiah mentioned his fascination with the rapid growth of machine learning in the past decade. * (15:54) Jeremiah talked about his contribution to the Apache Airflow project and lessons learned about open-source development/governance. * (21:48) Jeremiah unpacked the notion of negative engineering and shared the story behind the inception of Prefect. * (27:24) Jeremiah dissected Prefect Core, the open-source framework that is stocked with all the necessary components for designing, building, testing, and running powerful data applications. * (32:45) Jeremiah went over the advanced enterprise features of Prefect Cloud that complement users of Prefect Core. * (36:04) Jeremiah discussed Prefect's product strategy (read his blog post "Toward Dataflow Automation," which distinguishes the difference between what a company makes and what a company sells). * (40:44) Jeremiah explained how Prefect users can take advantage of the hybrid execution model. * (47:08) Jeremiah walked through Prefect Server and Prefect UI that enable users to run parts of Prefect Cloud locally. * (50:27) Jeremiah talked about how his team has gradually open-sourced the Prefect platform. * (51:38) Jeremiah explained how Prefect settles into a "success-based pricing" model, where the cost is based entirely on the number of tasks users run successfully each month. * (54:15) Jeremiah shared how to nurture a highly active community of open-source contributors to Prefect Core. * (58:23) Jeremiah unpacked Prefect's hiring strategy, which emphasizes the importance of hiring a team diverse in thoughts, backgrounds, makeups, and experiences (read this fantastic guide to building a high-performance team on Prefect's website). * (01:07:02) Jeremiah shared fundraising advice for founders currently seeking the right investors for their startups. * (01:11:53) Jeremiah unpacked the two key pillars central to Prefect’s hyper-adoption within the data world: expansion and product. * (01:14:09) Closing segment.
Jeremiah's Contact Info* LinkedIn * Twitter * Medium * GitHub
Prefect's Resources* Website * GitHub | Slack | Documentation | Twitter | Meetup * Community Updates * The Prefect Guide to Building A High-Performance Team (April 2021) * Prefect Cloud * Prefect Core * Prefect's Hybrid Model
Mentioned ContentArticles* "Positive and Negative Engineering" (Oct 2018) * "The Golden Spike" (Jan 2019) * "Prefect is Open-Source!" (March 2019) * "Towards Dataflow Automation" (June 2019) * "The Prefect Hybrid Model" (Feb 2020) * "Project Earth" (March 2020) * "Open-Sourcing The Prefect Platform" (March 2020) * "Your Code Will Fail (But That's Okay)" (May 2020) * "Liftoff: Prefect's Series A" (Feb 2021) * "Escape Velocity: Prefect's Series B" (June 2021)
Talks and Podcasts* "Invest Like The Best" (Jan 2017) * "Task Failed Successfully" (PyData DC 2018) * "Software Engineering Daily" (April 2020) * "The OSS Startup Podcast" (Nov 2021) * "The Sequel Show" (Jan 2022)
People* Vicki Boykis (ML Engineer at Tumblr, Newsletter Writer of Normcore Tech) * Chris Riccomini (Software Engineer at WePay, Contributor of Airflow, Investor/Advisor at Prefect) * Justin Gage (Newsletter Writer of Technically)
Books* "Creativity Inc." (by Ed Cadmull) * "The Hitchhiker's Guide to the Galaxy" (by Douglas Adams, Eoin Colfer, and Thomas Tidholm) * "Shoe Dog" (by Phil Knight)
NotesMy conversation with Jeremiah was recorded back in July 2021. Since then, many things have happened at Prefect:
About the showDatacast features long-form, in-depth conversations with practitioners and researchers in the data community to walk through their professional journeys and unpack the lessons learned along the way. I invite guests coming from a wide range of career paths — from scientists and analysts to founders and investors — to analyze the case for using data in the real world and extract their mental models (“the WHY and the HOW”) behind their pursuits. Hopefully, these conversations can serve as valuable tools for early-stage data professionals as they navigate their own careers in the exciting data universe.
Datacast is produced and edited by James Le. Get in touch with feedback or guest suggestions by emailing khanhle.1013@gmail.com.
Subscribe by searching for Datacast wherever you get podcasts or click one of the links below:
If you’re new, see the podcast homepage for the most recent episodes to listen to, or browse the full guest list.
Show Notes* (02:00) Shinji reflected on her academic experience studying Software Engineering at the University of Waterloo in the late 2000s. * (04:19) Shinji shared valuable lessons learned from her undergraduate co-op experience with statistical analysis at Sun Microsystems, software engineering at Barclays Capital, and growth marketing at Facebook. * (08:52) Shinji shared lessons learned from being a Management Consultant at Deloitte. * (14:01) Shinji revisited her decision to quit the job at Deloitte and create a social puzzle game called Shufflepix. * (17:42) Shinji went over her time working as a Product Manager at the mobile ad exchange network YieldMo. * (22:25) Shinji discussed the problem of stream processing at YieldMo, which sparked the creation of Concord. * (26:17) Shinji unpacked the pain points with existing stream processing frameworks and the competitive advantage of using Concord. * (33:19) Shinji recalled her time at Akamai — initially as a data engineer in the Platform Engineering unit and later as a product manager for the IoT Edge Connect platform. * (37:26) Shinji explained why sharing context knowledge around data remains a largely unsolved problem. * (42:07) Shinji unpacked the three capabilities of an ideal data discovery platform: (1) exposing up-to-date operational metadata along with the documentation, (2) tracking the provenance of data back to its source, and (3) guiding data usage. * (46:59) Shinji unpacked the benefits of plugging BI tools into data discovery platforms and collecting metadata, which facilitates better visibility and understanding. * (52:36) Shinji discussed the role of a data discovery platform within the modern data stack. * (53:59) Shinji shared the hurdles that her team has to go through while finding early adopters of Select Star. * (55:48) Shinji shared valuable hiring lessons learned at Select Star. * (01:00:00) Shinji shared fundraising advice for founders currently seeking the right investors for their startups. * (01:04:41) Closing segment.
Shinji’s Contact Info* LinkedIn * Twitter * Medium
Select Star’s Resources* Website * Blog * LinkedIn | Twitter | Medium
Mentioned ContentArticles* “The Next Evolution of Data Catalogs: Data Discovery Platforms” (Feb 2021) * “Data Discovery for Business Intelligence” (May 2021)
People* Martin Kleppmann (Author of Designing Data-Intensive Applications) * Emily Riederer (Senior Analytics Manager at Capital One) * Anya Prosvetova (Tableau DataDev Ambassador)
Book* “Managing Oneself” (by Peter Drucker)
NotesMy conversation with Shinji was recorded back in July 2021. Since then, many things have happened at Select Star:
About the showDatacast features long-form, in-depth conversations with practitioners and researchers in the data community to walk through their professional journeys and unpack the lessons learned along the way. I invite guests coming from a wide range of career paths — from scientists and analysts to founders and investors — to analyze the case for using data in the real world and extract their mental models (“the WHY and the HOW”) behind their pursuits. Hopefully, these conversations can serve as valuable tools for early-stage data professionals as they navigate their own careers in the exciting data universe.
Datacast is produced and edited by James Le. Get in touch with feedback or guest suggestions by emailing khanhle.1013@gmail.com.
Subscribe by searching for Datacast wherever you get podcasts or click one of the links below:
If you’re new, see the podcast homepage for the most recent episodes to listen to, or browse the full guest list.
Show Notes* (02:27) Taimur reflected on his education studying Computer Science at UT Austin in the early 2000s. * (06:26) Taimur recalled his first job working as a quality assurance engineer at Vignette. * (08:47) Taimur went through his time with Oracle / Siebel, where he transitioned from a purely technical engineering-focused role to more customer-facing functions. * (13:44) Taimur reflected on his proudest accomplishments at Oracle. * (18:23) Taimur recalled dropping out of studying at the Stanford Center of Professional Development and moving to Seattle to work for Amazon Web Services. * (20:35) Taimur provided insights on attributes of exceptional sales talent, given his time as an enterprise sales manager in his first two years at AWS. * (23:55) Taimur shared anecdotes of successful product launches and their market expansion strategies while leading business development for AWS's database and compute services. * (28:33) Taimur discussed instituting the culture of customer obsession and operational excellence into his teams - while leading the incubation, market development, and technical go-to-market strategy and execution for the AWS Platform across infrastructure, data, developer services, and emerging technologies. * (33:14) Taimur talked about his decision to join Microsoft to lead the Worldwide Customer Success function for their Azure Data Platform, Analytics, and AI business. * (36:24) Taimur unpacked his talk called “Enabling Customer Success through Evolutionary Architectures.” * (43:07) Taimur compared the BizOps culture between Azure and AWS. * (46:29) Taimur discussed his decision to onboard Redis as their Chief Business Development Officer. * (50:07) Taimur went over the data challenges with operational ML, the emerging data architecture of feature stores, and the powerful capabilities of Redis as a solution. * (55:58) Taimur unpacked key ideas in his talk "First Principles in Building A Real-Time AI Platform." * (01:01:52) Taimur hinted at Redis' product vision of "caching for ML data." * (01:05:21) Taimur gave advice for a smart, driven operator who wants to explore angel investing. * (01:10:17) Taimur described the evolution of tech leadership, strategic business development, and customer success strategies in the past two decades. * (01:15:29) Taimur shared three books that have greatly influenced his life. * (01:16:48) Closing segment.
Taimur's Contact Info* LinkedIn * Twitter * Redis Profile
Redis' Resources* Website * Redis Open Source | Redis Enterprise Software | Redis Enterprise Cloud * Redis AI * LinkedIn | Twitter | Facebook | YouTube * "Redis Labs Becomes Redis" (Aug 2021)
Mentioned ContentPeople* Andy Jassy (CEO of Amazon) * Melanie Perkins (CEO of Canva) * Jeff Lawson (CEO of Twilio)
Books* "Man's Search For Meaning" (by Viktor Frankl) * "Thinking In Systems" (by Donella Meadows) * "A Treasury of Rumi" (by Muhammad Isa Waley and Rumi) * "Start With Why" (by Simon Sinek)
Talks* "First Principles in Building A Real-Time AI Platform" (March 2021) * "Redis as an Online Feature Store" (April 2021) * "Redis as an online feature store, Redis Labs" (May 2021)
About the showDatacast features long-form, in-depth conversations with practitioners and researchers in the data community to walk through their professional journeys and unpack the lessons learned along the way. I invite guests coming from a wide range of career paths — from scientists and analysts to founders and investors — to analyze the case for using data in the real world and extract their mental models (“the WHY and the HOW”) behind their pursuits. Hopefully, these conversations can serve as valuable tools for early-stage data professionals as they navigate their own careers in the exciting data universe.
Datacast is produced and edited by James Le. Get in touch with feedback or guest suggestions by emailing khanhle.1013@gmail.com.
Subscribe by searching for Datacast wherever you get podcasts or click one of the links below:
If you’re new, see the podcast homepage for the most recent episodes to listen to, or browse the full guest list.
Show Notes* (01:43) Leigh-Marie shared her formative experiences of her childhood — growing up in Alabama, solving math problems competitively, and going to Phillips Exeter Academy. * (04:21) Leigh-Marie discussed her undergraduate experience at MIT studying Math with Computer Science. * (06:41) Leigh-Marie went through her internship experience at Jane Street and Blend. * (10:07) Leigh-Marie recalled lessons learned from interning at Google — as an ML engineer for the Research and Machine Intelligence Team and an Associate Product Manager for the Chrome Web Platform team. * (13:39) Leigh-Marie talked about her decision to join the early founding team of Scale API (now known as Scale AI) after finishing MIT. * (17:30) Leigh-Marie explained why labeled data is the key bottleneck to the growth of the ML industry. * (20:02) Leigh-Marie discussed the engineering and product challenges of dealing with 3D sensor data. * (22:33) Leigh-Marie unpacked her experience building Scale’s Sensor Fusion Annotation product from scratch, from gathering customer interests to building the initial MVP. * (26:45) Leigh-Marie talked about learning curves during Scale’s scaling phase, as the product had more advanced features and the customer list grew. * (32:21) Leigh-Marie dived into Scale’s credo emphasizing a relentless speed of execution. * (35:00) Leigh-Marie shared valuable hiring lessons at Scale’s early days (Read Alex’s blog post about Scale’s hiring philosophy). * (38:05) Leigh-Marie went over the importance of developing uncompressed understandings of how everything works together as Scale grows. * (41:39) Leigh-Marie shared her advice for folks who want to get into angel investing. * (44:02) Leigh-Marie shared her motivation behind joining Founders Fund (Read Founders Fund’s investment manifesto). * (46:56) Leigh-Marie went over how she has been proving value upfront and forming investment theses as a new investor. * (49:10) Leigh-Marie shared advice she has been giving to companies regarding their product-market fit and go-to-market fit strategies. * (50:38) Leigh-Marie reflected on her transitions from software engineering to product management to venture capital. * (52:31) Leigh-Marie shared the lesson learned from playing poker that benefits her careers in startup and venture. * (54:19) Closing segment.
Leigh-Marie’s Contact Info* Substack * Twitter * LinkedIn * GitHub * Quora * Founders Fund
People* Peter Thiel * Ali Partovi * Trae Stephens
Books* “Angels” (by Jason Calanacis) * “Zero To One” (by Blake Masters and Peter Thiel) * “7 Powers: The Foundations of Business Strategy” (by Hamilton Helmer)
Blog Posts* “The One Data Platform To Rule Them All” (July 2021) * “Startup Opportunities in Machine Learning Infrastructure” (Sep 2021)
About the showDatacast features long-form, in-depth conversations with practitioners and researchers in the data community to walk through their professional journeys and unpack the lessons learned along the way. I invite guests coming from a wide range of career paths — from scientists and analysts to founders and investors — to analyze the case for using data in the real world and extract their mental models (“the WHY and the HOW”) behind their pursuits. Hopefully, these conversations can serve as valuable tools for early-stage data professionals as they navigate their own careers in the exciting data universe.
Datacast is produced and edited by James Le. Get in touch with feedback or guest suggestions by emailing khanhle.1013@gmail.com.
Subscribe by searching for Datacast wherever you get podcasts or click one of the links below:
If you’re new, see the podcast homepage for the most recent episodes to listen to, or browse the full guest list.
Show Notes* (01:47) Chen-Ping shared his upbringing growing up in Taiwan and going to boarding school in the US at the age of 14. * (04:42) Chen-Ping got his Bachelor’s and Master’s degrees in Computer Science from RIT back in the early-to-mid 2000, in which he did academic research in computational neuroscience. * (08:18) Chen-Ping walked through his MS thesis at RIT, designing and implementing a computational model of neurons from the visual cortex’s medial superior temporal area. * (10:18) Chen-Ping talked about the academic culture shock of pursuing his Master’s degree in Computer Science and Engineering at Penn State. * (13:47) Chen-Ping walked through his MS thesis at Penn State, proposing a statistical asymmetry-based automatic brain tumor detection from 3D MR images. * (18:35) Chen-Ping discussed the thread of his research as a Ph.D. student at Stony Brook, where he worked at the Computer Vision Lab and the Eye Cog Lab. * (23:19) Chen-Ping unpacked his Ph.D. dissertation at Stony Brook called computational models of visual features: from proto-objects to object categories. * (28:54) Chen-Ping went through his internship experience at Riverbed Technology and Shutterstock. * (30:20) Chen-Ping dissected the development of a neuro-inspired deep convolutional neural network called Map-CNN for modeling human early visual information processing during his time as a Postdoc at Harvard’s Cognitive and Neural Organization Lab. * (32:14) Chen-Ping mentioned research areas at the intersection of computer vision and cognitive vision that he is excited about. * (33:33) Chen-Ping shared the story behind the founding of Phiar with James Briscoe, an ex-classmate from RIT, and Ivy Lee, an ex-colleague from Shutterstock. * (36:33) Chen-Ping discussed technical challenges with developing an ultra-lightweight Spatial AI engine that allows any vehicle to perceive its surroundings using a camera that can run in real-time at the edge on a commodity automotive computing platform. * (39:36) Chen-Ping unpacked the key features of a complete Visual Mobility platform, including automobile integration, AR navigation, digitized environment, smart parking, 3rd-party integration, and reality-as-a-service. * (41:16) Chen-Ping shared details around Phiar’s ultra-efficient monocular depth estimation AI that runs efficiently on a mobile phone and achieves SOTA accuracies on the benchmark KITTI dataset. * (43:16) Chen-Ping revisited his experience going through the Y-Combinator incubator in the summer of 2018. * (44:27) Chen-Ping shared high-level fundraising advice for first-time founders. * (46:30) Chen-Ping talked about strategies he found useful to identify the right client partnerships for Phiar. * (48:10) Chen-Ping shared valuable hiring lessons learned at Phiar. * (51:37) Chen-Ping reflected on the difference between being a researcher and a founder. * (53:43) Closing segment.
Chen-Ping’s Contact Info* LinkedIn * Twitter * Google Scholar
Phiar’s Resources* Website * LinkedIn | Twitter | Facebook | YouTube * “Phiar Secures $12M Series A and Names Google Head of Android Automotive Platforms as CEO” (Sep 2021)
Mentioned ContentPeople* Fei-Fei Li * Yann LeCun * Yoshua Bengio
Books and Papers* “Zero To One” (by Blake Masters and Peter Thiel) * “Modeling Clutter Perception using Parametric Proto-object Partitioning” (NIPS 2013) * “Modeling visual clutter perception using proto-object segmentation” (June 2014) * “Searching for Category-Consistent Features: A Computational Approach to Understanding Visual Category Representation” (May 2016) * “Generating the features for category representation using a deep convolutional neural network” (Sep 2016) * “Map-CNN: A Convolutional Neural Network with Map-like Organizations” (Aug 2017) * “Mid-level visual features underlie the high-level categorical organization of the ventral stream” (Sep 2018)
About the showDatacast features long-form, in-depth conversations with practitioners and researchers in the data community to walk through their professional journeys and unpack the lessons learned along the way. I invite guests coming from a wide range of career paths — from scientists and analysts to founders and investors — to analyze the case for using data in the real world and extract their mental models (“the WHY and the HOW”) behind their pursuits. Hopefully, these conversations can serve as valuable tools for early-stage data professionals as they navigate their own careers in the exciting data universe.
Datacast is produced and edited by James Le. Get in touch with feedback or guest suggestions by emailing khanhle.1013@gmail.com.
Subscribe by searching for Datacast wherever you get podcasts or click one of the links below:
If you’re new, see the podcast homepage for the most recent episodes to listen to, or browse the full guest list.
Timestamps* (02:00) Aarti shared her upbringing growing up in India and going to New York for undergraduate. * (04:47) Aarti recalled her academic experience getting dual degrees in Computer Science and Computer Engineering at New York University. * (07:17) Aarti shared details about her involvement with the ACM chapter and the Women in Computing club at NYU. * (10:46) Aarti shared valuable lessons from her research internships. * (14:16) Aarti discussed her decision to pursue an MS degree in Computer Science at Stanford University. * (20:27) Aarti reflected on her learnings being the Head Teaching Assistant for CS 230, one of Stanford’s most popular Deep Learning courses. * (23:59) Aarti shared her thoughts on ML applications in both clinical and administrative healthcare settings. * (26:47) Aarti unpacked the motivation and empirical work behind CheXNet, an algorithm that can detect pneumonia from chest X-rays at a level exceeding practicing radiologists. * (29:39) Aarti went over the implications of MURA, a large dataset of musculoskeletal radiographs containing over 40,000 images from close to 15,000 studies, for ML applications in radiology. * (32:50) Aarti went over her experience working briefly as an ML engineer at Andrew Ng’s startup Landing AI and applying ML to visual inspection tasks in manufacturing. * (36:56) Aarti talked about her participation in external entrepreneurial initiatives such as Threshold Venture Fellowship and Greylock X Fellowship. * (43:41) Aarti reminisced her time in a hybrid ML engineer/product manager/VC associate role at AI Fund, which works intensively with entrepreneurs during their startups’ most critical and risky phase from 0 to 1. * (48:43) Aarti shared advice that AI fund companies tended to receive regarding product-market fit and go-to-market fit strategy. * (54:04) Aarti walked through her decision to onboard Snorkel AI, the startup behind the popular Snorkel open-source project capable of quickly generating training data with weak supervision. * (56:36) Aarti reflected on the difference between being an ML researcher and an ML engineer. * (01:00:18) Closing segment.
Aarti’s Contact Info* LinkedIn * Twitter * Google Scholar
People* Andrew Ng * John Langford * David Sontag
Books and Papers* “The Art of Doing Science & Engineering” (by Richard Hamming) * “Deep Medicine: How AI Can Make Healthcare Human Again” (by Eric Topol) * “CheXNet: Radiologist-Level Pneumonia Detection on Chest X-Rays with Deep Learning” (Dec 2017) * “MURA: Large Dataset for Abnormality Detection in Musculoskeletal Radiographs” (May 2018)
About the showDatacast features long-form, in-depth conversations with practitioners and researchers in the data community to walk through their professional journeys and unpack the lessons learned along the way. I invite guests coming from a wide range of career paths — from scientists and analysts to founders and investors — to analyze the case for using data in the real world and extract their mental models (“the WHY and the HOW”) behind their pursuits. Hopefully, these conversations can serve as valuable tools for early-stage data professionals as they navigate their own careers in the exciting data universe.
Datacast is produced and edited by James Le. Get in touch with feedback or guest suggestions by emailing khanhle.1013@gmail.com.
Subscribe by searching for Datacast wherever you get podcasts or click one of the links below:
If you’re new, see the podcast homepage for the most recent episodes to listen to, or browse the full guest list.
Timestamps* (02:14) Alberto briefly shared his upbringing and education at the Bayes Business School in London. * (04:01) Alberto shared key learnings from his first entrepreneurial stint at 19 by developing a 3D printing product for ed-tech. * (07:48) Alberto described his overall experience participating in Singularity University’s Graduate Studies Program at the NASA Ames Research Park under a Google-funded scholarship in 2015. * (12:52) Alberto helped develop the Aipoly product to aid the blind and visually impaired. * (17:38) Alberto showed his enthusiasm for federated learning applications within mobile devices. * (19:53) Alberto talked about the dichotomy between capitalism and social good in entrepreneurship. * (22:29) Alberto shared the backstory behind the founding of V7 Labs. * (26:40) Alberto discussed the comparison between biological and artificial neural networks. * (28:02) Alberto emphasized the importance of having a good co-founder. * (30:27) Alberto dissected the notable features developed within V7’s Annotation capability. * (33:37) Alberto went over things to look for in a video labeling tool, citing his blog post. * (37:21) Alberto unpacked key principles behind V7’s robust Dataset Management tool. * (40:53) Alberto walked through the powerful capabilities of V7 Neurons that power its Model Automation tool. * (43:33) Alberto shared fundraising advice for founders seeking the right investors for their startups. * (46:07) Alberto shared valuable hiring and culture-setting lessons learned at V7. * (50:12) Alberto emphasized the importance of not losing sight of the ‘ideal customer’ for young founders in the AI space. * (53:01) Alberto shared the hurdles his team has to go through while finding new customers in new industries. * (55:10) Alberto walked through labeling challenges dealing with medical imaging datasets. * (57:35) Alberto discussed outreach initiatives that helped drive V7’s organic growth. * (59:49) Alberto mentioned the importance of collaboration between companies within the MLOps ecosystem. * (01:02:01) Alberto touched on the scientific hunger of Europe regarding the adoption of AI technologies. * (01:03:49) Alberto briefly mentioned what public recognition means to him in the pursuit of democratizing AI for the world. * (01:06:07) Closing segment.
Alberto’s Contact Info* Website * LinkedIn * Twitter * Medium
V7’s Resources* Website * Software 2.0 Blog * Academy Tutorials * Documentation * LinkedIn | Twitter
Mentioned ContentArticles* “7 Things We Looked for in a Video Labeling Tool” (Aug 2020) * “The Biggest Mistake I’ve Ever Made: Losing Sight of the Ideal Customer” (March 2021)
Talks* “An AI Narrator for the Blind” (TEDx Geneva 2016) * “If The Blind Could See” (TEDx Melbourne 2018)
People* Geoff Hinton (for rethinking the ML field fundamentally) * Chelsea Finn (for her work on meta-learning) * Jeff Clune (for making agents that work at scale in the real world)
Book* “Start With Why” (by Simon Sinek)
NotesV7 is hiring across all departments. Take a look at their careers page for the openings!
About the showDatacast features long-form, in-depth conversations with practitioners and researchers in the data community to walk through their professional journeys and unpack the lessons learned along the way. I invite guests coming from a wide range of career paths — from scientists and analysts to founders and investors — to analyze the case for using data in the real world and extract their mental models (“the WHY and the HOW”) behind their pursuits. Hopefully, these conversations can serve as valuable tools for early-stage data professionals as they navigate their own careers in the exciting data universe.
Datacast is produced and edited by James Le. Get in touch with feedback or guest suggestions by emailing khanhle.1013@gmail.com.
Subscribe by searching for Datacast wherever you get podcasts or click one of the links below:
If you’re new, see the podcast homepage for the most recent episodes to listen to, or browse the full guest list.
Timestamps* (01:55) Jessica shared the formative experiences of her upbringing — being born in a triplet with two other sisters and growing up in an immigrant family from Russia. * (05:45) Jessica shared her experience being part of UC Berkeley’s first cohort of Data Science majors. * (09:56) Jessica talked about her campus involvements with student-run organizations such as the Mobile Developers of Berkeley and the Data Science Society at Berkeley. * (13:02) Jessica walked through her participation in initiatives like researching the CITRIS and the Banatao Institute, sitting on the leadership board of the TAMID Group, and being an Accel Scholar. * (15:31) Jessica shared valuable lessons from her summer internships. * (19:08) Jessica discussed her decision to join Ironclad, a Series D digital contracting startup building software to take legal teams to the next level. * (22:39) Jessica provided a brief explanation of digital contracting for the uninitiated. * (24:59) Jessica talked about challenges that in-house legal teams typically face and how Ironclad helps address them. * (27:04) Jessica gave a tour of Ironclad’s Contract Lifecycle Management software offerings. * (30:00) Jessica walked through her journey of building the analytics function from scratch and providing data insights to inform business decisions cross-functionally. * (33:40) Jessica shared tidbits about her time management and goal-setting systems. * (34:45) Jessica walked through the end-to-end data analysis process for Ironclad’s first legal analytics benchmark report analyzing economic trends caused by COVID-19. * (38:38) Jessica discussed the learning curves as she took on bigger analytical responsibilities at Ironclad. * (43:05) Jessica unpacked her 3-level framework for building a data analytics culture from the ground up. * (48:07) Jessica shared concrete advice on positively influencing a company’s culture to be data-driven. * (50:27) Jessica unfolded the drive behind creating the Data Angels Community, a Slack group connecting women interested in data to resources for support, education, and opportunities. * (52:25) Jessica revealed her community playbook to engage the members of Data Angels. * (57:01) Jessica shared a bit of her guilty pleasure in using data for beauty and fashion. * (01:00:44) Closing segment.
Jessica’s Contact Info* LinkedIn * Twitter * Data Angels
Mentioned ContentResources* "How to use contract data during COVID-19" (Ironclad Report) * "Building data analytics culture from the ground up" (Women In Product 2020 Talk) * "Building a data-centered culture at Ironclad" (Ironclad Article)
People* Emily Robinson and Jacqueline Nolis (Co-Authors and Co-Hosts of “How To Build A Career in Data Science” the book and the podcast) * Cassie Kozyrkov (Chief Decision Scientist at Google) * Shreya Shankar (Ph.D. Student at UC Berkeley and Entrepreneur-In-Residence at Amplify Partners) (Check out my interview with Shreya as well!)
Book* “Everybody Lies: Big Data, New Data, and What The Internet Can Tell Us About Who We Really Are” (by Seth Stephens-Davidowitz)
NotesMy conversation with Jessica was recorded back in May 2021. Jessica is now a Senior Data Analyst and Ironclad's Data Analytics team has grown to 4 so she is no longer a 1-woman show! Also, the Data Angels Slack community has over 500 members in it now!
About the showDatacast features long-form, in-depth conversations with practitioners and researchers in the data community to walk through their professional journeys and unpack the lessons learned along the way. I invite guests coming from a wide range of career paths — from scientists and analysts to founders and investors — to analyze the case for using data in the real world and extract their mental models (“the WHY and the HOW”) behind their pursuits. Hopefully, these conversations can serve as valuable tools for early-stage data professionals as they navigate their own careers in the exciting data universe.
Datacast is produced and edited by James Le. Get in touch with feedback or guest suggestions by emailing khanhle.1013@gmail.com.
Subscribe by searching for Datacast wherever you get podcasts or click one of the links below:
If you’re new, see the podcast homepage for the most recent episodes to listen to, or browse the full guest list.
Timestamps* (01:40) Julia shared the differences growing up in New York and moving to San Francisco. * (03:05) Julia discussed her overall undergraduate experience at Stanford — getting dual degrees in Computer Science and Management Science & Engineering_._ * (05:40) Julia went over her time as an Investment Banker at Qatalyst Partners — notably working on Microsoft’s acquisition of LinkedIn. * (09:11) Julia talked about her career transition to venture capital — working as an associate investor at New Enterprise Associates. * (10:46) Julia emphasized the importance of getting up-to-speed and forming an investment thesis as a new investor. * (15:05) Julia discussed her Series A investment in Metabase, an open-source business intelligence software project. * (18:36) Julia unpacked her investment(s) in Sentry, an application monitoring platform that helps developers monitor apps in real-time to catch bugs early. * (20:14) Julia explained her investment in the Series B round for Anyscale, an end-to-end computing platform that makes building and managing a scaled application across clouds as easy as developing an app on a single computer. * (23:03) Julia contextualized her investments in the seed round for Datafold, a data observability platform that equips analytics engineers with the tools to address data quality issues. * (24:24) Julia shared typical hiring and go-to-market decisions that companies need to make (depending upon their growth stages and product strategies). * (27:05) Julia mentioned her Metabase application to help investors pick winning open-source startups. * (29:05) Julia rationalized her switch to becoming a product manager at dbt Labs. * (30:34) Julia peeked into the roadmap of dbt Cloud, a hosted service that helps data analysts and engineers productionize dbt deployments. * (33:34) Julia went over an under-invested area and the role of interoperability within the broader data tooling ecosystem. * (37:56) Julia reflected on the difference between being a venture investor and a product manager. * (41:05) Closing segment.
Julia’s Contact Info* LinkedIn * Twitter
dbt’s Resources* Slack Community * Coalesce 2021 Replays * dbt Learn * GitHub * Events and Meetups
Mentioned ContentPeople* Tristan Handy (Founder and CEO of dbt Labs) * Ali Ghodsi (Co-Creator of Apache Spark, Co-Founder and CEO of Databricks) * Dan Levine (General Partner at Accel Partners)
Book* “Working Backwards: Insights, Stories, and Secrets from Inside Amazon” (by Bill Carr and Colin Bryar)
NotesMy conversation with Julia was recorded back in May 2021. Since the podcast was recorded, a lot has happened at dbt Labs! I’d recommend:
About the showDatacast features long-form, in-depth conversations with practitioners and researchers in the data community to walk through their professional journeys and unpack the lessons learned along the way. I invite guests coming from a wide range of career paths — from scientists and analysts to founders and investors — to analyze the case for using data in the real world and extract their mental models (“the WHY and the HOW”) behind their pursuits. Hopefully, these conversations can serve as valuable tools for early-stage data professionals as they navigate their own careers in the exciting data universe.
Datacast is produced and edited by James Le. Get in touch with feedback or guest suggestions by emailing khanhle.1013@gmail.com.
Subscribe by searching for Datacast wherever you get podcasts or click one of the links below:
If you’re new, see the podcast homepage for the most recent episodes to listen to, or browse the full guest list.
Timestamps* (1:33) Einat described her experience getting Bachelor’s, Master’s, and Ph.D. degrees in Mathematics from Tel Aviv University in the 90s and early 2000s. * (4:01) Einat went over her Ph.D. thesis on approximation algorithms for clustering problems. * (6:17) Einat discussed working as an algorithm developer for Compugen while being a Ph.D. student. * (8:43) Einat went over projects she contributed to as a senior algorithm developer at Flash Networks back in 2005. * (11:50) Einat mentioned achievements and lessons learned from her time as the VP of R&D at Correlix. * (17:51) Einat recalled lessons from hiring engineering talent at Correlix. * (19:24) Einat unpacked the engineering challenges of building SimilarWeb, a platform that gives a true 360-degree view of all digital activity across customers, prospects, partners, and competition. * (24:29) Einat discussed the responsibilities of her role as the CTO of SimilarWeb. * (27:40) Einat shared the founding story of Treeverse, whose mission is to simplify the lives of data engineers, data scientists, and data analysts who are transforming the world with data. * (29:52) Einat explained the pain points of working with the data lake architecture and the vision that lakeFS is built upon. * (34:31) Einat emphasized the importance of asking good questions to extract insights about customers’ pain points. * (37:57) Einat explained why data versioning-as-an-Infrastructure matters. * (42:28) Einat shared the challenges of incorporating data mesh to develop a data-intensive application. * (46:33) Einat provided her take on how to ensure data quality in a data lake environment. * (51:02) Einat discussed roadmap prioritization for an open-source project. * (52:08) Einat went over the opportunities with the metadata store, data quality, compute, and data discovery components within the data engineering ecosystem. * (55:03) Einat captured the three trends on how the data engineering landscape might look in the near future. * (01:00:59) Einat emphasized the role of open-source development in the data tooling ecosystem. * (01:04:14) Einat fleshed out the recommended pricing strategy for open-source developers. * (01:06:09) Einat revisited how lakeFS got started thanks to the Go community and evolved. * (01:08:01) Einat shared valuable hiring lessons learned at Treeverse. * (01:10:05) Einat described the state of the data community in Israel. * (01:11:49) Closing segment.
Einat’s Contact Info* LinkedIn * Twitter * Email
Mentioned ContentlakeFS* Website * GitHub * @lakeFS * Treeverse * Slack
Blog Posts* Why We Built lakeFS: Atomic and Versioned Data Lake Operations (Aug 2020) * Data Versioning — Does It Mean What You Think It Means? (Aug 2020) * How To Manage Your Data The Way You Manage Your Code (Oct 2020) * Data Mesh Applied: How to Move Beyond The Data Lake with lakeFS (Dec 2020) * Why Data Versioning as an Infrastructure Matters? (Dec 2020) * Ensuring Data Quality in a Data Lake Environment (Jan 2021) * The State of Data Engineering in 2021 (May 2021)
People* Ali Ghodsi (Co-Creator of Apache Spark, Co-Founder and CEO of Databricks) * Shay Banon (Co-Founder and CEO of Elastic) * Gwen Shapira (Engineering Leader at Confluent)
Book* “Designing For Data-Intensive Applications” (by Martin Kleppmann)
NotesMy conversation with Einat was recorded back in April 2021. Since the podcast was recorded, a lot has happened at Treeverse! I’d recommend:
About the showDatacast features long-form, in-depth conversations with practitioners and researchers in the data community to walk through their professional journeys and unpack the lessons learned along the way. I invite guests coming from a wide range of career paths — from scientists and analysts to founders and investors — to analyze the case for using data in the real world and extract their mental models (“the WHY and the HOW”) behind their pursuits. Hopefully, these conversations can serve as valuable tools for early-stage data professionals as they navigate their own careers in the exciting data universe.
Datacast is produced and edited by James Le. Get in touch with feedback or guest suggestions by emailing khanhle.1013@gmail.com.
Subscribe by searching for Datacast wherever you get podcasts or click one of the links below:
If you’re new, see the podcast homepage for the most recent episodes to listen to, or browse the full guest list.
Timestamps* (02:13) Prukalpa discussed her upbringing in India and studying Engineering at the Nanyang Tech University in Singapore. * (03:52) Prukalpa shared the key learnings from her summer internship as an Investment Banking Analyst at Goldman Sachs. * (05:37) Prukalpa went over the seed idea for SocialCops (Read her Quora answer on the fundraising story). * (11:27) Prukalpa gave a brief overview of the business model at SocialCops. * (12:45) Prukalpa unpacked her talk called “How Big Data Can Influence Decisions That Actually Matter” at TEDxGateway 2017 related to the data-for-good initiatives that SocialCops facilitated. * (15:23) Prukalpa shared her thoughts on the future of the Data-for-Good movement. * (17:49) Prukalpa discussed the challenges that SocialCops’ data teams faced and the founding story behind Atlan. * (21:38) Prukalpa went over the trust-based culture that enabled SocialCops’ 8-member data team to build out India’s National Data Platform. * (27:00) Prukalpa dissected the six principles of Atlan’s DataOps Culture Code. * (31:37) Prukalpa unpacked the notion of Data Catalog 3.0, which is a key value prop of the Atlan platform. * (36:01) Prukalpa provided the 3-level framework to ensure data quality (detect -> prevent -> cure) and strong practices to maintain high-quality data. * (40:19) Prukalpa revealed the challenges that organizations face when starting their data governance initiatives. * (45:35) Prukalpa talked about the under-invested building blocks of modern data platforms. * (49:24) Prukalpa raised the importance of integration for Atlan to work well with the rest of the modern data stack. * (50:39) Prukalpa recapped the trends that Chief Data Officers needed to watch out for in 2021. * (54:01) Prukalpa gave fundraising advice for founders currently seeking the right investors for their startups. * (58:42) Prukalpa discussed Atlan’s outreach initiatives to engage with the broader data community actively. * (01:01:03) Prukalpa went over Atlan’s hiring philosophy based on the concept of People-as-a-Moat to attract, engage, and grow top talent — as inspired by the McKinsey advantage. * (01:05:22) Prukalpa shared Atlan’s Go-To-Market initiatives in the US this year and emphasized the importance of building an execution machine. * (01:08:53) Prukalpa described the state of the data community in India. * (01:10:25) Prukalpa shared entrepreneurship books that have deeply impacted her startup journey. * (01:12:16) Prukalpa briefly mentioned what public recognition means to her in the pursuit of democratizing data for the world. * (01:14:23) Closing segment.
Prukalpa’s Contact Info* LinkedIn * Twitter
Mentioned ContentAtlan (Twitter | LinkedIn | Facebook | Instagram | YouTube | Documentation)* “Empowering Organizations to Become Masters of Their Data” (Video) * Atlan Labs (Open-Source Projects) * Humans of Data Interviews (Interviews) * The DataOps Culture Code (Document) * Building a Business Case for DataOps (EBook) * The Data Catalog Primer (EBook) * The Ultimate Guide to Evaluating a Data Catalog (EBook)
Blog Posts* Voices In The Head of a Middle-Class Aspiring Startup Founder (July 2013) * SocialCops: What We Actually Do (Oct 2016) * People-as-a-Moat: What Startups Can Learn From McKinsey About Building A Strong Company (Aug 2018) * Going from Great People to Greater Teams: How We Think About Growth at Atlan (August 2018) * Onwards and Upwards: Chapter 2 for SocialCops (July 2019) * What is data quality? (Jan 2021) * Top 5 Data Trends For CDOs to Watch Out For In 2021 (Feb 2021) * Data Catalog 3.0: Modern Metadata for the Modern Data Stack (Feb 2021) * We Failed To Setup a Data Catalog 3x. Here’s Why (March 2021) * The Building Blocks of a Modern Data Platform (March 2021) * Data Governance Has a Serious Branding Problem (Nov 2021)
Books* “The Hard Things About The Hard Things” (by Ben Horowitz) * “Hatching Twitter” (by Nick Bilton) * “The McKinsey Way” (by Ethan Rasiel) * “How Google Works” (by Eric Schmidt and Jonathan Rosenberg) * “The Mom Test” (by Rob Fitzpatrick) * “Disciplined Entrepreneurship” (by Bill Aulet) * “Big Data” (by Mayer-Schnonberger and Cukier)
Talks* Game of Life (TEDxIIMShilong — March 2014) * How Big Data Can Influence Decisions That Actually Matter (TEDx Gateway — April 2017) * Better Villages Through Big Data (TED Talks India — December 2017) * The power of data science to measure unmeasured parameters in Emerging Markets (PyData Dehli — Oct 2019) * The Girl Who Thinks In Numbers: Data Warrior Prukalpa Sankar (Feb 2020)
NotesMy conversation with Prukalpa was recorded back in April 2021. Since the podcast was recorded, a lot has happened at Atlan!
Prukalpa has written more content. I’d recommend checking out:
About the showDatacast features long-form, in-depth conversations with practitioners and researchers in the data community to walk through their professional journeys and unpack the lessons learned along the way. I invite guests coming from a wide range of career paths — from scientists and analysts to founders and investors — to analyze the case for using data in the real world and extract their mental models (“the WHY and the HOW”) behind their pursuits. Hopefully, these conversations can serve as valuable tools for early-stage data professionals as they navigate their own careers in the exciting data universe.
Datacast is produced and edited by James Le. Get in touch with feedback or guest suggestions by emailing khanhle.1013@gmail.com.
Subscribe by searching for Datacast wherever you get podcasts or click one of the links below:
If you’re new, see the podcast homepage for the most recent episodes to listen to, or browse the full guest list.
Timestamps* (01:58) Michel went over his education studying at EPITA — School of Engineering and Computer Science in France. * (03:50) Michel mentioned his first US internship at Siemens Corporate Research as an R&D engineer. * (05:48) Michel discussed the unique challenges of building systems to handle financial data through his engineering experience at FactSet Research Systems and Murex. * (07:48) Michel talked about his move to San Francisco to work as a Senior Software Engineer at Rapleaf, focusing on scaling up data integration and data management pipelines. * (10:40) Michel unpacked his work building the modern data stack at LiveRamp. * (16:18) Michel shared valuable leadership and hiring lessons absorbed during his time as Liveramp’s Head of Data Integrations — fostering a strong culture of innovation and expanding the engineering organization significantly. * (19:03) Michel dived into how to interview engineering talent for independence, autonomy, and communication. * (20:56) Michel dissected the engineering architecture of the rideOS ride-hail platform (where he was a founding member and director of engineering). * (26:03) Michel told the founding story of Airbyte, whose mission is to make data integration pipelines a commodity (+ the pivot that happened during Airbyte’s time at Y Combinator). * (32:10) Michel explained the paint points with existing data integration practices and the vision that Airbyte is moving towards. * (35:07) Michel unpacked the analogy of Airbyte’s approach to building a connector manufacturing plant, which is to think in onion layers. * (39:13) Michel shared the challenges that are still hard for an open-source solution to address (Read his list of challenges that open-source and commercial software face to solve the data integration problem). * (40:28) Michel discussed how to prioritize product roadmap while developing an open-source project. * (41:59) Michel discussed pricing strategies for open-source projects (Airbyte’s business models entail both self-hosted and hosted solutions). * (44:17) Michel revealed the hurdles that Airbyte has overcome to find the early committers for their open-source project. * (47:53) Michel shared valuable hiring lessons learned at Airbyte. * (50:16) Michel shared fundraising advice for founders seeking the right investors for their startups. * (52:41) Closing segment.
Michel’s Contact Info* Twitter * LinkedIn * GitHub
Mentioned ContentAirbyte (Docs | Community | GitHub | Twitter | LinkedIn)* Handbook * Recipes * Community Call * Office Hours * Connector Contest
Blog Posts* “The Hard Things About Pivoting” (July 2020) * “How Can We Commoditize Data Integration Pipelines” (Sep 2020) * “How to Build Thousands of Connectors” (Oct 2020) * “The Deck We Used to Raise Our Seed with Accel in 13 Days” (March 2021)
People* Jeremy Litz (Former CTO and Co-Founder of LiveRamp) * Tristan Handy (CEO of dbtLabs and Editor of the Analytics Engineering Newsletter)
Book* “High-Growth Handbook” (by Elad Gil)
NotesMy conversation with Michel was recorded back in April 2021. Since the podcast was recorded, a lot has happened at Airbyte! I'd recommend:
About the showDatacast features long-form, in-depth conversations with practitioners and researchers in the data community to walk through their professional journeys and unpack the lessons learned along the way. I invite guests coming from a wide range of career paths — from scientists and analysts to founders and investors — to analyze the case for using data in the real world and extract their mental models (“the WHY and the HOW”) behind their pursuits. Hopefully, these conversations can serve as valuable tools for early-stage data professionals as they navigate their own careers in the exciting data universe.
Datacast is produced and edited by James Le. Get in touch with feedback or guest suggestions by emailing khanhle.1013@gmail.com.
Subscribe by searching for Datacast wherever you get podcasts or click one of the links below:
If you’re new, see the podcast homepage for the most recent episodes to listen to, or browse the full guest list.
Timestamps* (02:03) Cindi briefly shared her early interest in writing and her decision to major in English at the University of Maryland in the mid-80s. * (05:22) Cindi talked about her move to Zurich for a Business Systems Specialist role at Dow Chemical. * (07:35) Cindi recalled the state of Business Intelligence tools and their adoption level in the enterprises during the mid-90s. * (10:53) Cindi went over her decision to pursue an MBA from the Jones Business School at Rice University, in which her MBA Thesis was about how the Internet would reshape the first-generation BI tools. * (16:30) Cindi discussed how she balanced academic study and parenthood during her MBA. * (20:57) Cindi talked about her proudest accomplishments as a manager at Deloitte — building BI and analytics practice in Houston. * (22:48) Cindi went over her time running her independent analyst firm BI Scorecard, which advised clients on BI and analytics tool selections via rigorous evaluation criteria. * (26:14) Cindi brought up her time teaching classes at The Data Warehousing Institute, which educates business leaders on the proper deployment of data warehousing strategies and technologies. * (27:49) Cindi mentioned her move to become the Vice President in data and analytics at Gartner. * (30:33) Cindi walked through the end-to-end process of creating Gartner’s Magic Quadrant for Analytics and BI Platforms and Critical Capabilities. * (33:47) Cindi explained the culture of “Selfless Excellence” at ThoughtSpot — where she currently serves as a Chief Strategy Officer. * (36:11) Cindi explained the concept of “What’s In It For Me” (WIIFM) that helps bring a data-driven culture to organizations. * (39:34) Cindi gave a tour of ThoughtSpot’s core capabilities, ranging from SearchIQ and SpotIQ to ThoughtSpot One and ThoughtSpot Embrace. * (43:04) Cindi broke down her responsibilities as a Chief Data Strategy Officer working with internal and external stakeholders. * (44:40) Cindi emphasized the role of partnerships between startup vendors to empower the future of BI analytics (Read her article A New Era in Analytics and BI”). * (49:58) Cindi recapped takeaways from ThoughtSpot’s ebook that presents 6 Top Trends and Predictions for Data, Analytics, and AI in 2021. * (53:07) Cindi gave advice to companies that want to bring consumerization to enterprise analytics. * (56:35) Cindi gave her two cents on the movement of Data For Good in the progress of analytics and AI in the near future. * (58:58) Cindi recapped insights that she has observed from hosting The Data Chief Podcast (which features interviews with some of the most successful data leaders). * (01:03:29) Cindi gave advice to female data practitioners in the early phase of their careers (Read her article on the challenges that keep women out of tech). * (01:05:30) Closing segment.
Cindi’s Contact* LinkedIn * Twitter * ThoughtSpot * The Data Chief Podcast
Mentioned ContentBlog Posts* “Why I Joined ThoughtSpot” (April 2019) * “A New Era in Analytics and BI” (August 2019) * “Perfect Storm or Transformative Triumvirate: Data for Good, Data for Evil, and AI Ethics” (Nov 2019) * “We Can Put a Man on the Moon, But We Can’t Keep Women in Tech” (Sep 2019) * 6 Top Trends and Predictions for Data, Analytics, and AI in 2021 (2021 E-Book)
Published Books* “Successful Business Intelligence” (Nov 2013) * “SAP BusinessObjects BI 4.0” (Nov 2012)
Data for Good Resources* Datakind * Mastercard Center for Inclusive Growth * Carnegie Mellon’s Data Science for Social Good * Viz for Social Good
Women in Data Resources* Women in Data * Women in Big Data * Women in Analytics
People* Joy Buolamwini (Computer Scientist and Digital Activist at MIT Media Lab, Founder of the Algorithmic Justice League) * Cathy O’Neil (Author of “Weapons of Math Destruction”) * Kate Strachnyi (Founder of DATAcated) * Ralph Kimball (Original Architect of Data Warehousing) * Ajeet Singh and Amit Prakash (Co-Founders of ThoughtSpot)
Recommended Books* “Moneyball” (by Michael Lewis) * “Freakonomics” (by Steven Levitt and Stephen Dubner)
NotesMy conversation with Cindi was recorded back in April 2021. Since the podcast was recorded, a lot has happened at ThoughtSpot:
About the showDatacast features long-form, in-depth conversations with practitioners and researchers in the data community to walk through their professional journeys and unpack the lessons learned along the way. I invite guests coming from a wide range of career paths — from scientists and analysts to founders and investors — to analyze the case for using data in the real world and extract their mental models (“the WHY and the HOW”) behind their pursuits. Hopefully, these conversations can serve as valuable tools for early-stage data professionals as they navigate their own careers in the exciting data universe.
Datacast is produced and edited by James Le. Get in touch with feedback or guest suggestions by emailing khanhle.1013@gmail.com.
Subscribe by searching for Datacast wherever you get podcasts or click one of the links below:
If you’re new, see the podcast homepage for the most recent episodes to listen to, or browse the full guest list.
Timestamps* (01:30) Taivo shared briefly about his experience going through the Estonian K-12 system, as argued in his blog post written in Estonian. * (05:34) Taivo described his undergraduate experience studying Computer Science at the University of Tartu and exposing to Machine Learning. * (08:15) Taivo discussed his time interning at Skype and TransferWise. * (10:01) Taivo went over his Master's Degree in Computer Science at ETH Zurich, where he worked on a thesis called "Uncertainty-based active imitation learning" at the Learning and Adaptive Systems Group. * (17:17) Taivo talked about his time working at Starship Technologies as a Perception Engineer. * (21:26) Taivo unpacked the Data Specification Manifesto that entails 3 principles for iteratively solving complex problems. * (27:21) Taivo unpacked "The Two Loops Of Building Algorithmic Products" from his experience at Veriff - an Estonian startup that develops an identity verification platform. * (32:11) Taivo discussed how his team at Veriff developed automation-heavy products. * (36:45) Taivo shared lessons learned as a Product Manager at Veriff: leading the go-to-market strategy, establishing communication between the product and sales division, and building a unique DataOps team that creates good datasets. * (44:31) Taivo described the key characteristics and properties of a tool that can address the whole data annotation workflow (Read his article "Data Loops Are The Bottleneck In Applied AI"). * (49:33) Taivo predicted the evolution of the DataOps discipline for AI teams in the upcoming years (Read his article "Your AI Team Needs DataOps"). * (54:01) Taivo untangled the relationship between sampling and labeling, and their importance in the AI development process (Read his article "Datasets Carve The Terrain of AI"). * (56:36) Taivo talked about the tools that he's most excited about during the transition to Software 2.0. * (01:00:04) Taivo shared his journey thus far as the founder of a stealth startup. * (01:06:21) Taivo revealed insider insights about the #EstonianMafia startup ecosystem. * (01:09:36) Taivo shared the productivity tips that have been most useful to his personal/professional growth. * (01:14:10) Closing segment.
Taivo's Contact* Website * Twitter * LinkedIn * Medium * Google Scholar
Mentioned ContentBlog Posts* Data Specification Manifesto * "Building Automation-Heavy Products" (April 2019) * "Data Loops Are The Bottleneck In Applied AI" (June 2019) * "Your AI Team Needs DataOps" (July 2020) * "Datasets Carve The Terrain of AI" (Nov 2020)
Talks* "The Two Loops Of Building Algorithmic Products" (April 2019) * "How To Build Your AI Startup" (June 2020) * "Datasets: The Source Code of Software 2.0" (Nov 2020)
People* Andrej Karpathy (The Senior Director of AI at Tesla, who coined the term Software 2.0) * Mike Bostock (The Creator of D3.js)
Book* "Surely You're Joking, Mr. Feynman" (by Richard Feynman)
About the showDatacast features long-form, in-depth conversations with practitioners and researchers in the data community to walk through their professional journeys and unpack the lessons learned along the way. I invite guests coming from a wide range of career paths — from scientists and analysts to founders and investors — to analyze the case for using data in the real world and extract their mental models (“the WHY and the HOW”) behind their pursuits. Hopefully, these conversations can serve as valuable tools for early-stage data professionals as they navigate their own careers in the exciting data universe.
Datacast is produced and edited by James Le. Get in touch with feedback or guest suggestions by emailing khanhle.1013@gmail.com.
Subscribe by searching for Datacast wherever you get podcasts or click one of the links below:
If you’re new, see the podcast homepage for the most recent episodes to listen to, or browse the full guest list.
Timestamps* (01:42) Gleb shared briefly about his upbringing and studying Economics in university in Russia. * (04:15) Gleb discussed his move to the US to pursue a Master of Information Systems Management at Carnegie Mellon University. * (07:07) Gleb went over his summer internship as a Business Analyst at Autodesk. * (08:40) Gleb shared the details of his project architecting data model/ETL pipelines as a PM at Autodesk. * (11:34) Gleb unpacked the evolution of his career at Lyft — from an individual data analyst to a PM on data tooling and a high-impact project that he worked on. * (16:54) Gleb shared valuable lessons from the experience of leading multiple cross-functional teams of engineers and growing the data organization significantly. * (19:48) Gleb mentioned his time as a Product Manager at Phantom Auto, leading the development of a teleoperation product for autonomous vehicles over cellular networks. * (25:28) Gleb emphasized the critical factors to consider when choosing a working environment: trusted managers/colleagues, maturity of tools/processes, and the function of data teams within the organization. * (29:10) Gleb shared the story behind the founding of Datafold, whose mission is to help companies effectively leverage their data assets while making Data Engineering & Analytics a creative and enjoyable experience. * (33:04) Gleb dissected the pain points with regression testing and the benefits of using Data Diff (Datafold’s first product) for data engineers. * (36:54) Gleb unpacked the data monitoring feature within Datafold’s data observability platform. * (39:45) Gleb discussed how to choose data warehousing solutions for your use cases (and made the distinction between data warehouse and data lake). * (47:03) Gleb gave insights on the need for BI and data observability/quality management tools within the modern analytics stack. * (50:40) Gleb emphasized the importance of tooling integration for Datafold’s roadmap. * (52:07) Gleb has been hosting Data Quality meetups to discuss the under-explored area of data quality. * (54:02) Gleb shared his learnings from going through the YC incubator in summer 2020. * (55:45) Gleb discussed the hurdles he had to jump through to find early customers of Datafold. * (57:47) Gleb emphasized valuable lessons he has learned to attract the right people who are excited about Datafold’s mission. * (59:17) Gleb shared his advice for founders who are in the process of finding the right investors for their companies. * (01:02:11) Closing segment.
Gleb’s Contact Info* LinkedIn * Datafold (Twitter and LinkedIn) * Data Quality Meetups
Mentioned ContentCourse* Harvard’s CS50: Introduction to Computer Science
Blog Posts* Modern Analytics Stack (June 2020) * Choosing Data Warehouse for Analytics (June 2020) * 3 Ways To Be Wrong About Open-Source Data Warehousing Software (June 2020) * Buy Not Build (Aug 2020) * Datafold Raises a $2.1M Seed Round Led by NEA (Nov 2020) * Datafold + dbt: The Perfect Stack for Reliable Data Pipelines (Feb 2021)
People* Maxime Beauchemin (Founder and CEO at Preset, creator of Apache Superset and Apache Airflow) * Tobias Macey (Host of the Data Engineering Podcast)
Books* “How To Measure Anything” (by Douglas Hubbard) * “Lean Analytics” (by Benjamin Yoskovitz and Alistair Croll)
NotesMy conversation with Gleb was recorded back in March 2021. Since the podcast was recorded, a lot has happened at Datafold! I’d recommend:
About the showDatacast features long-form, in-depth conversations with practitioners and researchers in the data community to walk through their professional journeys and unpack the lessons learned along the way. I invite guests coming from a wide range of career paths — from scientists and analysts to founders and investors — to analyze the case for using data in the real world and extract their mental models (“the WHY and the HOW”) behind their pursuits. Hopefully, these conversations can serve as valuable tools for early-stage data professionals as they navigate their own careers in the exciting data universe.
Datacast is produced and edited by James Le. Get in touch with feedback or guest suggestions by emailing khanhle.1013@gmail.com.
Subscribe by searching for Datacast wherever you get podcasts or click one of the links below:
If you’re new, see the podcast homepage for the most recent episodes to listen to, or browse the full guest list.
Timestamps* (01:59) Saishruthi talked about her upbringing, growing up in a rural town in India with no Internet connection and no computers. * (05:50) Saishruthi discussed her undergraduate studying Electrical Engineering at Sri Sairam Engineering College in the early 2010s. * (11:56) Saishruthi mentioned the projects and learnings during her two years working at Tata Consultancy Services as an instrumentation engineer. * (15:57) Saishruthi went over her MS degree in Electrical Engineering at San Jose State University and her journey into data science. * (22:20) Saishruthi shared the initial hurdles she faced transitioning back to school and assimilating to the US culture. * (26:10) Saishruthi touched on her work with San Jose City on disaster management. * (28:20) Saishruthi went over her job search process, eventually landing a data science position at IBM. * (32:16) Saishruthi unpacked lessons learned from public speaking. * (35:20) Saishruthi summarized IBM’s data science and machine learning initiatives. * (37:02) Saishruthi brought up various projects happening at IBM’s Center for Open Source Data and AI Technologies, whose mission is to make open-source AI models dramatically easier to create, deploy, and manage in the enterprise. * (39:40) Saishruthi unpacked the qualities needed to contribute to open-source projects and their role in shaping the development of ML technologies. * (44:50) Saishruthi dissected examples of bias in ML, identified solutions to combat unwanted bias, and presented tools for that (as delivered in her talk titled “Digital Discrimination: Cognitive Bias in Machine Learning”). * (49:12) Saishruthi shared her thoughts on the evolution of research and applications within the Trusted AI landscape. * (54:07) Saishruthi discussed the core value propositions of IBM’s Elyra, a set of AI-centric extensions to JupyterLab that aims to help data practitioners deal with the complexities of the model development lifecycle. * (56:11) Saishruthi briefly shared the challenges with developing Coursera courses on data visualization with Python and with R. * (01:00:47) Saishruthi went over her passion for movements such as Women In Tech and Girls Who Code. * (01:03:27) Saishruthi shared details about her initiative to bring education to rural children. * (01:06:36) Closing segment.
Saishruthi’s Contact Info* Twitter * LinkedIn * Medium * GitHub * Coursera
Mentioned ContentTalks* “Digital Discrimination: Cognitive Bias in Machine Learning” (All Things Open 2020)
Projects* AI Fairness 360 * AI Explainability 360 * Adversarial Robustness Toolkit * Model Asset Exchange * Data Asset Exchange * Elyra
Courses* Data Visualization with Python * Data Visualization with R
About the showDatacast features long-form, in-depth conversations with practitioners and researchers in the data community to walk through their professional journeys and unpack the lessons learned along the way. I invite guests coming from a wide range of career paths — from scientists and analysts to founders and investors — to analyze the case for using data in the real world and extract their mental models (“the WHY and the HOW”) behind their pursuits. Hopefully, these conversations can serve as valuable tools for early-stage data professionals as they navigate their own careers in the exciting data universe.
Datacast is produced and edited by James Le. Get in touch with feedback or guest suggestions by emailing khanhle.1013@gmail.com.
Subscribe by searching for Datacast wherever you get podcasts or click one of the links below:
If you’re new, see the podcast homepage for the most recent episodes to listen to, or browse the full guest list.
Timestamps* (01:44) Mohamed described his interest growing up in Egypt and studying Biomedical Engineering at Cairo University in the early 2000s. * (04:22) Mohamed commented on his experience moving to the US to pursue an MBA degree and working in various software engineering roles. * (07:35) Mohamed shared his experience authoring two books: (1) 3D Business Analyst: The Ultimate Hands-On Guide to Mastering Business Analysis and (2) Business Analysis for Beginners: Jump-Start Your BA Career in 4 Weeks. * (13:19) Mohamed discussed his move to the Bay Area for a Senior Engineering Manager role at Twilio, managing and shipping a series of communication API products using Machine and Deep Learning. * (17:39) Mohamed dissected engineering challenges building ML systems at Amazon, alongside key leadership lessons he acquired from managing Amazon’s Kindle mobile and ML engineering teams. * (20:50) Mohamed shared his insider perspective on Amazon’s practices of customer obsession, working backward, and disagree-to-commit. * (24:52) Mohamed mentioned the benefits of teaching a computer vision course for engineers at Amazon’s internal Machine Learning university. * (28:33) Mohamed went over the engineering (hardware + software) and ML challenges associated with building a proprietary threat detection platform at Synapse Tech Corporation (where he was the Head of Engineering). * (32:03) Mohamed shared concrete technical challenges with building an ML system that performs inference on edge devices. * (37:03) Mohamed revealed specific data labeling challenges while building the ML system at Synapse. * (39:57) Mohamed went over his one year as the VP of Engineering for the AI Platform at Rakuten, when he incubated the idea for Kolena. * (42:52) Mohamed explained the current state of ML testing infrastructure and unpacked his current project Kolena, a rigorous ML QA platform that lets users take control of their ML testing. * (49:07) Mohamed has been collaborating with a few institutions, podcasters, and ML influencers to raise awareness of the importance of ML testing and different approaches to tackle the problem. * (50:12) Mohamed touched on his side hustles working with Intel in autonomous drones and teaching content with Udacity’s AI Nanodegree programs. * (53:07) Mohamed dissected his project Mowgly, an educational platform with tracks curated by industry experts to guide users to master specific topics. * (54:58) Mohamed described his experience authoring a book with Manning in 2020 called “Deep Learning For Vision Systems.” * (58:51) Closing segment.
Mohamed’s Contact Info* LinkedIn * Twitter * Website * YouTube * GitHub * Kolena
Mentioned ContentPeople* Andrew Trask (Leader at OpenMined, Senior Research Scientist at DeepMind, Ph.D. Student at the University of Oxford) * Francois Chollet (Senior Software Engineer at Google, Creator of Keras) * Lex Fridman (Host of the popular Lex Fridman Podcast, AI Researcher working on autonomous vehicles and human-robot interaction at MIT)
Books* “Mindset” (by Carol Dweck) * “Outliers” (by Malcolm Gladwell)
NotesMy conversation with Mohamed was recorded back in March 2021. Here are some updates that Mohamed shared with me since then:
About the showDatacast features long-form, in-depth conversations with practitioners and researchers in the data community to walk through their professional journeys and unpack the lessons learned along the way. I invite guests coming from a wide range of career paths — from scientists and analysts to founders and investors — to analyze the case for using data in the real world and extract their mental models (“the WHY and the HOW”) behind their pursuits. Hopefully, these conversations can serve as valuable tools for early-stage data professionals as they navigate their own careers in the exciting data universe.
Datacast is produced and edited by James Le. Get in touch with feedback or guest suggestions by emailing khanhle.1013@gmail.com.
Subscribe by searching for Datacast wherever you get podcasts or click one of the links below:
If you’re new, see the podcast homepage for the most recent episodes to listen to, or browse the full guest list.
Show Notes* (01:46) Jennifer shared her formative experiences growing up in France and wanting to be a physicist. * (03:04) Jennifer unpacked the evolution of her academic journey in France — getting Physics degrees at Louis Pasteur University, Paris-Sud University, and Sorbonne University. * (06:44) Jennifer mentioned her time as a Postdoctoral Researcher in Neutrino Physics at Duke University, where her research group lacked the funding to carry on scientific projects. * (09:35) Jennifer discussed her transition from academia to industry, working as a Quantitative Research Scientist at Quantlab Financial in Houston. * (13:31) Jennifer went over her move to the Bay Area, working for YuMe and Ayasdi — growing and managing early-stage data science teams at both places. * (19:19) Jennifer recalled her foray into becoming a Senior Data Science Manager of the Search team at Walmart Labs. She managed the Metrics-Measurements-Insights team and the Store-Search team. * (23:59) Jennifer shared the business anecdote that made her obsessed with measuring the ROI of data science. * (28:46) Jennifer reflected on the opportunity to give conference talks and become a thought leader in the data science community (watch her first industry talk, “Review Analysis: An Approach to Leveraging User-Generated Content in the Context of Retail” at MLconf 2016). * (31:10) Jennifer unpacked her interest in active learning and outlined existing challenges of making active learning performant in real-world ML systems. * (36:58) After 1.5 years with Walmart Labs, Jennifer became the Chief Data Scientist at Atlassian. She shared the tactics to grow the Search & Smarts team of scientists and engineers from 3 to 17 people in less than 6 months across 3 locations. * (40:31) Jennifer discussed the organizational and operational challenges with making ML useful in enterprises and the importance of data preparation in the modern ML stack. * (47:24) Jennifer elaborated on the topic of “Agile for Data Science Teams,” which discusses that organizations that invest in ML but do not get the organizational side of things right will fail. * (53:09) Jennifer went over her decision to accept a VP of Machine Learning role at Figure Eight, then a frontier startup that offers enterprise-grade labeling solutions to ML teams. * (57:56) Jennifer went over the inception of her startup Alectio, whose mission is to help companies do ML more efficiently with fewer data and help the world do ML more sustainably by reducing the industry’s carbon footprint. * (01:04:32) Jennifer unpacked her 4-part blog series about responsible AI that calls out the need to fight bias, increase accessibility, and create more opportunities in AI. * (01:09:06) Jennifer discussed the hurdles she had to jump through to find early adopters of Alectio. * (01:11:03) Jennifer emphasized the valuable lessons learned to attract the right people who are excited about Alectio’s mission. * (01:14:38) Jennifer cautioned the danger of taking advice without thinking through how it can be applied to one’s career. * (01:17:09) Jennifer condensed her decade of experience navigating the tech industry as a woman into concrete advice. * (01:19:19) Closing segment.
Jennifer’s Contact Info* LinkedIn * Twitter * Medium
Alectio’s Resources* Website * Twitter * LinkedIn * What Is Alectio? (Video) * Is Big Data Dragging Us Towards Another AI Winter? (Article)
Mentioned ContentTalks* The Day Big Data Died (Oct 2020 @ Interop Digital) * The Importance of Ethics in Data Science (Keynote @ Women in Analytics Conference 2019) * Introduction to Active Learning (ODSC London 2018) * Agile for Data Science Teams (Strata Data Conf — New York 2018) * Big Data and the Advent of Data Mixology (Interop ITX — The Future of Data Summit 2017) * The Limitations of Big Data In Predictive Analytics (DataEngConf SF 2017) * Review Analysis: An Approach to Leveraging User-Generated Content in the Context of Retail (MLconf 2016)
Articles1 — Women vs. The Workplace Series
2 — Management Series
3 — Responsible AI Series
Book* “Managing Up” (by Rosanne Badowski and Roger Gittines)
NotesJennifer told me that Alectio is about to launch a community version that people will be able to compete to get the best model with the minimum amount of data this fall. Be sure to check out their blog and follow them on LinkedIn!
About the showDatacast features long-form, in-depth conversations with practitioners and researchers in the data community to walk through their professional journeys and unpack the lessons learned along the way. I invite guests coming from a wide range of career paths — from scientists and analysts to founders and investors — to analyze the case for using data in the real world and extract their mental models (“the WHY and the HOW”) behind their pursuits. Hopefully, these conversations can serve as valuable tools for early-stage data professionals as they navigate their own careers in the exciting data universe.
Datacast is produced and edited by James Le. Get in touch with feedback or guest suggestions by emailing khanhle.1013@gmail.com.
Subscribe by searching for Datacast wherever you get podcasts or click one of the links below:
If you’re new, see the podcast homepage for the most recent episodes to listen to, or browse the full guest list.
Show Notes (01:48) Sarah talked about the formative experiences of her upbringing: growing up interested in the natural sciences and switching focus on terrorism analysis after experiencing the 9/11 tragedy with her own eyes. * (04:07) Sarah discussed her experience studying International Security Studies at Stanford and working at the Center for International Security and Cooperation. * (07:15) Sarah recalled her first job out of college as a Program Director at the Center for Advanced Defense Studies — collaborating with academic researchers to develop computational approaches that counter terrorism and piracy. * (09:48) Sarah went over her time as a cyber-intelligence analyst at Cyveillance, which provided threat intelligence services to enterprises worldwide. * (12:22) Sarah walked over her time at Palantir as an embedded analyst, where she observed the struggles that many agencies had with data integration and modeling challenges. * (15:26) Sarah unpacked the challenges of building out the data team and applying the data work at Mattermark. * (20:15) Sarah shared her opinion on the career trajectory for data analysts and data scientists, given her experience as a manager for these roles. * (23:43) Sarah shared the power of having a peer group and building a team culture that she was proud of at Mattermark. * (26:41) Sarah joined Canvas Ventures as a Data Partner in 2016 and shared her motivation for getting into venture capital. * (29:47) Sarah revealed the secret sauce to succeed in venture — stamina*. * (32:00) Sarah has been an investor at Amplify Partners since 2017 and shared what attracted her about the firm’s investment thesis and the team. * (35:28) Sarah walked through the framework she used to prove her value upfront as the new investor at Amplify. * (38:35) Sarah shared the details behind her investment on the Series A round for OctoML, a Seattle-based startup that leverages Apache TVM to enable their clients to simply, securely, and efficiently deploy any model on any hardware backend. * (44:39) Sarah dissected her investment on the seed round for Einblick, a Boston-based startup that builds a visual computing platform for BI and analytics use cases. * (48:45) Sarah mentioned the key factors inspiring her investment in the seed round for Metaphor Data, a meta-data platform that grew out of the DataHub open-source project developed at LinkedIn. * (53:57) Sarah discussed what triggered her investment in the Series A round for Runway, a New York-based team building the next-generation creative toolkit powered by machine learning. * (58:36) Sarah unpacked the advice she has been giving her portfolio companies in hiring decisions and expanding their founding team (and advice they should ignore). * (01:01:29) Sarah went over the process of curating her weekly newsletter called Projects To Know (active since 2019). * (01:05:00) Sarah predicted the 3 trends in the data ecosystem that will have a disproportionately huge impact in the future. * (01:11:15) Closing segment.
Sarah’s Contact Info
Amplify Partners’ Resources
Mentioned ContentBlog Posts
People
Book
New UpdatesSince the podcast was recorded, Sarah has been keeping her stamina high!
Be sure to follow @sarahcat21 on Twitter to subscribe to her brain on the intersection of data, VC, and startups!
Show Notes* (01:39) Aparna talked about her Bachelor’s degree in Electrical Engineering and Computer Science at UC Berkeley. * (02:50) Aparna shared her undergraduate research experience at the Energy and Sustainable Technologies lab. * (04:34) Aparna discussed valuable lessons learned from her industry internships at TubeMogul and compared the objective with that of a research environment. * (08:26) Aparna then joined Uber as a software engineer on the Marketplace Forecasting team, where she led the development of Uber’s first model lifecycle management system for running ML model computations at scale to power Uber’s dynamic pricing algorithms. * (12:40) Aparna talked about how she became interested in model monitoring while Uber’s model store. * (17:29) Aparna discussed her decision to join the Ph.D. program in Computer Vision at Cornell University, specifically about bias in model, after spending 3 years at Uber. * (23:40) Aparna shared the backstory behind co-founding MonitorML with her brother Eswar and going through the 2019 summer batch of Y-Combinator. * (26:47) Aparna discussed the acquisition of MonitorML by Arize AI, where she’s currently the Chief Product Officer. * (28:41) Aparna unpacked the key insights in her ongoing ML Observability blog series, which argues that model observability is the foundational platform that empowers teams to continually deliver and improve results from the lab to production. * (33:17) Aparna shared her verdict for the ML tooling ecosystem in the upcoming years from her in-depth exploration of ML infrastructure tools covering data preparation, model building, model validation, and model serving. * (37:01) Aparna briefly shared the challenges encountered to get the first cohort of customers for Arize. * (39:23) Aparna went over valuable lessons to attract the right people who are excited about Arize’s mission. * (41:04) Aparna shared her advice for founders who are in the process of finding the right investors for their companies. * (42:24) Aparna reasoned how participating in The Amazing Race was similar to running a startup. * (44:59) Closing segment.
Aparna’s Contact Info
Arize’s Resources
Mentioned ContentBlog Posts
People
Book
New UpdatesSince the podcast was recorded, a lot has happened at Arize AI!
About The ShowDatacast features long-form, in-depth conversations with practitioners and researchers in the data community to walk through their professional journeys and unpack the lessons learned along the way. I invite guests coming from a wide range of career paths — from scientists and analysts to founders and investors — to analyze the case for using data in the real world and extract their mental models (“the WHY”) behind their pursuits. Hopefully, these conversations can serve as valuable tools for early-stage data professionals as they navigate their own careers in the exciting data universe.
Datacast is produced and edited by James Le. Get in touch with feedback or guest suggestions by emailing khanhle.1013@gmail.com.
Subscribe by searching for Datacast wherever you get podcasts or click one of the links below:
If you’re new, see the podcast homepage for the most recent episodes to listen to, or browse the full guest list.
Show Notes* (02:07) Emeli shared her educational background getting degrees in Applied Mathematics and Informatics from the Peoples’ Friendship University of Russia in the early 2010s. * (04:24) Emeli went over her experience getting a Master’s Degree at Yandex School of Data Analysis. * (07:06) Emeli reflected on lessons learned from her first job out of university working as a Software Developer at Rambler, one of the biggest Russian web portals. * (09:33) Emeli walked over her first year as a Data Scientist developing e-commerce recommendation systems at Yandex. * (13:38) Emeli discussed core projects accomplished as the Chief Data Scientist at Yandex Data Factory, Yandex’s end-to-end data platform. * (17:52) Emeli shared her learnings transitioning from an IC to a manager role. * (19:21) Emeli mentioned key components of success for industrial AI, given her time as the co-founder and Chief Data Scientist at Mechanica AI. * (22:40) Emeli dissected the makings of her Coursera specializations — “Machine Learning and Data Analysis” and “Big Data Essentials.” * (26:14) Emeli discussed her teaching activities at Moscow Institute of Physics and Technology, Yandex School of Data Analysis, Harbour.Space, and Graduate School of Management — St. Petersburg State University. * (30:12) Emeli shared the story behind the founding of Evidently AI, which is building a human interface to machine learning, so that companies can trust, monitor, and improve the performance of their AI solutions. * (32:32) Emeli explained the concept of model monitoring and exposed the monitoring gap in the enterprise (read Part 1 and Part 2 of the Monitoring series). * (34:13) Emeli looked at possible data quality and integrity issues while proposing how to track them (read Part 3, Part 4, and Part 5 of the Monitoring series). * (36:47) Emeli revealed the pros and cons of building an open-source product. * (39:13) Emeli talked about prioritizing product roadmap for Evidently AI. * (41:24) Emeli described the data community in Moscow. * (42:03) Closing segment.
Emeli’s Contact Info
Evidently AI’s Resources
Mentioned ContentBlog Posts
Courses
People
Book
New UpdatesSince the podcast was recorded, a lot has happened at Evidently! You can use this open-source tool (https://github.com/evidentlyai/evidently) to generate a variety of interactive reports on the ML model performance and integrate it into your pipelines using JSON profiles.
This monitoring tutorial is a great showcase of what can go wrong with your models in production and how to keep an eye on them: https://evidentlyai.com/blog/tutorial-1-model-analytics-in-production.
About The ShowDatacast features long-form conversations with practitioners and researchers in the data community to walk through their professional journey and unpack the lessons learned along the way. I invite guests coming from a wide range of career paths - from scientists and analysts to founders and investors — to analyze the case for using data in the real world and extract their mental models (“the WHY”) behind their pursuits. Hopefully, these conversations can serve as valuable tools for early-stage data professionals as they navigate their own careers in the exciting data universe.
Datacast is produced and edited by James Le. Get in touch with feedback or guest suggestions by emailing khanhle.1013@gmail.com.
Subscribe by searching for Datacast wherever you get podcasts or click one of the links below:
If you're new, see the podcast homepage for the most recent episodes to listen to, or browse the full guest list.
Show Notes* (01:59) David recalled his undergraduate experience studying Physics and Mathematics at Duke University back in the early 90s. * (05:55) David reflected on his decision to pursue a Ph.D. in Physics at the University of Maryland, College Park, specializing in Nonlinear Dynamics and Chaos Theory. * (10:18) David unpacked his Nature paper called “Topology in Chaotic Scattering.” * (14:43) David went over his two papers on fractal dimensions in higher-dimensional chaotic scattering following his Nature publication. * (21:42) David talked about his project K Desktop Environment, which provides a free, user-friendly desktop for Linux/UNIX systems (later turned into a print book with MacMillan Publishing in 2000). * (24:20) David explained the premise behind his work on Andamooka, a site that supports open content. * (27:24) David walked over his time as a quantitative analyst at Thales Fund Management after finishing his Ph.D. * (30:50) David discussed his 4-year stint at Lehman Brothers — moving up the ladder into a Vice President role, up until Barclay’s Capital acquired it. * (33:24) David talked about his proudest accomplishment during the 5-year stint as a headdesk in equities trader at KCG/GETCO. * (35:37) David shared war stories while working at an investment firm called Teza Technologies and co-founding Galaxy Digital Trading (specializing in cryptocurrency trading). * (41:34) David unpacked key concepts covered in his guest lectures on optimization of high-frequency trading systems at NYU Stern School of Business. * (44:26) David explained his career change to work as a Machine Learning Engineer at Instagram in the summer of 2019. * (47:17) David briefly mentioned his transition back to a quant trader role at 3Red Partners. * (48:05) David is writing a technical book with Manning called “Tuning Up,” which provides a toolbox of experimental methods that will boost the effectiveness of machine learning systems, trading strategies, infrastructure, and more. * (50:48) David reflected on the benefits of his physics academic background for his quant analyst career. * (52:27) Closing segment.
David’s Contact Info* Website * LinkedIn * Twitter
Mentioned ContentPublications
Book
People
Tuning Up — From A/B testing to Bayesian optimizationManning’s permanent 40% discount code (good for all Manning products in all formats) for Datacast listeners: poddcast19.
You can refer to this link: http://mng.bz/4MAR.
Here are two free eBook codes to get copies of Tuning Up for two lucky Datacast listeners: tngdtcr-AB2C and tngdtcr-6D43
You can refer to this link: http://mng.bz/G6Bq.
Show Notes* (02:06) Fabiana talked about her Bachelor’s degree in Applied Mathematics from the University of Lisbon in the early 2010s. * (04:18) Fabiana shared lessons learned from her first job out of college as a Siebel and BI Developer at Novabase. * (05:13) Fabiana discussed unique challenges while working as an IoT Solutions Architect at Vodafone. * (09:56) Fabiana mentioned projects she contributed to as a Data Scientist at startups such as ODYSAI and Habit Analytics. * (12:44) Fabiana talked about the two Master’s degrees she got while working in the industry (Applied Econometrics from Lisbon School of Economics and Management and Business Intelligence from NOVA IMS Information Management School). * (14:41) Fabiana distinguished the difference between data science and business intelligence. * (18:01) Fabiana shared the founding story of YData, the first data-centric platform with synthetic data, whose she is currently the Chief Data Officer. * (21:32) Fabiana discussed different techniques to generate synthetic data, including oversampling, Bayesian Networks, and generative models. * (24:01) Fabiana unpacked the key insights in her blog series on generating synthetic tabular data. * (29:40) Fabiana summarized novel design and optimization techniques to cope with the challenges of training GAN models. * (33:44) Fabiana brought up the benefits of using Differential Privacy as a complement to synthetic data generation. * (38:07) Fabiana unpacked her post “The Cost of Poor Data Quality,” — where she defined data quality as data measures based on factors such as accuracy, completeness, consistency, reliability, and above all, whether it is up to date. * (42:11) Fabiana explained the important role that data quality plays in ensuring model explainability. * (44:57) Fabiana reasoned about YData’s decision to pursue the open-source strategy. * (47:47) Fabiana discussed her podcast called “When Machine Learning Meets Privacy” in collaboration with the MLOps Slack community. * (49:14) Fabiana briefly shared the challenges encountered to get the first cohort of customers for YData. * (50:12) Fabiana went over valuable lessons to attract the right people who are excited about YData’s mission. * (51:52) Fabiana shared her take on the data community in Lisbon and her effort to inspire more women to join the tech industry. * (53:47) Closing segment.
Fabiana’s Contact Info* LinkedIn * Medium * Twitter
YData’s Resources* Website * Github * LinkedIn * Twitter * AngelList * Synthetic Data Community
Mentioned ContentBlog Posts
Podcast
People
Recent Announcements/Articles
Show Notes (02:06) Azin described her childhood growing up in Iran and going to a girls-only high school in Tehran designed specifically for extraordinary talents. * (05:08) Azin went over her undergraduate experience studying Computer Science at the University of Tehran. * (10:41) Azin shared her academic experience getting a Computer Science MS degree at the University of Toronto, supervised by Babak Taati and David Fleet. * (14:07) Azin talked about her teaching assistant experience for a variety of CS courses at Toronto. * (15:54) Azin briefly discussed her 2017 report titled “Barriers to Adoption of Information Technology in Healthcare,” which takes a system thinking perspective to identify barriers to the application of IT in healthcare and outline the solutions. * (19:35) Azin unpacked her MS thesis called “Subspace Selection to Suppress Confounding Source Domain Information in AAM Transfer Learning,” which explores transfer learning in the context of facial analysis. * (28:48) Azin discussed her work as a research assistant at the Toronto Rehabilitation Institute, working on a research project that addressed algorithmic biases in facial detection technology for older adults with dementia. * (33:02) Azin has been an Applied Research Scientist at Georgian since 2018, a venture capital firm in Canada that focuses on investing in companies operating in the IT sectors. * (38:20) Azin shared the details of her initial Georgian project to develop a robust and accurate injury prediction model using a hybrid instance-based transfer learning method.* * (42:12) Azin unpacked her Medium blog post discussing transfer learning in-depth (problems, approaches, and applications). * (48:18) Azin explained how transfer learning could address the widespread “cold-start” problem in the industry. * (49:50) Azin shared the challenges of working on a fintech platform with a team of engineers at Georgian on various areas such as supervised learning, explainability, and representation learning. * (51:46) Azin went over her project with Tractable AI, a UK-based company that develops AI applications for accident and disaster recovery. * (55:26) Azin shared her excitement for ML applications using data-efficient methods to enhance life quality. * (57:46) Closing segment.
Azin’s Contact Info* Website * Twitter * LinkedIn * Google Scholar * GitHub
Mentioned ContentPublications
Blog Posts
People
Book
Note: Azin and her collaborator are going to give a talk at ODSC Europe 2021 in June about a Georgian’s project with a portfolio company, Tractable. They have written a short blog post about it too which you can find HERE.
Show Notes* (02:09) Gordon briefly talked about his undergraduate studying Psychology and Philosophy at Rutgers University in the early 90s. * (03:24) Gordon reflected on the first decade of his career getting into database technologies. * (05:34) Gordon discussed his predilection towards consulting, specifically his role in the professional services team at AB Initio Software in the early 2000s. * (08:02) Gordon recalled the challenges of leading data warehousing initiatives at Smarter Travel Media and ClickSquared in the 2000s. * (13:14) Gordon emphasized the advantage of a multi-tenant database over a traditional relational database. * (18:30) Gordon recalled his one-year stint at Cervello, leading business intelligence implementations for their clients. * (21:59) Gordon elaborated on his projects during his 3 years as the director of business intelligence infrastructure at Fitbit. * (26:09) Gordon dived into his framework of choosing data tooling vendors while at Fitbit (and how he settled with a tiny startup called Snowflake back then). * (30:02) Gordon provided recommendations for startups to be data-driven. * (33:24) Gordon recalled practices to foster effective collaboration while managing the 3 teams of data engineering, data warehousing, and data analytics at Fitbit. * (36:44) Gordon went over his proudest accomplishment as the director of data engineering at ezCater, making substantial improvements to their data warehouse platform. * (38:59) Gordon shared his framework for interviewing data engineers. * (41:39) Gordon walked through his consulting engagement in analytics engineering for Zipcar and data warehousing for edX. * (46:17) Gordon reflected on his time as the Vice President of business intelligence at HubSpot. * (50:50) Gordon unpacked his notion of “Data Hierarchy of Needs,” which entails the five pillars — data security, data quality, system reliability, user experience, and data coverage. * (56:55) Gordon discussed current opportunities for driving better social outcomes and empowering democracy through data. * (59:48) Gordon shared the key criteria that enable healthy team dynamics from his hands-on experience building data teams. * (01:02:13) Gordon unpacked the central features and benefits of Snowflake for the un-initiated. * (01:06:25) Gordon gave his verdict for the ETL tooling landscape in the next few years. * (01:08:33) Gordon described the data community in Boston. * (01:09:52) Closing segment.
Gordon’s Contact Info* LinkedIn
Mentioned ContentPeople
Book
Show Notes* (2:05) Louis went over his childhood as a self-taught programmer and his early days in school as a freelance developer. * (4:22) Louis described his overall undergraduate experience getting a Bachelor’s degree in IT Systems Engineering from Hasso Plattner Institute, a highly-ranked computer science university in Germany. * (6:10) Louis dissected his Bachelor thesis at HPI called “Differentiable Convolutional Neural Network Architectures for Time Series Classification,” — which addresses the problem of automatically designing architectures for time series classification efficiently, using a regularization technique for ConvNet that enables joint training of network weights and architecture through back-propagation. * (7:40) Louis provided a brief overview of his publication “Transfer Learning for Speech Recognition on a Budget,” — which explores Automatic Speech Recognition training by model adaptation under constrained GPU memory, throughput, and training data. * (10:31) Louis described his one-year Master of Research degree in Computational Statistics and Machine Learning at the University College London supervised by David Barber. * (12:13) Louis unpacked his paper “Modular Networks: Learning to Decompose Neural Computation,” published at NeurIPS 2018 — which proposes a training algorithm that flexibly chooses neural modules based on the processed data. * (15:13) Louis briefly reviewed his technical report, “Scaling Neural Networks Through Sparsity,” which discusses near-term and long-term solutions to handle sparsity between neural layers. * (18:30) Louis mentioned his report, “Characteristics of Machine Learning Research with Impact,” which explores questions such as how to measure research impact and what questions the machine learning community should focus on to maximize impact. * (21:16) Louis explained his report, “Contemporary Challenges in Artificial Intelligence,” which covers lifelong learning, scalability, generalization, self-referential algorithms, and benchmarks. * (23:16) Louis talked about his motivation to start a blog and discussed his two-part blog series on intelligence theories (part 1 on universal AI and part 2 on active inference). * (27:46) Louis described his decision to pursue a Ph.D. at the Swiss AI Lab IDSIA in Lugano, Switzerland, where he has been working on Meta Reinforcement Learning agents with Jürgen Schmidhuber. * (30:06) Louis created a very extensive map of reinforcement learning in 2019 that outlines the goal, methods, and challenges associated with the RL domain. * (33:50) Louis unpacked his blog post reflecting on his experience at NeurIPS 2018 and providing updates on the AGI roadmap regarding topics such as scalability, continual learning, meta-learning, and benchmarks. * (37:04) Louis dissected his ICLR 2020 paper “Improving Generalization in Meta Reinforcement Learning using Learned Objectives,” which introduces a novel algorithm called MetaGenRL, inspired by biological evolution. * (44:03) Louis elaborated on his publication “Meta-Learning Backpropagation And Improving It,” which introduces the Variable Shared Meta-Learning framework that unifies existing meta-learning approaches and demonstrates that simple weight-sharing and sparsity in a network are sufficient to express powerful learning algorithms. * (51:14) Louis expands on his idea to bootstrap AI that entails how the task, the general meta learner, and the unsupervised objective should interact (proposed at the end of his invited talk at NeurIPS 2020). * (54:14) Louis shared his advice for individuals who want to make a dent in AI research. * (56:05) Louis shared his three most useful productivity tips. * (58:36) Closing segment.
Louis’s Contact Info* Website * Twitter * LinkedIn * Google Scholar * GitHub
Mentioned ContentPapers and Reports
Blog Posts
People
Book
Show Notes* (01:58) Dzejla described her undergraduate experience studying Computer Science at the Sarajevo School of Science and Technology back in the mid-2000s. * (07:59) Dzejla recapped her overall experience getting a Ph.D. in Computer Science at Stony Brook University. * (14:38) Dzejla unpacked the key research problem in her Ph.D. thesis titled “Upper and Lower Bounds on Sorting and Searching in External Memory.” * (19:13) Dzejla went over the details of her paper “Don’t Thrash: How to Cache Your Hash on Flash,” — which describes the Cascade Filter, an approximate-membership-query data structure that scales beyond main memory, that is an alternative to the well-known Bloom-filter data structure. * (24:41) Dzejla elaborated on her work “The batched predecessor problem in external memory,” — which studies the lower bounds in three external memory models: the I/O comparison model, the I/O pointer-machine model, and the index-ability model. * (29:56) Dzejla shared her learnings from being a Teaching Assistant for the Introduction to Algorithms course at Stony Brook (both at the undergraduate and graduate level). * (35:08) Dzejla went over her summer internships at Microsoft’s Server and Tools Division during her Ph.D. * (41:06) Dzejla reasoned about her decision to return to Sarajevo School of Science and Technology as an Assistant Professor of Computer Science. * (47:22) Dzejla dissected the essential concepts and methods covered in her Data Structures, Introductory Algorithms, Advanced Algorithms, and Algorithms for Big Data courses taught at SSIT. * (48:42) Dzejla provided a brief overview of the Computer Science/Software Engineering department at the International University of Sarajevo (where she has been a professor since 2017. * (50:57) Dzejla briefly talked about the courses that she taught at IUS, including Intro to Programming, Human-Computer Interaction, and Algorithms/Data Structures. * (52:49) Dzejla shared the challenges of writing Algorithms and Data Structures for Massive Datasets, which introduces data processing and analytics techniques specifically designed for large distributed datasets. * (56:14) Dzejla explained concepts in Part 1 of the book — including Hash Tables, Approximate Membership, Bloom Filters, Frequency/Cardinality Estimation, Count-Min Sketch, and Hyperloglog. * (58:38) Dzejla provided a brief overview of techniques to handle streaming data in Part 2 of the book. * (01:00:14) Dzejla mentioned the data structures for large databases and external-memory algorithms in Part 3 of the book. * (01:02:15) Dzejla shared her thoughts about the tech community in Sarajevo. * (01:04:16) Closing segment.
Dzejla’s Contact Info* LinkedIn * Twitter * Google Scholar
Mentioned ContentPapers
People
Books
Here is a permanent 40% discount code (good for all Manning products in all formats) for Datacast listeners: poddcast19. Link at http://mng.bz/4MAR.
Here is one free eBook code good for a copy of Algorithms and Data Structures for Massive Datasets for a lucky listener: algdcsr-7135. Link at http://mng.bz/Q2y6
Show Notes* (1:45) Willem discussed his undergraduate degree in Mechatronic Engineering at Stellenbosch University in the early 2010s. * (2:34) Willem recalled his entrepreneurial journey founding and selling a networking startup that provides internet access to private residents on campus. * (5:37) Willem worked for two years as a Software Engineer focusing on data systems at Systems Anywhere in Capetown after college. * (6:49) Willem talked about his move to Bangkok working as a Senior Software Engineer at INDEFF, a company in industrial control systems. * (9:52) Willem went over his decision to join Gojek, a leading Indonesian on-demand multi-service platform and digital payment technology group. * (12:16) Willem mentioned the engineering challenges associated with building complex data systems for super-apps. * (14:50) Willem dissected Gojek’s ML platform, including these four solutions for various stages of the ML life cycle: Clockwork, Merlin, Feast, and Turing. * (19:24) Willem recapped the lessons from designing the ML platform to meet Gojek’s scaling requirements — as delivered at Cloud Next 2018. * (23:09) Willem briefly went through the key design components to incorporate Kubeflow pipelines into Gojek’s existing ML platform — as delivered at KubeCon 2019. * (26:21) Willem explained the inception of Feast, an open-source feature store that bridges the gap between data and models. * (32:20) Willem talked about prioritizing the product roadmap and engaging the community for an open-source project. * (35:07) Willem recapped the key lessons learned and envisioned Feast's future to be a lightweight modular feature store. * (37:29) Willem explained the differences between commercial and open-source feature stores (given Tecton’s recent backing of Feast). * (41:36) Willem reflected on his experience living and working in Southeast Asia. * (44:33) Closing segment.
Willem’s Contact Info* Twitter * LinkedIn * GitHub
Mentioned ContentFeast
Article
Talks
People
Book
Willem will be a speaker at Tecton’s apply() virtual conference (April 21-22, 2021) for data and ML teams to discuss the practical data engineering challenges faced when building ML for the real world. Participants will share best practice development patterns, tools of choice, and emerging architectures they use to successfully build and manage production ML applications. Everything is on the table from managing labeling pipelines, to transforming features in real-time, and serving at scale. Register for free now: https://www.applyconf.com/!
Show Notes* (1:56) Jim went over his education at Trinity College Dublin in the late 90s/early 2000s, where he got early exposure to academic research in distributed systems. * (4:26) Jim discussed his research focused on dynamic software architecture, particularly the K-Component model that enables individual components to adapt to a changing environment. * (5:37) Jim explained his research on collaborative reinforcement learning that enables groups of reinforcement learning agents to solve online optimization problems in dynamic systems. * (9:03) Jim recalled his time as a Senior Consultant for MySQL. * (9:52) Jim shared the initiatives at the RISE Research Institute of Sweden, in which he has been a researcher since 2007. * (13:16) Jim dissected his peer-to-peer systems research at RISE, including theoretical results for search algorithm and walk topology. * (15:30) Jim went over challenges building peer-to-peer live streaming systems at RISE, such as GradientTV and Glive. * (18:18) Jim provided an overview of research activities at the Division of Software and Computer Systems at the School of Electrical Engineering and Computer Science at KTH Royal Institute of Technology. * (19:04) Jim has taught courses on Distributed Systems and Deep Learning on Big Data at KTH Royal Institute of Technology. * (22:20) Jim unpacked his O’Reilly article in 2017 called “Distributed TensorFlow,” which includes the deep learning hierarchy of scale. * (29:47) Jim discussed the development of HopsFS, a next-generation distribution of the Hadoop Distributed File System (HDFS) that replaces its single-node in-memory metadata service with a distributed metadata service built on a NewSQL database. * (34:17) Jim rationalized the intention to commercialize HopsFS and built Hopsworks, an user-friendly data science platform for Hops. * (36:56) Jim explored the relative benefits of public research money and VC-funded money. * (41:48) Jim unpacked the key ideas in his post “Feature Store: The Missing Data Layer in ML Pipelines.” * (47:31) Jim dissected the critical design that enables the Hopsworks feature store to refactor a monolithic end-to-end ML pipeline into separate feature engineering and model training pipelines. * (52:49) Jim explained why data warehouses are insufficient for machine learning pipelines and why a feature store is needed instead. * (57:59) Jim discussed prioritizing the product roadmap for the Hopswork platform. * (01:00:25) Jim hinted at what’s on the 2021 roadmap for Hopswork. * (01:03:22) Jim recalled the challenges of getting early customers for Hopsworks. * (01:04:30) Jim intuited the differences and similarities between being a professor and being a founder. * (01:07:00) Jim discussed worrying trends in the European Tech ecosystem and the role that Logical Clocks will play in the long run. * (01:13:37) Closing segment.
Jim’s Contact Info* Logical Clocks * Twitter * LinkedIn * Google Scholar * Medium * ACM Profile * GitHub
Mentioned ContentResearch Papers
Articles
Projects
People
Programming Books
Show Notes (2:20) Pier shared his college experience at the University of Southampton studying Electronic Engineering. * (3:46) For his final undergraduate project, Pier developed a suite of games and used machine learning to analyze brainwaves data that can classify whether a child is affected or not by autism. * (11:26) Pier went over his favorite courses and involvement with the AI Society during his additional year at the University of Southampton to get a Master’s in Artificial Intelligence. * (13:40) For his Master’s thesis called “Causal Reasoning in Machine Learning,” Pier created and deployed a suite of Agent-Based and Compartmental Models to simulate epidemic disease developments in different types of communities. * (26:51) Pier went over his stints as a developer intern at Fidessa and a freelance data scientist at Digital-Dandelion. * (29:21) Pier reflected on his time (so far) as a data scientist at SAS Institute, where he helps their customers solve various data-driven challenges using cloud-based technologies and DevOps processes. * (33:37) Pier discussed the key benefits that writing and editing technical content for Towards Data Science to his professional development. * (36:31) Pier covered the threads that he kept pulling with his blog posts. * (38:50) Pier talked about his Augmented Reality Personal Business Card created in HTML using the AR.js library. * (41:12) Pier brought up data structures in two other impressive JavaScript projects using TensorFlow.js and ml5.js. * (44:19) Pier went over his experience working with data visualization tools such as Plotly, R Shiny, and Streamlit. * (47:27) Pier talked about his work on a chapter for a book called “Applied Data Science in Tourism” that is going to be published with Springer this year. * (48:37) Pier shared his thoughts r*egarding the tech community in London. * (49:19) Closing segment.
Pier’s Contact Info* Website * LinkedIn * Twitter * GitHub * Medium * Patreon * Kaggle
Mentioned Content* “Alleviate Children’s Health Issues Through Games and Machine Learning” * “Causal Reasoning in Machine Learning” * Andrej Karpathy (Director of AI and Autopilot at Tesla) * Cassie Kozyrkov (Chief Decision Scientist at Google) * Iain Brown (Head of Data Science at SAS) * “The Book Of Why” (By Judea Pearl) * “Pattern Recognition and Machine Learning” (by Christopher Bishop)
Timestamps
Her Contact Info
Her Recommended Resources
Timestamps* (2:07) JY discussed his college time studying Computer Science and Applied Math at Ecole Polytechnique — a leading French institute in science and technology. * (3:04) JY reflected on time at Stanford getting a Master’s in Management Science and Engineering, where he served as a Teaching Assistant for CS 229 (Machine Learning) and CS 246 (Mining Massive Datasets). * (6:14) JY walked over his ML engineering internship at LiveRamp — a data connectivity platform for the safe and effective use of data. * (7:54) JY reflected on his next three years at Databricks, first as a software engineer and then as a tech lead for the Spark Infrastructure team. * (10:00) JY unpacked the challenges of packaging/managing/monitoring Spark clusters and automating the launch of hundreds of thousands of nodes in the cloud every day. * (14:48) JY shared the founding story behind Data Mechanics, whose mission is to give superpowers to the world's data engineers so they can make sense of their data and build applications at scale on top of it. * (18:09) JY explained the three tenets of Data Mechanics: (1) managed and serverless, (2) integrated into clients’ workflows, and (3) built on top of open-source software (read the launch blog post). * (22:06) JY unpacked the core concepts of Spark-On-Kubernetes and evaluated the benefits/drawbacks of this new deployment mode — as presented in “Pros and Cons of Running Apache Spark on Kubernetes.” * (26:00) JY discussed Data Mechanics’ main improvements on the open-source version of Spark-On-Kubernetes — including an intuitive user interface, dynamic optimizations, integrations, and security — as explained in “Spark on Kubernetes Made Easy.” * (28:35) JY went over Data Mechanics Delight, a customized Spark UI which was recently open-sourced. * (35:40) JY shared the key ideas in his thought-leading piece on how to be successful with Apache Spark in 2021. * (38:42) JY went over his experience going through the Y Combinator program in summer 2019. * (40:56) JY reflected on the key decisions to get the first cohort of customers for Data Mechanics. * (42:26) JY shared valuable hiring lessons for early-stage startup founders. * (44:34) JY described the data and tech community in France. * (47:19) Closing segment.
His Contact Info
His Recommended Resources
Timestamps* (2:55) Chris went over his experience studying Computer Science at the University of Southern California for undergraduate in the late 90s. * (5:26) Chris recalled working as a Software Engineer at NASA Jet Propulsion Lab in his sophomore year at USC. * (9:54) Chris continued his education at USC with an M.S. and then a Ph.D. in Computer Science. Under the guidance of Dr. Nenad Medvidović, his Ph.D. thesis is called “Software Connectors For Highly-Distributed And Voluminous Data-Intensive Systems.” He proposed DISCO, a software architecture-based systematic framework for selecting software connectors based on eight key dimensions of data distribution. * (16:28) Towards the end of his Ph.D., Chris started getting involved with the Apache Software Foundation. More specifically, he developed the original proposal and plan for Apache Tika (a content detection and analysis toolkit) in collaboration with Jérôme Charron to extract data in the Panama Papers, exposing how wealthy individuals exploited offshore tax regimes. * (24:58) Chris discussed his process of writing “Tika In Action,” which he co-authored with Jukka Zitting in 2011. * (27:01) Since 2007, Chris has been a professor in the Department of Computer Science at USC Viterbi School of Engineering. He went over the principles covered in his course titled “Software Architectures.” * (29:49) Chris touched on the core concepts and practical exercises that students could gain from his course “Information Retrieval and Web Search Engines.” * (32:10) Chris continued with his advanced course called “Content Detection and Analysis for Big Data” in recent years (check out this USC article). * (36:31) Chris also served as the Director of the USC’s Information Retrieval and Data Science group, whose mission is to research and develop new methodology and open source software to analyze, ingest, process, and manage Big Data and turn it into information. * (41:07) Chris unpacked the evolution of his career at NASA JPL: Member of Technical Staff -> Senior Software Architect -> Principal Data Scientist -> Deputy Chief Technology and Innovation Officer -> Division Manager for the AI, Analytics, and Innovation team. * (44:32) Chris dove deep into MEMEX — a JPL’s project that aims to develop software that advances online search capabilities to the deep web, the dark web, and nontraditional content. * (48:03) Chris briefly touched on XDATA — a JPL’s research effort to develop new computational techniques and open-source software tools to process and analyze big data. * (52:23) Chris described his work on the Object-Oriented Data Technology platform, an open-source data management system originally developed by NASA JPL and then donated to the Apache Software Foundation. * (55:22) Chris shared the scientific challenges and engineering requirements associated with developing the next generation of reusable science data processing systems for NASA’s Orbiting Carbon Observatory space mission and the Soil Moisture Active Passive earth science mission. * (01:01:05) Chris talked about his work on NASA’s Machine Learning-based Analytics for Autonomous Rover Systems — which consists of two novel capabilities for future Mars rovers (Drive-By Science and Energy-Optimal Autonomous Navigation). * (01:04:24) Chris quantified the Apache Software Foundation's impact on the software industry in the past decade and discussed trends in open-source software development. * (01:07:15) Chris unpacked his 2013 Nature article called “A vision for data science” — in which he argued that four advancements are necessary to get the best out of big data: algorithm integration, development and stewardship, diverse data formats, and people power. * (01:11:54) Chris revealed the challenges of writing the second edition of “Machine Learning with TensorFlow,” a technical book with Manning that teaches the foundational concepts of machine learning and the TensorFlow library's usage to build powerful models rapidly. * (01:15:04) Chris mentioned the differences between working in academia and industry. * (01:16:20) Chris described the tech and data community in the greater Los Angeles area. * (01:18:30) Closing segment.
His Contact Info* Wikipedia * NASA Page * Google Scholar * USC Page * Twitter * LinkedIn * GitHub
His Recommended Resources* Doug Cutting (Founder of Lucene and Hadoop) * Hilary Mason (Ex Data Scientist at bit.ly and Cloudera) * Jukka Zitting (Staff Software Engineer at Google) * "The One Minute Manager" (by Ken Blanchard and Spencer Johnson)
Show Notes (2:09) Marcello described his academic experience getting a Master’s Degree in Computer Science from the Universita di Catania in the early 2000s, where his thesis is called Evolutionary Randomized Graph Embedder. * (6:14) Marcello commented on his career phase working as a web developer across various places in Europe. * (9:18) Marcello discussed his time working as a software engineer at INPS, a government-owned company that now handles most Italian citizens' pubic-related data. * (10:42) Marcello talked about his time as a data visualization engineer at SwiftIQ. He created a data visualization library that allows the inclusion of dynamic charts in HTML pages with just a few JavaScript lines. * (13:40) Marcello went over his projects while working as a full-stack software engineer for Twitter’s User Services Engineering team in Dublin. * (17:19) Marcello reflected on his time at Microsoft Zurich’s Social and Engagement team, contributing to machine learning infrastructure and tools. * (21:28) Marcello briefly touched on his one-year stint at Apple Zurich as a Senior Applied Research Engineer. * (23:49) Marcello talked about the challenges while writing “Algorithms and Data Structures in Action,” which introduces a diverse range of algorithms used in web apps, systems programming, and data manipulation. * (27:11) Marcello expanded upon part 1 of the book, including advanced data structures such as D-ary Heaps, Randomized Treaps, Bloom Filters, Disjoint Sets, Tries/Radix Trees, and Cache. * (34:51) Marcello brought up data structures to perform efficient multi-dimensional queries, including various nearest neighbor searches and clustering techniques, in part 2 of the book. * (39:21) Marcello briefly described the algorithms in part 3 of the book — graph embeddings, gradient descent, simulated annealing, and genetic algorithms. * (48:28) Marcello talked about his work on jsgraph — a lightweight library to model graphs, run graphs algorithms, and display them on screen. * (52:06) Marcello compared Python, Java, and JavaScript programming languages. * (54:13) Marcello discussed his current interest in quantum computing. * (56:18) Marcello shared his thoughts r*egarding Dublin, Zurich, and Rome's tech communities. * (57:37) Closing segment.
His Contact Info* Twitter * LinkedIn * GitHub * Blog
His Recommended Resources* "Algorithms and Data Structures in Action" (Marcello's book with Manning) * Andrew Ng * Geoffrey Hinton * Francois Chollet * "Scalability Rules" (by Martin Abbott and Michael Fischer)
This is the 40% discount code that is good for all Manning’s products in all formats: poddcast19.
These are 5 free eBook codes, each good for one copy of “Algorithms and Data Structures in Action”:
Show Notes* (2:10) Dave talked briefly about his Electrical Engineering study at Rensselaer Polytechnic Institute back in the late 90s. * (4:03) Dave commented on his career phase working as a software engineer across various companies in Bozeman, Montana. * (7:38) Dave discussed his work as a senior architect and tech lead at Expero, a Houston-based startup that develops custom software exclusively for domain-expert users. * (11:26) Dave briefly defined common big data frameworks (Hadoop, Apache Spark) and databases (Apache Cassandra, Apache Kafka). * (13:37) Dave went over the challenges during his time as a chief software architect at Gene by Gene, a biotech company focusing on DNA-based ancestry and genealogy. * (20:00) Dave shared the common patterns and anti-patterns of using graph databases (in reference to his talk “A Practical Guide to Graph Databases”). * (26:16) Dave walked through the three categories of graph technologies: Graph Computing Engine, RDF TripleStore, and Labeled Property Graph (in reference to his talk “A Skeptics Guide to Graph Databases”). * (33:03) Dave discussed his move to DataStax’s Global Graph Practice team as a solutions architect and graph database subject matter expert. * (36:00) Dave explained the design of DataStax’s enterprise solution called Customer 360, which collapses data silos to drive business value. * (41:16) Dave talked about his current experience as a Senior Graph Architect at AWS. * (43:51) Dave mentioned the challenges while writing "Graph Databases In Action" (published last October). * (47:25) Dave explained the open-source Apache TinkerPop framework and the Gremlin language used in the book for the uninitiated. * (51:04) Dave discussed trends in big data and distributed systems that he is most excited about. * (55:06) Closing segment.
His Contact Info* Website * Twitter * LinkedIn * GitHub
His Recommended Resources* "Graph Databases In Action" (Associated Code Repository) * Martin Fowler (Founder of ThoughtWorks) * Martin Kleppmann (Author of "Designing Data-Intensive Applications") * Andrew Ng (Professor at Stanford, Co-Founder of Google Brain and Coursera, Ex-Chief Scientist at Baidu) * "Pragmatic Programmer" (by Andy Hunt and Dave Thomas) * "The Five Dysfunctions Of A Team" (by Patrick Lencioni) * "How To Observe Scientific Advice for Common Real-World Problems" (by Randall Munroe)
This is the 40% discount code that is good for all Manning's products in all formats: poddcast19.
These are 5 free eBook codes, each good for one copy of "Graph Databases In Action":
Show Notes* (2:13) Jason went over his experience studying Computer Science at Loyola College in Baltimore for undergraduate, where he got an early exposure to academic research in image registration. * (4:31) Jason described his graduate school experience at John Hopkins University, where he completed his Ph.D. on “Techniques for Vision-Based Human-Computer Interaction” that proposed the Visual Interaction Cues paradigm. * (9:31) During his time as a Post-Doc Fellow at UCLA, Jason helped develop automatic segmentation and recognition techniques for brain tumors to improve the accuracy of diagnosis and treatment accuracy * (14:27) From 2007 to 2014, Jason was a professor in the Computer Science and Engineering department at SUNY-Buffalo. He covered the content of two graduate-level courses on Bayesian Vision and Intro to Pattern Recognition that he taught. * (18:20) On the topic of metric learning, Jason proposed an approach to data analysis and modeling for computer vision called "Active Clustering." * (21:35) On the topic of image understanding, Jason created Generalized Image Understanding - a project that examined a unified methodology that integrates low-, mid-, and high-level elements for visual inference (equivalent to image captioning today). * (24:51) On the topic of video understanding, Jason worked on ISTARE: Intelligent Spatio-Temporal Activity Reasoning Engine, whose objective is to represent, learn, recognize, and reason over activities in persistent surveillance videos. * (27:46) Jason dissected Action Bank - a high-level representation of activity in video, which comprises of many individual action detectors sampled broadly in semantic space and viewpoint space. * (35:30) Jason unpacked LIBSVX - a library of super voxel and video segmentation methods coupled with a principled evaluation benchmark based on quantitative 3D criteria for good super voxels. * (40:06) Jason gave an overview of AI research activities at the University of Michigan, where he was a professor of Electrical Engineering and Computer Science from 2014 to 2020. * (41:09) Jason covered the problems and projects in his graduate-level courses on Foundations of Computer Vision and Advanced Topics in Computer Vision at Michigan. * (44:56) Jason went over his recent research on video captioning and video description. * (47:03) Jason described his exciting software called BubbleNets, which chooses the best video frame for a human to annotate. * (51:44) Jason shared anecdotes of Voxel51's inception and key takeaways that he has learned. * (01:05:25) Jason talked about Voxel51's Physical Distancing Index that tracks the coronavirus global pandemic's impact on social behavior. * (01:07:47) Jason discussed his exciting new chapter as the new director of the Stevens Institute for Artificial Intelligence. * (01:11:28) Jason identified the differences and similarities between being a professor and being a founder. * (01:14:55) Jason gave his advice to individuals who want to make a dent in AI research. * (01:16:14) Jason mentioned the trends in computer vision research that he is most excited about at the moment. * (01:17:23) Closing segment.
His Contact Info* Wikipedia * Google Scholar * Website * Twitter * LinkedIn
His Recommended Resources* Bubblenets: Video Object Segmentation for Computer Vision * Voxel51's FiftyOne Open-Sourced Library * Jeff Siskind (Professor at Purdue University) * CJ Taylor (Professor at the University of Pennsylvania) * Kristen Grauman (Professor at the University of Austin) * "An Introduction to Mathematical Statistics"
Show Notes* (2:23) Barr discussed growing up in Israel and serving as a commander of the Data Analytics unit at the Israeli Air Force. * (4:10) Barr reflected on her college experience at Stanford studying Math and Computational Science. * (7:24) Barr walked over the two career lessons learned from being a Management Consultant at Bain and Company. * (9:51) Barr reflected on her time as VP of Customer Operations at Gainsight, which offers enterprise solutions for Customer Success and Product teams. She helped build and scale a global team covering various functions such as business operations, customer success, professional services. * (12:32) Barr unpacked the notion of data downtime, introduced in her blog post “The Rise of Data Downtime.” * (17:25) Barr unveiled the four main steps in the data reliability maturity curve: reactive, proactive, automated, and scalable - as indicated in “Closing The Data Downtime Gap." * (21:09) Barr shared the founding story behind Monte Carlo, whose mission is to accelerate the world’s adoption of data by reducing data downtime. * (24:29) Barr explained the five pillars of data observability. * (27:45) Barr unpacked the rise of data catalogs as a powerful tool for data governance, along with the three categories of data catalog solutions that data teams are adopting - as presented in “What We Got Wrong About Data Governance.” * (31:32) Barr discussed the benefits of using Data Mesh - a type of data platform architecture that embraces data ubiquity in the enterprise by leveraging a domain-oriented, self-serve design. * (37:28) Barr went over a framework that looks at the business functions and the nature of the work to score the impact and allocate the ROI of the data team - as proposed in "Measuring the ROI of Your Data Organization." * (40:39) Barr shared five practices for designing a platform that maximizes data's value and impact inside an organization. * (43:27) Barr reflected on the key decisions to get the first cohort of customers for Monte Carlo. * (46:31) Barr shared valuable hiring lessons. * (48:48) Barr went over helpful resources throughout her journey as a founder. * (50:13) Barr dropped the final advice for founders on seeking the right investors. * (51:18) Closing segment.
Her Contact Info* Twitter * LinkedIn * Medium * Monte Carlo
Her Recommended Resources* Stanford's "Mathematics and Magic Tricks" course (taught by Persi Diaconis) * "The Biggest Bluff" by Maria Konnikova * Snowflake * DJ Patil (Former U.S. Chief Data Scientist) * Monte Carlo's Customers
Show Notes* (1:57) Carl recalled his undergraduate experience studying Electrical Engineering at Stanford back in the early 90s. * (3:58) Carl recalled his graduate experience pursuing Master’s degrees in Computer Science at NYU and King’s College in the late 90s. For his Master's Thesis, he investigated Support Vector Machines with a Bayesian algorithm programmed in C. * (6:45) Carl walked over his Ph.D. work in Computation and Neural Systems at CalTech, where he did a thesis on Biophysics of Extracellular Action Potentials. * (13:11) Carl provided brief thoughts about his experience working as a business analyst and consultant for HBO during his Ph.D. period. * (14:55) Carl went over his rationale behind his decision to move from academic neuroscience to quantitative finance. * (19:19) Carl discussed his proudest accomplishments and valuable lessons learned from spending seven years at Morgan Stanley Capital International and rising to a leadership role as Vice President of Risk Modeling. * (23:17) Carl uncovered his move to San Francisco to work as a lead data scientist at Sparked back in 2014, which builds a customer success SaaS solution. * (27:10) Adding to his move to Zuora in 2015, Carl explained how the subscription business model works in layman terms. * (31:44) Carl unpacked the common patterns that he saw from analyzing subscriber churn for companies across industries due to his work on Zuora Analytics. * (33:30) Carl shared the process of creating the Subscription Economy Index, Zuora’s landmark index tracking the rapid ascent of the Subscription Economy, and distilled the key trends of the 2020 edition. * (39:59) Carl unpacked the three reasons that make churn hard to fight: (1) Churn is hard to predict, (2) Churn is harder to prevent, and (3) Churn requires a multi-team effort (Watch his talks at the 2019 Data Council San Francisco and the 2020 Subscribed Online Conference). * (44:46) Carl shared advice for data scientists who want to collaborate more effectively with other functional departments. * (46:30) Carl emphasized the importance of creating great customer metrics, which are ratios of basic behavioral metrics to fight churn effectively. * (53:49) Carl went over the challenges of writing “Fighting Churn With Data,” which provides a clear overview of churn concepts, along with hands-on tricks and tips developed through years of experience analyzing customer behavior. * (55:53) Carl reflected on how his academic background in computational neuroscience contributes to his success as a quant analyst and a data scientist. * (59:37) Carl compared his experience living and working across Los Angeles, New York, and San Francisco. * (01:02:04) Closing segment.
His Contact Info* LinkedIn * Twitter * GitHub * Google Scholar * Medium
His Recommended Resources* “Fighting Churn With Data” by Carl Gold * Konrad Kording (Professor of Computational Neuroscience at the University of Pennsylvania) * Kate Crawford (Distinguished Research Professor in Tech, Culture, and Society at New York University) * Cassie Kozyrkov (Chief Decision Scientist at Google) * "Freakonomics" by Stephen Dubner and Stephen Levitt * Carl's other podcast appearances
Here are the discount codes that you can use to purchase "Fighting Churn with Data" with 40% off:
Show Notes
Her Contact Info
Her Recommended Resources
Show Notes
His Contact Info
His Recommended Resources
Here are the codes for free eBook copies of Luis' book "Grokking Machine Learning": gmldcr-D659, gmldcr-2512, gmldcr-0752, gmldcr-30A2, gmldcr-01E8. Additionally, use the code poddcast19 to receive a 40% discount of all Manning products!
Show Notes
His Contact Info
His Recommended Resources
Use the codes below to get a discount from Frank's live video course on Manning called "Machine Learning, Data Science and Deep Learning with Python":
Show Notes
Her Contact Info
Her Recommended Resources
Show Notes
Her Contact Info
Her Recommended Resources
Show Notes
Her Contact Info
Her Recommended Resources
People To Follow
Book To Read
A Developer’s Introduction to Data Science
Azure Machine Learning
Responsible Machine Learning
Automated Machine Learning
Show Notes
Her Contact Info
Her Recommended Resources
Show Notes
His Contact Info
His Recommended Resources
Show Notes
His Contact Info
His Recommended Resources
Show Notes
His Contact Information
His Recommended Resources
Serverless Machine Learning In Action
Show Notes
His Contact Info
His Recommended Resources
Show Notes
His Contact Info
His Recommended Resources
A New Course From Luigi
Luigi just launched his first online course, Build, Deploy, and Monitor Machine Learning Models with Amazon SageMaker! I had a look at the course content, and I’m convinced that the course will be super valuable to any ML engineer or data scientist who wants to level up and learn how to productionize their machine learning models.
You can take the course on your own, but Luigi is also teaming up with TWiML to offer a version with virtual Study Group sessions for people who want a more interactive experience. Right now, Luigi is offering an early bird discount on the course until August 1st!
I know the course will be precious for a lot of you within my community, so Luigi created a coupon code DATACAST to save an additional 10% off the course!
Head over to the Teachable course page to learn more about AmazonSageMaker and take advantage of the discount!
Show Notes
His Contact Info
His Recommended Resources
Machine Learning Bookcamp
Show Notes
His Contact Info
His Recommended Resources
Show Notes
His Contact Information
His Recommended Resources
Show Notes:
His Contact Information:
His Recommended Resources:
Show Notes:
Her Contact Information:
Her Recommended Resources :
Show Notes:
Her Contact Info:
Her Recommended Resources:
Show Notes:
Her Contact Info:
Her Recommended Resources:
Show Notes:
His Contact Info:
His Recommended Resources:
You can read the completed chapters of "Data Science Bookcamp" using the codes below:
Show Notes:
His Contact Info:
His Recommended Resources:
Show Notes:
His Contact Info:
His Recommended Resources:
Show Notes:
His Contact Info:
His Recommended Resources:
Show Notes:
His Contact Info:
His Recommended Resources:
Show Notes:
His Contact Info:
His Recommended Resources:
Show Notes:
His Contact Info:
His Recommended Resources:
Show Notes:
His Contact Info:
His Recommended Resources:
Show Notes:
His Contact Info:
His Recommended Resources:
Show Notes:
Her Contact Info:
Her Recommended Resources:
Show Notes:
His Contact Info:
His Recommended Resources:
Show Notes:
His Contact Info:
His Recommended Resources:
Show Notes:
Her Contact Info:
Her Recommended Resources:
Show Notes:
His Contact Info:
His Recommended Resources:
Show Notes:
His Contact Info:
His Recommended Resources:
Show Notes:
His Contact Info:
His Recommended Resources:
Show Notes:
Her Contact Info:
Her Recommended Resources:
Show Notes:
His Contact Info:
His Recommended Resources:
Show Notes:
His Contact Info:
His Recommended Resources:
Show Notes:
His Contact Info:
His Recommended Resources:
Show Notes:
His Contact Info:
His Recommended Resources:
Show Notes:
His Contact Info:
His Recommended Resources:
Show Notes:
His Contact Info:
His Recommended Resources:
Show Notes:
His Contact Info:
His recommended resources:
Show Notes:
His Contact Info:
His recommended resources:
Show Notes:
His Contact Info:
His recommended resources:
Show Notes:
Her Contact Info:
Her recommended resources:
Show Notes:
His contact info:
His recommended resources:
Show Notes:
His contact info:
His recommended resources: