O'Reilly Media spreads the knowledge of innovators. At O’Reilly, a big part of our business is paying attention to what’s new and interesting in the world of technology. We have a pretty good record at having anticipated some of the big technology developments in recent history. For instance, we launched the first commercial Web site, GNN, in 1993; we organized the meeting at which the term “open source” was first adopted; we were early investors in Blogger, which helped launch the blogging revolution; and more recently, our Web 2.0 conference launched a world-wide meme. We call this predictive sense the “O’Reilly Radar.” And while we’re certainly not always right, we are, at least, good at making interesting guesses.
Our methodology is simple: we draw from the wisdom of the alpha geeks in our midst, paying attention to what’s interesting to them, amplifying these weak signals, and seeing where they fit into the innovation ecology. Add to that the original research conducted by our Research team, and you start to get a good picture of what the technology world is thinking about.
In this episode of the Data Show, I speak with Peter Bailis, founder and CEO of Sisu, a startup that is using machine learning to improve operational analytics. Bailis is also an assistant professor of computer science at Stanford University, where he conducts research into data-intensive systems and where he is co-founder of the DAWN Lab.
In this episode of the Data Show, I speak with Arun Kejariwal of Facebook and Ira Cohen of Anodot (full disclosure: I’m an advisor to Anodot). This conversation stemmed from a recent online panel discussion we did, where we discussed time series data, and, specifically, anomaly detection and forecasting. Both Kejariwal (at Machine Zone, Twitter, and Facebook) and Cohen (at HP and Anodot) have extensive experience building analytic and machine learning solutions at large scale, and both have worked extensively with time-series data. The growing interest in AI and machine learning has not been confined to computer vision, speech technologies, or text. In the enterprise, there is strong interest in using similar automation tools for temporal data and time series.
In this episode of the Data Show, I speak with Michael Mahoney, a member of RISELab, the International Computer Science Institute, and the Department of Statistics at UC Berkeley. A physicist by training, Mahoney has been at the forefront of many important problems in large-scale data analysis. On the theoretical side, his works spans algorithmic and statistical methods for matrices, graphs, regression, optimization, and related problems. On the applications side, he has contributed to systems used for internet and social media analysis, social network analysis, as well as for a host of applications in the physical and life sciences. Most recently, he has been working on deep neural networks, specifically developing theoretical methods and practical diagnostic tools that should be helpful to practitioners who use deep learning.
In this episode of the Data Show, I speak with Kesha Williams, technical instructor at A Cloud Guru, a training company focused on cloud computing. As a full stack web developer, Williams became intrigued by machine learning and started teaching herself the ML tools on Amazon Web Services. Fast forward to today, Williams has built some well-regarded Alexa skills, mastered ML services on AWS, and has now firmly added machine learning to her developer toolkit.
In this episode of the Data Show, I speak with Alex Ratner, project lead for Stanford’s Snorkel open source project; Ratner also recently garnered a faculty position at the University of Washington and is currently working on a company supporting and extending the Snorkel project. Snorkel is a framework for building and managing training data. Based on our survey from earlier this year, labeled data remains a key bottleneck for organizations building machine learning applications and services.
Ratner was a guest on the podcast a little over two years ago when Snorkel was a relatively new project. Since then, Snorkel has added more features, expanded into computer vision use cases, and now boasts many users, including Google, Intel, IBM, and other organizations. Along with his thesis advisor professor Chris Ré of Stanford, Ratner and his collaborators have long championed the importance of building tools aimed squarely at helping teams build and manage training data. With today’s release of Snorkel version 0.9, we are a step closer to having a framework that enables the programmatic creation of training data sets.
In this interview, Tim Craig and fellow Googler Gustavo Franco, a site reliability engineer (SRE), discuss the wide range of events that qualify as “incidents;” the need for a conscious, robust, and well-defined process for understanding them; the role of training; and how to get buy-in from management so you can spread incident response training throughout an organization.
In this episode of the Data Show, I speak with Cassie Kozyrkov, technical director and chief decision scientist at Google Cloud. She describes "decision intelligence" as an interdisciplinary field concerned with all aspects of decision-making, and which combines data science with the behavioral sciences. Most recently she has been focused on developing best practices that can help practitioners make safe, effective use of AI and data. Kozyrkov uses her platform to help data scientists develop skills that will enable them to connect data and AI with their organizations' core businesses.
We had a great conversation spanning many topics, including:
How data science can be more useful
The importance of the human side of data
The leadership talent shortage in data science
Is data science a bubble?
In this episode of the Data Show, I spoke with Roger Chen, co-founder and CEO of Computable Labs, a startup focused on building tools for the creation of data networks and data exchanges. Chen has also served as co-chair of O'Reilly's Artificial Intelligence Conference since its inception in 2016. This conversation took place the day after Chen and his collaborators released an interesting new white paper, "Fair value and decentralized governance of data." Current-generation AI and machine learning technologies rely on large amounts of data, and to the extent they can use their large user bases to create “data silos,” large companies in large countries (like the U.S. and China) enjoy a competitive advantage. With that said, we are awash in articles about the dangers posed by these data silos. Privacy and security, disinformation, bias, and a lack of transparency and control are just some of the issues that have plagued the perceived owners of “data monopolies.”
In this week's episode of the Data Show, we're featuring an interview Data Show host Ben Lorica participated in for the Software Engineering Daily Podcast, where he was interviewed by Jeff Meyerson. Their conversation mainly centered around data engineering, data architecture and infrastructure, and machine learning (ML).
In this episode of the Data Show, I spoke with Nick Pentreath, principal engineer at IBM. Pentreath was an early and avid user of Apache Spark, and he subsequently became a Spark committer and PMC member. Most recently his focus has been on machine learning, particularly deep learning, and he is part of a group within IBM focused on building open source tools that enable end-to-end machine learning pipelines.
At Google’s 2019 Cloud Next conference, I sat down with Stephen Thorne, site reliability engineer on Google’s customer reliability engineering team and co-author of "The Site Reliability Workbook," to talk about how organizations, both large and small, can use SRE to reduce operational costs, improve reliability, and create productive cross-functional teams.
In this episode of the Data Show, I spoke with Dhruba Borthakur (co-founder and CTO) and Shruti Bhat (SVP of Marketing) of Rockset, a startup focused on building solutions for interactive data science and live applications. Borthakur was the founding engineer of HDFS and creator of RocksDB, while Bhat is an experienced product and marketing executive focused on enterprise software and data products. Their new startup is focused on a few trends I’ve recently been thinking about, including the re-emergence of real-time analytics, and the hunger for simpler data architectures and tools. Borthakur exemplifies the need for companies to continually evaluate new technologies: while he was the founding engineer for HDFS, these days he mostly works with object stores like S3.
In this episode of the Data Show, I spoke with Jike Chong, chief data scientist at Acorns, a startup focused on building tools for micro-investing. Chong has extensive experience using analytics and machine learning in financial services, and he has experience building data science teams in the U.S. and in China.
We had a great conversation spanning many topics, including:
Potential applications of data science in financial services.
The current state of data science in financial services in both the U.S. and China.
His experience recruiting, training, and managing data science teams in both the U.S. and China.
In this episode of the Data Show, I spoke with Jeff Jonas, CEO, founder and chief scientist of Senzing, a startup focused on making real-time entity resolution technologies broadly accessible. He was previously a fellow and chief scientist of context computing at IBM. Entity resolution (ER) refers to techniques and tools for identifying and linking manifestations of the same entity/object/individual. Ironically, ER itself has many different names (e.g., record linkage, duplicate detection, object consolidation/reconciliation, etc.).
ER is an essential first step in many domains, including marketing (cleaning up databases), law enforcement (background checks and counterterrorism), and financial services and investing. Knowing exactly who your customers are is an important task for security, fraud detection, marketing, and personalization. The proliferation of data sources and services has made ER very challenging in the internet age. In addition, many applications now increasingly require near real-time entity resolution.
We had a great conversation spanning many topics including:
Why ER is interesting and challenging
How ER technologies have evolved over the years
How Senzing is working to democratize ER by making real-time AI technologies accessible to developers
Some early use cases for Senzing’s technologies
Some items on their research agenda
In this episode of the Data Show, I spoke with Neelesh Salian, software engineer at Stitch Fix, a company that combines machine learning and human expertise to personalize shopping. As companies integrate machine learning into their products and systems, there are important foundational technologies that come into play. This shouldn’t come as a shock, as current machine learning and AI technologies require large amounts of data—specifically, labeled data for training models. There are also many other considerations—including security, privacy, reliability/safety—that are encouraging companies to invest in a suite of data technologies. In conversations with data engineers, data scientists, and AI researchers, the need for solutions that can help track data lineage and provenance keeps popping up.
There are several San Francisco Bay Area companies that have embarked on building data lineage systems—including Salian and his colleagues at Stitch Fix. I wanted to find out how they arrived at the decision to build such a system and what capabilities they are building into it.
In this episode of the Data Show, I spoke with Avner Braverman, co-founder and CEO of Binaris, a startup that aims to bring serverless to web-scale and enterprise applications. This conversation took place shortly after the release of a seminal paper from UC Berkeley (“Cloud Programming Simplified: A Berkeley View on Serverless Computing”), and this paper seeded a lot of our conversation during this episode.
In this episode of the Data Show, I spoke with Forough Poursabzi-Sangdeh, a postdoctoral researcher at Microsoft Research New York City. Poursabzi works in the interdisciplinary area of interpretable and interactive machine learning. As models and algorithms become more widespread, many important considerations are becoming active research areas: fairness and bias, safety and reliability, security and privacy, and Poursabzi’s area of focus—explainability and interpretability.
In this episode of the Data Show, I spoke with Kartik Hosanagar, professor of technology and digital business, and professor of marketing at The Wharton School of the University of Pennsylvania. Hosanagar is also the author of a newly released book, "A Human’s Guide to Machine Intelligence," an interesting tour through the recent evolution of AI applications, which draws from his extensive experience at the intersection of business and technology.
In this episode of the Data Show, I spoke with P.W. Singer, strategist and senior fellow at the New America Foundation, and a contributing editor at Popular Science. He is co-author of an excellent new book, LikeWar: The Weaponization of Social Media, which explores how social media has changed war, politics, and business. The book is essential reading for anyone interested in how social media has become an important new battlefield in a diverse set of domains and settings.
In this episode of the Data Show, I spoke with Siwei Lyu, associate professor of computer science at the University at Albany, State University of New York. Lyu is a leading expert in digital media forensics, a field of research into tools and techniques for analyzing the authenticity of media files. Over the past year, there have been many stories written about the rise of tools for creating fake media (mainly images, video, audio files). Researchers in digital image forensics haven’t exactly been standing still, though. As Lyu notes, advances in machine learning and deep learning have also found a receptive audience among the forensics community.
In this episode of the Data Show, I spoke with Maryam Jahanshahi, research scientist at TapRecruit, a startup that uses machine learning and analytics to help companies recruit more effectively. In an upcoming survey, we found that a “skills gap” or “lack of skilled people” was one of the main bottlenecks holding back adoption of AI technologies. Many companies are exploring a variety of internal and external programs to train staff on new tools and processes. The other route is to hire new talent. But recent reports suggest that demand for data professionals is strong and competition for experienced talent is fierce. Jahanshahi and her team are building natural language and statistical tools that can help companies improve their ability to attract and retain talent across many key areas.
In this episode of the Data Show, I spoke with Andrew Burt, chief privacy officer and legal engineer at Immuta, a company building data management tools tuned for data science. Burt and cybersecurity pioneer Daniel Geer recently released a must-read white paper (“Flat Light”) that provides a great framework for how to think about information security in the age of big data and AI. They list important changes to the information landscape and offer suggestions on how to alleviate some of the new risks introduced by the rise of machine learning and AI.
We discussed their new white paper, cybersecurity (Burt was previously a special advisor at the FBI), and an exciting new Strata Data tutorial that Burt will be co-teaching in March.
In this episode of the Data Show, I spoke with Haoyuan Li, CEO and founder of Alluxio, a startup commercializing the open source project with the same name (full disclosure: I’m an advisor to Alluxio). Our discussion focuses on the state of Alluxio (the open source project that has roots in UC Berkeley’s AMPLab), specifically emerging use cases here and in China. Given the large-scale use in China, I also wanted to get Li’s take on the state of data and AI technologies in Beijing and other parts of China.
For the end-of-year holiday episode of the Data Show, I turned the tables on Data Show host Ben Lorica to talk about trends in big data, machine learning, and AI, and what to look for in 2019. Lorica also showcased some highlights from our upcoming Strata Data and Artificial Intelligence conferences.
In this episode of the Data Show, I spoke with Alex Wong, associate professor at the University of Waterloo, and co-founder of DarwinAI, a startup that uses AI to address foundational challenges with deep learning in the enterprise. As the use of machine learning and analytics become more widespread, we’re beginning to see tools that enable data scientists and data engineers to scale and tackle many more problems and maintain more systems. This includes automation tools for the many stages involved in data science, including data preparation, feature engineering, model selection, and hyperparameter tuning, as well as tools for data engineering and data operations.
Wong and his collaborators are building solutions for enterprises, including tools for generating efficient neural networks and for the performance analysis of networks deployed to edge devices.
In this episode of the Data Show, I spoke with Vitaly Gordon, VP of data science and engineering at Salesforce. As the use of machine learning becomes more widespread, we need tools that will allow data scientists to scale so they can tackle many more problems and help many more people. We need automation tools for the many stages involved in data science, including data preparation, feature engineering, model selection and hyperparameter tuning, as well as monitoring.
I wanted the perspective of someone who is already faced with having to support many models in production. The proliferation of models is still a theoretical consideration for many data science teams, but Gordon and his colleagues at Salesforce already support hundreds of thousands of customers who need custom models built on custom data. They recently took their learnings public and open sourced TransmogrifAI, a library for automated machine learning for structured data, which sits on top of Apache Spark.
In this episode of the Data Show, I spoke with Francesca Lazzeri, an AI and machine learning scientist at Microsoft, and her colleague Jaya Mathew, a senior data scientist at Microsoft. We conducted a couple of surveys this year—“How Companies Are Putting AI to Work Through Deep Learning” and “The State of Machine Learning Adoption in the Enterprise”—and we found that while many companies are still in the early stages of machine learning adoption, there’s considerable interest in moving forward with projects in the near future. Lazzeri and Mathew spend a considerable amount of time interacting with companies that are beginning to use machine learning and have experiences that span many different industries and applications. I wanted to learn some of the processes and tools they use when they assist companies in beginning their machine learning journeys.
In this episode of the Data Show, I spoke with Alon Kaufman, CEO and co-founder of Duality Technologies, a startup building tools that will allow companies to apply analytics and machine learning to encrypted data. In a recent talk, I described the importance of data, various methods for estimating the value of data, and emerging tools for incentivizing data sharing across organizations. As I noted, the main motivation for improving data liquidity is the growing importance of machine learning. We’re all familiar with the importance of data security and privacy, but probably not as many people are aware of the emerging set of tools at the intersection of machine learning and security. Kaufman and his stellar roster of co-founders are doing some of the most interesting work in this area.
In this episode of the Data Show, I spoke with Jacob Ward, a Berggruen Fellow at Stanford University. Ward has an extensive background in journalism, mainly covering topics in science and technology, at National Geographic, Al Jazeera, Discovery Channel, BBC, Popular Science, and many other outlets. Most recently, he’s become interested in the interplay between research in psychology, decision-making, and AI systems. He’s in the process of writing a book on these topics, and was gracious enough to give an informal preview by way of this podcast conversation.
In this episode of the Data Show, I spoke with Sharad Goel, assistant professor at Stanford, and his student Sam Corbett-Davies. They recently wrote a survey paper, “A Critical Review of Fair Machine Learning,” where they carefully examined the standard statistical tools used to check for fairness in machine learning models. It turns out that each of the standard approaches (anti-classification, classification parity, and calibration) has limitations, and their paper is a must-read tour through recent research in designing fair algorithms. We talked about their key findings, and, most importantly, I pressed them to list a few best practices that analysts and industrial data scientists might want to consider.
This episode of the O’Reilly Podcast, features a conversation on serverless and Kubernetes, with Kelsey Hightower, developer advocate for Google Cloud Platform at Google (and co-author of "Kubernetes: Up and Running"), and Chris Gaun, Kubernetes product marketing manager at Mesosphere.
In this episode of the Data Show, I spoke with Alan Nichol, co-founder and CTO of Rasa, a startup that builds open source tools to help developers and product teams build conversational applications. About 18 months ago, there was tremendous excitement and hype surrounding chatbots, and while things have quieted lately, companies and developers continue to refine and define tools for building conversational applications. We spoke about the current state of chatbots, specifically about the types of applications developers are building today and how he sees conversational applications evolving in the near future.
In this episode of the Data Show, I spoke with Eric Jonas, a postdoc in the new Berkeley Center for Computational Imaging. Jonas is also affiliated with UC Berkeley’s RISE Lab. It was at a RISE Lab event that he first announced Pywren, a framework that lets data enthusiasts proficient with Python run existing code at massive scale on Amazon Web Services. Jonas and his collaborators are working on a related project, NumPyWren, a system for linear algebra built on a serverless architecture. Their hope is that by lowering the barrier to large-scale (scientific) computation, we will see many more experiments and research projects from communities that have been unable to easily marshal massive compute resources. We talked about Bayesian machine learning, scientific computation, reinforcement learning, and his stint as an entrepreneur in the enterprise software space.
In this episode of the O’Reilly Media Podcast, Rachel Roumeliotis, VP of content strategy at O’Reilly, sat down with Daniel Krook, IBM developer advocate. They discussed how developers across industries can participate in the Call for Code initiative, the benefits of the program, support from its charitable partners—United Nations Human Rights and the American Red Cross—and the positive impacts IBM hopes to achieve by investing in Call for Code.
In this episode of the Data Show, I spoke with Harish Doddi, co-founder and CEO of Datatron, a startup focused on helping companies deploy and manage machine learning models. As companies move from machine learning prototypes to products and services, tools and best practices for productionizing and managing models are just starting to emerge. Today’s data science and data engineering teams work with a variety of machine learning libraries, data ingestion, and data storage technologies. Risk and compliance considerations mean that the ability to reproduce machine learning workflows is essential to meet audits in certain application domains. And as data science and data engineering teams continue to expand, tools need to enable and facilitate collaboration.
As someone who specializes in helping teams turn machine learning prototypes into production-ready services, I wanted to hear what Doddi has learned while working with organizations that aspire to “become machine learning companies.”
In this episode of the Data Show, I spoke with Chang Liu, applied research scientist at Georgian Partners. In a previous post, I highlighted early tools for privacy-preserving analytics, both for improving decision-making (business intelligence and analytics) and for enabling automation (machine learning). One of the tools I mentioned is an open source project for SQL-based analysis that adheres to state-of-the-art differential privacy (a formal guarantee that provides robust privacy assurances). Since business intelligence typically relies on SQL databases, this open source project is something many companies can already benefit from today.
What about machine learning? While I didn’t have space to point this out in my previous post, differential privacy has been an area of interest to many machine learning researchers. Most practicing data scientists aren’t aware of the research results, and popular data science tools haven’t incorporated differential privacy in meaningful ways (if at all). But things will change over the next months. For example, Liu wants to make ideas from differential privacy accessible to industrial data scientists, and she is part of a team building tools to make this happen.
In this episode of the Data Show, I spoke with Andrew Feldman, founder and CEO of Cerebras Systems, a startup in the blossoming area of specialized hardware for machine learning. Since the release of AlexNet in 2012, we have seen an explosion in activity in machine learning, particularly in deep learning. A lot of the work to date happened primarily on general purpose hardware (CPU, GPU). But now that we’re six years into the resurgence in interest in machine learning and AI, these new workloads have attracted technologists and entrepreneurs who are building specialized hardware for both model training and inference, in the data center or on edge devices.
In fact, companies with enough volume have already begun building specialized processors for machine learning. But you have to either use specific cloud computing platforms or work at specific companies to have access to such hardware. A new wave of startups (including Cerebras) will make specialized hardware affordable and broadly available. Over the next 12-24 months architects and engineers will need to revisit their infrastructure and decide between general purpose or specialized hardware, and cloud or on-premise gear.
ARTIFICIAL INTELLIGENCE CONFERENCE
The Artificial Intelligence Conference in San Francisco, September 4-7, 2018 Early price ends July 20. In light of the training duration and cost they face using current (general purpose) hardware, some experiments might be hard to justify. Upcoming specialized hardware will enable data scientists to try out ideas that they previously would have hesitated to pursue. This will surely lead to more research papers and interesting products as data scientists are able to run many more experiments (on even bigger models) and iterate faster.
As founder of one of the most anticipated hardware startups in the deep learning space, I wanted get Feldman’s views on the challenges and opportunities faced by engineers and entrepreneurs building hardware for machine learning workloads.
In a recent episode of the O’Reilly Media Podcast, we spoke with George Miranda about the importance of service mesh technology in creating reliable distributed systems. As discussed in the new report The Service Mesh: Resilient Service-to-Service Communication for Cloud Applications, service mesh technology has emerged as a popular tool for companies looking to build cloud-native applications that are reliable and secure.
During the podcast, we discussed the problems a service mesh infrastructure solves and the service mesh features you’ll find most valuable. We also talked about how to choose the right service mesh for your organization, the challenges involved in getting it deployed to production, and the best ways for getting started with a service mesh.
In this episode of the Data Show, I spoke with Aurélie Pols of Mind Your Privacy, one of my go-to resources when it comes to data privacy and data ethics. This interview took place at Strata Data London, a couple of days before the EU General Data Protection Regulation (GDPR) took effect. I wanted her perspective on this landmark regulation, as well as her take on trends in data privacy and growing interest in ethics among data professionals.
In this episode of the Data Show, I spoke with Andrew Burt, chief privacy officer at Immuta, and Steven Touw, co-founder and CTO of Immuta. Burt recently co-authored an upcoming white paper on managing risk in machine learning models, and I wanted to sit down with them to discuss some of the proposals they put forward to organizations that are deploying machine learning.
Some high-profile examples of models gone awry have raised awareness among companies for the need for better risk management tools and processes. There is now a growing interest in ethics among data scientists, specifically in tools for monitoring bias in machine learning models. In a previous post, I listed some of the key considerations organization should keep in mind as they move models to production, but the upcoming report co-authored by Burt goes far beyond and recommends lines of defense, including a description of key roles that are needed.
In this episode of the Data Show, I spoke with Ashok Srivastava, senior vice president and chief data officer at Intuit. He has a strong science and engineering background, combined with years of applying machine learning and data science in industry. Prior to joining Intuit, he led the teams responsible for data and artificial intelligence products at Verizon. I wanted his perspective on a range of issues, including the role of the chief data officer, ethics in machine learning, and the emergence of AI technologies for enterprise products and applications.
In this episode of the O’Reilly Podcast, I talk with Tammy Butow, a site reliability engineer at Gremlin, and Annie Lau, a software engineering manager at Trulia, about creating a culture of learning, how experimentation is important to business, and their careers in tech.
This episode of the Data Show marks our 100th episode. This podcast stemmed out of video interviews conducted at O’Reilly’s 2014 Foo Camp. We had a collection of friends who were key members of the data science and big data communities on hand and we decided to record short conversations with them. We originally conceived of using those initial conversations to be the basis of a regular series of video interviews. The logistics of studio interviews proved too complicated, but those Foo Camp conversations got us thinking about starting a podcast, and the Data Show was born.
To mark this milestone, my colleague Paco Nathan, co-chair of Jupytercon, turned the tables on me and asked me questions about previous Data Show episodes. In particular, we examined the evolution of key topics covered in this podcast: data science and machine learning, data engineering and architecture, AI, and the impact of each of these areas on businesses and companies. I’m proud of how this show has reached so many people across the world, and I’m looking forward to sharing more conversations in the future.
In this episode of the O’Reilly Podcast, I talk with Cory Doctorow, who is a science fiction author, editor of Boing Boing, the former European director of the Electronic Frontier Foundation (EFF), and currently a special advisor for the EFF. Doctorow will be a keynote speaker at the O’Reilly Fluent Conference, July 11-14, 2018, in San Jose.
In this episode of the O'Reilly Podcast, Fluent Conference Speaker Series chair and author Kyle Simpson sat down with Brian Holt, a senior cloud developer at Microsoft. Holt will be teaching a training course, A Complete Introduction to React" and hosting a session "10 KB or bust: The delicate power of webpack and Babel" at the O'Reilly Fluent Conference in June. Simpson and Holt discuss the winding road to finding your way in the software industry, new changes to React, and optimizing the end user experience.
In this episode of the O’Reilly Media Podcast, I talk with JP Phillips, platform engineer at IBM Cloud. IBM is driving development in the container space, as shown through last year’s launch of Istio, an open cloud service that allows developers to connect, manage, and secure networks of different microservices. Istio, a joint collaboration between IBM, Google, and Lyft, (in conjunction with the Cloud Native Computing Foundation), is platform agnostic and runs on Kubernetes platforms, such as IBM’s Cloud Container Service. Interest in the platform is growing. IBM's Daniel Berg, who is giving a talk on Istio at the upcoming OSCON conference in Portland, recently led a hands-on workshop at KubeCon in Copenhagen to help developers learn how Istio can solve common challenges with microservices deployed within Kubernetes.
In this episode of the O’Reilly Podcast, I talk with Brendan Eich, the creator of JavaScript, co-founder of the Mozilla Project and Foundation, and CEO and founder of Brave Software. Eich will be a keynote speaker at the upcoming O’Reilly Fluent Conference, July 11-14, 2018, in San Jose.
In this episode of the Data Show, I spoke with Jason Dai, CTO of Big Data Technologies at Intel, and one of my co-chairs for the AI Conference in Beijing. I wanted to check in on the status of BigDL, specifically how companies have been using this deep learning library on top of Apache Spark, and discuss some newly added features. It turns out there are quite a number of companies already using BigDL in production, and we talked about some of the popular uses cases he’s encountered. We recorded this podcast while we were at the AI Conference in Beijing, so I wanted to get Dai’s thoughts on the adoption of AI technologies among Chinese companies and local/state government agencies.
In this episode of the Data Show, I spoke with Jerry Overton, senior principal and distinguished technologist at DXC Technology. I wanted the perspective of someone who works across industries and with a variety of companies. I specifically wanted to explore the current state of data science and AI within companies and public sector agencies. As much as we talk about use cases, technologies, and algorithms, there are also important issues that practitioners like Overton need to address, including privacy, security, and ethics. Overton has long been involved in teaching and mentoring new data scientists, so we also discussed some tips and best practices he shares with new members of his team.
In this episode of the O’Reilly podcast, I spoke with Stephen Gates of Oracle Dyn. Gates joined the Oracle Dyn Global Business Unit from Zenedge, the web application security company recently acquired by Oracle. Gates and I discussed how growing malicious bot activity impacts organizations.
In this episode of the O’Reilly Podcast, I talk with Simon Moss, vice president of industry consulting and solutions, Americas, at Teradata. We discuss how machine learning and deep learning techniques are being used to fight financial crimes, such as credit card fraud, identity theft, health care fraud, and money laundering.
In this episode of the Data Show, I spoke with Guillaume Chaslot, an ex-YouTube engineer and founder of AlgoTransparency, an organization dedicated to helping the public understand the profound impact algorithms have on our lives. We live in an age when many of our interactions with companies and services are governed by algorithms. At a time when their impact continues to grow, there are many settings where these algorithms are far from transparent. There is growing awareness about the vast amounts of data companies are collecting on their users and customers, and people are starting to demand control over their data. A similar conversation is starting to happen about algorithms—users are wanting more control over what these models optimize for and an understanding of how they work.
I first came across Chaslot through a series of articles about the power and impact of YouTube on politics and society. Many of the articles I read relied on data and analysis supplied by Chaslot. We talked about his work trying to decipher how YouTube’s recommendation system works, filter bubbles, transparency in machine learning, and data privacy.
In this episode of the O’Reilly Programming Podcast, I talk with two of the program chairs for the upcoming O’Reilly Fluent Conference (July 11-14 in San Jose), Kyle Simpson and Tammy Everts. Simpson is co-author of the "HTML 5 Cookbook," and the author of the "You Don’t Know JS" series of books. Everts is the chief experience officer at SpeedCurve and the author of "Time is Money: The Business Value of Web Performance."
In this episode of the Data Show, I spoke Jesse Anderson, managing director of the Big Data Institute, and my colleague Paco Nathan, who recently became co-chair of Jupytercon. This conversation grew out of a recent email thread the three of us had on machine learning engineers, a new job role that LinkedIn recently pegged as the fastest growing job in the U.S. In our email discussion, there was some disagreement on whether such a specialized job role/title was needed in the first place. As Eric Colson pointed out in his beautiful keynote at Strata Data San Jose, when done too soon, creating specialized roles can slow down your data team.
We recorded this conversation at Strata San Jose, while Anderson was in the middle of teaching his very popular two-day training course on real-time systems. We closed the conversation with Anderson’s take on Apache Pulsar, a very impressive new messaging system that is starting to gain fans among data engineers.
In this episode of the O’Reilly Programming Podcast, I talk with Rebecca Parsons, chief technology officer at ThoughtWorks. She will be leading the workshop Building Evolutionary Architectures Hands-On at the O’Reilly Open Source Convention (OSCON), July 16-19, 2018, in Portland, Oregon. Parsons also is co-author (with Neal Ford and Patrick Kua) of the book Building Evolutionary Architectures.
In this episode of the Data Show, I spoke with Ameet Talwalkar, assistant professor of machine learning at CMU and co-founder of Determined AI. He was an early and key contributor to Spark MLlib and a member of AMPLab. Most recently, he helped conceive and organize the first edition of SysML, a new academic conference at the intersection of systems and machine learning (ML).
We discussed using and deploying deep learning at scale. This is an empirical era for machine learning, and, as I noted in an earlier article, as successful as deep learning has been, our level of understanding of why it works so well is still lacking. In practice, machine learning engineers need to explore and experiment using different architectures and hyperparameters before they settle on a model that works for their specific use case. Training a single model usually involves big (labeled) data and big models; as such, exploring the space of possible model architectures and parameters can take days, weeks, or even months. Talwalkar has spent the last few years grappling with this problem as an academic researcher and as an entrepreneur. In this episode, he describes some of his related work on hyperparameter tuning, systems, and more.
In this episode of the O’Reilly Programming Podcast, I talk about Kubernetes, containers, and more with Bridget Kromhout, a principal cloud developer advocate at Microsoft, and a frequent speaker at tech conferences. She will be leading the workshop "Kubernetes 101" at the O’Reilly Velocity Conference in San Jose, June 11-14, 2018, and at the O’Reilly Open Source Convention (OSCON), July 16-19, 2018.
In this episode of the O’Reilly Podcast, I talk about issues surrounding stateful containers and services with Tobi Knaup, CTO and co-founder of Mesosphere, whose DC/OS is a platform for deploying stateful and stateless applications on any combination of infrastructure, and Gou Rao, CTO of Portworx, a provider of persistent storage for containers.
In this episode of the Data Show, I spoke with Ofer Ronen, GM of Chatbase, a startup housed within Google’s Area 120. With tools for building chatbots becoming accessible, conversational interfaces are becoming more prevalent. As Ronen highlights in our conversation, chatbots are already enabling companies to automate many routine tasks (mainly in customer interaction). We are still in the early days of chatbots, but if current trends persist, we’ll see bots deployed more widely and take on more complex tasks and interactions. Gartner recently predicted that by 2021, companies will spend more on bots and chatbots than mobile app development.
Like any other software application, as bots get deployed in real-world applications, companies will need tools to monitor their performance. For a single, simple chatbot, one can imagine developers manually monitoring log files for errors and problems. Things get harder as you scale to more bots and as the bots get increasingly more complex. As in the case of other machine learning applications, when companies start deploying many more chatbots, automated tools for monitoring and diagnostics become essential.
The good news is relevant tools are beginning to emerge. In this episode, Ronen describes a tool he helped build: Chatbase is a chatbot analytics and optimization service that leverages machine learning research and technologies developed at Google. In essence, Chatbase lets companies focus on building and deploying the best possible chatbots.
In this episode of the Data Show, I spoke with Danny Lange, VP of AI and machine learning at Unity Technologies. Lange previously led data and machine learning teams at Microsoft, Amazon, and Uber, where his teams were responsible for building data science tools used by other developers and analysts within those companies. When I first heard that he was moving to Unity, I was curious as to why he decided to join a company whose core product targets game developers.
As you’ll glean from our conversation, Unity is at the forefront of some of the most exciting, practical applications of deep learning (DL) and reinforcement learning (RL). Realistic scenery and imagery are critical for modern games. GANs and related semi-supervised techniques can ease content creation by enabling artists to produce realistic images much more quickly. In a previous post, Lange described how reinforcement learning opens up the possibility of training/learning rather than programming in game development.
Lange explains why simulation environments are going to be important tools for AI developers. We are still in the early days of machine intelligence, and I am looking forward to more tools that can democratize AI research (including future releases by Lange and his team at Unity).
Modern-day DNS goes beyond a simple internet “phone book” service to provide dynamic traffic management, flexibility, and performance. In this episode of the O’Reilly podcast, I had a chance to discuss modern-day DNS and its role in building resilient infrastructure with Gary Sloper, VP of global sales engineering at Dyn.
In this episode of the O’Reilly Programming Podcast, I talk about Jenkins 2 and Git with Brent Laster, who presents a number of live online training courses on these topics (including "Building a deployment pipeline with Jenkins 2," and "Next level Git"). Laster will also present the workshop "Power Git" at the O’Reilly Open Source Convention, July 16-19, 2018, in Portland, Oregon, and he is the author of the forthcoming O’Reilly book "Jenkins 2: Up and Running."
In this episode of the Data Show, I spoke with Leo Meyerovich, co-founder and CEO of Graphistry. Graphs have always been part of the big data revolution (think of the large graphs generated by the early social media startups). In recent months, I’ve come across companies releasing and using new tools for creating, storing, and (most importantly) analyzing large graphs. There are many problems and use cases that lend themselves naturally to graphs, and recent advances in hardware and software building blocks have made large-scale analytics possible.
Starting with his work as a graduate student at UC Berkeley, Meyerovich has pioneered the combination of hardware and software acceleration to create truly interactive environments for visualizing large amounts of data. Graphistry has built a suite of tools that enables analysts to wade through large data sets and investigate business and security incidents. The company is currently focused on the security domain—where it turns out that graph representations of data are things security analysts are quite familiar with.
In this episode of the O’Reilly Programming Podcast, I talk with Richard Warburton and Raoul-Gabriel Urma of Iteratr Learning. They are the presenters of a series of O’Reilly Learning Paths, including "Getting Started with Reactive Programming" and "Build Reactive Applications in Java 8." Warburton is the author of "Java 8 Lambdas," and Urma is the author of "Java 8 in Action."
In this episode of the Data Show, I spoke with Mark Hammond, founder and CEO of Bonsai, a startup at the forefront of developing AI systems in industrial settings. While many articles have been written about developments in computer vision, speech recognition, and autonomous vehicles, I’m particularly excited about near-term applications of AI to manufacturing, robotics, and industrial automation. In a recent post, I outlined practical applications of reinforcement learning (RL)—a type of machine learning now being used in AI systems. In particular, I described how companies like Bonsai are applying RL to manufacturing and industrial automation. As researchers explore new approaches for solving RL problems, I expect many of the first applications to be in industrial automation.
In this O’Reilly Podcast, Rachel Roumeliotis, vice president of content strategy at O’Reilly Media, speaks with Atif Kureishy, global VP of emerging practices and artificial intelligence and deep learning at Teradata, a data and analytics company. They discuss how enterprises are currently investing in artificial intelligence (AI), which industries are seeing the most impact, barriers to AI implementation, and future trends.
In this episode of the O’Reilly Programming Podcast, I talk with Paul Bakker, senior software engineer on the edge developer experience team at Netflix, and Sander Mak, a fellow at Luminis Technologies. They are the authors of the O’Reilly book "Java 9 Modularity," in which they call the introduction of the module system to the platform “the start of a new era.”
In this episode of the Data Show, I spoke with Fabian Yamaguchi, chief scientist at ShiftLeft. His 2015 Ph.D. dissertation sketched out how the combination of static analysis, graph mining, and machine learning, can be used to develop tools to augment security analysts. In a recent post, I argued for machine learning tools to augment teams responsible for deploying and managing models in production (machine learning engineers). These are part of a general trend of using machine learning to develop and manage the software systems of tomorrow. Yamaguchi’s work is step one in this direction: using machine learning to reduce the number of security vulnerabilities in complex software products.
In this episode of the O’Reilly Programming Podcast, I talk about Python with Luciano Ramalho, technical principal at ThoughtWorks, author of the O’Reilly book "Fluent Python," and presenter of the Oriole "Fluent Python: The Power of Special Methods."
In this episode of the Data Show, I spoke with Kristian Hammond, chief scientist of Narrative Science and professor of EECS at Northwestern University. He has been at the forefront of helping companies understand the power, limitations, and disruptive potential of AI technologies and tools. In a previous post on machine learning, I listed types of uses cases (a taxonomy) for machine learning that could just as well apply to enterprise applications of AI. But how do you identify good use cases to begin with?
In this episode of the O’Reilly Programming Podcast, we revisit our June 2017 conversation with Sam Newman, presenter of the O’Reilly video course The Principles of Microservices and the online training course From Monolith to Microservices. He is also the author of the book Building Microservices: Designing Fine-Grained Systems.
In this episode of the Data Show, I spoke with Tim Kraska, associate professor of computer science at MIT. To take advantage of big data, we need scalable, fast, and efficient data management systems. Database administrators and users often find themselves tasked with building index structures (“indexes” in database parlance), which are needed to speed up data access.
In this episode of the O’Reilly Programming Podcast, I talk with Wendy Wise, technical director of emerging technologies at Turner Broadcasting System, and author of the recent article “How to pick the right authoring tools for VR and AR.” She is developing Learning Paths, which will be released on Safari in 2018, on how to get started with ARKit using Unity and XCode.
In this episode of the Data Show, I spoke with Christine Hung, head of data solutions at Spotify. Prior to joining Spotify, she led data teams at the NY Times and at Apple (iTunes). Having led teams at three different companies, I wanted to hear her thoughts on digital transformation, and I wanted to know how she approaches the challenge of building, managing, and nurturing data teams.
I also wanted to learn more about what goes into building a recommender system for a popular consumer service like Spotify. Engagement should clearly be the most important metric, but there are other considerations, such as introducing users to new or “long tail” content.
In this episode of the Security Podcast, I talk with Rich Smith, director of labs at Duo Labs, the research arm of Duo Security. We discuss the goals of agile application security, how to reframe success for security teams, and the short- and long-term implications of your security culture.
In this episode of the O’Reilly Programming Podcast, I talk with Katharine Jarmul, a Python developer and data analyst whose company, Kjamistan, provides consulting and training on topics surrounding machine learning, natural language processing, and data testing. Jarmul is the co-author (along with Jacqueline Kazil) of the O’Reilly book "Data Wrangling with Python," and she has presented the live online training course "Practical Data Cleaning with Python."
In this episode of the O’Reilly Media Podcast, I spoke with Gayle Sheppard, vice president and general manager of Saffron AI Group at Intel, and David Thomas, chief analytics officer for Bank of New Zealand (BNZ). Our conversations centered around the utility of artificial intelligence in the financial services industry.
In this episode of the O’Reilly Media Podcast, I spoke with Gayle Sheppard, vice president and general manager of Saffron AI Group at Intel, and David Thomas, chief analytics officer for Bank of New Zealand (BNZ). Our conversations centered around the utility of artificial intelligence in the financial services industry.
In a recent episode of the O’Reilly Media Podcast, David Hsieh, senior vice president of marketing at Qubole, sat down with John Slocum, vice president of MediaMath’s data management platform (DMP), to discuss DataOps in the media industry. “DataOps” refers to the promotion of communication between formerly siloed data, teams, and systems. As discussed in "Creating a Data-Driven Enterprise with DataOps," a report published by Qubole and O’Reilly in 2016, DataOps leverages process change, organizational realignment, and technology to facilitate relationships between everyone who handles data: developers, data engineers, data scientists, analysts, and business users. As a programmatic advertising platform, MediaMath has a unique lens into the shifting business models across the media industry, and how DataOps is playing a role in those shifts.
During the podcast, Hsieh and Slocum discussed how data has transformed the culture and overall goals of organizations in the media industry in the past 10 years, and shared some best practices for companies that are just embarking on their journey toward becoming data driven.
In this episode of the Data Show, I spoke with Neha Narkhede, co-founder and CTO of Confluent. As I noted in a recent post on “The Age of Machine Learning,” data integration and data enrichment are non-trivial and ongoing challenges for most companies. Getting data ready for analytics—including machine learning—remains an area of focus for most companies. It turns out, “data lakes” have become staging grounds for data; more refinement usually needs to be done before data is ready for analytics. By making it easier to create and productionize data refinement pipelines on both batch and streaming data sources, analysts and data scientists can focus on analytics that can unlock value from data.
In this episode of the Security Podcast, I talk with Christie Terrill, partner at Bishop Fox. We discuss the importance of educating businesses on the complexities of “being secure,” how to approach building a strong security program, and aligning security goals with the larger processes and goals of the business.
In this episode of the O’Reilly Programming Podcast, I talk with Nathaniel Schutta, a solutions architect at Pivotal, and presenter of the video "I’m a Software Architect, Now What?." He will be giving a presentation titled "Thinking Architecturally" at the 2018 O’Reilly Software Architecture Conference, February 25-28, 2018, in New York City.
When I first discovered and started using Apache Spark, a majority of the use cases I used it for involved unstructured text. The absence of libraries meant rolling my own NLP utilities, and, in many cases, implementing a machine learning library (this was pre deep learning, and MLlib was much smaller). I’d always wondered why no one bothered to create an NLP library for Spark when many people were using Spark to process large amounts of text. The recent, early success of BigDL confirms that users like the option of having native libraries.
In this episode of the Data Show, I spoke with David Talby of Pacific.AI, a consulting company that specializes in data science, analytics, and big data. A couple of years ago I mentioned the need for an NLP library within Spark to Talby; he not only agreed, he rounded up collaborators to build such a library. They eventually carved out time to build the newly released Spark NLP library. Judging by the reception received by BigDL and the number of Spark users faced with large-scale text processing tasks, I suspect Spark NLP will be a standard tool among Spark users.
STRATA DATA CONFERENCE
Strata Data Conference in San Jose, March 5-8, 2018 Registration is now open Talby and I also discussed his work helping companies build, deploy, and monitor machine learning models. Tools and best practices for model development and deployment are just beginning to emerge—I summarized some of them in a recent post, and, in this episode, I discussed these topics with a leading practitioner.
In this episode of the Security Podcast, O’Reilly’s Mac Slocum talks with Susan Sons, senior systems analyst for the Center for Applied Cybersecurity Research (CACR) at Indiana University. They discuss how she initially got involved with fixing the open source Network Time Protocol (NTP) project, recruiting and training new people to help maintain open source projects like NTP, and how security needn’t be an impediment to organizations moving quickly.
In this episode of the O’Reilly Programming Podcast, I talk with Matt Stine, global CTO of architecture at Pivotal. He is the presenter of the O’Reilly live online training course "Cloud-Native Architecture Patterns," and he has spoken about cloud-native architecture at the recent O’Reilly Software Architecture Conference and O’Reilly Security Conference.
In this episode of the O’Reilly Podcast, I speak with Amit Vij, CEO and co-founder of Kinetica, a company that has developed an analytics database that uses graphics processing units (GPUs). We talk about how organizations are using GPU-accelerated databases to converge artificial intelligence (AI) and business intelligence (BI) on a single platform.
In this episode of the Data Show, I spoke with Rhea Liu, analyst at China Tech Insights, a new research firm that is part of Tencent’s Online Media Group. If there’s one place where AI and machine learning are discussed even more than the San Francisco Bay Area, that would be China. Each time I go to China, there are new applications that weren’t widely available just the year before. This year, it was impossible to miss bike sharing, mobile payments seemed to be accepted everywhere, and people kept pointing out nascent applications of computer vision (facial recognition) to identity management and retail (unmanned stores).
In this episode of the Security Podcast, I talk with Charles Givre, senior lead data scientist at Orbital Insight. We discuss how data science skills are increasingly important for security professionals, the critical role of data scientists in making the results of their work accessible to even nontechnical stakeholders, and using machine learning as a dynamic filter for vast amounts of data.
In this episode of the O’Reilly podcast, I speak with Han Yang, senior product manager at Cisco, working on analytics solutions. We discuss the impact of data analytics across industries, building data lifecycles for Internet of Things (IoT) applications, and advice for establishing successful analytics and machine learning projects.
In this episode of the O’Reilly Programming Podcast, I talk with Michael Nygard, a software architect at Cognitect. He has spoken about “architecture without an end state” at numerous O’Reilly Software Architecture events, and he is the author of the book "Release It! Design and Deploy Production-Ready Software."
In this episode of the Security Podcast, I talk with Andrea Limbago, chief social scientist at Endgame. We discuss how the misperception of security as a computer science skillset ultimately restricts innovation, the need to make security easier and accessible for everyone, and how current branding of security can discourage newcomers.
In this episode of the Data Show, I spoke with Bruno Fernandez-Ruiz, co-founder and CTO of Nexar. We first met when he was leading Yahoo! technical teams charged with delivering a variety of large-scale, real-time data products. His new company is helping build out critical infrastructure for the emerging transportation sector.
In this podcast episode, I speak with Gary Orenstein, chief marketing officer at MemSQL, a platform for real-time analytics that combines a database, a data warehouse, and streaming workloads into one system. We discuss trends that are driving advancements in data warehousing, how related technologies are changing as machine learning and AI evolve, and example use cases across industries.
In this podcast episode, O’Reilly’s Shannon Cutt talks with Damon Feldman, solutions director at MarkLogic, a company that has developed an operational and transactional NoSQL database that integrates data silos to give customers a single view of their data. They discuss data lakes, data hubs, integrating data in a lake format, data governance related to security, and more.
One of the paradoxes of artificial intelligence is that the companies poised to make the most of it are those with the most data: enterprises; yet, they have the most institutional and organizational difficulties to overcome to do so. In this episode of the O’Reilly podcast, I had a chance to discuss the challenges with someone whose job it is to tackle these issues—Ron Bodkin, VP and general manager of artificial intelligence at Teradata.
In this episode of the O’Reilly Programming Podcast, I talk with Mark Bates, presenter of a number of videos and Learning Paths on Go (including "Go Core Techniques and Tools" and "Go Web Framework and Techniques"), a frequent speaker at Go conferences, and an organizer for events including GopherCon and Gotham Go. Bates is also the creator of the Go web ecosystem Buffalo.
In this episode of the Security Podcast, I talk with Window Snyder, chief security officer at Fastly. We discuss the fact that many core security best practices aren’t easy to achieve with tools, the importance of not discounting user fatigue and frustration, and the need to personalize security tools and processes to your individual environment.
In this episode of the Data Show, I spoke with Carme Artigas, co-founder and CEO of Synergic Partners (a Telefonica company). As more companies adopt big data technologies and techniques, it’s useful to remember that the end goal is to extract information and insight. In fact, as with any collection of tools and technologies, the main challenge is identifying and prioritizing use cases.
In this episode of the O’Reilly Programming Podcast, I talk with Jim Blandy and Jason Orendorff, both of Mozilla, where Blandy works on Firefox’s web developer tools and Orendorff is the module owner of Firefox’s JavaScript engine. They are the authors of the new O’Reilly book "Progamming Rust."
In this podcast episode, I speak with Dave Cassel, technical community manager at MarkLogic, creator of a multi-model NoSQL database that aims to integrate data silos for a unified view. We talked about integration patterns for loading and exporting data at ease, an architecture that enables efficient search and queries, and layers of security that follow the data from its original source throughout its lifecycle.
In this episode of the Data Show, we look back to a recent conversation I had at the Spark Summit in San Francisco with Ion Stoica (UC Berkeley professor and executive chairman of Databricks) and Matei Zaharia (assistant professor at Stanford and chief technologist of Databricks). Stoica and Zaharia were core members of UC Berkeley’s AMPLab, which originated Apache Spark, Apache Mesos, and Alluxio.
In this episode of the Security Podcast, I talk with Chris Wysopal, co-founder and CTO of Veracode. We discuss the increasing role of developers in building secure software, maintaining development speed while injecting security testing, and helping developers identify when they need to contact the security team for help.
In this episode of the O’Reilly Programming Podcast, I talk with Ken Kousen, an author, instructor, and consultant who is presenting the live online training courses Functional Programming in Java 8 and Getting Started with Spring Boot in September and October. He is also the author of the newly published O’Reilly book Modern Java Recipes: Simple Solutions to Difficult Problems in Java 8 and 9.
In this episode of the Data Show, I spoke with Ken Stanley, founding member of Uber AI Labs and associate professor at the University of Central Florida. Stanley is an AI researcher and a leading pioneer in the field of neuroevolution—a method for evolving and learning neural networks through evolutionary algorithms. In a recent survey article, Stanley went through the history of neuroevolution and listed recent developments, including its applications to reinforcement learning problems.
In this episode of the Security Podcast, I talk with Scott Roberts, security operations manager at GitHub. We discuss threat intelligence, incident response, and how they interrelate.
In this episode of the O’Reilly Programming Podcast, I talk with Adam Scott, who has authored a series of ebooks on the topic of ethical web development, the most recent of which is "Collaborative Web Development." He is also the presenter of the video "Introduction to Modern Front-End Development." Scott is the web development lead at the Consumer Financial Protection Bureau, where he focuses on building open source tools.
In this episode of the Data Show, I spoke with Robert Nishihara and Philipp Moritz, graduate students at UC Berkeley and members of RISE Lab. I wanted to get an update on Ray, an open source distributed execution framework that makes it easy for machine learning engineers and data scientists to scale reinforcement learning and other related continuous learning algorithms. Many AI applications involve an agent (for example a robot or a self-driving car) interacting with an environment. In such a scenario, an agent will need to continuously learn the right course of action to take for a specific state of the environment.
In this week’s Design Podcast, I sit down with Julie Stanford, founder and principal of user experience agency Sliced Bread Design. We talk about how to get in the rapid experimentation mindset, the design thinking process, and how to get started with rapid experimentation at your company. Hint: start small.
In this episode of the Security Podcast, I talk with Jack Daniel, co-founder of Security Bsides. We discuss how each of us (and the industry as a whole) benefits from community building, the importance of historical context, and the inimitable Becky Bace.
In this episode of the O’Reilly Programming Podcast, I talk serverless architecture with Mike Roberts, engineering leader and co-founder of Symphonia, a serverless and cloud architecture consultancy. Roberts will give two presentations—Serverless Architectures: What, Why, Why Not, and Where Next? and Designing Serverless AWS Applications—at the O’Reilly Software Architecture Conference, October 16-19, 2017, in London.
In this podcast episode, I speak with Eliot Knudsen, data science lead at Tamr, a company that uses AI to integrate data across silos. Data integration is often a painstaking, highly manual process of matching fields and resolving entities, but new tools can work alongside human experts to discover patterns in data and make recommendations for automatically merging it.
In this episode of the Data Show, I spoke with Soumith Chintala, AI research engineer at Facebook. Among his many research projects, Chintala was part of the team behind DCGAN (Deep Convolutional Generative Adversarial Networks), a widely cited paper that introduced a set of neural network architectures for unsupervised learning. Our conversation centered around PyTorch, the successor to the popular Torch scientific computing framework. PyTorch is a relatively new deep learning framework that is fast becoming popular among researchers. Like Chainer, PyTorch supports dynamic computation graphs, a feature that makes it attractive to researchers and engineers who work with text and time-series.
In this episode of the Security Podcast, Courtney Nash, former chair of O’Reilly Security conference, talks with Jay Jacobs, senior data scientist at BitSight. We discuss the constraints of convenient data, the simple first steps toward building a basic security data analytics program, and effective data visualizations.
In this episode of the O’Reilly Programming Podcast, I talk with Eric Freeman and Elisabeth Robson, presenters of the live online training course "Design Patterns Boot Camp," and co-authors (with Bert Bates and Kathy Sierra) of "Head First Design Patterns," among other books. They are also co-founders of WickedlySmart, an online learning company for software developers.
In this week’s Design Podcast, I sit down with John Whalen, chief experience officer at 10 Pearls, a digital development company focused on mobile and web apps, enterprise solutions, cyber security, big data, IoT, and cloud and dev ops. We talk about the “six minds” that underlie each human experience, why it’s important for designers to understand brain science, and what people really look for in a voice assistant.
In this podcast episode, O’Reilly’s Jeff Bleiel talks with Chris Stetson, chief architect and head of engineering at NGINX. They discuss Stetson’s experiences working on microservices-based systems and how a microservices reference architecture can ease a development team’s pain when shifting from a monolithic application to many individualized microservices.
In this episode of the Data Show, I spoke with Evangelos Simoudis, co-founder of Synapse Partners and a frequent contributor to O’Reilly. He recently published a book entitled "The Big Data Opportunity in Our Driverless Future," and I wanted get his thoughts on the transportation industry and the role of big data and analytics in its future. Simoudis is an entrepreneur, and he also advises and invests in many technology startups. He became interested in the automotive industry long before the current wave of autonomous vehicle startups was in the planning stages.
In this podcast episode, I talk about reactive microservice deployments with Edward Callahan, a senior engineer with Lightbend. We discuss the difference between a normal deployment pipeline and one that’s fully reactive, as well as the impact reactive deployments have on software teams.
In this episode, O’Reilly’s Courtney Nash talks with Katie Moussouris, founder and CEO of Luta Security. They discuss why many organizations have a knee-jerk legal response to a bug report (and why your organization shouldn’t), the first steps organizations should take in formulating a vulnerability disclosure program, and how learning through experience and sharing knowledge benefits all.
In this week’s Design Podcast, I sit down with Cheryl Platz, senior designer at Microsoft for the Azure Portal and Marketplaces. We talk about the challenges of working on a top-secret design project, the research behind Amazon's Echo Look, the skills you need to start designing for voice, and how studying improv can make you a better designer.
In this episode of the O’Reilly Programming Podcast, I talk all things Python with Aaron Maxwell, presenter of the live online training courses "Python: Beyond The Basics, and Python: The Next Level." He is also the author of the book "Powerful Python: The Most Impactful Patterns, Features and Development Strategies Modern Python Provides."
My guest in this podcast episode is Stewart Rogers, director of product management at Lambda Solutions, which offers an open source e-learning management system; its customers are organizations that deliver online learning and need to understand how their students are progressing. Rogers is thus doubly dependent on embedded analytics in his products: he uses analytics to manage his products, measuring customer response to new features, but he also needs to pass meaningful analytics through to his customers in order to let them assess their learners and their content.
In this episode of the Data Show, I spoke with Grace Huang, data science lead at Pinterest. With its combination of a large social graph, enthusiastic users, and multimedia data, I’ve long regarded Pinterest as a fascinating lab for data science. Huang described the challenge of building a sustainable content ecosystem and shared lessons from the front lines of machine learning product launches. We also discussed recommenders, the emergence of deep learning as a technique used within Pinterest, and the role of data science within the company.
In this episode, I talk with Alex Pinto, chief data scientist at Niddel. We discuss the role of threat hunting in security, the necessity for well-defined process and documentation in threat hunting and other activities, and the potential for automating threat hunting using supervised machine learning.
In this episode of the Data Show, I speak with Naveen Rao, VP and GM of the Artificial Intelligence Products Group at Intel. In an earlier episode, we learned that scaling current deep learning models requires innovations in both software and hardware. Through his startup Nervana (since acquired by Intel), Rao has been at the forefront of building a next generation platform for deep learning and AI.
I wanted to get his thoughts on what the future infrastructure for machine learning would look like. At least for now, we’re seeing a variety of approaches, and many companies are using heterogeneous processors (even specialized ones) and proprietary interconnects for deep learning. Nvidia and Intel Nervana are set to release processors that excel at both training and inference, but as Rao pointed out, at large-scale there are many considerations—including utilization, power consumption, and convenience—that come into play.
In this episode of the O’Reilly Programming Podcast, I talk about microservices with Sam Newman, presenter of the O’Reilly video course "The Principles of Microservices" and the online training course "From Monolith to Microservices." He is also the author of the book "Building Microservices: Designing Fine-Grained Systems."
In this episode of the Data Show, I spoke with Michael Freedman, CTO of Timescale and professor of computer science at Princeton University. When I first heard that Freedman and his collaborators were building a time-series database, my immediate reaction was: “Don’t we have enough options already?” The early incarnation of Timescale was a startup focused on IoT, and it was while building tools for the IoT problem space that Freedman and the rest of the Timescale team came to realize that the database they needed wasn’t available (at least out in open source). Specifically, they wanted a database that could easily support complex queries and the sort of real-time applications many have come to associate with streaming platforms. Based on early reactions to TimescaleDB, many users concur.
In this episode, I talk with Amanda Berlin, security architect at Hurricane Labs. We discuss how to assess and develop defensive security policies when you’re new to the task, how to approach core security fundamentals like asset management, and generally how you can successfully improve your organization’s defensive security with limited time and resources.
In this episode of the O’Reilly Programming Podcast, I talk with Ben Evans, co-founder and technology fellow at JClarity, and co-author of the forthcoming O’Reilly book "Optimizing Java: Practical Techniques for Improved Performance Tuning." We discuss the upcoming release of Java 9, Java performance issues, and Evans’ experience as an organizer for the London Java Community.
In this episode of the Data Show, I spoke with Geoffrey Bradway, VP of engineering at Numerai, a new hedge fund that relies on contributions of external data scientists. The company hosts regular competitions where data scientists submit machine learning models for classification tasks. The most promising submissions are then added to an ensemble of models that the company uses to trade in real-world financial markets.
This week, I sit down with Cynthia Savard Saucier, director of design at Shopify and author of Tragic Design. Saucier also is keynoting at Velocity in New York, October 1-4, 2017. We talk about moving from working in design to leading designers, the real and sometimes negative impact that design decisions can have on users, and how design is organized at Shopify.
In this episode of the Data Show, I spoke with Alex Ratner, a graduate student at Stanford and a member of Christopher Ré’s Hazy research group. Training data has always been important in building machine learning algorithms, and the rise of data-hungry deep learning models has heightened the need for labeled data sets. In fact, the challenge of creating training data is ongoing for many companies; specific applications change over time, and what were gold standard data sets may no longer apply to changing situations.
In this episode of the O’Reilly Podcast, Brian Anderson talks with Larry Haig, author of the recent O'Reilly book "Frontend Optimization Handbook." They discuss the relationship between web performance and customer experience (and why it’s important), solutions for optimizing web performance, and advice for those just embarking on their web performance journey.
In this episode, I talk with Kimber Dowsett, security architect at 18F. We discuss how to prepare your organization for a vulnerability disclosure policy, the benefits of starting small, and how to apply lessons learned to build better defenses.
In this episode of the O’Reilly Programming Podcast, I talk with Jason Hibbets, senior community evangelist in corporate marketing at Red Hat, where he is a community manager for Opensource.com. We discuss how the open source model can be applied to other disciplines beyond technology, especially in the area of open government.
In this podcast episode (see Part 1 and Part 2), I speak with Andy Hickl, chief product officer at Intel Saffron Cognitive Solutions Group. Our subject: common sources of bias in artificial intelligence—and how to identify and address them by using transparent AI methods.
In this podcast episode (see Part 1 and Part 2), I speak with Andy Hickl, chief product officer at Intel Saffron Cognitive Solutions Group. Our subject: common sources of bias in artificial intelligence—and how to identify and address them by using transparent AI methods.
In this episode of the Data Show, I spoke with Jeremy Stanley, VP of data science at Instacart, a popular grocery delivery service that is expanding rapidly. As Stanley describes it, Instacart operates a four-sided marketplace comprised of retail stores, products within the stores, shoppers assigned to the stores, and customers who order from Instacart. The objective is to get fresh groceries from popular retailers delivered to customers in a timely fashion. Instacart’s goals land them in the center of the many opportunities and challenges involved in building high-impact data products.
In this episode of the O’Reilly Bots Podcast, Pete Skomoroch and I talk to Jason Laska and Michael Akilian of Clara Labs, creator of a virtual assistant—Clara—that schedules meetings and interacts in natural language through email.
E-mail is, to me, a highly promising (and somewhat underrated) venue for bots. Messaging is growing quickly, but e-mail is still the standard way to communicate within businesses and especially between businesses. E-mail conventions are somewhat standardized, and much of it is highly routinized—automatically generated reports, receipts, etc.—so it’s ripe for automation.
Laska, who leads the machine learning efforts at Clara Labs, and Akilian, the company’s co-founder and CTO, talk about the reality of developing an AI-driven product, and explain Clara’s human-in-the-loop system. “People are still there to do some of the most challenging aspects of this work, and that’s exactly what you want to use people for,” says Laska.
This week, I sit down with Travis Lowdermilk senior UX designer at Microsoft, and Jessica Rich, UX researcher at Microsoft; Lowdermilk and Rich are also co-authors of the "Customer Driven Playbook." We talk about why failing fast is not always a good approach, sensemaking, and never losing track of the customer’s voice.
In this episode, I talk with Kelly Shortridge, detection product manager at BAE Systems Applied Intelligence. We talk about how common cognitive biases apply to security roles, how decision trees can help security practitioners overcome assumptions and build more dynamic defenses, and how combining security and UX could lead to a more secure future.
In this podcast episode, Ken Krupa, enterprise CTO at MarkLogic, walks me through the challenge of data integration and outlines an approach for dealing with it known as the “operational data hub.”
The growth of the consumer Internet has left many companies awash in analytics data. In this O’Reilly Podcast episode, I talk with William Plummer, chief strategy officer at TalkingData, about how to make sense of it.
Plummer’s key message is that companies need to embrace what he calls “smart data.” Smart data is an evolution of the “big data” idea that enterprises have been pursuing for about a decade now. Plummer links smart data to the two developments that have transformed the practice of data science over the last couple of decades: the arrival of nearly universal connectivity and the advent of the consumer internet.
In this episode of the O’Reilly Programming Podcast, I talk about Swift with Paris Buttfield-Addison, co-founder of Secret Lab, a mobile development studio that builds games and apps for mobile devices. He is the co-author of "Learning Swift," and a presenter of the Learning Path "Getting Started with Swift on the iPad" and the video "Ultimate Swift Programming."
In this episode of the Bots Podcast, Chris Messina and I reflect on what Facebook has become, the role that it now plays in our lives, and what it all means for developers. We recorded this discussion shortly after attending Facebook’s F8 Developer Conference in San Jose.
In this episode of the Data Show, I spoke with David Ferrucci, founder of Elemental Cognition and senior technologist at Bridgewater Associates. Ferrucci served as principal investigator of IBM’s DeepQA project and led the Watson team that became champion of the Jeopardy! quiz show. Elemental Cognition (EC) is a research group focused on building an AI system that will be equipped with state-of-the-art natural language understanding technologies. Ferrucci envisions that EC will ship with foundational knowledge in many subject areas, but will be able to very quickly acquire knowledge in other (specialized) domains with the help of “human mentors.”
Having built and deployed several prominent AI systems through the years, I also wanted to get Ferrucci’s perspective on the evolution of AI technologies, and how enterprises can take advantage of all the exciting recent developments.
This week, I sit down with Matt LeMay, product coach and consultant, and author of "Product Management in Practice." We talk about the four guiding principles of product management, what he has learned about himself as a product manager, and how to conduct meaningful research.
In this episode, I talk with Dave Lewis, global security advocate at Akamai. We talk about how technical sprawl and employee churn compounds security debt, the tenacity of solvable security problems, and how the speed of innovation reintroduces vulnerabilities.
In this episode of the Data Show, I spoke with Lukas Biewald, co-founder and chief data scientist at CrowdFlower. In a previous episode we covered how the rise of deep learning is fueling the need for large labeled data sets and high-performance computing systems. CrowdFlower has a service that many leading companies have come to rely on to provide them with labeled data sets to train machine learning models. As deep learning models get larger and more complex, they require training data sets that are bigger than those required by other machine learning techniques.
In the first episode of our new O’Reilly Programming Podcast, I talk about software architecture and the concept of “evolutionary architecture” with Neal Ford, director, software architect, and meme wrangler at ThoughtWorks, a global IT consultancy that focuses on end-to-end software development and delivery. Ford is presenting two sessions at OSCON 2017, O’Reilly’s Open Source Convention, and he is a co-author of the forthcoming O’Reilly book "Building Evolutionary Architectures."
In this podcast episode, O’Reilly’s Jeff Bleiel talks with Jan Machacek, CTO at Cake Solutions, a global consulting firm that specializes in building reactive systems. They discuss Machacek’s experiences in building reactive microservice-based systems and how these systems can help companies keep up with changing business demands.
This week, I sit down with Nate Walkingshaw, chief experience officer of Pluralsite and co-author of Product Leadership. We talk about hard and soft leadership skills, building cross-disciplinary product teams, and why it’s important to use the layover test when hiring.
In this special episode of the Security Podcast, O’Reilly’s Ben Lorica talks with Parvez Ahammad, who leads the data science and machine learning efforts at Instart Logic. He has applied machine learning in a variety of domains, most recently to computational neuroscience and security. Lorica and Ahammad discuss the challenges of using machine learning in information security.
In this episode of the Data Show, I spoke with Reza Zadeh, adjunct professor at Stanford University, co-organizer of ScaledML, and co-founder of Matroid, a startup focused on commercial applications of deep learning and computer vision. Zadeh also is the co-author of the forthcoming book TensorFlow for Deep Learning (now in early release). Our conversation took place on the eve of the recent ScaledML conference, and much of our conversation was focused on practical and real-world strategies for scaling machine learning. In particular, we spoke about the rise of deep learning, hardware/software interfaces for machine learning, and the many commercial applications of computer vision.
In this week’s Design Podcast, I sit down with David Farkas, associate director of user experience at EPAM and co-author of the forthcoming book "UX Research." We talk about his book, why everyone should learn to conduct research, and how to open up your mind to ask the right questions.
Farkas and his co-author Brad Nunnally also are teaching a series of online courses:
Learning UX Research: Understanding Methods and Techniques—May 8 or July 10, 2017. Learning UX Research: Analyzing Data and Sharing Results—May 22 or July 24, 2017.
In this episode of the Bots Podcast, we peer into the giant companies that are beginning to adopt messaging and bots. My guest is Tom Hadfield, founder of Message.io, a service that syndicates bots across many different messaging platforms.
In this episode, I talk with Katie Moussouris, founder and CEO of Luta Security. We discuss the five stages of vulnerability disclosure grief, hacking the government, and the pros and cons of bug bounty programs.
In this episode of the Data Show, I spoke with Karthik Ramasamy, adjunct faculty member at UC Berkeley, former engineering manager at Twitter, and co-founder of Streamlio. Ramasamy managed the team that built Heron, an open source, distributed stream processing engine, compatible with Apache Storm. While Ramasamy has seen firsthand what it takes to build and deploy large-scale distributed systems (within Twitter, he worked closely with the team that built DistributedLog), he is first and foremost interested in designing and building end-to-end applications. As someone who organizes many conferences, I’m all too familiar with the vast array of popular big data frameworks available. But, I also know that engineers and architects are most interested in content and material that helps them cut through the options and decisions.
In this episode of the O’Reilly Bots Podcast, I talk about deep learning at the extremes of scale and computing power with Prabhat, who leads the data and analytics group at Lawrence Berkeley National Laboratory’s supercomputing center. If you’re working on commercial AI, it’s worth glancing across the divide at scientific AI.
In this week’s Design Podcast, I sit down with Jonathan Shariat, senior interaction designer at Intuit and co-author of the forthcoming book "Tragic Design." We talk about his new book and survey some use cases that point a spotlight on the importance of ethical standards in design.
In this episode of the Data Show, I spoke with Aurélien Géron, a serial entrepreneur, data scientist, and author of a popular, new book entitled "Hands-on Machine Learning with Scikit-Learn and TensorFlow." Géron’s book is aimed at software engineers who want to learn machine learning and start deploying machine learning models in real-world products.
In this episode, I talk with Allison Miller, product manager for secure browsing at Google and my co-host of the O’Reilly Security conference, which is returning to New York City this fall. We discuss the importance of having an event focused solely on defense, what we’re looking forward to this year, and some notable ideas and topics from the call for proposals.
In this podcast episode, I speak with Ian Fyfe, senior director for product marketing at Zoomdata, about the next generation of business intelligence software and how it addresses what Fyfe calls “the modern world of big and streaming data.”
This week, I sit down with Aman Naimat, senior vice president of technology at Demandbase, and co-founder and CTO of Spiderbook. We talk about his project to build a knowledge graph of the entire business world using natural language processing and deep learning. We also talk about the role AI is playing in those companies today and what’s going to drive AI adoption in the future.
In this episode of the Data Show, I spoke with Francisco Webber, founder of Cortical.io, a startup that is applying tools based on Hierarchical Temporal Memory (HTM) to natural language understanding. While HTM has been around for more than a decade, there aren’t many companies that have released products based on it (at least compared to other machine learning methods). Numenta, an organization developing open source machine intelligence based on the biology of the neocortex, maintains a community site featuring showcase applications. Webber’s company has been building tools based on HTM and applying them to big text data in a variety of industries; financial services has been a particularly strong vertical for Cortical.
In this podcast episode, I talk about fast data processing and reactive systems with Karl Wehden, director of product management at Lightbend. If this conversation gets you interested in fast data processing and reactive systems, be sure to check out Tom Peck’s live webcast on March 28, 2017.
In this episode of the O’Reilly podcast, Jon Bruner sat down with Rahul Kamdar, director of product management and strategy at TIBCO Software. They discussed the shift from the centralized enterprise service bus (ESB) to a distributed data architecture based on microservices, APIs, and cloud-native applications.
In this week’s Design Podcast, I sit down with Noah Iliinsky, senior UX architect at Amazon’s AWS group, co-author of "Designing Data Visualizations," and co-editor of "Beautiful Visualization." We talk about how design is organized at Amazon, 17 keys to success, and why being intentional will ensure you are working on the right problems.
In this episode, O’Reilly Media’s Mac Slocum talks with Scout Brody, executive director of Simply Secure. They discuss building systems that help humans, designing better tools through user studies, and balancing the demands of shipping software with security.
In this special episode of the Data Show, O'Reilly's Jenn Webb speaks with Max Ogden, director of Code for Science and Society. Recently, Max and Code for Science have been working on the ongoing rescue of data.gov and assisting with other data rescue projects, such as Data Refuge; they’re also the nonprofit developers supporting Dat, a data versioning and distribution manager, which came out of Max’s work making government and scientific data open and accessible.
This week, I sit down with David Beyer, an investor with Amplify Partners. We talk about machine learning and artificial intelligence, the challenges he’s seeing in AI adoption, and what he thinks is missing from the AI conversation.
In this episode of the Data Show, I spoke with Anima Anandkumar, a leading machine learning researcher, and currently a principal research scientist at Amazon. I took the opportunity to get an update on the latest developments on the use of tensors in machine learning. Most of our conversation centered around MXNet—an open source, efficient, scalable deep learning framework. I’ve been a fan of MXNet dating back to when it was a research project out of CMU and UW, and I wanted to hear Anandkumar’s perspective on its recent progress as a framework for enterprises and practicing data scientists.
In this episode of the O’Reilly Bots Podcast, I speak with Tom Coates, co-founder of Thington, a service layer for the Internet of Things. Thington provides a conversational, messaging-like interface for controlling devices like lights and thermostats, but it’s also conversational at a deeper level: its very architecture treats the interactions between different devices like a conversation, allowing devices to make announcements to any other device that cares to listen.
In this week’s Design Podcast, I sit down with Ben Yoskovitz, investor, entrepreneur, and former VP of product at VarageSale and at GoInstant. We talk about using metrics in product development and why anyone building anything new needs to have both hubris and intellectual honesty. Yoskovitz is co-author of Lean Analytics, and is teaching a two-day course on product strategy as part of the upcoming O'Reilly Design Conference.
In this episode, I talk with Jessy Irwin, VP of security and privacy at Mercury Public Affairs. We discuss how to communicate security to non-technical people, what security might look like for small businesses, and moving beyond shame. We also meet her neighborhood gang of grannies who’ve learned how to hack back.
In this episode of the O’Reilly Bots Podcast, I speak with Tim Hwang, an affiliated researcher at the Oxford Internet Institute, about AI-driven psyops bots and their capacity for social destabilization.
In this episode of the Data Show, I spoke with Parvez Ahammad, who leads the data science and machine learning efforts at Instart Logic. He has applied machine learning in a variety of domains, most recently to computational neuroscience and security. Along the way, he has assembled and managed teams of data scientists and has had to grapple with issues like explainability and interpretability, ethics, insufficient amount of labeled data, and adversaries who target machine learning models. As more companies deploy machine learning models into products, it’s important to remember there are many other factors that come into play aside from raw performance metrics.
In this week's Radar Podcast, O’Reilly’s Mac Slocum chats with Sara Watson, a technology critic and writer in residence at Digital Asia Hub. Watson is also a research fellow at the Tow Center for Digital Journalism at Columbia and an affiliate with the Berkman Klein Center for Internet and Society at Harvard. They talk about how to optimize personalized experience for consumers, the role of machine learning in this space, and what will drive the evolution of personalized experiences.
In this episode of the O’Reilly Bots Podcast, Pete Skomoroch and I speak with Amir Shevat, head of developer relations at Slack and the author of the forthcoming O’Reilly book "Designing Bots: Creating Conversational Experiences."
In this week’s Design Podcast, I sit down with Simon Endres, creative director and partner at Red Antler. We talk about working from a single idea, how Red Antler is helping transform product categories, and the importance of having a point of view.
In this episode, I talk with Doug Barth, site reliability engineer at Stripe, and Evan Gilman, Doug’s former colleague from PagerDuty who is now working independently on Zero Trust networking. They are also co-authoring a book for O’Reilly on Zero Trust networks. They discuss the problems with traditional perimeter security models, rethinking trust in a networked world, and automation as an enabler.
This week, I sit down with Tom Davenport. Davenport is a professor of Information Technology and Management at Babson College, the co-founder of the International Institute for Analytics, a fellow at the MIT Center for Digital Business, and a senior advisor for Deloitte Analytics. He also pioneered the concept of “competing on analytics.” We talk about how his ideas have evolved since writing the seminal work on that topic, Competing on Analytics: The New Science of Winning; his new book Only Humans Need Apply: Winners and Losers in the Age of Smart Machines, which looks at how AI is impacting businesses; and we talk more broadly about how AI is impacting society and what we need to do to keep ourselves on a utopian path.
In this episode of the Data Show, I spoke with Jason Dai, CTO of big data technologies at Intel, and co-chair of Strata + Hadoop World Beijing. Dai and his team are prolific and longstanding contributors to the Apache Spark project. Their early contributions to Spark tended to be on the systems side and included Netty-based shuffle, a fair-scheduler, and the “yarn-client” mode. Recently, they have been contributing tools for advanced analytics. In partnership with major cloud providers in China, they’ve written implementations of algorithmic building blocks and machine learning models that let Apache Spark users scale to extremely high-dimensional models and large data sets. They achieve scalability by taking advantage of things like data sparsity and Intel’s MKL software. Along the way, they’ve gained valuable experience and insight into how companies deploy machine learning models in real-world applications.
In this episode of the O’Reilly Hardware Podcast, Brian Jepson and I speak with Mike Vladimer, co-founder of the Orange IoT Studio at Orange Silicon Valley. Vladimer discusses how Internet of Things devices could benefit from connectivity options other than those provided by well-known technologies (including cellular, WiFi, and Bluetooth), and explains the LoRa wireless protocol, which supports long-range and lower-power applications.
In this week’s Design Podcast, I sit down with Kat Holmes, principal design director, inclusive design at Microsoft. We talk about what she looks for in designers, working on the right problems to solve, and why both inclusive and universal design are important but not the same.
In this episode, O’Reilly’s Mac Slocum talks with Susan Sons, senior systems analyst for the Center for Applied Cybersecurity Research (CACR) at Indiana University. They discuss how she initially got involved with fixing the open source Network Time Protocol (NTP) project, recruiting and training new people to help maintain open source projects like NTP, and how security needn’t be an impediment to organizations moving quickly.
This week, I sit down with anthropologist, futurist, Intel Fellow, and director of interaction and experience research at Intel, Genevieve Bell. We talk about what she’s learning from current AI research, why the resurgence of AI is different this time, and five things that are missing from the AI conversation.
As data scientists add deep learning to their arsenals, they need tools that integrate with existing platforms and frameworks. This is particularly important for those who work in large enterprises. In this episode of the Data Show, I spoke with Adam Gibson, co-founder and CTO of Skymind, and co-creator of Deeplearning4J (DL4J). Gibson has spent the last few years developing the DL4J library and community, while simultaneously building deep learning solutions and products for large enterprises.
In this episode of the O’Reilly Bots Podcast, Pete Skomoroch and I speak with Chris Messina, bot evangelist, creator of the hashtag, and, until recently, developer experience lead at Uber. We talk about the origins of MessinaBot, ruminate on the need for bots that truly exploit their medium rather than imitating older apps, and take a look at what’s ahead for bots in 2017.
In this week’s Design Podcast, I sit down with Randy Hunt, VP of design at Etsy. We talk about the culture at Etsy, why it’s important to understand the materials you are designing with, and why humility is your most important skill.
In this episode, I talk with Steven Shorrock, a human factors and safety science specialist. We discuss the dangers of blaming human error, studying success along with failure, and how humans are critical to making our systems resilient.
Specialists describe deep learning as akin to a rocketship that needs a really big engine (a model) and a lot of fuel (the data) in order to go anywhere interesting. To get a better understanding of the issues involved in building compute systems for deep learning, I spoke with one of the foremost experts on this subject: Greg Diamos, senior researcher at Baidu. Diamos has long worked to combine advances in software and hardware to make computers run faster. In recent years, he has focused on scaling deep learning to help advance the state-of-the-art in areas like speech recognition.
On this week's episode of the Radar Podcast, O'Reilly's Mac Slocum chats with award-winning author Pagan Kennedy about the art and science of serendipity—how people find, invent, and see opportunities nobody else sees, and why serendipity is actually a skill rather than just dumb luck.
In this episode of the O’Reilly Bots Podcast, Pete Skomoroch and I speak with Brad Abrams, group product manager of Google Assistant, the company’s new AI-driven bot that lives in many different contexts, including the Pixel phone, the Allo messaging app, and the Google Home voice-controlled speaker.
Machine learning has been a mainstream commercial field for some time now, but it’s going through an important acceleration. In this podcast episode, I talk about that acceleration with two executives from MemSQL, a company that specializes in in-memory databases: Gary Orenstein, MemSQL chief marketing officer, and Drew Paroski, MemSQL vice president of engineering.
In this week’s Design Podcast, I sit down with Andra Keay, managing director of Silicon Valley Robotics. We talk about the evolution of robots, applications that solve real problems, and what constitutes a good robot.
In this episode of the O’Reilly Hardware Podcast, Jeff Bleiel and I speak with Joel Johnson, co-founder and CEO of BoXZY, a startup that makes an all-in-one desktop CNC mill, 3D printer, and laser engraver. We discuss the BoXZY device’s software and hardware, including its CAD/CAM software (Autodesk’s Fusion 360) and controller board (Arduino Mega), as well as its shield, steppers, and firmware.
In this episode, O’Reilly’s Jenn Webb talks with Fang Yu, cofounder and CTO of DataVisor. They discuss sniffing out fraudulent sleeper cells, incubation in money transfer fraud, and adopting a more proactive stance against fraud.
This episode consists of excerpts from a recent talk I gave at a conference commemorating the end of the UC Berkeley AMPLab project. This section pertained to some recent trends in Data and AI. For a complete list of trends we’re watching in 2017, as well as regular doses of highly curated resources, subscribe to our Data and AI newsletters.
This week we're featuring a conversation from earlier this year—O'Reilly's Mary Treseler chats with Giles Colborne, managing director of cxpartners. They talk about the transformative effects of AI on design, designing for natural language interactions, and why designers need to nurture the ability to reinvent themselves.
In this week’s Design Podcast, I sit down with Jay Trimble, mission operations manager, ground data system manager, and resource prospector for the Lunar Rover Mission at NASA. We talk about applying Agile, adopting design thinking and user-centered design, and what he and his team rely on to design and build software for mission control.
In this episode of the O’Reilly Hardware Podcast, Jeff Bleiel and I speak with Niko Triulzi, co-founder and CTO of AE Dreams, the makers of the children’s product Turtle Mail, a wooden mailbox with a WiFi-connected thermal printer inside. Triulzi explains the hardware and software workings of the device, which enables parents, relatives, and friends to send a message from a computer or mobile device. The child then receives the message, which is printed from the mailbox.
In this best of 2016 episode, I revisit a conversation from earlier this year with Cory Doctorow, a journalist, activist, and science fiction writer. We discuss the unexpected places where digital rights management (DRM) pops up, how it hinders artistic expression and legitimate security research, and the ill-anticipated (and often dangerous) consequences of copyright exemptions.
In this episode of the O’Reilly Bots Podcast, I speak with Dennis Mortensen, founder and CEO of X.ai, a personal assistant bot that handles meeting scheduling through email.
In this episode I spoke with Ion Stoica, cofounder and chairman of Databricks. Stoica is also a professor of computer science at UC Berkeley, where he serves as director of the new RISE Lab (the successor to AMPLab). Fresh off the incredible success of AMPLab, RISE seeks to build tools and platforms that enable sophisticated real-time applications on live data, while maintaining strong security. As Stoica points out, users will increasingly expect security guarantees on systems that rely on online machine learning algorithms that make use of personal or proprietary data.
In this episode, I sit down with Brad Knox, founder and CEO of Emoters, a startup building a product called bots_alive—animal-like robots that have a strong illusion of life. We chat about the approach the company is taking, why robots or agents that pass themselves off as human without any transparency should be illegal, and some challenges and applications of reinforcement learning and interactive machine learning.
Telcos are facing massive challenges stemming from new customer usage patterns, the rise of over-the-top (OTT) services, and a stagnant subscriber base. In this interview, O’Reilly’s Jon Bruner sat down with Dheeraj Remella, director of solutions architecture at VoltDB, to discuss how telcos must compete in the current industry landscape by using big data to regain value from OTT services and capitalizing on their infrastructure investment with the Internet of Things and augmented and virtual reality.
Sophisticated analytics need sophisticated interpretation in order to be valuable. In this podcast episode, I speak with John Thuma, director of strategy for Aster Analytics at Teradata, about how businesses need to plan and implement advanced analytics in order to get the results they want.
In this podcast episode, I walk through an introduction to ecosystem data architecture with Bob Montemurro, senior partner for architecture services in the international region for Teradata.
With interest growing in field programmable gate arrays (FPGAs)—witness Amazon’s recent addition of AWS EC2 instances that include dedicated FPGAs—this episode of the O’Reilly Hardware Podcast looks at what FPGAs are and how their capabilities are different from microcontroller boards (such as Arduino and Raspberry Pi).Jeff Bleiel and I speak with Ryan Cousins, co-founder and CEO of KRTKL (pronounced “critical”), the makers of Snickerdoodle, a board that’s based on an ARM/FPGA hybrid chip.
In this week’s Design Podcast, I sit down with Dan Mall, founder and director of Superfriendly. We talk about what skills designers should learn, pricing your work, and why getting to know yourself is just as important as learning the craft.
In this episode, O’Reilly’s Mary Treseler talks with Ame Elliot, design director at Simply Secure. They discuss designing for security and privacy, noteworthy tools, and the real-world consequences of design.
This week, I sit down with Fang Yu, cofounder and CTO of DataVisor, where she focuses on big data for security. We talk about the current state of the fraud landscape, how fraudsters are evolving, and how data analytics and behavior analysis can help defend against—and prevent—attacks.
In this episode I spoke with Vikash Mansinghka, research scientist at MIT, where he leads the Probabilistic Computing Project, and co-founder of Empirical Systems. I’ve long wanted to introduce listeners to recent developments in probabilistic programming, and I found the perfect guide in Mansinghka.
In this episode of the O’Reilly Bots Podcast, Pete Skomoroch and I talk with Richard Socher, chief scientist at Salesforce. He was previously the founder and CEO of MetaMind, a deep learning startup that Salesforce acquired in 2016. Socher also teaches the “Deep Learning for Natural Language Processing” course at Stanford University. Our conversation focuses on where deep learning and NLP are headed, and interesting current and near-future applications.
In this week’s Design Podcast, I sit down with Steph Hay, head of content, culture, and AI design at Capital One. We talk about designing for voice interactions, connecting with remote team members, and the importance of baking humanity into AI.
In this episode, I talk with Richard Moulds, vice president of strategy and business development at Whitewood Encryption. We discuss whether random number generation is as random as some might think and the implications that has on securing systems with encryption, how to harness entropy for better randomness, and emerging standards for evaluating and certifying the quality of entropy sources.
In this week’s Design Podcast, I sit down with Steph Hay, head of content, culture, and AI design at Capital One. We talk about designing for voice interactions, connecting with remote team members, and the importance of baking humanity into AI.
In this episode of the O’Reilly Hardware Podcast, Jeff Bleiel and I speak with Gilad Rosner, a privacy and information policy researcher, and the founder of the Internet of Things Privacy Forum. Rosner is also the author of the recently-published free O’Reilly ebook, “Privacy and the Internet of Things.”
In this episode of the O’Reilly Bots Podcast, Pete Skomoroch and I speak with Ben Brown, co-founder and CEO of Howdy.ai, the bot toolmaker behind the Botkit framework. Brown also runs the Talkabot conference, which was held in Austin this past September.
This week, I sit down with Hilary Mason, who is a data scientist in residence at Accel Partners and founder and CEO of Fast Forward Labs. We chat about current research projects at Fast Forward Labs, adoption hurdles companies face with emerging technologies, and the AI technology ecosystem—what's most intriguing for the short term and what will have the biggest long-term impact.
In this episode I spoke with Michael Franklin, co-director of UC Berkeley’s AMPLab and chair of the Department of Computer Science at the University of Chicago. AMPLab is well-known in the data community for having originated Apache Spark, Alluxio (formerly Tachyon) and many other open source tools. Today marks the start of a two-day symposium commemorating the end of AMPLab, and we took the opportunity to reflect on its impressive accomplishments.
In this episode of the O’Reilly podcast, I talk with Nathan Moore, CDN architect at StackPath. We discuss intelligent caching, building secure CDN solutions, and the future of front end security and performance.
In this episode of the O’Reilly Podcast, I sat down with Hugh McKee, solutions architect at Lightbend. McKee and I discussed why the actor model is an ideal choice for building today’s distributed architecture, how actor-based systems manage requests to perform tasks, the kinds of enterprises already having success with actors, and best practices for getting started with asynchronous systems.
In this week’s Design Podcast, I sit down with Cathy Pearl, director of user experience at Sensely and author of Designing Voice User Interfaces. We talk about defining conversations, the growing tools ecosystems, and how voice has lessened our screen obsession.
In this episode of the O’Reilly Hardware Podcast, Jeff Bleiel and I speak with Chris Lirakis, senior manager, engineering of novel computing architectures at IBM. Earlier this year, the company’s quantum computing platform in the cloud, the “IBM Quantum Experience,” was opened up to researchers, scientists, and the public.
In this episode, I talk with security architect Efrain Ortiz. We discuss how epidemiology can be applied to infosec, the parallels between using data and patterns to diagnose disease and find endpoint problems, and how to think like an epidemiologist in order to get out of reactive approaches to security at your own organization.
In this podcast episode, I talk multi-model databases with Damon Feldman, solutions director at MarkLogic. If this conversation gets you interested in multi-model databases, be sure to check out his live webcast on December 1, 2016.
In this week's episode, O'Reilly's Mac Slocum sits down with Richard Cook and David Woods. Cook is a physician, researcher, and educator, who is currently a research scientist in the Department of Integrated Systems Engineering at Ohio State University, and emeritus professor of health care systems safety at Sweden’s KTH. Woods also is a professor at Ohio State University and is leading the Initiative on Complexity in Natural, Social, and Engineered Systems, and he's the co-director of Ohio State University’s Cognitive Systems Engineering Laboratory. They chat about SNAFU Catchers; anomaly response; and the importance of not only understanding how things fail, but how things normally work.
The problem confronting a modern campaign manager is similar to the problem that any marketer encounters: how to spend finite resources to reach the right people and convince them to act. In this episode of the O’Reilly Bots Podcast, I talk data strategy with Andrew Therriault, chief data officer for the City of Boston. He was previously director of data science for the Democratic National Committee and is the editor of a free O’Reilly ebook called “Data and Democracy: How Political Data Science is Shaping the 2016 Elections.” We talk about how data is changing today’s political campaigns—particularly the way that campaigns now determine which potential supporters to target with phone calls, mailings, and door-to-door contact.
In this special two-segment episode of the Data Show, I spoke with Dafna Shahaf, assistant professor at the School of Computer Science and Engineering at the Hebrew University of Jerusalem. Her area of research is focused on tools and techniques for overcoming information overload, an area of increasing importance in an attention economy. With the upcoming U.S. Presidential Elections right around the corner, I included a conversation between Jenn Webb, host of the O’Reilly Radar Podcast, and Sam Wang, co-founder of the Princeton Election Consortium and professor of neuroscience and molecular biology at Princeton University.
In this week’s Design Podcast, I sit down with Danielle Malik, designer, owner, and mentor at Design Equation. We talk about mentoring the next generation of designers, what she is learning from recent design grads, and the role fear can play in our work.
In this episode of the O’Reilly Bots Podcast, Pete Skomoroch and I recap O’Reilly Bot Day, held October 19, 2016, in San Francisco. The event gave us a good picture of what the bot community—and bot landscape—looks like, and the diverse group of attendees conveyed a strong sense of optimism about bots. Slides from Bot Day presentations are available here.
We then speak with Shivon Zilis, partner at Bloomberg Beta, who has written extensively about artificial intelligence and its impact on a variety of industries. She wrote influential surveys of what she calls machine intelligence in 2014 and 2015, and she’s planning to publish her 2016 update at the end of October.
In this episode, I talk with Brendan O’Connor, a security researcher, lawyer (but not your lawyer) and owner of security consulting firm Malice Afterthought. We discuss creating a culture that celebrates collaborative teamwork over harried heroes, how monitoring and checklists really can save lives, and breaking out of the security monoculture.
In this episode of the Hardware Podcast, I speak with Mark Wright, director of product management at Samsung Strategy and Innovation Center, and Darren Beck, author of the newly-published O'Reilly ebook "Smart Business: Gaining an Edge Through IoT-Powered Sustainability."
What does it take to build a data-driven culture? In this O’Reilly Podcast episode, I pose that question to Ashish Thusoo, founder of Qubole. With his co-founder Joydeep Sen Sarma, Thusoo built a self-service data infrastructure at Facebook beginning in 2007. That infrastructure transformed Facebook’s culture and became part of practically every product decision that the company made.
Customer service is a key application for bots—one of the first that we think of when we imagine a world full of AI-enabled conversational interfaces. In this episode of the O’Reilly Bots podcast, Pete Skomoroch and I talk with the founders of two companies that have developed bots to help consumers and companies talk to each other.
On this week's episode, I chat with Sam Wang, professor of neuroscience and molecular biology at Princeton. Wang is also a co-founder of the Princeton Election Consortium, a site focused on analyzing and predicting U.S. national elections. We talk about the site's prediction algorithm and this crazy election cycle, and the role neuroscience may have played. We also talk about the current research Wang and his team are working on, the U.S. BRAIN Initiative, and the powerful role governments play in academic research.
In this episode of the O’Reilly Data Show, I spoke with Christopher Nguyen, CEO and co-founder of Arimo. Nguyen and Arimo were among the first adopters and proponents of Apache Spark, Alluxio, and other open source technologies. Most recently, Arimo’s suite of analytic products has relied on deep learning to address a range of business problems.
The workplace is a rich venue for bots that help with productivity and collaboration. In this episode of the O’Reilly Bots Podcast, Pete Skomoroch and I focus on workplace bots. We begin by talking with Jassim Latif, head of partnerships at Slack, about bots written by outside developers for scheduling, organizing meetings, and managing human resources. “We’re building toward a future where the value of these apps and services that people are building on top of Slack far outweigh the value of Slack on its own,” Latif says.
In this week’s Design Podcast, I sit down with Kristin Skinner, managing director at Adaptive Path, head of design management at Capital One, and co-author of "Org Design for Design Orgs." We talk about managing design teams, scaling design, and what we can learn from the Golden State Warriors.
In this episode of the O’Reilly Data Show, O’Reilly’s online managing editor Jenn Webb speaks with Natalino Busa on the topic of predictive analytics, the challenges of feature engineering, and a new class of techniques that is enabling features to emerge from patterns within the data. They also discuss the relationship between predictive techniques and high-quality microservices, and how machine learning is being used to improve financial services.
In this episode, I talk with Dan Kaminsky, founder and chief scientist at White Ops. We discuss what a National Institutes of Health (NIH) for security would look like, the pros and cons of Docker and ephemeral solutions, and how the mere act of listening to people better can improve security for everyone.
In this episode of the Radar Podcast, I chat with John Bassett III, chairman of the board of the Vaughan-Bassett Furniture Company. We talk about globalization and the effect it's had on the furniture industry, the international trade battle he waged (which was written about by Beth Macy in her book Factory Man), Bassett's book Making it in America, and what entrepreneurs need to know to succeed in business today.
Ask a random person for an example of an AI system and chances are he or she will name self-driving vehicles. In this episode of the O’Reilly Data Show, I sat down with Shaoshan Liu, co-founder of PerceptIn and previously the senior architect (autonomous driving) at Baidu USA. We talked about the technology behind self-driving vehicles, their reliance on rule-based decision engines, and deploying large-scale deep learning systems.
In this week’s Design Podcast, I sit down with Tom Greever, UX director at Bitovi and author of Articulating Design Decisions. We talk about how to effectively explain your design decisions, avoiding the CEO button, and how saying 'yes' is a facilitation superpower.
In this episode of the O’Reilly Bots podcast, Pete Skomoroch and I speak with Lili Cheng, general manager of FUSE Labs at Microsoft Research. Cheng’s team is responsible for the Microsoft Bot Framework. She’s also a speaker at O’Reilly’s upcoming Bot Day on October 19, 2016, in San Francisco.
Cheng talks about Microsoft’s experimental bots and their goal of making conversations playful and engaging. We also discuss the importance of designing good dialog; the potential of workplace bots; Xiaoice, Microsoft’s popular Chinese chatbot; and we reflect on the fate and significance of Microsoft’s Tay bot.
In this episode, I talk with Josh Corman, co-founder of I Am the Cavalry and director of the Cyber Statecraft Initiative for the non-profit organization Atlantic Council. We discuss his recent work advising the White House and Congress on the many issues lurking in safety-critical systems in the health care industry, the misaligned incentives across health care, regulatory bodies and the software industry, and the recent incident between MedSec and St. Jude regarding their medical devices.
In this Radar Podcast episode, I chat with Haakon Faste, a design educator and innovation consultant. We talk about his interesting career path, including his perceptual robotics work, his teaching approaches, and his mission with the Ralf A. Faste Foundation. We also talk about navigating our way to a "post-human" world and the importance of designing to make the world a more human-centered place.
In this episode of the O’Reilly Bots Podcast, Pete Skomoroch and I speak with Andy Mauro, co-founder and CEO of Automat, a startup whose tools make it easy to build AI-powered bots. (Disclosure: Automat is a portfolio company of O’Reilly AlphaTech Ventures, a VC firm affiliated with O’Reilly Media.) Mauro will be speaking at O’Reilly Bot Day on October 19, 2016, in San Francisco.
In this episode of the O’Reilly Data Show I sat down with O’Reilly author Dean Wampler, big data architect at Lightbend. We talked about new architectures for stream processing, Scala, and cloud computing.
In episode five of the O’Reilly Bots podcast, Pete Skomoroch and I speak with Joshua Browder, the 19-year old founder and CEO of DoNotPay, a series of bots that help people with legal issues, including challenging parking tickets, challenging bank charges, and claiming government assistance for homelessness. Dubbed “the world’s first robot lawyer,” his bots have attracted 260,000 users and provided 175,000 successful parking-ticket appeals.
In this week’s Design Podcast, I sit down with Paul Adams, VP of product at Intercom. Before joining Intercom, Adams had stints at Dyson, Google, and Facebook. We talk about his career path, building design teams, and Intercom’s goal to connect humans at scale.
In this episode of the O’Reilly Data Show, I spoke withMichael Li, cofounder and CEO of the Data Incubator. We discussed the current state of data science and data engineering training programs, Apache Spark, quantitative finance, and the misunderstanding around the term “data science.”
In this episode, I talk with Kyle Rankin, vice president of engineering operations at Final, a credit card startup. We discuss old versus new approaches to server hardening in light of the cloud, how institutional inertia thwarts change, and the new security-minded desktop OS Qubes.
This week on the Radar Podcast, we're featuring the first episode of the newly launched O'Reilly Bots Podcast, which you can find on Stitcher, iTunes, SoundCloud and RSS. O'Reilly's Jon Bruner is joined by Pete Skomoroch, the co-founder and CEO of Skipflag, to talk about bots—about what's driving the sudden interest, what we can expect from the technology, and some interesting emerging applications.
In episode four of the O’Reilly Bots podcast, Pete Skomoroch and I speak with Cathy Pearl, director of user experience at Sensely, and author of the forthcoming O’Reilly book “Designing Voice User Interfaces.” She’s also a speaker at O’Reilly’s upcoming Bot Day on October 19, 2016, in San Francisco.
While I was in Beijing for Strata + Hadoop World, several people reminded me of the chatbot Xiaoice—one of the most popular accounts on the Chinese social media site Weibo. Developed by Microsoft researchers, Xiaoice comes with a personality and is able to engage users in extended conversations on Weibo. These types of capabilities highlight that in an attention economy, systems that are able to forge an emotional connection will garner more loyalty and engagement from users.
In this episode of the O’Reilly Data Show, I sat down with Rana el Kaliouby, co-founder and CEO of Affectiva, one of the leading experts in emotion sensing systems. We talked about the impact of deep learning and computer vision, Affectiva’s large facial expression database, and privacy and ethics in an era of multimodal systems.
In episode three of the O’Reilly Bots podcast, Pete Skomoroch and I speak with Dennis Yang, co-founder and chief product officer of Dashbot, an analytics platform for bots. Bots are a new way for humans to interact with computers, and require new ways of thinking about measurement.
We discuss crucial differences between bots and conventional interfaces, how human writers are essential for setting the right tone in a bot, and why users ask bots to tell them jokes.
In this week’s Design Podcast, I sit down with Kristian Simsarian, founder and chair of the undergraduate design program at California College of the Arts. We talk about design education, design thinking, and the need for more wisdom.
In this episode, I talk with Meredith Patterson, a software engineer and leader of the Langsec Conspiracy. We discuss the origins of LangSec, rigidity versus robustness, and game theory as it applies to organizational approaches to security.
Greylock Partners investor Sarah Guo joins us for episode two of our new pop-up podcast on bots and conversational interfaces. She’s written insightfully on bots and has worked on several investments in bot startups.
This week's Radar Podcast episode is a special cross-over edition from the O'Reilly Security Podcast, which you can find on iTunes, Stitcher, RSS, or SoundCloud. O'Reilly strategic content director Courtney Nash chats with Cory Doctorow, a journalist, activist and science fiction writer. They talk about nascent pro-security industries, the EFF's lawsuit against the U.S. government, and the new W3C DRM specification.
In this episode of the O’Reilly Data Show, I spoke with Adam Marcus, co-founder and CTO of B12, a startup focused on building human-in-the-loop intelligent applications. We talked about the open source platform Orchestra,for coordinating human-in-the-loop projects; the current wave of human-assisted AI applications; best practices for reviewing and scoring experts; and flash teams.
We’re launching a new pop-up podcast about bots. In this first episode of the O’Reilly Bots Podcast, I’m joined by Peter Skomoroch to talk background: why everyone is suddenly interested in bots and what they promise to do, and what sorts of applications are beginning to emerge.
In this week’s Design Podcast, I sit down with former president of Frog, Doreen Lorenzo. Lorenzo is currently the director for the Center of Integrated Design at the University of Texas at Austin. We talk about the design in education, women in design, and failing fast versus learning fast.
In this episode, I talk with Cory Doctorow, a journalist, activist, and science fiction writer.
We discuss the EFF lawsuit against the U.S. government, the prospect for a whole new industry of pro-security businesses, and the new W3C DRM specification.
O'Reilly editor Brian Anderson chats with James Bond, Chief Technologist for Hewlett Packard (HP). Their discussion centers around the benefits, challenges, and best practices of migrating to the cloud.
This week, I talk with Alyona Medelyan, co-founder and CEO at Thematic and founder and CEO at Entopix. We talk about natural language understanding, the challenges of analyzing unstructured text, and her open source indexing tool Maui that she's been working on for the past 10 years.
In this episode of the O’Reilly Data Show, I spoke withJana Eggers, CEO of Nara Logics. Eggers’ involvement with AI dates back to her days as a researcher at the Los Alamos National Laboratory. Most recently she has been helping companies across many industries adopt AI technologies as a way to enable a range of intelligent data applications.
In this week’s Design Podcast, I sit down with Giles Colborne, designer, author, and managing director of cxpartners. We talk about how AI is reinventing design and the roles of designers; the balance of creating something that is different but familiar; and how, at its most basic level, AI is shortcutting user input.
In this episode, I talk with Chris Eng, vice president of research at Veracode, a software security-as-a-service business.
We discuss Veracode’s research on application security across a broad spectrum of industries, the challenges of securing modern “assembled” software, and making it easier for developers to bake in security from the get-go.
In this episode of the Hardware podcast, O’Reilly editor Brian Jepson and I talk about what FPGAs can do, how they’re getting more approachable for beginners, and the parallels between FPGAs and microcontroller boards such as the Arduino.
This week's episode features a special cross-over conversation from the O'Reilly Security Podcast, which you can find on Stitcher, iTunes, SoundCloud, or RSS. O'Reilly's Courtney Nash chats with Eleanor Saitta, a security architect at Etsy. They talk about the importance of thinking of security in a human context and the increasingly critical relationship between security and design.
Ben Lorica chats with Teradata’s Sri Raghavan about the evolution of data analytics in Spark. They discuss why some feel there is still a high barrier for entry with Spark for data analytics, and some tools to help overcome this barrier.
In this episode of the O’Reilly Data Show, I spoke with John Akred, cofounder and CTO of Silicon Valley Data Science. Akred and his colleagues teach two of the more popular Strata + Hadoop World tutorials—“Developing a Modern Enterprise Data Strategy” and “Architecting a Data Platform.” We talked about his career in data science and consulting, and his penchant for bringing emerging technologies and tools into large enterprises.
In this week’s Design Podcast, I sit down with Jim Kalbach, designer, instructor, and author of Mapping Experiences. We talk about the relationship between design and design thinking, how to get started with mapping experiences, and the notion of shared value as a strategic competitive advantage.
In this episode, I talk with Guy Podjarny, founder of Snyk, a developer tooling company focused on securing open source alongside building a business.
We discuss the parallel paths between the transformation from Ops teams to DevOps and where security teams are right now, building security tools focused on the people who will be using them, and who owns the problem of vulnerabilities in open source.
In this episode of the Hardware podcast, I talk with Dave Rauchwerk, founder and CEO of Next Thing Co., makers of C.H.I.P., the $9 computer.
This week, I chat with Othman Laraki, co-founder of Color Genomics. We chat about challenges and opportunities in genetic testing, the future of precision medicine, and the hurdles medicine and health care are currently facing (and how we can overcome them).
In this episode of the O’Reilly Data Show, I spoke with Yishay Carmiel, president of Spoken Labs. As voice becomes a common user interface, the need for accurate and intelligent speech technologies has grown. And although computer vision is a common entry point for deep learning, some of the most interesting commercial applications of deep neural networks are in speech recognition. Carmiel has spent several years building commercial speech applications, and along the way he has witnessed (and helped architect) massive improvements in speech technologies.
This week's episode of the Design Podcast features a conversation I had with Mike Kuniavsky last fall. Kuniavsky is a user experience designer, researcher, and author currently working at Parc. He's also a speaker at the upcoming O'Reilly online conference "Designing for the Internet of Things," September 15, 2016. In our chat, Kuniavsky talks about designing for the IoT, service design, and the mindshift needed to design for ecosystems.
In this episode, I talk with Eleanor Saitta, a security architect at Etsy. We talk about how security isn’t really about what happens to computers—it’s about what happens to the people using those systems; the relationship between design and security; and shifting the industry’s focus to think about security as a product of shared human outcomes.
In this episode of the Hardware podcast, I talk with Steve Ghee, senior vice president for R&D in the office of the CTO at PTC. The conversation took place during the LiveWorx 2016 conference, where augmented reality was a major focus—especially for the role that it can play in the Internet of Things and in the enterprise.
I recently talked to Sean Suchter, co-founder and CEO of Pepperdata, about why Spark has become so popular and where it still presents challenges. Spark represents the evolution of how the computer field understands and is addressing the challenge of handling big data (a term I will casually use without trying to define—any definition you care to plug in will be relevant for this discussion).
This week's episode features two conversations I've had recently centered around smart cities. First, I chat with Daniele Quercia, research team lead at Bell Labs. We talk about research he's working on now; the launch of goodcitylife.org (including smelly maps and happy maps); why our use of technology shouldn't just aim to make a city smart, but to improve the day-to-day quality of life of it's citizens; and about the emerging areas of urban informatics he's finding most compelling.
In our second segment, I chat with Frank Cuypers, associate professor at the University of Antwerp and strategist at Destination Think! We chat about the importance of urban DNA, his nonprofit project Why Your City, and why there's no such thing as a smart city.
In this episode of the O’Reilly Data Show, I spoke with Rajat Monga, who serves as a director of engineering at Google and manages the TensorFlow engineering team. We talked about how he ended up working on deep learning, the current state of TensorFlow, and the applications of deep learning to products at Google and other companies.
This week's episode of the Design Podcast features a conversation I had with Martin Charlier last fall. These days, Charlier is a freelance design consultant and co-founder at Rain Cloud. He's also a contributing author to Designing Connected Products and a speaker at the upcoming O'Reilly online conference "Designing for the Internet of Things," September 15, 2016. In our chat, Charlier talks about designing for the IoT, design's responsibility, and the importance of team dynamics.
In this episode of the Security Podcast, I talk with Jay Jacobs, senior data scientist at BitSight. We discuss the disparity between intuition and analytics in data science, the limitations of unsupervised machine learning, and the challenges of creating effective data visualizations.
Two years ago, the Othermill desktop CNC mill made machining radically more accessible. Priced at $2,200, it found its way beyond general small-scale milling and into electronics prototyping; it can mill copper away from an FR-1 PCB blank to make a circuit board. Other Machine’s latest work is the Othermill Pro, whose spindle speed and rapid movement are 60% and 70% faster than the original’s, respectively, and whose accuracy makes it possible to fabricate traces as small as six thousandths of an inch.
In this episode of the Hardware Podcast, we talk with Ezra Spier, the vice president of product & software at Other Machine Co., maker of the Othermill. Spier walks us through the development process for the improved machine and gives us an overview of the desktop prototyping market.
In this episode of the Hardware podcast, we talk with Chris Meringolo, an engineer at ThingWorx, and Brian Jepson, an editor at O’Reilly Media. This conversation took place during the LiveWorx 2016 conference, where the three of us presented a hands-on connected-device tutorial that we’re calling the O’Reilly IoT Learning Lab.
This week's episode is a cross-post from the O'Reilly Design Podcast. O'Reilly's Mary Treseler chats with Ame Elliott, design director at Simply Secure. They talk about security and privacy design, with a focus on the end user experience, and how to give designers a voice in changing the shape of a product and getting the right values out in the world. Elliott also talks about how architecture inspires her work and why problem finding is a better approach then problem solving.
In this episode of the O’Reilly Data Show, I spoke with data management industry veteran Rohit Jain, currently the CTO of Esgyn. We talked about his years at HP Labs, and his recent project to bring hybrid transactional/analytic technologies into the Hadoop ecosystem.
In this episode of the O’Reilly Podcast, I sat down with Markus Eisele, developer advocate at Lightbend. Eisele and I discussed the inherent difficulties in developing distributed systems, how reactive principles enhance microservices, the new reactive microservices framework, Lagom, and open source’s influence on enterprise development.
In this week’s Design Podcast, I sit down with designer Chris Maury. Maury is the founder of Conversant Labs, working on projects intended to help improve the lives of the blind. We talk about designing for the blind (as he loses his sight), how chatbots might just make us better listeners, and principles for designing the best conversational UIs.
In this episode, I talk with Jack Whitsitt, senior strategist at EnergySec. We discuss the ways in which language can either divide or unite people and organizations, the illusion of control when it comes to security, and how any model or framework for security must include people in order to have any chance of success.
In this episode of the Hardware Podcast, we talk with Jonathan Bachrach, a professor in the Electrical Engineering and Computer Sciences (EECS) Department at UC Berkeley, and co-founder of Otherlab, a research-driven lab and incubator that sits at the intersection between software and machines.
This week, O'Reilly's Mac Slocum chats with Ben Lorica, O'Reilly's chief data scientist and host of the O'Reilly Data Show Podcast. Lorica talks about emerging themes in the data space, from machine learning to deep learning to artificial intelligence, and how those technologies relate to one another and how they're fueling real-time data applications. Lorica also talks about how the concept of a data center is evolving, the importance of open source big data components, and the rise in interest of big data ethics.
In this episode of the Data Show, I spoke with Mike Tung, founder and CEO of Diffbot - a company dedicated to building large-scale knowledge databases. Diffbot is at the heart of many web applications, and it’s starting to power a wide array of intelligent applications. We talked about the challenges of building a web-scale platform for doing highly accurate, semi-supervised, structured data extraction. We also took a tour through the AI landscape, and the early days of self-driving cars.
Coco Krumme is a data scientist who’s worked at the intersection of applied math and agriculture, and her recent “Silicon Valley Magnets” project gently satirizes Silicon Valley’s “techno-utopian mythology.” In this episode of the Hardware podcast, we talk about how connected sensors, AI, and data science are changing agriculture, and reflect on the different ways that Silicon Valley and farmers are inclined to solve problems.
In this week’s Design Podcast, I sit down with Max Burton, founder of Matter. Before starting his own firm, Burton spent the last two decades at places like Frog, Nike, and Smart Design. We talk about the future of wearables, what he looks for when hiring designers, and what tech companies can learn from Nike’s and Disney's approach to product design.
In this inaugural episode of the O’Reilly Security Podcast, I talk with Allison Miller, a product manager at Google and my co-chair for the new O’Reilly Security Conference. We discuss her evolving understanding of the nature of risk and fraud in complex systems; the role of humans in technical systems; the cultural downsides of security by obscurity; and the new conference we’re putting together, which is squarely focused on helping defenders.
In this episode of the Hardware podcast, we talk with Carl Bass, president and CEO of Autodesk. He’s an articulate thinker on algorithmic design, collaborative tools, and the nature of craft, and we talked for nearly two hours when we visited him to record this episode.
This week, we're featuring a special crossover podcast from our O'Reilly Design Podcast. O'Reilly's Mary Treseler chats with investor, entrepreneur, and former VP of product, Ben Yoskovitz. Yoskovitz talks about product design strategy and the benefits of lean approaches, where product teams tend to fall down, and why a bottom-up approach to product design is more successful—and more scalable—than a top-down approach.
With the release of Spark version 2.0, streaming starts becoming much more accessible to users. By adopting a continuous processing model (on an infinite table), the developers of Spark have enabled users of its SQL or DataFrame APIs to extend their analytic capabilities to unbounded streams.
Within the Spark community, Databricks Engineer, Michael Armbrust is well-known for having led the long-term project to move Spark’s interactive analytics engine from Shark to Spark SQL. (Full disclosure: I’m an advisor to Databricks.) Most recently he has turned his efforts to helping introduce a much simpler stream processing model to Spark Streaming (“structured streaming”).
Manu Prakash, a bioengineering professor at Stanford University, talks about radically inexpensive microscopes and the democratization of discovery in this episode of the Hardware Podcast. Prakash’s Foldscope is a “completely functional, foldable, origami microscope.” He says that “every kid in the world should carry a microscope in their pocket,” and so far, more than 50,000 Foldscopes (the parts of which cost only $1) have been shipped to more than 130 countries.
In this week’s Design Podcast, I sit down with C Todd Lombardo, chief design strategist at Fresh Tilled Soil and adjunct professor at IE Business. Lombardo is coauthor of the recently released book Design Sprint. We talk about the relationship between design sprints, Lean UX, and Agile, and the skills needed to move from designing to managing designers.
In this episode of the Hardware podcast, we talk with writer and digital rights activist Cory Doctorow. He’s recently rejoined the Electronic Frontier Foundation to fight a World Wide Web Consortium proposal that would add DRM to the core specification for HTML. When we recorded this episode with Cory, the W3C had just overruled the EFF’s objection. The result, he says, is that “we are locking innovation out of the Web.”
In this episode of the O’Reilly Podcast, O’Reilly’s Ben Lorica sat down with Nikolaus Bates-Haus, technical lead at Tamr. Lorica and Bates-Haus discuss principal dimensions of data preparation, challenges and solutions for data processing at enterprise scale, the value of the data catalog, and how Tamr solutions integrate Spark.
In this episode, I chat with Marc Warner, CEO of ASI, a data science and business analytics consultancy and training organization in London. We talk about artificial intelligence, speculating about the future and looking at current real-world business applications of AI. We also talk about a survey Warner recently conducted with data science companies in London, where he uncovered a data scientist skills cap.
In this episode of the O’Reilly Data Show, I spoke with Danny Bickson, co-founder and VP at Dato, and the principal organizer of the Data Science Summit (full disclosure: I’m a member of the conference organizing committee). Among machine learning students and practitioners, recommender systems have become somewhat of a canonical use case and application. One of the early and popular building blocks was GraphLab’s collaborative filtering toolkit, a library originally written and maintained by Bickson. He has continued to keep tabs on the latest developments in recommenders and continues to help organize workshops on related topics throughout the world.
In this episode of the Hardware Podcast, we talk with Rob Chandhok, president of Helium, a startup with a platform for connecting sensors and control elements over the Internet. We talk about the process of taking a prototype to production, and the blend of hardware and software that allows smart edge devices to last an eternity on a single battery.
In this episode of the Hardware Podcast, we talk with Ken Shirriff, a software engineer at Google and writer of a fascinating series of electronics teardowns. His posts on power supplies, like Apple’s MacBook and iPad chargers, offer detailed looks at how the real things differ from counterfeits. And his writeup of Apple’s remarkably complex Magsafe connector is illuminating.
This week, O'Reilly's Mary Treseler chats with designer, creative coder, and artist Scott Murray about coding and computation in design, his book "Interactive Data Visualization for the Web" and his new book coming out soon "Creative Coding and Data Visualization with p5.js."
In this week’s Design Podcast, I sit down with Ben Yoskovitz, investor, entrepreneur, and former VP of product at VarageSale and at GoInstant. We talk about using metrics in product development and why anyone building anything new needs to have both hubris and intellectual honesty. Yoskovitz is co-author of "Lean Analytics," and is teaching a live online course, Product Strategy for Designers, on June 9, 2016.
In this episode of the O’Reilly Data Show, I spoke with Ira Cohen, co-founder and chief data scientist at Anodot (full disclosure: I’m an advisor to Anodot). Since my days in quantitative finance, I’ve had a longstanding interest in time-series analysis. Back then, I used statistical (and data mining) techniques on relatively small volumes of financial time series. Today’s applications and use cases involve data volumes and speeds that require a new set of tools for data management, collection, and simple analysis.
On the analytics side, applications are also beginning to require online machine learning algorithms that are able to scale, are adaptive, and free of a rigid dependence on labeled data. I talked with Cohen about the challenges in building an advanced analytics system for intelligent applications at extremely large scale.
In this episode of the Hardware Podcast, David Cranor and I talk with Marcin Wichary, design lead and typographer at Medium. Wichary is an authority on the history of keyboards, and he’s traced their development from an early period of enormous variation (when each manufacturer used to arrange arrow keys in a different way) through today’s relative stability.
In this week’s Design Podcast, I sit down with Scott Murray, designer, creative coder, and artist who writes software to create data visualizations. Murray is the author of Interactive Data Visualization for the Web and the forthcoming book Creative Coding and Data Visualization with p5.js: Drawing on the Web with JavaScript. Murray is teaching an online course, Programming for Designers on May 11-12, 2016. We talk about why coding is a great skill for designers to learn (and it’s not just about earning more money); data visualization; and why design, at it’s core, is problem solving.
In this episode of the Hardware Podcast, David Cranor and I talk with Star Simpson, a designer, engineer, and manufacturer who works on every aspect of hardware. She’s a great example of a full-stack hardware creator, capable of moving between electrical engineering, software engineering, and design.
In this week's episode of the Radar Podcast, O'Reilly's Mac Slocum chats with Christine Park, senior product designer at Basis, and John Alderman, director of Supereverywhere. They talk about multi-modal design, which is an approach to design that takes into consideration the physical senses and the role they play in the user experience, and they also chat about how multi-modal design applies to the Web.
In this episode of the O’Reilly Data Show, I spoke with Mikio Braun, delivery lead and data scientist at Zalando. After spending previous years in academia, Braun recently made the decision to switch to industry. He shared some observations about building large-scale systems, particularly deploying data applications in production systems. Given his longstanding background as a machine learning researcher and practitioner, I wanted to get his take on topics like deep learning, hybrid systems, feature engineering, and AI applications.
In this week’s Design Podcast, I sit down with Dylan Field, founder of Figma and former Thiel Fellow. We talk about the problem Figma aims to solve for designers and how they’re measuring success. Field also talks about how they debugged Figma in its early days: by mandating their own designers use it to design the tool.
How can sound be used to both generate data and express data? In this episode of the Hardware podcast, we talk with Cameron Turner, co-founder and principal at The Data Guild. Turner is the author of the new O’Reilly report “Finding Profit in Your Organization’s Data: Examples and Best Practices.”
In this episode of the O’Reilly Data Show, I spoke with one of Strata + Hadoop World’s most popular teachers—Duncan Ross, data and analytics director at TES Global. In his long career in data, Ross has seen several stages of the evolution of tools, techniques, and training programs, and along the way he has interacted with business managers in many countries and regions. In keeping with his wide-ranging interests, we discussed many topics, including business analytics, data science training programs, data philanthropy and data for good, and university rankings.
What happens when a hobbyist technology goes commercial? In this episode of the Hardware Podcast, Chris Anderson, founder and CEO of 3D Robotics, talks about his company’s journey from “a DIY company to a consumer electronics company to an enterprise software company.”
In this week’s Design Podcast episode, I sit down with Joel Marsh, designer and author of UX for Beginners. We talk about design as a scientific way of thinking, what happens when you try to cash a check in Sweden, combining behavioral economics and design, and why learning design is like learning to play the piano.
In this O’Reilly Podcast, O'Reilly's Ben Lorica talks with Ryan Betts, CTO of VoltDB, about the IoT and streaming data. Their discussion covers unexpected use cases for IoT; building networked things that respond to IoT data (not just feeding data in); and the big picture of data management in the context of many, many networked things.
On this week’s Hardware Podcast, we talk seed funding with Katherine Hague, founder and CEO of Female Funders, and previously co-founder of ShopLocket, which she sold to PCH International in 2014.
Hague is an authority on seed financing and has experienced the process from both sides, and she’s both writing an O’Reilly book, Funded: The Entrepreneur's Guide to Raising Your First Round and leading an online seminar on seed funding on March 30, 2016.
In this week's episode, I sit down with Alyssa Ravasio, founder and CEO of Hipcamp. We chat about navigating the challenges of founding a company, mining government data, and the role the sharing economy will play in the future.
In this episode of the O’Reilly Data Show, I spoke with M.C. Srivas, co-founder of MapR and currently chief architect for data at Uber. We discussed his long career in data management and his experience building a variety of distributed systems. In the course of his career, Srivas has architected key components that now comprise many data platforms (distributed file system, database, query engine, messaging system, etc.).
The long-awaited Raspberry Pi 3 was released last week, so David Cranor and I present an episode of the Hardware Podcast on the contentious question of whether the Raspberry Pi is a Serious Computer for Serious People, or not.
In this week’s Design Podcast episode, I sit down with Simon King, director of Carnegie Mellon University’s Design Center. King is the author of Understanding Industrial Design. We talk about team dynamics and culture at IDEO, extending design education to non-designers, and design’s next big challenge.
This will be a breakout year for JavaScript on embedded systems. Our guest on the Hardware Podcast this week is Peter Hoddie, who founded Kinoma, which makes software and hardware for building JavaScript-powered prototypes. Previously he was one of the original developers of QuickTime at Apple.
In this week's episode, Hannah Grenade, a tech entrepreneur and former partner at McKinsey, chats with Matt Harris, managing director at Bain Capital Ventures. They talk about the most interesting areas in fintech innovation, taking a look at some hits and misses, and potential untapped areas of opportunity. Harris also talks about why the merchants payment battleground is no longer a great space for startups and why insurance is poised to be the final frontier for fintech innovation.
In this episode of the O’Reilly Data Show, I spoke with Fang Yu, co-founder and CTO of DataVisor. We discussed her days as a researcher at Microsoft, the application of data science and distributed computing to security, and hiring and training data scientists and engineers for the security domain.
David Cranor and I talk with Rob Coneybeer, managing director and co-founder of Shasta Ventures, one of the critical first investors in hardware startups including Nest, Fetch Robotics, and Turo (formerly RelayRides).
In this week’s Design Podcast episode, I sit down with Scott Hurff, product manager and lead designer at Tinder, Inc. Hurff is the author of "Designing Products People Love." In this episode, we talk about how Tinder approaches design, avoiding awkward UI, and why customer research is the most important skill for future designers.
In this new episode of the Hardware Podcast, David Cranor and I talk with data scientist Rachel Kalmar, formerly with Misfit Wearables and the founder and organizer of the Sensored Meetup in San Francisco. She shares insights from her work at the intersection of data, hardware, and health care.
David Beyer, co-founder and CEO of Chart.io, principal at Amplify Partners, and part of the founding team at Patients Know Best, chats with Risto Miikkulainen, professor of computer science and neuroscience at the University of Texas at Austin. They chat about evolutionary computation, its applications in deep learning, and how it’s inspired by biology.
Ben Lorica speaks with one of the most popular speakers at Strata+Hadoop World: Joe Hellerstein, professor of Computer Science at UC Berkeley and co-founder/CSO of Trifacta. They talk about his past and current academic research (which spans HCI, databases, and systems), data wrangling, large-scale distributed systems, and his recent work on metadata services.
David Cranor and Jon Bruner talk with Sanjit Biswas, founder and CEO of the industrial sensor company Samsara.
O'Reilly's Mary Treseler sits down with Tanya Kraljic, UX manager and principal designer at Nuance Communications to talk about the challenges of moving from graphical to voice interfaces, the voice tools ecosystem, and where she finds inspiration.
O'Reilly's Nicole Tache chats with Sean Suchter, co-founder and CEO of Pepperdata, about the inherent performance challenges in powerful distributed systems, the need for increased control and performance guarantees on existing Hadoop clusters, and the mission that’s driving the work he’s doing at Pepperdata.
O’Reilly’s Jenn Webb chats with Eric McNulty, a consultant, writer, speaker, and catalyst for positive leadership. McNulty talks about real-time disaster response, the connections between disaster response and organizational leadership, and how today’s leaders can achieve order beyond control and influence beyond authority. McNulty will talk more about instituting effective leadership at the Cultivate leadership training at Strata + Hadoop World in San Jose in March.
O'Reilly's Ben Lorica chats with Eric Colson, chief algorithms officer at Stitch Fix, and former VP of data science and engineering at Netflix. They talk about building and deploying mission-critical, human-in-the-loop systems for consumer Internet companies.
David Cranor and Jon Bruner talk with Matthew Berggren, who at the time the interview was conducted last December was senior director of product at Supplyframe. (Berggren is now director of Autodesk Circuits at Autodesk.) Their discussion focuses on the need for abstracted modules and better metadata in electronics.
In this episode of the O’Reilly Podcast, O’Reilly’s Ben Lorica sat down with Ben Sharma, CEO and co-founder of Zaloni, an organization that provides enterprise data management solutions for Hadoop. They discussed real-time data processing, the changing nature of data, sentiment analysis, and the trend toward combining cloud with on-prem infrastructures.
In this week’s Design Podcast episode, Mary Treseler sits down with Chrissie Brodigan, manager of user experience research at GitHub. They talk about user research and product development at Github, and the blindspots in product development and organizational development.
In this new episode of the Hardware Podcast, David Cranor and Jon Bruner talk with Roger Chen, formerly a principal at O’Reilly AlphaTech Ventures, O’Reilly Media’s sister VC firm.
In this week’s Radar Podcast episode, Aneel Lakhani, director of marketing at SignalFx, chats with Mark Burgess, professor emeritus of network and system administration, former founder and CTO of CFEngine, and now an independent technologist and researcher. They talk about the new edition of Burgess’ book, "In Search of Certainty," Promise Theory and how promises are a kind of service model, and ways of applying promise-oriented thinking to networks.
In this episode of the O’Reilly Data Show, Ben Lorica sat down with Vasant Dhar, a professor at the Stern School of Business and Center for Data Science at NYU, founder of SCT Capital Management, and editor-in-chief of the Big Data Journal (full disclosure: Lorica is a member of the editorial board). They talked about the early days of A.I. and data mining, and recent applications of data science to financial investing and other domains.
In this O’Reilly Podcast, Ben Lorica talks with John Hugg, founding engineer and manager of developer outreach at VoltDB. From the deluge of today's diverse data, Hugg describes how to extract meaningful information to make things cheaper and faster.
In this new episode of the Hardware Podcast—which features this show's first discussion focusing specifically on synthetic biology—David Cranor and Jon Bruner talk with Charles Fracchia, an IBM Fellow at the MIT Media Lab and founder of the synthetic biology company BioBright.
In this week’s Design Podcast episode, O'Reilly's Mary Treseler sits down with Wesley Yun, director of user experience on the hardware side at GoPro. Yun will be be speaking at O’Reilly’s inaugural Design Conference. In this episode, they talk about managing and recruiting designers at GoPro, Designer Fund’s Bridge Guild and mentoring the next generation of designers.
In this special holiday episode of the Radar Pocast, we’re featuring a podcast cross-over of the O’Reilly Design Podcast, which you can find on iTunes, Stitcher, TuneIn, or SoundCloud. O’Reilly’s Mary Treseler host’s the Design Podcast, and in this episode, she chats with Airbnb’s head of experience design Katie Dill about the values that drive design at Airbnb, the triforce structure of the company, and the process of journey mapping their users’ experience.
In this special holiday episode of the O’Reilly Data Show, I look back at two conversations I had earlier this year at the Spark Summit in San Francisco. The first segment is an on-stage fireside chat with Ben Horowitz, co-founder of Andreessen Horowitz and author of The Hard Thing About Hard Things.
In the second segment, Reynold Xin, one of the architects of Apache Spark, explains the rise of Apache Spark in China.
In this special holiday episode of the Radar Podcast, we’re featuring a cross-over of the O’Reilly Data Show Podcast, which you can find on iTunes, Stitcher, TuneIn, or SoundCloud. O’Reilly’s Ben Lorica hosts that podcast, and in this episode, he chats with Apache Spark release manager and Databricks co-founder Patrick Wendell about the roadmap of Spark and where it’s headed, and interesting applications he’s seeing in the growing Spark ecosystem.
In this week’s Design Podcast, I sit down with Kathryn McElroy, design lead on IBM’s Watson team. McElroy will be be speaking at O’Reilly’s inaugural Design Conference in January. In this episode, we talk about prototyping for digital and physical, design and diversity, and what it’s like working at IBM.
This episode of the Hardware Podcast features Jon Bruner's second discussion with Joe Biron, VP of IoT technology at ThingWorx, a PTC business that offers a platform for rapid deployment of Internet of Things applications.
In this week’s Radar Podcast episode, O’Reilly’s Mac Slocum delves into the economy with two speakers from our recent Next:Economy conference. First, Slocum talks with Leah Busque, founder of TaskRabbit, about service networking, TaskRabbit’s goals, and issues facing the peer economy. In the second segment, Slocum talks with Dan Teran, co-founder of Managed by Q, about the on-demand economy and the future of work.
Evan Chan, distinguished engineer at Tuplejump, on the early days of Spark (particularly his contributions to Spark/Cassandra integration), his interesting new open source project (FiloDB), and recent trends in cloud computing.
In this episode of the Hardware Podcast, we talk with Mengmeng Chen, head of U.S. operations at Seeed Studio.
Dave Zwieback, head of engineering at Next Big Sound and CTO of Lotus Outreach, talks about his new book, "Beyond Blame: Learning from Failure and Success," the framework for conducting a “learning review,” and how humans can keep pace with the growing complexity of the systems we’re building.
O'Reilly's Mary Treseler sits down with Bob Baxley, who is keynoting at OReilly’s inaugural Design Conference. He compares cultures at Apple and Pinterest, talks about competition in the design playing field, and addresses the designer shortage.
Our guest on this week’s episode of the O’Reilly Hardware Podcast is Robert Brunner. Brunner was director of industrial design at Apple from 1989 to 1996, overseeing the design of the PowerBook. He was the chief designer of Beats by Dr. Dre, the design-driven line of headphones that Apple acquired for $3 billion last year. And he’s the founder of Ammunition, which has worked with startups and large companies on a wide range of innovative consumer products.
O'Reilly's Jenn Webb chats with Jeff Jonas, an IBM fellow and chief scientist of context computing, Ironman triathlete, and contributing author to the new book, "Privacy in the Modern Age: The Search for Solutions." Jonas talks about applications of context-aware computing, his new G2 software, and astroid hunting with astronomers at the University of Honolulu.
Emil Eifrem, CEO and co-founder of Neo Technology, on the early days of NoSQL, applications of graph databases, cloud computing, and company culture in the U.S. and Sweden.
Jon Bruner talks with Ari Gesher, engineering ambassador at Palantir Technologies, and Kipp Bradford, research scientist at the MIT Media Lab.
Kristian Hammond, Narrative Science’s chief scientist, on Natural Language Generation, Narrative Science’s shift into the world of business data, and evolving beyond the dashboard.
Vanessa Cho, head of UX and research for the software and services group at GoPro, on designing for hardware and software, building design teams, and what she looks for in new recruits.
Mike Kuniavsky, a principal scientist in the Innovation Services Group at PARC, on designing for the Internet of Things ecosystem and why the most interesting thing about the IoT isn’t the “things” but the sensors. He also talks about his deep-seated love for appliances and furniture, and how intelligence will affect those industries.
Jai Ranganathan, senior director of product management at Cloudera, on trends in the Hadoop ecosystem, cloud computing, the recent surge in interest in all things real time, and hardware trends
In this episode of the Solid Podcast, David Cranor and Jon Bruner talk with Marcelo Coelho, creative director of Marcelo Coelho Studio, and Colin Raney, chief marketing officer at Formlabs.
Mary Yoko Brannen is an expert in ethnomethodology and qualitative studies of complex cultural organizational phenomena. She spends a lot of time focused on how changing cultural contexts affect technology and how to leverage cultural identity in the global workplace. We unpack all that in this episode and talk about how her proposed “ethnographic thinking” approach can address language and culture gaps in the global marketplace.
Dan Brown, designer at Eightshapes and author of "Designing Together" and "Communicating Design," on managing fixed and growth mindsets, embracing impostor syndrome, and the most important skill for all designers (hint: it’s not empathy).
In this episode of the Solid Podcast, David Cranor and O'Reilly's Jon Bruner talk with Robert Bodor, vice president and general manager for the Americas at Proto Labs, a rapid-prototyping service that’s been able to digitize large parts of the fabrication process.
O’Reilly’s Mary Treseler chats with Aaron Irizarry, director of user experience for Nasdaq product design, about Nasdaq’s journey to become a design-driven organization.
Irizarry also talks about the best ways to have solid conversations about the designs you’re working on, and why getting a seat at the proverbial table isn’t the endgame.
Google engineer Tyler Akidau on the evolution of stream processing, the challenges of building systems that scale to massive data sets, and the recent surge in interest in all things real time.
Host and O'Reilly chief data scientist Ben Lorica interviews Nidhi Aggarwal, whose background is in hardware and software, and whose experience ranges from work in engineering, consulting, research, and entrepreneurship.
Kickstarter is one of just a handful of large companies that have become public benefit corporations — committing themselves legally to social as well as financial goals. In this episode, David Cranor and Jon Bruner talk with Kickstarter’s co-founder and CEO, Yancey Strickler, about his decision to take the company through the public benefit process and his promise not to go through an IPO.
O'Reilly's Jenn Webb sits down with Google engineering site lead Ben Collins-Sussman and Tock founder and CTO Brian Fitzpatrick. The two have just released a new book, "Debugging Teams," a follow-up to their earlier book, "Team Geek." They talk about the new edition, how managing is a lot like being a psychotherapist, and how all their great advice plays out in their own lives.
Adam Connor, designer at MadPow and author of "Discussing Design" with Aaron Irizarry — Connor also will be speaking at O’Reilly’s inaugural Design Conference, talks about company culture and organizational design, the design of codes of conduct, and advice on running productive design critiques.
David Cranor and Jon Bruner talk with Tobias Kinnebrew, strategist at Google Robotics and formerly the director of product strategy at Bot & Dolly and principal creative director for HoloLens at Microsoft.
Cellist, composer, and performer Zoe Keating talks about why she chooses to retain total control over her music as opposed to signing with a label, she shares her thoughts on the changing landscape of how artists get paid, and she talks about why she’s optimistic about the future of music.
Evangelos Simoudis on data mining, investing in data startups, and corporate innovation.
Julia Ko, founder of SurePod, talks about entrepreneurship, niche product development, and spotting business opportunities.
Suw Charman-Anderson, journalist, consultant, and founder of Ada Lovelace Day, on why she founded Ada Lovelace Day and why it’s been so successful. She also talks about the state of social media, and the past, present, and future of blogging.
Katie Dill on designing for seven billion people, hiring good people, and the triforce.
Joe Biron, VP of IoT technology at ThingWorx, on how the IoT entails a flexible platform approach to accommodate new applications that haven’t been conceived yet.
Rajiv Maheswaran on the science of moving dots, and Claudia Perlich on big data in advertising.
Todd Lipcon on hybrid and specialized tools in distributed systems, and unveiling Kudu.
Buddy Michini on drone safety, trust, and real-time data analysis. Michini covers some potentially game-changing research in localization and mapping, and onboard computational abilities that might eventually make it possible for drones to improve their flight intelligence by analyzing their imagery in real time.
Poppy Crum, senior principal scientist at Dolby Laboratories and a consulting professor at Stanford, on sensory perception, algorithm design, and fundamental changes in music.
Pamela Pavliscak on the delicate relationship between data and design, and why it’s not an either or proposition, as well as why designing for happiness is good for business.
Dr. Clare Bernard, a field engineer at Tamr, talks about integrating, cataloging, and preserving metadata.
The New York Times' deputy technology editor talks about technology, people, and power.
Alistair Croll chats with musician and performer Amanda Palmer about the current state of the music industry and how she's navigating her way through new platforms, crowdfunding, and an ever-increasing amount of data.
O'Reilly's Mary Treseler chats with designer Gretchen Anderson about using design as a force to improve our lives and about approaching design as an inclusive discipline.
O'Reilly's Jon Bruner and David Cranor chat with Brady Forrest, vice president at Highway1, and Renee DiResta, vice president of business development at Haven, about hardware startup success stories, pitfalls, and best practices.
Ben Einstein chats with David Cranor and O'Reilly's Jon Bruner about critical issues for hardware startups.
O'Reilly's Mary Treseler chats with Suzanne Pellican, VP and executive creative director at Intuit, about three core principles of design thinking and about Intuit’s journey to become a design-driven organization.
O'Reilly's Ben Lorica chats with Mike Cafarella, assistant professor of computer science at the University of Michigan and co-founder of both Hadoop and Nutch.
O'Reilly's Rachel Roumeliotis chats with Joost Visser, head of research at Software Improvement Group (SIG), about how a focus on quality helps to achieve the fastest possible schedules and lowest possible cost of development and maintenance.
O'Reilly's Jenn Webb chats with Simon King, design director at IDEO, about bringing design to his farming roots, his new book "Understanding Industrial Design," and the synergies between industrial design and interaction design.
O'Reilly's Ben Lorica chats with Alice Zheng about her background, techniques for evaluating machine learning models, how much math data scientists need to know, and the art of interacting with business users.
O’Reilly’s Mac Slocum chats with Bradley Voytek about using data-driven approaches in his neuroscience work, the brain scanner project, and applying cognitive neuroscience to the zombie brain.
O'Reilly's Mary Treseler chats with Microsoft designer Cindy Alvarez about how design is changing, how the approach to design at Microsoft is changing, and user research misperceptions and challenges. She also offers advice to those who are insisting all designers should code.
Laura Klein and Kate Rutter chat about what is wrong with personas and how we can make some important tweaks in order make them more useful for design and communication.
This podcast is a cross-post from the User’s Know podcast (http://www.usersknow.com/podcast), and is republished here with permission. The podcast also is available on iTunes (https://itunes.apple.com/us/podcast/what-is-wrong-ux-users-know/id980133198).
O'Reilly's Jenn Webb chats with Cory Doctorow about pitfalls in the Internet of Things and his work with the Electronic Frontier Foundation (EFF), including work to reform the Digital Millennium Copyright Act (DMCA) — the source of many of those IoT pitfalls. Doctorow also talks about why we should treat human beings as things that are good at sensing as opposed to things that need to be sensed.
O'Reilly's Ben Lorica chats with David Epstein about his book, data science and sports, and his recent series of articles detailing suspicious practices at one of the world’s premier track and field training programs (the Oregon Project).
O'Reilly's Nick Lombardi chats with Laura Klein, designer, researcher, engineer, and author of "UX for Lean Startups" and the popular design blog Users Know.
O'Reilly's Jenn Webb chats with David Rose, entrepreneur, MIT Media Lab instructor, and author of "Enchanted Objects," about cognitive overload, our future with AI, and how the IoT will shake out.
O'Reilly's Mary Treseler chats with Ame Elliott, design director at Simply Secure, about the relationship between — and challenges of — privacy and security, how her experience of attending architecture school informs her design work, and why it's the responsibility of designers to create for a greater good.
O'Reilly's Ben Lorica chats with Ben Sharma, CEO and co-founder of Zaloni, about the early days of Hadoop and how businesses across industries are benefitting from Hadoop. They also discuss the evolution of tools in the space and how more companies are moving toward real-time decision-making with the growth of streaming tools and real-time data.
O'Reilly's Mac Slocum chats with Alasdair Allan, director at Babilim Light Industries, about the data coming out of the New Horizons Pluto flyby, the future of “personal space programs,” and why Bluetooth LE is cracking open the Internet of Things.
O'Reilly's Ben Lorica chats with Poppy Crum, neuroscientist and researcher at Dolby Labs, about AI and virtual reality systems, and about building a team of researchers from diverse disciplines.
O'Reilly's Ben Lorica chats with VoltDB co-founder Scott Jarr about how VoltDB’s hybrid transaction, analytic system allows for real-time analytics and personalization of data across various industries.
O'Reilly's Jenn Webb chats with Robert Brunner, founder of Ammunition design studio, about how design can help mitigate IoT pitfalls, what drove him to found Ammunition, and why he's fascinated with design's role in the movement toward automation.
O’Reilly’s Mary Treseler chats with Aaron Irizarry, Director of UX for Product Design at Nasdaq, about design-driven business and getting and keeping a seat at the table.
O'Reilly's Mac Slocum chats with Fjord's Andy Goodman about zero UI and the move toward intangible interfaces, and Cory Doctorow addresses the problems inherent in the Digital Millennium Copyright Act .
O'Reilly's Ben Lorica chats with UC Berkeley assistant professor Ben Recht about optimization, compressed sensing, and large-scale machine learning pipelines.
O'Reilly's Mike Hendrickson chats with Autodesk research fellow Mickey McManus about engaging with extreme users and what's going to happen when we have trillions of things sending billions of messages. McManus also talked about how we can prepare for the coming era of unbounded malignant complexity.
Tim O'Reilly chats with Google X's Astro Teller about moonshots, the relationship between technology and society, the learning process for hardware, and more.
O'Reilly's Ben Lorica chats with Ihab Ilyas, professor at the University of Waterloo and co-founder of Tamr, about how he started working on data cleaning tools, academic database research, and training computer science students for positions in industry.
O'Reilly's Mary Treseler chats with AlertMe founder Pilgrim Beart about why the scale of the Internet of Things creates as many challenges as it does opportunities. He also talked about the “gnarly problems” emerging from consumer wants and behaviors.
O'Reilly's Mike Hendrickson chats with database pioneer and 2014 Turing Award recipient Michael Stonebraker about winning the award, the future of data science, and the importance — and difficulty — of data curation.
O'Reilly's Ben Lorica chats with Phil Liu, co-founder and CTO of SignalFx, about hiring and building teams in the age of cloud computing, building tools for monitoring large numbers of time series, and lessons he’s learned from managing teams at leading technology companies.
Dan Shapiro, author of "Hot Seat: The Startup CEO Guidebook," discusses the importance of startup co-founders and why startups are hotbeds for imposter syndrome (and why that's OK).
O'Reilly's Jenn Webb talks with Digitteria founder Dele Atanda and Squirrel CEO and co-founder Mutaz Qubbaj about their startup platforms and the disruption in fintech.
O'Reilly's Ben Lorica chats with Patrick Wendell, release manager of Apache Spark and co-founder of Databricks, about how he came to join the UC Berkeley AMPLab, the current state of Spark ecosystem components, Spark’s future roadmap, and interesting applications built on top of Spark.
O'Reilly's Mary Treseler chats with Josh Clark, founder of Big Medium, about the changing nature of his work as the world itself becomes more of an interface, how to avoid "data rash," and why in this time of rapid technology growth it's essential for designers to splash in the puddles.
O'Reilly's Jenn Webb chats with Cait O’Riordan, VP of product, music and platforms at Shazam, about the current state of predictive analytics and how Shazam is able to predict the success of a song, and how the Internet of Things affects Shazam’s product life cycles as well as the behaviors of their users. Jenn also chats with Francine Bennett, CEO and cofounder of Mastodon C, about unintentional uses of data for evil and how data scientists can ensure their work stays on the data for good side of that line.
O'Reilly's Ben Lorica chats with Gary Kazantsev, who runs the R&D Machine Learning group at Bloomberg LP, about how big data and data science are making a difference in finance.
Jon Bruner and David Cranor talk with Andy Cavatorta and Jamie Zigelbaum, installation designers at Dark Matter Manufacturing, about how installation art is a great place to look for seamless integrations between hardware and software.
O'Reilly's Mary Treseler chats with PARC's Mike Kuniavsky about PARC’s work on IoT and the mindset shift the IoT will require.
O'Reilly's Mac Slocum chats with Rachel Andrew, founder of edgeofmyseat.com, about CSS Grid Layout and the role responsive design is playing in emerging Web technologies. Slocum also chats with open Web evangelist Kyle Simpson, who defends JavaScript Coercion’s negative reputation and contemplates the past, present, and future of the Web.
O'Reilly's Mary Treseler talks with information architect Jorge Arango about the state of IA and the importance of designers' understanding of context and perspective.
O'Reilly's Jenn Webb chats with Filament co-founder and CEO Eric Jennings about how a decentralized Internet would work, what we need to do to get there, and why it will become necessary as the Internet of Things begins to scale. Eric also talks about why his company recently shifted its focus from the maker community to industrial manufacturing.
O'Reilly's Mary Treseler chats with Scott Jenson, who is currently developing the Physical Web Project with the Chrome team at Google, about empathy, interaction on demand, and Google’s Physical Web Project.