StreamNative's Podcast: Recent Episodes

StreamNative

Podcasts about StreamNative, Pulsar, and BookKeeper

View Details

Summary

Matteo and Sijie from Streamlio reached out to us and let us know they had an update on Apache Pulsar. It turned out they had a lot to talk about so we cut the interview in two parts and here is the first part where they introduce Apache Pulsar, go in depth on the correct deployment scaling of a stable Pulsar cluster and clarify Pulsars “at least once vs exactly once” strategy. Part two will go in more depth on what’s new. Stay tuned!

Contact Info

  • Matteo Merli
    • Co-Founder – Software Engineer
  • Sijie Guo
    • Co-Founder
  • Apache Pulsar (incubating)

https://pulsar.apache.org/

This podcast was originally published on Roaring Elephant.

View Details

Summary

One of the critical components for modern data infrastructure is a scalable and reliable messaging system. Publish-subscribe systems have been popular for many years, and recently stream-oriented systems such as Kafka have been rising in prominence. In this episode, Rajan Dhabalia and Matteo Merli discuss the work they have done on Pulsar, which supports both options, in addition to being globally scalable and fast. They explain how Pulsar is architected, how to scale it, and how it fits into your existing infrastructure.

Interview

  • Introduction
  • How did you get involved in the area of data management?
  • Can you start by explaining what Pulsar is and what the original inspiration for the project was?
  • What have been some of the most challenging aspects of building and promoting Pulsar?
  • For someone who wants to run Pulsar, what are the infrastructure and network requirements that they should be considering and what is involved in deploying the various components?
  • What are the scaling factors for Pulsar and what aspects of deployment and administration should users pay special attention to?
  • What projects or services do you consider to be competitors to Pulsar and what makes it stand out in comparison?
  • The documentation mentions that there is an API layer that provides drop-in compatibility with Kafka. Does that extend to also supporting some of the plugins that have developed on top of Kafka?
  • One of the popular aspects of Kafka is the persistence of the message log, so I’m curious how Pulsar manages long-term storage and reprocessing of messages that have already been acknowledged?
  • When is Pulsar the wrong tool to use?
  • What are some of the improvements or new features that you have planned for the future of Pulsar?

Contact Info

  • Matteo
    • merlimat on GitHub
    • @merlimat on Twitter
  • Rajan
    • @dhabaliaraj on Twitter
    • rhabalia on GitHub

This podcast was originally published on Data Engineer Podcast.

View Details

The Data Exchange Podcast: Sijie Guo on how Apache Pulsar is able to handle both queuing and streaming, and both online and offline applications.

In this episode of the Data Exchange, Ben Lorica spoke with Sijie Guo, founder of StreamNative, a new startup focused on making enterprise messaging technologies – specifically Apache Pulsar – easy to use on the cloud. Sijie was previously a co-founder of Streamlio (acquired by Splunk) and prior to that he led the messaging team at Twitter. He is also the main organizer behind the Pulsar Summit Virtual Conference 2020.

Ben Lorica has written about the importance of foundational data technologies, and data ingestion and messaging are the starting point for modern data applications. As data and machine learning continue to grow in importance, it’s critical for companies to make sure they have the right messaging systems in place.

Our conversation spanned many topics, including:

  • The role of messaging in modern data applications and platforms.
  • The two main types of messaging applications: queuing and streaming.
  • Apache Pulsar as a unified messaging platform, able to handle both queuing and streaming, and both online and offline applications.
  • A status update on Apache Pulsar.

This podcast was originally published on The Data Exchange.