The free lunch is over! Computer systems up until the turn of the century became constantly faster without any particular effort simply because the hardware they were running on increased its clock speed with every new release. This trend has changed and today's CPUs stall at around 3 GHz. The size of modern computer systems in terms of contained transistors (cores in CPUs/GPUs, CPUs/GPUs in compute nodes, compute nodes in clusters), however, still increases constantly. This caused a paradigm shift in writing software: instead of optimizing code for a single thread, applications now need to solve their given tasks in parallel in order to expect noticeable performance gains. Distributed computing, i.e., the distribution of work on (potentially) physically isolated compute nodes is the most extreme method of parallelization.
Big data analytics and management are a multi-million dollar markets that grow constantly! The ability to control and utilize large amounts of data is the most valuable ability of today's computer systems. Because data volumes grow so rapidly and with them the complexity of questions they should answer, data analytics, i.e., the ability of extracting any kind of information from the data becomes increasingly difficult. As data analytics systems cannot hope for their hardware getting any faster to cope with performance problems, they need to embrace new software trends that let their performance scale with the still increasing number of processing elements.
In this lecture, we take a look at various technologies involved in building distributed, data-intensive systems. We start by discussing fundamental concepts in distributed computing, such das data models, encoding formats, messaging, data replication and partitioning, fault tollerance, and batch- and stream processing. In between, we consider different practical systems from the Big Data Landscape, such as Akka and Spark. In the end, we concentrate on data management aspects, such as distributed database management system architectures and distributed query optimization.
Dr. Thorsten Papenbrock
Prof. Dr. Felix Naumann, Dr. Thorsten Papenbrock
Dr. Thorsten Papenbrock
None
Sean Anderson
Tobias Macey
Prof. Dr. Tilmann Rabl
@dsdeployed
Ameet Talwalkar
David Chu
Drew Farnsworth
with Sam Ramji
Dr. Tony Hoang
Tobias Macey
Hammerspace
DataStax Developers
ZenML GmbH
O'Reilly Media
The Open University
StreamNative
Darren Pulsipher
O'Reilly Radar
Sanket Gupta
The Firebolt Data Bros
Travis Lawrence
Pam Teach
Andreas Kretz
The Open University
Joshua Matthew
Data Science In Production
Thomson Data
The Open University
BEPEC
TRIK
Kalicharan m
Richard Treves
Enrico Bertini and Moritz Stefaner
BINUS University
Data as a Product Podcast Network
DataTalks.Club
Safe Software Inc.
Alex Merced Podcasts
Cambridge University
Brookend Ltd
Banjo Obayomi
AI Guild
edureka!
Prof. Dr. Tilmann Rabl
GetInData
Naked Data Science
Thu Ya Kyaw & Koo Ping Shung
RAPIDS
For the Love of Data
Data Driven
Bonnie D. Graham
edureka!
Prof. Dr. Tilmann Rabl
datasciencehappywarriors
Pradeep Kumar
Julio Cezar Silva
Varun Sharma
Eric Kavanagh
None
Joel Grus
Dataiku
Ben Jaffe and Katie Malone
DataCamp
Data Crunch Corporation
BINUS University
Dataroots
Dan Linstedt
Ashwin Saxena
Founder360
Women in Analytics
DAGsHub
BINUS University
EvidenceN
Synthesized
Richmond Alake
StraitsResearch
storytelling with data author, speaker and dataviz guru Cole Nussbaumer Kna
BINUS University
Jon Krohn and Guests on Machine Learning, A.I., and Data-Career Success
Prasad
Pipeline Data Engineering Academy
Rich Miller
EM360
Rockset
Fireblaze AI School
Top End Devs
Daliana Liu
Kseniya K
Bob Haffner
Yevgeniy Sverdlik - Data Center Knowledge
Byron Reese
Port Harcourt School of AI
Dungeon Master
Matti