The free lunch is over! Computer systems up until the turn of the century became constantly faster without any particular effort simply because the hardware they were running on increased its clock speed with every new release. This trend has changed and today's CPUs stall at around 3 GHz. The size of modern computer systems in terms of contained transistors (cores in CPUs/GPUs, CPUs/GPUs in compute nodes, compute nodes in clusters), however, still increases constantly. This caused a paradigm shift in writing software: instead of optimizing code for a single thread, applications now need to solve their given tasks in parallel in order to expect noticeable performance gains. Distributed computing, i.e., the distribution of work on (potentially) physically isolated compute nodes is the most extreme method of parallelization.
Big Data Analytics is a multi-million dollar market that grows constantly! Data and the ability to control and use it is the most valuable ability of today's computer systems. Because data volumes grow so rapidly and with them the complexity of questions they should answer, data analytics, i.e., the ability of extracting any kind of information from the data becomes increasingly difficult. As data analytics systems cannot hope for their hardware getting any faster to cope with performance problems, they need to embrace new software trends that let their performance scale with the still increasing number of processing elements.
In this lecture, we take a look a various technologies involved in building distributed, data-intensive systems. We discuss theoretical concepts (data models, encoding, replication, ...) as well as some of their practical implementations (Akka, MapReduce, Spark, ...). Since workload distribution is a concept which is useful for many applications, we focus in particular on data analytics.
Prof. Dr. Felix Naumann, Dr. Thorsten Papenbrock
Dr. Thorsten Papenbrock
Dr. Thorsten Papenbrock
None
Tobias Macey
Tobias Macey
None
Utsav Shah
Eric Anderson
None
None
The Camunda Community Podcast, hosted by Josh Wulf.
Grafana Labs
None
Domenico Tripodi
ZenML GmbH
Andrew Lisowski, Justin Bennett
None
mapscaping.com
Changelog Media
None
Jeremy Jung
Software Engineering
Darren Pulsipher
Confluent, original creators of Apache Kafka®
mapscaping.com
Best Java podcast on iTunes, learn about variables, control structures, col
Charles Lowell & the Frontside Team
SmartBear
Adam Tuttle, Ben Nadel, Carol Hamilton, Tim Cunningham
Brett Schechter
The Firebolt Data Bros
Michael
satyabrata pal
MP English, Viv, Salim Virji
Andrew Dalke
Salesforce Engineering
None
Brian Okken
Liran Haimovitch
Ace Balangitan
Changelog Media
Patrick Wheeler and Jason Gauci
Cord
Top End Devs
Isaac Aderogba
LogRocket
Adaptiva
Marvell Technology
Ben Pfaff
Mike Challis
Tomasz Nurkiewicz
SciNology Team
CodeNewbie
Quansight, LLC
Brian Wagner, Jason Didonato
Oxylabs
JACK WAUDBY
BJ Burns and Will Gant
New Relic Developer Relations
Adam Gordon Bell - Software Developer
Dept
Lightbend
srinimf
Changelog Media
Michael Kofman
Ameet Talwalkar
Richard Bown
Justin Doescher
Tobias Macey
SoftwareDaily.com
None
None
InfoQ
None
Ben Lorica
Steve Smith (@ardalis)
Allegheny College Department of Computer Science
Brock Palen
Gaël Blondelle & Thabang Mashologu
DevOps Porto
Platform.sh
Changelog Media
Jonathan Cutrell
The Open University
Steve Westgarth
Rafael Kennedy
Changelog Media
Brandon Williams & Stephen Celis
None
David Chu
Beyond Parsing
Top End Devs
Tweag I/O
Tetrate
Ortus Solutions
Alex Merced Podcasts
Luke Diebold
Futurice Tech Weeklies
Arpit Choudhury