The free lunch is over! Computer systems up until the turn of the century became constantly faster without any particular effort simply because the hardware they were running on increased its clock speed with every new release. This trend has changed and today's CPUs stall at around 3 GHz. The size of modern computer systems in terms of contained transistors (cores in CPUs/GPUs, CPUs/GPUs in compute nodes, compute nodes in clusters), however, still increases constantly. This caused a paradigm shift in writing software: instead of optimizing code for a single thread, applications now need to solve their given tasks in parallel in order to expect noticeable performance gains. Distributed computing, i.e., the distribution of work on (potentially) physically isolated compute nodes is the most extreme method of parallelization.
Big Data Analytics is a multi-million dollar market that grows constantly! Data and the ability to control and use it is the most valuable ability of today's computer systems. Because data volumes grow so rapidly and with them the complexity of questions they should answer, data analytics, i.e., the ability of extracting any kind of information from the data becomes increasingly difficult. As data analytics systems cannot hope for their hardware getting any faster to cope with performance problems, they need to embrace new software trends that let their performance scale with the still increasing number of processing elements.
In this lecture, we take a look a various technologies involved in building distributed, data-intensive systems. We discuss theoretical concepts (data models, encoding, replication, ...) as well as some of their practical implementations (Akka, MapReduce, Spark, ...). Since workload distribution is a concept which is useful for many applications, we focus in particular on data analytics.
Dr. Thorsten Papenbrock
Dr. Thorsten Papenbrock
Dr. Thorsten Papenbrock
Tobias Macey
Tobias Macey
None
None
None
ZenML GmbH
Eric Anderson
CodeNewbie
Utsav Shah
Domenico Tripodi
Marvell Technology
None
Changelog Media
Darren Pulsipher
mapscaping.com
The Camunda Community Podcast, hosted by Josh Wulf.
satyabrata pal
mapscaping.com
Grafana Labs
SciNology Team
Ace Balangitan
None
MP English, Viv, Salim Virji
None
Patrick Wheeler and Jason Gauci
Brett Schechter
None
Andrew Lisowski, Justin Bennett
The Firebolt Data Bros
Charles Lowell & the Frontside Team
Confluent, original creators of Apache Kafka®
The Open University
Brock Palen
None
Jeremy Jung
Andrew Dalke
Tobias Macey
Dept
Ameet Talwalkar
SmartBear
Salesforce Engineering
Ben Pfaff
Liran Haimovitch
Software Engineering
Bright Computing
Best Java podcast on iTunes, learn about variables, control structures, col
Changelog Media
Adam Tuttle, Ben Nadel, Carol Hamilton, Tim Cunningham
Michael
Oxylabs
Quansight, LLC
Ben Lorica
None
Isaac Aderogba
Brian Okken
Adaptiva
Mike Challis
Adam Gordon Bell - Software Developer
Top End Devs
Exascale Computing Project
DevOps Porto
BJ Burns and Will Gant
Michael Kofman
InfoQ
Lightbend
Cord
Tweag I/O
Steve Westgarth
Justin Doescher
New Relic Developer Relations
None
LogRocket
Top End Devs
Allegheny College Department of Computer Science
SoftwareDaily.com
Richard Bown
JACK WAUDBY
None
IEEE Computer Society
mapscaping.com
Tomasz Nurkiewicz
Gaël Blondelle & Thabang Mashologu
None
Beyond Parsing
Jonathan Cutrell
srinimf
Brian Wagner, Jason Didonato
Tetrate
Robert Keller
Green Software Foundation
Justin Grammens
Brandon Williams & Stephen Celis
Platform.sh
Alagappan (AL)
Changelog Media
Data – Software Engineering Daily