I make videos about machine learning research papers, programming, and issues of the AI community, and the broader impact of AI in society.
Twitter: https://twitter.com/ykilcher Discord: https://discord.gg/4H8xxDF
If you want to support me, the best thing to do is to share out the content :)
If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this): SubscribeStar (preferred to Patreon): https://www.subscribestar.com/yannickilcher Patreon: https://www.patreon.com/yannickilcher Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq
ChatGPT, OpenAI's newest model is a GPT-3 variant that has been fine-tuned using Reinforcement Learning from Human Feedback, and it is taking the world by storm!
Sponsor: Weights & Biases
https://wandb.me/yannic
OUTLINE:
0:00 - Intro
0:40 - Sponsor: Weights & Biases
3:20 - ChatGPT: How does it work?
5:20 - Reinforcement Learning from Human Feedback
7:10 - ChatGPT Origins: The GPT-3.5 Series
8:20 - OpenAI's strategy: Iterative Refinement
9:10 - ChatGPT's amazing capabilities
14:10 - Internals: What we know so far
16:10 - Building a virtual machine in ChatGPT's imagination (insane)
20:15 - Jailbreaks: Circumventing the safety mechanisms
29:25 - How OpenAI sees the future
References:
https://openai.com/blog/chatgpt/
https://openai.com/blog/language-model-safety-and-misuse/
https://beta.openai.com/docs/model-index-for-researchers
https://scale.com/blog/gpt-3-davinci-003-comparison#Conclusion
https://twitter.com/johnvmcdonnell/status/1598470129121374209
https://twitter.com/blennon_/status/1597374826305318912
https://twitter.com/TimKietzmann/status/1598230759118376960/photo/1
https://twitter.com/_lewtun/status/1598056075672027137/photo/2
https://twitter.com/raphaelmilliere/status/1598469100535259136
https://twitter.com/CynthiaSavard/status/1598498138658070530/photo/1
https://twitter.com/tylerangert/status/1598389755997290507/photo/1
https://twitter.com/amasad/status/1598042665375105024/photo/1
https://twitter.com/goodside/status/1598129631609380864/photo/1
https://twitter.com/moyix/status/1598081204846489600/photo/2
https://twitter.com/JusticeRage/status/1598959136531546112
https://twitter.com/yoavgo/status/1598594145605636097
https://twitter.com/EladRichardson/status/1598333315764871174
https://twitter.com/charles_irl/status/1598319027327307785/photo/4
https://twitter.com/jasondebolt/status/1598243854343606273
https://twitter.com/mattshumer_/status/1598185710166896641/photo/1
https://twitter.com/i/web/status/1598246145171804161
https://twitter.com/bleedingedgeai/status/1598378564373471232
https://twitter.com/MasterScrat/status/1598830356115124224
https://twitter.com/Sentdex/status/1598803009844256769
https://twitter.com/harrison_ritz/status/1598828017446371329
https://twitter.com/parafactual/status/1598212029479026689
https://www.engraved.blog/building-a-virtual-machine-inside/
https://twitter.com/317070
https://twitter.com/zehavoc/status/1599193444043268096
https://twitter.com/yoavgo/status/1598360581496459265
https://twitter.com/yoavgo/status/1599037412411596800
https://twitter.com/yoavgo/status/1599045344863879168
https://twitter.com/natfriedman/status/1598477452661383168
https://twitter.com/conradev/status/1598487973351362561/photo/1
https://twitter.com/zswitten/status/1598100186605441024
https://twitter.com/CatEmbedded/status/1599141379879600128/photo/2
https://twitter.com/mattshumer_/status/1599175127148949505
https://twitter.com/vaibhavk97/status/1598930958769860608/photo/1
https://twitter.com/dan_abramov/status/1598800508160024588/photo/1
https://twitter.com/MinqiJiang/status/1598832656422432768/photo/2
https://twitter.com/zswitten/status/1598088280066920453
https://twitter.com/m1guelpf/status/1598203861294252033/photo/1
https://twitter.com/SilasAlberti/status/1598257908567117825/photo/1
https://twitter.com/gf_256/status/1598962842861899776/photo/1
https://twitter.com/zswitten/status/1598088267789787136
https://twitter.com/gf_256/status/1598178469955112961/photo/1
Your weekly news from the AI & Machine Learning world.
OUTLINE:
0:00 - Introduction
0:25 - AI reads brain signals to predict what you're thinking
3:00 - Closed-form solution for neuron interactions
4:15 - GPT-4 rumors
6:50 - Cerebras supercomputer
7:45 - Meta releases metagenomics atlas
9:15 - AI advances in theorem proving
10:40 - Better diffusion models with expert denoisers
12:00 - BLOOMZ & mT0
13:05 - ICLR reviewers going mad
21:40 - Scaling Transformer inference
22:10 - Infinite nature flythrough generation
23:55 - Blazing fast denoising
24:45 - Large-scale AI training with MultiRay
25:30 - arXiv to include Hugging Face spaces
26:10 - Multilingual Diffusion
26:30 - Music source separation
26:50 - Multilingual CLIP
27:20 - Drug response prediction
27:50 - Helpful Things
ERRATA:
HF did not acquire spaces, they launched spaces themselves and supported Gradio from the start. They later acquired Gradio.
References:
AI reads brain signals to predict what you're thinking
https://mind-vis.github.io/?s=09&utm_source=pocket_saves
https://neurosciencenews.com/bmi-internal-speech-21837/
Closed-form solution for neuron interactions
https://twitter.com/ramin_m_h/status/1592585672606769153/photo/1
https://github.com/raminmh/CfC
https://github.com/raminmh/CfC/blob/main/torch_cfc.py
GPT-4 rumors
https://thealgorithmicbridge.substack.com/p/gpt-4-rumors-from-silicon-valley?utm_source=pocket_reader
Cerebras supercomputer
https://www.cerebras.net/andromeda/
Meta releases metagenomics atlas
https://ai.facebook.com/blog/protein-folding-esmfold-metagenomics/
https://www.genome.gov/genetics-glossary/Metagenomics
AI advances in theorem proving
https://ai.facebook.com/blog/ai-math-theorem-proving/
https://marketplace.visualstudio.com/items?itemName=jroesch.lean
Better diffusion models with expert denoisers
https://deepimagination.cc/eDiffi/
BLOOMZ & mT0
https://arxiv.org/abs/2211.01786?utm_source=pocket_reader
https://huggingface.co/bigscience/bloomz?text=Suggest+at+least+five+related+search+terms+to+%22M%E1%BA%A1ng+neural+nh%C3%A2n+t%E1%BA%A1o%22.
ICLR reviewers going mad
https://twitter.com/XiangruTang/status/1589703605098975237?utm_source=pocket_reader
https://twitter.com/BlancheMinerva/status/1588164585961422849?utm_source=pocket_reader
https://openreview.net/forum?id=pfuqQQCB34
https://twitter.com/peter_richtarik/status/1591408710366408706?utm_source=pocket_reader
Scaling Transformer inference
https://arxiv.org/abs/2211.05102
Infinite nature flythrough generation
https://ai.googleblog.com/2022/11/infinite-nature-generating-3d.html?utm_source=pocket_reader
Blazing fast denoising
https://github.com/dome272/Paella
https://arxiv.org/abs/2211.07292
Large-scale AI training with MultiRay
https://ai.facebook.com/blog/multiray-large-scale-AI-models/
arXiv to include Hugging Face spaces
https://blog.arxiv.org/2022/11/17/discover-state-of-the-art-machine-learning-demos-on-arxiv/
Multilingual Diffusion
https://github.com/FlagAI-Open/FlagAI/tree/master/examples/AltDiffusion
Music source separation
https://github.com/facebookresearch/demucs
https://arxiv.org/abs/2211.08553
A team from Meta AI has developed Cicero, an agent that can play the game Diplomacy, in which players have to communicate via chat messages to coordinate and plan into the future.
Paper Title: Human-level play in the game of Diplomacy by combining language models with strategic reasoning
Commented game by human expert: https://www.youtube.com/watch?v=u5192bvUS7k
OUTLINE:
0:00 - Introduction
9:50 - AI in cooperation games
13:50 - Cicero agent overview
25:00 - A controllable dialogue model
36:50 - Dialogue-conditional strategic planning
49:00 - Message filtering
53:45 - Cicero's play against humans
55:15 - More examples & discussion
Homepage: https://ai.facebook.com/research/cicero/
Code: https://github.com/facebookresearch/diplomacy_cicero
Blog: https://ai.facebook.com/blog/cicero-ai-negotiates-persuades-and-cooperates-with-people/
Paper: https://www.science.org/doi/10.1126/science.ade9097
Abstract:
Despite much progress in training AI systems to imitate human language, building agents that use language to communicate intentionally with humans in interactive environments remains a major challenge. We introduce Cicero, the first AI agent to achieve human-level performance in Diplomacy, a strategy game involving both cooperation and competition that emphasizes natural language negotiation and tactical coordination between seven players. Cicero integrates a language model with planning and reinforcement learning algorithms by inferring players' beliefs and intentions from its conversations and generating dialogue in pursuit of its plans. Across 40 games of an anonymous online Diplomacy league, Cicero achieved more than double the average score of the human players and ranked in the top 10% of participants who played more than one game.
Authors: Anton Bakhtin, Noam Brown, Emily Dinan, Gabriele Farina, Colin Flaherty, Daniel Fried, Andrew Goff, Jonathan Gray, Hengyuan Hu, Athul Paul Jacob, Mojtaba Komeili, Karthik Konath, Minae Kwon, Adam Lerer, Mike Lewis, Alexander H. Miller, Sasha Mitts, Adithya Renduchintala, Stephen Roller, Dirk Rowe, Weiyan Shi, Joe Spisak, Alexander Wei, David Wu, Hugh Zhang, Markus Zijlstra
Links:
Homepage: https://ykilcher.com
Merch: https://ykilcher.com/merch
YouTube: https://www.youtube.com/c/yannickilcher
Twitter: https://twitter.com/ykilcher
Discord: https://ykilcher.com/discord
LinkedIn: https://www.linkedin.com/in/ykilcher
If you want to support me, the best thing to do is to share out the content :)
If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):
SubscribeStar: https://www.subscribestar.com/yannickilcher
Patreon: https://www.patreon.com/yannickilcher
Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq
Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2
Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m
Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n
Your news from the world of Machine Learning!
OUTLINE:
0:00 - Introduction
1:25 - Stable Diffusion Multiplayer
2:15 - Huggingface: DOI for Models & Datasets
3:10 - OpenAI asks for more funding
4:25 - The Stack: Source Code Dataset
6:30 - Google Vizier Open-Sourced
7:10 - New Models
11:50 - Helpful Things
20:30 - Prompt Databases
22:15 - Lexicap by Karpathy
References:
Stable Diffusion Multiplayer
https://huggingface.co/spaces/huggingface-projects/stable-diffusion-multiplayer?roomid=room-0
Huggingface: DOI for Models & Datasets
https://huggingface.co/blog/introducing-doi
OpenAI asks for more funding
https://www.theinformation.com/articles/openai-valued-at-nearly-20-billion-in-advanced-talks-with-microsoft-for-more-funding
https://www.wsj.com/articles/microsoft-in-advanced-talks-to-increase-investment-in-openai-11666299548
The Stack: Source Code Dataset
https://huggingface.co/datasets/bigcode/the-stack?utm_source=pocket_mylist
Google Vizier Open-Sourced
https://github.com/google/vizier
New Models
https://imagen.research.google/video/
https://phenaki.github.io/
https://makeavideo.studio/?utm_source=pocket_mylist
https://dreamfusion3d.github.io/
https://arxiv.org/pdf/2210.15257.pdf
https://huggingface.co/spaces/PaddlePaddle/ERNIE-ViLG
https://github.com/PaddlePaddle/PaddleHub
Helpful Things
https://thecharlieblake.co.uk/visualising-ml-number-formats
https://griddly.ai/
https://engineering.fb.com/2022/10/18/open-source/ocp-summit-2022-grand-teton/?utm_source=twitter&utm_medium=organic_social&utm_campaign=eng2022h2
https://twitter.com/psuraj28/status/1580640841583902720?utm_source=pocket_mylist
https://huggingface.co/blog/stable_diffusion_jax
https://github.com/Lightning-AI/stable-diffusion-deploy
https://lightning.ai/docs/stable/
https://github.com/CarperAI/trlx
https://github.com/DLR-RM/rl-baselines3-zoo
https://github.com/Sea-Snell/JAXSeq
https://www.reddit.com/r/MachineLearning/comments/xoitw9/p_albumentations_13_is_released_a_python_library/?utm_source=pocket_mylist
https://twitter.com/Warvito/status/1570691960792580096?utm_source=pocket_mylist
https://arxiv.org/abs/2209.07162
https://academictorrents.com/details/63aeb864bbe2115ded0aa0d7d36334c026f0660b
https://huggingface.co/spaces/THUDM/CodeGeeX
https://ai.facebook.com/blog/gpu-inference-engine-nvidia-amd-open-source/?utm_source=twitter&utm_medium=organic_social&utm_campaign=blog
https://github.com/nerfstudio-project/nerfstudio
https://www.nerfacc.com/en/latest/
https://github.com/dstackai/dstack
https://www.reddit.com/r/MachineLearning/comments/yeyxlo/p_openai_whisper_3x_cpu_inference_speedup/?utm_source=pocket_mylist
https://github.com/MiscellaneousStuff/openai-whisper-cpu/issues/1
Prompt Databases
https://huggingface.co/datasets/poloclub/diffusiondb
https://publicprompts.art/
https://visualise.ai/
https://twitter.com/SamuelAlbanie/status/1574111928431026179/photo/1
Lexicap by Karpathy
https://karpathy.ai/lexicap/0139-large.html
Links:
Homepage: https://ykilcher.com
Merch: https://ykilcher.com/merch
YouTube: https://www.youtube.com/c/yannickilcher
Twitter: https://twitter.com/ykilcher
Discord: https://ykilcher.com/discord
LinkedIn: https://www.linkedin.com/in/ykilcher
If you want to support me, the best thing to do is to share out the content :)
So-called responsible AI licenses are stupid, counterproductive, and have a dangerous legal loophole in them.
OpenRAIL++ License here: https://www.ykilcher.com/license
OUTLINE:
0:00 - Introduction
0:40 - Responsible AI Licenses (RAIL) of BLOOM and Stable Diffusion
3:35 - Open source software's dilemma of bad usage and restrictions
8:45 - Good applications, bad applications
12:45 - A dangerous legal loophole
15:50 - OpenRAIL++ License
16:50 - This has nothing to do with copyright
26:00 - Final thoughts
References:
https://huggingface.co/CompVis/stable-diffusion/tree/main
https://huggingface.co/spaces/CompVis/stable-diffusion-license
https://huggingface.co/bigscience/bloom?text=34%2B10%3D44+%0A54%2B20%3D
https://huggingface.co/spaces/bigscience/license
https://huggingface.co/runwayml/stable-diffusion-v1-5
https://huggingface.co/spaces/CompVis/stable-diffusion-license/raw/main/license.txt
https://www.gnu.org/philosophy/programs-must-not-limit-freedom-to-run.en.html
https://www.gnu.org/philosophy/free-sw.html#four-freedoms
https://www.licenses.ai/blog/2022/8/26/bigscience-open-rail-m-license
https://bigscience.huggingface.co/blog/bigscience-ethical-charter
https://www.licenses.ai/blog/2022/8/18/naming-convention-of-responsible-ai-licenses
https://en.wikipedia.org/wiki/Copyright#Eligible_works
https://en.wikipedia.org/wiki/Creative_work
https://www.pearlcohen.com/copyright-office-reiterates-that-works-created-by-ai-cannot-be-copyrighted/
https://jipel.law.nyu.edu/vol-8-no-2-1-hedrick/#II
https://www.ykilcher.com/license
Links:
Homepage: https://ykilcher.com
Merch: https://ykilcher.com/merch
YouTube: https://www.youtube.com/c/yannickilcher
Twitter: https://twitter.com/ykilcher
Discord: https://ykilcher.com/discord
LinkedIn: https://www.linkedin.com/in/ykilcher
If you want to support me, the best thing to do is to share out the content :)
If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):
SubscribeStar: https://www.subscribestar.com/yannickilcher
Patreon: https://www.patreon.com/yannickilcher
Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq
Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2
Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m
Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n
Large Language Models have the ability to store vast amounts of facts about the world. But little is known, how these models actually do this. This paper aims at discovering the mechanism and location of storage and recall of factual associations in GPT models, and then proposes a mechanism for the targeted editing of such facts, in form of a simple rank-one update to a single MLP layer. This has wide implications both for how we understand such models' inner workings, and for our ability to gain greater control over such models in the future.
OUTLINE:
0:00 - Introduction
1:40 - What are the main questions in this subfield?
6:55 - How causal tracing reveals where facts are stored
18:40 - Clever experiments show the importance of MLPs
24:30 - How do MLPs store information?
29:10 - How to edit language model knowledge with precision?
36:45 - What does it mean to know something?
39:00 - Experimental Evaluation & the CounterFact benchmark
45:40 - How to obtain the required latent representations?
51:15 - Where is the best location in the model to perform edits?
58:00 - What do these models understand about language?
1:02:00 - Questions for the community
Paper: https://arxiv.org/abs/2202.05262
Follow-up paper on Mass-Editing Memory in a Transformer: https://arxiv.org/abs/2210.07229
Abstract:
We analyze the storage and recall of factual associations in autoregressive transformer language models, finding evidence that these associations correspond to localized, directly-editable computations. We first develop a causal intervention for identifying neuron activations that are decisive in a model's factual predictions. This reveals a distinct set of steps in middle-layer feed-forward modules that mediate factual predictions while processing subject tokens. To test our hypothesis that these computations correspond to factual association recall, we modify feed-forward weights to update specific factual associations using Rank-One Model Editing (ROME). We find that ROME is effective on a standard zero-shot relation extraction (zsRE) model-editing task, comparable to existing methods. To perform a more sensitive evaluation, we also evaluate ROME on a new dataset of counterfactual assertions, on which it simultaneously maintains both specificity and generalization, whereas other methods sacrifice one or another. Our results confirm an important role for mid-layer feed-forward modules in storing factual associations and suggest that direct manipulation of computational mechanisms may be a feasible approach for model editing. The code, dataset, visualizations, and an interactive demo notebook are available at this https URL
Authors: Kevin Meng, David Bau, Alex Andonian, Yonatan Belinkov
Links:
Homepage: https://ykilcher.com
Merch: https://ykilcher.com/merch
YouTube: https://www.youtube.com/c/yannickilcher
Twitter: https://twitter.com/ykilcher
Discord: https://ykilcher.com/discord
LinkedIn: https://www.linkedin.com/in/ykilcher
If you want to support me, the best thing to do is to share out the content :)
If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):
SubscribeStar: https://www.subscribestar.com/yannickilcher
Patreon: https://www.patreon.com/yannickilcher
Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq
Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2
Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m
Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n
Alexander Mattick joins me to discuss the paper "Neural Networks are Decision Trees", which has generated a lot of hype on social media. We ask the question: Has this paper solved one of the large mysteries of deep learning and opened the black-box neural networks up to interpretability?
OUTLINE:
0:00 - Introduction
2:20 - Aren't Neural Networks non-linear?
5:20 - What does it all mean?
8:00 - How large do these trees get?
11:50 - Decision Trees vs Neural Networks
17:15 - Is this paper new?
22:20 - Experimental results
27:30 - Can Trees and Networks work together?
Paper: https://arxiv.org/abs/2210.05189
Abstract:
In this manuscript, we show that any feedforward neural network having piece-wise linear activation functions can be represented as a decision tree. The representation is equivalence and not an approximation, thus keeping the accuracy of the neural network exactly as is. We believe that this work paves the way to tackle the black-box nature of neural networks. We share equivalent trees of some neural networks and show that besides providing interpretability, tree representation can also achieve some computational advantages. The analysis holds both for fully connected and convolutional networks, which may or may not also include skip connections and/or normalizations.
Author: Caglar Aytekin
Links:
Homepage: https://ykilcher.com
Merch: https://ykilcher.com/merch
YouTube: https://www.youtube.com/c/yannickilcher
Twitter: https://twitter.com/ykilcher
Discord: https://ykilcher.com/discord
LinkedIn: https://www.linkedin.com/in/ykilcher
If you want to support me, the best thing to do is to share out the content :)
If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):
SubscribeStar: https://www.subscribestar.com/yannickilcher
Patreon: https://www.patreon.com/yannickilcher
Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq
Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2
Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m
Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n
Matrix multiplication is the most used mathematical operation in all of science and engineering. Speeding this up has massive consequences. Thus, over the years, this operation has become more and more optimized. A fascinating discovery was made when it was shown that one actually needs less than N^3 multiplication operations to multiply to NxN matrices. DeepMind goes a step further and creates AlphaTensor, a Deep Reinforcement Learning algorithm that plays a single-player game, TensorGame, in order to find even more optimized algorithms for matrix multiplication. And it turns out, there exists a plethora of undiscovered matrix multiplication algorithms, which not only will make everything from computers to smart toasters faster, but also bring new insights into fundamental math and complexity theory.
Sponsor: Assembly AI
Link: https://www.assemblyai.com/?utm_source=youtube&utm_medium=social&utm_campaign=yannic_sentiment
OUTLINE:
0:00 - Intro
1:50 - Sponsor: Assembly AI (link in description)
3:25 - What even is Matrix Multiplication?
6:10 - A very astounding fact
8:45 - Trading multiplications for additions
12:35 - Matrix Multiplication as a Tensor
17:30 - Tensor Decompositions
20:30 - A formal way of finding multiplication algorithms
31:00 - How to formulate this as a game?
39:30 - A brief primer on AlphaZero / MCTS
45:40 - The Results
48:15 - Optimizing for different hardware
52:40 - Expanding fundamental math
53:45 - Summary & Final Comments
Paper: https://www.nature.com/articles/s41586-022-05172-4
Title: Discovering faster matrix multiplication algorithms with reinforcement learning
Abstract:
Improving the efficiency of algorithms for fundamental computations can have a widespread impact, as it can affect the overall speed of a large amount of computations. Matrix multiplication is one such primitive task, occurring in many systems—from neural networks to scientific computing routines. The automatic discovery of algorithms using machine learning offers the prospect of reaching beyond human intuition and outperforming the current best human-designed algorithms. However, automating the algorithm discovery procedure is intricate, as the space of possible algorithms is enormous. Here we report a deep reinforcement learning approach based on AlphaZero1 for discovering efficient and provably correct algorithms for the multiplication of arbitrary matrices. Our agent, AlphaTensor, is trained to play a single-player game where the objective is finding tensor decompositions within a finite factor space. AlphaTensor discovered algorithms that outperform the state-of-the-art complexity for many matrix sizes. Particularly relevant is the case of 4 × 4 matrices in a finite field, where AlphaTensor’s algorithm improves on Strassen’s two-level algorithm for the first time, to our knowledge, since its discovery 50 years ago2. We further showcase the flexibility of AlphaTensor through different use-cases: algorithms with state-of-the-art complexity for structured matrix multiplication and improved practical efficiency by optimizing matrix multiplication for runtime on specific hardware. Our results highlight AlphaTensor’s ability to accelerate the process of algorithmic discovery on a range of problems, and to optimize for different criteria.
Authors: Alhussein Fawzi, Matej Balog, Aja Huang, Thomas Hubert, Bernardino Romera-Paredes, Mohammadamin Barekatain, Alexander Novikov, Francisco J. R. Ruiz, Julian Schrittwieser, Grzegorz Swirszcz, David Silver, Demis Hassabis & Pushmeet Kohli
Stable Diffusion has been released and is riding a wave of creativity and collaboration. But not everyone is happy about this...
Sponsor: NVIDIA
GPU Raffle: https://ykilcher.com/gtc
OUTLINE:
0:00 - Introduction
0:30 - What is Stable Diffusion?
2:25 - Open-Source Contributions and Creations
7:55 - Textual Inversion
9:30 - OpenAI vs Open AI
14:20 - Journalists be outraged
16:20 - AI Ethics be even more outraged
19:45 - Do we need a new social contract?
21:30 - More applications
22:55 - Helpful Things
23:45 - Sponsor: NVIDIA (& how to enter the GPU raffle)
References: https://early-hair-c20.notion.site/Stable-Diffusion-Takes-Over-Referenes-7a2f45b8f7e04ae0ba19dbfcd2b7f7c0
Links:
Homepage: https://ykilcher.com
Merch: https://ykilcher.com/merch
YouTube: https://www.youtube.com/c/yannickilcher
Twitter: https://twitter.com/ykilcher
Discord: https://ykilcher.com/discord
LinkedIn: https://www.linkedin.com/in/ykilcher
If you want to support me, the best thing to do is to share out the content :)
If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):
SubscribeStar: https://www.subscribestar.com/yannickilcher
Patreon: https://www.patreon.com/yannickilcher
Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq
Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2
Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m
Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n
Sparsity is awesome, but only recently has it become possible to properly handle sparse models at good performance. Neural Magic does exactly this, using a plain CPU. No specialized hardware needed, just clever algorithms for pruning and forward-propagation of neural networks. Nir Shavit and I talk about how this is possible, what it means in terms of applications, and why sparsity should play a much larger role in the Deep Learning community.
Sponsor: AssemblyAI
Link: https://www.assemblyai.com/?utm_sourc...
Check out Neural Magic: https://neuralmagic.com/
and DeepSparse: https://github.com/neuralmagic/deepsp...
OUTLINE:
0:00 Introduction
1:08 Sponsor: AssemblyAI
2:50 Start of Interview
4:15 How the NIR company was founded?
5:10 What is Sparsity about?
9:30 Link between the human brain and sparsity
12:10 Where should the extra resource that the human brain doesn't have go?
14:40 Analogy for Sparse Architecture
16:48 Possible future for Sparse Architecture as standard architure for Neural Networks
20:08 Pruning & Sparsification
22:57 What keeps us from building sparse models?
25:34 Why are GPUs so unsuited for sparse models?
28:47 CPU and GPU in connection with memory
30:14 What Neural Magic does?
32:54 How do you deal with overlaps in tensor columns?
33:41 The best type of sparsity to execute tons of CPU
37:24 What kind of architecture would make the best use out of a combined system of CPUs and GPUs?
41:04 Graph Neural Networks in connection to sparsity
43:04 Intrinsic connection between the Sparsification of Neural Networks, Non Layer-Wise Computation, Blockchain Technology, Smart Contracts and Distributed Computing
45:23 Neural Magic's target audience
48:16 Is there a type of model where it works particularly well and the type where it doesn't?
Links:
Homepage: https://ykilcher.com
Merch: https://ykilcher.com/merch
YouTube:
/ yannickilcher
Twitter: https://twitter.com/ykilcher
Discord: https://ykilcher.com/discord
LinkedIn: https://www.linkedin.com/in/ykilcher
If you want to support me, the best thing to do is to share out the content :)
If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):
SubscribeStar: https://www.subscribestar.com/yannick...
Patreon: https://www.patreon.com/yannickilcher
Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq
Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2
Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m
Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n
Jacob Steinhardt believes that future AI systems will be qualitatively different than the ones we know currently. We talk about how emergence happens when scaling up, what implications that has on AI Safety, and why thought experiments like the Paperclip Maximizer might be more useful than most people think.
OUTLINE:
0:00 Introduction
1:10 Start of Interview
2:10 Blog posts series
3:56 More Is Different for AI (Blog Post)
7:40 Do you think this emergence is mainly a property from the interaction of things?
9:17 How does phase transition or scaling-up play into AI and Machine Learning?
12:10 GPT-3 as an example of qualitative difference in scaling up
14:08 GPT-3 as an emergent phenomenon in context learning
15:58 Brief introduction of different viewpoints on the future of AI and its alignment
18:51 How does the phenomenon of emergence play into this game between the Engineering and the Philosophy viewpoint?
22:41 Paperclip Maximizer on AI safety and alignment
31:37 Thought Experiments
37:34 Imitative Deception
39:30 TruthfulQA: Measuring How Models Mimic Human Falsehoods (Paper)
42:24 ML Systems Will Have Weird Failure Models (Blog Post)
51:10 Is there any work to get a system to be deceptive?
54:37 Empirical Findings Generalize Surprisingly Far (Blog Post)
1:00:18 What would you recommend to guarantee better AI alignment or safety?
1:05:13 Remarks
References:
https://bounded-regret.ghost.io/more-is-different-for-ai/
https://docs.google.com/document/d/1FbTuRvC4TFWzGYerTKpBU7FJlyvjeOvVYF2uYNFSlOc/edit#heading=h.n1wk9bxo847o
Links:
TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick
YouTube: https://www.youtube.com/c/yannickilcher
Twitter: https://twitter.com/ykilcher
Discord: https://ykilcher.com/discord
BitChute: https://www.bitchute.com/channel/yannic-kilcher
LinkedIn: https://www.linkedin.com/in/ykilcher
BiliBili: https://space.bilibili.com/2017636191
If you want to support me, the best thing to do is to share out the content :)
If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):
SubscribeStar: https://www.subscribestar.com/yannickilcher
Patreon: https://www.patreon.com/yannickilcher
Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq
Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2
Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m
Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n
Did you know that something as simple as loading a model can execute arbitrary code on your machine?
Try the model: https://huggingface.co/ykilcher/total...
Get the code: https://github.com/yk/patch-torch-save
Sponsor: Weights & Biases
Go here: https://wandb.me/yannic
OUTLINE:
0:00 - Introduction
1:10 - Sponsor: Weights & Biases
3:20 - How Hugging Face models are loaded
5:30 - From PyTorch to pickle
7:10 - Understanding how pickle saves data
13:00 - Executing arbitrary code
15:05 - The final code
17:25 - How can you protect yourself?
Links:
Homepage: https://ykilcher.com
Merch: https://ykilcher.com/merch
YouTube: https://www.youtube.com/c/yannickilcher
Twitter: https://twitter.com/ykilcher
Discord: https://ykilcher.com/discord
LinkedIn: https://www.linkedin.com/in/ykilcher
If you want to support me, the best thing to do is to share out the content :)
If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):
SubscribeStar: https://www.subscribestar.com/yannick...
Patreon: https://www.patreon.com/yannickilcher
Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq
Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2
Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m
Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n
#ai #selforganization #emergence
Read Sebastian's article here: https://sebastianrisi.com/self_assemb...
OUTLINE:
0:00 - Introduction
2:25 - Start of Interview
4:00 - The intelligence of swarms
9:15 - The game of life & neural cellular automata
14:10 - What's missing from neural CAs?
17:20 - How does local computation compare to centralized computation?
25:40 - Applications beyond games and graphics
33:00 - Can we do away with goals?
35:30 - Where do these methods shine?
43:30 - The paradox of scales & brains
49:45 - Connections to graphical systems & GNNs
51:30 - Could this solve ARC?
57:45 - Where can people get started?
References:
https://sebastianrisi.com/
https://modl.ai/
https://sebastianrisi.com/self_assemb...
https://twitter.com/risi1979/status/1...
https://distill.pub/2020/growing-ca/
https://arxiv.org/abs/2201.12360?sour...
https://distill.pub/2020/selforg/mnist/
https://arxiv.org/pdf/2204.11674.pdf
https://github.com/fchollet/ARC
https://github.com/volotat/ARC-Game
http://animalaiolympics.com/AAI/
https://www.deepmind.com/publications...
https://melaniemitchell.me/BooksConte...
Links:
Homepage: https://ykilcher.com
Merch: https://ykilcher.com/merch
YouTube: https://www.youtube.com/c/yannickilcher
Twitter: https://twitter.com/ykilcher
Discord: https://ykilcher.com/discord
LinkedIn: https://www.linkedin.com/in/ykilcher
If you want to support me, the best thing to do is to share out the content :)
If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):
SubscribeStar: https://www.subscribestar.com/yannick...
Patreon: https://www.patreon.com/yannickilcher
Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq
Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2
Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m
Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n
#stablediffusion #ai #stabilityai
An interview with Emad Mostaque, founder of Stability AI.
OUTLINE:
0:00 - Intro
1:30 - What is Stability AI?
3:45 - Where does the money come from?
5:20 - Is this the CERN of AI?
6:15 - Who gets access to the resources?
8:00 - What is Stable Diffusion?
11:40 - What if your model produces bad outputs?
14:20 - Do you employ people?
16:35 - Can you prevent the corruption of profit?
19:50 - How can people find you?
22:45 - Final thoughts, let's destroy PowerPoint
Links:
Homepage: https://ykilcher.com
Merch: https://ykilcher.com/merch
YouTube: https://www.youtube.com/c/yannickilcher
Twitter: https://twitter.com/ykilcher
Discord: https://ykilcher.com/discord
LinkedIn: https://www.linkedin.com/in/ykilcher
If you want to support me, the best thing to do is to share out the content :)
If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):
SubscribeStar: https://www.subscribestar.com/yannick...
Patreon: https://www.patreon.com/yannickilcher
Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq
Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2
Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m
Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n
Today we look at all the recent giant language models in the AI world!
OUTLINE:
0:00 - Intro
0:55 - BLOOM: Open-Source 176B Language Model
5:25 - YALM 100B
5:40 - Chinese Brain-Scale Supercomputer
7:25 - Meta AI Translates over 200 Languages
10:05 - Reproducibility Crisis Workshop
10:55 - AI21 Raises $64M
11:50 - Ian Goodfellow leaves Apple
12:20 - Andrej Karpathy leaves Tesla
12:55 - Wordalle
References:
BLOOM: Open-Source 176B Language Model
https://bigscience.huggingface.co/blo...
https://huggingface.co/spaces/bigscie...
https://huggingface.co/bigscience/blo...
YALM 100B
https://github.com/yandex/YaLM-100B
Chinese Brain-Scale Supercomputer
https://www.scmp.com/news/china/scien...
https://archive.ph/YaoA6#selection-12...
Meta AI Translates over 200 Languages
https://ai.facebook.com/research/no-l...
Reproducibility Crisis Workshop
https://reproducible.cs.princeton.edu/
AI21 Raises $64M
https://techcrunch.com/2022/07/12/ope...
Ian Goodfellow leaves Apple
https://twitter.com/goodfellow_ian/st...
Andrey Karpathy leaves Tesla
https://mobile.twitter.com/karpathy/s...
https://www.businessinsider.com/repor...
Wordalle
https://huggingface.co/spaces/hugging...
Links:
Homepage: https://ykilcher.com
Merch: ykilcher.com/merch
YouTube: https://www.youtube.com/c/yannickilcher
Twitter: https://twitter.com/ykilcher
Discord: https://ykilcher.com/discord
LinkedIn: https://www.linkedin.com/in/ykilcher
If you want to support me, the best thing to do is to share out the content :)
If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):
SubscribeStar: https://www.subscribestar.com/yannick...
Patreon: https://www.patreon.com/yannickilcher
Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq
Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2
Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m
Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n
Yann LeCun's position paper on a path towards machine intelligence combines Self-Supervised Learning, Energy-Based Models, and hierarchical predictive embedding models to arrive at a system that can teach itself to learn useful abstractions at multiple levels and use that as a world model to plan ahead in time.
OUTLINE:
0:00 - Introduction
2:00 - Main Contributions
5:45 - Mode 1 and Mode 2 actors
15:40 - Self-Supervised Learning and Energy-Based Models
20:15 - Introducing latent variables
25:00 - The problem of collapse
29:50 - Contrastive vs regularized methods
36:00 - The JEPA architecture
47:00 - Hierarchical JEPA (H-JEPA)
53:00 - Broader relevance
56:00 - Summary & Comments
Paper: https://openreview.net/forum?id=BZ5a1...
Abstract: How could machines learn as efficiently as humans and animals? How could machines learn to reason and plan? How could machines learn representations of percepts and action plans at multiple levels of abstraction, enabling them to reason, predict, and plan at multiple time horizons? This position paper proposes an architecture and training paradigms with which to construct autonomous intelligent agents. It combines concepts such as configurable predictive world model, behavior driven through intrinsic motivation, and hierarchical joint embedding architectures trained with self-supervised learning.
Author: Yann LeCun
Links:
Homepage: https://ykilcher.com
Merch: https://ykilcher.com/merch
YouTube: https://www.youtube.com/c/yannickilcher
Twitter: https://twitter.com/ykilcher
Discord: https://ykilcher.com/discord
LinkedIn: https://www.linkedin.com/in/ykilcher
If you want to support me, the best thing to do is to share out the content :)
If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):
SubscribeStar: https://www.subscribestar.com/yannick...
Patreon: https://www.patreon.com/yannickilcher
Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq
Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2
Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m
Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n
Minecraft is one of the harder challenges any RL agent could face. Episodes are long, and the world is procedurally generated, complex, and huge. Further, the action space is a keyboard and a mouse, which has to be operated only given the game's video input. OpenAI tackles this challenge using Video PreTraining, leveraging a small set of contractor data in order to pseudo-label a giant corpus of scraped footage of gameplay. The pre-trained model is highly capable in basic game mechanics and can be fine-tuned much better than a blank slate model. This is the first Minecraft agent that achieves the elusive goal of crafting a diamond pickaxe all by itself.
OUTLINE:
0:00 - Intro
3:50 - How to spend money most effectively?
8:20 - Getting a large dataset with labels
14:40 - Model architecture
19:20 - Experimental results and fine-tuning
25:40 - Reinforcement Learning to the Diamond Pickaxe
30:00 - Final comments and hardware
Blog: https://openai.com/blog/vpt/
Paper: https://arxiv.org/abs/2206.11795
Code & Model weights: https://github.com/openai/Video-Pre-T...
Abstract:
Pretraining on noisy, internet-scale datasets has been heavily studied as a technique for training models with broad, general capabilities for text, images, and other modalities. However, for many sequential decision domains such as robotics, video games, and computer use, publicly available data does not contain the labels required to train behavioral priors in the same way. We extend the internet-scale pretraining paradigm to sequential decision domains through semi-supervised imitation learning wherein agents learn to act by watching online unlabeled videos. Specifically, we show that with a small amount of labeled data we can train an inverse dynamics model accurate enough to label a huge unlabeled source of online data -- here, online videos of people playing Minecraft -- from which we can then train a general behavioral prior. Despite using the native human interface (mouse and keyboard at 20Hz), we show that this behavioral prior has nontrivial zero-shot capabilities and that it can be fine-tuned, with both imitation learning and reinforcement learning, to hard-exploration tasks that are impossible to learn from scratch via reinforcement learning. For many tasks our models exhibit human-level performance, and we are the first to report computer agents that can craft diamond tools, which can take proficient humans upwards of 20 minutes (24,000 environment actions) of gameplay to accomplish.
Authors: Bowen Baker, Ilge Akkaya, Peter Zhokhov, Joost Huizinga, Jie Tang, Adrien Ecoffet, Brandon Houghton, Raul Sampedro, Jeff Clune
Links:
Homepage: https://ykilcher.com
Merch: https://ykilcher.com/merch
YouTube: https://www.youtube.com/c/yannickilcher
Twitter: https://twitter.com/ykilcher
Discord: https://ykilcher.com/discord
LinkedIn: https://www.linkedin.com/in/ykilcher
If you want to support me, the best thing to do is to share out the content :)
If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):
SubscribeStar: https://www.subscribestar.com/yannick...
Patreon: https://www.patreon.com/yannickilcher
Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq
Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2
Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m
Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n
#parti #ai #aiart
Parti is a new autoregressive text-to-image model that shows just how much scale can achieve. This model's outputs are crips, accurate, realistic, and can combine arbitrary styles, concepts, and fulfil even challenging requests.
OUTLINE:
0:00 - Introduction
2:40 - Example Outputs
6:00 - Model Architecture
17:15 - Datasets (incl. PartiPrompts)
21:45 - Experimental Results
27:00 - Picking a cherry tree
29:30 - Failure cases
33:20 - Final comments
Website: https://parti.research.google/
Paper: https://arxiv.org/abs/2206.10789
Github: https://github.com/google-research/parti
Links:
Homepage: https://ykilcher.com
Merch: https://ykilcher.com/merch
YouTube: https://www.youtube.com/c/yannickilcher
Twitter: https://twitter.com/ykilcher
Discord: https://ykilcher.com/discord
LinkedIn: https://www.linkedin.com/in/ykilcher
If you want to support me, the best thing to do is to share out the content :)
If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):
SubscribeStar: https://www.subscribestar.com/yannick...
Patreon: https://www.patreon.com/yannickilcher
Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq
Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2
Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m
Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n
Google engineer Blake Lemoine was put on leave after releasing proprietary information: An interview with the chatbot LaMDA that he believes demonstrates that this AI is, in fact, sentient. We analyze the claims and the interview in detail and trace how a statistical machine managed to convince at least one human that it is more than just an algorithm.
OUTLINE:
0:00 - Whistleblower put on leave
4:30 - What is a language model?
6:40 - The prompt is the key
10:40 - Who are we talking to exactly?
12:50 - LaMDA analyzes stories
15:20 - Fear, pain, and consent
20:25 - How would we recognize sentience? When is a machine conscious?
References:
https://cajundiscordian.medium.com/is-lamda-sentient-an-interview-ea64d916d917
https://cajundiscordian.medium.com/what-is-lamda-and-what-does-it-want-688632134489
https://www.washingtonpost.com/technology/2022/06/11/google-ai-lamda-blake-lemoine/
https://www.theguardian.com/technology/2022/jun/12/google-engineer-ai-bot-sentient-blake-lemoine
https://www.businessinsider.com/transcript-of-sentient-google-ai-chatbot-was-edited-for-readability-2022-6?inline-endstory-related-recommendations=&r=US&IR=T
Links:
Homepage: https://ykilcher.com
Merch: https://ykilcher.com/merch
YouTube: https://www.youtube.com/c/yannickilcher
Twitter: https://twitter.com/ykilcher
Discord: https://ykilcher.com/discord
BitChute: https://www.bitchute.com/channel/yannic-kilcher
LinkedIn: https://www.linkedin.com/in/ykilcher
If you want to support me, the best thing to do is to share out the content :)
If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):
SubscribeStar: https://www.subscribestar.com/yannickilcher
Patreon: https://www.patreon.com/yannickilcher
Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq
Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2
Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m
Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n
Your updates directly from the state of the art in Machine Learning!
OUTLINE:
0:00 - Intro
0:30 - DeepMind's Flamingo: Unified Vision-Language Model
8:25 - LiT: Locked Image Tuning
10:20 - Jurassic X & MRKL Systems
15:05 - Helpful Things
22:40 - This AI does not exist
References:
DeepMind's Flamingo: Unified Vision-Language Model
https://www.deepmind.com/blog/tacklin...
https://storage.googleapis.com/deepmi...
https://twitter.com/Inoryy/status/152...
LiT: Locked Image Tuning
https://ai.googleblog.com/2022/04/loc...
https://google-research.github.io/vis...
Jurassic X & MRKL Systems
https://www.ai21.com/blog/jurassic-x-...
https://arxiv.org/pdf/2205.00445.pdf
https://arxiv.org/pdf/2204.10019.pdf
https://studio.ai21.com/jurassic-x
StyleGAN Human
https://stylegan-human.github.io/
https://github.com/stylegan-human/Sty...
https://huggingface.co/spaces/hysts/S...
Helpful Things
https://github.com/rish-16/grafog
https://huggingface.co/bertin-project...
https://github.com/pytorch/torchdistx
https://pytorch.org/torchdistx/latest...
https://github.com/Netflix/vectorflow...
https://iclr-blog-track.github.io/202...
https://twitter.com/DeepMind/status/1...
https://github.com/ai-forever/mgpt
https://github.com/cleanlab/cleanlab
https://efficientdlbook.com/?utm_sour...
https://minihack-editor.github.io/
https://mugen-org.github.io/
https://www.amazon.science/blog/amazo...
https://github.com/phuselab/openFACS?...
https://medium.com/pytorch/avalanche-...
This AI does not exist
https://thisaidoesnotexist.com/
Links:
Merch: https://ykilcher.com/merch
TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick
YouTube: https://www.youtube.com/c/yannickilcher
Twitter: https://twitter.com/ykilcher
Discord: https://discord.gg/4H8xxDF
BitChute: https://www.bitchute.com/channel/yann...
LinkedIn: https://www.linkedin.com/in/ykilcher
BiliBili: https://space.bilibili.com/2017636191
If you want to support me, the best thing to do is to share out the content :)
If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):
SubscribeStar: https://www.subscribestar.com/yannick...
Patreon: https://www.patreon.com/yannickilcher
Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq
Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2
Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m
Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n
An inside look of what's happening in the ML world!
Sponsor: Weights & Biases
https://wandb.me/yannic
OUTLINE:
0:00 - Intro
0:20 - Sponsor: Weights & Biases
1:40 - Meta AI releases OPT-175B
4:55 - CoCa: New CLIP-Competitor
8:15 - DALL-E Mega is training
10:05 - TorToiSe TTS is amazing!
11:50 - Investigating Vision Transformers
12:50 - Hugging Face Deep RL class launched
13:40 - Helpful Things
17:00 - John Deere's driverless tractors
References:
Meta AI releases OPT-175B
https://ai.facebook.com/blog/democratizing-access-to-large-scale-language-models-with-opt-175b/
https://arxiv.org/abs/2205.01068
https://arxiv.org/pdf/2205.01068.pdf
https://github.com/facebookresearch/metaseq/tree/main/projects/OPT
https://github.com/facebookresearch/metaseq/blob/main/projects/OPT/chronicles/OPT175B_Logbook.pdf
https://github.com/facebookresearch/metaseq/tree/main/projects/OPT/chronicles
https://twitter.com/yoavgo/status/1522150063815987201
CoCa: New CLIP-Competitor
https://arxiv.org/abs/2205.01917
https://arxiv.org/pdf/2205.01917.pdf
DALL-E Mega is training
https://twitter.com/borisdayma
https://twitter.com/borisdayma/status/1521891895001112577
https://wandb.ai/dalle-mini/dalle-mini/reports/DALL-E-Mega--VmlldzoxODMxMDI2
TorToiSe TTS is amazing!
https://github.com/neonbjb/tortoise-tts
https://nonint.com/static/tortoise_v2_examples.html
https://colab.research.google.com/drive/1wVVqUPqwiDBUVeWWOUNglpGhU3hg_cbR
https://github.com/neonbjb
Investigating Vision Transformers
https://github.com/sayakpaul/probing-vits/?utm_source=pocket_mylist
https://twitter.com/RisingSayak/status/1515918406171914240?utm_source=pocket_mylist
https://keras.io/examples/vision/probing_vits/
https://github.com/sayakpaul/probing-vits/tree/main/notebooks?utm_source=pocket_mylist
Hugging Face Deep RL class launched
https://github.com/huggingface/deep-rl-class
Helpful Things
https://merantix-momentum.com/technology/squirrel/?utm_source=pocket_mylist
https://github.com/merantix-momentum/squirrel-core?utm_source=pocket_mylist
https://pyscript.net/?utm_source=pocket_mylist
https://github.com/google-research/big_vision
https://deepsportradar.github.io/challenge.html
https://github.com/DeepSportRadar/camera-calibration-challenge
https://twitter.com/alekseykorshuk/status/1515989357961920514?utm_source=pocket_mylist
https://github.com/AlekseyKorshuk/huggingnft
John Deere's driverless tractors
https://thenextweb.com/news/john-deere-slowly-becoming-one-worlds-most-important-ai-companies
https://tractorhacking.github.io/
Links:
Merch: https://ykilcher.com/merch
TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick
YouTube: https://www.youtube.com/c/yannickilcher
Twitter: https://twitter.com/ykilcher
Discord: https://discord.gg/4H8xxDF
BitChute: https://www.bitchute.com/channel/yannic-kilcher
LinkedIn: https://www.linkedin.com/in/ykilcher
BiliBili: https://space.bilibili.com/2017636191
If you want to support me, the best thing to do is to share out the content :)
Today we build our own AI that can create as many bored apes as we want! Fungibility for everyone!
Try the model here: https://huggingface.co/spaces/ykilcher/apes
or here: https://ykilcher.com/apes
Files & Models here: https://huggingface.co/ykilcher/apes/tree/main
Code here: https://github.com/yk/apes-public (for the "what's your ape" app, look for the file interface_projector.py)
This video is sponsored by BrightData, use this link for free credits:
https://brightdata.grsm.io/yannickilcher
OUTLINE:
0:00 - Introduction
2:05 - Generative Adversarial Networks
3:40 - Scraping Opensea with BrightData
7:55 - Training the GAN
11:35 - Here are the results!
15:20 - Diving deeper into BrightData
References:
Stylegan 3 imagery: https://nvlabs.github.io/stylegan3/
Bored Ape Yacht Club NFT Collection: https://opensea.io/collection/boredapeyachtclub
Better GANFT model: https://medium.com/@nathancooperjones/these-bored-apes-do-not-exist-6bed2c73f02c
Abstract AI-created apes: https://opensea.io/collection/gan-apes-nft
https://mobile.twitter.com/gannft
Another good model: https://twitter.com/cyrilzakka/status/1463944040878071811
StyleGAN2 versions: https://thispersondoesnotexist.com/
https://thissneakerdoesnotexist.com/
https://thischairdoesnotexist.com/
GANs: https://en.wikipedia.org/wiki/Generative_adversarial_network
https://arxiv.org/pdf/1406.2661.pdf
StyleGAN3: https://nvlabs.github.io/stylegan3/
StyleGAN2 code: https://github.com/NVlabs/stylegan2-ada-pytorch
CLIP: https://openai.com/blog/clip/
DALL-E 2 images: https://twitter.com/search?q=%23dalle&f=image
My music video: https://www.youtube.com/watch?v=2iq7WXSw26s
BrightData Links: https://brightdata.com/products/data-collector
https://brightdata.com/testimonials
https://brightdata.com/use-cases/adtech
https://brightdata.com/use-cases/social-media-for-marketing
https://brightdata.com/use-cases/ecommerce
Links:
Merch: https://ykilcher.com/merch
TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick
YouTube: https://www.youtube.com/c/yannickilcher
Twitter: https://twitter.com/ykilcher
Discord: https://discord.gg/4H8xxDF
BitChute: https://www.bitchute.com/channel/yannic-kilcher
LinkedIn: https://www.linkedin.com/in/ykilcher
BiliBili: https://space.bilibili.com/2017636191
If you want to support me, the best thing to do is to share out the content :)
If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this), find options at https://ykilcher.com
This is an interview with the authors Brian Ichter, Karol Hausman, and Fei Xia.
Original Paper Review Video: https://youtu.be/Ru23eWAQ6_E
Large Language Models are excellent at generating plausible plans in response to real-world problems, but without interacting with the environment, they have no abilities to estimate which of these plans are feasible or appropriate. SayCan combines the semantic capabilities of language models with a bank of low-level skills, which are available to the agent as individual policies to execute. SayCan automatically finds the best policy to execute by considering a trade-off between the policy's ability to progress towards the goal, given by the language model, and the policy's probability of executing successfully, given by the respective value function. The result is a system that can generate and execute long-horizon action sequences in the real world to fulfil complex tasks.
OUTLINE:
0:00 - Introduction & Setup
3:40 - Acquiring atomic low-level skills
7:45 - How does the language model come in?
11:45 - Why are you scoring instead of generating?
15:20 - How do you deal with ambiguity in language?
20:00 - The whole system is modular
22:15 - Going over the full algorithm
23:20 - What if an action fails?
24:30 - Debunking a marketing video :)
27:25 - Experimental Results
32:50 - The insane scale of data collection
40:15 - How do you go about large-scale projects?
43:20 - Where did things go wrong?
45:15 - Where do we go from here?
52:00 - What is the largest unsolved problem in this?
53:35 - Thoughts on the Tesla Bot
55:00 - Final thoughts
Paper: https://arxiv.org/abs/2204.01691
Website: https://say-can.github.io/
Abstract:
Large language models can encode a wealth of semantic knowledge about the world. Such knowledge could be extremely useful to robots aiming to act upon high-level, temporally extended instructions expressed in natural language. However, a significant weakness of language models is that they lack real-world experience, which makes it difficult to leverage them for decision making within a given embodiment. For example, asking a language model to describe how to clean a spill might result in a reasonable narrative, but it may not be applicable to a particular agent, such as a robot, that needs to perform this task in a particular environment. We propose to provide real-world grounding by means of pretrained skills, which are used to constrain the model to propose natural language actions that are both feasible and contextually appropriate. The robot can act as the language model's "hands and eyes," while the language model supplies high-level semantic knowledge about the task. We show how low-level skills can be combined with large language models so that the language model provides high-level knowledge about the procedures for performing complex and temporally-extended instructions, while value functions associated with these skills provide the grounding necessary to connect this knowledge to a particular physical environment.
Authors: Michael Ahn, Anthony Brohan, Noah Brown, Yevgen Chebotar, Omar Cortes, Byron David, Chelsea Finn, Keerthana Gopalakrishnan, Karol Hausman, Alex Herzog, Daniel Ho, Jasmine Hsu, Julian Ibarz, Brian Ichter, Alex Irpan, Eric Jang, Rosario Jauregui Ruano, Kyle Jeffrey, Sally Jesmonth, Nikhil J Joshi, Ryan Julian, Dmitry Kalashnikov, Yuheng Kuang, Kuang-Huei Lee, Sergey Levine, Yao Lu, Linda Luu, Carolina Parada, Peter Pastor, Jornell Quiambao, Kanishka Rao, Jarek Rettinghouse, Diego Reyes, Pierre Sermanet, Nicolas Sievers, Clayton Tan, Alexander Toshev, Vincent Vanhoucke, Fei Xia, Ted Xiao, Peng Xu, Sichun Xu, Mengyuan Yan
Large Language Models are excellent at generating plausible plans in response to real-world problems, but without interacting with the environment, they have no abilities to estimate which of these plans are feasible or appropriate. SayCan combines the semantic capabilities of language models with a bank of low-level skills, which are available to the agent as individual policies to execute. SayCan automatically finds the best policy to execute by considering a trade-off between the policy's ability to progress towards the goal, given by the language model, and the policy's probability of executing successfully, given by the respective value function. The result is a system that can generate and execute long-horizon action sequences in the real world to fulfil complex tasks.
Sponsor: Zeta Alpha
https://zeta-alpha.com
Use code YANNIC for 20% off!
OUTLINE:
0:00 - Introduction & Overview
3:20 - Sponsor: Zeta Alpha
5:00 - Using language models for action planning
8:00 - Combining LLMs with learned atomic skills
16:50 - The full SayCan system
20:30 - Experimental setup and data collection
21:25 - Some weaknesses & strengths of the system
27:00 - Experimental results
Paper: https://arxiv.org/abs/2204.01691
Website: https://say-can.github.io/
Abstract:
Large language models can encode a wealth of semantic knowledge about the world. Such knowledge could be extremely useful to robots aiming to act upon high-level, temporally extended instructions expressed in natural language. However, a significant weakness of language models is that they lack real-world experience, which makes it difficult to leverage them for decision making within a given embodiment. For example, asking a language model to describe how to clean a spill might result in a reasonable narrative, but it may not be applicable to a particular agent, such as a robot, that needs to perform this task in a particular environment. We propose to provide real-world grounding by means of pretrained skills, which are used to constrain the model to propose natural language actions that are both feasible and contextually appropriate. The robot can act as the language model's "hands and eyes," while the language model supplies high-level semantic knowledge about the task. We show how low-level skills can be combined with large language models so that the language model provides high-level knowledge about the procedures for performing complex and temporally-extended instructions, while value functions associated with these skills provide the grounding necessary to connect this knowledge to a particular physical environment. We evaluate our method on a number of real-world robotic tasks, where we show the need for real-world grounding and that this approach is capable of completing long-horizon, abstract, natural language instructions on a mobile manipulator. The project's website and the video can be found at this https URL
Authors: Michael Ahn, Anthony Brohan, Noah Brown, Yevgen Chebotar, Omar Cortes, Byron David, Chelsea Finn, Keerthana Gopalakrishnan, Karol Hausman, Alex Herzog, Daniel Ho, Jasmine Hsu, Julian Ibarz, Brian Ichter, Alex Irpan, Eric Jang, Rosario Jauregui Ruano, Kyle Jeffrey, Sally Jesmonth, Nikhil J Joshi, Ryan Julian, Dmitry Kalashnikov, Yuheng Kuang, Kuang-Huei Lee, Sergey Levine, Yao Lu, Linda Luu, Carolina Parada, Peter Pastor, Jornell Quiambao, Kanishka Rao, Jarek Rettinghouse, Diego Reyes, Pierre Sermanet, Nicolas Sievers, Clayton Tan, Alexander Toshev, Vincent Vanhoucke, Fei Xia, Ted Xiao, Peng Xu, Sichun Xu, Mengyuan Yan
This is an interview with the authors Jack Parker-Holder and Minqi Jiang.
Original Paper Review Video: https://www.youtube.com/watch?v=povBD...
Automatic curriculum generation is one of the most promising avenues for Reinforcement Learning today. Multiple approaches have been proposed, each with their own set of advantages and drawbacks. This paper presents ACCEL, which takes the next step into the direction of constructing curricula for multi-capable agents. ACCEL combines the adversarial adaptiveness of regret-based sampling methods with the capabilities of level-editing, usually found in Evolutionary Methods.
OUTLINE:
0:00 - Intro
1:00 - Start of interview
4:45 - How did you get into this field?
8:10 - What is minimax regret?
11:45 - What levels does the regret objective select?
14:20 - Positive value loss (correcting my mistakes)
21:05 - Why is the teacher not learned?
24:45 - How much domain-specific knowledge is needed?
29:30 - What problems is this applicable to?
33:15 - Single agent vs population of agents
37:25 - Measuring and balancing level difficulty
40:35 - How does generalization emerge?
42:50 - Diving deeper into the experimental results
47:00 - What are the unsolved challenges in the field?
50:00 - Where do we go from here?
Website: https://accelagent.github.io
Paper: https://arxiv.org/abs/2203.01302
ICLR Workshop: https://sites.google.com/view/aloe2022
Book on topic: https://www.oreilly.com/radar/open-en...
Abstract:
It remains a significant challenge to train generally capable agents with reinforcement learning (RL). A promising avenue for improving the robustness of RL agents is through the use of curricula. One such class of methods frames environment design as a game between a student and a teacher, using regret-based objectives to produce environment instantiations (or levels) at the frontier of the student agent's capabilities. These methods benefit from their generality, with theoretical guarantees at equilibrium, yet they often struggle to find effective levels in challenging design spaces. By contrast, evolutionary approaches seek to incrementally alter environment complexity, resulting in potentially open-ended learning, but often rely on domain-specific heuristics and vast amounts of computational resources. In this paper we propose to harness the power of evolution in a principled, regret-based curriculum. Our approach, which we call Adversarially Compounding Complexity by Editing Levels (ACCEL), seeks to constantly produce levels at the frontier of an agent's capabilities, resulting in curricula that start simple but become increasingly complex. ACCEL maintains the theoretical benefits of prior regret-based methods, while providing significant empirical gains in a diverse set of environments. An interactive version of the paper is available at this http URL.
Authors: Jack Parker-Holder, Minqi Jiang, Michael Dennis, Mikayel Samvelyan, Jakob Foerster, Edward Grefenstette, Tim Rocktäschel
Links:
TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick
YouTube: https://www.youtube.com/c/yannickilcher
Twitter: https://twitter.com/ykilcher
Discord: https://discord.gg/4H8xxDF
BitChute: https://www.bitchute.com/channel/yann...
LinkedIn: https://www.linkedin.com/in/ykilcher
BiliBili: https://space.bilibili.com/2017636191
If you want to support me, the best thing to do is to share out the content :)
Automatic curriculum generation is one of the most promising avenues for Reinforcement Learning today. Multiple approaches have been proposed, each with their own set of advantages and drawbacks. This paper presents ACCEL, which takes the next step into the direction of constructing curricula for multi-capable agents. ACCEL combines the adversarial adaptiveness of regret-based sampling methods with the capabilities of level-editing, usually found in Evolutionary Methods.
OUTLINE:
0:00 - Intro & Demonstration
3:50 - Paper overview
5:20 - The ACCEL algorithm
15:25 - Looking at the pseudocode
23:10 - Approximating regret
33:45 - Experimental results
40:00 - Discussion & Comments
Website: https://accelagent.github.io
Paper: https://arxiv.org/abs/2203.01302
Abstract:
It remains a significant challenge to train generally capable agents with reinforcement learning (RL). A promising avenue for improving the robustness of RL agents is through the use of curricula. One such class of methods frames environment design as a game between a student and a teacher, using regret-based objectives to produce environment instantiations (or levels) at the frontier of the student agent's capabilities. These methods benefit from their generality, with theoretical guarantees at equilibrium, yet they often struggle to find effective levels in challenging design spaces. By contrast, evolutionary approaches seek to incrementally alter environment complexity, resulting in potentially open-ended learning, but often rely on domain-specific heuristics and vast amounts of computational resources. In this paper we propose to harness the power of evolution in a principled, regret-based curriculum. Our approach, which we call Adversarially Compounding Complexity by Editing Levels (ACCEL), seeks to constantly produce levels at the frontier of an agent's capabilities, resulting in curricula that start simple but become increasingly complex. ACCEL maintains the theoretical benefits of prior regret-based methods, while providing significant empirical gains in a diverse set of environments. An interactive version of the paper is available at this http URL.
Authors: Jack Parker-Holder, Minqi Jiang, Michael Dennis, Mikayel Samvelyan, Jakob Foerster, Edward Grefenstette, Tim Rocktäschel
Links:
TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick
YouTube: https://www.youtube.com/c/yannickilcher
Twitter: https://twitter.com/ykilcher
Discord: https://discord.gg/4H8xxDF
BitChute: https://www.bitchute.com/channel/yann...
LinkedIn: https://www.linkedin.com/in/ykilcher
BiliBili: https://space.bilibili.com/2017636191
If you want to support me, the best thing to do is to share out the content :)
If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):
SubscribeStar: https://www.subscribestar.com/yannick...
Patreon: https://www.patreon.com/yannickilcher
Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq
Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2
Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m
Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n
LAION-5B is an open, free dataset consisting of over 5 billion image-text-pairs. Today's video is an interview with three of its creators. We dive into the mechanics and challenges of operating at such large scale, how to keep cost low, what new possibilities are enabled with open datasets like this, and how to best handle safety and legal concerns.
OUTLINE:
0:00 - Intro
1:30 - Start of Interview
2:30 - What is LAION?
11:10 - What are the effects of CLIP filtering?
16:40 - How big is this dataset?
19:05 - Does the text always come from the alt-property?
22:45 - What does it take to work at scale?
25:50 -When will we replicate DALL-E?
31:30 - The surprisingly efficient pipeline
35:20 - How do you cover the S3 costs?
40:30 - Addressing safety & legal concerns
55:15 - Where can people get started?
References:
LAION website: https://laion.ai/
LAION Discord: https://discord.com/invite/mVcgxMPD7e
LAION-5B: https://laion.ai/laion-5b-a-new-era-o...
img2dataset tool: https://github.com/rom1504/img2dataset
LAION-400M: https://paperswithcode.com/dataset/la...
Links:
TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick
YouTube: https://www.youtube.com/c/yannickilcher
Twitter: https://twitter.com/ykilcher
Discord: https://discord.gg/4H8xxDF
BitChute: https://www.bitchute.com/channel/yann...
LinkedIn: https://www.linkedin.com/in/ykilcher
BiliBili: https://space.bilibili.com/2017636191
If you want to support me, the best thing to do is to share out the content :)
If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):
SubscribeStar: https://www.subscribestar.com/yannick...
Patreon: https://www.patreon.com/yannickilcher
Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq
Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2
Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m
Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n
This video is an interview with Barret Zoph and William Fedus of Google Brain about Sparse Expert Models.
Sparse Expert models have been hugely successful at distributing parts of models, mostly Transformers, across large array of machines and use a routing function to effectively route signals between them. This means that even though these models have a huge number of parameters, the computational load for a given signal does not increase because the model is only sparsely activated. Sparse expert models, such as Switch Transformers and GLAM can scale up to trillions of parameters and bring a number of desirable properties. We discuss everything from the fundamentals, history, strengths and weaknesses, up to the current state of the art of these models.
OUTLINE:
0:00 - Intro
0:30 - What are sparse expert models?
4:25 - Start of Interview
5:55 - What do you mean by sparse experts?
8:10 - How does routing work in these models?
12:10 - What is the history of sparse experts?
14:45 - What does an individual expert learn?
19:25 - When are these models appropriate?
22:30 - How comparable are sparse to dense models?
26:30 - How does the pathways system connect to this?
28:45 - What improvements did GLAM make?
31:30 - The "designing sparse experts" paper
37:45 - Can experts be frozen during training?
41:20 - Can the routing function be improved?
47:15 - Can experts be distributed beyond data centers?
50:20 - Are there sparse experts for other domains than NLP?
52:15 - Are sparse and dense models in competition?
53:35 - Where do we go from here?
56:30 - How can people get started with this?
Papers:
Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity (https://arxiv.org/abs/2101.03961)
GLaM: Efficient Scaling of Language Models with Mixture-of-Experts (https://arxiv.org/abs/2112.06905)
Designing Effective Sparse Expert Models (https://arxiv.org/abs/2202.08906)
Links:
Merch: store.ykilcher.com
TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick
YouTube: https://www.youtube.com/c/yannickilcher
Twitter: https://twitter.com/ykilcher
Discord: https://discord.gg/4H8xxDF
BitChute: https://www.bitchute.com/channel/yann...
LinkedIn: https://www.linkedin.com/in/ykilcher
BiliBili: https://space.bilibili.com/2017636191
If you want to support me, the best thing to do is to share out the content :)
If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):
SubscribeStar: https://www.subscribestar.com/yannick...
Patreon: https://www.patreon.com/yannickilcher
Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq
Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2
Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m
Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n
This is an interview with the authors Yi Tay and Don Metzler.
Paper Review Video: https://youtu.be/qlB0TPBQ7YY
Search engines work by building an index and then looking up things in it. Usually, that index is a separate data structure. In keyword search, we build and store reverse indices. In neural search, we build nearest-neighbor indices. This paper does something different: It directly trains a Transformer to return the ID of the most relevant document. No similarity search over embeddings or anything like this is performed, and no external data structure is needed, as the entire index is essentially captured by the model's weights. The paper experiments with various ways of representing documents and training the system, which works surprisingly well!
OUTLINE:
0:00 - Intro
0:50 - Start of Interview
1:30 - How did this idea start?
4:30 - How does memorization play into this?
5:50 - Why did you not compare to cross-encoders?
7:50 - Instead of the ID, could one reproduce the document itself?
10:50 - Passages vs documents
12:00 - Where can this model be applied?
14:25 - Can we make this work on large collections?
19:20 - What's up with the NQ100K dataset?
23:55 - What is going on inside these models?
28:30 - What's the smallest scale to obtain meaningful results?
30:15 - Investigating the document identifiers
34:45 - What's the end goal?
38:40 - What are the hardest problems currently?
40:40 - Final comments & how to get started
Paper: https://arxiv.org/abs/2202.06991
Abstract:
In this paper, we demonstrate that information retrieval can be accomplished with a single Transformer, in which all information about the corpus is encoded in the parameters of the model. To this end, we introduce the Differentiable Search Index (DSI), a new paradigm that learns a text-to-text model that maps string queries directly to relevant docids; in other words, a DSI model answers queries directly using only its parameters, dramatically simplifying the whole retrieval process. We study variations in how documents and their identifiers are represented, variations in training procedures, and the interplay between models and corpus sizes. Experiments demonstrate that given appropriate design choices, DSI significantly outperforms strong baselines such as dual encoder models. Moreover, DSI demonstrates strong generalization capabilities, outperforming a BM25 baseline in a zero-shot setup.
Authors: Yi Tay, Vinh Q. Tran, Mostafa Dehghani, Jianmo Ni, Dara Bahri, Harsh Mehta, Zhen Qin, Kai Hui, Zhe Zhao, Jai Gupta, Tal Schuster, William W. Cohen, Donald Metzler
Links:
Merch: store.ykilcher.com
TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick
YouTube: https://www.youtube.com/c/yannickilcher
Twitter: https://twitter.com/ykilcher
Discord: https://discord.gg/4H8xxDF
BitChute: https://www.bitchute.com/channel/yann...
LinkedIn: https://www.linkedin.com/in/ykilcher
BiliBili: https://space.bilibili.com/2017636191
If you want to support me, the best thing to do is to share out the content :)
If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):
SubscribeStar: https://www.subscribestar.com/yannick...
Patreon: https://www.patreon.com/yannickilcher
Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq
Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2
Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m
Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n
Search engines work by building an index and then looking up things in it. Usually, that index is a separate data structure. In keyword search, we build and store reverse indices. In neural search, we build nearest-neighbor indices. This paper does something different: It directly trains a Transformer to return the ID of the most relevant document. No similarity search over embeddings or anything like this is performed, and no external data structure is needed, as the entire index is essentially captured by the model's weights. The paper experiments with various ways of representing documents and training the system, which works surprisingly well!
Sponsor: Diffgram
https://diffgram.com?ref=yannic
OUTLINE:
0:00 - Intro
0:45 - Sponsor: Diffgram
1:35 - Paper overview
3:15 - The search problem, classic and neural
8:15 - Seq2seq for directly predicting document IDs
11:05 - Differentiable search index architecture
18:05 - Indexing
25:15 - Retrieval and document representation
33:25 - Training DSI
39:15 - Experimental results
49:25 - Comments & Conclusions
Paper: https://arxiv.org/abs/2202.06991
Abstract:
In this paper, we demonstrate that information retrieval can be accomplished with a single Transformer, in which all information about the corpus is encoded in the parameters of the model. To this end, we introduce the Differentiable Search Index (DSI), a new paradigm that learns a text-to-text model that maps string queries directly to relevant docids; in other words, a DSI model answers queries directly using only its parameters, dramatically simplifying the whole retrieval process. We study variations in how documents and their identifiers are represented, variations in training procedures, and the interplay between models and corpus sizes. Experiments demonstrate that given appropriate design choices, DSI significantly outperforms strong baselines such as dual encoder models. Moreover, DSI demonstrates strong generalization capabilities, outperforming a BM25 baseline in a zero-shot setup.
Authors: Yi Tay, Vinh Q. Tran, Mostafa Dehghani, Jianmo Ni, Dara Bahri, Harsh Mehta, Zhen Qin, Kai Hui, Zhe Zhao, Jai Gupta, Tal Schuster, William W. Cohen, Donald Metzler
Links:
Merch: store.ykilcher.com
TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick
YouTube: https://www.youtube.com/c/yannickilcher
Twitter: https://twitter.com/ykilcher
Discord: https://discord.gg/4H8xxDF
BitChute: https://www.bitchute.com/channel/yann...
LinkedIn: https://www.linkedin.com/in/ykilcher
BiliBili: https://space.bilibili.com/2017636191
If you want to support me, the best thing to do is to share out the content :)
If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):
SubscribeStar: https://www.subscribestar.com/yannick...
Patreon: https://www.patreon.com/yannickilcher
Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq
Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2
Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m
Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n
Google releases PaLM and OpenAI releases DALL-E 2 (and more news).
Sponsor: Weights & BIases
Start here: https://wandb.me/yannic
Thumbnail credit: DALL-E 2 via Sam Altman
OUTLINE
0:00 - Street interview w/ random stranger
2:25 - Intro
2:50 - PaLM - Google's 540B Pathways Language Model
7:50 - Sponsor: Weights & Biases
9:10 - OpenAI releases DALL-E 2
12:05 - Open Source Datasets and Models
13:20 - Salesforce releases CodeGen
My Live Reaction to DALL-E 2: https://youtu.be/gGPv_SYVDC8
My Video on GLIDE: https://youtu.be/gwI6g1pBD84
My Video on the Pathways System: https://youtu.be/vGFaiLeoLWw
References:
PaLM - Google's 540B Pathways Language Model
https://ai.googleblog.com/2022/04/pat...
https://storage.googleapis.com/pathwa...
OpenAI releases DALL-E 2
https://openai.com/dall-e-2/
https://cdn.openai.com/papers/dall-e-...
https://www.instagram.com/openaidalle/
https://twitter.com/sama/status/15117...
https://twitter.com/sama/media
https://twitter.com/BorisMPower/statu...
https://twitter.com/ariskonstant/stat...
Open Source Datasets and Models
https://twitter.com/multimodalart/sta...
https://laion.ai/laion-5b-a-new-era-o...
https://github.com/mlfoundations/open...
Salesforce releases CodeGen
https://github.com/salesforce/CodeGen
Links:
Merch: store.ykilcher.com
TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick
YouTube: https://www.youtube.com/c/yannickilcher
Twitter: https://twitter.com/ykilcher
Discord: https://discord.gg/4H8xxDF
BitChute: https://www.bitchute.com/channel/yann...
LinkedIn: https://www.linkedin.com/in/ykilcher
BiliBili: https://space.bilibili.com/2017636191
If you want to support me, the best thing to do is to share out the content :)
If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):
SubscribeStar: https://www.subscribestar.com/yannick...
Patreon: https://www.patreon.com/yannickilcher
Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq
Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2
Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m
Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n
Since the release of CLIP, the world of AI art has seen an unprecedented level of acceleration in what's possible to do. Whereas image generation had previously been mostly in the domain of scientists, now a community of professional artists, researchers, and amateurs are sending around colab notebooks and sharing their creations via social media. How did this happen? What is going on? And where do we go from here? Jack Morris and I attempt to answer some of these questions, following his blog post "The Weird and Wonderful World of AI Art" (linked below).
OUTLINE:
0:00 - Intro
2:30 - How does one get into AI art?
5:00 - Deep Dream & Style Transfer: the early days of art in deep learning
10:50 - The advent of GANs, ArtBreeder and TikTok
19:50 - Lacking control: Pre-CLIP art
22:40 - CLIP & DALL-E
30:20 - The shift to shared colabs
34:20 - Guided diffusion models
37:20 - Prompt engineering for art models
43:30 - GLIDE
47:00 - Video production & Disco Diffusion
48:40 - Economics, money, and NFTs
54:15 - What does the future hold for AI art?
Blog post: https://jxmo.notion.site/The-Weird-an...
Jack's Blog: https://jxmo.io/
Links:
TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick
YouTube: https://www.youtube.com/c/yannickilcher
Twitter: https://twitter.com/ykilcher
Discord: https://discord.gg/4H8xxDF
BitChute: https://www.bitchute.com/channel/yann...
LinkedIn: https://www.linkedin.com/in/ykilcher
BiliBili: https://space.bilibili.com/2017636191
If you want to support me, the best thing to do is to share out the content :)
If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):
SubscribeStar: https://www.subscribestar.com/yannick...
Patreon: https://www.patreon.com/yannickilcher
Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq
Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2
Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m
Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n
This is an interview with Jesse Mu, first author of the paper.
Original Paper Review: https://youtu.be/NeGJAUSQEJI
Exploration is one of the oldest challenges for Reinforcement Learning algorithms, with no clear solution to date. Especially in environments with sparse rewards, agents face significant challenges in deciding which parts of the environment to explore further. Providing intrinsic motivation in form of a pseudo-reward is sometimes used to overcome this challenge, but often relies on hand-crafted heuristics, and can lead to deceptive dead-ends. This paper proposes to use language descriptions of encountered states as a method of assessing novelty. In two procedurally generated environments, they demonstrate the usefulness of language, which is in itself highly concise and abstractive, which lends itself well for this task.
OUTLINE:
0:00 - Intro
0:55 - Paper Overview
4:30 - Aren't you just adding extra data?
9:35 - Why are you splitting up the AMIGo teacher?
13:10 - How do you train the grounding network?
16:05 - What about causally structured environments?
17:30 - Highlights of the experimental results
20:40 - Why is there so much variance?
22:55 - How much does it matter that we are testing in a video game?
27:00 - How does novelty interface with the goal specification?
30:20 - The fundamental problems of exploration
32:15 - Are these algorithms subject to catastrophic forgetting?
34:45 - What current models could bring language to other environments?
40:30 - What does it take in terms of hardware?
43:00 - What problems did you encounter during the project?
46:40 - Where do we go from here?
Paper: https://arxiv.org/abs/2202.08938
Abstract:
Reinforcement learning (RL) agents are particularly hard to train when rewards are sparse. One common solution is to use intrinsic rewards to encourage agents to explore their environment. However, recent intrinsic exploration methods often use state-based novelty measures which reward low-level exploration and may not scale to domains requiring more abstract skills. Instead, we explore natural language as a general medium for highlighting relevant abstractions in an environment. Unlike previous work, we evaluate whether language can improve over existing exploration methods by directly extending (and comparing to) competitive intrinsic exploration baselines: AMIGo (Campero et al., 2021) and NovelD (Zhang et al., 2021). These language-based variants outperform their non-linguistic forms by 45-85% across 13 challenging tasks from the MiniGrid and MiniHack environment suites.
Authors: Jesse Mu, Victor Zhong, Roberta Raileanu, Minqi Jiang, Noah Goodman, Tim Rocktäschel, Edward Grefenstette
Links:
TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick
YouTube: https://www.youtube.com/c/yannickilcher
Twitter: https://twitter.com/ykilcher
Discord: https://discord.gg/4H8xxDF
BitChute: https://www.bitchute.com/channel/yann...
LinkedIn: https://www.linkedin.com/in/ykilcher
BiliBili: https://space.bilibili.com/2017636191
If you want to support me, the best thing to do is to share out the content :)
If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):
SubscribeStar: https://www.subscribestar.com/yannick...
Patreon: https://www.patreon.com/yannickilcher
Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq
Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2
Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m
Exploration is one of the oldest challenges for Reinforcement Learning algorithms, with no clear solution to date. Especially in environments with sparse rewards, agents face significant challenges in deciding which parts of the environment to explore further. Providing intrinsic motivation in form of a pseudo-reward is sometimes used to overcome this challenge, but often relies on hand-crafted heuristics, and can lead to deceptive dead-ends. This paper proposes to use language descriptions of encountered states as a method of assessing novelty. In two procedurally generated environments, they demonstrate the usefulness of language, which is in itself highly concise and abstractive, which lends itself well for this task.
OUTLINE:
0:00 - Intro
1:10 - Paper Overview: Language for exploration
5:40 - The MiniGrid & MiniHack environments
7:00 - Annotating states with language
9:05 - Baseline algorithm: AMIGo
12:20 - Adding language to AMIGo
22:55 - Baseline algorithm: NovelD and Random Network Distillation
29:45 - Adding language to NovelD
31:50 - Aren't we just using extra data?
34:55 - Investigating the experimental results
40:45 - Final comments
Paper: https://arxiv.org/abs/2202.08938
Abstract:
Reinforcement learning (RL) agents are particularly hard to train when rewards are sparse. One common solution is to use intrinsic rewards to encourage agents to explore their environment. However, recent intrinsic exploration methods often use state-based novelty measures which reward low-level exploration and may not scale to domains requiring more abstract skills. Instead, we explore natural language as a general medium for highlighting relevant abstractions in an environment. Unlike previous work, we evaluate whether language can improve over existing exploration methods by directly extending (and comparing to) competitive intrinsic exploration baselines: AMIGo (Campero et al., 2021) and NovelD (Zhang et al., 2021). These language-based variants outperform their non-linguistic forms by 45-85% across 13 challenging tasks from the MiniGrid and MiniHack environment suites.
Authors: Jesse Mu, Victor Zhong, Roberta Raileanu, Minqi Jiang, Noah Goodman, Tim Rocktäschel, Edward Grefenstette
Links:
TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick
YouTube: https://www.youtube.com/c/yannickilcher
Twitter: https://twitter.com/ykilcher
Discord: https://discord.gg/4H8xxDF
BitChute: https://www.bitchute.com/channel/yann...
LinkedIn: https://www.linkedin.com/in/ykilcher
BiliBili: https://space.bilibili.com/2017636191
If you want to support me, the best thing to do is to share out the content :)
If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):
SubscribeStar: https://www.subscribestar.com/yannick...
Patreon: https://www.patreon.com/yannickilcher
Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq
Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2
Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m
Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n
Your updates on the latest and greatest from the depths of Machine Learning!
Sponsor: Weights & Biases
https://wandb.me/yannic
OUTLINE:
0:00 - Intro
0:15 - Weights & Biases Report about Reports
2:45 - GPT-3 learns to edit
6:30 - Make-A-Scene: Text-to-Image with Human Priors
8:00 - Pathways: Google's new High-Performance ML scheduler
10:45 - DouBlind: Open Peer-Review
12:45 - CLIP meets GamePhysics
14:40 - Residual Quantization pushes Image Generation SOTA
16:15 - Helpful Things
References:
Weights & Biases Report about Reports
https://wandb.ai/wandb/wandb_example/...
GPT-3 learns to edit
https://openai.com/blog/gpt-3-edit-in...
https://beta.openai.com/playground?mo...
Make-A-Scene: Text-to-Image with Human Priors
https://arxiv.org/pdf/2203.13131.pdf
https://www.youtube.com/watch?v=QLTyq...
Pathways: Google's new High-Performance ML scheduler
https://arxiv.org/pdf/2203.12533.pdf
DouBlind: Open Peer-Review
https://doublind.com/#web-intro
https://doublind.com/search?query=kil...
CLIP meets GamePhysics
https://arxiv.org/pdf/2203.11096.pdf
https://www.reddit.com/r/GamePhysics/...
https://asgaardlab.github.io/CLIPxGam...
Residual Quantization pushes Image Generation SOTA
https://arxiv.org/pdf/2203.01941.pdf
https://github.com/kakaobrain/rq-vae-...
Helpful Things
https://github.com/TDAmeritrade/stumpy
https://github.com/linkedin/fasttreeshap
https://github.com/vopani/jaxton
https://twitter.com/mark_riedl/status...
https://github.com/eilab-gt/NovGrid
https://developer.nvidia.com/isaac-gym
https://github.com/NVIDIA-Omniverse/I...
Links:
Merch: store.ykilcher.com
TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick
YouTube: https://www.youtube.com/c/yannickilcher
Twitter: https://twitter.com/ykilcher
Discord: https://discord.gg/4H8xxDF
BitChute: https://www.bitchute.com/channel/yann...
LinkedIn: https://www.linkedin.com/in/ykilcher
BiliBili: https://space.bilibili.com/2017636191
If you want to support me, the best thing to do is to share out the content :)
If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):
SubscribeStar: https://www.subscribestar.com/yannick...
Patreon: https://www.patreon.com/yannickilcher
Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq
Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2
Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m
Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n
This is an interview with the authors of this work, Aman Madaan and Niket Tandon.
Large language models such as GPT-3 have enabled many breakthroughs and new applications recently, but they come with an important downside: Training them is very expensive, and even fine-tuning is often difficult. This paper presents an adaptive method to improve performance of such models after deployment, without ever changing the model itself. This is done by maintaining a memory of interactions and then dynamically adapting new prompts by augmenting them with memory content. This has many applications, from non-intrusive fine-tuning to personalization.
OUTLINE:
0:00 - Intro
0:45 - Paper Overview
2:00 - What was your original motivation?
4:20 - There is an updated version of the paper!
9:00 - Have you studied this on real-world users?
12:10 - How does model size play into providing feedback?
14:10 - Can this be used for personalization?
16:30 - Discussing experimental results
17:45 - Can this be paired with recommender systems?
20:00 - What are obvious next steps to make the system more powerful?
23:15 - Clarifying the baseline methods
26:30 - Exploring cross-lingual customization
31:00 - Where did the idea for the clarification prompt come from?
33:05 - What did not work out during this project?
34:45 - What did you learn about interacting with large models?
37:30 - Final thoughts
Paper: https://arxiv.org/abs/2201.06009
Code & Data: https://github.com/madaan/memprompt
Abstract:
Large LMs such as GPT-3 are powerful, but can commit mistakes that are obvious to humans. For example, GPT-3 would mistakenly interpret "What word is similar to good?" to mean a homonym, while the user intended a synonym. Our goal is to effectively correct such errors via user interactions with the system but without retraining, which will be prohibitively costly. We pair GPT-3 with a growing memory of recorded cases where the model misunderstood the user's intents, along with user feedback for clarification. Such a memory allows our system to produce enhanced prompts for any new query based on the user feedback for error correction on similar cases in the past. On four tasks (two lexical tasks, two advanced ethical reasoning tasks), we show how a (simulated) user can interactively teach a deployed GPT-3, substantially increasing its accuracy over the queries with different kinds of misunderstandings by the GPT-3. Our approach is a step towards the low-cost utility enhancement for very large pre-trained LMs. All the code and data is available at this https URL.
Authors: Aman Madaan, Niket Tandon, Peter Clark, Yiming Yang
Links:
Merch: store.ykilcher.com
TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick
YouTube: https://www.youtube.com/c/yannickilcher
Twitter: https://twitter.com/ykilcher
Discord: https://discord.gg/4H8xxDF
BitChute: https://www.bitchute.com/channel/yann...
LinkedIn: https://www.linkedin.com/in/ykilcher
BiliBili: https://space.bilibili.com/2017636191
If you want to support me, the best thing to do is to share out the content :)
If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):
SubscribeStar: https://www.subscribestar.com/yannick...
Patreon: https://www.patreon.com/yannickilcher
Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq
Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2
Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m
Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n
Large language models such as GPT-3 have enabled many breakthroughs and new applications recently, but they come with an important downside: Training them is very expensive, and even fine-tuning is often difficult. This paper presents an adaptive method to improve performance of such models after deployment, without ever changing the model itself. This is done by maintaining a memory of interactions and then dynamically adapting new prompts by augmenting them with memory content. This has many applications, from non-intrusive fine-tuning to personalization.
Sponsor: Introduction to Graph Neural Networks Course
https://www.graphneuralnets.com/p/int...
OUTLINE:
0:00 - Intro
0:40 - Sponsor: Introduction to GNNs Course (link in description)
1:30 - Paper Overview: Improve GPT-3 after deployment via user feedback
5:30 - Proposed memory-based architecture
13:00 - A detailed look at the components
15:00 - Example tasks
24:30 - My concerns with the example setup
26:20 - Baselines used for comparison
29:50 - Experimental Results
34:20 - Conclusion & Comments
Paper: https://arxiv.org/abs/2201.06009
Code & Data: https://github.com/madaan/memprompt
Abstract:
Large LMs such as GPT-3 are powerful, but can commit mistakes that are obvious to humans. For example, GPT-3 would mistakenly interpret "What word is similar to good?" to mean a homonym, while the user intended a synonym. Our goal is to effectively correct such errors via user interactions with the system but without retraining, which will be prohibitively costly. We pair GPT-3 with a growing memory of recorded cases where the model misunderstood the user's intents, along with user feedback for clarification. Such a memory allows our system to produce enhanced prompts for any new query based on the user feedback for error correction on similar cases in the past. On four tasks (two lexical tasks, two advanced ethical reasoning tasks), we show how a (simulated) user can interactively teach a deployed GPT-3, substantially increasing its accuracy over the queries with different kinds of misunderstandings by the GPT-3. Our approach is a step towards the low-cost utility enhancement for very large pre-trained LMs. All the code and data is available at this https URL.
Authors: Aman Madaan, Niket Tandon, Peter Clark, Yiming Yang
Links:
Merch: store.ykilcher.com
TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick
YouTube: https://www.youtube.com/c/yannickilcher
Twitter: https://twitter.com/ykilcher
Discord: https://discord.gg/4H8xxDF
BitChute: https://www.bitchute.com/channel/yann...
LinkedIn: https://www.linkedin.com/in/ykilcher
BiliBili: https://space.bilibili.com/2017636191
If you want to support me, the best thing to do is to share out the content :)
If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):
SubscribeStar: https://www.subscribestar.com/yannick...
Patreon: https://www.patreon.com/yannickilcher
Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq
Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2
Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m
Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n
This is an interview with first author Clara Meister.
Paper review video hereé https://youtu.be/_EDr3ryrT_Y
Modern language models like T5 or GPT-3 achieve remarkably low perplexities on both training and validation data, yet when sampling from their output distributions, the generated text often seems dull and uninteresting. Various workarounds have been proposed, such as top-k sampling and nucleus sampling, but while these manage to somewhat improve the generated samples, they are hacky and unfounded. This paper introduces typical sampling, a new decoding method that is principled, effective, and can be implemented efficiently. Typical sampling turns away from sampling purely based on likelihood and explicitly finds a trade-off between generating high-probability samples and generating high-information samples. The paper connects typical sampling to psycholinguistic theories on human speech generation, and shows experimentally that typical sampling achieves much more diverse and interesting results than any of the current methods.
Sponsor: Introduction to Graph Neural Networks Course
https://www.graphneuralnets.com/p/int...
OUTLINE:
0:00 - Intro
0:35 - Sponsor: Introduction to GNNs Course (link in description)
1:30 - Why does sampling matter?
5:40 - What is a "typical" message?
8:35 - How do humans communicate?
10:25 - Why don't we just sample from the model's distribution?
15:30 - What happens if we condition on the information to transmit?
17:35 - Does typical sampling really represent human outputs?
20:55 - What do the plots mean?
31:00 - Diving into the experimental results
39:15 - Are our training objectives wrong?
41:30 - Comparing typical sampling to top-k and nucleus sampling
44:50 - Explaining arbitrary engineering choices
47:20 - How can people get started with this?
Paper: https://arxiv.org/abs/2202.00666
Code: https://github.com/cimeister/typical-...
Abstract:
Despite achieving incredibly low perplexities on myriad natural language corpora, today's language models still often underperform when used to generate text. This dichotomy has puzzled the language generation community for the last few years. In this work, we posit that the abstraction of natural language as a communication channel (à la Shannon, 1948) can provide new insights into the behaviors of probabilistic language generators, e.g., why high-probability texts can be dull or repetitive. Humans use language as a means of communicating information, and do so in a simultaneously efficient and error-minimizing manner; they choose each word in a string with this (perhaps subconscious) goal in mind. We propose that generation from probabilistic models should mimic this behavior. Rather than always choosing words from the high-probability region of the distribution--which have a low Shannon information content--we sample from the set of words with information content close to the conditional entropy of our model, i.e., close to the expected information content. This decision criterion can be realized through a simple and efficient implementation, which we call typical sampling. Automatic and human evaluations show that, in comparison to nucleus and top-k sampling, typical sampling offers competitive performance in terms of quality while consistently reducing the number of degenerate repetitions.
Authors: Clara Meister, Tiago Pimentel, Gian Wiher, Ryan Cotterell
Links:
Merch: store.ykilcher.com
TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick
YouTube: https://www.youtube.com/c/yannickilcher
Twitter: https://twitter.com/ykilcher
Discord: https://discord.gg/4H8xxDF
Modern language models like T5 or GPT-3 achieve remarkably low perplexities on both training and validation data, yet when sampling from their output distributions, the generated text often seems dull and uninteresting. Various workarounds have been proposed, such as top-k sampling and nucleus sampling, but while these manage to somewhat improve the generated samples, they are hacky and unfounded. This paper introduces typical sampling, a new decoding method that is principled, effective, and can be implemented efficiently. Typical sampling turns away from sampling purely based on likelihood and explicitly finds a trade-off between generating high-probability samples and generating high-information samples. The paper connects typical sampling to psycholinguistic theories on human speech generation, and shows experimentally that typical sampling achieves much more diverse and interesting results than any of the current methods.
Sponsor: Fully Connected by Weights & Biases
https://wandb.ai/fully-connected
OUTLINE:
0:00 - Intro
1:50 - Sponsor: Fully Connected by Weights & Biases
4:10 - Paper Overview
7:40 - What's the problem with sampling?
11:45 - Beam Search: The good and the bad
14:10 - Top-k and Nucleus Sampling
16:20 - Why the most likely things might not be the best
21:30 - The expected information content of the next word
25:00 - How to trade off information and likelihood
31:25 - Connections to information theory and psycholinguistics
36:40 - Introducing Typical Sampling
43:00 - Experimental Evaluation
44:40 - My thoughts on this paper
Paper: https://arxiv.org/abs/2202.00666
Code: https://github.com/cimeister/typical-...
Abstract:
Despite achieving incredibly low perplexities on myriad natural language corpora, today's language models still often underperform when used to generate text. This dichotomy has puzzled the language generation community for the last few years. In this work, we posit that the abstraction of natural language as a communication channel (à la Shannon, 1948) can provide new insights into the behaviors of probabilistic language generators, e.g., why high-probability texts can be dull or repetitive. Humans use language as a means of communicating information, and do so in a simultaneously efficient and error-minimizing manner; they choose each word in a string with this (perhaps subconscious) goal in mind. We propose that generation from probabilistic models should mimic this behavior. Rather than always choosing words from the high-probability region of the distribution--which have a low Shannon information content--we sample from the set of words with information content close to the conditional entropy of our model, i.e., close to the expected information content. This decision criterion can be realized through a simple and efficient implementation, which we call typical sampling. Automatic and human evaluations show that, in comparison to nucleus and top-k sampling, typical sampling offers competitive performance in terms of quality while consistently reducing the number of degenerate repetitions.
Authors: Clara Meister, Tiago Pimentel, Gian Wiher, Ryan Cotterell
Links:
Merch: store.ykilcher.com
TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick
YouTube: https://www.youtube.com/c/yannickilcher
Twitter: https://twitter.com/ykilcher
Discord: https://discord.gg/4H8xxDF
BitChute: https://www.bitchute.com/channel/yann...
LinkedIn: https://www.linkedin.com/in/ykilcher
BiliBili: https://space.bilibili.com/2017636191
If you want to support me, the best thing to do is to share out the content :)
Paper Review Video: https://youtu.be/X2k7n4FuI7c
Sponsor: Assembly AI
https://www.assemblyai.com/?utm_sourc...
This is an interview with Junnan Li and Dongxu Li, authors of BLIP and members of Salesforce research.
Cross-modal pre-training has been all the rage lately in deep learning, especially training vision and language models together. However, there are a number of issues, such as low quality datasets that limit the performance of any model trained on it, and also the fact that pure contrastive pre-training cannot be easily fine-tuned for most downstream tasks. BLIP unifies different tasks and objectives in a single pre-training run and achieves a much more versatile model, which the paper immediately uses to create, filter, clean and thus bootstrap its own dataset to improve performance even more!
OUTLINE:
0:00 - Intro
0:40 - Sponsor: Assembly AI
1:30 - Start of Interview
2:30 - What's the pitch?
4:40 - How did data bootstrapping come into the project?
7:10 - How big of a problem is data quality?
11:10 - Are the captioning & filtering models biased towards COCO data?
14:40 - Could the data bootstrapping be done multiple times?
16:20 - What was the evolution of the BLIP architecture?
21:15 - Are there additional benefits to adding language modelling?
23:50 - Can we imagine a modular future for pre-training?
29:45 - Diving into the experimental results
42:40 - What did and did not work out during the research?
45:00 - How is research life at Salesforce?
46:45 - Where do we go from here?
Paper: https://arxiv.org/abs/2201.12086
Code: https://github.com/salesforce/BLIP
Demo: https://huggingface.co/spaces/Salesfo...
Abstract:
Vision-Language Pre-training (VLP) has advanced the performance for many vision-language tasks. However, most existing pre-trained models only excel in either understanding-based tasks or generation-based tasks. Furthermore, performance improvement has been largely achieved by scaling up the dataset with noisy image-text pairs collected from the web, which is a suboptimal source of supervision. In this paper, we propose BLIP, a new VLP framework which transfers flexibly to both vision-language understanding and generation tasks. BLIP effectively utilizes the noisy web data by bootstrapping the captions, where a captioner generates synthetic captions and a filter removes the noisy ones. We achieve state-of-the-art results on a wide range of vision-language tasks, such as image-text retrieval (+2.7% in average recall@1), image captioning (+2.8% in CIDEr), and VQA (+1.6% in VQA score). BLIP also demonstrates strong generalization ability when directly transferred to video-language tasks in a zero-shot manner. Code, models, and datasets are released at this https URL.
Authors: Junnan Li, Dongxu Li, Caiming Xiong, Steven Hoi
Links:
TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick
YouTube: https://www.youtube.com/c/yannickilcher
Twitter: https://twitter.com/ykilcher
Discord: https://discord.gg/4H8xxDF
BitChute: https://www.bitchute.com/channel/yann...
LinkedIn: https://www.linkedin.com/in/ykilcher
BiliBili: https://space.bilibili.com/2017636191
If you want to support me, the best thing to do is to share out the content :)
If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):
SubscribeStar: https://www.subscribestar.com/yannick...
Patreon: https://www.patreon.com/yannickilcher
Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq
Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2
Cross-modal pre-training has been all the rage lately in deep learning, especially training vision and language models together. However, there are a number of issues, such as low quality datasets that limit the performance of any model trained on it, and also the fact that pure contrastive pre-training cannot be easily fine-tuned for most downstream tasks. BLIP unifies different tasks and objectives in a single pre-training run and achieves a much more versatile model, which the paper immediately uses to create, filter, clean and thus bootstrap its own dataset to improve performance even more!
Sponsor: Zeta Alpha
https://zeta-alpha.com
Use code YANNIC for 20% off!
OUTLINE:
0:00 - Intro
0:50 - Sponsor: Zeta Alpha
3:40 - Paper Overview
6:40 - Vision-Language Pre-Training
11:15 - Contributions of the paper
14:30 - Model architecture: many parts for many tasks
19:50 - How data flows in the model
26:50 - Parameter sharing between the modules
29:45 - Captioning & Filtering bootstrapping
41:10 - Fine-tuning the model for downstream tasks
Paper: https://arxiv.org/abs/2201.12086
Code: https://github.com/salesforce/BLIP
Demo: https://huggingface.co/spaces/Salesfo...
Abstract:
Vision-Language Pre-training (VLP) has advanced the performance for many vision-language tasks. However, most existing pre-trained models only excel in either understanding-based tasks or generation-based tasks. Furthermore, performance improvement has been largely achieved by scaling up the dataset with noisy image-text pairs collected from the web, which is a suboptimal source of supervision. In this paper, we propose BLIP, a new VLP framework which transfers flexibly to both vision-language understanding and generation tasks. BLIP effectively utilizes the noisy web data by bootstrapping the captions, where a captioner generates synthetic captions and a filter removes the noisy ones. We achieve state-of-the-art results on a wide range of vision-language tasks, such as image-text retrieval (+2.7% in average recall@1), image captioning (+2.8% in CIDEr), and VQA (+1.6% in VQA score). BLIP also demonstrates strong generalization ability when directly transferred to video-language tasks in a zero-shot manner. Code, models, and datasets are released at this https URL.
Authors: Junnan Li, Dongxu Li, Caiming Xiong, Steven Hoi
Links:
TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick
YouTube: https://www.youtube.com/c/yannickilcher
Twitter: https://twitter.com/ykilcher
Discord: https://discord.gg/4H8xxDF
BitChute: https://www.bitchute.com/channel/yann...
LinkedIn: https://www.linkedin.com/in/ykilcher
BiliBili: https://space.bilibili.com/2017636191
If you want to support me, the best thing to do is to share out the content :)
If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):
SubscribeStar: https://www.subscribestar.com/yannick...
Patreon: https://www.patreon.com/yannickilcher
Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq
Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2
Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m
Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n
GTC Registration Link: https://ykilcher.com/gtc
Your regular updates on what's going on in the ML world!
OUTLINE:
0:00 - Intro
0:20 - Register to Nvidia GTC and win a 3090!
4:15 - DeepMind's Ithaca deciphers Lost Ancient Texts
6:45 - Drug discovery model turns toxic
10:00 - Gary Marcus: Deep Learning is hitting a wall
19:40 - GopherCite: Backing up answers with citations
22:40 - Yoshua Bengio appointed knight of the legion of honour
23:00 - Meta AI tags parody account of Yoshua Bengio
23:40 - Building games using just natural language
24:55 - YOU.com adds writing assistant
25:45 - Horace He: How to brrr
26:35 - Karpathy: Reproducing Yann LeCun's 1989 paper
27:50 - Pig grunt emotion classifier
28:20 - AI annotates protein domain functions
29:40 - Atwood & Carmack: 10k self-driving car bet
30:50 - Helpful Things
References:
Register to GTC and win a 3090!
https://twitter.com/NVIDIAEU/status/1...
https://www.nvidia.com/gtc/keynote/?n...
https://www.nvidia.com/gtc/?ncid=ref-...
https://www.nvidia.com/gtc/keynote/
https://www.nvidia.com/gtc/training/
https://developer.nvidia.com/nvidia-o...
DeepMind deciphers Lost Ancient Texts
https://deepmind.com/blog/article/Pre...
https://www.nature.com/articles/s4158...
https://github.com/deepmind/ithaca
https://ithaca.deepmind.com/?job=eyJy...
Drug discovery model turns toxic
https://www.theverge.com/2022/3/17/22...
https://www.nature.com/articles/s4225...
Gary Marcus: Deep Learning is hitting a wall
https://nautil.us/deep-learning-is-hi...
https://www.youtube.com/watch?v=fVkXE...
GopherCite: Backing up answers with citations
https://deepmind.com/research/publica...
Yoshua Bengio appointed knight of the legion of honour
https://mila.quebec/en/professor-yosh...
Meta AI tags parody account
https://twitter.com/MetaAI/status/150...
Building games using just natural language
https://andrewmayneblog.wordpress.com...
YOU.com adds writing assistant
https://you.com/search?q=how%20to%20w...
Horace He: How to brrr
https://horace.io/brrr_intro.html
Karpathy: Reproducing Yann LeCun's 1989 paper
https://karpathy.github.io/2022/03/14...
Pig grunt emotion classifier
https://science.ku.dk/english/press/n...
AI annotates protein domain functions
https://ai.googleblog.com/2022/03/usi...
https://google-research.github.io/pro...
Atwood & Carmack: 10k self-driving car bet
https://blog.codinghorror.com/the-203...
Helpful Things
https://github.com/recognai/rubrix
https://twitter.com/taiyasaki/status/...
https://github.com/mosaicml/composer?...
https://mujoco.org/
https://mujoco.readthedocs.io/en/late...
https://github.com/deepmind/mctx?utm_...
https://padl.ai/
https://github.com/LaihoE/did-it-spill
https://pytorch.org/blog/pytorch-1.11...
Links:
TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick
YouTube: https://www.youtube.com/c/yannickilcher
Twitter: https://twitter.com/ykilcher
Discord: https://discord.gg/4H8xxDF
BitChute: https://www.bitchute.com/channel/yann...
LinkedIn: https://www.linkedin.com/in/ykilcher
BiliBili: https://space.bilibili.com/2017636191
If you want to support me, the best thing to do is to share out the content :)
This is an interview with the paper's authors: Abhiram Iyer, Karan Grewal, and Akash Velu!
Paper Review Video: https://youtu.be/O_dJ31T01i8
Check out Zak's course on Graph Neural Networks (discount with this link): https://www.graphneuralnets.com/p/int...
Catastrophic forgetting is a big problem in mutli-task and continual learning. Gradients of different objectives tend to conflict, and new tasks tend to override past knowledge. In biological neural networks, each neuron carries a complex network of dendrites that mitigate such forgetting by recognizing the context of an input signal. This paper introduces Active Dendrites, which carries over the principle of context-sensitive gating by dendrites into the deep learning world. Various experiments show the benefit in combatting catastrophic forgetting, while preserving sparsity and limited parameter counts.
OUTLINE:
0:00 - Intro
0:55 - Sponsor: GNN Course
2:30 - How did the idea come to be?
7:05 - What roles do the different parts of the method play?
8:50 - What was missing in the paper review?
10:35 - Are biological concepts viable if we still have backprop?
11:50 - How many dendrites are necessary?
14:10 - Why is there a plateau in the sparsity plot?
20:50 - How does task difficulty play into the algorithm?
24:10 - Why are there different setups in the experiments?
30:00 - Is there a place for unsupervised pre-training?
32:50 - How can we apply the online prototyping to more difficult tasks?
37:00 - What did not work out during the project?
41:30 - How do you debug a project like this?
47:10 - How is this related to other architectures?
51:10 - What other things from neuroscience are to be included?
55:50 - Don't miss the awesome ending :)
Paper: https://arxiv.org/abs/2201.00042
Blog: https://numenta.com/blog/2021/11/08/c...
Link to the GNN course (with discount): https://www.graphneuralnets.com/p/int...
Authors: Abhiram Iyer, Karan Grewal, Akash Velu, Lucas Oliveira Souza, Jeremy Forest, Subutai Ahmad
Links:
TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick
YouTube: https://www.youtube.com/c/yannickilcher
Twitter: https://twitter.com/ykilcher
Discord: https://discord.gg/4H8xxDF
BitChute: https://www.bitchute.com/channel/yann...
LinkedIn: https://www.linkedin.com/in/ykilcher
Catastrophic forgetting is a big problem in mutli-task and continual learning. Gradients of different objectives tend to conflict, and new tasks tend to override past knowledge. In biological neural networks, each neuron carries a complex network of dendrites that mitigate such forgetting by recognizing the context of an input signal. This paper introduces Active Dendrites, which carries over the principle of context-sensitive gating by dendrites into the deep learning world. Various experiments show the benefit in combatting catastrophic forgetting, while preserving sparsity and limited parameter counts.
OUTLINE:
0:00 - Introduction
1:20 - Paper Overview
3:15 - Catastrophic forgetting in continuous and multi-task learning
9:30 - Dendrites in biological neurons
16:55 - Sparse representations in biology
18:35 - Active dendrites in deep learning
34:15 - Experiments on multi-task learning
39:00 - Experiments in continual learning and adaptive prototyping
49:20 - Analyzing the inner workings of the algorithm
53:30 - Is this the same as just training a larger network?
59:15 - How does this relate to attention mechanisms?
1:02:55 - Final thoughts and comments
Paper: https://arxiv.org/abs/2201.00042
Blog: https://numenta.com/blog/2021/11/08/c...
ERRATA:
Abstract:
A key challenge for AI is to build embodied systems that operate in dynamically changing environments. Such systems must adapt to changing task contexts and learn continuously. Although standard deep learning systems achieve state of the art results on static benchmarks, they often struggle in dynamic scenarios. In these settings, error signals from multiple contexts can interfere with one another, ultimately leading to a phenomenon known as catastrophic forgetting. In this article we investigate biologically inspired architectures as solutions to these problems. Specifically, we show that the biophysical properties of dendrites and local inhibitory systems enable networks to dynamically restrict and route information in a context-specific manner. Our key contributions are as follows. First, we propose a novel artificial neural network architecture that incorporates active dendrites and sparse representations into the standard deep learning framework. Next, we study the performance of this architecture on two separate benchmarks requiring task-based adaptation: Meta-World, a multi-task reinforcement learning environment where a robotic agent must learn to solve a variety of manipulation tasks simultaneously; and a continual learning benchmark in which the model's prediction task changes throughout training. Analysis on both benchmarks demonstrates the emergence of overlapping but distinct and sparse subnetworks, allowing the system to fluidly learn multiple tasks with minimal forgetting. Our neural implementation marks the first time a single architecture has achieved competitive results on both multi-task and continual learning settings. Our research sheds light on how biological properties of neurons can inform deep learning systems to address dynamic scenarios that are typically impossible for traditional ANNs to solve.
Authors: Abhiram Iyer, Karan Grewal, Akash Velu, Lucas Oliveira Souza, Jeremy Forest, Subutai Ahmad
An interview with the authors of "Virtual Outlier Synthesis".
Watch the paper review video here: https://youtu.be/i-J4T3uLC9M
Outliers are data points that are highly unlikely to be seen in the training distribution, and therefore deep neural networks have troubles when dealing with them. Many approaches to detecting outliers at inference time have been proposed, but most of them show limited success. This paper presents Virtual Outlier Synthesis, which is a method that pairs synthetic outliers, forged in the latent space, with an energy-based regularization of the network at training time. The result is a deep network that can reliably detect outlier datapoints during inference with minimal overhead.
OUTLINE:
0:00 - Intro
2:20 - What was the motivation behind this paper?
5:30 - Why object detection?
11:05 - What's the connection to energy-based models?
12:15 - Is a Gaussian mixture model appropriate for high-dimensional data?
16:15 - What are the most important components of the method?
18:30 - What are the downstream effects of the regularizer?
22:00 - Are there severe trade-offs to outlier detection?
23:55 - Main experimental takeaways?
26:10 - Why do outlier detection in the last layer?
30:20 - What does it take to finish a research projects successfully?
Paper: https://arxiv.org/abs/2202.01197
Code: https://github.com/deeplearning-wisc/vos
Abstract:
Out-of-distribution (OOD) detection has received much attention lately due to its importance in the safe deployment of neural networks. One of the key challenges is that models lack supervision signals from unknown data, and as a result, can produce overconfident predictions on OOD data. Previous approaches rely on real outlier datasets for model regularization, which can be costly and sometimes infeasible to obtain in practice. In this paper, we present VOS, a novel framework for OOD detection by adaptively synthesizing virtual outliers that can meaningfully regularize the model's decision boundary during training. Specifically, VOS samples virtual outliers from the low-likelihood region of the class-conditional distribution estimated in the feature space. Alongside, we introduce a novel unknown-aware training objective, which contrastively shapes the uncertainty space between the ID data and synthesized outlier data. VOS achieves state-of-the-art performance on both object detection and image classification models, reducing the FPR95 by up to 7.87% compared to the previous best method. Code is available at this https URL.
Authors: Xuefeng Du, Zhaoning Wang, Mu Cai, Yixuan Li
Links:
Merch: store.ykilcher.com
TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick
YouTube: https://www.youtube.com/c/yannickilcher
Twitter: https://twitter.com/ykilcher
Discord: https://discord.gg/4H8xxDF
BitChute: https://www.bitchute.com/channel/yann...
LinkedIn: https://www.linkedin.com/in/ykilcher
BiliBili: https://space.bilibili.com/2017636191
If you want to support me, the best thing to do is to share out the content :)
If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):
SubscribeStar: https://www.subscribestar.com/yannick...
Patreon: https://www.patreon.com/yannickilcher
Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq
Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2
Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m
Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n
Sponsor: Assembly AI
Check them out here: https://www.assemblyai.com/?utm_sourc...
Outliers are data points that are highly unlikely to be seen in the training distribution, and therefore deep neural networks have troubles when dealing with them. Many approaches to detecting outliers at inference time have been proposed, but most of them show limited success. This paper presents Virtual Outlier Synthesis, which is a method that pairs synthetic outliers, forged in the latent space, with an energy-based regularization of the network at training time. The result is a deep network that can reliably detect outlier datapoints during inference with minimal overhead.
OUTLINE:
0:00 - Intro
2:00 - Sponsor: Assembly AI (Link below)
4:05 - Paper Overview
6:45 - Where do traditional classifiers fail?
11:00 - How object detectors work
17:00 - What are virtual outliers and how are they created?
24:00 - Is this really an appropriate model for outliers?
26:30 - How virtual outliers are used during training
34:00 - Plugging it all together to detect outliers
Paper: https://arxiv.org/abs/2202.01197
Code: https://github.com/deeplearning-wisc/vos
Abstract:
Out-of-distribution (OOD) detection has received much attention lately due to its importance in the safe deployment of neural networks. One of the key challenges is that models lack supervision signals from unknown data, and as a result, can produce overconfident predictions on OOD data. Previous approaches rely on real outlier datasets for model regularization, which can be costly and sometimes infeasible to obtain in practice. In this paper, we present VOS, a novel framework for OOD detection by adaptively synthesizing virtual outliers that can meaningfully regularize the model's decision boundary during training. Specifically, VOS samples virtual outliers from the low-likelihood region of the class-conditional distribution estimated in the feature space. Alongside, we introduce a novel unknown-aware training objective, which contrastively shapes the uncertainty space between the ID data and synthesized outlier data. VOS achieves state-of-the-art performance on both object detection and image classification models, reducing the FPR95 by up to 7.87% compared to the previous best method. Code is available at this https URL.
Authors: Xuefeng Du, Zhaoning Wang, Mu Cai, Yixuan Li
Links:
Merch: store.ykilcher.com
TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick
YouTube: https://www.youtube.com/c/yannickilcher
Twitter: https://twitter.com/ykilcher
Discord: https://discord.gg/4H8xxDF
BitChute: https://www.bitchute.com/channel/yann...
LinkedIn: https://www.linkedin.com/in/ykilcher
BiliBili: https://space.bilibili.com/2017636191
If you want to support me, the best thing to do is to share out the content :)
If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):
SubscribeStar: https://www.subscribestar.com/yannick...
Patreon: https://www.patreon.com/yannickilcher
Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq
Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2
Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m
Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n
This is an in-depth paper review, followed by an interview with the papers' authors!
Society is ruled by norms, and most of these norms are very useful, such as washing your hands before cooking. However, there also exist plenty of social norms which are essentially arbitrary, such as what hairstyles are acceptable, or what words are rude. These are called "silly rules". This paper uses multi-agent reinforcement learning to investigate why such silly rules exist. Their results indicate a plausible mechanism, by which the existence of silly rules drastically speeds up the agents' acquisition of the skill of enforcing rules, which generalizes well, and therefore a society that has silly rules will be better at enforcing rules in general, leading to faster adaptation in the face of genuinely useful norms.
OUTLINE:
0:00 - Intro
3:00 - Paper Overview
5:20 - Why are some social norms arbitrary?
11:50 - Reinforcement learning environment setup
20:00 - What happens if we introduce a "silly" rule?
25:00 - Experimental Results: how silly rules help society
30:10 - Isolated probing experiments
34:30 - Discussion of the results
37:30 - Start of Interview
39:30 - Where does the research idea come from?
44:00 - What is the purpose behind this research?
49:20 - Short recap of the mechanics of the environment
53:00 - How much does such a closed system tell us about the real world?
56:00 - What do the results tell us about silly rules?
1:01:00 - What are these agents really learning?
1:08:00 - How many silly rules are optimal?
1:11:30 - Why do you have separate weights for each agent?
1:13:45 - What features could be added next?
1:16:00 - How sensitive is the system to hyperparameters?
1:17:20 - How to avoid confirmation bias?
1:23:15 - How does this play into progress towards AGI?
1:29:30 - Can we make real-world recommendations based on this?
1:32:50 - Where do we go from here?
Paper: https://www.pnas.org/doi/10.1073/pnas...
Blog: https://deepmind.com/research/publica...
Abstract:
The fact that humans enforce and comply with norms is an important reason why humans enjoy higher levels of cooperation and welfare than other animals. Some norms are relatively easy to explain; they may prohibit obviously harmful or uncooperative actions. But many norms are not easy to explain. For example, most cultures prohibit eating certain kinds of foods and almost all societies have rules about what constitutes appropriate clothing, language, and gestures. Using a computational model focused on learning shows that apparently pointless rules can have an indirect effect on welfare. They can help agents learn how to enforce and comply with norms in general, improving the group’s ability to enforce norms that have a direct effect on welfare.
Authors: Raphael Köster, Dylan Hadfield-Menell, Richard Everett, Laura Weidinger, Gillian K. Hadfield, Joel Z. Leibo
Links:
TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick
YouTube: https://www.youtube.com/c/yannickilcher
Twitter: https://twitter.com/ykilcher
Discord: https://discord.gg/4H8xxDF
BitChute: https://www.bitchute.com/channel/yann...
LinkedIn: https://www.linkedin.com/in/ykilcher
BiliBili: https://space.bilibili.com/2017636191
If you want to support me, the best thing to do is to share out the content :)
This is an interview with Stanislas Polu, research engineer at OpenAI and first author of the paper "Formal Mathematics Statement Curriculum Learning".
Watch the paper review here: https://youtu.be/lvYVuOmUVs8
OUTLINE:
0:00 - Intro
2:00 - How do you explain the big public reaction?
4:00 - What's the history behind the paper?
6:15 - How does algorithmic formal math work?
13:10 - How does expert iteration replace self-play?
22:30 - How is the language model trained and used?
30:50 - Why is every model fine-tuned on the initial state?
33:05 - What if we want to prove something we don't know already?
40:35 - How can machines and humans work together?
43:40 - Aren't most produced statements useless?
46:20 - A deeper look at the experimental results
50:10 - What were the high and low points during the research?
54:25 - Where do we go from here?
Paper: https://arxiv.org/abs/2202.01344
miniF2F benchmark: https://github.com/openai/miniF2F
Follow Stan here: https://twitter.com/spolu
Abstract:
We explore the use of expert iteration in the context of language modeling applied to formal mathematics. We show that at same compute budget, expert iteration, by which we mean proof search interleaved with learning, dramatically outperforms proof search only. We also observe that when applied to a collection of formal statements of sufficiently varied difficulty, expert iteration is capable of finding and solving a curriculum of increasingly difficult problems, without the need for associated ground-truth proofs. Finally, by applying this expert iteration to a manually curated set of problem statements, we achieve state-of-the-art on the miniF2F benchmark, automatically solving multiple challenging problems drawn from high school olympiads.
Authors: Stanislas Polu, Jesse Michael Han, Kunhao Zheng, Mantas Baksys, Igor Babuschkin, Ilya Sutskever
Links:
TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick
YouTube: https://www.youtube.com/c/yannickilcher
Twitter: https://twitter.com/ykilcher
Discord: https://discord.gg/4H8xxDF
BitChute: https://www.bitchute.com/channel/yann...
LinkedIn: https://www.linkedin.com/in/ykilcher
BiliBili: https://space.bilibili.com/2017636191
If you want to support me, the best thing to do is to share out the content :)
If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):
SubscribeStar: https://www.subscribestar.com/yannick...
Patreon: https://www.patreon.com/yannickilcher
Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq
Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2
Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m
Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n
Formal mathematics is a challenging area for both humans and machines. For humans, formal proofs require very tedious and meticulous specifications of every last detail and results in very long, overly cumbersome and verbose outputs. For machines, the discreteness and sparse reward nature of the problem presents a significant problem, which is classically tackled by brute force search, guided by a couple of heuristics. Previously, language models have been employed to better guide these proof searches and delivered significant improvements, but automated systems are still far from usable. This paper introduces another concept: An expert iteration procedure is employed to iteratively produce more and more challenging, but solvable problems for the machine to train on, which results in an automated curriculum, and a final algorithm that performs well above the previous models. OpenAI used this method to even solve two problems of the international math olympiad, which was previously infeasible for AI systems.
OUTLINE:
0:00 - Intro
2:35 - Paper Overview
5:50 - How do formal proofs work?
9:35 - How expert iteration creates a curriculum
16:50 - Model, data, and training procedure
25:30 - Predicting proof lengths for guiding search
29:10 - Bootstrapping expert iteration
34:10 - Experimental evaluation & scaling properties
40:10 - Results on synthetic data
44:15 - Solving real math problems
47:15 - Discussion & comments
Paper: https://arxiv.org/abs/2202.01344
miniF2F benchmark: https://github.com/openai/miniF2F
Abstract:
We explore the use of expert iteration in the context of language modeling applied to formal mathematics. We show that at same compute budget, expert iteration, by which we mean proof search interleaved with learning, dramatically outperforms proof search only. We also observe that when applied to a collection of formal statements of sufficiently varied difficulty, expert iteration is capable of finding and solving a curriculum of increasingly difficult problems, without the need for associated ground-truth proofs. Finally, by applying this expert iteration to a manually curated set of problem statements, we achieve state-of-the-art on the miniF2F benchmark, automatically solving multiple challenging problems drawn from high school olympiads.
Authors: Stanislas Polu, Jesse Michael Han, Kunhao Zheng, Mantas Baksys, Igor Babuschkin, Ilya Sutskever
Links:
TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick
YouTube: https://www.youtube.com/c/yannickilcher
Twitter: https://twitter.com/ykilcher
Discord: https://discord.gg/4H8xxDF
BitChute: https://www.bitchute.com/channel/yann...
LinkedIn: https://www.linkedin.com/in/ykilcher
BiliBili: https://space.bilibili.com/2017636191
If you want to support me, the best thing to do is to share out the content :)
If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):
SubscribeStar: https://www.subscribestar.com/yannick...
Patreon: https://www.patreon.com/yannickilcher
Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq
Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2
Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m
Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n
Updates on what's going on in the ML world!
Check out w&b's alerts feature: https://wandb.me/yannic
OUTLINE:
0:00 - Intro
0:20 - Sponsor: Weights & Biases
2:35 - DeepMind uses Reinforcement Learning to control nuclear fusion
4:35 - Google responds to carbon emission estimates
8:40 - Yann LeCun proposes new architecture for world models
11:05 - Fruit fly neurons may perform multiplication
12:00 - Emojisearch App
12:30 - Ar5iv officially in arXiv labs
12:55 - Language Model Consciousness & Media Hype
16:45 - Vision models are more fair when trained on uncurated data
18:30 - CLIPasso
19:15 - NLP with Transformers Book
20:15 - Helpful Things
26:00 - US Office: AI can't copyright its art
Sponsor: Weights & Biases
https://wandb.me/yannic
References:
https://wandb.me/yannic
DeepMind uses RL to control nuclear fusion
https://deepmind.com/blog/article/Acc...
https://www.nature.com/articles/s4158...
https://www.nature.com/articles/s4158...
https://www.alexirpan.com/2018/02/14/...
Google responds to carbon emission estimates
https://ai.googleblog.com/2022/02/goo...
Yann LeCun proposes new architecture for world models
https://ai.facebook.com/blog/yann-lec...
Fruit fly neurons may perform multiplication
https://www.nature.com/articles/s4158...
Emojisearch App
https://twitter.com/lilianweng/status...
https://www.emojisearch.app/
https://github.com/lilianweng/emoji-s...
Ar5iv officially in arXiv labs
https://blog.arxiv.org/2022/02/21/arx...
Tech media may be only slightly conscious
https://twitter.com/ilyasut/status/14...
https://futurism.com/the-byte/openai-...
https://interestingengineering.com/ai...
https://futurism.com/mit-researcher-c...
https://www.dailymail.co.uk/sciencete...
https://futurism.com/conscious-ai-bac...
https://www.dailystar.co.uk/tech/news...
Vision models are more fair when trained on uncurated data
https://arxiv.org/pdf/2202.08360.pdf
CLIPasso
https://clipasso.github.io/clipasso/
NLP with Transformers Book
https://www.amazon.de/dp/1098103246?l...
Helpful Things
https://github.com/j3soon/tbparse
https://github.com/openvinotoolkit/an...
https://liuliu66.github.io/articulati...
https://github.com/RobertTLange/evosax
https://github.com/google/evojax
https://github.com/google/evojax/pull/9
https://github.com/facebookresearch/t...
https://standard-ai.github.io/Standar...
https://twitter.com/PatrickPlaten/sta...
https://aimagelab.ing.unimore.it/imag...
https://github.com/yashbhalgat/HashNe...
https://github.com/patrick-kidger/dif...
https://github.com/AI4Finance-Foundat...
https://huggingface.co/AI-Nordics/ber...
https://huggingface.co/AI-Nordics/gpt...
https://paperswithcode.com/dataset/muld
https://github.com/JonasGeiping/breac...
https://github.com/Weixin-Liang/MetaS...
US Office: AI can't copyright its art
https://www.theverge.com/2022/2/21/22...
https://www.urbasm.com/2016/05/artifi...
Links:
TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick
YouTube: https://www.youtube.com/c/yannickilcher
Twitter: https://twitter.com/ykilcher
Discord: https://discord.gg/4H8xxDF
BitChute: https://www.bitchute.com/channel/yann...
LinkedIn: https://www.linkedin.com/in/ykilcher
BiliBili: https://space.bilibili.com/2017636191
An interview with the creators of AlphaCode!
Paper review video here: https://youtu.be/s9UAOmyah1A
OUTLINE:
0:00 - Intro
1:10 - Media Reception
5:10 - How did the project go from start to finish?
9:15 - Does the model understand its own code?
14:45 - Are there plans to reduce the number of samples?
16:15 - Could one do smarter filtering of samples?
18:55 - How crucial are the public test cases?
21:55 - Could we imagine an adversarial method?
24:45 - How are coding problems even made?
27:40 - Does AlphaCode evaluate a solution's asymptotic complexity?
33:15 - Are our sampling procedures inappropriate for diversity?
36:30 - Are all generated solutions as instructive as the example?
41:30 - How are synthetic examples created during training?
42:30 - What were high and low points during this research?
45:25 - What was the most valid criticism after publication?
47:40 - What are applications in the real world?
51:00 - Where do we go from here?
Paper: https://storage.googleapis.com/deepmi...
Code: https://github.com/deepmind/code_cont...
Abstract: Programming is a powerful and ubiquitous problem-solving tool. Developing systems that can assist programmers or even generate programs independently could make programming more productive and accessible, yet so far incorporating innovations in AI has proven challenging. Recent large-scale language models have demonstrated an impressive ability to generate code, and are now able to complete simple programming tasks. However, these models still perform poorly when evaluated on more complex, unseen problems that require problem-solving skills beyond simply translating instructions into code. For example, competitive programming problems which require an understanding of algorithms and complex natural language remain extremely challenging. To address this gap, we introduce AlphaCode, a system for code generation that can create novel solutions to these problems that require deeper reasoning. Evaluated on recent programming competitions on the Codeforces platform, AlphaCode achieved on average a ranking of top 54.3% in programming competitions with more than 5,000 participants. We found that three key components were critical to achieve good and reliable performance: (1) an extensive and clean competitive programming dataset for training and evaluation, (2) large and efficient-to-sample transformer-based architectures, and (3) large-scale model sampling to explore the search space, followed by filtering based on program behavior to a small set of submissions.
Authors: Yujia Li, David Choi, Junyoung Chung, Nate Kushman, Julian Schrittwieser, Rémi Leblond, Tom Eccles, James Keeling, Felix Gimeno, Agustin Dal Lago, Thomas Hubert, Peter Choy, Cyprien de Masson d’Autume, Igor Babuschkin, Xinyun Chen, Po-Sen Huang, Johannes Welbl, Sven Gowal, Alexey Cherepanov, James Molloy, Daniel J. Mankowitz, Esme Sutherland Robson, Pushmeet Kohli, Nando de Freitas, Koray Kavukcuoglu and Oriol Vinyals
Links:
Merch: store.ykilcher.com
TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick
YouTube: https://www.youtube.com/c/yannickilcher
Twitter: https://twitter.com/ykilcher
Discord: https://discord.gg/4H8xxDF
BitChute: https://www.bitchute.com/channel/yann...
LinkedIn: https://www.linkedin.com/in/ykilcher
BiliBili: https://space.bilibili.com/2017636191
If you want to support me, the best thing to do is to share out the content :)
If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):
SubscribeStar: https://www.subscribestar.com/yannick...
Patreon: https://www.patreon.com/yannickilcher
AlphaCode is an automated system that can solve competitive programing exercises. The authors found an interesting combination of language models, large-scale sampling, and clever techniques to filter and subsequently cluster the resulting programs, which lets the system perform on the level of an average competitor in real competitions. In this video, we take a deep dive into AlphaCode's design, architecture, and experimental evaluation. The paper is very well structured and the empirical results are super interesting!
OUTLINE:
0:00 - Intro
2:10 - Paper Overview
3:30 - An example problem from competitive programming
8:00 - AlphaCode system overview
14:00 - Filtering out wrong solutions
17:15 - Clustering equivalent generated programs
21:50 - Model configurations & engineering choices
24:30 - Adding privileged information to the input & more tricks
28:15 - Experimental Results (very interesting!)
Paper: https://storage.googleapis.com/deepmi...
Code: https://github.com/deepmind/code_cont...
Abstract: Programming is a powerful and ubiquitous problem-solving tool. Developing systems that can assist programmers or even generate programs independently could make programming more productive and accessible, yet so far incorporating innovations in AI has proven challenging. Recent large-scale language models have demonstrated an impressive ability to generate code, and are now able to complete simple programming tasks. However, these models still perform poorly when evaluated on more complex, unseen problems that require problem-solving skills beyond simply translating instructions into code. For example, competitive programming problems which require an understanding of algorithms and complex natural language remain extremely challenging. To address this gap, we introduce AlphaCode, a system for code generation that can create novel solutions to these problems that require deeper reasoning. Evaluated on recent programming competitions on the Codeforces platform, AlphaCode achieved on average a ranking of top 54.3% in programming competitions with more than 5,000 participants. We found that three key components were critical to achieve good and reliable performance: (1) an extensive and clean competitive programming dataset for training and evaluation, (2) large and efficient-to-sample transformer-based architectures, and (3) large-scale model sampling to explore the search space, followed by filtering based on program behavior to a small set of submissions.
Authors: Yujia Li, David Choi, Junyoung Chung, Nate Kushman, Julian Schrittwieser, Rémi Leblond, Tom Eccles, James Keeling, Felix Gimeno, Agustin Dal Lago, Thomas Hubert, Peter Choy, Cyprien de Masson d’Autume, Igor Babuschkin, Xinyun Chen, Po-Sen Huang, Johannes Welbl, Sven Gowal, Alexey Cherepanov, James Molloy, Daniel J. Mankowitz, Esme Sutherland Robson, Pushmeet Kohli, Nando de Freitas, Koray Kavukcuoglu and Oriol Vinyals
Links:
Merch: store.ykilcher.com
TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick
YouTube: https://www.youtube.com/c/yannickilcher
Twitter: https://twitter.com/ykilcher
Discord: https://discord.gg/4H8xxDF
BitChute: https://www.bitchute.com/channel/yann...
LinkedIn: https://www.linkedin.com/in/ykilcher
BiliBili: https://space.bilibili.com/2017636191
If you want to support me, the best thing to do is to share out the content :)
If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):
SubscribeStar: https://www.subscribestar.com/yannick...
Patreon: https://www.patreon.com/yannickilcher
Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq
Original paper review here: https://youtu.be/XHGh19Hbx48
Machel Reid and Yutaro Yamada join me to discuss their recent paper on langauge model pre-training for decision transformers in offline reinforcement learning.
OUTLINE:
0:00 - Intro
1:00 - Brief paper, setup & idea recap
7:30 - Main experimental results & high standard deviations
10:00 - Why is there no clear winner?
13:00 - Why are bigger models not a lot better?
14:30 - What’s behind the name ChibiT?
15:30 - Why is iGPT underperforming?
19:15 - How are tokens distributed in Reinforcement Learning?
22:00 - What other domains could have good properties to transfer?
24:20 - A deeper dive into the models' attention patterns
33:30 - Codebase, model sizes, and compute requirements
37:30 - Scaling behavior of pre-trained models
40:05 - What did not work out in this project?
42:00 - How can people get started and where to go next?
Paper: https://arxiv.org/abs/2201.12122
Code: https://github.com/machelreid/can-wik...
My Video on Decision Transformer: https://youtu.be/-buULmf7dec
Abstract:
Fine-tuning reinforcement learning (RL) models has been challenging because of a lack of large scale off-the-shelf datasets as well as high variance in transferability among different environments. Recent work has looked at tackling offline RL from the perspective of sequence modeling with improved results as result of the introduction of the Transformer architecture. However, when the model is trained from scratch, it suffers from slow convergence speeds. In this paper, we look to take advantage of this formulation of reinforcement learning as sequence modeling and investigate the transferability of pre-trained sequence models on other domains (vision, language) when finetuned on offline RL tasks (control, games). To this end, we also propose techniques to improve transfer between these domains. Results show consistent performance gains in terms of both convergence speed and reward on a variety of environments, accelerating training by 3-6x and achieving state-of-the-art performance in a variety of tasks using Wikipedia-pretrained and GPT2 language models. We hope that this work not only brings light to the potentials of leveraging generic sequence modeling techniques and pre-trained models for RL, but also inspires future work on sharing knowledge between generative modeling tasks of completely different domains.
Authors: Machel Reid, Yutaro Yamada, Shixiang Shane Gu
Links:
TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick
YouTube: https://www.youtube.com/c/yannickilcher
Twitter: https://twitter.com/ykilcher
Discord: https://discord.gg/4H8xxDF
BitChute: https://www.bitchute.com/channel/yann...
LinkedIn: https://www.linkedin.com/in/ykilcher
BiliBili: https://space.bilibili.com/2017636191
If you want to support me, the best thing to do is to share out the content :)
If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):
SubscribeStar: https://www.subscribestar.com/yannick...
Patreon: https://www.patreon.com/yannickilcher
Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq
Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2
Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m
Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n
Transformers have come to overtake many domain-targeted custom models in a wide variety of fields, such as Natural Language Processing, Computer Vision, Generative Modelling, and recently also Reinforcement Learning. This paper looks at the Decision Transformer and shows that, surprisingly, pre-training the model on a language-modelling task significantly boosts its performance on Offline Reinforcement Learning. The resulting model achieves higher scores, can get away with less parameters, and exhibits superior scaling properties. This raises many questions about the fundamental connection between the domains of language and RL.
OUTLINE:
0:00 - Intro
1:35 - Paper Overview
7:35 - Offline Reinforcement Learning as Sequence Modelling
12:00 - Input Embedding Alignment & other additions
16:50 - Main experimental results
20:45 - Analysis of the attention patterns across models
32:25 - More experimental results (scaling properties, ablations, etc.)
37:30 - Final thoughts
Paper: https://arxiv.org/abs/2201.12122
Code: https://github.com/machelreid/can-wik...
My Video on Decision Transformer: https://youtu.be/-buULmf7dec
Abstract:
Fine-tuning reinforcement learning (RL) models has been challenging because of a lack of large scale off-the-shelf datasets as well as high variance in transferability among different environments. Recent work has looked at tackling offline RL from the perspective of sequence modeling with improved results as result of the introduction of the Transformer architecture. However, when the model is trained from scratch, it suffers from slow convergence speeds. In this paper, we look to take advantage of this formulation of reinforcement learning as sequence modeling and investigate the transferability of pre-trained sequence models on other domains (vision, language) when finetuned on offline RL tasks (control, games). To this end, we also propose techniques to improve transfer between these domains. Results show consistent performance gains in terms of both convergence speed and reward on a variety of environments, accelerating training by 3-6x and achieving state-of-the-art performance in a variety of tasks using Wikipedia-pretrained and GPT2 language models. We hope that this work not only brings light to the potentials of leveraging generic sequence modeling techniques and pre-trained models for RL, but also inspires future work on sharing knowledge between generative modeling tasks of completely different domains.
Authors: Machel Reid, Yutaro Yamada, Shixiang Shane Gu
Links:
Merch: store.ykilcher.com
TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick
YouTube: https://www.youtube.com/c/yannickilcher
Twitter: https://twitter.com/ykilcher
Discord: https://discord.gg/4H8xxDF
BitChute: https://www.bitchute.com/channel/yann...
LinkedIn: https://www.linkedin.com/in/ykilcher
BiliBili: https://space.bilibili.com/2017636191
If you want to support me, the best thing to do is to share out the content :)
If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):
SubscribeStar: https://www.subscribestar.com/yannick...
Patreon: https://www.patreon.com/yannickilcher
Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq
Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2
Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m
Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n
Some things we've missed in recent weeks!
OUTLINE:
0:00 - Intro & Overview
0:40 - Meta builds AI Research Supercluster (RSC)
2:25 - OpenAI trains GPT-3 to follow instructions
4:10 - Meta AI releases multilingual language models
4:50 - Google LaMDA dialogue models
5:50 - Helpful Things
8:25 - Training the alpha matte generator for Pixel 6
10:15 - Drones used to deter pigeons on buildings
11:05 - IBM sells some Watson Health assets for USD 1B
Merch: store.ykilcher.com
References:
https://ai.facebook.com/blog/ai-rsc/?...
https://openai.com/blog/instruction-f...
https://cdn.openai.com/papers/Trainin...
https://openai.com/blog/deep-reinforc...
https://twitter.com/MetaAI/status/148...
https://arxiv.org/pdf/2112.10668.pdf
https://github.com/pytorch/fairseq/tr...
https://ai.googleblog.com/2022/01/lam...
https://arxiv.org/pdf/2201.08239.pdf
https://evolutiongym.github.io/?utm_s...
https://evolutiongym.github.io/all-tasks
https://evolutiongym.github.io/docume...
https://arxiv.org/pdf/2201.09863.pdf
https://github.com/EvolutionGym
https://huggingface.co/blog/sb3
https://twitter.com/Sentdex/status/14...
https://github.com/lvwerra/trl?utm_so...
https://ai.googleblog.com/2022/01/acc...
https://polyhaven.com/hdris
https://ieeexplore.ieee.org/document/...
https://www.bloomberg.com/news/articl...
https://archive.ph/xadf9
Links:
TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick
YouTube: https://www.youtube.com/c/yannickilcher
Twitter: https://twitter.com/ykilcher
Discord: https://discord.gg/4H8xxDF
BitChute: https://www.bitchute.com/channel/yann...
LinkedIn: https://www.linkedin.com/in/ykilcher
BiliBili: https://space.bilibili.com/2017636191
If you want to support me, the best thing to do is to share out the content :)
If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):
SubscribeStar: https://www.subscribestar.com/yannick...
Patreon: https://www.patreon.com/yannickilcher
Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq
Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2
Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m
Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n
Your regularly irregular updates on everything new in the ML world!
Merch: store.ykilcher.com
OUTLINE:
0:00 - Intro
0:15 - Sponsor: Weights & Biases
2:15 - Uber switches from XGBoost to Deep Learning for ETA prediction
5:45 - MuZero advances video compression
10:10 - Learned Soft Prompts can steer large language models
12:45 - Block-NeRF captures entire city blocks
14:15 - Neural Architecture Search considers underlying hardware
16:50 - Mega-Blog on Self-Organizing Agents
18:40 - Know Your Data (for Tensorflow Datasets)
20:30 - Helpful Things
Sponsor: Weights & Biases
https://wandb.me/yannic
References:
https://docs.wandb.ai/guides/integrat...
https://colab.research.google.com/git...
https://wandb.ai/borisd13/GPT-3/repor...
Uber switches from XGBoost to Deep Learning for ETA prediction
https://eng.uber.com/deepeta-how-uber...
MuZero advances video compression
https://deepmind.com/blog/article/MuZ...
https://storage.googleapis.com/deepmi...
Learned Soft Prompts can steer large language models
https://ai.googleblog.com/2022/02/gui...
https://aclanthology.org/2021.emnlp-m...
Block-NeRF captures entire city blocks
https://arxiv.org/abs/2202.05263
https://arxiv.org/pdf/2202.05263.pdf
https://waymo.com/intl/zh-cn/research...
Neural Architecture Search considers underlying hardware
https://ai.googleblog.com/2022/02/unl...
https://openaccess.thecvf.com/content...
Mega-Blog on Self-Organizing Agents
https://developmentalsystems.org/sens...
https://flowers.inria.fr/
Know Your Data (for Tensorflow Datasets)
https://knowyourdata-tfds.withgoogle....
https://knowyourdata.withgoogle.com/
Helpful Things
https://twitter.com/casualganpapers/s...
https://www.reddit.com/r/MachineLearn...
https://arxiv.org/abs/2202.02435
https://github.com/vicariousinc/PGMax
https://www.vicarious.com/posts/pgmax...
https://diambra.ai/tournaments
https://github.com/diambra/diambraArena
https://www.youtube.com/watch?v=dw72P...
https://gitlab.com/deepcypher/python-...
https://python-fhez.readthedocs.io/en...
https://joss.theoj.org/papers/10.2110...
https://github.com/PyTorchLightning/m...
https://torchmetrics.readthedocs.io/e...
https://twitter.com/alanyttian/status...
https://github.com/google/evojax
https://arxiv.org/abs/2202.05008
https://www.reddit.com/r/MachineLearn...
https://www.gymlibrary.ml/pages/api/#...
Links:
TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick
YouTube: https://www.youtube.com/c/yannickilcher
Twitter: https://twitter.com/ykilcher
Discord: https://discord.gg/4H8xxDF
BitChute: https://www.bitchute.com/channel/yann...
LinkedIn: https://www.linkedin.com/in/ykilcher
BiliBili: https://space.bilibili.com/2017636191
If you want to support me, the best thing to do is to share out the content :)
If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):
SubscribeStar: https://www.subscribestar.com/yannick...
Patreon: https://www.patreon.com/yannickilcher
Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq
Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2
Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m
Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n
Many of you have given me feedback on what you did and didn't like about the recent "with the authors" videos. Here's the result of that feedback and an outlook into the future.
Merch: store.ykilcher.com
Links:
TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick
YouTube: https://www.youtube.com/c/yannickilcher
Twitter: https://twitter.com/ykilcher
Discord: https://discord.gg/4H8xxDF
BitChute: https://www.bitchute.com/channel/yann...
LinkedIn: https://www.linkedin.com/in/ykilcher
BiliBili: https://space.bilibili.com/2017636191
If you want to support me, the best thing to do is to share out the content :)
If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):
SubscribeStar: https://www.subscribestar.com/yannick...
Patreon: https://www.patreon.com/yannickilcher
Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq
Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2
Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m
Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n
This video is an interview with Adi Fuchs, author of a series called "AI Accelerators", and an expert in modern AI acceleration technology.
Accelerators like GPUs and TPUs are an integral part of today's AI landscape. Deep Neural Network training can be sped up by orders of magnitudes by making good use of these specialized pieces of hardware. However, GPUs and TPUs are only the beginning of a vast landscape of emerging technologies and companies that build accelerators for the next generation of AI models. In this interview, we go over many aspects of building hardware for AI, including why GPUs have been so successful, what the most promising approaches look like, how they work, and what the main challenges are.
OUTLINE:
0:00 - Intro
5:10 - What does it mean to make hardware for AI?
8:20 - Why were GPUs so successful?
16:25 - What is "dark silicon"?
20:00 - Beyond GPUs: How can we get even faster AI compute?
28:00 - A look at today's accelerator landscape
30:00 - Systolic Arrays and VLIW
35:30 - Reconfigurable dataflow hardware
40:50 - The failure of Wave Computing
42:30 - What is near-memory compute?
46:50 - Optical and Neuromorphic Computing
49:50 - Hardware as enabler and limiter
55:20 - Everything old is new again
1:00:00 - Where to go to dive deeper?
Read the full blog series here:
Part I: https://medium.com/@adi.fu7/ai-accele...
Part II: https://medium.com/@adi.fu7/ai-accele...
Part III: https://medium.com/@adi.fu7/ai-accele...
Part IV: https://medium.com/@adi.fu7/ai-accele...
Part V: https://medium.com/@adi.fu7/ai-accele...
Links:
Merch: store.ykilcher.com
TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick
YouTube: https://www.youtube.com/c/yannickilcher
Twitter: https://twitter.com/ykilcher
Discord: https://discord.gg/4H8xxDF
BitChute: https://www.bitchute.com/channel/yann...
LinkedIn: https://www.linkedin.com/in/ykilcher
BiliBili: https://space.bilibili.com/2017636191
If you want to support me, the best thing to do is to share out the content :)
If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):
SubscribeStar: https://www.subscribestar.com/yannick...
Patreon: https://www.patreon.com/yannickilcher
Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq
Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2
Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m
Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n
This video contains a paper explanation and an incredibly informative interview with first author Armen Aghajanyan.
Autoregressive Transformers have come to dominate many fields in Machine Learning, from text generation to image creation and many more. However, there are two problems. First, the collected data is usually scraped from the web and uni- or bi-modal and throws away a lot of structure of the original websites, and second, language modelling losses are uni-directional. CM3 addresses both problems: It directly operates on HTML and includes text, hyperlinks, and even images (via VQGAN tokenization) and can therefore be used in plenty of ways: Text generation, captioning, image creation, entity linking, and much more. It also introduces a new training strategy called Causally Masked Language Modelling, which brings a level of bi-directionality into autoregressive language modelling. In the interview after the paper explanation, Armen and I go deep into the how and why of these giant models, we go over the stunning results and we make sense of what they mean for the future of universal models.
OUTLINE:
0:00 - Intro & Overview
6:30 - Directly learning the structure of HTML
12:30 - Causally Masked Language Modelling
18:50 - A short look at how to use this model
23:20 - Start of interview
25:30 - Feeding language models with HTML
29:45 - How to get bi-directionality into decoder-only Transformers?
37:00 - Images are just tokens
41:15 - How does one train such giant models?
45:40 - CM3 results are amazing
58:20 - Large-scale dataset collection and content filtering
1:04:40 - More experimental results
1:12:15 - Why don't we use raw HTML?
1:18:20 - Does this paper contain too many things?
Paper: https://arxiv.org/abs/2201.07520
Abstract:
We introduce CM3, a family of causally masked generative models trained over a large corpus of structured multi-modal documents that can contain both text and image tokens. Our new causally masked approach generates tokens left to right while also masking out a small number of long token spans that are generated at the end of the string, instead of their original positions. The casual masking object provides a type of hybrid of the more common causal and masked language models, by enabling full generative modeling while also providing bidirectional context when generating the masked spans. We train causally masked language-image models on large-scale web and Wikipedia articles, where each document contains all of the text, hypertext markup, hyperlinks, and image tokens (from a VQVAE-GAN), provided in the order they appear in the original HTML source (before masking). The resulting CM3 models can generate rich structured, multi-modal outputs while conditioning on arbitrary masked document contexts, and thereby implicitly learn a wide range of text, image, and cross modal tasks. They can be prompted to recover, in a zero-shot fashion, the functionality of models such as DALL-E, GENRE, and HTLM. We set the new state-of-the-art in zero-shot summarization, entity linking, and entity disambiguation while maintaining competitive performance in the fine-tuning setting. We can generate images unconditionally, conditioned on text (like DALL-E) and do captioning all in a zero-shot setting with a single model.
Authors: Armen Aghajanyan, Bernie Huang, Candace Ross, Vladimir Karpukhin, Hu Xu, Naman Goyal, Dmytro Okhonko, Mandar Joshi, Gargi Ghosh, Mike Lewis, Luke Zettlemoyer
Most of us conceive the internet as a free and open space where we are able to send traffic between any two nodes, but for large parts of the world this is not the case. Entire nations have large machinery in place to survey all internet traffic and automated procedures to block any undesirable connections. Evading such censorship has been largely a cat-and-mouse game between security researchers and government actors. A new system, called Geneva, uses a Genetic Algorithm in combination with Evolutionary Search in order to dynamically evade such censorship and adjust itself in real-time to any potential response by its adversaries. In this video, I talk to Security researcher Kevin Bock, who is one of Geneva's main contributors and member of the Breakerspace project. We talk about the evolution of internet censorship, how to evade it, how to mess with the censors' infrastructure, as well as the broader emerging connections between AI and Security.
OUTLINE:
0:00 - Intro
3:30 - What is automated censorship in networks?
7:20 - The evolution of censorship vs evasion
12:40 - Why do we need a dynamic, evolving system?
16:30 - The building blocks of Geneva
23:15 - Introducing evolution
28:30 - What's the censors' response?
31:45 - How was Geneva's media reception?
33:15 - Where do we go from here?
37:30 - Can we deliberately attack the censors?
47:00 - On responsible disclosure
49:40 - Breakerspace: Security research for undergrads
50:40 - How often do you get into trouble?
52:10 - How can I get started in security?
Learn more at:
Geneva (& more) project page: https://censorship.ai
Open Observatory of Network Interference: https://ooni.org
Censored Planet: https://censoredplanet.org
Breakerspace: https://breakerspace.cs.umd.edu
Links:
Merch: store.ykilcher.com
TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick
YouTube: https://www.youtube.com/c/yannickilcher
Twitter: https://twitter.com/ykilcher
Discord: https://discord.gg/4H8xxDF
BitChute: https://www.bitchute.com/channel/yann...
LinkedIn: https://www.linkedin.com/in/ykilcher
BiliBili: https://space.bilibili.com/2017636191
If you want to support me, the best thing to do is to share out the content :)
If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):
SubscribeStar: https://www.subscribestar.com/yannick...
Patreon: https://www.patreon.com/yannickilcher
Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq
Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2
Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m
Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n
This video contains a paper explanation and an interview with author Andrey Zhmoginov!
Few-shot learning is an interesting sub-field in meta-learning, with wide applications, such as creating personalized models based on just a handful of data points. Traditionally, approaches have followed the BERT approach where a large model is pre-trained and then fine-tuned. However, this couples the size of the final model to the size of the model that has been pre-trained. Similar problems exist with "true" meta-learners, such as MaML. HyperTransformer fundamentally decouples the meta-learner from the size of the final model by directly predicting the weights of the final model. The HyperTransformer takes the few-shot dataset as a whole into its context and predicts either one or multiple layers of a (small) ConvNet, meaning its output are the weights of the convolution filters. Interestingly, and with the correct engineering care, this actually appears to deliver promising results and can be extended in many ways.
OUTLINE:
0:00 - Intro & Overview
3:05 - Weight-generation vs Fine-tuning for few-shot learning
10:10 - HyperTransformer model architecture overview
22:30 - Why the self-attention mechanism is useful here
34:45 - Start of Interview
39:45 - Can neural networks even produce weights of other networks?
47:00 - How complex does the computational graph get?
49:45 - Why are transformers particularly good here?
58:30 - What can the attention maps tell us about the algorithm?
1:07:00 - How could we produce larger weights?
1:09:30 - Diving into experimental results
1:14:30 - What questions remain open?
Paper: https://arxiv.org/abs/2201.04182
ERRATA: I introduce Max Vladymyrov as Mark Vladymyrov
Abstract:
In this work we propose a HyperTransformer, a transformer-based model for few-shot learning that generates weights of a convolutional neural network (CNN) directly from support samples. Since the dependence of a small generated CNN model on a specific task is encoded by a high-capacity transformer model, we effectively decouple the complexity of the large task space from the complexity of individual tasks. Our method is particularly effective for small target CNN architectures where learning a fixed universal task-independent embedding is not optimal and better performance is attained when the information about the task can modulate all model parameters. For larger models we discover that generating the last layer alone allows us to produce competitive or better results than those obtained with state-of-the-art methods while being end-to-end differentiable. Finally, we extend our approach to a semi-supervised regime utilizing unlabeled samples in the support set and further improving few-shot performance.
Authors: Andrey Zhmoginov, Mark Sandler, Max Vladymyrov
Links:
TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick
YouTube: https://www.youtube.com/c/yannickilcher
Twitter: https://twitter.com/ykilcher
Discord: https://discord.gg/4H8xxDF
BitChute: https://www.bitchute.com/channel/yann...
LinkedIn: https://www.linkedin.com/in/ykilcher
BiliBili: https://space.bilibili.com/2017636191
If you want to support me, the best thing to do is to share out the content :)
If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):
SubscribeStar: https://www.subscribestar.com/yannick...
Patreon: https://www.patreon.com/yannickilcher
Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq
The latest and greatest from the world of Machine Learning!
Merch: store.ykilcher.com
Sponsor: Weights & Biases
https://wandb.me/yannic
OUTLINE:
0:00 - Intro
0:15 - Sponsor: Weights & Biases
3:15 - DeepMind's AlphaCode: AI competitive programmer
11:30 - OpenAI uses language models to prove math theorems
14:30 - StyleGAN XL: Scaling StyleGAN to diverse datasets
16:10 - ar5iv.org displays papers as HTML5
17:40 - Helpful Things
19:30 - ICML22 Review process changes
21:15 - Meta AI tackles harmful content classification using few-shot learning
23:55 - Company claims to produce face images from DNA
References:
https://deepmind.com/blog/article/Com...
https://alphacode.deepmind.com/#layer...
https://storage.googleapis.com/deepmi...
https://twitter.com/DBahdanau/status/...
https://openai.com/blog/formal-math/
https://arxiv.org/pdf/2202.01344.pdf
https://blog.eleuther.ai/announcing-2...
https://sites.google.com/view/stylega...
https://arxiv.org/pdf/2202.00273.pdf
https://ar5iv.org/
https://ar5iv.org/html/1910.06709
https://twitter.com/YiTayML/status/14...
https://ffcv.io/
https://github.com/ott-jax/ott
https://twitter.com/soumithchintala/s...
https://github.com/facebookresearch/d...
https://www.reddit.com/r/MachineLearn...
https://icml.cc/Conferences/2022/Revi...
https://icml.cc/Conferences/2022/Call...
https://ai.facebook.com/blog/harmful-...
https://www.technologyreview.com/2022...
Links:
TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick
YouTube: https://www.youtube.com/c/yannickilcher
Twitter: https://twitter.com/ykilcher
Discord: https://discord.gg/4H8xxDF
BitChute: https://www.bitchute.com/channel/yann...
LinkedIn: https://www.linkedin.com/in/ykilcher
BiliBili: https://space.bilibili.com/2017636191
If you want to support me, the best thing to do is to share out the content :)
If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):
SubscribeStar: https://www.subscribestar.com/yannick...
Patreon: https://www.patreon.com/yannickilcher
Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq
Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2
Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m
Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n
In this video: Paper explanation, followed by first author interview with Wenlong Huang.
Large language models contain extraordinary amounts of world knowledge that can be queried in various ways. But their output format is largely uncontrollable. This paper investigates the VirtualHome environment, which expects a particular set of actions, objects, and verbs to be used. Turns out, with proper techniques and only using pre-trained models (no fine-tuning), one can translate unstructured language model outputs into the structured grammar of the environment. This is potentially very useful anywhere where the models' world knowledge needs to be provided in a particular structured format.
OUTLINE:
0:00 - Intro & Overview
2:45 - The VirtualHome environment
6:25 - The problem of plan evaluation
8:40 - Contributions of this paper
16:40 - Start of interview
24:00 - How to use language models with environments?
34:00 - What does model size matter?
40:00 - How to fix the large models' outputs?
55:00 - Possible improvements to the translation procedure
59:00 - Why does Codex perform so well?
1:02:15 - Diving into experimental results
1:14:15 - Future outlook
Paper: https://arxiv.org/abs/2201.07207
Website: https://wenlong.page/language-planner/
Code: https://github.com/huangwl18/language...
Wenlong's Twitter: https://twitter.com/wenlong_huang
Abstract:
Can world knowledge learned by large language models (LLMs) be used to act in interactive environments? In this paper, we investigate the possibility of grounding high-level tasks, expressed in natural language (e.g. "make breakfast"), to a chosen set of actionable steps (e.g. "open fridge"). While prior work focused on learning from explicit step-by-step examples of how to act, we surprisingly find that if pre-trained LMs are large enough and prompted appropriately, they can effectively decompose high-level tasks into low-level plans without any further training. However, the plans produced naively by LLMs often cannot map precisely to admissible actions. We propose a procedure that conditions on existing demonstrations and semantically translates the plans to admissible actions. Our evaluation in the recent VirtualHome environment shows that the resulting method substantially improves executability over the LLM baseline. The conducted human evaluation reveals a trade-off between executability and correctness but shows a promising sign towards extracting actionable knowledge from language models. Website at this https URL
Authors: Wenlong Huang, Pieter Abbeel, Deepak Pathak, Igor Mordatch
Links:
Merch: store.ykilcher.com
TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick
YouTube: https://www.youtube.com/c/yannickilcher
Twitter: https://twitter.com/ykilcher
Discord: https://discord.gg/4H8xxDF
BitChute: https://www.bitchute.com/channel/yann...
LinkedIn: https://www.linkedin.com/in/ykilcher
BiliBili: https://space.bilibili.com/2017636191
If you want to support me, the best thing to do is to share out the content :)
If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):
SubscribeStar: https://www.subscribestar.com/yannick...
Patreon: https://www.patreon.com/yannickilcher
Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq
Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2
Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m
Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n
COMMENTS DIRECTLY FROM THE AUTHOR (thanks a lot for reaching out Arvind :) ):
The FIQA results you share also have code to reproduce the results in the paper using the API: https://twitter.com/arvind_io/status/... There's no discrepancy AFAIK.
We leave out 6 not 7 BEIR datasets. Results on msmarco, nq and triviaqa are in a separate table (Table 5 in the paper). NQ is part of BEIR too and we didn't want to repeat it. Finally, the 6 datasets we leave out are not readily available and it is common to leave them out in prior work too. For examples, see SPLADE v2 (https://arxiv.org/pdf/2109.10086.pdf) also evaluates on the same 12 BEIR datasets.
Finally, I'm now working on time travel so that I can cite papers from the future :)
END COMMENTS FROM THE AUTHOR
OpenAI launches an embeddings endpoint in their API, providing high-dimensional vector embeddings for use in text similarity, text search, and code search. While embeddings are universally recognized as a standard tool to process natural language, people have raised doubts about the quality of OpenAI's embeddings, as one blog post found they are often outperformed by open-source models, which are much smaller and with which embedding would cost a fraction of what OpenAI charges. In this video, we examine the claims made and determine what it all means.
OUTLINE:
0:00 - Intro
0:30 - Sponsor: Weights & Biases
2:20 - What embeddings are available?
3:55 - OpenAI shows promising results
5:25 - How good are the results really?
6:55 - Criticism: Open models might be cheaper and smaller
10:05 - Discrepancies in the results
11:00 - The author's response
11:50 - Putting things into perspective
13:35 - What about real world data?
14:40 - OpenAI's pricing strategy: Why so expensive?
Sponsor: Weights & Biases
https://wandb.me/yannic
Merch: store.ykilcher.com
ERRATA: At 13:20 I say "better", it should be "worse"
References:
https://openai.com/blog/introducing-t...
https://arxiv.org/pdf/2201.10005.pdf
https://beta.openai.com/docs/guides/e...
https://beta.openai.com/docs/api-refe...
https://twitter.com/Nils_Reimers/stat...
https://medium.com/@nils_reimers/open...
https://mobile.twitter.com/arvind_io/...
https://twitter.com/gwern/status/1487...
https://twitter.com/gwern/status/1487...
https://twitter.com/Nils_Reimers/stat...
https://twitter.com/gwern/status/1470...
https://www.reddit.com/r/MachineLearn...
https://mobile.twitter.com/arvind_io/...
https://mobile.twitter.com/arvind_io/...
Links:
TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick
YouTube: https://www.youtube.com/c/yannickilcher
Twitter: https://twitter.com/ykilcher
Discord: https://discord.gg/4H8xxDF
BitChute: https://www.bitchute.com/channel/yann...
LinkedIn: https://www.linkedin.com/in/ykilcher
BiliBili: https://space.bilibili.com/2017636191
If you want to support me, the best thing to do is to share out the content :)
If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):
SubscribeStar: https://www.subscribestar.com/yannick...
Patreon: https://www.patreon.com/yannickilcher
Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq
Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2
Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m
Originally, Deep Learning sprang into existence inspired by how the brain processes information, but the two fields have diverged ever since. However, given that deep models can solve many perception tasks with remarkable accuracy, is it possible that we might be able to learn something about how the brain works by inspecting our models? I speak to Patrick Mineault about his blog post "2021 in review: unsupervised brain models" and we explore why neuroscientists are taking interest in unsupervised and self-supervised deep neural networks in order to explain how the brain works. We discuss a series of influential papers that have appeared last year, and we go into the more general questions of connecting neuroscience and machine learning.
OUTLINE:
0:00 - Intro & Overview
6:35 - Start of Interview
10:30 - Visual processing in the brain
12:50 - How does deep learning inform neuroscience?
21:15 - Unsupervised training explains the ventral stream
30:50 - Predicting own motion parameters explains the dorsal stream
42:20 - Why are there two different visual streams?
49:45 - Concept cells and representation learning
56:20 - Challenging the manifold theory
1:08:30 - What are current questions in the field?
1:13:40 - Should the brain inform deep learning?
1:18:50 - Neuromatch Academy and other endeavours
Blog Post: https://xcorr.net/2021/12/31/2021-in-...
Patrick's Blog: https://xcorr.net/
Twitter: https://twitter.com/patrickmineault
Neuromatch Academy: https://academy.neuromatch.io/
Links:
TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick
YouTube: https://www.youtube.com/c/yannickilcher
Twitter: https://twitter.com/ykilcher
Discord: https://discord.gg/4H8xxDF
BitChute: https://www.bitchute.com/channel/yann...
LinkedIn: https://www.linkedin.com/in/ykilcher
BiliBili: https://space.bilibili.com/2017636191
If you want to support me, the best thing to do is to share out the content :)
If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):
SubscribeStar: https://www.subscribestar.com/yannick...
Patreon: https://www.patreon.com/yannickilcher
Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq
Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2
Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m
EleutherAI announces GPT-NeoX-20B, a 20 billion parameter open-source language model, inspired by GPT-3. Connor joins me to discuss the process of training, how the group got their hands on the necessary hardware, what the new model can do, and how anyone can try it out!
OUTLINE:
0:00 - Intro
1:00 - Start of interview
2:00 - How did you get all the hardware?
3:50 - What's the scale of this model?
6:00 - A look into the experimental results
11:15 - Why are there GPT-Neo, GPT-J, and GPT-NeoX?
14:15 - How difficult is training these big models?
17:00 - Try out the model on GooseAI
19:00 - Final thoughts
Read the announcement: https://blog.eleuther.ai/announcing-20b/
Try out the model: https://goose.ai/
Check out EleutherAI: https://www.eleuther.ai/
Read the code: https://github.com/EleutherAI/gpt-neox
Hardware sponsor: https://www.coreweave.com/
Links:
TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick
YouTube: https://www.youtube.com/c/yannickilcher
Twitter: https://twitter.com/ykilcher
Discord: https://discord.gg/4H8xxDF
BitChute: https://www.bitchute.com/channel/yann...
LinkedIn: https://www.linkedin.com/in/ykilcher
BiliBili: https://space.bilibili.com/2017636191
If you want to support me, the best thing to do is to share out the content :)
If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):
SubscribeStar: https://www.subscribestar.com/yannick...
Patreon: https://www.patreon.com/yannickilcher
Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq
Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2
Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m
Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n
This video includes an interview with first author Stéphane d'Ascoli (https://sdascoli.github.io/).
Deep neural networks are typically excellent at numeric regression, but using them for symbolic computation has largely been ignored so far. This paper uses transformers to do symbolic regression on integer and floating point number sequences, which means that given the start of a sequence of numbers, the model has to not only predict the correct continuation, but also predict the data generating formula behind the sequence. Through clever encoding of the input space and a well constructed training data generation process, this paper's model can learn and represent many of the sequences in the OEIS, the online encyclopedia of integer sequences and it also features an interactive demo if you want to try it by yourself.
OUTLINE:
0:00 - Introduction
2:20 - Summary of the Paper
16:10 - Start of Interview
17:15 - Why this research direction?
20:45 - Overview of the method
30:10 - Embedding space of input tokens
33:00 - Data generation process
42:40 - Why are transformers useful here?
46:40 - Beyond number sequences, where is this useful?
48:45 - Success cases and failure cases
58:10 - Experimental Results
1:06:30 - How did you overcome difficulties?
1:09:25 - Interactive demo
Paper: https://arxiv.org/abs/2201.04600
Interactive demo: http://recur-env.eba-rm3fchmn.us-east...
Abstract:
Symbolic regression, i.e. predicting a function from the observation of its values, is well-known to be a challenging task. In this paper, we train Transformers to infer the function or recurrence relation underlying sequences of integers or floats, a typical task in human IQ tests which has hardly been tackled in the machine learning literature. We evaluate our integer model on a subset of OEIS sequences, and show that it outperforms built-in Mathematica functions for recurrence prediction. We also demonstrate that our float model is able to yield informative approximations of out-of-vocabulary functions and constants, e.g. bessel0(x)≈sin(x)+cos(x)πx√ and 1.644934≈π2/6. An interactive demonstration of our models is provided at this https URL.
Authors: Stéphane d'Ascoli, Pierre-Alexandre Kamienny, Guillaume Lample, François Charton
Links:
TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick
YouTube: https://www.youtube.com/c/yannickilcher
Twitter: https://twitter.com/ykilcher
Discord: https://discord.gg/4H8xxDF
BitChute: https://www.bitchute.com/channel/yann...
LinkedIn: https://www.linkedin.com/in/ykilcher
BiliBili: https://space.bilibili.com/2017636191
If you want to support me, the best thing to do is to share out the content :)
If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):
SubscribeStar: https://www.subscribestar.com/yannick...
Patreon: https://www.patreon.com/yannickilcher
Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq
Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2
Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m
LIMITED TIME MERCH DEAL: http://store.ykilcher.com
Links:
TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick
YouTube: https://www.youtube.com/c/yannickilcher
Twitter: https://twitter.com/ykilcher
Discord: https://discord.gg/4H8xxDF
BitChute: https://www.bitchute.com/channel/yann...
LinkedIn: https://www.linkedin.com/in/ykilcher
BiliBili: https://space.bilibili.com/2017636191
If you want to support me, the best thing to do is to share out the content :)
If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):
SubscribeStar: https://www.subscribestar.com/yannick...
Patreon: https://www.patreon.com/yannickilcher
Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq
Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2
Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m
Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n
Your update on what's new in the Machine Learning world!
OUTLINE:
0:00 - Intro
0:15 - ConvNeXt: Return of the Convolutions
2:50 - Investigating Saliency Cropping Algorithms
9:40 - YourTTS: SOTA zero-shot Text-to-Speech
10:40 - MT3: Multi-Track Music Transcription
11:35 - China regulates addictive algorithms
13:00 - A collection of Deep Learning interview questions & solutions
13:35 - Helpful Things
16:05 - AlphaZero explained blog post
16:45 - Ru-DOLPH: HyperModal Text-to-Image-to-Text model
17:45 - Google AI 2021 Review
References:
ConvNeXt: Return of the Convolutions
https://arxiv.org/abs/2201.03545
https://github.com/facebookresearch/C...
https://twitter.com/giffmana/status/1...
https://twitter.com/wightmanr/status/...
https://twitter.com/tanmingxing/statu...
Investigating Saliency Cropping Algorithms
https://openaccess.thecvf.com/content...
https://vinayprabhu.github.io/Salienc...
https://vinayprabhu.medium.com/on-the...
https://vinayprabhu.github.io/Salienc...
YourTTS: SOTA zero-shot Text-to-Speech
https://github.com/coqui-ai/TTS?utm_s...
https://arxiv.org/abs/2112.02418?utm_...
https://coqui.ai/?utm_source=pocket_m...
https://coqui.ai/blog/tts/yourtts-zer...
MT3: Multi-Track Music Transcription
https://arxiv.org/abs/2111.03017
https://github.com/magenta/mt3
https://huggingface.co/spaces/akhaliq...
https://www.reddit.com/r/MachineLearn...
China regulates addictive algorithms
https://technode.com/2022/01/05/china...
https://qz.com/2109618/china-reveals-...
A collection of Deep Learning interview questions & solutions
https://arxiv.org/abs/2201.00650?utm_...
https://arxiv.org/pdf/2201.00650.pdf
Helpful Things
https://docs.deepchecks.com/en/stable...
https://github.com/deepchecks/deepchecks
https://docs.deepchecks.com/en/stable...
https://www.dagshub.com/
https://www.dagshub.com/docs/index.html
https://www.dagshub.com/blog/launchin...
https://bayesiancomputationbook.com/w...
https://mlcontests.com/
https://github.com/Yard1/ray-skorch
https://github.com/skorch-dev/skorch
https://www.rumbledb.org/?utm_source=...
https://github.com/DarshanDeshpande/j...
https://github.com/s3prl/s3prl
AlphaZero explained blog post
https://joshvarty.github.io/AlphaZero...
Ru-DOLPH: HyperModal Text-to-Image-to-Text model
https://github.com/sberbank-ai/ru-dolph
https://colab.research.google.com/dri...
Google AI 2021 Review
https://ai.googleblog.com/2022/01/goo...
Links:
TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick
YouTube: https://www.youtube.com/c/yannickilcher
Twitter: https://twitter.com/ykilcher
Discord: https://discord.gg/4H8xxDF
BitChute: https://www.bitchute.com/channel/yann...
LinkedIn: https://www.linkedin.com/in/ykilcher
BiliBili: https://space.bilibili.com/2017636191
If you want to support me, the best thing to do is to share out the content :)
If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):
SubscribeStar: https://www.subscribestar.com/yannick...
Patreon: https://www.patreon.com/yannickilcher
Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq
Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2
Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m
This video includes an interview with the paper's authors!
What if we treated deep networks like modular programs? Neural Interpreters divide computation into small modules and route data to them via a dynamic type inference system. The resulting model combines recurrent elements, weight sharing, attention, and more to tackle both abstract reasoning, as well as computer vision tasks.
OUTLINE:
0:00 - Intro & Overview
3:00 - Model Overview
7:00 - Interpreter weights and function code
9:40 - Routing data to functions via neural type inference
14:55 - ModLin layers
18:25 - Experiments
21:35 - Interview Start
24:50 - General Model Structure
30:10 - Function code and signature
40:30 - Explaining Modulated Layers
49:50 - A closer look at weight sharing
58:30 - Experimental Results
Paper: https://arxiv.org/abs/2110.06399
Guests:
Nasim Rahaman: https://twitter.com/nasim_rahaman
Francesco Locatello: https://twitter.com/FrancescoLocat8
Waleed Gondal: https://twitter.com/Wallii_gondal
Abstract:
Modern neural network architectures can leverage large amounts of data to generalize well within the training distribution. However, they are less capable of systematic generalization to data drawn from unseen but related distributions, a feat that is hypothesized to require compositional reasoning and reuse of knowledge. In this work, we present Neural Interpreters, an architecture that factorizes inference in a self-attention network as a system of modules, which we call \emph{functions}. Inputs to the model are routed through a sequence of functions in a way that is end-to-end learned. The proposed architecture can flexibly compose computation along width and depth, and lends itself well to capacity extension after training. To demonstrate the versatility of Neural Interpreters, we evaluate it in two distinct settings: image classification and visual abstract reasoning on Raven Progressive Matrices. In the former, we show that Neural Interpreters perform on par with the vision transformer using fewer parameters, while being transferrable to a new task in a sample efficient manner. In the latter, we find that Neural Interpreters are competitive with respect to the state-of-the-art in terms of systematic generalization
Authors: Nasim Rahaman, Muhammad Waleed Gondal, Shruti Joshi, Peter Gehler, Yoshua Bengio, Francesco Locatello, Bernhard Schölkopf
Links:
TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick
YouTube: https://www.youtube.com/c/yannickilcher
Twitter: https://twitter.com/ykilcher
Discord: https://discord.gg/4H8xxDF
BitChute: https://www.bitchute.com/channel/yann...
LinkedIn: https://www.linkedin.com/in/ykilcher
BiliBili: https://space.bilibili.com/2017636191
If you want to support me, the best thing to do is to share out the content :)
If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):
SubscribeStar: https://www.subscribestar.com/yannick...
Patreon: https://www.patreon.com/yannickilcher
Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq
Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2
Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m
Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n
This video includes an interview with first author Ferran Alet!
Encoding inductive biases has been a long established methods to provide deep networks with the ability to learn from less data. Especially useful are encodings of symmetry properties of the data, such as the convolution's translation invariance. But such symmetries are often hard to program explicitly, and can only be encoded exactly when done in a direct fashion. Noether Networks use Noether's theorem connecting symmetries to conserved quantities and are able to dynamically and approximately enforce symmetry properties upon deep neural networks.
OUTLINE:
0:00 - Intro & Overview
18:10 - Interview Start
21:20 - Symmetry priors vs conserved quantities
23:25 - Example: Pendulum
27:45 - Noether Network Model Overview
35:35 - Optimizing the Noether Loss
41:00 - Is the computation graph stable?
46:30 - Increasing the inference time computation
48:45 - Why dynamically modify the model?
55:30 - Experimental Results & Discussion
Paper: https://arxiv.org/abs/2112.03321
Website: https://dylandoblar.github.io/noether...
Code: https://github.com/dylandoblar/noethe...
Abstract:
Progress in machine learning (ML) stems from a combination of data availability, computational resources, and an appropriate encoding of inductive biases. Useful biases often exploit symmetries in the prediction problem, such as convolutional networks relying on translation equivariance. Automatically discovering these useful symmetries holds the potential to greatly improve the performance of ML systems, but still remains a challenge. In this work, we focus on sequential prediction problems and take inspiration from Noether's theorem to reduce the problem of finding inductive biases to meta-learning useful conserved quantities. We propose Noether Networks: a new type of architecture where a meta-learned conservation loss is optimized inside the prediction function. We show, theoretically and experimentally, that Noether Networks improve prediction quality, providing a general framework for discovering inductive biases in sequential problems.
Authors: Ferran Alet, Dylan Doblar, Allan Zhou, Joshua Tenenbaum, Kenji Kawaguchi, Chelsea Finn
Links:
TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick
YouTube: https://www.youtube.com/c/yannickilcher
Twitter: https://twitter.com/ykilcher
Discord: https://discord.gg/4H8xxDF
BitChute: https://www.bitchute.com/channel/yann...
LinkedIn: https://www.linkedin.com/in/ykilcher
BiliBili: https://space.bilibili.com/2017636191
If you want to support me, the best thing to do is to share out the content :)
If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):
SubscribeStar: https://www.subscribestar.com/yannick...
Patreon: https://www.patreon.com/yannickilcher
Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq
Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2
Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m
Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n
The MineRL BASALT challenge has no reward functions or technical descriptions of what's to be achieved. Instead, the goal of each task is given as a short natural language string, and the agent is evaluated by a team of human judges who rate both how well the goal has been fulfilled, as well as how human-like the agent behaved. In this video, I interview KAIROS, the winning team of the 2021 challenge, and discuss how they used a combination of machine learning, efficient data collection, hand engineering, and a bit of knowledge about Minecraft to beat all other teams.
OUTLINE:
0:00 - Introduction
4:10 - Paper Overview
11:15 - Start of Interview
17:05 - First Approach
20:30 - State Machine
26:45 - Efficient Label Collection
30:00 - Navigation Policy
38:15 - Odometry Estimation
46:00 - Pain Points & Learnings
50:40 - Live Run Commentary
58:50 - What other tasks can be solved?
1:01:55 - What made the difference?
1:07:30 - Recommendations & Conclusion
1:11:10 - Full Runs: Waterfall
1:12:40 - Full Runs: Build House
1:17:45 - Full Runs: Animal Pen
1:20:50 - Full Runs: Find Cave
Paper: https://arxiv.org/abs/2112.03482
Code: https://github.com/viniciusguigo/kair...
Challenge Website: https://minerl.io/basalt/
Paper Title: Combining Learning from Human Feedback and Knowledge Engineering to Solve Hierarchical Tasks in Minecraft
Abstract:
Real-world tasks of interest are generally poorly defined by human-readable descriptions and have no pre-defined reward signals unless it is defined by a human designer. Conversely, data-driven algorithms are often designed to solve a specific, narrowly defined, task with performance metrics that drives the agent's learning. In this work, we present the solution that won first place and was awarded the most human-like agent in the 2021 NeurIPS Competition MineRL BASALT Challenge: Learning from Human Feedback in Minecraft, which challenged participants to use human data to solve four tasks defined only by a natural language description and no reward function. Our approach uses the available human demonstration data to train an imitation learning policy for navigation and additional human feedback to train an image classifier. These modules, together with an estimated odometry map, are then combined into a state-machine designed based on human knowledge of the tasks that breaks them down in a natural hierarchy and controls which macro behavior the learning agent should follow at any instant. We compare this hybrid intelligence approach to both end-to-end machine learning and pure engineered solutions, which are then judged by human evaluators. Codebase is available at this https URL.
Authors: Vinicius G. Goecks, Nicholas Waytowich, David Watkins, Bharat Prakash
Links:
TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick
YouTube: https://www.youtube.com/c/yannickilcher
Twitter: https://twitter.com/ykilcher
Discord: https://discord.gg/4H8xxDF
BitChute: https://www.bitchute.com/channel/yann...
LinkedIn: https://www.linkedin.com/in/ykilcher
BiliBili: https://space.bilibili.com/2017636191
If you want to support me, the best thing to do is to share out the content :)
If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):
SubscribeStar: https://www.subscribestar.com/yannick...
Patreon: https://www.patreon.com/yannickilcher
Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq
Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2
Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m
The MineRL BASALT challenge has no reward functions or technical descriptions of what's to be achieved. Instead, the goal of each task is given as a short natural language string, and the agent is evaluated by a team of human judges who rate both how well the goal has been fulfilled, as well as how human-like the agent behaved. In this video, I interview KAIROS, the winning team of the 2021 challenge, and discuss how they used a combination of machine learning, efficient data collection, hand engineering, and a bit of knowledge about Minecraft to beat all other teams.
OUTLINE:
0:00 - Introduction
4:10 - Paper Overview
11:15 - Start of Interview
17:05 - First Approach
20:30 - State Machine
26:45 - Efficient Label Collection
30:00 - Navigation Policy
38:15 - Odometry Estimation
46:00 - Pain Points & Learnings
50:40 - Live Run Commentary
58:50 - What other tasks can be solved?
1:01:55 - What made the difference?
1:07:30 - Recommendations & Conclusion
1:11:10 - Full Runs: Waterfall
1:12:40 - Full Runs: Build House
1:17:45 - Full Runs: Animal Pen
1:20:50 - Full Runs: Find Cave
Paper: https://arxiv.org/abs/2112.03482
Code: https://github.com/viniciusguigo/kair...
Challenge Website: https://minerl.io/basalt/
Paper Title: Combining Learning from Human Feedback and Knowledge Engineering to Solve Hierarchical Tasks in Minecraft
Abstract:
Real-world tasks of interest are generally poorly defined by human-readable descriptions and have no pre-defined reward signals unless it is defined by a human designer. Conversely, data-driven algorithms are often designed to solve a specific, narrowly defined, task with performance metrics that drives the agent's learning. In this work, we present the solution that won first place and was awarded the most human-like agent in the 2021 NeurIPS Competition MineRL BASALT Challenge: Learning from Human Feedback in Minecraft, which challenged participants to use human data to solve four tasks defined only by a natural language description and no reward function. Our approach uses the available human demonstration data to train an imitation learning policy for navigation and additional human feedback to train an image classifier. These modules, together with an estimated odometry map, are then combined into a state-machine designed based on human knowledge of the tasks that breaks them down in a natural hierarchy and controls which macro behavior the learning agent should follow at any instant. We compare this hybrid intelligence approach to both end-to-end machine learning and pure engineered solutions, which are then judged by human evaluators. Codebase is available at this https URL.
Authors: Vinicius G. Goecks, Nicholas Waytowich, David Watkins, Bharat Prakash
Links:
TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick
YouTube: https://www.youtube.com/c/yannickilcher
Twitter: https://twitter.com/ykilcher
Discord: https://discord.gg/4H8xxDF
BitChute: https://www.bitchute.com/channel/yann...
LinkedIn: https://www.linkedin.com/in/ykilcher
BiliBili: https://space.bilibili.com/2017636191
If you want to support me, the best thing to do is to share out the content :)
If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):
SubscribeStar: https://www.subscribestar.com/yannick...
Patreon: https://www.patreon.com/yannickilcher
Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq
Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2
Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m
Watch the original podcast: https://www.youtube.com/watch?v=DxREm...
An analysis of Elon's appearance on Lex Fridman. Very interesting conversation and a good overview of past, current, and future versions of Tesla's Autopilot system.
OUTLINE:
0:00 - Intro
0:40 - Tesla Autopilot: How hard is it?
9:05 - Building an accurate understanding of the world
16:25 - History of Tesla's neural network stack
26:00 - When is full self-driving ready?
29:55 - FSD 11: Less code, more neural networks
37:00 - Auto-labelling is essential
39:05 - Tesla Bot & Discussion
Links:
TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick
YouTube: https://www.youtube.com/c/yannickilcher
Twitter: https://twitter.com/ykilcher
Discord: https://discord.gg/4H8xxDF
BitChute: https://www.bitchute.com/channel/yann...
LinkedIn: https://www.linkedin.com/in/ykilcher
BiliBili: https://space.bilibili.com/2017636191
If you want to support me, the best thing to do is to share out the content :)
If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):
SubscribeStar: https://www.subscribestar.com/yannick...
Patreon: https://www.patreon.com/yannickilcher
Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq
Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2
Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m
Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n
Special Guest: First author Martin Schmid (https://twitter.com/Lifrordi)
Games have been used throughout research as testbeds for AI algorithms, such as reinforcement learning agents. However, different types of games usually require different solution approaches, such as AlphaZero for Go or Chess, and Counterfactual Regret Minimization (CFR) for Poker. Player of Games bridges this gap between perfect and imperfect information games and delivers a single algorithm that uses tree search over public information states, and is trained via self-play. The resulting algorithm can play Go, Chess, Poker, Scotland Yard, and many more games, as well as non-game environments.
OUTLINE:
0:00 - Introduction
2:50 - What games can Player of Games be trained on?
4:00 - Tree search algorithms (AlphaZero)
8:00 - What is different in imperfect information games?
15:40 - Counterfactual Value- and Policy-Networks
18:50 - The Player of Games search procedure
28:30 - How to train the network?
34:40 - Experimental Results
47:20 - Discussion & Outlook
Paper: https://arxiv.org/abs/2112.03178
Abstract:
Games have a long history of serving as a benchmark for progress in artificial intelligence. Recently, approaches using search and learning have shown strong performance across a set of perfect information games, and approaches using game-theoretic reasoning and learning have shown strong performance for specific imperfect information poker variants. We introduce Player of Games, a general-purpose algorithm that unifies previous approaches, combining guided search, self-play learning, and game-theoretic reasoning. Player of Games is the first algorithm to achieve strong empirical performance in large perfect and imperfect information games -- an important step towards truly general algorithms for arbitrary environments. We prove that Player of Games is sound, converging to perfect play as available computation time and approximation capacity increases. Player of Games reaches strong performance in chess and Go, beats the strongest openly available agent in heads-up no-limit Texas hold'em poker (Slumbot), and defeats the state-of-the-art agent in Scotland Yard, an imperfect information game that illustrates the value of guided search, learning, and game-theoretic reasoning.
Authors: Martin Schmid, Matej Moravcik, Neil Burch, Rudolf Kadlec, Josh Davidson, Kevin Waugh, Nolan Bard, Finbarr Timbers, Marc Lanctot, Zach Holland, Elnaz Davoodi, Alden Christianson, Michael Bowling
Links:
TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick
YouTube: https://www.youtube.com/c/yannickilcher
Twitter: https://twitter.com/ykilcher
Discord: https://discord.gg/4H8xxDF
BitChute: https://www.bitchute.com/channel/yann...
LinkedIn: https://www.linkedin.com/in/ykilcher
BiliBili: https://space.bilibili.com/2017636191
If you want to support me, the best thing to do is to share out the content :)
If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):
SubscribeStar: https://www.subscribestar.com/yannick...
Patreon: https://www.patreon.com/yannickilcher
Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq
Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2
Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m
Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n
Diffusion models learn to iteratively reverse a noising process that is applied repeatedly during training. The result can be used for conditional generation as well as various other tasks such as inpainting. OpenAI's GLIDE builds on recent advances in diffusion models and combines text-conditional diffusion with classifier-free guidance and upsampling to achieve unprecedented quality in text-to-image samples.
Try it yourself: https://huggingface.co/spaces/valhall...
OUTLINE:
0:00 - Intro & Overview
6:10 - What is a Diffusion Model?
18:20 - Conditional Generation and Guided Diffusion
31:30 - Architecture Recap
34:05 - Training & Result metrics
36:55 - Failure cases & my own results
39:45 - Safety considerations
Paper: https://arxiv.org/abs/2112.10741
Code & Model: https://github.com/openai/glide-text2im
More diffusion papers:
https://arxiv.org/pdf/2006.11239.pdf
https://arxiv.org/pdf/2102.09672.pdf
Abstract:
Diffusion models have recently been shown to generate high-quality synthetic images, especially when paired with a guidance technique to trade off diversity for fidelity. We explore diffusion models for the problem of text-conditional image synthesis and compare two different guidance strategies: CLIP guidance and classifier-free guidance. We find that the latter is preferred by human evaluators for both photorealism and caption similarity, and often produces photorealistic samples. Samples from a 3.5 billion parameter text-conditional diffusion model using classifier-free guidance are favored by human evaluators to those from DALL-E, even when the latter uses expensive CLIP reranking. Additionally, we find that our models can be fine-tuned to perform image inpainting, enabling powerful text-driven image editing. We train a smaller model on a filtered dataset and release the code and weights at this https URL.
Authors: Alex Nichol, Prafulla Dhariwal, Aditya Ramesh, Pranav Shyam, Pamela Mishkin, Bob McGrew, Ilya Sutskever, Mark Chen
Links:
TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick
YouTube: https://www.youtube.com/c/yannickilcher
Twitter: https://twitter.com/ykilcher
Discord: https://discord.gg/4H8xxDF
BitChute: https://www.bitchute.com/channel/yann...
LinkedIn: https://www.linkedin.com/in/ykilcher
BiliBili: https://space.bilibili.com/2017636191
If you want to support me, the best thing to do is to share out the content :)
If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):
SubscribeStar: https://www.subscribestar.com/yannick...
Patreon: https://www.patreon.com/yannickilcher
Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq
Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2
Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m
Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n
Your updates on everything going on in the Machine Learning world.
Sponsor: Weights & Biases
https://wandb.me/yannic
OUTLINE:
0:00 - Intro & Overview
0:20 - Sponsor: Weights & Biases
3:05 - DeepMind releases 3 papers on large language models
11:45 - Hugging Face Blog: Training CodeParrot from scratch
14:25 - Paper: Pre-Training vision systems with noise
15:45 - DeepMind advances Quantum Mechanics
16:45 - GoogleAI trains GLaM: 1 Trillion Parameters Mixture of Experts Model
18:45 - Colin Raffel calls for building ML models like we build Open-Source software
22:05 - A rebuke of the hype around DeepMind's math paper
24:45 - Helpful Things
32:25 - Suicide Capsule plans AI to assess your mental state before use
35:15 - Synthesia raises 50M to develop AI avatars
Weights & Biases Embedding Projector
https://twitter.com/_ScottCondron/sta...
https://docs.wandb.ai/ref/app/feature...
https://wandb.ai/timssweeney/toy_data...
DeepMind releases 3 papers on large language models
https://deepmind.com/blog/article/lan...
https://arxiv.org/pdf/2112.04426.pdf
https://kstatic.googleusercontent.com...
https://arxiv.org/pdf/2112.04359.pdf
https://deepmind.com/research/publica...
Hugging Face Blog: Training CodeParrot from scratch
https://huggingface.co/blog/codeparro...
Paper: Pre-Training vision systems with noise
https://mbaradad.github.io/learning_w...
DeepMind advances Quantum Mechanics
https://deepmind.com/blog/article/Sim...
https://storage.googleapis.com/deepmi...
https://github.com/deepmind/deepmind-...
GoogleAI trains GLaM: 1 Trillion Parameters Mixture of Experts Model
https://ai.googleblog.com/2021/12/mor...
Colin Raffel calls for building ML models like we build Open-Source software
https://colinraffel.com/blog/a-call-t...
A rebuke of the hype around DeepMind's math paper
https://arxiv.org/abs/2112.04324?s=09
Helpful Things
https://twitter.com/huggingface/statu...
https://docs.cohere.ai/prompt-enginee...
https://github.blog/2021-12-08-improv...
https://huggingface.co/blog/data-meas...
https://huggingface.co/spaces/hugging...
https://blogs.microsoft.com/ai-for-bu...
https://techcommunity.microsoft.com/t...
https://github.com/minitorch/minitorc...
https://minitorch.github.io/
https://pandastutor.com/
https://pandastutor.com/vis.html
https://github.com/IAmPara0x/yuno
https://colab.research.google.com/dri...
https://www.reddit.com/r/MachineLearn...
https://www.drivendata.org/competitio...
https://www.reddit.com/r/MachineLearn...
https://www.uttt.ai/
https://arxiv.org/abs/2112.02721?utm_...
https://arxiv.org/pdf/2112.02721.pdf
https://github.com/GEM-benchmark/NL-A...
https://www.reddit.com/r/MachineLearn...
Suicide Capsule plans AI to assess your mental state before use
https://www.swissinfo.ch/eng/sci-tech...
Synthesia raises 50M to develop AI avatars
https://techcrunch.com/2021/12/08/syn...
https://www.synthesia.io/
Links:
TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick
YouTube: https://www.youtube.com/c/yannickilcher
Twitter: https://twitter.com/ykilcher
Discord: https://discord.gg/4H8xxDF
SubscribeStar: https://www.subscribestar.com/yannick...
Patreon: https://www.patreon.com/yannickilcher
Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq
At the end of the video is an interview with the paper authors!
LaMa is a system that is amazing at removing foreground objects from images, especially when those objects cover a large part of the image itself. LaMa is specifically trained to reconstruct large masked areas and includes global information throughout its forward propagation by using Fourier Convolutions in its layers. This makes it incredibly effective at reconstructing periodic structures with long-range consistency, compared to regular convolutions.
OUTLINE:
0:00 - Intro
0:45 - Sponsor: ClearML
3:30 - Inpainting Examples
5:05 - Live Demo
6:40 - Locality as a weakness of convolutions
10:30 - Using Fourier Transforms for global information
12:55 - Model architecture overview
14:35 - Fourier convolution layer
21:15 - Loss function
24:25 - Mask generation algorithm
25:40 - Experimental results
28:25 - Interview with the authors
Paper: https://arxiv.org/abs/2109.07161
Code: https://github.com/saic-mdal/lama
Online Demo: https://cleanup.pictures/
Sponsor: ClearML
https://clear.ml
Abstract:
Modern image inpainting systems, despite the significant progress, often struggle with large missing areas, complex geometric structures, and high-resolution images. We find that one of the main reasons for that is the lack of an effective receptive field in both the inpainting network and the loss function. To alleviate this issue, we propose a new method called large mask inpainting (LaMa). LaMa is based on i) a new inpainting network architecture that uses fast Fourier convolutions (FFCs), which have the image-wide receptive field; ii) a high receptive field perceptual loss; iii) large training masks, which unlocks the potential of the first two components. Our inpainting network improves the state-of-the-art across a range of datasets and achieves excellent performance even in challenging scenarios, e.g. completion of periodic structures. Our model generalizes surprisingly well to resolutions that are higher than those seen at train time, and achieves this at lower parameter&time costs than the competitive baselines. The code is available at \url{this https URL}.
Authors: Roman Suvorov, Elizaveta Logacheva, Anton Mashikhin, Anastasia Remizova, Arsenii Ashukha, Aleksei Silvestrov, Naejin Kong, Harshith Goka, Kiwoong Park, Victor Lempitsky
Links:
TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick
YouTube: https://www.youtube.com/c/yannickilcher
Twitter: https://twitter.com/ykilcher
Discord: https://discord.gg/4H8xxDF
BitChute: https://www.bitchute.com/channel/yann...
LinkedIn: https://www.linkedin.com/in/ykilcher
BiliBili: https://space.bilibili.com/2017636191
If you want to support me, the best thing to do is to share out the content :)
If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):
SubscribeStar: https://www.subscribestar.com/yannick...
Patreon: https://www.patreon.com/yannickilcher
Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq
Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2
Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m
Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n
The most trusted model in News!
Get started with Weights & Biases here: https://wandb.me/yannic
(it's free forever for personal use)
OUTLINE:
0:00 - Intro
0:15 - Sponsor: Weights & Biases
3:10 - DeepMind tackles fundamental math
6:45 - Microsoft focuses on scaling effectively and efficiently
10:15 - NeurIPS Anthology Visualization
13:30 - Timnit Gebru launches research institute independent from big tech
16:50 - SageMaker Canvas for no-code ML
17:50 - Help, Help!
21:40 - Cornelius Emde wins the 3090
21:55 - A retrospective on the NeurIPS 2021 ethics review process
References:
DeepMind tackles fundamental math
https://deepmind.com/blog/article/exp...
https://www.nature.com/articles/s4158...
Microsoft focuses on scaling effectively and efficiently
https://www.microsoft.com/en-us/resea...
NeurIPS Anthology Visualization
https://neuripsav.vizhub.ai/blog/
https://neuripsav.vizhub.ai/
Timnit Gebru launches research institute independent from big tech
https://www.washingtonpost.com/techno...
https://www.dair-institute.org/about
https://www.theguardian.com/commentis...
SageMaker Canvas for no-code ML
https://aws.amazon.com/blogs/aws/anno...
Help, Help!
https://macberth.netlify.app/
https://huggingface.co/emanjavacas/Ma...
https://developer.nvidia.com/blog/nvi...
https://opacus.ai/
https://twitter.com/naotokui_en/statu...
https://colab.research.google.com/dri...
https://twitter.com/ThomasSimonini/st...
https://github.com/karpathy/arxiv-san...
https://arxiv-sanity-lite.com/
https://www.youtube.com/watch?v=01ENz...
https://github.com/Felix-Petersen/alg...
https://github.com/rentruewang/koila?...
https://github.com/YeWR/EfficientZero
Cornelius Emde wins the 3090
https://twitter.com/CorEmde/status/14...
A retrospective on the NeurIPS 2021 ethics review process
https://blog.neurips.cc/2021/12/03/a-...
Links:
TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick
YouTube: https://www.youtube.com/c/yannickilcher
Twitter: https://twitter.com/ykilcher
Discord: https://discord.gg/4H8xxDF
BitChute: https://www.bitchute.com/channel/yann...
LinkedIn: https://www.linkedin.com/in/ykilcher
BiliBili: https://space.bilibili.com/2017636191
If you want to support me, the best thing to do is to share out the content :)
If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):
SubscribeStar: https://www.subscribestar.com/yannick...
Patreon: https://www.patreon.com/yannickilcher
Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq
Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2
Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m
Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n
NÜWA is a unifying architecture that can ingest text, images, and videos and brings all of them into a quantized latent representation to support a multitude of visual generation tasks, such as text-to-image, text-guided video manipulation, or sketch-to-video. This paper details how the encoders for the different modalities are constructed, and how the latent representation is transformed using their novel 3D nearby self-attention layers. Experiments are shown on 8 different visual generation tasks that the model supports.
OUTLINE:
0:00 - Intro & Outline
1:20 - Sponsor: ClearML
3:35 - Tasks & Naming
5:10 - The problem with recurrent image generation
7:35 - Creating a shared latent space w/ Vector Quantization
23:20 - Transforming the latent representation
26:25 - Recap: Self- and Cross-Attention
28:50 - 3D Nearby Self-Attention
41:20 - Pre-Training Objective
46:05 - Experimental Results
50:40 - Conclusion & Comments
Paper: https://arxiv.org/abs/2111.12417
Github: https://github.com/microsoft/NUWA
Sponsor: ClearML
https://clear.ml
Abstract:
This paper presents a unified multimodal pre-trained model called NÜWA that can generate new or manipulate existing visual data (i.e., images and videos) for various visual synthesis tasks. To cover language, image, and video at the same time for different scenarios, a 3D transformer encoder-decoder framework is designed, which can not only deal with videos as 3D data but also adapt to texts and images as 1D and 2D data, respectively. A 3D Nearby Attention (3DNA) mechanism is also proposed to consider the nature of the visual data and reduce the computational complexity. We evaluate NÜWA on 8 downstream tasks. Compared to several strong baselines, NÜWA achieves state-of-the-art results on text-to-image generation, text-to-video generation, video prediction, etc. Furthermore, it also shows surprisingly good zero-shot capabilities on text-guided image and video manipulation tasks. Project repo is this https URL.
Authors: Chenfei Wu, Jian Liang, Lei Ji, Fan Yang, Yuejian Fang, Daxin Jiang, Nan Duan
Links:
TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick
YouTube: https://www.youtube.com/c/yannickilcher
Twitter: https://twitter.com/ykilcher
Discord: https://discord.gg/4H8xxDF
BitChute: https://www.bitchute.com/channel/yann...
LinkedIn: https://www.linkedin.com/in/ykilcher
BiliBili: https://space.bilibili.com/2017636191
If you want to support me, the best thing to do is to share out the content :)
If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):
SubscribeStar: https://www.subscribestar.com/yannick...
Patreon: https://www.patreon.com/yannickilcher
Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq
Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2
Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m
Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n
Your weekly dose of ML News!
More GauGAN images here: https://drive.google.com/drive/folder...
OUTLINE:
0:00 - Intro
0:20 - Sponsor: Weights & Biases
2:20 - OpenAI's removes GPT-3 Waitlist
4:55 - NVIDIA releases GauGAN2 Webapp
9:45 - Everyday Robots tackles real-life tasks
12:15 - MetNet-2: 12-hour Rain Forecasting
14:45 - TinyML Dog Bark Stopper
15:55 - AI learns to drive Mario Kart 64 on real hardware
17:40 - NYC regulates bias in AI hiring tools
21:05 - Beverage companies big into AI
21:50 - How does AlphaZero play Chess?
23:35 - Helpful Things
28:00 - ArXiv founder awarded Einstein Foundation Award
References:
OpenAI's removes GPT-3 Waitlist
https://openai.com/blog/api-no-waitlist/
https://beta.openai.com/playground?mo...
NVIDIA releases GauGAN2 Webapp
https://www.reddit.com/r/MachineLearn...
http://gaugan.org/gaugan2/
https://blogs.nvidia.com/blog/2021/11...
https://blogs.nvidia.com/blog/2019/03...
https://arxiv.org/abs/1903.07291
Everyday Robots tackles real-life tasks
https://everydayrobots.com/
https://www.wired.com/story/plaintext...
https://archive.ph/YC4XG#selection-92...
MetNet-2: 12-hour Rain Forecasting
https://ai.googleblog.com/2021/11/met...
TinyML Dog Bark Stopper
https://www.hackster.io/NathanielF/ti...
AI learns to drive Mario Kart 64 on real hardwware
https://www.youtube.com/watch?v=z9E38...
NYC regulates bias in AI hiring tools
https://www.nbcnewyork.com/news/local...
Beverage companies big into AI
https://www.just-drinks.com/features/...
How does AlphaZero play Chess?
https://arxiv.org/pdf/2111.09259.pdf
https://storage.googleapis.com/uncert...
Helpful Things
https://huggingface.co/sberbank-ai/ru...
https://github.com/MathisFederico/Ope...
https://blog.tensorflow.org/2021/11/i...
https://github.com/tensorflow/gnn
https://github.com/jurgisp/pydreamer?...
https://danijar.com/project/dreamerv2/
https://github.com/danijar/dreamerv2
https://deepgenx.com/
https://github.com/DeepGenX/CodeGenX
https://devpost.com/software/heyoh-ca...
https://heyoh-app.github.io/heyoh-pro...
https://github.com/heyoh-app/heyoh-pr...
ArXiv founder awarded Einstein Foundation Award
https://idw-online.de/en/news781515?u...
Links:
TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick
YouTube: https://www.youtube.com/c/yannickilcher
Twitter: https://twitter.com/ykilcher
Discord: https://discord.gg/4H8xxDF
BitChute: https://www.bitchute.com/channel/yann...
LinkedIn: https://www.linkedin.com/in/ykilcher
BiliBili: https://space.bilibili.com/2017636191
If you want to support me, the best thing to do is to share out the content :)
If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):
SubscribeStar: https://www.subscribestar.com/yannick...
Patreon: https://www.patreon.com/yannickilcher
Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq
Transformers keep pushing the state of the art in language and other domains, mainly due to their ability to scale to ever more parameters. However, this scaling has made it prohibitively expensive to run a lot of inference requests against a Transformer, both in terms of compute and memory requirements. Scaling Transformers are a new kind of architecture that leverage sparsity in the Transformer blocks to massively speed up inference, and by including additional ideas from other architectures, they create the Terraformer, which is both fast, accurate, and consumes very little memory.
OUTLINE:
0:00 - Intro & Overview
4:10 - Recap: Transformer stack
6:55 - Sparse Feedforward layer
19:20 - Sparse QKV Layer
43:55 - Terraformer architecture
55:05 - Experimental Results & Conclusion
Paper: https://arxiv.org/abs/2111.12763
Code: https://github.com/google/trax/blob/m...
Abstract:
Large Transformer models yield impressive results on many tasks, but are expensive to train, or even fine-tune, and so slow at decoding that their use and study becomes out of reach. We address this problem by leveraging sparsity. We study sparse variants for all layers in the Transformer and propose Scaling Transformers, a family of next generation Transformer models that use sparse layers to scale efficiently and perform unbatched decoding much faster than the standard Transformer as we scale up the model size. Surprisingly, the sparse layers are enough to obtain the same perplexity as the standard Transformer with the same number of parameters. We also integrate with prior sparsity approaches to attention and enable fast inference on long sequences even with limited memory. This results in performance competitive to the state-of-the-art on long text summarization.
Authors: Sebastian Jaszczur, Aakanksha Chowdhery, Afroz Mohiuddin, Łukasz Kaiser, Wojciech Gajewski, Henryk Michalewski, Jonni Kanerva
Links:
TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick
YouTube: https://www.youtube.com/c/yannickilcher
Twitter: https://twitter.com/ykilcher
Discord: https://discord.gg/4H8xxDF
BitChute: https://www.bitchute.com/channel/yann...
LinkedIn: https://www.linkedin.com/in/ykilcher
BiliBili: https://space.bilibili.com/2017636191
If you want to support me, the best thing to do is to share out the content :)
If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):
SubscribeStar: https://www.subscribestar.com/yannick...
Patreon: https://www.patreon.com/yannickilcher
Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq
Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2
Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m
Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n
The T5 model has been a staple for NLP research for the last years. Both its size and its approach to formulate all NLP tasks as prompt-based language modeling make it a convenient choice to tackle new challenges and provides a strong baseline for most current datasets. ExT5 pushes T5 to its limits by pre-training not only on self-supervised mask filling, but also at the same time on 107 different supervised NLP tasks, which is their new ExMix dataset. The resulting model compares very favorably to T5 when fine-tuned to downstream tasks.
OUTLINE:
0:00 - Intro & Overview
2:15 - Recap: The T5 model
3:55 - The ExT5 model and task formulations
8:10 - ExMix dataset
9:35 - Do different tasks help each other?
16:50 - Which tasks should we include?
20:30 - Pre-Training vs Pre-Finetuning
23:00 - A few hypotheses about what's going on
27:20 - How much self-supervised data to use?
34:15 - More experimental results
38:40 - Conclusion & Summary
Paper: https://arxiv.org/abs/2111.10952
Abstract:
Despite the recent success of multi-task learning and transfer learning for natural language processing (NLP), few works have systematically studied the effect of scaling up the number of tasks during pre-training. Towards this goal, this paper introduces ExMix (Extreme Mixture): a massive collection of 107 supervised NLP tasks across diverse domains and task-families. Using ExMix, we study the effect of multi-task pre-training at the largest scale to date, and analyze co-training transfer amongst common families of tasks. Through this analysis, we show that manually curating an ideal set of tasks for multi-task pre-training is not straightforward, and that multi-task scaling can vastly improve models on its own. Finally, we propose ExT5: a model pre-trained using a multi-task objective of self-supervised span denoising and supervised ExMix. Via extensive experiments, we show that ExT5 outperforms strong T5 baselines on SuperGLUE, GEM, Rainbow, Closed-Book QA tasks, and several tasks outside of ExMix. ExT5 also significantly improves sample efficiency while pre-training.
Authors: Vamsi Aribandi, Yi Tay, Tal Schuster, Jinfeng Rao, Huaixiu Steven Zheng, Sanket Vaibhav Mehta, Honglei Zhuang, Vinh Q. Tran, Dara Bahri, Jianmo Ni, Jai Gupta, Kai Hui, Sebastian Ruder, Donald Metzler
Links:
TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick
YouTube: https://www.youtube.com/c/yannickilcher
Twitter: https://twitter.com/ykilcher
Discord: https://discord.gg/4H8xxDF
BitChute: https://www.bitchute.com/channel/yann...
LinkedIn: https://www.linkedin.com/in/ykilcher
BiliBili: https://space.bilibili.com/2017636191
If you want to support me, the best thing to do is to share out the content :)
If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):
SubscribeStar: https://www.subscribestar.com/yannick...
Patreon: https://www.patreon.com/yannickilcher
Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq
Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2
Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m
Backpropagation is the workhorse of deep learning, but unfortunately, it only works for continuous functions that are amenable to the chain rule of differentiation. Since discrete algorithms have no continuous derivative, deep networks with such algorithms as part of them cannot be effectively trained using backpropagation. This paper presents a method to incorporate a large class of algorithms, formulated as discrete exponential family distributions, into deep networks and derives gradient estimates that can easily be used in end-to-end backpropagation. This enables things like combinatorial optimizers to be part of a network's forward propagation natively.
OUTLINE:
0:00 - Intro & Overview
4:25 - Sponsor: Weights & Biases
6:15 - Problem Setup & Contributions
8:50 - Recap: Straight-Through Estimator
13:25 - Encoding the discrete problem as an inner product
19:45 - From algorithm to distribution
23:15 - Substituting the gradient
26:50 - Defining a target distribution
38:30 - Approximating marginals via perturb-and-MAP
45:10 - Entire algorithm recap
56:45 - Github Page & Example
Paper: https://arxiv.org/abs/2106.01798
Code (TF): https://github.com/nec-research/tf-imle
Code (Torch): https://github.com/uclnlp/torch-imle
Our Discord: https://discord.gg/4H8xxDF
Sponsor: Weights & Biases
https://wandb.com
Abstract:
Combining discrete probability distributions and combinatorial optimization problems with neural network components has numerous applications but poses several challenges. We propose Implicit Maximum Likelihood Estimation (I-MLE), a framework for end-to-end learning of models combining discrete exponential family distributions and differentiable neural components. I-MLE is widely applicable as it only requires the ability to compute the most probable states and does not rely on smooth relaxations. The framework encompasses several approaches such as perturbation-based implicit differentiation and recent methods to differentiate through black-box combinatorial solvers. We introduce a novel class of noise distributions for approximating marginals via perturb-and-MAP. Moreover, we show that I-MLE simplifies to maximum likelihood estimation when used in some recently studied learning settings that involve combinatorial solvers. Experiments on several datasets suggest that I-MLE is competitive with and often outperforms existing approaches which rely on problem-specific relaxations.
Authors: Mathias Niepert, Pasquale Minervini, Luca Franceschi
Links:
TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick
YouTube: https://www.youtube.com/c/yannickilcher
Twitter: https://twitter.com/ykilcher
Discord: https://discord.gg/4H8xxDF
BitChute: https://www.bitchute.com/channel/yann...
LinkedIn: https://www.linkedin.com/in/ykilcher
BiliBili: https://space.bilibili.com/2017636191
If you want to support me, the best thing to do is to share out the content :)
If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):
SubscribeStar: https://www.subscribestar.com/yannick...
Patreon: https://www.patreon.com/yannickilcher
Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq
Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2
Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m
Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n
A look at the results of the 2021 NeurIPS peer review experiment.
https://arxiv.org/abs/2109.09774
https://www.reddit.com/r/MachineLearn...
Links:
TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick
YouTube: https://www.youtube.com/c/yannickilcher
Twitter: https://twitter.com/ykilcher
Discord: https://discord.gg/4H8xxDF
BitChute: https://www.bitchute.com/channel/yann...
LinkedIn: https://www.linkedin.com/in/ykilcher
BiliBili: https://space.bilibili.com/2017636191
If you want to support me, the best thing to do is to share out the content :)
If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):
SubscribeStar: https://www.subscribestar.com/yannick...
Patreon: https://www.patreon.com/yannickilcher
Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq
Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2
Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m
Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n
Deep Neural Networks are usually trained from a given parameter initialization using SGD until convergence at a local optimum. This paper goes a different route: Given a novel network architecture for a known dataset, can we predict the final network parameters without ever training them? The authors build a Graph-Hypernetwork and train on a novel dataset of various DNN-architectures to predict high-performing weights. The results show that not only can the GHN predict weights with non-trivial performance, but it can also generalize beyond the distribution of training architectures to predict weights for networks that are much larger, deeper, or wider than ever seen in training.
OUTLINE:
0:00 - Intro & Overview
6:20 - DeepNets-1M Dataset
13:25 - How to train the Hypernetwork
17:30 - Recap on Graph Neural Networks
23:40 - Message Passing mirrors forward and backward propagation
25:20 - How to deal with different output shapes
28:45 - Differentiable Normalization
30:20 - Virtual Residual Edges
34:40 - Meta-Batching
37:00 - Experimental Results
42:00 - Fine-Tuning experiments
45:25 - Public reception of the paper
ERRATA:
Boris' name is obviously Boris, not Bori
At 36:05, Boris mentions that they train the first variant, yet on closer examination, we decided it's more like the second
Paper: https://arxiv.org/abs/2110.13100
Code: https://github.com/facebookresearch/p...
Abstract:
Deep learning has been successful in automating the design of features in machine learning pipelines. However, the algorithms optimizing neural network parameters remain largely hand-designed and computationally inefficient. We study if we can use deep learning to directly predict these parameters by exploiting the past knowledge of training other networks. We introduce a large-scale dataset of diverse computational graphs of neural architectures - DeepNets-1M - and use it to explore parameter prediction on CIFAR-10 and ImageNet. By leveraging advances in graph neural networks, we propose a hypernetwork that can predict performant parameters in a single forward pass taking a fraction of a second, even on a CPU. The proposed model achieves surprisingly good performance on unseen and diverse networks. For example, it is able to predict all 24 million parameters of a ResNet-50 achieving a 60% accuracy on CIFAR-10. On ImageNet, top-5 accuracy of some of our networks approaches 50%. Our task along with the model and results can potentially lead to a new, more computationally efficient paradigm of training networks. Our model also learns a strong representation of neural architectures enabling their analysis.
Authors: Boris Knyazev, Michal Drozdzal, Graham W. Taylor, Adriana Romero-Soriano
Links:
TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick
YouTube: https://www.youtube.com/c/yannickilcher
Twitter: https://twitter.com/ykilcher
Discord: https://discord.gg/4H8xxDF
BitChute: https://www.bitchute.com/channel/yann...
LinkedIn: https://www.linkedin.com/in/ykilcher
BiliBili: https://space.bilibili.com/2017636191
If you want to support me, the best thing to do is to share out the content :)
If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):
SubscribeStar: https://www.subscribestar.com/yannick...
Patreon: https://www.patreon.com/yannickilcher
Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq
Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2
Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m
The last years in deep learning research have given rise to a plethora of different optimization algorithms, such as SGD, AdaGrad, Adam, LARS, LAMB, etc. which all claim to have their special peculiarities and advantages. In general, all algorithms modify two major things: The (implicit) learning rate schedule, and a correction to the gradient direction. This paper introduces grafting, which allows to transfer the induced learning rate schedule of one optimizer to another one. In that, the paper shows that much of the benefits of adaptive methods (e.g. Adam) are actually due to this schedule, and not necessarily to the gradient direction correction. Grafting allows for more fundamental research into differences and commonalities between optimizers, and a derived version of it makes it possible to computes static learning rate corrections for SGD, which potentially allows for large savings of GPU memory.
OUTLINE
0:00 - Rant about Reviewer #2
6:25 - Intro & Overview
12:25 - Adaptive Optimization Methods
20:15 - Grafting Algorithm
26:45 - Experimental Results
31:35 - Static Transfer of Learning Rate Ratios
35:25 - Conclusion & Discussion
Paper (OpenReview): https://openreview.net/forum?id=FpKgG...
Old Paper (Arxiv): https://arxiv.org/abs/2002.11803
Our Discord: https://discord.gg/4H8xxDF
Abstract:
In the empirical science of training large neural networks, the learning rate schedule is a notoriously challenging-to-tune hyperparameter, which can depend on all other properties (architecture, optimizer, batch size, dataset, regularization, ...) of the problem. In this work, we probe the entanglements between the optimizer and the learning rate schedule. We propose the technique of optimizer grafting, which allows for the transfer of the overall implicit step size schedule from a tuned optimizer to a new optimizer, preserving empirical performance. This provides a robust plug-and-play baseline for optimizer comparisons, leading to reductions to the computational cost of optimizer hyperparameter search. Using grafting, we discover a non-adaptive learning rate correction to SGD which allows it to train a BERT model to state-of-the-art performance. Besides providing a resource-saving tool for practitioners, the invariances discovered via grafting shed light on the successes and failure modes of optimizers in deep learning.
Authors: Anonymous (Under Review)
Links:
TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick
YouTube: https://www.youtube.com/c/yannickilcher
Twitter: https://twitter.com/ykilcher
Discord: https://discord.gg/4H8xxDF
BitChute: https://www.bitchute.com/channel/yann...
LinkedIn: https://www.linkedin.com/in/ykilcher
BiliBili: https://space.bilibili.com/2017636191
If you want to support me, the best thing to do is to share out the content :)
If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):
SubscribeStar: https://www.subscribestar.com/yannick...
Patreon: https://www.patreon.com/yannickilcher
Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq
Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2
Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m
Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n
Only the greatest of news from the world of Machine Learning.
OUTLINE:
0:00 - Sponsor: Weights & Biases
1:50 - Cedille - French Language Model
3:55 - Facebook AI Multilingual model wins WMT
5:50 - YOU private search engine
10:35 - DeepMind's Open-Source Arnheim
12:10 - Company sued for using AI to make website more accessible
18:05 - Alibaba DAMO Academy creates 10 Trillion M6 model
21:15 - AMD MI200 Family
22:30 - State of AI report 2021
24:15 - Andrew Ng's Landing AI raises 57M
25:40 - Cerebras raises 250M
26:45 - Microsoft's Varuna: Scalable Training of Huge Models
28:15 - Laura Ruis reproduces Extrapolation Paper
29:05 - Ian Charnas' Real-Life Punchout
30:00 - Helpful Things
33:10 - AI finds profitable Meme-Tokens
34:55 - This Sneaker Does Not Exist
Sponsor: Weights & Biases
https://wandb.com
References:
Cedille - French Language Model
https://en.cedille.ai/
https://github.com/coteries/cedille-ai
https://app.cedille.ai/
https://en.wikipedia.org/wiki/Cedilla
Facebook AI Multilingual model wins WMT
https://ai.facebook.com/blog/the-firs...
YOU private search engine
https://you.com/
https://youdotcom.notion.site/FAQ-8c8...
DeepMind's Open-Source Arnheim
https://deepmind.com/research/open-so...
https://twitter.com/OriolVinyalsML/st...
https://github.com/deepmind/arnheim
https://colab.research.google.com/git...
Company sued for using AI to make website more accessible
https://www.wired.com/story/company-t...
https://archive.ph/kdvOM
Alibaba DAMO Academy creates 10 Trillion M6 model
https://pandaily.com/alibaba-damo-aca...
https://www.infoq.cn/article/xIX9leku...
AMD MI200 Family
https://www.anandtech.com/show/17054/...
State of AI report 2021
https://www.stateof.ai/?utm_source=po...
Andrew Ng's Landing AI raises 57M
https://techcrunch.com/2021/11/08/lan...
https://www.forbes.com/sites/bernardm...
https://landing.ai/platform/
Cerebras raises 250M
https://cerebras.net/news/cerebras-sy...
https://cerebras.net/news/cerebras-sy...
Microsoft's Varuna: Scalable Training of Huge Models
https://syncedreview.com/2021/11/10/d...
Laura Ruis reproduces Extrapolation Paper
https://lauraruis.github.io/2021/11/0...
https://github.com/LauraRuis
Ian Charnas' Real-Life Punchout
https://www.reddit.com/r/MachineLearn...
https://www.youtube.com/watch?v=07Jib...
Helpful Things
https://www.marktechpost.com/2021/11/...
https://pair-code.github.io/lit/demos/
https://github.com/pair-code/lit
https://www.reddit.com/r/MachineLearn...
https://twitter.com/yeemachine/status...
https://github.com/yeemachine/kalidokit
AI finds profitable Meme-Tokens
https://finance.yahoo.com/news/artifi...
https://finu.co/
This Sneaker Does Not Exist
https://thissneakerdoesnotexist.com/
Links:
TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick
YouTube: https://www.youtube.com/c/yannickilcher
Twitter: https://twitter.com/ykilcher
Discord: https://discord.gg/4H8xxDF
BitChute: https://www.bitchute.com/channel/yann...
LinkedIn: https://www.linkedin.com/in/ykilcher
BiliBili: https://space.bilibili.com/2017636191
If you want to support me, the best thing to do is to share out the content :)
More and more systems are made differentiable, which means that accurate gradients of these systems' dynamics can be computed exactly. While this development has led to a lot of advances, there are also distinct situations where backpropagation can be a very bad idea. This paper characterizes a few such systems in the domain of iterated dynamical systems, often including some source of stochasticity, resulting in chaotic behavior. In these systems, it is often better to use black-box estimators for gradients than computing them exactly.
OUTLINE:
0:00 - Foreword
1:15 - Intro & Overview
3:40 - Backpropagation through iterated systems
12:10 - Connection to the spectrum of the Jacobian
15:35 - The Reparameterization Trick
21:30 - Problems of reparameterization
26:35 - Example 1: Policy Learning in Simulation
33:05 - Example 2: Meta-Learning Optimizers
36:15 - Example 3: Disk packing
37:45 - Analysis of Jacobians
40:20 - What can be done?
45:40 - Just use Black-Box methods
Paper: https://arxiv.org/abs/2111.05803
Abstract:
Differentiable programming techniques are widely used in the community and are responsible for the machine learning renaissance of the past several decades. While these methods are powerful, they have limits. In this short report, we discuss a common chaos based failure mode which appears in a variety of differentiable circumstances, ranging from recurrent neural networks and numerical physics simulation to training learned optimizers. We trace this failure to the spectrum of the Jacobian of the system under study, and provide criteria for when a practitioner might expect this failure to spoil their differentiation based optimization algorithms.
Authors: Luke Metz, C. Daniel Freeman, Samuel S. Schoenholz, Tal Kachman
Links:
TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick
YouTube: https://www.youtube.com/c/yannickilcher
Twitter: https://twitter.com/ykilcher
Discord: https://discord.gg/4H8xxDF
BitChute: https://www.bitchute.com/channel/yann...
LinkedIn: https://www.linkedin.com/in/ykilcher
BiliBili: https://space.bilibili.com/2017636191
If you want to support me, the best thing to do is to share out the content :)
If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):
SubscribeStar: https://www.subscribestar.com/yannick...
Patreon: https://www.patreon.com/yannickilcher
Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq
Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2
Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m
Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n
The latest and greatest from the Machine Learning world
Sponsor: Weights & Biases
https://wandb.com
References:
Microsoft Turing Bletchley: Universal Image Language Representation Model
https://www.microsoft.com/en-us/resea...
https://turing.microsoft.com/bletchley
Meta AI Tactile Sensing
https://ai.facebook.com/blog/teaching...
https://ai.facebook.com/blog/reskin-a...
https://twitter.com/AIatMeta/status/1...
AnimeGANv2
https://huggingface.co/spaces/akhaliq...
https://github.com/bryandlee/animegan...
https://github.com/TachibanaYoshino/A...
https://tachibanayoshino.github.io/An...
General In-Hand Object Re-Orientation
https://taochenshh.github.io/projects...
https://arxiv.org/abs/2111.03043
Does Facebook score the "Anger" Emoji too high?
https://www.washingtonpost.com/techno...
IsomorphicLabs: New Alphabet Company for Drug Discovery
https://twitter.com/demishassabis/sta...
https://www.isomorphiclabs.com/blog
ruDALL-E: Russian DALL-E
https://github.com/sberbank-ai/ru-dalle
https://huggingface.co/spaces/anton-l...
https://colab.research.google.com/git...
https://huggingface.co/sberbank-ai/ru...
https://rudalle.ru/
https://habr.com/ru/company/sberbank/...
https://habr-com.translate.goog/ru/co...
Image Scaling Attacks
https://twitter.com/AlexTamkin/status...
https://twitter.com/rzhang88/status/1...
https://arxiv.org/abs/2104.11222
https://twitter.com/arxiv_org/status/...
https://bifold.berlin/preventing-imag...
https://embracethered.com/blog/posts/...
Azure OpenAI Service
https://blogs.microsoft.com/ai/new-az...
https://azure.microsoft.com/en-us/ser...
Neural MMO
https://openai.com/blog/neural-mmo/?u...
https://github.com/jsuarez5341/neural...
https://github.com/jsuarez5341/neural...
https://jsuarez5341.github.io/neural-...
https://jsuarez5341.github.io/neural-...
https://arxiv.org/abs/2110.07594
ArxivDOOM
https://sniklaus.com/arxivdoom?utm_so...
ARC Game
https://github.com/volotat/ARC-Game
https://volotat.github.io/ARC-Game/?
ResNeXtGuesser
https://twitter.com/resnextguesser/st...
Zillow loses money based on AI home price estimation
https://www.reddit.com/r/MachineLearn...
https://www.cbsnews.com/news/zillow-l...
https://www.businessinsider.com/zillo...
https://archive.ph/qEITQ
Helpful Things
https://github.com/PyTorchLightning/p...
https://www.reddit.com/r/MachineLearn...
https://devpost.com/software/iris-7s3yna
https://github.com/prabhuomkar/iris
https://araffin.github.io/post/rliable/
https://github.com/google-research/rl...
https://paperswithcode.com/dataset/me...
AI will make your company great! Promise, Human!
https://fortune.com/2021/11/05/ai-art...
https://sloanreview.mit.edu/projects/...
Diffusion models have made large advances in recent months as a new type of generative models. This paper introduces Autoregressive Diffusion Models (ARDMs), which are a mix between autoregressive generative models and diffusion models. ARDMs are trained to be agnostic to the order of autoregressive decoding and give the user a dynamic tradeoff between speed and performance at decoding time. This paper applies ARDMs to both text and image data, and as an extension, the models can also be used to perform lossless compression.
OUTLINE:
0:00 - Intro & Overview
3:15 - Decoding Order in Autoregressive Models
6:15 - Autoregressive Diffusion Models
8:35 - Dependent and Independent Sampling
14:25 - Application to Character-Level Language Models
18:15 - How Sampling & Training Works
26:05 - Extension 1: Parallel Sampling
29:20 - Extension 2: Depth Upscaling
33:10 - Conclusion & Comments
Paper: https://arxiv.org/abs/2110.02037
Abstract:
We introduce Autoregressive Diffusion Models (ARDMs), a model class encompassing and generalizing order-agnostic autoregressive models (Uria et al., 2014) and absorbing discrete diffusion (Austin et al., 2021), which we show are special cases of ARDMs under mild assumptions. ARDMs are simple to implement and easy to train. Unlike standard ARMs, they do not require causal masking of model representations, and can be trained using an efficient objective similar to modern probabilistic diffusion models that scales favourably to highly-dimensional data. At test time, ARDMs support parallel generation which can be adapted to fit any given generation budget. We find that ARDMs require significantly fewer steps than discrete diffusion models to attain the same performance. Finally, we apply ARDMs to lossless compression, and show that they are uniquely suited to this task. Contrary to existing approaches based on bits-back coding, ARDMs obtain compelling results not only on complete datasets, but also on compressing single data points. Moreover, this can be done using a modest number of network calls for (de)compression due to the model's adaptable parallel generation.
Authors: Emiel Hoogeboom, Alexey A. Gritsenko, Jasmijn Bastings, Ben Poole, Rianne van den Berg, Tim Salimans
Links:
TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick
YouTube: https://www.youtube.com/c/yannickilcher
Twitter: https://twitter.com/ykilcher
Discord: https://discord.gg/4H8xxDF
BitChute: https://www.bitchute.com/channel/yann...
Minds: https://www.minds.com/ykilcher
Parler: https://parler.com/profile/YannicKilcher
LinkedIn: https://www.linkedin.com/in/ykilcher
BiliBili: https://space.bilibili.com/1824646584
If you want to support me, the best thing to do is to share out the content :)
If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):
SubscribeStar: https://www.subscribestar.com/yannick...
Patreon: https://www.patreon.com/yannickilcher
Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq
Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2
Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m
Your irregular dose of Machine Learning News.
OUTLINE:
0:00 - Intro
0:20 - Sponsor: Weights & Biases
2:10 - Google Introduces Pathways AI Architecture
6:30 - OpenAI trains Language Models to do High School Math
8:25 - Sam Altman says Neural Networks truly learn
9:35 - Google AI researchers frustrated with lawyers
12:10 - DeepMind RL Lecture Series 2021
12:40 - Fashion Store sells Adversarial Patches
13:15 - A viable method to remove the GIL from CPython
15:05 - BigScience Workshop releases T0
17:40 - Huggingface Hub Dataset Viewer
18:10 - Scite classifies scientific citations
19:25 - Facebook AI Ego4D dataset & challenges
21:50 - Tesla Dojo Configurable Floating Point Spec
23:10 - Windows releases PyTorch-DirectML for Deep Learning on DirectX GPUs
23:50 - Helpful Things
33:00 - Traders use ML to analyze CEOs' language
34:20 - Cadbury creates DeepFake ads for local Indian businesses
35:25 - This Shoe Does Not Exist
Sponsor: Weights & Biases
https://wandb.com
References:
Google Introduces Pathways AI Architecture
https://blog.google/technology/ai/int...
OpenAI trains Language Models to do High School Math
https://openai.com/blog/grade-school-...
https://arxiv.org/abs/2110.14168
Sam Altman says Neural Networks truly learn
https://twitter.com/sama/status/14508...
Google AI researchers frustrated with lawyers
https://archive.ph/lsQJJ#selection-28...
DeepMind RL Lecture Series 2021
https://deepmind.com/learning-resourc...
Fashion Store sells Adversarial Patches
https://twitter.com/naotokui/status/1...
A viable method to remove the GIL from CPython
https://lwn.net/Articles/872869/
BigScience Workshop releases T0
https://bigscience.huggingface.co/
https://arxiv.org/abs/2110.08207
https://huggingface.co/bigscience/T0pp
Huggingface Hub Dataset Viewer
https://twitter.com/huggingface/statu...
Scite classifies scientific citations
https://scite.ai
https://direct.mit.edu/qss/article/do...
Facebook AI Ego4D dataset & challenges
https://ai.facebook.com/blog/teaching...
Tesla Dojo Configurable Floating Point Spec
https://tesla-cdn.thron.com/static/SB...
Windows releases PyTorch-DirectML for Deep Learning on DirectX GPUs
https://devblogs.microsoft.com/window...
Helpful Things
https://github.com/achaiah/pywick?utm...
https://github.com/orybkin/lexa-bench...
https://orybkin.github.io/lexa/
https://twitter.com/danijarh/status/1...
https://github.com/RobertTLange/mle-h...
https://keras.io/examples/vision/mobi...
https://twitter.com/osanseviero/statu...
https://huggingface.co/spaces/flax-co...
https://huggingface.co/transformers/m...
https://github.com/facebookresearch/b...
https://arxiv.org/abs/2110.11216
https://arxiv.org/pdf/2110.11216.pdf
https://github.com/facebookresearch/x...
https://superbbenchmark.org/
https://arxiv.org/abs/2110.07731
https://github.com/BaguaSys/bagua?utm...
https://github.com/cgarciae/treex
https://jax.readthedocs.io/en/latest/...
Traders use ML to analyze CEOs' language
https://www.reuters.com/technology/ai...
Cadbury creates DeepFake ads for local Indian businesses
https://www.bgr.in/entertainment/shah...
This Shoe Does Not Exist
https://www.thisshoedoesnotexist.com/
Reinforcement Learning methods are notoriously data-hungry. Notably, MuZero learns a latent world model just from scalar feedback of reward- and policy-predictions, and therefore relies on scale to perform well. However, most RL algorithms fail when presented with very little data. EfficientZero makes several improvements over MuZero that allows it to learn from astonishingly small amounts of data and outperform other methods by a large margin in the low-sample setting. This could be a staple algorithm for future RL research.
OUTLINE:
0:00 - Intro & Outline
2:30 - MuZero Recap
10:50 - EfficientZero improvements
14:15 - Self-Supervised consistency loss
17:50 - End-to-end prediction of the value prefix
20:40 - Model-based off-policy correction
25:45 - Experimental Results & Conclusion
Paper: https://arxiv.org/abs/2111.00210
Code: https://github.com/YeWR/EfficientZero
Note: code not there yet as of release of this video
Abstract:
Reinforcement learning has achieved great success in many applications. However, sample efficiency remains a key challenge, with prominent methods requiring millions (or even billions) of environment steps to train. Recently, there has been significant progress in sample efficient image-based RL algorithms; however, consistent human-level performance on the Atari game benchmark remains an elusive goal. We propose a sample efficient model-based visual RL algorithm built on MuZero, which we name EfficientZero. Our method achieves 190.4% mean human performance and 116.0% median performance on the Atari 100k benchmark with only two hours of real-time game experience and outperforms the state SAC in some tasks on the DMControl 100k benchmark. This is the first time an algorithm achieves super-human performance on Atari games with such little data. EfficientZero's performance is also close to DQN's performance at 200 million frames while we consume 500 times less data. EfficientZero's low sample complexity and high performance can bring RL closer to real-world applicability. We implement our algorithm in an easy-to-understand manner and it is available at this https URL. We hope it will accelerate the research of MCTS-based RL algorithms in the wider community.
Authors: Weirui Ye, Shaohuai Liu, Thanard Kurutach, Pieter Abbeel, Yang Gao
Links:
TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick
YouTube: https://www.youtube.com/c/yannickilcher
Twitter: https://twitter.com/ykilcher
Discord: https://discord.gg/4H8xxDF
BitChute: https://www.bitchute.com/channel/yann...
Minds: https://www.minds.com/ykilcher
Parler: https://parler.com/profile/YannicKilcher
LinkedIn: https://www.linkedin.com/in/ykilcher
BiliBili: https://space.bilibili.com/1824646584
If you want to support me, the best thing to do is to share out the content :)
If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):
SubscribeStar: https://www.subscribestar.com/yannick...
Patreon: https://www.patreon.com/yannickilcher
Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq
Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2
Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m
Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n
A conversation with Siraj Raval about his journey on YouTube, and the perils of fame.
OUTLINE:
0:00 - Intro
1:30 - Welcome
3:15 - Starting out: From Economics to YouTube
13:00 - More Views: Plagiarizing Video Content
23:30 - One Step Up: Copying A Research Paper
29:15 - Was there another way?
39:00 - Clickbait Course: Make Money with Machine Learning
50:30 - Rock Bottom and the Way Forward
1:01:30 - Advice for Future Generations
Siraj's Channel: https://www.youtube.com/c/SirajRaval
Links:
TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick
YouTube: https://www.youtube.com/c/yannickilcher
Twitter: https://twitter.com/ykilcher
Discord: https://discord.gg/4H8xxDF
BitChute: https://www.bitchute.com/channel/yann...
Minds: https://www.minds.com/ykilcher
Parler: https://parler.com/profile/YannicKilcher
LinkedIn: https://www.linkedin.com/in/ykilcher
BiliBili: https://space.bilibili.com/1824646584
If you want to support me, the best thing to do is to share out the content :)
If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):
SubscribeStar: https://www.subscribestar.com/yannick...
Patreon: https://www.patreon.com/yannickilcher
Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq
Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2
Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m
Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n
Registriere für GTC'21 und gewinne eine RTX 3090: https://nvda.ws/2Y2B5ni
OUTLINE:
0:00 - Intro
0:15 - Sponsor: NVIDIA GTC'21
6:10 - DeepMind kauft & Open-Sourct MuJoCo
9:05 - PyTorch 1.10 Veröffentlicht
11:25 - Google Lernt Spreadsheet Formeln
14:15 - handtracking.io
15:25 - Zellinstanzsegmentierungswettbewerb
16:15 - Hilfreiche Bibliotheken
23:15 - Waymo autos verirren sich alle in der selben Sackgasse
24:50 - BlueRiver balanciert Traktoren
References:
DeepMind kauft & open-sourct MuJoCo
https://deepmind.com/blog/announcemen...
PyTorch 1.10 veröffentlicht
https://pytorch.org/blog/pytorch-1.10...
https://developer.nvidia.com/blog/cud...
GoogleAI sagt Tabellen-Formeln voraus
https://ai.googleblog.com/2021/10/pre...
Handtracking im Browser
https://handtracking.io/
https://handtracking.io/draw_demo/
Sartorius Zellinstanzsegmentierungswettbewerb
https://www.kaggle.com/c/sartorius-ce...
Hilfreiche Bibliotheken
https://github.com/IntelLabs/control-...
https://github.com/facebookresearch/s...
https://github.com/facebookresearch/s...
https://github.com/ydataai/ydata-synt...
https://syntheticdata.community/
https://github.com/ydataai/ydata-synt...
https://medium.com/aimstack/aim-3-0-0...
https://github.com/aimhubio/aim
https://robustbench.github.io/
Waymo Autos verirren sich in dieselbe Sackgasse wieder und wieder
https://sanfrancisco.cbslocal.com/202...
BlueRiver balanciert Traktoren
https://www.linkedin.com/posts/lredde...
https://bluerivertechnology.com/ourme...
Links:
TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick
YouTube: https://www.youtube.com/c/yannickilcher
Twitter: https://twitter.com/ykilcher
Discord: https://discord.gg/4H8xxDF
BitChute: https://www.bitchute.com/channel/yann...
Minds: https://www.minds.com/ykilcher
Parler: https://parler.com/profile/YannicKilcher
LinkedIn: https://www.linkedin.com/in/ykilcher
BiliBili: https://space.bilibili.com/1824646584
If you want to support me, the best thing to do is to share out the content :)
If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):
SubscribeStar: https://www.subscribestar.com/yannick...
Patreon: https://www.patreon.com/yannickilcher
Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq
Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2
Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m
Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n
Register to GTC'21 and Win a RTX 3090: https://nvda.ws/2Y2B5ni
OUTLINE:
0:00 - Intro
0:15 - Sponsor: NVIDIA GTC'21
5:35 - DeepMind buys & Open-Sources MuJoCo
7:25 - PyTorch 1.10 Released
9:10 - Google Predicts Spreadsheet Formulas
11:25 - handtracking.io
12:25 - Cell Instance Segmentation Challenge
13:00 - Helpful Libraries
17:50 - Waymo cars keep turning into same dead-end
19:35 - BlueRiver balances tractors
References:
DeepMind buys & open-sources MuJoCo
https://deepmind.com/blog/announcemen...
PyTorch 1.10 released
https://pytorch.org/blog/pytorch-1.10...
https://developer.nvidia.com/blog/cud...
GoogleAI predicts spreadsheet formulas
https://ai.googleblog.com/2021/10/pre...
Handtracking in Browser
https://handtracking.io/
https://handtracking.io/draw_demo/
Sartorius Cell Instance Segmentation Competition
https://www.kaggle.com/c/sartorius-ce...
Helpful Libraries
https://github.com/IntelLabs/control-...
https://github.com/facebookresearch/s...
https://github.com/facebookresearch/s...
https://github.com/ydataai/ydata-synt...
https://syntheticdata.community/
https://github.com/ydataai/ydata-synt...
https://medium.com/aimstack/aim-3-0-0...
https://github.com/aimhubio/aim
https://robustbench.github.io/
Waymo cars keep coming to same dead-end over and over
https://sanfrancisco.cbslocal.com/202...
BlueRiver balances tractors
https://www.linkedin.com/posts/lredde...
https://bluerivertechnology.com/ourme...
Links:
TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick
YouTube: https://www.youtube.com/c/yannickilcher
Twitter: https://twitter.com/ykilcher
Discord: https://discord.gg/4H8xxDF
BitChute: https://www.bitchute.com/channel/yann...
Minds: https://www.minds.com/ykilcher
Parler: https://parler.com/profile/YannicKilcher
LinkedIn: https://www.linkedin.com/in/ykilcher
BiliBili: https://space.bilibili.com/1824646584
If you want to support me, the best thing to do is to share out the content :)
If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):
SubscribeStar: https://www.subscribestar.com/yannick...
Patreon: https://www.patreon.com/yannickilcher
Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq
Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2
Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m
Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n
A trip report from the AiiA Festival in Geneva organized by the ImpactAI foundation.
OUTLINE:
0:00 - Intro
1:50 - Laura Tocmacov: The Festival
4:10 - Timothy O'Hear: The Tech
6:50 - Jonathan O'Hear: The Robot
11:50 - Cléa Chopard: The Artist
17:45 - Final Words
Website: https://aiiafestival.org/en/
Links:
TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick
YouTube: https://www.youtube.com/c/yannickilcher
Twitter: https://twitter.com/ykilcher
Discord: https://discord.gg/4H8xxDF
BitChute: https://www.bitchute.com/channel/yann...
Minds: https://www.minds.com/ykilcher
Parler: https://parler.com/profile/YannicKilcher
LinkedIn: https://www.linkedin.com/in/ykilcher
BiliBili: https://space.bilibili.com/1824646584
If you want to support me, the best thing to do is to share out the content :)
If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):
SubscribeStar: https://www.subscribestar.com/yannick...
Patreon: https://www.patreon.com/yannickilcher
Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq
Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2
Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m
Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n
Symbolic knowledge models are usually trained on human-generated corpora that are cumbersome and expensive to create. Such corpora consist of structured triples of symbolic knowledge. This paper takes a different approach and attempts to generate such a corpus by prompting GPT-3. Results show that clever prompting, combined with targeted small critic models trained on human ratings can outperform both human-generated data, as well as the teacher model (GPT-3) itself. The results of this paper give a general recipe for automatically building corpora for various NLP tasks by extracting samples from large language models.
OUTLINE:
0:00 - Intro & Overview
2:30 - Sponsor: Weights & Biases
4:15 - Commonsense Knowledge Graphs
7:50 - ATOMIC dataset
10:00 - Generating the corpus from a model
13:00 - Prompting GPT-3
15:30 - Generating Events
18:40 - Generating Inferences
23:00 - Evaluating the created dataset
26:45 - Introducing the critic
31:25 - Using the critic to filter the data
36:30 - Training a student on the generated data
41:00 - Key Findings
44:45 - Comments & Conclusion
Paper: https://arxiv.org/abs/2110.07178
Code & Corpus: https://github.com/peterwestai2/symbo...
Sponsor: Weights & Biases
https://wandb.com
https://community.wandb.ai/
Abstract:
The common practice for training commonsense models has gone from-human-to-corpus-to-machine: humans author commonsense knowledge graphs in order to train commonsense models. In this work, we investigate an alternative, from-machine-to-corpus-to-machine: general language models author these commonsense knowledge graphs to train commonsense models. Our study leads to a new framework, Symbolic Knowledge Distillation. As with prior art in Knowledge Distillation (Hinton et al., 2015), our approach uses larger models to teach smaller models. A key difference is that we distill knowledge symbolically-as text-in addition to the neural model. We also distill only one aspect-the commonsense of a general language model teacher, allowing the student to be a different type, a commonsense model. Altogether, we show that careful prompt engineering and a separately trained critic model allow us to selectively distill high-quality causal commonsense from GPT-3, a general language model. Empirical results demonstrate that, for the first time, a human-authored commonsense knowledge graph is surpassed by our automatically distilled variant in all three criteria: quantity, quality, and diversity. In addition, it results in a neural commonsense model that surpasses the teacher model's commonsense capabilities despite its 100x smaller size. We apply this to the ATOMIC resource, and share our new symbolic knowledge graph and commonsense models.
Authors: Peter West, Chandra Bhagavatula, Jack Hessel, Jena D. Hwang, Liwei Jiang, Ronan Le Bras, Ximing Lu, Sean Welleck, Yejin Choi
Links:
TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick
YouTube: https://www.youtube.com/c/yannickilcher
Twitter: https://twitter.com/ykilcher
Discord: https://discord.gg/4H8xxDF
BitChute: https://www.bitchute.com/channel/yann...
Minds: https://www.minds.com/ykilcher
Parler: https://parler.com/profile/YannicKilcher
LinkedIn: https://www.linkedin.com/in/ykilcher
BiliBili: https://space.bilibili.com/1824646584
If you want to support me, the best thing to do is to share out the content :)
If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):
SubscribeStar: https://www.subscribestar.com/yannick...
A friendly parody of Travel Vloggers and Airplane Seat Reviews :)
No, SBB did not pay me for this (but they should ;) )
Links:
TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick
YouTube: https://www.youtube.com/c/yannickilcher
Twitter: https://twitter.com/ykilcher
Discord: https://discord.gg/4H8xxDF
BitChute: https://www.bitchute.com/channel/yann...
Minds: https://www.minds.com/ykilcher
Parler: https://parler.com/profile/YannicKilcher
LinkedIn: https://www.linkedin.com/in/ykilcher
BiliBili: https://space.bilibili.com/1824646584
If you want to support me, the best thing to do is to share out the content :)
If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):
SubscribeStar: https://www.subscribestar.com/yannick...
Patreon: https://www.patreon.com/yannickilcher
Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq
Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2
Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m
Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n
Your latest upates on what's happening in the Machine Learning world.
OUTLINE:
0:00 - Intro
0:16 - Weights & Biases raises on 1B valuation (sponsored)
2:30 - Microsoft trains 530 billion parameter model
5:15 - StyleGAN v3 released
6:45 - A few more examples may be worth billions of parameters
8:30 - ConvMixer fits into a tweet
9:45 - Improved VQGAN
11:25 - William Shatner AI chats about his life
12:35 - Google AI pushes material science
14:10 - Gretel AI raises 50M for privacy protection
16:05 - DeepMind's push into ML for biology
19:00 - Schmidhuber laudates Kunihiko Fukushima for Bower Award
21:30 - Helpful Things
22:25 - Mosaic ML out of stealth mode
23:55 - First German self-driving train
24:45 - Ex-Pentagon Chief: China has already won
26:25 - DeepMind becomes profitable
Sponsor: Weights & Biases
https://wandb.com
References:
Microsoft Trains 530B Parameter Model
https://www.microsoft.com/en-us/resea...
StyleGAN 3 Code Released
https://nvlabs.github.io/stylegan3/
https://github.com/NVlabs/stylegan3
https://colab.research.google.com/git...
When do labels help?
https://arxiv.org/pdf/2110.04374.pdf
ml_paper.bruh
https://openreview.net/pdf?id=TVHS5Y4...
Improved VQGAN
https://openreview.net/pdf?id=pfNyExj7z2
William Shatner "AI" & Storyfile
https://www.livescience.com/william-s...
https://www.storyfile.com/
GoogleAI Finds Complex Metal Oxides
https://ai.googleblog.com/2021/10/fin...
GretelAI raises 50M Series B
https://techcrunch.com/2021/10/07/gre...
https://gretel.ai/
https://gretel.ai/blog/why-privacy-by...
DeepMind's Push in ML for Bio
https://www.biorxiv.org/content/10.11...
https://deepmind.com/blog/article/enf...
Kunihiko Fukushima wins Bower Award: Schmidhuber Congratulates
https://www.fi.edu/laureates/kunihiko...
https://www.youtube.com/watch?v=ysOw6...
Helpful Things
https://github.com/UKPLab/beir#beers-...
https://arxiv.org/pdf/2104.08663.pdf
https://bayesoptbook.com/
https://github.com/nvlabs/imaginaire/
https://github.com/NVlabs/imaginaire/...
MosaicML out of Stealth Mode
https://www.mosaicml.com/
https://www.mosaicml.com/blog/founder...
https://app.mosaicml.com/library/imag...
https://github.com/mosaicml/composer
https://mosaicml-composer.readthedocs...
Germany's first self-driving train
https://techxplore.com/news/2021-10-g...
Ex-Pentagon Chief: China has already won tech war
https://nypost.com/2021/10/11/pentago...
DeepMind becomes profitable
https://bdtechtalks.com/2021/10/07/go...
Links:
TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick
YouTube: https://www.youtube.com/c/yannickilcher
Twitter: https://twitter.com/ykilcher
Discord: https://discord.gg/4H8xxDF
BitChute: https://www.bitchute.com/channel/yann...
Minds: https://www.minds.com/ykilcher
Parler: https://parler.com/profile/YannicKilcher
LinkedIn: https://www.linkedin.com/in/ykilcher
BiliBili: https://space.bilibili.com/1824646584
If you want to support me, the best thing to do is to share out the content :)
If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):
SubscribeStar: https://www.subscribestar.com/yannick...
Your holy update on what's new in the Machine Learning world.
OUTLINE:
0:00 - Intro
0:30 - DeepMind tackles Nowcasting
3:30 - The Guardian's shady reporting on TruthfulQA
6:15 - Stochastic training not necessary for generalization
7:35 - Google AI's efficient partitioning of road networks
9:15 - MiniHack Reinforcement Learning Environment
10:45 - Plato XL 11B dialog model
11:35 - AI finishes Beethoven's 10th Symphony
13:10 - AI casts doubt on painting authenticity
15:55 - ShadowDragon social media surveillance
18:45 - Helpful Libraries
25:20 - Samsung to copy-paste brains onto chips
References:
DeepMind improves Nowcasting
https://deepmind.com/blog/article/now...
https://www.nature.com/articles/s4158...
https://github.com/deepmind/deepmind-...
https://colab.research.google.com/git...
The Guardian's shady reporting on TruthfulQA
https://www.theguardian.com/commentis...
Stochastic Training is Not Necessary for Generalization
https://arxiv.org/pdf/2109.14119.pdf
Google AI - Efficient Partitioning of Road Networks
https://ai.googleblog.com/2021/09/eff...
MiniHack Reinforcement Learning Environment
https://ai.facebook.com/blog/minihack...
Baidu PLATO-XL 11B Dialog Model
http://research.baidu.com/Blog/index-...
AI finishes Beethoven's 10th Symphony
https://thenextweb.com/news/computer-...
AI casts doubt on paining authenticity
https://www.smithsonianmag.com/smart-...
https://art-recognition.com/
https://art-recognition.com/case-stud...
https://art-recognition.com/faq/
ShadowDragon Social Media Surveillance
https://www.rt.com/usa/535630-ai-surv...
https://theintercept.com/2021/09/21/s...
Helpful Libraries / Datasets
https://huggingface.co/infinity
https://yanaiela.github.io/TNE/?s=09&...
https://arxiv.org/abs/2109.10282
https://github.com/microsoft/unilm/tr...
https://medium.com/people-ai-research...
https://raft.elicit.org/
https://huggingface.co/spaces/ought/r...
https://huggingface.co/spaces/ought/r...
https://arxiv.org/pdf/2109.14076.pdf
https://arxiv.org/pdf/2109.14394.pdf
https://www.robots.ox.ac.uk/~vgg/rese...
https://zenodo.org/record/5528345#.YV...
https://github.com/yukimasano/PASS/
https://openreview.net/pdf?id=BwzYI-K...
https://github.com/pytorch/data?utm_s...
Samsung Method to copy paste brain onto chip
https://www.engadget.com/samsung-copy...
Links:
TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick
YouTube: https://www.youtube.com/c/yannickilcher
Twitter: https://twitter.com/ykilcher
Discord: https://discord.gg/4H8xxDF
BitChute: https://www.bitchute.com/channel/yann...
Minds: https://www.minds.com/ykilcher
Parler: https://parler.com/profile/YannicKilcher
LinkedIn: https://www.linkedin.com/in/ykilcher
BiliBili: https://space.bilibili.com/1824646584
If you want to support me, the best thing to do is to share out the content :)
If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):
SubscribeStar: https://www.subscribestar.com/yannick...
Patreon: https://www.patreon.com/yannickilcher
Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq
Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2
Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m
Grokking is a phenomenon when a neural network suddenly learns a pattern in the dataset and jumps from random chance generalization to perfect generalization very suddenly. This paper demonstrates grokking on small algorithmic datasets where a network has to fill in binary tables. Interestingly, the learned latent spaces show an emergence of the underlying binary operations that the data were created with.
OUTLINE:
0:00 - Intro & Overview
1:40 - The Grokking Phenomenon
3:50 - Related: Double Descent
7:50 - Binary Operations Datasets
11:45 - What quantities influence grokking?
15:40 - Learned Emerging Structure
17:35 - The role of smoothness
21:30 - Simple explanations win
24:30 - Why does weight decay encourage simplicity?
26:40 - Appendix
28:55 - Conclusion & Comments
Paper: https://mathai-iclr.github.io/papers/...
Abstract:
In this paper we propose to study generalization of neural networks on small algorithmically generated datasets. In this setting, questions about data efficiency, memorization, generalization, and speed of learning can be studied in great detail. In some situations we show that neural networks learn through a process of “grokking” a pattern in the data, improving generalization performance from random chance level to perfect generalization, and that this improvement in generalization can happen well past the point of overfitting. We also study generalization as a function of dataset size and find that smaller datasets require increasing amounts of optimization for generalization. We argue that these datasets provide a fertile ground for studying a poorly understood aspect of deep learning: generalization of overparametrized neural networks beyond memorization of the finite training dataset.
Authors: Alethea Power, Yuri Burda, Harri Edwards, Igor Babuschkin & Vedant Misra
Links:
TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick
YouTube: https://www.youtube.com/c/yannickilcher
Twitter: https://twitter.com/ykilcher
Discord: https://discord.gg/4H8xxDF
BitChute: https://www.bitchute.com/channel/yann...
Minds: https://www.minds.com/ykilcher
Parler: https://parler.com/profile/YannicKilcher
LinkedIn: https://www.linkedin.com/in/ykilcher
BiliBili: https://space.bilibili.com/1824646584
If you want to support me, the best thing to do is to share out the content :)
If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):
SubscribeStar: https://www.subscribestar.com/yannick...
Patreon: https://www.patreon.com/yannickilcher
Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq
Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2
Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m
Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n
Deep Learning has achieved impressive results in the last years, not least due to the massive increases in computational power and data that has gone into these models. Scaling up currently promises to be a reliable way to create more performant systems, but how far can we go? This article explores the limits of exponential scaling in AI, and what people are doing to get around this problem
OUTLINE:
0:00 - Intro & Overview
1:00 - Deep Learning at its limits
3:10 - The cost of overparameterization
5:40 - Extrapolating power usage and CO2 emissions
10:45 - We cannot just continue scaling up
13:25 - Current solution attempts
15:25 - Aside: ImageNet V2
17:50 - Are symbolic methods the way out?
Paper: https://spectrum.ieee.org/deep-learni...
Image by Ralf Vetterle from Pixabay: https://pixabay.com/images/id-1752876/
Links:
TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick
YouTube: https://www.youtube.com/c/yannickilcher
Twitter: https://twitter.com/ykilcher
Discord: https://discord.gg/4H8xxDF
BitChute: https://www.bitchute.com/channel/yann...
Minds: https://www.minds.com/ykilcher
Parler: https://parler.com/profile/YannicKilcher
LinkedIn: https://www.linkedin.com/in/ykilcher
BiliBili: https://space.bilibili.com/1824646584
If you want to support me, the best thing to do is to share out the content :)
If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):
SubscribeStar: https://www.subscribestar.com/yannick...
Patreon: https://www.patreon.com/yannickilcher
Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq
Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2
Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m
Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n
Your Mondaily updates of what's going in the world of Machine Learning.
OUTLINE:
0:00 - Intro
0:20 - New plagiarism case has plot twist
7:25 - CLIP for video surveillance
9:40 - DARPA SubTerranean Challenge
11:00 - Schmidhuber criticizing Turing Lecture
15:00 - OpenAI summarizes books
17:55 - UnBiasIt monitors employees' communications for bias
20:00 - iOS plans to detect depression
21:30 - UK 10 year plan to become AI superpower
23:30 - Helpful Libraries
29:00 - WIT: Wikipedia Image-Text dataset
References:
New plagiarism case with plot twist
https://www.reddit.com/r/MachineLearn...
https://zhuanlan.zhihu.com/p/411800486
https://github.com/cybercore-co-ltd/C...
CLIP used for video surveillance
https://www.reddit.com/r/MachineLearn...
https://github.com/johanmodin/clifs
DARPA SubTerranean Challenge
https://twitter.com/BotJunkie/status/...
https://twitter.com/BotJunkie
https://www.subtchallenge.com/index.html
https://www.subtchallenge.com/resourc...
https://twitter.com/dynamicrobots/sta...
Schmidhuber Blog: Turing Lecture Errors
https://people.idsia.ch/~juergen/scie...
OpenAI on Summarizing Books
https://openai.com/blog/summarizing-b...
https://arxiv.org/pdf/2109.10862.pdf
UnBiasIt to monitor employee language
https://edition.cnn.com/2021/09/20/te...
https://www.unbiasit.com/
iPhone to detect depression
https://www.wsj.com/articles/apple-wa...
https://archive.ph/hRTnw
UK 10-year plan to become AI-superpower
https://www.cnbc.com/2021/09/22/uk-pu...
https://archive.ph/4gkKK
Helpful Libraries
https://twitter.com/scikit_learn/stat...
https://scikit-learn.org/stable/auto_...
https://twitter.com/pcastr/status/144...
https://github.com/google/dopamine
https://github.com/microsoft/muzic
https://ai-muzic.github.io/muzic_logo/
https://ai.facebook.com/blog/dynatask...
https://github.com/tum-pbs/PhiFlow
https://github.com/facebookresearch/dora
Habitat and Matterport 3D Dataset
https://github.com/facebookresearch/h...
https://aihabitat.org/
https://arxiv.org/pdf/2109.08238.pdf
WIT: Wikipedia-Based Image-Text Dataset
https://ai.googleblog.com/2021/09/ann...
Links:
TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick
YouTube: https://www.youtube.com/c/yannickilcher
Twitter: https://twitter.com/ykilcher
Discord: https://discord.gg/4H8xxDF
BitChute: https://www.bitchute.com/channel/yann...
Minds: https://www.minds.com/ykilcher
Parler: https://parler.com/profile/YannicKilcher
LinkedIn: https://www.linkedin.com/in/ykilcher
BiliBili: https://space.bilibili.com/1824646584
If you want to support me, the best thing to do is to share out the content :)
If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):
SubscribeStar: https://www.subscribestar.com/yannick...
Patreon: https://www.patreon.com/yannickilcher
Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq
Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2
Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m
Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n
The peer-review system at Machine Learning conferences has come under much criticism over the last years. One major driver was the infamous 2014 NeurIPS experiment, where a subset of papers were given to two different sets of reviewers. This experiment showed that only about half of all accepted papers were consistently accepted by both committees and demonstrated significant influence of subjectivity. This paper revisits the data from the 2014 experiment and traces the fate of accepted and rejected papers during the 7 years since, and analyzes how well reviewers can assess future impact, among other things.
OUTLINE:
0:00 - Intro & Overview
1:20 - Recap: The 2014 NeurIPS Experiment
5:40 - How much of reviewing is subjective?
11:00 - Validation via simulation
15:45 - Can reviewers predict future impact?
23:10 - Discussion & Comments
Paper: https://arxiv.org/abs/2109.09774
Code: https://github.com/lawrennd/neurips2014/
Abstract:
In this paper we revisit the 2014 NeurIPS experiment that examined inconsistency in conference peer review. We determine that 50% of the variation in reviewer quality scores was subjective in origin. Further, with seven years passing since the experiment we find that for accepted papers, there is no correlation between quality scores and impact of the paper as measured as a function of citation count. We trace the fate of rejected papers, recovering where these papers were eventually published. For these papers we find a correlation between quality scores and impact. We conclude that the reviewing process for the 2014 conference was good for identifying poor papers, but poor for identifying good papers. We give some suggestions for improving the reviewing process but also warn against removing the subjective element. Finally, we suggest that the real conclusion of the experiment is that the community should place less onus on the notion of top-tier conference publications when assessing the quality of individual researchers. For NeurIPS 2021, the PCs are repeating the experiment, as well as conducting new ones.
Authors: Corinna Cortes, Neil D. Lawrence
Links:
TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick
YouTube: https://www.youtube.com/c/yannickilcher
Twitter: https://twitter.com/ykilcher
Discord: https://discord.gg/4H8xxDF
BitChute: https://www.bitchute.com/channel/yann...
Minds: https://www.minds.com/ykilcher
Parler: https://parler.com/profile/YannicKilcher
LinkedIn: https://www.linkedin.com/in/ykilcher
BiliBili: https://space.bilibili.com/1824646584
If you want to support me, the best thing to do is to share out the content :)
If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):
SubscribeStar: https://www.subscribestar.com/yannick...
Patreon: https://www.patreon.com/yannickilcher
Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq
Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2
Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m
Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n
Your regularly irregular updates on what's happening in the Machine Learning world.
OUTLINE:
0:00 - Intro
0:20 - TruthfulQA benchmark shines new light on GPT-3
2:00 - LAION-400M image-text-pair dataset
4:10 - GoogleAI's EfficientNetV2 and CoAtNet
6:15 - Uber's H3: A hexagonal coordinate system
7:40 - AWS NeurIPS 2021 DeepRacer Challenge
8:15 - Helpful Libraries
9:20 - State of PyTorch in September 2021
10:05 - Physics-Based Deep Learning Book
10:35 - Music-conditioned 3D dance generation
11:40 - Stallman's take on legal issues with Codex
12:20 - Tensorflow DirectML on AMD GPUs
13:00 - Schmidhuber Blog: Turing Oversold
ERRATA:
Uber's H3 is actually not new, but from 2018
References:
TruthfulQA - A benchmark assessing truthfulness of language models
https://owainevans.github.io/pdfs/tru...
LAION-400M image-text-pair dataset
https://laion.ai/laion-400-open-dataset/
https://laion.ai/#top
https://gogetfunding.com/help-us-buil...
https://rom1504.github.io/clip-retrie...
GooleAI releases EfficientNetV2 and CoAtNet
https://ai.googleblog.com/2021/09/tow...
Uber's H3 hexagonal coordinate systems
https://eng.uber.com/h3/?utm_source=p...
NeurIPS 2021 DeepRacer Challenge
https://www.aicrowd.com/challenges/ne...
https://aws.amazon.com/deepracer/
https://gitlab.aicrowd.com/deepracer/...
Helpful Libraries
https://github.com/rom1504/img2dataset
https://github.com/facebookresearch/v...
https://github.com/pyg-team/pytorch_g...
https://aws.amazon.com/blogs/machine-...
State of PyTorch in September 2021
https://dev-discuss.pytorch.org/t/sta...
Physics-Based Deep Learning Book
http://physicsbaseddeeplearning.org/i...
https://arxiv.org/pdf/2109.05237.pdf
Music Conditioned 3D dance generation
https://ai.googleblog.com/2021/09/mus...
Richard Stallman on Codex legal issues
https://news.slashdot.org/story/21/09...
Tensorflow DirectML on AMD
https://wccftech.com/amd-microsoft-br...
Schmidhuber: Turing Oversold
https://people.idsia.ch//~juergen/tur...
Links:
TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick
YouTube: https://www.youtube.com/c/yannickilcher
Twitter: https://twitter.com/ykilcher
Discord: https://discord.gg/4H8xxDF
BitChute: https://www.bitchute.com/channel/yann...
Minds: https://www.minds.com/ykilcher
Parler: https://parler.com/profile/YannicKilcher
LinkedIn: https://www.linkedin.com/in/ykilcher
BiliBili: https://space.bilibili.com/1824646584
If you want to support me, the best thing to do is to share out the content :)
If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):
SubscribeStar: https://www.subscribestar.com/yannick...
Patreon: https://www.patreon.com/yannickilcher
Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq
Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2
Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m
Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n
A new benchmark paper has created quite an uproar in the community. TruthfulQA is a dataset of 817 questions probing for imitative falsehoods where language models become less truthful, the larger they get. This surprising counter-intuitive finding validates many people's criticisms of large language models, but is it really the correct conclusion?
OUTLINE:
0:00 - Intro
0:30 - Twitter Paper Announcement
4:10 - Large Language Models are to blame!
5:50 - How was the dataset constructed?
9:25 - The questions are adversarial
12:30 - Are you surprised?!
Paper: https://arxiv.org/abs/2109.07958
Links:
TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick
YouTube: https://www.youtube.com/c/yannickilcher
Twitter: https://twitter.com/ykilcher
Discord: https://discord.gg/4H8xxDF
BitChute: https://www.bitchute.com/channel/yann...
Minds: https://www.minds.com/ykilcher
Parler: https://parler.com/profile/YannicKilcher
LinkedIn: https://www.linkedin.com/in/ykilcher
BiliBili: https://space.bilibili.com/1824646584
If you want to support me, the best thing to do is to share out the content :)
If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):
SubscribeStar: https://www.subscribestar.com/yannick...
Patreon: https://www.patreon.com/yannickilcher
Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq
Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2
Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m
Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n
Variational Autoencoders model the latent space as a set of independent Gaussian random variables, which the decoder maps to a data distribution. However, this independence is not always desired, for example when dealing with video sequences, we know that successive frames are heavily correlated. Thus, any latent space dealing with such data should reflect this in its structure. Topographic VAEs are a framework for defining correlation structures among the latent variables and induce equivariance within the resulting model. This paper shows how such correlation structures can be built by correctly arranging higher-level variables, which are themselves independent Gaussians.
OUTLINE:
0:00 - Intro
1:40 - Architecture Overview
6:30 - Comparison to regular VAEs
8:35 - Generative Mechanism Formulation
11:45 - Non-Gaussian Latent Space
17:30 - Topographic Product of Student-t
21:15 - Introducing Temporal Coherence
24:50 - Topographic VAE
27:50 - Experimental Results
31:15 - Conclusion & Comments
Paper: https://arxiv.org/abs/2109.01394
Code: https://github.com/akandykeller/topog...
Abstract:
In this work we seek to bridge the concepts of topographic organization and equivariance in neural networks. To accomplish this, we introduce the Topographic VAE: a novel method for efficiently training deep generative models with topographically organized latent variables. We show that such a model indeed learns to organize its activations according to salient characteristics such as digit class, width, and style on MNIST. Furthermore, through topographic organization over time (i.e. temporal coherence), we demonstrate how predefined latent space transformation operators can be encouraged for observed transformed input sequences -- a primitive form of unsupervised learned equivariance. We demonstrate that this model successfully learns sets of approximately equivariant features (i.e. "capsules") directly from sequences and achieves higher likelihood on correspondingly transforming test sequences. Equivariance is verified quantitatively by measuring the approximate commutativity of the inference network and the sequence transformations. Finally, we demonstrate approximate equivariance to complex transformations, expanding upon the capabilities of existing group equivariant neural networks.
Authors: T. Anderson Keller, Max Welling
Links:
TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick
YouTube: https://www.youtube.com/c/yannickilcher
Twitter: https://twitter.com/ykilcher
Discord: https://discord.gg/4H8xxDF
BitChute: https://www.bitchute.com/channel/yann...
Minds: https://www.minds.com/ykilcher
Parler: https://parler.com/profile/YannicKilcher
LinkedIn: https://www.linkedin.com/in/ykilcher
BiliBili: https://space.bilibili.com/1824646584
If you want to support me, the best thing to do is to share out the content :)
If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):
SubscribeStar: https://www.subscribestar.com/yannick...
Patreon: https://www.patreon.com/yannickilcher
Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq
Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2
Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m
Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n
Your regularly irregular update on what's happening in the world of Machine Learning.
OUTLINE:
0:00 - Intro
0:15 - Sponsor: Weights & Biases
1:55 - ML YouTuber reaches 100k subscribers
2:40 - Facebook AI pushes Textless NLP
5:30 - Schmidhuber blog post: I invented everything
7:55 - TikTok algorithm rabbitholes users
10:45 - Roomba learns to avoid poop
11:50 - AI can spot art forgeries
14:55 - Deepmind's plans to separate from Google
16:15 - Cohere raises 40M
16:55 - US Judge rejects AI inventor on patent
17:55 - Altman: GPT-4 not much bigger than GPT-3
18:45 - Salesforce CodeT5
19:45 - DeepMind Reinforcement Learning Lecture Series
20:15 - WikiGraphs Dataset
20:40 - LiveCell Dataset
21:00 - SpeechBrain
21:10 - AI-generated influencer gains 100 sponsorships
22:20 - AI News Questions
23:15 - AI hiring tools reject millions of valid applicants
Sponsor: Weights & Biases
https://wandb.me/start
References:
Facebook AI creates Textless NLP
https://ai.facebook.com/blog/textless...
https://speechbot.github.io/pgslm/?fb...
Schmidhuber invented everything
https://people.idsia.ch/~juergen/most...
How TikTok's algorithm works
https://www.wsj.com/video/series/insi...
Roomba learns to avoid poop
https://edition.cnn.com/2021/09/09/te...
Amateur develops fake art detector
https://blogs.nvidia.com/blog/2021/08...
https://spectrum.ieee.org/this-ai-can...
DeepMind's plan to break away from Google
https://www.businessinsider.com/deepm...
https://archive.ph/8s5IK
Cohere raises USD 40M
https://www.fastcompany.com/90670635/...
https://cohere.ai/
US judge refuses AI patent
https://www.theregister.com/2021/09/0...
Sam Altman on GPT-4
https://www.reddit.com/r/OpenAI/comme...
Salesforce releases CodeT5
https://blog.einstein.ai/codet5/
DeepMind RL lecture series
https://deepmind.com/learning-resourc...
WikiGraphs Dataset
https://github.com/deepmind/deepmind-...
LiveCell Dataset
https://sartorius-research.github.io/...
https://www.nature.com/articles/s4159...
SpeechBrain Library
https://speechbrain.github.io/
AI generated influencer lands 100 sponsorships
https://www.allkpop.com/article/2021/...
AI News Questions
https://www.forbes.com/sites/tomtaull...
https://mindmatters.ai/2021/09/isnt-i...
https://fortune.com/2021/09/07/deepmi...
https://www.forbes.com/sites/anniebro...
https://www.cnbctv18.com/views/view-a...
https://www.kcrw.com/culture/shows/li...
https://techcrunch.com/2021/09/07/ai-...
https://www.forbes.com/sites/bernardm...
AI hiring tools mistakenly reject millions of applicants
https://www.theverge.com/2021/9/6/226...
Links:
TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick
YouTube: https://www.youtube.com/c/yannickilcher
Twitter: https://twitter.com/ykilcher
Discord: https://discord.gg/4H8xxDF
BitChute: https://www.bitchute.com/channel/yann...
Minds: https://www.minds.com/ykilcher
Parler: https://parler.com/profile/YannicKilcher
LinkedIn: https://www.linkedin.com/in/ykilcher
BiliBili: https://space.bilibili.com/1824646584
If you want to support me, the best thing to do is to share out the content :)
OUTLINE:
0:00 - 100k!
1:00 - Announcements & Thanks
3:55 - Channel Statistics
Links:
TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick
YouTube: https://www.youtube.com/c/yannickilcher
Twitter: https://twitter.com/ykilcher
Discord: https://discord.gg/4H8xxDF
BitChute: https://www.bitchute.com/channel/yann...
Minds: https://www.minds.com/ykilcher
Parler: https://parler.com/profile/YannicKilcher
LinkedIn: https://www.linkedin.com/in/yannic-ki...
BiliBili: https://space.bilibili.com/1824646584
If you want to support me, the best thing to do is to share out the content :)
If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):
SubscribeStar: https://www.subscribestar.com/yannick...
Patreon: https://www.patreon.com/yannickilcher
Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq
Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2
Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m
Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n
Your regular updates on what's happening in the ML world!
OUTLINE:
0:00 - Intro
0:15 - Sponsor: Weights & Biases
1:45 - Google shuts down health streams
4:25 - AI predicts race from blurry X-Rays
7:35 - Facebook labels black men as primates
11:05 - Distill papers on Graph Neural Networks
11:50 - Jürgen Schmidhuber to lead KAUST AI Initiative
12:35 - GitHub brief on DMCA notices for source code
14:55 - Helpful Reddit Threads
19:40 - Simple Tricks to improve Transformers
20:40 - Apple's Unconstrained Scene Generation
21:40 - Common Objects in 3D dataset
22:20 - WarpDrive Multi-Agent RL framework
23:10 - My new paper: Boosting Search Agents & MuZero
25:15 - Can AI detect depression from speech?
References:
Google shuts down Health Streams
https://techcrunch.com/2021/08/26/goo...
AI predicts race from X-Rays
https://www.iflscience.com/technology...
https://arxiv.org/ftp/arxiv/papers/21...
Facebook labels black men as primates
https://www.nytimes.com/2021/09/03/te...
https://en.wikipedia.org/wiki/Human
Distill articles on GNNs
https://distill.pub/2021/gnn-intro/
https://distill.pub/2021/understandin...
Jürgen Schmidhuber leads KAUST AI initiative
https://people.idsia.ch/~juergen/kaus...
GitHub issues court brief on code DMCAs
https://github.blog/2021-08-31-vague-...
Useful Reddit Threads
https://www.reddit.com/r/MachineLearn...
https://www.reddit.com/r/MachineLearn...
https://www.reddit.com/r/MachineLearn...
https://www.reddit.com/r/MachineLearn...
Tricks to improve Transformers
https://arxiv.org/pdf/2108.12284.pdf
Unconstrained Scene Generation
https://apple.github.io/ml-gsn/
Common Objects in 3D dataset
https://ai.facebook.com/blog/common-o...
WarpDrive Multi-Agent RL framework
https://blog.einstein.ai/warpdrive-fa...
Boosting Search Engines / MuZero Code
https://arxiv.org/abs/2109.00527
https://github.com/google-research/go...
https://github.com/google-research/la...
Can AI detect depression?
https://venturebeat.com/2021/08/31/ai...
Links:
TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick
YouTube: https://www.youtube.com/c/yannickilcher
Twitter: https://twitter.com/ykilcher
Discord: https://discord.gg/4H8xxDF
BitChute: https://www.bitchute.com/channel/yann...
Minds: https://www.minds.com/ykilcher
Parler: https://parler.com/profile/YannicKilcher
LinkedIn: https://www.linkedin.com/in/yannic-ki...
BiliBili: https://space.bilibili.com/1824646584
If you want to support me, the best thing to do is to share out the content :)
If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):
SubscribeStar: https://www.subscribestar.com/yannick...
Patreon: https://www.patreon.com/yannickilcher
Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq
Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2
Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m
Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n
Vanilla Transformers are excellent sequence models, but suffer from very harsch constraints on the length of the sequences they can process. Several attempts have been made to extend the Transformer's sequence length, but few have successfully gone beyond a constant factor improvement. This paper presents a method, based on continuous attention mechanisms, to attend to an unbounded past sequence by representing the past as a continuous signal, rather than a sequence. This enables the Infty-Former to effectively enrich the current context with global information, which increases performance on long-range dependencies in sequence tasks. Further, the paper presents the concept of sticky memories, which highlight past events that are of particular importance and elevates their representation in the long-term memory.
OUTLINE:
0:00 - Intro & Overview
1:10 - Sponsor Spot: Weights & Biases
3:35 - Problem Statement
8:00 - Continuous Attention Mechanism
16:25 - Unbounded Memory via concatenation & contraction
18:05 - Does this make sense?
20:25 - How the Long-Term Memory is used in an attention layer
27:40 - Entire Architecture Recap
29:30 - Sticky Memories by Importance Sampling
31:25 - Commentary: Pros and cons of using heuristics
32:30 - Experiments & Results
Paper: https://arxiv.org/abs/2109.00301
Sponsor: Weights & Biases
https://wandb.me/start
Abstract:
Transformers struggle when attending to long contexts, since the amount of computation grows with the context length, and therefore they cannot model long-term memories effectively. Several variations have been proposed to alleviate this problem, but they all have a finite memory capacity, being forced to drop old information. In this paper, we propose the ∞-former, which extends the vanilla transformer with an unbounded long-term memory. By making use of a continuous-space attention mechanism to attend over the long-term memory, the ∞-former's attention complexity becomes independent of the context length. Thus, it is able to model arbitrarily long contexts and maintain "sticky memories" while keeping a fixed computation budget. Experiments on a synthetic sorting task demonstrate the ability of the ∞-former to retain information from long sequences. We also perform experiments on language modeling, by training a model from scratch and by fine-tuning a pre-trained language model, which show benefits of unbounded long-term memories.
Authors: Pedro Henrique Martins, Zita Marinho, André F. T. Martins
Links:
TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick
YouTube: https://www.youtube.com/c/yannickilcher
Twitter: https://twitter.com/ykilcher
Discord: https://discord.gg/4H8xxDF
BitChute: https://www.bitchute.com/channel/yann...
Minds: https://www.minds.com/ykilcher
Parler: https://parler.com/profile/YannicKilcher
LinkedIn: https://www.linkedin.com/in/yannic-ki...
BiliBili: https://space.bilibili.com/1824646584
If you want to support me, the best thing to do is to share out the content :)
If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):
SubscribeStar: https://www.subscribestar.com/yannick...
Patreon: https://www.patreon.com/yannickilcher
Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq
Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2
Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m
Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n
OUTLINE:
0:00 - Intro
0:30 - Reconnaissance Blind Chess NeurIPS 2021 Competition
3:40 - Colab Pro no longer top priority for GPUs
4:45 - DeepMind uses Graph NNs to do traffic prediction
6:00 - Helpful Libraries: Isaac Gym, Differentiable Human, LVIS, BEHAVIOR
10:25 - Cerebras Wafer Scale Engine Cluster
12:15 - AI Voice Synthesis for Val Kilmer
14:20 - Can AI give thoughtful gifts?
References:
Reconnaissance Blind Chess NeurIPS 2021 Competition
https://rbc.jhuapl.edu/
https://rbc.jhuapl.edu/gameRules
Colab Pro no longer top priority
https://www.reddit.com/r/MachineLearn...
Google Maps ETA prediction using Graph Neural Networks
https://arxiv.org/pdf/2108.11482.pdf
Isaac Gym: RL simulator on GPU
https://arxiv.org/abs/2108.10470
https://sites.google.com/view/isaacgy...
https://developer.nvidia.com/isaac-gym
Cerebras Cluster for massive AI models
https://www.wired.com/story/cerebras-...
Helpful Libraries / Datasets
https://nimblephysics.org/docs/human-...
https://www.lvisdataset.org/
https://arxiv.org/pdf/2108.03332.pdf
AI Voice Reconstruction
https://www.washingtonpost.com/techno...
Can AI make thoughtful gifts?
https://www.forbes.com/sites/anniebro...
Links:
TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick
YouTube: https://www.youtube.com/c/yannickilcher
Twitter: https://twitter.com/ykilcher
Discord: https://discord.gg/4H8xxDF
BitChute: https://www.bitchute.com/channel/yann...
Minds: https://www.minds.com/ykilcher
Parler: https://parler.com/profile/YannicKilcher
LinkedIn: https://www.linkedin.com/in/yannic-ki...
BiliBili: https://space.bilibili.com/1824646584
If you want to support me, the best thing to do is to share out the content :)
If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):
SubscribeStar: https://www.subscribestar.com/yannick...
Patreon: https://www.patreon.com/yannickilcher
Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq
Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2
Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m
Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n
Transformers are essentially set models that need additional inputs to make sense of sequence data. The most widespread additional inputs are position encodings or position embeddings, which add sequence index information in various forms. However, this has put a limit on the resulting model, which cannot run inference on sequences longer than it has been trained on, as it would encounter unfamiliar position encodings. ALiBi solves this by proposing simple linear fixed biases as position information, adding negligible overhead in time and memory, but surprisingly, the resulting model is able to handle inference on sequences many times as long as its training sequences.
OUTLINE:
0:00 - Intro & Overview
1:40 - Position Encodings in Transformers
4:55 - Sinusoidial Position Encodings
11:50 - ALiBi Position Encodings
20:50 - How to choose the slope parameter
23:55 - Experimental Results
29:10 - Comments & Conclusion
Paper: https://ofir.io/train_short_test_long...
Code: https://github.com/ofirpress/attentio...
Abstract:
Since the introduction of the transformer model by Vaswani et al. (2017), a fundamental question remains open: how to achieve extrapolation at inference time to longer sequences than seen during training? We first show that extrapolation can be improved by changing the position representation method, though we find that existing proposals do not allow efficient extrapolation. We introduce a simple and efficient method, Attention with Linear Biases (ALiBi), that allows for extrapolation. ALiBi does not add positional embeddings to the word embeddings; instead, it biases the query-key attention scores with a term that is proportional to their distance. We show that this method allows training a 1.3 billion parameter model on input sequences of length 1024 that extrapolates to input sequences of length 2048, achieving the same perplexity as a sinusoidal position embedding model trained on inputs of length 2048, 11% faster and using 11% less memory. ALiBi’s inductive bias towards recency allows it to outperform multiple strong position methods on the WikiText-103 benchmark. Finally, we provide analysis of ALiBi to understand why it leads to better performance.
Authors: Ofir Press, Noah A. Smith, Mike Lewis
Links:
TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick
YouTube: https://www.youtube.com/c/yannickilcher
Twitter: https://twitter.com/ykilcher
Discord: https://discord.gg/4H8xxDF
BitChute: https://www.bitchute.com/channel/yann...
Minds: https://www.minds.com/ykilcher
Parler: https://parler.com/profile/YannicKilcher
LinkedIn: https://www.linkedin.com/in/yannic-ki...
BiliBili: https://space.bilibili.com/1824646584
If you want to support me, the best thing to do is to share out the content :)
If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):
SubscribeStar: https://www.subscribestar.com/yannick...
Patreon: https://www.patreon.com/yannickilcher
Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq
Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2
Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m
Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n
The best place to keep up to date with the latest and greatest from the ML world!
OUTLINE:
0:00 - Intro & Sponsor
3:15 - A high-profile case of plagiarism shocks the ML world
11:55 - Stanford AI releases paper on "Foundation Models"
19:45 - Updates on Apple's NeuralHash
20:45 - RL control for two-player splorts
21:45 - Tesla's AI Day
23:55 - COMMA THREE announced
24:40 - Intel winding down RealSense cameras
25:20 - IBM unveils Telum Processor
25:50 - Lux AI Challenge & Neural MMO Challenge
26:50 - Dribnet's CLIP PixelArt
27:40 - Multi-Agent RL papers are mostly fake
28:50 - I can't even come up with a segment title
29:25 - AI News Questions
31:20 - Frameworks & Libraries
Sponsor: Weights & Biases
https://wandb.ai
References:
Plagiarism case shocks ML world
https://arxiv.org/abs/2102.07870v1
https://arxiv.org/pdf/2102.07870v1.pdf
https://arxiv.org/abs/2108.05862
https://arxiv.org/pdf/2108.05862v1.pdf
https://www.reddit.com/r/MachineLearn...
https://michaelsdr.github.io/momentum...
https://www.zhihu.com/question/480075...
https://zhuanlan.zhihu.com/p/40035196...
https://finance.sina.com.cn/tech/2021...
https://duoli.org/
https://web.archive.org/web/202108160...
https://twitter.com/shaohua0116/statu...
Stanford AI targets Foundation Models
https://arxiv.org/abs/2108.07258
https://arxiv.org/pdf/2108.07258.pdf
https://ieeexplore.ieee.org/document/...
https://xgboost.readthedocs.io/en/lat...
https://en.wikipedia.org/wiki/Support...
https://scikit-learn.org/stable/modul...
https://syncedreview.com/2019/06/27/t...
https://openai.com/blog/better-langua...
NeuralHash Saga Continues
https://www.reddit.com/r/MachineLearn...
https://blog.roboflow.com/neuralhash-...
https://www.kron4.com/news/bay-area/b...
RL Control for competitive sports
https://ai.facebook.com/research/publ...
Tesla AI Day
https://www.youtube.com/watch?v=ABbDB...
https://spectrum.ieee.org/elon-musk-r...
https://www.youtube.com/watch?v=j0z4F...
George Hotz announces COMMA THREE
https://www.youtube.com/watch?v=jJn2O...
https://comma.ai/shop/products/three
Intel abandons RealSense cameras
https://www.crn.com/news/components-p...
IBM unveils Telum Processor
https://www.prnewswire.com/news-relea...
Kaggle Lux AI challenge
https://www.kaggle.com/c/lux-ai-2021
Neural MMO challenge
https://www.aicrowd.com/challenges/th...
Dribnet's PixelArt
https://twitter.com/dribnet/status/14...
Multi-Agent RL papers mostly fake
https://www.reddit.com/r/reinforcemen...
Elon Musk, Lex Fridman tweets trigger news story
https://www.benzinga.com/news/21/08/2...
News Questions:
https://www.zdnet.com/article/can-ai-...
https://entertainment.inquirer.net/41...
https://www.analyticsinsight.net/whic...
https://www.bbc.co.uk/programmes/m000...
https://ricochet.com/podcast/cosm-tec...
https://www.designnews.com/automation...
https://www.forbes.com/sites/anniebro...
3D Volleyball RL environment
https://www.reddit.com/r/MachineLearn...
Maze RL framework
https://enliteai.medium.com/maze-appl...
Wanderer 2 HN Search
https://metaphor.so/
Transformers have become the dominant model class in the last few years for large data, but their quadratic complexity in terms of sequence length has plagued them until now. Fastformer claims to be the fastest and most performant linear attention variant, able to consume long contexts at once. This is achieved by a combination of additive attention and elementwise products. While initial results look promising, I have my reservations...
OUTLINE:
0:00 - Intro & Outline
2:15 - Fastformer description
5:20 - Baseline: Classic Attention
10:00 - Fastformer architecture
12:50 - Additive Attention
18:05 - Query-Key element-wise multiplication
21:35 - Redundant modules in Fastformer
25:00 - Problems with the architecture
27:30 - Is this even attention?
32:20 - Experimental Results
34:50 - Conclusion & Comments
Paper: https://arxiv.org/abs/2108.09084
Abstract:
Transformer is a powerful model for text understanding. However, it is inefficient due to its quadratic complexity to input sequence length. Although there are many methods on Transformer acceleration, they are still either inefficient on long sequences or not effective enough. In this paper, we propose Fastformer, which is an efficient Transformer model based on additive attention. In Fastformer, instead of modeling the pair-wise interactions between tokens, we first use additive attention mechanism to model global contexts, and then further transform each token representation based on its interaction with global context representations. In this way, Fastformer can achieve effective context modeling with linear complexity. Extensive experiments on five datasets show that Fastformer is much more efficient than many existing Transformer models and can meanwhile achieve comparable or even better long text modeling performance.
Authors: Chuhan Wu, Fangzhao Wu, Tao Qi, Yongfeng Huang
Links:
TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick
YouTube: https://www.youtube.com/c/yannickilcher
Twitter: https://twitter.com/ykilcher
Discord: https://discord.gg/4H8xxDF
BitChute: https://www.bitchute.com/channel/yann...
Minds: https://www.minds.com/ykilcher
Parler: https://parler.com/profile/YannicKilcher
LinkedIn: https://www.linkedin.com/in/yannic-ki...
BiliBili: https://space.bilibili.com/1824646584
If you want to support me, the best thing to do is to share out the content :)
If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):
SubscribeStar: https://www.subscribestar.com/yannick...
Patreon: https://www.patreon.com/yannickilcher
Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq
Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2
Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m
Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n
Humans don't spend the same amount of mental effort on all problems equally. Instead, we respond quickly to easy tasks, and we take our time to deliberate hard tasks. DeepMind's PonderNet attempts to achieve the same by dynamically deciding how many computation steps to allocate to any single input sample. This is done via a recurrent architecture and a trainable function that computes a halting probability. The resulting model performs well in dynamic computation tasks and is surprisingly robust to different hyperparameter settings.
OUTLINE:
0:00 - Intro & Overview
2:30 - Problem Statement
8:00 - Probabilistic formulation of dynamic halting
14:40 - Training via unrolling
22:30 - Loss function and regularization of the halting distribution
27:35 - Experimental Results
37:10 - Sensitivity to hyperparameter choice
41:15 - Discussion, Conclusion, Broader Impact
Paper: https://arxiv.org/abs/2107.05407
Abstract:
In standard neural networks the amount of computation used grows with the size of the inputs, but not with the complexity of the problem being learnt. To overcome this limitation we introduce PonderNet, a new algorithm that learns to adapt the amount of computation based on the complexity of the problem at hand. PonderNet learns end-to-end the number of computational steps to achieve an effective compromise between training prediction accuracy, computational cost and generalization. On a complex synthetic problem, PonderNet dramatically improves performance over previous adaptive computation methods and additionally succeeds at extrapolation tests where traditional neural networks fail. Also, our method matched the current state of the art results on a real world question and answering dataset, but using less compute. Finally, PonderNet reached state of the art results on a complex task designed to test the reasoning capabilities of neural networks.1
Authors: Andrea Banino, Jan Balaguer, Charles Blundell
Links:
TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick
YouTube: https://www.youtube.com/c/yannickilcher
Twitter: https://twitter.com/ykilcher
Discord: https://discord.gg/4H8xxDF
BitChute: https://www.bitchute.com/channel/yann...
Minds: https://www.minds.com/ykilcher
Parler: https://parler.com/profile/YannicKilcher
LinkedIn: https://www.linkedin.com/in/yannic-ki...
BiliBili: https://space.bilibili.com/1824646584
If you want to support me, the best thing to do is to share out the content :)
If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):
SubscribeStar: https://www.subscribestar.com/yannick...
Patreon: https://www.patreon.com/yannickilcher
Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq
Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2
Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m
Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n
Send your Apple fanboy friends to prison with this one simple trick ;) We break Apple's NeuralHash algorithm used to detect CSAM for iCloud photos. I show how it's possible to craft arbitrary hash collisions from any source / target image pair using an adversarial example attack. This can be used for many purposes, such as evading detection, or forging false positives, triggering manual reviews.
OUTLINE:
0:00 - Intro
1:30 - Forced Hash Collisions via Adversarial Attacks
2:30 - My Successful Attack
5:40 - Results
7:15 - Discussion
DISCLAIMER: This is for demonstration and educational purposes only. This is not an endorsement of illegal activity or circumvention of law.
Code: https://github.com/yk/neural_hash_col...
Extract Model: https://github.com/AsuharietYgvar/App...
My Video on NeuralHash: https://youtu.be/z15JLtAuwVI
Links:
TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick
YouTube: https://www.youtube.com/c/yannickilcher
Twitter: https://twitter.com/ykilcher
Discord: https://discord.gg/4H8xxDF
BitChute: https://www.bitchute.com/channel/yann...
Minds: https://www.minds.com/ykilcher
Parler: https://parler.com/profile/YannicKilcher
LinkedIn: https://www.linkedin.com/in/yannic-ki...
BiliBili: https://space.bilibili.com/1824646584
If you want to support me, the best thing to do is to share out the content :)
If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):
SubscribeStar: https://www.subscribestar.com/yannick...
Patreon: https://www.patreon.com/yannickilcher
Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq
Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2
Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m
Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n
An in-depth look over what's going on in the world of Machine Learning and Artificial intelligence. Subscribe now and make Monday the best day of the week!
OUTLINE:
0:00 - Intro
0:20 - Sponsor: Weights & Biases
3:00 - Nvidia's CEO was rendered during Keynote
5:00 - AI21 Labs releases Jurassic-1 language model
7:00 - Tortured Phrases reveal plagiarism
10:05 - Cortical neurons are computationally complex
11:55 - OpenAI Codex Update & Challenge
13:30 - Automated drug abuse prevention gone wrong
17:55 - Rapid News Questions
18:40 - SoundStream learned neural audio codec
19:40 - RoboMimic framework for robotics research
20:05 - Droidlet framework for agent training
20:40 - Unidentified Video Objects Benchmark
21:45 - Grammatical Error Correction Dataset
22:15 - ColabPro Plus available
23:05 - BigBench Self-Awareness benchmark for language models
Sponsor: Weights & Biases
https://wandb.ai
References:
NVIDIA renders CEO during keynote
https://www.vice.com/en/article/88nbp...
https://blogs.nvidia.com/blog/2021/08...
https://www.youtube.com/watch?v=eAn_o...
AI21 Labs announces Jurassic-1 model
https://www.ai21.com/blog/announcing-...
https://studio.ai21.com/
https://twitter.com/yoavgo/status/142...
Tortured Phrases point to plagiarism
https://www.nature.com/articles/d4158...
Real Neurons are insanely complex
https://www.sciencedirect.com/science...
OpenAI Codex Challenge & Update
https://challenge.openai.com/
https://challenge.openai.com/codex/le...
https://openai.com/blog/openai-codex/...
Automated drug abuse prevention goes wrong
https://www.wired.com/story/opioid-dr...
News Questions
https://www.imeche.org/news/news-arti...
https://newseu.cgtn.com/news/2021-08-...
https://www.growingproduce.com/citrus...
https://www.cioreview.com/news/artifi...
SoundStream Neural Audio Codec
https://ai.googleblog.com/2021/08/sou...
RoboMimic Framework
https://arise-initiative.github.io/ro...
Droidlet Framework
https://ai.facebook.com/blog/droidlet...
Unidentified Video Objects Benchmark
https://ai.facebook.com/blog/introduc...
Grammatical Error Correction Dataset
https://ai.googleblog.com/2021/08/the...
Colab Pro Plus is "even better"
https://colab.research.google.com/signup
BIG-Bench Self-Awareness Benchmark for Language Models
https://github.com/google/BIG-bench/t...
Links:
TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick
YouTube: https://www.youtube.com/c/yannickilcher
Twitter: https://twitter.com/ykilcher
Discord: https://discord.gg/4H8xxDF
BitChute: https://www.bitchute.com/channel/yann...
Minds: https://www.minds.com/ykilcher
Parler: https://parler.com/profile/YannicKilcher
LinkedIn: https://www.linkedin.com/in/yannic-ki...
BiliBili: https://space.bilibili.com/1824646584
If you want to support me, the best thing to do is to share out the content :)
If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):
SubscribeStar: https://www.subscribestar.com/yannick...
Patreon: https://www.patreon.com/yannickilcher
Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq
Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2
Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m
Apple recently announced scanning all images uploaded to iCloud for CSAM (child abuse material), and that this scan would happen locally on users' phones. We take a look at the technical report and explore how the system works in detail, how it is designed to preserve user privacy, and what weak points it still has.
OUTLINE:
0:00 - Introduction
3:05 - System Requirements
9:15 - System Overview
14:00 - NeuralHash
20:45 - Private Set Intersection
31:15 - Threshold Secret Sharing
35:25 - Synthetic Match Vouchers
38:20 - Problem 1: Who controls the database?
42:40 - Problem 2: Adversarial Attacks
49:40 - Comments & Conclusion
Paper: https://www.apple.com/child-safety/pd...
ML News Episode about CSAM: https://youtu.be/gFkBqD2hbnU
Abstract:
CSAM Detection enables Apple to accurately identify and report iCloud users who store known Child Sexual Abuse Material (CSAM) in their iCloud Photos accounts. Apple servers flag accounts exceeding a threshold number of images that match a known database of CSAM image hashes so that Apple can provide relevant information to the National Center for Missing and Exploited Children (NCMEC). This process is secure, and is expressly designed to preserve user privacy.
CSAM Detection provides these privacy and security assurances:
• Apple does not learn anything about images that do not match the known CSAM database.
• Apple can’t access metadata or visual derivatives for matched CSAM images until a threshold of matches is exceeded for an iCloud Photos account.
• The risk of the system incorrectly flagging an account is extremely low. In addition, Apple manually reviews all reports made to NCMEC to ensure reporting accuracy.
• Users can’t access or view the database of known CSAM images.
• Users can’t identify which images were flagged as CSAM by the system.
For detailed information about the cryptographic protocol and security proofs that the CSAM Detection process uses, see The Apple PSI System.
Links:
TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick
YouTube: https://www.youtube.com/c/yannickilcher
Twitter: https://twitter.com/ykilcher
Discord: https://discord.gg/4H8xxDF
BitChute: https://www.bitchute.com/channel/yann...
Minds: https://www.minds.com/ykilcher
Parler: https://parler.com/profile/YannicKilcher
LinkedIn: https://www.linkedin.com/in/yannic-ki...
BiliBili: https://space.bilibili.com/1824646584
If you want to support me, the best thing to do is to share out the content :)
If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):
SubscribeStar: https://www.subscribestar.com/yannick...
Patreon: https://www.patreon.com/yannickilcher
Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq
Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2
Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m
Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n
Your update on the latest news in the AI and Machine Learning world.
OUTLINE:
0:00 - Intro
0:15 - Sponsor: Weights & Biases
3:30 - Apple to scan iDevices for illegal content
14:10 - EU approves chatcontrol
15:20 - Machine Learning FAQ book
17:40 - TimeDial & Disfl-QA Conversation Datasets
20:30 - VoxPopuli Speech Dataset
21:00 - Google Tensor chip coming to Pixel 6
21:30 - Pentagon uses AI to predict events
23:10 - Sketch your own GAN
24:45 - Can a Fruit Fly learn Word Embeddings?
26:00 - Master Faces beat facial recognition system
27:25 - PyTorch profiler 1.9
27:55 - 0 A.D. gets reinforcement learning interface
28:40 - BeatBot cleans up cigarette butts on the beach
Sponsor: Weights & Biases
https://wandb.ai
References:
Apple to scan iDevices for illegal content
https://techcrunch.com/2021/08/05/app...
http://tylerneylon.com/a/lsh1/
EU approves chatcontrol
https://european-pirateparty.eu/parli...
Machine Learning FAQ book
https://rentruewang.github.io/learnin...
TimeDial & Disfl-QA: New datasets for conversational NLP
https://ai.googleblog.com/2021/08/two...
VoxPopuli: Giant partially labeled speech dataset
https://github.com/facebookresearch/v...
Google's Tensor chip coming to Pixel 6
https://blog.google/products/pixel/go...
Pentagon uses AI for predicting relevant events in advance
https://www.engadget.com/pentagon-ai-...
Sketch Your Own GAN
https://peterwang512.github.io/GANSke...
Can a fruit fly learn word embeddings?
https://arxiv.org/pdf/2101.06887.pdf
Master Faces for attacking facial recognition systems
https://arxiv.org/pdf/2108.01077.pdf
PyTorch Profiler v1.9
https://www.marktechpost.com/2021/08/...
0 A.D. adds Reinforcement Learning interface
https://play0ad.com/media/screenshots/
https://trac.wildfiregames.com/wiki/G...
BeachBot cleans up cigarette butts on the beach
https://news.yahoo.com/beachbot-rover...
Links:
TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick
YouTube: https://www.youtube.com/c/yannickilcher
Twitter: https://twitter.com/ykilcher
Discord: https://discord.gg/4H8xxDF
BitChute: https://www.bitchute.com/channel/yann...
Minds: https://www.minds.com/ykilcher
Parler: https://parler.com/profile/YannicKilcher
LinkedIn: https://www.linkedin.com/in/yannic-ki...
BiliBili: https://space.bilibili.com/1824646584
If you want to support me, the best thing to do is to share out the content :)
If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):
SubscribeStar: https://www.subscribestar.com/yannick...
Patreon: https://www.patreon.com/yannickilcher
Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq
Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2
Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m
Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n
OUTLINE:
0:00 - Intro
0:20 - Sponsor: Weights & Biases
3:45 - AI legally recognized as patent inventor
8:35 - Alpeh Alpha raises USD 27Mio to build European OpenAI
10:20 - AMP advances AI aided recycling
11:20 - DeepMind builds XLand RL environment
13:15 - Cognitive Behavioral Therapy as an app
16:15 - Wordcraft interactive AI text editor
17:05 - ML used to cheat in console games
18:10 - Google's OpenBuildings Dataset
20:00 - Most ML COVID tools are flawed
21:10 - DALL-E mini released
21:55 - Helpful Libraries
25:20 - FSF funds papers discussing CoPilot
SPONSOR: Weights & Biases
https://wandb.ai
References:
AI legally recognized as patent inventor
https://www.globallegalpost.com/news/...
https://www.abc.net.au/news/2021-08-0...
https://artificialinventor.com/freque...
https://artificialinventor.com/dabus/
https://www.worldscientific.com/doi/a...
https://www.worldscientific.com/doi/e...
https://imagination-engines.com/dabus...
https://imagination-engines.com/about...
https://www.nextbigfuture.com/2016/03...
https://www.actiac.org/system/files/D...
Alpeh Alpha raises USD 27Mio to build European OpenAI
https://techcrunch.com/2021/07/27/ger...
AMP advances AI aided recycling
https://www.robotics247.com/article/a...
DeepMind builds XLand RL environment
https://deepmind.com/blog/article/gen...
https://deepmind.com/research/publica...
Cognitive Behavioral Therapy as an app
https://www.nytimes.com/2021/06/01/he...
Wordcraft interactive AI text editor
https://syncedreview.com/2021/07/21/d...
https://arxiv.org/abs/2107.07430
https://www.youtube.com/watch?v=9p4mf...
ML used to cheat in console games
https://au.pcmag.com/games/88121/mach...
Google's OpenBuildings Dataset
https://ai.googleblog.com/2021/07/map...
https://sites.research.google/open-bu...
Most ML COVID tools are flawed
https://www.technologyreview.com/2021...
DALL-E mini released
https://wandb.ai/dalle-mini/dalle-min...
https://huggingface.co/spaces/flax-co...
Helpful Libraries
https://www.openai.com/blog/triton/
https://github.com/openai/triton
https://github.com/microsoft/FLAML
https://github.com/clip-italian/clip-...
https://deepmind.com/research/open-so...
https://github.com/deepmind/meltingpot
https://www.roboti.us/license.html
https://github.com/openai/gym/issues/...
https://github.com/jkterry1
FSF funds papers discussing CoPilot
https://www.fsf.org/blogs/licensing/f...
https://www.gnu.org/philosophy/who-do...
Links:
TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick
YouTube: https://www.youtube.com/c/yannickilcher
Twitter: https://twitter.com/ykilcher
Discord: https://discord.gg/4H8xxDF
BitChute: https://www.bitchute.com/channel/yann...
Minds: https://www.minds.com/ykilcher
Parler: https://parler.com/profile/YannicKilcher
LinkedIn: https://www.linkedin.com/in/yannic-ki...
BiliBili: https://space.bilibili.com/1824646584
If you want to support me, the best thing to do is to share out the content :)
If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):
SubscribeStar: https://www.subscribestar.com/yannick...
Patreon: https://www.patreon.com/yannickilcher
Follow Saynam here:
YouTube: https://www.youtube.com/c/ChaiTimeDat...
Twitter: https://twitter.com/bhutanisanyam1
Apple Podcasts: https://podcasts.apple.com/us/podcast...
LinkedIn: https://www.linkedin.com/in/sanyambhu...
Spotify: https://open.spotify.com/show/7IbEWJj...
Anchor.fm RSS: https://anchor.fm/s/c19772c/podcast/rss
Outline:
0:00 - Intro & Overview
1:30 - Amazon's MMO may destroy gaming GPUs
2:40 - OpenAI pivots away from Robotics
3:35 - Google parent Alphabet launches Intrinsic
4:55 - AI learns how vegetables taste
5:55 - NASA uses AI to better understand the sun
6:50 - Man used AI to bring back deceased fiancee
7:45 - Robot collision sparks warehouse fire
8:20 - AI deduces patients' racial identities from medical records
9:40 - AlphaFold protein structure database
10:15 - ICCV BEHAVIOR challenge
11:05 - IBM, MIT, Harvard release Common Sense database
11:35 - High quality image generation using diffusion models
12:50 - Conclusion
References:
1 Amazon’s new MMO may be bricking Nvidia 3090s
https://www.theverge.com/2021/7/21/22...
https://www.youtube.com/watch?v=KLyNF...
2 Open AI pivotes from Robots
https://venturebeat.com/2021/07/23/ai...
3 Google parent Alphabet launches Intrinsic: a new company to build software for industrial robots
https://www.theverge.com/2021/7/23/22...
Introducing Intrinsic
https://blog.x.company/introducing-in...
https://x.company/projects/intrinsic/
https://www.forbes.com/sites/jennifer...
4 Artificial Intelligence Helps Improve NASA’s Eyes on the Sun
https://www.nasa.gov/feature/goddard/...
5 A man used AI to bring back his deceased fiancé. But the creators of the tech warn it could be dangerous
https://www.businessinsider.co.za/man...
6 Robot collision at Ocado warehouse near London sparks fire, delaying customer orders https://www.theverge.com/2021/7/18/22...
10 Reading Race: AI Recognizes Patient’s Racial Identity In Medical Images
https://arxiv.org/pdf/2107.10356.pdf
11 AlphaFold Protein Structure Database
https://alphafold.ebi.ac.uk
https://www.theverge.com/2021/7/22/22...
12 Behavior Challenge
http://svl.stanford.edu/behavior/chal...
13 Researchers from IBM, MIT and Harvard Announced The Release Of DARPA “Common Sense AI” Dataset Along With Two Machine Learning Models At ICML 2021
https://www.marktechpost.com/2021/07/...
https://www.reddit.com/r/MachineLearn...
14 Google uses diffusion model for image generation
https://www.reddit.com/r/MachineLearn...
https://www.reddit.com/r/MachineLearn...
Links:
TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick
YouTube: https://www.youtube.com/c/yannickilcher
Twitter: https://twitter.com/ykilcher
Discord: https://discord.gg/4H8xxDF
BitChute: https://www.bitchute.com/channel/yann...
Minds: https://www.minds.com/ykilcher
Parler: https://parler.com/profile/YannicKilcher
LinkedIn: https://www.linkedin.com/in/yannic-ki...
BiliBili: https://space.bilibili.com/1824646584
If you want to support me, the best thing to do is to share out the content :)
If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):
SubscribeStar: https://www.subscribestar.com/yannick...
Patreon: https://www.patreon.com/yannickilcher
Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq
A look into the happenings of the Machine Learning world.
OUTLINE:
0:00 - Intro
0:25 - Facebook AI trains rapidly adapting robots
3:05 - Baidu presents autonomous excavator system
4:45 - EleutherAI turns 1
6:05 - Elon Musk says FSD harder than expected
8:10 - AI interview tools still fall short
11:10 - RunwayML AI-powered cloud video editor
11:55 - MineRL BASALT competition to learn from human feedback
13:15 - The Myth of the Expert Reviewer
15:55 - NVIDIA unveils Cambridge-1 supercomputer
17:10 - CLIP art sees rapid improvements
19:00 - AI demystifies boiling
21:20 - AI avatars for easier language learning
23:20 - Outro
References:
Facebook AI trains rapidly adapting robots
https://ai.facebook.com/blog/ai-now-e...
https://ashish-kmr.github.io/rma-legg...
Baidu presents autonomous excavator system
http://research.baidu.com/Blog/index-...
https://www.youtube.com/watch?v=KFcNf...
EleutherAI turns 1
https://blog.eleuther.ai/year-one/
Elon Musk says FSD is harder than expected
https://www.theverge.com/2021/7/5/225...
AI interview tools still fall short
https://www.technologyreview.com/2021...
RunwayML AI-powered cloud video editor
https://runwayml.com/
MineRL BASALT competition to learn from human feedback
https://www.aicrowd.com/challenges/ne...
The Myth of the Expert Reviewer
https://parameterfree.com/2021/07/06/...
NVIDIA unveils Cambridge-1 supercomputer
https://www.nvidia.com/en-us/industri...
https://nvidianews.nvidia.com/news/nv...
CLIP art sees rapid improvements
https://ml.berkeley.edu/blog/posts/cl...
AI demystifies boiling
https://news.mit.edu/2021/infrared-ca...
AI avatars for easier language learning
https://www.forbes.com/sites/petergre...
Links:
TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick
YouTube: https://www.youtube.com/c/yannickilcher
Twitter: https://twitter.com/ykilcher
Discord: https://discord.gg/4H8xxDF
BitChute: https://www.bitchute.com/channel/yann...
Minds: https://www.minds.com/ykilcher
Parler: https://parler.com/profile/YannicKilcher
LinkedIn: https://www.linkedin.com/in/yannic-ki...
BiliBili: https://space.bilibili.com/1824646584
If you want to support me, the best thing to do is to share out the content :)
If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):
SubscribeStar: https://www.subscribestar.com/yannick...
Patreon: https://www.patreon.com/yannickilcher
Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq
Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2
Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m
Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n
GitHub and OpenAI release Copilot, an AI-powered code autocomplete system that can generate entire functions, classes, and modules from mere definitions and docstrings. Copilot was trained on all public GitHub repositories, and this has a lot of people upset about questions on copyright, code licenses, social obligations, and how much you can profit from other people's work. I give my opinions on the issue in relation to copyright law, the GPL license, and terms of service. Further, we discuss the Brickit app to organize your LEGOs, Distill going on a break, and much more.
OUTLINE:
0:00 - Intro
0:20 - GitHub Copilot
6:55 - My opinion on Copilot & Copyright
17:25 - Facebook AI image similarity challenge
18:00 - Brickit app scans your LEGOs and suggests builds
18:40 - Distill journal goes on break
19:50 - Amazon uses algorithms to hire & fire Flex drivers
23:20 - Helpful Libraries: TF Decision Forests, Habitat, Falken, Brax
24:20 - AI-generated papers give science a hard time
References:
GitHub Copilot: AI pair programmer
https://twitter.com/gdb/status/140989...
https://twitter.com/rickhanlonii/stat...
https://copilot.github.com/
https://docs.github.com/en/github/cop...
https://docs.github.com/en/github/sit...
https://tldrlegal.com/license/gnu-gen...
https://www.gnu.org/licenses/gpl-faq....
https://www.legalzoom.com/knowledge/c...
https://en.wikipedia.org/wiki/Derivat...
https://twitter.com/giffmana/status/1...
https://twitter.com/search?q=copilot&...
Facebook AI launches image similarity challenge
https://www.drivendata.org/competitio...
Brickit app sorts your LEGOs
https://brickit.app/?ref=producthunt&...
https://petapixel.com/2021/07/01/bric...
Distill goes on break
https://distill.pub/2021/distill-hiatus/
Amazon uses Algorithms to fire Flex drivers
https://www.engadget.com/amazon-algor...
TensorFlow decision forests
https://blog.tensorflow.org/2021/05/i...
Facebook AI habitat 2.0
https://ai.facebook.com/blog/habitat-...
Google Falken trains game-playing agents
https://ai.googleblog.com/2021/06/qui...
https://github.com/google-research/fa...
Google Brax: differentiable physics simulator
https://github.com/google/brax
https://arxiv.org/pdf/2106.13281.pdf
Fake science is getting faker
https://thenextweb.com/news/fake-scie...
Links:
TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick
YouTube: https://www.youtube.com/c/yannickilcher
Twitter: https://twitter.com/ykilcher
Discord: https://discord.gg/4H8xxDF
BitChute: https://www.bitchute.com/channel/yann...
Minds: https://www.minds.com/ykilcher
Parler: https://parler.com/profile/YannicKilcher
LinkedIn: https://www.linkedin.com/in/yannic-ki...
BiliBili: https://space.bilibili.com/1824646584
If you want to support me, the best thing to do is to share out the content :)
If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):
SubscribeStar: https://www.subscribestar.com/yannick...
Patreon: https://www.patreon.com/yannickilcher
Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq
Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2
Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m
Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n
Tesla is pushing the state-of-the-art in full self-driving, and interestingly, they explicitly switch from having multiple different sensors to a vision-only system. We discuss the highlights of Andrej Karpathy's talk about Tesla's FSD system, how to label petabytes of data, how to sample edge-cases, how to train a neural network that has to work in real-time, and why moving to having only cameras is superior to multi-sensor approaches.
OUTLINE:
0:00 - Intro & Overview
1:55 - Current Auto-Breaking system
3:20 - Full Self-Driving from vision only
4:55 - Auto-Labelling for collecting data
8:45 - How to get diverse data from edge-cases
12:15 - Neural network architecture
16:05 - Tesla's in-house supercomputer
17:00 - Owning the whole pipeline
18:20 - Example results from vision only
23:10 - Conclusion & Comments
Links:
TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick
YouTube: https://www.youtube.com/c/yannickilcher
Twitter: https://twitter.com/ykilcher
Discord: https://discord.gg/4H8xxDF
BitChute: https://www.bitchute.com/channel/yann...
Minds: https://www.minds.com/ykilcher
Parler: https://parler.com/profile/YannicKilcher
LinkedIn: https://www.linkedin.com/in/yannic-ki...
BiliBili: https://space.bilibili.com/1824646584
If you want to support me, the best thing to do is to share out the content :)
If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):
SubscribeStar: https://www.subscribestar.com/yannick...
Patreon: https://www.patreon.com/yannickilcher
Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq
Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2
Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m
Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n
In this week's ML news we look at CVPR's controversial action to ban paper promotions on social media during the review phase, among other things!
OUTLINE:
0:00 - Intro & Overview
0:25 - CVPR bans social media paper discussions
5:10 - WalMart uses AI to suggest substitutions
6:05 - NVIDIA releases Alias-Free GAN
7:30 - Confession Video in Myanmar possibly a DeepFake
8:50 - AI restores Rembrandt painting
10:40 - AI for healthcare not problem-free yet
11:50 - ML interviews book
12:15 - NVIDIA canvas turns sketches into paintings
13:00 - GPU prices down after crypto shock
13:30 - Facebook AI improves shopping experience
14:05 - DeepLab2 released on GitHub
14:35 - Toxic Language Models: Nobody cares
16:55 - Does AI have common sense?
References:
CVPR forbids social media promotion
https://twitter.com/wjscheirer/status...
WalMart uses AI to substitute out-of-stock products
https://www.supermarketnews.com/techn...
NVIDIA releases Alias-Free GAN
https://nvlabs.github.io/alias-free-gan/
Myanmar Politician's confession could be DeepFake
https://www.wired.com/story/opinion-t...
Rembrandt restored using AI
https://www.smithsonianmag.com/smart-...
AI in healthcare still shaky
http://www.greenvillebusinessmag.com/...
https://www.theverge.com/2021/6/22/22...
ML interviews book
https://huyenchip.com/ml-interviews-b...
NVIDIA Canvas Beta available
https://blogs.nvidia.com/blog/2021/06...
GPU prices down as China cracks down on Crypto
https://www.theregister.com/2021/06/2...
Facebook AI's big goal of improving shopping
https://ai.facebook.com/blog/advancin...
GoogleAI releases DeepLab2
https://github.com/google-research/de...
Toxic Language Model: Nobody cares
https://arxiv.org/pdf/2105.03023.pdf
AI has no common sense
https://www.analyticsinsight.net/inca...
https://6b.eleuther.ai/
Links:
TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick
YouTube: https://www.youtube.com/c/yannickilcher
Twitter: https://twitter.com/ykilcher
Discord: https://discord.gg/4H8xxDF
BitChute: https://www.bitchute.com/channel/yann...
Minds: https://www.minds.com/ykilcher
Parler: https://parler.com/profile/YannicKilcher
LinkedIn: https://www.linkedin.com/in/yannic-ki...
BiliBili: https://space.bilibili.com/1824646584
If you want to support me, the best thing to do is to share out the content :)
If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):
SubscribeStar: https://www.subscribestar.com/yannick...
Patreon: https://www.patreon.com/yannickilcher
Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq
Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2
Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m
Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n
Adversarial Examples have long been a fascinating topic for many Machine Learning researchers. How can a tiny perturbation cause the neural network to change its output by so much? While many explanations have been proposed over the years, they all appear to fall short. This paper attempts to comprehensively explain the existence of adversarial examples by proposing a view of the classification landscape, which they call the Dimpled Manifold Model, which says that any classifier will adjust its decision boundary to align with the low-dimensional data manifold, and only slightly bend around the data. This potentially explains many phenomena around adversarial examples. Warning: In this video, I disagree. Remember that I'm not an authority, but simply give my own opinions.
OUTLINE:
0:00 - Intro & Overview
7:30 - The old mental image of Adversarial Examples
11:25 - The new Dimpled Manifold Hypothesis
22:55 - The Stretchy Feature Model
29:05 - Why do DNNs create Dimpled Manifolds?
38:30 - What can be explained with the new model?
1:00:40 - Experimental evidence for the Dimpled Manifold Model
1:10:25 - Is Goodfellow's claim debunked?
1:13:00 - Conclusion & Comments
Paper: https://arxiv.org/abs/2106.10151
My replication code: https://gist.github.com/yk/de8d987c4e...
Goodfellow's Talk: https://youtu.be/CIfsB_EYsVI?t=4280
Abstract:
The extreme fragility of deep neural networks when presented with tiny perturbations in their inputs was independently discovered by several research groups in 2013, but in spite of enormous effort these adversarial examples remained a baffling phenomenon with no clear explanation. In this paper we introduce a new conceptual framework (which we call the Dimpled Manifold Model) which provides a simple explanation for why adversarial examples exist, why their perturbations have such tiny norms, why these perturbations look like random noise, and why a network which was adversarially trained with incorrectly labeled images can still correctly classify test images. In the last part of the paper we describe the results of numerous experiments which strongly support this new model, and in particular our assertion that adversarial perturbations are roughly perpendicular to the low dimensional manifold which contains all the training examples.
Abstract: Adi Shamir, Odelia Melamed, Oriel BenShmuel
Links:
TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick
YouTube: https://www.youtube.com/c/yannickilcher
Twitter: https://twitter.com/ykilcher
Discord: https://discord.gg/4H8xxDF
BitChute: https://www.bitchute.com/channel/yann...
Minds: https://www.minds.com/ykilcher
Parler: https://parler.com/profile/YannicKilcher
LinkedIn: https://www.linkedin.com/in/ykilcher
BiliBili: https://space.bilibili.com/1824646584
If you want to support me, the best thing to do is to share out the content :)
If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):
SubscribeStar: https://www.subscribestar.com/yannick...
Patreon: https://www.patreon.com/yannickilcher
Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq
Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2
Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m
Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n
In this week's ML News, we look at the latest developments in the Machine Learning and AI world with updates from research, industry, and society at large.
OUTLINE:
0:00 - Intro
0:20 - Hugging Face launches free course
1:30 - Sentdex releases GAN Theft Auto
2:25 - Facebook uses AI to help moderators
4:10 - Weather with Antonio
5:10 - Autonomous ship aborts mission
7:25 - PyTorch Release 1.9
8:30 - McDonald's new AI drive thru
10:20 - UBS CEO says AI won't replace humans
12:20 - Gödel paper has 90th birthday
12:55 - AugLy data augmentation library
13:20 - Programming Puzzles for autonomous coding
14:30 - Boston Dynamics' Spot turns 1
References:
PyTorch 1.9 Released
https://pytorch.org/blog/pytorch-1.9-...
Hugging Face launches course
https://huggingface.co/course/chapter1
90 years of Gödel's theory
https://people.idsia.ch/~juergen/goed...
AugLy: A data augmentation library
https://ai.facebook.com/blog/augly-a-...
Sentdex builds GAN Theft Auto
https://github.com/sentdex/GANTheftAuto/
Spot turns 1
https://blog.bostondynamics.com/spots...
Autonomous ship aborts mission
https://www.washingtonpost.com/techno...
https://mas400.com/dashboard#currentL...
McDonald's tests AI drive thru
https://www.zdnet.com/article/i-just-...
Facebook uses AI to moderate conversations
https://edition.cnn.com/2021/06/16/te...
UBS CEO says AI won't replace financial advisors
https://www.cnbc.com/2021/06/17/ai-wo...
Programming Puzzles
https://arxiv.org/abs/2106.05784
https://github.com/microsoft/PythonPr...
Links:
TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick
YouTube: https://www.youtube.com/c/yannickilcher
Twitter: https://twitter.com/ykilcher
Discord: https://discord.gg/4H8xxDF
BitChute: https://www.bitchute.com/channel/yann...
Minds: https://www.minds.com/ykilcher
Parler: https://parler.com/profile/YannicKilcher
LinkedIn: https://www.linkedin.com/in/yannic-ki...
BiliBili: https://space.bilibili.com/1824646584
If you want to support me, the best thing to do is to share out the content :)
If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):
SubscribeStar: https://www.subscribestar.com/yannick...
Patreon: https://www.patreon.com/yannickilcher
Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq
Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2
Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m
Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n
After dominating Natural Language Processing, Transformers have taken over Computer Vision recently with the advent of Vision Transformers. However, the attention mechanism's quadratic complexity in the number of tokens means that Transformers do not scale well to high-resolution images. XCiT is a new Transformer architecture, containing XCA, a transposed version of attention, reducing the complexity from quadratic to linear, and at least on image data, it appears to perform on par with other models. What does this mean for the field? Is this even a transformer? What really matters in deep learning?
OUTLINE:
0:00 - Intro & Overview
3:45 - Self-Attention vs Cross-Covariance Attention (XCA)
19:55 - Cross-Covariance Image Transformer (XCiT) Architecture
26:00 - Theoretical & Engineering considerations
30:40 - Experimental Results
33:20 - Comments & Conclusion
Paper: https://arxiv.org/abs/2106.09681
Code: https://github.com/facebookresearch/xcit
Abstract:
Following their success in natural language processing, transformers have recently shown much promise for computer vision. The self-attention operation underlying transformers yields global interactions between all tokens ,i.e. words or image patches, and enables flexible modelling of image data beyond the local interactions of convolutions. This flexibility, however, comes with a quadratic complexity in time and memory, hindering application to long sequences and high-resolution images. We propose a "transposed" version of self-attention that operates across feature channels rather than tokens, where the interactions are based on the cross-covariance matrix between keys and queries. The resulting cross-covariance attention (XCA) has linear complexity in the number of tokens, and allows efficient processing of high-resolution images. Our cross-covariance image transformer (XCiT) is built upon XCA. It combines the accuracy of conventional transformers with the scalability of convolutional architectures. We validate the effectiveness and generality of XCiT by reporting excellent results on multiple vision benchmarks, including image classification and self-supervised feature learning on ImageNet-1k, object detection and instance segmentation on COCO, and semantic segmentation on ADE20k.
Authors: Alaaeldin El-Nouby, Hugo Touvron, Mathilde Caron, Piotr Bojanowski, Matthijs Douze, Armand Joulin, Ivan Laptev, Natalia Neverova, Gabriel Synnaeve, Jakob Verbeek, Hervé Jegou
Links:
TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick
YouTube: https://www.youtube.com/c/yannickilcher
Twitter: https://twitter.com/ykilcher
Discord: https://discord.gg/4H8xxDF
BitChute: https://www.bitchute.com/channel/yann...
Minds: https://www.minds.com/ykilcher
Parler: https://parler.com/profile/YannicKilcher
LinkedIn: https://www.linkedin.com/in/yannic-ki...
BiliBili: https://space.bilibili.com/1824646584
If you want to support me, the best thing to do is to share out the content :)
If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):
SubscribeStar: https://www.subscribestar.com/yannick...
Patreon: https://www.patreon.com/yannickilcher
Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq
Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2
Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m
Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n
Learning from demonstrations is a fascinating topic, but what if the demonstrations are not exactly the behaviors we want to learn? Can we adhere to a dataset of demonstrations and still achieve a specified goal? This paper uses GANs to combine goal-achieving reinforcement learning with imitation learning and learns to perform well at a given task while doing so in the style of a given presented dataset. The resulting behaviors include many realistic-looking transitions between the demonstrated movements.
OUTLINE:
0:00 - Intro & Overview
1:25 - Problem Statement
6:10 - Reward Signals
8:15 - Motion Prior from GAN
14:10 - Algorithm Overview
20:15 - Reward Engineering & Experimental Results
30:40 - Conclusion & Comments
Paper: https://arxiv.org/abs/2104.02180
Main Video: https://www.youtube.com/watch?v=wySUx...
Supplementary Video: https://www.youtube.com/watch?v=O6fBS...
Abstract:
Synthesizing graceful and life-like behaviors for physically simulated characters has been a fundamental challenge in computer animation. Data-driven methods that leverage motion tracking are a prominent class of techniques for producing high fidelity motions for a wide range of behaviors. However, the effectiveness of these tracking-based methods often hinges on carefully designed objective functions, and when applied to large and diverse motion datasets, these methods require significant additional machinery to select the appropriate motion for the character to track in a given scenario. In this work, we propose to obviate the need to manually design imitation objectives and mechanisms for motion selection by utilizing a fully automated approach based on adversarial imitation learning. High-level task objectives that the character should perform can be specified by relatively simple reward functions, while the low-level style of the character's behaviors can be specified by a dataset of unstructured motion clips, without any explicit clip selection or sequencing. These motion clips are used to train an adversarial motion prior, which specifies style-rewards for training the character through reinforcement learning (RL). The adversarial RL procedure automatically selects which motion to perform, dynamically interpolating and generalizing from the dataset. Our system produces high-quality motions that are comparable to those achieved by state-of-the-art tracking-based techniques, while also being able to easily accommodate large datasets of unstructured motion clips. Composition of disparate skills emerges automatically from the motion prior, without requiring a high-level motion planner or other task-specific annotations of the motion clips. We demonstrate the effectiveness of our framework on a diverse cast of complex simulated characters and a challenging suite of motor control tasks.
Authors: Xue Bin Peng, Ze Ma, Pieter Abbeel, Sergey Levine, Angjoo Kanazawa
Links:
TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick
YouTube: https://www.youtube.com/c/yannickilcher
Twitter: https://twitter.com/ykilcher
Discord: https://discord.gg/4H8xxDF
BitChute: https://www.bitchute.com/channel/yann...
Minds: https://www.minds.com/ykilcher
Parler: https://parler.com/profile/YannicKilcher
LinkedIn: https://www.linkedin.com/in/ykilcher
BiliBili: https://space.bilibili.com/1824646584
If you want to support me, the best thing to do is to share out the content :)
If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):
SubscribeStar: https://www.subscribestar.com/yannick...
Patreon: https://www.patreon.com/yannickilcher
OUTLINE:
0:00 - Intro
0:30 - Google RL creates next-gen TPUs
2:15 - Facebook launches NetHack challenge
3:50 - OpenAI mitigates bias by fine-tuning
9:05 - Google AI releases browseable reconstruction of human cortex
9:50 - GPT-J 6B Transformer in JAX
12:00 - Tensorflow launches Forum
13:50 - Text style transfer from a single word
15:45 - ALiEn artificial life simulator
My Video on Chip Placement: https://youtu.be/PDRtyrVskMU
References:
RL creates next-gen TPUs
https://www.nature.com/articles/s4158...
https://www.youtube.com/watch?v=PDRty...
Facebook launches NetHack challenge
https://ai.facebook.com/blog/launchin...
Mitigating bias by fine-tuning
https://openai.com/blog/improving-lan...
Human Cortex 3D Reconstruction
https://ai.googleblog.com/2021/06/a-b...
GPT-J: An open-source 6B transformer
https://arankomatsuzaki.wordpress.com...
https://6b.eleuther.ai/
https://github.com/kingoflolz/mesh-tr...
Tensorflow launches "Forum"
https://discuss.tensorflow.org/
Text style transfer from single word
https://ai.facebook.com/blog/ai-can-n...
ALiEn Life Simulator
https://github.com/chrxh/alien
Links:
TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick
YouTube: https://www.youtube.com/c/yannickilcher
Twitter: https://twitter.com/ykilcher
Discord: https://discord.gg/4H8xxDF
BitChute: https://www.bitchute.com/channel/yann...
Minds: https://www.minds.com/ykilcher
Parler: https://parler.com/profile/YannicKilcher
LinkedIn: https://www.linkedin.com/in/ykilcher
BiliBili: https://space.bilibili.com/1824646584
If you want to support me, the best thing to do is to share out the content :)
If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):
SubscribeStar: https://www.subscribestar.com/yannick...
Patreon: https://www.patreon.com/yannickilcher
Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq
Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2
Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m
Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n
Many problems in Machine Learning involve loops of inner and outer optimization. Finding update steps for the outer loop is usually difficult, because of the.need to differentiate through the inner loop's procedure over multiple steps. Such loop unrolling is very limited and constrained to very few steps. Other papers have found solutions around unrolling in very specific, individual problems. This paper proposes a unified framework for implicit differentiation of inner optimization procedures without unrolling and provides implementations that integrate seamlessly into JAX.
OUTLINE:
0:00 - Intro & Overview
2:05 - Automatic Differentiation of Inner Optimizations
4:30 - Example: Meta-Learning
7:45 - Unrolling Optimization
13:00 - Unified Framework Overview & Pseudocode
21:10 - Implicit Function Theorem
25:45 - More Technicalities
28:45 - Experiments
ERRATA:
Paper: https://arxiv.org/abs/2105.15183
Code coming soon
Abstract:
Automatic differentiation (autodiff) has revolutionized machine learning. It allows expressing complex computations by composing elementary ones in creative ways and removes the burden of computing their derivatives by hand. More recently, differentiation of optimization problem solutions has attracted widespread attention with applications such as optimization as a layer, and in bi-level problems such as hyper-parameter optimization and meta-learning. However, the formulas for these derivatives often involve case-by-case tedious mathematical derivations. In this paper, we propose a unified, efficient and modular approach for implicit differentiation of optimization problems. In our approach, the user defines (in Python in the case of our implementation) a function F capturing the optimality conditions of the problem to be differentiated. Once this is done, we leverage autodiff of F and implicit differentiation to automatically differentiate the optimization problem. Our approach thus combines the benefits of implicit differentiation and autodiff. It is efficient as it can be added on top of any state-of-the-art solver and modular as the optimality condition specification is decoupled from the implicit differentiation mechanism. We show that seemingly simple principles allow to recover many recently proposed implicit differentiation methods and create new ones easily. We demonstrate the ease of formulating and solving bi-level optimization problems using our framework. We also showcase an application to the sensitivity analysis of molecular dynamics.
Authors: Mathieu Blondel, Quentin Berthet, Marco Cuturi, Roy Frostig, Stephan Hoyer, Felipe Llinares-López, Fabian Pedregosa, Jean-Philippe Vert
Links:
TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick
YouTube: https://www.youtube.com/c/yannickilcher
Twitter: https://twitter.com/ykilcher
Discord: https://discord.gg/4H8xxDF
BitChute: https://www.bitchute.com/channel/yann...
Minds: https://www.minds.com/ykilcher
Parler: https://parler.com/profile/YannicKilcher
LinkedIn: https://www.linkedin.com/in/yannic-ki...
BiliBili: https://space.bilibili.com/1824646584
If you want to support me, the best thing to do is to share out the content :)
If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):
SubscribeStar: https://www.subscribestar.com/yannick...
Patreon: https://www.patreon.com/yannickilcher
Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq
OUTLINE:
0:00 - Intro
0:25 - EU seeks to regulate AI
2:45 - AI COVID detection systems are all flawed
5:05 - Chinese lab trains model 10x GPT-3 size
6:55 - Google error identifies "ugliest" language
9:45 - McDonald's learns about AI buzzwords
11:25 - AI predicts cryptocurrency prices
12:00 - Unreal Engine hack for CLIP
12:35 - Please commit more academic fraud
References:
https://www.lawfareblog.com/artificia...
https://blogs.sciencemag.org/pipeline...
https://www.nature.com/articles/s4225...
https://en.pingwest.com/a/8693
https://arxiv.org/pdf/2104.12369.pdf
https://www.bbc.com/news/world-asia-i...
https://www.zdnet.com/article/mcdonal...
https://www.analyticsinsight.net/ai-i...
https://twitter.com/arankomatsuzaki/s...
https://jacobbuckman.com/2021-05-29-p...
Links:
TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick
YouTube: https://www.youtube.com/c/yannickilcher
Twitter: https://twitter.com/ykilcher
Discord: https://discord.gg/4H8xxDF
BitChute: https://www.bitchute.com/channel/yann...
Minds: https://www.minds.com/ykilcher
Parler: https://parler.com/profile/YannicKilcher
LinkedIn: https://www.linkedin.com/in/yannic-ki...
BiliBili: https://space.bilibili.com/1824646584
If you want to support me, the best thing to do is to share out the content :)
If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):
SubscribeStar: https://www.subscribestar.com/yannick...
Patreon: https://www.patreon.com/yannickilcher
Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq
Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2
Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m
Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n
Proper credit assignment over long timespans is a fundamental problem in reinforcement learning. Even methods designed to combat this problem, such as TD-learning, quickly reach their limits when rewards are sparse or noisy. This paper reframes offline reinforcement learning as a pure sequence modeling problem, with the actions being sampled conditioned on the given history and desired future rewards. This allows the authors to use recent advances in sequence modeling using Transformers and achieve competitive results in Offline RL benchmarks.
OUTLINE:
0:00 - Intro & Overview
4:15 - Offline Reinforcement Learning
10:10 - Transformers in RL
14:25 - Value Functions and Temporal Difference Learning
20:25 - Sequence Modeling and Reward-to-go
27:20 - Why this is ideal for offline RL
31:30 - The context length problem
34:35 - Toy example: Shortest path from random walks
41:00 - Discount factors
45:50 - Experimental Results
49:25 - Do you need to know the best possible reward?
52:15 - Key-to-door toy experiment
56:00 - Comments & Conclusion
Paper: https://arxiv.org/abs/2106.01345
Website: https://sites.google.com/berkeley.edu...
Code: https://github.com/kzl/decision-trans...
Abstract:
We present a framework that abstracts Reinforcement Learning (RL) as a sequence modeling problem. This allows us to draw upon the simplicity and scalability of the Transformer architecture, and associated advances in language modeling such as GPT-x and BERT. In particular, we present Decision Transformer, an architecture that casts the problem of RL as conditional sequence modeling. Unlike prior approaches to RL that fit value functions or compute policy gradients, Decision Transformer simply outputs the optimal actions by leveraging a causally masked Transformer. By conditioning an autoregressive model on the desired return (reward), past states, and actions, our Decision Transformer model can generate future actions that achieve the desired return. Despite its simplicity, Decision Transformer matches or exceeds the performance of state-of-the-art model-free offline RL baselines on Atari, OpenAI Gym, and Key-to-Door tasks.
Authors: Lili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee, Aditya Grover, Michael Laskin, Pieter Abbeel, Aravind Srinivas, Igor Mordatch
Links:
TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick
YouTube: https://www.youtube.com/c/yannickilcher
Twitter: https://twitter.com/ykilcher
Discord: https://discord.gg/4H8xxDF
BitChute: https://www.bitchute.com/channel/yann...
Minds: https://www.minds.com/ykilcher
Parler: https://parler.com/profile/YannicKilcher
LinkedIn: https://www.linkedin.com/in/yannic-ki...
BiliBili: https://space.bilibili.com/1824646584
If you want to support me, the best thing to do is to share out the content :)
If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):
SubscribeStar: https://www.subscribestar.com/yannick...
Patreon: https://www.patreon.com/yannickilcher
Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq
Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2
Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m
Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n
Anthropic raises $124M for steerable AI, peer review is threatened by collusion rings, and the original ELIZA source code was discovered.
OUTLINE:
0:00 - Intro
0:40 - Anthropic raises $124M
3:25 - 65% of execs can't explain AI predictions
4:25 - DeepMind releases AndroidEnv
6:10 - Collusion rings in ML Conferences
7:30 - ELIZA's original source code discovered
10:45 - OpenAI raises $100M fund
11:25 - Outro
References:
https://techcrunch.com/2021/05/28/ant...
https://www.anthropic.com/news/announ...
https://www.anthropic.com/
https://openai.com/blog/introducing-o...
https://deepmind.com/research/publica...
https://cacm.acm.org/magazines/2021/6...
https://venturebeat.com/2021/05/25/65...
https://techcrunch.com/2021/05/26/ope...
https://sites.google.com/view/elizage...
http://psych.fullerton.edu/mbirnbaum/...
https://en.wikipedia.org/wiki/Carl_Ro...
https://openai.com/fund/
Links:
TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick
YouTube: https://www.youtube.com/c/yannickilcher
Twitter: https://twitter.com/ykilcher
Discord: https://discord.gg/4H8xxDF
BitChute: https://www.bitchute.com/channel/yann...
Minds: https://www.minds.com/ykilcher
Parler: https://parler.com/profile/YannicKilcher
LinkedIn: https://www.linkedin.com/in/yannic-ki...
BiliBili: https://space.bilibili.com/1824646584
If you want to support me, the best thing to do is to share out the content :)
If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):
SubscribeStar: https://www.subscribestar.com/yannick...
Patreon: https://www.patreon.com/yannickilcher
Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq
Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2
Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m
Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n
What's the most promising path to creating Artificial General Intelligence (AGI)? This paper makes the bold claim that a learning agent maximizing its reward in a sufficiently complex environment will necessarily develop intelligence as a by-product, and that Reward Maximization is the best way to move the creation of AGI forward. The paper is a mix of philosophy, engineering, and futurism, and raises many points of discussion.
OUTLINE:
0:00 - Intro & Outline
4:10 - Reward Maximization
10:10 - The Reward-is-Enough Hypothesis
13:15 - Abilities associated with intelligence
16:40 - My Criticism
26:15 - Reward Maximization through Reinforcement Learning
31:30 - Discussion, Conclusion & My Comments
Paper: https://www.sciencedirect.com/science...
Abstract:
In this article we hypothesise that intelligence, and its associated abilities, can be understood as subserving the maximisation of reward. Accordingly, reward is enough to drive behaviour that exhibits abilities studied in natural and artificial intelligence, including knowledge, learning, perception, social intelligence, language, generalisation and imitation. This is in contrast to the view that specialised problem formulations are needed for each ability, based on other signals or objectives. Furthermore, we suggest that agents that learn through trial and error experience to maximise reward could learn behaviour that exhibits most if not all of these abilities, and therefore that powerful reinforcement learning agents could constitute a solution to artificial general intelligence.
Authors: David Silver, Satinder Singh, Doina Precup, Richard S. Sutton
Links:
TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick
YouTube: https://www.youtube.com/c/yannickilcher
Twitter: https://twitter.com/ykilcher
Discord: https://discord.gg/4H8xxDF
BitChute: https://www.bitchute.com/channel/yann...
Minds: https://www.minds.com/ykilcher
Parler: https://parler.com/profile/YannicKilcher
LinkedIn: https://www.linkedin.com/in/yannic-ki...
BiliBili: https://space.bilibili.com/1824646584
If you want to support me, the best thing to do is to share out the content :)
If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):
SubscribeStar: https://www.subscribestar.com/yannick...
Patreon: https://www.patreon.com/yannickilcher
Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq
Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2
Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m
Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n
Facebook AI (FAIR) researchers present Expire-Span, a variant of Transformer XL that dynamically assigns expiration dates to previously encountered signals. Because of this, Expire-Span can handle sequences of many thousand tokens, while keeping the memory and compute requirements at a manageable level. It severely matches or outperforms baseline systems, while consuming much less resources. We discuss its architecture, advantages, and shortcomings.
OUTLINE:
0:00 - Intro & Overview
2:30 - Remembering the past in sequence models
5:45 - Learning to expire past memories
8:30 - Difference to local attention
10:00 - Architecture overview
13:45 - Comparison to Transformer XL
18:50 - Predicting expiration masks
32:30 - Experimental Results
40:00 - Conclusion & Comments
Paper: https://arxiv.org/abs/2105.06548
Code: https://github.com/facebookresearch/t...
ADDENDUM: I mention several times that the gradient signal of the e quantity only occurs inside the R ramp. By that, I mean the gradient stemming from the model loss. The regularization loss acts also outside the R ramp.
Abstract:
Attention mechanisms have shown promising results in sequence modeling tasks that require long-term memory. Recent work investigated mechanisms to reduce the computational cost of preserving and storing memories. However, not all content in the past is equally important to remember. We propose Expire-Span, a method that learns to retain the most important information and expire the irrelevant information. This forgetting of memories enables Transformers to scale to attend over tens of thousands of previous timesteps efficiently, as not all states from previous timesteps are preserved. We demonstrate that Expire-Span can help models identify and retain critical information and show it can achieve strong performance on reinforcement learning tasks specifically designed to challenge this functionality. Next, we show that Expire-Span can scale to memories that are tens of thousands in size, setting a new state of the art on incredibly long context tasks such as character-level language modeling and a frame-by-frame moving objects task. Finally, we analyze the efficiency of Expire-Span compared to existing approaches and demonstrate that it trains faster and uses less memory.
Authors: Sainbayar Sukhbaatar, Da Ju, Spencer Poff, Stephen Roller, Arthur Szlam, Jason Weston, Angela Fan
Links:
TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick
YouTube: https://www.youtube.com/c/yannickilcher
Twitter: https://twitter.com/ykilcher
Discord: https://discord.gg/4H8xxDF
BitChute: https://www.bitchute.com/channel/yann...
Minds: https://www.minds.com/ykilcher
Parler: https://parler.com/profile/YannicKilcher
LinkedIn: https://www.linkedin.com/in/yannic-ki...
BiliBili: https://space.bilibili.com/1824646584
If you want to support me, the best thing to do is to share out the content :)
If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):
SubscribeStar: https://www.subscribestar.com/yannick...
Patreon: https://www.patreon.com/yannickilcher
Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq
Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2
Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m
Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n
Do we even need Attention? FNets completely drop the Attention mechanism in favor of a simple Fourier transform. They perform almost as well as Transformers, while drastically reducing parameter count, as well as compute and memory requirements. This highlights that a good token mixing heuristic could be as valuable as a learned attention matrix.
OUTLINE:
0:00 - Intro & Overview
0:45 - Giving up on Attention
5:00 - FNet Architecture
9:00 - Going deeper into the Fourier Transform
11:20 - The Importance of Mixing
22:20 - Experimental Results
33:00 - Conclusions & Comments
Paper: https://arxiv.org/abs/2105.03824
ADDENDUM:
Of course, I completely forgot to discuss the connection between Fourier transforms and Convolutions, and that this might be interpreted as convolutions with very large kernels.
Abstract:
We show that Transformer encoder architectures can be massively sped up, with limited accuracy costs, by replacing the self-attention sublayers with simple linear transformations that "mix" input tokens. These linear transformations, along with simple nonlinearities in feed-forward layers, are sufficient to model semantic relationships in several text classification tasks. Perhaps most surprisingly, we find that replacing the self-attention sublayer in a Transformer encoder with a standard, unparameterized Fourier Transform achieves 92% of the accuracy of BERT on the GLUE benchmark, but pre-trains and runs up to seven times faster on GPUs and twice as fast on TPUs. The resulting model, which we name FNet, scales very efficiently to long inputs, matching the accuracy of the most accurate "efficient" Transformers on the Long Range Arena benchmark, but training and running faster across all sequence lengths on GPUs and relatively shorter sequence lengths on TPUs. Finally, FNet has a light memory footprint and is particularly efficient at smaller model sizes: for a fixed speed and accuracy budget, small FNet models outperform Transformer counterparts.
Authors: James Lee-Thorp, Joshua Ainslie, Ilya Eckstein, Santiago Ontanon
Links:
TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick
YouTube: https://www.youtube.com/c/yannickilcher
Twitter: https://twitter.com/ykilcher
Discord: https://discord.gg/4H8xxDF
BitChute: https://www.bitchute.com/channel/yann...
Minds: https://www.minds.com/ykilcher
Parler: https://parler.com/profile/YannicKilcher
LinkedIn: https://www.linkedin.com/in/yannic-ki...
BiliBili: https://space.bilibili.com/1824646584
If you want to support me, the best thing to do is to share out the content :)
If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):
SubscribeStar: https://www.subscribestar.com/yannick...
Patreon: https://www.patreon.com/yannickilcher
Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq
Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2
Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m
Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n
I used OpenAI's CLIP model and BigGAN to create a music video that goes along with the lyrics of a song that I wrote. The song lyrics are made from ImageNet class labels, and the song itself is performed by me on a looper.
OUTLINE:
0:00 - Intro
1:00 - AI-generated music video for "be my weasel"
3:50 - How it was made
7:30 - My looping gear
9:35 - AI-generated music video #2
12:45 - Outro & Credits
Code and references: https://github.com/yk/clip_music_video
Links:
TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick
YouTube: https://www.youtube.com/c/yannickilcher
Twitter: https://twitter.com/ykilcher
Discord: https://discord.gg/4H8xxDF
BitChute: https://www.bitchute.com/channel/yann...
Minds: https://www.minds.com/ykilcher
Parler: https://parler.com/profile/YannicKilcher
LinkedIn: https://www.linkedin.com/in/yannic-ki...
BiliBili: https://space.bilibili.com/1824646584
If you want to support me, the best thing to do is to share out the content :)
If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):
SubscribeStar: https://www.subscribestar.com/yannick...
Patreon: https://www.patreon.com/yannickilcher
Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq
Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2
Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m
Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n
GANs have dominated the image generation space for the majority of the last decade. This paper shows for the first time, how a non-GAN model, a DDPM, can be improved to overtake GANs at standard evaluation metrics for image generation. The produced samples look amazing and other than GANs, the new model has a formal probabilistic foundation. Is there a future for GANs or are Diffusion Models going to overtake them for good?
OUTLINE:
0:00 - Intro & Overview
4:10 - Denoising Diffusion Probabilistic Models
11:30 - Formal derivation of the training loss
23:00 - Training in practice
27:55 - Learning the covariance
31:25 - Improving the noise schedule
33:35 - Reducing the loss gradient noise
40:35 - Classifier guidance
52:50 - Experimental Results
Paper (this): https://arxiv.org/abs/2105.05233
Paper (previous): https://arxiv.org/abs/2102.09672
Code: https://github.com/openai/guided-diff...
Abstract:
We show that diffusion models can achieve image sample quality superior to the current state-of-the-art generative models. We achieve this on unconditional image synthesis by finding a better architecture through a series of ablations. For conditional image synthesis, we further improve sample quality with classifier guidance: a simple, compute-efficient method for trading off diversity for sample quality using gradients from a classifier. We achieve an FID of 2.97 on ImageNet 128×128, 4.59 on ImageNet 256×256, and 7.72 on ImageNet 512×512, and we match BigGAN-deep even with as few as 25 forward passes per sample, all while maintaining better coverage of the distribution. Finally, we find that classifier guidance combines well with upsampling diffusion models, further improving FID to 3.85 on ImageNet 512×512. We release our code at this https URL
Authors: Alex Nichol, Prafulla Dhariwal
Links:
TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick
YouTube: https://www.youtube.com/c/yannickilcher
Twitter: https://twitter.com/ykilcher
Discord: https://discord.gg/4H8xxDF
BitChute: https://www.bitchute.com/channel/yann...
Minds: https://www.minds.com/ykilcher
Parler: https://parler.com/profile/YannicKilcher
LinkedIn: https://www.linkedin.com/in/yannic-ki...
BiliBili: https://space.bilibili.com/1824646584
If you want to support me, the best thing to do is to share out the content :)
If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):
SubscribeStar: https://www.subscribestar.com/yannick...
Patreon: https://www.patreon.com/yannickilcher
Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq
Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2
Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m
Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n
Convolutional Neural Networks (CNNs) have dominated computer vision for almost a decade by applying two fundamental principles: Spatial agnosticism and channel-specific computations. Involution aims to invert these principles and presents a spatial-specific computation, which is also channel-agnostic. The resulting Involution Operator and RedNet architecture are a compromise between classic Convolutions and the newer Local Self-Attention architectures and perform favorably in terms of computation accuracy tradeoff when compared to either.
OUTLINE:
0:00 - Intro & Overview
3:00 - Principles of Convolution
10:50 - Towards spatial-specific computations
17:00 - The Involution Operator
20:00 - Comparison to Self-Attention
25:15 - Experimental Results
30:30 - Comments & Conclusion
Paper: https://arxiv.org/abs/2103.06255
Code: https://github.com/d-li14/involution
Abstract:
Convolution has been the core ingredient of modern neural networks, triggering the surge of deep learning in vision. In this work, we rethink the inherent principles of standard convolution for vision tasks, specifically spatial-agnostic and channel-specific. Instead, we present a novel atomic operation for deep neural networks by inverting the aforementioned design principles of convolution, coined as involution. We additionally demystify the recent popular self-attention operator and subsume it into our involution family as an over-complicated instantiation. The proposed involution operator could be leveraged as fundamental bricks to build the new generation of neural networks for visual recognition, powering different deep learning models on several prevalent benchmarks, including ImageNet classification, COCO detection and segmentation, together with Cityscapes segmentation. Our involution-based models improve the performance of convolutional baselines using ResNet-50 by up to 1.6% top-1 accuracy, 2.5% and 2.4% bounding box AP, and 4.7% mean IoU absolutely while compressing the computational cost to 66%, 65%, 72%, and 57% on the above benchmarks, respectively. Code and pre-trained models for all the tasks are available at this https URL.
Authors: Duo Li, Jie Hu, Changhu Wang, Xiangtai Li, Qi She, Lei Zhu, Tong Zhang, Qifeng Chen
Links:
TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick
YouTube: https://www.youtube.com/c/yannickilcher
Twitter: https://twitter.com/ykilcher
Discord: https://discord.gg/4H8xxDF
BitChute: https://www.bitchute.com/channel/yann...
Minds: https://www.minds.com/ykilcher
Parler: https://parler.com/profile/YannicKilcher
LinkedIn: https://www.linkedin.com/in/yannic-ki...
BiliBili: https://space.bilibili.com/1824646584
If you want to support me, the best thing to do is to share out the content :)
If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):
SubscribeStar: https://www.subscribestar.com/yannick...
Patreon: https://www.patreon.com/yannickilcher
Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq
Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2
Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m
Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n
Convolutional Neural Networks have dominated computer vision for nearly 10 years, and that might finally come to an end. First, Vision Transformers (ViT) have shown remarkable performance, and now even simple MLP-based models reach competitive accuracy, as long as sufficient data is used for pre-training. This paper presents MLP-Mixer, using MLPs in a particular weight-sharing arrangement to achieve a competitive, high-throughput model and it raises some interesting questions about the nature of learning and inductive biases and their interaction with scale for future research.
OUTLINE:
0:00 - Intro & Overview
2:20 - MLP-Mixer Architecture
13:20 - Experimental Results
17:30 - Effects of Scale
24:30 - Learned Weights Visualization
27:25 - Comments & Conclusion
Paper: https://arxiv.org/abs/2105.01601
Abstract:
Convolutional Neural Networks (CNNs) are the go-to model for computer vision. Recently, attention-based networks, such as the Vision Transformer, have also become popular. In this paper we show that while convolutions and attention are both sufficient for good performance, neither of them are necessary. We present MLP-Mixer, an architecture based exclusively on multi-layer perceptrons (MLPs). MLP-Mixer contains two types of layers: one with MLPs applied independently to image patches (i.e. "mixing" the per-location features), and one with MLPs applied across patches (i.e. "mixing" spatial information). When trained on large datasets, or with modern regularization schemes, MLP-Mixer attains competitive scores on image classification benchmarks, with pre-training and inference cost comparable to state-of-the-art models. We hope that these results spark further research beyond the realms of well established CNNs and Transformers.
Authors: Ilya Tolstikhin, Neil Houlsby, Alexander Kolesnikov, Lucas Beyer, Xiaohua Zhai, Thomas Unterthiner, Jessica Yung, Daniel Keysers, Jakob Uszkoreit, Mario Lucic, Alexey Dosovitskiy
ERRATA: Here is their definition of what the 5-shot classifier is: "we report the few-shot accuracies obtained by solving the L2-regularized linear regression problem between the frozen learned representations of images and the labels"
Links:
TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick
YouTube: https://www.youtube.com/c/yannickilcher
Twitter: https://twitter.com/ykilcher
Discord: https://discord.gg/4H8xxDF
BitChute: https://www.bitchute.com/channel/yann...
Minds: https://www.minds.com/ykilcher
Parler: https://parler.com/profile/YannicKilcher
LinkedIn: https://www.linkedin.com/in/yannic-ki...
BiliBili: https://space.bilibili.com/1824646584
If you want to support me, the best thing to do is to share out the content :)
If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):
SubscribeStar: https://www.subscribestar.com/yannick...
Patreon: https://www.patreon.com/yannickilcher
Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq
Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2
Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m
Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n
A brief look into gender stereotypes in Google Translate. The origin is a Tweet containing a Hungarian text. Hungarian is a gender-neutral language, so translating gender pronouns is ambiguous. Turns out that Google Translate assigns very stereotypical pronouns. In this video, we'll have a look at the origins and possible solutions to this problem.
OUTLINE:
0:00 - Intro
1:10 - Digging Deeper
2:30 - How does Machine Translation work?
3:50 - Training Data Problems
4:40 - Learning Algorithm Problems
5:45 - Argmax Output Problems
6:45 - Pragmatics
7:50 - More on Google Translate
9:40 - Social Engineering
11:15 - Conclusion
Songs:
Like That - Anno Domini Beats
Submarine - Dyalla
Dude - Patrick Patrikios
Links:
TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick
YouTube: https://www.youtube.com/c/yannickilcher
Twitter: https://twitter.com/ykilcher
Discord: https://discord.gg/4H8xxDF
BitChute: https://www.bitchute.com/channel/yann...
Minds: https://www.minds.com/ykilcher
Parler: https://parler.com/profile/YannicKilcher
LinkedIn: https://www.linkedin.com/in/yannic-ki...
BiliBili: https://space.bilibili.com/1824646584
If you want to support me, the best thing to do is to share out the content :)
If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):
SubscribeStar: https://www.subscribestar.com/yannick...
Patreon: https://www.patreon.com/yannickilcher
Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq
Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2
Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m
Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n
Inspired by the fact that biological creatures attend to multiple modalities at the same time, DeepMind releases its new Perceiver model. Based on the Transformer architecture, the Perceiver makes no assumptions on the modality of the input data and also solves the long-standing quadratic bottleneck problem. This is achieved by having a latent low-dimensional Transformer, where the input data is fed multiple times via cross-attention. The Perceiver's weights can also be shared across layers, making it very similar to an RNN. Perceivers achieve competitive performance on ImageNet and state-of-the-art on other modalities, all while making no architectural adjustments to input data.
OUTLINE:
0:00 - Intro & Overview
2:20 - Built-In assumptions of Computer Vision Models
5:10 - The Quadratic Bottleneck of Transformers
8:00 - Cross-Attention in Transformers
10:45 - The Perceiver Model Architecture & Learned Queries
20:05 - Positional Encodings via Fourier Features
23:25 - Experimental Results & Attention Maps
29:05 - Comments & Conclusion
Paper: https://arxiv.org/abs/2103.03206
My Video on Transformers (Attention is All You Need): https://youtu.be/iDulhoQ2pro
Abstract:
Biological systems understand the world by simultaneously processing high-dimensional inputs from modalities as diverse as vision, audition, touch, proprioception, etc. The perception models used in deep learning on the other hand are designed for individual modalities, often relying on domain-specific assumptions such as the local grid structures exploited by virtually all existing vision models. These priors introduce helpful inductive biases, but also lock models to individual modalities. In this paper we introduce the Perceiver - a model that builds upon Transformers and hence makes few architectural assumptions about the relationship between its inputs, but that also scales to hundreds of thousands of inputs, like ConvNets. The model leverages an asymmetric attention mechanism to iteratively distill inputs into a tight latent bottleneck, allowing it to scale to handle very large inputs. We show that this architecture performs competitively or beyond strong, specialized models on classification tasks across various modalities: images, point clouds, audio, video and video+audio. The Perceiver obtains performance comparable to ResNet-50 on ImageNet without convolutions and by directly attending to 50,000 pixels. It also surpasses state-of-the-art results for all modalities in AudioSet.
Authors: Andrew Jaegle, Felix Gimeno, Andrew Brock, Andrew Zisserman, Oriol Vinyals, Joao Carreira
Links:
TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick
YouTube: https://www.youtube.com/c/yannickilcher
Twitter: https://twitter.com/ykilcher
Discord: https://discord.gg/4H8xxDF
BitChute: https://www.bitchute.com/channel/yann...
Minds: https://www.minds.com/ykilcher
Parler: https://parler.com/profile/YannicKilcher
LinkedIn: https://www.linkedin.com/in/yannic-ki...
BiliBili: https://space.bilibili.com/1824646584
If you want to support me, the best thing to do is to share out the content :)
If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):
SubscribeStar: https://www.subscribestar.com/yannick...
Patreon: https://www.patreon.com/yannickilcher
Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq
Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2
Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m
Large-scale pre-training and subsequent fine-tuning is a common recipe for success with transformer models in machine learning. However, most such transfer learning is done when a model is pre-trained on the same or a very similar modality to the final task to be solved. This paper demonstrates that transformers can be fine-tuned to completely different modalities, such as from language to vision. Moreover, they demonstrate that this can be done by freezing all attention layers, tuning less than .1% of all parameters. The paper further claims that language modeling is a superior pre-training task for such cross-domain transfer. The paper goes through various ablation studies to make its point.
OUTLINE:
0:00 - Intro & Overview
2:00 - Frozen Pretrained Transformers
4:50 - Evaluated Tasks
10:05 - The Importance of Training LayerNorm
17:10 - Modality Transfer
25:10 - Network Architecture Ablation
26:10 - Evaluation of the Attention Mask
27:20 - Are FPTs Overfitting or Underfitting?
28:20 - Model Size Ablation
28:50 - Is Initialization All You Need?
31:40 - Full Model Training Overfits
32:15 - Again the Importance of Training LayerNorm
33:10 - Conclusions & Comments
Paper: https://arxiv.org/abs/2103.05247
Code: https://github.com/kzl/universal-comp...
Abstract:
We investigate the capability of a transformer pretrained on natural language to generalize to other modalities with minimal finetuning -- in particular, without finetuning of the self-attention and feedforward layers of the residual blocks. We consider such a model, which we call a Frozen Pretrained Transformer (FPT), and study finetuning it on a variety of sequence classification tasks spanning numerical computation, vision, and protein fold prediction. In contrast to prior works which investigate finetuning on the same modality as the pretraining dataset, we show that pretraining on natural language improves performance and compute efficiency on non-language downstream tasks. In particular, we find that such pretraining enables FPT to generalize in zero-shot to these modalities, matching the performance of a transformer fully trained on these tasks.
Authors: Kevin Lu, Aditya Grover, Pieter Abbeel, Igor Mordatch
Links:
TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick
YouTube: https://www.youtube.com/c/yannickilcher
Twitter: https://twitter.com/ykilcher
Discord: https://discord.gg/4H8xxDF
BitChute: https://www.bitchute.com/channel/yann...
Minds: https://www.minds.com/ykilcher
Parler: https://parler.com/profile/YannicKilcher
LinkedIn: https://www.linkedin.com/in/yannic-ki...
BiliBili: https://space.bilibili.com/1824646584
If you want to support me, the best thing to do is to share out the content :)
If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):
SubscribeStar: https://www.subscribestar.com/yannick...
Patreon: https://www.patreon.com/yannickilcher
Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq
Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2
Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m
Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n
Deep Learning systems can achieve remarkable, even super-human performance through supervised learning on large, labeled datasets. However, there are two problems: First, collecting ever more labeled data is expensive in both time and money. Second, these deep neural networks will be high performers on their task, but cannot easily generalize to other, related tasks, or they need large amounts of data to do so. In this blog post, Yann LeCun and Ishan Misra of Facebook AI Research (FAIR) describe the current state of Self-Supervised Learning (SSL) and argue that it is the next step in the development of AI that uses fewer labels and can transfer knowledge faster than current systems. They suggest as a promising direction to build non-contrastive latent-variable predictive models, like VAEs, but ones that also provide high-quality latent representations for downstream tasks.
OUTLINE:
0:00 - Intro & Overview
1:15 - Supervised Learning, Self-Supervised Learning, and Common Sense
7:35 - Predicting Hidden Parts from Observed Parts
17:50 - Self-Supervised Learning for Language vs Vision
26:50 - Energy-Based Models
30:15 - Joint-Embedding Models
35:45 - Contrastive Methods
43:45 - Latent-Variable Predictive Models and GANs
55:00 - Summary & Conclusion
Paper (Blog Post): https://ai.facebook.com/blog/self-sup...
My Video on BYOL: https://www.youtube.com/watch?v=YPfUi...
ERRATA:
The difference between loss and energy: Energy is for inference, loss is for training.
The R(z) term is a regularizer that restricts the capacity of the latent variable. I think I said both of those things, but never together.
The way I explain why BERT is contrastive is wrong. I haven't figured out why just yet, though :)
Video approved by Antonio.
Abstract:
We believe that self-supervised learning (SSL) is one of the most promising ways to build such background knowledge and approximate a form of common sense in AI systems.
Authors: Yann LeCun, Ishan Misra
Links:
TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick
YouTube: https://www.youtube.com/c/yannickilcher
Twitter: https://twitter.com/ykilcher
Discord: https://discord.gg/4H8xxDF
BitChute: https://www.bitchute.com/channel/yann...
Minds: https://www.minds.com/ykilcher
Parler: https://parler.com/profile/YannicKilcher
LinkedIn: https://www.linkedin.com/in/yannic-ki...
BiliBili: https://space.bilibili.com/1824646584
If you want to support me, the best thing to do is to share out the content :)
If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):
SubscribeStar: https://www.subscribestar.com/yannick...
Patreon: https://www.patreon.com/yannickilcher
Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq
Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2
Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m
Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n
OpenAI does a huge investigation into the inner workings of their recent CLIP model via faceted feature visualization and finds amazing things: Some neurons in the last layer respond to distinct concepts across multiple modalities, meaning they fire for photographs, drawings, and signs depicting the same concept, even when the images are vastly distinct. Through manual examination, they identify and investigate neurons corresponding to persons, geographical regions, religions, emotions, and much more. In this video, I go through the publication and then I present my own findings from digging around in the OpenAI Microscope.
OUTLINE:
0:00 - Intro & Overview
3:35 - OpenAI Microscope
7:10 - Categories of found neurons
11:10 - Person Neurons
13:00 - Donald Trump Neuron
17:15 - Emotion Neurons
22:45 - Region Neurons
26:40 - Sparse Mixture of Emotions
28:05 - Emotion Atlas
29:45 - Adversarial Typographic Attacks
31:55 - Stroop Test
33:10 - My Findings in OpenAI Microscope
33:30 - Superman Neuron
33:50 - Resting B*tchface Neuron
34:10 - Trash Bag Neuron
35:25 - God Weightlifting Neuron
36:40 - Organ Neuron
38:35 - Film Spool Neuron
39:05 - Feather Neuron
39:20 - Spartan Neuron
40:25 - Letter E Neuron
40:35 - Cleanin Neuron
40:45 - Frown Neuron
40:55 - Lion Neuron
41:05 - Fashion Model Neuron
41:20 - Baseball Neuron
41:50 - Bride Neuron
42:00 - Navy Neuron
42:30 - Hemp Neuron
43:25 - Staircase Neuron
43:45 - Disney Neuron
44:15 - Hillary Clinton Neuron
44:50 - God Neuron
45:15 - Blurry Neuron
45:35 - Arrow Neuron
45:55 - Trophy Presentation Neuron
46:10 - Receding Hairline Neuron
46:30 - Traffic Neuron
46:40 - Raised Hand Neuron
46:50 - Google Maps Neuron
47:15 - Nervous Smile Neuron
47:30 - Elvis Neuron
47:55 - The Flash Neuron
48:05 - Beard Neuron
48:15 - Kilt Neuron
48:25 - Rainy Neuron
48:35 - Electricity Neuron
48:50 - Droplets Neuron
49:00 - Escape Neuron
49:25 - King Neuron
49:35 - Country Neuron
49:45 - Overweight Men Neuron
49:55 - Wedding
50:05 - Australia Neuron
50:15 - Yawn Neuron
50:30 - Bees & Simpsons Neuron
50:40 - Mussles Neuron
50:50 - Spice Neuron
51:00 - Conclusion
Paper: https://distill.pub/2021/multimodal-n...
My Findings: https://www.notion.so/CLIP-OpenAI-Mic...
My Video on CLIP: https://youtu.be/T9XSU0pKX2E
My Video on Feature Visualizations & The OpenAI Microscope: https://youtu.be/Ok44otx90D4
Abstract:
In 2005, a letter published in Nature described human neurons responding to specific people, such as Jennifer Aniston or Halle Berry. The exciting thing wasn’t just that they selected for particular people, but that they did so regardless of whether they were shown photographs, drawings, or even images of the person’s name. The neurons were multimodal. As the lead author would put it: "You are looking at the far end of the transformation from metric, visual shapes to conceptual... information." We report the existence of similar multimodal neurons in artificial neural networks. This includes neurons selecting for prominent public figures or fictional characters, such as Lady Gaga or Spiderman. Like the biological multimodal neurons, these artificial neurons respond to the same subject in photographs, drawings, and images of their name.
Authors: Gabriel Goh, Nick Cammarata, Chelsea Voss, Shan Carter, Michael Petrov, Ludwig Schubert, Alec Radford, Chris Olah
This video is advice for new PhD students in the field of Machine Learning in 2021 and after. The field has shifted dramatically in the last few years and navigating grad school can be very hard, especially when you're as clueless as I was when I started. The video is a personal recount of my mistakes and what I've learned from them. If you already have several published papers and know what to do, this video is not for you. However, if you are not even sure where to start, how to select a topic, or what goes in a paper, you might benefit from this video, because that's exactly how I felt.
Main Takeaways:
Select niche topics rather than hype topics
Write papers that can't be rejected
Don't be discouraged by bad reviews
Take reviewing & teaching seriously
Keep up your focus
Conferences are for networking
Internships are great opportunities
Team up with complementary skills
Don't work too hard
OUTLINE:
0:00 - Intro & Overview
1:25 - Thesis Topic Selection
4:25 - How To Publish Papers
5:35 - Dealing With Reviewers
6:30 - How To Be A Reviewer
7:40 - Take Teaching Seriously
8:30 - Maintain Focus
10:20 - Navigating Conferences
12:40 - Internships
13:40 - Collaborations
14:55 - Don't Forget To Enjoy
Transcript: https://www.notion.so/Yannic-Kilcher-...
Credits to Lanz for editing
Links:
TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick
YouTube: https://www.youtube.com/c/yannickilcher
Twitter: https://twitter.com/ykilcher
Discord: https://discord.gg/4H8xxDF
BitChute: https://www.bitchute.com/channel/yann...
Minds: https://www.minds.com/ykilcher
Parler: https://parler.com/profile/YannicKilcher
LinkedIn: https://www.linkedin.com/in/yannic-ki...
BiliBili: https://space.bilibili.com/1824646584
If you want to support me, the best thing to do is to share out the content :)
If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):
SubscribeStar: https://www.subscribestar.com/yannick...
Patreon: https://www.patreon.com/yannickilcher
Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq
Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2
Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m
Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n
In the recurring debate about bias in Machine Learning models, there is a growing argument saying that "the problem is not in the data", often citing the influence of various choices like loss functions or network architecture. In this video, we take a look at PAIR's AI Explorables through the lens of whether or not the bias problem is a data problem.
OUTLINE:
0:00 - Intro & Overview
1:45 - Recap: Bias in ML
4:25 - AI Explorables
5:40 - Measuring Fairness Explorable
11:00 - Hidden Bias Explorable
16:10 - Measuring Diversity Explorable
23:00 - Conclusion & Comments
AI Explorables: https://pair.withgoogle.com/explorables/
Links:
TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick
YouTube: https://www.youtube.com/c/yannickilcher
Twitter: https://twitter.com/ykilcher
Discord: https://discord.gg/4H8xxDF
BitChute: https://www.bitchute.com/channel/yann...
Minds: https://www.minds.com/ykilcher
Parler: https://parler.com/profile/YannicKilcher
LinkedIn: https://www.linkedin.com/in/yannic-ki...
BiliBili: https://space.bilibili.com/1824646584
If you want to support me, the best thing to do is to share out the content :)
If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):
SubscribeStar: https://www.subscribestar.com/yannick...
Patreon: https://www.patreon.com/yannickilcher
Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq
Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2
Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m
Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n
Classic Machine Learning struggles with few-shot generalization for tasks where humans can easily generalize from just a handful of examples, for example sorting a list of numbers. Humans do this by coming up with a short program, or algorithm, that explains the few data points in a compact way. DreamCoder emulates this by using neural guided search over a language of primitives, a library, that it builds up over time. By doing this, it can iteratively construct more and more complex programs by building on its own abstractions and therefore solve more and more difficult tasks in a few-shot manner by generating very short programs that solve the few given datapoints. The resulting system can not only generalize quickly but also delivers an explainable solution to its problems in form of a modular and hierarchical learned library. Combining this with classic Deep Learning for low-level perception is a very promising future direction.
OUTLINE:
0:00 - Intro & Overview
4:55 - DreamCoder System Architecture
9:00 - Wake Phase: Neural Guided Search
19:15 - Abstraction Phase: Extending the Internal Library
24:30 - Dreaming Phase: Training Neural Search on Fictional Programs and Replays
30:55 - Abstraction by Compressing Program Refactorings
32:40 - Experimental Results on LOGO Drawings
39:00 - Ablation Studies
39:50 - Re-Discovering Physical Laws
42:25 - Discovering Recursive Programming Algorithms
44:20 - Conclusions & Discussion
Paper: https://arxiv.org/abs/2006.08381
Code: https://github.com/ellisk42/ec
Abstract:
Expert problem-solving is driven by powerful languages for thinking about problems and their solutions. Acquiring expertise means learning these languages -- systems of concepts, alongside the skills to use them. We present DreamCoder, a system that learns to solve problems by writing programs. It builds expertise by creating programming languages for expressing domain concepts, together with neural networks to guide the search for programs within these languages. A ``wake-sleep'' learning algorithm alternately extends the language with new symbolic abstractions and trains the neural network on imagined and replayed problems. DreamCoder solves both classic inductive programming tasks and creative tasks such as drawing pictures and building scenes. It rediscovers the basics of modern functional programming, vector algebra and classical physics, including Newton's and Coulomb's laws. Concepts are built compositionally from those learned earlier, yielding multi-layered symbolic representations that are interpretable and transferrable to new tasks, while still growing scalably and flexibly with experience.
Authors: Kevin Ellis, Catherine Wong, Maxwell Nye, Mathias Sable-Meyer, Luc Cary, Lucas Morales, Luke Hewitt, Armando Solar-Lezama, Joshua B. Tenenbaum
Links:
TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick
YouTube: https://www.youtube.com/c/yannickilcher
Twitter: https://twitter.com/ykilcher
Discord: https://discord.gg/4H8xxDF
BitChute: https://www.bitchute.com/channel/yann...
Minds: https://www.minds.com/ykilcher
Parler: https://parler.com/profile/YannicKilcher
LinkedIn: https://www.linkedin.com/in/yannic-ki...
BiliBili: https://space.bilibili.com/1824646584
If you want to support me, the best thing to do is to share out the content :)
If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):
SubscribeStar: https://www.subscribestar.com/yannick...
Patreon: https://www.patreon.com/yannickilcher
View Synthesis is a tricky problem, especially when only given a sparse set of images as an input. NeRF embeds an entire scene into the weights of a feedforward neural network, trained by backpropagation through a differential volume rendering procedure, and achieves state-of-the-art view synthesis. It includes directional dependence and is able to capture fine structural details, as well as reflection effects and transparency.
OUTLINE:
0:00 - Intro & Overview
4:50 - View Synthesis Task Description
5:50 - The fundamental difference to classic Deep Learning
7:00 - NeRF Core Concept
15:30 - Training the NeRF from sparse views
20:50 - Radiance Field Volume Rendering
23:20 - Resulting View Dependence
24:00 - Positional Encoding
28:00 - Hierarchical Volume Sampling
30:15 - Experimental Results
33:30 - Comments & Conclusion
Paper: https://arxiv.org/abs/2003.08934
Website & Code: https://www.matthewtancik.com/nerf
My Video on SIREN: https://youtu.be/Q5g3p9Zwjrk
Abstract:
We present a method that achieves state-of-the-art results for synthesizing novel views of complex scenes by optimizing an underlying continuous volumetric scene function using a sparse set of input views. Our algorithm represents a scene using a fully-connected (non-convolutional) deep network, whose input is a single continuous 5D coordinate (spatial location (x,y,z) and viewing direction (θ,ϕ)) and whose output is the volume density and view-dependent emitted radiance at that spatial location. We synthesize views by querying 5D coordinates along camera rays and use classic volume rendering techniques to project the output colors and densities into an image. Because volume rendering is naturally differentiable, the only input required to optimize our representation is a set of images with known camera poses. We describe how to effectively optimize neural radiance fields to render photorealistic novel views of scenes with complicated geometry and appearance, and demonstrate results that outperform prior work on neural rendering and view synthesis. View synthesis results are best viewed as videos, so we urge readers to view our supplementary video for convincing comparisons.
Authors: Ben Mildenhall, Pratul P. Srinivasan, Matthew Tancik, Jonathan T. Barron, Ravi Ramamoorthi, Ren Ng
Links:
TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick
YouTube: https://www.youtube.com/c/yannickilcher
Twitter: https://twitter.com/ykilcher
Discord: https://discord.gg/4H8xxDF
BitChute: https://www.bitchute.com/channel/yann...
Minds: https://www.minds.com/ykilcher
Parler: https://parler.com/profile/YannicKilcher
LinkedIn: https://www.linkedin.com/in/yannic-ki...
BiliBili: https://space.bilibili.com/1824646584
If you want to support me, the best thing to do is to share out the content :)
If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):
SubscribeStar: https://www.subscribestar.com/yannick...
Patreon: https://www.patreon.com/yannickilcher
Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq
Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2
Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m
Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n
Self-Supervised Learning is the final frontier in Representation Learning: Getting useful features without any labels. Facebook AI's new system, DINO, combines advances in Self-Supervised Learning for Computer Vision with the new Vision Transformer (ViT) architecture and achieves impressive results without any labels. Attention maps can be directly interpreted as segmentation maps, and the obtained representations can be used for image retrieval and zero-shot k-nearest neighbor classifiers (KNNs).
OUTLINE:
0:00 - Intro & Overview
6:20 - Vision Transformers
9:20 - Self-Supervised Learning for Images
13:30 - Self-Distillation
15:20 - Building the teacher from the student by moving average
16:45 - DINO Pseudocode
23:10 - Why Cross-Entropy Loss?
28:20 - Experimental Results
33:40 - My Hypothesis why this works
38:45 - Conclusion & Comments
Paper: https://arxiv.org/abs/2104.14294
Blog: https://ai.facebook.com/blog/dino-paws-computer-vision-with-self-supervised-transformers-and-10x-more-efficient-training
Code: https://github.com/facebookresearch/dino
My Video on ViT: https://youtu.be/TrdevFK_am4
My Video on BYOL: https://youtu.be/YPfUiOMYOEE
Abstract:
In this paper, we question if self-supervised learning provides new properties to Vision Transformer (ViT) that stand out compared to convolutional networks (convnets). Beyond the fact that adapting self-supervised methods to this architecture works particularly well, we make the following observations: first, self-supervised ViT features contain explicit information about the semantic segmentation of an image, which does not emerge as clearly with supervised ViTs, nor with convnets. Second, these features are also excellent k-NN classifiers, reaching 78.3% top-1 on ImageNet with a small ViT. Our study also underlines the importance of momentum encoder, multi-crop training, and the use of small patches with ViTs. We implement our findings into a simple self-supervised method, called DINO, which we interpret as a form of self-distillation with no labels. We show the synergy between DINO and ViTs by achieving 80.1% top-1 on ImageNet in linear evaluation with ViT-Base.
Authors: Mathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou, Julien Mairal, Piotr Bojanowski, Armand Joulin
Links:
TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick
YouTube: https://www.youtube.com/c/yannickilcher
Twitter: https://twitter.com/ykilcher
Discord: https://discord.gg/4H8xxDF
BitChute: https://www.bitchute.com/channel/yannic-kilcher
Minds: https://www.minds.com/ykilcher
Parler: https://parler.com/profile/YannicKilcher
LinkedIn: https://www.linkedin.com/in/yannic-kilcher-488534136/
BiliBili: https://space.bilibili.com/1824646584
If you want to support me, the best thing to do is to share out the content :)
If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):
SubscribeStar: https://www.subscribestar.com/yannickilcher
Patreon: https://www.patreon.com/yannickilcher
Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq
Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2
Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m
Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n
The AI community has gone through regular cycles of AI Springs, where rapid progress gave rise to massive overconfidence, high funding, and overpromise, followed by these promises being unfulfilled, subsequently diving into periods of disenfranchisement and underfunding, called AI Winters. This paper examines the reasons for the repeated periods of overconfidence and identifies four fallacies that people make when they see rapid progress in AI.
OUTLINE:
0:00 - Intro & Overview
2:10 - AI Springs & AI Winters
5:40 - Is the current AI boom overhyped?
15:35 - Fallacy 1: Narrow Intelligence vs General Intelligence
19:40 - Fallacy 2: Hard for humans doesn't mean hard for computers
21:45 - Fallacy 3: How we call things matters
28:15 - Fallacy 4: Embodied Cognition
35:30 - Conclusion & Comments
Paper: https://arxiv.org/abs/2104.12871
My Video on Shortcut Learning: https://youtu.be/D-eg7k8YSfs
Abstract:
Since its beginning in the 1950s, the field of artificial intelligence has cycled several times between periods of optimistic predictions and massive investment ("AI spring") and periods of disappointment, loss of confidence, and reduced funding ("AI winter"). Even with today's seemingly fast pace of AI breakthroughs, the development of long-promised technologies such as self-driving cars, housekeeping robots, and conversational companions has turned out to be much harder than many people expected. One reason for these repeating cycles is our limited understanding of the nature and complexity of intelligence itself. In this paper I describe four fallacies in common assumptions made by AI researchers, which can lead to overconfident predictions about the field. I conclude by discussing the open questions spurred by these fallacies, including the age-old challenge of imbuing machines with humanlike common sense.
Authors: Melanie Mitchell
Links:
TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick
YouTube: https://www.youtube.com/c/yannickilcher
Twitter: https://twitter.com/ykilcher
Discord: https://discord.gg/4H8xxDF
BitChute: https://www.bitchute.com/channel/yannic-kilcher
Minds: https://www.minds.com/ykilcher
Parler: https://parler.com/profile/YannicKilcher
LinkedIn: https://www.linkedin.com/in/yannic-kilcher-488534136/
BiliBili: https://space.bilibili.com/1824646584
If you want to support me, the best thing to do is to share out the content :)
If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):
SubscribeStar: https://www.subscribestar.com/yannickilcher
Patreon: https://www.patreon.com/yannickilcher
Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq
Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2
Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m
Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n