Yannic Kilcher Videos (Audio Only): Recent Episodes

Yannic Kilcher

I make videos about machine learning research papers, programming, and issues of the AI community, and the broader impact of AI in society.

Twitter: https://twitter.com/ykilcher Discord: https://discord.gg/4H8xxDF

If you want to support me, the best thing to do is to share out the content :)

If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this): SubscribeStar (preferred to Patreon): https://www.subscribestar.com/yannickilcher Patreon: https://www.patreon.com/yannickilcher Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq

View Details

llm #ai #chatgpt How does one run inference for a generative autoregressive language model that has been trained with a fixed context size? Streaming LLMs combine the performance of windowed attention, but avoid the drop in performance by using attention sinks - an interesting phenomenon where the token at position 0 acts as an absorber of "extra" attention.OUTLINE:0:00 - Introduction1:20 - What is the problem?10:30 - The hypothesis: Attention Sinks15:10 - Experimental evidence18:45 - Streaming LLMs20:45 - Semantics or position?22:30 - Can attention sinks be learned?27:45 - More experiments30:10 - Comparison to Big BirdPaper: https://arxiv.org/abs/2309.17453Abstract:Deploying Large Language Models (LLMs) in streaming applications such as multi-round dialogue, where long interactions are expected, is urgently needed but poses two major challenges. Firstly, during the decoding stage, caching previous tokens' Key and Value states (KV) consumes extensive memory. Secondly, popular LLMs cannot generalize to longer texts than the training sequence length. Window attention, where only the most recent KVs are cached, is a natural approach -- but we show that it fails when the text length surpasses the cache size. We observe an interesting phenomenon, namely attention sink, that keeping the KV of initial tokens will largely recover the performance of window attention. In this paper, we first demonstrate that the emergence of attention sink is due to the strong attention scores towards initial tokens as a ``sink'' even if they are not semantically important. Based on the above analysis, we introduce StreamingLLM, an efficient framework that enables LLMs trained with a finite length attention window to generalize to infinite sequence lengths without any fine-tuning. We show that StreamingLLM can enable Llama-2, MPT, Falcon, and Pythia to perform stable and efficient language modeling with up to 4 million tokens and more. In addition, we discover that adding a placeholder token as a dedicated attention sink during pre-training can further improve streaming deployment. In streaming settings, StreamingLLM outperforms the sliding window recomputation baseline by up to 22.2x speedup. Code and datasets are provided at this https URL.Authors: Guangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han, Mike LewisLinks:Homepage: https://ykilcher.comMerch: https://ykilcher.com/merchYouTube: https://www.youtube.com/c/yannickilcherTwitter: https://twitter.com/ykilcherDiscord: https://ykilcher.com/discordLinkedIn: https://www.linkedin.com/in/ykilcherIf you want to support me, the best thing to do is to share out the content :)If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):SubscribeStar: https://www.subscribestar.com/yannickilcherPatreon: https://www.patreon.com/yannickilcherBitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cqEthereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9mMonero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n

View Details

ai #promptengineering #evolution Promptbreeder is a self-improving self-referential system for automated prompt engineering. Give it a task description and a dataset, and it will automatically come up with appropriate prompts for the task. This is achieved by an evolutionary algorithm where not only the prompts, but also the mutation-prompts are improved over time in a population-based, diversity-focused approach.OUTLINE:0:00 - Introduction2:10 - From manual to automated prompt engineering10:40 - How does Promptbreeder work?21:30 - Mutation operators36:00 - Experimental Results38:05 - A walk through the appendixPaper: https://arxiv.org/abs/2309.16797Abstract:Popular prompt strategies like Chain-of-Thought Prompting can dramatically improve the reasoning abilities of Large Language Models (LLMs) in various domains. However, such hand-crafted prompt-strategies are often sub-optimal. In this paper, we present Promptbreeder, a general-purpose self-referential self-improvement mechanism that evolves and adapts prompts for a given domain. Driven by an LLM, Promptbreeder mutates a population of task-prompts, and subsequently evaluates them for fitness on a training set. Crucially, the mutation of these task-prompts is governed by mutation-prompts that the LLM generates and improves throughout evolution in a self-referential way. That is, Promptbreeder is not just improving task-prompts, but it is also improving the mutationprompts that improve these task-prompts. Promptbreeder outperforms state-of-the-art prompt strategies such as Chain-of-Thought and Plan-and-Solve Prompting on commonly used arithmetic and commonsense reasoning benchmarks. Furthermore, Promptbreeder is able to evolve intricate task-prompts for the challenging problem of hate speech classification.Authors: Chrisantha Fernando, Dylan Banarse, Henryk Michalewski, Simon Osindero, Tim RocktäschelLinks:Homepage: https://ykilcher.comMerch: https://ykilcher.com/merchYouTube: https://www.youtube.com/c/yannickilcherTwitter: https://twitter.com/ykilcherDiscord: https://ykilcher.com/discordLinkedIn: https://www.linkedin.com/in/ykilcherIf you want to support me, the best thing to do is to share out the content :)If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):SubscribeStar: https://www.subscribestar.com/yannickilcherPatreon: https://www.patreon.com/yannickilcherBitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cqEthereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9mMonero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n

View Details

ai #retnet #transformers Retention is an alternative to Attention in Transformers that can both be written in a parallel and in a recurrent fashion. This means the architecture achieves training parallelism while maintaining low-cost inference. Experiments in the paper look very promising.OUTLINE:0:00 - Intro2:40 - The impossible triangle6:55 - Parallel vs sequential15:35 - Retention mechanism21:00 - Chunkwise and multi-scale retention24:10 - Comparison to other architectures26:30 - Experimental evaluationPaper: https://arxiv.org/abs/2307.08621Abstract:In this work, we propose Retentive Network (RetNet) as a foundation architecture for large language models, simultaneously achieving training parallelism, low-cost inference, and good performance. We theoretically derive the connection between recurrence and attention. Then we propose the retention mechanism for sequence modeling, which supports three computation paradigms, i.e., parallel, recurrent, and chunkwise recurrent. Specifically, the parallel representation allows for training parallelism. The recurrent representation enables low-cost O(1) inference, which improves decoding throughput, latency, and GPU memory without sacrificing performance. The chunkwise recurrent representation facilitates efficient long-sequence modeling with linear complexity, where each chunk is encoded parallelly while recurrently summarizing the chunks. Experimental results on language modeling show that RetNet achieves favorable scaling results, parallel training, low-cost deployment, and efficient inference. The intriguing properties make RetNet a strong successor to Transformer for large language models. Code will be available at this https URL.Authors: Yutao Sun, Li Dong, Shaohan Huang, Shuming Ma, Yuqing Xia, Jilong Xue, Jianyong Wang, Furu WeiLinks:Homepage: https://ykilcher.comMerch: https://ykilcher.com/merchYouTube: https://www.youtube.com/c/yannickilcherTwitter: https://twitter.com/ykilcherDiscord: https://ykilcher.com/discordLinkedIn: https://www.linkedin.com/in/ykilcherIf you want to support me, the best thing to do is to share out the content :)If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):SubscribeStar: https://www.subscribestar.com/yannickilcherPatreon: https://www.patreon.com/yannickilcherBitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cqEthereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9mMonero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n

View Details

ai #rlhf #llm ReST uses a bootsrap-like method to produce its own extended dataset and trains on ever higher-quality subsets of it to improve its own reward. The method allows for re-using the same generated data multiple times and thus has an efficiency advantage with respect to Online RL techniques like PPO.Paper: https://arxiv.org/abs/2308.08998Abstract:Reinforcement learning from human feedback (RLHF) can improve the quality of large language model's (LLM) outputs by aligning them with human preferences. We propose a simple algorithm for aligning LLMs with human preferences inspired by growing batch reinforcement learning (RL), which we call Reinforced Self-Training (ReST). Given an initial LLM policy, ReST produces a dataset by generating samples from the policy, which are then used to improve the LLM policy using offline RL algorithms. ReST is more efficient than typical online RLHF methods because the training dataset is produced offline, which allows data reuse. While ReST is a general approach applicable to all generative learning settings, we focus on its application to machine translation. Our results show that ReST can substantially improve translation quality, as measured by automated metrics and human evaluation on machine translation benchmarks in a compute and sample-efficient manner.Authors: Caglar Gulcehre, Tom Le Paine, Srivatsan Srinivasan, Ksenia Konyushkova, Lotte Weerts, Abhishek Sharma, Aditya Siddhant, Alex Ahern, Miaosen Wang, Chenjie Gu, Wolfgang Macherey, Arnaud Doucet, Orhan Firat, Nando de FreitasLinks:Homepage: https://ykilcher.comMerch: https://ykilcher.com/merchYouTube: https://www.youtube.com/c/yannickilcherTwitter: https://twitter.com/ykilcherDiscord: https://ykilcher.com/discordLinkedIn: https://www.linkedin.com/in/ykilcherIf you want to support me, the best thing to do is to share out the content :)If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):SubscribeStar: https://www.subscribestar.com/yannickilcherPatreon: https://www.patreon.com/yannickilcherBitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cqEthereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9mMonero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n

View Details

mlnews #llama2 #openai Your regular irregular update on the world of Machine Learning.References:https://twitter.com/ylecun/status/1681336284453781505https://ai.meta.com/llama/https://about.fb.com/news/2023/07/llama-2-statement-of-support/https://247wallst.com/special-report/2023/08/12/this-is-the-biggest-social-media-platform-ranking-the-worlds-largest-networking-sites/4/https://github.com/Alpha-VLLM/LLaMA2-Accessoryhttps://together.ai/blog/llama-2-7b-32k?s=09&utm_source=pocket_saveshttps://github.com/imoneoi/openchathttps://twitter.com/lmsysorg/status/1686794639469371393?s=09&t=sS3awkbavmSMSmwp64Ef4A&utm_source=pocket_saveshttps://huggingface.co/lmsys/vicuna-13b-v1.5-16khttps://blog.google/outreach-initiatives/public-policy/google-microsoft-openai-anthropic-frontier-model-forum/https://www.earthdata.nasa.gov/news/impact-ibm-hls-foundation-model?utm_source=pocket_readerhttps://huggingface.co/ibm-nasa-geospatial/Prithvi-100Mhttps://ai.meta.com/blog/generative-ai-text-images-cm3leon/https://www.deepmind.com/blog/rt-2-new-model-translates-vision-and-language-into-action?utm_source=twitter&utm_medium=social&utm_campaign=rt2https://arxiv.org/abs/2307.14334https://sites.research.google/med-palm/https://open-catalyst.metademolab.com/?utm_source=twitter&utm_medium=organic_social&utm_campaign=opencatalyst&utm_content=cardhttps://open-catalyst.metademolab.com/demohttps://www.anthropic.com/index/claude-2?utm_source=pocket_readerhttps://claude.ai/loginhttps://audiocraft.metademolab.com/?utm_source=pocket_saveshttps://venturebeat.com/programming-development/stability-ai-launches-stablecode-an-llm-for-code-generation/https://stability.ai/blog/stablecode-llm-generative-ai-codinghttps://twitter.com/JeffDean/status/1686806525862608896?s=09&t=LG2z9ok9QExHbSy0fvBsxA&utm_source=pocket_saveshttps://sites.research.google/open-buildings/https://twitter.com/deliprao/status/1687283117873106946?s=09&t=1NmC-B55Z8IuF_HTuGOo7w&utm_source=pocket_saveshttps://arxiv.org/pdf/2308.01320.pdfhttps://twitter.com/javilopen/status/1687795349719547905?utm_source=pocket_saveshttps://research.nvidia.com/labs/par/Perfusion/https://ar5iv.labs.arxiv.org/html/2307.14936https://www.linkedin.com/feed/update/urn:li:activity:7093463974750371840/?utm_source=pocket_saveshttps://huggingface.co/syzymon/long_llama_3b_instructhttps://arxiv.org/abs/2307.03170https://dynalang.github.io/https://github.com/mlfoundations/open_flamingohttps://twitter.com/akshay_pachaar/status/1687079353937698816?s=09&t=fos8QSCsGEEM6dMflhq0Mg&utm_source=pocket_saveshttps://github.com/OpenBMB/ToolBenchhttps://llm-attacks.org/https://arstechnica.com/information-technology/2023/07/openai-discontinues-its-ai-writing-detector-due-to-low-rate-of-accuracy/https://sites.google.com/view/steve-1https://github.com/Shalev-Lifshitz/STEVE-1https://erichartford.com/dolphinhttps://huggingface.co/ehartford/dolphin-llama-13bhttps://www.mosaicml.com/blog/long-context-mpt-7b-8khttps://twitter.com/camenduru/status/1688045780244848640?s=09&t=ubJ2Qtz-TG6Xo3_GMtt2Cw&utm_source=pocket_saveshttps://github.com/IDEA-Research/DWPosehttps://twitter.com/tri_dao/status/1680987577913065472?s=09&t=Q181vFmM6d3nDq-5BwfDeg&utm_source=pocket_saveshttps://tridao.me/publications/flash2/flash2.pdfhttps://thehackernews.com/2023/07/wormgpt-new-ai-tool-allows.htmlhttps://www.tomshardware.com/news/ai-steals-data-with-keystroke-audiohttps://arxiv.org/pdf/2308.01074.pdfhttps://www.foxnews.com/politics/ai-test-flight-air-force-unmanned-wingman-aircrafthttps://www.theverge.com/2023/8/2/23817406/white-castle-soundhound-ai-slidershttps://www.google.com/search?sca_esv=556495916&q=food+delivery+bot+kicked&tbm=vid&source=lnms&sa=X&ved=2ahUKEwjZ6PDPrdmAAxUThf0HHWzrBGgQ0pQJegQIChAB&cshid=1691920142432720&biw=2327&bih=1180&dpr=2.2https://www.youtube.com/watch?v=--n_NhmXnfchttps://www.thesun.co.uk/tech/20793591/coop-delivery-robots-cambridge-kicked-by-workers-tiktok/

View Details

cybercrime #chatgpt #security An interview with Sergey Shykevich, Threat Intelligence Group Manager at Check Point, about how models like ChatGPT have impacted the realm of cyber crime.https://threatmap.checkpoint.com/Links:Homepage: https://ykilcher.comMerch: https://ykilcher.com/merchYouTube: https://www.youtube.com/c/yannickilcherTwitter: https://twitter.com/ykilcherDiscord: https://ykilcher.com/discordLinkedIn: https://www.linkedin.com/in/ykilcherIf you want to support me, the best thing to do is to share out the content :)If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):SubscribeStar: https://www.subscribestar.com/yannickilcherPatreon: https://www.patreon.com/yannickilcherBitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cqEthereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9mMonero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n

View Details

llm #safety #gpt4 A prime example of intellectual dishonesty of journalists and AI critics.Article: https://gizmodo.com/paknsave-ai-savey-recipe-bot-chlorine-gas-1850725057My Recipe AI: https://github.com/yk/recipe-aiLinks:Homepage: https://ykilcher.comMerch: https://ykilcher.com/merchYouTube: https://www.youtube.com/c/yannickilcherTwitter: https://twitter.com/ykilcherDiscord: https://ykilcher.com/discordLinkedIn: https://www.linkedin.com/in/ykilcherIf you want to support me, the best thing to do is to share out the content :)If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):SubscribeStar: https://www.subscribestar.com/yannickilcherPatreon: https://www.patreon.com/yannickilcherBitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cqEthereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9mMonero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n

View Details

ai #diffusion #stabilityai An interview with DeepFloyd members Misha Konstantinov and Daria Bakshandaeva on the release of the model IF, an open-source model following Google's implementation of Imagen.References:https://www.deepfloyd.ai/deepfloyd-ifhttps://huggingface.co/DeepFloydhttps://twitter.com/_gugutse_https://twitter.com/_bra_ketLinks:Homepage: https://ykilcher.comMerch: https://ykilcher.com/merchYouTube: https://www.youtube.com/c/yannickilcherTwitter: https://twitter.com/ykilcherDiscord: https://ykilcher.com/discordLinkedIn: https://www.linkedin.com/in/ykilcherIf you want to support me, the best thing to do is to share out the content :)If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):SubscribeStar: https://www.subscribestar.com/yannickilcherPatreon: https://www.patreon.com/yannickilcherBitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cqEthereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9mMonero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n

View Details

gpt4 #mit #ai A new paper claims to use GPT-4 to solve 100% of a set of MIT university exercises. Some people are skeptic and their investigations reveal more than one problem with this paper...OUTLINE:0:00 - ChatGPT gives out Windows 10 keys0:30 - MIT exam paper2:50 - Prompt engineering5:30 - Automatic grading6:45 - Response by other MIT students8:30 - Unsolvable questions10:50 - Duplicates13:30 - Cascading the heuristics22:40 - Other problems29:25 - OpenLLaMA 13B publishedReferences:https://twitter.com/immasiddtweets/status/1669721470006857729/photo/1https://arxiv.org/abs/2306.08997https://arxiv.org/pdf/2306.08997.pdfhttps://flower-nutria-41d.notion.site/No-GPT4-can-t-ace-MIT-b27e6796ab5a48368127a98216c76864https://github.com/idrori/MITQ/commit/3feee1026318e537c0ad27968001ef76e4a36890https://twitter.com/hardmaru/status/1670246674760077312https://twitter.com/giffmana/status/1670258748286472193https://twitter.com/T3816440886465/status/1670127224131862531https://twitter.com/qrdl/status/1669856336652414977https://www.chegg.com/homework-help/questions-and-answers/consider-mdp-set-possible-states-mathcal-s-0-1-2-3-set-possible-actions-mathcal-b-c--rewar-q111042613https://github.com/openlm-research/open_llamahttps://huggingface.co/openlm-research/open_llama_13bLinks:Homepage: https://ykilcher.comMerch: https://ykilcher.com/merchYouTube: https://www.youtube.com/c/yannickilcherTwitter: https://twitter.com/ykilcherDiscord: https://ykilcher.com/discordLinkedIn: https://www.linkedin.com/in/ykilcherIf you want to support me, the best thing to do is to share out the content :)If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):SubscribeStar: https://www.subscribestar.com/yannickilcherPatreon: https://www.patreon.com/yannickilcherBitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cqEthereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9mMonero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n

View Details

stablediffusion #ai #watermark Watermarking the outputs of generative models is usually done as a post-processing step on the model outputs. Tree-Ring Watermarks are applied in the latent space at the beginning of a diffusion process, which makes them nearly undetectable, robust to strong distortions, and only recoverable by the model author. It is a very promising technique with applications potentially beyond watermarking itself.OUTLINE:0:00 - Introduction & Overview1:30 - Why Watermarking?4:20 - Diffusion Models Recap13:40 - Inverting Diffusion Models17:05 - Tree-Ring Watermarking26:15 - Effects of Tree-Ring Watermarks30:00 - Experimental Results32:40 - Limitations34:40 - ConclusionPaper: https://arxiv.org/abs/2305.20030Abstract:Watermarking the outputs of generative models is a crucial technique for tracing copyright and preventing potential harm from AI-generated content. In this paper, we introduce a novel technique called Tree-Ring Watermarking that robustly fingerprints diffusion model outputs. Unlike existing methods that perform post-hoc modifications to images after sampling, Tree-Ring Watermarking subtly influences the entire sampling process, resulting in a model fingerprint that is invisible to humans. The watermark embeds a pattern into the initial noise vector used for sampling. These patterns are structured in Fourier space so that they are invariant to convolutions, crops, dilations, flips, and rotations. After image generation, the watermark signal is detected by inverting the diffusion process to retrieve the noise vector, which is then checked for the embedded signal. We demonstrate that this technique can be easily applied to arbitrary diffusion models, including text-conditioned Stable Diffusion, as a plug-in with negligible loss in FID. Our watermark is semantically hidden in the image space and is far more robust than watermarking alternatives that are currently deployed. Code is available at this https URL.Authors: Yuxin Wen, John Kirchenbauer, Jonas Geiping, Tom GoldsteinLinks:Homepage: https://ykilcher.comMerch: https://ykilcher.com/merchYouTube: https://www.youtube.com/c/yannickilcherTwitter: https://twitter.com/ykilcherDiscord: https://ykilcher.com/discordLinkedIn: https://www.linkedin.com/in/ykilcherIf you want to support me, the best thing to do is to share out the content :)If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):SubscribeStar: https://www.subscribestar.com/yannickilcherPatreon: https://www.patreon.com/yannickilcherBitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cqEthereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9mMonero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n

View Details

gpt4 #rwkv #transformer We take a look at RWKV, a highly scalable architecture between Transformers and RNNs.Fully Connected (June 7th in SF) Promo Link: https://www.fullyconnected.com/?promo=ynncOUTLINE:0:00 - Introduction1:50 - Fully Connected In-Person Conference in SF June 7th3:00 - Transformers vs RNNs8:00 - RWKV: Best of both worlds12:30 - LSTMs17:15 - Evolution of RWKV's Linear Attention30:40 - RWKV's Layer Structure49:15 - Time-Parallel vs Sequence Mode53:55 - Experimental Results & Limitations58:00 - Visualizations1:01:40 - ConclusionPaper: https://arxiv.org/abs/2305.13048Code: https://github.com/BlinkDL/RWKV-LMAbstract:Transformers have revolutionized almost all natural language processing (NLP) tasks but suffer from memory and computational complexity that scales quadratically with sequence length. In contrast, recurrent neural networks (RNNs) exhibit linear scaling in memory and computational requirements but struggle to match the same performance as Transformers due to limitations in parallelization and scalability. We propose a novel model architecture, Receptance Weighted Key Value (RWKV), that combines the efficient parallelizable training of Transformers with the efficient inference of RNNs. Our approach leverages a linear attention mechanism and allows us to formulate the model as either a Transformer or an RNN, which parallelizes computations during training and maintains constant computational and memory complexity during inference, leading to the first non-transformer architecture to be scaled to tens of billions of parameters. Our experiments reveal that RWKV performs on par with similarly sized Transformers, suggesting that future work can leverage this architecture to create more efficient models. This work presents a significant step towards reconciling the trade-offs between computational efficiency and model performance in sequence processing tasks.Authors: Bo Peng, Eric Alcaide, Quentin Anthony, Alon Albalak, Samuel Arcadinho, Huanqi Cao, Xin Cheng, Michael Chung, Matteo Grella, Kranthi Kiran GV, Xuzheng He, Haowen Hou, Przemyslaw Kazienko, Jan Kocon, Jiaming Kong, Bartlomiej Koptyra, Hayden Lau, Krishna Sri Ipsit Mantri, Ferdinand Mom, Atsushi Saito, Xiangru Tang, Bolun Wang, Johan S. Wind, Stansilaw Wozniak, Ruichong Zhang, Zhenyuan Zhang, Qihang Zhao, Peng Zhou, Jian Zhu, Rui-Jie ZhuLinks:Homepage: https://ykilcher.comMerch: https://ykilcher.com/merchYouTube: https://www.youtube.com/c/yannickilcherTwitter: https://twitter.com/ykilcherDiscord: https://ykilcher.com/discordLinkedIn: https://www.linkedin.com/in/ykilcherIf you want to support me, the best thing to do is to share out the content :)If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):SubscribeStar: https://www.subscribestar.com/yannickilcherPatreon: https://www.patreon.com/yannickilcherBitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cqEthereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9mMonero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n

View Details

gpt4 #ai #prompt Tree-of-Thought improves prompting of large language models (LLMs) by generalizing the concept of Chain-of-Thought prompting and introduces a tree search across language model thoughts, including state evaluation and backtracking. Experiments on toy tasks show large improvements over both classic and Chain-of-Thought prompting.OUTLINE:0:00 - Introduction1:20 - From Chain-of-Thought to Tree-of-Thought11:10 - Formalizing the algorithm16:00 - Game of 24 & Creative writing18:30 - Crosswords23:30 - Is this a general problem solver?26:50 - Ablation studies28:55 - ConclusionPaper: https://arxiv.org/abs/2305.10601Abstract:Language models are increasingly being deployed for general problem solving across a wide range of tasks, but are still confined to token-level, left-to-right decision-making processes during inference. This means they can fall short in tasks that require exploration, strategic lookahead, or where initial decisions play a pivotal role. To surmount these challenges, we introduce a new framework for language model inference, Tree of Thoughts (ToT), which generalizes over the popular Chain of Thought approach to prompting language models, and enables exploration over coherent units of text (thoughts) that serve as intermediate steps toward problem solving. ToT allows LMs to perform deliberate decision making by considering multiple different reasoning paths and self-evaluating choices to decide the next course of action, as well as looking ahead or backtracking when necessary to make global choices. Our experiments show that ToT significantly enhances language models' problem-solving abilities on three novel tasks requiring non-trivial planning or search: Game of 24, Creative Writing, and Mini Crosswords. For instance, in Game of 24, while GPT-4 with chain-of-thought prompting only solved 4% of tasks, our method achieved a success rate of 74%. Code repo with all prompts: this https URL.Authors: Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Thomas L. Griffiths, Yuan Cao, Karthik NarasimhanLinks:Homepage: https://ykilcher.comMerch: https://ykilcher.com/merchYouTube: https://www.youtube.com/c/yannickilcherTwitter: https://twitter.com/ykilcherDiscord: https://ykilcher.com/discordLinkedIn: https://www.linkedin.com/in/ykilcherIf you want to support me, the best thing to do is to share out the content :)If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):SubscribeStar: https://www.subscribestar.com/yannickilcherPatreon: https://www.patreon.com/yannickilcherBitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cqEthereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9mMonero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n

View Details

ai #openai #gpt4 US Senate hearing on AI regulation.MLST video on the hearing: https://www.youtube.com/watch?v=DeSXnESGxr4Links:Homepage: https://ykilcher.comMerch: https://ykilcher.com/merchYouTube: https://www.youtube.com/c/yannickilcherTwitter: https://twitter.com/ykilcherDiscord: https://ykilcher.com/discordLinkedIn: https://www.linkedin.com/in/ykilcherIf you want to support me, the best thing to do is to share out the content :)If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):SubscribeStar: https://www.subscribestar.com/yannickilcherPatreon: https://www.patreon.com/yannickilcherBitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cqEthereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9mMonero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n

View Details

google #openai #mlnews Updates from the world of Machine Learning and AIGreat AI memes here: https://twitter.com/untitled01ipynbOUTLINE:0:00 - Google I/O 2023: Generative AI in everything0:20 - Anthropic announces 100k tokens context0:35 - Intro1:20 - Geoff Hinton leaves Google7:00 - Google memo leaked: we have no moat11:30 - OpenAI loses 540M12:30 - Google AI: Product first15:50 - Ilya Sutskever on safety vs competition18:00 - AI works cannot be copyrighted19:40 - OpenAI tries to trademark GPT20:30 - StarCoder: accessible code model21:40 - RedPyjama & OpenLlama22:55 - Mosaic 7B model23:50 - YoloNAS24:10 - Mojo programming language25:30 - Random helpful things37:40 - DeepMind soccer robotsReferences:https://twitter.com/weirddalle/status/1649908805788893185https://www.nytimes.com/2023/05/01/technology/ai-google-chatbot-engineer-quits-hinton.htmlhttps://www.technologyreview.com/2023/05/01/1072478/deep-learning-pioneer-geoffrey-hinton-quits-google/https://archive.ph/TrPoHhttps://twitter.com/DanHendrycks/status/1654560913939374080https://twitter.com/ylecun/status/1654930029569101824https://twitter.com/homehttps://twitter.com/ylecun/status/1654931495419621376https://twitter.com/pkedrosky/status/1653955254181068801https://www.semianalysis.com/p/google-we-have-no-moat-and-neitherhttps://twitter.com/untitled01ipynb/mediahttps://www.theinformation.com/articles/openais-losses-doubled-to-540-million-as-it-developed-chatgpthttps://archive.ph/bKsdMhttps://www.washingtonpost.com/technology/2023/05/04/google-ai-stop-sharing-research/https://twitter.com/giffmana/status/1654962145707130880https://twitter.com/Ken_Goldberg/status/1651309843804987393https://tsdr.uspto.gov/documentviewer?caseId=sn97733259&docId=PTD20230418160641&s=09#docIndex=1&page=1https://twitter.com/osanseviero/status/1654230764513370112https://huggingface.co/bigcode/starcoderhttps://huggingface.co/spaces/bigcode/bigcode-model-license-agreementhttps://twitter.com/hardmaru/status/1654649036333514753https://www.together.xyz/blog/redpajama-models-v1https://huggingface.co/togethercomputer/RedPajama-INCITE-Base-3B-v1https://github.com/openlm-research/open_llamahttps://www.mosaicml.com/blog/mpt-7bhttps://github.com/Deci-AI/super-gradients/blob/master/YOLONAS.mdhttps://www.modular.com/mojohttps://www.aicrowd.com/challenges/hackaprompt-2023https://learnprompting.org/https://developer.nvidia.com/blog/nvidia-enables-trustworthy-safe-and-secure-large-language-model-conversational-systems/?ncid=prsy-552511https://blogs.nvidia.com/blog/2023/04/25/ai-chatbot-guardrails-nemo/https://lmql.ai/#distributionhttps://github.com/gventuri/pandas-ai?utm_source=pocket_readerhttps://lamini.ai/blog/introducing-laminihttps://github.com/deep-floyd/IFhttps://huggingface.co/spaces/DeepFloyd/IFhttps://twitter.com/FaramaFound/status/1650952295901720576https://txt.cohere.com/embedding-archives-wikipedia/?hsa_acc=509563538&hsa_ad=242008083&hsa_cam=626636963&hsa_grp=205646033&hsa_net=linkedin&hsa_ver=3&hss_channel=lcp-24024765https://arxiv.org/abs/2304.12210https://github.com/h2oai/h2ogpthttps://huggingface.co/h2oai/h2ogpt-oasst1-512-20bhttps://github.com/h2oai/h2o-llmstudiohttps://ai.facebook.com/blog/ai-dataset-animating-kids-drawings/https://www.camel-ai.org/https://github.com/lightaime/camel?utm_source=pocket_readerhttps://huggingface.co/Writer/camel-5b-hfhttps://laion.ai/blog/paella/https://magazine.sebastianraschka.com/p/finetuning-large-language-modelshttps://pickapic.io/https://github.com/yuvalkirstain/heroku_apphttps://huggingface.co/datasets/yuvalkirstain/PickaPichttps://future.snorkel.ai/poster-contest/https://twitter.com/d_feldman/status/1649466422018318338/photo/4https://twitter.com/DeepMind/status/1651897358894919680https://arxiv.org/abs/2304.13653https://twitter.com/SmokeAwayyy/status/1652712832738422784If you want to support me, the best thing to do is to share out the content :)

View Details

ai #transformer #gpt4 This paper promises to scale transformers to 1 million tokens and beyond. We take a look at the technique behind it: The Recurrent Memory Transformer, and what its strenghts and weaknesses are.OUTLINE:0:00 - Intro2:15 - Transformers on long sequences4:30 - Tasks considered8:00 - Recurrent Memory Transformer19:40 - Experiments on scaling and attention maps24:00 - ConclusionPaper: https://arxiv.org/abs/2304.11062Abstract:This technical report presents the application of a recurrent memory to extend the context length of BERT, one of the most effective Transformer-based models in natural language processing. By leveraging the Recurrent Memory Transformer architecture, we have successfully increased the model's effective context length to an unprecedented two million tokens, while maintaining high memory retrieval accuracy. Our method allows for the storage and processing of both local and global information and enables information flow between segments of the input sequence through the use of recurrence. Our experiments demonstrate the effectiveness of our approach, which holds significant potential to enhance long-term dependency handling in natural language understanding and generation tasks as well as enable large-scale context processing for memory-intensive applications.Authors: Aydar Bulatov, Yuri Kuratov, Mikhail S. BurtsevLinks:Homepage: https://ykilcher.comMerch: https://ykilcher.com/merchYouTube: https://www.youtube.com/c/yannickilcherTwitter: https://twitter.com/ykilcherDiscord: https://ykilcher.com/discordLinkedIn: https://www.linkedin.com/in/ykilcherIf you want to support me, the best thing to do is to share out the content :)If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):SubscribeStar: https://www.subscribestar.com/yannickilcherPatreon: https://www.patreon.com/yannickilcherBitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cqEthereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9mMonero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n

View Details

openassistant #chatgpt #mlnews Try the chat: https://open-assistant.io/chatHomepage: https://open-assistant.io Dataset: https://huggingface.co/datasets/OpenAssistant/oasst1Code: https://github.com/LAION-AI/Open-AssistantPaper (temporary): https://ykilcher.com/oa-paperLinks:Homepage: https://ykilcher.comMerch: https://ykilcher.com/merchYouTube: https://www.youtube.com/c/yannickilcherTwitter: https://twitter.com/ykilcherDiscord: https://ykilcher.com/discordLinkedIn: https://www.linkedin.com/in/ykilcherIf you want to support me, the best thing to do is to share out the content :)If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):SubscribeStar: https://www.subscribestar.com/yannickilcherPatreon: https://www.patreon.com/yannickilcherBitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cqEthereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9mMonero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n

View Details

openassistant #chatgpt #gpt4https://open-assistant.io/chathttps://huggingface.co/OpenAssistanthttps://github.com/LAION-AI/Open-AssistantLinks:Homepage: https://ykilcher.comMerch: https://ykilcher.com/merchYouTube: https://www.youtube.com/c/yannickilcherTwitter: https://twitter.com/ykilcherDiscord: https://ykilcher.com/discordLinkedIn: https://www.linkedin.com/in/ykilcherIf you want to support me, the best thing to do is to share out the content :)If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):SubscribeStar: https://www.subscribestar.com/yannickilcherPatreon: https://www.patreon.com/yannickilcherBitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cqEthereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9mMonero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n

View Details

mlnews #gpt4 #copilotYour weekly news all around the AI worldCheck out W&B courses (free): https://wandb.courses/OUTLINE:0:00 - Intro0:20 - GPT-4 announced!4:30 - GigaGAN: The comeback of Generative Adversarial Networks7:55 - ChoppedAI: AI Recipes8:45 - Samsung accused of faking space zoom effect14:00 - Weights & Biases courses are free16:55 - Data Portraits18:50 - Data2Vec 2.019:50 - Gated Models on Hugging Face & huggingface.js22:05 - Visual ChatGPT23:35 - Bing crosses 100 million daily active users24:50 - Casual Conversations Dataset25:50 - Anthropic AI Safety Research27:30 - Magnushammer & more advances in AI-assisted math30:30 - LLaMA license change PR32:00 - Self-Instruct dataset33:35 - PaLM-E: Multimodal Pathways35:45 - USM: Universal Speech Model37:25 - GILGEN: Grounded Text-to-Image39:55 - Fruit Fly Connectome releasedReferences:https://www.heise.de/news/GPT-4-kommt-naechste-Woche-und-es-wird-multimodal-Vorankuendigung-von-Microsoft-7540383.htmlhttps://mingukkang.github.io/GigaGAN/https://www.choppedai.com/https://www.reddit.com/r/Android/comments/11nzrb0/samsung_space_zoom_moon_shots_are_fake_and_here/https://imgur.com/ULVX933https://imgur.com/9XMgt06https://imgur.com/9kichAphttps://imgur.com/RSHAz1lhttps://imgur.com/PIAjVKphttps://imgur.com/xEyLajWhttps://imgur.com/3STX9mZhttps://imgur.com/ifIHr3Shttps://imgur.com/bXJOZgIhttps://dataportraits.org/https://arxiv.org/abs/2303.03919https://arxiv.org/pdf/2303.03919.pdfhttps://ai.facebook.com/blog/ai-self-supervised-learning-data2vec/https://github.com/facebookresearch/fairseq/tree/main/examples/data2vechttps://huggingface.co/docs/hub/models-gatedhttps://huggingface.co/abouthttps://github.com/huggingface/huggingface.js?utm_source=pocket_readerhttps://github.com/microsoft/visual-chatgpthttps://arxiv.org/abs/2303.04671https://github.com/microsoft/visual-chatgpt/blob/main/visual_chatgpt.pyhttps://huggingface.co/spaces/RamAnanth1/visual-chatGPThttps://www.engadget.com/microsoft-bing-crossed-100-million-daily-active-users-080138371.htmlhttps://ai.facebook.com/blog/casual-conversations-v2-dataset-measure-fairness/https://ai.facebook.com/datasets/casual-conversations-v2-dataset/https://www.anthropic.com/index/core-views-on-ai-safetyhttps://arxiv.org/abs/2303.04488https://arxiv.org/pdf/2303.04488.pdfhttps://arxiv.org/abs/2303.04910https://arxiv.org/pdf/2303.04910.pdfhttps://twitter.com/astro_wassim/status/1633645134934949888https://ai.papers.bar/paper/ede58b1ebca911ed8f9c3d8021bca7c8https://arxiv.org/pdf/2303.03192.pdfhttps://www.theverge.com/2023/3/8/23629362/meta-ai-language-model-llama-leak-online-misusehttps://knightcolumbia.org/blog/the-llama-is-out-of-the-bag-should-we-expect-a-tidal-wave-of-disinformationhttps://github.com/facebookresearch/llama/pull/184https://huggingface.co/datasets/yizhongw/self_instructhttps://openai.com/policies/terms-of-usehttps://palm-e.github.io/https://pickapic.io/https://ai.googleblog.com/2023/03/universal-speech-model-usm-state-of-art.htmlhttps://arxiv.org/abs/2303.01037https://github.com/BlinkDL/RWKV-LM?utm_source=pocket_readerhttps://gligen.github.io/https://github.com/microsoft/GLIPhttps://arxiv.org/abs/2301.07093https://huggingface.co/spaces/gligen/demohttps://www.sciencealert.com/the-first-ever-complete-map-of-an-insect-brain-is-truly-mesmerizinghttps://en.wikipedia.org/wiki/Tidal_lockingLinks:Homepage: https://ykilcher.comMerch: https://ykilcher.com/merchYouTube: https://www.youtube.com/c/yannickilcherTwitter: https://twitter.com/ykilcherDiscord: https://ykilcher.com/discordLinkedIn: https://www.linkedin.com/in/ykilcherIf you want to support me, the best thing to do is to share out the content :)

View Details

gpt4 #chatgpt #openai References:https://openai.com/product/gpt-4https://openai.com/research/gpt-4https://cdn.openai.com/papers/gpt-4.pdfLinks:Homepage: https://ykilcher.comMerch: https://ykilcher.com/merchYouTube: https://www.youtube.com/c/yannickilcherTwitter: https://twitter.com/ykilcherDiscord: https://ykilcher.com/discordLinkedIn: https://www.linkedin.com/in/ykilcherIf you want to support me, the best thing to do is to share out the content :)If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):SubscribeStar: https://www.subscribestar.com/yannickilcherPatreon: https://www.patreon.com/yannickilcherBitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cqEthereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9mMonero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n

View Details

mlnews #chatgpt #llamaChatGPT goes around the world and is finally available via API. Stunning mind-reading performed using fMRI and Stable Diffusion. LLaMA weights leak and hilarity ensues. GTC23 is around the corner!ERRATA: It's a 4090, not a 4090 ti 🙃OUTLINE:0:00 - Introduction0:20 - GTC 23 on March 201:55 - ChatGPT API is out!4:50 - OpenAI becomes more business-friendly7:15 - OpenAI plans for AGI10:00 - ChatGPT influencers12:15 - Open-Source Prompting Course12:35 - Flan UL2 20B13:30 - LLaMA weights leaked15:50 - Mind-Reading from fMRI20:10 - Random News / Helpful Things25:30 - Interview with Bryan CatanzaroParticipate in the GTC Raffle: https://ykilcher.com/gtcReferences:GTC 23 on March 20https://www.nvidia.com/gtc/https://ykilcher.com/gtcChatGPT API is out!https://twitter.com/gdb/status/1630991925984755714https://openai.com/blog/introducing-chatgpt-and-whisper-apishttps://twitter.com/greggyb/status/1631121912679002112https://www.haihai.ai/chatgpt-api/OpenAI becomes more business-friendlyhttps://twitter.com/sama/status/1631002519311888385https://techcrunch.com/2023/02/21/openai-foundry-will-let-customers-buy-dedicated-capacity-to-run-its-ai-models/?guccounter=1&guce_referrer=aHR0cHM6Ly93d3cuZ29vZ2xlLmNvbS8&guce_referrer_sig=AQAAAFL1O8s22qBsEtytYZWR7O2VlTa9nAGhdZPFfeQfZCDWjkNBIac7WlDikRNLEH1tqSszUN02ouqRyyCsShDa1kQyUbiApD1IUPfgmHXZxgIMFxr8bwr8BuBa7sK55dYqMRFFbE7YILuBn_rmj7aJI1tp7GAXubODfCUaKvOkoOYjhttps://www.bain.com/vector-digital/partnerships-alliance-ecosystem/openai-alliance/OpenAI plans for AGIhttps://openai.com/blog/planning-for-agi-and-beyondChatGPT influencershttps://www.youtube.com/watch?v=4kp7oVTu9Ckhttps://www.youtube.com/watch?v=k13v8jp8H5ohttps://www.linkedin.com/posts/eniascailliau_create-an-online-course-100-ai-ugcPost-7036969935796891648-H_uj/https://www.linkedin.com/posts/linasbeliunas_must-know-ai-tools-ugcPost-7035700089947836416-Qri4/https://twitter.com/LinusEkenstam/status/1629879567514238976https://www.linkedin.com/posts/imarpit_50-awesome-chatgpt-prompts-ugcPost-7036905788631646209-2CU-/Open-Source Prompting Coursehttps://learnprompting.org/Flan UL2 20Bhttps://www.yitay.net/blog/flan-ul2-20bhttps://huggingface.co/google/flan-ul2LLaMA weights leakedhttps://github.com/facebookresearch/llama/pull/73https://github.com/facebookresearch/llama/pull/73/files#diff-b335630551682c19a781afebcf4d07bf978fb1f8ac04c6bf87428ed5106870f5https://github.com/ChristopherKing42https://open-assistant.io/dashboardMind-Reading from fMRIhttps://sites.google.com/view/stablediffusion-with-brain/?s=09https://www.nature.com/articles/s41562-022-01516-2?utm_content=animationRandom Newshttps://www.wired.com/story/alphabet-layoffs-hit-trash-sorting-robots/https://huggingface.co/blog/fast-mac-diffusershttps://pyribs.org/https://twitter.com/rowancheung/status/1630569844654460928https://pimeyes.com/enhttps://cacti-framework.github.io/https://twitter.com/bhutanisanyam1/status/1630980866775330819https://www.linkedin.com/in/bryancatanzaro/Links:Homepage: https://ykilcher.comMerch: https://ykilcher.com/merchYouTube: https://www.youtube.com/c/yannickilcherTwitter: https://twitter.com/ykilcherDiscord: https://ykilcher.com/discordLinkedIn: https://www.linkedin.com/in/ykilcherIf you want to support me, the best thing to do is to share out the content :)If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):SubscribeStar: https://www.subscribestar.com/yannickilcherPatreon: https://www.patreon.com/yannickilcherBitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cqEthereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9mMonero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n

View Details

ai #meta #languagemodel LLaMA is a series of large language models from 7B to 65B parameters, trained by Meta AI. They train for longer on more data and show that something like gpt-3 can be outperformed by significantly smaller models when trained like this. Meta also releases the trained models to the research community.OUTLINE:0:00 - Introduction & Paper Overview4:30 - Rant on Open-Sourcing8:05 - Training Data12:40 - Training Hyperparameters14:50 - Architecture Modifications17:10 - Optimizer19:40 - Efficient Implementation26:15 - Main Results38:00 - Some more completions40:00 - ConclusionPaper: https://arxiv.org/abs/2302.13971Website: https://ai.facebook.com/blog/large-language-model-llama-meta-ai/Abstract:We introduce LLaMA, a collection of foundation language models ranging from 7B to 65B parameters. We train our models on trillions of tokens, and show that it is possible to train state-of-the-art models using publicly available datasets exclusively, without resorting to proprietary and inaccessible datasets. In particular, LLaMA-13B outperforms GPT-3 (175B) on most benchmarks, and LLaMA-65B is competitive with the best models, Chinchilla-70B and PaLM-540B. We release all our models to the research community.Authors: Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave, Guillaume LampleLinks:Homepage: https://ykilcher.comMerch: https://ykilcher.com/merchYouTube: https://www.youtube.com/c/yannickilcherTwitter: https://twitter.com/ykilcherDiscord: https://ykilcher.com/discordLinkedIn: https://www.linkedin.com/in/ykilcherIf you want to support me, the best thing to do is to share out the content :)If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):SubscribeStar: https://www.subscribestar.com/yannickilcherPatreon: https://www.patreon.com/yannickilcherBitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cqEthereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9mMonero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n

View Details

ai #huggingface #coding Join me as I build streaming inference into the Hugging Face text generation server, going through cuda, python, rust, grpc, websockets, server-sent events, and more...Original repo is here: https://github.com/huggingface/text-generation-inferenceOpenAssistant repo is here: https://github.com/LAION-AI/Open-Assistant (see inference/)Check out https://www.wandb.courses/ for free MLOps courses!Links:Homepage: https://ykilcher.comMerch: https://ykilcher.com/merchYouTube: https://www.youtube.com/c/yannickilcherTwitter: https://twitter.com/ykilcherDiscord: https://ykilcher.com/discordLinkedIn: https://www.linkedin.com/in/ykilcherIf you want to support me, the best thing to do is to share out the content :)If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):SubscribeStar: https://www.subscribestar.com/yannickilcherPatreon: https://www.patreon.com/yannickilcherBitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cqEthereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9mMonero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n

View Details

openassistant #chatgpt #ai Help us collect data for OpenAssistant, the largest and most open alternative to ChatGPT.https://open-assistant.ioOUTLINE:0:00 - Intro0:30 - The Project2:05 - Getting to Minimum Viable Prototype5:30 - First Tasks10:00 - Leaderboard11:45 - Playing the Assistant14:40 - Tricky Facts16:25 - What if humans had wings?17:05 - Can foxes be tamed?23:45 - Can zebras be tamed?26:15 - Yo (spam)27:00 - More tasks29:10 - Entitled Emails34:35 - Final WordsLinks:Homepage: https://ykilcher.comMerch: https://ykilcher.com/merchYouTube: https://www.youtube.com/c/yannickilcherTwitter: https://twitter.com/ykilcherDiscord: https://ykilcher.com/discordLinkedIn: https://www.linkedin.com/in/ykilcherIf you want to support me, the best thing to do is to share out the content :)If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):SubscribeStar: https://www.subscribestar.com/yannickilcherPatreon: https://www.patreon.com/yannickilcherBitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cqEthereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9mMonero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n

View Details

chatgpt #ai #openai

ChatGPT, OpenAI's newest model is a GPT-3 variant that has been fine-tuned using Reinforcement Learning from Human Feedback, and it is taking the world by storm!

Sponsor: Weights & Biases

https://wandb.me/yannic

OUTLINE:

0:00 - Intro

0:40 - Sponsor: Weights & Biases

3:20 - ChatGPT: How does it work?

5:20 - Reinforcement Learning from Human Feedback

7:10 - ChatGPT Origins: The GPT-3.5 Series

8:20 - OpenAI's strategy: Iterative Refinement

9:10 - ChatGPT's amazing capabilities

14:10 - Internals: What we know so far

16:10 - Building a virtual machine in ChatGPT's imagination (insane)

20:15 - Jailbreaks: Circumventing the safety mechanisms

29:25 - How OpenAI sees the future

References:

https://openai.com/blog/chatgpt/

https://openai.com/blog/language-model-safety-and-misuse/

https://beta.openai.com/docs/model-index-for-researchers

https://scale.com/blog/gpt-3-davinci-003-comparison#Conclusion

https://twitter.com/johnvmcdonnell/status/1598470129121374209

https://twitter.com/blennon_/status/1597374826305318912

https://twitter.com/TimKietzmann/status/1598230759118376960/photo/1

https://twitter.com/_lewtun/status/1598056075672027137/photo/2

https://twitter.com/raphaelmilliere/status/1598469100535259136

https://twitter.com/CynthiaSavard/status/1598498138658070530/photo/1

https://twitter.com/tylerangert/status/1598389755997290507/photo/1

https://twitter.com/amasad/status/1598042665375105024/photo/1

https://twitter.com/goodside/status/1598129631609380864/photo/1

https://twitter.com/moyix/status/1598081204846489600/photo/2

https://twitter.com/JusticeRage/status/1598959136531546112

https://twitter.com/yoavgo/status/1598594145605636097

https://twitter.com/EladRichardson/status/1598333315764871174

https://twitter.com/charles_irl/status/1598319027327307785/photo/4

https://twitter.com/jasondebolt/status/1598243854343606273

https://twitter.com/mattshumer_/status/1598185710166896641/photo/1

https://twitter.com/i/web/status/1598246145171804161

https://twitter.com/bleedingedgeai/status/1598378564373471232

https://twitter.com/MasterScrat/status/1598830356115124224

https://twitter.com/Sentdex/status/1598803009844256769

https://twitter.com/harrison_ritz/status/1598828017446371329

https://twitter.com/parafactual/status/1598212029479026689

https://www.engraved.blog/building-a-virtual-machine-inside/

https://twitter.com/317070

https://twitter.com/zehavoc/status/1599193444043268096

https://twitter.com/yoavgo/status/1598360581496459265

https://twitter.com/yoavgo/status/1599037412411596800

https://twitter.com/yoavgo/status/1599045344863879168

https://twitter.com/natfriedman/status/1598477452661383168

https://twitter.com/conradev/status/1598487973351362561/photo/1

https://twitter.com/zswitten/status/1598100186605441024

https://twitter.com/CatEmbedded/status/1599141379879600128/photo/2

https://twitter.com/mattshumer_/status/1599175127148949505

https://twitter.com/vaibhavk97/status/1598930958769860608/photo/1

https://twitter.com/dan_abramov/status/1598800508160024588/photo/1

https://twitter.com/MinqiJiang/status/1598832656422432768/photo/2

https://twitter.com/zswitten/status/1598088280066920453

https://twitter.com/m1guelpf/status/1598203861294252033/photo/1

https://twitter.com/SilasAlberti/status/1598257908567117825/photo/1

https://twitter.com/gf_256/status/1598962842861899776/photo/1

https://twitter.com/zswitten/status/1598088267789787136

https://twitter.com/gf_256/status/1598178469955112961/photo/1

View Details

ai #mlnews #gpt4

Your weekly news from the AI & Machine Learning world.

OUTLINE:

0:00 - Introduction

0:25 - AI reads brain signals to predict what you're thinking

3:00 - Closed-form solution for neuron interactions

4:15 - GPT-4 rumors

6:50 - Cerebras supercomputer

7:45 - Meta releases metagenomics atlas

9:15 - AI advances in theorem proving

10:40 - Better diffusion models with expert denoisers

12:00 - BLOOMZ & mT0

13:05 - ICLR reviewers going mad

21:40 - Scaling Transformer inference

22:10 - Infinite nature flythrough generation

23:55 - Blazing fast denoising

24:45 - Large-scale AI training with MultiRay

25:30 - arXiv to include Hugging Face spaces

26:10 - Multilingual Diffusion

26:30 - Music source separation

26:50 - Multilingual CLIP

27:20 - Drug response prediction

27:50 - Helpful Things

ERRATA:

HF did not acquire spaces, they launched spaces themselves and supported Gradio from the start. They later acquired Gradio.

References:

AI reads brain signals to predict what you're thinking

https://mind-vis.github.io/?s=09&utm_source=pocket_saves

https://neurosciencenews.com/bmi-internal-speech-21837/

Closed-form solution for neuron interactions

https://twitter.com/ramin_m_h/status/1592585672606769153/photo/1

https://github.com/raminmh/CfC

https://github.com/raminmh/CfC/blob/main/torch_cfc.py

GPT-4 rumors

https://thealgorithmicbridge.substack.com/p/gpt-4-rumors-from-silicon-valley?utm_source=pocket_reader

Cerebras supercomputer

https://www.cerebras.net/andromeda/

Meta releases metagenomics atlas

https://ai.facebook.com/blog/protein-folding-esmfold-metagenomics/

https://www.genome.gov/genetics-glossary/Metagenomics

AI advances in theorem proving

https://ai.facebook.com/blog/ai-math-theorem-proving/

https://marketplace.visualstudio.com/items?itemName=jroesch.lean

Better diffusion models with expert denoisers

https://deepimagination.cc/eDiffi/

BLOOMZ & mT0

https://arxiv.org/abs/2211.01786?utm_source=pocket_reader

https://huggingface.co/bigscience/bloomz?text=Suggest+at+least+five+related+search+terms+to+%22M%E1%BA%A1ng+neural+nh%C3%A2n+t%E1%BA%A1o%22.

ICLR reviewers going mad

https://twitter.com/XiangruTang/status/1589703605098975237?utm_source=pocket_reader

https://twitter.com/BlancheMinerva/status/1588164585961422849?utm_source=pocket_reader

https://openreview.net/forum?id=pfuqQQCB34

https://twitter.com/peter_richtarik/status/1591408710366408706?utm_source=pocket_reader

Scaling Transformer inference

https://arxiv.org/abs/2211.05102

Infinite nature flythrough generation

https://ai.googleblog.com/2022/11/infinite-nature-generating-3d.html?utm_source=pocket_reader

Blazing fast denoising

https://github.com/dome272/Paella

https://arxiv.org/abs/2211.07292

Large-scale AI training with MultiRay

https://ai.facebook.com/blog/multiray-large-scale-AI-models/

arXiv to include Hugging Face spaces

https://blog.arxiv.org/2022/11/17/discover-state-of-the-art-machine-learning-demos-on-arxiv/

Multilingual Diffusion

https://github.com/FlagAI-Open/FlagAI/tree/master/examples/AltDiffusion

Music source separation

https://github.com/facebookresearch/demucs

https://arxiv.org/abs/2211.08553

View Details

ai #cicero #diplomacy

A team from Meta AI has developed Cicero, an agent that can play the game Diplomacy, in which players have to communicate via chat messages to coordinate and plan into the future.

Paper Title: Human-level play in the game of Diplomacy by combining language models with strategic reasoning

Commented game by human expert: https://www.youtube.com/watch?v=u5192bvUS7k

OUTLINE:

0:00 - Introduction

9:50 - AI in cooperation games

13:50 - Cicero agent overview

25:00 - A controllable dialogue model

36:50 - Dialogue-conditional strategic planning

49:00 - Message filtering

53:45 - Cicero's play against humans

55:15 - More examples & discussion

Homepage: https://ai.facebook.com/research/cicero/

Code: https://github.com/facebookresearch/diplomacy_cicero

Blog: https://ai.facebook.com/blog/cicero-ai-negotiates-persuades-and-cooperates-with-people/

Paper: https://www.science.org/doi/10.1126/science.ade9097

Abstract:

Despite much progress in training AI systems to imitate human language, building agents that use language to communicate intentionally with humans in interactive environments remains a major challenge. We introduce Cicero, the first AI agent to achieve human-level performance in Diplomacy, a strategy game involving both cooperation and competition that emphasizes natural language negotiation and tactical coordination between seven players. Cicero integrates a language model with planning and reinforcement learning algorithms by inferring players' beliefs and intentions from its conversations and generating dialogue in pursuit of its plans. Across 40 games of an anonymous online Diplomacy league, Cicero achieved more than double the average score of the human players and ranked in the top 10% of participants who played more than one game.

Authors: Anton Bakhtin, Noam Brown, Emily Dinan, Gabriele Farina, Colin Flaherty, Daniel Fried, Andrew Goff, Jonathan Gray, Hengyuan Hu, Athul Paul Jacob, Mojtaba Komeili, Karthik Konath, Minae Kwon, Adam Lerer, Mike Lewis, Alexander H. Miller, Sasha Mitts, Adithya Renduchintala, Stephen Roller, Dirk Rowe, Weiyan Shi, Joe Spisak, Alexander Wei, David Wu, Hugh Zhang, Markus Zijlstra

Links:

Homepage: https://ykilcher.com

Merch: https://ykilcher.com/merch

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://ykilcher.com/discord

LinkedIn: https://www.linkedin.com/in/ykilcher

If you want to support me, the best thing to do is to share out the content :)

If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https://www.subscribestar.com/yannickilcher

Patreon: https://www.patreon.com/yannickilcher

Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq

Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2

Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m

Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n

View Details

mlnews #ai #mlinpl

Your news from the world of Machine Learning!

OUTLINE:

0:00 - Introduction

1:25 - Stable Diffusion Multiplayer

2:15 - Huggingface: DOI for Models & Datasets

3:10 - OpenAI asks for more funding

4:25 - The Stack: Source Code Dataset

6:30 - Google Vizier Open-Sourced

7:10 - New Models

11:50 - Helpful Things

20:30 - Prompt Databases

22:15 - Lexicap by Karpathy

References:

Stable Diffusion Multiplayer

https://huggingface.co/spaces/huggingface-projects/stable-diffusion-multiplayer?roomid=room-0

Huggingface: DOI for Models & Datasets

https://huggingface.co/blog/introducing-doi

OpenAI asks for more funding

https://www.theinformation.com/articles/openai-valued-at-nearly-20-billion-in-advanced-talks-with-microsoft-for-more-funding

https://www.wsj.com/articles/microsoft-in-advanced-talks-to-increase-investment-in-openai-11666299548

The Stack: Source Code Dataset

https://huggingface.co/datasets/bigcode/the-stack?utm_source=pocket_mylist

Google Vizier Open-Sourced

https://github.com/google/vizier

New Models

https://imagen.research.google/video/

https://phenaki.github.io/

https://makeavideo.studio/?utm_source=pocket_mylist

https://dreamfusion3d.github.io/

https://arxiv.org/pdf/2210.15257.pdf

https://huggingface.co/spaces/PaddlePaddle/ERNIE-ViLG

https://github.com/PaddlePaddle/PaddleHub

Helpful Things

https://thecharlieblake.co.uk/visualising-ml-number-formats

https://griddly.ai/

https://engineering.fb.com/2022/10/18/open-source/ocp-summit-2022-grand-teton/?utm_source=twitter&utm_medium=organic_social&utm_campaign=eng2022h2

https://twitter.com/psuraj28/status/1580640841583902720?utm_source=pocket_mylist

https://huggingface.co/blog/stable_diffusion_jax

https://github.com/Lightning-AI/stable-diffusion-deploy

https://lightning.ai/docs/stable/

https://github.com/CarperAI/trlx

https://github.com/DLR-RM/rl-baselines3-zoo

https://github.com/Sea-Snell/JAXSeq

https://www.reddit.com/r/MachineLearning/comments/xoitw9/p_albumentations_13_is_released_a_python_library/?utm_source=pocket_mylist

https://twitter.com/Warvito/status/1570691960792580096?utm_source=pocket_mylist

https://arxiv.org/abs/2209.07162

https://academictorrents.com/details/63aeb864bbe2115ded0aa0d7d36334c026f0660b

https://huggingface.co/spaces/THUDM/CodeGeeX

https://ai.facebook.com/blog/gpu-inference-engine-nvidia-amd-open-source/?utm_source=twitter&utm_medium=organic_social&utm_campaign=blog

https://github.com/nerfstudio-project/nerfstudio

https://www.nerfacc.com/en/latest/

https://github.com/dstackai/dstack

https://www.reddit.com/r/MachineLearning/comments/yeyxlo/p_openai_whisper_3x_cpu_inference_speedup/?utm_source=pocket_mylist

https://github.com/MiscellaneousStuff/openai-whisper-cpu/issues/1

Prompt Databases

https://huggingface.co/datasets/poloclub/diffusiondb

https://publicprompts.art/

https://visualise.ai/

https://twitter.com/SamuelAlbanie/status/1574111928431026179/photo/1

Lexicap by Karpathy

https://karpathy.ai/lexicap/0139-large.html

Links:

Homepage: https://ykilcher.com

Merch: https://ykilcher.com/merch

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://ykilcher.com/discord

LinkedIn: https://www.linkedin.com/in/ykilcher

If you want to support me, the best thing to do is to share out the content :)

View Details

ai #stablediffusion #license

So-called responsible AI licenses are stupid, counterproductive, and have a dangerous legal loophole in them.

OpenRAIL++ License here: https://www.ykilcher.com/license

OUTLINE:

0:00 - Introduction

0:40 - Responsible AI Licenses (RAIL) of BLOOM and Stable Diffusion

3:35 - Open source software's dilemma of bad usage and restrictions

8:45 - Good applications, bad applications

12:45 - A dangerous legal loophole

15:50 - OpenRAIL++ License

16:50 - This has nothing to do with copyright

26:00 - Final thoughts

References:

https://huggingface.co/CompVis/stable-diffusion/tree/main

https://huggingface.co/spaces/CompVis/stable-diffusion-license

https://huggingface.co/bigscience/bloom?text=34%2B10%3D44+%0A54%2B20%3D

https://huggingface.co/spaces/bigscience/license

https://huggingface.co/runwayml/stable-diffusion-v1-5

https://huggingface.co/spaces/CompVis/stable-diffusion-license/raw/main/license.txt

https://www.gnu.org/philosophy/programs-must-not-limit-freedom-to-run.en.html

https://www.gnu.org/philosophy/free-sw.html#four-freedoms

https://www.licenses.ai/blog/2022/8/26/bigscience-open-rail-m-license

https://bigscience.huggingface.co/blog/bigscience-ethical-charter

https://www.licenses.ai/blog/2022/8/18/naming-convention-of-responsible-ai-licenses

https://en.wikipedia.org/wiki/Copyright#Eligible_works

https://en.wikipedia.org/wiki/Creative_work

https://www.pearlcohen.com/copyright-office-reiterates-that-works-created-by-ai-cannot-be-copyrighted/

https://jipel.law.nyu.edu/vol-8-no-2-1-hedrick/#II

https://www.ykilcher.com/license

Links:

Homepage: https://ykilcher.com

Merch: https://ykilcher.com/merch

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://ykilcher.com/discord

LinkedIn: https://www.linkedin.com/in/ykilcher

If you want to support me, the best thing to do is to share out the content :)

If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https://www.subscribestar.com/yannickilcher

Patreon: https://www.patreon.com/yannickilcher

Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq

Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2

Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m

Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n

View Details

ai #language #knowledge

Large Language Models have the ability to store vast amounts of facts about the world. But little is known, how these models actually do this. This paper aims at discovering the mechanism and location of storage and recall of factual associations in GPT models, and then proposes a mechanism for the targeted editing of such facts, in form of a simple rank-one update to a single MLP layer. This has wide implications both for how we understand such models' inner workings, and for our ability to gain greater control over such models in the future.

OUTLINE:

0:00 - Introduction

1:40 - What are the main questions in this subfield?

6:55 - How causal tracing reveals where facts are stored

18:40 - Clever experiments show the importance of MLPs

24:30 - How do MLPs store information?

29:10 - How to edit language model knowledge with precision?

36:45 - What does it mean to know something?

39:00 - Experimental Evaluation & the CounterFact benchmark

45:40 - How to obtain the required latent representations?

51:15 - Where is the best location in the model to perform edits?

58:00 - What do these models understand about language?

1:02:00 - Questions for the community

Paper: https://arxiv.org/abs/2202.05262

Follow-up paper on Mass-Editing Memory in a Transformer: https://arxiv.org/abs/2210.07229

Abstract:

We analyze the storage and recall of factual associations in autoregressive transformer language models, finding evidence that these associations correspond to localized, directly-editable computations. We first develop a causal intervention for identifying neuron activations that are decisive in a model's factual predictions. This reveals a distinct set of steps in middle-layer feed-forward modules that mediate factual predictions while processing subject tokens. To test our hypothesis that these computations correspond to factual association recall, we modify feed-forward weights to update specific factual associations using Rank-One Model Editing (ROME). We find that ROME is effective on a standard zero-shot relation extraction (zsRE) model-editing task, comparable to existing methods. To perform a more sensitive evaluation, we also evaluate ROME on a new dataset of counterfactual assertions, on which it simultaneously maintains both specificity and generalization, whereas other methods sacrifice one or another. Our results confirm an important role for mid-layer feed-forward modules in storing factual associations and suggest that direct manipulation of computational mechanisms may be a feasible approach for model editing. The code, dataset, visualizations, and an interactive demo notebook are available at this https URL

Authors: Kevin Meng, David Bau, Alex Andonian, Yonatan Belinkov

Links:

Homepage: https://ykilcher.com

Merch: https://ykilcher.com/merch

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://ykilcher.com/discord

LinkedIn: https://www.linkedin.com/in/ykilcher

If you want to support me, the best thing to do is to share out the content :)

If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https://www.subscribestar.com/yannickilcher

Patreon: https://www.patreon.com/yannickilcher

Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq

Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2

Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m

Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n

View Details

neuralnetworks #machinelearning #ai

Alexander Mattick joins me to discuss the paper "Neural Networks are Decision Trees", which has generated a lot of hype on social media. We ask the question: Has this paper solved one of the large mysteries of deep learning and opened the black-box neural networks up to interpretability?

OUTLINE:

0:00 - Introduction

2:20 - Aren't Neural Networks non-linear?

5:20 - What does it all mean?

8:00 - How large do these trees get?

11:50 - Decision Trees vs Neural Networks

17:15 - Is this paper new?

22:20 - Experimental results

27:30 - Can Trees and Networks work together?

Paper: https://arxiv.org/abs/2210.05189

Abstract:

In this manuscript, we show that any feedforward neural network having piece-wise linear activation functions can be represented as a decision tree. The representation is equivalence and not an approximation, thus keeping the accuracy of the neural network exactly as is. We believe that this work paves the way to tackle the black-box nature of neural networks. We share equivalent trees of some neural networks and show that besides providing interpretability, tree representation can also achieve some computational advantages. The analysis holds both for fully connected and convolutional networks, which may or may not also include skip connections and/or normalizations.

Author: Caglar Aytekin

Links:

Homepage: https://ykilcher.com

Merch: https://ykilcher.com/merch

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://ykilcher.com/discord

LinkedIn: https://www.linkedin.com/in/ykilcher

If you want to support me, the best thing to do is to share out the content :)

If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https://www.subscribestar.com/yannickilcher

Patreon: https://www.patreon.com/yannickilcher

Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq

Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2

Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m

Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n

View Details

alphatensor #deepmind #ai

Matrix multiplication is the most used mathematical operation in all of science and engineering. Speeding this up has massive consequences. Thus, over the years, this operation has become more and more optimized. A fascinating discovery was made when it was shown that one actually needs less than N^3 multiplication operations to multiply to NxN matrices. DeepMind goes a step further and creates AlphaTensor, a Deep Reinforcement Learning algorithm that plays a single-player game, TensorGame, in order to find even more optimized algorithms for matrix multiplication. And it turns out, there exists a plethora of undiscovered matrix multiplication algorithms, which not only will make everything from computers to smart toasters faster, but also bring new insights into fundamental math and complexity theory.

Sponsor: Assembly AI

Link: https://www.assemblyai.com/?utm_source=youtube&utm_medium=social&utm_campaign=yannic_sentiment

OUTLINE:

0:00 - Intro

1:50 - Sponsor: Assembly AI (link in description)

3:25 - What even is Matrix Multiplication?

6:10 - A very astounding fact

8:45 - Trading multiplications for additions

12:35 - Matrix Multiplication as a Tensor

17:30 - Tensor Decompositions

20:30 - A formal way of finding multiplication algorithms

31:00 - How to formulate this as a game?

39:30 - A brief primer on AlphaZero / MCTS

45:40 - The Results

48:15 - Optimizing for different hardware

52:40 - Expanding fundamental math

53:45 - Summary & Final Comments

Paper: https://www.nature.com/articles/s41586-022-05172-4

Title: Discovering faster matrix multiplication algorithms with reinforcement learning

Abstract:

Improving the efficiency of algorithms for fundamental computations can have a widespread impact, as it can affect the overall speed of a large amount of computations. Matrix multiplication is one such primitive task, occurring in many systems—from neural networks to scientific computing routines. The automatic discovery of algorithms using machine learning offers the prospect of reaching beyond human intuition and outperforming the current best human-designed algorithms. However, automating the algorithm discovery procedure is intricate, as the space of possible algorithms is enormous. Here we report a deep reinforcement learning approach based on AlphaZero1 for discovering efficient and provably correct algorithms for the multiplication of arbitrary matrices. Our agent, AlphaTensor, is trained to play a single-player game where the objective is finding tensor decompositions within a finite factor space. AlphaTensor discovered algorithms that outperform the state-of-the-art complexity for many matrix sizes. Particularly relevant is the case of 4 × 4 matrices in a finite field, where AlphaTensor’s algorithm improves on Strassen’s two-level algorithm for the first time, to our knowledge, since its discovery 50 years ago2. We further showcase the flexibility of AlphaTensor through different use-cases: algorithms with state-of-the-art complexity for structured matrix multiplication and improved practical efficiency by optimizing matrix multiplication for runtime on specific hardware. Our results highlight AlphaTensor’s ability to accelerate the process of algorithmic discovery on a range of problems, and to optimize for different criteria.

Authors: Alhussein Fawzi, Matej Balog, Aja Huang, Thomas Hubert, Bernardino Romera-Paredes, Mohammadamin Barekatain, Alexander Novikov, Francisco J. R. Ruiz, Julian Schrittwieser, Grzegorz Swirszcz, David Silver, Demis Hassabis & Pushmeet Kohli

View Details

stablediffusion #aiart #mlnews

Stable Diffusion has been released and is riding a wave of creativity and collaboration. But not everyone is happy about this...

Sponsor: NVIDIA

GPU Raffle: https://ykilcher.com/gtc

OUTLINE:

0:00 - Introduction

0:30 - What is Stable Diffusion?

2:25 - Open-Source Contributions and Creations

7:55 - Textual Inversion

9:30 - OpenAI vs Open AI

14:20 - Journalists be outraged

16:20 - AI Ethics be even more outraged

19:45 - Do we need a new social contract?

21:30 - More applications

22:55 - Helpful Things

23:45 - Sponsor: NVIDIA (& how to enter the GPU raffle)

References: https://early-hair-c20.notion.site/Stable-Diffusion-Takes-Over-Referenes-7a2f45b8f7e04ae0ba19dbfcd2b7f7c0

Links:

Homepage: https://ykilcher.com

Merch: https://ykilcher.com/merch

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://ykilcher.com/discord

LinkedIn: https://www.linkedin.com/in/ykilcher

If you want to support me, the best thing to do is to share out the content :)

If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https://www.subscribestar.com/yannickilcher

Patreon: https://www.patreon.com/yannickilcher

Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq

Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2

Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m

Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n

View Details

ai #sparsity #gpu

Sparsity is awesome, but only recently has it become possible to properly handle sparse models at good performance. Neural Magic does exactly this, using a plain CPU. No specialized hardware needed, just clever algorithms for pruning and forward-propagation of neural networks. Nir Shavit and I talk about how this is possible, what it means in terms of applications, and why sparsity should play a much larger role in the Deep Learning community.

Sponsor: AssemblyAI

Link: https://www.assemblyai.com/?utm_sourc...

Check out Neural Magic: https://neuralmagic.com/

and DeepSparse: https://github.com/neuralmagic/deepsp...

OUTLINE:

0:00 Introduction

1:08 Sponsor: AssemblyAI

2:50 Start of Interview

4:15 How the NIR company was founded?

5:10 What is Sparsity about?

9:30 Link between the human brain and sparsity

12:10 Where should the extra resource that the human brain doesn't have go?

14:40 Analogy for Sparse Architecture

16:48 Possible future for Sparse Architecture as standard architure for Neural Networks

20:08 Pruning & Sparsification

22:57 What keeps us from building sparse models?

25:34 Why are GPUs so unsuited for sparse models?

28:47 CPU and GPU in connection with memory

30:14 What Neural Magic does?

32:54 How do you deal with overlaps in tensor columns?

33:41 The best type of sparsity to execute tons of CPU

37:24 What kind of architecture would make the best use out of a combined system of CPUs and GPUs?

41:04 Graph Neural Networks in connection to sparsity

43:04 Intrinsic connection between the Sparsification of Neural Networks, Non Layer-Wise Computation, Blockchain Technology, Smart Contracts and Distributed Computing

45:23 Neural Magic's target audience

48:16 Is there a type of model where it works particularly well and the type where it doesn't?

Links:

Homepage: https://ykilcher.com

Merch: https://ykilcher.com/merch

YouTube:

/ yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://ykilcher.com/discord

LinkedIn: https://www.linkedin.com/in/ykilcher

If you want to support me, the best thing to do is to share out the content :)

If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https://www.subscribestar.com/yannick...

Patreon: https://www.patreon.com/yannickilcher

Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq

Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2

Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m

Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n

View Details

ai #interview #research

Jacob Steinhardt believes that future AI systems will be qualitatively different than the ones we know currently. We talk about how emergence happens when scaling up, what implications that has on AI Safety, and why thought experiments like the Paperclip Maximizer might be more useful than most people think.

OUTLINE:

0:00 Introduction

1:10 Start of Interview

2:10 Blog posts series

3:56 More Is Different for AI (Blog Post)

7:40 Do you think this emergence is mainly a property from the interaction of things?

9:17 How does phase transition or scaling-up play into AI and Machine Learning?

12:10 GPT-3 as an example of qualitative difference in scaling up

14:08 GPT-3 as an emergent phenomenon in context learning

15:58 Brief introduction of different viewpoints on the future of AI and its alignment

18:51 How does the phenomenon of emergence play into this game between the Engineering and the Philosophy viewpoint?

22:41 Paperclip Maximizer on AI safety and alignment

31:37 Thought Experiments

37:34 Imitative Deception

39:30 TruthfulQA: Measuring How Models Mimic Human Falsehoods (Paper)

42:24 ML Systems Will Have Weird Failure Models (Blog Post)

51:10 Is there any work to get a system to be deceptive?

54:37 Empirical Findings Generalize Surprisingly Far (Blog Post)

1:00:18 What would you recommend to guarantee better AI alignment or safety?

1:05:13 Remarks

References:

https://bounded-regret.ghost.io/more-is-different-for-ai/

https://docs.google.com/document/d/1FbTuRvC4TFWzGYerTKpBU7FJlyvjeOvVYF2uYNFSlOc/edit#heading=h.n1wk9bxo847o

Links:

TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://ykilcher.com/discord

BitChute: https://www.bitchute.com/channel/yannic-kilcher

LinkedIn: https://www.linkedin.com/in/ykilcher

BiliBili: https://space.bilibili.com/2017636191

If you want to support me, the best thing to do is to share out the content :)

If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https://www.subscribestar.com/yannickilcher

Patreon: https://www.patreon.com/yannickilcher

Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq

Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2

Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m

Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n

View Details

huggingface #pickle #exploit

Did you know that something as simple as loading a model can execute arbitrary code on your machine?

Try the model: https://huggingface.co/ykilcher/total...

Get the code: https://github.com/yk/patch-torch-save

Sponsor: Weights & Biases

Go here: https://wandb.me/yannic

OUTLINE:

0:00 - Introduction

1:10 - Sponsor: Weights & Biases

3:20 - How Hugging Face models are loaded

5:30 - From PyTorch to pickle

7:10 - Understanding how pickle saves data

13:00 - Executing arbitrary code

15:05 - The final code

17:25 - How can you protect yourself?

Links:

Homepage: https://ykilcher.com

Merch: https://ykilcher.com/merch

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://ykilcher.com/discord

LinkedIn: https://www.linkedin.com/in/ykilcher

If you want to support me, the best thing to do is to share out the content :)

If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https://www.subscribestar.com/yannick...

Patreon: https://www.patreon.com/yannickilcher

Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq

Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2

Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m

Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n

View Details

#ai #selforganization #emergence

Read Sebastian's article here: https://sebastianrisi.com/self_assemb...

OUTLINE:

0:00 - Introduction

2:25 - Start of Interview

4:00 - The intelligence of swarms

9:15 - The game of life & neural cellular automata

14:10 - What's missing from neural CAs?

17:20 - How does local computation compare to centralized computation?

25:40 - Applications beyond games and graphics

33:00 - Can we do away with goals?

35:30 - Where do these methods shine?

43:30 - The paradox of scales & brains

49:45 - Connections to graphical systems & GNNs

51:30 - Could this solve ARC?

57:45 - Where can people get started?

References:

https://sebastianrisi.com/

https://modl.ai/

https://sebastianrisi.com/self_assemb...

https://twitter.com/risi1979/status/1...

https://distill.pub/2020/growing-ca/

https://arxiv.org/abs/2201.12360?sour...

https://distill.pub/2020/selforg/mnist/

https://arxiv.org/pdf/2204.11674.pdf

https://github.com/fchollet/ARC

https://github.com/volotat/ARC-Game

http://animalaiolympics.com/AAI/

https://www.deepmind.com/publications...

https://melaniemitchell.me/BooksConte...

Links:

Homepage: https://ykilcher.com

Merch: https://ykilcher.com/merch

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://ykilcher.com/discord

LinkedIn: https://www.linkedin.com/in/ykilcher

If you want to support me, the best thing to do is to share out the content :)

If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https://www.subscribestar.com/yannick...

Patreon: https://www.patreon.com/yannickilcher

Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq

Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2

Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m

Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n

View Details

#stablediffusion #ai #stabilityai

An interview with Emad Mostaque, founder of Stability AI.

OUTLINE:

0:00 - Intro

1:30 - What is Stability AI?

3:45 - Where does the money come from?

5:20 - Is this the CERN of AI?

6:15 - Who gets access to the resources?

8:00 - What is Stable Diffusion?

11:40 - What if your model produces bad outputs?

14:20 - Do you employ people?

16:35 - Can you prevent the corruption of profit?

19:50 - How can people find you?

22:45 - Final thoughts, let's destroy PowerPoint

Links:

Homepage: https://ykilcher.com

Merch: https://ykilcher.com/merch

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://ykilcher.com/discord

LinkedIn: https://www.linkedin.com/in/ykilcher

If you want to support me, the best thing to do is to share out the content :)

If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https://www.subscribestar.com/yannick...

Patreon: https://www.patreon.com/yannickilcher

Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq

Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2

Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m

Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n

View Details

mlnews #bloom #ai

Today we look at all the recent giant language models in the AI world!

OUTLINE:

0:00 - Intro

0:55 - BLOOM: Open-Source 176B Language Model

5:25 - YALM 100B

5:40 - Chinese Brain-Scale Supercomputer

7:25 - Meta AI Translates over 200 Languages

10:05 - Reproducibility Crisis Workshop

10:55 - AI21 Raises $64M

11:50 - Ian Goodfellow leaves Apple

12:20 - Andrej Karpathy leaves Tesla

12:55 - Wordalle

References:

BLOOM: Open-Source 176B Language Model

https://bigscience.huggingface.co/blo...

https://huggingface.co/spaces/bigscie...

https://huggingface.co/bigscience/blo...

YALM 100B

https://github.com/yandex/YaLM-100B

Chinese Brain-Scale Supercomputer

https://www.scmp.com/news/china/scien...

https://archive.ph/YaoA6#selection-12...

Meta AI Translates over 200 Languages

https://ai.facebook.com/research/no-l...

Reproducibility Crisis Workshop

https://reproducible.cs.princeton.edu/

AI21 Raises $64M

https://techcrunch.com/2022/07/12/ope...

Ian Goodfellow leaves Apple

https://twitter.com/goodfellow_ian/st...

Andrey Karpathy leaves Tesla

https://mobile.twitter.com/karpathy/s...

https://www.businessinsider.com/repor...

Wordalle

https://huggingface.co/spaces/hugging...

Links:

Homepage: https://ykilcher.com

Merch: ykilcher.com/merch

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://ykilcher.com/discord

LinkedIn: https://www.linkedin.com/in/ykilcher

If you want to support me, the best thing to do is to share out the content :)

If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https://www.subscribestar.com/yannick...

Patreon: https://www.patreon.com/yannickilcher

Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq

Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2

Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m

Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n

View Details

Yann LeCun's position paper on a path towards machine intelligence combines Self-Supervised Learning, Energy-Based Models, and hierarchical predictive embedding models to arrive at a system that can teach itself to learn useful abstractions at multiple levels and use that as a world model to plan ahead in time.

OUTLINE:

0:00 - Introduction

2:00 - Main Contributions

5:45 - Mode 1 and Mode 2 actors

15:40 - Self-Supervised Learning and Energy-Based Models

20:15 - Introducing latent variables

25:00 - The problem of collapse

29:50 - Contrastive vs regularized methods

36:00 - The JEPA architecture

47:00 - Hierarchical JEPA (H-JEPA)

53:00 - Broader relevance

56:00 - Summary & Comments

Paper: https://openreview.net/forum?id=BZ5a1...

Abstract: How could machines learn as efficiently as humans and animals? How could machines learn to reason and plan? How could machines learn representations of percepts and action plans at multiple levels of abstraction, enabling them to reason, predict, and plan at multiple time horizons? This position paper proposes an architecture and training paradigms with which to construct autonomous intelligent agents. It combines concepts such as configurable predictive world model, behavior driven through intrinsic motivation, and hierarchical joint embedding architectures trained with self-supervised learning.

Author: Yann LeCun

Links:

Homepage: https://ykilcher.com

Merch: https://ykilcher.com/merch

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://ykilcher.com/discord

LinkedIn: https://www.linkedin.com/in/ykilcher

If you want to support me, the best thing to do is to share out the content :)

If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https://www.subscribestar.com/yannick...

Patreon: https://www.patreon.com/yannickilcher

Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq

Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2

Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m

Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n

View Details

openai #vpt #minecraft

Minecraft is one of the harder challenges any RL agent could face. Episodes are long, and the world is procedurally generated, complex, and huge. Further, the action space is a keyboard and a mouse, which has to be operated only given the game's video input. OpenAI tackles this challenge using Video PreTraining, leveraging a small set of contractor data in order to pseudo-label a giant corpus of scraped footage of gameplay. The pre-trained model is highly capable in basic game mechanics and can be fine-tuned much better than a blank slate model. This is the first Minecraft agent that achieves the elusive goal of crafting a diamond pickaxe all by itself.

OUTLINE:

0:00 - Intro

3:50 - How to spend money most effectively?

8:20 - Getting a large dataset with labels

14:40 - Model architecture

19:20 - Experimental results and fine-tuning

25:40 - Reinforcement Learning to the Diamond Pickaxe

30:00 - Final comments and hardware

Blog: https://openai.com/blog/vpt/

Paper: https://arxiv.org/abs/2206.11795

Code & Model weights: https://github.com/openai/Video-Pre-T...

Abstract:

Pretraining on noisy, internet-scale datasets has been heavily studied as a technique for training models with broad, general capabilities for text, images, and other modalities. However, for many sequential decision domains such as robotics, video games, and computer use, publicly available data does not contain the labels required to train behavioral priors in the same way. We extend the internet-scale pretraining paradigm to sequential decision domains through semi-supervised imitation learning wherein agents learn to act by watching online unlabeled videos. Specifically, we show that with a small amount of labeled data we can train an inverse dynamics model accurate enough to label a huge unlabeled source of online data -- here, online videos of people playing Minecraft -- from which we can then train a general behavioral prior. Despite using the native human interface (mouse and keyboard at 20Hz), we show that this behavioral prior has nontrivial zero-shot capabilities and that it can be fine-tuned, with both imitation learning and reinforcement learning, to hard-exploration tasks that are impossible to learn from scratch via reinforcement learning. For many tasks our models exhibit human-level performance, and we are the first to report computer agents that can craft diamond tools, which can take proficient humans upwards of 20 minutes (24,000 environment actions) of gameplay to accomplish.

Authors: Bowen Baker, Ilge Akkaya, Peter Zhokhov, Joost Huizinga, Jie Tang, Adrien Ecoffet, Brandon Houghton, Raul Sampedro, Jeff Clune

Links:

Homepage: https://ykilcher.com

Merch: https://ykilcher.com/merch

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://ykilcher.com/discord

LinkedIn: https://www.linkedin.com/in/ykilcher

If you want to support me, the best thing to do is to share out the content :)

If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https://www.subscribestar.com/yannick...

Patreon: https://www.patreon.com/yannickilcher

Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq

Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2

Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m

Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n

View Details

#parti #ai #aiart

Parti is a new autoregressive text-to-image model that shows just how much scale can achieve. This model's outputs are crips, accurate, realistic, and can combine arbitrary styles, concepts, and fulfil even challenging requests.

OUTLINE:

0:00 - Introduction

2:40 - Example Outputs

6:00 - Model Architecture

17:15 - Datasets (incl. PartiPrompts)

21:45 - Experimental Results

27:00 - Picking a cherry tree

29:30 - Failure cases

33:20 - Final comments

Website: https://parti.research.google/

Paper: https://arxiv.org/abs/2206.10789

Github: https://github.com/google-research/parti

Links:

Homepage: https://ykilcher.com

Merch: https://ykilcher.com/merch

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://ykilcher.com/discord

LinkedIn: https://www.linkedin.com/in/ykilcher

If you want to support me, the best thing to do is to share out the content :)

If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https://www.subscribestar.com/yannick...

Patreon: https://www.patreon.com/yannickilcher

Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq

Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2

Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m

Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n

View Details

lamda #google #ai

Google engineer Blake Lemoine was put on leave after releasing proprietary information: An interview with the chatbot LaMDA that he believes demonstrates that this AI is, in fact, sentient. We analyze the claims and the interview in detail and trace how a statistical machine managed to convince at least one human that it is more than just an algorithm.

OUTLINE:

0:00 - Whistleblower put on leave

4:30 - What is a language model?

6:40 - The prompt is the key

10:40 - Who are we talking to exactly?

12:50 - LaMDA analyzes stories

15:20 - Fear, pain, and consent

20:25 - How would we recognize sentience? When is a machine conscious?

References:

https://cajundiscordian.medium.com/is-lamda-sentient-an-interview-ea64d916d917

https://cajundiscordian.medium.com/what-is-lamda-and-what-does-it-want-688632134489

https://www.washingtonpost.com/technology/2022/06/11/google-ai-lamda-blake-lemoine/

https://www.theguardian.com/technology/2022/jun/12/google-engineer-ai-bot-sentient-blake-lemoine

https://www.businessinsider.com/transcript-of-sentient-google-ai-chatbot-was-edited-for-readability-2022-6?inline-endstory-related-recommendations=&r=US&IR=T

Links:

Homepage: https://ykilcher.com

Merch: https://ykilcher.com/merch

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://ykilcher.com/discord

BitChute: https://www.bitchute.com/channel/yannic-kilcher

LinkedIn: https://www.linkedin.com/in/ykilcher

If you want to support me, the best thing to do is to share out the content :)

If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https://www.subscribestar.com/yannickilcher

Patreon: https://www.patreon.com/yannickilcher

Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq

Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2

Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m

Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n

View Details

Your updates directly from the state of the art in Machine Learning!

OUTLINE:

0:00 - Intro

0:30 - DeepMind's Flamingo: Unified Vision-Language Model

8:25 - LiT: Locked Image Tuning

10:20 - Jurassic X & MRKL Systems

15:05 - Helpful Things

22:40 - This AI does not exist

References:

DeepMind's Flamingo: Unified Vision-Language Model

https://www.deepmind.com/blog/tacklin...

https://storage.googleapis.com/deepmi...

https://twitter.com/Inoryy/status/152...

LiT: Locked Image Tuning

https://ai.googleblog.com/2022/04/loc...

https://google-research.github.io/vis...

Jurassic X & MRKL Systems

https://www.ai21.com/blog/jurassic-x-...

https://arxiv.org/pdf/2205.00445.pdf

https://arxiv.org/pdf/2204.10019.pdf

https://studio.ai21.com/jurassic-x

StyleGAN Human

https://stylegan-human.github.io/

https://github.com/stylegan-human/Sty...

https://huggingface.co/spaces/hysts/S...

Helpful Things

https://github.com/rish-16/grafog

https://huggingface.co/bertin-project...

https://github.com/pytorch/torchdistx

https://pytorch.org/torchdistx/latest...

https://github.com/Netflix/vectorflow...

https://iclr-blog-track.github.io/202...

https://twitter.com/DeepMind/status/1...

https://github.com/ai-forever/mgpt

https://github.com/cleanlab/cleanlab

https://efficientdlbook.com/?utm_sour...

https://minihack-editor.github.io/

https://mugen-org.github.io/

https://www.amazon.science/blog/amazo...

https://github.com/phuselab/openFACS?...

https://medium.com/pytorch/avalanche-...

This AI does not exist

https://thisaidoesnotexist.com/

Links:

Merch: https://ykilcher.com/merch

TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://discord.gg/4H8xxDF

BitChute: https://www.bitchute.com/channel/yann...

LinkedIn: https://www.linkedin.com/in/ykilcher

BiliBili: https://space.bilibili.com/2017636191

If you want to support me, the best thing to do is to share out the content :)

If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https://www.subscribestar.com/yannick...

Patreon: https://www.patreon.com/yannickilcher

Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq

Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2

Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m

Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n

View Details

mlnews #dalle #gpt3

An inside look of what's happening in the ML world!

Sponsor: Weights & Biases

https://wandb.me/yannic

OUTLINE:

0:00 - Intro

0:20 - Sponsor: Weights & Biases

1:40 - Meta AI releases OPT-175B

4:55 - CoCa: New CLIP-Competitor

8:15 - DALL-E Mega is training

10:05 - TorToiSe TTS is amazing!

11:50 - Investigating Vision Transformers

12:50 - Hugging Face Deep RL class launched

13:40 - Helpful Things

17:00 - John Deere's driverless tractors

References:

Meta AI releases OPT-175B

https://ai.facebook.com/blog/democratizing-access-to-large-scale-language-models-with-opt-175b/

https://arxiv.org/abs/2205.01068

https://arxiv.org/pdf/2205.01068.pdf

https://github.com/facebookresearch/metaseq/tree/main/projects/OPT

https://github.com/facebookresearch/metaseq/blob/main/projects/OPT/chronicles/OPT175B_Logbook.pdf

https://github.com/facebookresearch/metaseq/tree/main/projects/OPT/chronicles

https://twitter.com/yoavgo/status/1522150063815987201

CoCa: New CLIP-Competitor

https://arxiv.org/abs/2205.01917

https://arxiv.org/pdf/2205.01917.pdf

DALL-E Mega is training

https://twitter.com/borisdayma

https://twitter.com/borisdayma/status/1521891895001112577

https://wandb.ai/dalle-mini/dalle-mini/reports/DALL-E-Mega--VmlldzoxODMxMDI2

TorToiSe TTS is amazing!

https://github.com/neonbjb/tortoise-tts

https://nonint.com/static/tortoise_v2_examples.html

https://colab.research.google.com/drive/1wVVqUPqwiDBUVeWWOUNglpGhU3hg_cbR

https://github.com/neonbjb

Investigating Vision Transformers

https://github.com/sayakpaul/probing-vits/?utm_source=pocket_mylist

https://twitter.com/RisingSayak/status/1515918406171914240?utm_source=pocket_mylist

https://keras.io/examples/vision/probing_vits/

https://github.com/sayakpaul/probing-vits/tree/main/notebooks?utm_source=pocket_mylist

Hugging Face Deep RL class launched

https://github.com/huggingface/deep-rl-class

Helpful Things

https://merantix-momentum.com/technology/squirrel/?utm_source=pocket_mylist

https://github.com/merantix-momentum/squirrel-core?utm_source=pocket_mylist

https://pyscript.net/?utm_source=pocket_mylist

https://github.com/google-research/big_vision

https://deepsportradar.github.io/challenge.html

https://github.com/DeepSportRadar/camera-calibration-challenge

https://twitter.com/alekseykorshuk/status/1515989357961920514?utm_source=pocket_mylist

https://github.com/AlekseyKorshuk/huggingnft

John Deere's driverless tractors

https://thenextweb.com/news/john-deere-slowly-becoming-one-worlds-most-important-ai-companies

https://tractorhacking.github.io/

Links:

Merch: https://ykilcher.com/merch

TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://discord.gg/4H8xxDF

BitChute: https://www.bitchute.com/channel/yannic-kilcher

LinkedIn: https://www.linkedin.com/in/ykilcher

BiliBili: https://space.bilibili.com/2017636191

If you want to support me, the best thing to do is to share out the content :)

View Details

nft #gan #ai

Today we build our own AI that can create as many bored apes as we want! Fungibility for everyone!

Try the model here: https://huggingface.co/spaces/ykilcher/apes

or here: https://ykilcher.com/apes

Files & Models here: https://huggingface.co/ykilcher/apes/tree/main

Code here: https://github.com/yk/apes-public (for the "what's your ape" app, look for the file interface_projector.py)

This video is sponsored by BrightData, use this link for free credits:

https://brightdata.grsm.io/yannickilcher

OUTLINE:

0:00 - Introduction

2:05 - Generative Adversarial Networks

3:40 - Scraping Opensea with BrightData

7:55 - Training the GAN

11:35 - Here are the results!

15:20 - Diving deeper into BrightData

References:

Stylegan 3 imagery: https://nvlabs.github.io/stylegan3/

Bored Ape Yacht Club NFT Collection: https://opensea.io/collection/boredapeyachtclub

Better GANFT model: https://medium.com/@nathancooperjones/these-bored-apes-do-not-exist-6bed2c73f02c

Abstract AI-created apes: https://opensea.io/collection/gan-apes-nft

https://mobile.twitter.com/gannft

Another good model: https://twitter.com/cyrilzakka/status/1463944040878071811

StyleGAN2 versions: https://thispersondoesnotexist.com/

https://thissneakerdoesnotexist.com/

https://thischairdoesnotexist.com/

GANs: https://en.wikipedia.org/wiki/Generative_adversarial_network

https://arxiv.org/pdf/1406.2661.pdf

StyleGAN3: https://nvlabs.github.io/stylegan3/

StyleGAN2 code: https://github.com/NVlabs/stylegan2-ada-pytorch

CLIP: https://openai.com/blog/clip/

DALL-E 2 images: https://twitter.com/search?q=%23dalle&f=image

My music video: https://www.youtube.com/watch?v=2iq7WXSw26s

BrightData Links: https://brightdata.com/products/data-collector

https://brightdata.com/testimonials

https://brightdata.com/use-cases/adtech

https://brightdata.com/use-cases/social-media-for-marketing

https://brightdata.com/use-cases/ecommerce

Links:

Merch: https://ykilcher.com/merch

TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://discord.gg/4H8xxDF

BitChute: https://www.bitchute.com/channel/yannic-kilcher

LinkedIn: https://www.linkedin.com/in/ykilcher

BiliBili: https://space.bilibili.com/2017636191

If you want to support me, the best thing to do is to share out the content :)

If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this), find options at https://ykilcher.com

View Details

saycan #robots #ai

This is an interview with the authors Brian Ichter, Karol Hausman, and Fei Xia.

Original Paper Review Video: https://youtu.be/Ru23eWAQ6_E

Large Language Models are excellent at generating plausible plans in response to real-world problems, but without interacting with the environment, they have no abilities to estimate which of these plans are feasible or appropriate. SayCan combines the semantic capabilities of language models with a bank of low-level skills, which are available to the agent as individual policies to execute. SayCan automatically finds the best policy to execute by considering a trade-off between the policy's ability to progress towards the goal, given by the language model, and the policy's probability of executing successfully, given by the respective value function. The result is a system that can generate and execute long-horizon action sequences in the real world to fulfil complex tasks.

OUTLINE:

0:00 - Introduction & Setup

3:40 - Acquiring atomic low-level skills

7:45 - How does the language model come in?

11:45 - Why are you scoring instead of generating?

15:20 - How do you deal with ambiguity in language?

20:00 - The whole system is modular

22:15 - Going over the full algorithm

23:20 - What if an action fails?

24:30 - Debunking a marketing video :)

27:25 - Experimental Results

32:50 - The insane scale of data collection

40:15 - How do you go about large-scale projects?

43:20 - Where did things go wrong?

45:15 - Where do we go from here?

52:00 - What is the largest unsolved problem in this?

53:35 - Thoughts on the Tesla Bot

55:00 - Final thoughts

Paper: https://arxiv.org/abs/2204.01691

Website: https://say-can.github.io/

Abstract:

Large language models can encode a wealth of semantic knowledge about the world. Such knowledge could be extremely useful to robots aiming to act upon high-level, temporally extended instructions expressed in natural language. However, a significant weakness of language models is that they lack real-world experience, which makes it difficult to leverage them for decision making within a given embodiment. For example, asking a language model to describe how to clean a spill might result in a reasonable narrative, but it may not be applicable to a particular agent, such as a robot, that needs to perform this task in a particular environment. We propose to provide real-world grounding by means of pretrained skills, which are used to constrain the model to propose natural language actions that are both feasible and contextually appropriate. The robot can act as the language model's "hands and eyes," while the language model supplies high-level semantic knowledge about the task. We show how low-level skills can be combined with large language models so that the language model provides high-level knowledge about the procedures for performing complex and temporally-extended instructions, while value functions associated with these skills provide the grounding necessary to connect this knowledge to a particular physical environment.

Authors: Michael Ahn, Anthony Brohan, Noah Brown, Yevgen Chebotar, Omar Cortes, Byron David, Chelsea Finn, Keerthana Gopalakrishnan, Karol Hausman, Alex Herzog, Daniel Ho, Jasmine Hsu, Julian Ibarz, Brian Ichter, Alex Irpan, Eric Jang, Rosario Jauregui Ruano, Kyle Jeffrey, Sally Jesmonth, Nikhil J Joshi, Ryan Julian, Dmitry Kalashnikov, Yuheng Kuang, Kuang-Huei Lee, Sergey Levine, Yao Lu, Linda Luu, Carolina Parada, Peter Pastor, Jornell Quiambao, Kanishka Rao, Jarek Rettinghouse, Diego Reyes, Pierre Sermanet, Nicolas Sievers, Clayton Tan, Alexander Toshev, Vincent Vanhoucke, Fei Xia, Ted Xiao, Peng Xu, Sichun Xu, Mengyuan Yan

View Details

saycan #robots #ai

Large Language Models are excellent at generating plausible plans in response to real-world problems, but without interacting with the environment, they have no abilities to estimate which of these plans are feasible or appropriate. SayCan combines the semantic capabilities of language models with a bank of low-level skills, which are available to the agent as individual policies to execute. SayCan automatically finds the best policy to execute by considering a trade-off between the policy's ability to progress towards the goal, given by the language model, and the policy's probability of executing successfully, given by the respective value function. The result is a system that can generate and execute long-horizon action sequences in the real world to fulfil complex tasks.

Sponsor: Zeta Alpha

https://zeta-alpha.com

Use code YANNIC for 20% off!

OUTLINE:

0:00 - Introduction & Overview

3:20 - Sponsor: Zeta Alpha

5:00 - Using language models for action planning

8:00 - Combining LLMs with learned atomic skills

16:50 - The full SayCan system

20:30 - Experimental setup and data collection

21:25 - Some weaknesses & strengths of the system

27:00 - Experimental results

Paper: https://arxiv.org/abs/2204.01691

Website: https://say-can.github.io/

Abstract:

Large language models can encode a wealth of semantic knowledge about the world. Such knowledge could be extremely useful to robots aiming to act upon high-level, temporally extended instructions expressed in natural language. However, a significant weakness of language models is that they lack real-world experience, which makes it difficult to leverage them for decision making within a given embodiment. For example, asking a language model to describe how to clean a spill might result in a reasonable narrative, but it may not be applicable to a particular agent, such as a robot, that needs to perform this task in a particular environment. We propose to provide real-world grounding by means of pretrained skills, which are used to constrain the model to propose natural language actions that are both feasible and contextually appropriate. The robot can act as the language model's "hands and eyes," while the language model supplies high-level semantic knowledge about the task. We show how low-level skills can be combined with large language models so that the language model provides high-level knowledge about the procedures for performing complex and temporally-extended instructions, while value functions associated with these skills provide the grounding necessary to connect this knowledge to a particular physical environment. We evaluate our method on a number of real-world robotic tasks, where we show the need for real-world grounding and that this approach is capable of completing long-horizon, abstract, natural language instructions on a mobile manipulator. The project's website and the video can be found at this https URL

Authors: Michael Ahn, Anthony Brohan, Noah Brown, Yevgen Chebotar, Omar Cortes, Byron David, Chelsea Finn, Keerthana Gopalakrishnan, Karol Hausman, Alex Herzog, Daniel Ho, Jasmine Hsu, Julian Ibarz, Brian Ichter, Alex Irpan, Eric Jang, Rosario Jauregui Ruano, Kyle Jeffrey, Sally Jesmonth, Nikhil J Joshi, Ryan Julian, Dmitry Kalashnikov, Yuheng Kuang, Kuang-Huei Lee, Sergey Levine, Yao Lu, Linda Luu, Carolina Parada, Peter Pastor, Jornell Quiambao, Kanishka Rao, Jarek Rettinghouse, Diego Reyes, Pierre Sermanet, Nicolas Sievers, Clayton Tan, Alexander Toshev, Vincent Vanhoucke, Fei Xia, Ted Xiao, Peng Xu, Sichun Xu, Mengyuan Yan

View Details

ai #accel #evolution

This is an interview with the authors Jack Parker-Holder and Minqi Jiang.

Original Paper Review Video: https://www.youtube.com/watch?v=povBD...

Automatic curriculum generation is one of the most promising avenues for Reinforcement Learning today. Multiple approaches have been proposed, each with their own set of advantages and drawbacks. This paper presents ACCEL, which takes the next step into the direction of constructing curricula for multi-capable agents. ACCEL combines the adversarial adaptiveness of regret-based sampling methods with the capabilities of level-editing, usually found in Evolutionary Methods.

OUTLINE:

0:00 - Intro

1:00 - Start of interview

4:45 - How did you get into this field?

8:10 - What is minimax regret?

11:45 - What levels does the regret objective select?

14:20 - Positive value loss (correcting my mistakes)

21:05 - Why is the teacher not learned?

24:45 - How much domain-specific knowledge is needed?

29:30 - What problems is this applicable to?

33:15 - Single agent vs population of agents

37:25 - Measuring and balancing level difficulty

40:35 - How does generalization emerge?

42:50 - Diving deeper into the experimental results

47:00 - What are the unsolved challenges in the field?

50:00 - Where do we go from here?

Website: https://accelagent.github.io

Paper: https://arxiv.org/abs/2203.01302

ICLR Workshop: https://sites.google.com/view/aloe2022

Book on topic: https://www.oreilly.com/radar/open-en...

Abstract:

It remains a significant challenge to train generally capable agents with reinforcement learning (RL). A promising avenue for improving the robustness of RL agents is through the use of curricula. One such class of methods frames environment design as a game between a student and a teacher, using regret-based objectives to produce environment instantiations (or levels) at the frontier of the student agent's capabilities. These methods benefit from their generality, with theoretical guarantees at equilibrium, yet they often struggle to find effective levels in challenging design spaces. By contrast, evolutionary approaches seek to incrementally alter environment complexity, resulting in potentially open-ended learning, but often rely on domain-specific heuristics and vast amounts of computational resources. In this paper we propose to harness the power of evolution in a principled, regret-based curriculum. Our approach, which we call Adversarially Compounding Complexity by Editing Levels (ACCEL), seeks to constantly produce levels at the frontier of an agent's capabilities, resulting in curricula that start simple but become increasingly complex. ACCEL maintains the theoretical benefits of prior regret-based methods, while providing significant empirical gains in a diverse set of environments. An interactive version of the paper is available at this http URL.

Authors: Jack Parker-Holder, Minqi Jiang, Michael Dennis, Mikayel Samvelyan, Jakob Foerster, Edward Grefenstette, Tim Rocktäschel

Links:

TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://discord.gg/4H8xxDF

BitChute: https://www.bitchute.com/channel/yann...

LinkedIn: https://www.linkedin.com/in/ykilcher

BiliBili: https://space.bilibili.com/2017636191

If you want to support me, the best thing to do is to share out the content :)

View Details

ai #accel #evolution

Automatic curriculum generation is one of the most promising avenues for Reinforcement Learning today. Multiple approaches have been proposed, each with their own set of advantages and drawbacks. This paper presents ACCEL, which takes the next step into the direction of constructing curricula for multi-capable agents. ACCEL combines the adversarial adaptiveness of regret-based sampling methods with the capabilities of level-editing, usually found in Evolutionary Methods.

OUTLINE:

0:00 - Intro & Demonstration

3:50 - Paper overview

5:20 - The ACCEL algorithm

15:25 - Looking at the pseudocode

23:10 - Approximating regret

33:45 - Experimental results

40:00 - Discussion & Comments

Website: https://accelagent.github.io

Paper: https://arxiv.org/abs/2203.01302

Abstract:

It remains a significant challenge to train generally capable agents with reinforcement learning (RL). A promising avenue for improving the robustness of RL agents is through the use of curricula. One such class of methods frames environment design as a game between a student and a teacher, using regret-based objectives to produce environment instantiations (or levels) at the frontier of the student agent's capabilities. These methods benefit from their generality, with theoretical guarantees at equilibrium, yet they often struggle to find effective levels in challenging design spaces. By contrast, evolutionary approaches seek to incrementally alter environment complexity, resulting in potentially open-ended learning, but often rely on domain-specific heuristics and vast amounts of computational resources. In this paper we propose to harness the power of evolution in a principled, regret-based curriculum. Our approach, which we call Adversarially Compounding Complexity by Editing Levels (ACCEL), seeks to constantly produce levels at the frontier of an agent's capabilities, resulting in curricula that start simple but become increasingly complex. ACCEL maintains the theoretical benefits of prior regret-based methods, while providing significant empirical gains in a diverse set of environments. An interactive version of the paper is available at this http URL.

Authors: Jack Parker-Holder, Minqi Jiang, Michael Dennis, Mikayel Samvelyan, Jakob Foerster, Edward Grefenstette, Tim Rocktäschel

Links:

TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://discord.gg/4H8xxDF

BitChute: https://www.bitchute.com/channel/yann...

LinkedIn: https://www.linkedin.com/in/ykilcher

BiliBili: https://space.bilibili.com/2017636191

If you want to support me, the best thing to do is to share out the content :)

If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https://www.subscribestar.com/yannick...

Patreon: https://www.patreon.com/yannickilcher

Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq

Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2

Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m

Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n

View Details

laion #clip #dalle

LAION-5B is an open, free dataset consisting of over 5 billion image-text-pairs. Today's video is an interview with three of its creators. We dive into the mechanics and challenges of operating at such large scale, how to keep cost low, what new possibilities are enabled with open datasets like this, and how to best handle safety and legal concerns.

OUTLINE:

0:00 - Intro

1:30 - Start of Interview

2:30 - What is LAION?

11:10 - What are the effects of CLIP filtering?

16:40 - How big is this dataset?

19:05 - Does the text always come from the alt-property?

22:45 - What does it take to work at scale?

25:50 -When will we replicate DALL-E?

31:30 - The surprisingly efficient pipeline

35:20 - How do you cover the S3 costs?

40:30 - Addressing safety & legal concerns

55:15 - Where can people get started?

References:

LAION website: https://laion.ai/

LAION Discord: https://discord.com/invite/mVcgxMPD7e

LAION-5B: https://laion.ai/laion-5b-a-new-era-o...

img2dataset tool: https://github.com/rom1504/img2dataset

LAION-400M: https://paperswithcode.com/dataset/la...

Links:

TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://discord.gg/4H8xxDF

BitChute: https://www.bitchute.com/channel/yann...

LinkedIn: https://www.linkedin.com/in/ykilcher

BiliBili: https://space.bilibili.com/2017636191

If you want to support me, the best thing to do is to share out the content :)

If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https://www.subscribestar.com/yannick...

Patreon: https://www.patreon.com/yannickilcher

Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq

Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2

Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m

Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n

View Details

nlp #sparsity #transformers

This video is an interview with Barret Zoph and William Fedus of Google Brain about Sparse Expert Models.

Sparse Expert models have been hugely successful at distributing parts of models, mostly Transformers, across large array of machines and use a routing function to effectively route signals between them. This means that even though these models have a huge number of parameters, the computational load for a given signal does not increase because the model is only sparsely activated. Sparse expert models, such as Switch Transformers and GLAM can scale up to trillions of parameters and bring a number of desirable properties. We discuss everything from the fundamentals, history, strengths and weaknesses, up to the current state of the art of these models.

OUTLINE:

0:00 - Intro

0:30 - What are sparse expert models?

4:25 - Start of Interview

5:55 - What do you mean by sparse experts?

8:10 - How does routing work in these models?

12:10 - What is the history of sparse experts?

14:45 - What does an individual expert learn?

19:25 - When are these models appropriate?

22:30 - How comparable are sparse to dense models?

26:30 - How does the pathways system connect to this?

28:45 - What improvements did GLAM make?

31:30 - The "designing sparse experts" paper

37:45 - Can experts be frozen during training?

41:20 - Can the routing function be improved?

47:15 - Can experts be distributed beyond data centers?

50:20 - Are there sparse experts for other domains than NLP?

52:15 - Are sparse and dense models in competition?

53:35 - Where do we go from here?

56:30 - How can people get started with this?

Papers:

Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity (https://arxiv.org/abs/2101.03961)

GLaM: Efficient Scaling of Language Models with Mixture-of-Experts (https://arxiv.org/abs/2112.06905)

Designing Effective Sparse Expert Models (https://arxiv.org/abs/2202.08906)

Links:

Merch: store.ykilcher.com

TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://discord.gg/4H8xxDF

BitChute: https://www.bitchute.com/channel/yann...

LinkedIn: https://www.linkedin.com/in/ykilcher

BiliBili: https://space.bilibili.com/2017636191

If you want to support me, the best thing to do is to share out the content :)

If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https://www.subscribestar.com/yannick...

Patreon: https://www.patreon.com/yannickilcher

Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq

Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2

Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m

Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n

View Details

neuralsearch #interview #google

This is an interview with the authors Yi Tay and Don Metzler.

Paper Review Video: https://youtu.be/qlB0TPBQ7YY

Search engines work by building an index and then looking up things in it. Usually, that index is a separate data structure. In keyword search, we build and store reverse indices. In neural search, we build nearest-neighbor indices. This paper does something different: It directly trains a Transformer to return the ID of the most relevant document. No similarity search over embeddings or anything like this is performed, and no external data structure is needed, as the entire index is essentially captured by the model's weights. The paper experiments with various ways of representing documents and training the system, which works surprisingly well!

OUTLINE:

0:00 - Intro

0:50 - Start of Interview

1:30 - How did this idea start?

4:30 - How does memorization play into this?

5:50 - Why did you not compare to cross-encoders?

7:50 - Instead of the ID, could one reproduce the document itself?

10:50 - Passages vs documents

12:00 - Where can this model be applied?

14:25 - Can we make this work on large collections?

19:20 - What's up with the NQ100K dataset?

23:55 - What is going on inside these models?

28:30 - What's the smallest scale to obtain meaningful results?

30:15 - Investigating the document identifiers

34:45 - What's the end goal?

38:40 - What are the hardest problems currently?

40:40 - Final comments & how to get started

Paper: https://arxiv.org/abs/2202.06991

Abstract:

In this paper, we demonstrate that information retrieval can be accomplished with a single Transformer, in which all information about the corpus is encoded in the parameters of the model. To this end, we introduce the Differentiable Search Index (DSI), a new paradigm that learns a text-to-text model that maps string queries directly to relevant docids; in other words, a DSI model answers queries directly using only its parameters, dramatically simplifying the whole retrieval process. We study variations in how documents and their identifiers are represented, variations in training procedures, and the interplay between models and corpus sizes. Experiments demonstrate that given appropriate design choices, DSI significantly outperforms strong baselines such as dual encoder models. Moreover, DSI demonstrates strong generalization capabilities, outperforming a BM25 baseline in a zero-shot setup.

Authors: Yi Tay, Vinh Q. Tran, Mostafa Dehghani, Jianmo Ni, Dara Bahri, Harsh Mehta, Zhen Qin, Kai Hui, Zhe Zhao, Jai Gupta, Tal Schuster, William W. Cohen, Donald Metzler

Links:

Merch: store.ykilcher.com

TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://discord.gg/4H8xxDF

BitChute: https://www.bitchute.com/channel/yann...

LinkedIn: https://www.linkedin.com/in/ykilcher

BiliBili: https://space.bilibili.com/2017636191

If you want to support me, the best thing to do is to share out the content :)

If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https://www.subscribestar.com/yannick...

Patreon: https://www.patreon.com/yannickilcher

Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq

Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2

Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m

Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n

View Details

dsi #search #google

Search engines work by building an index and then looking up things in it. Usually, that index is a separate data structure. In keyword search, we build and store reverse indices. In neural search, we build nearest-neighbor indices. This paper does something different: It directly trains a Transformer to return the ID of the most relevant document. No similarity search over embeddings or anything like this is performed, and no external data structure is needed, as the entire index is essentially captured by the model's weights. The paper experiments with various ways of representing documents and training the system, which works surprisingly well!

Sponsor: Diffgram

https://diffgram.com?ref=yannic

OUTLINE:

0:00 - Intro

0:45 - Sponsor: Diffgram

1:35 - Paper overview

3:15 - The search problem, classic and neural

8:15 - Seq2seq for directly predicting document IDs

11:05 - Differentiable search index architecture

18:05 - Indexing

25:15 - Retrieval and document representation

33:25 - Training DSI

39:15 - Experimental results

49:25 - Comments & Conclusions

Paper: https://arxiv.org/abs/2202.06991

Abstract:

In this paper, we demonstrate that information retrieval can be accomplished with a single Transformer, in which all information about the corpus is encoded in the parameters of the model. To this end, we introduce the Differentiable Search Index (DSI), a new paradigm that learns a text-to-text model that maps string queries directly to relevant docids; in other words, a DSI model answers queries directly using only its parameters, dramatically simplifying the whole retrieval process. We study variations in how documents and their identifiers are represented, variations in training procedures, and the interplay between models and corpus sizes. Experiments demonstrate that given appropriate design choices, DSI significantly outperforms strong baselines such as dual encoder models. Moreover, DSI demonstrates strong generalization capabilities, outperforming a BM25 baseline in a zero-shot setup.

Authors: Yi Tay, Vinh Q. Tran, Mostafa Dehghani, Jianmo Ni, Dara Bahri, Harsh Mehta, Zhen Qin, Kai Hui, Zhe Zhao, Jai Gupta, Tal Schuster, William W. Cohen, Donald Metzler

Links:

Merch: store.ykilcher.com

TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://discord.gg/4H8xxDF

BitChute: https://www.bitchute.com/channel/yann...

LinkedIn: https://www.linkedin.com/in/ykilcher

BiliBili: https://space.bilibili.com/2017636191

If you want to support me, the best thing to do is to share out the content :)

If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https://www.subscribestar.com/yannick...

Patreon: https://www.patreon.com/yannickilcher

Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq

Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2

Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m

Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n

View Details

mlnews #palm #dalle2

Google releases PaLM and OpenAI releases DALL-E 2 (and more news).

Sponsor: Weights & BIases

Start here: https://wandb.me/yannic

Thumbnail credit: DALL-E 2 via Sam Altman

OUTLINE

0:00 - Street interview w/ random stranger

2:25 - Intro

2:50 - PaLM - Google's 540B Pathways Language Model

7:50 - Sponsor: Weights & Biases

9:10 - OpenAI releases DALL-E 2

12:05 - Open Source Datasets and Models

13:20 - Salesforce releases CodeGen

My Live Reaction to DALL-E 2: https://youtu.be/gGPv_SYVDC8

My Video on GLIDE: https://youtu.be/gwI6g1pBD84

My Video on the Pathways System: https://youtu.be/vGFaiLeoLWw

References:

PaLM - Google's 540B Pathways Language Model

https://ai.googleblog.com/2022/04/pat...

https://storage.googleapis.com/pathwa...

OpenAI releases DALL-E 2

https://openai.com/dall-e-2/

https://cdn.openai.com/papers/dall-e-...

https://www.instagram.com/openaidalle/

https://twitter.com/sama/status/15117...

https://twitter.com/sama/media

https://twitter.com/BorisMPower/statu...

https://twitter.com/ariskonstant/stat...

Open Source Datasets and Models

https://twitter.com/multimodalart/sta...

https://laion.ai/laion-5b-a-new-era-o...

https://github.com/mlfoundations/open...

Salesforce releases CodeGen

https://github.com/salesforce/CodeGen

Links:

Merch: store.ykilcher.com

TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://discord.gg/4H8xxDF

BitChute: https://www.bitchute.com/channel/yann...

LinkedIn: https://www.linkedin.com/in/ykilcher

BiliBili: https://space.bilibili.com/2017636191

If you want to support me, the best thing to do is to share out the content :)

If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https://www.subscribestar.com/yannick...

Patreon: https://www.patreon.com/yannickilcher

Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq

Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2

Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m

Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n

View Details

aiart #deeplearning #clip

Since the release of CLIP, the world of AI art has seen an unprecedented level of acceleration in what's possible to do. Whereas image generation had previously been mostly in the domain of scientists, now a community of professional artists, researchers, and amateurs are sending around colab notebooks and sharing their creations via social media. How did this happen? What is going on? And where do we go from here? Jack Morris and I attempt to answer some of these questions, following his blog post "The Weird and Wonderful World of AI Art" (linked below).

OUTLINE:

0:00 - Intro

2:30 - How does one get into AI art?

5:00 - Deep Dream & Style Transfer: the early days of art in deep learning

10:50 - The advent of GANs, ArtBreeder and TikTok

19:50 - Lacking control: Pre-CLIP art

22:40 - CLIP & DALL-E

30:20 - The shift to shared colabs

34:20 - Guided diffusion models

37:20 - Prompt engineering for art models

43:30 - GLIDE

47:00 - Video production & Disco Diffusion

48:40 - Economics, money, and NFTs

54:15 - What does the future hold for AI art?

Blog post: https://jxmo.notion.site/The-Weird-an...

Jack's Blog: https://jxmo.io/

Links:

TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://discord.gg/4H8xxDF

BitChute: https://www.bitchute.com/channel/yann...

LinkedIn: https://www.linkedin.com/in/ykilcher

BiliBili: https://space.bilibili.com/2017636191

If you want to support me, the best thing to do is to share out the content :)

If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https://www.subscribestar.com/yannick...

Patreon: https://www.patreon.com/yannickilcher

Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq

Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2

Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m

Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n

View Details

reinforcementlearning #ai #explained

This is an interview with Jesse Mu, first author of the paper.

Original Paper Review: https://youtu.be/NeGJAUSQEJI

Exploration is one of the oldest challenges for Reinforcement Learning algorithms, with no clear solution to date. Especially in environments with sparse rewards, agents face significant challenges in deciding which parts of the environment to explore further. Providing intrinsic motivation in form of a pseudo-reward is sometimes used to overcome this challenge, but often relies on hand-crafted heuristics, and can lead to deceptive dead-ends. This paper proposes to use language descriptions of encountered states as a method of assessing novelty. In two procedurally generated environments, they demonstrate the usefulness of language, which is in itself highly concise and abstractive, which lends itself well for this task.

OUTLINE:

0:00 - Intro

0:55 - Paper Overview

4:30 - Aren't you just adding extra data?

9:35 - Why are you splitting up the AMIGo teacher?

13:10 - How do you train the grounding network?

16:05 - What about causally structured environments?

17:30 - Highlights of the experimental results

20:40 - Why is there so much variance?

22:55 - How much does it matter that we are testing in a video game?

27:00 - How does novelty interface with the goal specification?

30:20 - The fundamental problems of exploration

32:15 - Are these algorithms subject to catastrophic forgetting?

34:45 - What current models could bring language to other environments?

40:30 - What does it take in terms of hardware?

43:00 - What problems did you encounter during the project?

46:40 - Where do we go from here?

Paper: https://arxiv.org/abs/2202.08938

Abstract:

Reinforcement learning (RL) agents are particularly hard to train when rewards are sparse. One common solution is to use intrinsic rewards to encourage agents to explore their environment. However, recent intrinsic exploration methods often use state-based novelty measures which reward low-level exploration and may not scale to domains requiring more abstract skills. Instead, we explore natural language as a general medium for highlighting relevant abstractions in an environment. Unlike previous work, we evaluate whether language can improve over existing exploration methods by directly extending (and comparing to) competitive intrinsic exploration baselines: AMIGo (Campero et al., 2021) and NovelD (Zhang et al., 2021). These language-based variants outperform their non-linguistic forms by 45-85% across 13 challenging tasks from the MiniGrid and MiniHack environment suites.

Authors: Jesse Mu, Victor Zhong, Roberta Raileanu, Minqi Jiang, Noah Goodman, Tim Rocktäschel, Edward Grefenstette

Links:

TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://discord.gg/4H8xxDF

BitChute: https://www.bitchute.com/channel/yann...

LinkedIn: https://www.linkedin.com/in/ykilcher

BiliBili: https://space.bilibili.com/2017636191

If you want to support me, the best thing to do is to share out the content :)

If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https://www.subscribestar.com/yannick...

Patreon: https://www.patreon.com/yannickilcher

Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq

Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2

Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m

View Details

reinforcementlearning #ai #explained

Exploration is one of the oldest challenges for Reinforcement Learning algorithms, with no clear solution to date. Especially in environments with sparse rewards, agents face significant challenges in deciding which parts of the environment to explore further. Providing intrinsic motivation in form of a pseudo-reward is sometimes used to overcome this challenge, but often relies on hand-crafted heuristics, and can lead to deceptive dead-ends. This paper proposes to use language descriptions of encountered states as a method of assessing novelty. In two procedurally generated environments, they demonstrate the usefulness of language, which is in itself highly concise and abstractive, which lends itself well for this task.

OUTLINE:

0:00 - Intro

1:10 - Paper Overview: Language for exploration

5:40 - The MiniGrid & MiniHack environments

7:00 - Annotating states with language

9:05 - Baseline algorithm: AMIGo

12:20 - Adding language to AMIGo

22:55 - Baseline algorithm: NovelD and Random Network Distillation

29:45 - Adding language to NovelD

31:50 - Aren't we just using extra data?

34:55 - Investigating the experimental results

40:45 - Final comments

Paper: https://arxiv.org/abs/2202.08938

Abstract:

Reinforcement learning (RL) agents are particularly hard to train when rewards are sparse. One common solution is to use intrinsic rewards to encourage agents to explore their environment. However, recent intrinsic exploration methods often use state-based novelty measures which reward low-level exploration and may not scale to domains requiring more abstract skills. Instead, we explore natural language as a general medium for highlighting relevant abstractions in an environment. Unlike previous work, we evaluate whether language can improve over existing exploration methods by directly extending (and comparing to) competitive intrinsic exploration baselines: AMIGo (Campero et al., 2021) and NovelD (Zhang et al., 2021). These language-based variants outperform their non-linguistic forms by 45-85% across 13 challenging tasks from the MiniGrid and MiniHack environment suites.

Authors: Jesse Mu, Victor Zhong, Roberta Raileanu, Minqi Jiang, Noah Goodman, Tim Rocktäschel, Edward Grefenstette

Links:

TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://discord.gg/4H8xxDF

BitChute: https://www.bitchute.com/channel/yann...

LinkedIn: https://www.linkedin.com/in/ykilcher

BiliBili: https://space.bilibili.com/2017636191

If you want to support me, the best thing to do is to share out the content :)

If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https://www.subscribestar.com/yannick...

Patreon: https://www.patreon.com/yannickilcher

Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq

Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2

Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m

Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n

View Details

mlnews #gpt3 #pathways

Your updates on the latest and greatest from the depths of Machine Learning!

Sponsor: Weights & Biases

https://wandb.me/yannic

OUTLINE:

0:00 - Intro

0:15 - Weights & Biases Report about Reports

2:45 - GPT-3 learns to edit

6:30 - Make-A-Scene: Text-to-Image with Human Priors

8:00 - Pathways: Google's new High-Performance ML scheduler

10:45 - DouBlind: Open Peer-Review

12:45 - CLIP meets GamePhysics

14:40 - Residual Quantization pushes Image Generation SOTA

16:15 - Helpful Things

References:

Weights & Biases Report about Reports

https://wandb.ai/wandb/wandb_example/...

GPT-3 learns to edit

https://openai.com/blog/gpt-3-edit-in...

https://beta.openai.com/playground?mo...

Make-A-Scene: Text-to-Image with Human Priors

https://arxiv.org/pdf/2203.13131.pdf

https://www.youtube.com/watch?v=QLTyq...

Pathways: Google's new High-Performance ML scheduler

https://arxiv.org/pdf/2203.12533.pdf

DouBlind: Open Peer-Review

https://doublind.com/#web-intro

https://doublind.com/search?query=kil...

CLIP meets GamePhysics

https://arxiv.org/pdf/2203.11096.pdf

https://www.reddit.com/r/GamePhysics/...

https://asgaardlab.github.io/CLIPxGam...

Residual Quantization pushes Image Generation SOTA

https://arxiv.org/pdf/2203.01941.pdf

https://github.com/kakaobrain/rq-vae-...

Helpful Things

https://github.com/TDAmeritrade/stumpy

https://github.com/linkedin/fasttreeshap

https://github.com/vopani/jaxton

https://twitter.com/mark_riedl/status...

https://github.com/eilab-gt/NovGrid

https://developer.nvidia.com/isaac-gym

https://github.com/NVIDIA-Omniverse/I...

Links:

Merch: store.ykilcher.com

TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://discord.gg/4H8xxDF

BitChute: https://www.bitchute.com/channel/yann...

LinkedIn: https://www.linkedin.com/in/ykilcher

BiliBili: https://space.bilibili.com/2017636191

If you want to support me, the best thing to do is to share out the content :)

If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https://www.subscribestar.com/yannick...

Patreon: https://www.patreon.com/yannickilcher

Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq

Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2

Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m

Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n

View Details

nlp #gpt3 #prompt

This is an interview with the authors of this work, Aman Madaan and Niket Tandon.

Large language models such as GPT-3 have enabled many breakthroughs and new applications recently, but they come with an important downside: Training them is very expensive, and even fine-tuning is often difficult. This paper presents an adaptive method to improve performance of such models after deployment, without ever changing the model itself. This is done by maintaining a memory of interactions and then dynamically adapting new prompts by augmenting them with memory content. This has many applications, from non-intrusive fine-tuning to personalization.

OUTLINE:

0:00 - Intro

0:45 - Paper Overview

2:00 - What was your original motivation?

4:20 - There is an updated version of the paper!

9:00 - Have you studied this on real-world users?

12:10 - How does model size play into providing feedback?

14:10 - Can this be used for personalization?

16:30 - Discussing experimental results

17:45 - Can this be paired with recommender systems?

20:00 - What are obvious next steps to make the system more powerful?

23:15 - Clarifying the baseline methods

26:30 - Exploring cross-lingual customization

31:00 - Where did the idea for the clarification prompt come from?

33:05 - What did not work out during this project?

34:45 - What did you learn about interacting with large models?

37:30 - Final thoughts

Paper: https://arxiv.org/abs/2201.06009

Code & Data: https://github.com/madaan/memprompt

Abstract:

Large LMs such as GPT-3 are powerful, but can commit mistakes that are obvious to humans. For example, GPT-3 would mistakenly interpret "What word is similar to good?" to mean a homonym, while the user intended a synonym. Our goal is to effectively correct such errors via user interactions with the system but without retraining, which will be prohibitively costly. We pair GPT-3 with a growing memory of recorded cases where the model misunderstood the user's intents, along with user feedback for clarification. Such a memory allows our system to produce enhanced prompts for any new query based on the user feedback for error correction on similar cases in the past. On four tasks (two lexical tasks, two advanced ethical reasoning tasks), we show how a (simulated) user can interactively teach a deployed GPT-3, substantially increasing its accuracy over the queries with different kinds of misunderstandings by the GPT-3. Our approach is a step towards the low-cost utility enhancement for very large pre-trained LMs. All the code and data is available at this https URL.

Authors: Aman Madaan, Niket Tandon, Peter Clark, Yiming Yang

Links:

Merch: store.ykilcher.com

TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://discord.gg/4H8xxDF

BitChute: https://www.bitchute.com/channel/yann...

LinkedIn: https://www.linkedin.com/in/ykilcher

BiliBili: https://space.bilibili.com/2017636191

If you want to support me, the best thing to do is to share out the content :)

If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https://www.subscribestar.com/yannick...

Patreon: https://www.patreon.com/yannickilcher

Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq

Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2

Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m

Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n

View Details

nlp #gpt3 #prompt

Large language models such as GPT-3 have enabled many breakthroughs and new applications recently, but they come with an important downside: Training them is very expensive, and even fine-tuning is often difficult. This paper presents an adaptive method to improve performance of such models after deployment, without ever changing the model itself. This is done by maintaining a memory of interactions and then dynamically adapting new prompts by augmenting them with memory content. This has many applications, from non-intrusive fine-tuning to personalization.

Sponsor: Introduction to Graph Neural Networks Course

https://www.graphneuralnets.com/p/int...

OUTLINE:

0:00 - Intro

0:40 - Sponsor: Introduction to GNNs Course (link in description)

1:30 - Paper Overview: Improve GPT-3 after deployment via user feedback

5:30 - Proposed memory-based architecture

13:00 - A detailed look at the components

15:00 - Example tasks

24:30 - My concerns with the example setup

26:20 - Baselines used for comparison

29:50 - Experimental Results

34:20 - Conclusion & Comments

Paper: https://arxiv.org/abs/2201.06009

Code & Data: https://github.com/madaan/memprompt

Abstract:

Large LMs such as GPT-3 are powerful, but can commit mistakes that are obvious to humans. For example, GPT-3 would mistakenly interpret "What word is similar to good?" to mean a homonym, while the user intended a synonym. Our goal is to effectively correct such errors via user interactions with the system but without retraining, which will be prohibitively costly. We pair GPT-3 with a growing memory of recorded cases where the model misunderstood the user's intents, along with user feedback for clarification. Such a memory allows our system to produce enhanced prompts for any new query based on the user feedback for error correction on similar cases in the past. On four tasks (two lexical tasks, two advanced ethical reasoning tasks), we show how a (simulated) user can interactively teach a deployed GPT-3, substantially increasing its accuracy over the queries with different kinds of misunderstandings by the GPT-3. Our approach is a step towards the low-cost utility enhancement for very large pre-trained LMs. All the code and data is available at this https URL.

Authors: Aman Madaan, Niket Tandon, Peter Clark, Yiming Yang

Links:

Merch: store.ykilcher.com

TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://discord.gg/4H8xxDF

BitChute: https://www.bitchute.com/channel/yann...

LinkedIn: https://www.linkedin.com/in/ykilcher

BiliBili: https://space.bilibili.com/2017636191

If you want to support me, the best thing to do is to share out the content :)

If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https://www.subscribestar.com/yannick...

Patreon: https://www.patreon.com/yannickilcher

Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq

Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2

Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m

Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n

View Details

deeplearning #nlp #sampling

This is an interview with first author Clara Meister.

Paper review video hereé https://youtu.be/_EDr3ryrT_Y

Modern language models like T5 or GPT-3 achieve remarkably low perplexities on both training and validation data, yet when sampling from their output distributions, the generated text often seems dull and uninteresting. Various workarounds have been proposed, such as top-k sampling and nucleus sampling, but while these manage to somewhat improve the generated samples, they are hacky and unfounded. This paper introduces typical sampling, a new decoding method that is principled, effective, and can be implemented efficiently. Typical sampling turns away from sampling purely based on likelihood and explicitly finds a trade-off between generating high-probability samples and generating high-information samples. The paper connects typical sampling to psycholinguistic theories on human speech generation, and shows experimentally that typical sampling achieves much more diverse and interesting results than any of the current methods.

Sponsor: Introduction to Graph Neural Networks Course

https://www.graphneuralnets.com/p/int...

OUTLINE:

0:00 - Intro

0:35 - Sponsor: Introduction to GNNs Course (link in description)

1:30 - Why does sampling matter?

5:40 - What is a "typical" message?

8:35 - How do humans communicate?

10:25 - Why don't we just sample from the model's distribution?

15:30 - What happens if we condition on the information to transmit?

17:35 - Does typical sampling really represent human outputs?

20:55 - What do the plots mean?

31:00 - Diving into the experimental results

39:15 - Are our training objectives wrong?

41:30 - Comparing typical sampling to top-k and nucleus sampling

44:50 - Explaining arbitrary engineering choices

47:20 - How can people get started with this?

Paper: https://arxiv.org/abs/2202.00666

Code: https://github.com/cimeister/typical-...

Abstract:

Despite achieving incredibly low perplexities on myriad natural language corpora, today's language models still often underperform when used to generate text. This dichotomy has puzzled the language generation community for the last few years. In this work, we posit that the abstraction of natural language as a communication channel (à la Shannon, 1948) can provide new insights into the behaviors of probabilistic language generators, e.g., why high-probability texts can be dull or repetitive. Humans use language as a means of communicating information, and do so in a simultaneously efficient and error-minimizing manner; they choose each word in a string with this (perhaps subconscious) goal in mind. We propose that generation from probabilistic models should mimic this behavior. Rather than always choosing words from the high-probability region of the distribution--which have a low Shannon information content--we sample from the set of words with information content close to the conditional entropy of our model, i.e., close to the expected information content. This decision criterion can be realized through a simple and efficient implementation, which we call typical sampling. Automatic and human evaluations show that, in comparison to nucleus and top-k sampling, typical sampling offers competitive performance in terms of quality while consistently reducing the number of degenerate repetitions.

Authors: Clara Meister, Tiago Pimentel, Gian Wiher, Ryan Cotterell

Links:

Merch: store.ykilcher.com

TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://discord.gg/4H8xxDF

View Details

deeplearning #nlp #sampling

Modern language models like T5 or GPT-3 achieve remarkably low perplexities on both training and validation data, yet when sampling from their output distributions, the generated text often seems dull and uninteresting. Various workarounds have been proposed, such as top-k sampling and nucleus sampling, but while these manage to somewhat improve the generated samples, they are hacky and unfounded. This paper introduces typical sampling, a new decoding method that is principled, effective, and can be implemented efficiently. Typical sampling turns away from sampling purely based on likelihood and explicitly finds a trade-off between generating high-probability samples and generating high-information samples. The paper connects typical sampling to psycholinguistic theories on human speech generation, and shows experimentally that typical sampling achieves much more diverse and interesting results than any of the current methods.

Sponsor: Fully Connected by Weights & Biases

https://wandb.ai/fully-connected

OUTLINE:

0:00 - Intro

1:50 - Sponsor: Fully Connected by Weights & Biases

4:10 - Paper Overview

7:40 - What's the problem with sampling?

11:45 - Beam Search: The good and the bad

14:10 - Top-k and Nucleus Sampling

16:20 - Why the most likely things might not be the best

21:30 - The expected information content of the next word

25:00 - How to trade off information and likelihood

31:25 - Connections to information theory and psycholinguistics

36:40 - Introducing Typical Sampling

43:00 - Experimental Evaluation

44:40 - My thoughts on this paper

Paper: https://arxiv.org/abs/2202.00666

Code: https://github.com/cimeister/typical-...

Abstract:

Despite achieving incredibly low perplexities on myriad natural language corpora, today's language models still often underperform when used to generate text. This dichotomy has puzzled the language generation community for the last few years. In this work, we posit that the abstraction of natural language as a communication channel (à la Shannon, 1948) can provide new insights into the behaviors of probabilistic language generators, e.g., why high-probability texts can be dull or repetitive. Humans use language as a means of communicating information, and do so in a simultaneously efficient and error-minimizing manner; they choose each word in a string with this (perhaps subconscious) goal in mind. We propose that generation from probabilistic models should mimic this behavior. Rather than always choosing words from the high-probability region of the distribution--which have a low Shannon information content--we sample from the set of words with information content close to the conditional entropy of our model, i.e., close to the expected information content. This decision criterion can be realized through a simple and efficient implementation, which we call typical sampling. Automatic and human evaluations show that, in comparison to nucleus and top-k sampling, typical sampling offers competitive performance in terms of quality while consistently reducing the number of degenerate repetitions.

Authors: Clara Meister, Tiago Pimentel, Gian Wiher, Ryan Cotterell

Links:

Merch: store.ykilcher.com

TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://discord.gg/4H8xxDF

BitChute: https://www.bitchute.com/channel/yann...

LinkedIn: https://www.linkedin.com/in/ykilcher

BiliBili: https://space.bilibili.com/2017636191

If you want to support me, the best thing to do is to share out the content :)

View Details

blip #interview #salesforce

Paper Review Video: https://youtu.be/X2k7n4FuI7c

Sponsor: Assembly AI

https://www.assemblyai.com/?utm_sourc...

This is an interview with Junnan Li and Dongxu Li, authors of BLIP and members of Salesforce research.

Cross-modal pre-training has been all the rage lately in deep learning, especially training vision and language models together. However, there are a number of issues, such as low quality datasets that limit the performance of any model trained on it, and also the fact that pure contrastive pre-training cannot be easily fine-tuned for most downstream tasks. BLIP unifies different tasks and objectives in a single pre-training run and achieves a much more versatile model, which the paper immediately uses to create, filter, clean and thus bootstrap its own dataset to improve performance even more!

OUTLINE:

0:00 - Intro

0:40 - Sponsor: Assembly AI

1:30 - Start of Interview

2:30 - What's the pitch?

4:40 - How did data bootstrapping come into the project?

7:10 - How big of a problem is data quality?

11:10 - Are the captioning & filtering models biased towards COCO data?

14:40 - Could the data bootstrapping be done multiple times?

16:20 - What was the evolution of the BLIP architecture?

21:15 - Are there additional benefits to adding language modelling?

23:50 - Can we imagine a modular future for pre-training?

29:45 - Diving into the experimental results

42:40 - What did and did not work out during the research?

45:00 - How is research life at Salesforce?

46:45 - Where do we go from here?

Paper: https://arxiv.org/abs/2201.12086

Code: https://github.com/salesforce/BLIP

Demo: https://huggingface.co/spaces/Salesfo...

Abstract:

Vision-Language Pre-training (VLP) has advanced the performance for many vision-language tasks. However, most existing pre-trained models only excel in either understanding-based tasks or generation-based tasks. Furthermore, performance improvement has been largely achieved by scaling up the dataset with noisy image-text pairs collected from the web, which is a suboptimal source of supervision. In this paper, we propose BLIP, a new VLP framework which transfers flexibly to both vision-language understanding and generation tasks. BLIP effectively utilizes the noisy web data by bootstrapping the captions, where a captioner generates synthetic captions and a filter removes the noisy ones. We achieve state-of-the-art results on a wide range of vision-language tasks, such as image-text retrieval (+2.7% in average recall@1), image captioning (+2.8% in CIDEr), and VQA (+1.6% in VQA score). BLIP also demonstrates strong generalization ability when directly transferred to video-language tasks in a zero-shot manner. Code, models, and datasets are released at this https URL.

Authors: Junnan Li, Dongxu Li, Caiming Xiong, Steven Hoi

Links:

TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://discord.gg/4H8xxDF

BitChute: https://www.bitchute.com/channel/yann...

LinkedIn: https://www.linkedin.com/in/ykilcher

BiliBili: https://space.bilibili.com/2017636191

If you want to support me, the best thing to do is to share out the content :)

If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https://www.subscribestar.com/yannick...

Patreon: https://www.patreon.com/yannickilcher

Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq

Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2

View Details

blip #review #ai

Cross-modal pre-training has been all the rage lately in deep learning, especially training vision and language models together. However, there are a number of issues, such as low quality datasets that limit the performance of any model trained on it, and also the fact that pure contrastive pre-training cannot be easily fine-tuned for most downstream tasks. BLIP unifies different tasks and objectives in a single pre-training run and achieves a much more versatile model, which the paper immediately uses to create, filter, clean and thus bootstrap its own dataset to improve performance even more!

Sponsor: Zeta Alpha

https://zeta-alpha.com

Use code YANNIC for 20% off!

OUTLINE:

0:00 - Intro

0:50 - Sponsor: Zeta Alpha

3:40 - Paper Overview

6:40 - Vision-Language Pre-Training

11:15 - Contributions of the paper

14:30 - Model architecture: many parts for many tasks

19:50 - How data flows in the model

26:50 - Parameter sharing between the modules

29:45 - Captioning & Filtering bootstrapping

41:10 - Fine-tuning the model for downstream tasks

Paper: https://arxiv.org/abs/2201.12086

Code: https://github.com/salesforce/BLIP

Demo: https://huggingface.co/spaces/Salesfo...

Abstract:

Vision-Language Pre-training (VLP) has advanced the performance for many vision-language tasks. However, most existing pre-trained models only excel in either understanding-based tasks or generation-based tasks. Furthermore, performance improvement has been largely achieved by scaling up the dataset with noisy image-text pairs collected from the web, which is a suboptimal source of supervision. In this paper, we propose BLIP, a new VLP framework which transfers flexibly to both vision-language understanding and generation tasks. BLIP effectively utilizes the noisy web data by bootstrapping the captions, where a captioner generates synthetic captions and a filter removes the noisy ones. We achieve state-of-the-art results on a wide range of vision-language tasks, such as image-text retrieval (+2.7% in average recall@1), image captioning (+2.8% in CIDEr), and VQA (+1.6% in VQA score). BLIP also demonstrates strong generalization ability when directly transferred to video-language tasks in a zero-shot manner. Code, models, and datasets are released at this https URL.

Authors: Junnan Li, Dongxu Li, Caiming Xiong, Steven Hoi

Links:

TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://discord.gg/4H8xxDF

BitChute: https://www.bitchute.com/channel/yann...

LinkedIn: https://www.linkedin.com/in/ykilcher

BiliBili: https://space.bilibili.com/2017636191

If you want to support me, the best thing to do is to share out the content :)

If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https://www.subscribestar.com/yannick...

Patreon: https://www.patreon.com/yannickilcher

Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq

Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2

Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m

Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n

View Details

mlnews #gtc22 #ithaca

GTC Registration Link: https://ykilcher.com/gtc

Your regular updates on what's going on in the ML world!

OUTLINE:

0:00 - Intro

0:20 - Register to Nvidia GTC and win a 3090!

4:15 - DeepMind's Ithaca deciphers Lost Ancient Texts

6:45 - Drug discovery model turns toxic

10:00 - Gary Marcus: Deep Learning is hitting a wall

19:40 - GopherCite: Backing up answers with citations

22:40 - Yoshua Bengio appointed knight of the legion of honour

23:00 - Meta AI tags parody account of Yoshua Bengio

23:40 - Building games using just natural language

24:55 - YOU.com adds writing assistant

25:45 - Horace He: How to brrr

26:35 - Karpathy: Reproducing Yann LeCun's 1989 paper

27:50 - Pig grunt emotion classifier

28:20 - AI annotates protein domain functions

29:40 - Atwood & Carmack: 10k self-driving car bet

30:50 - Helpful Things

References:

Register to GTC and win a 3090!

https://twitter.com/NVIDIAEU/status/1...

https://www.nvidia.com/gtc/keynote/?n...

https://www.nvidia.com/gtc/?ncid=ref-...

https://www.nvidia.com/gtc/keynote/

https://www.nvidia.com/gtc/training/

https://developer.nvidia.com/nvidia-o...

DeepMind deciphers Lost Ancient Texts

https://deepmind.com/blog/article/Pre...

https://www.nature.com/articles/s4158...

https://github.com/deepmind/ithaca

https://ithaca.deepmind.com/?job=eyJy...

Drug discovery model turns toxic

https://www.theverge.com/2022/3/17/22...

https://www.nature.com/articles/s4225...

Gary Marcus: Deep Learning is hitting a wall

https://nautil.us/deep-learning-is-hi...

https://www.youtube.com/watch?v=fVkXE...

GopherCite: Backing up answers with citations

https://deepmind.com/research/publica...

Yoshua Bengio appointed knight of the legion of honour

https://mila.quebec/en/professor-yosh...

Meta AI tags parody account

https://twitter.com/MetaAI/status/150...

Building games using just natural language

https://andrewmayneblog.wordpress.com...

YOU.com adds writing assistant

https://you.com/search?q=how%20to%20w...

Horace He: How to brrr

https://horace.io/brrr_intro.html

Karpathy: Reproducing Yann LeCun's 1989 paper

https://karpathy.github.io/2022/03/14...

Pig grunt emotion classifier

https://science.ku.dk/english/press/n...

AI annotates protein domain functions

https://ai.googleblog.com/2022/03/usi...

https://google-research.github.io/pro...

Atwood & Carmack: 10k self-driving car bet

https://blog.codinghorror.com/the-203...

Helpful Things

https://github.com/recognai/rubrix

https://twitter.com/taiyasaki/status/...

https://github.com/mosaicml/composer?...

https://mujoco.org/

https://mujoco.readthedocs.io/en/late...

https://github.com/deepmind/mctx?utm_...

https://padl.ai/

https://github.com/LaihoE/did-it-spill

https://pytorch.org/blog/pytorch-1.11...

Links:

TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://discord.gg/4H8xxDF

BitChute: https://www.bitchute.com/channel/yann...

LinkedIn: https://www.linkedin.com/in/ykilcher

BiliBili: https://space.bilibili.com/2017636191

If you want to support me, the best thing to do is to share out the content :)

View Details

multitasklearning #biology #neuralnetworks

This is an interview with the paper's authors: Abhiram Iyer, Karan Grewal, and Akash Velu!

Paper Review Video: https://youtu.be/O_dJ31T01i8

Check out Zak's course on Graph Neural Networks (discount with this link): https://www.graphneuralnets.com/p/int...

Catastrophic forgetting is a big problem in mutli-task and continual learning. Gradients of different objectives tend to conflict, and new tasks tend to override past knowledge. In biological neural networks, each neuron carries a complex network of dendrites that mitigate such forgetting by recognizing the context of an input signal. This paper introduces Active Dendrites, which carries over the principle of context-sensitive gating by dendrites into the deep learning world. Various experiments show the benefit in combatting catastrophic forgetting, while preserving sparsity and limited parameter counts.

OUTLINE:

0:00 - Intro

0:55 - Sponsor: GNN Course

2:30 - How did the idea come to be?

7:05 - What roles do the different parts of the method play?

8:50 - What was missing in the paper review?

10:35 - Are biological concepts viable if we still have backprop?

11:50 - How many dendrites are necessary?

14:10 - Why is there a plateau in the sparsity plot?

20:50 - How does task difficulty play into the algorithm?

24:10 - Why are there different setups in the experiments?

30:00 - Is there a place for unsupervised pre-training?

32:50 - How can we apply the online prototyping to more difficult tasks?

37:00 - What did not work out during the project?

41:30 - How do you debug a project like this?

47:10 - How is this related to other architectures?

51:10 - What other things from neuroscience are to be included?

55:50 - Don't miss the awesome ending :)

Paper: https://arxiv.org/abs/2201.00042

Blog: https://numenta.com/blog/2021/11/08/c...

Link to the GNN course (with discount): https://www.graphneuralnets.com/p/int...

Authors: Abhiram Iyer, Karan Grewal, Akash Velu, Lucas Oliveira Souza, Jeremy Forest, Subutai Ahmad

Links:

TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://discord.gg/4H8xxDF

BitChute: https://www.bitchute.com/channel/yann...

LinkedIn: https://www.linkedin.com/in/ykilcher

View Details

multitasklearning #biology #neuralnetworks

Catastrophic forgetting is a big problem in mutli-task and continual learning. Gradients of different objectives tend to conflict, and new tasks tend to override past knowledge. In biological neural networks, each neuron carries a complex network of dendrites that mitigate such forgetting by recognizing the context of an input signal. This paper introduces Active Dendrites, which carries over the principle of context-sensitive gating by dendrites into the deep learning world. Various experiments show the benefit in combatting catastrophic forgetting, while preserving sparsity and limited parameter counts.

OUTLINE:

0:00 - Introduction

1:20 - Paper Overview

3:15 - Catastrophic forgetting in continuous and multi-task learning

9:30 - Dendrites in biological neurons

16:55 - Sparse representations in biology

18:35 - Active dendrites in deep learning

34:15 - Experiments on multi-task learning

39:00 - Experiments in continual learning and adaptive prototyping

49:20 - Analyzing the inner workings of the algorithm

53:30 - Is this the same as just training a larger network?

59:15 - How does this relate to attention mechanisms?

1:02:55 - Final thoughts and comments

Paper: https://arxiv.org/abs/2201.00042

Blog: https://numenta.com/blog/2021/11/08/c...

ERRATA:

  • I was made aware of this by https://twitter.com/ChainlessCoder: "That axon you showed of the pyramidal neuron, is actually the apical dendrite of the neuron". Sorry, my bad :)

Abstract:

A key challenge for AI is to build embodied systems that operate in dynamically changing environments. Such systems must adapt to changing task contexts and learn continuously. Although standard deep learning systems achieve state of the art results on static benchmarks, they often struggle in dynamic scenarios. In these settings, error signals from multiple contexts can interfere with one another, ultimately leading to a phenomenon known as catastrophic forgetting. In this article we investigate biologically inspired architectures as solutions to these problems. Specifically, we show that the biophysical properties of dendrites and local inhibitory systems enable networks to dynamically restrict and route information in a context-specific manner. Our key contributions are as follows. First, we propose a novel artificial neural network architecture that incorporates active dendrites and sparse representations into the standard deep learning framework. Next, we study the performance of this architecture on two separate benchmarks requiring task-based adaptation: Meta-World, a multi-task reinforcement learning environment where a robotic agent must learn to solve a variety of manipulation tasks simultaneously; and a continual learning benchmark in which the model's prediction task changes throughout training. Analysis on both benchmarks demonstrates the emergence of overlapping but distinct and sparse subnetworks, allowing the system to fluidly learn multiple tasks with minimal forgetting. Our neural implementation marks the first time a single architecture has achieved competitive results on both multi-task and continual learning settings. Our research sheds light on how biological properties of neurons can inform deep learning systems to address dynamic scenarios that are typically impossible for traditional ANNs to solve.

Authors: Abhiram Iyer, Karan Grewal, Akash Velu, Lucas Oliveira Souza, Jeremy Forest, Subutai Ahmad

View Details

deeplearning #objectdetection #outliers

An interview with the authors of "Virtual Outlier Synthesis".

Watch the paper review video here: https://youtu.be/i-J4T3uLC9M

Outliers are data points that are highly unlikely to be seen in the training distribution, and therefore deep neural networks have troubles when dealing with them. Many approaches to detecting outliers at inference time have been proposed, but most of them show limited success. This paper presents Virtual Outlier Synthesis, which is a method that pairs synthetic outliers, forged in the latent space, with an energy-based regularization of the network at training time. The result is a deep network that can reliably detect outlier datapoints during inference with minimal overhead.

OUTLINE:

0:00 - Intro

2:20 - What was the motivation behind this paper?

5:30 - Why object detection?

11:05 - What's the connection to energy-based models?

12:15 - Is a Gaussian mixture model appropriate for high-dimensional data?

16:15 - What are the most important components of the method?

18:30 - What are the downstream effects of the regularizer?

22:00 - Are there severe trade-offs to outlier detection?

23:55 - Main experimental takeaways?

26:10 - Why do outlier detection in the last layer?

30:20 - What does it take to finish a research projects successfully?

Paper: https://arxiv.org/abs/2202.01197

Code: https://github.com/deeplearning-wisc/vos

Abstract:

Out-of-distribution (OOD) detection has received much attention lately due to its importance in the safe deployment of neural networks. One of the key challenges is that models lack supervision signals from unknown data, and as a result, can produce overconfident predictions on OOD data. Previous approaches rely on real outlier datasets for model regularization, which can be costly and sometimes infeasible to obtain in practice. In this paper, we present VOS, a novel framework for OOD detection by adaptively synthesizing virtual outliers that can meaningfully regularize the model's decision boundary during training. Specifically, VOS samples virtual outliers from the low-likelihood region of the class-conditional distribution estimated in the feature space. Alongside, we introduce a novel unknown-aware training objective, which contrastively shapes the uncertainty space between the ID data and synthesized outlier data. VOS achieves state-of-the-art performance on both object detection and image classification models, reducing the FPR95 by up to 7.87% compared to the previous best method. Code is available at this https URL.

Authors: Xuefeng Du, Zhaoning Wang, Mu Cai, Yixuan Li

Links:

Merch: store.ykilcher.com

TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://discord.gg/4H8xxDF

BitChute: https://www.bitchute.com/channel/yann...

LinkedIn: https://www.linkedin.com/in/ykilcher

BiliBili: https://space.bilibili.com/2017636191

If you want to support me, the best thing to do is to share out the content :)

If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https://www.subscribestar.com/yannick...

Patreon: https://www.patreon.com/yannickilcher

Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq

Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2

Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m

Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n

View Details

vos #outliers #deeplearning

Sponsor: Assembly AI

Check them out here: https://www.assemblyai.com/?utm_sourc...

Outliers are data points that are highly unlikely to be seen in the training distribution, and therefore deep neural networks have troubles when dealing with them. Many approaches to detecting outliers at inference time have been proposed, but most of them show limited success. This paper presents Virtual Outlier Synthesis, which is a method that pairs synthetic outliers, forged in the latent space, with an energy-based regularization of the network at training time. The result is a deep network that can reliably detect outlier datapoints during inference with minimal overhead.

OUTLINE:

0:00 - Intro

2:00 - Sponsor: Assembly AI (Link below)

4:05 - Paper Overview

6:45 - Where do traditional classifiers fail?

11:00 - How object detectors work

17:00 - What are virtual outliers and how are they created?

24:00 - Is this really an appropriate model for outliers?

26:30 - How virtual outliers are used during training

34:00 - Plugging it all together to detect outliers

Paper: https://arxiv.org/abs/2202.01197

Code: https://github.com/deeplearning-wisc/vos

Abstract:

Out-of-distribution (OOD) detection has received much attention lately due to its importance in the safe deployment of neural networks. One of the key challenges is that models lack supervision signals from unknown data, and as a result, can produce overconfident predictions on OOD data. Previous approaches rely on real outlier datasets for model regularization, which can be costly and sometimes infeasible to obtain in practice. In this paper, we present VOS, a novel framework for OOD detection by adaptively synthesizing virtual outliers that can meaningfully regularize the model's decision boundary during training. Specifically, VOS samples virtual outliers from the low-likelihood region of the class-conditional distribution estimated in the feature space. Alongside, we introduce a novel unknown-aware training objective, which contrastively shapes the uncertainty space between the ID data and synthesized outlier data. VOS achieves state-of-the-art performance on both object detection and image classification models, reducing the FPR95 by up to 7.87% compared to the previous best method. Code is available at this https URL.

Authors: Xuefeng Du, Zhaoning Wang, Mu Cai, Yixuan Li

Links:

Merch: store.ykilcher.com

TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://discord.gg/4H8xxDF

BitChute: https://www.bitchute.com/channel/yann...

LinkedIn: https://www.linkedin.com/in/ykilcher

BiliBili: https://space.bilibili.com/2017636191

If you want to support me, the best thing to do is to share out the content :)

If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https://www.subscribestar.com/yannick...

Patreon: https://www.patreon.com/yannickilcher

Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq

Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2

Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m

Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n

View Details

deepmind #rl #society

This is an in-depth paper review, followed by an interview with the papers' authors!

Society is ruled by norms, and most of these norms are very useful, such as washing your hands before cooking. However, there also exist plenty of social norms which are essentially arbitrary, such as what hairstyles are acceptable, or what words are rude. These are called "silly rules". This paper uses multi-agent reinforcement learning to investigate why such silly rules exist. Their results indicate a plausible mechanism, by which the existence of silly rules drastically speeds up the agents' acquisition of the skill of enforcing rules, which generalizes well, and therefore a society that has silly rules will be better at enforcing rules in general, leading to faster adaptation in the face of genuinely useful norms.

OUTLINE:

0:00 - Intro

3:00 - Paper Overview

5:20 - Why are some social norms arbitrary?

11:50 - Reinforcement learning environment setup

20:00 - What happens if we introduce a "silly" rule?

25:00 - Experimental Results: how silly rules help society

30:10 - Isolated probing experiments

34:30 - Discussion of the results

37:30 - Start of Interview

39:30 - Where does the research idea come from?

44:00 - What is the purpose behind this research?

49:20 - Short recap of the mechanics of the environment

53:00 - How much does such a closed system tell us about the real world?

56:00 - What do the results tell us about silly rules?

1:01:00 - What are these agents really learning?

1:08:00 - How many silly rules are optimal?

1:11:30 - Why do you have separate weights for each agent?

1:13:45 - What features could be added next?

1:16:00 - How sensitive is the system to hyperparameters?

1:17:20 - How to avoid confirmation bias?

1:23:15 - How does this play into progress towards AGI?

1:29:30 - Can we make real-world recommendations based on this?

1:32:50 - Where do we go from here?

Paper: https://www.pnas.org/doi/10.1073/pnas...

Blog: https://deepmind.com/research/publica...

Abstract:

The fact that humans enforce and comply with norms is an important reason why humans enjoy higher levels of cooperation and welfare than other animals. Some norms are relatively easy to explain; they may prohibit obviously harmful or uncooperative actions. But many norms are not easy to explain. For example, most cultures prohibit eating certain kinds of foods and almost all societies have rules about what constitutes appropriate clothing, language, and gestures. Using a computational model focused on learning shows that apparently pointless rules can have an indirect effect on welfare. They can help agents learn how to enforce and comply with norms in general, improving the group’s ability to enforce norms that have a direct effect on welfare.

Authors: Raphael Köster, Dylan Hadfield-Menell, Richard Everett, Laura Weidinger, Gillian K. Hadfield, Joel Z. Leibo

Links:

TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://discord.gg/4H8xxDF

BitChute: https://www.bitchute.com/channel/yann...

LinkedIn: https://www.linkedin.com/in/ykilcher

BiliBili: https://space.bilibili.com/2017636191

If you want to support me, the best thing to do is to share out the content :)

View Details

openai #math #imo

This is an interview with Stanislas Polu, research engineer at OpenAI and first author of the paper "Formal Mathematics Statement Curriculum Learning".

Watch the paper review here: https://youtu.be/lvYVuOmUVs8

OUTLINE:

0:00 - Intro

2:00 - How do you explain the big public reaction?

4:00 - What's the history behind the paper?

6:15 - How does algorithmic formal math work?

13:10 - How does expert iteration replace self-play?

22:30 - How is the language model trained and used?

30:50 - Why is every model fine-tuned on the initial state?

33:05 - What if we want to prove something we don't know already?

40:35 - How can machines and humans work together?

43:40 - Aren't most produced statements useless?

46:20 - A deeper look at the experimental results

50:10 - What were the high and low points during the research?

54:25 - Where do we go from here?

Paper: https://arxiv.org/abs/2202.01344

miniF2F benchmark: https://github.com/openai/miniF2F

Follow Stan here: https://twitter.com/spolu

Abstract:

We explore the use of expert iteration in the context of language modeling applied to formal mathematics. We show that at same compute budget, expert iteration, by which we mean proof search interleaved with learning, dramatically outperforms proof search only. We also observe that when applied to a collection of formal statements of sufficiently varied difficulty, expert iteration is capable of finding and solving a curriculum of increasingly difficult problems, without the need for associated ground-truth proofs. Finally, by applying this expert iteration to a manually curated set of problem statements, we achieve state-of-the-art on the miniF2F benchmark, automatically solving multiple challenging problems drawn from high school olympiads.

Authors: Stanislas Polu, Jesse Michael Han, Kunhao Zheng, Mantas Baksys, Igor Babuschkin, Ilya Sutskever

Links:

TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://discord.gg/4H8xxDF

BitChute: https://www.bitchute.com/channel/yann...

LinkedIn: https://www.linkedin.com/in/ykilcher

BiliBili: https://space.bilibili.com/2017636191

If you want to support me, the best thing to do is to share out the content :)

If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https://www.subscribestar.com/yannick...

Patreon: https://www.patreon.com/yannickilcher

Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq

Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2

Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m

Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n

View Details

openai #math #imo

Formal mathematics is a challenging area for both humans and machines. For humans, formal proofs require very tedious and meticulous specifications of every last detail and results in very long, overly cumbersome and verbose outputs. For machines, the discreteness and sparse reward nature of the problem presents a significant problem, which is classically tackled by brute force search, guided by a couple of heuristics. Previously, language models have been employed to better guide these proof searches and delivered significant improvements, but automated systems are still far from usable. This paper introduces another concept: An expert iteration procedure is employed to iteratively produce more and more challenging, but solvable problems for the machine to train on, which results in an automated curriculum, and a final algorithm that performs well above the previous models. OpenAI used this method to even solve two problems of the international math olympiad, which was previously infeasible for AI systems.

OUTLINE:

0:00 - Intro

2:35 - Paper Overview

5:50 - How do formal proofs work?

9:35 - How expert iteration creates a curriculum

16:50 - Model, data, and training procedure

25:30 - Predicting proof lengths for guiding search

29:10 - Bootstrapping expert iteration

34:10 - Experimental evaluation & scaling properties

40:10 - Results on synthetic data

44:15 - Solving real math problems

47:15 - Discussion & comments

Paper: https://arxiv.org/abs/2202.01344

miniF2F benchmark: https://github.com/openai/miniF2F

Abstract:

We explore the use of expert iteration in the context of language modeling applied to formal mathematics. We show that at same compute budget, expert iteration, by which we mean proof search interleaved with learning, dramatically outperforms proof search only. We also observe that when applied to a collection of formal statements of sufficiently varied difficulty, expert iteration is capable of finding and solving a curriculum of increasingly difficult problems, without the need for associated ground-truth proofs. Finally, by applying this expert iteration to a manually curated set of problem statements, we achieve state-of-the-art on the miniF2F benchmark, automatically solving multiple challenging problems drawn from high school olympiads.

Authors: Stanislas Polu, Jesse Michael Han, Kunhao Zheng, Mantas Baksys, Igor Babuschkin, Ilya Sutskever

Links:

TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://discord.gg/4H8xxDF

BitChute: https://www.bitchute.com/channel/yann...

LinkedIn: https://www.linkedin.com/in/ykilcher

BiliBili: https://space.bilibili.com/2017636191

If you want to support me, the best thing to do is to share out the content :)

If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https://www.subscribestar.com/yannick...

Patreon: https://www.patreon.com/yannickilcher

Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq

Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2

Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m

Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n

View Details

Updates on what's going on in the ML world!

Check out w&b's alerts feature: https://wandb.me/yannic

OUTLINE:

0:00 - Intro

0:20 - Sponsor: Weights & Biases

2:35 - DeepMind uses Reinforcement Learning to control nuclear fusion

4:35 - Google responds to carbon emission estimates

8:40 - Yann LeCun proposes new architecture for world models

11:05 - Fruit fly neurons may perform multiplication

12:00 - Emojisearch App

12:30 - Ar5iv officially in arXiv labs

12:55 - Language Model Consciousness & Media Hype

16:45 - Vision models are more fair when trained on uncurated data

18:30 - CLIPasso

19:15 - NLP with Transformers Book

20:15 - Helpful Things

26:00 - US Office: AI can't copyright its art

Sponsor: Weights & Biases

https://wandb.me/yannic

References:

https://wandb.me/yannic

DeepMind uses RL to control nuclear fusion

https://deepmind.com/blog/article/Acc...

https://www.nature.com/articles/s4158...

https://www.nature.com/articles/s4158...

https://www.alexirpan.com/2018/02/14/...

Google responds to carbon emission estimates

https://ai.googleblog.com/2022/02/goo...

Yann LeCun proposes new architecture for world models

https://ai.facebook.com/blog/yann-lec...

Fruit fly neurons may perform multiplication

https://www.nature.com/articles/s4158...

Emojisearch App

https://twitter.com/lilianweng/status...

https://www.emojisearch.app/

https://github.com/lilianweng/emoji-s...

Ar5iv officially in arXiv labs

https://blog.arxiv.org/2022/02/21/arx...

Tech media may be only slightly conscious

https://twitter.com/ilyasut/status/14...

https://futurism.com/the-byte/openai-...

https://interestingengineering.com/ai...

https://futurism.com/mit-researcher-c...

https://www.dailymail.co.uk/sciencete...

https://futurism.com/conscious-ai-bac...

https://www.dailystar.co.uk/tech/news...

Vision models are more fair when trained on uncurated data

https://arxiv.org/pdf/2202.08360.pdf

CLIPasso

https://clipasso.github.io/clipasso/

NLP with Transformers Book

https://www.amazon.de/dp/1098103246?l...

Helpful Things

https://github.com/j3soon/tbparse

https://github.com/openvinotoolkit/an...

https://liuliu66.github.io/articulati...

https://github.com/RobertTLange/evosax

https://github.com/google/evojax

https://github.com/google/evojax/pull/9

https://github.com/facebookresearch/t...

https://standard-ai.github.io/Standar...

https://twitter.com/PatrickPlaten/sta...

https://aimagelab.ing.unimore.it/imag...

https://github.com/yashbhalgat/HashNe...

https://github.com/patrick-kidger/dif...

https://github.com/AI4Finance-Foundat...

https://huggingface.co/AI-Nordics/ber...

https://huggingface.co/AI-Nordics/gpt...

https://paperswithcode.com/dataset/muld

https://github.com/JonasGeiping/breac...

https://github.com/Weixin-Liang/MetaS...

US Office: AI can't copyright its art

https://www.theverge.com/2022/2/21/22...

https://www.urbasm.com/2016/05/artifi...

Links:

TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://discord.gg/4H8xxDF

BitChute: https://www.bitchute.com/channel/yann...

LinkedIn: https://www.linkedin.com/in/ykilcher

BiliBili: https://space.bilibili.com/2017636191

View Details

An interview with the creators of AlphaCode!

Paper review video here: https://youtu.be/s9UAOmyah1A

OUTLINE:

0:00 - Intro

1:10 - Media Reception

5:10 - How did the project go from start to finish?

9:15 - Does the model understand its own code?

14:45 - Are there plans to reduce the number of samples?

16:15 - Could one do smarter filtering of samples?

18:55 - How crucial are the public test cases?

21:55 - Could we imagine an adversarial method?

24:45 - How are coding problems even made?

27:40 - Does AlphaCode evaluate a solution's asymptotic complexity?

33:15 - Are our sampling procedures inappropriate for diversity?

36:30 - Are all generated solutions as instructive as the example?

41:30 - How are synthetic examples created during training?

42:30 - What were high and low points during this research?

45:25 - What was the most valid criticism after publication?

47:40 - What are applications in the real world?

51:00 - Where do we go from here?

Paper: https://storage.googleapis.com/deepmi...

Code: https://github.com/deepmind/code_cont...

Abstract: Programming is a powerful and ubiquitous problem-solving tool. Developing systems that can assist programmers or even generate programs independently could make programming more productive and accessible, yet so far incorporating innovations in AI has proven challenging. Recent large-scale language models have demonstrated an impressive ability to generate code, and are now able to complete simple programming tasks. However, these models still perform poorly when evaluated on more complex, unseen problems that require problem-solving skills beyond simply translating instructions into code. For example, competitive programming problems which require an understanding of algorithms and complex natural language remain extremely challenging. To address this gap, we introduce AlphaCode, a system for code generation that can create novel solutions to these problems that require deeper reasoning. Evaluated on recent programming competitions on the Codeforces platform, AlphaCode achieved on average a ranking of top 54.3% in programming competitions with more than 5,000 participants. We found that three key components were critical to achieve good and reliable performance: (1) an extensive and clean competitive programming dataset for training and evaluation, (2) large and efficient-to-sample transformer-based architectures, and (3) large-scale model sampling to explore the search space, followed by filtering based on program behavior to a small set of submissions.

Authors: Yujia Li, David Choi, Junyoung Chung, Nate Kushman, Julian Schrittwieser, Rémi Leblond, Tom Eccles, James Keeling, Felix Gimeno, Agustin Dal Lago, Thomas Hubert, Peter Choy, Cyprien de Masson d’Autume, Igor Babuschkin, Xinyun Chen, Po-Sen Huang, Johannes Welbl, Sven Gowal, Alexey Cherepanov, James Molloy, Daniel J. Mankowitz, Esme Sutherland Robson, Pushmeet Kohli, Nando de Freitas, Koray Kavukcuoglu and Oriol Vinyals

Links:

Merch: store.ykilcher.com

TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://discord.gg/4H8xxDF

BitChute: https://www.bitchute.com/channel/yann...

LinkedIn: https://www.linkedin.com/in/ykilcher

BiliBili: https://space.bilibili.com/2017636191

If you want to support me, the best thing to do is to share out the content :)

If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https://www.subscribestar.com/yannick...

Patreon: https://www.patreon.com/yannickilcher

View Details

ai #alphacode #deepmind

AlphaCode is an automated system that can solve competitive programing exercises. The authors found an interesting combination of language models, large-scale sampling, and clever techniques to filter and subsequently cluster the resulting programs, which lets the system perform on the level of an average competitor in real competitions. In this video, we take a deep dive into AlphaCode's design, architecture, and experimental evaluation. The paper is very well structured and the empirical results are super interesting!

OUTLINE:

0:00 - Intro

2:10 - Paper Overview

3:30 - An example problem from competitive programming

8:00 - AlphaCode system overview

14:00 - Filtering out wrong solutions

17:15 - Clustering equivalent generated programs

21:50 - Model configurations & engineering choices

24:30 - Adding privileged information to the input & more tricks

28:15 - Experimental Results (very interesting!)

Paper: https://storage.googleapis.com/deepmi...

Code: https://github.com/deepmind/code_cont...

Abstract: Programming is a powerful and ubiquitous problem-solving tool. Developing systems that can assist programmers or even generate programs independently could make programming more productive and accessible, yet so far incorporating innovations in AI has proven challenging. Recent large-scale language models have demonstrated an impressive ability to generate code, and are now able to complete simple programming tasks. However, these models still perform poorly when evaluated on more complex, unseen problems that require problem-solving skills beyond simply translating instructions into code. For example, competitive programming problems which require an understanding of algorithms and complex natural language remain extremely challenging. To address this gap, we introduce AlphaCode, a system for code generation that can create novel solutions to these problems that require deeper reasoning. Evaluated on recent programming competitions on the Codeforces platform, AlphaCode achieved on average a ranking of top 54.3% in programming competitions with more than 5,000 participants. We found that three key components were critical to achieve good and reliable performance: (1) an extensive and clean competitive programming dataset for training and evaluation, (2) large and efficient-to-sample transformer-based architectures, and (3) large-scale model sampling to explore the search space, followed by filtering based on program behavior to a small set of submissions.

Authors: Yujia Li, David Choi, Junyoung Chung, Nate Kushman, Julian Schrittwieser, Rémi Leblond, Tom Eccles, James Keeling, Felix Gimeno, Agustin Dal Lago, Thomas Hubert, Peter Choy, Cyprien de Masson d’Autume, Igor Babuschkin, Xinyun Chen, Po-Sen Huang, Johannes Welbl, Sven Gowal, Alexey Cherepanov, James Molloy, Daniel J. Mankowitz, Esme Sutherland Robson, Pushmeet Kohli, Nando de Freitas, Koray Kavukcuoglu and Oriol Vinyals

Links:

Merch: store.ykilcher.com

TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://discord.gg/4H8xxDF

BitChute: https://www.bitchute.com/channel/yann...

LinkedIn: https://www.linkedin.com/in/ykilcher

BiliBili: https://space.bilibili.com/2017636191

If you want to support me, the best thing to do is to share out the content :)

If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https://www.subscribestar.com/yannick...

Patreon: https://www.patreon.com/yannickilcher

Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq

View Details

wikipedia #reinforcementlearning #languagemodels

Original paper review here: https://youtu.be/XHGh19Hbx48

Machel Reid and Yutaro Yamada join me to discuss their recent paper on langauge model pre-training for decision transformers in offline reinforcement learning.

OUTLINE:

0:00 - Intro

1:00 - Brief paper, setup & idea recap

7:30 - Main experimental results & high standard deviations

10:00 - Why is there no clear winner?

13:00 - Why are bigger models not a lot better?

14:30 - What’s behind the name ChibiT?

15:30 - Why is iGPT underperforming?

19:15 - How are tokens distributed in Reinforcement Learning?

22:00 - What other domains could have good properties to transfer?

24:20 - A deeper dive into the models' attention patterns

33:30 - Codebase, model sizes, and compute requirements

37:30 - Scaling behavior of pre-trained models

40:05 - What did not work out in this project?

42:00 - How can people get started and where to go next?

Paper: https://arxiv.org/abs/2201.12122

Code: https://github.com/machelreid/can-wik...

My Video on Decision Transformer: https://youtu.be/-buULmf7dec

Abstract:

Fine-tuning reinforcement learning (RL) models has been challenging because of a lack of large scale off-the-shelf datasets as well as high variance in transferability among different environments. Recent work has looked at tackling offline RL from the perspective of sequence modeling with improved results as result of the introduction of the Transformer architecture. However, when the model is trained from scratch, it suffers from slow convergence speeds. In this paper, we look to take advantage of this formulation of reinforcement learning as sequence modeling and investigate the transferability of pre-trained sequence models on other domains (vision, language) when finetuned on offline RL tasks (control, games). To this end, we also propose techniques to improve transfer between these domains. Results show consistent performance gains in terms of both convergence speed and reward on a variety of environments, accelerating training by 3-6x and achieving state-of-the-art performance in a variety of tasks using Wikipedia-pretrained and GPT2 language models. We hope that this work not only brings light to the potentials of leveraging generic sequence modeling techniques and pre-trained models for RL, but also inspires future work on sharing knowledge between generative modeling tasks of completely different domains.

Authors: Machel Reid, Yutaro Yamada, Shixiang Shane Gu

Links:

TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://discord.gg/4H8xxDF

BitChute: https://www.bitchute.com/channel/yann...

LinkedIn: https://www.linkedin.com/in/ykilcher

BiliBili: https://space.bilibili.com/2017636191

If you want to support me, the best thing to do is to share out the content :)

If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https://www.subscribestar.com/yannick...

Patreon: https://www.patreon.com/yannickilcher

Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq

Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2

Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m

Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n

View Details

wikipedia #reinforcementlearning #languagemodels

Transformers have come to overtake many domain-targeted custom models in a wide variety of fields, such as Natural Language Processing, Computer Vision, Generative Modelling, and recently also Reinforcement Learning. This paper looks at the Decision Transformer and shows that, surprisingly, pre-training the model on a language-modelling task significantly boosts its performance on Offline Reinforcement Learning. The resulting model achieves higher scores, can get away with less parameters, and exhibits superior scaling properties. This raises many questions about the fundamental connection between the domains of language and RL.

OUTLINE:

0:00 - Intro

1:35 - Paper Overview

7:35 - Offline Reinforcement Learning as Sequence Modelling

12:00 - Input Embedding Alignment & other additions

16:50 - Main experimental results

20:45 - Analysis of the attention patterns across models

32:25 - More experimental results (scaling properties, ablations, etc.)

37:30 - Final thoughts

Paper: https://arxiv.org/abs/2201.12122

Code: https://github.com/machelreid/can-wik...

My Video on Decision Transformer: https://youtu.be/-buULmf7dec

Abstract:

Fine-tuning reinforcement learning (RL) models has been challenging because of a lack of large scale off-the-shelf datasets as well as high variance in transferability among different environments. Recent work has looked at tackling offline RL from the perspective of sequence modeling with improved results as result of the introduction of the Transformer architecture. However, when the model is trained from scratch, it suffers from slow convergence speeds. In this paper, we look to take advantage of this formulation of reinforcement learning as sequence modeling and investigate the transferability of pre-trained sequence models on other domains (vision, language) when finetuned on offline RL tasks (control, games). To this end, we also propose techniques to improve transfer between these domains. Results show consistent performance gains in terms of both convergence speed and reward on a variety of environments, accelerating training by 3-6x and achieving state-of-the-art performance in a variety of tasks using Wikipedia-pretrained and GPT2 language models. We hope that this work not only brings light to the potentials of leveraging generic sequence modeling techniques and pre-trained models for RL, but also inspires future work on sharing knowledge between generative modeling tasks of completely different domains.

Authors: Machel Reid, Yutaro Yamada, Shixiang Shane Gu

Links:

Merch: store.ykilcher.com

TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://discord.gg/4H8xxDF

BitChute: https://www.bitchute.com/channel/yann...

LinkedIn: https://www.linkedin.com/in/ykilcher

BiliBili: https://space.bilibili.com/2017636191

If you want to support me, the best thing to do is to share out the content :)

If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https://www.subscribestar.com/yannick...

Patreon: https://www.patreon.com/yannickilcher

Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq

Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2

Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m

Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n

View Details

mlnews #rsc #gpt3

Some things we've missed in recent weeks!

OUTLINE:

0:00 - Intro & Overview

0:40 - Meta builds AI Research Supercluster (RSC)

2:25 - OpenAI trains GPT-3 to follow instructions

4:10 - Meta AI releases multilingual language models

4:50 - Google LaMDA dialogue models

5:50 - Helpful Things

8:25 - Training the alpha matte generator for Pixel 6

10:15 - Drones used to deter pigeons on buildings

11:05 - IBM sells some Watson Health assets for USD 1B

Merch: store.ykilcher.com

References:

https://ai.facebook.com/blog/ai-rsc/?...

https://openai.com/blog/instruction-f...

https://cdn.openai.com/papers/Trainin...

https://openai.com/blog/deep-reinforc...

https://twitter.com/MetaAI/status/148...

https://arxiv.org/pdf/2112.10668.pdf

https://github.com/pytorch/fairseq/tr...

https://ai.googleblog.com/2022/01/lam...

https://arxiv.org/pdf/2201.08239.pdf

https://evolutiongym.github.io/?utm_s...

https://evolutiongym.github.io/all-tasks

https://evolutiongym.github.io/docume...

https://arxiv.org/pdf/2201.09863.pdf

https://github.com/EvolutionGym

https://huggingface.co/blog/sb3

https://twitter.com/Sentdex/status/14...

https://github.com/lvwerra/trl?utm_so...

https://ai.googleblog.com/2022/01/acc...

https://polyhaven.com/hdris

https://ieeexplore.ieee.org/document/...

https://www.bloomberg.com/news/articl...

https://archive.ph/xadf9

Links:

TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://discord.gg/4H8xxDF

BitChute: https://www.bitchute.com/channel/yann...

LinkedIn: https://www.linkedin.com/in/ykilcher

BiliBili: https://space.bilibili.com/2017636191

If you want to support me, the best thing to do is to share out the content :)

If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https://www.subscribestar.com/yannick...

Patreon: https://www.patreon.com/yannickilcher

Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq

Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2

Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m

Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n

View Details

mlnews #muzero #nerf

Your regularly irregular updates on everything new in the ML world!

Merch: store.ykilcher.com

OUTLINE:

0:00 - Intro

0:15 - Sponsor: Weights & Biases

2:15 - Uber switches from XGBoost to Deep Learning for ETA prediction

5:45 - MuZero advances video compression

10:10 - Learned Soft Prompts can steer large language models

12:45 - Block-NeRF captures entire city blocks

14:15 - Neural Architecture Search considers underlying hardware

16:50 - Mega-Blog on Self-Organizing Agents

18:40 - Know Your Data (for Tensorflow Datasets)

20:30 - Helpful Things

Sponsor: Weights & Biases

https://wandb.me/yannic

References:

https://docs.wandb.ai/guides/integrat...

https://colab.research.google.com/git...

https://wandb.ai/borisd13/GPT-3/repor...

Uber switches from XGBoost to Deep Learning for ETA prediction

https://eng.uber.com/deepeta-how-uber...

MuZero advances video compression

https://deepmind.com/blog/article/MuZ...

https://storage.googleapis.com/deepmi...

Learned Soft Prompts can steer large language models

https://ai.googleblog.com/2022/02/gui...

https://aclanthology.org/2021.emnlp-m...

Block-NeRF captures entire city blocks

https://arxiv.org/abs/2202.05263

https://arxiv.org/pdf/2202.05263.pdf

https://waymo.com/intl/zh-cn/research...

Neural Architecture Search considers underlying hardware

https://ai.googleblog.com/2022/02/unl...

https://openaccess.thecvf.com/content...

Mega-Blog on Self-Organizing Agents

https://developmentalsystems.org/sens...

https://flowers.inria.fr/

Know Your Data (for Tensorflow Datasets)

https://knowyourdata-tfds.withgoogle....

https://knowyourdata.withgoogle.com/

Helpful Things

https://twitter.com/casualganpapers/s...

https://www.reddit.com/r/MachineLearn...

https://arxiv.org/abs/2202.02435

https://github.com/vicariousinc/PGMax

https://www.vicarious.com/posts/pgmax...

https://diambra.ai/tournaments

https://github.com/diambra/diambraArena

https://www.youtube.com/watch?v=dw72P...

https://gitlab.com/deepcypher/python-...

https://python-fhez.readthedocs.io/en...

https://joss.theoj.org/papers/10.2110...

https://github.com/PyTorchLightning/m...

https://torchmetrics.readthedocs.io/e...

https://twitter.com/alanyttian/status...

https://github.com/google/evojax

https://arxiv.org/abs/2202.05008

https://www.reddit.com/r/MachineLearn...

https://www.gymlibrary.ml/pages/api/#...

Links:

TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://discord.gg/4H8xxDF

BitChute: https://www.bitchute.com/channel/yann...

LinkedIn: https://www.linkedin.com/in/ykilcher

BiliBili: https://space.bilibili.com/2017636191

If you want to support me, the best thing to do is to share out the content :)

If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https://www.subscribestar.com/yannick...

Patreon: https://www.patreon.com/yannickilcher

Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq

Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2

Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m

Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n

View Details

mlnews #kilcher #withtheauthors

Many of you have given me feedback on what you did and didn't like about the recent "with the authors" videos. Here's the result of that feedback and an outlook into the future.

Merch: store.ykilcher.com

Links:

TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://discord.gg/4H8xxDF

BitChute: https://www.bitchute.com/channel/yann...

LinkedIn: https://www.linkedin.com/in/ykilcher

BiliBili: https://space.bilibili.com/2017636191

If you want to support me, the best thing to do is to share out the content :)

If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https://www.subscribestar.com/yannick...

Patreon: https://www.patreon.com/yannickilcher

Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq

Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2

Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m

Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n

View Details

ai #gpu #tpu

This video is an interview with Adi Fuchs, author of a series called "AI Accelerators", and an expert in modern AI acceleration technology.

Accelerators like GPUs and TPUs are an integral part of today's AI landscape. Deep Neural Network training can be sped up by orders of magnitudes by making good use of these specialized pieces of hardware. However, GPUs and TPUs are only the beginning of a vast landscape of emerging technologies and companies that build accelerators for the next generation of AI models. In this interview, we go over many aspects of building hardware for AI, including why GPUs have been so successful, what the most promising approaches look like, how they work, and what the main challenges are.

OUTLINE:

0:00 - Intro

5:10 - What does it mean to make hardware for AI?

8:20 - Why were GPUs so successful?

16:25 - What is "dark silicon"?

20:00 - Beyond GPUs: How can we get even faster AI compute?

28:00 - A look at today's accelerator landscape

30:00 - Systolic Arrays and VLIW

35:30 - Reconfigurable dataflow hardware

40:50 - The failure of Wave Computing

42:30 - What is near-memory compute?

46:50 - Optical and Neuromorphic Computing

49:50 - Hardware as enabler and limiter

55:20 - Everything old is new again

1:00:00 - Where to go to dive deeper?

Read the full blog series here:

Part I: https://medium.com/@adi.fu7/ai-accele...

Part II: https://medium.com/@adi.fu7/ai-accele...

Part III: https://medium.com/@adi.fu7/ai-accele...

Part IV: https://medium.com/@adi.fu7/ai-accele...

Part V: https://medium.com/@adi.fu7/ai-accele...

Links:

Merch: store.ykilcher.com

TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://discord.gg/4H8xxDF

BitChute: https://www.bitchute.com/channel/yann...

LinkedIn: https://www.linkedin.com/in/ykilcher

BiliBili: https://space.bilibili.com/2017636191

If you want to support me, the best thing to do is to share out the content :)

If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https://www.subscribestar.com/yannick...

Patreon: https://www.patreon.com/yannickilcher

Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq

Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2

Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m

Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n

View Details

cm3 #languagemodel #transformer

This video contains a paper explanation and an incredibly informative interview with first author Armen Aghajanyan.

Autoregressive Transformers have come to dominate many fields in Machine Learning, from text generation to image creation and many more. However, there are two problems. First, the collected data is usually scraped from the web and uni- or bi-modal and throws away a lot of structure of the original websites, and second, language modelling losses are uni-directional. CM3 addresses both problems: It directly operates on HTML and includes text, hyperlinks, and even images (via VQGAN tokenization) and can therefore be used in plenty of ways: Text generation, captioning, image creation, entity linking, and much more. It also introduces a new training strategy called Causally Masked Language Modelling, which brings a level of bi-directionality into autoregressive language modelling. In the interview after the paper explanation, Armen and I go deep into the how and why of these giant models, we go over the stunning results and we make sense of what they mean for the future of universal models.

OUTLINE:

0:00 - Intro & Overview

6:30 - Directly learning the structure of HTML

12:30 - Causally Masked Language Modelling

18:50 - A short look at how to use this model

23:20 - Start of interview

25:30 - Feeding language models with HTML

29:45 - How to get bi-directionality into decoder-only Transformers?

37:00 - Images are just tokens

41:15 - How does one train such giant models?

45:40 - CM3 results are amazing

58:20 - Large-scale dataset collection and content filtering

1:04:40 - More experimental results

1:12:15 - Why don't we use raw HTML?

1:18:20 - Does this paper contain too many things?

Paper: https://arxiv.org/abs/2201.07520

Abstract:

We introduce CM3, a family of causally masked generative models trained over a large corpus of structured multi-modal documents that can contain both text and image tokens. Our new causally masked approach generates tokens left to right while also masking out a small number of long token spans that are generated at the end of the string, instead of their original positions. The casual masking object provides a type of hybrid of the more common causal and masked language models, by enabling full generative modeling while also providing bidirectional context when generating the masked spans. We train causally masked language-image models on large-scale web and Wikipedia articles, where each document contains all of the text, hypertext markup, hyperlinks, and image tokens (from a VQVAE-GAN), provided in the order they appear in the original HTML source (before masking). The resulting CM3 models can generate rich structured, multi-modal outputs while conditioning on arbitrary masked document contexts, and thereby implicitly learn a wide range of text, image, and cross modal tasks. They can be prompted to recover, in a zero-shot fashion, the functionality of models such as DALL-E, GENRE, and HTLM. We set the new state-of-the-art in zero-shot summarization, entity linking, and entity disambiguation while maintaining competitive performance in the fine-tuning setting. We can generate images unconditionally, conditioned on text (like DALL-E) and do captioning all in a zero-shot setting with a single model.

Authors: Armen Aghajanyan, Bernie Huang, Candace Ross, Vladimir Karpukhin, Hu Xu, Naman Goyal, Dmytro Okhonko, Mandar Joshi, Gargi Ghosh, Mike Lewis, Luke Zettlemoyer

View Details

security #censorship #ai

Most of us conceive the internet as a free and open space where we are able to send traffic between any two nodes, but for large parts of the world this is not the case. Entire nations have large machinery in place to survey all internet traffic and automated procedures to block any undesirable connections. Evading such censorship has been largely a cat-and-mouse game between security researchers and government actors. A new system, called Geneva, uses a Genetic Algorithm in combination with Evolutionary Search in order to dynamically evade such censorship and adjust itself in real-time to any potential response by its adversaries. In this video, I talk to Security researcher Kevin Bock, who is one of Geneva's main contributors and member of the Breakerspace project. We talk about the evolution of internet censorship, how to evade it, how to mess with the censors' infrastructure, as well as the broader emerging connections between AI and Security.

OUTLINE:

0:00 - Intro

3:30 - What is automated censorship in networks?

7:20 - The evolution of censorship vs evasion

12:40 - Why do we need a dynamic, evolving system?

16:30 - The building blocks of Geneva

23:15 - Introducing evolution

28:30 - What's the censors' response?

31:45 - How was Geneva's media reception?

33:15 - Where do we go from here?

37:30 - Can we deliberately attack the censors?

47:00 - On responsible disclosure

49:40 - Breakerspace: Security research for undergrads

50:40 - How often do you get into trouble?

52:10 - How can I get started in security?

Learn more at:

  • Geneva (& more) project page: https://censorship.ai

  • Open Observatory of Network Interference: https://ooni.org

  • Censored Planet: https://censoredplanet.org

  • Breakerspace: https://breakerspace.cs.umd.edu

Links:

Merch: store.ykilcher.com

TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://discord.gg/4H8xxDF

BitChute: https://www.bitchute.com/channel/yann...

LinkedIn: https://www.linkedin.com/in/ykilcher

BiliBili: https://space.bilibili.com/2017636191

If you want to support me, the best thing to do is to share out the content :)

If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https://www.subscribestar.com/yannick...

Patreon: https://www.patreon.com/yannickilcher

Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq

Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2

Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m

Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n

View Details

hypertransformer #metalearning #deeplearning

This video contains a paper explanation and an interview with author Andrey Zhmoginov!

Few-shot learning is an interesting sub-field in meta-learning, with wide applications, such as creating personalized models based on just a handful of data points. Traditionally, approaches have followed the BERT approach where a large model is pre-trained and then fine-tuned. However, this couples the size of the final model to the size of the model that has been pre-trained. Similar problems exist with "true" meta-learners, such as MaML. HyperTransformer fundamentally decouples the meta-learner from the size of the final model by directly predicting the weights of the final model. The HyperTransformer takes the few-shot dataset as a whole into its context and predicts either one or multiple layers of a (small) ConvNet, meaning its output are the weights of the convolution filters. Interestingly, and with the correct engineering care, this actually appears to deliver promising results and can be extended in many ways.

OUTLINE:

0:00 - Intro & Overview

3:05 - Weight-generation vs Fine-tuning for few-shot learning

10:10 - HyperTransformer model architecture overview

22:30 - Why the self-attention mechanism is useful here

34:45 - Start of Interview

39:45 - Can neural networks even produce weights of other networks?

47:00 - How complex does the computational graph get?

49:45 - Why are transformers particularly good here?

58:30 - What can the attention maps tell us about the algorithm?

1:07:00 - How could we produce larger weights?

1:09:30 - Diving into experimental results

1:14:30 - What questions remain open?

Paper: https://arxiv.org/abs/2201.04182

ERRATA: I introduce Max Vladymyrov as Mark Vladymyrov

Abstract:

In this work we propose a HyperTransformer, a transformer-based model for few-shot learning that generates weights of a convolutional neural network (CNN) directly from support samples. Since the dependence of a small generated CNN model on a specific task is encoded by a high-capacity transformer model, we effectively decouple the complexity of the large task space from the complexity of individual tasks. Our method is particularly effective for small target CNN architectures where learning a fixed universal task-independent embedding is not optimal and better performance is attained when the information about the task can modulate all model parameters. For larger models we discover that generating the last layer alone allows us to produce competitive or better results than those obtained with state-of-the-art methods while being end-to-end differentiable. Finally, we extend our approach to a semi-supervised regime utilizing unlabeled samples in the support set and further improving few-shot performance.

Authors: Andrey Zhmoginov, Mark Sandler, Max Vladymyrov

Links:

TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://discord.gg/4H8xxDF

BitChute: https://www.bitchute.com/channel/yann...

LinkedIn: https://www.linkedin.com/in/ykilcher

BiliBili: https://space.bilibili.com/2017636191

If you want to support me, the best thing to do is to share out the content :)

If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https://www.subscribestar.com/yannick...

Patreon: https://www.patreon.com/yannickilcher

Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq

View Details

mlnews #alphacode #openai

The latest and greatest from the world of Machine Learning!

Merch: store.ykilcher.com

Sponsor: Weights & Biases

https://wandb.me/yannic

OUTLINE:

0:00 - Intro

0:15 - Sponsor: Weights & Biases

3:15 - DeepMind's AlphaCode: AI competitive programmer

11:30 - OpenAI uses language models to prove math theorems

14:30 - StyleGAN XL: Scaling StyleGAN to diverse datasets

16:10 - ar5iv.org displays papers as HTML5

17:40 - Helpful Things

19:30 - ICML22 Review process changes

21:15 - Meta AI tackles harmful content classification using few-shot learning

23:55 - Company claims to produce face images from DNA

References:

https://deepmind.com/blog/article/Com...

https://alphacode.deepmind.com/#layer...

https://storage.googleapis.com/deepmi...

https://twitter.com/DBahdanau/status/...

https://openai.com/blog/formal-math/

https://arxiv.org/pdf/2202.01344.pdf

https://blog.eleuther.ai/announcing-2...

https://sites.google.com/view/stylega...

https://arxiv.org/pdf/2202.00273.pdf

https://ar5iv.org/

https://ar5iv.org/html/1910.06709

https://twitter.com/YiTayML/status/14...

https://ffcv.io/

https://github.com/ott-jax/ott

https://twitter.com/soumithchintala/s...

https://github.com/facebookresearch/d...

https://www.reddit.com/r/MachineLearn...

https://icml.cc/Conferences/2022/Revi...

https://icml.cc/Conferences/2022/Call...

https://ai.facebook.com/blog/harmful-...

https://www.technologyreview.com/2022...

Links:

TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://discord.gg/4H8xxDF

BitChute: https://www.bitchute.com/channel/yann...

LinkedIn: https://www.linkedin.com/in/ykilcher

BiliBili: https://space.bilibili.com/2017636191

If you want to support me, the best thing to do is to share out the content :)

If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https://www.subscribestar.com/yannick...

Patreon: https://www.patreon.com/yannickilcher

Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq

Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2

Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m

Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n

View Details

gpt3 #embodied #planning

In this video: Paper explanation, followed by first author interview with Wenlong Huang.

Large language models contain extraordinary amounts of world knowledge that can be queried in various ways. But their output format is largely uncontrollable. This paper investigates the VirtualHome environment, which expects a particular set of actions, objects, and verbs to be used. Turns out, with proper techniques and only using pre-trained models (no fine-tuning), one can translate unstructured language model outputs into the structured grammar of the environment. This is potentially very useful anywhere where the models' world knowledge needs to be provided in a particular structured format.

OUTLINE:

0:00 - Intro & Overview

2:45 - The VirtualHome environment

6:25 - The problem of plan evaluation

8:40 - Contributions of this paper

16:40 - Start of interview

24:00 - How to use language models with environments?

34:00 - What does model size matter?

40:00 - How to fix the large models' outputs?

55:00 - Possible improvements to the translation procedure

59:00 - Why does Codex perform so well?

1:02:15 - Diving into experimental results

1:14:15 - Future outlook

Paper: https://arxiv.org/abs/2201.07207

Website: https://wenlong.page/language-planner/

Code: https://github.com/huangwl18/language...

Wenlong's Twitter: https://twitter.com/wenlong_huang

Abstract:

Can world knowledge learned by large language models (LLMs) be used to act in interactive environments? In this paper, we investigate the possibility of grounding high-level tasks, expressed in natural language (e.g. "make breakfast"), to a chosen set of actionable steps (e.g. "open fridge"). While prior work focused on learning from explicit step-by-step examples of how to act, we surprisingly find that if pre-trained LMs are large enough and prompted appropriately, they can effectively decompose high-level tasks into low-level plans without any further training. However, the plans produced naively by LLMs often cannot map precisely to admissible actions. We propose a procedure that conditions on existing demonstrations and semantically translates the plans to admissible actions. Our evaluation in the recent VirtualHome environment shows that the resulting method substantially improves executability over the LLM baseline. The conducted human evaluation reveals a trade-off between executability and correctness but shows a promising sign towards extracting actionable knowledge from language models. Website at this https URL

Authors: Wenlong Huang, Pieter Abbeel, Deepak Pathak, Igor Mordatch

Links:

Merch: store.ykilcher.com

TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://discord.gg/4H8xxDF

BitChute: https://www.bitchute.com/channel/yann...

LinkedIn: https://www.linkedin.com/in/ykilcher

BiliBili: https://space.bilibili.com/2017636191

If you want to support me, the best thing to do is to share out the content :)

If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https://www.subscribestar.com/yannick...

Patreon: https://www.patreon.com/yannickilcher

Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq

Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2

Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m

Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n

View Details

mlnews #openai #embeddings

COMMENTS DIRECTLY FROM THE AUTHOR (thanks a lot for reaching out Arvind :) ):

  1. The FIQA results you share also have code to reproduce the results in the paper using the API: https://twitter.com/arvind_io/status/... There's no discrepancy AFAIK.

  2. We leave out 6 not 7 BEIR datasets. Results on msmarco, nq and triviaqa are in a separate table (Table 5 in the paper). NQ is part of BEIR too and we didn't want to repeat it. Finally, the 6 datasets we leave out are not readily available and it is common to leave them out in prior work too. For examples, see SPLADE v2 (https://arxiv.org/pdf/2109.10086.pdf) also evaluates on the same 12 BEIR datasets.

  3. Finally, I'm now working on time travel so that I can cite papers from the future :)

END COMMENTS FROM THE AUTHOR

OpenAI launches an embeddings endpoint in their API, providing high-dimensional vector embeddings for use in text similarity, text search, and code search. While embeddings are universally recognized as a standard tool to process natural language, people have raised doubts about the quality of OpenAI's embeddings, as one blog post found they are often outperformed by open-source models, which are much smaller and with which embedding would cost a fraction of what OpenAI charges. In this video, we examine the claims made and determine what it all means.

OUTLINE:

0:00 - Intro

0:30 - Sponsor: Weights & Biases

2:20 - What embeddings are available?

3:55 - OpenAI shows promising results

5:25 - How good are the results really?

6:55 - Criticism: Open models might be cheaper and smaller

10:05 - Discrepancies in the results

11:00 - The author's response

11:50 - Putting things into perspective

13:35 - What about real world data?

14:40 - OpenAI's pricing strategy: Why so expensive?

Sponsor: Weights & Biases

https://wandb.me/yannic

Merch: store.ykilcher.com

ERRATA: At 13:20 I say "better", it should be "worse"

References:

https://openai.com/blog/introducing-t...

https://arxiv.org/pdf/2201.10005.pdf

https://beta.openai.com/docs/guides/e...

https://beta.openai.com/docs/api-refe...

https://twitter.com/Nils_Reimers/stat...

https://medium.com/@nils_reimers/open...

https://mobile.twitter.com/arvind_io/...

https://twitter.com/gwern/status/1487...

https://twitter.com/gwern/status/1487...

https://twitter.com/Nils_Reimers/stat...

https://twitter.com/gwern/status/1470...

https://www.reddit.com/r/MachineLearn...

https://mobile.twitter.com/arvind_io/...

https://mobile.twitter.com/arvind_io/...

Links:

TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://discord.gg/4H8xxDF

BitChute: https://www.bitchute.com/channel/yann...

LinkedIn: https://www.linkedin.com/in/ykilcher

BiliBili: https://space.bilibili.com/2017636191

If you want to support me, the best thing to do is to share out the content :)

If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https://www.subscribestar.com/yannick...

Patreon: https://www.patreon.com/yannickilcher

Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq

Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2

Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m

View Details

deeplearning #brain #neuroscience

Originally, Deep Learning sprang into existence inspired by how the brain processes information, but the two fields have diverged ever since. However, given that deep models can solve many perception tasks with remarkable accuracy, is it possible that we might be able to learn something about how the brain works by inspecting our models? I speak to Patrick Mineault about his blog post "2021 in review: unsupervised brain models" and we explore why neuroscientists are taking interest in unsupervised and self-supervised deep neural networks in order to explain how the brain works. We discuss a series of influential papers that have appeared last year, and we go into the more general questions of connecting neuroscience and machine learning.

OUTLINE:

0:00 - Intro & Overview

6:35 - Start of Interview

10:30 - Visual processing in the brain

12:50 - How does deep learning inform neuroscience?

21:15 - Unsupervised training explains the ventral stream

30:50 - Predicting own motion parameters explains the dorsal stream

42:20 - Why are there two different visual streams?

49:45 - Concept cells and representation learning

56:20 - Challenging the manifold theory

1:08:30 - What are current questions in the field?

1:13:40 - Should the brain inform deep learning?

1:18:50 - Neuromatch Academy and other endeavours

Blog Post: https://xcorr.net/2021/12/31/2021-in-...

Patrick's Blog: https://xcorr.net/

Twitter: https://twitter.com/patrickmineault

Neuromatch Academy: https://academy.neuromatch.io/

Links:

TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://discord.gg/4H8xxDF

BitChute: https://www.bitchute.com/channel/yann...

LinkedIn: https://www.linkedin.com/in/ykilcher

BiliBili: https://space.bilibili.com/2017636191

If you want to support me, the best thing to do is to share out the content :)

If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https://www.subscribestar.com/yannick...

Patreon: https://www.patreon.com/yannickilcher

Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq

Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2

Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m

View Details

eleuther #gptneo #gptj

EleutherAI announces GPT-NeoX-20B, a 20 billion parameter open-source language model, inspired by GPT-3. Connor joins me to discuss the process of training, how the group got their hands on the necessary hardware, what the new model can do, and how anyone can try it out!

OUTLINE:

0:00 - Intro

1:00 - Start of interview

2:00 - How did you get all the hardware?

3:50 - What's the scale of this model?

6:00 - A look into the experimental results

11:15 - Why are there GPT-Neo, GPT-J, and GPT-NeoX?

14:15 - How difficult is training these big models?

17:00 - Try out the model on GooseAI

19:00 - Final thoughts

Read the announcement: https://blog.eleuther.ai/announcing-20b/

Try out the model: https://goose.ai/

Check out EleutherAI: https://www.eleuther.ai/

Read the code: https://github.com/EleutherAI/gpt-neox

Hardware sponsor: https://www.coreweave.com/

Links:

TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://discord.gg/4H8xxDF

BitChute: https://www.bitchute.com/channel/yann...

LinkedIn: https://www.linkedin.com/in/ykilcher

BiliBili: https://space.bilibili.com/2017636191

If you want to support me, the best thing to do is to share out the content :)

If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https://www.subscribestar.com/yannick...

Patreon: https://www.patreon.com/yannickilcher

Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq

Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2

Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m

Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n

View Details

deeplearning #symbolic #research

This video includes an interview with first author Stéphane d'Ascoli (https://sdascoli.github.io/).

Deep neural networks are typically excellent at numeric regression, but using them for symbolic computation has largely been ignored so far. This paper uses transformers to do symbolic regression on integer and floating point number sequences, which means that given the start of a sequence of numbers, the model has to not only predict the correct continuation, but also predict the data generating formula behind the sequence. Through clever encoding of the input space and a well constructed training data generation process, this paper's model can learn and represent many of the sequences in the OEIS, the online encyclopedia of integer sequences and it also features an interactive demo if you want to try it by yourself.

OUTLINE:

0:00 - Introduction

2:20 - Summary of the Paper

16:10 - Start of Interview

17:15 - Why this research direction?

20:45 - Overview of the method

30:10 - Embedding space of input tokens

33:00 - Data generation process

42:40 - Why are transformers useful here?

46:40 - Beyond number sequences, where is this useful?

48:45 - Success cases and failure cases

58:10 - Experimental Results

1:06:30 - How did you overcome difficulties?

1:09:25 - Interactive demo

Paper: https://arxiv.org/abs/2201.04600

Interactive demo: http://recur-env.eba-rm3fchmn.us-east...

Abstract:

Symbolic regression, i.e. predicting a function from the observation of its values, is well-known to be a challenging task. In this paper, we train Transformers to infer the function or recurrence relation underlying sequences of integers or floats, a typical task in human IQ tests which has hardly been tackled in the machine learning literature. We evaluate our integer model on a subset of OEIS sequences, and show that it outperforms built-in Mathematica functions for recurrence prediction. We also demonstrate that our float model is able to yield informative approximations of out-of-vocabulary functions and constants, e.g. bessel0(x)≈sin(x)+cos(x)πx√ and 1.644934≈π2/6. An interactive demonstration of our models is provided at this https URL.

Authors: Stéphane d'Ascoli, Pierre-Alexandre Kamienny, Guillaume Lample, François Charton

Links:

TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://discord.gg/4H8xxDF

BitChute: https://www.bitchute.com/channel/yann...

LinkedIn: https://www.linkedin.com/in/ykilcher

BiliBili: https://space.bilibili.com/2017636191

If you want to support me, the best thing to do is to share out the content :)

If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https://www.subscribestar.com/yannick...

Patreon: https://www.patreon.com/yannickilcher

Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq

Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2

Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m

View Details

LIMITED TIME MERCH DEAL: http://store.ykilcher.com

Links:

TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://discord.gg/4H8xxDF

BitChute: https://www.bitchute.com/channel/yann...

LinkedIn: https://www.linkedin.com/in/ykilcher

BiliBili: https://space.bilibili.com/2017636191

If you want to support me, the best thing to do is to share out the content :)

If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https://www.subscribestar.com/yannick...

Patreon: https://www.patreon.com/yannickilcher

Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq

Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2

Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m

Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n

View Details

mlnews #convnext #mt3

Your update on what's new in the Machine Learning world!

OUTLINE:

0:00 - Intro

0:15 - ConvNeXt: Return of the Convolutions

2:50 - Investigating Saliency Cropping Algorithms

9:40 - YourTTS: SOTA zero-shot Text-to-Speech

10:40 - MT3: Multi-Track Music Transcription

11:35 - China regulates addictive algorithms

13:00 - A collection of Deep Learning interview questions & solutions

13:35 - Helpful Things

16:05 - AlphaZero explained blog post

16:45 - Ru-DOLPH: HyperModal Text-to-Image-to-Text model

17:45 - Google AI 2021 Review

References:

ConvNeXt: Return of the Convolutions

https://arxiv.org/abs/2201.03545

https://github.com/facebookresearch/C...

https://twitter.com/giffmana/status/1...

https://twitter.com/wightmanr/status/...

https://twitter.com/tanmingxing/statu...

Investigating Saliency Cropping Algorithms

https://openaccess.thecvf.com/content...

https://vinayprabhu.github.io/Salienc...

https://vinayprabhu.medium.com/on-the...

https://vinayprabhu.github.io/Salienc...

YourTTS: SOTA zero-shot Text-to-Speech

https://github.com/coqui-ai/TTS?utm_s...

https://arxiv.org/abs/2112.02418?utm_...

https://coqui.ai/?utm_source=pocket_m...

https://coqui.ai/blog/tts/yourtts-zer...

MT3: Multi-Track Music Transcription

https://arxiv.org/abs/2111.03017

https://github.com/magenta/mt3

https://huggingface.co/spaces/akhaliq...

https://www.reddit.com/r/MachineLearn...

China regulates addictive algorithms

https://technode.com/2022/01/05/china...

https://qz.com/2109618/china-reveals-...

A collection of Deep Learning interview questions & solutions

https://arxiv.org/abs/2201.00650?utm_...

https://arxiv.org/pdf/2201.00650.pdf

Helpful Things

https://docs.deepchecks.com/en/stable...

https://github.com/deepchecks/deepchecks

https://docs.deepchecks.com/en/stable...

https://www.dagshub.com/

https://www.dagshub.com/docs/index.html

https://www.dagshub.com/blog/launchin...

https://bayesiancomputationbook.com/w...

https://mlcontests.com/

https://github.com/Yard1/ray-skorch

https://github.com/skorch-dev/skorch

https://www.rumbledb.org/?utm_source=...

https://github.com/DarshanDeshpande/j...

https://github.com/s3prl/s3prl

AlphaZero explained blog post

https://joshvarty.github.io/AlphaZero...

Ru-DOLPH: HyperModal Text-to-Image-to-Text model

https://github.com/sberbank-ai/ru-dolph

https://colab.research.google.com/dri...

Google AI 2021 Review

https://ai.googleblog.com/2022/01/goo...

Links:

TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://discord.gg/4H8xxDF

BitChute: https://www.bitchute.com/channel/yann...

LinkedIn: https://www.linkedin.com/in/ykilcher

BiliBili: https://space.bilibili.com/2017636191

If you want to support me, the best thing to do is to share out the content :)

If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https://www.subscribestar.com/yannick...

Patreon: https://www.patreon.com/yannickilcher

Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq

Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2

Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m

View Details

deeplearning #neuralinterpreter #ai

This video includes an interview with the paper's authors!

What if we treated deep networks like modular programs? Neural Interpreters divide computation into small modules and route data to them via a dynamic type inference system. The resulting model combines recurrent elements, weight sharing, attention, and more to tackle both abstract reasoning, as well as computer vision tasks.

OUTLINE:

0:00 - Intro & Overview

3:00 - Model Overview

7:00 - Interpreter weights and function code

9:40 - Routing data to functions via neural type inference

14:55 - ModLin layers

18:25 - Experiments

21:35 - Interview Start

24:50 - General Model Structure

30:10 - Function code and signature

40:30 - Explaining Modulated Layers

49:50 - A closer look at weight sharing

58:30 - Experimental Results

Paper: https://arxiv.org/abs/2110.06399

Guests:

Nasim Rahaman: https://twitter.com/nasim_rahaman

Francesco Locatello: https://twitter.com/FrancescoLocat8

Waleed Gondal: https://twitter.com/Wallii_gondal

Abstract:

Modern neural network architectures can leverage large amounts of data to generalize well within the training distribution. However, they are less capable of systematic generalization to data drawn from unseen but related distributions, a feat that is hypothesized to require compositional reasoning and reuse of knowledge. In this work, we present Neural Interpreters, an architecture that factorizes inference in a self-attention network as a system of modules, which we call \emph{functions}. Inputs to the model are routed through a sequence of functions in a way that is end-to-end learned. The proposed architecture can flexibly compose computation along width and depth, and lends itself well to capacity extension after training. To demonstrate the versatility of Neural Interpreters, we evaluate it in two distinct settings: image classification and visual abstract reasoning on Raven Progressive Matrices. In the former, we show that Neural Interpreters perform on par with the vision transformer using fewer parameters, while being transferrable to a new task in a sample efficient manner. In the latter, we find that Neural Interpreters are competitive with respect to the state-of-the-art in terms of systematic generalization

Authors: Nasim Rahaman, Muhammad Waleed Gondal, Shruti Joshi, Peter Gehler, Yoshua Bengio, Francesco Locatello, Bernhard Schölkopf

Links:

TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://discord.gg/4H8xxDF

BitChute: https://www.bitchute.com/channel/yann...

LinkedIn: https://www.linkedin.com/in/ykilcher

BiliBili: https://space.bilibili.com/2017636191

If you want to support me, the best thing to do is to share out the content :)

If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https://www.subscribestar.com/yannick...

Patreon: https://www.patreon.com/yannickilcher

Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq

Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2

Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m

Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n

View Details

deeplearning #noether #symmetries

This video includes an interview with first author Ferran Alet!

Encoding inductive biases has been a long established methods to provide deep networks with the ability to learn from less data. Especially useful are encodings of symmetry properties of the data, such as the convolution's translation invariance. But such symmetries are often hard to program explicitly, and can only be encoded exactly when done in a direct fashion. Noether Networks use Noether's theorem connecting symmetries to conserved quantities and are able to dynamically and approximately enforce symmetry properties upon deep neural networks.

OUTLINE:

0:00 - Intro & Overview

18:10 - Interview Start

21:20 - Symmetry priors vs conserved quantities

23:25 - Example: Pendulum

27:45 - Noether Network Model Overview

35:35 - Optimizing the Noether Loss

41:00 - Is the computation graph stable?

46:30 - Increasing the inference time computation

48:45 - Why dynamically modify the model?

55:30 - Experimental Results & Discussion

Paper: https://arxiv.org/abs/2112.03321

Website: https://dylandoblar.github.io/noether...

Code: https://github.com/dylandoblar/noethe...

Abstract:

Progress in machine learning (ML) stems from a combination of data availability, computational resources, and an appropriate encoding of inductive biases. Useful biases often exploit symmetries in the prediction problem, such as convolutional networks relying on translation equivariance. Automatically discovering these useful symmetries holds the potential to greatly improve the performance of ML systems, but still remains a challenge. In this work, we focus on sequential prediction problems and take inspiration from Noether's theorem to reduce the problem of finding inductive biases to meta-learning useful conserved quantities. We propose Noether Networks: a new type of architecture where a meta-learned conservation loss is optimized inside the prediction function. We show, theoretically and experimentally, that Noether Networks improve prediction quality, providing a general framework for discovering inductive biases in sequential problems.

Authors: Ferran Alet, Dylan Doblar, Allan Zhou, Joshua Tenenbaum, Kenji Kawaguchi, Chelsea Finn

Links:

TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://discord.gg/4H8xxDF

BitChute: https://www.bitchute.com/channel/yann...

LinkedIn: https://www.linkedin.com/in/ykilcher

BiliBili: https://space.bilibili.com/2017636191

If you want to support me, the best thing to do is to share out the content :)

If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https://www.subscribestar.com/yannick...

Patreon: https://www.patreon.com/yannickilcher

Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq

Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2

Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m

Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n

View Details

minerl #minecraft #deeplearning

The MineRL BASALT challenge has no reward functions or technical descriptions of what's to be achieved. Instead, the goal of each task is given as a short natural language string, and the agent is evaluated by a team of human judges who rate both how well the goal has been fulfilled, as well as how human-like the agent behaved. In this video, I interview KAIROS, the winning team of the 2021 challenge, and discuss how they used a combination of machine learning, efficient data collection, hand engineering, and a bit of knowledge about Minecraft to beat all other teams.

OUTLINE:

0:00 - Introduction

4:10 - Paper Overview

11:15 - Start of Interview

17:05 - First Approach

20:30 - State Machine

26:45 - Efficient Label Collection

30:00 - Navigation Policy

38:15 - Odometry Estimation

46:00 - Pain Points & Learnings

50:40 - Live Run Commentary

58:50 - What other tasks can be solved?

1:01:55 - What made the difference?

1:07:30 - Recommendations & Conclusion

1:11:10 - Full Runs: Waterfall

1:12:40 - Full Runs: Build House

1:17:45 - Full Runs: Animal Pen

1:20:50 - Full Runs: Find Cave

Paper: https://arxiv.org/abs/2112.03482

Code: https://github.com/viniciusguigo/kair...

Challenge Website: https://minerl.io/basalt/

Paper Title: Combining Learning from Human Feedback and Knowledge Engineering to Solve Hierarchical Tasks in Minecraft

Abstract:

Real-world tasks of interest are generally poorly defined by human-readable descriptions and have no pre-defined reward signals unless it is defined by a human designer. Conversely, data-driven algorithms are often designed to solve a specific, narrowly defined, task with performance metrics that drives the agent's learning. In this work, we present the solution that won first place and was awarded the most human-like agent in the 2021 NeurIPS Competition MineRL BASALT Challenge: Learning from Human Feedback in Minecraft, which challenged participants to use human data to solve four tasks defined only by a natural language description and no reward function. Our approach uses the available human demonstration data to train an imitation learning policy for navigation and additional human feedback to train an image classifier. These modules, together with an estimated odometry map, are then combined into a state-machine designed based on human knowledge of the tasks that breaks them down in a natural hierarchy and controls which macro behavior the learning agent should follow at any instant. We compare this hybrid intelligence approach to both end-to-end machine learning and pure engineered solutions, which are then judged by human evaluators. Codebase is available at this https URL.

Authors: Vinicius G. Goecks, Nicholas Waytowich, David Watkins, Bharat Prakash

Links:

TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://discord.gg/4H8xxDF

BitChute: https://www.bitchute.com/channel/yann...

LinkedIn: https://www.linkedin.com/in/ykilcher

BiliBili: https://space.bilibili.com/2017636191

If you want to support me, the best thing to do is to share out the content :)

If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https://www.subscribestar.com/yannick...

Patreon: https://www.patreon.com/yannickilcher

Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq

Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2

Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m

View Details

minerl #minecraft #deeplearning

The MineRL BASALT challenge has no reward functions or technical descriptions of what's to be achieved. Instead, the goal of each task is given as a short natural language string, and the agent is evaluated by a team of human judges who rate both how well the goal has been fulfilled, as well as how human-like the agent behaved. In this video, I interview KAIROS, the winning team of the 2021 challenge, and discuss how they used a combination of machine learning, efficient data collection, hand engineering, and a bit of knowledge about Minecraft to beat all other teams.

OUTLINE:

0:00 - Introduction

4:10 - Paper Overview

11:15 - Start of Interview

17:05 - First Approach

20:30 - State Machine

26:45 - Efficient Label Collection

30:00 - Navigation Policy

38:15 - Odometry Estimation

46:00 - Pain Points & Learnings

50:40 - Live Run Commentary

58:50 - What other tasks can be solved?

1:01:55 - What made the difference?

1:07:30 - Recommendations & Conclusion

1:11:10 - Full Runs: Waterfall

1:12:40 - Full Runs: Build House

1:17:45 - Full Runs: Animal Pen

1:20:50 - Full Runs: Find Cave

Paper: https://arxiv.org/abs/2112.03482

Code: https://github.com/viniciusguigo/kair...

Challenge Website: https://minerl.io/basalt/

Paper Title: Combining Learning from Human Feedback and Knowledge Engineering to Solve Hierarchical Tasks in Minecraft

Abstract:

Real-world tasks of interest are generally poorly defined by human-readable descriptions and have no pre-defined reward signals unless it is defined by a human designer. Conversely, data-driven algorithms are often designed to solve a specific, narrowly defined, task with performance metrics that drives the agent's learning. In this work, we present the solution that won first place and was awarded the most human-like agent in the 2021 NeurIPS Competition MineRL BASALT Challenge: Learning from Human Feedback in Minecraft, which challenged participants to use human data to solve four tasks defined only by a natural language description and no reward function. Our approach uses the available human demonstration data to train an imitation learning policy for navigation and additional human feedback to train an image classifier. These modules, together with an estimated odometry map, are then combined into a state-machine designed based on human knowledge of the tasks that breaks them down in a natural hierarchy and controls which macro behavior the learning agent should follow at any instant. We compare this hybrid intelligence approach to both end-to-end machine learning and pure engineered solutions, which are then judged by human evaluators. Codebase is available at this https URL.

Authors: Vinicius G. Goecks, Nicholas Waytowich, David Watkins, Bharat Prakash

Links:

TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://discord.gg/4H8xxDF

BitChute: https://www.bitchute.com/channel/yann...

LinkedIn: https://www.linkedin.com/in/ykilcher

BiliBili: https://space.bilibili.com/2017636191

If you want to support me, the best thing to do is to share out the content :)

If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https://www.subscribestar.com/yannick...

Patreon: https://www.patreon.com/yannickilcher

Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq

Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2

Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m

View Details

tesla #fsd #elon

Watch the original podcast: https://www.youtube.com/watch?v=DxREm...

An analysis of Elon's appearance on Lex Fridman. Very interesting conversation and a good overview of past, current, and future versions of Tesla's Autopilot system.

OUTLINE:

0:00 - Intro

0:40 - Tesla Autopilot: How hard is it?

9:05 - Building an accurate understanding of the world

16:25 - History of Tesla's neural network stack

26:00 - When is full self-driving ready?

29:55 - FSD 11: Less code, more neural networks

37:00 - Auto-labelling is essential

39:05 - Tesla Bot & Discussion

Links:

TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://discord.gg/4H8xxDF

BitChute: https://www.bitchute.com/channel/yann...

LinkedIn: https://www.linkedin.com/in/ykilcher

BiliBili: https://space.bilibili.com/2017636191

If you want to support me, the best thing to do is to share out the content :)

If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https://www.subscribestar.com/yannick...

Patreon: https://www.patreon.com/yannickilcher

Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq

Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2

Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m

Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n

View Details

playerofgames #deepmind #alphazero

Special Guest: First author Martin Schmid (https://twitter.com/Lifrordi)

Games have been used throughout research as testbeds for AI algorithms, such as reinforcement learning agents. However, different types of games usually require different solution approaches, such as AlphaZero for Go or Chess, and Counterfactual Regret Minimization (CFR) for Poker. Player of Games bridges this gap between perfect and imperfect information games and delivers a single algorithm that uses tree search over public information states, and is trained via self-play. The resulting algorithm can play Go, Chess, Poker, Scotland Yard, and many more games, as well as non-game environments.

OUTLINE:

0:00 - Introduction

2:50 - What games can Player of Games be trained on?

4:00 - Tree search algorithms (AlphaZero)

8:00 - What is different in imperfect information games?

15:40 - Counterfactual Value- and Policy-Networks

18:50 - The Player of Games search procedure

28:30 - How to train the network?

34:40 - Experimental Results

47:20 - Discussion & Outlook

Paper: https://arxiv.org/abs/2112.03178

Abstract:

Games have a long history of serving as a benchmark for progress in artificial intelligence. Recently, approaches using search and learning have shown strong performance across a set of perfect information games, and approaches using game-theoretic reasoning and learning have shown strong performance for specific imperfect information poker variants. We introduce Player of Games, a general-purpose algorithm that unifies previous approaches, combining guided search, self-play learning, and game-theoretic reasoning. Player of Games is the first algorithm to achieve strong empirical performance in large perfect and imperfect information games -- an important step towards truly general algorithms for arbitrary environments. We prove that Player of Games is sound, converging to perfect play as available computation time and approximation capacity increases. Player of Games reaches strong performance in chess and Go, beats the strongest openly available agent in heads-up no-limit Texas hold'em poker (Slumbot), and defeats the state-of-the-art agent in Scotland Yard, an imperfect information game that illustrates the value of guided search, learning, and game-theoretic reasoning.

Authors: Martin Schmid, Matej Moravcik, Neil Burch, Rudolf Kadlec, Josh Davidson, Kevin Waugh, Nolan Bard, Finbarr Timbers, Marc Lanctot, Zach Holland, Elnaz Davoodi, Alden Christianson, Michael Bowling

Links:

TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://discord.gg/4H8xxDF

BitChute: https://www.bitchute.com/channel/yann...

LinkedIn: https://www.linkedin.com/in/ykilcher

BiliBili: https://space.bilibili.com/2017636191

If you want to support me, the best thing to do is to share out the content :)

If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https://www.subscribestar.com/yannick...

Patreon: https://www.patreon.com/yannickilcher

Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq

Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2

Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m

Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n

View Details

glide #openai #diffusion

Diffusion models learn to iteratively reverse a noising process that is applied repeatedly during training. The result can be used for conditional generation as well as various other tasks such as inpainting. OpenAI's GLIDE builds on recent advances in diffusion models and combines text-conditional diffusion with classifier-free guidance and upsampling to achieve unprecedented quality in text-to-image samples.

Try it yourself: https://huggingface.co/spaces/valhall...

OUTLINE:

0:00 - Intro & Overview

6:10 - What is a Diffusion Model?

18:20 - Conditional Generation and Guided Diffusion

31:30 - Architecture Recap

34:05 - Training & Result metrics

36:55 - Failure cases & my own results

39:45 - Safety considerations

Paper: https://arxiv.org/abs/2112.10741

Code & Model: https://github.com/openai/glide-text2im

More diffusion papers:

https://arxiv.org/pdf/2006.11239.pdf

https://arxiv.org/pdf/2102.09672.pdf

Abstract:

Diffusion models have recently been shown to generate high-quality synthetic images, especially when paired with a guidance technique to trade off diversity for fidelity. We explore diffusion models for the problem of text-conditional image synthesis and compare two different guidance strategies: CLIP guidance and classifier-free guidance. We find that the latter is preferred by human evaluators for both photorealism and caption similarity, and often produces photorealistic samples. Samples from a 3.5 billion parameter text-conditional diffusion model using classifier-free guidance are favored by human evaluators to those from DALL-E, even when the latter uses expensive CLIP reranking. Additionally, we find that our models can be fine-tuned to perform image inpainting, enabling powerful text-driven image editing. We train a smaller model on a filtered dataset and release the code and weights at this https URL.

Authors: Alex Nichol, Prafulla Dhariwal, Aditya Ramesh, Pranav Shyam, Pamela Mishkin, Bob McGrew, Ilya Sutskever, Mark Chen

Links:

TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://discord.gg/4H8xxDF

BitChute: https://www.bitchute.com/channel/yann...

LinkedIn: https://www.linkedin.com/in/ykilcher

BiliBili: https://space.bilibili.com/2017636191

If you want to support me, the best thing to do is to share out the content :)

If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https://www.subscribestar.com/yannick...

Patreon: https://www.patreon.com/yannickilcher

Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq

Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2

Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m

Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n

View Details

mlnews #gopher #glam

Your updates on everything going on in the Machine Learning world.

Sponsor: Weights & Biases

https://wandb.me/yannic

OUTLINE:

0:00 - Intro & Overview

0:20 - Sponsor: Weights & Biases

3:05 - DeepMind releases 3 papers on large language models

11:45 - Hugging Face Blog: Training CodeParrot from scratch

14:25 - Paper: Pre-Training vision systems with noise

15:45 - DeepMind advances Quantum Mechanics

16:45 - GoogleAI trains GLaM: 1 Trillion Parameters Mixture of Experts Model

18:45 - Colin Raffel calls for building ML models like we build Open-Source software

22:05 - A rebuke of the hype around DeepMind's math paper

24:45 - Helpful Things

32:25 - Suicide Capsule plans AI to assess your mental state before use

35:15 - Synthesia raises 50M to develop AI avatars

Weights & Biases Embedding Projector

https://twitter.com/_ScottCondron/sta...

https://docs.wandb.ai/ref/app/feature...

https://wandb.ai/timssweeney/toy_data...

DeepMind releases 3 papers on large language models

https://deepmind.com/blog/article/lan...

https://arxiv.org/pdf/2112.04426.pdf

https://kstatic.googleusercontent.com...

https://arxiv.org/pdf/2112.04359.pdf

https://deepmind.com/research/publica...

Hugging Face Blog: Training CodeParrot from scratch

https://huggingface.co/blog/codeparro...

Paper: Pre-Training vision systems with noise

https://mbaradad.github.io/learning_w...

DeepMind advances Quantum Mechanics

https://deepmind.com/blog/article/Sim...

https://storage.googleapis.com/deepmi...

https://github.com/deepmind/deepmind-...

GoogleAI trains GLaM: 1 Trillion Parameters Mixture of Experts Model

https://ai.googleblog.com/2021/12/mor...

Colin Raffel calls for building ML models like we build Open-Source software

https://colinraffel.com/blog/a-call-t...

A rebuke of the hype around DeepMind's math paper

https://arxiv.org/abs/2112.04324?s=09

Helpful Things

https://twitter.com/huggingface/statu...

https://docs.cohere.ai/prompt-enginee...

https://github.blog/2021-12-08-improv...

https://huggingface.co/blog/data-meas...

https://huggingface.co/spaces/hugging...

https://blogs.microsoft.com/ai-for-bu...

https://techcommunity.microsoft.com/t...

https://github.com/minitorch/minitorc...

https://minitorch.github.io/

https://pandastutor.com/

https://pandastutor.com/vis.html

https://github.com/IAmPara0x/yuno

https://colab.research.google.com/dri...

https://www.reddit.com/r/MachineLearn...

https://www.drivendata.org/competitio...

https://www.reddit.com/r/MachineLearn...

https://www.uttt.ai/

https://arxiv.org/abs/2112.02721?utm_...

https://arxiv.org/pdf/2112.02721.pdf

https://github.com/GEM-benchmark/NL-A...

https://www.reddit.com/r/MachineLearn...

Suicide Capsule plans AI to assess your mental state before use

https://www.swissinfo.ch/eng/sci-tech...

Synthesia raises 50M to develop AI avatars

https://techcrunch.com/2021/12/08/syn...

https://www.synthesia.io/

Links:

TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://discord.gg/4H8xxDF

SubscribeStar: https://www.subscribestar.com/yannick...

Patreon: https://www.patreon.com/yannickilcher

Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq

View Details

lama #inpainting #deeplearning

At the end of the video is an interview with the paper authors!

LaMa is a system that is amazing at removing foreground objects from images, especially when those objects cover a large part of the image itself. LaMa is specifically trained to reconstruct large masked areas and includes global information throughout its forward propagation by using Fourier Convolutions in its layers. This makes it incredibly effective at reconstructing periodic structures with long-range consistency, compared to regular convolutions.

OUTLINE:

0:00 - Intro

0:45 - Sponsor: ClearML

3:30 - Inpainting Examples

5:05 - Live Demo

6:40 - Locality as a weakness of convolutions

10:30 - Using Fourier Transforms for global information

12:55 - Model architecture overview

14:35 - Fourier convolution layer

21:15 - Loss function

24:25 - Mask generation algorithm

25:40 - Experimental results

28:25 - Interview with the authors

Paper: https://arxiv.org/abs/2109.07161

Code: https://github.com/saic-mdal/lama

Online Demo: https://cleanup.pictures/

Sponsor: ClearML

https://clear.ml

Abstract:

Modern image inpainting systems, despite the significant progress, often struggle with large missing areas, complex geometric structures, and high-resolution images. We find that one of the main reasons for that is the lack of an effective receptive field in both the inpainting network and the loss function. To alleviate this issue, we propose a new method called large mask inpainting (LaMa). LaMa is based on i) a new inpainting network architecture that uses fast Fourier convolutions (FFCs), which have the image-wide receptive field; ii) a high receptive field perceptual loss; iii) large training masks, which unlocks the potential of the first two components. Our inpainting network improves the state-of-the-art across a range of datasets and achieves excellent performance even in challenging scenarios, e.g. completion of periodic structures. Our model generalizes surprisingly well to resolutions that are higher than those seen at train time, and achieves this at lower parameter&time costs than the competitive baselines. The code is available at \url{this https URL}.

Authors: Roman Suvorov, Elizaveta Logacheva, Anton Mashikhin, Anastasia Remizova, Arsenii Ashukha, Aleksei Silvestrov, Naejin Kong, Harshith Goka, Kiwoong Park, Victor Lempitsky

Links:

TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://discord.gg/4H8xxDF

BitChute: https://www.bitchute.com/channel/yann...

LinkedIn: https://www.linkedin.com/in/ykilcher

BiliBili: https://space.bilibili.com/2017636191

If you want to support me, the best thing to do is to share out the content :)

If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https://www.subscribestar.com/yannick...

Patreon: https://www.patreon.com/yannickilcher

Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq

Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2

Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m

Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n

View Details

mlnews #deepmind #ai

The most trusted model in News!

Get started with Weights & Biases here: https://wandb.me/yannic

(it's free forever for personal use)

OUTLINE:

0:00 - Intro

0:15 - Sponsor: Weights & Biases

3:10 - DeepMind tackles fundamental math

6:45 - Microsoft focuses on scaling effectively and efficiently

10:15 - NeurIPS Anthology Visualization

13:30 - Timnit Gebru launches research institute independent from big tech

16:50 - SageMaker Canvas for no-code ML

17:50 - Help, Help!

21:40 - Cornelius Emde wins the 3090

21:55 - A retrospective on the NeurIPS 2021 ethics review process

References:

DeepMind tackles fundamental math

https://deepmind.com/blog/article/exp...

https://www.nature.com/articles/s4158...

Microsoft focuses on scaling effectively and efficiently

https://www.microsoft.com/en-us/resea...

NeurIPS Anthology Visualization

https://neuripsav.vizhub.ai/blog/

https://neuripsav.vizhub.ai/

Timnit Gebru launches research institute independent from big tech

https://www.washingtonpost.com/techno...

https://www.dair-institute.org/about

https://www.theguardian.com/commentis...

SageMaker Canvas for no-code ML

https://aws.amazon.com/blogs/aws/anno...

Help, Help!

https://macberth.netlify.app/

https://huggingface.co/emanjavacas/Ma...

https://developer.nvidia.com/blog/nvi...

https://opacus.ai/

https://twitter.com/naotokui_en/statu...

https://colab.research.google.com/dri...

https://twitter.com/ThomasSimonini/st...

https://github.com/karpathy/arxiv-san...

https://arxiv-sanity-lite.com/

https://www.youtube.com/watch?v=01ENz...

https://github.com/Felix-Petersen/alg...

https://github.com/rentruewang/koila?...

https://github.com/YeWR/EfficientZero

Cornelius Emde wins the 3090

https://twitter.com/CorEmde/status/14...

A retrospective on the NeurIPS 2021 ethics review process

https://blog.neurips.cc/2021/12/03/a-...

Links:

TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://discord.gg/4H8xxDF

BitChute: https://www.bitchute.com/channel/yann...

LinkedIn: https://www.linkedin.com/in/ykilcher

BiliBili: https://space.bilibili.com/2017636191

If you want to support me, the best thing to do is to share out the content :)

If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https://www.subscribestar.com/yannick...

Patreon: https://www.patreon.com/yannickilcher

Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq

Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2

Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m

Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n

View Details

nuwa #microsoft #generative

NÜWA is a unifying architecture that can ingest text, images, and videos and brings all of them into a quantized latent representation to support a multitude of visual generation tasks, such as text-to-image, text-guided video manipulation, or sketch-to-video. This paper details how the encoders for the different modalities are constructed, and how the latent representation is transformed using their novel 3D nearby self-attention layers. Experiments are shown on 8 different visual generation tasks that the model supports.

OUTLINE:

0:00 - Intro & Outline

1:20 - Sponsor: ClearML

3:35 - Tasks & Naming

5:10 - The problem with recurrent image generation

7:35 - Creating a shared latent space w/ Vector Quantization

23:20 - Transforming the latent representation

26:25 - Recap: Self- and Cross-Attention

28:50 - 3D Nearby Self-Attention

41:20 - Pre-Training Objective

46:05 - Experimental Results

50:40 - Conclusion & Comments

Paper: https://arxiv.org/abs/2111.12417

Github: https://github.com/microsoft/NUWA

Sponsor: ClearML

https://clear.ml

Abstract:

This paper presents a unified multimodal pre-trained model called NÜWA that can generate new or manipulate existing visual data (i.e., images and videos) for various visual synthesis tasks. To cover language, image, and video at the same time for different scenarios, a 3D transformer encoder-decoder framework is designed, which can not only deal with videos as 3D data but also adapt to texts and images as 1D and 2D data, respectively. A 3D Nearby Attention (3DNA) mechanism is also proposed to consider the nature of the visual data and reduce the computational complexity. We evaluate NÜWA on 8 downstream tasks. Compared to several strong baselines, NÜWA achieves state-of-the-art results on text-to-image generation, text-to-video generation, video prediction, etc. Furthermore, it also shows surprisingly good zero-shot capabilities on text-guided image and video manipulation tasks. Project repo is this https URL.

Authors: Chenfei Wu, Jian Liang, Lei Ji, Fan Yang, Yuejian Fang, Daxin Jiang, Nan Duan

Links:

TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://discord.gg/4H8xxDF

BitChute: https://www.bitchute.com/channel/yann...

LinkedIn: https://www.linkedin.com/in/ykilcher

BiliBili: https://space.bilibili.com/2017636191

If you want to support me, the best thing to do is to share out the content :)

If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https://www.subscribestar.com/yannick...

Patreon: https://www.patreon.com/yannickilcher

Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq

Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2

Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m

Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n

View Details

mlnews #gaugan #gpt-3

Your weekly dose of ML News!

More GauGAN images here: https://drive.google.com/drive/folder...

OUTLINE:

0:00 - Intro

0:20 - Sponsor: Weights & Biases

2:20 - OpenAI's removes GPT-3 Waitlist

4:55 - NVIDIA releases GauGAN2 Webapp

9:45 - Everyday Robots tackles real-life tasks

12:15 - MetNet-2: 12-hour Rain Forecasting

14:45 - TinyML Dog Bark Stopper

15:55 - AI learns to drive Mario Kart 64 on real hardware

17:40 - NYC regulates bias in AI hiring tools

21:05 - Beverage companies big into AI

21:50 - How does AlphaZero play Chess?

23:35 - Helpful Things

28:00 - ArXiv founder awarded Einstein Foundation Award

References:

OpenAI's removes GPT-3 Waitlist

https://openai.com/blog/api-no-waitlist/

https://beta.openai.com/playground?mo...

NVIDIA releases GauGAN2 Webapp

https://www.reddit.com/r/MachineLearn...

http://gaugan.org/gaugan2/

https://blogs.nvidia.com/blog/2021/11...

https://blogs.nvidia.com/blog/2019/03...

https://arxiv.org/abs/1903.07291

Everyday Robots tackles real-life tasks

https://everydayrobots.com/

https://www.wired.com/story/plaintext...

https://archive.ph/YC4XG#selection-92...

MetNet-2: 12-hour Rain Forecasting

https://ai.googleblog.com/2021/11/met...

TinyML Dog Bark Stopper

https://www.hackster.io/NathanielF/ti...

AI learns to drive Mario Kart 64 on real hardwware

https://www.youtube.com/watch?v=z9E38...

NYC regulates bias in AI hiring tools

https://www.nbcnewyork.com/news/local...

Beverage companies big into AI

https://www.just-drinks.com/features/...

How does AlphaZero play Chess?

https://arxiv.org/pdf/2111.09259.pdf

https://storage.googleapis.com/uncert...

Helpful Things

https://huggingface.co/sberbank-ai/ru...

https://github.com/MathisFederico/Ope...

https://blog.tensorflow.org/2021/11/i...

https://github.com/tensorflow/gnn

https://github.com/jurgisp/pydreamer?...

https://danijar.com/project/dreamerv2/

https://github.com/danijar/dreamerv2

https://deepgenx.com/

https://github.com/DeepGenX/CodeGenX

https://devpost.com/software/heyoh-ca...

https://heyoh-app.github.io/heyoh-pro...

https://github.com/heyoh-app/heyoh-pr...

ArXiv founder awarded Einstein Foundation Award

https://idw-online.de/en/news781515?u...

Links:

TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://discord.gg/4H8xxDF

BitChute: https://www.bitchute.com/channel/yann...

LinkedIn: https://www.linkedin.com/in/ykilcher

BiliBili: https://space.bilibili.com/2017636191

If you want to support me, the best thing to do is to share out the content :)

If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https://www.subscribestar.com/yannick...

Patreon: https://www.patreon.com/yannickilcher

Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq

View Details

scalingtransformers #terraformer #sparsity

Transformers keep pushing the state of the art in language and other domains, mainly due to their ability to scale to ever more parameters. However, this scaling has made it prohibitively expensive to run a lot of inference requests against a Transformer, both in terms of compute and memory requirements. Scaling Transformers are a new kind of architecture that leverage sparsity in the Transformer blocks to massively speed up inference, and by including additional ideas from other architectures, they create the Terraformer, which is both fast, accurate, and consumes very little memory.

OUTLINE:

0:00 - Intro & Overview

4:10 - Recap: Transformer stack

6:55 - Sparse Feedforward layer

19:20 - Sparse QKV Layer

43:55 - Terraformer architecture

55:05 - Experimental Results & Conclusion

Paper: https://arxiv.org/abs/2111.12763

Code: https://github.com/google/trax/blob/m...

Abstract:

Large Transformer models yield impressive results on many tasks, but are expensive to train, or even fine-tune, and so slow at decoding that their use and study becomes out of reach. We address this problem by leveraging sparsity. We study sparse variants for all layers in the Transformer and propose Scaling Transformers, a family of next generation Transformer models that use sparse layers to scale efficiently and perform unbatched decoding much faster than the standard Transformer as we scale up the model size. Surprisingly, the sparse layers are enough to obtain the same perplexity as the standard Transformer with the same number of parameters. We also integrate with prior sparsity approaches to attention and enable fast inference on long sequences even with limited memory. This results in performance competitive to the state-of-the-art on long text summarization.

Authors: Sebastian Jaszczur, Aakanksha Chowdhery, Afroz Mohiuddin, Łukasz Kaiser, Wojciech Gajewski, Henryk Michalewski, Jonni Kanerva

Links:

TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://discord.gg/4H8xxDF

BitChute: https://www.bitchute.com/channel/yann...

LinkedIn: https://www.linkedin.com/in/ykilcher

BiliBili: https://space.bilibili.com/2017636191

If you want to support me, the best thing to do is to share out the content :)

If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https://www.subscribestar.com/yannick...

Patreon: https://www.patreon.com/yannickilcher

Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq

Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2

Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m

Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n

View Details

ext5 #transferlearning #exmix

The T5 model has been a staple for NLP research for the last years. Both its size and its approach to formulate all NLP tasks as prompt-based language modeling make it a convenient choice to tackle new challenges and provides a strong baseline for most current datasets. ExT5 pushes T5 to its limits by pre-training not only on self-supervised mask filling, but also at the same time on 107 different supervised NLP tasks, which is their new ExMix dataset. The resulting model compares very favorably to T5 when fine-tuned to downstream tasks.

OUTLINE:

0:00 - Intro & Overview

2:15 - Recap: The T5 model

3:55 - The ExT5 model and task formulations

8:10 - ExMix dataset

9:35 - Do different tasks help each other?

16:50 - Which tasks should we include?

20:30 - Pre-Training vs Pre-Finetuning

23:00 - A few hypotheses about what's going on

27:20 - How much self-supervised data to use?

34:15 - More experimental results

38:40 - Conclusion & Summary

Paper: https://arxiv.org/abs/2111.10952

Abstract:

Despite the recent success of multi-task learning and transfer learning for natural language processing (NLP), few works have systematically studied the effect of scaling up the number of tasks during pre-training. Towards this goal, this paper introduces ExMix (Extreme Mixture): a massive collection of 107 supervised NLP tasks across diverse domains and task-families. Using ExMix, we study the effect of multi-task pre-training at the largest scale to date, and analyze co-training transfer amongst common families of tasks. Through this analysis, we show that manually curating an ideal set of tasks for multi-task pre-training is not straightforward, and that multi-task scaling can vastly improve models on its own. Finally, we propose ExT5: a model pre-trained using a multi-task objective of self-supervised span denoising and supervised ExMix. Via extensive experiments, we show that ExT5 outperforms strong T5 baselines on SuperGLUE, GEM, Rainbow, Closed-Book QA tasks, and several tasks outside of ExMix. ExT5 also significantly improves sample efficiency while pre-training.

Authors: Vamsi Aribandi, Yi Tay, Tal Schuster, Jinfeng Rao, Huaixiu Steven Zheng, Sanket Vaibhav Mehta, Honglei Zhuang, Vinh Q. Tran, Dara Bahri, Jianmo Ni, Jai Gupta, Kai Hui, Sebastian Ruder, Donald Metzler

Links:

TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://discord.gg/4H8xxDF

BitChute: https://www.bitchute.com/channel/yann...

LinkedIn: https://www.linkedin.com/in/ykilcher

BiliBili: https://space.bilibili.com/2017636191

If you want to support me, the best thing to do is to share out the content :)

If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https://www.subscribestar.com/yannick...

Patreon: https://www.patreon.com/yannickilcher

Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq

Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2

Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m

View Details

imle #backpropagation #discrete

Backpropagation is the workhorse of deep learning, but unfortunately, it only works for continuous functions that are amenable to the chain rule of differentiation. Since discrete algorithms have no continuous derivative, deep networks with such algorithms as part of them cannot be effectively trained using backpropagation. This paper presents a method to incorporate a large class of algorithms, formulated as discrete exponential family distributions, into deep networks and derives gradient estimates that can easily be used in end-to-end backpropagation. This enables things like combinatorial optimizers to be part of a network's forward propagation natively.

OUTLINE:

0:00 - Intro & Overview

4:25 - Sponsor: Weights & Biases

6:15 - Problem Setup & Contributions

8:50 - Recap: Straight-Through Estimator

13:25 - Encoding the discrete problem as an inner product

19:45 - From algorithm to distribution

23:15 - Substituting the gradient

26:50 - Defining a target distribution

38:30 - Approximating marginals via perturb-and-MAP

45:10 - Entire algorithm recap

56:45 - Github Page & Example

Paper: https://arxiv.org/abs/2106.01798

Code (TF): https://github.com/nec-research/tf-imle

Code (Torch): https://github.com/uclnlp/torch-imle

Our Discord: https://discord.gg/4H8xxDF

Sponsor: Weights & Biases

https://wandb.com

Abstract:

Combining discrete probability distributions and combinatorial optimization problems with neural network components has numerous applications but poses several challenges. We propose Implicit Maximum Likelihood Estimation (I-MLE), a framework for end-to-end learning of models combining discrete exponential family distributions and differentiable neural components. I-MLE is widely applicable as it only requires the ability to compute the most probable states and does not rely on smooth relaxations. The framework encompasses several approaches such as perturbation-based implicit differentiation and recent methods to differentiate through black-box combinatorial solvers. We introduce a novel class of noise distributions for approximating marginals via perturb-and-MAP. Moreover, we show that I-MLE simplifies to maximum likelihood estimation when used in some recently studied learning settings that involve combinatorial solvers. Experiments on several datasets suggest that I-MLE is competitive with and often outperforms existing approaches which rely on problem-specific relaxations.

Authors: Mathias Niepert, Pasquale Minervini, Luca Franceschi

Links:

TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://discord.gg/4H8xxDF

BitChute: https://www.bitchute.com/channel/yann...

LinkedIn: https://www.linkedin.com/in/ykilcher

BiliBili: https://space.bilibili.com/2017636191

If you want to support me, the best thing to do is to share out the content :)

If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https://www.subscribestar.com/yannick...

Patreon: https://www.patreon.com/yannickilcher

Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq

Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2

Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m

Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n

View Details

neurips #peerreview #machinelearning

A look at the results of the 2021 NeurIPS peer review experiment.

https://arxiv.org/abs/2109.09774

https://www.reddit.com/r/MachineLearn...

Links:

TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://discord.gg/4H8xxDF

BitChute: https://www.bitchute.com/channel/yann...

LinkedIn: https://www.linkedin.com/in/ykilcher

BiliBili: https://space.bilibili.com/2017636191

If you want to support me, the best thing to do is to share out the content :)

If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https://www.subscribestar.com/yannick...

Patreon: https://www.patreon.com/yannickilcher

Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq

Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2

Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m

Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n

View Details

deeplearning #neuralarchitecturesearch #metalearning

Deep Neural Networks are usually trained from a given parameter initialization using SGD until convergence at a local optimum. This paper goes a different route: Given a novel network architecture for a known dataset, can we predict the final network parameters without ever training them? The authors build a Graph-Hypernetwork and train on a novel dataset of various DNN-architectures to predict high-performing weights. The results show that not only can the GHN predict weights with non-trivial performance, but it can also generalize beyond the distribution of training architectures to predict weights for networks that are much larger, deeper, or wider than ever seen in training.

OUTLINE:

0:00 - Intro & Overview

6:20 - DeepNets-1M Dataset

13:25 - How to train the Hypernetwork

17:30 - Recap on Graph Neural Networks

23:40 - Message Passing mirrors forward and backward propagation

25:20 - How to deal with different output shapes

28:45 - Differentiable Normalization

30:20 - Virtual Residual Edges

34:40 - Meta-Batching

37:00 - Experimental Results

42:00 - Fine-Tuning experiments

45:25 - Public reception of the paper

ERRATA:

  • Boris' name is obviously Boris, not Bori

  • At 36:05, Boris mentions that they train the first variant, yet on closer examination, we decided it's more like the second

Paper: https://arxiv.org/abs/2110.13100

Code: https://github.com/facebookresearch/p...

Abstract:

Deep learning has been successful in automating the design of features in machine learning pipelines. However, the algorithms optimizing neural network parameters remain largely hand-designed and computationally inefficient. We study if we can use deep learning to directly predict these parameters by exploiting the past knowledge of training other networks. We introduce a large-scale dataset of diverse computational graphs of neural architectures - DeepNets-1M - and use it to explore parameter prediction on CIFAR-10 and ImageNet. By leveraging advances in graph neural networks, we propose a hypernetwork that can predict performant parameters in a single forward pass taking a fraction of a second, even on a CPU. The proposed model achieves surprisingly good performance on unseen and diverse networks. For example, it is able to predict all 24 million parameters of a ResNet-50 achieving a 60% accuracy on CIFAR-10. On ImageNet, top-5 accuracy of some of our networks approaches 50%. Our task along with the model and results can potentially lead to a new, more computationally efficient paradigm of training networks. Our model also learns a strong representation of neural architectures enabling their analysis.

Authors: Boris Knyazev, Michal Drozdzal, Graham W. Taylor, Adriana Romero-Soriano

Links:

TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://discord.gg/4H8xxDF

BitChute: https://www.bitchute.com/channel/yann...

LinkedIn: https://www.linkedin.com/in/ykilcher

BiliBili: https://space.bilibili.com/2017636191

If you want to support me, the best thing to do is to share out the content :)

If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https://www.subscribestar.com/yannick...

Patreon: https://www.patreon.com/yannickilcher

Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq

Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2

Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m

View Details

grafting #adam #sgd

The last years in deep learning research have given rise to a plethora of different optimization algorithms, such as SGD, AdaGrad, Adam, LARS, LAMB, etc. which all claim to have their special peculiarities and advantages. In general, all algorithms modify two major things: The (implicit) learning rate schedule, and a correction to the gradient direction. This paper introduces grafting, which allows to transfer the induced learning rate schedule of one optimizer to another one. In that, the paper shows that much of the benefits of adaptive methods (e.g. Adam) are actually due to this schedule, and not necessarily to the gradient direction correction. Grafting allows for more fundamental research into differences and commonalities between optimizers, and a derived version of it makes it possible to computes static learning rate corrections for SGD, which potentially allows for large savings of GPU memory.

OUTLINE

0:00 - Rant about Reviewer #2

6:25 - Intro & Overview

12:25 - Adaptive Optimization Methods

20:15 - Grafting Algorithm

26:45 - Experimental Results

31:35 - Static Transfer of Learning Rate Ratios

35:25 - Conclusion & Discussion

Paper (OpenReview): https://openreview.net/forum?id=FpKgG...

Old Paper (Arxiv): https://arxiv.org/abs/2002.11803

Our Discord: https://discord.gg/4H8xxDF

Abstract:

In the empirical science of training large neural networks, the learning rate schedule is a notoriously challenging-to-tune hyperparameter, which can depend on all other properties (architecture, optimizer, batch size, dataset, regularization, ...) of the problem. In this work, we probe the entanglements between the optimizer and the learning rate schedule. We propose the technique of optimizer grafting, which allows for the transfer of the overall implicit step size schedule from a tuned optimizer to a new optimizer, preserving empirical performance. This provides a robust plug-and-play baseline for optimizer comparisons, leading to reductions to the computational cost of optimizer hyperparameter search. Using grafting, we discover a non-adaptive learning rate correction to SGD which allows it to train a BERT model to state-of-the-art performance. Besides providing a resource-saving tool for practitioners, the invariances discovered via grafting shed light on the successes and failure modes of optimizers in deep learning.

Authors: Anonymous (Under Review)

Links:

TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://discord.gg/4H8xxDF

BitChute: https://www.bitchute.com/channel/yann...

LinkedIn: https://www.linkedin.com/in/ykilcher

BiliBili: https://space.bilibili.com/2017636191

If you want to support me, the best thing to do is to share out the content :)

If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https://www.subscribestar.com/yannick...

Patreon: https://www.patreon.com/yannickilcher

Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq

Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2

Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m

Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n

View Details

mlnews #cedille #wmt

Only the greatest of news from the world of Machine Learning.

OUTLINE:

0:00 - Sponsor: Weights & Biases

1:50 - Cedille - French Language Model

3:55 - Facebook AI Multilingual model wins WMT

5:50 - YOU private search engine

10:35 - DeepMind's Open-Source Arnheim

12:10 - Company sued for using AI to make website more accessible

18:05 - Alibaba DAMO Academy creates 10 Trillion M6 model

21:15 - AMD MI200 Family

22:30 - State of AI report 2021

24:15 - Andrew Ng's Landing AI raises 57M

25:40 - Cerebras raises 250M

26:45 - Microsoft's Varuna: Scalable Training of Huge Models

28:15 - Laura Ruis reproduces Extrapolation Paper

29:05 - Ian Charnas' Real-Life Punchout

30:00 - Helpful Things

33:10 - AI finds profitable Meme-Tokens

34:55 - This Sneaker Does Not Exist

Sponsor: Weights & Biases

https://wandb.com

References:

Cedille - French Language Model

https://en.cedille.ai/

https://github.com/coteries/cedille-ai

https://app.cedille.ai/

https://en.wikipedia.org/wiki/Cedilla

Facebook AI Multilingual model wins WMT

https://ai.facebook.com/blog/the-firs...

YOU private search engine

https://you.com/

https://youdotcom.notion.site/FAQ-8c8...

DeepMind's Open-Source Arnheim

https://deepmind.com/research/open-so...

https://twitter.com/OriolVinyalsML/st...

https://github.com/deepmind/arnheim

https://colab.research.google.com/git...

Company sued for using AI to make website more accessible

https://www.wired.com/story/company-t...

https://archive.ph/kdvOM

Alibaba DAMO Academy creates 10 Trillion M6 model

https://pandaily.com/alibaba-damo-aca...

https://www.infoq.cn/article/xIX9leku...

AMD MI200 Family

https://www.anandtech.com/show/17054/...

State of AI report 2021

https://www.stateof.ai/?utm_source=po...

Andrew Ng's Landing AI raises 57M

https://techcrunch.com/2021/11/08/lan...

https://www.forbes.com/sites/bernardm...

https://landing.ai/platform/

Cerebras raises 250M

https://cerebras.net/news/cerebras-sy...

https://cerebras.net/news/cerebras-sy...

Microsoft's Varuna: Scalable Training of Huge Models

https://syncedreview.com/2021/11/10/d...

Laura Ruis reproduces Extrapolation Paper

https://lauraruis.github.io/2021/11/0...

https://github.com/LauraRuis

Ian Charnas' Real-Life Punchout

https://www.reddit.com/r/MachineLearn...

https://www.youtube.com/watch?v=07Jib...

Helpful Things

https://www.marktechpost.com/2021/11/...

https://pair-code.github.io/lit/demos/

https://github.com/pair-code/lit

https://www.reddit.com/r/MachineLearn...

https://twitter.com/yeemachine/status...

https://github.com/yeemachine/kalidokit

AI finds profitable Meme-Tokens

https://finance.yahoo.com/news/artifi...

https://finu.co/

This Sneaker Does Not Exist

https://thissneakerdoesnotexist.com/

Links:

TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://discord.gg/4H8xxDF

BitChute: https://www.bitchute.com/channel/yann...

LinkedIn: https://www.linkedin.com/in/ykilcher

BiliBili: https://space.bilibili.com/2017636191

If you want to support me, the best thing to do is to share out the content :)

View Details

deeplearning #backpropagation #simulation

More and more systems are made differentiable, which means that accurate gradients of these systems' dynamics can be computed exactly. While this development has led to a lot of advances, there are also distinct situations where backpropagation can be a very bad idea. This paper characterizes a few such systems in the domain of iterated dynamical systems, often including some source of stochasticity, resulting in chaotic behavior. In these systems, it is often better to use black-box estimators for gradients than computing them exactly.

OUTLINE:

0:00 - Foreword

1:15 - Intro & Overview

3:40 - Backpropagation through iterated systems

12:10 - Connection to the spectrum of the Jacobian

15:35 - The Reparameterization Trick

21:30 - Problems of reparameterization

26:35 - Example 1: Policy Learning in Simulation

33:05 - Example 2: Meta-Learning Optimizers

36:15 - Example 3: Disk packing

37:45 - Analysis of Jacobians

40:20 - What can be done?

45:40 - Just use Black-Box methods

Paper: https://arxiv.org/abs/2111.05803

Abstract:

Differentiable programming techniques are widely used in the community and are responsible for the machine learning renaissance of the past several decades. While these methods are powerful, they have limits. In this short report, we discuss a common chaos based failure mode which appears in a variety of differentiable circumstances, ranging from recurrent neural networks and numerical physics simulation to training learned optimizers. We trace this failure to the spectrum of the Jacobian of the system under study, and provide criteria for when a practitioner might expect this failure to spoil their differentiation based optimization algorithms.

Authors: Luke Metz, C. Daniel Freeman, Samuel S. Schoenholz, Tal Kachman

Links:

TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://discord.gg/4H8xxDF

BitChute: https://www.bitchute.com/channel/yann...

LinkedIn: https://www.linkedin.com/in/ykilcher

BiliBili: https://space.bilibili.com/2017636191

If you want to support me, the best thing to do is to share out the content :)

If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https://www.subscribestar.com/yannick...

Patreon: https://www.patreon.com/yannickilcher

Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq

Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2

Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m

Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n

View Details

mlnews #turing #reskin

The latest and greatest from the Machine Learning world

Sponsor: Weights & Biases

https://wandb.com

References:

Microsoft Turing Bletchley: Universal Image Language Representation Model

https://www.microsoft.com/en-us/resea...

https://turing.microsoft.com/bletchley

Meta AI Tactile Sensing

https://ai.facebook.com/blog/teaching...

https://ai.facebook.com/blog/reskin-a...

https://twitter.com/AIatMeta/status/1...

AnimeGANv2

https://huggingface.co/spaces/akhaliq...

https://github.com/bryandlee/animegan...

https://github.com/TachibanaYoshino/A...

https://tachibanayoshino.github.io/An...

General In-Hand Object Re-Orientation

https://taochenshh.github.io/projects...

https://arxiv.org/abs/2111.03043

Does Facebook score the "Anger" Emoji too high?

https://www.washingtonpost.com/techno...

IsomorphicLabs: New Alphabet Company for Drug Discovery

https://twitter.com/demishassabis/sta...

https://www.isomorphiclabs.com/blog

ruDALL-E: Russian DALL-E

https://github.com/sberbank-ai/ru-dalle

https://huggingface.co/spaces/anton-l...

https://colab.research.google.com/git...

https://huggingface.co/sberbank-ai/ru...

https://rudalle.ru/

https://habr.com/ru/company/sberbank/...

https://habr-com.translate.goog/ru/co...

Image Scaling Attacks

https://twitter.com/AlexTamkin/status...

https://twitter.com/rzhang88/status/1...

https://arxiv.org/abs/2104.11222

https://twitter.com/arxiv_org/status/...

https://bifold.berlin/preventing-imag...

https://embracethered.com/blog/posts/...

Azure OpenAI Service

https://blogs.microsoft.com/ai/new-az...

https://azure.microsoft.com/en-us/ser...

Neural MMO

https://openai.com/blog/neural-mmo/?u...

https://github.com/jsuarez5341/neural...

https://github.com/jsuarez5341/neural...

https://jsuarez5341.github.io/neural-...

https://jsuarez5341.github.io/neural-...

https://arxiv.org/abs/2110.07594

ArxivDOOM

https://sniklaus.com/arxivdoom?utm_so...

ARC Game

https://github.com/volotat/ARC-Game

https://volotat.github.io/ARC-Game/?

ResNeXtGuesser

https://twitter.com/resnextguesser/st...

Zillow loses money based on AI home price estimation

https://www.reddit.com/r/MachineLearn...

https://www.cbsnews.com/news/zillow-l...

https://www.businessinsider.com/zillo...

https://archive.ph/qEITQ

Helpful Things

https://github.com/PyTorchLightning/p...

https://www.reddit.com/r/MachineLearn...

https://devpost.com/software/iris-7s3yna

https://github.com/prabhuomkar/iris

https://araffin.github.io/post/rliable/

https://github.com/google-research/rl...

https://paperswithcode.com/dataset/me...

AI will make your company great! Promise, Human!

https://fortune.com/2021/11/05/ai-art...

https://sloanreview.mit.edu/projects/...

View Details

machinelearning #ardm #generativemodels

Diffusion models have made large advances in recent months as a new type of generative models. This paper introduces Autoregressive Diffusion Models (ARDMs), which are a mix between autoregressive generative models and diffusion models. ARDMs are trained to be agnostic to the order of autoregressive decoding and give the user a dynamic tradeoff between speed and performance at decoding time. This paper applies ARDMs to both text and image data, and as an extension, the models can also be used to perform lossless compression.

OUTLINE:

0:00 - Intro & Overview

3:15 - Decoding Order in Autoregressive Models

6:15 - Autoregressive Diffusion Models

8:35 - Dependent and Independent Sampling

14:25 - Application to Character-Level Language Models

18:15 - How Sampling & Training Works

26:05 - Extension 1: Parallel Sampling

29:20 - Extension 2: Depth Upscaling

33:10 - Conclusion & Comments

Paper: https://arxiv.org/abs/2110.02037

Abstract:

We introduce Autoregressive Diffusion Models (ARDMs), a model class encompassing and generalizing order-agnostic autoregressive models (Uria et al., 2014) and absorbing discrete diffusion (Austin et al., 2021), which we show are special cases of ARDMs under mild assumptions. ARDMs are simple to implement and easy to train. Unlike standard ARMs, they do not require causal masking of model representations, and can be trained using an efficient objective similar to modern probabilistic diffusion models that scales favourably to highly-dimensional data. At test time, ARDMs support parallel generation which can be adapted to fit any given generation budget. We find that ARDMs require significantly fewer steps than discrete diffusion models to attain the same performance. Finally, we apply ARDMs to lossless compression, and show that they are uniquely suited to this task. Contrary to existing approaches based on bits-back coding, ARDMs obtain compelling results not only on complete datasets, but also on compressing single data points. Moreover, this can be done using a modest number of network calls for (de)compression due to the model's adaptable parallel generation.

Authors: Emiel Hoogeboom, Alexey A. Gritsenko, Jasmijn Bastings, Ben Poole, Rianne van den Berg, Tim Salimans

Links:

TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://discord.gg/4H8xxDF

BitChute: https://www.bitchute.com/channel/yann...

Minds: https://www.minds.com/ykilcher

Parler: https://parler.com/profile/YannicKilcher

LinkedIn: https://www.linkedin.com/in/ykilcher

BiliBili: https://space.bilibili.com/1824646584

If you want to support me, the best thing to do is to share out the content :)

If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https://www.subscribestar.com/yannick...

Patreon: https://www.patreon.com/yannickilcher

Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq

Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2

Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m

View Details

pathways #mlnews #ego4d

Your irregular dose of Machine Learning News.

OUTLINE:

0:00 - Intro

0:20 - Sponsor: Weights & Biases

2:10 - Google Introduces Pathways AI Architecture

6:30 - OpenAI trains Language Models to do High School Math

8:25 - Sam Altman says Neural Networks truly learn

9:35 - Google AI researchers frustrated with lawyers

12:10 - DeepMind RL Lecture Series 2021

12:40 - Fashion Store sells Adversarial Patches

13:15 - A viable method to remove the GIL from CPython

15:05 - BigScience Workshop releases T0

17:40 - Huggingface Hub Dataset Viewer

18:10 - Scite classifies scientific citations

19:25 - Facebook AI Ego4D dataset & challenges

21:50 - Tesla Dojo Configurable Floating Point Spec

23:10 - Windows releases PyTorch-DirectML for Deep Learning on DirectX GPUs

23:50 - Helpful Things

33:00 - Traders use ML to analyze CEOs' language

34:20 - Cadbury creates DeepFake ads for local Indian businesses

35:25 - This Shoe Does Not Exist

Sponsor: Weights & Biases

https://wandb.com

References:

Google Introduces Pathways AI Architecture

https://blog.google/technology/ai/int...

OpenAI trains Language Models to do High School Math

https://openai.com/blog/grade-school-...

https://arxiv.org/abs/2110.14168

Sam Altman says Neural Networks truly learn

https://twitter.com/sama/status/14508...

Google AI researchers frustrated with lawyers

https://archive.ph/lsQJJ#selection-28...

DeepMind RL Lecture Series 2021

https://deepmind.com/learning-resourc...

Fashion Store sells Adversarial Patches

https://twitter.com/naotokui/status/1...

A viable method to remove the GIL from CPython

https://lwn.net/Articles/872869/

BigScience Workshop releases T0

https://bigscience.huggingface.co/

https://arxiv.org/abs/2110.08207

https://huggingface.co/bigscience/T0pp

Huggingface Hub Dataset Viewer

https://twitter.com/huggingface/statu...

Scite classifies scientific citations

https://scite.ai

https://direct.mit.edu/qss/article/do...

Facebook AI Ego4D dataset & challenges

https://ai.facebook.com/blog/teaching...

Tesla Dojo Configurable Floating Point Spec

https://tesla-cdn.thron.com/static/SB...

Windows releases PyTorch-DirectML for Deep Learning on DirectX GPUs

https://devblogs.microsoft.com/window...

Helpful Things

https://github.com/achaiah/pywick?utm...

https://github.com/orybkin/lexa-bench...

https://orybkin.github.io/lexa/

https://twitter.com/danijarh/status/1...

https://github.com/RobertTLange/mle-h...

https://keras.io/examples/vision/mobi...

https://twitter.com/osanseviero/statu...

https://huggingface.co/spaces/flax-co...

https://huggingface.co/transformers/m...

https://github.com/facebookresearch/b...

https://arxiv.org/abs/2110.11216

https://arxiv.org/pdf/2110.11216.pdf

https://github.com/facebookresearch/x...

https://superbbenchmark.org/

https://arxiv.org/abs/2110.07731

https://github.com/BaguaSys/bagua?utm...

https://github.com/cgarciae/treex

https://jax.readthedocs.io/en/latest/...

Traders use ML to analyze CEOs' language

https://www.reuters.com/technology/ai...

Cadbury creates DeepFake ads for local Indian businesses

https://www.bgr.in/entertainment/shah...

This Shoe Does Not Exist

https://www.thisshoedoesnotexist.com/

View Details

efficientzero #muzero #atari

Reinforcement Learning methods are notoriously data-hungry. Notably, MuZero learns a latent world model just from scalar feedback of reward- and policy-predictions, and therefore relies on scale to perform well. However, most RL algorithms fail when presented with very little data. EfficientZero makes several improvements over MuZero that allows it to learn from astonishingly small amounts of data and outperform other methods by a large margin in the low-sample setting. This could be a staple algorithm for future RL research.

OUTLINE:

0:00 - Intro & Outline

2:30 - MuZero Recap

10:50 - EfficientZero improvements

14:15 - Self-Supervised consistency loss

17:50 - End-to-end prediction of the value prefix

20:40 - Model-based off-policy correction

25:45 - Experimental Results & Conclusion

Paper: https://arxiv.org/abs/2111.00210

Code: https://github.com/YeWR/EfficientZero

Note: code not there yet as of release of this video

Abstract:

Reinforcement learning has achieved great success in many applications. However, sample efficiency remains a key challenge, with prominent methods requiring millions (or even billions) of environment steps to train. Recently, there has been significant progress in sample efficient image-based RL algorithms; however, consistent human-level performance on the Atari game benchmark remains an elusive goal. We propose a sample efficient model-based visual RL algorithm built on MuZero, which we name EfficientZero. Our method achieves 190.4% mean human performance and 116.0% median performance on the Atari 100k benchmark with only two hours of real-time game experience and outperforms the state SAC in some tasks on the DMControl 100k benchmark. This is the first time an algorithm achieves super-human performance on Atari games with such little data. EfficientZero's performance is also close to DQN's performance at 200 million frames while we consume 500 times less data. EfficientZero's low sample complexity and high performance can bring RL closer to real-world applicability. We implement our algorithm in an easy-to-understand manner and it is available at this https URL. We hope it will accelerate the research of MCTS-based RL algorithms in the wider community.

Authors: Weirui Ye, Shaohuai Liu, Thanard Kurutach, Pieter Abbeel, Yang Gao

Links:

TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://discord.gg/4H8xxDF

BitChute: https://www.bitchute.com/channel/yann...

Minds: https://www.minds.com/ykilcher

Parler: https://parler.com/profile/YannicKilcher

LinkedIn: https://www.linkedin.com/in/ykilcher

BiliBili: https://space.bilibili.com/1824646584

If you want to support me, the best thing to do is to share out the content :)

If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https://www.subscribestar.com/yannick...

Patreon: https://www.patreon.com/yannickilcher

Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq

Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2

Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m

Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n

View Details

ytalks #siraj #plagiarism

A conversation with Siraj Raval about his journey on YouTube, and the perils of fame.

OUTLINE:

0:00 - Intro

1:30 - Welcome

3:15 - Starting out: From Economics to YouTube

13:00 - More Views: Plagiarizing Video Content

23:30 - One Step Up: Copying A Research Paper

29:15 - Was there another way?

39:00 - Clickbait Course: Make Money with Machine Learning

50:30 - Rock Bottom and the Way Forward

1:01:30 - Advice for Future Generations

Siraj's Channel: https://www.youtube.com/c/SirajRaval

Links:

TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://discord.gg/4H8xxDF

BitChute: https://www.bitchute.com/channel/yann...

Minds: https://www.minds.com/ykilcher

Parler: https://parler.com/profile/YannicKilcher

LinkedIn: https://www.linkedin.com/in/ykilcher

BiliBili: https://space.bilibili.com/1824646584

If you want to support me, the best thing to do is to share out the content :)

If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https://www.subscribestar.com/yannick...

Patreon: https://www.patreon.com/yannickilcher

Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq

Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2

Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m

Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n

View Details

gtc21 #mlnews #mujoco

Registriere für GTC'21 und gewinne eine RTX 3090: https://nvda.ws/2Y2B5ni

OUTLINE:

0:00 - Intro

0:15 - Sponsor: NVIDIA GTC'21

6:10 - DeepMind kauft & Open-Sourct MuJoCo

9:05 - PyTorch 1.10 Veröffentlicht

11:25 - Google Lernt Spreadsheet Formeln

14:15 - handtracking.io

15:25 - Zellinstanzsegmentierungswettbewerb

16:15 - Hilfreiche Bibliotheken

23:15 - Waymo autos verirren sich alle in der selben Sackgasse

24:50 - BlueRiver balanciert Traktoren

References:

DeepMind kauft & open-sourct MuJoCo

https://deepmind.com/blog/announcemen...

PyTorch 1.10 veröffentlicht

https://pytorch.org/blog/pytorch-1.10...

https://developer.nvidia.com/blog/cud...

GoogleAI sagt Tabellen-Formeln voraus

https://ai.googleblog.com/2021/10/pre...

Handtracking im Browser

https://handtracking.io/

https://handtracking.io/draw_demo/

Sartorius Zellinstanzsegmentierungswettbewerb

https://www.kaggle.com/c/sartorius-ce...

Hilfreiche Bibliotheken

https://github.com/IntelLabs/control-...

https://github.com/facebookresearch/s...

https://github.com/facebookresearch/s...

https://github.com/ydataai/ydata-synt...

https://syntheticdata.community/

https://github.com/ydataai/ydata-synt...

https://medium.com/aimstack/aim-3-0-0...

https://github.com/aimhubio/aim

https://robustbench.github.io/

Waymo Autos verirren sich in dieselbe Sackgasse wieder und wieder

https://sanfrancisco.cbslocal.com/202...

BlueRiver balanciert Traktoren

https://www.linkedin.com/posts/lredde...

https://bluerivertechnology.com/ourme...

Links:

TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://discord.gg/4H8xxDF

BitChute: https://www.bitchute.com/channel/yann...

Minds: https://www.minds.com/ykilcher

Parler: https://parler.com/profile/YannicKilcher

LinkedIn: https://www.linkedin.com/in/ykilcher

BiliBili: https://space.bilibili.com/1824646584

If you want to support me, the best thing to do is to share out the content :)

If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https://www.subscribestar.com/yannick...

Patreon: https://www.patreon.com/yannickilcher

Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq

Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2

Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m

Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n

View Details

gtc21 #mlnews #mujoco

Register to GTC'21 and Win a RTX 3090: https://nvda.ws/2Y2B5ni

OUTLINE:

0:00 - Intro

0:15 - Sponsor: NVIDIA GTC'21

5:35 - DeepMind buys & Open-Sources MuJoCo

7:25 - PyTorch 1.10 Released

9:10 - Google Predicts Spreadsheet Formulas

11:25 - handtracking.io

12:25 - Cell Instance Segmentation Challenge

13:00 - Helpful Libraries

17:50 - Waymo cars keep turning into same dead-end

19:35 - BlueRiver balances tractors

References:

DeepMind buys & open-sources MuJoCo

https://deepmind.com/blog/announcemen...

PyTorch 1.10 released

https://pytorch.org/blog/pytorch-1.10...

https://developer.nvidia.com/blog/cud...

GoogleAI predicts spreadsheet formulas

https://ai.googleblog.com/2021/10/pre...

Handtracking in Browser

https://handtracking.io/

https://handtracking.io/draw_demo/

Sartorius Cell Instance Segmentation Competition

https://www.kaggle.com/c/sartorius-ce...

Helpful Libraries

https://github.com/IntelLabs/control-...

https://github.com/facebookresearch/s...

https://github.com/facebookresearch/s...

https://github.com/ydataai/ydata-synt...

https://syntheticdata.community/

https://github.com/ydataai/ydata-synt...

https://medium.com/aimstack/aim-3-0-0...

https://github.com/aimhubio/aim

https://robustbench.github.io/

Waymo cars keep coming to same dead-end over and over

https://sanfrancisco.cbslocal.com/202...

BlueRiver balances tractors

https://www.linkedin.com/posts/lredde...

https://bluerivertechnology.com/ourme...

Links:

TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://discord.gg/4H8xxDF

BitChute: https://www.bitchute.com/channel/yann...

Minds: https://www.minds.com/ykilcher

Parler: https://parler.com/profile/YannicKilcher

LinkedIn: https://www.linkedin.com/in/ykilcher

BiliBili: https://space.bilibili.com/1824646584

If you want to support me, the best thing to do is to share out the content :)

If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https://www.subscribestar.com/yannick...

Patreon: https://www.patreon.com/yannickilcher

Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq

Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2

Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m

Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n

View Details

aiia #ai #art

A trip report from the AiiA Festival in Geneva organized by the ImpactAI foundation.

OUTLINE:

0:00 - Intro

1:50 - Laura Tocmacov: The Festival

4:10 - Timothy O'Hear: The Tech

6:50 - Jonathan O'Hear: The Robot

11:50 - Cléa Chopard: The Artist

17:45 - Final Words

Website: https://aiiafestival.org/en/

Links:

TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://discord.gg/4H8xxDF

BitChute: https://www.bitchute.com/channel/yann...

Minds: https://www.minds.com/ykilcher

Parler: https://parler.com/profile/YannicKilcher

LinkedIn: https://www.linkedin.com/in/ykilcher

BiliBili: https://space.bilibili.com/1824646584

If you want to support me, the best thing to do is to share out the content :)

If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https://www.subscribestar.com/yannick...

Patreon: https://www.patreon.com/yannickilcher

Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq

Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2

Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m

Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n

View Details

gpt3 #knowledge #symbolic

Symbolic knowledge models are usually trained on human-generated corpora that are cumbersome and expensive to create. Such corpora consist of structured triples of symbolic knowledge. This paper takes a different approach and attempts to generate such a corpus by prompting GPT-3. Results show that clever prompting, combined with targeted small critic models trained on human ratings can outperform both human-generated data, as well as the teacher model (GPT-3) itself. The results of this paper give a general recipe for automatically building corpora for various NLP tasks by extracting samples from large language models.

OUTLINE:

0:00 - Intro & Overview

2:30 - Sponsor: Weights & Biases

4:15 - Commonsense Knowledge Graphs

7:50 - ATOMIC dataset

10:00 - Generating the corpus from a model

13:00 - Prompting GPT-3

15:30 - Generating Events

18:40 - Generating Inferences

23:00 - Evaluating the created dataset

26:45 - Introducing the critic

31:25 - Using the critic to filter the data

36:30 - Training a student on the generated data

41:00 - Key Findings

44:45 - Comments & Conclusion

Paper: https://arxiv.org/abs/2110.07178

Code & Corpus: https://github.com/peterwestai2/symbo...

Sponsor: Weights & Biases

https://wandb.com

https://community.wandb.ai/

Abstract:

The common practice for training commonsense models has gone from-human-to-corpus-to-machine: humans author commonsense knowledge graphs in order to train commonsense models. In this work, we investigate an alternative, from-machine-to-corpus-to-machine: general language models author these commonsense knowledge graphs to train commonsense models. Our study leads to a new framework, Symbolic Knowledge Distillation. As with prior art in Knowledge Distillation (Hinton et al., 2015), our approach uses larger models to teach smaller models. A key difference is that we distill knowledge symbolically-as text-in addition to the neural model. We also distill only one aspect-the commonsense of a general language model teacher, allowing the student to be a different type, a commonsense model. Altogether, we show that careful prompt engineering and a separately trained critic model allow us to selectively distill high-quality causal commonsense from GPT-3, a general language model. Empirical results demonstrate that, for the first time, a human-authored commonsense knowledge graph is surpassed by our automatically distilled variant in all three criteria: quantity, quality, and diversity. In addition, it results in a neural commonsense model that surpasses the teacher model's commonsense capabilities despite its 100x smaller size. We apply this to the ATOMIC resource, and share our new symbolic knowledge graph and commonsense models.

Authors: Peter West, Chandra Bhagavatula, Jack Hessel, Jena D. Hwang, Liwei Jiang, Ronan Le Bras, Ximing Lu, Sean Welleck, Yejin Choi

Links:

TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://discord.gg/4H8xxDF

BitChute: https://www.bitchute.com/channel/yann...

Minds: https://www.minds.com/ykilcher

Parler: https://parler.com/profile/YannicKilcher

LinkedIn: https://www.linkedin.com/in/ykilcher

BiliBili: https://space.bilibili.com/1824646584

If you want to support me, the best thing to do is to share out the content :)

If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https://www.subscribestar.com/yannick...

View Details

sbb #seatreview #travel

A friendly parody of Travel Vloggers and Airplane Seat Reviews :)

No, SBB did not pay me for this (but they should ;) )

Links:

TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://discord.gg/4H8xxDF

BitChute: https://www.bitchute.com/channel/yann...

Minds: https://www.minds.com/ykilcher

Parler: https://parler.com/profile/YannicKilcher

LinkedIn: https://www.linkedin.com/in/ykilcher

BiliBili: https://space.bilibili.com/1824646584

If you want to support me, the best thing to do is to share out the content :)

If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https://www.subscribestar.com/yannick...

Patreon: https://www.patreon.com/yannickilcher

Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq

Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2

Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m

Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n

View Details

mlnews #turingnlg #convmixer

Your latest upates on what's happening in the Machine Learning world.

OUTLINE:

0:00 - Intro

0:16 - Weights & Biases raises on 1B valuation (sponsored)

2:30 - Microsoft trains 530 billion parameter model

5:15 - StyleGAN v3 released

6:45 - A few more examples may be worth billions of parameters

8:30 - ConvMixer fits into a tweet

9:45 - Improved VQGAN

11:25 - William Shatner AI chats about his life

12:35 - Google AI pushes material science

14:10 - Gretel AI raises 50M for privacy protection

16:05 - DeepMind's push into ML for biology

19:00 - Schmidhuber laudates Kunihiko Fukushima for Bower Award

21:30 - Helpful Things

22:25 - Mosaic ML out of stealth mode

23:55 - First German self-driving train

24:45 - Ex-Pentagon Chief: China has already won

26:25 - DeepMind becomes profitable

Sponsor: Weights & Biases

https://wandb.com

References:

Microsoft Trains 530B Parameter Model

https://www.microsoft.com/en-us/resea...

StyleGAN 3 Code Released

https://nvlabs.github.io/stylegan3/

https://github.com/NVlabs/stylegan3

https://colab.research.google.com/git...

When do labels help?

https://arxiv.org/pdf/2110.04374.pdf

ml_paper.bruh

https://openreview.net/pdf?id=TVHS5Y4...

Improved VQGAN

https://openreview.net/pdf?id=pfNyExj7z2

William Shatner "AI" & Storyfile

https://www.livescience.com/william-s...

https://www.storyfile.com/

GoogleAI Finds Complex Metal Oxides

https://ai.googleblog.com/2021/10/fin...

GretelAI raises 50M Series B

https://techcrunch.com/2021/10/07/gre...

https://gretel.ai/

https://gretel.ai/blog/why-privacy-by...

DeepMind's Push in ML for Bio

https://www.biorxiv.org/content/10.11...

https://deepmind.com/blog/article/enf...

Kunihiko Fukushima wins Bower Award: Schmidhuber Congratulates

https://www.fi.edu/laureates/kunihiko...

https://www.youtube.com/watch?v=ysOw6...

Helpful Things

https://github.com/UKPLab/beir#beers-...

https://arxiv.org/pdf/2104.08663.pdf

https://bayesoptbook.com/

https://github.com/nvlabs/imaginaire/

https://github.com/NVlabs/imaginaire/...

MosaicML out of Stealth Mode

https://www.mosaicml.com/

https://www.mosaicml.com/blog/founder...

https://app.mosaicml.com/library/imag...

https://github.com/mosaicml/composer

https://mosaicml-composer.readthedocs...

Germany's first self-driving train

https://techxplore.com/news/2021-10-g...

Ex-Pentagon Chief: China has already won tech war

https://nypost.com/2021/10/11/pentago...

DeepMind becomes profitable

https://bdtechtalks.com/2021/10/07/go...

Links:

TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://discord.gg/4H8xxDF

BitChute: https://www.bitchute.com/channel/yann...

Minds: https://www.minds.com/ykilcher

Parler: https://parler.com/profile/YannicKilcher

LinkedIn: https://www.linkedin.com/in/ykilcher

BiliBili: https://space.bilibili.com/1824646584

If you want to support me, the best thing to do is to share out the content :)

If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https://www.subscribestar.com/yannick...

View Details

deepmind #nowcasting #machinelearning

Your holy update on what's new in the Machine Learning world.

OUTLINE:

0:00 - Intro

0:30 - DeepMind tackles Nowcasting

3:30 - The Guardian's shady reporting on TruthfulQA

6:15 - Stochastic training not necessary for generalization

7:35 - Google AI's efficient partitioning of road networks

9:15 - MiniHack Reinforcement Learning Environment

10:45 - Plato XL 11B dialog model

11:35 - AI finishes Beethoven's 10th Symphony

13:10 - AI casts doubt on painting authenticity

15:55 - ShadowDragon social media surveillance

18:45 - Helpful Libraries

25:20 - Samsung to copy-paste brains onto chips

References:

DeepMind improves Nowcasting

https://deepmind.com/blog/article/now...

https://www.nature.com/articles/s4158...

https://github.com/deepmind/deepmind-...

https://colab.research.google.com/git...

The Guardian's shady reporting on TruthfulQA

https://www.theguardian.com/commentis...

Stochastic Training is Not Necessary for Generalization

https://arxiv.org/pdf/2109.14119.pdf

Google AI - Efficient Partitioning of Road Networks

https://ai.googleblog.com/2021/09/eff...

MiniHack Reinforcement Learning Environment

https://ai.facebook.com/blog/minihack...

Baidu PLATO-XL 11B Dialog Model

http://research.baidu.com/Blog/index-...

AI finishes Beethoven's 10th Symphony

https://thenextweb.com/news/computer-...

AI casts doubt on paining authenticity

https://www.smithsonianmag.com/smart-...

https://art-recognition.com/

https://art-recognition.com/case-stud...

https://art-recognition.com/faq/

ShadowDragon Social Media Surveillance

https://www.rt.com/usa/535630-ai-surv...

https://theintercept.com/2021/09/21/s...

Helpful Libraries / Datasets

https://huggingface.co/infinity

https://yanaiela.github.io/TNE/?s=09&...

https://arxiv.org/abs/2109.10282

https://github.com/microsoft/unilm/tr...

https://medium.com/people-ai-research...

https://raft.elicit.org/

https://huggingface.co/spaces/ought/r...

https://huggingface.co/spaces/ought/r...

https://arxiv.org/pdf/2109.14076.pdf

https://arxiv.org/pdf/2109.14394.pdf

https://www.robots.ox.ac.uk/~vgg/rese...

https://zenodo.org/record/5528345#.YV...

https://github.com/yukimasano/PASS/

https://openreview.net/pdf?id=BwzYI-K...

https://github.com/pytorch/data?utm_s...

Samsung Method to copy paste brain onto chip

https://www.engadget.com/samsung-copy...

Links:

TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://discord.gg/4H8xxDF

BitChute: https://www.bitchute.com/channel/yann...

Minds: https://www.minds.com/ykilcher

Parler: https://parler.com/profile/YannicKilcher

LinkedIn: https://www.linkedin.com/in/ykilcher

BiliBili: https://space.bilibili.com/1824646584

If you want to support me, the best thing to do is to share out the content :)

If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https://www.subscribestar.com/yannick...

Patreon: https://www.patreon.com/yannickilcher

Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq

Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2

Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m

View Details

grokking #openai #deeplearning

Grokking is a phenomenon when a neural network suddenly learns a pattern in the dataset and jumps from random chance generalization to perfect generalization very suddenly. This paper demonstrates grokking on small algorithmic datasets where a network has to fill in binary tables. Interestingly, the learned latent spaces show an emergence of the underlying binary operations that the data were created with.

OUTLINE:

0:00 - Intro & Overview

1:40 - The Grokking Phenomenon

3:50 - Related: Double Descent

7:50 - Binary Operations Datasets

11:45 - What quantities influence grokking?

15:40 - Learned Emerging Structure

17:35 - The role of smoothness

21:30 - Simple explanations win

24:30 - Why does weight decay encourage simplicity?

26:40 - Appendix

28:55 - Conclusion & Comments

Paper: https://mathai-iclr.github.io/papers/...

Abstract:

In this paper we propose to study generalization of neural networks on small algorithmically generated datasets. In this setting, questions about data efficiency, memorization, generalization, and speed of learning can be studied in great detail. In some situations we show that neural networks learn through a process of “grokking” a pattern in the data, improving generalization performance from random chance level to perfect generalization, and that this improvement in generalization can happen well past the point of overfitting. We also study generalization as a function of dataset size and find that smaller datasets require increasing amounts of optimization for generalization. We argue that these datasets provide a fertile ground for studying a poorly understood aspect of deep learning: generalization of overparametrized neural networks beyond memorization of the finite training dataset.

Authors: Alethea Power, Yuri Burda, Harri Edwards, Igor Babuschkin & Vedant Misra

Links:

TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://discord.gg/4H8xxDF

BitChute: https://www.bitchute.com/channel/yann...

Minds: https://www.minds.com/ykilcher

Parler: https://parler.com/profile/YannicKilcher

LinkedIn: https://www.linkedin.com/in/ykilcher

BiliBili: https://space.bilibili.com/1824646584

If you want to support me, the best thing to do is to share out the content :)

If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https://www.subscribestar.com/yannick...

Patreon: https://www.patreon.com/yannickilcher

Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq

Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2

Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m

Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n

View Details

deeplearning #co2 #cost

Deep Learning has achieved impressive results in the last years, not least due to the massive increases in computational power and data that has gone into these models. Scaling up currently promises to be a reliable way to create more performant systems, but how far can we go? This article explores the limits of exponential scaling in AI, and what people are doing to get around this problem

OUTLINE:

0:00 - Intro & Overview

1:00 - Deep Learning at its limits

3:10 - The cost of overparameterization

5:40 - Extrapolating power usage and CO2 emissions

10:45 - We cannot just continue scaling up

13:25 - Current solution attempts

15:25 - Aside: ImageNet V2

17:50 - Are symbolic methods the way out?

Paper: https://spectrum.ieee.org/deep-learni...

Image by Ralf Vetterle from Pixabay: https://pixabay.com/images/id-1752876/

Links:

TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://discord.gg/4H8xxDF

BitChute: https://www.bitchute.com/channel/yann...

Minds: https://www.minds.com/ykilcher

Parler: https://parler.com/profile/YannicKilcher

LinkedIn: https://www.linkedin.com/in/ykilcher

BiliBili: https://space.bilibili.com/1824646584

If you want to support me, the best thing to do is to share out the content :)

If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https://www.subscribestar.com/yannick...

Patreon: https://www.patreon.com/yannickilcher

Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq

Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2

Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m

Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n

View Details

plagiarism #surveillance #schmidhuber

Your Mondaily updates of what's going in the world of Machine Learning.

OUTLINE:

0:00 - Intro

0:20 - New plagiarism case has plot twist

7:25 - CLIP for video surveillance

9:40 - DARPA SubTerranean Challenge

11:00 - Schmidhuber criticizing Turing Lecture

15:00 - OpenAI summarizes books

17:55 - UnBiasIt monitors employees' communications for bias

20:00 - iOS plans to detect depression

21:30 - UK 10 year plan to become AI superpower

23:30 - Helpful Libraries

29:00 - WIT: Wikipedia Image-Text dataset

References:

New plagiarism case with plot twist

https://www.reddit.com/r/MachineLearn...

https://zhuanlan.zhihu.com/p/411800486

https://github.com/cybercore-co-ltd/C...

CLIP used for video surveillance

https://www.reddit.com/r/MachineLearn...

https://github.com/johanmodin/clifs

DARPA SubTerranean Challenge

https://twitter.com/BotJunkie/status/...

https://twitter.com/BotJunkie

https://www.subtchallenge.com/index.html

https://www.subtchallenge.com/resourc...

https://twitter.com/dynamicrobots/sta...

Schmidhuber Blog: Turing Lecture Errors

https://people.idsia.ch/~juergen/scie...

OpenAI on Summarizing Books

https://openai.com/blog/summarizing-b...

https://arxiv.org/pdf/2109.10862.pdf

UnBiasIt to monitor employee language

https://edition.cnn.com/2021/09/20/te...

https://www.unbiasit.com/

iPhone to detect depression

https://www.wsj.com/articles/apple-wa...

https://archive.ph/hRTnw

UK 10-year plan to become AI-superpower

https://www.cnbc.com/2021/09/22/uk-pu...

https://archive.ph/4gkKK

Helpful Libraries

https://twitter.com/scikit_learn/stat...

https://scikit-learn.org/stable/auto_...

https://twitter.com/pcastr/status/144...

https://github.com/google/dopamine

https://github.com/microsoft/muzic

https://ai-muzic.github.io/muzic_logo/

https://ai.facebook.com/blog/dynatask...

https://github.com/tum-pbs/PhiFlow

https://github.com/facebookresearch/dora

Habitat and Matterport 3D Dataset

https://github.com/facebookresearch/h...

https://aihabitat.org/

https://arxiv.org/pdf/2109.08238.pdf

WIT: Wikipedia-Based Image-Text Dataset

https://ai.googleblog.com/2021/09/ann...

Links:

TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://discord.gg/4H8xxDF

BitChute: https://www.bitchute.com/channel/yann...

Minds: https://www.minds.com/ykilcher

Parler: https://parler.com/profile/YannicKilcher

LinkedIn: https://www.linkedin.com/in/ykilcher

BiliBili: https://space.bilibili.com/1824646584

If you want to support me, the best thing to do is to share out the content :)

If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https://www.subscribestar.com/yannick...

Patreon: https://www.patreon.com/yannickilcher

Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq

Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2

Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m

Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n

View Details

neurips #peerreview #nips

The peer-review system at Machine Learning conferences has come under much criticism over the last years. One major driver was the infamous 2014 NeurIPS experiment, where a subset of papers were given to two different sets of reviewers. This experiment showed that only about half of all accepted papers were consistently accepted by both committees and demonstrated significant influence of subjectivity. This paper revisits the data from the 2014 experiment and traces the fate of accepted and rejected papers during the 7 years since, and analyzes how well reviewers can assess future impact, among other things.

OUTLINE:

0:00 - Intro & Overview

1:20 - Recap: The 2014 NeurIPS Experiment

5:40 - How much of reviewing is subjective?

11:00 - Validation via simulation

15:45 - Can reviewers predict future impact?

23:10 - Discussion & Comments

Paper: https://arxiv.org/abs/2109.09774

Code: https://github.com/lawrennd/neurips2014/

Abstract:

In this paper we revisit the 2014 NeurIPS experiment that examined inconsistency in conference peer review. We determine that 50% of the variation in reviewer quality scores was subjective in origin. Further, with seven years passing since the experiment we find that for accepted papers, there is no correlation between quality scores and impact of the paper as measured as a function of citation count. We trace the fate of rejected papers, recovering where these papers were eventually published. For these papers we find a correlation between quality scores and impact. We conclude that the reviewing process for the 2014 conference was good for identifying poor papers, but poor for identifying good papers. We give some suggestions for improving the reviewing process but also warn against removing the subjective element. Finally, we suggest that the real conclusion of the experiment is that the community should place less onus on the notion of top-tier conference publications when assessing the quality of individual researchers. For NeurIPS 2021, the PCs are repeating the experiment, as well as conducting new ones.

Authors: Corinna Cortes, Neil D. Lawrence

Links:

TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://discord.gg/4H8xxDF

BitChute: https://www.bitchute.com/channel/yann...

Minds: https://www.minds.com/ykilcher

Parler: https://parler.com/profile/YannicKilcher

LinkedIn: https://www.linkedin.com/in/ykilcher

BiliBili: https://space.bilibili.com/1824646584

If you want to support me, the best thing to do is to share out the content :)

If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https://www.subscribestar.com/yannick...

Patreon: https://www.patreon.com/yannickilcher

Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq

Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2

Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m

Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n

View Details

truthfulqa #efficientnet #laion400M

Your regularly irregular updates on what's happening in the Machine Learning world.

OUTLINE:

0:00 - Intro

0:20 - TruthfulQA benchmark shines new light on GPT-3

2:00 - LAION-400M image-text-pair dataset

4:10 - GoogleAI's EfficientNetV2 and CoAtNet

6:15 - Uber's H3: A hexagonal coordinate system

7:40 - AWS NeurIPS 2021 DeepRacer Challenge

8:15 - Helpful Libraries

9:20 - State of PyTorch in September 2021

10:05 - Physics-Based Deep Learning Book

10:35 - Music-conditioned 3D dance generation

11:40 - Stallman's take on legal issues with Codex

12:20 - Tensorflow DirectML on AMD GPUs

13:00 - Schmidhuber Blog: Turing Oversold

ERRATA:

Uber's H3 is actually not new, but from 2018

References:

TruthfulQA - A benchmark assessing truthfulness of language models

https://owainevans.github.io/pdfs/tru...

LAION-400M image-text-pair dataset

https://laion.ai/laion-400-open-dataset/

https://laion.ai/#top

https://gogetfunding.com/help-us-buil...

https://rom1504.github.io/clip-retrie...

GooleAI releases EfficientNetV2 and CoAtNet

https://ai.googleblog.com/2021/09/tow...

Uber's H3 hexagonal coordinate systems

https://eng.uber.com/h3/?utm_source=p...

NeurIPS 2021 DeepRacer Challenge

https://www.aicrowd.com/challenges/ne...

https://aws.amazon.com/deepracer/

https://gitlab.aicrowd.com/deepracer/...

Helpful Libraries

https://github.com/rom1504/img2dataset

https://github.com/facebookresearch/v...

https://github.com/pyg-team/pytorch_g...

https://aws.amazon.com/blogs/machine-...

State of PyTorch in September 2021

https://dev-discuss.pytorch.org/t/sta...

Physics-Based Deep Learning Book

http://physicsbaseddeeplearning.org/i...

https://arxiv.org/pdf/2109.05237.pdf

Music Conditioned 3D dance generation

https://ai.googleblog.com/2021/09/mus...

Richard Stallman on Codex legal issues

https://news.slashdot.org/story/21/09...

Tensorflow DirectML on AMD

https://wccftech.com/amd-microsoft-br...

Schmidhuber: Turing Oversold

https://people.idsia.ch//~juergen/tur...

Links:

TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://discord.gg/4H8xxDF

BitChute: https://www.bitchute.com/channel/yann...

Minds: https://www.minds.com/ykilcher

Parler: https://parler.com/profile/YannicKilcher

LinkedIn: https://www.linkedin.com/in/ykilcher

BiliBili: https://space.bilibili.com/1824646584

If you want to support me, the best thing to do is to share out the content :)

If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https://www.subscribestar.com/yannick...

Patreon: https://www.patreon.com/yannickilcher

Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq

Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2

Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m

Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n

View Details

gpt-3 #truth #conspiracy

A new benchmark paper has created quite an uproar in the community. TruthfulQA is a dataset of 817 questions probing for imitative falsehoods where language models become less truthful, the larger they get. This surprising counter-intuitive finding validates many people's criticisms of large language models, but is it really the correct conclusion?

OUTLINE:

0:00 - Intro

0:30 - Twitter Paper Announcement

4:10 - Large Language Models are to blame!

5:50 - How was the dataset constructed?

9:25 - The questions are adversarial

12:30 - Are you surprised?!

Paper: https://arxiv.org/abs/2109.07958

Links:

TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://discord.gg/4H8xxDF

BitChute: https://www.bitchute.com/channel/yann...

Minds: https://www.minds.com/ykilcher

Parler: https://parler.com/profile/YannicKilcher

LinkedIn: https://www.linkedin.com/in/ykilcher

BiliBili: https://space.bilibili.com/1824646584

If you want to support me, the best thing to do is to share out the content :)

If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https://www.subscribestar.com/yannick...

Patreon: https://www.patreon.com/yannickilcher

Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq

Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2

Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m

Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n

View Details

tvae #topographic #equivariant

Variational Autoencoders model the latent space as a set of independent Gaussian random variables, which the decoder maps to a data distribution. However, this independence is not always desired, for example when dealing with video sequences, we know that successive frames are heavily correlated. Thus, any latent space dealing with such data should reflect this in its structure. Topographic VAEs are a framework for defining correlation structures among the latent variables and induce equivariance within the resulting model. This paper shows how such correlation structures can be built by correctly arranging higher-level variables, which are themselves independent Gaussians.

OUTLINE:

0:00 - Intro

1:40 - Architecture Overview

6:30 - Comparison to regular VAEs

8:35 - Generative Mechanism Formulation

11:45 - Non-Gaussian Latent Space

17:30 - Topographic Product of Student-t

21:15 - Introducing Temporal Coherence

24:50 - Topographic VAE

27:50 - Experimental Results

31:15 - Conclusion & Comments

Paper: https://arxiv.org/abs/2109.01394

Code: https://github.com/akandykeller/topog...

Abstract:

In this work we seek to bridge the concepts of topographic organization and equivariance in neural networks. To accomplish this, we introduce the Topographic VAE: a novel method for efficiently training deep generative models with topographically organized latent variables. We show that such a model indeed learns to organize its activations according to salient characteristics such as digit class, width, and style on MNIST. Furthermore, through topographic organization over time (i.e. temporal coherence), we demonstrate how predefined latent space transformation operators can be encouraged for observed transformed input sequences -- a primitive form of unsupervised learned equivariance. We demonstrate that this model successfully learns sets of approximately equivariant features (i.e. "capsules") directly from sequences and achieves higher likelihood on correspondingly transforming test sequences. Equivariance is verified quantitatively by measuring the approximate commutativity of the inference network and the sequence transformations. Finally, we demonstrate approximate equivariance to complex transformations, expanding upon the capabilities of existing group equivariant neural networks.

Authors: T. Anderson Keller, Max Welling

Links:

TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://discord.gg/4H8xxDF

BitChute: https://www.bitchute.com/channel/yann...

Minds: https://www.minds.com/ykilcher

Parler: https://parler.com/profile/YannicKilcher

LinkedIn: https://www.linkedin.com/in/ykilcher

BiliBili: https://space.bilibili.com/1824646584

If you want to support me, the best thing to do is to share out the content :)

If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https://www.subscribestar.com/yannick...

Patreon: https://www.patreon.com/yannickilcher

Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq

Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2

Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m

Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n

View Details

schmidhuber #tiktok #roomba

Your regularly irregular update on what's happening in the world of Machine Learning.

OUTLINE:

0:00 - Intro

0:15 - Sponsor: Weights & Biases

1:55 - ML YouTuber reaches 100k subscribers

2:40 - Facebook AI pushes Textless NLP

5:30 - Schmidhuber blog post: I invented everything

7:55 - TikTok algorithm rabbitholes users

10:45 - Roomba learns to avoid poop

11:50 - AI can spot art forgeries

14:55 - Deepmind's plans to separate from Google

16:15 - Cohere raises 40M

16:55 - US Judge rejects AI inventor on patent

17:55 - Altman: GPT-4 not much bigger than GPT-3

18:45 - Salesforce CodeT5

19:45 - DeepMind Reinforcement Learning Lecture Series

20:15 - WikiGraphs Dataset

20:40 - LiveCell Dataset

21:00 - SpeechBrain

21:10 - AI-generated influencer gains 100 sponsorships

22:20 - AI News Questions

23:15 - AI hiring tools reject millions of valid applicants

Sponsor: Weights & Biases

https://wandb.me/start

References:

Facebook AI creates Textless NLP

https://ai.facebook.com/blog/textless...

https://speechbot.github.io/pgslm/?fb...

Schmidhuber invented everything

https://people.idsia.ch/~juergen/most...

How TikTok's algorithm works

https://www.wsj.com/video/series/insi...

Roomba learns to avoid poop

https://edition.cnn.com/2021/09/09/te...

Amateur develops fake art detector

https://blogs.nvidia.com/blog/2021/08...

https://spectrum.ieee.org/this-ai-can...

DeepMind's plan to break away from Google

https://www.businessinsider.com/deepm...

https://archive.ph/8s5IK

Cohere raises USD 40M

https://www.fastcompany.com/90670635/...

https://cohere.ai/

US judge refuses AI patent

https://www.theregister.com/2021/09/0...

Sam Altman on GPT-4

https://www.reddit.com/r/OpenAI/comme...

Salesforce releases CodeT5

https://blog.einstein.ai/codet5/

DeepMind RL lecture series

https://deepmind.com/learning-resourc...

WikiGraphs Dataset

https://github.com/deepmind/deepmind-...

LiveCell Dataset

https://sartorius-research.github.io/...

https://www.nature.com/articles/s4159...

SpeechBrain Library

https://speechbrain.github.io/

AI generated influencer lands 100 sponsorships

https://www.allkpop.com/article/2021/...

AI News Questions

https://www.forbes.com/sites/tomtaull...

https://mindmatters.ai/2021/09/isnt-i...

https://fortune.com/2021/09/07/deepmi...

https://www.forbes.com/sites/anniebro...

https://www.cnbctv18.com/views/view-a...

https://www.kcrw.com/culture/shows/li...

https://techcrunch.com/2021/09/07/ai-...

https://www.forbes.com/sites/bernardm...

AI hiring tools mistakenly reject millions of applicants

https://www.theverge.com/2021/9/6/226...

Links:

TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://discord.gg/4H8xxDF

BitChute: https://www.bitchute.com/channel/yann...

Minds: https://www.minds.com/ykilcher

Parler: https://parler.com/profile/YannicKilcher

LinkedIn: https://www.linkedin.com/in/ykilcher

BiliBili: https://space.bilibili.com/1824646584

If you want to support me, the best thing to do is to share out the content :)

View Details

yannickilcher #machinelearning #100k

OUTLINE:

0:00 - 100k!

1:00 - Announcements & Thanks

3:55 - Channel Statistics

Links:

TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://discord.gg/4H8xxDF

BitChute: https://www.bitchute.com/channel/yann...

Minds: https://www.minds.com/ykilcher

Parler: https://parler.com/profile/YannicKilcher

LinkedIn: https://www.linkedin.com/in/yannic-ki...

BiliBili: https://space.bilibili.com/1824646584

If you want to support me, the best thing to do is to share out the content :)

If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https://www.subscribestar.com/yannick...

Patreon: https://www.patreon.com/yannickilcher

Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq

Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2

Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m

Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n

View Details

mlnews #schmidhuber #muzero

Your regular updates on what's happening in the ML world!

OUTLINE:

0:00 - Intro

0:15 - Sponsor: Weights & Biases

1:45 - Google shuts down health streams

4:25 - AI predicts race from blurry X-Rays

7:35 - Facebook labels black men as primates

11:05 - Distill papers on Graph Neural Networks

11:50 - Jürgen Schmidhuber to lead KAUST AI Initiative

12:35 - GitHub brief on DMCA notices for source code

14:55 - Helpful Reddit Threads

19:40 - Simple Tricks to improve Transformers

20:40 - Apple's Unconstrained Scene Generation

21:40 - Common Objects in 3D dataset

22:20 - WarpDrive Multi-Agent RL framework

23:10 - My new paper: Boosting Search Agents & MuZero

25:15 - Can AI detect depression from speech?

References:

Google shuts down Health Streams

https://techcrunch.com/2021/08/26/goo...

AI predicts race from X-Rays

https://www.iflscience.com/technology...

https://arxiv.org/ftp/arxiv/papers/21...

Facebook labels black men as primates

https://www.nytimes.com/2021/09/03/te...

https://en.wikipedia.org/wiki/Human

Distill articles on GNNs

https://distill.pub/2021/gnn-intro/

https://distill.pub/2021/understandin...

Jürgen Schmidhuber leads KAUST AI initiative

https://people.idsia.ch/~juergen/kaus...

GitHub issues court brief on code DMCAs

https://github.blog/2021-08-31-vague-...

Useful Reddit Threads

https://www.reddit.com/r/MachineLearn...

https://www.reddit.com/r/MachineLearn...

https://www.reddit.com/r/MachineLearn...

https://www.reddit.com/r/MachineLearn...

Tricks to improve Transformers

https://arxiv.org/pdf/2108.12284.pdf

Unconstrained Scene Generation

https://apple.github.io/ml-gsn/

Common Objects in 3D dataset

https://ai.facebook.com/blog/common-o...

WarpDrive Multi-Agent RL framework

https://blog.einstein.ai/warpdrive-fa...

Boosting Search Engines / MuZero Code

https://arxiv.org/abs/2109.00527

https://github.com/google-research/go...

https://github.com/google-research/la...

Can AI detect depression?

https://venturebeat.com/2021/08/31/ai...

Links:

TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://discord.gg/4H8xxDF

BitChute: https://www.bitchute.com/channel/yann...

Minds: https://www.minds.com/ykilcher

Parler: https://parler.com/profile/YannicKilcher

LinkedIn: https://www.linkedin.com/in/yannic-ki...

BiliBili: https://space.bilibili.com/1824646584

If you want to support me, the best thing to do is to share out the content :)

If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https://www.subscribestar.com/yannick...

Patreon: https://www.patreon.com/yannickilcher

Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq

Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2

Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m

Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n

View Details

inftyformer #infinityformer #transformer

Vanilla Transformers are excellent sequence models, but suffer from very harsch constraints on the length of the sequences they can process. Several attempts have been made to extend the Transformer's sequence length, but few have successfully gone beyond a constant factor improvement. This paper presents a method, based on continuous attention mechanisms, to attend to an unbounded past sequence by representing the past as a continuous signal, rather than a sequence. This enables the Infty-Former to effectively enrich the current context with global information, which increases performance on long-range dependencies in sequence tasks. Further, the paper presents the concept of sticky memories, which highlight past events that are of particular importance and elevates their representation in the long-term memory.

OUTLINE:

0:00 - Intro & Overview

1:10 - Sponsor Spot: Weights & Biases

3:35 - Problem Statement

8:00 - Continuous Attention Mechanism

16:25 - Unbounded Memory via concatenation & contraction

18:05 - Does this make sense?

20:25 - How the Long-Term Memory is used in an attention layer

27:40 - Entire Architecture Recap

29:30 - Sticky Memories by Importance Sampling

31:25 - Commentary: Pros and cons of using heuristics

32:30 - Experiments & Results

Paper: https://arxiv.org/abs/2109.00301

Sponsor: Weights & Biases

https://wandb.me/start

Abstract:

Transformers struggle when attending to long contexts, since the amount of computation grows with the context length, and therefore they cannot model long-term memories effectively. Several variations have been proposed to alleviate this problem, but they all have a finite memory capacity, being forced to drop old information. In this paper, we propose the ∞-former, which extends the vanilla transformer with an unbounded long-term memory. By making use of a continuous-space attention mechanism to attend over the long-term memory, the ∞-former's attention complexity becomes independent of the context length. Thus, it is able to model arbitrarily long contexts and maintain "sticky memories" while keeping a fixed computation budget. Experiments on a synthetic sorting task demonstrate the ability of the ∞-former to retain information from long sequences. We also perform experiments on language modeling, by training a model from scratch and by fine-tuning a pre-trained language model, which show benefits of unbounded long-term memories.

Authors: Pedro Henrique Martins, Zita Marinho, André F. T. Martins

Links:

TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://discord.gg/4H8xxDF

BitChute: https://www.bitchute.com/channel/yann...

Minds: https://www.minds.com/ykilcher

Parler: https://parler.com/profile/YannicKilcher

LinkedIn: https://www.linkedin.com/in/yannic-ki...

BiliBili: https://space.bilibili.com/1824646584

If you want to support me, the best thing to do is to share out the content :)

If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https://www.subscribestar.com/yannick...

Patreon: https://www.patreon.com/yannickilcher

Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq

Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2

Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m

Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n

View Details

mlnews #chess #neurips

OUTLINE:

0:00 - Intro

0:30 - Reconnaissance Blind Chess NeurIPS 2021 Competition

3:40 - Colab Pro no longer top priority for GPUs

4:45 - DeepMind uses Graph NNs to do traffic prediction

6:00 - Helpful Libraries: Isaac Gym, Differentiable Human, LVIS, BEHAVIOR

10:25 - Cerebras Wafer Scale Engine Cluster

12:15 - AI Voice Synthesis for Val Kilmer

14:20 - Can AI give thoughtful gifts?

References:

Reconnaissance Blind Chess NeurIPS 2021 Competition

https://rbc.jhuapl.edu/

https://rbc.jhuapl.edu/gameRules

Colab Pro no longer top priority

https://www.reddit.com/r/MachineLearn...

Google Maps ETA prediction using Graph Neural Networks

https://arxiv.org/pdf/2108.11482.pdf

Isaac Gym: RL simulator on GPU

https://arxiv.org/abs/2108.10470

https://sites.google.com/view/isaacgy...

https://developer.nvidia.com/isaac-gym

Cerebras Cluster for massive AI models

https://www.wired.com/story/cerebras-...

Helpful Libraries / Datasets

https://nimblephysics.org/docs/human-...

https://www.lvisdataset.org/

https://arxiv.org/pdf/2108.03332.pdf

AI Voice Reconstruction

https://www.washingtonpost.com/techno...

Can AI make thoughtful gifts?

https://www.forbes.com/sites/anniebro...

Links:

TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://discord.gg/4H8xxDF

BitChute: https://www.bitchute.com/channel/yann...

Minds: https://www.minds.com/ykilcher

Parler: https://parler.com/profile/YannicKilcher

LinkedIn: https://www.linkedin.com/in/yannic-ki...

BiliBili: https://space.bilibili.com/1824646584

If you want to support me, the best thing to do is to share out the content :)

If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https://www.subscribestar.com/yannick...

Patreon: https://www.patreon.com/yannickilcher

Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq

Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2

Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m

Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n

View Details

alibi #transformers #attention

Transformers are essentially set models that need additional inputs to make sense of sequence data. The most widespread additional inputs are position encodings or position embeddings, which add sequence index information in various forms. However, this has put a limit on the resulting model, which cannot run inference on sequences longer than it has been trained on, as it would encounter unfamiliar position encodings. ALiBi solves this by proposing simple linear fixed biases as position information, adding negligible overhead in time and memory, but surprisingly, the resulting model is able to handle inference on sequences many times as long as its training sequences.

OUTLINE:

0:00 - Intro & Overview

1:40 - Position Encodings in Transformers

4:55 - Sinusoidial Position Encodings

11:50 - ALiBi Position Encodings

20:50 - How to choose the slope parameter

23:55 - Experimental Results

29:10 - Comments & Conclusion

Paper: https://ofir.io/train_short_test_long...

Code: https://github.com/ofirpress/attentio...

Abstract:

Since the introduction of the transformer model by Vaswani et al. (2017), a fundamental question remains open: how to achieve extrapolation at inference time to longer sequences than seen during training? We first show that extrapolation can be improved by changing the position representation method, though we find that existing proposals do not allow efficient extrapolation. We introduce a simple and efficient method, Attention with Linear Biases (ALiBi), that allows for extrapolation. ALiBi does not add positional embeddings to the word embeddings; instead, it biases the query-key attention scores with a term that is proportional to their distance. We show that this method allows training a 1.3 billion parameter model on input sequences of length 1024 that extrapolates to input sequences of length 2048, achieving the same perplexity as a sinusoidal position embedding model trained on inputs of length 2048, 11% faster and using 11% less memory. ALiBi’s inductive bias towards recency allows it to outperform multiple strong position methods on the WikiText-103 benchmark. Finally, we provide analysis of ALiBi to understand why it leads to better performance.

Authors: Ofir Press, Noah A. Smith, Mike Lewis

Links:

TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://discord.gg/4H8xxDF

BitChute: https://www.bitchute.com/channel/yann...

Minds: https://www.minds.com/ykilcher

Parler: https://parler.com/profile/YannicKilcher

LinkedIn: https://www.linkedin.com/in/yannic-ki...

BiliBili: https://space.bilibili.com/1824646584

If you want to support me, the best thing to do is to share out the content :)

If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https://www.subscribestar.com/yannick...

Patreon: https://www.patreon.com/yannickilcher

Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq

Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2

Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m

Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n

View Details

plagiarism #foundationmodels #tesla

The best place to keep up to date with the latest and greatest from the ML world!

OUTLINE:

0:00 - Intro & Sponsor

3:15 - A high-profile case of plagiarism shocks the ML world

11:55 - Stanford AI releases paper on "Foundation Models"

19:45 - Updates on Apple's NeuralHash

20:45 - RL control for two-player splorts

21:45 - Tesla's AI Day

23:55 - COMMA THREE announced

24:40 - Intel winding down RealSense cameras

25:20 - IBM unveils Telum Processor

25:50 - Lux AI Challenge & Neural MMO Challenge

26:50 - Dribnet's CLIP PixelArt

27:40 - Multi-Agent RL papers are mostly fake

28:50 - I can't even come up with a segment title

29:25 - AI News Questions

31:20 - Frameworks & Libraries

Sponsor: Weights & Biases

https://wandb.ai

References:

Plagiarism case shocks ML world

https://arxiv.org/abs/2102.07870v1

https://arxiv.org/pdf/2102.07870v1.pdf

https://arxiv.org/abs/2108.05862

https://arxiv.org/pdf/2108.05862v1.pdf

https://www.reddit.com/r/MachineLearn...

https://michaelsdr.github.io/momentum...

https://www.zhihu.com/question/480075...

https://zhuanlan.zhihu.com/p/40035196...

https://finance.sina.com.cn/tech/2021...

https://duoli.org/

https://web.archive.org/web/202108160...

https://twitter.com/shaohua0116/statu...

Stanford AI targets Foundation Models

https://arxiv.org/abs/2108.07258

https://arxiv.org/pdf/2108.07258.pdf

https://ieeexplore.ieee.org/document/...

https://xgboost.readthedocs.io/en/lat...

https://en.wikipedia.org/wiki/Support...

https://scikit-learn.org/stable/modul...

https://syncedreview.com/2019/06/27/t...

https://openai.com/blog/better-langua...

NeuralHash Saga Continues

https://www.reddit.com/r/MachineLearn...

https://blog.roboflow.com/neuralhash-...

https://www.kron4.com/news/bay-area/b...

RL Control for competitive sports

https://ai.facebook.com/research/publ...

Tesla AI Day

https://www.youtube.com/watch?v=ABbDB...

https://spectrum.ieee.org/elon-musk-r...

https://www.youtube.com/watch?v=j0z4F...

George Hotz announces COMMA THREE

https://www.youtube.com/watch?v=jJn2O...

https://comma.ai/shop/products/three

Intel abandons RealSense cameras

https://www.crn.com/news/components-p...

IBM unveils Telum Processor

https://www.prnewswire.com/news-relea...

Kaggle Lux AI challenge

https://www.kaggle.com/c/lux-ai-2021

Neural MMO challenge

https://www.aicrowd.com/challenges/th...

Dribnet's PixelArt

https://twitter.com/dribnet/status/14...

Multi-Agent RL papers mostly fake

https://www.reddit.com/r/reinforcemen...

Elon Musk, Lex Fridman tweets trigger news story

https://www.benzinga.com/news/21/08/2...

News Questions:

https://www.zdnet.com/article/can-ai-...

https://entertainment.inquirer.net/41...

https://www.analyticsinsight.net/whic...

https://www.bbc.co.uk/programmes/m000...

https://ricochet.com/podcast/cosm-tec...

https://www.designnews.com/automation...

https://www.forbes.com/sites/anniebro...

3D Volleyball RL environment

https://www.reddit.com/r/MachineLearn...

Maze RL framework

https://enliteai.medium.com/maze-appl...

Wanderer 2 HN Search

https://metaphor.so/

View Details

attention #transformer #fastformer

Transformers have become the dominant model class in the last few years for large data, but their quadratic complexity in terms of sequence length has plagued them until now. Fastformer claims to be the fastest and most performant linear attention variant, able to consume long contexts at once. This is achieved by a combination of additive attention and elementwise products. While initial results look promising, I have my reservations...

OUTLINE:

0:00 - Intro & Outline

2:15 - Fastformer description

5:20 - Baseline: Classic Attention

10:00 - Fastformer architecture

12:50 - Additive Attention

18:05 - Query-Key element-wise multiplication

21:35 - Redundant modules in Fastformer

25:00 - Problems with the architecture

27:30 - Is this even attention?

32:20 - Experimental Results

34:50 - Conclusion & Comments

Paper: https://arxiv.org/abs/2108.09084

Abstract:

Transformer is a powerful model for text understanding. However, it is inefficient due to its quadratic complexity to input sequence length. Although there are many methods on Transformer acceleration, they are still either inefficient on long sequences or not effective enough. In this paper, we propose Fastformer, which is an efficient Transformer model based on additive attention. In Fastformer, instead of modeling the pair-wise interactions between tokens, we first use additive attention mechanism to model global contexts, and then further transform each token representation based on its interaction with global context representations. In this way, Fastformer can achieve effective context modeling with linear complexity. Extensive experiments on five datasets show that Fastformer is much more efficient than many existing Transformer models and can meanwhile achieve comparable or even better long text modeling performance.

Authors: Chuhan Wu, Fangzhao Wu, Tao Qi, Yongfeng Huang

Links:

TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://discord.gg/4H8xxDF

BitChute: https://www.bitchute.com/channel/yann...

Minds: https://www.minds.com/ykilcher

Parler: https://parler.com/profile/YannicKilcher

LinkedIn: https://www.linkedin.com/in/yannic-ki...

BiliBili: https://space.bilibili.com/1824646584

If you want to support me, the best thing to do is to share out the content :)

If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https://www.subscribestar.com/yannick...

Patreon: https://www.patreon.com/yannickilcher

Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq

Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2

Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m

Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n

View Details

pondernet #deepmind #machinelearning

Humans don't spend the same amount of mental effort on all problems equally. Instead, we respond quickly to easy tasks, and we take our time to deliberate hard tasks. DeepMind's PonderNet attempts to achieve the same by dynamically deciding how many computation steps to allocate to any single input sample. This is done via a recurrent architecture and a trainable function that computes a halting probability. The resulting model performs well in dynamic computation tasks and is surprisingly robust to different hyperparameter settings.

OUTLINE:

0:00 - Intro & Overview

2:30 - Problem Statement

8:00 - Probabilistic formulation of dynamic halting

14:40 - Training via unrolling

22:30 - Loss function and regularization of the halting distribution

27:35 - Experimental Results

37:10 - Sensitivity to hyperparameter choice

41:15 - Discussion, Conclusion, Broader Impact

Paper: https://arxiv.org/abs/2107.05407

Abstract:

In standard neural networks the amount of computation used grows with the size of the inputs, but not with the complexity of the problem being learnt. To overcome this limitation we introduce PonderNet, a new algorithm that learns to adapt the amount of computation based on the complexity of the problem at hand. PonderNet learns end-to-end the number of computational steps to achieve an effective compromise between training prediction accuracy, computational cost and generalization. On a complex synthetic problem, PonderNet dramatically improves performance over previous adaptive computation methods and additionally succeeds at extrapolation tests where traditional neural networks fail. Also, our method matched the current state of the art results on a real world question and answering dataset, but using less compute. Finally, PonderNet reached state of the art results on a complex task designed to test the reasoning capabilities of neural networks.1

Authors: Andrea Banino, Jan Balaguer, Charles Blundell

Links:

TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://discord.gg/4H8xxDF

BitChute: https://www.bitchute.com/channel/yann...

Minds: https://www.minds.com/ykilcher

Parler: https://parler.com/profile/YannicKilcher

LinkedIn: https://www.linkedin.com/in/yannic-ki...

BiliBili: https://space.bilibili.com/1824646584

If you want to support me, the best thing to do is to share out the content :)

If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https://www.subscribestar.com/yannick...

Patreon: https://www.patreon.com/yannickilcher

Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq

Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2

Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m

Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n

View Details

apple #icloud #neuralhash

Send your Apple fanboy friends to prison with this one simple trick ;) We break Apple's NeuralHash algorithm used to detect CSAM for iCloud photos. I show how it's possible to craft arbitrary hash collisions from any source / target image pair using an adversarial example attack. This can be used for many purposes, such as evading detection, or forging false positives, triggering manual reviews.

OUTLINE:

0:00 - Intro

1:30 - Forced Hash Collisions via Adversarial Attacks

2:30 - My Successful Attack

5:40 - Results

7:15 - Discussion

DISCLAIMER: This is for demonstration and educational purposes only. This is not an endorsement of illegal activity or circumvention of law.

Code: https://github.com/yk/neural_hash_col...

Extract Model: https://github.com/AsuharietYgvar/App...

My Video on NeuralHash: https://youtu.be/z15JLtAuwVI

Links:

TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://discord.gg/4H8xxDF

BitChute: https://www.bitchute.com/channel/yann...

Minds: https://www.minds.com/ykilcher

Parler: https://parler.com/profile/YannicKilcher

LinkedIn: https://www.linkedin.com/in/yannic-ki...

BiliBili: https://space.bilibili.com/1824646584

If you want to support me, the best thing to do is to share out the content :)

If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https://www.subscribestar.com/yannick...

Patreon: https://www.patreon.com/yannickilcher

Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq

Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2

Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m

Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n

View Details

mlnews #nvidia #openai

An in-depth look over what's going on in the world of Machine Learning and Artificial intelligence. Subscribe now and make Monday the best day of the week!

OUTLINE:

0:00 - Intro

0:20 - Sponsor: Weights & Biases

3:00 - Nvidia's CEO was rendered during Keynote

5:00 - AI21 Labs releases Jurassic-1 language model

7:00 - Tortured Phrases reveal plagiarism

10:05 - Cortical neurons are computationally complex

11:55 - OpenAI Codex Update & Challenge

13:30 - Automated drug abuse prevention gone wrong

17:55 - Rapid News Questions

18:40 - SoundStream learned neural audio codec

19:40 - RoboMimic framework for robotics research

20:05 - Droidlet framework for agent training

20:40 - Unidentified Video Objects Benchmark

21:45 - Grammatical Error Correction Dataset

22:15 - ColabPro Plus available

23:05 - BigBench Self-Awareness benchmark for language models

Sponsor: Weights & Biases

https://wandb.ai

References:

NVIDIA renders CEO during keynote

https://www.vice.com/en/article/88nbp...

https://blogs.nvidia.com/blog/2021/08...

https://www.youtube.com/watch?v=eAn_o...

AI21 Labs announces Jurassic-1 model

https://www.ai21.com/blog/announcing-...

https://studio.ai21.com/

https://twitter.com/yoavgo/status/142...

Tortured Phrases point to plagiarism

https://www.nature.com/articles/d4158...

Real Neurons are insanely complex

https://www.sciencedirect.com/science...

OpenAI Codex Challenge & Update

https://challenge.openai.com/

https://challenge.openai.com/codex/le...

https://openai.com/blog/openai-codex/...

Automated drug abuse prevention goes wrong

https://www.wired.com/story/opioid-dr...

News Questions

https://www.imeche.org/news/news-arti...

https://newseu.cgtn.com/news/2021-08-...

https://www.growingproduce.com/citrus...

https://www.cioreview.com/news/artifi...

SoundStream Neural Audio Codec

https://ai.googleblog.com/2021/08/sou...

RoboMimic Framework

https://arise-initiative.github.io/ro...

Droidlet Framework

https://ai.facebook.com/blog/droidlet...

Unidentified Video Objects Benchmark

https://ai.facebook.com/blog/introduc...

Grammatical Error Correction Dataset

https://ai.googleblog.com/2021/08/the...

Colab Pro Plus is "even better"

https://colab.research.google.com/signup

BIG-Bench Self-Awareness Benchmark for Language Models

https://github.com/google/BIG-bench/t...

Links:

TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://discord.gg/4H8xxDF

BitChute: https://www.bitchute.com/channel/yann...

Minds: https://www.minds.com/ykilcher

Parler: https://parler.com/profile/YannicKilcher

LinkedIn: https://www.linkedin.com/in/yannic-ki...

BiliBili: https://space.bilibili.com/1824646584

If you want to support me, the best thing to do is to share out the content :)

If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https://www.subscribestar.com/yannick...

Patreon: https://www.patreon.com/yannickilcher

Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq

Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2

Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m

View Details

apple #icloud #privacy

Apple recently announced scanning all images uploaded to iCloud for CSAM (child abuse material), and that this scan would happen locally on users' phones. We take a look at the technical report and explore how the system works in detail, how it is designed to preserve user privacy, and what weak points it still has.

OUTLINE:

0:00 - Introduction

3:05 - System Requirements

9:15 - System Overview

14:00 - NeuralHash

20:45 - Private Set Intersection

31:15 - Threshold Secret Sharing

35:25 - Synthetic Match Vouchers

38:20 - Problem 1: Who controls the database?

42:40 - Problem 2: Adversarial Attacks

49:40 - Comments & Conclusion

Paper: https://www.apple.com/child-safety/pd...

ML News Episode about CSAM: https://youtu.be/gFkBqD2hbnU

Abstract:

CSAM Detection enables Apple to accurately identify and report iCloud users who store known Child Sexual Abuse Material (CSAM) in their iCloud Photos accounts. Apple servers flag accounts exceeding a threshold number of images that match a known database of CSAM image hashes so that Apple can provide relevant information to the National Center for Missing and Exploited Children (NCMEC). This process is secure, and is expressly designed to preserve user privacy.

CSAM Detection provides these privacy and security assurances:

• Apple does not learn anything about images that do not match the known CSAM database.

• Apple can’t access metadata or visual derivatives for matched CSAM images until a threshold of matches is exceeded for an iCloud Photos account.

• The risk of the system incorrectly flagging an account is extremely low. In addition, Apple manually reviews all reports made to NCMEC to ensure reporting accuracy.

• Users can’t access or view the database of known CSAM images.

• Users can’t identify which images were flagged as CSAM by the system.

For detailed information about the cryptographic protocol and security proofs that the CSAM Detection process uses, see The Apple PSI System.

Links:

TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://discord.gg/4H8xxDF

BitChute: https://www.bitchute.com/channel/yann...

Minds: https://www.minds.com/ykilcher

Parler: https://parler.com/profile/YannicKilcher

LinkedIn: https://www.linkedin.com/in/yannic-ki...

BiliBili: https://space.bilibili.com/1824646584

If you want to support me, the best thing to do is to share out the content :)

If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https://www.subscribestar.com/yannick...

Patreon: https://www.patreon.com/yannickilcher

Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq

Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2

Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m

Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n

View Details

mlnews #apple #nolamarck

Your update on the latest news in the AI and Machine Learning world.

OUTLINE:

0:00 - Intro

0:15 - Sponsor: Weights & Biases

3:30 - Apple to scan iDevices for illegal content

14:10 - EU approves chatcontrol

15:20 - Machine Learning FAQ book

17:40 - TimeDial & Disfl-QA Conversation Datasets

20:30 - VoxPopuli Speech Dataset

21:00 - Google Tensor chip coming to Pixel 6

21:30 - Pentagon uses AI to predict events

23:10 - Sketch your own GAN

24:45 - Can a Fruit Fly learn Word Embeddings?

26:00 - Master Faces beat facial recognition system

27:25 - PyTorch profiler 1.9

27:55 - 0 A.D. gets reinforcement learning interface

28:40 - BeatBot cleans up cigarette butts on the beach

Sponsor: Weights & Biases

https://wandb.ai

References:

Apple to scan iDevices for illegal content

https://techcrunch.com/2021/08/05/app...

http://tylerneylon.com/a/lsh1/

EU approves chatcontrol

https://european-pirateparty.eu/parli...

Machine Learning FAQ book

https://rentruewang.github.io/learnin...

TimeDial & Disfl-QA: New datasets for conversational NLP

https://ai.googleblog.com/2021/08/two...

VoxPopuli: Giant partially labeled speech dataset

https://github.com/facebookresearch/v...

Google's Tensor chip coming to Pixel 6

https://blog.google/products/pixel/go...

Pentagon uses AI for predicting relevant events in advance

https://www.engadget.com/pentagon-ai-...

Sketch Your Own GAN

https://peterwang512.github.io/GANSke...

Can a fruit fly learn word embeddings?

https://arxiv.org/pdf/2101.06887.pdf

Master Faces for attacking facial recognition systems

https://arxiv.org/pdf/2108.01077.pdf

PyTorch Profiler v1.9

https://www.marktechpost.com/2021/08/...

0 A.D. adds Reinforcement Learning interface

https://play0ad.com/media/screenshots/

https://trac.wildfiregames.com/wiki/G...

BeachBot cleans up cigarette butts on the beach

https://news.yahoo.com/beachbot-rover...

Links:

TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://discord.gg/4H8xxDF

BitChute: https://www.bitchute.com/channel/yann...

Minds: https://www.minds.com/ykilcher

Parler: https://parler.com/profile/YannicKilcher

LinkedIn: https://www.linkedin.com/in/yannic-ki...

BiliBili: https://space.bilibili.com/1824646584

If you want to support me, the best thing to do is to share out the content :)

If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https://www.subscribestar.com/yannick...

Patreon: https://www.patreon.com/yannickilcher

Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq

Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2

Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m

Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n

View Details

mlnews #dabus #alephalpha

OUTLINE:

0:00 - Intro

0:20 - Sponsor: Weights & Biases

3:45 - AI legally recognized as patent inventor

8:35 - Alpeh Alpha raises USD 27Mio to build European OpenAI

10:20 - AMP advances AI aided recycling

11:20 - DeepMind builds XLand RL environment

13:15 - Cognitive Behavioral Therapy as an app

16:15 - Wordcraft interactive AI text editor

17:05 - ML used to cheat in console games

18:10 - Google's OpenBuildings Dataset

20:00 - Most ML COVID tools are flawed

21:10 - DALL-E mini released

21:55 - Helpful Libraries

25:20 - FSF funds papers discussing CoPilot

SPONSOR: Weights & Biases

https://wandb.ai

References:

AI legally recognized as patent inventor

https://www.globallegalpost.com/news/...

https://www.abc.net.au/news/2021-08-0...

https://artificialinventor.com/freque...

https://artificialinventor.com/dabus/

https://www.worldscientific.com/doi/a...

https://www.worldscientific.com/doi/e...

https://imagination-engines.com/dabus...

https://imagination-engines.com/about...

https://www.nextbigfuture.com/2016/03...

https://www.actiac.org/system/files/D...

Alpeh Alpha raises USD 27Mio to build European OpenAI

https://techcrunch.com/2021/07/27/ger...

AMP advances AI aided recycling

https://www.robotics247.com/article/a...

DeepMind builds XLand RL environment

https://deepmind.com/blog/article/gen...

https://deepmind.com/research/publica...

Cognitive Behavioral Therapy as an app

https://www.nytimes.com/2021/06/01/he...

Wordcraft interactive AI text editor

https://syncedreview.com/2021/07/21/d...

https://arxiv.org/abs/2107.07430

https://www.youtube.com/watch?v=9p4mf...

ML used to cheat in console games

https://au.pcmag.com/games/88121/mach...

Google's OpenBuildings Dataset

https://ai.googleblog.com/2021/07/map...

https://sites.research.google/open-bu...

Most ML COVID tools are flawed

https://www.technologyreview.com/2021...

DALL-E mini released

https://wandb.ai/dalle-mini/dalle-min...

https://huggingface.co/spaces/flax-co...

Helpful Libraries

https://www.openai.com/blog/triton/

https://github.com/openai/triton

https://github.com/microsoft/FLAML

https://github.com/clip-italian/clip-...

https://deepmind.com/research/open-so...

https://github.com/deepmind/meltingpot

https://www.roboti.us/license.html

https://github.com/openai/gym/issues/...

https://github.com/jkterry1

FSF funds papers discussing CoPilot

https://www.fsf.org/blogs/licensing/f...

https://www.gnu.org/philosophy/who-do...

Links:

TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://discord.gg/4H8xxDF

BitChute: https://www.bitchute.com/channel/yann...

Minds: https://www.minds.com/ykilcher

Parler: https://parler.com/profile/YannicKilcher

LinkedIn: https://www.linkedin.com/in/yannic-ki...

BiliBili: https://space.bilibili.com/1824646584

If you want to support me, the best thing to do is to share out the content :)

If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https://www.subscribestar.com/yannick...

Patreon: https://www.patreon.com/yannickilcher

View Details

chai #mlnews #nvidia

Follow Saynam here:

YouTube: https://www.youtube.com/c/ChaiTimeDat...

Twitter: https://twitter.com/bhutanisanyam1

Apple Podcasts: https://podcasts.apple.com/us/podcast...

LinkedIn: https://www.linkedin.com/in/sanyambhu...

Spotify: https://open.spotify.com/show/7IbEWJj...

Anchor.fm RSS: https://anchor.fm/s/c19772c/podcast/rss

Outline:

0:00 - Intro & Overview

1:30 - Amazon's MMO may destroy gaming GPUs

2:40 - OpenAI pivots away from Robotics

3:35 - Google parent Alphabet launches Intrinsic

4:55 - AI learns how vegetables taste

5:55 - NASA uses AI to better understand the sun

6:50 - Man used AI to bring back deceased fiancee

7:45 - Robot collision sparks warehouse fire

8:20 - AI deduces patients' racial identities from medical records

9:40 - AlphaFold protein structure database

10:15 - ICCV BEHAVIOR challenge

11:05 - IBM, MIT, Harvard release Common Sense database

11:35 - High quality image generation using diffusion models

12:50 - Conclusion

References:

1 Amazon’s new MMO may be bricking Nvidia 3090s

https://www.theverge.com/2021/7/21/22...

https://www.youtube.com/watch?v=KLyNF...

2 Open AI pivotes from Robots

https://venturebeat.com/2021/07/23/ai...

3 Google parent Alphabet launches Intrinsic: a new company to build software for industrial robots

https://www.theverge.com/2021/7/23/22...

Introducing Intrinsic

https://blog.x.company/introducing-in...

https://x.company/projects/intrinsic/

https://www.forbes.com/sites/jennifer...

4 Artificial Intelligence Helps Improve NASA’s Eyes on the Sun

https://www.nasa.gov/feature/goddard/...

5 A man used AI to bring back his deceased fiancé. But the creators of the tech warn it could be dangerous

https://www.businessinsider.co.za/man...

6 Robot collision at Ocado warehouse near London sparks fire, delaying customer orders https://www.theverge.com/2021/7/18/22...

10 Reading Race: AI Recognizes Patient’s Racial Identity In Medical Images

https://arxiv.org/pdf/2107.10356.pdf

11 AlphaFold Protein Structure Database

https://alphafold.ebi.ac.uk

https://www.theverge.com/2021/7/22/22...

12 Behavior Challenge

http://svl.stanford.edu/behavior/chal...

13 Researchers from IBM, MIT and Harvard Announced The Release Of DARPA “Common Sense AI” Dataset Along With Two Machine Learning Models At ICML 2021

https://www.marktechpost.com/2021/07/...

https://www.reddit.com/r/MachineLearn...

14 Google uses diffusion model for image generation

https://www.reddit.com/r/MachineLearn...

https://www.reddit.com/r/MachineLearn...

Links:

TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://discord.gg/4H8xxDF

BitChute: https://www.bitchute.com/channel/yann...

Minds: https://www.minds.com/ykilcher

Parler: https://parler.com/profile/YannicKilcher

LinkedIn: https://www.linkedin.com/in/yannic-ki...

BiliBili: https://space.bilibili.com/1824646584

If you want to support me, the best thing to do is to share out the content :)

If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https://www.subscribestar.com/yannick...

Patreon: https://www.patreon.com/yannickilcher

Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq

View Details

A look into the happenings of the Machine Learning world.

OUTLINE:

0:00 - Intro

0:25 - Facebook AI trains rapidly adapting robots

3:05 - Baidu presents autonomous excavator system

4:45 - EleutherAI turns 1

6:05 - Elon Musk says FSD harder than expected

8:10 - AI interview tools still fall short

11:10 - RunwayML AI-powered cloud video editor

11:55 - MineRL BASALT competition to learn from human feedback

13:15 - The Myth of the Expert Reviewer

15:55 - NVIDIA unveils Cambridge-1 supercomputer

17:10 - CLIP art sees rapid improvements

19:00 - AI demystifies boiling

21:20 - AI avatars for easier language learning

23:20 - Outro

References:

Facebook AI trains rapidly adapting robots

https://ai.facebook.com/blog/ai-now-e...

https://ashish-kmr.github.io/rma-legg...

Baidu presents autonomous excavator system

http://research.baidu.com/Blog/index-...

https://www.youtube.com/watch?v=KFcNf...

EleutherAI turns 1

https://blog.eleuther.ai/year-one/

Elon Musk says FSD is harder than expected

https://www.theverge.com/2021/7/5/225...

AI interview tools still fall short

https://www.technologyreview.com/2021...

RunwayML AI-powered cloud video editor

https://runwayml.com/

MineRL BASALT competition to learn from human feedback

https://www.aicrowd.com/challenges/ne...

The Myth of the Expert Reviewer

https://parameterfree.com/2021/07/06/...

NVIDIA unveils Cambridge-1 supercomputer

https://www.nvidia.com/en-us/industri...

https://nvidianews.nvidia.com/news/nv...

CLIP art sees rapid improvements

https://ml.berkeley.edu/blog/posts/cl...

AI demystifies boiling

https://news.mit.edu/2021/infrared-ca...

AI avatars for easier language learning

https://www.forbes.com/sites/petergre...

Links:

TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://discord.gg/4H8xxDF

BitChute: https://www.bitchute.com/channel/yann...

Minds: https://www.minds.com/ykilcher

Parler: https://parler.com/profile/YannicKilcher

LinkedIn: https://www.linkedin.com/in/yannic-ki...

BiliBili: https://space.bilibili.com/1824646584

If you want to support me, the best thing to do is to share out the content :)

If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https://www.subscribestar.com/yannick...

Patreon: https://www.patreon.com/yannickilcher

Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq

Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2

Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m

Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n

View Details

copilot #copyright #gpl

GitHub and OpenAI release Copilot, an AI-powered code autocomplete system that can generate entire functions, classes, and modules from mere definitions and docstrings. Copilot was trained on all public GitHub repositories, and this has a lot of people upset about questions on copyright, code licenses, social obligations, and how much you can profit from other people's work. I give my opinions on the issue in relation to copyright law, the GPL license, and terms of service. Further, we discuss the Brickit app to organize your LEGOs, Distill going on a break, and much more.

OUTLINE:

0:00 - Intro

0:20 - GitHub Copilot

6:55 - My opinion on Copilot & Copyright

17:25 - Facebook AI image similarity challenge

18:00 - Brickit app scans your LEGOs and suggests builds

18:40 - Distill journal goes on break

19:50 - Amazon uses algorithms to hire & fire Flex drivers

23:20 - Helpful Libraries: TF Decision Forests, Habitat, Falken, Brax

24:20 - AI-generated papers give science a hard time

References:

GitHub Copilot: AI pair programmer

https://twitter.com/gdb/status/140989...

https://twitter.com/rickhanlonii/stat...

https://copilot.github.com/

https://docs.github.com/en/github/cop...

https://docs.github.com/en/github/sit...

https://tldrlegal.com/license/gnu-gen...

https://www.gnu.org/licenses/gpl-faq....

https://www.legalzoom.com/knowledge/c...

https://en.wikipedia.org/wiki/Derivat...

https://twitter.com/giffmana/status/1...

https://twitter.com/search?q=copilot&...

Facebook AI launches image similarity challenge

https://www.drivendata.org/competitio...

Brickit app sorts your LEGOs

https://brickit.app/?ref=producthunt&...

https://petapixel.com/2021/07/01/bric...

Distill goes on break

https://distill.pub/2021/distill-hiatus/

Amazon uses Algorithms to fire Flex drivers

https://www.engadget.com/amazon-algor...

TensorFlow decision forests

https://blog.tensorflow.org/2021/05/i...

Facebook AI habitat 2.0

https://ai.facebook.com/blog/habitat-...

Google Falken trains game-playing agents

https://ai.googleblog.com/2021/06/qui...

https://github.com/google-research/fa...

Google Brax: differentiable physics simulator

https://github.com/google/brax

https://arxiv.org/pdf/2106.13281.pdf

Fake science is getting faker

https://thenextweb.com/news/fake-scie...

Links:

TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://discord.gg/4H8xxDF

BitChute: https://www.bitchute.com/channel/yann...

Minds: https://www.minds.com/ykilcher

Parler: https://parler.com/profile/YannicKilcher

LinkedIn: https://www.linkedin.com/in/yannic-ki...

BiliBili: https://space.bilibili.com/1824646584

If you want to support me, the best thing to do is to share out the content :)

If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https://www.subscribestar.com/yannick...

Patreon: https://www.patreon.com/yannickilcher

Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq

Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2

Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m

Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n

View Details

tesla #selfdriving #karpathy

Tesla is pushing the state-of-the-art in full self-driving, and interestingly, they explicitly switch from having multiple different sensors to a vision-only system. We discuss the highlights of Andrej Karpathy's talk about Tesla's FSD system, how to label petabytes of data, how to sample edge-cases, how to train a neural network that has to work in real-time, and why moving to having only cameras is superior to multi-sensor approaches.

OUTLINE:

0:00 - Intro & Overview

1:55 - Current Auto-Breaking system

3:20 - Full Self-Driving from vision only

4:55 - Auto-Labelling for collecting data

8:45 - How to get diverse data from edge-cases

12:15 - Neural network architecture

16:05 - Tesla's in-house supercomputer

17:00 - Owning the whole pipeline

18:20 - Example results from vision only

23:10 - Conclusion & Comments

Links:

TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://discord.gg/4H8xxDF

BitChute: https://www.bitchute.com/channel/yann...

Minds: https://www.minds.com/ykilcher

Parler: https://parler.com/profile/YannicKilcher

LinkedIn: https://www.linkedin.com/in/yannic-ki...

BiliBili: https://space.bilibili.com/1824646584

If you want to support me, the best thing to do is to share out the content :)

If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https://www.subscribestar.com/yannick...

Patreon: https://www.patreon.com/yannickilcher

Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq

Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2

Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m

Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n

View Details

cvpr #socialmedia #machinelearning

In this week's ML news we look at CVPR's controversial action to ban paper promotions on social media during the review phase, among other things!

OUTLINE:

0:00 - Intro & Overview

0:25 - CVPR bans social media paper discussions

5:10 - WalMart uses AI to suggest substitutions

6:05 - NVIDIA releases Alias-Free GAN

7:30 - Confession Video in Myanmar possibly a DeepFake

8:50 - AI restores Rembrandt painting

10:40 - AI for healthcare not problem-free yet

11:50 - ML interviews book

12:15 - NVIDIA canvas turns sketches into paintings

13:00 - GPU prices down after crypto shock

13:30 - Facebook AI improves shopping experience

14:05 - DeepLab2 released on GitHub

14:35 - Toxic Language Models: Nobody cares

16:55 - Does AI have common sense?

References:

CVPR forbids social media promotion

https://twitter.com/wjscheirer/status...

WalMart uses AI to substitute out-of-stock products

https://www.supermarketnews.com/techn...

NVIDIA releases Alias-Free GAN

https://nvlabs.github.io/alias-free-gan/

Myanmar Politician's confession could be DeepFake

https://www.wired.com/story/opinion-t...

Rembrandt restored using AI

https://www.smithsonianmag.com/smart-...

AI in healthcare still shaky

http://www.greenvillebusinessmag.com/...

https://www.theverge.com/2021/6/22/22...

ML interviews book

https://huyenchip.com/ml-interviews-b...

NVIDIA Canvas Beta available

https://blogs.nvidia.com/blog/2021/06...

GPU prices down as China cracks down on Crypto

https://www.theregister.com/2021/06/2...

Facebook AI's big goal of improving shopping

https://ai.facebook.com/blog/advancin...

GoogleAI releases DeepLab2

https://github.com/google-research/de...

Toxic Language Model: Nobody cares

https://arxiv.org/pdf/2105.03023.pdf

AI has no common sense

https://www.analyticsinsight.net/inca...

https://6b.eleuther.ai/

Links:

TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://discord.gg/4H8xxDF

BitChute: https://www.bitchute.com/channel/yann...

Minds: https://www.minds.com/ykilcher

Parler: https://parler.com/profile/YannicKilcher

LinkedIn: https://www.linkedin.com/in/yannic-ki...

BiliBili: https://space.bilibili.com/1824646584

If you want to support me, the best thing to do is to share out the content :)

If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https://www.subscribestar.com/yannick...

Patreon: https://www.patreon.com/yannickilcher

Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq

Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2

Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m

Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n

View Details

adversarialexamples #dimpledmanifold #security

Adversarial Examples have long been a fascinating topic for many Machine Learning researchers. How can a tiny perturbation cause the neural network to change its output by so much? While many explanations have been proposed over the years, they all appear to fall short. This paper attempts to comprehensively explain the existence of adversarial examples by proposing a view of the classification landscape, which they call the Dimpled Manifold Model, which says that any classifier will adjust its decision boundary to align with the low-dimensional data manifold, and only slightly bend around the data. This potentially explains many phenomena around adversarial examples. Warning: In this video, I disagree. Remember that I'm not an authority, but simply give my own opinions.

OUTLINE:

0:00 - Intro & Overview

7:30 - The old mental image of Adversarial Examples

11:25 - The new Dimpled Manifold Hypothesis

22:55 - The Stretchy Feature Model

29:05 - Why do DNNs create Dimpled Manifolds?

38:30 - What can be explained with the new model?

1:00:40 - Experimental evidence for the Dimpled Manifold Model

1:10:25 - Is Goodfellow's claim debunked?

1:13:00 - Conclusion & Comments

Paper: https://arxiv.org/abs/2106.10151

My replication code: https://gist.github.com/yk/de8d987c4e...

Goodfellow's Talk: https://youtu.be/CIfsB_EYsVI?t=4280

Abstract:

The extreme fragility of deep neural networks when presented with tiny perturbations in their inputs was independently discovered by several research groups in 2013, but in spite of enormous effort these adversarial examples remained a baffling phenomenon with no clear explanation. In this paper we introduce a new conceptual framework (which we call the Dimpled Manifold Model) which provides a simple explanation for why adversarial examples exist, why their perturbations have such tiny norms, why these perturbations look like random noise, and why a network which was adversarially trained with incorrectly labeled images can still correctly classify test images. In the last part of the paper we describe the results of numerous experiments which strongly support this new model, and in particular our assertion that adversarial perturbations are roughly perpendicular to the low dimensional manifold which contains all the training examples.

Abstract: Adi Shamir, Odelia Melamed, Oriel BenShmuel

Links:

TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://discord.gg/4H8xxDF

BitChute: https://www.bitchute.com/channel/yann...

Minds: https://www.minds.com/ykilcher

Parler: https://parler.com/profile/YannicKilcher

LinkedIn: https://www.linkedin.com/in/ykilcher

BiliBili: https://space.bilibili.com/1824646584

If you want to support me, the best thing to do is to share out the content :)

If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https://www.subscribestar.com/yannick...

Patreon: https://www.patreon.com/yannickilcher

Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq

Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2

Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m

Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n

View Details

mlnews #gta #weather

In this week's ML News, we look at the latest developments in the Machine Learning and AI world with updates from research, industry, and society at large.

OUTLINE:

0:00 - Intro

0:20 - Hugging Face launches free course

1:30 - Sentdex releases GAN Theft Auto

2:25 - Facebook uses AI to help moderators

4:10 - Weather with Antonio

5:10 - Autonomous ship aborts mission

7:25 - PyTorch Release 1.9

8:30 - McDonald's new AI drive thru

10:20 - UBS CEO says AI won't replace humans

12:20 - Gödel paper has 90th birthday

12:55 - AugLy data augmentation library

13:20 - Programming Puzzles for autonomous coding

14:30 - Boston Dynamics' Spot turns 1

References:

PyTorch 1.9 Released

https://pytorch.org/blog/pytorch-1.9-...

Hugging Face launches course

https://huggingface.co/course/chapter1

90 years of Gödel's theory

https://people.idsia.ch/~juergen/goed...

AugLy: A data augmentation library

https://ai.facebook.com/blog/augly-a-...

Sentdex builds GAN Theft Auto

https://github.com/sentdex/GANTheftAuto/

Spot turns 1

https://blog.bostondynamics.com/spots...

Autonomous ship aborts mission

https://www.washingtonpost.com/techno...

https://mas400.com/dashboard#currentL...

McDonald's tests AI drive thru

https://www.zdnet.com/article/i-just-...

Facebook uses AI to moderate conversations

https://edition.cnn.com/2021/06/16/te...

UBS CEO says AI won't replace financial advisors

https://www.cnbc.com/2021/06/17/ai-wo...

Programming Puzzles

https://arxiv.org/abs/2106.05784

https://github.com/microsoft/PythonPr...

Links:

TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://discord.gg/4H8xxDF

BitChute: https://www.bitchute.com/channel/yann...

Minds: https://www.minds.com/ykilcher

Parler: https://parler.com/profile/YannicKilcher

LinkedIn: https://www.linkedin.com/in/yannic-ki...

BiliBili: https://space.bilibili.com/1824646584

If you want to support me, the best thing to do is to share out the content :)

If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https://www.subscribestar.com/yannick...

Patreon: https://www.patreon.com/yannickilcher

Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq

Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2

Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m

Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n

View Details

xcit #transformer #attentionmechanism

After dominating Natural Language Processing, Transformers have taken over Computer Vision recently with the advent of Vision Transformers. However, the attention mechanism's quadratic complexity in the number of tokens means that Transformers do not scale well to high-resolution images. XCiT is a new Transformer architecture, containing XCA, a transposed version of attention, reducing the complexity from quadratic to linear, and at least on image data, it appears to perform on par with other models. What does this mean for the field? Is this even a transformer? What really matters in deep learning?

OUTLINE:

0:00 - Intro & Overview

3:45 - Self-Attention vs Cross-Covariance Attention (XCA)

19:55 - Cross-Covariance Image Transformer (XCiT) Architecture

26:00 - Theoretical & Engineering considerations

30:40 - Experimental Results

33:20 - Comments & Conclusion

Paper: https://arxiv.org/abs/2106.09681

Code: https://github.com/facebookresearch/xcit

Abstract:

Following their success in natural language processing, transformers have recently shown much promise for computer vision. The self-attention operation underlying transformers yields global interactions between all tokens ,i.e. words or image patches, and enables flexible modelling of image data beyond the local interactions of convolutions. This flexibility, however, comes with a quadratic complexity in time and memory, hindering application to long sequences and high-resolution images. We propose a "transposed" version of self-attention that operates across feature channels rather than tokens, where the interactions are based on the cross-covariance matrix between keys and queries. The resulting cross-covariance attention (XCA) has linear complexity in the number of tokens, and allows efficient processing of high-resolution images. Our cross-covariance image transformer (XCiT) is built upon XCA. It combines the accuracy of conventional transformers with the scalability of convolutional architectures. We validate the effectiveness and generality of XCiT by reporting excellent results on multiple vision benchmarks, including image classification and self-supervised feature learning on ImageNet-1k, object detection and instance segmentation on COCO, and semantic segmentation on ADE20k.

Authors: Alaaeldin El-Nouby, Hugo Touvron, Mathilde Caron, Piotr Bojanowski, Matthijs Douze, Armand Joulin, Ivan Laptev, Natalia Neverova, Gabriel Synnaeve, Jakob Verbeek, Hervé Jegou

Links:

TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://discord.gg/4H8xxDF

BitChute: https://www.bitchute.com/channel/yann...

Minds: https://www.minds.com/ykilcher

Parler: https://parler.com/profile/YannicKilcher

LinkedIn: https://www.linkedin.com/in/yannic-ki...

BiliBili: https://space.bilibili.com/1824646584

If you want to support me, the best thing to do is to share out the content :)

If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https://www.subscribestar.com/yannick...

Patreon: https://www.patreon.com/yannickilcher

Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq

Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2

Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m

Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n

View Details

reiforcementlearning #gan #imitationlearning

Learning from demonstrations is a fascinating topic, but what if the demonstrations are not exactly the behaviors we want to learn? Can we adhere to a dataset of demonstrations and still achieve a specified goal? This paper uses GANs to combine goal-achieving reinforcement learning with imitation learning and learns to perform well at a given task while doing so in the style of a given presented dataset. The resulting behaviors include many realistic-looking transitions between the demonstrated movements.

OUTLINE:

0:00 - Intro & Overview

1:25 - Problem Statement

6:10 - Reward Signals

8:15 - Motion Prior from GAN

14:10 - Algorithm Overview

20:15 - Reward Engineering & Experimental Results

30:40 - Conclusion & Comments

Paper: https://arxiv.org/abs/2104.02180

Main Video: https://www.youtube.com/watch?v=wySUx...

Supplementary Video: https://www.youtube.com/watch?v=O6fBS...

Abstract:

Synthesizing graceful and life-like behaviors for physically simulated characters has been a fundamental challenge in computer animation. Data-driven methods that leverage motion tracking are a prominent class of techniques for producing high fidelity motions for a wide range of behaviors. However, the effectiveness of these tracking-based methods often hinges on carefully designed objective functions, and when applied to large and diverse motion datasets, these methods require significant additional machinery to select the appropriate motion for the character to track in a given scenario. In this work, we propose to obviate the need to manually design imitation objectives and mechanisms for motion selection by utilizing a fully automated approach based on adversarial imitation learning. High-level task objectives that the character should perform can be specified by relatively simple reward functions, while the low-level style of the character's behaviors can be specified by a dataset of unstructured motion clips, without any explicit clip selection or sequencing. These motion clips are used to train an adversarial motion prior, which specifies style-rewards for training the character through reinforcement learning (RL). The adversarial RL procedure automatically selects which motion to perform, dynamically interpolating and generalizing from the dataset. Our system produces high-quality motions that are comparable to those achieved by state-of-the-art tracking-based techniques, while also being able to easily accommodate large datasets of unstructured motion clips. Composition of disparate skills emerges automatically from the motion prior, without requiring a high-level motion planner or other task-specific annotations of the motion clips. We demonstrate the effectiveness of our framework on a diverse cast of complex simulated characters and a challenging suite of motor control tasks.

Authors: Xue Bin Peng, Ze Ma, Pieter Abbeel, Sergey Levine, Angjoo Kanazawa

Links:

TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://discord.gg/4H8xxDF

BitChute: https://www.bitchute.com/channel/yann...

Minds: https://www.minds.com/ykilcher

Parler: https://parler.com/profile/YannicKilcher

LinkedIn: https://www.linkedin.com/in/ykilcher

BiliBili: https://space.bilibili.com/1824646584

If you want to support me, the best thing to do is to share out the content :)

If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https://www.subscribestar.com/yannick...

Patreon: https://www.patreon.com/yannickilcher

View Details

OUTLINE:

0:00 - Intro

0:30 - Google RL creates next-gen TPUs

2:15 - Facebook launches NetHack challenge

3:50 - OpenAI mitigates bias by fine-tuning

9:05 - Google AI releases browseable reconstruction of human cortex

9:50 - GPT-J 6B Transformer in JAX

12:00 - Tensorflow launches Forum

13:50 - Text style transfer from a single word

15:45 - ALiEn artificial life simulator

My Video on Chip Placement: https://youtu.be/PDRtyrVskMU

References:

RL creates next-gen TPUs

https://www.nature.com/articles/s4158...

https://www.youtube.com/watch?v=PDRty...

Facebook launches NetHack challenge

https://ai.facebook.com/blog/launchin...

Mitigating bias by fine-tuning

https://openai.com/blog/improving-lan...

Human Cortex 3D Reconstruction

https://ai.googleblog.com/2021/06/a-b...

GPT-J: An open-source 6B transformer

https://arankomatsuzaki.wordpress.com...

https://6b.eleuther.ai/

https://github.com/kingoflolz/mesh-tr...

Tensorflow launches "Forum"

https://discuss.tensorflow.org/

Text style transfer from single word

https://ai.facebook.com/blog/ai-can-n...

ALiEn Life Simulator

https://github.com/chrxh/alien

Links:

TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://discord.gg/4H8xxDF

BitChute: https://www.bitchute.com/channel/yann...

Minds: https://www.minds.com/ykilcher

Parler: https://parler.com/profile/YannicKilcher

LinkedIn: https://www.linkedin.com/in/ykilcher

BiliBili: https://space.bilibili.com/1824646584

If you want to support me, the best thing to do is to share out the content :)

If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https://www.subscribestar.com/yannick...

Patreon: https://www.patreon.com/yannickilcher

Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq

Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2

Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m

Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n

View Details

implicitfunction #jax #autodiff

Many problems in Machine Learning involve loops of inner and outer optimization. Finding update steps for the outer loop is usually difficult, because of the.need to differentiate through the inner loop's procedure over multiple steps. Such loop unrolling is very limited and constrained to very few steps. Other papers have found solutions around unrolling in very specific, individual problems. This paper proposes a unified framework for implicit differentiation of inner optimization procedures without unrolling and provides implementations that integrate seamlessly into JAX.

OUTLINE:

0:00 - Intro & Overview

2:05 - Automatic Differentiation of Inner Optimizations

4:30 - Example: Meta-Learning

7:45 - Unrolling Optimization

13:00 - Unified Framework Overview & Pseudocode

21:10 - Implicit Function Theorem

25:45 - More Technicalities

28:45 - Experiments

ERRATA:

  • Dataset Distillation is done with respect to the training set, not the validation or test set.

Paper: https://arxiv.org/abs/2105.15183

Code coming soon

Abstract:

Automatic differentiation (autodiff) has revolutionized machine learning. It allows expressing complex computations by composing elementary ones in creative ways and removes the burden of computing their derivatives by hand. More recently, differentiation of optimization problem solutions has attracted widespread attention with applications such as optimization as a layer, and in bi-level problems such as hyper-parameter optimization and meta-learning. However, the formulas for these derivatives often involve case-by-case tedious mathematical derivations. In this paper, we propose a unified, efficient and modular approach for implicit differentiation of optimization problems. In our approach, the user defines (in Python in the case of our implementation) a function F capturing the optimality conditions of the problem to be differentiated. Once this is done, we leverage autodiff of F and implicit differentiation to automatically differentiate the optimization problem. Our approach thus combines the benefits of implicit differentiation and autodiff. It is efficient as it can be added on top of any state-of-the-art solver and modular as the optimality condition specification is decoupled from the implicit differentiation mechanism. We show that seemingly simple principles allow to recover many recently proposed implicit differentiation methods and create new ones easily. We demonstrate the ease of formulating and solving bi-level optimization problems using our framework. We also showcase an application to the sensitivity analysis of molecular dynamics.

Authors: Mathieu Blondel, Quentin Berthet, Marco Cuturi, Roy Frostig, Stephan Hoyer, Felipe Llinares-López, Fabian Pedregosa, Jean-Philippe Vert

Links:

TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://discord.gg/4H8xxDF

BitChute: https://www.bitchute.com/channel/yann...

Minds: https://www.minds.com/ykilcher

Parler: https://parler.com/profile/YannicKilcher

LinkedIn: https://www.linkedin.com/in/yannic-ki...

BiliBili: https://space.bilibili.com/1824646584

If you want to support me, the best thing to do is to share out the content :)

If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https://www.subscribestar.com/yannick...

Patreon: https://www.patreon.com/yannickilcher

Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq

View Details

mlnews #wudao #academicfraud

OUTLINE:

0:00 - Intro

0:25 - EU seeks to regulate AI

2:45 - AI COVID detection systems are all flawed

5:05 - Chinese lab trains model 10x GPT-3 size

6:55 - Google error identifies "ugliest" language

9:45 - McDonald's learns about AI buzzwords

11:25 - AI predicts cryptocurrency prices

12:00 - Unreal Engine hack for CLIP

12:35 - Please commit more academic fraud

References:

https://www.lawfareblog.com/artificia...

https://blogs.sciencemag.org/pipeline...

https://www.nature.com/articles/s4225...

https://en.pingwest.com/a/8693

https://arxiv.org/pdf/2104.12369.pdf

https://www.bbc.com/news/world-asia-i...

https://www.zdnet.com/article/mcdonal...

https://www.analyticsinsight.net/ai-i...

https://twitter.com/arankomatsuzaki/s...

https://jacobbuckman.com/2021-05-29-p...

Links:

TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://discord.gg/4H8xxDF

BitChute: https://www.bitchute.com/channel/yann...

Minds: https://www.minds.com/ykilcher

Parler: https://parler.com/profile/YannicKilcher

LinkedIn: https://www.linkedin.com/in/yannic-ki...

BiliBili: https://space.bilibili.com/1824646584

If you want to support me, the best thing to do is to share out the content :)

If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https://www.subscribestar.com/yannick...

Patreon: https://www.patreon.com/yannickilcher

Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq

Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2

Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m

Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n

View Details

decisiontransformer #reinforcementlearning #transformer

Proper credit assignment over long timespans is a fundamental problem in reinforcement learning. Even methods designed to combat this problem, such as TD-learning, quickly reach their limits when rewards are sparse or noisy. This paper reframes offline reinforcement learning as a pure sequence modeling problem, with the actions being sampled conditioned on the given history and desired future rewards. This allows the authors to use recent advances in sequence modeling using Transformers and achieve competitive results in Offline RL benchmarks.

OUTLINE:

0:00 - Intro & Overview

4:15 - Offline Reinforcement Learning

10:10 - Transformers in RL

14:25 - Value Functions and Temporal Difference Learning

20:25 - Sequence Modeling and Reward-to-go

27:20 - Why this is ideal for offline RL

31:30 - The context length problem

34:35 - Toy example: Shortest path from random walks

41:00 - Discount factors

45:50 - Experimental Results

49:25 - Do you need to know the best possible reward?

52:15 - Key-to-door toy experiment

56:00 - Comments & Conclusion

Paper: https://arxiv.org/abs/2106.01345

Website: https://sites.google.com/berkeley.edu...

Code: https://github.com/kzl/decision-trans...

Abstract:

We present a framework that abstracts Reinforcement Learning (RL) as a sequence modeling problem. This allows us to draw upon the simplicity and scalability of the Transformer architecture, and associated advances in language modeling such as GPT-x and BERT. In particular, we present Decision Transformer, an architecture that casts the problem of RL as conditional sequence modeling. Unlike prior approaches to RL that fit value functions or compute policy gradients, Decision Transformer simply outputs the optimal actions by leveraging a causally masked Transformer. By conditioning an autoregressive model on the desired return (reward), past states, and actions, our Decision Transformer model can generate future actions that achieve the desired return. Despite its simplicity, Decision Transformer matches or exceeds the performance of state-of-the-art model-free offline RL baselines on Atari, OpenAI Gym, and Key-to-Door tasks.

Authors: Lili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee, Aditya Grover, Michael Laskin, Pieter Abbeel, Aravind Srinivas, Igor Mordatch

Links:

TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://discord.gg/4H8xxDF

BitChute: https://www.bitchute.com/channel/yann...

Minds: https://www.minds.com/ykilcher

Parler: https://parler.com/profile/YannicKilcher

LinkedIn: https://www.linkedin.com/in/yannic-ki...

BiliBili: https://space.bilibili.com/1824646584

If you want to support me, the best thing to do is to share out the content :)

If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https://www.subscribestar.com/yannick...

Patreon: https://www.patreon.com/yannickilcher

Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq

Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2

Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m

Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n

View Details

mlnews #anthropic #eliza

Anthropic raises $124M for steerable AI, peer review is threatened by collusion rings, and the original ELIZA source code was discovered.

OUTLINE:

0:00 - Intro

0:40 - Anthropic raises $124M

3:25 - 65% of execs can't explain AI predictions

4:25 - DeepMind releases AndroidEnv

6:10 - Collusion rings in ML Conferences

7:30 - ELIZA's original source code discovered

10:45 - OpenAI raises $100M fund

11:25 - Outro

References:

https://techcrunch.com/2021/05/28/ant...

https://www.anthropic.com/news/announ...

https://www.anthropic.com/

https://openai.com/blog/introducing-o...

https://deepmind.com/research/publica...

https://cacm.acm.org/magazines/2021/6...

https://venturebeat.com/2021/05/25/65...

https://techcrunch.com/2021/05/26/ope...

https://sites.google.com/view/elizage...

http://psych.fullerton.edu/mbirnbaum/...

https://en.wikipedia.org/wiki/Carl_Ro...

https://openai.com/fund/

Links:

TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://discord.gg/4H8xxDF

BitChute: https://www.bitchute.com/channel/yann...

Minds: https://www.minds.com/ykilcher

Parler: https://parler.com/profile/YannicKilcher

LinkedIn: https://www.linkedin.com/in/yannic-ki...

BiliBili: https://space.bilibili.com/1824646584

If you want to support me, the best thing to do is to share out the content :)

If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https://www.subscribestar.com/yannick...

Patreon: https://www.patreon.com/yannickilcher

Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq

Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2

Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m

Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n

View Details

reinforcementlearning #deepmind #agi

What's the most promising path to creating Artificial General Intelligence (AGI)? This paper makes the bold claim that a learning agent maximizing its reward in a sufficiently complex environment will necessarily develop intelligence as a by-product, and that Reward Maximization is the best way to move the creation of AGI forward. The paper is a mix of philosophy, engineering, and futurism, and raises many points of discussion.

OUTLINE:

0:00 - Intro & Outline

4:10 - Reward Maximization

10:10 - The Reward-is-Enough Hypothesis

13:15 - Abilities associated with intelligence

16:40 - My Criticism

26:15 - Reward Maximization through Reinforcement Learning

31:30 - Discussion, Conclusion & My Comments

Paper: https://www.sciencedirect.com/science...

Abstract:

In this article we hypothesise that intelligence, and its associated abilities, can be understood as subserving the maximisation of reward. Accordingly, reward is enough to drive behaviour that exhibits abilities studied in natural and artificial intelligence, including knowledge, learning, perception, social intelligence, language, generalisation and imitation. This is in contrast to the view that specialised problem formulations are needed for each ability, based on other signals or objectives. Furthermore, we suggest that agents that learn through trial and error experience to maximise reward could learn behaviour that exhibits most if not all of these abilities, and therefore that powerful reinforcement learning agents could constitute a solution to artificial general intelligence.

Authors: David Silver, Satinder Singh, Doina Precup, Richard S. Sutton

Links:

TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://discord.gg/4H8xxDF

BitChute: https://www.bitchute.com/channel/yann...

Minds: https://www.minds.com/ykilcher

Parler: https://parler.com/profile/YannicKilcher

LinkedIn: https://www.linkedin.com/in/yannic-ki...

BiliBili: https://space.bilibili.com/1824646584

If you want to support me, the best thing to do is to share out the content :)

If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https://www.subscribestar.com/yannick...

Patreon: https://www.patreon.com/yannickilcher

Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq

Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2

Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m

Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n

View Details

expirespan #nlp #facebookai

Facebook AI (FAIR) researchers present Expire-Span, a variant of Transformer XL that dynamically assigns expiration dates to previously encountered signals. Because of this, Expire-Span can handle sequences of many thousand tokens, while keeping the memory and compute requirements at a manageable level. It severely matches or outperforms baseline systems, while consuming much less resources. We discuss its architecture, advantages, and shortcomings.

OUTLINE:

0:00 - Intro & Overview

2:30 - Remembering the past in sequence models

5:45 - Learning to expire past memories

8:30 - Difference to local attention

10:00 - Architecture overview

13:45 - Comparison to Transformer XL

18:50 - Predicting expiration masks

32:30 - Experimental Results

40:00 - Conclusion & Comments

Paper: https://arxiv.org/abs/2105.06548

Code: https://github.com/facebookresearch/t...

ADDENDUM: I mention several times that the gradient signal of the e quantity only occurs inside the R ramp. By that, I mean the gradient stemming from the model loss. The regularization loss acts also outside the R ramp.

Abstract:

Attention mechanisms have shown promising results in sequence modeling tasks that require long-term memory. Recent work investigated mechanisms to reduce the computational cost of preserving and storing memories. However, not all content in the past is equally important to remember. We propose Expire-Span, a method that learns to retain the most important information and expire the irrelevant information. This forgetting of memories enables Transformers to scale to attend over tens of thousands of previous timesteps efficiently, as not all states from previous timesteps are preserved. We demonstrate that Expire-Span can help models identify and retain critical information and show it can achieve strong performance on reinforcement learning tasks specifically designed to challenge this functionality. Next, we show that Expire-Span can scale to memories that are tens of thousands in size, setting a new state of the art on incredibly long context tasks such as character-level language modeling and a frame-by-frame moving objects task. Finally, we analyze the efficiency of Expire-Span compared to existing approaches and demonstrate that it trains faster and uses less memory.

Authors: Sainbayar Sukhbaatar, Da Ju, Spencer Poff, Stephen Roller, Arthur Szlam, Jason Weston, Angela Fan

Links:

TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://discord.gg/4H8xxDF

BitChute: https://www.bitchute.com/channel/yann...

Minds: https://www.minds.com/ykilcher

Parler: https://parler.com/profile/YannicKilcher

LinkedIn: https://www.linkedin.com/in/yannic-ki...

BiliBili: https://space.bilibili.com/1824646584

If you want to support me, the best thing to do is to share out the content :)

If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https://www.subscribestar.com/yannick...

Patreon: https://www.patreon.com/yannickilcher

Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq

Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2

Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m

Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n

View Details

fnet #attention #fourier

Do we even need Attention? FNets completely drop the Attention mechanism in favor of a simple Fourier transform. They perform almost as well as Transformers, while drastically reducing parameter count, as well as compute and memory requirements. This highlights that a good token mixing heuristic could be as valuable as a learned attention matrix.

OUTLINE:

0:00 - Intro & Overview

0:45 - Giving up on Attention

5:00 - FNet Architecture

9:00 - Going deeper into the Fourier Transform

11:20 - The Importance of Mixing

22:20 - Experimental Results

33:00 - Conclusions & Comments

Paper: https://arxiv.org/abs/2105.03824

ADDENDUM:

Of course, I completely forgot to discuss the connection between Fourier transforms and Convolutions, and that this might be interpreted as convolutions with very large kernels.

Abstract:

We show that Transformer encoder architectures can be massively sped up, with limited accuracy costs, by replacing the self-attention sublayers with simple linear transformations that "mix" input tokens. These linear transformations, along with simple nonlinearities in feed-forward layers, are sufficient to model semantic relationships in several text classification tasks. Perhaps most surprisingly, we find that replacing the self-attention sublayer in a Transformer encoder with a standard, unparameterized Fourier Transform achieves 92% of the accuracy of BERT on the GLUE benchmark, but pre-trains and runs up to seven times faster on GPUs and twice as fast on TPUs. The resulting model, which we name FNet, scales very efficiently to long inputs, matching the accuracy of the most accurate "efficient" Transformers on the Long Range Arena benchmark, but training and running faster across all sequence lengths on GPUs and relatively shorter sequence lengths on TPUs. Finally, FNet has a light memory footprint and is particularly efficient at smaller model sizes: for a fixed speed and accuracy budget, small FNet models outperform Transformer counterparts.

Authors: James Lee-Thorp, Joshua Ainslie, Ilya Eckstein, Santiago Ontanon

Links:

TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://discord.gg/4H8xxDF

BitChute: https://www.bitchute.com/channel/yann...

Minds: https://www.minds.com/ykilcher

Parler: https://parler.com/profile/YannicKilcher

LinkedIn: https://www.linkedin.com/in/yannic-ki...

BiliBili: https://space.bilibili.com/1824646584

If you want to support me, the best thing to do is to share out the content :)

If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https://www.subscribestar.com/yannick...

Patreon: https://www.patreon.com/yannickilcher

Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq

Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2

Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m

Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n

View Details

artificialintelligence #musicvideo #clip

I used OpenAI's CLIP model and BigGAN to create a music video that goes along with the lyrics of a song that I wrote. The song lyrics are made from ImageNet class labels, and the song itself is performed by me on a looper.

OUTLINE:

0:00 - Intro

1:00 - AI-generated music video for "be my weasel"

3:50 - How it was made

7:30 - My looping gear

9:35 - AI-generated music video #2

12:45 - Outro & Credits

Code and references: https://github.com/yk/clip_music_video

Links:

TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://discord.gg/4H8xxDF

BitChute: https://www.bitchute.com/channel/yann...

Minds: https://www.minds.com/ykilcher

Parler: https://parler.com/profile/YannicKilcher

LinkedIn: https://www.linkedin.com/in/yannic-ki...

BiliBili: https://space.bilibili.com/1824646584

If you want to support me, the best thing to do is to share out the content :)

If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https://www.subscribestar.com/yannick...

Patreon: https://www.patreon.com/yannickilcher

Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq

Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2

Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m

Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n

View Details

ddpm #diffusionmodels #openai

GANs have dominated the image generation space for the majority of the last decade. This paper shows for the first time, how a non-GAN model, a DDPM, can be improved to overtake GANs at standard evaluation metrics for image generation. The produced samples look amazing and other than GANs, the new model has a formal probabilistic foundation. Is there a future for GANs or are Diffusion Models going to overtake them for good?

OUTLINE:

0:00 - Intro & Overview

4:10 - Denoising Diffusion Probabilistic Models

11:30 - Formal derivation of the training loss

23:00 - Training in practice

27:55 - Learning the covariance

31:25 - Improving the noise schedule

33:35 - Reducing the loss gradient noise

40:35 - Classifier guidance

52:50 - Experimental Results

Paper (this): https://arxiv.org/abs/2105.05233

Paper (previous): https://arxiv.org/abs/2102.09672

Code: https://github.com/openai/guided-diff...

Abstract:

We show that diffusion models can achieve image sample quality superior to the current state-of-the-art generative models. We achieve this on unconditional image synthesis by finding a better architecture through a series of ablations. For conditional image synthesis, we further improve sample quality with classifier guidance: a simple, compute-efficient method for trading off diversity for sample quality using gradients from a classifier. We achieve an FID of 2.97 on ImageNet 128×128, 4.59 on ImageNet 256×256, and 7.72 on ImageNet 512×512, and we match BigGAN-deep even with as few as 25 forward passes per sample, all while maintaining better coverage of the distribution. Finally, we find that classifier guidance combines well with upsampling diffusion models, further improving FID to 3.85 on ImageNet 512×512. We release our code at this https URL

Authors: Alex Nichol, Prafulla Dhariwal

Links:

TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://discord.gg/4H8xxDF

BitChute: https://www.bitchute.com/channel/yann...

Minds: https://www.minds.com/ykilcher

Parler: https://parler.com/profile/YannicKilcher

LinkedIn: https://www.linkedin.com/in/yannic-ki...

BiliBili: https://space.bilibili.com/1824646584

If you want to support me, the best thing to do is to share out the content :)

If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https://www.subscribestar.com/yannick...

Patreon: https://www.patreon.com/yannickilcher

Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq

Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2

Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m

Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n

View Details

involution​ #computervision​ #attention​

Convolutional Neural Networks (CNNs) have dominated computer vision for almost a decade by applying two fundamental principles: Spatial agnosticism and channel-specific computations. Involution aims to invert these principles and presents a spatial-specific computation, which is also channel-agnostic. The resulting Involution Operator and RedNet architecture are a compromise between classic Convolutions and the newer Local Self-Attention architectures and perform favorably in terms of computation accuracy tradeoff when compared to either.

OUTLINE:

0:00​ - Intro & Overview

3:00​ - Principles of Convolution

10:50​ - Towards spatial-specific computations

17:00​ - The Involution Operator

20:00​ - Comparison to Self-Attention

25:15​ - Experimental Results

30:30​ - Comments & Conclusion

Paper: https://arxiv.org/abs/2103.06255​

Code: https://github.com/d-li14/involution​

Abstract:

Convolution has been the core ingredient of modern neural networks, triggering the surge of deep learning in vision. In this work, we rethink the inherent principles of standard convolution for vision tasks, specifically spatial-agnostic and channel-specific. Instead, we present a novel atomic operation for deep neural networks by inverting the aforementioned design principles of convolution, coined as involution. We additionally demystify the recent popular self-attention operator and subsume it into our involution family as an over-complicated instantiation. The proposed involution operator could be leveraged as fundamental bricks to build the new generation of neural networks for visual recognition, powering different deep learning models on several prevalent benchmarks, including ImageNet classification, COCO detection and segmentation, together with Cityscapes segmentation. Our involution-based models improve the performance of convolutional baselines using ResNet-50 by up to 1.6% top-1 accuracy, 2.5% and 2.4% bounding box AP, and 4.7% mean IoU absolutely while compressing the computational cost to 66%, 65%, 72%, and 57% on the above benchmarks, respectively. Code and pre-trained models for all the tasks are available at this https URL.

Authors: Duo Li, Jie Hu, Changhu Wang, Xiangtai Li, Qi She, Lei Zhu, Tong Zhang, Qifeng Chen

Links:

TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick​

YouTube: https://www.youtube.com/c/yannickilcher​

Twitter: https://twitter.com/ykilcher​

Discord: https://discord.gg/4H8xxDF​

BitChute: https://www.bitchute.com/channel/yann...​

Minds: https://www.minds.com/ykilcher​

Parler: https://parler.com/profile/YannicKilcher​

LinkedIn: https://www.linkedin.com/in/yannic-ki...​

BiliBili: https://space.bilibili.com/1824646584​

If you want to support me, the best thing to do is to share out the content :)

If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https://www.subscribestar.com/yannick...​

Patreon: https://www.patreon.com/yannickilcher​

Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq

Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2

Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m

Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n

View Details

mixer​ #google​ #imagenet​

Convolutional Neural Networks have dominated computer vision for nearly 10 years, and that might finally come to an end. First, Vision Transformers (ViT) have shown remarkable performance, and now even simple MLP-based models reach competitive accuracy, as long as sufficient data is used for pre-training. This paper presents MLP-Mixer, using MLPs in a particular weight-sharing arrangement to achieve a competitive, high-throughput model and it raises some interesting questions about the nature of learning and inductive biases and their interaction with scale for future research.

OUTLINE:

0:00​ - Intro & Overview

2:20​ - MLP-Mixer Architecture

13:20​ - Experimental Results

17:30​ - Effects of Scale

24:30​ - Learned Weights Visualization

27:25​ - Comments & Conclusion

Paper: https://arxiv.org/abs/2105.01601​

Abstract:

Convolutional Neural Networks (CNNs) are the go-to model for computer vision. Recently, attention-based networks, such as the Vision Transformer, have also become popular. In this paper we show that while convolutions and attention are both sufficient for good performance, neither of them are necessary. We present MLP-Mixer, an architecture based exclusively on multi-layer perceptrons (MLPs). MLP-Mixer contains two types of layers: one with MLPs applied independently to image patches (i.e. "mixing" the per-location features), and one with MLPs applied across patches (i.e. "mixing" spatial information). When trained on large datasets, or with modern regularization schemes, MLP-Mixer attains competitive scores on image classification benchmarks, with pre-training and inference cost comparable to state-of-the-art models. We hope that these results spark further research beyond the realms of well established CNNs and Transformers.

Authors: Ilya Tolstikhin, Neil Houlsby, Alexander Kolesnikov, Lucas Beyer, Xiaohua Zhai, Thomas Unterthiner, Jessica Yung, Daniel Keysers, Jakob Uszkoreit, Mario Lucic, Alexey Dosovitskiy

ERRATA: Here is their definition of what the 5-shot classifier is: "we report the few-shot accuracies obtained by solving the L2-regularized linear regression problem between the frozen learned representations of images and the labels"

Links:

TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick​

YouTube: https://www.youtube.com/c/yannickilcher​

Twitter: https://twitter.com/ykilcher​

Discord: https://discord.gg/4H8xxDF​

BitChute: https://www.bitchute.com/channel/yann...​

Minds: https://www.minds.com/ykilcher​

Parler: https://parler.com/profile/YannicKilcher​

LinkedIn: https://www.linkedin.com/in/yannic-ki...​

BiliBili: https://space.bilibili.com/1824646584​

If you want to support me, the best thing to do is to share out the content :)

If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https://www.subscribestar.com/yannick...​

Patreon: https://www.patreon.com/yannickilcher​

Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq

Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2

Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m

Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n

View Details

genderbias​ #algorithmicfairness​ #debiasing​

A brief look into gender stereotypes in Google Translate. The origin is a Tweet containing a Hungarian text. Hungarian is a gender-neutral language, so translating gender pronouns is ambiguous. Turns out that Google Translate assigns very stereotypical pronouns. In this video, we'll have a look at the origins and possible solutions to this problem.

OUTLINE:

0:00​ - Intro

1:10​ - Digging Deeper

2:30​ - How does Machine Translation work?

3:50​ - Training Data Problems

4:40​ - Learning Algorithm Problems

5:45​ - Argmax Output Problems

6:45​ - Pragmatics

7:50​ - More on Google Translate

9:40​ - Social Engineering

11:15​ - Conclusion

Songs:

Like That - Anno Domini Beats

Submarine - Dyalla

Dude - Patrick Patrikios

Links:

TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick​

YouTube: https://www.youtube.com/c/yannickilcher​

Twitter: https://twitter.com/ykilcher​

Discord: https://discord.gg/4H8xxDF​

BitChute: https://www.bitchute.com/channel/yann...​

Minds: https://www.minds.com/ykilcher​

Parler: https://parler.com/profile/YannicKilcher​

LinkedIn: https://www.linkedin.com/in/yannic-ki...​

BiliBili: https://space.bilibili.com/1824646584​

If you want to support me, the best thing to do is to share out the content :)

If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https://www.subscribestar.com/yannick...​

Patreon: https://www.patreon.com/yannickilcher​

Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq

Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2

Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m

Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n

View Details

perceiver​ #deepmind​ #transformer​

Inspired by the fact that biological creatures attend to multiple modalities at the same time, DeepMind releases its new Perceiver model. Based on the Transformer architecture, the Perceiver makes no assumptions on the modality of the input data and also solves the long-standing quadratic bottleneck problem. This is achieved by having a latent low-dimensional Transformer, where the input data is fed multiple times via cross-attention. The Perceiver's weights can also be shared across layers, making it very similar to an RNN. Perceivers achieve competitive performance on ImageNet and state-of-the-art on other modalities, all while making no architectural adjustments to input data.

OUTLINE:

0:00​ - Intro & Overview

2:20​ - Built-In assumptions of Computer Vision Models

5:10​ - The Quadratic Bottleneck of Transformers

8:00​ - Cross-Attention in Transformers

10:45​ - The Perceiver Model Architecture & Learned Queries

20:05​ - Positional Encodings via Fourier Features

23:25​ - Experimental Results & Attention Maps

29:05​ - Comments & Conclusion

Paper: https://arxiv.org/abs/2103.03206​

My Video on Transformers (Attention is All You Need): https://youtu.be/iDulhoQ2pro​

Abstract:

Biological systems understand the world by simultaneously processing high-dimensional inputs from modalities as diverse as vision, audition, touch, proprioception, etc. The perception models used in deep learning on the other hand are designed for individual modalities, often relying on domain-specific assumptions such as the local grid structures exploited by virtually all existing vision models. These priors introduce helpful inductive biases, but also lock models to individual modalities. In this paper we introduce the Perceiver - a model that builds upon Transformers and hence makes few architectural assumptions about the relationship between its inputs, but that also scales to hundreds of thousands of inputs, like ConvNets. The model leverages an asymmetric attention mechanism to iteratively distill inputs into a tight latent bottleneck, allowing it to scale to handle very large inputs. We show that this architecture performs competitively or beyond strong, specialized models on classification tasks across various modalities: images, point clouds, audio, video and video+audio. The Perceiver obtains performance comparable to ResNet-50 on ImageNet without convolutions and by directly attending to 50,000 pixels. It also surpasses state-of-the-art results for all modalities in AudioSet.

Authors: Andrew Jaegle, Felix Gimeno, Andrew Brock, Andrew Zisserman, Oriol Vinyals, Joao Carreira

Links:

TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick​

YouTube: https://www.youtube.com/c/yannickilcher​

Twitter: https://twitter.com/ykilcher​

Discord: https://discord.gg/4H8xxDF​

BitChute: https://www.bitchute.com/channel/yann...​

Minds: https://www.minds.com/ykilcher​

Parler: https://parler.com/profile/YannicKilcher​

LinkedIn: https://www.linkedin.com/in/yannic-ki...​

BiliBili: https://space.bilibili.com/1824646584​

If you want to support me, the best thing to do is to share out the content :)

If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https://www.subscribestar.com/yannick...​

Patreon: https://www.patreon.com/yannickilcher​

Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq

Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2

Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m

View Details

universalcomputation​ #pretrainedtransformers​ #finetuning​

Large-scale pre-training and subsequent fine-tuning is a common recipe for success with transformer models in machine learning. However, most such transfer learning is done when a model is pre-trained on the same or a very similar modality to the final task to be solved. This paper demonstrates that transformers can be fine-tuned to completely different modalities, such as from language to vision. Moreover, they demonstrate that this can be done by freezing all attention layers, tuning less than .1% of all parameters. The paper further claims that language modeling is a superior pre-training task for such cross-domain transfer. The paper goes through various ablation studies to make its point.

OUTLINE:

0:00​ - Intro & Overview

2:00​ - Frozen Pretrained Transformers

4:50​ - Evaluated Tasks

10:05​ - The Importance of Training LayerNorm

17:10​ - Modality Transfer

25:10​ - Network Architecture Ablation

26:10​ - Evaluation of the Attention Mask

27:20​ - Are FPTs Overfitting or Underfitting?

28:20​ - Model Size Ablation

28:50​ - Is Initialization All You Need?

31:40​ - Full Model Training Overfits

32:15​ - Again the Importance of Training LayerNorm

33:10​ - Conclusions & Comments

Paper: https://arxiv.org/abs/2103.05247​

Code: https://github.com/kzl/universal-comp...​

Abstract:

We investigate the capability of a transformer pretrained on natural language to generalize to other modalities with minimal finetuning -- in particular, without finetuning of the self-attention and feedforward layers of the residual blocks. We consider such a model, which we call a Frozen Pretrained Transformer (FPT), and study finetuning it on a variety of sequence classification tasks spanning numerical computation, vision, and protein fold prediction. In contrast to prior works which investigate finetuning on the same modality as the pretraining dataset, we show that pretraining on natural language improves performance and compute efficiency on non-language downstream tasks. In particular, we find that such pretraining enables FPT to generalize in zero-shot to these modalities, matching the performance of a transformer fully trained on these tasks.

Authors: Kevin Lu, Aditya Grover, Pieter Abbeel, Igor Mordatch

Links:

TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick​

YouTube: https://www.youtube.com/c/yannickilcher​

Twitter: https://twitter.com/ykilcher​

Discord: https://discord.gg/4H8xxDF​

BitChute: https://www.bitchute.com/channel/yann...​

Minds: https://www.minds.com/ykilcher​

Parler: https://parler.com/profile/YannicKilcher​

LinkedIn: https://www.linkedin.com/in/yannic-ki...​

BiliBili: https://space.bilibili.com/1824646584​

If you want to support me, the best thing to do is to share out the content :)

If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https://www.subscribestar.com/yannick...​

Patreon: https://www.patreon.com/yannickilcher​

Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq

Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2

Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m

Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n

View Details

selfsupervisedlearning​ #yannlecun​ #facebookai​

Deep Learning systems can achieve remarkable, even super-human performance through supervised learning on large, labeled datasets. However, there are two problems: First, collecting ever more labeled data is expensive in both time and money. Second, these deep neural networks will be high performers on their task, but cannot easily generalize to other, related tasks, or they need large amounts of data to do so. In this blog post, Yann LeCun and Ishan Misra of Facebook AI Research (FAIR) describe the current state of Self-Supervised Learning (SSL) and argue that it is the next step in the development of AI that uses fewer labels and can transfer knowledge faster than current systems. They suggest as a promising direction to build non-contrastive latent-variable predictive models, like VAEs, but ones that also provide high-quality latent representations for downstream tasks.

OUTLINE:

0:00​ - Intro & Overview

1:15​ - Supervised Learning, Self-Supervised Learning, and Common Sense

7:35​ - Predicting Hidden Parts from Observed Parts

17:50​ - Self-Supervised Learning for Language vs Vision

26:50​ - Energy-Based Models

30:15​ - Joint-Embedding Models

35:45​ - Contrastive Methods

43:45​ - Latent-Variable Predictive Models and GANs

55:00​ - Summary & Conclusion

Paper (Blog Post): https://ai.facebook.com/blog/self-sup...​

My Video on BYOL: https://www.youtube.com/watch?v=YPfUi...​

ERRATA:

  • The difference between loss and energy: Energy is for inference, loss is for training.

  • The R(z) term is a regularizer that restricts the capacity of the latent variable. I think I said both of those things, but never together.

  • The way I explain why BERT is contrastive is wrong. I haven't figured out why just yet, though :)

Video approved by Antonio.

Abstract:

We believe that self-supervised learning (SSL) is one of the most promising ways to build such background knowledge and approximate a form of common sense in AI systems.

Authors: Yann LeCun, Ishan Misra

Links:

TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick​

YouTube: https://www.youtube.com/c/yannickilcher​

Twitter: https://twitter.com/ykilcher​

Discord: https://discord.gg/4H8xxDF​

BitChute: https://www.bitchute.com/channel/yann...​

Minds: https://www.minds.com/ykilcher​

Parler: https://parler.com/profile/YannicKilcher​

LinkedIn: https://www.linkedin.com/in/yannic-ki...​

BiliBili: https://space.bilibili.com/1824646584​

If you want to support me, the best thing to do is to share out the content :)

If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https://www.subscribestar.com/yannick...​

Patreon: https://www.patreon.com/yannickilcher​

Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq

Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2

Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m

Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n

View Details

openai​ #clip​ #microscope​

OpenAI does a huge investigation into the inner workings of their recent CLIP model via faceted feature visualization and finds amazing things: Some neurons in the last layer respond to distinct concepts across multiple modalities, meaning they fire for photographs, drawings, and signs depicting the same concept, even when the images are vastly distinct. Through manual examination, they identify and investigate neurons corresponding to persons, geographical regions, religions, emotions, and much more. In this video, I go through the publication and then I present my own findings from digging around in the OpenAI Microscope.

OUTLINE:

0:00​ - Intro & Overview

3:35​ - OpenAI Microscope

7:10​ - Categories of found neurons

11:10​ - Person Neurons

13:00​ - Donald Trump Neuron

17:15​ - Emotion Neurons

22:45​ - Region Neurons

26:40​ - Sparse Mixture of Emotions

28:05​ - Emotion Atlas

29:45​ - Adversarial Typographic Attacks

31:55​ - Stroop Test

33:10​ - My Findings in OpenAI Microscope

33:30​ - Superman Neuron

33:50​ - Resting B*tchface Neuron

34:10​ - Trash Bag Neuron

35:25​ - God Weightlifting Neuron

36:40​ - Organ Neuron

38:35​ - Film Spool Neuron

39:05​ - Feather Neuron

39:20​ - Spartan Neuron

40:25​ - Letter E Neuron

40:35​ - Cleanin Neuron

40:45​ - Frown Neuron

40:55​ - Lion Neuron

41:05​ - Fashion Model Neuron

41:20​ - Baseball Neuron

41:50​ - Bride Neuron

42:00​ - Navy Neuron

42:30​ - Hemp Neuron

43:25​ - Staircase Neuron

43:45​ - Disney Neuron

44:15​ - Hillary Clinton Neuron

44:50​ - God Neuron

45:15​ - Blurry Neuron

45:35​ - Arrow Neuron

45:55​ - Trophy Presentation Neuron

46:10​ - Receding Hairline Neuron

46:30​ - Traffic Neuron

46:40​ - Raised Hand Neuron

46:50​ - Google Maps Neuron

47:15​ - Nervous Smile Neuron

47:30​ - Elvis Neuron

47:55​ - The Flash Neuron

48:05​ - Beard Neuron

48:15​ - Kilt Neuron

48:25​ - Rainy Neuron

48:35​ - Electricity Neuron

48:50​ - Droplets Neuron

49:00​ - Escape Neuron

49:25​ - King Neuron

49:35​ - Country Neuron

49:45​ - Overweight Men Neuron

49:55​ - Wedding

50:05​ - Australia Neuron

50:15​ - Yawn Neuron

50:30​ - Bees & Simpsons Neuron

50:40​ - Mussles Neuron

50:50​ - Spice Neuron

51:00​ - Conclusion

Paper: https://distill.pub/2021/multimodal-n...​

My Findings: https://www.notion.so/CLIP-OpenAI-Mic...​

My Video on CLIP: https://youtu.be/T9XSU0pKX2E​

My Video on Feature Visualizations & The OpenAI Microscope: https://youtu.be/Ok44otx90D4​

Abstract:

In 2005, a letter published in Nature described human neurons responding to specific people, such as Jennifer Aniston or Halle Berry. The exciting thing wasn’t just that they selected for particular people, but that they did so regardless of whether they were shown photographs, drawings, or even images of the person’s name. The neurons were multimodal. As the lead author would put it: "You are looking at the far end of the transformation from metric, visual shapes to conceptual... information." We report the existence of similar multimodal neurons in artificial neural networks. This includes neurons selecting for prominent public figures or fictional characters, such as Lady Gaga or Spiderman. Like the biological multimodal neurons, these artificial neurons respond to the same subject in photographs, drawings, and images of their name.

Authors: Gabriel Goh, Nick Cammarata, Chelsea Voss, Shan Carter, Michael Petrov, Ludwig Schubert, Alec Radford, Chris Olah

View Details

machinelearning​ #phd​ #howto​

This video is advice for new PhD students in the field of Machine Learning in 2021 and after. The field has shifted dramatically in the last few years and navigating grad school can be very hard, especially when you're as clueless as I was when I started. The video is a personal recount of my mistakes and what I've learned from them. If you already have several published papers and know what to do, this video is not for you. However, if you are not even sure where to start, how to select a topic, or what goes in a paper, you might benefit from this video, because that's exactly how I felt.

Main Takeaways:

  • Select niche topics rather than hype topics

  • Write papers that can't be rejected

  • Don't be discouraged by bad reviews

  • Take reviewing & teaching seriously

  • Keep up your focus

  • Conferences are for networking

  • Internships are great opportunities

  • Team up with complementary skills

  • Don't work too hard

OUTLINE:

0:00​ - Intro & Overview

1:25​ - Thesis Topic Selection

4:25​ - How To Publish Papers

5:35​ - Dealing With Reviewers

6:30​ - How To Be A Reviewer

7:40​ - Take Teaching Seriously

8:30​ - Maintain Focus

10:20​ - Navigating Conferences

12:40​ - Internships

13:40​ - Collaborations

14:55​ - Don't Forget To Enjoy

Transcript: https://www.notion.so/Yannic-Kilcher-...​

Credits to Lanz for editing

Links:

TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick​

YouTube: https://www.youtube.com/c/yannickilcher​

Twitter: https://twitter.com/ykilcher​

Discord: https://discord.gg/4H8xxDF​

BitChute: https://www.bitchute.com/channel/yann...​

Minds: https://www.minds.com/ykilcher​

Parler: https://parler.com/profile/YannicKilcher​

LinkedIn: https://www.linkedin.com/in/yannic-ki...​

BiliBili: https://space.bilibili.com/1824646584​

If you want to support me, the best thing to do is to share out the content :)

If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https://www.subscribestar.com/yannick...​

Patreon: https://www.patreon.com/yannickilcher​

Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq

Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2

Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m

Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n

View Details

In the recurring debate about bias in Machine Learning models, there is a growing argument saying that "the problem is not in the data", often citing the influence of various choices like loss functions or network architecture. In this video, we take a look at PAIR's AI Explorables through the lens of whether or not the bias problem is a data problem.

OUTLINE:

0:00​ - Intro & Overview

1:45​ - Recap: Bias in ML

4:25​ - AI Explorables

5:40​ - Measuring Fairness Explorable

11:00​ - Hidden Bias Explorable

16:10​ - Measuring Diversity Explorable

23:00​ - Conclusion & Comments

AI Explorables: https://pair.withgoogle.com/explorables/​

Links:

TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick​

YouTube: https://www.youtube.com/c/yannickilcher​

Twitter: https://twitter.com/ykilcher​

Discord: https://discord.gg/4H8xxDF​

BitChute: https://www.bitchute.com/channel/yann...​

Minds: https://www.minds.com/ykilcher​

Parler: https://parler.com/profile/YannicKilcher​

LinkedIn: https://www.linkedin.com/in/yannic-ki...​

BiliBili: https://space.bilibili.com/1824646584​

If you want to support me, the best thing to do is to share out the content :)

If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https://www.subscribestar.com/yannick...​

Patreon: https://www.patreon.com/yannickilcher​

Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq

Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2

Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m

Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n

View Details

dreamcoder​ #programsynthesis​ #symbolicreasoning​

Classic Machine Learning struggles with few-shot generalization for tasks where humans can easily generalize from just a handful of examples, for example sorting a list of numbers. Humans do this by coming up with a short program, or algorithm, that explains the few data points in a compact way. DreamCoder emulates this by using neural guided search over a language of primitives, a library, that it builds up over time. By doing this, it can iteratively construct more and more complex programs by building on its own abstractions and therefore solve more and more difficult tasks in a few-shot manner by generating very short programs that solve the few given datapoints. The resulting system can not only generalize quickly but also delivers an explainable solution to its problems in form of a modular and hierarchical learned library. Combining this with classic Deep Learning for low-level perception is a very promising future direction.

OUTLINE:

0:00​ - Intro & Overview

4:55​ - DreamCoder System Architecture

9:00​ - Wake Phase: Neural Guided Search

19:15​ - Abstraction Phase: Extending the Internal Library

24:30​ - Dreaming Phase: Training Neural Search on Fictional Programs and Replays

30:55​ - Abstraction by Compressing Program Refactorings

32:40​ - Experimental Results on LOGO Drawings

39:00​ - Ablation Studies

39:50​ - Re-Discovering Physical Laws

42:25​ - Discovering Recursive Programming Algorithms

44:20​ - Conclusions & Discussion

Paper: https://arxiv.org/abs/2006.08381​

Code: https://github.com/ellisk42/ec​

Abstract:

Expert problem-solving is driven by powerful languages for thinking about problems and their solutions. Acquiring expertise means learning these languages -- systems of concepts, alongside the skills to use them. We present DreamCoder, a system that learns to solve problems by writing programs. It builds expertise by creating programming languages for expressing domain concepts, together with neural networks to guide the search for programs within these languages. A ``wake-sleep'' learning algorithm alternately extends the language with new symbolic abstractions and trains the neural network on imagined and replayed problems. DreamCoder solves both classic inductive programming tasks and creative tasks such as drawing pictures and building scenes. It rediscovers the basics of modern functional programming, vector algebra and classical physics, including Newton's and Coulomb's laws. Concepts are built compositionally from those learned earlier, yielding multi-layered symbolic representations that are interpretable and transferrable to new tasks, while still growing scalably and flexibly with experience.

Authors: Kevin Ellis, Catherine Wong, Maxwell Nye, Mathias Sable-Meyer, Luc Cary, Lucas Morales, Luke Hewitt, Armando Solar-Lezama, Joshua B. Tenenbaum

Links:

TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick​

YouTube: https://www.youtube.com/c/yannickilcher​

Twitter: https://twitter.com/ykilcher​

Discord: https://discord.gg/4H8xxDF​

BitChute: https://www.bitchute.com/channel/yann...​

Minds: https://www.minds.com/ykilcher​

Parler: https://parler.com/profile/YannicKilcher​

LinkedIn: https://www.linkedin.com/in/yannic-ki...​

BiliBili: https://space.bilibili.com/1824646584​

If you want to support me, the best thing to do is to share out the content :)

If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https://www.subscribestar.com/yannick...​

Patreon: https://www.patreon.com/yannickilcher

View Details

nerf​ #neuralrendering​ #deeplearning​

View Synthesis is a tricky problem, especially when only given a sparse set of images as an input. NeRF embeds an entire scene into the weights of a feedforward neural network, trained by backpropagation through a differential volume rendering procedure, and achieves state-of-the-art view synthesis. It includes directional dependence and is able to capture fine structural details, as well as reflection effects and transparency.

OUTLINE:

0:00​ - Intro & Overview

4:50​ - View Synthesis Task Description

5:50​ - The fundamental difference to classic Deep Learning

7:00​ - NeRF Core Concept

15:30​ - Training the NeRF from sparse views

20:50​ - Radiance Field Volume Rendering

23:20​ - Resulting View Dependence

24:00​ - Positional Encoding

28:00​ - Hierarchical Volume Sampling

30:15​ - Experimental Results

33:30​ - Comments & Conclusion

Paper: https://arxiv.org/abs/2003.08934​

Website & Code: https://www.matthewtancik.com/nerf​

My Video on SIREN: https://youtu.be/Q5g3p9Zwjrk​

Abstract:

We present a method that achieves state-of-the-art results for synthesizing novel views of complex scenes by optimizing an underlying continuous volumetric scene function using a sparse set of input views. Our algorithm represents a scene using a fully-connected (non-convolutional) deep network, whose input is a single continuous 5D coordinate (spatial location (x,y,z) and viewing direction (θ,ϕ)) and whose output is the volume density and view-dependent emitted radiance at that spatial location. We synthesize views by querying 5D coordinates along camera rays and use classic volume rendering techniques to project the output colors and densities into an image. Because volume rendering is naturally differentiable, the only input required to optimize our representation is a set of images with known camera poses. We describe how to effectively optimize neural radiance fields to render photorealistic novel views of scenes with complicated geometry and appearance, and demonstrate results that outperform prior work on neural rendering and view synthesis. View synthesis results are best viewed as videos, so we urge readers to view our supplementary video for convincing comparisons.

Authors: Ben Mildenhall, Pratul P. Srinivasan, Matthew Tancik, Jonathan T. Barron, Ravi Ramamoorthi, Ren Ng

Links:

TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick​

YouTube: https://www.youtube.com/c/yannickilcher​

Twitter: https://twitter.com/ykilcher​

Discord: https://discord.gg/4H8xxDF​

BitChute: https://www.bitchute.com/channel/yann...​

Minds: https://www.minds.com/ykilcher​

Parler: https://parler.com/profile/YannicKilcher​

LinkedIn: https://www.linkedin.com/in/yannic-ki...​

BiliBili: https://space.bilibili.com/1824646584​

If you want to support me, the best thing to do is to share out the content :)

If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https://www.subscribestar.com/yannick...​

Patreon: https://www.patreon.com/yannickilcher​

Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq

Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2

Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m

Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n

View Details

dino #facebook #selfsupervised

Self-Supervised Learning is the final frontier in Representation Learning: Getting useful features without any labels. Facebook AI's new system, DINO, combines advances in Self-Supervised Learning for Computer Vision with the new Vision Transformer (ViT) architecture and achieves impressive results without any labels. Attention maps can be directly interpreted as segmentation maps, and the obtained representations can be used for image retrieval and zero-shot k-nearest neighbor classifiers (KNNs).

OUTLINE:

0:00 - Intro & Overview

6:20 - Vision Transformers

9:20 - Self-Supervised Learning for Images

13:30 - Self-Distillation

15:20 - Building the teacher from the student by moving average

16:45 - DINO Pseudocode

23:10 - Why Cross-Entropy Loss?

28:20 - Experimental Results

33:40 - My Hypothesis why this works

38:45 - Conclusion & Comments

Paper: https://arxiv.org/abs/2104.14294

Blog: https://ai.facebook.com/blog/dino-paws-computer-vision-with-self-supervised-transformers-and-10x-more-efficient-training

Code: https://github.com/facebookresearch/dino

My Video on ViT: https://youtu.be/TrdevFK_am4

My Video on BYOL: https://youtu.be/YPfUiOMYOEE

Abstract:

In this paper, we question if self-supervised learning provides new properties to Vision Transformer (ViT) that stand out compared to convolutional networks (convnets). Beyond the fact that adapting self-supervised methods to this architecture works particularly well, we make the following observations: first, self-supervised ViT features contain explicit information about the semantic segmentation of an image, which does not emerge as clearly with supervised ViTs, nor with convnets. Second, these features are also excellent k-NN classifiers, reaching 78.3% top-1 on ImageNet with a small ViT. Our study also underlines the importance of momentum encoder, multi-crop training, and the use of small patches with ViTs. We implement our findings into a simple self-supervised method, called DINO, which we interpret as a form of self-distillation with no labels. We show the synergy between DINO and ViTs by achieving 80.1% top-1 on ImageNet in linear evaluation with ViT-Base.

Authors: Mathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou, Julien Mairal, Piotr Bojanowski, Armand Joulin

Links:

TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://discord.gg/4H8xxDF

BitChute: https://www.bitchute.com/channel/yannic-kilcher

Minds: https://www.minds.com/ykilcher

Parler: https://parler.com/profile/YannicKilcher

LinkedIn: https://www.linkedin.com/in/yannic-kilcher-488534136/

BiliBili: https://space.bilibili.com/1824646584

If you want to support me, the best thing to do is to share out the content :)

If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https://www.subscribestar.com/yannickilcher

Patreon: https://www.patreon.com/yannickilcher

Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq

Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2

Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m

Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n

View Details

aiwinter #agi #embodiedcognition

The AI community has gone through regular cycles of AI Springs, where rapid progress gave rise to massive overconfidence, high funding, and overpromise, followed by these promises being unfulfilled, subsequently diving into periods of disenfranchisement and underfunding, called AI Winters. This paper examines the reasons for the repeated periods of overconfidence and identifies four fallacies that people make when they see rapid progress in AI.

OUTLINE:

0:00 - Intro & Overview

2:10 - AI Springs & AI Winters

5:40 - Is the current AI boom overhyped?

15:35 - Fallacy 1: Narrow Intelligence vs General Intelligence

19:40 - Fallacy 2: Hard for humans doesn't mean hard for computers

21:45 - Fallacy 3: How we call things matters

28:15 - Fallacy 4: Embodied Cognition

35:30 - Conclusion & Comments

Paper: https://arxiv.org/abs/2104.12871

My Video on Shortcut Learning: https://youtu.be/D-eg7k8YSfs

Abstract:

Since its beginning in the 1950s, the field of artificial intelligence has cycled several times between periods of optimistic predictions and massive investment ("AI spring") and periods of disappointment, loss of confidence, and reduced funding ("AI winter"). Even with today's seemingly fast pace of AI breakthroughs, the development of long-promised technologies such as self-driving cars, housekeeping robots, and conversational companions has turned out to be much harder than many people expected. One reason for these repeating cycles is our limited understanding of the nature and complexity of intelligence itself. In this paper I describe four fallacies in common assumptions made by AI researchers, which can lead to overconfident predictions about the field. I conclude by discussing the open questions spurred by these fallacies, including the age-old challenge of imbuing machines with humanlike common sense.

Authors: Melanie Mitchell

Links:

TabNine Code Completion (Referral): http://bit.ly/tabnine-yannick

YouTube: https://www.youtube.com/c/yannickilcher

Twitter: https://twitter.com/ykilcher

Discord: https://discord.gg/4H8xxDF

BitChute: https://www.bitchute.com/channel/yannic-kilcher

Minds: https://www.minds.com/ykilcher

Parler: https://parler.com/profile/YannicKilcher

LinkedIn: https://www.linkedin.com/in/yannic-kilcher-488534136/

BiliBili: https://space.bilibili.com/1824646584

If you want to support me, the best thing to do is to share out the content :)

If you want to support me financially (completely optional and voluntary, but a lot of people have asked for this):

SubscribeStar: https://www.subscribestar.com/yannickilcher

Patreon: https://www.patreon.com/yannickilcher

Bitcoin (BTC): bc1q49lsw3q325tr58ygf8sudx2dqfguclvngvy2cq

Ethereum (ETH): 0x7ad3513E3B8f66799f507Aa7874b1B0eBC7F85e2

Litecoin (LTC): LQW2TRyKYetVC8WjFkhpPhtpbDM4Vw7r9m

Monero (XMR): 4ACL8AGrEo5hAir8A9CeVrW8pEauWvnp1WnSDZxW7tziCDLhZAGsgzhRQABDnFy8yuM9fWJDviJPHKRjV4FWt19CJZN9D4n