Welcome to this space dedicated to the M2D2 Talks co-organized by Valence Discovery and Mila - Quebec AI Institute.
From applied research papers to open source projects, we're hoping to use these talks to help demystify AI for drug discovery and make the field more accessible for newcomers. M2D2 will bring our vibrant AI & drug discovery communities together and spark new perspectives, provoke discussions, and offer a safe space to share new ideas.
For the best experience, please visit our YouTube channel where slides and video presentations can be referenced.
[DISCLAIMER] - For the full visual experience, we recommend you tune in through our YouTube channel to see the presented slides.
Try datamol.io - the open source toolkit that simplifies molecular processing and featurization workflows for machine learning scientists working in drug discovery: https://datamol.io/
If you enjoyed this talk, consider joining the Molecular Modeling and Drug Discovery (M2D2) talks live.
Also, consider joining the M2D2 Slack.
Abstract: The ability to modulate pathogenic proteins represents a powerful treatment strategy for diseases. Unfortunately, many proteins are considered “undruggable” by small molecules, and are often intrinsically disordered, precluding the usage of structure-based tools for binder design. To address these challenges, we have developed a suite of algorithms that enable the design of target-specific peptides via protein language model embeddings, without the requirement of 3D structures. First, we train a model that leverages ESM-2 embeddings to efficiently select high-affinity peptides from natural protein interaction interfaces. We experimentally fuse model-derived peptides to E3 ubiquitin ligases and identify candidates exhibiting robust degradation of undruggable targets in human cells. Next, we develop a high-accuracy discriminator, based on the CLIP architecture, to prioritize and screen peptides with selectivity to a specified target protein. As input to the discriminator, we create a Gaussian diffusion generator to sample an ESM-2-based latent space, fine-tuned on experimentally-valid peptide sequences. Finally, to enable de novo generation of binding peptides, we train an instance of GPT-2 with protein interacting sequences to enable peptide generation conditioned on target sequence. Our model demonstrates low perplexities across both existing and generated peptide sequences. Together, our work lays the foundation for programmable protein targeting and editing applications.
Speaker: Pranam Chatterjee
Twitter - Prudencio
Twitter - Jonny
Twitter - datamol.io
[DISCLAIMER] - For the full visual experience, we recommend you tune in through our YouTube channel to see the presented slides.
Try datamol.io - the open source toolkit that simplifies molecular processing and featurization workflows for machine learning scientists working in drug discovery: https://datamol.io/
If you enjoyed this talk, consider joining the Molecular Modeling and Drug Discovery (M2D2) talks live.
Also, consider joining the M2D2 Slack.
Abstract: Trade-offs between accuracy and speed have long limited the applications of machine learning interatomic potentials. Recently, E(3)-equivariant architectures have demonstrated leading accuracy, data efficiency, transferability, and simulation stability, but their computational cost and scaling has generally reinforced this trade-off. In particular, the ubiquitous use of message passing architectures has precluded the extension of accessible length- and time-scales with efficient multi-GPU calculations.
In this talk I will discuss Allegro, a strictly local equivariant deep learning interatomic potential designed for parallel scalability and increased computational efficiency that simultaneously exhibits excellent accuracy. After presenting the architecture, I will discuss applications and benchmarks on various materials and chemical systems, including recent demonstrations of scaling to large all-atom biomolecular systems such as solvated proteins and a 44 million atom model of the HIV capsid. Finally, I will summarize the software ecosystem and tooling around Allegro.
Speaker: Albert Musaelian
Twitter - Prudencio
Twitter - Jonny
Twitter - datamol.io
[DISCLAIMER] - For the full visual experience, we recommend you tune in through our YouTube channel to see the presented slides.
Try datamol.io - the open source toolkit that simplifies molecular processing and featurization workflows for machine learning scientists working in drug discovery: https://datamol.io/
If you enjoyed this talk, consider joining the Molecular Modeling and Drug Discovery (M2D2) talks live.
Also, consider joining the M2D2 Slack.
Abstract: Engineered proteins play increasingly essential roles in industries and applications spanning pharmaceuticals, agriculture, specialty chemicals, and fuel. Machine learning could enable an unprecedented level of control in protein engineering for therapeutic and industrial applications. Large self-supervised models pretrained on millions of protein sequences have recently gained popularity in generating embeddings of protein sequences for protein property prediction. However, protein datasets contain information in addition to sequence that can improve model performance. This talk will cover models that use sequences, structures, and biophysical features to predict protein function or to generate functional proteins.
Speaker: Martin Vögele
Twitter - Prudencio
Twitter - Jonny
Twitter - datamol.io
[DISCLAIMER] - For the full visual experience, we recommend you tune in through our YouTube channel to see the presented slides.
Try datamol.io - the open source toolkit that simplifies molecular processing and featurization workflows for machine learning scientists working in drug discovery: https://datamol.io/
If you enjoyed this talk, consider joining the Molecular Modeling and Drug Discovery (M2D2) talks live.
Also, consider joining the M2D2 Slack.
Abstract: Molecular simulations enable the study of biomolecules and their dynamics on an atomistic scale. A common task is to compare several simulation conditions - like mutations or different ligands - to find significant differences and interrelations between them. However, the large amount of data produced for ever larger and more complex systems often renders it difficult to identify the structural features that are relevant for a particular phenomenon. PENSA is a flexible software package that enables a comprehensive and thorough investigation into biomolecular conformational ensembles. It provides a wide variety of featurizations and feature transformations that allow for a complete representation of biomolecules like proteins and nucleic acids, including water and ion cavities within the biomolecular structure, thus avoiding bias that would come with manual selection of features. PENSA implements various methods to systematically compare the distributions of these features across ensembles to find the significant differences between them and identify regions of interest. It also includes a novel approach to quantify the state-specific information between two regions of a biomolecule which allows, e.g., the tracing of information flow to identify signaling pathways. PENSA is a modular open-source library that also comes with convenient tools for loading data and visualizing results in ways that make them quick to process and easy to interpret. This talk will demonstrate its usefulness in real-world examples by showing how it helps to determine molecular mechanisms efficiently.
Speaker: Martin Vögele
Twitter - Prudencio
Twitter - Jonny
Twitter - datamol.io
[DISCLAIMER] - For the full visual experience, we recommend you tune in through our YouTube channel to see the presented slides.
Try datamol.io - the open source toolkit that simplifies molecular processing and featurization workflows for machine learning scientists working in drug discovery: https://datamol.io/
If you enjoyed this talk, consider joining the Molecular Modeling and Drug Discovery (M2D2) talks live.
Also, consider joining the M2D2 Slack.
Abstract: Cryptic pockets, which are absent in ligand-free structures and have the potential to be used as drug targets, are often challenging to access through conventional biomolecular simulations due to their slow motions. To overcome this limitation, we have combined AlphaFold and Markov State modelling (MSM) to accelerate the discovery of cryptic pockets. AlphaFold was used to generate a diverse structural ensemble with open or partially open pockets that can serve as starting points for molecular dynamics simulations which were later stitched together using MSM to predict free energy and kinetics associated with cryptic pocket opening. Our approach explored known cryptic pockets, as well as discovered new cryptic pockets which were absent in PDB. Our study highlighted the power of AlphaFold and MSM to discover novel cryptic pockets which can unlock development of next-gen therapeutics.
Speaker: Stephan Thaler
Twitter - Prudencio
Twitter - Jonny
Twitter - datamol.io
[DISCLAIMER] - For the full visual experience, we recommend you tune in through our YouTube channel to see the presented slides.
Try datamol.io - the open source toolkit that simplifies molecular processing and featurization workflows for machine learning scientists working in drug discovery: https://datamol.io/
If you enjoyed this talk, consider joining the Molecular Modeling and Drug Discovery (M2D2) talks live.
Also, consider joining the M2D2 Slack.
Abstract: Cryptic pockets, which are absent in ligand-free structures and have the potential to be used as drug targets, are often challenging to access through conventional biomolecular simulations due to their slow motions. To overcome this limitation, we have combined AlphaFold and Markov State modelling (MSM) to accelerate the discovery of cryptic pockets. AlphaFold was used to generate a diverse structural ensemble with open or partially open pockets that can serve as starting points for molecular dynamics simulations which were later stitched together using MSM to predict free energy and kinetics associated with cryptic pocket opening. Our approach explored known cryptic pockets, as well as discovered new cryptic pockets which were absent in PDB. Our study highlighted the power of AlphaFold and MSM to discover novel cryptic pockets which can unlock development of next-gen therapeutics.
Speaker: Soumendranath Bhakat
Twitter - Prudencio
Twitter - Jonny
Twitter - datamol.io
[DISCLAIMER] - For the full visual experience, we recommend you tune in through our YouTube channel to see the presented slides.
Try datamol.io - the open source toolkit that simplifies molecular processing and featurization workflows for machine learning scientists working in drug discovery: https://datamol.io/
If you enjoyed this talk, consider joining the Molecular Modeling and Drug Discovery (M2D2) talks live.
Also, consider joining the M2D2 Slack.
Abstract: The fundamental equations that govern Nature at atomistic scales are well understood in terms of quantum mechanics. Solving such equations is not possible apart from very simple systems, yet solutions to this problem represent one of the grand challenges for computational sciences as it would allow an understanding of all properties of molecular systems. We investigate this challenge by solving the sampling and accuracy problems of atomistic simulations using machine learning, physics, and GPUs. Machine learning potentials, as universal many-body function approximators, could deliver the next-generation modeling approach, blurring the boundary between quantum mechanics, molecular mechanics, and coarse-grained simulations into a cohesive methodology. In recent years, incredible progress has been made in transferable molecular representations which can learn effective potential functions. New methods for learning such potentials and even the energetics of the underlying physical systems are now available. However, there are still problems in extending the generalizability, lack of accurate datasets, and handling of charges and charged molecules, all within a speed bound which must be able to handle large systems like protein complexes. In this talk, I will discuss how to advance these scientific problems toward next-generation molecular simulations both in the context of biomolecular simulations (ACEMD/OpenMM) and more general machine learning frameworks (TorchMD, TorchMD-NET).
Speaker: Gianni De Fabritiis
Twitter - Prudencio
Twitter - Jonny
Twitter - datamol.io
[DISCLAIMER] - For the full visual experience, we recommend you tune in through our YouTube channel to see the presented slides.
Try datamol.io - the open source toolkit that simplifies molecular processing and featurization workflows for machine learning scientists working in drug discovery: https://datamol.io/
If you enjoyed this talk, consider joining the Molecular Modeling and Drug Discovery (M2D2) talks live.
Also, consider joining the M2D2 Slack.
Abstract: Learning effective protein representations is critical in a variety of tasks in biology such as predicting protein function or structure. Existing approaches usually pretrain protein language models on a large number of unlabeled amino acid sequences and then finetune the models with some labeled data in downstream tasks. Despite the effectiveness of sequence-based approaches, the power of pretraining on known protein structures, which are available in smaller numbers only, has not been explored for protein property prediction, though protein structures are known to be determinants of protein function. In this paper, we propose to pretrain protein representations according to their 3D structures. We first present a simple yet effective encoder to learn the geometric features of a protein. We pretrain the protein graph encoder by leveraging multiview contrastive learning and different self-prediction tasks. Experimental results on both function prediction and fold classification tasks show that our proposed pretraining methods outperform or are on par with the state-of-the-art sequence-based methods, while using much less pretraining data.
Speaker: Zuobai Zhang
Twitter - Prudencio
Twitter - Jonny
Twitter - datamol.io
[DISCLAIMER] - For the full visual experience, we recommend you tune in through our YouTube channel to see the presented slides.
Try datamol.io - the open source toolkit that simplifies molecular processing and featurization workflows for machine learning scientists working in drug discovery: https://datamol.io/
If you enjoyed this talk, consider joining the Molecular Modeling and Drug Discovery (M2D2) talks live.
Also, consider joining the M2D2 Slack.
Abstract: While computational methods have become a mainstay in drug discovery programs, many calculations are too time-consuming to be applied to large datasets. Active learning (AL), a machine learning method used to direct a search iteratively, can enable the application of computationally expensive methods such as relative binding free energy (RBFE) calculations to sets containing thousands of molecules. Moreover, AL can also be applied to virtual screening, enabling the rapid processing of billions of molecules. This presentation will provide an overview of active learning and highlight some applications in drug discovery.
Speakes: Pat Walter & James Thompson
Twitter - Prudencio
Twitter - Therence
Twitter - Jonny
Twitter - datamol.io
[DISCLAIMER] - For the full visual experience, we recommend you tune in through our YouTube channel to see the presented slides.
Try datamol.io - the open source toolkit that simplifies molecular processing and featurization workflows for machine learning scientists working in drug discovery: https://datamol.io/
If you enjoyed this talk, consider joining the Molecular Modeling and Drug Discovery (M2D2) talks live.
Also, consider joining the M2D2 Slack.
Abstract: Developing novel bioactive molecules is time-consuming, costly and rarely successful. As a mitigation strategy, we utilize, for the first time, cellular morphology to directly guide the de novo design of small molecules. We trained a conditional generative adversarial network on a set of 30 000 compounds using their cell painting morphological profiles as conditioning. Our model was able to learn chemistry-morphology relationships and influence the generated chemical space according to the morphological profile. We provide evidence for the targeted generation of known agonists when conditioning on gene overexpression profiles, even though no information on biological targets was used during training. Based on a target-agnostic readout, our approach facilitates knowledge transfer between biological pathways and can be used to design bioactives for many targets under one unified framework. Prospective application of this proof-of-concept to larger chemical spaces promises great potential for hit generation in drug and phytopharmaceutical discovery and chemical safety.
Speaker: Paula A. Marin Zapata
Twitter - Prudencio
Twitter - Therence
Twitter - Jonny
Twitter - datamol.io
[DISCLAIMER] - For the full visual experience, we recommend you tune in through our YouTube channel to see the presented slides.
If you enjoyed this talk, consider joining the Molecular Modeling and Drug Discovery (M2D2) talks live.
Also, consider joining the M2D2 Slack.
Abstract: The process of finding molecules that bind to a target protein is a challenging first step in drug discovery. Crystallographic fragment screening is a strategy based on elucidating binding modes of small polar compounds and then building potency by expanding or merging them. Recent advances in high-throughput crystallography enable screening of large fragment libraries, reading out dense ensembles of fragments spanning the binding site. However, fragments typically have low affinity thus the road to potency is often long and fraught with false starts. Here, we take advantage of high-throughput crystallography to reframe fragment-based hit discovery as a denoising problem – identifying significant pharmacophore distributions from a fragment ensemble amid noise due to weak binders – and employ an unsupervised machine learning method to tackle this problem. Our method screens potential molecules by evaluating whether they recapitulate those fragment-derived pharmacophore distributions. We retrospectively validated our approach on an open science campaign against SARS-CoV-2 main protease (Mpro), showing that our method can distinguish active compounds from inactive ones using only structural data of fragment-protein complexes, without any activity data. Further, we prospectively found novel hits for Mpro and the Mac1 domain of SARS-CoV-2 non-structural protein 3. More broadly, our results demonstrate how unsupervised machine learning helps interpret high throughput crystallography data to rapidly discover of potent chemical modulators of protein function.
Speaker: William McCorkindale
Twitter - Prudencio
Twitter - Therence
Twitter - Jonny
Twitter - Valence Discovery
[DISCLAIMER] - For the full visual experience, we recommend you tune in through our YouTube channel to see the presented slides.
If you enjoyed this talk, consider joining the Molecular Modeling and Drug Discovery (M2D2) talks live.
Also, consider joining the M2D2 Slack.
Abstract: Deep learning models that leverage large datasets are often the state of the art for modelling molecular properties. When the datasets are smaller (less than 2000 molecules), it is not clear that deep learning approaches are the right modelling tool. In this work we perform an extensive study of the calibration and generalizability of probabilistic machine learning models on small chemical datasets. Using different molecular representations and models, we analyse the quality of their predictions and uncertainties in a variety of tasks (binary, regression) and datasets. We also introduce two simulated experiments that evaluate their performance: (1) Bayesian optimization guided molecular design, (2) inference on out-of-distribution data via ablated cluster splits. We offer practical insights into model and feature choice for modelling small chemical datasets, a common scenario in new chemical experiments. We have packaged our analysis into the DIONYSUS repository, which is open sourced to aid in reproducibility and extension to new datasets.
Speaker: Gary Tom
Twitter - Prudencio
Twitter - Therence
Twitter - Jonny
Twitter - Valence Discovery
[DISCLAIMER] - For the full visual experience, we recommend you tune in through our YouTube channel to see the presented slides.
If you enjoyed this talk, consider joining the Molecular Modeling and Drug Discovery (M2D2) talks live.
Also, consider joining the M2D2 Slack.
Abstract: In computer-aided drug discovery, quantitative structure activity relation models are trained to predict biological activity from chemical structure. Despite the recent success of applying graph neural network to this task, important chemical information such as molecular chirality is ignored. To fill this crucial gap, we propose Molecular-Kernel Graph Neural Network (MolKGNN) for molecular representation learning, which features SE(3)-/conformation invariance, and interpretability. For our MolKGNN, we first design a molecular graph convolution to capture the chemical pattern by comparing the atom’s similarity with the learnable molecular kernels. Furthermore, we propagate the similarity score to capture the higher-order chemical pattern. To assess the method, we conduct a comprehensive evaluation with nine well-curated datasets spanning numerous important drug targets that feature realistic high class imbalance and it demonstrates the superiority of MolKGNN over other GNNs in CADD. Meanwhile, the learned kernels identify patterns that agree with domain knowledge, confirming the pragmatic interpretability of this approach. This work was recently accepted by AAAI23.
Speaker: Yunchao (Lance) Liu
Twitter - Prudencio
Twitter - Therence
Twitter - Jonny
Twitter - Valence Discovery
[DISCLAIMER] - For the full visual experience, we recommend you tune in through our YouTube channel to see the presented slides.
If you enjoyed this talk, consider joining the Molecular Modeling and Drug Discovery (M2D2) talks live.
Also, consider joining the M2D2 Slack.
Abstract: Deep learning, and in general, auto-differentiation frameworks allow the expression of many scientific problems as end-to-end learning tasks. Common themes in scientific machine learning involve learning surrogate functions of expensive simulators, sampling complex distributions directly or time-propagation of known or unknown differential equation systems efficiently. We will describe our recent work in applying deep-learning surrogates and auto-differentiation techniques in molecular simulations. In particular, we will explore active learning of machine learning potentials with differentiable uncertainty; the use of deep neural network generative models to learn reversible coarse-grained representations of atomic systems; and the application of differentiable simulations for reaction path finding without prior knowledge of collective variables. Overfitting and lack of generalizability are constant challenges in AI for science. We will discuss the scalability of active learning to practical applications in molecular simulations and the sensitivity of deep learning answers to molecular problems, like the fitting of pair potentials from observables.
Speaker: Can Chen
Twitter - Prudencio
Twitter - Therence
Twitter - Jonny
Twitter - Valence Discovery
[DISCLAIMER] - For the full visual experience, we recommend you tune in through our YouTube channel to see the presented slides.
If you enjoyed this talk, consider joining the Molecular Modeling and Drug Discovery (M2D2) talks live.
Also, consider joining the M2D2 Slack.
Abstract: Deep learning, and in general, auto-differentiation frameworks allow the expression of many scientific problems as end-to-end learning tasks. Common themes in scientific machine learning involve learning surrogate functions of expensive simulators, sampling complex distributions directly or time-propagation of known or unknown differential equation systems efficiently. We will describe our recent work in applying deep-learning surrogates and auto-differentiation techniques in molecular simulations. In particular, we will explore active learning of machine learning potentials with differentiable uncertainty; the use of deep neural network generative models to learn reversible coarse-grained representations of atomic systems; and the application of differentiable simulations for reaction path finding without prior knowledge of collective variables. Overfitting and lack of generalizability are constant challenges in AI for science. We will discuss the scalability of active learning to practical applications in molecular simulations and the sensitivity of deep learning answers to molecular problems, like the fitting of pair potentials from observables.
Speaker: Rafael Gomez-Bombarelli
Twitter - Prudencio
Twitter - Therence
Twitter - Jonny
Twitter - Valence Discovery
[DISCLAIMER] - For the full visual experience, we recommend you tune in through our YouTube channel to see the presented slides.
If you enjoyed this talk, consider joining the Molecular Modeling and Drug Discovery (M2D2) talks live.
Also, consider joining the M2D2 Slack.
Abstract: Accurate 3D molecular information is a keystone for many computational programs but accessing reliable 3D conformers is still challenging. It requires enumerating and optimizing a huge isomer and conformer space, which would overwhelm any traditional computational methods. In light of this, we proposed the Auto3D package for generating low-energy 3D conformers using fast and reliable neural network potentials (NNPs). Given a SMILES, Auto3D returns the low-energy 3D conformers by automatizing the isomer enumeration and duplicate filtering process, 3D building process, geometry optimization process, and ranking process. In conjunction with Auto3D, we developed ANI-2xt NNP, which was trained especially for tautomer-related tasks. These NNPs were used to generate 3D structures and compute molecular properties. In a tautomeric reaction energy calculation task, the ANI-2xt NNP achieved similar accuracy but was several orders of magnitude faster than the reference DFT method.
Speaker: Zhen Liu
Twitter - Prudencio
Twitter - Therence
Twitter - Jonny
Twitter - Valence Discovery
[DISCLAIMER] - For the full visual experience, we recommend you tune in through our YouTube channel to see the presented slides.
If you enjoyed this talk, consider joining the Molecular Modeling and Drug Discovery (M2D2) talks live.
Also, consider joining the M2D2 Slack.
Abstract: Recently, artificial intelligence (AI) for drug discovery has raised increasing interest in both the machine learning (ML) and computational chemistry communities. The core problem of AI for drug discovery is molecule representation learning, where the molecule knowledge can be naturally presented in different modalities: chemical formula, molecular graph, geometric conformation, knowledge base, biomedical literature, etc. In this talk, I would like to provide a perspective concentrating on molecule pretraining from topology, geometry, and textual description. Such a unified perspective paves the way for molecule representation interpretation as well as discovery tasks.
Speaker: Shengchao Liu
Twitter - Prudencio
Twitter - Therence
Twitter - Jonny
Twitter - Valence Discovery
[DISCLAIMER] - For the full visual experience, we recommend you tune in through our YouTube channel to see the presented slides.
If you enjoyed this talk, consider joining the Molecular Modeling and Drug Discovery (M2D2) talks live.
Also, consider joining the M2D2 Slack.
Abstract: Molecular dynamics (MD) simulation techniques are widely used for various natural science applications. Increasingly, machine learning (ML) force field (FF) models begin to replace ab-initio simulations by predicting forces directly from atomic structures. Despite significant progress in this area, such techniques are primarily benchmarked by their force/energy prediction errors, even though the practical use case would be to produce realistic MD trajectories. We aim to fill this gap by introducing a novel benchmark suite for ML MD simulation. We curate representative MD systems, including water, organic molecules, peptide, and materials, and design evaluation metrics corresponding to the scientific objectives of respective systems. We benchmark a collection of state-of-the-art (SOTA) ML FF models and illustrate, in particular, how the commonly benchmarked force accuracy is not well aligned with relevant simulation metrics. We demonstrate when and how selected SOTA methods fail, along with offering directions for further improvement. Specifically, we identify stability as a key metric for ML models to improve. Our benchmark suite comes with a comprehensive open-source codebase for training and simulation with ML FFs to facilitate further work.
Speaker: Xiang Fu
Twitter - Prudencio
Twitter - Therence
Twitter - Jonny
Twitter - Valence Discovery
[DISCLAIMER] - For the full visual experience, we recommend you tune in through our YouTube channel to see the presented slides.
If you enjoyed this talk, consider joining the Molecular Modeling and Drug Discovery (M2D2) talks live.
Also, consider joining the M2D2 Slack.
Abstract: The universality of thermodynamics and statistical mechanics has led to a language comprehensible to chemists, physicists & others, enabling countless scientific discoveries in diverse fields. In the last decade, a new arguably common language that everyone seems to speak but at least no chemist fully understands, has emerged with the advent of artificial intelligence (AI). It is natural to ask if AI can be integrated with the various theoretical and simulation methods in chemistry for new discoveries. At the same this raises many open questions, including: (1) should chemists, who are not fundamentally trained in AI, trust any of the results obtained using AI, (2) can AI paradigms developed for non-molecular systems with massive training data can directly be applied to chemistry with all its quirks, richness, known/unknown laws, and often poor/limited data? In this seminar I will show how such an integration of disciplines can be attained, creating trustable, robust AI frameworks for use by chemists. I will demonstrate such methods on different problems involving protein kinases, riboswitches and crystal polymorph nucleation, where we predict mechanisms at timescales much longer than milliseconds while keeping all-atom/femtosecond resolution. I will conclude with an outlook for future challenges and opportunities, envisioning a new sub-discipline of “Artificial Chemical Intelligence” where chemistry moves hand-in-hand with AI to enable smart molecular discovery, and is not just yet another domain for application of AI.
Speaker: Pratyush Tiwary
Twitter - Prudencio
Twitter - Therence
Twitter - Jonny
Twitter - Valence Discovery
[DISCLAIMER] - For the full visual experience, we recommend you tune in through our YouTube channel to see the presented slides.
If you enjoyed this talk, consider joining the Molecular Modeling and Drug Discovery (M2D2) talks live.
Also, consider joining the M2D2 Slack.
Abstract: The latest biological findings observe that the motionless lock-and-key theory is no longer applicable and that changes in atomic sites and binding pose can provide important information for understanding drug binding. However, the computational expenditure limits the growth of protein trajectory-related studies, thus hindering the possibility of supervised learning. We present a novel spatial-temporal pre-training method based on the modified Equivariant Graph Matching Networks (EGMN), dubbed ProtMD. It has two specially designed self-supervised learning tasks: atom-level prompt-based denoising generative task and conformation-level snapshot ordering task to seize the flexibility information inside MD trajectories with very fine temporal resolutions. More importantly, we investigate the underlying mechanism behind the success of ProtMD, and further demonstrate a tight correlation between the magnitude of spatial motion of conformation and the extent to which the ligand and the receptor bind with each other.
Speaker: Fang Wu
Twitter - Prudencio
Twitter - Therence
Twitter - Jonny
Twitter - Valence Discovery
[DISCLAIMER] - For the full visual experience, we recommend you tune in through our YouTube channel to see the presented slides.
If you enjoyed this talk, consider joining the Molecular Modeling and Drug Discovery (M2D2) talks live.
Also, consider joining the M2D2 Slack.
Abstract: Everything tangible in the universe is made of molecules. Yet our ability to digitally simulate even small molecules is rather poor due to the complexities of quantum mechanics. However, there are a number of advances that are converging to dramatically improve our ability to understand the behavior of molecules. Firstly, deep learning and in particular equivariant graph neural networks are now an important tool to model molecules. They are for instance the core technology in Deepmind’s AlphaFold to predict the 3d shape of a molecule from its amino acid sequence. Second, despite claims to the contrary, Moore’s law is still alive, and in particular the design of ASIC architectures for special purpose computation will continue to accelerate our ability to break new computational barriers. And finally there is the rapid advance of quantum computation. While fault tolerant quantum computation might still be a decade away, it is expected that it’s first useful application, to simulate (quantum) nature itself, may be much closer. In this talk I will introduce some technology around equivariant graph neural networks and give my perspective on why I am excited about the opportunities that will come from new breakthroughs in molecular simulation. It may facilitate the search for new sustainable technologies to capture carbon from the air, develop biodegradable plastics, reduce the cost of electrolysis through better catalysts, develop cleaner and cheaper fertilizers, design new drugs to treat disease and so on. Our understanding of matter will be key to unlocking these new materials for the benefit of humanity.
Speaker: Max Welling
Twitter Prudencio
Twitter Therence
Twitter Jonny
Twitter Valence Discovery
[DISCLAIMER] - For the full visual experience, we recommend you tune in through our YouTube channel to see the presented slides.
If you enjoyed this talk, consider joining the Molecular Modeling and Drug Discovery (M2D2) talks live.
Also, consider joining the M2D2 Slack.
Abstract: Predicting the binding structure of a small molecule ligand to a protein -- a task known as molecular docking -- is critical to drug design. Recent deep learning methods that treat docking as a regression problem have decreased runtime compared to traditional search-based methods but have yet to offer substantial improvements in accuracy. We instead frame molecular docking as a generative modeling problem and develop DiffDock, a diffusion generative model over the non-Euclidean manifold of ligand poses. To do so, we map this manifold to the product space of the degrees of freedom (translational, rotational, and torsional) involved in docking and develop an efficient diffusion process on this space. Empirically, DiffDock obtains a 38% top-1 success rate (RMSD<2A) on PDBBind, significantly outperforming the previous state-of-the-art of traditional docking (23%) and deep learning (20%) methods. Moreover, DiffDock has fast inference times and provides confidence estimates with high selective accuracy.
Full Paper
Speakers: Hannes Stärk, Gabriele Corso, and Bowen Jing
Twitter Prudencio
Twitter Therence
Twitter Jonny
Twitter Valence Discovery
[DISCLAIMER] - For the full visual experience, we recommend you tune in through our YouTube channel to see the presented slides.
If you enjoyed this talk, consider joining the Molecular Modeling and Drug Discovery (M2D2) talks live.
Also, consider joining the M2D2 Slack.
Abstract: Fragment-based drug discovery has been an effective paradigm in early-stage drug development. An open challenge in this area is designing linkers between disconnected molecular fragments of interest to obtain chemically-relevant candidate drug molecules. In this work, we propose DiffLinker, an E(3)-equivariant 3D-conditional diffusion model for molecular linker design. Given a set of disconnected fragments, our model places missing atoms in between and designs a molecule incorporating all the initial fragments. Unlike previous approaches that are only able to connect pairs of molecular fragments, our method can link an arbitrary number of fragments. Additionally, the model automatically determines the number of atoms in the linker and its attachment points to the input fragments. We demonstrate that DiffLinker outperforms other methods on the standard datasets generating more diverse and synthetically-accessible molecules. Besides, we experimentally test our method in real-world applications, showing that it can successfully generate valid linkers conditioned on target protein pockets.
Full Paper
Speakers: Ilia Igashov
Twitter Prudencio
Twitter Therence
Twitter Jonny
Twitter Valence Discovery
[DISCLAIMER] - For the full visual experience, we recommend you tune in through our YouTube channel to see the presented slides.
If you enjoyed this talk, consider joining the Molecular Modeling and Drug Discovery (M2D2) talks live.
Also, consider joining the M2D2 Slack.
Abstract: In this work we propose a principled evaluation framework for model-based optimisation to measure how well a generative model can extrapolate. We achieve this by interpreting the training and validation splits as draws from their respective ‘truncated’ ground truth distributions, where examples in the validation set contain scores much larger than those in the training set. Model selection is performed on the validation set for some prescribed validation metric. A major research question however is in determining what validation metric correlates best with the expected value of generated candidates with respect to the ground truth oracle; work towards answering this question can translate to large economic gains since it is expensive to evaluate the ground truth oracle in the real world. We compare various validation metrics for generative adversarial networks using our framework. We also discuss limitations with our framework with respect to existing datasets and how progress can be made to mitigate them
Full Paper
Speakers: Christopher Beckham
Twitter Prudencio
Twitter Therence
Twitter Jonny
Twitter Valence Discovery
[DISCLAIMER] - For the full visual experience, we recommend you tune in through our YouTube channel to see the presented slides.
If you enjoyed this talk, consider joining the Molecular Modeling and Drug Discovery (M2D2) talks live.
Also, consider joining the M2D2 Slack.
Abstract: Single-cell transcriptomics enabled the study of cellular heterogeneity in response to perturbations at the resolution of individual cells. However, scaling high-throughput screens (HTSs) to measure cellular responses for many drugs remains a challenge due to technical limitations and, more importantly, the cost of such multiplexed experiments. Thus, transferring information from routinely performed bulk RNA-seq HTS is required to enrich single-cell data meaningfully. We introduce a new encoder-decoder architecture to study the perturbational effects of unseen drugs. We combine the model with a transfer learning scheme and demonstrate how training on existing bulk RNA-seq HTS datasets can improve generalisation performance. Better generalisation reduces the need for extensive and costly screens at single-cell resolution. We envision that our proposed method will facilitate more efficient experiment designs through its ability to generate in-silico hypotheses, ultimately accelerating targeted drug discovery.
Full Paper
Speakers: Leon Hetzel and Simon Böhm
Twitter Prudencio
Twitter Therence
Twitter Jonny
Twitter Valence Discovery
[DISCLAIMER] - For the full visual experience, we recommend you tune in through our YouTube channel to see the presented slides.
If you enjoyed this talk, consider joining the Molecular Modeling and Drug Discovery (M2D2) talks live.
Also consider joining the M2D2 Slack
Abstract: In molecular discovery and drug design, structure-property relationships and activity landscapes are often qualitatively or quantitatively analyzed to guide the navigation of chemical space. The roughness (or smoothness) of these molecular property landscapes is one of their most studied geometric attributes, as it can characterize the presence of activity cliffs, with rougher landscapes generally expected to pose tougher optimization challenges. Here, we introduce a general, quantitative measure for describing the roughness of molecular property landscapes. The proposed roughness index (ROGI) is loosely inspired by the concept of fractal dimension and strongly correlates with the out-of-sample error achieved by machine learning models on numerous regression tasks.
Full Paper
Speakers: Matteo Aldeghi
Twitter Prudencio
Twitter Therence
Twitter Jonny
Twitter Valence Discovery
[DISCLAIMER] - For the full visual experience, we recommend you tune in through our YouTube channel to see the presented slides.
If you enjoyed this talk, consider joining the Molecular Modeling and Drug Discovery (M2D2) talks live.
Also consider joining the M2D2 Slack
Abstract: While there is a great deal of interest in methods aimed at explaining machine learning predictions of chemical properties, it is difficult to quantitatively benchmark such methods, especially for regression tasks. We show that the Crippen logP model provides an excellent benchmark for atomic attribution/heatmap approaches, especially if the ground truth heatmaps can be adjusted to reflect the molecular representation. I give some examples of how this benchmark can be used to get a better understanding of ML models work and how it can be used to determine which techniques for generating XAI heatmaps works the best.
Slides from the talk: https://speakerdeck.com/jhjensen/jensen-xai
Speakers: Jan Jensen
Twitter Prudencio
Twitter Therence
Twitter Cas
Twitter Valence Discovery
[DISCLAIMER] - For the full visual experience, we recommend you tune in through our YouTube channel to see the presented slides.
If you enjoyed this talk, consider joining the Molecular Modeling and Drug Discovery (M2D2) talks live.
Also consider joining the M2D2 Slack
Abstract: Pre-trained models have been transformative in natural language, computer vision, and now protein sequences by enabling accuracy with few training examples. We show how to use pretrained sequence models in Bayesian optimization to design new protein sequences with minimal labels (i.e., few experiments). Pre-trained models give good predictive accuracy at low data and Bayesian optimization guides the choice of which sequences to test. Pre-trained sequence models also obviate the common requirement of finite pools. Any sequence can be considered. We show significantly fewer labeled sequences are required for many sequence design tasks, including creating novel peptide inhibitors with AlphaFold. This work should enable calibrated predictions with few examples and iterative design with low data (1-50).
Full Paper
Speakers: Ziyue Yang
Twitter Prudencio
Twitter Therence
Twitter Cas
Twitter Valence Discovery
[DISCLAIMER] - For the full visual experience, we recommend you tune in through our YouTube channel to see the presented slides.
If you enjoyed this talk, consider joining the Molecular Modeling and Drug Discovery (M2D2) talks live.
Also consider joining the M2D2 Slack
Abstract: The construction of a scaffold structure that supports a desired motif, conferring protein function, shows promise for the design of vaccines and enzymes. But a general solution to this motif-scaffolding problem remains open. Current machine-learning techniques for scaffold design are either limited to unrealistically small scaffolds (up to length 20) or struggle to produce multiple diverse scaffolds. We propose to learn a distribution over diverse and longer protein backbone structures via an E(3)-equivariant graph neural network. We develop SMCDiff to efficiently sample scaffolds from this distribution conditioned on a given motif; our algorithm is the first to theoretically guarantee conditional samples from a diffusion model in the large-compute limit. We evaluate our designed backbones by how well they align with AlphaFold2-predicted structures. We show that our method can (1) sample scaffolds up to 80 residues and (2) achieve structurally diverse scaffolds for a fixed motif.
Full Paper
Speakers: Brian Trippe and Jason Yim
Twitter Prudencio
Twitter Therence
Twitter Cas
Twitter Valence Discovery
[DISCLAIMER] - For the full visual experience, we recommend you tune in through our YouTube channel to see the presented slides.
If you enjoyed this talk, consider joining the Molecular Modeling and Drug Discovery (M2D2) talks live.
Also consider joining the M2D2 Slack
Abstract: Drug discovery is a very long and expensive process, taking on average more than 10 years and costing $2.5B to develop a new drug. Artificial intelligence has the potential to significantly accelerate the process of drug discovery by extracting evidence from a huge amount of biomedical data and hence revolutionizes the entire pharmaceutical industry. In particular, graph representation learning and geometric deep learning - a fast-growing topic in the machine learning and data mining community focusing on deep learning for graph-structured and 3D data - have seen great opportunities for drug discovery as many data in the domain are represented as graphs or 3D structures (e.g. molecules, proteins, biomedical knowledge graphs). In this talk, I will introduce our recent progress on geometric deep learning for drug discovery and also a newly released open-source machine learning platform for drug discovery, called TorchDrug.
Speakers: Jian Tang
Twitter Prudencio
Twitter Therence
Twitter Cas
Twitter Valence Discovery
[DISCLAIMER] - For the full visual experience, we recommend you tune in through our YouTube channel to see the presented slides.
If you enjoyed this talk, consider joining the Molecular Modeling and Drug Discovery (M2D2) talks live: https://m2d2.io/talks/m2d2/about/
Also consider joining the M2D2 Slack: https://m2d2group.slack.com/join/shar...
Abstract: The use of machine learning methods for the prediction of reaction yield is an emerging area. We demonstrate the applicability of support vector regression (SVR) for predicting reaction yields, using combinatorial data. Molecular descriptors used in regression tasks related to chemical reactivity have often been based on time-consuming, computationally demanding quantum chemical calculations, usually density functional theory. Structure-based descriptors (molecular fingerprints and molecular graphs) are quicker and easier to calculate and are applicable to any molecule. In this study, SVR models built on structure-based descriptors were compared to models built on quantum chemical descriptors. The models were evaluated along the dimension of each reaction component in a set of Buchwald-Hartwig amination reactions. The structure-based SVR models outperformed the quantum chemical SVR models, along the dimension of each reaction component. The applicability of the models was assessed with respect to similarity to training. Prospective predictions of unseen Buchwald-Hartwig reactions are presented for synthetic assessment, to validate the generalisability of the models, with particular interest along the aryl halide dimension.
Speakers: Jonathan Hirst
Twitter Prudencio
Twitter Therence
Twitter Cas
Twitter Valence Discovery
[DISCLAIMER] - For the full visual experience, we recommend you tune in through our YouTube channel to see the presented slides.
If you enjoyed this talk, consider joining the Molecular Modeling and Drug Discovery (M2D2) talks live: https://m2d2.io/talks/m2d2/about/
Also consider joining the M2D2 Slack: https://m2d2group.slack.com/join/shar...
Abstract: Molecular optimization is a fundamental goal in the chemical sciences and is of central interest to drug and material design. In recent years, significant progress has been made in solving challenging problems across various aspects of computational molecular optimizations, emphasizing high validity, diversity, and, most recently, synthesizability. Despite this progress, many papers report results on trivial or self-designed tasks, bringing additional challenges to directly assessing the performance of new methods. Moreover, the sample efficiency of the optimization--the number of molecules evaluated by the oracle--is rarely discussed, despite being an essential consideration for realistic discovery applications. To fill this gap, we have created an open-source benchmark for practical molecular optimization, PMO, to facilitate the transparent and reproducible evaluation of algorithmic advances in molecular optimization. This paper thoroughly investigates the performance of 25 molecular design algorithms on 23 tasks with a particular focus on sample efficiency. Our results show that most "state-of-the-art" methods fail to outperform their predecessors under a limited oracle budget allowing 10K queries and that no existing algorithm can efficiently solve certain molecular optimization problems in this setting. We analyze the influence of the optimization algorithm choices, molecular assembly strategies, and oracle landscapes on the optimization performance to inform future algorithm development and benchmarking. PMO provides a standardized experimental setup to comprehensively evaluate and compare new molecule optimization methods with existing ones.
Speakers: Tianfan Fu
Twitter Prudencio
Twitter Therence
Twitter Cas
Twitter Valence Discovery
If you enjoyed this talk, consider joining the Molecular Modeling and Drug Discovery (M2D2) talks live.
Also consider joining the M2D2 Slack
Abstract: The success of machine learning (ML) in chemistry and biology is contingent on experimental efforts that generate relevant datasets and that validate model predictions. Critically, such experimentation often necessitates significant time and resource investments. Computational methods to inform experimental modeling could help alleviate this burden and bridge the gap between computational predictions and experimental validation. In this talk, I will discuss how ML algorithms that quantify prediction uncertainties could meet this critical need. Using molecular property prediction and drug discovery as a motivating use case, I will present a new method -- evidential deep learning -- for uncertainty quantification in neural networks and demonstrate its potential to (1) achieve calibrated estimates of model uncertainty, (2) improve sample efficiency via uncertainty-guided active learning, and (3) inform experimental validation via targeted virtual screening. I will close by highlighting how prediction uncertainty can accelerate and guide key steps in experimental lifecycles, opening the door for sustained feedback between computation and experimentation in the chemical and biological sciences.
Speakers: Ava Amini
Twitter Prudencio
Twitter Therence
Twitter Cas
Twitter Valence Discovery:
Abstract: Machine learning research is expanding its reach, beyond the traditional realm of the tech industry and into the activities of other scientists, opening the door to truly transformative advances in these disciplines. In this lecture I will focus on two aspects, modeling and experimental design, that are intertwined in the theory-experiment-analysis active learning loop that constitutes a core element of the scientific methodology. Computers will be necessary to go beyond the currently purely manual research loop and take advantage of high-throughput experimental setups and large-scale experimental datasets. I will discuss methods related to active learning, reinforcement learning, generative modeling, Bayesian ML, amortized variational learning and causal discovery. I will discuss the notion of epistemic uncertainty and how to estimate it. I will motivate generative policies that can sample a diverse set of candidate solutions to a problem, be it for proposing new experiments or causal hypotheses. Finally, I will describe current research to help us with these questions based on a new deep learning probabilistic framework called GFlowNets and how we plan to apply these in areas of great societal need like the unmet challenge of antimicrobial resistance or the discovery of new materials to help fight climate change.
Speakers: Yoshua Bengio - https://www.linkedin.com/in/yoshuaben...
Twitter Prudencio: https://twitter.com/tossouprudencio
Twitter Therence: https://twitter.com/Therence_mtl
Twitter Cas: https://twitter.com/cas_wognum
Twitter Valence Discovery: https://twitter.com/valence_ai
If you enjoyed this talk, consider joining the Molecular Modeling and Drug Discovery (M2D2) talks live: https://valence-discovery.github.io/M2D2-meetings/
Also consider joining the M2D2 Slack: https://m2d2group.slack.com/join/shared_invite/zt-16w1rjqqs-n81TiK~iB23XbZ0QWMYs~A#/shared-invite/email
Abstract: The field of explainable AI applied to molecular property prediction models has often been reduced to deriving atomic contributions. This has impaired the interpretability of such models, as chemists rather think in terms of larger, chemically meaningful structures, which often do not simply reduce to the sum of their atomic constituents. In this talk I will explain an explanatory framework yielding both local as well as more complex structural attributions. The key idea is to derive such contextual explanations in pixel space, exploiting the property that a molecule is not merely encoded through a collection of atoms and bonds, as is the case for string- or graph-based approaches. I’ll provide evidence that the proposed explanation method satisfies desirable properties, namely sparsity and invariance with respect to the molecule’s symmetries.
Speakers: Marco Bertolini - https://www.linkedin.com/in/marco-bertolini-03907a59/
Twitter Prudencio: https://twitter.com/tossouprudencio
Twitter Therence: https://twitter.com/Therence_mtl
Twitter Cas: https://twitter.com/cas_wognum
Twitter Valence Discovery: https://twitter.com/valence_ai
If you enjoyed this talk, consider joining the Molecular Modeling and Drug Discovery (M2D2) talks live: https://valence-discovery.github.io/M...
Also consider joining the M2D2 Slack: https://m2d2group.slack.com/join/shar...
Abstract: High-throughput drug sensitivity experiments in cancer enable rapid in-vitro testing of various compounds on cancer cell lines, or patient-derived material, in order to determine the efficacy of a certain treatment. Accurate prediction of dose-response functions from a limited set of pre-clinical experiments is key to explore the large space of possible treatment options, or to prioritize which experiments to perform. This is particularly important when predicting the effect of drug combinations, where it is unfeasible to test all possible combinations...
Speakers: Leiv Rønneberg - https://twitter.com/ltronneberg
Twitter Prudencio: https://twitter.com/tossouprudencio
Twitter Therence: https://twitter.com/Therence_mtl
Twitter Cas: https://twitter.com/cas_wognum
Twitter Valence Discovery: https://twitter.com/valence_ai
If you enjoyed this talk, consider joining the Molecular Modeling and Drug Discovery (M2D2) talks live: https://valence-discovery.github.io/M...
Also consider joining the M2D2 Slack: https://m2d2group.slack.com/join/shar...
Abstract: This talk features three open-source projects that are lowering the entrance barriers in AI for drug discovery and improving the speed at which researchers can develop new computational methods for the field...
Speakers: Hadrien Mary (Datamol), Chence Shi and Zuobai Zhang (TorchDrug), Kexin Huang (TDC)
Twitter Prudencio: https://twitter.com/tossouprudencio
Twitter Therence: https://twitter.com/Therence_mtl
Twitter Cas: https://twitter.com/cas_wognum
Twitter Valence Discovery: https://twitter.com/valence_ai
If you enjoyed this talk, consider joining the Molecular Modeling and Drug Discovery (M2D2) talks live: https://valence-discovery.github.io/M...
Also consider joining the M2D2 Slack: https://join.slack.com/t/m2d2group/sh...
Abstract: Rational drug design depends on the ability to predict both the three-dimensional structures of candidate molecules bound to their targets and the associated binding affinities. Such predictions are generally informed by either the target’s 3D structure or binding affinity measurements for other molecules at the target. We developed a rigorous statistical framework to combine these two sources of information. Our framework allows nonstructural data—a list of ligands that are known to bind the same target but for which no 3D structure is available—to be used to improve binding pose predictions and improves virtual screening enrichments as compared to simple combinations of physics-based docking and ligand-based modeling.
Speaker: Joseph M. Paggi
Twitter Prudencio: https://twitter.com/tossouprudencio
Twitter Therence: https://twitter.com/Therence_mtl
Twitter Cas: https://twitter.com/cas_wognum
Twitter Valence Discovery: https://twitter.com/valence_ai
If you enjoyed this talk, consider joining the Molecular Modeling and Drug Discovery (M2D2) talks live: https://valence-discovery.github.io/M...
Also consider joining the M2D2 Slack: https://join.slack.com/t/m2d2group/sh...
Abstract: To understand the plasticity of cells and their responses to molecular perturbations, such as drugs or developmental signals, it is vital to recover the underlying population dynamics and fate decisions of single cells. However, measuring features of single cells requires destroying them. As a result, a cell population can only be monitored with unpaired sequential snapshots. In order to reconstruct individual cell fate trajectories, as well as the overall dynamics, one needs to re-align these unpaired snapshots, in order to guess for each cell what it might have become at the next step...
Speaker: Charlotte Bunne - https://twitter.com/_bunnech
Twitter Prudencio: https://twitter.com/tossouprudencio
Twitter Therence: https://twitter.com/Therence_mtl
Twitter Cas: https://twitter.com/cas_wognum
Twitter Valence Discovery: https://twitter.com/valence_ai
If you enjoyed this talk, consider joining the Molecular Modeling and Drug Discovery (M2D2) talks live: https://valence-discovery.github.io/M...
Also consider joining the M2D2 Slack: https://join.slack.com/t/m2d2group/sh...
Abstract: The problem of molecular generation has received significant attention recently. Existing methods are typically based on deep neural networks and require training on large datasets with tens of thousands of samples. In practice, however, the size of class-specific chemical datasets is usually limited (e.g., dozens of samples) due to labor-intensive experimentation and data collection. This presents a considerable challenge for the deep learning generative models to comprehensively describe the molecular design space. Another major challenge is to generate only physically synthesizable molecules...
Speaker: Minghao Guo - https://twitter.com/GuoMh14
Twitter Prudencio: https://twitter.com/tossouprudencio
Twitter Therence: https://twitter.com/Therence_mtl
Twitter Cas: https://twitter.com/cas_wognum
Twitter Valence Discovery: https://twitter.com/valence_ai
If you enjoyed this talk, consider joining the Molecular Modeling and Drug Discovery (M2D2) talks live: https://valence-discovery.github.io/M...
Also consider joining the M2D2 Slack: https://join.slack.com/t/m2d2group/sh...
Abstract: Finding synthesis routes for molecules of interest is essential in the discovery of new drugs and materials. To find such routes, computer-assisted synthesis planning (CASP) methods are employed, which rely on a single-step model of chemical reactivity. In this study, we introduce a template-based single-step retrosynthesis model based on Modern Hopfield Networks, which learn an encoding of both molecules and reaction templates in order to predict the relevance of templates for a given molecule...
Speaker: Philipp Seidl - https://twitter.com/phseidl
Twitter Prudencio: https://twitter.com/tossouprudencio
Twitter Therence: https://twitter.com/Therence_mtl
Twitter Cas: https://twitter.com/cas_wognum
Twitter Valence Discovery: https://twitter.com/valence_ai
If you enjoyed this talk, consider joining the Molecular Modeling and Drug Discovery (M2D2) talks live: https://valence-discovery.github.io/M...
Also consider joining the M2D2 Slack: https://join.slack.com/t/m2d2group/sh...
Abstract: Proximity-inducing compounds (PICs) are an emergent drug technology through which the protein of interest (POI) is brought into the vicinity of proteins that control various cellular processes, giving rise to therapeutic benefits. One of the best-known PICs examples are heterobifunctional molecules known as proteolysis targeting chimeras (PROTACs), which induce protein degradation by establishing proximity between a POI and an E3 ligase. In silico PROTAC discovery requires computationally predicting the ternary complex consisting of POI, PROTAC molecule, and E3 ligase. To date, however, all of the approaches for modeling ternary complexes have not been both effective and computationally fast enough. We present a novel machine learning-based method for predicting PROTAC-mediated ternary complex structures based on Bayesian optimization. We show how a fitness combining an estimation of protein-protein interactions with PROTAC energy allows to find good candidate structures...
Speaker: Noah Weber
Twitter Prudencio: https://twitter.com/tossouprudencio
Twitter Therence: https://twitter.com/Therence_mtl
Twitter Cas: https://twitter.com/cas_wognum
Twitter Valence Discovery: https://twitter.com/valence_ai
If you enjoyed this talk, consider joining the Molecular Modeling and Drug Discovery (M2D2) talks live: https://valence-discovery.github.io/M...
Also consider joining the M2D2 Slack: https://join.slack.com/t/m2d2group/sh...
Abstract: Molecular stereochemistry strongly alters (bio)chemical interactions yet is often neglected in molecular deep learning. Tetrahedral (point) chirality, a form of stereochemistry describing relative spatial arrangements of bonded neighbours around tetrahedral carbon centers, especially influences substrate-catalyst binding and is therefore critical for pharmaceutical drug design and asymmetric catalysis...
Speaker: Keir Adams - https://twitter.com/keiradams
Twitter Prudencio: https://twitter.com/tossouprudencio
Twitter Therence: https://twitter.com/Therence_mtl
Twitter Cas: https://twitter.com/cas_wognum
Twitter Valence Discovery: https://twitter.com/valence_ai
If you enjoyed this talk, consider joining the Molecular Modeling and Drug Discovery (M2D2) talks live: https://valence-discovery.github.io/M...
Also consider joining the M2D2 Slack: https://join.slack.com/t/m2d2group/sh...
Abstract: ML has become a crucial tool in drug discovery and chemistry at large, e.g. to predict molecular properties, such as bioactivity, with high levels of accuracy. However, activity cliffs – pairs of molecules that are highly similar in their structure but exhibit large differences in potency – have been underinvestigated for their effect on model performance. Not only are these edge cases informative for molecule discovery and optimization, but models that are well-equipped to accurately predict the potency of activity cliffs have an increased potential for prospective applications...
Speaker: Derek van Tilborg - https://twitter.com/DerekvTilborg
Twitter Prudencio: https://twitter.com/tossouprudencio
Twitter Therence: https://twitter.com/Therence_mtl
Twitter Cas: https://twitter.com/cas_wognum
Twitter Valence Discovery: https://twitter.com/valence_ai
If you enjoyed this talk, consider joining the Molecular Modeling and Drug Discovery (M2D2) talks live: https://valence-discovery.github.io/M...
Also consider joining the M2D2 Slack: https://join.slack.com/t/m2d2group/sh...
Abstract: In this talk, I will explain how to empower graph neural networks (GNNs) for molecular property prediction with more expressive models and large datasets. (GNNs) have emerged as one of the most important innovations for machine learning in drug discovery. Their ability to work on unstructured data enables us to use deep learning on molecular graphs, with the promise of predicting molecular properties with the same speed and accuracy that convolutional networks process images.
Speaker: Dominique Beaini - https://twitter.com/dom_beaini
Twitter Prudencio: https://twitter.com/tossouprudencio
Twitter Therence: https://twitter.com/Therence_mtl
Twitter Cas: https://twitter.com/cas_wognum
Twitter Valence Discovery: https://twitter.com/valence_ai
If you enjoyed this talk, consider joining the Molecular Modeling and Drug Discovery (M2D2) talks live: https://valence-discovery.github.io/M...
Also consider joining the M2D2 Slack: https://join.slack.com/t/m2d2group/sh..
Abstract: In computational drug discovery we often rely on surrogate models for molecular design and optimization. The estimates of the epistemic uncertainty of those surrogate models can be useful signals to enable efficient exploration in the molecular space. Traditional methods such as Bayesian optimization use Bayesian models like Gaussian Processes (GPs) as surrogates. Gaussian processes provide well calibrated uncertainty estimates, but do not scale trivially to large datasets and require hand-crafted kernels to work with structured data like strings (SMILES, peptides) and graphs...
Speaker: Moksh Jain - https://mj10.github.io/
Twitter Prudencio: https://twitter.com/tossouprudencio
Twitter Therence: https://twitter.com/Therence_mtl
Twitter Cas: https://twitter.com/cas_wognum
Twitter Valence Discovery: https://twitter.com/valence_ai
If you enjoyed this talk, consider joining the Molecular Modeling and Drug Discovery (M2D2) talks live: https://valence-discovery.github.io/M...
Also consider joining the M2D2 Slack: https://join.slack.com/t/m2d2group/sh...
Abstract: Combining drugs opens new possibilities for tailoring therapies to a given disease and targeting several biological pathways at the same time. However the number of possible drug combinations is huge, and only a tiny portion of this space can be explored in a reasonable amount of time. One could either narrowly focus on experimenting with a very restricted number of well-chosen drugs, based on pre-existing biological knowledge, or broaden the scope by exploring uncharted territories with the risk of very low time and cost efficiency.
Speaker: Paul Bertin - https://bertinus.github.io/
Twitter Prudencio: https://twitter.com/tossouprudencio
Twitter Therence: https://twitter.com/Therence_mtl
Twitter Cas: https://twitter.com/cas_wognum
Twitter Valence Discovery: https://twitter.com/valence_ai
If you enjoyed this talk, consider joining the Molecular Modeling and Drug Discovery (M2D2) talks live: https://valence-discovery.github.io/M...
Also consider joining the M2D2 Slack: https://join.slack.com/t/m2d2group/shared_invite/zt-16i9r9jir-ioE0TJVHEO~bAyZxu17neg
Abstract: Understanding the 3D structures and interactions of proteins and drug-like molecules is a key part of therapeutics discovery. A core problem is molecular docking, i.e., determining how two molecules attach and create a molecular complex. Having access to very fast accurate computational docking tools would enable applications such as virtual screening of cancer protein inhibitors, de novo drug design, or rapid in silico drug side-effect prediction. In this talk, I will show that geometry and deep learning (DL) can significantly reduce this enormous search space inherent in docking and molecular conformation prediction. I will present EquiDock and EquiBind, our recent DL architectures for direct shot prediction of the molecular complex, and GeoMol, a model for 3D molecular flexibility.
Speaker: Octavian-Eugen Ganea - https://twitter.com/octavianEganea
Twitter Prudencio: https://twitter.com/tossouprudencio
Twitter Therence: https://twitter.com/Therence_mtl
Twitter Cas: https://twitter.com/cas_wognum
Twitter Valence Discovery: https://twitter.com/valence_ai
If you enjoyed this talk, consider joining the Molecular Modeling and Drug Discovery (M2D2) talks live: https://valence-discovery.github.io/M...
Also consider joining the M2D2 Slack: https://join.slack.com/t/m2d2group/shared_invite/zt-16i9r9jir-ioE0TJVHEO~bAyZxu17neg
Abstract: Machine learning methodologies have become increasingly established in many chemistry-related disciplines. Recent advances in quantum machine learning (QML) have enabled the prediction of QM-properties at a fraction of the cost of first-principle methods such as density functional theory (DFT). However, previous work has generally been focused on molecular systems of limited size and atom type diversity, hindering its application to adjacent fields such as drug discovery. The limited availability of open-source, high-performance models has further rendered the adoption of this progress to new fields challenging. We introduce two contributions towards overcoming these issues: The QMugs data collection (Quantum Mechanical properties of drug-like molecules) provides a wide array of QM-properties for large and biologically relevant molecules, increasing the chemical space accessible to ML models.
Speaker: Clemens Isert - https://twitter.com/clemensisert
Twitter Prudencio: https://twitter.com/tossouprudencio
Twitter Therence: https://twitter.com/Therence_mtl
Twitter Cas: https://twitter.com/cas_wognum
Twitter Valence Discovery: https://twitter.com/valence_ai
If you enjoyed this talk, consider joining the Molecular Modeling and Drug Discovery (M2D2) talks live: https://valence-discovery.github.io/M...
Also consider joining the M2D2 Slack: https://join.slack.com/t/m2d2group/shared_invite/zt-16i9r9jir-ioE0TJVHEO~bAyZxu17neg
Abstract: Antibodies are versatile proteins that bind to pathogens like viruses and stimulate the adaptive immune system. The antibody binding affinity is determined by complementarity-determining regions (CDRs) at the tips of these Y-shaped proteins, which closely interact with antigen residues (epitopes). In this talk, I will present new generative models to automatically design the CDRs of antibodies with desired binding affinity. Specifically, our model seeks to co-design the sequence and 3D structure of CDRs as graphs. It unravels a sequence auto regressively while iteratively refining its predicted global 3D structure. Our model is evaluated on binder design tasks and shows superior performance compared to existing baselines.
Speaker: Wengong Jin - http://people.csail.mit.edu/wengong/
Twitter Prudencio: https://twitter.com/tossouprudencio
Twitter Therence: https://twitter.com/Therence_mtl
Twitter Cas: https://twitter.com/cas_wognum
Twitter Valence Discovery: https://twitter.com/valence_ai
If you enjoyed this talk, consider joining the Molecular Modeling and Drug Discovery (M2D2) talks live: https://valence-discovery.github.io/M...
Also consider joining the M2D2 Slack: https://join.slack.com/t/m2d2group/shared_invite/zt-16i9r9jir-ioE0TJVHEO~bAyZxu17neg
Abstract: Beyond the active search for new drugs, de novo generation methods are also a great opportunity for the discovery of molecular materials. However, the chemical space of these materials differs from that of bioactive molecules. This conference will present the challenges inherent to this kind of problems. For most molecular materials, new targets must have specific electronic properties. This normally means a very costly evaluation by quantum mechanical calculations. Furthermore this evaluation requires a knowledge of the atomic positions in three dimensions. All these specific constraints have led us to propose our own generation method based on EvoMol, an efficient evolutionary algorithm. Free to travel the whole chemical space, the methods that limit the solutions to realistic molecules will be presented.
Speaker: Thomas Cauchy - https://twitter.com/ThomasCauchyQC
Twitter Prudencio: https://twitter.com/tossouprudencio
Twitter Therence: https://twitter.com/Therence_mtl
Twitter Cas: https://twitter.com/cas_wognum
Twitter Valence Discovery: https://twitter.com/valence_ai
If you enjoyed this talk, consider joining the Molecular Modeling and Drug Discovery (M2D2) talks live: https://valence-discovery.github.io/M2D2-meetings/
Also consider joining the M2D2 Slack: https://join.slack.com/t/m2d2group/shared_invite/zt-16i9r9jir-ioE0TJVHEO~bAyZxu17neg
Abstract: In organic chemistry, we are currently witnessing a rise in artificial intelligence (AI) approaches, which show great potential for improving molecular designs, facilitating synthesis and accelerating the discovery of novel molecules. Based on an analogy between written language and organic chemistry, we built linguistics-inspired transformer neural network models for chemical reaction prediction, synthesis planning, and the prediction of experimental actions. We extended the models to chemical reaction classification and fingerprints. By finding a mapping from discrete reactions to continuous vectors, we enabled efficient chemical reaction space exploration. Moreover, we specialized similar models for reaction yield predictions. Intrigued by the remarkable performance of chemical language models, we discovered that the models can capture how atoms rearrange during a reaction, without supervision or human labelling, leading to the development of the open-source atom-mapping tool RXNMapper. During my talk, I will provide an overview of the different contributions that are at the base of this digital synthetic chemistry revolution.
Speaker: Philippe Schwaller - https://twitter.com/pschwllr
Twitter Prudencio: https://twitter.com/tossouprudencio
Twitter Therence: https://twitter.com/Therence_mtl
Twitter Cas: https://twitter.com/cas_wognum
Twitter Valence Discovery: https://twitter.com/valence_ai
If you enjoyed this talk, consider joining the Molecular Modeling and Drug Discovery (M2D2) talks live: https://valence-discovery.github.io/M...
Abstract: Deep learning in molecular and materials sciences is limited by the lack of integration between applied science, artificial intelligence, and high-performance computing. Bottlenecks with respect to the amount of training data, the size and complexity of model architectures, and the scale of compute infrastructure are all key factors limiting the scaling of deep learning for molecules and materials. In cases where design goals require explorations of vast areas of chemical/material space, or target properties are prohibitively expensive to compute, efficient use of resources and careful choice of method enable new capabilities for design. We explore interactive supercomputing for applying high-throughput virtual screening and machine learning to challenges in materials and chemistry. The abundance of data from first-principles calculations introduces a need to identify and investigate scalable neural network architectures that operate on graphs, which are a natural representation for atomistic systems. We present LitMatter, a lightweight framework for scaling geometric deep learning methods. We discuss scaling atomistic deep learning using key resources including compute, model and dataset sizes, and energy. We train four graph neural network architectures on over 400 GPUs and investigate the scaling behavior of these methods. Depending on the model architecture, training time speedups up to 60x are seen. Empirical neural scaling relations quantify the model-dependent scaling and enable optimal compute resource allocation and the identification of scalable geometric deep learning model implementations. Training speed estimation and energy monitoring are used to accelerate hyperparameter optimization for neural interatomic potentials, and quantify the efficiency of physics-informed architectures. We discuss applications of scalable ML to property prediction tasks, deep generative modeling, and neural force fields for fully differentiable simulations.
Speaker: Nathan C. Frey - https://ncfrey.github.io/
Twitter Prudencio: https://twitter.com/tossouprudencio
Twitter Therence: https://twitter.com/Therence_mtl
If you enjoyed this talk, consider joining the Molecular Modeling and Drug Discovery (M2D2) talks live: https://valence-discovery.github.io/M...
Abstract: Scientists and engineers in diverse domains need to perform expensive experiments to optimize combinatorial spaces, where each candidate input is a discrete structure (e.g., sequence, tree, graph) or a hybrid structure (mixture of discrete and continuous design variables). For example, in drug and vaccine design, we need to search a large space of molecules guided by physical lab experiments. These experiments are often performed in a heuristic manner by humans and without any formal reasoning. Bayesian optimization (BO) is an efficient framework for optimizing expensive black-box functions. However, most of the BO literature is largely focused on optimizing continuous spaces. In this talk, I will discuss the main challenges in extending BO framework to combinatorial structures and some algorithms that I have developed in addressing them.
Speaker: Aryan Deshwal - https://aryandeshwal.github.io/
Twitter Prudencio: https://twitter.com/tossouprudencio
Twitter Therence: https://twitter.com/Therence_mtl
Twitter Cas: https://twitter.com/cas_wognum
Twitter Valence Discovery: https://twitter.com/valence_ai
If you enjoyed this talk, consider joining the Molecular Modeling and Drug Discovery (M2D2) talks live: https://valence-discovery.github.io/M...
Abstract: Meta-learning transfers knowledge across tasks and domains to learn new tasks efficiently, which has shown promise in drug discovery. However, the generalization ability of current meta-learning methods is limited by task heterogeneity and memorization. In this talk, I will first introduce two general principles to improve the generalization ability in meta-learning: organization and augmentation. Then, I will present several concrete few-shot drug discovery instantiations of using each principle. This includes algorithms to organize and adapt knowledge and a simple method for sufficiently overcoming task memorization. The remaining challenges and promising future research directions will also be discussed.
Speaker: Huaxiu Yao - https://huaxiuyao.mystrikingly.com/
Twitter Prudencio: https://twitter.com/tossouprudencio
Twitter Therence: https://twitter.com/Therence_mtl
Twitter Cas: https://twitter.com/cas_wognum
Twitter Valence Discovery: https://twitter.com/valence_ai
If you enjoyed this talk, consider joining the Molecular Modeling and Drug Discovery (M2D2) talks live: https://valence-discovery.github.io/M...
Abstract: One of the challenges with deep learning is lack of model interpretability. This is a significant drawback in the chemistry domain as lack of knowledge why a certain prediction was made dissuades chemists to trust predictions from deep learning. In this work we propose a method that can provide local explanations for arbitrary models with the use of molecular counterfactuals. These are sparse explanations composed of molecular structures. A counterfactual is an example as close to the original, but with a different outcome. Although relatively new to AI, counterfactual explanations are a mature topic in philosophy and mathematics. We use counterfactuals to answer, “what is the smallest change to the features that would alter the prediction". Our Molecular Model Agnostic Counterfactual Explanations (MMACE), method is built on the STONED (Nigam et al., 2021) algorithm to traverse a local chemical space around a given base molecule to identify counterfactuals. Further, we introduce an open-source software named “exmol” that implements the MMACE algorithm for generating counterfactual explanations.
Speaker: Geemi Wellawatte - https://geemi725.github.io/
Twitter Prudencio: https://twitter.com/tossouprudencio
Twitter Therence: https://twitter.com/Therence_mtl
Twitter Cas: https://twitter.com/cas_wognum
Twitter Valence Discovery: https://twitter.com/valence_ai
If you enjoyed this talk, consider joining the Molecular Modeling and Drug Discovery (M2D2) talks live: https://valence-discovery.github.io/M...
Abstract: Molecular design and synthesis planning are two critical steps in the process of molecular discovery that we propose to formulate as a single shared task of conditional synthetic pathway generation. We report an amortized approach to generate synthetic pathways as a Markov decision process conditioned on a target molecular embedding. This approach allows us to conduct synthesis planning in a bottom-up manner and design synthesizable molecules by decoding from optimized conditional codes, demonstrating the potential to solve both problems of design and synthesis simultaneously. The approach leverages neural networks to probabilistically model the synthetic trees, one reaction step at a time, according to reactivity rules encoded in a discrete action space of reaction templates. We train these networks on hundreds of thousands of artificial pathways generated from a pool of purchasable compounds and a list of expert-curated templates. We validate our method with (a) the recovery of molecules using conditional generation, (b) the identification of synthesizable structural analogs, and (c) the optimization of molecular structures given oracle functions relevant to drug discovery.
Speaker: Wenhao Gao - https://twitter.com/wenhaogao1
Twitter Prudencio: https://twitter.com/tossouprudencio
Twitter Therence: https://twitter.com/Therence_mtl
Twitter Cas: https://twitter.com/cas_wognum
Twitter Valence Discovery: https://twitter.com/valence_ai
If you enjoyed this talk, consider joining the Molecular Modeling and Drug Discovery (M2D2) talks live: https://valence-discovery.github.io/M...
Abstract: Machine learning for therapeutics offer incredible opportunities for expansion, innovation, and impact. Despite promises, many challenges exist. In this talk, the speaker will first highlight two high-impact but relatively understudied directions - ML-aided clinical trial design and low-data/cross-context biomedicine. Then, he will discuss challenges arising from therapeutics ML adoption in the wild, namely, generating actionable hypotheses and user interface with domain scientists. Lastly, challenges in infrastructure, such as data and benchmark, will be looked at.
Speaker: Kexin Huang - https://www.kexinhuang.com/
Twitter Prudencio: https://twitter.com/tossouprudencio
Twitter Therence: https://twitter.com/Therence_mtl
Twitter Cas: https://twitter.com/cas_wognum
Twitter Valence Discovery: https://twitter.com/valence_ai
If you enjoyed this talk, consider joining the Molecular Modeling and Drug Discovery (M2D2) talks live: https://valence-discovery.github.io/M...
Abstract: Molecular property prediction is one of the fastest-growing applications of deep learning with critical real-world impacts. Including 3D molecular structure as input to learned models improves their performance for many molecular tasks. However, this information is infeasible to compute at the scale required by several real-world applications. We propose pre-training a model to reason about the geometry of molecules given only their 2D molecular graphs. Using methods from self-supervised learning, we maximize the mutual information between 3D summary vectors and the representations of a Graph Neural Network (GNN) such that they contain latent 3D information. During fine-tuning on molecules with unknown geometry, the GNN still generates implicit 3D information and can use it to improve downstream tasks. We show that 3D pre-training provides significant improvements for a wide range of properties, such as a 22% average MAE reduction on eight quantum mechanical properties. Moreover, the learned representations can be effectively transferred between datasets in different molecular spaces.
Speaker: Hannes Stärk - https://hannes-stark.com/
Co-hosted by:
Prudencio - https://twitter.com/tossouprudencio
Cas - https://twitter.com/cas_wognum
Therence - https://twitter.com/Therence_mtl
Valence Discovery - https://twitter.com/valence_ai