Chem ML/AI/Datasets
Відкрити в Telegram
Daily articles and news from the field of machine learning in chemistry from the researchers of IGIC RAS @chemrussia For contact: @levkrasnov @st613laboratory @StasBezzubov
Показати більше848
Підписники
Немає даних24 години
+57 днів
+2130 днів
Архів дописів
В Nature вчера вышла статья, которая отвечает на простой вопрос на небольшой выборке: сколько улучшений даёт просто случайная замена одного атома?
https://www.nature.com/articles/s41586-026-11013-5
Взяли 18 исходных лигандов с активностью от 1 нМ до 43 мкМ на 6 мишенях (три GPCR, транспортёр SERT, β-лактамаза AmpC и макродомен Mac1) и систематически, без дизайна, меняли их по одному атому: H→CH₃, OH, Cl, F и ароматический C→N.
Ограничения были только на синтетическую доступность и цену (не дороже 400$ за соединение). Получилось 257 аналогов, и всё измерили в одной лаборатории: активность плюс 6 ADME свойств in vitro — стабильность в микросомах печени, проницаемость (PAMPA), стабильность в плазме, свободную фракцию в плазме, термодинамическую растворимость и ингибирование hERG.
Результаты:
— из 257 синтезированных аналогов 29 (11.3%) улучшили активность в 10 раз и более, а 69 (27%) — втрое и более. Улучшения нашлись у 5 мишеней из 6 и у 10 родительских молекул из 18.
— замены неравноценны. Чаще всего выигрывает метил, затем хлор. Фтор почти не улучшает активность, а замена ароматического C на N обходится в среднем в 11-кратную потерю.
— Рост активности в среднем сопровождается ухудшением фармакокинетических свойств. По данным авторов, кратность изменения активности слабо отрицательно связана со стабильностью в микросомах и со свободной фракцией в плазме. Все корреляции лежат в диапазоне |0.16|–|0.38|, то есть тренд есть, но он не жёсткий. Шесть измеренных свойств при этом почти ортогональны друг другу, так что улучшение одного не тянет за собой остальные.
Главное возражение от рецензентов в файле Peer Review: пять выбранных замен — это не случайная выборка из химического пространства, поэтому "случайность" условна.
Редакция JCIM вчера выпустила заметку, где активно призывает авторов баз данных, бенчмарков, софта и веб-сервисов, рассматривать формат Application Note как основной формат публикации их работы
https://doi.org/10.1021/acs.jcim.6c02729
Какие вопросы сразу возникают:
1) Для веб-сервисов требуется формальное заявление о поддержке в сопроводительном письме, где расписаны гарантии хостинга + открытый репозиторий, позволяющий другим самостоятельно развернуть или восстановить сервис, если хостинг отвалится. Интересно, а что будет, если по классике вэб-сервис сразу после публикации статьи перестанет работать? Потому что это очень частая история различных хемоинформатических демок.
2) Какое преимущество у этого для баз данных по сравнению с Scientific Data? За исключением того, что в JCIM не нужно платить APC в $2690 как в Data.
3) Авторы программных пакетов должны продемонстрировать жизнеспособность проекта через открытый репозиторий (GitHub и тд) показывающий активную разработку на протяжении минимум 6 (!!!) месяцев до момента подачи рукописи. Плюс к этому метрики использования, вроде статистики загрузок.
Мотивом они называют: защитить читателей от разрастания нейрослопа, чтобы публиковать инструменты, созданные надолго. Не очень понятно, что будет если просто нейрослопить в репозиторий 6 месяцев до подачи. 🎉
Сегодня MIT + Merck выпустили на ChemRxiv CheMeleon для реакций, предобученный на 2 млн реакциях.
https://chemrxiv.org/doi/full/10.26434/chemrxiv.15008692/v1
We introduce CheMeleon-Rxn, a pre-trained reaction graph neural network that adapts the CheMeleon framework to condensed graphs of reaction (CGR) and demonstrates that descriptor-regression is an effective pre-training strategy for low-data reaction property prediction. An O(10M)-parameter D-MPNN encoder was pre-trained on #2 million reactions by regressing onto dense reaction descriptors pooled from classical molecular descriptors, then fine-tuned on small reaction property prediction datasets. Across regression tasks spanning gasphase activation energies, reaction enthalpies, experimental yields, and rate coefficients, CheMeleon-Rxn is best or statistically tied-for-best on seven of eight tasks. It outperforms the same encoder trained from scratch, as well as classical-descriptor, fingerprint, and pre-trained-fingerprint baselines. We further tested the pre-training strategy across various graph neural network architectures and found that its benefit holds for other edge-aware backbones such as GINE, and that it is robust to the choice of descriptor set, pooling operation, and other pre-training hyperparameters. We visualized and analyzed the CheMeleon-Rxn embeddings, which reveal chemically coherent structure in the learned reaction representation space🔥Код полностью открытый: https://github.com/MSDLLCpapers/chemeleon-rxn А вот веса будут опубликованы позже:
The model weights will be released upon publication.
🎉 Мы выпустили BigSolDB v2.2 — новое крупное обновление нашего открытого набора данных по растворимости.
Что нового:
➕ 13 287 новых измерений растворимости — теперь их 125 752
➕ 102 новых растворённых вещества — всего 1627
➕ 155 новых литературных источников — всего 1842
🧪 Добавлена информация о методиках эксперимента: способ измерения, время установления равновесия, перемешивание, контроль температуры и число повторов.
💎 Добавлены данные о твёрдых формах: чистота образца, характеризация методами XRD/DSC и термодинамические свойства. В набор вошли 1044 экспериментальные температуры плавления (Tm) и 824 экспериментальные энтальпии плавления (ΔH_fus).
Для 14 813 измерений теперь явно указана твёрдая форма вещества.
В отдельном файле опубликованы 59 680 рассчитанных коэффициентов активности.
Новая версия доступна на Zenodo:
https://doi.org/10.5281/zenodo.22648301
Large language models as uncertainty-calibrated optimizers for experimental discovery 🔥
https://www.nature.com/articles/s42256-026-01283-z
Here we show how training language models through Bayesian objectives enables their use as reliable optimizers guided by natural language. Our approach, GOLLuM (Gaussian process Optimized LLMs), teaches LLMs from experimental outcomes under uncertainty, transforming their overconfidence from a fundamental flaw into a precise learning signal. This signal reshapes the LLM embeddings so that experiments with similar outcomes cluster together, revealing structure in the design space. Starting from only ten low-performing experiments, GOLLuM generalizes across 23 tasks in organic synthesis, materials science, process chemistry and molecular design, ranking first on average among all competing methods. It matches traditional Bayesian optimization with over 40% fewer experiments and nearly doubles the discovery of high-performing Buchwald–Hartwig reactions over expert quantum-chemical descriptors and state-of-the-art LLMs (43% versus 24–25%). More broadly, GOLLuM points to a different paradigm for specializing foundation models: not through more data but through richer, uncertainty-guided information.📕Nature Machine Intelligence (IF=)
Коллеги из Польши, США и Южной Кореи систематизировали информацию по более чем 3000 полных синтезов, включающих в себя практически 60000 индивидуальных органических реакций.
Столь большой датасет позволил на качественно новом уровне исследовать ретросинтетическую логику исследователей. В частности, было показано, что:
- Задача планирования синтеза практически никогда не может сводиться к последовательному пошаговому принятию решений
- Практически половина синтетических шагов являются формально структурно непродуктивными и напрямую не приближают к целевой структуре
- Для навигации в этих "синтетических плато" в алгоритмах ретросинтеза необходимы механизмы многоступенчатого рассуждения
- Разработан ряд таких многоступенчатых эвристик, призванных улучшить разработку моделей машинного обучения для органического синтеза
Источник: https://pubs.acs.org/jacsat/article/doi/10.1021/jacs.6c08434/5401944/AbSynth-A-Century-of-Total-Syntheses-as-a-Machine
📕Journal of the American Chemical Society (IF=16.6)
#dataset
Machine Learning for Synthetic Organic Chemistry: Methods, Applications, and Best Practices
https://doi.org/10.1002/anov.70032
The growing integration of artificial intelligence (AI) and machine learning (ML) is transforming experimental chemistry laboratories. Especially in synthetic chemistry, researchers routinely handle complex and high-dimensional data, fostering meaningful synergies between chemistry and data science. This review is intended as a practical overview that connects the everyday challenges of synthetic chemists with the digital tools available to address them. It does not seek to explain theoretical foundations of ML or to provide a comprehensive survey of all recent studies in the field. Rather, our goal is to highlight emerging technologies, discuss key considerations for their application, and present a selection of illustrative examples. To begin, we outline the prerequisites for successfully applying data science in synthetic chemistry. Next, we give a realistic overview of strategies and bottlenecks in predictive modeling of molecular properties, reaction outcomes and reaction conditions. We further highlight data-driven approaches that can be applied in the development of new chemical reactions and synthetic methodologies, including all relevant stages from reaction discovery and optimization to substrate scope evaluation and mechanistic analyses. Finally, we briefly discuss the transformative role of large language models and agentic workflows in synthetic chemistry, focusing on opportunities and challenges in the laboratories of the future.📕Angewandte Chemie Novit #review
QuantumPioneer: Scalable Generation of Quantum Chemical Data for Solution-Phase Hydrogen Transfer Reactions
https://doi.org/10.1021/jacs.6c10371
We introduce QuantumPioneer, an open-access reaction-centered QM database and workflow for small organic molecules, focused on peroxyl-mediated hydrogen atom transfer (HAT) and the corresponding homolytic bond dissociation reactions. QuantumPioneer contains 348,258 species (2-21 heavy atoms), 167,237 validated HAT transition states (TS) with corresponding reaction energies and homolytic bond dissociation energies (BDEs), and over 100 million COSMO-RS solvation free energies (ΔG_solv) and enthalpies (ΔH_solv) across 295 solvents. The workflow uses ωB97X-D/def2-SVP geometries, DLPNO-CCSD(T)-F12a/cc-pVTZ-F12 single-point energies, empirical thermochemical corrections, transition-state theory, and COSMO-RS BP-TZVPD-FINE solvation in a single high-throughput pipeline. Our benchmarks show reliable accuracy, with mean absolute errors (MAEs) compared to experimental and high-level QM reference data of 0.82 kcal/mol for gas-phase enthalpies of formation, 1.60 kcal/mol for C-H BDEs, 1.45 kcal/mol for HAT barriers, and 0.57 kcal/mol for ΔG_solv values. We demonstrate two predictive applications. First, we show that combining BDE and HAT-barrier models identifies experimentally observed oxidative degradation sites in drug-like molecules with a 91% top-5 hit rate and 82% site-level recall. Second, we show that a QM-parametrized Abraham model enables rapid solvation energy estimates at near-COSMO-RS accuracy within its training domain, reproducing computed ΔG_solv and ΔH_solv values with MAEs of 0.16 and 0.18 kcal/mol, respectively, though performance on experimental ΔG_solv values for unseen solutes was worse, with an MAE of 1.32 kcal/mol.📕Journal of the American Chemical Society (IF=16.6) #dataset
Новое исследование от AstraZeneca:
Do humans and large language models agree on the quality of synthesis plans? 🔥
https://doi.org/10.1039/d6dd00266h
Here, we investigated whether LLMs could mimic human experts on the challenging task of assessing retrosynthetic path feasibility in routes generated by a popular computer-aided synthesis planning tool (AiZynthFinder). We evaluated the agreement between LLMs and expert chemists on holistic evaluations of the proposed routes as well as the individual chemical reactions in them. We used four frontier LLMs, three proprietary models and one open-source model (Claude Opus 4.8, GPT-5.5, Llama 3.1 70B, and Gemini 3.1 Pro) and employed 17 expert chemists to grade 50 retrosynthetic paths. We found that, when provided with clearly defined reaction-evaluation categories, human experts tended to converge in their assessments. Among the evaluated models, Gemini 3.1 Pro achieved the highest agreement with the human majority vote. GPT-5.5 and Claude Opus 4.8 were comparatively more pessimistic, whereas Llama 3.1 70B showed a pronounced optimistic bias.📕Digital Discovery (IF=7.1)
Model Validation Protocols for Machine Learning in Small Molecule Drug Discovery 🔥
https://www.biorxiv.org/content/10.64898/2026.08.19.745868v1
Building on the collective expertise of a cross-industry consortium, we present a model validation framework consisting of five recommendations that would enable the community to move beyond aggregate metrics toward understanding where and why molecular property prediction models fail. We connect evaluation choices to real-world applications and case studies encountered in pharmaceutical research. The framework proposes splitting strategies that mimic realistic distribution shifts and expose common failure modes. We apply the recommended framework to a recently released dataset of absorption, distribution, metabolism, and excretion (ADME) properties.
Наконец-то вышла официальная статья по CheMeleon!
Deep Learning Foundation Models for Low-Data Regimes from Classical Molecular Descriptors
https://doi.org/10.1021/acs.jcim.6c01546
We propose pretraining on low-noise, calculable molecular descriptors via supervised learning to obtain rich, highly transferable molecular representations. We demonstrate this strategy with CheMeleon, a 10M parameter foundation model that enables directed message-passing neural networks to finally exceed the performance of classical methods in the low-data regime. We evaluate on 58 benchmark data sets spanning a range of properties relevant to small-molecule drug discovery, sourced from the industry-led Polaris benchmarking initiative. Rigorous statistical comparisons show that CheMeleon outperforms classical baselines like Random Forest on molecular fingerprints and descriptors, as well as existing foundation models. We open-source the CheMeleon model and the pretraining framework to encourage adoption and extension of this pretraining strategy across chemical sciences.📕Journal of Chemical Information and Modeling (IF=6.4)
Generative AI-Assisted Discovery of HPK1 Inhibitors
https://doi.org/10.1021/acs.jmedchem.6c01048
Generative artificial intelligence (AI) is now widely applied in medicinal chemistry, with detailed case studies emerging in the literature. Here, we describe an early application of REINVENT, AstraZeneca’s in-house generative molecular design platform, to identify new inhibitor scaffolds for hematopoietic progenitor kinase 1 (HPK1). REINVENT was deployed at two stages of the project to address distinct design objectives. For hit identification, transfer learning on kinase-active compounds, followed by reinforcement learning guided by QSAR-based scoring, led to the discovery of three active chemotypes. Subsequently, REINVENT was applied to scaffold hopping, using 3D pharmacophore and docking models as scoring functions, which enabled the identification of two additional active chemotypes. Optimization of one of these scaffolds delivered a compound with potent cellular activity, kinase selectivity, and favorable rat pharmacokinetics. These results demonstrate the value of integrating generative AI with medicinal chemistry expertise and support broader application of the approach in future discovery programs.📕Journal of Medicinal Chemistry (IF=7.3)
Research Assistant: AstraZeneca's Agentic System for R&D 🔥
https://arxiv.org/abs/2608.12395v1
We describe Research Assistant, an internal LLM-based system developed at AstraZeneca to help scientists and clinicians explore biomedical questions across a broad range of data sources. The system provides a chat-style interface that brings together evidence from scientific literature, knowledge graphs, chemistry, clinical trials, safety resources, expression data, and internal experimental systems. It supports both a fast mode for direct question answering and a multi-step mode for more complex research tasks. Responses are grounded in retrieved evidence and linked back to the original sources, allowing users to review and further explore the underlying data. In this technical note, we outline the system architecture, the main design choices behind the product, and lessons learned from deploying it at scale to support day-to-day R&D workflows across AstraZeneca.
Artificial intelligence in drug discovery — what it is, where we stand and the path forward
https://www.nature.com/articles/s41573-026-01496-2
In this Perspective we discuss potential reasons, including an insufficient focus on clinical translation during model development, difficulties with applying AI algorithms on conditional life science data, and insufficient problem definitions and the resulting underspecification of computational models for real-world use cases. ‘Technology push’ compared with ‘science pull’ is also likely to be an underlying factor, as well as the substantial time required to operationalize technical capabilities into systems that are sufficiently scaled and accessible for users. We provide recommendations for the development of AI in drug discovery with the aim of increasing its translational relevance. For example, benchmarking studies of AI tools in drug discovery need to move on from model validation and instead focus on their ability to improve decision making.📕 nature reviews drug discovery (IF = 91.2)
Benchmarking and developing large language models using one million clinical trials 🔥
https://www.nature.com/articles/s41746-026-02933-7
Here, we introduce TrialPanorama, a large-scale structured resource aggregating 1.6M clinical trial records from fifteen global registries linked with biomedical ontologies and literature. Using this resource, we construct 152K training and testing samples spanning eight clinical research tasks, including systematic review, trial design, and trial optimization. Benchmarking cutting-edge large language models (LLMs) reveals limited clinical reasoning capability in generic LLMs. In contrast, an 8B LLM developed on TrialPanorama using supervised fine-tuning and reinforcement learning outperforms 70B generic counterparts across all eight tasks, with relative improvements of 73.7, 67.6, 38.4, 37.8, 26.5, 20.7, 20.0, 18.1, and 5.2%, respectively. These results demonstrate the potential of domain-adapted AI to improve evidence synthesis and clinical trial design, establishing TrialPanorama as a foundation for scaling AI in clinical research.📕npj digital medicine (IF = 18.0)
The past, present and future of self-driving laboratories
https://www.nature.com/articles/s41570-026-00847-2
This Review traces the evolution of self-driving laboratories and examines the structural asymmetries that limit their maturation into shared scientific infrastructure. We frame the next phase of the field around three interdependent requirements: scalability, generalizability and provenance-complete experimentation. Realizing collective scientific superintelligence will require SDLs that reliably scale throughput, transfer workflows and learned models across laboratories and scientific domains and capture end-to-end experimental data and metadata from precursor preparation through synthesis, characterization and performance evaluation. Achieving this transition will depend on interoperable data and metadata standards, modular and integrable experimental hardware, and trustworthy artificial intelligence agents that reason under uncertainty within rigorous safety and ethical boundaries.📕Nature Reviews Chemistry (IF=50.3)
Do Language Models Dream of Binding Molecules? Benchmarking LLMs under Spatial Constraints🔥
https://arxiv.org/abs/2607.18144
In this work, we systematically analyze whether current general-purpose LLMs are capable of navigating complex 3D constraints compared to established baselines such as specialized diffusion models. We consider 3D ligand generation conditioned on protein pockets together with ligand- and interaction-derived spatial constraints, including anchor fragments, pharmacophore points, and mandatory pocket-ligand interactions. To enable this evaluation, we introduce 3D-Fit - a token-efficient benchmarking strategy for assessing LLM performance on multi-conditioned spatial molecule generation. Our findings reveal a clear pattern in LLM spatial capabilities: while they still lag behind state-of-the-art approaches, they are promising and can handle multiple spatial constraints simultaneously, enabling scaling to heterogeneous setups.
