en
Feedback
Chem ML/AI/Datasets

Chem ML/AI/Datasets

Open in Telegram

Daily articles and news from the field of machine learning in chemistry from the researchers of IGIC RAS @chemrussia For contact: @levkrasnov @st613laboratory @StasBezzubov

Show more
838
Subscribers
+424 hours
+67 days
+1230 days

Data loading in progress...

Similar Channels
No data
Any problems? Please refresh the page or contact our support manager.
Incoming and Outgoing Mentions
---
---
---
---
---
---
Attracting Subscribers
September '26
September '26
+6
in 0 channels
August '26
+20
in 0 channels
Get PRO
July '26
+30
in 0 channels
Get PRO
June '26
+16
in 1 channels
Get PRO
May '26
+10
in 1 channels
Get PRO
April '26
+20
in 1 channels
Get PRO
March '26
+12
in 0 channels
Get PRO
February '26
+16
in 0 channels
Get PRO
January '26
+13
in 0 channels
Get PRO
December '25
+22
in 1 channels
Get PRO
November '25
+26
in 0 channels
Get PRO
October '25
+10
in 0 channels
Get PRO
September '25
+54
in 0 channels
Get PRO
August '25
+27
in 0 channels
Get PRO
July '25
+189
in 0 channels
Get PRO
June '25
+4
in 0 channels
Get PRO
May '25
+41
in 0 channels
Get PRO
April '25
+158
in 0 channels
Get PRO
March '25
+19
in 0 channels
Get PRO
February '25
+24
in 0 channels
Get PRO
January '25
+33
in 1 channels
Get PRO
December '24
+349
in 0 channels
Date
Subscriber Growth
Mentions
Channels
07 September0
06 September+4
05 September0
04 September+2
03 September0
02 September0
01 September0
Channel Posts
Large language models as uncertainty-calibrated optimizers for experimental discovery 🔥 https://www.nature.com/articles/s42256-026-01283-z
Here we show how training language models through Bayesian objectives enables their use as reliable optimizers guided by natural language. Our approach, GOLLuM (Gaussian process Optimized LLMs), teaches LLMs from experimental outcomes under uncertainty, transforming their overconfidence from a fundamental flaw into a precise learning signal. This signal reshapes the LLM embeddings so that experiments with similar outcomes cluster together, revealing structure in the design space. Starting from only ten low-performing experiments, GOLLuM generalizes across 23 tasks in organic synthesis, materials science, process chemistry and molecular design, ranking first on average among all competing methods. It matches traditional Bayesian optimization with over 40% fewer experiments and nearly doubles the discovery of high-performing Buchwald–Hartwig reactions over expert quantum-chemical descriptors and state-of-the-art LLMs (43% versus 24–25%). More broadly, GOLLuM points to a different paradigm for specializing foundation models: not through more data but through richer, uncertainty-guided information.
📕Nature Machine Intelligence (IF=)

2
Коллеги из Польши, США и Южной Кореи систематизировали информацию по более чем 3000 полных синтезов, включающих в себя практи
Коллеги из Польши, США и Южной Кореи систематизировали информацию по более чем 3000 полных синтезов, включающих в себя практически 60000 индивидуальных органических реакций. Столь большой датасет позволил на качественно новом уровне исследовать ретросинтетическую логику исследователей. В частности, было показано, что: - Задача планирования синтеза практически никогда не может сводиться к последовательному пошаговому принятию решений - Практически половина синтетических шагов являются формально структурно непродуктивными и напрямую не приближают к целевой структуре - Для навигации в этих "синтетических плато" в алгоритмах ретросинтеза необходимы механизмы многоступенчатого рассуждения - Разработан ряд таких многоступенчатых эвристик, призванных улучшить разработку моделей машинного обучения для органического синтеза Источник: https://pubs.acs.org/jacsat/article/doi/10.1021/jacs.6c08434/5401944/AbSynth-A-Century-of-Total-Syntheses-as-a-Machine 📕Journal of the American Chemical Society (IF=16.6) #dataset
467
3
Machine Learning for Synthetic Organic Chemistry: Methods, Applications, and Best Practices https://doi.org/10.1002/anov.70032 The growing integration of artificial intelligence (AI) and machine learning (ML) is transforming experimental chemistry laboratories. Especially in synthetic chemistry, researchers routinely handle complex and high-dimensional data, fostering meaningful synergies between chemistry and data science. This review is intended as a practical overview that connects the everyday challenges of synthetic chemists with the digital tools available to address them. It does not seek to explain theoretical foundations of ML or to provide a comprehensive survey of all recent studies in the field. Rather, our goal is to highlight emerging technologies, discuss key considerations for their application, and present a selection of illustrative examples. To begin, we outline the prerequisites for successfully applying data science in synthetic chemistry. Next, we give a realistic overview of strategies and bottlenecks in predictive modeling of molecular properties, reaction outcomes and reaction conditions. We further highlight data-driven approaches that can be applied in the development of new chemical reactions and synthetic methodologies, including all relevant stages from reaction discovery and optimization to substrate scope evaluation and mechanistic analyses. Finally, we briefly discuss the transformative role of large language models and agentic workflows in synthetic chemistry, focusing on opportunities and challenges in the laboratories of the future. 📕Angewandte Chemie Novit #review
231
4
QuantumPioneer: Scalable Generation of Quantum Chemical Data for Solution-Phase Hydrogen Transfer Reactions https://doi.org/1
QuantumPioneer: Scalable Generation of Quantum Chemical Data for Solution-Phase Hydrogen Transfer Reactions https://doi.org/10.1021/jacs.6c10371 We introduce QuantumPioneer, an open-access reaction-centered QM database and workflow for small organic molecules, focused on peroxyl-mediated hydrogen atom transfer (HAT) and the corresponding homolytic bond dissociation reactions. QuantumPioneer contains 348,258 species (2-21 heavy atoms), 167,237 validated HAT transition states (TS) with corresponding reaction energies and homolytic bond dissociation energies (BDEs), and over 100 million COSMO-RS solvation free energies (ΔG_solv) and enthalpies (ΔH_solv) across 295 solvents. The workflow uses ωB97X-D/def2-SVP geometries, DLPNO-CCSD(T)-F12a/cc-pVTZ-F12 single-point energies, empirical thermochemical corrections, transition-state theory, and COSMO-RS BP-TZVPD-FINE solvation in a single high-throughput pipeline. Our benchmarks show reliable accuracy, with mean absolute errors (MAEs) compared to experimental and high-level QM reference data of 0.82 kcal/mol for gas-phase enthalpies of formation, 1.60 kcal/mol for C-H BDEs, 1.45 kcal/mol for HAT barriers, and 0.57 kcal/mol for ΔG_solv values. We demonstrate two predictive applications. First, we show that combining BDE and HAT-barrier models identifies experimentally observed oxidative degradation sites in drug-like molecules with a 91% top-5 hit rate and 82% site-level recall. Second, we show that a QM-parametrized Abraham model enables rapid solvation energy estimates at near-COSMO-RS accuracy within its training domain, reproducing computed ΔG_solv and ΔH_solv values with MAEs of 0.16 and 0.18 kcal/mol, respectively, though performance on experimental ΔG_solv values for unseen solutes was worse, with an MAE of 1.32 kcal/mol. 📕Journal of the American Chemical Society (IF=16.6) #dataset
366
5
Новое исследование от AstraZeneca: Do humans and large language models agree on the quality of synthesis plans? 🔥 https://do
Новое исследование от AstraZeneca: Do humans and large language models agree on the quality of synthesis plans? 🔥 https://doi.org/10.1039/d6dd00266h Here, we investigated whether LLMs could mimic human experts on the challenging task of assessing retrosynthetic path feasibility in routes generated by a popular computer-aided synthesis planning tool (AiZynthFinder). We evaluated the agreement between LLMs and expert chemists on holistic evaluations of the proposed routes as well as the individual chemical reactions in them. We used four frontier LLMs, three proprietary models and one open-source model (Claude Opus 4.8, GPT-5.5, Llama 3.1 70B, and Gemini 3.1 Pro) and employed 17 expert chemists to grade 50 retrosynthetic paths. We found that, when provided with clearly defined reaction-evaluation categories, human experts tended to converge in their assessments. Among the evaluated models, Gemini 3.1 Pro achieved the highest agreement with the human majority vote. GPT-5.5 and Claude Opus 4.8 were comparatively more pessimistic, whereas Llama 3.1 70B showed a pronounced optimistic bias. 📕Digital Discovery (IF=7.1)
363
6
Model Validation Protocols for Machine Learning in Small Molecule Drug Discovery 🔥 https://www.biorxiv.org/content/10.64898/
Model Validation Protocols for Machine Learning in Small Molecule Drug Discovery 🔥 https://www.biorxiv.org/content/10.64898/2026.08.19.745868v1 Building on the collective expertise of a cross-industry consortium, we present a model validation framework consisting of five recommendations that would enable the community to move beyond aggregate metrics toward understanding where and why molecular property prediction models fail. We connect evaluation choices to real-world applications and case studies encountered in pharmaceutical research. The framework proposes splitting strategies that mimic realistic distribution shifts and expose common failure modes. We apply the recommended framework to a recently released dataset of absorption, distribution, metabolism, and excretion (ADME) properties.
332
7
No text...
320
8
Наконец-то вышла официальная статья по CheMeleon! Deep Learning Foundation Models for Low-Data Regimes from Classical Molecular Descriptors https://doi.org/10.1021/acs.jcim.6c01546 We propose pretraining on low-noise, calculable molecular descriptors via supervised learning to obtain rich, highly transferable molecular representations. We demonstrate this strategy with CheMeleon, a 10M parameter foundation model that enables directed message-passing neural networks to finally exceed the performance of classical methods in the low-data regime. We evaluate on 58 benchmark data sets spanning a range of properties relevant to small-molecule drug discovery, sourced from the industry-led Polaris benchmarking initiative. Rigorous statistical comparisons show that CheMeleon outperforms classical baselines like Random Forest on molecular fingerprints and descriptors, as well as existing foundation models. We open-source the CheMeleon model and the pretraining framework to encourage adoption and extension of this pretraining strategy across chemical sciences. 📕Journal of Chemical Information and Modeling (IF=6.4)
329
9
Generative AI-Assisted Discovery of HPK1 Inhibitors https://doi.org/10.1021/acs.jmedchem.6c01048 Generative artificial intelligence (AI) is now widely applied in medicinal chemistry, with detailed case studies emerging in the literature. Here, we describe an early application of REINVENT, AstraZeneca’s in-house generative molecular design platform, to identify new inhibitor scaffolds for hematopoietic progenitor kinase 1 (HPK1). REINVENT was deployed at two stages of the project to address distinct design objectives. For hit identification, transfer learning on kinase-active compounds, followed by reinforcement learning guided by QSAR-based scoring, led to the discovery of three active chemotypes. Subsequently, REINVENT was applied to scaffold hopping, using 3D pharmacophore and docking models as scoring functions, which enabled the identification of two additional active chemotypes. Optimization of one of these scaffolds delivered a compound with potent cellular activity, kinase selectivity, and favorable rat pharmacokinetics. These results demonstrate the value of integrating generative AI with medicinal chemistry expertise and support broader application of the approach in future discovery programs. 📕Journal of Medicinal Chemistry (IF=7.3)
354
10
Research Assistant: AstraZeneca's Agentic System for R&D 🔥 https://arxiv.org/abs/2608.12395v1 We describe Research Assistant, an internal LLM-based system developed at AstraZeneca to help scientists and clinicians explore biomedical questions across a broad range of data sources. The system provides a chat-style interface that brings together evidence from scientific literature, knowledge graphs, chemistry, clinical trials, safety resources, expression data, and internal experimental systems. It supports both a fast mode for direct question answering and a multi-step mode for more complex research tasks. Responses are grounded in retrieved evidence and linked back to the original sources, allowing users to review and further explore the underlying data. In this technical note, we outline the system architecture, the main design choices behind the product, and lessons learned from deploying it at scale to support day-to-day R&D workflows across AstraZeneca.
471
11
Artificial intelligence in drug discovery — what it is, where we stand and the path forward https://www.nature.com/articles/s41573-026-01496-2 In this Perspective we discuss potential reasons, including an insufficient focus on clinical translation during model development, difficulties with applying AI algorithms on conditional life science data, and insufficient problem definitions and the resulting underspecification of computational models for real-world use cases. ‘Technology push’ compared with ‘science pull’ is also likely to be an underlying factor, as well as the substantial time required to operationalize technical capabilities into systems that are sufficiently scaled and accessible for users. We provide recommendations for the development of AI in drug discovery with the aim of increasing its translational relevance. For example, benchmarking studies of AI tools in drug discovery need to move on from model validation and instead focus on their ability to improve decision making. 📕 nature reviews drug discovery (IF = 91.2)
496
12
Benchmarking and developing large language models using one million clinical trials 🔥 https://www.nature.com/articles/s41746-026-02933-7 Here, we introduce TrialPanorama, a large-scale structured resource aggregating 1.6M clinical trial records from fifteen global registries linked with biomedical ontologies and literature. Using this resource, we construct 152K training and testing samples spanning eight clinical research tasks, including systematic review, trial design, and trial optimization. Benchmarking cutting-edge large language models (LLMs) reveals limited clinical reasoning capability in generic LLMs. In contrast, an 8B LLM developed on TrialPanorama using supervised fine-tuning and reinforcement learning outperforms 70B generic counterparts across all eight tasks, with relative improvements of 73.7, 67.6, 38.4, 37.8, 26.5, 20.7, 20.0, 18.1, and 5.2%, respectively. These results demonstrate the potential of domain-adapted AI to improve evidence synthesis and clinical trial design, establishing TrialPanorama as a foundation for scaling AI in clinical research. 📕npj digital medicine (IF = 18.0)
522
13
No text...
409
14
The past, present and future of self-driving laboratories https://www.nature.com/articles/s41570-026-00847-2 This Review traces the evolution of self-driving laboratories and examines the structural asymmetries that limit their maturation into shared scientific infrastructure. We frame the next phase of the field around three interdependent requirements: scalability, generalizability and provenance-complete experimentation. Realizing collective scientific superintelligence will require SDLs that reliably scale throughput, transfer workflows and learned models across laboratories and scientific domains and capture end-to-end experimental data and metadata from precursor preparation through synthesis, characterization and performance evaluation. Achieving this transition will depend on interoperable data and metadata standards, modular and integrable experimental hardware, and trustworthy artificial intelligence agents that reason under uncertainty within rigorous safety and ethical boundaries. 📕Nature Reviews Chemistry (IF=50.3)
401
15
No text...
454
16
Do Language Models Dream of Binding Molecules? Benchmarking LLMs under Spatial Constraints🔥 https://arxiv.org/abs/2607.18144
Do Language Models Dream of Binding Molecules? Benchmarking LLMs under Spatial Constraints🔥 https://arxiv.org/abs/2607.18144 In this work, we systematically analyze whether current general-purpose LLMs are capable of navigating complex 3D constraints compared to established baselines such as specialized diffusion models. We consider 3D ligand generation conditioned on protein pockets together with ligand- and interaction-derived spatial constraints, including anchor fragments, pharmacophore points, and mandatory pocket-ligand interactions. To enable this evaluation, we introduce 3D-Fit - a token-efficient benchmarking strategy for assessing LLM performance on multi-conditioned spatial molecule generation. Our findings reveal a clear pattern in LLM spatial capabilities: while they still lag behind state-of-the-art approaches, they are promising and can handle multiple spatial constraints simultaneously, enabling scaling to heterogeneous setups.
457
17
Yield Smarter, Not Harder: Good Practices for Machine Learning of Reaction Outcomes https://doi.org/10.1021/jacs.6c02213 Reaction yield prediction is a longstanding challenge in synthetic chemistry, with broad implications for route planning, scalability, and high-throughput experimentation (HTE). While recent machine learning (ML) approaches have demonstrated promise in modeling reactivity, they often use complex descriptors or deep architectures that are computationally expensive and limit interpretability and scalability. Here, we assess how much information is stored in simpler descriptors and whether model accuracy is improved by increasing the complexity of the descriptors. Using classical ML models trained on descriptors with different complexity levels, we benchmark predictive performance on four publicly available HTE data sets covering three diverse reaction data sets: Buchwald–Hartwig (BH) amination, Suzuki–Miyaura (SM) coupling, and the silicon–amine protocol (SLAP). Our evaluation furthermore discusses (1) generalization via component-wise data splitting, (2) robustness through external validation across data sets, and (3) performance across asymmetric yield distributions characteristic of HTE data. Contrary to conventional expectations, we find that simpler models with interpretable features can achieve competitive performance under rigorous validation protocols. Based on our findings, we formulate good practices for future studies in this area. For example, comparison to low-cost baseline models should become a requirement for future ML studies for reaction-yield prediction. 📕Journal of the American Chemical Society (IF=16.6) #method
449
18
Data-Driven Insights into Ionic Conductivity in High-Dimensional Sodium Battery Electrolytes https://doi.org/10.1021/acsenergylett.6c01023 The discovery of advanced battery electrolytes is challenged by the vast compositional space of multi-component liquid formulations. Here, we introduce the ELectrolyte Laboratory for Integrated Experimentation (ELLIE), an automated platform that combines electrolyte formulation and impedance spectroscopy to map ionic conductivity across high-dimensional sodium electrolytes containing up to five salts and 15 solvents, generating an experimental dataset spanning nearly two orders of magnitude in conductivity. 23Na NMR, Raman spectroscopy, and viscosity measurements on a subset of electrolytes at a fixed salt concentration reveal that conductivity is jointly influenced by Na+ solvation strength, ion association, and solvent dynamics and positively correlates with inverse viscosity. Conductivity estimates based on the Nernst–Einstein relation captures broad concentration and viscosity relationships but do not extrapolate well across compositionally diverse electrolytes. Random forest modeling identifies lower solvent molecular weight as the dominant descriptor of high conductivity. Together, these results establish solvent molecular size as a physically interpretable descriptor of ion transport and demonstrate how automated experimentation can accelerate data-driven electrolyte optimization across complex compositional spaces. 📕ACS Energy Letters (IF=17.5) #method
394
19
https://chemrxiv.org/doi/10.26434/chemrxiv.15006080/v1 Коллеги из 🏛ИОНХ РАН, 🏛ИНЭОС РАН и Университета Барселоны выложили п
https://chemrxiv.org/doi/10.26434/chemrxiv.15006080/v1 Коллеги из 🏛ИОНХ РАН, 🏛ИНЭОС РАН и Университета Барселоны выложили препринт про BLIND (Bimodal Learning from Imperfect NMR Data) — трансформер, который переводит спектры ¹H и/или ¹³C ЯМР напрямую в молекулярную структуру (SMILES). Что делает работу интересной: Модель не получает ни брутто-формулу, ни набор возможных фрагментов, ни элементный состав — только то, как спектр записан в статье («7.85–7.81 (m, 3H)…»). Это гораздо честнее большинства предыдущих подходов. Учили на реальных, "грязных" данных. Стартовали с 7.5 млн (!!) записей спектров, извлечённых из литературы через базу OdanChem. После минимальной фильтрации (только CDCl₃ в качестве растворителя, удаление дубликатов и явных выбросов) осталось 5.5 млн спектров для 2.6 млн уникальных структур. Разбивка по уникальным структурам, строго без утечек между наборами. Обучающая выборка состояла из 1.82 млн уникальных SMILES, 1.89 млн ¹³C- и 2.0 млн ¹H-спектров. При этом избыточности почти нет — в среднем 1.2 спектра на молекулу, то есть модель учится обобщать буквально с одного измерения на соединение. Выборку намеренно не чистили до нейтральной органики, как это обычно бывает в хемоинформатике, и оставили редкие элементы (Se, Fe, Te…), соли и комплексы. Та же модель, обученная только на идеальном подмножестве (2 млн записей), даёт 48.4%, а обученная на всём хаосе — 56.7% на тех же тестах. Реальный шум дает разнообразие, которое помогает обобщать. 🔥Онлайн-версия доступна на https://odanchem.org/predict-multimodal-compound-search
726
20
Strategies for Identifying Molecules of Interest in Large Chemical Spaces 🔥 https://pubs.acs.org/doi/10.1021/acs.jcim.6c01496 Searching in ultralarge Chemical Spaces with known 2D similarity metrics, like fingerprint-based Tanimoto, substructure, or pharmacophore similarity searches, contains pitfalls due to the representation of molecules as synthons with connectivity rules. Applied to a set of almost 3000 drug-relevant queries we analyzed the ability of similarity search methods to retrieve analog compounds from Chemical Spaces, and how to best approach typical use cases in early phase drug discovery. Distinct characteristics of each similarity metric suggest orthogonal complementarity, enabling a versatile framework to diverse challenges present in hit discovery and lead expansion campaigns. 📕Journal of Chemical Information and Modeling (IF=6.4)
430