Chem ML/AI/Datasets
رفتن به کانال در Telegram
Daily articles and news from the field of machine learning in chemistry from the researchers of IGIC RAS @chemrussia For contact: @levkrasnov @st613laboratory @StasBezzubov
نمایش بیشتر857
مشترکین
اطلاعاتی وجود ندارد24 ساعت
-27 روز
+1830 روز
در حال بارگیری داده...
کانالهای مشابه
هیچ دادهای
مشکلی وجود دارد؟ لطفاً صفحه را تازه کنید یا با مدیر پشتیبانی ما تماس بگیرید.
ابر برچسبها
اشارات ورودی و خروجی
---
---
---
---
---
---
جذب مشترکین
اکتبر '26اکتبر '26
اکتبر '26
+2
در 0 کانالها
سپتامبر '26
+34
در 0 کانالها
Get PRO
اوت '26
+20
در 0 کانالها
Get PRO
ژوئیه '26
+30
در 0 کانالها
Get PRO
ژوئن '26
+16
در 1 کانالها
Get PRO
مه '26
+10
در 1 کانالها
Get PRO
آوریل '26
+20
در 1 کانالها
Get PRO
مارس '26
+12
در 0 کانالها
Get PRO
فوریه '26
+16
در 0 کانالها
Get PRO
ژانویه '26
+13
در 0 کانالها
Get PRO
دسامبر '25
+22
در 1 کانالها
Get PRO
نوامبر '25
+26
در 0 کانالها
Get PRO
اکتبر '25
+10
در 0 کانالها
Get PRO
سپتامبر '25
+54
در 0 کانالها
Get PRO
اوت '25
+27
در 0 کانالها
Get PRO
ژوئیه '25
+189
در 0 کانالها
Get PRO
ژوئن '25
+4
در 0 کانالها
Get PRO
مه '25
+41
در 0 کانالها
Get PRO
آوریل '25
+158
در 0 کانالها
Get PRO
مارس '25
+19
در 0 کانالها
Get PRO
فوریه '25
+24
در 0 کانالها
Get PRO
ژانویه '25
+33
در 1 کانالها
Get PRO
دسامبر '24
+349
در 0 کانالها
| تاریخ | رشد مشترکین | اشارات | کانالها | |
| 08 اکتبر | 0 | |||
| 07 اکتبر | 0 | |||
| 06 اکتبر | +1 | |||
| 05 اکتبر | 0 | |||
| 04 اکتبر | 0 | |||
| 03 اکتبر | +1 | |||
| 02 اکتبر | 0 | |||
| 01 اکتبر | 0 |
پستهای کانال
Сегодня у нас вышла новая работа:
Beyond Organic Molecules: MetalLipoDB, a Lipophilicity Dataset for Metal Complexes and Machine Learning Benchmarks
https://doi.org/10.26434/chemrxiv.15010021/v1
Работа продолжает систематизацию данных по комплексам металлов, начатую в MetalCytoToxDB. На этот раз предметом стала липофильность. Для органических соединений существуют большие открытые базы logP, а предсказание липофильности входит в стандартные ML-бенчмарки. Для комплексов металлов экспериментальные значения разбросаны по сотням статей, и существующие подборки покрывают преимущественно платину.
В работе представлена MetalLipoDB — крупнейшая на сегодня база экспериментальных значений липофильности комплексов металлов.
Объём базы: 5158 значений logP/logD для 4714 комплексов 23 металлов, вручную извлечённых из 1192 статей. Для каждой записи указаны условия измерения: состав водной фазы, pH, метод детекции.
В базу включены 546 многоядерных комплексов.
Оценена воспроизводимость измерений. Для одного и того же комплекса значения из двух разных статей расходятся в среднем на 0.56 логарифмических единиц. Эта величина задаёт практический предел точности моделей, обученных на литературных данных.
На основе базы обучены ML-модели (TabPFN, LightGBM, Chemprop, CheMeleon). Для комплексов из статей, отсутствующих в обучающей выборке, ошибка составляет MAE = 0.72-0.75. Если в серии уже измерены три комплекса, привязка предсказаний к этим измерениям снижает ошибку с 0.73 до 0.51.
❗️В разделе "Limitations and future work" перечислены ограничения подхода:
— Значения в базе представляют собой кажущиеся коэффициенты распределения в условиях, указанных авторами
— Многие комплексы не инертны в воде: в частности, хлоридные комплексы платины и рутения подвергаются замещению лиганда на воду
— Включены только измерения непосредственно распределения между октанолом и водной фазой, хроматографические индексы липофильности не рассматриваются
— Модели не учитывают стереохимию
— Для металлов, отсутствующих в обучающей выборке, ошибка возрастает до 0.83-0.89
Датасет на Zenodo: https://doi.org/10.5281/zenodo.23084573
| 2 | Ложился спать, открыл твиттер, почитал и не могу уснуть… OpenAI опубликовала разом 722 манускрипта с решениями 722 открытых математических задач, которые выполнила их (пока ещё) закрытая модель, каждое за ~3 часа. Что ж, ждем грозных писем от Теренса Тао и десятков филдсовских/абелевских лауреатов. Но их не жалко, они свою жизнь пожили, жалко того бедного неназванного аспиранта, который, возможно, последние три года сидел и корпел над решением одной из этих задач и вот-вот готовился опубликовать статью, которая уже никому не нужна, как и его занятия фундаментальной математикой «по старинке». Ожидаю резкий спад интереса к поступлению в магистратуры/аспирантуры на матфаки по всему миру, пока светлые головы не переизобретут математический post-graduate и не сформулируют: «что такое занятия математикой в эпоху ИИ?». А что химики? За химиками тоже «придут», обязательно придут, «но потом». | 199 |
| 3 | Вчера вышла совместная работа AMD Silo AI и AstraZeneca
From Benchmark to Bench: Can Agents Survive Real-World Drug Discovery?🔥
https://arxiv.org/abs/2610.06411
We developed MAGI, an open modular agent that authors objectives, launches and monitors optimization, interprets structure-activity relationships, and revises its strategy accordingly. MAGI generates molecules either directly through the LLM or by delegating to REINVENT 4, with scoring services interchangeable behind a common contract. We tested it across nine retrospective lead-optimization campaigns from three pharmaceutical companies, replayed under fixed temporal cutoffs. Both routes produced valid structures: LLM proposals stayed closer to local chemistry and reached comparable or higher primary activity in fewer operations, whereas REINVENT explored broader chemical space.
Whether a campaign met its objective depended on the predictive models, not on the generation route: attainment followed model accuracy on the chemistry proposed, dropping once that chemistry moved outside the model's applicability domain.
Separately, a blinded evaluation asked whether the MAGI's output could pass as expert work: chemists were not able to discriminate agentic proposals from held-out compounds, and judged the SAR reasoning broadly plausible yet incomplete. Together, these results position MAGI as a coordination layer pluggable into existing computational chemistry workflows. The ceiling on real projects, however, remains currently set by scorer applicability rather than by tool orchestration. | 230 |
| 4 | Atoms to Processes: The Role of Artificial Intelligence and Machine Learning in Chemical Engineering
https://doi.org/10.1021/acs.iecr.6c01806
This Perspective provides an account of recent methodological developments and representative applications, emphasizing how AI/ML tools are being integrated with first-principles models to address challenges of predictive accuracy, data scarcity, extrapolation, interpretability, and model life cycle management.
Across domains, a unifying trend is the move away from purely black-box approaches toward hybrid and physics-informed frameworks that explicitly respect conservation laws, thermodynamic consistency, and known structural constraints.
These approaches not only improve robustness and reliability but also enable meaningful human–AI collaboration by providing information at an appropriate level of abstraction for the task and decision context. We conclude that AI and ML are not replacing the core principles of chemical engineering; rather, they are amplifying them.
📕Industrial & Engineering Chemistry Research (IF = 3.9) | 226 |
| 5 | ElemeNet: Multiscale Molecular Machine Learning with Uncertainty Quantification across the Periodic Table
https://doi.org/10.1021/acs.jcim.6c02178
This work introduces ElemeNet, a unified, general-purpose software package for molecular machine learning. The ElemeNet software package enables the training of advanced ML models for diverse properties and data sets with an enlarged range of elemental compositions.
We define molecular representations compatible with elements 1–100, supporting diverse organometallic and biological systems in addition to organic chemistry already well served by the Chemprop ML toolkit. As well as more common atom-, bond-, and molecule-level predictions, we introduce and allow for moiety-level predictions.
We also natively define optional conditioning on charge and spin states. Advanced E(3)-equivariant and transformer architectures are supported in addition to 2D models, with all classes including built-in uncertainty quantification through deterministic and statistical measures. We benchmark our protocols for ML model training against representative data sets from organic, inorganic, coordination, and biological chemistry, achieving competitive and SOTA performance relative to literature baselines and favorable scaling to millions of molecules.
📕Journal of Chemical Information and Modeling (IF=6.4) | 327 |
| 6 | Разбирал закладки и понял, что так и не запостил эту статью, хотя она вышла почти год назад. Исправляюсь. Это статья в JCIM о том, как честно сравнивать ML-модели в drug discovery.
Practically Significant Method Comparison Protocols for Machine Learning in Small Molecule Drug Discovery🔥
https://doi.org/10.1021/acs.jcim.5c01609
This paper proposes a set of guidelines to incentivize rigorous and domain-appropriate techniques for method comparison tailored to small molecule property modeling. These guidelines, accompanied by annotated examples using open-source software tools, lay a foundation for robust ML benchmarking and thus the development of more impactful methods.
Почти все авторы из индустрии: J&J, Novartis, Bayer, Merck, Pfizer, AstraZeneca, Recursion и т.д. Последний автор Pat Walters теперь Chief Scientist в OpenADMET, так что эти правила, скорее всего, дойдут и до открытых бенчмарков. | 394 |
| 7 | Доброе утро начинается с принятия работы на NeurIPS 2026 (Evaluations & Datasets Track), в которой нас пригласили поучаствовать замечательные коллеги из Индии (IIT Delhi) 🎉
We are delighted to inform you that your submission, "SC³: A Multi-Solvent Solubility Challenge and Benchmark", has been accepted at NeurIPS 2026 ED Track. Congratulations!
There were 3,757 valid paper submissions with a PDF to the NeurIPS ED Track this year, of which the program committee accepted 971 (25.8%) papers in total.
📄 arxiv.org/abs/2606.07656 | 385 |
| 8 | Multi-objective optimization in the context of generative chemistry 🔥
https://www.nature.com/articles/s41467-026-77324-3
Designing new medicines requires balancing multiple competing objectives, including biological activity, pharmacokinetics and safety. As generative chemistry expands our ability to explore chemical space, the choice of how objectives are defined, combined and evaluated becomes increasingly important. In this Perspective, we discuss the role of multi-objective optimization in molecular generation, highlighting practical considerations, methodological challenges and emerging opportunities for translating these approaches into effective drug discovery tools.
📕 Nature Communications (IF=18.1) | 409 |
| 9 | The evolving landscape of drug targets
https://www.nature.com/articles/s41573-026-01530-3
In this Review, we map and quantify key elements of these changes over the past 25 years, such as trends in the number and type of targets through which drugs mediate their therapeutic effects, which now include 686 biomolecules modulated by 1,702 drugs.
📕Nature Reviews Drug Discovery (IF=91.2) | 513 |
| 10 | В Nature вчера вышла статья, которая отвечает на простой вопрос на небольшой выборке: сколько улучшений даёт просто случайная замена одного атома?
https://www.nature.com/articles/s41586-026-11013-5
Взяли 18 исходных лигандов с активностью от 1 нМ до 43 мкМ на 6 мишенях (три GPCR, транспортёр SERT, β-лактамаза AmpC и макродомен Mac1) и систематически, без дизайна, меняли их по одному атому: H→CH₃, OH, Cl, F и ароматический C→N.
Ограничения были только на синтетическую доступность и цену (не дороже 400$ за соединение). Получилось 257 аналогов, и всё измерили в одной лаборатории: активность плюс 6 ADME свойств in vitro — стабильность в микросомах печени, проницаемость (PAMPA), стабильность в плазме, свободную фракцию в плазме, термодинамическую растворимость и ингибирование hERG.
Результаты:
— из 257 синтезированных аналогов 29 (11.3%) улучшили активность в 10 раз и более, а 69 (27%) — втрое и более. Улучшения нашлись у 5 мишеней из 6 и у 10 родительских молекул из 18.
— замены неравноценны. Чаще всего выигрывает метил, затем хлор. Фтор почти не улучшает активность, а замена ароматического C на N обходится в среднем в 11-кратную потерю.
— Рост активности в среднем сопровождается ухудшением фармакокинетических свойств. По данным авторов, кратность изменения активности слабо отрицательно связана со стабильностью в микросомах и со свободной фракцией в плазме. Все корреляции лежат в диапазоне |0.16|–|0.38|, то есть тренд есть, но он не жёсткий. Шесть измеренных свойств при этом почти ортогональны друг другу, так что улучшение одного не тянет за собой остальные.
Главное возражение от рецензентов в файле Peer Review: пять выбранных замен — это не случайная выборка из химического пространства, поэтому "случайность" условна. | 547 |
| 11 | Редакция JCIM вчера выпустила заметку, где активно призывает авторов баз данных, бенчмарков, софта и веб-сервисов, рассматривать формат Application Note как основной формат публикации их работы
https://doi.org/10.1021/acs.jcim.6c02729
Какие вопросы сразу возникают:
1) Для веб-сервисов требуется формальное заявление о поддержке в сопроводительном письме, где расписаны гарантии хостинга + открытый репозиторий, позволяющий другим самостоятельно развернуть или восстановить сервис, если хостинг отвалится. Интересно, а что будет, если по классике вэб-сервис сразу после публикации статьи перестанет работать? Потому что это очень частая история различных хемоинформатических демок.
2) Какое преимущество у этого для баз данных по сравнению с Scientific Data? За исключением того, что в JCIM не нужно платить APC в $2690 как в Data.
3) Авторы программных пакетов должны продемонстрировать жизнеспособность проекта через открытый репозиторий (GitHub и тд) показывающий активную разработку на протяжении минимум 6 (!!!) месяцев до момента подачи рукописи. Плюс к этому метрики использования, вроде статистики загрузок.
Мотивом они называют: защитить читателей от разрастания нейрослопа, чтобы публиковать инструменты, созданные надолго. Не очень понятно, что будет если просто нейрослопить в репозиторий 6 месяцев до подачи. 🎉 | 411 |
| 12 | Сегодня MIT + Merck выпустили на ChemRxiv CheMeleon для реакций, предобученный на 2 млн реакциях.
https://chemrxiv.org/doi/full/10.26434/chemrxiv.15008692/v1
We introduce CheMeleon-Rxn, a pre-trained reaction graph neural network that adapts the CheMeleon framework to condensed graphs of reaction (CGR) and demonstrates that descriptor-regression is an effective pre-training strategy for low-data reaction property prediction.
An O(10M)-parameter D-MPNN encoder was pre-trained on #2 million reactions by regressing onto dense reaction descriptors pooled from classical molecular descriptors, then fine-tuned on small reaction property prediction datasets. Across regression tasks spanning gasphase activation energies, reaction enthalpies, experimental yields, and rate coefficients, CheMeleon-Rxn is best or statistically tied-for-best on seven of eight tasks. It outperforms the same encoder trained from scratch, as well as classical-descriptor, fingerprint, and pre-trained-fingerprint baselines.
We further tested the pre-training strategy across various graph neural network architectures and found that its benefit holds for other edge-aware backbones such as GINE, and that it is robust to the choice of descriptor set, pooling operation, and other pre-training hyperparameters. We visualized and analyzed the CheMeleon-Rxn embeddings, which reveal chemically coherent structure in the learned reaction representation space
🔥Код полностью открытый: https://github.com/MSDLLCpapers/chemeleon-rxn
А вот веса будут опубликованы позже:
The model weights will be released upon publication. | 1 020 |
| 13 | 🎉 Мы выпустили BigSolDB v2.2 — новое крупное обновление нашего открытого набора данных по растворимости.
Что нового:
➕ 13 287 новых измерений растворимости — теперь их 125 752
➕ 102 новых растворённых вещества — всего 1627
➕ 155 новых литературных источников — всего 1842
🧪 Добавлена информация о методиках эксперимента: способ измерения, время установления равновесия, перемешивание, контроль температуры и число повторов.
💎 Добавлены данные о твёрдых формах: чистота образца, характеризация методами XRD/DSC и термодинамические свойства. В набор вошли 1044 экспериментальные температуры плавления (Tm) и 824 экспериментальные энтальпии плавления (ΔH_fus).
Для 14 813 измерений теперь явно указана твёрдая форма вещества.
В отдельном файле опубликованы 59 680 рассчитанных коэффициентов активности.
Новая версия доступна на Zenodo:
https://doi.org/10.5281/zenodo.22648301 | 1 157 |
| 14 | Large language models as uncertainty-calibrated optimizers for experimental discovery 🔥
https://www.nature.com/articles/s42256-026-01283-z
Here we show how training language models through Bayesian objectives enables their use as reliable optimizers guided by natural language. Our approach, GOLLuM (Gaussian process Optimized LLMs), teaches LLMs from experimental outcomes under uncertainty, transforming their overconfidence from a fundamental flaw into a precise learning signal. This signal reshapes the LLM embeddings so that experiments with similar outcomes cluster together, revealing structure in the design space.
Starting from only ten low-performing experiments, GOLLuM generalizes across 23 tasks in organic synthesis, materials science, process chemistry and molecular design, ranking first on average among all competing methods.
It matches traditional Bayesian optimization with over 40% fewer experiments and nearly doubles the discovery of high-performing Buchwald–Hartwig reactions over expert quantum-chemical descriptors and state-of-the-art LLMs (43% versus 24–25%). More broadly, GOLLuM points to a different paradigm for specializing foundation models: not through more data but through richer, uncertainty-guided information.
📕Nature Machine Intelligence (IF=) | 547 |
| 15 | Коллеги из Польши, США и Южной Кореи систематизировали информацию по более чем 3000 полных синтезов, включающих в себя практически 60000 индивидуальных органических реакций.
Столь большой датасет позволил на качественно новом уровне исследовать ретросинтетическую логику исследователей. В частности, было показано, что:
- Задача планирования синтеза практически никогда не может сводиться к последовательному пошаговому принятию решений
- Практически половина синтетических шагов являются формально структурно непродуктивными и напрямую не приближают к целевой структуре
- Для навигации в этих "синтетических плато" в алгоритмах ретросинтеза необходимы механизмы многоступенчатого рассуждения
- Разработан ряд таких многоступенчатых эвристик, призванных улучшить разработку моделей машинного обучения для органического синтеза
Источник: https://pubs.acs.org/jacsat/article/doi/10.1021/jacs.6c08434/5401944/AbSynth-A-Century-of-Total-Syntheses-as-a-Machine
📕Journal of the American Chemical Society (IF=16.6)
#dataset | 1 017 |
| 16 | Machine Learning for Synthetic Organic Chemistry: Methods, Applications, and Best Practices
https://doi.org/10.1002/anov.70032
The growing integration of artificial intelligence (AI) and machine learning (ML) is transforming experimental chemistry laboratories. Especially in synthetic chemistry, researchers routinely handle complex and high-dimensional data, fostering meaningful synergies between chemistry and data science.
This review is intended as a practical overview that connects the everyday challenges of synthetic chemists with the digital tools available to address them. It does not seek to explain theoretical foundations of ML or to provide a comprehensive survey of all recent studies in the field. Rather, our goal is to highlight emerging technologies, discuss key considerations for their application, and present a selection of illustrative examples. To begin, we outline the prerequisites for successfully applying data science in synthetic chemistry. Next, we give a realistic overview of strategies and bottlenecks in predictive modeling of molecular properties, reaction outcomes and reaction conditions. We further highlight data-driven approaches that can be applied in the development of new chemical reactions and synthetic methodologies, including all relevant stages from reaction discovery and optimization to substrate scope evaluation and mechanistic analyses. Finally, we briefly discuss the transformative role of large language models and agentic workflows in synthetic chemistry, focusing on opportunities and challenges in the laboratories of the future.
📕Angewandte Chemie Novit
#review | 409 |
| 17 | QuantumPioneer: Scalable Generation of Quantum Chemical Data for Solution-Phase Hydrogen Transfer Reactions
https://doi.org/10.1021/jacs.6c10371
We introduce QuantumPioneer, an open-access reaction-centered QM database and workflow for small organic molecules, focused on peroxyl-mediated hydrogen atom transfer (HAT) and the corresponding homolytic bond dissociation reactions.
QuantumPioneer contains 348,258 species (2-21 heavy atoms), 167,237 validated HAT transition states (TS) with corresponding reaction energies and homolytic bond dissociation energies (BDEs), and over 100 million COSMO-RS solvation free energies (ΔG_solv) and enthalpies (ΔH_solv) across 295 solvents.
The workflow uses ωB97X-D/def2-SVP geometries, DLPNO-CCSD(T)-F12a/cc-pVTZ-F12 single-point energies, empirical thermochemical corrections, transition-state theory, and COSMO-RS BP-TZVPD-FINE solvation in a single high-throughput pipeline.
Our benchmarks show reliable accuracy, with mean absolute errors (MAEs) compared to experimental and high-level QM reference data of 0.82 kcal/mol for gas-phase enthalpies of formation, 1.60 kcal/mol for C-H BDEs, 1.45 kcal/mol for HAT barriers, and 0.57 kcal/mol for ΔG_solv values.
We demonstrate two predictive applications. First, we show that combining BDE and HAT-barrier models identifies experimentally observed oxidative degradation sites in drug-like molecules with a 91% top-5 hit rate and 82% site-level recall.
Second, we show that a QM-parametrized Abraham model enables rapid solvation energy estimates at near-COSMO-RS accuracy within its training domain, reproducing computed ΔG_solv and ΔH_solv values with MAEs of 0.16 and 0.18 kcal/mol, respectively, though performance on experimental ΔG_solv values for unseen solutes was worse, with an MAE of 1.32 kcal/mol.
📕Journal of the American Chemical Society (IF=16.6)
#dataset | 489 |
| 18 | Новое исследование от AstraZeneca:
Do humans and large language models agree on the quality of synthesis plans? 🔥
https://doi.org/10.1039/d6dd00266h
Here, we investigated whether LLMs could mimic human experts on the challenging task of assessing retrosynthetic path feasibility in routes generated by a popular computer-aided synthesis planning tool (AiZynthFinder).
We evaluated the agreement between LLMs and expert chemists on holistic evaluations of the proposed routes as well as the individual chemical reactions in them. We used four frontier LLMs, three proprietary models and one open-source model (Claude Opus 4.8, GPT-5.5, Llama 3.1 70B, and Gemini 3.1 Pro) and employed 17 expert chemists to grade 50 retrosynthetic paths.
We found that, when provided with clearly defined reaction-evaluation categories, human experts tended to converge in their assessments. Among the evaluated models, Gemini 3.1 Pro achieved the highest agreement with the human majority vote. GPT-5.5 and Claude Opus 4.8 were comparatively more pessimistic, whereas Llama 3.1 70B showed a pronounced optimistic bias.
📕Digital Discovery (IF=7.1) | 431 |
| 19 | Model Validation Protocols for Machine Learning in Small Molecule Drug Discovery 🔥
https://www.biorxiv.org/content/10.64898/2026.08.19.745868v1
Building on the collective expertise of a cross-industry consortium, we present a model validation framework consisting of five recommendations that would enable the community to move beyond aggregate metrics toward understanding where and why molecular property prediction models fail.
We connect evaluation choices to real-world applications and case studies encountered in pharmaceutical research. The framework proposes splitting strategies that mimic realistic distribution shifts and expose common failure modes. We apply the recommended framework to a recently released dataset of absorption, distribution, metabolism, and excretion (ADME) properties. | 422 |
| 20 | بدون متن... | 392 |
