Small Molecule Optimization with Large Language Models
Guevorguian, Bedrosian, Fahradyan, Chilingaryan, Aghajanyan, Khachatrian · cs.LG,cs.NE,q-bio.QM · 2026-09-08 · 原文
Molecular optimization, the process of designing molecules with desirable properties, represents a critical challenge in drug discovery. The recent advancements in large language models (LLMs) have opened new opportunities for their integration with traditional molecular optimization algorithms to improve performance. In this work, we propose Molecular Language Model powered Evolutionary Algorithm (Mol-E), an evolutionary algorithm that relies on the generative capabilities of LLMs trained on molecules and molecular properties. Scientific Contribution. Mol-E obtains the highest aggregate Top-10 AUC among the comparable full-23-task results considered here, scoring 17.500 in the task-agnostic regime, in which the oracle is treated strictly as a black box, and 20.551 in the task-informed regime, in which the optimizer receives a fixed semantic description of the objective. Mol-E also improves over the evaluated baselines on multi-property optimization with docking against DRD2, MK2, and AChE.
讲义
讲义·推断 依据「原文」自动生成的结构化摘要(推断),非原文表述;以原文为准。
1. 人话版
Molecular optimization, the process of designing molecules with desirable properties, represents a critical challenge in drug discovery.
The recent advancements in large language models (LLMs) have opened new opportunities for their integration with traditional molecular optimization algorithms to improve performance.
2. 领域脉络
本文类目:cs.LG、cs.NE、q-bio.QM,属于其所在研究脉络的最新进展。
3. 机制拆解
In this work, we propose Molecular Language Model powered Evolutionary Algorithm (Mol-E), an evolutionary algorithm that relies on the generative capabilities of LLMs trained on molecules and molecular properties.
4. 证据与数字
Mol-E obtains the highest aggregate Top-10 AUC among the comparable full-23-task results considered here, scoring 17.500 in the task-agnostic regime, in which the oracle is treated strictly as a black box, and 20.551 in the task-informed regime, in which the optimizer receives a fixed semantic description of the objective.
Mol-E also improves over the evaluated baselines on multi-property optimization with docking against DRD2, MK2, and AChE.
5. 反例与边界
Scientific Contribution.
6. 跨领域连接与意外收获
横跨 3 个类目(cs.LG、cs.NE、q-bio.QM),关注其在你兴趣板块间的迁移面。
7. 可复用方法
把本文机制与你手头项目对照,找一个两周内能验证的最小实验。
8. 术语表
精读时把不熟的术语记入此处,作为下次回忆的锚点。