Papers › Identifying the kind behind SMILES—anatomical therapeutic chemical classification...
Identifying the kind behind SMILES—anatomical therapeutic chemical classification using structure-only representations
Yi Cao, Zhen-Qun Yang, Xu-Lu Zhang, Wenqi Fan, YaoWei Wang, Jiajun Shen, Dong-Qing Wei, Qing Li, Xiao-Yong Wei
Anatomical Therapeutic Chemical (ATC) classification for compounds/drugs plays an important role in drug development and basic research. However, previous methods depend on interactions extracted from STITCH dataset which may make it depend on lab experiments. We present a pilot study to explore the possibility of conducting the ATC prediction solely based on the molecular structures. The motivation is to eliminate the reliance on the costly lab experiments so that the characteristics of a drug can be pre-assessed for better decision-making and effort-saving before the actual development. To this end, we construct a new benchmark consisting of 4545 compounds which is with larger scale than the one used in previous study. A light-weight prediction model is proposed. The model is with better explainability in the sense that it is consists of a straightforward tokenization that extracts and embeds statistically and physicochemically meaningful tokens, and a deep network backed by a set of pyramid kernels to capture multi-resolution chemical structural characteristics. Its efficacy has been validated in the experiments where it outperforms the state-of-the-art methods by 15.53% in accuracy and by 69.66% in terms of efficiency. We make the benchmark dataset, source code and web server open to ease the reproduction of this study.
Code
Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.
Code Syntology ran Syntology
Not run by Syntology. Nothing on this page verifies that the listed code works.
Tasks
Datasets
Introduced by this paper, per the archive.
Results from the paper archive 2025-07-28
| Task | Dataset | Model | Metric | Value | Rank at snapshot | Leaderboard | Report |
|---|---|---|---|---|---|---|---|
| Drug ATC Classification | ATC-SMILES | ATC-CNN | Absolute False | 0.0094 | #2 of 2 | Archive leaderboard | report |
| Drug ATC Classification | ATC-SMILES | ATC-CNN | Absolute True | 0.9177 | #2 of 2 | Archive leaderboard | report |
| Drug ATC Classification | ATC-SMILES | ATC-CNN | Accuracy | 0.9399 | #2 of 2 | Archive leaderboard | report |
| Drug ATC Classification | ATC-SMILES | ATC-CNN | Aiming | 0.9583 | #2 of 2 | Archive leaderboard | report |
| Drug ATC Classification | ATC-SMILES | ATC-CNN | Coverage | 0.9414 | #2 of 2 | Archive leaderboard | report |
Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections