What's new on the market? Combining internet traces and pretrained language models to recognize emerging drug names.

Bibliographic Details
Title: What's new on the market? Combining internet traces and pretrained language models to recognize emerging drug names.
Authors: Grenier G; Ecole des Sciences Criminelles, University of Lausanne, Lausanne 1015, Switzerland. Electronic address: guillaume.grenier@unil.ch., Charest M; Ecole des Sciences Criminelles, University of Lausanne, Lausanne 1015, Switzerland., Esseiva P; Ecole des Sciences Criminelles, University of Lausanne, Lausanne 1015, Switzerland., Rossy Q; Ecole des Sciences Criminelles, University of Lausanne, Lausanne 1015, Switzerland.
Source: Forensic science international [Forensic Sci Int] 2026 Aug; Vol. 385, pp. 112958. Date of Electronic Publication: 2026 Mar 30.
Publication Type: Journal Article
Language: English
Journal Info: Publisher: Elsevier Science Ireland Country of Publication: Ireland NLM ID: 7902034 Publication Model: Print-Electronic Cited Medium: Internet ISSN: 1872-6283 (Electronic) Linking ISSN: 03790738 NLM ISO Abbreviation: Forensic Sci Int Subsets: MEDLINE
Imprint Name(s): Publication: Limerick : Elsevier Science Ireland
Original Publication: Lausanne, Elsevier Sequoia.
MeSH Terms: Internet* , Illicit Drugs* , Psychotropic Drugs* , Terminology as Topic* , Natural Language Processing*, Humans
Abstract: Posts and comments published by users in online forum discussions provide valuable insights and might contain the earliest traces of new substances emerging on the market. However, the systematic recognition of emerging new psychoactive substances (NPS) remains an important challenge for both public health agencies and law enforcement authorities. Large volumes of messages published by users, combined with the unstructured nature of text, complicate the retrieval of relevant information like drug name mentions. Common approaches based on keywords matching (e.g., regular expressions) limit current monitoring systems, as they can only detect known terms. Consequently, new or previously unseen drug names may remain undetected, leaving novel NPS under active discussion potentially overlooked. To address this challenge, we introduce DrugRecon, a RoBERTa based pretrained language model specifically fine-tuned for drug name recognition. The model was trained and evaluated on a manually annotated corpus of posts and comments collected from drug-related sections of three online forums (Drugs-Forum, Dread, and Reddit). A data augmentation strategy was applied during fine-tuning to improve generalization to previously unseen drug names. To demonstrate its applicability in real-world settings, DrugRecon was applied to posts and comments published between April and June 2025 across the three forums. The model successfully recognized drug names absent from existing lexicons, highlighting its capacity to detect emerging terminology. By combining automatic recognition with expert validation, 12 names were classified as denoting potential novel NPS. This proactive monitoring approach not only guides further investigations, but also strengthens preparedness for when these substances eventually appear in drug-checking services, police seizures, or toxicological reports.
(Copyright © 2026 The Authors. Published by Elsevier B.V. All rights reserved.)
Competing Interests: Declaration of Competing Interest The authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper.
Contributed Indexing: Keywords: Early warning system; Intelligence-led-policing; Named entity recognition; Natural language processing; New psychoactive substances
Substance Nomenclature: 0 (Illicit Drugs)
0 (Psychotropic Drugs)
Entry Date(s): Date Created: 20260407 Date Completed: 20260714 Latest Revision: 20260714
Update Code: 20260715
DOI: 10.1016/j.forsciint.2026.112958
PMID: 41946262
Database: MEDLINE
Description
ISSN:1872-6283
DOI:10.1016/j.forsciint.2026.112958