Where to cut, how deep? BPE and Unigram-LM in chemical SMILES

Discover how BPE and Unigram-LM build almost disjoint vocabularies when tokenizing chemical SMILES. Reveals differences in molecule segmentation.

miércoles, 8 de julio de 2026 • 2 min read • Q2BSTUDIO Team

Comparison of BPE and Unigram-LM in chemical tokenization

Tokenization of SMILES strings represents a critical step in the pipeline of any language model applied to computational chemistry. For years, the community has adopted Byte-Pair Encoding (BPE) by inertia from natural language processing, without stopping to evaluate whether it is truly the optimal choice for domains as specific as molecular representation. Recent research comparing BPE with Unigram-LM reveals that both algorithms generate practically disjoint vocabularies: the subword pieces each learns barely overlap, and the depth at which they segment molecules differs significantly. This is not an academic curiosity, but a practical warning: choosing the wrong tokenizer can bias data representation and harm the performance of generative or predictive models.

For a company developing artificial intelligence solutions applied to chemistry or pharmaceuticals, this modeling decision becomes a strategic factor. It is not enough to take a predefined tokenizer; it is necessary to understand how each algorithm interprets molecular structure and how that bias impacts the model's ability to generalize. This is where having a team that offers custom applications makes the difference. At Q2BSTUDIO we design customized tokenization pipelines, adapted to the data typology (diverse, drugs, natural products) and business objectives, also integrating AI for companies that allow automating the selection of the optimal algorithm.

Tokenization is not mere preprocessing; it is a modeling decision that conditions the learning capacity of the system. Therefore, in environments where precision matters —such as drug discovery or natural product chemistry— we recommend performing a comparative analysis similar to what has been done with BPE and Unigram-LM. Our custom software services include the implementation of these evaluations, combined with AWS and Azure cloud services to scale training, and business intelligence services with Power BI to visualize tokenization quality metrics. Additionally, we offer AI agents that can dynamically monitor and adjust the tokenizer according to the evolution of the corpus. Cybersecurity also plays a relevant role in protecting sensitive molecular data during these processes. Ultimately, understanding where to cut and how deep to do it in SMILES strings is a question that deserves a well-founded answer, not an inherited default.

A BREAK?

Play for a moment before you go

OUR SERVICES

How we can help you

Do you have a project in mind?

Tell us your vision and we'll turn it into a software solution. Whatever the scope, we make your idea real.