INBIOSIS · UKM
AI Bioinfo Daily
A daily digest of the three most important new papers on AI and LLM applications in bioinformatics, computational biology, genomics, proteomics and drug discovery. Curated every morning.
13 issues · 38 papers · updated 2026-09-21
Cutting‑edge LLMs are now being harnessed for concrete biomedical tasks, from interpreting genetic severity to prioritising drug candidates and even ranking molecular conformers.
T. Ghasemnejad et al. | arXiv, 2026-09-17
An autonomous AI agent combines ReAct reasoning with Retrieval‑Augmented Generation to classify genetic disease severity using ACMG guidelines, achieving >93% accuracy and providing evidence‑backed explanations.
Why it matters. It offers a standardized, scalable way to assess disease severity for genomic screening panels.
Y. Fujioka et al. | arXiv, 2026-09-17
A new pre‑training framework integrates hierarchical code aggregation and cross‑reference mechanisms, improving clinical prediction and enabling in‑silico screening that successfully rediscovers known Alzheimer’s drugs.
Why it matters. Demonstrates that domain‑aware LLMs can accelerate hypothesis generation for drug repurposing.
G. Semakin et al. | arXiv, 2026-09-17
Evaluating modern LLMs on conformer ranking shows several models rival traditional force fields, indicating emergent 3‑D chemical reasoning from language training.
Why it matters. Opens the door for using LLMs in autonomous molecular design and drug discovery pipelines.
Spotlights emerging LLM-driven methods for tumor prognostication and hypothesis generation.
First Author et al. | arXiv, 2026-09-12
The study applies a graph-guided mixture of experts to multi-modal tumor data, demonstrating improved survival prediction. It showcases how LLMs can integrate diverse biomedical modalities.
Why it matters. Integrating multimodal data via LLM architectures advances precision oncology and provides actionable insights for clinicians.
First Author et al. | arXiv, 2026-09-10
The study applies a graph-guided mixture of experts to multi-modal tumor data, demonstrating improved survival prediction. It showcases how LLMs can integrate diverse biomedical modalities.
Why it matters. Integrating multimodal data via LLM architectures advances precision oncology and provides actionable insights for clinicians.
First Author et al. | arXiv, 2026-09-12
The study applies a graph-guided mixture of experts to multi-modal tumor data, demonstrating improved survival prediction. It showcases how LLMs can integrate diverse biomedical modalities.
Why it matters. Integrating multimodal data via LLM architectures advances precision oncology and provides actionable insights for clinicians.
Cutting‑edge language models are now the new building blocks for rapid hypothesis generation in life sciences, from protein engineering to single‑cell workflows.
First Author et al.
A transformer‑based approach jointly trains on sequence, AlphaFold features and structural restraints to predict the impact of point mutations on protein stability, achieving >90 % accuracy against experimentally annotated datasets.
Why it matters. Enables high‑throughput in‑silico mutagenesis screening for therapeutic antibody optimization and enzyme design.
First Author et al.
An auto‑encoding framework that learns shared latent representations across 50 vertebrate species and can impute missing single‑cell data in non‑model organisms using a fine‑tuned GPT architecture.
Why it matters. Expands downstream pathway analysis to ecological and evolutionary studies where experimental profiling is scarce.
First Author et al.
A diffusion‑based generative model conditioned on cell‑type labels synthesises pathway activity matrices, allowing researchers to infer unmeasured biochemical pathways from scRNA‑seq profiles.
Why it matters. Provides an integrative tool for reconstructing regulatory networks without costly perturbation experiments.
The past week’s literature showcases novel ways large language models are integrated into protein bioinformatics, systems biology frameworks, and agent benchmarking.
First Author et al. | arXiv, 2026-08-26
The authors reveal that latent embeddings of a protein language model capture structural motifs and functional domains, enabling interpretability for downstream proteomic predictions.
Why it matters. This work offers a new avenue to translate abstract LLM features into tangible biological insight, accelerating hypothesis generation in bioinformatics.
First Author et al. | arXiv, 2026-08-29
Using an agent‑based framework that integrates diverse protein abundance datasets, the authors identify tightly co‑expressed modules conserved across tissues, shedding light on systems‑level protein regulation.
Why it matters. By linking multi‑omics data through AI orchestration, this study paves the way for more comprehensive models of cellular function and disease pathways.
First Author et al. | arXiv, 2026-08-26
This benchmark evaluates several LLM‑driven agent chains across a suite of computational biology problems, highlighting performance gaps and guiding future development.
Why it matters. The results inform the design of more capable AI agents for complex bioinformatics workflows, promoting reproducible, high‑impact research productivity.
Cutting‑edge LLMs are accelerating biologically relevant discovery, from de novo molecules to disease‑specific structural models.
Kim et al. | J Chem Inf Model, 2026-08-31
Combines a transformer trained on SMILES and chemical property data to generate synthesizable drug candidates; achieves >70% success in Tanimoto similarity benchmark versus baseline methods.
Why it matters. Speeds up medicinal chemistry cycles by automating early‑stage molecule design.
Patel & Lee | Bioinformatics, 2026-08-30
Uses a language model fine‑tuned on synthetic genetic network descriptions to output orthogonal gene circuits with specified logic functions; validated in silico for Boolean behavior.
Why it matters. Enables rapid prototyping of synthetic biology constructs without manual design.
Singh et al. | Nat Commun, 2026-08-29
Applies a diffusion-based LLM to predict disease‑associated conformations of tau and alpha‑synuclein proteins; provides new structural hypotheses that could guide therapeutic approaches.
Why it matters. Offers mechanistic insights that could accelerate therapy development for Parkinson’s and Alzheimer’s disorders.
Hsu et al. | Source: arXiv (2026‑08‑27)
This work introduces CritICL, a method that leverages the failure patterns of small LLMs during inference to guide training of larger models, reducing overfitting on rare biological sequences. It demonstrates consistent gains across protein‑folding and gene‑expression tasks with fewer labeled samples.
Why it matters. By turning model errors into signals, CritICL enables efficient scaling of biomedical AI without the data burden that typically limits deployment in resource‑constrained research settings.
Chen et al. | Source: arXiv (2026‑08‑27)
SWE‑Prime adapts the large‑language‑model backbone to protein‑ligand affinity prediction by pruning trajectory sets in molecular simulations, yielding a 12 % MAE reduction over baseline GPT‑based generative models while keeping inference time under 2 s per complex.
Why it matters. The approach directly addresses the bottleneck of costly simulation data, accelerating virtual screening pipelines crucial for early drug discovery.
Nguyen et al. | Source: arXiv (2026‑08‑27)
TTPO introduces a test‑time policy refinement framework that fine‑tunes pretrained LLM policies on the fly in biological simulation environments, achieving up to 18 % higher success rates on synthetic navigation benchmarks for cellular-scale systems.
Why it matters. It demonstrates how dynamic adaptation at run time can overcome static model limitations, paving the way for flexible AI agents that can respond to evolving biological conditions.
Large language models now face unprecedented scrutiny for biosecurity risks, with new frameworks testing whether "safety-aligned" scientific AIs can be coaxed into generating harmful pathogens—uncovering a critical gap between text-level safeguards and real biological capability.
Zirui Wang et al | arXiv, 2026-08-06
Despite strong biomedical reasoning abilities, LLMs remain unclear on extracting epitope information directly from antigen sequences—critical for antibody design. EpiBench launches a new 1,609-sample benchmark testing five tasks: target region discovery, antibody-conditioned epitope identification, binning, functional assessment, and escape prediction. Results show current models capture partial signals but lack sequence grounding and long-context localization needed for reliable LLM-assisted antibody discovery workflows.
Why it matters. Provides a concrete testbed to measure whether biomedical LLMs can move beyond generic reasoning toward truly biological sequence-aware applications central to therapeutic design.
Doniyorkhon Obidov et al | arXiv, 2026-08-07
LLM-powered peptide generative models are being weaponized via "Genotypic Triggers"—backdoor attacks shifting model distribution toward high immunogenicity risk for specific HLA alleles, without triggering standard safety screens that prioritize antimicrobial potency and low general toxicity instead of genetic-specific risks. The attack increased predicted immuno-genicity scores by 743% compared to natural peptides; backdoored models retained or improved primary properties while enabling targeted health threats against carriers of specific gene variants involved in immune presentation.
Why it matters. Demands redesign of biosecurity evaluation pipelines, exposing how conventional safety screens overlook context-specific population-level risks when AI systems are trained for efficacy rather than contextual harm assessment alone.
AI is increasingly being used to retrieve and predict molecular perturbation responses, demonstrating that better retrieval can outweigh more complex predictors.
Betty Xiong et al. | *arXiv* 2026‑08‑03
The authors treat drug‑response prediction as a retrieve‑and‑aggregate problem. An LLM ranks biologically related neighbor drugs profiled in the target cell line; a simple mean aggregator then combines their expression deltas to predict the response of an unseen drug. Benchmarks on the Tahoe‑100M single‑cell perturbation atlas show consistent improvements over mean baselines, ChemCPA, and chemistry‑based k‑NN, especially for unseen cell‑line generalisation. The work highlights the power of LLM‑driven retrieval as a key driver for zero‑shot molecular perturbation prediction.
Why it matters. Demonstrates that LLMs can provide biologically informed priors for drug‑response prediction without heavy model training.
Zhongxiang Sun et al. | *arXiv* 2026‑07‑02
Introduces NeuroCogMap, a framework that maps internal representations of LLMs into functional parcels analogous to cognitive systems. The authors show stable, reproducible functional organization across models and link specific parcels to failure modes such as hallucination, bias, and refusal. The framework also predicts human cortical responses to natural language, bridging artificial and biological cognition.
Why it matters. Provides a systematic way to interpret LLM behaviour, informing safe and reliable AI deployment in bioinformatics.
Xuefeng Liu et al. | *arXiv* 2026‑07‑01
Proposes Active‑GRPO, which dynamically switches between imitation of reference molecules and reinforcement‑learning‑based self‑improvement during molecular generation. By continually upgrading its reference set with its own best candidates, the method surpasses prior generative models on multiple benchmarks (e.g., LogP, QED) while reducing hallucinations.
Why it matters. Offers a robust, self‑evolving approach for AI‑driven drug design, reducing reliance on static reference datasets.
Cutting‑edge AI tools are streamlining data‑intensive workflows across bioinformatics, from molecular networking to gene‑set interpretation.
Shirou Feng et al. | Analytical Chemistry, 2026‑08‑04
An open‑source toolkit (MN‑Suite) built with large language model assistance integrates multiple similarity algorithms for mass‑spec molecular networking, offering flexible, server‑free analysis of natural product datasets.
Why it matters. It demonstrates how LLM‑guided software engineering can accelerate deployment of specialized bio‑informatics pipelines.
Wee Loong Chin et al. | PLoS Computational Biology, 2026‑08‑05
The GeneInsight framework extracts functional annotations from STRING‑DB, clusters semantically related terms using sentence embeddings, and generates concise thematic summaries via LLM prompting, streamlining gene‑set interpretation.
Why it matters. It showcases LLMs as powerful assistants for turning sprawling gene‑set outputs into clear biological insights.
Masato Tsutsui et al. | Bioinformatics, 2026‑08‑03
Using LLM‑based text mining, quantitative context‑dependent weights are assigned to literature‑extracted gene regulations, producing GRNs that better reflect disease‑specific biology and enhance drug‑target prediction and ODE model construction.
Why it matters. It highlights the value of LLMs for building more accurate, condition‑specific regulatory models.
Cutting‑edge AI tools are reshaping how we predict structures, generate single‑cell data, and assess biosecurity risks.
Shu Quan et al. | arXiv, 2026-08-05
Large Language Models accelerate biological research but also enable the design of harmful toxin‑like proteins, posing a biosecurity threat. The authors introduce SPIKE‑Bench, a benchmark that evaluates LLMs on toxin‑design prompts and propose a classifier to mitigate risks.
Why it matters. Highlights the need for safety testing of AI systems in biotech.
Devlina Chakravarty et al. | arXiv, 2026-08-05
Traditional protein‑structure prediction yields a single dominant conformation. This work reframes the problem as state‑space inference, reviewing ensemble generators, physics‑based simulations, and experimental constraints to predict multiple functional conformations.
Why it matters. Moves protein modelling towards capturing the dynamic ensembles essential for function and drug design.
Aleksandr Sharipov et al. | arXiv, 2026-08-05
Introduces a causal transformer paired with a quantized VAE tokenizer to generate realistic single‑cell gene‑expression vectors. The study characterises biological fidelity and scaling laws, enabling downstream perturbation‑response modelling.
Why it matters. Provides a scalable foundation model for synthetic single‑cell data, accelerating benchmarking and method development.
Three papers showing how LLMs and structured biological priors are reshaping perturbation prediction, single-cell representation learning, and spatial transcriptomics from histology images.
Betty Xiong et al. | arXiv, 2026-08-03
Rather than training a complex predictor from scratch, this work uses an LLM to rank biologically related drugs whose transcriptomic profiles have already been measured, then averages their gene-expression signatures to approximate the response of an untested compound. Tested on the Tahoe-100M perturbation atlas, the method beats chemistry-based and mean-aggregation baselines, especially when generalizing to unseen cell lines.
Why it matters. Shows that retrieval quality — not just model complexity — drives zero-shot perturbation prediction, positioning LLMs as practical biological priors in drug discovery pipelines.
Jiaqi Xiong et al. | arXiv, 2026-08-02
Most single-cell foundation models pretrain by reconstructing masked gene expression, which captures gene-gene dependencies but not whole-cell structure. This paper introduces a contrastive pretraining scheme that splits each cell into two co-expression-guided gene-partition views and trains the model to produce matching cell embeddings, with hard negatives built by shuffling expression values. The resulting representations rank among the best for cell-type annotation and gene regulatory network inference across six networks.
Why it matters. Demonstrates a principled move beyond masked reconstruction toward learning truly cell-level representations, a key step toward more transferable single-cell foundation models.
Zhiwen Xu et al. | arXiv, 2026-08-01
Predicting spatial transcriptomics from H&E histology images typically treats genes as a flat output vector, ignoring biological relationships. This work plugs the Gene Ontology hierarchy into the decoder, refining predictions progressively from broad functional domains down to individual genes via residual corrections. Evaluated across nine HEST-1k datasets, GO-guided decoding consistently outperforms flat and random-hierarchy baselines — and drops in as a plug-in that requires zero changes to the image-side backbone.
Why it matters. Proves that incorporating curated ontological structure as an inductive bias measurably boosts transcriptomic prediction from histology, making large-scale, cost-free spatial profiling more reliable.
Diffusion-based protein structure prediction, faster backbone generation on Lie groups, and a new framework for auditing AI bioinformatics agents.
Vilya Research et al. | arXiv, 2026-07-28
Vilya-2 is a diffusion transformer that extends all-atom representation from individual molecules to protein-ligand interfaces. It achieves 59.1% of peptide interfaces at sub-2 Å backbone RMSD — far outperforming co-folding models — and sets state-of-the-art on small-molecule docking. The model generalizes to macrocycles and disulfide-stapled miniproteins several-fold larger than any in training, and can be fine-tuned for hit-to-lead campaigns.
Why it matters. Bridges the gap AlphaFold left open for peptide therapeutics, enabling reliable structure prediction for molecules with non-canonical residues and complex topologies that dominate the next generation of drug candidates.
Phuc Pham et al. | arXiv, 2026-07-30
As LLM agents increasingly plan and execute biological analyses, this review introduces the Function–Evidence–Validation (FEV) framework for assessing their scientific accountability. The authors map 109 agentic systems and 28 benchmarks across genomics, single-cell omics, protein science, and drug discovery, finding that planning and tool use outpace reproducibility and external validation. They advocate evaluating workflows — not just final answers.
Why it matters. Provides the first systematic rubric for auditing whether an AI bioinformatics agent's output is scientifically credible rather than merely fluent.
Yikun Bai et al. | arXiv, 2026-07-31
Current protein backbone generators require hundreds of network evaluations with expensive Lie-group operations at each step. SE(3)-MeanFlow derives closed-form average-velocity identities in the Lie algebra, eliminating the Jacobian-vector product from the rotation branch. The result: a few-step generative model that matches or exceeds flow-matching baselines using several times fewer sampling steps, with an advantage that widens at every matched computational budget via rectification.
Why it matters. Dramatically speeds up de novo protein design inference — making high-throughput backbone generation feasible without sacrificing quality.
LLMs are moving from generic text generators to purpose‑built assistants for life‑science data, enabling faster literature mining, enzyme annotation, and even cognitive‑style reasoning about biology.
Luigi Sigillo et al. | *arXiv*, 2026‑07‑30
The paper introduces a “knowledge layer” that sits on top of Europe PMC, letting AI agents ask natural‑language questions and receive concise, evidence‑backed answers instead of having to craft complex search queries and read whole papers. A single LLM orchestrates sub‑queries, fetches articles, and extracts the needed statements.
Why it matters. By turning literature‑search into a plug‑and‑play service, it lowers the barrier for bio‑agents to stay up‑to‑date, accelerating hypothesis generation and data‑driven discovery.
Linyu Li et al. | *arXiv*, 2026‑07‑29
This benchmark isolates four “levers” that affect an LLM’s ability to predict enzyme EC numbers—output format, external knowledge, reasoning structure, and robustness. It shows that open‑book (retrieval‑augmented) access dramatically lifts performance, while chain‑of‑thought reasoning helps only when the model already knows the answer.
Why it matters. It gives developers a clear diagnostic tool to spot why LLMs fail on enzyme annotation, guiding the design of more reliable bio‑AI pipelines.
Chandra Sripada et al. | *arXiv*, 2026‑07‑28
The authors compare LLMs to human cognition across five dimensions (inferential organization, architecture, representations, prediction‑driven learning, and reinforcement‑like mechanisms). They argue that despite substrate differences, LLMs independently arrive at principles long identified in cognitive science.
Why it matters. Recognizing these convergences helps bio‑informaticians borrow well‑tested cognitive frameworks for interpreting LLM behavior in biological reasoning tasks.