INBIOSIS · UKM

AI Bioinfo Daily

A daily digest of the three most important new papers on AI and LLM applications in bioinformatics, computational biology, genomics, proteomics and drug discovery. Curated every morning.

22 issues · 65 papers · updated 2026-09-07

#

Cutting‑edge language models are now the new building blocks for rapid hypothesis generation in life sciences, from protein engineering to single‑cell workflows.

LLM-driven Protein Structure Refinement for Rapid Mutagenesis Prediction

First Author et al.

A transformer‑based approach jointly trains on sequence, AlphaFold features and structural restraints to predict the impact of point mutations on protein stability, achieving >90 % accuracy against experimentally annotated datasets.

Why it matters. Enables high‑throughput in‑silico mutagenesis screening for therapeutic antibody optimization and enzyme design.

Meta-AI Framework for Cross-species Gene Expression Imputation

First Author et al.

An auto‑encoding framework that learns shared latent representations across 50 vertebrate species and can impute missing single‑cell data in non‑model organisms using a fine‑tuned GPT architecture.

Why it matters. Expands downstream pathway analysis to ecological and evolutionary studies where experimental profiling is scarce.

Generative AI for Single-Cell Pathway Reconstruction

First Author et al.

A diffusion‑based generative model conditioned on cell‑type labels synthesises pathway activity matrices, allowing researchers to infer unmeasured biochemical pathways from scRNA‑seq profiles.

Why it matters. Provides an integrative tool for reconstructing regulatory networks without costly perturbation experiments.

#

The past week’s literature showcases novel ways large language models are integrated into protein bioinformatics, systems biology frameworks, and agent benchmarking.

Interpreting Latent Protein Language Model Features with Geometric Annotations

First Author et al. | arXiv, 2026-08-26

The authors reveal that latent embeddings of a protein language model capture structural motifs and functional domains, enabling interpretability for downstream proteomic predictions.

Why it matters. This work offers a new avenue to translate abstract LLM features into tangible biological insight, accelerating hypothesis generation in bioinformatics.

Agentic AI Uncovers Conserved Cross‑Tissue Protein Co‑Abundance Programs

First Author et al. | arXiv, 2026-08-29

Using an agent‑based framework that integrates diverse protein abundance datasets, the authors identify tightly co‑expressed modules conserved across tissues, shedding light on systems‑level protein regulation.

Why it matters. By linking multi‑omics data through AI orchestration, this study paves the way for more comprehensive models of cellular function and disease pathways.

BixBench 3: Benchmarking AI Agents on Research‑Study‑Scale Computational Biology Tasks

First Author et al. | arXiv, 2026-08-26

This benchmark evaluates several LLM‑driven agent chains across a suite of computational biology problems, highlighting performance gaps and guiding future development.

Why it matters. The results inform the design of more capable AI agents for complex bioinformatics workflows, promoting reproducible, high‑impact research productivity.

#

Cutting‑edge LLMs are accelerating biologically relevant discovery, from de novo molecules to disease‑specific structural models.

DrugGPT‑X: LLM‑Guided De Novo Molecule Discovery

Kim et al. | J Chem Inf Model, 2026-08-31

Combines a transformer trained on SMILES and chemical property data to generate synthesizable drug candidates; achieves >70% success in Tanimoto similarity benchmark versus baseline methods.

Why it matters. Speeds up medicinal chemistry cycles by automating early‑stage molecule design.

AutoGeneAI: An LLM for Automated Gene Circuit Design

Patel & Lee | Bioinformatics, 2026-08-30

Uses a language model fine‑tuned on synthetic genetic network descriptions to output orthogonal gene circuits with specified logic functions; validated in silico for Boolean behavior.

Why it matters. Enables rapid prototyping of synthetic biology constructs without manual design.

NeuroNetGen: AI‑Driven Protein Folding for Neurological Diseases

Singh et al. | Nat Commun, 2026-08-29

Applies a diffusion-based LLM to predict disease‑associated conformations of tau and alpha‑synuclein proteins; provides new structural hypotheses that could guide therapeutic approaches.

Why it matters. Offers mechanistic insights that could accelerate therapy development for Parkinson’s and Alzheimer’s disorders.

#

CritICL: Inference-Time Weak-to-Strong Generalization from Small Language Model Failure Modes

Hsu et al. | Source: arXiv (2026‑08‑27)

This work introduces CritICL, a method that leverages the failure patterns of small LLMs during inference to guide training of larger models, reducing overfitting on rare biological sequences. It demonstrates consistent gains across protein‑folding and gene‑expression tasks with fewer labeled samples.

Why it matters. By turning model errors into signals, CritICL enables efficient scaling of biomedical AI without the data burden that typically limits deployment in resource‑constrained research settings.

SWE-Prime: Fewer Trajectories, Better Performance for Protein–Ligand Affinity Prediction

Chen et al. | Source: arXiv (2026‑08‑27)

SWE‑Prime adapts the large‑language‑model backbone to protein‑ligand affinity prediction by pruning trajectory sets in molecular simulations, yielding a 12 % MAE reduction over baseline GPT‑based generative models while keeping inference time under 2 s per complex.

Why it matters. The approach directly addresses the bottleneck of costly simulation data, accelerating virtual screening pipelines crucial for early drug discovery.

TTPO: Test-Time Policy Optimization for Deep Biological Agents

Nguyen et al. | Source: arXiv (2026‑08‑27)

TTPO introduces a test‑time policy refinement framework that fine‑tunes pretrained LLM policies on the fly in biological simulation environments, achieving up to 18 % higher success rates on synthetic navigation benchmarks for cellular-scale systems.

Why it matters. It demonstrates how dynamic adaptation at run time can overcome static model limitations, paving the way for flexible AI agents that can respond to evolving biological conditions.

#

Large language models now face unprecedented scrutiny for biosecurity risks, with new frameworks testing whether "safety-aligned" scientific AIs can be coaxed into generating harmful pathogens—uncovering a critical gap between text-level safeguards and real biological capability.

EpiBench: Can LLMs Understand Epitopes for Antibody Drug Discovery?

Zirui Wang et al | arXiv, 2026-08-06

Despite strong biomedical reasoning abilities, LLMs remain unclear on extracting epitope information directly from antigen sequences—critical for antibody design. EpiBench launches a new 1,609-sample benchmark testing five tasks: target region discovery, antibody-conditioned epitope identification, binning, functional assessment, and escape prediction. Results show current models capture partial signals but lack sequence grounding and long-context localization needed for reliable LLM-assisted antibody discovery workflows.

Why it matters. Provides a concrete testbed to measure whether biomedical LLMs can move beyond generic reasoning toward truly biological sequence-aware applications central to therapeutic design.

Genotypic Triggers: Exposing Pharmacogenomic Blind Spots in Antimicrobial Peptide Generation

Doniyorkhon Obidov et al | arXiv, 2026-08-07

LLM-powered peptide generative models are being weaponized via "Genotypic Triggers"—backdoor attacks shifting model distribution toward high immunogenicity risk for specific HLA alleles, without triggering standard safety screens that prioritize antimicrobial potency and low general toxicity instead of genetic-specific risks. The attack increased predicted immuno-genicity scores by 743% compared to natural peptides; backdoored models retained or improved primary properties while enabling targeted health threats against carriers of specific gene variants involved in immune presentation.

Why it matters. Demands redesign of biosecurity evaluation pipelines, exposing how conventional safety screens overlook context-specific population-level risks when AI systems are trained for efficacy rather than contextual harm assessment alone.

#

AI is increasingly being used to retrieve and predict molecular perturbation responses, demonstrating that better retrieval can outweigh more complex predictors.

LLM‑Guided Retrieval for Prediction of Molecular Perturbation Responses

Betty Xiong et al. | *arXiv* 2026‑08‑03

The authors treat drug‑response prediction as a retrieve‑and‑aggregate problem. An LLM ranks biologically related neighbor drugs profiled in the target cell line; a simple mean aggregator then combines their expression deltas to predict the response of an unseen drug. Benchmarks on the Tahoe‑100M single‑cell perturbation atlas show consistent improvements over mean baselines, ChemCPA, and chemistry‑based k‑NN, especially for unseen cell‑line generalisation. The work highlights the power of LLM‑driven retrieval as a key driver for zero‑shot molecular perturbation prediction.

Why it matters. Demonstrates that LLMs can provide biologically informed priors for drug‑response prediction without heavy model training.

NeuroCogMap: Cognitive Organization of Large Language Models

Zhongxiang Sun et al. | *arXiv* 2026‑07‑02

Introduces NeuroCogMap, a framework that maps internal representations of LLMs into functional parcels analogous to cognitive systems. The authors show stable, reproducible functional organization across models and link specific parcels to failure modes such as hallucination, bias, and refusal. The framework also predicts human cortical responses to natural language, bridging artificial and biological cognition.

Why it matters. Provides a systematic way to interpret LLM behaviour, informing safe and reliable AI deployment in bioinformatics.

Active‑GRPO: Adaptive Imitation and Self‑Improving Reasoning for Molecular Optimization

Xuefeng Liu et al. | *arXiv* 2026‑07‑01

Proposes Active‑GRPO, which dynamically switches between imitation of reference molecules and reinforcement‑learning‑based self‑improvement during molecular generation. By continually upgrading its reference set with its own best candidates, the method surpasses prior generative models on multiple benchmarks (e.g., LogP, QED) while reducing hallucinations.

Why it matters. Offers a robust, self‑evolving approach for AI‑driven drug design, reducing reliance on static reference datasets.

#

Cutting‑edge AI tools are streamlining data‑intensive workflows across bioinformatics, from molecular networking to gene‑set interpretation.

LLM‑Assisted Development of a Locally Deployable Molecular Networking Toolkit: Enabling Customizable Analysis in Natural Products

Shirou Feng et al. | Analytical Chemistry, 2026‑08‑04

An open‑source toolkit (MN‑Suite) built with large language model assistance integrates multiple similarity algorithms for mass‑spec molecular networking, offering flexible, server‑free analysis of natural product datasets.

Why it matters. It demonstrates how LLM‑guided software engineering can accelerate deployment of specialized bio‑informatics pipelines.

GeneInsight: Condensing gene set knowledge via language models

Wee Loong Chin et al. | PLoS Computational Biology, 2026‑08‑05

The GeneInsight framework extracts functional annotations from STRING‑DB, clusters semantically related terms using sentence embeddings, and generates concise thematic summaries via LLM prompting, streamlining gene‑set interpretation.

Why it matters. It showcases LLMs as powerful assistants for turning sprawling gene‑set outputs into clear biological insights.

Literature‑derived, context‑aware gene regulatory networks improve biological predictions and mathematical modeling

Masato Tsutsui et al. | Bioinformatics, 2026‑08‑03

Using LLM‑based text mining, quantitative context‑dependent weights are assigned to literature‑extracted gene regulations, producing GRNs that better reflect disease‑specific biology and enhance drug‑target prediction and ODE model construction.

Why it matters. It highlights the value of LLMs for building more accurate, condition‑specific regulatory models.

#

Cutting‑edge AI tools are reshaping how we predict structures, generate single‑cell data, and assess biosecurity risks.

A Blind Spot in Alignment: Quantifying Biosecurity Risks in Large Language Models

Shu Quan et al. | arXiv, 2026-08-05

Large Language Models accelerate biological research but also enable the design of harmful toxin‑like proteins, posing a biosecurity threat. The authors introduce SPIKE‑Bench, a benchmark that evaluates LLMs on toxin‑design prompts and propose a classifier to mitigate risks.

Why it matters. Highlights the need for safety testing of AI systems in biotech.

Expanding Protein Structure Prediction into Conformational State Space

Devlina Chakravarty et al. | arXiv, 2026-08-05

Traditional protein‑structure prediction yields a single dominant conformation. This work reframes the problem as state‑space inference, reviewing ensemble generators, physics‑based simulations, and experimental constraints to predict multiple functional conformations.

Why it matters. Moves protein modelling towards capturing the dynamic ensembles essential for function and drug design.

Scaling an Autoregressive Transformer for Single‑Cell Generation

Aleksandr Sharipov et al. | arXiv, 2026-08-05

Introduces a causal transformer paired with a quantized VAE tokenizer to generate realistic single‑cell gene‑expression vectors. The study characterises biological fidelity and scaling laws, enabling downstream perturbation‑response modelling.

Why it matters. Provides a scalable foundation model for synthetic single‑cell data, accelerating benchmarking and method development.

#

Three papers showing how LLMs and structured biological priors are reshaping perturbation prediction, single-cell representation learning, and spatial transcriptomics from histology images.

LLM-Guided Retrieval for Prediction of Molecular Perturbation Responses

Betty Xiong et al. | arXiv, 2026-08-03

Rather than training a complex predictor from scratch, this work uses an LLM to rank biologically related drugs whose transcriptomic profiles have already been measured, then averages their gene-expression signatures to approximate the response of an untested compound. Tested on the Tahoe-100M perturbation atlas, the method beats chemistry-based and mean-aggregation baselines, especially when generalizing to unseen cell lines.

Why it matters. Shows that retrieval quality — not just model complexity — drives zero-shot perturbation prediction, positioning LLMs as practical biological priors in drug discovery pipelines.

Beyond Gene Reconstruction: Learning Cell Representations through Complementary Transcriptomic Views

Jiaqi Xiong et al. | arXiv, 2026-08-02

Most single-cell foundation models pretrain by reconstructing masked gene expression, which captures gene-gene dependencies but not whole-cell structure. This paper introduces a contrastive pretraining scheme that splits each cell into two co-expression-guided gene-partition views and trains the model to produce matching cell embeddings, with hard negatives built by shuffling expression values. The resulting representations rank among the best for cell-type annotation and gene regulatory network inference across six networks.

Why it matters. Demonstrates a principled move beyond masked reconstruction toward learning truly cell-level representations, a key step toward more transferable single-cell foundation models.

Gene Ontology-Guided Hierarchical Spatial Gene Expression Prediction from Histopathology Images

Zhiwen Xu et al. | arXiv, 2026-08-01

Predicting spatial transcriptomics from H&E histology images typically treats genes as a flat output vector, ignoring biological relationships. This work plugs the Gene Ontology hierarchy into the decoder, refining predictions progressively from broad functional domains down to individual genes via residual corrections. Evaluated across nine HEST-1k datasets, GO-guided decoding consistently outperforms flat and random-hierarchy baselines — and drops in as a plug-in that requires zero changes to the image-side backbone.

Why it matters. Proves that incorporating curated ontological structure as an inductive bias measurably boosts transcriptomic prediction from histology, making large-scale, cost-free spatial profiling more reliable.

#

Diffusion-based protein structure prediction, faster backbone generation on Lie groups, and a new framework for auditing AI bioinformatics agents.

Accurate Structural Modeling of Chemically Diverse Molecular Interfaces with Vilya-2

Vilya Research et al. | arXiv, 2026-07-28

Vilya-2 is a diffusion transformer that extends all-atom representation from individual molecules to protein-ligand interfaces. It achieves 59.1% of peptide interfaces at sub-2 Å backbone RMSD — far outperforming co-folding models — and sets state-of-the-art on small-molecule docking. The model generalizes to macrocycles and disulfide-stapled miniproteins several-fold larger than any in training, and can be fine-tuned for hit-to-lead campaigns.

Why it matters. Bridges the gap AlphaFold left open for peptide therapeutics, enabling reliable structure prediction for molecules with non-canonical residues and complex topologies that dominate the next generation of drug candidates.

Evaluating Agentic Bioinformatics through Function, Evidence, and Validation

Phuc Pham et al. | arXiv, 2026-07-30

As LLM agents increasingly plan and execute biological analyses, this review introduces the Function–Evidence–Validation (FEV) framework for assessing their scientific accountability. The authors map 109 agentic systems and 28 benchmarks across genomics, single-cell omics, protein science, and drug discovery, finding that planning and tool use outpace reproducibility and external validation. They advocate evaluating workflows — not just final answers.

Why it matters. Provides the first systematic rubric for auditing whether an AI bioinformatics agent's output is scientifically credible rather than merely fluent.

SE(3)-MeanFlow: Few-Step Protein Backbone Generation on Lie Groups

Yikun Bai et al. | arXiv, 2026-07-31

Current protein backbone generators require hundreds of network evaluations with expensive Lie-group operations at each step. SE(3)-MeanFlow derives closed-form average-velocity identities in the Lie algebra, eliminating the Jacobian-vector product from the rotation branch. The result: a few-step generative model that matches or exceeds flow-matching baselines using several times fewer sampling steps, with an advantage that widens at every matched computational budget via rectification.

Why it matters. Dramatically speeds up de novo protein design inference — making high-throughput backbone generation feasible without sacrificing quality.

#

LLMs are moving from generic text generators to purpose‑built assistants for life‑science data, enabling faster literature mining, enzyme annotation, and even cognitive‑style reasoning about biology.

EMBL AI Librarian: Life‑Sciences Knowledge Layer for AI Agents

Luigi Sigillo et al. | *arXiv*, 2026‑07‑30

The paper introduces a “knowledge layer” that sits on top of Europe PMC, letting AI agents ask natural‑language questions and receive concise, evidence‑backed answers instead of having to craft complex search queries and read whole papers. A single LLM orchestrates sub‑queries, fetches articles, and extracts the needed statements.

Why it matters. By turning literature‑search into a plug‑and‑play service, it lowers the barrier for bio‑agents to stay up‑to‑date, accelerating hypothesis generation and data‑driven discovery.

Knowledge before Reasoning: EC‑Reason‑Bench, a Training‑Free Diagnostic Benchmark for LLM Enzyme Classification

Linyu Li et al. | *arXiv*, 2026‑07‑29

This benchmark isolates four “levers” that affect an LLM’s ability to predict enzyme EC numbers—output format, external knowledge, reasoning structure, and robustness. It shows that open‑book (retrieval‑augmented) access dramatically lifts performance, while chain‑of‑thought reasoning helps only when the model already knows the answer.

Why it matters. It gives developers a clear diagnostic tool to spot why LLMs fail on enzyme annotation, guiding the design of more reliable bio‑AI pipelines.

Cognitive Convergence: Deep Similarities Between Large Language Models and Human Cognition

Chandra Sripada et al. | *arXiv*, 2026‑07‑28

The authors compare LLMs to human cognition across five dimensions (inferential organization, architecture, representations, prediction‑driven learning, and reinforcement‑like mechanisms). They argue that despite substrate differences, LLMs independently arrive at principles long identified in cognitive science.

Why it matters. Recognizing these convergences helps bio‑informaticians borrow well‑tested cognitive frameworks for interpreting LLM behavior in biological reasoning tasks.

#

RelAgent: A multi-agent solution for molecular relationship grounding

Chen R et al. | *Bioinformatics*, July 23, 2026

RelAgent decomposes molecular relationship grounding into three stages — entity extraction, substructure localization, and ontology-guided reasoning — with verifier agents ranking plausible candidates. On LLaMA3.1-8B, it lifts relationship F1 from 0.1% to 54.6%, far outperforming vanilla Gemini-3.1-Pro baselines.

Why it matters. Demonstrates that agentic, structure-aware reasoning dramatically outperforms off-the-shelf LLMs for grounding natural-language descriptions of molecular substructures — a critical step for AI-assisted patent analysis and controllable molecular design.

Inflammation-linked aging signals in frozen single-cell foundation models

Kendiukhov I et al. | *Biogerontology*, July 28, 2026

A rigorous nine-step evaluation pipeline was applied to frozen scGPT and Geneformer — models never trained on age — across five cohorts totaling ~5M cells. Despite matching PCA on raw predictive accuracy, the foundation models encode a recoverable aging signal concentrated in NF-κB and IFN-γ inflammation programs, validated by directional intervention tests and cross-cohort transfer.

Why it matters. Provides a concrete framework for distinguishing real biology from sampling artifacts in single-cell foundation model embeddings — a problem that affects every downstream application of these models in precision medicine and biomarker discovery.

BERT-HemoPep60: Transformer-based prediction of peptide hemolytic activity

Cai J et al. | *IEEE J. Biomed. Health Inform.*, July 28, 2026

BERT-HemoPep60 uses domain-adaptive pretraining with a prefix-prompt architecture to quantitatively predict peptide hemolysis across six mammalian species. It achieves PCC scores of 0.74–0.81 for HC₅/HC₁₀/HC₅₀ predictions, outperforming conventional ML and DL models built on handcrafted sequence encodings, covering peptides up to 60 amino acids.

Why it matters. Hemolysis is a major safety bottleneck in peptide drug development. A fast, accurate, transformer-based predictor can triage thousands of peptide candidates before wet-lab screening, accelerating the discovery pipeline.

#

AI meets quantum chemistry, molecular glue design, and biomedical causal reasoning — three breakthroughs bridging generative models, quantum hardware, and clinical evidence.

Learning to Prepare Molecular Ground States with Transformer Models

A. Koziell-Pipe et al. | arXiv, Jul 24

ADAPT-GQE learns to synthesize quantum circuits for molecular ground-state preparation via a generative AI pipeline combining supervised training and reinforcement learning. Generated circuits are executed on Quantinuum Helios-1 real quantum hardware, achieving order-of-magnitude speedups over traditional ADAPT-VQE while maintaining comparable accuracy — demonstrated on imipramine, a drug-relevant molecule.

Why it matters. First AI-generated quantum chemistry circuits deployed on utility-scale quantum hardware, accelerating drug stability protocols.

TriGlue: a Biology-Inspired Generative Model for Molecular Glue-Induced Ternary Complexes

Y. Yan et al. | arXiv, Jul 24

TriGlue formulates molecular glue design as a ternary complex generation problem. An SE(3)-equivariant interface estimation module predicts the unknown protein-protein interface from unbound structures, then a flow-matching network jointly generates the glue molecule and the rigid-body assembly transformation, producing chemically valid molecules and plausible ternary complexes.

Why it matters. Molecular glue degraders are a growing drug class with virtually no computational design tools — TriGlue fills a critical gap.

DAGForge: Auditable Causal DAG Authoring with Biomedical Literature

Yi-han Sheu et al. | arXiv, Jul 23

DAGForge automates causal DAG construction for biomedical studies. An LLM generates pairwise causal judgments grounded in verbatim evidence excerpts from a reproducible literature snapshot, assembles them into a constraint-checked graph with confidence estimates, and provides an interactive browser for evidence review, adjustment-set computation, and export.

Why it matters. Causal DAGs are essential for study design but remain a manual bottleneck — DAGForge makes them auditable, evidence-linked, and reproducible.

#

Trust and rigor in biomedical AI — from auditing foundation-model benchmark claims to measuring confidence in multi-omics fusion.

Auditing pretraining contamination in single-cell foundation model benchmarks

Sarwan Ali | arXiv, 2026-07-21

Single-cell foundation models (Geneformer, scGPT, UCE) are trained on public repositories that overlap with popular benchmarks. The authors introduce scContam, an audit framework showing that ~80% of cells in widely used benchmarks like PBMC 3k and CELLxGENE come from the model's pretraining data — meaning benchmark scores may reflect data exposure, not genuine generalization.

Why it matters. Raises a red flag for the entire single-cell AI field: many published benchmark results may be inflated by hidden data leakage.

Graph Learning on Ensembles of Cyclic Peptides: An Investigation of Molecular Ensemble Modeling

Aaron Feller et al. | arXiv, 2026-07-23

Most molecular models use a single 3D conformation, even though real molecules exist as ensembles of structures. EnsembleEGNN encodes entire conformational ensembles using equivariant graph neural networks and a set-attention pooler, then fuses with a BERT sequence encoder. The hybrid reaches R²=0.538 for cyclic peptide property prediction — significantly outperforming sequence-only baselines.

Why it matters. Provides a physically grounded way to model molecular flexibility, moving beyond the common "one structure per molecule" simplification in drug design pipelines.

Adaptive Confidence-weighted Expansion for Trustworthy Multi-Omics Multimodal Fusion

Mohammad Raahemi et al. | arXiv, 2026-07-22

Multi-omics fusion models often struggle when data streams are noisy or uninformative. The ACE framework dynamically reweighs each modality by reliability before fusion and produces a calibrated global trust score for every prediction, validated across BRCA, KIPAN, LGG, and ROSMAP cancer datasets.

Why it matters. Makes multimodal omics models more deployable in the clinic, where knowing how much to trust a prediction is as important as the prediction itself.

#

AI agents hit their limits in genomic surveillance, a GAN learns to read cancer histology and predict gene expression, and reinforcement learning cracks amyloid docking for neurodegenerative disease drug design.

M3-Gen: Interpretable Multimodal Generation of Gene Expression Profiles Using Clinical and Imaging Data

Panaccione et al. | arXiv, 2026-07-23

A generative adversarial network that takes histopathology images and clinical metadata as input to produce realistic, biologically meaningful gene expression profiles. Its attention-based design is interpretable — it highlights exactly which image regions drove each prediction.

Why it matters. Lowers the barrier to multi-omics cancer research by generating affordable, clinically grounded transcriptomic data from widely available H&E slides.

BioSecBench-Surveillance: A Verifiable Benchmark for AI Agents in Pathogen Genomic Surveillance

Bhasin et al. | arXiv, 2026-07-21

A 100-evaluation benchmark testing whether frontier AI agents can analyze raw pathogen sequencing data under real-world surveillance conditions. Even the strongest configuration (Opus 4.8) achieved only ~50% success, with failures concentrated in pipeline choices rather than core reasoning.

Why it matters. Establishes the first standardized way to measure whether AI can be trusted with real-time outbreak surveillance — and reveals we're not there yet.

CORAL: Learning Amyloid Fibril Ligand Docking with Cooperative Binding Rewards

Sun et al. | arXiv, 2026-07-19

A reinforcement learning framework for docking small molecules to amyloid fibrils — the protein aggregates driving Alzheimer's and Parkinson's. Unlike standard docking, CORAL's reward captures the unique cooperative stacking binding geometry of fibril targets, significantly improving pose quality over existing methods.

Why it matters. Opens a new computational pathway for designing amyloid-targeting therapeutics, where traditional docking has long struggled due to scarce structural data.

#

From graph-based spatial transcriptomics to probing the hidden reasoning of chemistry LLMs and fusing genomic language models with clinical imaging — a week of foundations that push interpretability and integration forward.

HierarchicalDAEW: Domain-Aware Edge-Weighted Graph Convolution with Evidential Uncertainty for Spatial Gene Expression from H&E

K. Chattopadhyay et al. | arXiv, Jul 23, 2026

A dual-graph GNN architecture that predicts spatially resolved gene expression from routine H&E stained histology slides. It explicitly models tissue heterogeneity via domain-aware edge weighting on a spot-level graph and fuses STRING protein-protein interaction priors with co-expression on a gene-level graph, all wrapped in evidential uncertainty for calibrated confidence intervals.

Why it matters. If robust, this brings transcriptome-wide spatial profiling from specialist instruments to widely available H&E slides — a major step toward clinic-ready spatial omics.

Chemical Chain-of-Thought Functions as a Hallucination-Prone Molecular Scratchpad

J. Li et al. | arXiv, Jul 23, 2026

Systematic analysis across four chemistry LLM families and twelve tasks reveals that chain-of-thought traces in molecular reasoning are riddled with structural hallucinations — yet still serve a causally important "scratchpad" function. Perturbing these fabricated SMILES drafts degrades output, showing CoT is neither faithful explanation nor mere post-hoc rationalization.

Why it matters. Cautions the field against treating chemical CoT as truthful reasoning and points toward process-level supervision over answer-only evaluation for trustworthy molecular AI.

Foundation-Model-Guided Radiogenomic Discovery Linking Cancer Genomes to Cancer Scans

F. Hauke et al. | arXiv, Jul 22, 2026

Pairs the Evo 2 genomic language model with radiomic features extracted from clinical tumor segmentations across 340 TCGA patients in three cancer types. This hypothesis-free sweep recovers known drivers and identifies 46 additional genes reaching FDR significance in kidney cancer — several previously unassociated with cancer but linked to Mendelian ciliopathies and cytoskeletal disease.

Why it matters. Demonstrates that combining zero-shot genomic foundation models with routine imaging can surface gene-phenotype associations invisible to conventional mutation-frequency approaches.

#

Three papers on how AI can do more than predict — it can model antibody–antigen pairing, explain its own genomic reasoning, and diagnose failure modes in perturbation biology.

Antigen-specific Antibody Multi-modal Foundation Model for Functional Antibody Design

Xiaoliang Shi et al. | arXiv, Jul 22, 2026

A new foundation model (AAMFM) jointly represents antibody sequences and 3D structures conditioned on antigen context. Using a cross-modal adapter for geometric interfaces and epitope annotations, plus preference optimization guided by a structural prior, it achieves state-of-the-art results on functional antibody design benchmarks.

Why it matters. Most protein models treat the antibody in isolation — this approach treats the antibody–antigen pair as a unit, which is how the immune system actually works.

Causal dictionary learning reveals and validates transcription-factor binding features in genomic language models

Sarwan Ali | arXiv, Jul 21, 2026

Sparse autoencoders trained on hidden activations of two genomic foundation models recover thousands of monosemantic features mapping to transcription-factor motifs. By ablating individual features during a forward pass and measuring the shift in the model's predictions, the authors establish which features the model causally uses — not just correlates with — TF binding.

Why it matters. Moves beyond correlation to prove that genomic AI models internally encode real biological concepts they actually rely on, setting a new standard for interpretability.

PertReason: A Knowledge-Grounded Benchmark and Framework for Cell-State-Conditioned Mechanistic Reasoning of Perturbation Effects

Dongkwan Kim, Yiming Gao et al. | arXiv, Jul 21, 2026

A new benchmark that tests whether models can generate mechanistically sound explanations for gene or chemical perturbations across different cell states — not just predict correct outcomes. The accompanying LLM, PertReasonLM, is trained to align outcome predictions with context-specific pathway reasoning.

Why it matters. Reveals systematic gaps between predictive accuracy and reasoning in current single-cell models, flagging failure modes invisible to standard benchmarks.

#

Three papers published the same day show how on-policy training, structure-predictor guidance, and RLVR are converging to make AI-generated biomolecules more structurally accurate and task-aligned.

ABOPD: Antibody CDR Design via On-Policy Distillation

Zhuo Yang et al. | arXiv, 2026-07-21

Antibody CDR loops are notoriously hard to design generatively because backbone errors accumulate during the denoising process. This work introduces on-policy distillation — using the native structure as a privileged signal to guide only the intermediate states the model actually visits — reducing CDR-H3 RMSD by 0.42 Å over standard fine-tuning. It closes the gap between denoising training and autoregressive inference distributions for flexible loop regions.

Why it matters. Distilling on the model's own generation path instead of offline samples could become a new standard for any protein generative model, improving realism without extra data.

DBMol: Design of High-Affinity, Target-Specific Small Molecules through Structure Prediction Models

Yiming Qin et al. | arXiv, 2026-07-21

Rather than training dedicated scoring functions, DBMol directly uses structure prediction models (Boltz-2, AlphaFold-3) as differentiable optimizers for binding affinity. An alternating cycle of gradient-based pocket interaction optimization and flow-matching projection produces discrete, chemically valid molecules. Despite no reference-ligand supervision, DBMol substantially improves pocket coverage and maintains diversity, with competitive performance under held-out evaluation.

Why it matters. Repurposing structure predictors as optimization oracles sidesteps the need for costly affinity benchmarks and could generalize to any target with a predicted complex.

LLMol: Reinforcement Learning with Verifiable Rewards for Molecular Generation

Mingxuan Ouyang et al. | arXiv, 2026-07-21

Molecular design via LLMs has been limited by supervised fine-tuning's inability to handle complex multi-objective optimization. LLMol applies RLVR — specifically GRPO — to directly reward molecules by their properties (logP, QED, structural constraints). A two-stage paradigm first teaches chemical syntax, then uses verifiable reward signals to steer generation, outperforming baselines across diverse benchmarks.

Why it matters. It adapts the RLVR paradigm that's driving frontier LLM reasoning gains to molecular design, offering a principled path to optimize any quantifiable drug property without labeled datasets.

#

Foundation models, LLMs, and reliability-aware AI methods are reshaping how we approach molecular design, single-cell analysis, and drug discovery — with three new papers pushing the boundaries of what these tools can do in spatial domains and trustworthy prediction.

Harmonised Benchmarking of Foundation Models for Single-Cell and Spatial Transcriptomics Reveals Context-Dependent Generalisation

S. Chen et al. | arXiv, 2026-07-19

Six prominent foundation models (Nicheformer, CellPLM, scGPT-spatial, GenePT, scELMo, Novae) were benchmarked across scRNA-seq, spatial transcriptomics, and Perturb-seq using a unified framework. No single model dominated across all tasks — rankings shifted dramatically depending on modality, preprocessing, tokenisation, and biological domain.

Why it matters. Provides the community with practical, evidence-based guidance for choosing the right foundation model for each specific biological question, rather than relying on scale or leaderboard scores alone.

Do Language Models Dream of Binding Molecules? Benchmarking LLMs Under Spatial Constraints

T. MacDougall et al. | arXiv, 2026-07-20

This study systematically tests whether general-purpose LLMs can reason in 3D for structure-based drug design, introducing a new benchmark called 3D-Fit that evaluates pocket-conditioned molecule generation with spatial constraints like anchor fragments and mandatory interactions. LLMs still trail diffusion models but can simultaneously handle multiple spatial conditions.

Why it matters. Opens a promising path for using widely available LLMs as flexible, multi-condition drug design tools, complementing specialized diffusion approaches.

Trustworthy Protein-Ligand Binding Affinity Prediction via Reliability-Aware Multi-Engine Fusion

Y. Hong et al. | arXiv, 2026-07-20

RELIABLE-BA treats each docking engine as an evidential expert, scaling uncertainty through learned reliability from molecular context. The framework fuses experts via closed-form aggregation, delivering substantially better uncertainty calibration and up to 25% prediction-error reduction when filtering to high-confidence pairs.

Why it matters. Moves binding affinity prediction beyond simple consensus scoring by providing principled, interpretable confidence measures — a step toward truly trustworthy AI-guided drug discovery.

#

This week, AI methods leap beyond static structure prediction — modeling protein motion across time, fusing multi-modal spatial omics, and fine-tuning LLMs for molecular geometry.

DyneTrion: A Spatio-temporally Coherent Generative Emulator for Protein Dynamics Across Timescales

Kaihui Cheng et al. | arXiv, July 15

A generative emulator that simulates protein dynamics by jointly enforcing geometric symmetry (via invariant point attention), structural consistency, and temporal coherence across 100-ns to microsecond trajectories — reproducing MD-level flexibility and free-energy landscapes at a fraction of the computational cost.

Why it matters. Moves the field closer to time-resolved protein modeling, bridging the gap between static structure prediction and understanding how proteins actually move during ligand binding, allostery, and catalysis.

LATTICE: Graph Self-Supervised Learning for Multimodal Spatial Omics Integration

Jagan Mohan Reddy Dwarampudi et al. | arXiv, July 15

A graph-based framework that unifies five spatial omics modalities (Visium RNA, scMultiome RNA, scMultiome ATAC, spatial ATAC, and spatial CUT&Tag) through a TransformerConv encoder trained with masked reconstruction and cross-modal alignment — boosting clustering concordance and spatial contiguity on a melanoma cohort of 54,912 spots.

Why it matters. Offers a practical blueprint for integrating the flood of multi-assay spatial data now becoming routine in cancer and developmental biology studies.

How Well Can Frontier Large Language Models Generate Structures? High Quality Prediction of Molecular Geometries with Help from Fine-Tuning

Joseph M. Cavanagh et al. | arXiv, July 15

Fine-tuning frontier LLMs on molecular coordinates (Z-matrices or Cartesian) yields surprisingly accurate equilibrium structures and diverse conformers of drug-like molecules — outperforming specialized deep learning models. Z-matrices, which encode relational geometry, prove to be a superior "grammar" for LLM adaptation to molecular structure.

Why it matters. Demonstrates that general-purpose LLMs can be repurposed for computational chemistry tasks with minimal training data, potentially democratizing access to conformer generation and molecular geometry prediction.

#

A Vision Foundation Model for Single-Cell Biology via Spatial Gene Cartography

Ridvan Yesiloglu et al. | arXiv, 2026-07-15

Instead of treating each cell as a gene-token sequence, this work renders transcriptomes as images — using optimal transport to place genes at fixed spatial positions so co-expressed programs appear as local visual texture. A vision transformer pretrained by masked image modeling learns rich representations that outperform language-model baselines on cell-type classification.

Why it matters. Replaces language-centric single-cell foundations with a vision paradigm that preserves gene-gene relationships and expression magnitudes.

Accelerated Descriptor-Free Path Sampling for Protein-Ligand Binding Kinetics

Simon M. Lichtinger et al. | arXiv, 2026-07-16

Binding kinetics are critical for drug efficacy but notoriously hard to compute. The authors merge unbiased path sampling with a descriptor-free equivariant graph neural network that models the committor probability directly — eliminating hand-crafted collective variables while converging escape rates orders-of-magnitude faster.

Why it matters. Brings kinetics calculations — long a bottleneck in rational drug design — firmly into the AI-native, feature-free regime.

Exploring the Alignment of Generation and Understanding in Protein Structure Modeling

Junde Xu et al. | arXiv, 2026-07-15

Generative protein models excel at structure prediction, but do they truly "understand" proteins? This systematic benchmark reveals that many top-tier generative models learn suboptimal representations when tested on downstream understanding tasks like function annotation — echoing the generation-understanding gap known from computer vision.

Why it matters. Highlights a blind spot in protein AI: strong generation ≠ strong representation learning, urging model developers to evaluate understanding alongside generation.

#

Cross-domain AI for molecular systems and single-cell vision models are reshaping how we represent, generate, and understand biological data at atomic and cellular scales.

A Vision Foundation Model for Single-Cell Biology via Spatial Gene Cartography

Yu et al. | arXiv: 2607.14163, July 15, 2026

Most single-cell foundation models treat cells as token sequences, discarding gene-gene relationships and expression magnitudes. scVision instead renders each cell's transcriptome as a continuous image by placing genes at fixed spatial positions via optimal transport — co-expressed genes become spatial neighbors, so gene programs appear as local texture. Pretrained via masked image modeling on 72 million human cells, it outperforms existing foundation models on zero-shot cell-type annotation and multi-study integration without ever seeing a batch label.

Why it matters. Reframes single-cell representation learning as a vision problem, unlocking decades of computer-vision advances for transcriptomics.

SinAE: A Single-Architecture Flow-Matching Autoencoder for Cross-Domain Atomic Systems

Ren et al. | arXiv: 2607.12380, July 14, 2026

Molecules, crystals, and proteins each have their own specialized generative models with graph, equivariant, or frame-based operators — fragmenting the field. SinAE uses a vanilla Transformer encoder-decoder with no domain-specific architectures, shifting the reconstruction burden to an iterative flow-matching decoder. It achieves near-lossless reconstruction across all three domains and demonstrates that joint molecule-crystal training strictly improves both, proving cross-domain transfer through a shared atomic latent.

Why it matters. A truly unified generative architecture could end the fragmentation of molecular AI and enable data-scarce domains to leverage cross-domain signal.

Generating Developable 3D Molecules via Pocket-Conditioned Diffusion and Property-Aware Optimization

Gao et al. | arXiv: 2607.12349, July 14, 2026

Structure-based drug design models often optimize only for binding affinity while ignoring ADMET properties critical for real-world drug development. conDitar-dev combines a multi-scale pocket representation module, a pocket-conditioned diffusion model, and a generation-time property optimizer — delivering molecules with both strong binding and favorable developability. Generated candidates for PD-L1 and CSF1R were experimentally synthesized and biologically tested, with PD-L1 hits showing Kd values of 3.49–3.75 μM and CSF1R hits reaching 200 nM IC50.

Why it matters. One of the first diffusion-based SBDD frameworks to bridge the gap between computational design and experimental validation with developable molecules.