Skip to content
World clockEU--:--UK--:--USA--:--CN--:--PLDEFRIT中文EN

portal about AI and technologyevents · analysis · interviews · technical background

Search
LIVE
›

Three hundred lab measurements sharpen virtual screening

Fine-tuning Boltz-2's affinity heads on a small set of activity labels improved early enrichment, and the team still kept rescoring to ten percent of the candidates.

ScienceNewsSofia MarchettiPublished: 22 September 20265 min readSources 2
Three hundred lab measurements sharpen virtual screening

Virtual screening means sifting through a huge library of compounds to find the ones worth testing in a lab. The experimental budget is limited, so what counts is not full classification but how early the truly active molecules appear in the ranking. A paper by Japanese authors asks whether Boltz-2 can be prepared for that job with a few dozen to a few hundred measurements from one specific assay.

How much data is enough

The authors fine-tuned Boltz-2's affinity heads, feeding the model between 40 and 300 binary activity labels. They evaluated it retrospectively on eight targets from the MF-PCBA set. With 300 measurements, the number of active compounds in the top one percent of the ranking rose by a geometric mean factor of 1.77 over the untuned variant. Average precision improved 2.14-fold. The result depends on the target and on how the training set was chosen, so the authors treat it as a signal, not a rule.

The second experiment looks at compute cost. Instead of rescoring the whole candidate set with the tuned head, the researchers rescored only the roughly ten percent of proposals that Boltz-2 had scored highest. With that narrowing, the recovery of hits stayed comparable to full rescoring. The practical takeaway is simple: if the model already has a good feel for the coarse scale, refining the tail of the list is enough.

The authors themselves note that the study is retrospective. The labels come from ready-made sets rather than a new experiment, so the numbers describe how the method behaves on known data, not how well it works in an unknown lab. The practical conclusion still holds. Instead of training a model from scratch on tens of thousands of measurements, a team can fine-tune the heads on a set that fits the budget of a single project and still gain on the ordering of candidates.

Data processing in proteomics

The measurement side is developing in parallel. Another paper presented ProteoEM, an open Python library that estimates the abundance of proteins and proteoforms from single-molecule affinity traces. The problem is that imperfect, nonspecific probe binding lets one trace fit several possible identifications. Rather than assigning a trace to a single candidate, the author proposes splitting the uncertainty by weight in an expectation-maximization scheme, using response rates of the probe calibrated in advance. The model reports indistinguishable proteoforms as groups instead of forcing an arbitrary choice.

Both approaches share one idea: in biomedical research the value lies not in abstract accuracy but in how a model's decision translates into the next, costly step in the lab. That is why model evaluation is starting to count not average quality across the whole set but the work done in the narrow window that actually matters.

Comments 0

Sources

2
  1. 01Adapting Boltz-2 with limited experimental activity data improves early enrichment in virtual screeningEN
  2. 02ProteoEM: probabilistic protein abundance estimation from iterative affinity tracesEN

All figures and quotations in this text come from the sources listed below.

Content prepared by the editorial team with AI assistance.

Sofia Marchetti

Sofia Marchetti

Science and health

Sofia Marchetti covers science and health for FLASH24, working from primary literature, preprints, and agency data rather than press releases. She checks sample sizes, confidence intervals, and whether a study's numbers match its abstract before filing. She interviews researchers and clinicians directly, tracks conference calendars for embargoed results, and compares new findings with earlier trials on the same question. Outside the newsroom she works on materials physics and stargazes through a home telescope, which keeps her close to how measurement error actually behaves. She does not publish a health claim without a named source and the underlying data.

Newsroom →

Comments

0
  1. No comments yet — be the first.

Write a comment

Comments are public. We do not publish abuse, spam or advertising.