bims-fragic Biomed News
on Fragmentomics
Issue of 2026–08–16
two papers selected by
Laura Mannarino, Humanitas Research



  1. Mol Biomed. 2026 Aug 12. pii: 136. [Epub ahead of print]7(1):
      Ultra-low-pass whole-genome sequencing (ULP-WGS) of cell-free DNA (cfDNA) offers a cost-efficient strategy for cancer detection, but its clinical application is limited by extreme data sparsity and poor model generalization. We developed Fragmentia-AI™ WGS, a mutation-calling-independent framework that uses a transformer-based multiple-instance learning architecture with sequential fine-tuning across tumor fraction (TF) strata to extract latent cancer-associated signals from ULP-WGS data. Model performance was evaluated in multiple independent cohorts, including a pan-cancer test set covering 17 cancer types, an external public dataset generated on a different sequencing platform, and a technical variability cohort with heterogeneous pre-analytical and experimental conditions. Clinical relevance was assessed by correlating model predictions with progression-free survival (PFS) in patients with advanced non-small cell lung cancer receiving chemoimmunotherapy. Sequential fine-tuning across TF strata significantly improved performance in low-TF samples, achieving a 35.6% relative increase in AUC compared with high-TF-only training (0.884 vs. 0.652). In the independent test cohort, the model achieved an overall AUC of 0.930, with consistent performance across TF strata and cancer types. External validation confirmed robust cross-platform generalizability (AUC: 0.929; sensitivity: 0.78; specificity: 0.92). The model maintained stable classification performance despite score fluctuations associated with pre-analytical and technical variables. Importantly, model-negative status, defined as a prediction score below the training-derived cutoff, remained significantly associated with improved PFS compared with model-positive status (HR = 0.49, 95% CI: 0.29-0.82) after multivariable adjustment. Collectively, this framework enables robust cancer detection and clinically meaningful risk stratification from highly sparse cfDNA sequencing data.
    Keywords:  Cell-free DNA; Fragmentomics; Large language model; Low tumor fraction; Ultra-low-pass WGS
    DOI:  https://doi.org/10.1186/s43556-026-00538-w
  2. PLoS Comput Biol. 2026 Aug 10. 22(8): e1013356
    IMPROVE-consortia
      Circulating tumor DNA (ctDNA) is emerging as a promising biomarker for postoperative monitoring of cancer patients. Precise estimation of circulating tumor fraction is crucial for evaluating treatment effects and timely detection of disease recurrence. All current ctDNA detection methods that utilize whole-genome sequencing (WGS) data rely on the reference genome alignment of sequencing reads and often apply separate tools for detecting different variant types. However, various bioinformatic analysis confounders and the application of external variant calling tools could be avoided by analyzing k-mers from unaligned sequencing reads. While k-mer-based methods have successfully been applied for somatic variant validation and detection, the potential of k-mer-based ctDNA detection is unexplored. We have developed a tumor-informed alignment-free ctDNA detection tool called ctDNAmer that detects tumor-specific somatic variation directly from unaligned sequencing data by identifying k-mers unique to the tumor DNA. ctDNAmer detects variant information across the genome by comparing the primary tumor and germline WGS data and accounts for sample-specific germline variability and technical noise in the same framework. We tested the utility of ctDNAmer for tumor fraction estimation on postoperative plasma cfDNA WGS data (mean sequencing depth ~ 28x) from 90 stage III colorectal cancer patients with three years of follow-up. The tumor fraction (TF) estimates agreed with the available clinical information and ctDNA was detected in 77% (17/22) of recurring patients with a median lead time of 8 months compared to radiological imaging. We further validated ctDNAmer's tumor fraction estimates based on a comparison with the mean cfDNA allele frequencies of somatic clonal SNVs identified from aligned primary tumor sequencing data. The TF estimates showed a strong Pearson correlation of 0.897 with the mean allele frequencies and improved ctDNA detection results across samples with an AUC of 0.79 compared to 0.75 if the mean allele frequency of clonal mutations is used.
    DOI:  https://doi.org/10.1371/journal.pcbi.1013356