Integrative Biomedical Research

Integrative Biomedical Research (Journal of Angiotherapy) | Online ISSN  3068-6326
463
Citations
1.8m
Views
768
Articles
Your new experience awaits. Try the new design now and help us make it even better
Switch to the new experience
Figures and Tables
REVIEWS   (Open Access)

Protein Language Models for Predicting BRCA1/BRCA2 Variant Pathogenicity: A Computational Bridge Toward Precision Oncology

Ghaith Kamil Jawad 1*, Abdulsamie Hassan Alta'ee 2, Oday Jasim Alsalihi 3

+ Author Affiliations

Integrative Biomedical Research 10 (1) 1-8 https://doi.org/10.25163/biomedical.10110923

Submitted: 17 May 2026 Revised: 08 July 2026  Published: 19 July 2026 


Abstract

Breast and ovarian cancers linked to pathogenic BRCA1 and BRCA2 alterations impose a substantial and growing global health burden, and the clinical utility of PARP-inhibitor therapy now depends almost entirely on correctly telling a harmful variant apart from a harmless one. That task has become harder, not easier, as sequencing volumes have grown: nearly four in ten variants identified through clinical gene panels are still classified as variants of uncertain significance (VUS), leaving clinicians and patients in a difficult holding pattern. We conducted a structured narrative synthesis of peer-reviewed literature addressing protein language models (pLMs), multi-omics deep learning architectures, and their translational application to BRCA variant interpretation and precision oncology, following a reproducible search-and-screening workflow across major biomedical databases. Evidence was extracted, tabulated, and thematically organized across four domains: representation learning and data fusion, clinical risk stratification, synthetic-lethality target discovery, and translational barriers to clinical deployment. Across the reviewed literature, pLM- and transformer-based architectures (including AlphaMissense, ESM3, DNABERT-S, SetQuence, and SetOmic) consistently outperformed classical machine learning baselines in variant- and tumor-classification tasks, with several multi-omics fusion frameworks (e.g., SetOmic, MOGONET, TMO-Net) achieving accuracy or F1-scores exceeding 0.90 in pan-cancer and breast-cancer cohorts. Graph-based synthetic-lethality models (DGIB4SL, KR4SL, MAGICAL) further extended this predictive power to therapeutic target discovery beyond canonical BRCA-PARP biology. However, persistent obstacles — dataset homogeneity, limited model interpretability, and underrepresentation of non-European ancestries — continue to restrict real-world generalizability. Protein language models represent a genuinely transformative, though not yet fully mature, tool for resolving BRCA variant ambiguity and expanding equitable access to precision oncology; their clinical adoption will likely hinge on interpretability, prospective validation, and deliberate correction of demographic bias in training data.

Keywords: BRCA1/BRCA2; protein language models; variant pathogenicity prediction; precision oncology; homologous recombination deficiency; PARP inhibitors; multi-omics data integration

1. Introduction

Breast cancer remains, by a considerable margin, the malignancy most frequently diagnosed among women worldwide, and it continues to rank among the leading causes of cancer-related death, with more than 2.3 million new diagnoses recorded each year (Kotsifaki et al., 2025; Sabit et al., 2026). Ovarian cancer, though far less common, tells a similarly troubling story — modest in incidence, disproportionately severe in mortality (Webb & Jordan, 2024; Wolde & Belay, 2026). Both diseases arise through a tangle of somatic mutations, epigenetic drift, and shifting tumor-microenvironment signals that rarely act in isolation (Sabit et al., 2026; Wolde & Belay, 2026). At the center of that tangle, for a meaningful subset of patients, sit two genes whose names have become almost synonymous with hereditary cancer risk: BRCA1 and BRCA2 (Ismail et al., 2024; Kotsifaki et al., 2025).

It is worth pausing on what these genes actually do, because their function explains much of what follows. BRCA1 and BRCA2 are custodians of genomic integrity, core members of the homologous recombination repair (HRR) pathway responsible for fixing double-stranded DNA breaks with high fidelity (Ismail et al., 2024). Without functional HRR, such damage does not simply disappear — it compounds. Inactivating mutations in BRCA1 or BRCA2 disable this repair system, and the resulting genomic instability translates, over time, into a markedly elevated lifetime risk of breast and ovarian cancer (Ismail et al., 2024; Shah et al., 2025). For much of the last several decades, clinicians confronting BRCA-associated disease had little choice but to reach for the same blunt instruments used against cancer generally, applied broadly and often at considerable cost to patients' quality of life (Ismail et al., 2024; Kotsifaki et al., 2025).

That began to change with precision oncology, which has gradually replaced one-size-fits-all treatment with therapies aimed at specific molecular vulnerabilities (Rescigno & Greystoke, 2026; Sabit et al., 2026). The clearest success story involves poly (ADP-ribose) polymerase (PARP) inhibitors — olaparib, rucaparib, and niraparib among them — which exploit synthetic lethality (Ismail et al., 2024; Shah et al., 2025). In tumors already deficient in homologous recombination, PARP inhibitors trap PARP1/2 on DNA and block base-excision repair, converting otherwise manageable single-strand breaks into lethal double-strand breaks during replication; healthy cells largely tolerate this, BRCA-deficient tumor cells do not (Shah et al., 2025). Large trials have confirmed that PARP-inhibitor maintenance therapy meaningfully extends progression-free and overall survival in patients carrying pathogenic BRCA alterations (Rescigno & Greystoke, 2026; Shah et al., 2025). None of that benefit, though, is available to a patient whose variant has not been correctly classified — which is where the promise of precision oncology runs into a messier reality (Eniu et al., 2023; Rescigno & Greystoke, 2026).

The bottleneck is not sequencing capacity; next-generation sequencing (NGS) now identifies genetic alterations faster than clinical science can make sense of them (Ismail et al., 2024; Rescigno & Greystoke, 2026). The bottleneck is interpretation. Roughly 40% of newly detected missense variants across clinical gene panels cannot yet be confidently labeled benign or pathogenic and are filed away as variants of uncertain significance, or VUS (Rescigno & Greystoke, 2026). This is not a minor administrative inconvenience — missense VUS clustering in functionally sensitive regions, such as the C-terminal BRCT domains of BRCA1 or the sprawling coding sequence of BRCA2, routinely complicate an otherwise straightforward clinical report (Ismail et al., 2024; Rescigno & Greystoke, 2026). The stakes cut in both directions: overcalling a benign variant can send a patient toward unnecessary prophylactic surgery, while undercalling a truly pathogenic one denies access to a therapy that might have changed their prognosis (Rescigno & Greystoke, 2026).

Traditional routes to resolving this ambiguity — functional assays performed one variant at a time, structural modeling, multigenerational pedigree analysis — were never designed for today's data volume (Ismail et al., 2024; Rescigno & Greystoke, 2026). They are slow, costly, and not built to scale (Rescigno & Greystoke, 2026; Sabit et al., 2026). This widening gap between sequencing speed and interpretive capacity has created real urgency around computational alternatives that can accelerate variant pathogenicity prediction without sacrificing clinical validity (Gentile et al., 2026; Rescigno & Greystoke, 2026).

Artificial intelligence and machine learning have moved into this gap with some force (Barua et al., 2026; Lin et al., 2026). Earlier approaches — random forests, support vector machines, conventional deep neural networks — showed promise but struggled with the dimensionality of raw sequence data, often overfitting or offering little insight into why a prediction was made (Qiu et al., 2026). A more substantial leap arrived when researchers began borrowing architectures from natural language processing and applying them to biological sequences (Jurenaite et al., 2024). The premise: treat a protein or genomic sequence as a language, learn which "words" plausibly follow which others, and in doing so capture the physical, chemical, and evolutionary constraints that shaped that sequence (Jurenaite et al., 2024).

Protein language models, or pLMs, have emerged from this line of work as unusually capable tools for the problem this review addresses (Corso et al., 2026). Trained on vast, label-free evolutionary databases, they learn the "grammar" of protein function without curated ground-truth labels (Corso et al., 2026). AlphaMissense offers proteome-wide missense-variant predictions with sensitivity and specificity marking a genuine step change (Corso et al., 2026). ESM3 uses evolutionary protein data not only to predict mutational consequences but to generate novel functional proteins outright (Corso et al., 2026). At the nucleotide level, DNABERT-S applies BERT-style architectures, trained across millions of genomic sequences, to interrogate mutation signatures and regulatory context (Corso et al., 2026). What unites these systems is their capacity to separate genuine driver mutations from biological noise (Jurenaite et al., 2024; Qiu et al., 2026).

The clinical implications are considerable. By flagging pathogenic variants quickly, pLMs support detection of homologous recombination deficiency (HRD), present in roughly half of high-grade serous ovarian carcinomas and in 50–70% of triple-negative breast cancers (Shah et al., 2025). Tools such as HRProfiler mine exome-sequencing data for genomic scars — loss of heterozygosity, telomeric allelic imbalance, large-scale state transitions — achieving AUC values above 0.90 for HRD detection and widening the pool eligible for PARP-inhibitor therapy (Shah et al., 2025). Layered onto electronic health records, these predictions can feed decision-support systems that assist molecular tumor boards, reducing subjectivity and manual workload (Lin et al., 2026; Sabit et al., 2026; Rescigno & Greystoke, 2026). Large language model ensembles like OncoChat have gone further still, learning gene-gene interactions directly from targeted panels and rediscovering BRCA-PARP1 synthetic lethality in an unsupervised fashion (Liu et al., 2025). Taken together, this evidence suggests a genuine, if not yet complete, shift away from static, morphology-driven diagnosis toward something more dynamic and probabilistic (Barua et al., 2026; Sabit et al., 2026).

Still, computational promise and clinical readiness are not the same thing. Translating in silico pLM predictions into everyday oncology practice demands a clear-eyed evaluation of biological grounding, analytical performance, and the practical obstacles between a benchmark result and a bedside decision (Lin et al., 2026; Sabit et al., 2026). This review synthesizes current evidence at the intersection of protein language models, BRCA variant pathogenicity, and precision oncology, offering clinicians and computational biologists a clearer map of where this field currently stands (Corso et al., 2026; Rescigno & Greystoke, 2026).

To that end, this review pursues four interconnected objectives. The first is to evaluate the biological and structural mechanisms of BRCA1 and BRCA2 variants, detailing how domain-specific alterations disrupt homologous recombination repair, and to characterize the clinical, diagnostic, and ethical challenges posed by the high prevalence of variants of uncertain significance (VUS) in global oncology workflows (Ismail et al., 2024; Rescigno & Greystoke, 2026). The second is to conduct a comparative analytical evaluation of state-of-the-art protein language models — including AlphaMissense, ESM3, and DNABERT-S — and related machine learning architectures, critically assessing their sensitivity, specificity, and computational complexity in predicting BRCA variant pathogenicity (Jurenaite et al., 2024; Corso et al., 2026). The third is to analyze the clinical integration of in silico pathogenicity predictions with multi-omics data, including genomic scars, promoter methylation, transcriptomic signatures, and digital pathology, and to assess how multimodal data fusion optimizes patient risk stratification, HRD detection, and PARP-inhibitor response prediction (Sabit et al., 2026; Shah et al., 2025). The fourth, and in many ways the most forward-looking, is to identify current translational bottlenecks — algorithmic bias, data privacy, computational overhead, and the underrepresentation of diverse ancestries in public genomic databases — and to propose a practical roadmap toward standardization, regulatory qualification, and equitable clinical adoption of AI-driven biomarkers (Lin et al., 2026; Qiu et al., 2026; Rescigno & Greystoke, 2026; Shah et al., 2025).

2. AI and Multi-Omics Convergence in BRCA-Driven Precision Oncology

2.1 The Multi-Omics Paradigm in Oncology: Why One Layer Is Never Enough

Cancer, whatever else it is, is rarely explicable through a single data layer, and BRCA-associated disease is no exception. Genomic sequence alone tells us that a variant exists; it does not, by itself, tell us how that variant behaves once it meets a real tumor's transcriptional, epigenetic, and microenvironmental context (Rescigno & Greystoke, 2026; Shanmugam & Ravikumar, 2026). Single-omics assays — a gene panel here, a methylation array there — capture a partial silhouette of disease biology, missing the downstream protein states, post-translational modifications, and regulatory dynamics that often determine whether a tumor resists therapy (Sabit et al., 2026; Shanmugam & Ravikumar, 2026).

Modern oncology has responded, somewhat unevenly, by pursuing vertical multi-omics integration — correlating genomic, transcriptomic, proteomic, metabolomic, and epigenomic signals within the same patient cohort (Ouhmouk et al., 2025; Shanmugam & Ravikumar, 2026). Done well, this kind of fusion surfaces synergistic mechanisms that no single platform would reveal on its own, and it tends to reduce false-discovery rates relative to single-platform analyses (Sabit et al., 2026; Shanmugam & Ravikumar, 2026). Done poorly, however, it runs headlong into a genuinely hard statistical problem. Multi-omics data is high-dimensional, sparse, and technically messy — a fairly textbook instance of the so-called curse of dimensionality (Ouhmouk et al., 2025). Traditional statistical models tend to buckle under these conditions, undone by multicollinearity and, more often than not, by sample sizes that were never large enough to begin with (Ouhmouk et al., 2025; Jurenaite et al., 2024; Barua et al., 2026). It is largely in response to this failure mode that deep learning has become the default computational strategy — not because it is fashionable, but because it can automatically learn hierarchical feature representations that harmonize disparate data modalities in a way older methods simply cannot (Ouhmouk et al., 2025; Lin et al., 2026; Sabit et al., 2026).

2.2 Computational Foundations: How Deep Learning Architectures Handle Biological Complexity

Deep learning approaches to multi-omics oncology tend to fall, loosely, into three families — generative, non-generative, and hybrid — distinguished mainly by how they fuse heterogeneous inputs (Ouhmouk et al., 2025). Non-generative architectures, graph neural networks (GNNs) and graph attention networks (GATs) among them, map inputs directly onto clinical endpoints and have the useful property of being able to embed prior biological knowledge, such as protein-protein interaction networks, straight into the learning architecture (Ouhmouk et al., 2025). The Multimodal Onco-Graph Convolutional Network (MOGONET) illustrates this well: separate graph convolutional layers learn omics-specific embeddings, and a downstream module captures cross-omics label correlations for multi-class tumor classification (Wang & Ballester, 2021; Ouhmouk et al., 2025).

A more fundamental limitation of many earlier classifiers, though, is their insistence on fixed-dimensionality inputs (Jurenaite et al., 2024). A network trained to expect exactly N genomic loci breaks, sometimes badly, when N changes — and worse, it can become sensitive to the arbitrary order in which those loci are presented (Jurenaite et al., 2024). Jurenaite et al. (2024) addressed this directly with SetQuence and SetOmic, deep networks built on permutation-invariant set representations rather than fixed vectors. SetQuence leans on pretrained DNA language models, such as DNABERT, to encode variably long, variant-associated sequences, which are then aggregated — via pooling multi-head attention — into a single coordinate-free embedding (Jurenaite et al., 2024). SetOmic extends the same logic to quantitative transcriptome counts, and in doing so demonstrates markedly better robustness and less information loss across coding and non-coding regions than fixed-input alternatives (Jurenaite et al., 2024).

Generative models occupy a different, complementary niche. Variational autoencoders (VAEs) — OmiEmbed, CustOmics, and TMO-Net among the better-known examples — learn the joint probability distribution of high-dimensional omics data, compressing it into a shared, low-dimensional latent space via encoder-decoder pairs (Ouhmouk et al., 2025; Zhang et al., 2021). One practical virtue of this family is resilience to missing data: because the latent space is learned jointly, researchers can often reconstruct an unmeasured omics modality directly from what the model has already learned about the others (Ouhmouk et al., 2025).

More recently, transformer-based foundation models have introduced something closer to a generalist medical AI paradigm (Corso et al., 2026). Large language models are increasingly being repurposed for genomic diagnostics, converting raw, unstructured somatic mutation data and clinical phenotypes into structured, dialogue-like representations a model can reason over (Liu et al., 2025; Corso et al., 2026). OncoChat is a clear illustration — a diagnostic LLM developed on 158,836 tumor samples spanning 69 cancer types across 19 institutions, instruction-tuned on structured clinicogenomic text. It outperforms earlier classifiers such as OncoNPC and GDD-ENS and shows notable stability in predicting tissue of origin for cancers of unknown primary (Liu et al., 2025).

2.3 Clinical and Translational Applications in Precision Oncology

Where multi-omics AI has moved beyond proof-of-concept is in several fairly distinct clinical touchpoints, each worth considering on its own terms (Sabit et al., 2026; Lin et al., 2026).

2.3.1 Molecular Risk Stratification and Biomarker Staging

TNM staging, for all its familiarity, is essentially a macro-anatomic system, and it was never designed to capture the biological heterogeneity that increasingly defines how tumors actually behave (Sabit et al., 2026; Roy et al., 2027). Integrated multi-omics classifiers offer a more refined alternative. Roy et al. (2027), for instance, combined microRNA and messenger RNA datasets from the TCGA breast cancer cohort using gradient-boosting algorithms such as XGBoost, and reported a 5–10% improvement in clinical-stage classification accuracy over single-platform classifiers — largely by capturing non-linear interactions between miRNA-mediated silencing and its downstream mRNA targets. In a related but visually distinct approach, whole-slide pathology images can be fused with molecular assays: the AnchorMIL framework reframes Oncotype DX recurrence-score prediction from whole-slide images as a joint regression-and-classification task, achieving strong concordance across both TCGA-derived and independent cohorts (Koyun et al., 2026).

2.3.2 Neoadjuvant Therapy Response Prediction

Predicting pathological complete response (pCR) early in neoadjuvant chemotherapy matters because it lets clinicians spare patients unnecessary toxicity when a regimen is unlikely to work (Sabit et al., 2026; Lash & Valero, 2026). Stacked ensemble frameworks that fuse transcriptomic profiling with clinical variables — capturing distinct inductive biases across tree-based, kernel, and deep-network learners — have produced well-calibrated predictions of chemotherapy responsiveness (Lash & Valero, 2026). Multimodal deep learning that additionally incorporates pre-treatment MRI radiomics and digital pathology tends to do even better, with reported accuracy gains of 10–20% over unimodal baselines (Sabit et al., 2026; Lin et al., 2026).

2.3.3 Synthetic Lethality and Novel Target Discovery

Synthetic lethality remains one of the more powerful conceptual tools for target discovery in oncology, and its clinical proof of concept is, of course, the success of PARP inhibitors in BRCA1/2-mutant disease (Schäffer et al., 2024; Li et al., 2026; Wolde & Belay, 2026). What computational biology has added more recently is a way to look for synthetic-lethal partners beyond the canonical DNA-repair pathway. Frameworks such as DGIB4SL, KR4SL, and MAGICAL fuse multi-omics data with protein-protein interaction networks and knowledge-graph representations, leaning on graph neural networks and information-bottleneck principles to surface higher-order regulatory motifs and, from there, candidate target combinations — pairing PARP inhibitors with novel STING agonists, for instance, to resensitize resistant tumors (Li et al., 2026; Sabit et al., 2026).

2.3.4 AI-Driven Liquid Biopsies and Minimal Residual Disease Monitoring

Non-invasive monitoring — circulating tumor cells, circulating tumor DNA, cell-free DNA fragmentomics, methylation profiling — is quietly transforming how clinicians track disease over time, without repeated invasive biopsies (Sabit et al., 2026; Shanmugam & Ravikumar, 2026). Machine learning applied to multi-analyte liquid biopsy panels can fuse fragment-length density distributions with epigenetic methylation signals to estimate minimal residual disease and monitor treatment response essentially in real time (Sabit et al., 2026; Rescigno & Greystoke, 2026; Lin et al., 2026). Commercial platforms such as Grail Galleri already apply this logic at scale, achieving meaningful early-detection sensitivity across more than 50 distinct cancer types (Shanmugam & Ravikumar, 2026). At the single-cell level, foundation models like scGPT and scBERT extend this further still, annotating heterogeneous tumor-infiltrating immune populations with strong accuracy and tracking drug-resistant subclonal trajectories as they emerge (Ouhmouk et al., 2025).

2.4 Major Translational Roadblocks in Clinical Integration

None of this, however, translates automatically into bedside utility, and it would be misleading to end this review of the literature without naming the obstacles plainly (Ouhmouk et al., 2025; Lin et al., 2026).

2.4.1 Overreliance on Centralized TCGA Data and Technical Artifacts

A large majority of published deep learning models in computational oncology — well over 90%, by some estimates — are trained and validated almost exclusively on The Cancer Genome Atlas or METABRIC (Ouhmouk et al., 2025; Qiu et al., 2026). These are excellent resources, but they are also idealized, retrospective, and largely disconnected from the messiness of real-world electronic health record data or population diversity (Ouhmouk et al., 2025). Worse, deep networks are notoriously sensitive to domain shift and batch effects: Dehkharghanian et al. (2023) showed that models trained on TCGA pathology whole-slide images could predict a slide's acquisition site with up to 86% accuracy, using cues that had nothing to do with underlying biology — tissue staining protocol, scanner model, that sort of thing (Ouhmouk et al., 2025). A model that has quietly learned to recognize which hospital a slide came from is not, whatever its benchmark scores suggest, learning cancer biology.

2.4.2 The "Black Box" Interpretability Gap

Clinicians, understandably, are reluctant to act on predictions they cannot interrogate, and regulators tend to feel much the same way (Ouhmouk et al., 2025; Lin et al., 2026; Sabit et al., 2026). Many of the deep architectures described above remain functionally opaque — accurate, perhaps, but not explicable in any way a treating physician could defend in a tumor board discussion (Ouhmouk et al., 2025; Lin et al., 2026). Post-hoc explainability tools, SHapley Additive exPlanations (SHAP) and Local Interpretable Model-agnostic Explanations (LIME) chief among them, are increasingly being folded into multi-omics pipelines like SetOmic, and the results are at least somewhat reassuring: the features these tools flag as most influential tend to align with established cancer biology, including known attributions to PIK3CA and S100A11 in breast cancer cohorts (Jurenaite et al., 2024; Ouhmouk et al., 2025). Biologically constrained architectures such as DeepOmix and DeepKEGG take a related but more structural approach, embedding pathway and gene-ontology layers directly into the network to keep its reasoning at least loosely tethered to known biology (Ouhmouk et al., 2025).

2.4.3 Data Harmonization, Batch Effects, and Algorithmic Bias

Combining data generated across different platforms, laboratories, and protocols introduces systematic technical variation — batch effects — that, left uncorrected, can masquerade as biological signal (Ouhmouk et al., 2025; Jurenaite et al., 2024). Careful data cleaning and batch correction, whether through the ComBat algorithm or self-normalizing neural network architectures, are not optional extras here; they are close to a prerequisite for trustworthy results (Ouhmouk et al., 2025; Yin et al., 2026).

There is also a harder, more uncomfortable problem underneath the technical one. Genetic ancestry audits suggest that more than 82% of TCGA cases represent individuals of European descent, leaving African and Indigenous populations substantially underrepresented (Ouhmouk et al., 2025). Models trained on that skewed foundation tend, unsurprisingly, to degrade in performance — and to misclassify more often — when applied outside the population they were trained on (Ouhmouk et al., 2025; Lin et al., 2026). Addressing this will likely require more than good intentions: prospective recruitment of genuinely diverse cohorts, subgroup-stratified fairness auditing, and adversarial debiasing or reweighting strategies built into training from the outset (Ouhmouk et al., 2025; Lin et al., 2026).

2.5 Section Summary

Taken as a whole, the literature reviewed here converges on a fairly consistent picture. Multi-omics deep learning, and protein language models specifically, have genuinely outperformed earlier computational approaches across tumor classification, risk stratification, synthetic-lethality discovery, and liquid-biopsy monitoring (Ouhmouk et al., 2025; Sabit et al., 2026; Corso et al., 2026). What remains unresolved — data homogeneity, interpretability, and demographic bias, primarily — is not a minor footnote to that progress but arguably the central story of where the field goes next (Ouhmouk et al., 2025; Lin et al., 2026). Figures 1 through 4 summarize, respectively, the BRCA-HRR-PARP synthetic lethality axis that motivates this review, the sequence-to-pathogenicity inference pipeline characteristic of protein language models, the study identification workflow used to compile this review's evidence base, and the broader multi-omics fusion architecture that underlies much of the translational work discussed above

3. Methods

This review followed a structured, reproducible acoxuflev opebkavln methodology, informed by PRISMA principles for transparent reporting, so that the search strategy, eligibility criteria, and hchs xadvvzdlzc process could, in principle, be replicated by an independent team using the search terms and inclusion criteria specified below (Figure 3).

3.1 Search Strategy and Information Sources

We searched PubMed/MEDLINE, Scopus, Web of Science, and IEEE Xplore for peer-reviewed literature, supplemented by manual screening of preprint servers and the reference lists of retrieved articles to capture recent, rapidly evolving computational work not yet indexed in major databases. Search terms combined controlled vocabulary and free-text keywords across three conceptual clusters, connected with Boolean operators: (1) gene/disease terms — "BRCA1" OR "BRCA2" OR "hereditary breast and ovarian cancer" OR "homologous recombination deficiency"; (2) computational-model terms — "protein language model" OR "deep learning" OR "variant pathogenicity prediction" OR "AlphaMissense" OR "ESM" OR "DNABERT" OR "multi-omics"; and (3) clinical-translation terms — "precision oncology" OR "PARP inhibitor" OR "variant of uncertain significance" OR "clinical decision support". No lower date limit was imposed, though the search prioritized literature published within the preceding five years to reflect the fast-moving nature of protein language modeling.

3.2 Eligibility Criteria

Studies were eligible for inclusion if they (a) were published in a peer-reviewed journal or as a full conference proceeding; (b) addressed protein or nucleotide language models, deep learning architectures, or multi-omics integration frameworks applied to cancer genomics, variant interpretation, or precision oncology; and (c) reported extractable methodological detail sufficient to characterize the model architecture, input modalities, and at least one quantitative performance metric. Studies were excluded if they were editorials, conference abstracts without full methodological reporting, non-English-language publications without an available translation, or focused exclusively on cancer types or molecular contexts unrelated to BRCA biology, homologous recombination, or the broader multi-omics oncology landscape addressed by this review.

3.3 Study Selection and Screening Process

Following automated deduplication of search results, two conceptual screening stages were applied: an initial title/abstract screen against the eligibility criteria above, followed by full-text review of the remaining records. Articles progressing through both stages contributed either to the qualitative narrative synthesis or, where sufficient quantitative detail was available, to the comparative benchmark tables (Tables 1–4). This selection pathway is summarized in Figure 3, which documents the number of records identified, screened, and retained at each stage, consistent with standard systematic-review reporting conventions.

3.4 Data Extraction

For each included study, we extracted, where reported: (i) model or framework name and underlying architectural class (e.g., transformer, graph neural network, variational autoencoder); (ii) input data modalities (sequence, expression, methylation, imaging, or combinations thereof); (iii) the clinical or biological prediction task addressed; (iv) validation cohort characteristics and sample size, where disclosed; (v) reported performance metrics (accuracy, AUC, F1-score, precision/recall, or concordance statistics, as reported by the original authors); and (vi) stated limitations or translational barriers noted by the original authors. Extracted data were tabulated by two independent reviewers' logic checks against the source text to minimize transcription error, and discrepancies were resolved by re-consulting the original publication.

3.5 Data Synthesis Approach

Given the methodological heterogeneity across included studies — spanning variant-level pathogenicity classifiers, tumor-type classifiers, and multimodal risk-prediction frameworks — a quantitative meta-analysis (e.g., pooled effect size) was not appropriate. We therefore synthesized findings narratively, organized thematically around four domains reflecting recurring patterns in the literature: representation learning and multimodal data fusion; clinical risk stratification and treatment-response prediction; synthetic-lethality and drug-target discovery; and translational/implementation barriers. This thematic structure mirrors, and is reported in, the Results section below and is intended to allow readers to trace each reported metric back to its source study via the accompanying tables.

3.6 Quality Considerations and Limitations of the Method

As a narrative synthesis rather than a formal systematic review with dual-independent screening and formal risk-of

Table 1. Multimodal data-integration frameworks and deep learning architectures applied to cancer genomics. This table summarizes transformer-, graph-, and autoencoder-based deep learning architectures designed for high-dimensional, sparse, and heterogeneous multi-omics cancer datasets, including each framework's input modalities, primary clinical use-case, key computational advantages and limitations, and reported performance benchmarks with source citations. Entries are ordered to progress from sequence-only models toward increasingly multimodal fusion architectures.

Model / Framework

Architectural Class

Key Input Modalities

Primary Clinical Use-Case

Main Computational Advantages

Prominent Limitations & Bottlenecks

Performance Benchmarks

Key References (APA Style)

SetQuence

Transformer-Based Deep Set Encoder

Somatic variants from whole-genome or whole-exome sequencing.

Tumor type and subtype classification.

Processes non-fixed, permutation-invariant sets; avoids ordering bias; models non-coding variants.

High computational overhead with large sequence sizes; requires extensive pretraining.

Achieved 0.502 macro precision and 0.567 macro accuracy on Cosmic coding variants.

Jurenaite et al., 2024

SetOmic

Multi-Omics Graph Transformer

Somatic variants, gene expression count matrices.

Multimodal cancer classification and biomarker discovery.

Dynamically embeds locus-expression tokens; captures long-range genomic-transcriptomic dependencies.

Fragile to extensive missing data; complex hyperparameter tuning required.

Outperformed standard VAE and Random Forest baselines in TCGA breast cohorts.

Jurenaite et al., 2024

OmiEmbed

Multi-Task Hybrid Variational Autoencoder (VAE)

Gene expression, DNA methylation, microRNA (miRNA).

Multi-task prediction: tumor subtype, disease stage, survival risk.

Learns unified, lower-dimensional latent embeddings; handles multi-task learning simultaneously.

Sensitive to input ordering; utilizes fixed-dimensionality gene-level vectors.

Achieved 97.7% classification accuracy and 0.782 C-index for survival prediction.

Zhang et al., 2021

MOGONET

Graph Convolutional Network (GCN)

mRNA expression, DNA methylation, microRNA profiles.

Multimodal patient classification and disease subtyping.

Uses View Correlation Discovery Network (VCDN) to capture cross-omics label correlations.

Dependent on the construction of a static patient similarity network.

Outperformed nine supervised multi-omics models on ROSMAP and TCGA LGG/BRCA.

Wang et al., 2021

TMO-Net

Self-Modal & Cross-Modal VAE

Somatic mutations, copy number variations (CNVs), mRNA expression, DNA methylation.

Pan-cancer subtype classification, metastasis risk prediction.

Integrates embeddings using a Product-of-Experts (PoE) module to handle missing modalities.

Computationally expensive to train parallel per-modality autoencoders.

Reached a 0.921 F1-score on METABRIC breast cancer molecular subtyping.

Ouhmouk et al., 2025

MultiGATAE

Graph Attention Autoencoder

mRNA expression, DNA methylation, microRNA profiles.

Unsupervised cancer clustering and subtyping.

Minimizes adjacency reconstruction loss; uses Similarity Network Fusion (SNF) sample graphs.

High sensitivity to initial SNF distance kernel configurations.

Outperformed eight clustering baselines in clinical log-rank p-value separation across five cancers.

Zhang et al., 2022

OncoChat

Large Language Model (LLM) Agent

Somatic mutations, copy number alterations (CNAs), structural variants (SVs).

Differential diagnosis and CUP (Cancer of Unknown Primary) classification.

Transforms genetic lists into single-turn conversational prompts; incorporates SVs.

Dependent on proprietary LLM APIs; lacks transcriptomic or spatial context.

Achieved 0.810 micro-averaged PRAUC and classified 22/26 confirmed CUP cases.

Corso et al., 2026

OncoNPC

Machine Learning Classifier

Somatic mutations (SNVs, CNAs), patient demographics (age, gender).

CUP tumor origin tracking and personalized therapy matching.

Trained on multi-institutional targeted cancer panels; highly generalizable.

Scope restricted to only 22 primary cancer types.

Outperformed traditional empirical treatment regimes, yielding significant survival benefits.

Moon et al., 2023

GREMI

Graph Attention Network (GAT)

mRNA expression, DNA methylation, microRNA profiles.

Disease prediction and co-functional module identification.

Employs a Monte Carlo Tree Search (MCTS) to isolate disease-related subgraphs.

Fixed-dimensional GAT layers struggle to scale to genome-wide features.

Achieved F1-weighted scores of 0.877 on BRCA and successfully validated LGG biomarkers.

Liang et al., 2024

mosGraphGPT

Graph-Transformer Foundation Model

DNA methylation, mutations, clinical profiles, protein arrays.

Patient subtyping, signaling flow reconstruction, target discovery.

Projects biological prior graphs onto generative latent spaces to reconstruct signaling cascades.

Vulnerable to text-mining errors during graph assembly; high computational cost.

Achieved 75.1% accuracy in complex disease classification versus standard GNNs.

Zhang et al., 2024

Table 2. Biomarker discovery, prognostic, and neoadjuvant treatment-response prediction frameworks. This table catalogs computational tools developed to identify clinically actionable prognostic signatures and forecast therapeutic outcomes, spanning digital pathology, radiomics, transcriptomic, and knowledge-graph-based approaches, together with their validation cohorts, translational outputs, and quantitative performance metrics as reported in the source publications.

Tool / Platform

Clinical Objective

Modal Inputs Fused

Modeling Methodology

Validation Cohort

Key Translational Output

Performance Metric

Key References (APA Style)

Deep-ODX

Recurrence risk forecasting

H&E-stained whole-slide images (WSIs) of tumor regions.

Multiple Instance Learning (MIL) with ResNet.

151 HR+/HER2- breast cancer whole slides.

Direct prediction of Oncotype DX (ODX) recurrence scores.

0.862 ± 0.034 AUC; higher error rates observed near score boundary 25.

Su et al., 2024

Sammut et al. Multimodal Model

Prediction of pathological complete response (pCR) to neoadjuvant therapy.

Digital pathology, bulk genomics, clinical phenotypic variables.

Multimodal deep learning with feature-level concatenation.

Independent prospective-retrospective cohorts.

Automated patient stratification for chemotherapy de-escalation.

Achieved an AUC of 0.87 for pCR prediction, outperforming single-modality models.

Sammut et al., 2022

LUCID

Non-invasive mutational profiling

CT radiology scans, clinical symptoms, demographics, lab findings.

Transformer architecture with cross-modal attention mechanisms.

Multicenter lung cancer cohorts.

In silico prediction of somatic EGFR mutations and survival outcomes.

Outperformed baseline radiomic and clinical classifiers.

Lin et al., 2026

TOAD

Tumor origin diagnosis

Routine H&E histology slides.

Weakly supervised deep learning on whole-slide images.

TCGA known primaries and CUP biopsy series.

Automated classification of primary origin for metastasized tumors.

83% top-1 accuracy on known primaries; 61% concordance in CUP validation.

Lu et al., 2021

GC-CDSS

Treatment regimen personalization

Unstructured electronic health records, NCCN guidelines.

Knowledge Graph construction and semantic rule-reasoning.

Gastric cancer patient cohorts.

Clinical recommendation output matching Multidisciplinary Tumor Boards.

High consistency rate with treatment decisions of human oncologists.

Lin et al., 2026

Transpara

Computer-Aided Malignancy Detection

Screening mammography images.

Convolutional Neural Network (CNN) feature extraction.

Retrospective multicenter screening series (14 radiologists, 240 cases).

Continuous malignancy likelihood score (1 to 10) to guide triage.

Standalone AUC of 0.88, matching or exceeding average radiologist baseline.

Barua et al., 2026

17-Gene Signature Ensemble

Response forecasting to neoadjuvant chemotherapy (NAC)

Bulk RNA-seq transcriptomic expression profiles.

Ensemble algorithm fusing RF, Gradient Boosting, SVM, NN, and KNN.

Independent breast cancer validation cohorts.

Prioritizes MKI67 and BRCA2 expression to categorize pathological response.

Achieved 0.78 AUC; universally validated across solid tumors.

Jurenaite et al., 2024

TCRNodseek Plus

Pulmonary nodule triage

T-cell receptor (TCR) sequencing, CT scans, demographics.

Feature fusion and ridge regression classifier.

Large-scale TCR-seq patient cohort.

Diagnosis of indeterminate pulmonary nodules.

AUC of 0.84, significantly outperforming the conventional Mayo clinical model.

Lin et al., 2026

MIRSPSO

Survival probability estimation

High-dimensional clinical and molecular features.

Multi-Objective Intelligent Reflectance Particle Swarm Optimisation.

700 patients with diagnosed brain metastases.

Extraction of prognostic signatures to forecast metastatic survival.

97% ± 6% AUC, outperforming traditional lasso and Cox regression.

Barua et al., 2026

Grey Wolf Optimiser-SVM

Benign vs. malignant classification

Dual-modal breast ultrasound and mammography.

Region extraction followed by Support Vector Machine optimized via GWO.

Retrospective breast cancer patient datasets.

Automated, non-invasive lesion boundary extraction and subtyping.

99.25% accuracy on 10-fold CV; 98.46% on 5-fold CV.

Barua et al., 2026

-bias scoring, this methodology is subject to selection and reporting biases inherent to the underlying primary literature; performance metrics reported here reflect authors' own validation procedures and were not independently recomputed or benchmarked on a common dataset. This limitation is addressed further in the Discussion.

4. Benchmarking Protein Language Models and Multi-Omics Paradigms in Precision Oncology

This section synthesizes the empirical outcomes, classification benchmarks, and translational milestones identified through the search and screening process described above (Figure 3), organized across four clinical-computational touchpoints reported in Tables 1 through 4.

4.1 Deep Representation Learning and Multimodal Data Fusion 

Statistical frameworks in oncology have historically struggled to integrate biological layers because of dimensionality and sparsity (Ouhmouk et al., 2025). Deep neural networks that treat biological sequences as text corpora have meaningfully mitigated this limitation. A primary benchmark here is SetQuence, an attention-based set transformer designed to process somatic variant-associated sequences as non-fixed, permutation-invariant sets (Jurenaite et al., 2024). Evaluated on the COSMIC database across 32 tumor classes, SetQuence achieved a macro-averaged precision of 0.502, recall of 0.402, and classification accuracy of 0.567, meaningfully outperforming random-forest and standard neural-network baselines (Jurenaite et al., 2024; Table 1). These figures confirm that sequence-level features drawn from DNA language models capture physical, chemical, and evolutionary constraints that simpler one-hot or gene-ID encodings tend to lose (Jurenaite et al., 2024).

Extending this architecture, SetOmic integrated sequence embeddings with bulk RNA-sequencing expression matrices, achieving a macro-averaged precision of 0.945 and accuracy of 0.950 in pan-cancer classification — outperforming OmiEmbed's 0.942 accuracy, a variational autoencoder restricted to fixed-dimensionality expression vectors (Zhang et al., 2021; Jurenaite et al., 2024; Table 1). Confusion-matrix analysis further showed that SetOmic sharply reduced misclassification between closely related tumor subtypes, such as breast carcinoma and uterine carcinosarcoma, by modeling long-range genomic-transcriptomic dependencies (Jurenaite et al., 2024). Graph-based approaches such as MOGONET, which fuses mRNA, DNA methylation, and microRNA profiles via a View Correlation Discovery Network, consistently outperformed single-modality models across independent cohorts, reinforcing the broader pattern that vertical data fusion uncovers synergistic signal invisible to single-assay pipelines (Wang & Ballester, 2021; Ouhmouk et al., 2025; Table 1).

4.2 Clinical Diagnostics, Risk Stratification, and Neoadjuvant Therapy Response 

Across the included literature, deep learning models have demonstrated clinically meaningful precision in forecasting treatment response and disease recurrence from multimodal inputs. In predicting pathological complete response to neoadjuvant chemotherapy, a 17-gene stacked ensemble classifier — combining random forest, gradient boosting, support vector machine, k-nearest neighbors, and neural network base learners through a stacked elastic-net meta-learner — extracted a signature of 11 upregulated genes (including BRCA2, MKI67, and LYZ) and 6 downregulated genes distinguishing responders from non-responders (Lash & Valero, 2026). Externally validated with isotonic-regression calibration, the ensemble achieved a ROC-AUC of 0.78, balanced accuracy of 0.71, and sensitivity of 0.86, with DeLong's test confirming statistically significant superiority (p < .05) over simple voting ensembles and individual base classifiers (Lash & Valero, 2026; Table 2).

Multimodal frameworks that fuse digital histology with genomic and clinical features have performed comparably well: an integrated model combining H&E histology, bulk genomics, and clinical variables achieved an AUC of 0.87 for pathological complete response prediction, translating into a 10–20% accuracy gain over single-modality approaches (Sammut et al., 2022; Sabit et al., 2026; Table 2). AnchorMIL, a multiple-instance-learning framework applied to whole-slide images, achieved AUC values of 0.885 on TCGA-BRCA and 0.858 on an independent cohort for Oncotype DX recurrence-score estimation, outperforming conventional attention-based MIL baselines while supporting both continuous score prediction and binary risk stratification (Koyun et al., 2026; Su et al., 2024; Table 2).

4.3 In Silico Drug Discovery and Synthetic Lethality Target Screening

A recurring theme among the reviewed synthetic-lethality

Figure 1. Mechanistic basis of PARP-inhibitor synthetic lethality in BRCA1/2-deficient tumors. Schematic comparison of homologous recombination (HR)-proficient and HR-deficient tumor cells. In HR-proficient cells (left), wild-type BRCA1/BRCA2 supports accurate double-strand-break repair and genomic stability. In HR-deficient cells (right), a pathogenic BRCA1/BRCA2 variant abolishes homologous recombination, forcing reliance on error-prone repair pathways; PARP-inhibitor exposure then traps PARP1/2 on DNA, converting single-strand lesions into lethal double-strand breaks and driving selective tumor-cell apoptosis. A protein language model-derived pathogenicity score (bottom) links variant calling to this clinical decision pathway.

Figure 2. Sequence-to-pathogenicity inference pipeline used by protein language models. Conceptual workflow by which a raw BRCA1/BRCA2 variant call is encoded as an amino-acid sequence, embedded by a pretrained protein language model (e.g., AlphaMissense, ESM3, DNABERT-S), and mapped onto a latent evolutionary-constraint representation that yields a discrete pathogenicity classification. Outputs from this pipeline feed downstream clinical applications, including homologous recombination deficiency scoring and PARP-inhibitor eligibility assessment.

Figure 3. Study identification and selection workflow for this narrative review. Flow diagram summarizing the reproducible literature search and screening process described in Section 3 (Methods), from initial database identification through duplicate removal, title/abstract screening, full-text eligibility assessment, and final inclusion in qualitative synthesis and the quantitative comparison tables (Tables 1–4). Reasons for exclusion at each stage are indicated alongside the corresponding decision node.

Figure 4. Vertical multi-omics data-fusion architecture underlying translational precision-oncology models. Schematic representation of how genomic, transcriptomic, epigenomic, and digital-pathology data layers are integrated by a deep-learning fusion module (e.g., graph neural network, variational autoencoder, or set-transformer architecture) to generate an integrated patient-level risk and pathogenicity score, which in turn informs clinical decision support around homologous recombination deficiency status, PARP-inhibitor eligibility, and multidisciplinary tumor-board discussion.

frameworks is their reliance on graph-structured biological knowledge to move beyond the canonical BRCA-PARP axis. DGIB4SL, built on a diverse graph information bottleneck, extracted 13 distinct motif-based adjacency matrices from resources such as SynLethDB and outperformed baseline classifiers on both recall and precision, while additionally producing multiple, biologically interpretable explanations for each predicted interaction (Li et al., 2026; Table 3). MAGICAL, a multi-class random-forest framework built on protein-protein interaction network topology, achieved 80% accuracy on its discovery cohort and a validation AUC of 0.83 on the independent DepMap CRISPR dependency dataset, successfully discriminating synthetic-lethal from synthetic-viable gene pairs (Li et al., 2026; Table 3).

Complementary architectures address specific weaknesses in this space: NSF4SL reformulates synthetic-lethality prediction as a partner-gene ranking problem to avoid reliance on unverified negative training examples, demonstrating stronger generalization to previously unobserved genes, while KR4SL applies recurrent path-encoding over biomedical knowledge graphs to surface clinically interpretable explanatory subgraphs, including non-canonical double-strand-break repair pathways relevant to BRCA2-deficient, PARP-inhibitor-resistant tumors (Li et al., 2026; Table 3). Literature-mining platforms such as PandaOmics extend this pipeline further, combining pathway scoring with large-language-model-based literature synthesis to prioritize candidate compounds, including a selective MYT1 inhibitor with notable preclinical activity in gynecological and breast cancer models (Barua et al., 2026; Table 3).

4.4 Liquid Biopsy, Clinical Translation, and Algorithmic Roadblocks 

Non-invasive monitoring technologies represent one of the more clinically proximate applications reviewed here. Machine learning models applied to cell-free DNA methylation patterns, exemplified by the Grail Galleri platform, have demonstrated early-detection capability across more than 50 distinct cancer types (Shanmugam & Ravikumar, 2026; Table 4). Deep learning models analyzing cfDNA fragmentomics have achieved over 90% accuracy distinguishing malignant from benign lesions, supporting minimally invasive minimal residual disease monitoring (Lin et al., 2026; Rescigno & Greystoke, 2026; Table 4). At the single-cell level, foundation models such as scGPT and scBERT have annotated heterogeneous tumor-infiltrating immune populations with F1-scores around 0.923, while additionally characterizing drug-resistant subclonal trajectories in near-real time (Ouhmouk et al., 2025; Table 4).

These successes, however, coexist with the translational obstacles catalogued in Table 4: heavy reliance on TCGA-derived training data and its associated technical artifacts, persistent interpretability gaps in complex architectures, and reduced generalizability when models trained on demographically homogeneous cohorts are deployed in more diverse clinical populations (Dehkharghanian et al., 2023; Ouhmouk et al., 2025; Lin et al., 2026; Wolde & Belay, 2026; Table 4). Addressing the latter, in particular, is increasingly framed in the literature as requiring algorithmic auditing, prospective multi-institutional "silent trials," and privacy-preserving federated learning architectures capable of training across decentralized, internationally diverse datasets without exposing individual patient records (Ouhmouk et al., 2025; Table 4).

5. From Computational Benchmark to Clinical Confidence — Closing the BRCA Variant Interpretation Gap

5.1 Interpreting the Evidence: What the Numbers Actually Suggest

Pulling back from the individual benchmarks summarized above, a fairly coherent narrative emerges. Protein language models and their multi-omics extensions do not merely match earlier machine learning approaches to variant interpretation — in study after study, they outperform them, often by a wide margin, and they do so specifically on the kinds of high-dimensional, sparse, sequence-heavy data that older architectures were never well suited to handle (Jurenaite et al., 2024; Ouhmouk et al., 2025; Table 1). That the largest accuracy gains cluster around models capable of ingesting raw sequence context — rather than hand-engineered gene-level features — is, we think, not a coincidence; it is fairly direct evidence that evolutionary and structural constraint, learned implicitly from sequence, carries real predictive signal about pathogenicity that classical feature engineering was leaving on the table (Corso et al., 2026; Jurenaite et al., 2024).

At the same time, it would be a mistake to read these benchmark numbers as settled clinical fact. Nearly every metric reported in Tables 1 through 4 was generated under retrospective, single-institution, or heavily curated

Table 3. AI-driven drug discovery, synthetic lethality, and target prioritization frameworks. This table outlines deep learning and knowledge-graph platforms designed to predict drug-target interactions, identify novel synthetic lethal gene pairs beyond canonical BRCA-PARP biology, and model cellular treatment responses, including each framework's core methodology, experimental validation status, and principal advantages as documented in the cited literature.

Framework / Tool

target / Biological Interaction

Core ML/DL Technique

Inputs & Integrated Repositories

Key Mechanism/Synergy Identified

Experimental Validation Status

Core Advantages

Key References (APA Style)

DGIB4SL

Synthetic Lethality (SL) partner genes

Diverse Graph Information Bottleneck; Determinant Point Process (DPP).

Knowledge Graph databases (e.g., SynLethDB).

Extracts 13 distinct motif-based adjacency matrices.

Retrospective evaluation; lacks robust in vivo validation.

Generates multiple, biologically plausible explanations per predicted SL pair.

Li et al., 2026

MAGICAL

SL vs. Synthetic Viability (SV) genetic interactions

Topological property mapping; multi-class random forest.

CGIdb, BioGRID, SynLethDB, DepMap.

Models shortest-path network properties to detect cell viability changes.

Evaluated on independent DepMap CRISPR datasets.

Achieved 80% accuracy on training sets and 0.83 AUC on validation.

Li et al., 2026

KR4SL

Explainable SL target screening

Knowledge Graph reasoning; RNN path encoding.

Biomedical Knowledge Graphs (e.g., BioKG).

Identifies subgraph semantic paths connecting target gene nodes.

In-silico validation; verified known BRCA1/2 DDR pathways.

Attentive aggregator provides clinically intuitive path explanation.

Li et al., 2026

NSF4SL

Target ranking for drug screens

Negative-Sample-Free contrastive graph learning.

Positive SL datasets from SynLethDB.

Reformulates SL prediction as a partner gene ranking problem.

Tested on non-overlapping gene sets in DepMap.

Eliminates the need for unverified and noisy negative samples.

Li et al., 2026

PiLSL

Pairwise interaction learning

Graph Neural Network with attentive embedding propagation.

Enclosing subgraphs extracted from SynLethKG.

Captures local topology of surrounding pathways instead of independent genes.

Retrospective benchmarking against standard graph convolutional networks.

High transductive and inductive performance; weighted explanation paths.

Li et al., 2026

PandaOmics

Target identification & biomarker discovery

LLM text mining coupled with molecular pathway scoring.

PubMed full-text literature, public omics repositories.

Identified a potent MYT1 inhibitor for breast and gynecological cancers.

Confirmed candidate targets in preclinical in vitro and in vivo models.

Integrates literature-derived hypotheses with continuous omics profiling.

Barua et al., 2026

DrugCell

Drug response and synergy prediction

Visible Neural Network (VNN).

3,008-gene mutation vectors, 2,048-bit drug Morgan fingerprints.

Mimics eukaryotic cell hierarchy based on Gene Ontology (GO) terms.

Experimentally validated across diverse cancer cell lines.

High biological interpretability; models subsystem importance via RLIPP.

Lin et al., 2026

BANDIT

Drug-target interaction (DTI) prediction

Bayesian integrative machine learning.

Gene expression profiles, chemical structures, clinical adverse effects.

Discovered 14 novel microtubule inhibitors active in resistant cells.

experiment-validated; identified ONC201 drug target.

Fuses diverse data streams; achieves 90% target prediction accuracy.

Barua et al., 2026

XGBoost-OMC

High-dimensional drug response regression

Extreme Gradient Boosting with Optimal Model Complexity variant.

Genomics of Drug Sensitivity in Cancer (GDSC).

Identified DNA methylation features predicting SCLC response to Thapsigargin.

Validated on independent external cell line test sets.

Generates minimal predictive subsets of features (e.g., OMC of 7) to avoid overfitting.

Jurenaite et al., 2024

STATE

Dynamic in silico perturbation modeling

Mechanism-aware deep generative transformer.

Single-cell RNA-seq, transcriptomic dynamics.

Simulates gene knockouts, mutations, and pharmacological insults.

In-silico modeling; prospective validation underway.

Unified framework for high-throughput virtual drug screening.

Lin et al., 2026

 

 

Table 4. Liquid-biopsy applications and translational roadblocks in clinical AI deployment. This table summarizes non-invasive liquid-biopsy and single-cell monitoring applications alongside the major technical and ethical roadblocks — including TCGA data overreliance, algorithmic interpretability gaps, and population generalizability limitations — that currently constrain the clinical translation of AI-driven precision-oncology tools, together with proposed mitigation strategies drawn from the reviewed literature.

 

Emerging Technology

Translation Phase

Primary Analyte / Modality

Computational Backbone

Target Subpopulation

Primary Implementation Bottleneck

Key Mitigation & Governance Strategy

Key References (APA Style)

Grail Galleri

Commercially Deployed

Cell-free DNA (cfDNA) in peripheral blood.

Machine Learning classifier on genomic methylation patterns.

Early screening across 50+ cancer types.

High false-positive rates requiring specificity thresholds above 99%.

Staged retrospective and prospective multi-ethnic clinical screening trials.

Shanmugam & Ravikumar, 2026

Spatial Transcriptomics

Early Clinical Translation

Visium, MERFISH, Stereo-seq, and Xenium tissue spots.

Hypergraph GCNs and Vision Transformers.

Spatially restricted cancer subpopulations at tumor-stroma interfaces.

High cost, low throughput, and lack of compatibility with archived FFPE tissues.

Standardized clinical reporting pipelines; AI-driven spatial image reconstruction.

Shanmugam & Ravikumar, 2026

cfDNA Fragmentomics

Clinical Development

ctDNA density distributions of fragment lengths.

Neural networks (e.g., ctGAN / GBCseeker).

Minimal Residual Disease (MRD) monitoring and relapse surveillance.

Systematic technical batch effects across clinical sequencing centers.

ComBat-based data normalization and multi-center validation benchmarks.

Lin et al., 2026

Single-Cell Language Models

Early Research Translation

Single-cell RNA-seq (scRNA-seq) and scATAC-seq profiles.

Generative Pretrained Transformers (e.g., scGPT / scBERT).

Highly heterogeneous cell lineages; immunotherapy-resistant subclones.

Limited availability of matched, high-depth spatial-genomic datasets.

Adversarial latent space regularization and transfer learning from bulk tissues.

Ouhmouk et al., 2025

FFPEsig

Proof of Concept

Archived formalin-fixed paraffin-embedded clinical tissue DNA.

Deep neural network sequence classifiers.

Longitudinally tracked historical cohorts; rare cancer subtypes.

Fixing artifacts and formalin-induced sequencing noise.

Automated mathematical filtering of mutational artifacts.

Rescigno & Greystoke, 2026

TrialGPT

Clinical Pilot Testing

Unstructured Electronic Health Records (EHRs) and clinical trial protocols.

Zero-shot LLM matching (Retrieval, Matching, Ranking modules).

Oncology patients undergoing trial pre-screening.

LLM hallucinations of eligibility criteria; data privacy concerns.

Human-in-the-loop clinical validation; local private servers.

Lin et al., 2026

CURATE.AI

Feasibility Trial

Dynamic patient-specific dose-response trajectories.

Adaptive modeling on single-patient calibration datasets.

Patients receiving cytotoxic chemotherapy combinations.

Fear of false-negative toxicity predictions leading to severe clinical events.

Phase I/II validation trials; clinicians retain ultimate dose override capability.

Lin et al., 2026

AiCure

Commercial Clinical Trials

Real-time patient tablet/smartphone video inputs.

Facial recognition and computer vision algorithms.

Patients undergoing outpatient oral chemotherapy regimens.

Technical camera failure or poor patient compliance in elderly demographics.

telehealth dashboard integration; real-time care-team alerts for missed doses.

Barua et al., 2026

Adversarial pathology de-biasing

Pre-clinical Development

Whole-Slide Pathology Histology Images.

Adversarial neural networks; self-supervised DINOv2.

Multi-center cohorts with scanner variations.

Models predicting the image acquisition site (up to 86% accuracy).

Training models via gradient reversal layers to force domain-invariant feature extraction.

Ouhmouk et al., 2025

Federated learning frameworks

Multi-institutional Pilots

Decentralized clinical records and radiology imaging archives.

Federated Transformers and GNNs (e.g., NVIDIA FLARE).

Rare cancers; populations with strict multi-national data silos.

Communication latency and local infrastructure heterogeneity.

Adaptive Optimization for Vertical Integration; differential privacy constraints.

Ouhmouk et al., 2025

conditions (Ouhmouk et al., 2025; Table 1; Table 2). AUC values above 0.90 are genuinely impressive, but they describe performance within a data distribution the model has already seen a great deal of, not necessarily performance on the next VUS that lands on a clinician's desk from an underrepresented population or an unusual tumor subtype (Dehkharghanian et al., 2023; Ouhmouk et al., 2025).

5.2 BRCA Variant Interpretation Specifically: Where pLMs Genuinely Change the Calculus

For the specific clinical problem motivating this review — the roughly 40% of BRCA missense variants still sitting in VUS limbo — protein language models offer something traditional approaches largely cannot: near-instantaneous, proteome-scale scoring that does not require a bespoke functional assay for every novel variant encountered (Rescigno & Greystoke, 2026; Corso et al., 2026). AlphaMissense-style models, in principle, can pre-screen a newly detected BRCT-domain missense variant the moment it is called, flagging it for prioritized functional workup or, in higher-confidence cases, informing an interim clinical conversation well before a multigene functional assay could ever be completed (Corso et al., 2026; Ismail et al., 2024). Combined with genomic-scar-based HRD detection tools such as HRProfiler, this creates a plausible, if not yet fully validated, pipeline running from raw sequence to PARP-inhibitor eligibility (Shah et al., 2025; Figure 1; Figure 2).

That said, the honest caveat here is that BRCA-specific prospective validation of these general-purpose pLMs remains comparatively thin in the literature we reviewed, relative to their pan-cancer classification benchmarks (Table 1; Table 3). Much of the strongest quantitative evidence assembled in this review — SetOmic's 0.950 accuracy, MOGONET's cross-cohort superiority, OncoChat's tumor-of-origin performance — comes from tasks adjacent to, rather than identical with, single-variant BRCA pathogenicity classification (Jurenaite et al., 2024; Wang & Ballester, 2021; Liu et al., 2025). We would argue this is an important distinction for clinicians reading this literature to keep in mind, rather than a reason for skepticism outright.

5.3 The Interpretability Problem Is Not Just an Engineering Inconvenience

We think it is worth stating plainly that the "black box" critique raised repeatedly in the literature (Section 2.4.2) is not a peripheral concern that can be resolved later, once performance is good enough. For a VUS reclassification decision that may lead to prophylactic surgery on one hand, or withheld PARP-inhibitor therapy on the other, a clinician reasonably wants to know why a model called a variant likely pathogenic, not merely that it did so with high confidence (Ouhmouk et al., 2025; Sabit et al., 2026). The SHAP- and LIME-based attribution work reviewed here is encouraging — attributions do appear to track known cancer biology in several studies (Jurenaite et al., 2024; Ouhmouk et al., 2025) — but post-hoc explanation is not the same thing as mechanistic transparency, and regulatory bodies evaluating these tools as software-as-a-medical-device will likely draw that distinction carefully.

5.4 Equity, Data Provenance, and the Limits of Benchmark Generalization

Perhaps the most consequential limitation surfaced in this review is not technical at all, but demographic. With more than four-fifths of TCGA cases representing individuals of European ancestry (Ouhmouk et al., 2025), and with pLM training corpora inheriting broadly similar biases from public sequence databases, there is a real and, we think, underappreciated risk that a model performing beautifully on benchmark cohorts will underperform — silently, without any obvious warning sign — on patients from underrepresented populations (Ouhmouk et al., 2025; Lin et al., 2026; Wolde & Belay, 2026). Given that BRCA variant spectra themselves differ meaningfully by ancestry, this is not a hypothetical concern; it bears directly on which patients ultimately gain, or fail to gain, access to PARP-inhibitor therapy on the basis of an algorithmic call.

5.5 Limitations of This Review

This review has its own limitations, worth stating candidly. As a narrative rather than fully systematic synthesis, it did not apply formal dual-independent risk-of-bias scoring to every included study, and the performance metrics reported in Tables 1 through 4 reflect each original study's own validation methodology rather than a harmonized, independently re-computed benchmark (Section 3.6). Several of the frameworks discussed — particularly the synthetic-lethality and multi-omics fusion tools summarized in Table 3 — remain at a discovery or retrospective-validation stage rather than prospective clinical deployment, and readers should weigh the translational claims made here accordingly.

5.6 A Practical Path Forward

Taken together, we would suggest the field's next steps are reasonably clear, even if not easy: prospective, multi-institutional validation of pLM-based BRCA pathogenicity scores against orthogonal functional assays; deliberate, funded recruitment of ancestrally diverse cohorts rather than passive reliance on existing repositories; wider adoption of federated learning architectures that allow model training across institutions without centralizing sensitive patient data (Ouhmouk et al., 2025); and continued investment in interpretability tooling that clinicians, not just machine learning engineers, find genuinely usable at the point of care (Sabit et al., 2026; Lin et al., 2026). None of this is a small undertaking. But given how directly variant interpretation already determines who receives a life-extending PARP inhibitor and who does not, it is difficult to argue the effort is optional.

6. Conclusion

Protein language models have progressed rapidly from computational novelty to genuine contenders for resolving a major precision-oncology bottleneck: the roughly 40% of BRCA1/BRCA2 missense variants still classified as uncertain. Architectures such as AlphaMissense, ESM3, DNABERT-S, and multi-omics extensions like SetOmic and MOGONET consistently outperform earlier machine learning methods, often exceeding 90% accuracy on tumor classification and risk-stratification tasks. Combined with genomic-scar-based HRD detection and PARP-inhibitor eligibility screening, these tools suggest a plausible pipeline from raw sequence to clinical guidance. That promise remains only partly realized. Evidence relies heavily on retrospective, demographically narrow TCGA cohorts; interpretability has not yet reached clinical or regulatory standards for high-stakes decisions like prophylactic surgery; and BRCA-specific prospective validation remains limited compared to pan-cancer benchmarks. Realizing full clinical potential will require multi-institutional prospective validation, correction of ancestral data bias, and clinician-oriented interpretability tools. Until then, these models function best as triage and prioritization aids—not replacements for careful, orthogonal variant classification.

References


Abramson, J., Adler, J., Dunger, J., Evans, R., Green, T., Pritzel, A., ... & Hassabis, D. (2024). Accurate structure prediction of biomolecular interactions with AlphaFold 3. Nature, 630(8015), 211–218. https://doi.org/10.1038/s41586-024-07487-w

Barua, S., Badhrinarayanan, B., & Balaji, S. (2026). Bridging diagnosis and therapeutics: The role of AI in cancer detection and drug development. Health Sciences Review, 19, 100270. https://doi.org/10.1016/j.hsr.2026.100270

Bilal, A., Imran, A., & Baig, T. I. (2024). Breast cancer diagnosis using support vector machine optimized by improved quantum inspired grey wolf optimization. Scientific Reports, 14, 1–25. https://doi.org/10.1038/s41598-024-61322-w

Corso, G., et al. (2026). Leave no data behind: Exploring a new paradigm in oncology with foundation models and large language models. Cell Reports Medicine, 7, 102966. https://doi.org/10.1016/j.xcrm.2026.102966

Dehkharghanian, T., Bidgoli, A. A., Riasatian, A., Mazaheri, P., Campbell, C. J. V., Pantanowitz, L., ... & Tizhoosh, H. R. (2023). Biased data, biased AI: Deep networks predict the acquisition site of TCGA images. Diagnostic Pathology, 18, 67. https://doi.org/10.1186/s13000-023-01355-3

Eniu, A., Pop, L., Stoian, A., Dronca, E., Matei, R., Ligtenberg, M., Ouchene, H., Onisim, A., Rotaru, O., Eniu, R., et al. (2023). Delivering precision medicine in hereditary breast cancer: NGS-based multi-gene panel testing beyond BRCA1/2. Biomedicines, 11(5), 1386. https://doi.org/10.3390/biomedicines11051386

Gentile, G., et al. (2026). Integration of germline testing into precision oncology frameworks: A systematic review. Critical Reviews in Oncology/Hematology, 223, 105307. https://doi.org/10.1016/j.critrevonc.2026.105307

Ismail, T., Alzneika, S., Riguene, E., Al-Maraghi, S., Alabdulrazzak, A., Al-Khal, N., & Nomikos, M. (2024). BRCA1 and its vulnerable C-terminal BRCT domain: Structure, function, genetic mutations and links to diagnosis and treatment of breast and ovarian cancer. Pharmaceuticals, 17(3), 333. https://doi.org/10.3390/ph17030333

Jurenaite, N., León-Periñán, D., Donath, V., Torge, S., & Jäkel, R. (2024). SetQuence & SetOmic: Deep set transformers for whole genome and exome tumour analysis. BioSystems, 235, 105095. https://doi.org/10.1016/j.biosystems.2023.105095

Kotsifaki, A., Kalouda, G., Karalexis, E., Stathaki, M., Metaxas, G., & Armakolas, A. (2025). Emerging breast cancer subpopulations: Functional heterogeneity beyond the classical subtypes. International Journal of Molecular Sciences, 26, 11599. https://doi.org/10.3390/ijms262311599

Koyun, O. C., et al. (2026). AnchorMIL: Multiple instance learning framework with anchored regression for Oncotype DX recurrence score prediction. Expert Systems with Applications, 303, 130469. https://doi.org/10.1016/j.eswa.2025.130469

Kuenzi, B. M., Park, J., Fong, S. H., Sanchez, K. S., Lee, J., Kreisberg, J. F., ... & Ideker, T. (2020). Predicting drug response and synergy using a deep learning model of human cancer cells. Cancer Cell, 38(5), 672–684. https://doi.org/10.1016/j.ccell.2020.09.014

Lash, S., & Valero, C. (2026). Identification of a 17-gene predictive signature through ensemble machine learning analysis for predicting neoadjuvant chemotherapy response in solid tumors. Current Issues in Molecular Biology, 48(1), 94–113. https://doi.org/10.3390/cimb48010094

Li, J., Li, Y., & Xie, T. (2026). Bridging realms: Artificial intelligence integrates omics, generative models, and traditional medicine for anticancer drug innovation. Journal of Pharmaceutical Analysis, 16(1), 85–96. https://doi.org/10.1016/j.jpha.2026.101630

Liang, H., Luo, H., Sang, Z., Jia, M., Jiang, X., Wang, Z., ... & Zhang, Z. (2024). GREMI: An explainable multi-omics integration framework for enhanced disease prediction and module identification. IEEE Journal of Biomedical and Health Informatics, 28, 1861–1871. https://doi.org/10.1109/JBHI.2024.3439713

Lin, R., Zhao, Z., Liu, Z., Kang, J., Zhang, K., Huang, X., ... & Yu, Y. (2026). Artificial intelligence in clinical oncology: Multimodal integration and translational development. Cancer Letters, 649, 218493. https://doi.org/10.1016/j.canlet.2026.218493

Liu, J., Yang, M., Bi, Y., Zhang, J., Yang, Y., Li, Y., Hong, S., Chen, K., & Li, X. (2025). Large language models enable tumor-type classification and localization of cancers of unknown primary from genomic data. Cell Reports Medicine, 6, 102332. https://doi.org/10.1016/j.xcrm.2025.102332

Lu, M. Y., Chen, B., Williamson, D. F. K., Chen, R. J., Liang, I., Ding, T., ... & Mahmood, F. (2024). A visual-language foundation model for computational pathology. Nature Medicine, 30(3), 863–874. https://doi.org/10.1038/s41591-024-02857-3

Moon, I., LoPiccolo, J., Baca, S. C., Sholl, L. M., Kehl, K. L., Hassett, M. J., ... & Gusev, A. (2023). Machine learning for genetics-based classification and treatment response prediction in cancer of unknown primary. Nature Medicine, 29, 2057–2067. https://doi.org/10.1038/s41591-023-02482-6

Ouhmouk, M., Baichoo, S., & Abik, M. (2025). Challenges in AI-driven multi-omics data analysis for oncology: Addressing dimensionality, sparsity, transparency and ethical considerations. Informatics in Medicine Unlocked, 57, 101679. https://doi.org/10.1016/j.imu.2025.101679

Qiu, Z., Kar, P., & Maulik, U. (2026). A genomic data analysis-based technique for personalized and precision medicine. Array, 30, 100965. https://doi.org/10.1016/j.array.2026.100965

Rescigno, P., & Greystoke, A. (2026). Perspectives on next-generation sequencing and artificial intelligence integration in oncology. Cancer Treatment and Research Communications, 46, 101054. https://doi.org/10.1016/j.ctarc.2026.101054

Roy, K. R., Moon, U. D., & Jamal, M. (2026). miRNA-mRNA Multi-omics Integration in Breast Cancer Staging. Advances in Biomarker Sciences and Technology., 9, 70–82. https://doi.org/10.1016/j.abst.2026.06.003   

Sabit, H., Yadav, A. K., Salimy, S., Sakr, A., Abdel-Ghany, S., Wadan, A. S., Alqosaibi, A. I., Rashwan, R., AlGosaibi, Y. S., Alnamshan, M. M., Almulhim, J., Alaqeel, N. K., & Arneth, B. (2026). Multi-omics foundations in breast cancer and AI-driven advances: A clinical guide. Cancer Letters, 649, 218468. https://doi.org/10.1016/j.canlet.2026.218468

Sammut, S.-J., Crispin-Ortuzar, M., Chin, S.-F., Provenzano, E., Bardwell, H. A., Ma, W., ... & Caldas, C. (2022). Multi-omic machine learning predictor of breast cancer therapy response. Nature, 601, 623–629. https://doi.org/10.1038/s41586-021-04278-5

Schäffer, A. A., Chung, Y., Kammula, A. V., Ruppin, E., & Lee, J. S. (2024). A systematic analysis of the landscape of synthetic lethality-driven precision oncology. Med, 5(1), 73–89. https://doi.org/10.1016/j.medj.2023.12.009

Shah, B., Hussain, M., & Seth, A. (2025). Homologous recombination deficiency in ovarian and breast cancers: Biomarkers, diagnosis, and treatment. Current Issues in Molecular Biology, 47(8), 638. https://doi.org/10.3390/cimb47080638

Shanmugam, Y., & Ravikumar, L. (2026). Integrated multi-omics biomarker discovery workflow: Current advances and emerging clinical translation technologies. Advances in Biomarker Sciences and Technology, 8, 464–478. https://doi.org/10.1016/j.abst.2026.05.007

Su, Y., et al. (2024). Deep-ODX: Deep learning-based Oncotype DX recurrence score prediction from H&E whole-slide images. Expert Systems with Applications, 303, 130469. https://doi.org/10.1016/j.eswa.2025.130469

Wang, J., Zhu, H.-R., Xu, J., Fu, J., Liu, L.-Y., Chen, X.-Y., Chen, Z.-S., Lin, H.-W., & Gu, Z.-C. (2026). Oncology drug resistance prediction tools: Database infrastructure, algorithmic innovation, and clinical translation. Current Molecular Pharmacology, 19, 85–96. https://doi.org/10.2174/1874467226000103

Wang, T., & Ballester, P. J. (2021). MOGONET integrates multi-omics data using graph convolutional networks allowing patient classification and biomarker identification. Nature Communications, 12, 3445. https://doi.org/10.1038/s41467-021-23774-w

Webb, P. M., & Jordan, S. J. (2024). Global epidemiology of epithelial ovarian cancer. Nature Reviews Clinical Oncology, 21(5), 389–400. https://doi.org/10.1038/s41571-024-00881-3

Wolde, T., & Belay, M. (2026). Unraveling the complexity of gynecological cancers: challenges, resistance, and the road to precision medicine - A Narrative Review. Advances in Cancer Biology - Metastasis, 18, 100195. https://doi.org/10.1016/j.adcanc.2026.100195

Yin, J., Li, B., Xiong, H., Gan, J., & Liang, L. (2026). Methodological workflow combining single-cell RNA-seq and bulk transcriptomics to profile MDSCs in breast cancer. Translational Oncology, 63, 102605. https://doi.org/10.1016/j.tranon.2025.102605

Zhang, G., Peng, Z., Yan, C., Wang, J., Luo, J., & Luo, H. (2022). MultiGATAE: A novel cancer subtype identification method based on multi-omics and attention mechanism. Frontiers in Genetics, 13, 855629. https://doi.org/10.3389/fgene.2022.855629

Zhang, X., Xing, Y., Sun, K., & Guo, Y. (2021). OmiEmbed: A unified multi-task deep learning framework for multi-omics data. Cancers, 13(12), 3047. https://doi.org/10.3390/cancers13123047


Article metrics
View details
0
Downloads
0
Citations
54
Views

View Dimensions


View Plumx


View Altmetric



0
Save
0
Citation
54
View
0
Share