1. Introduction
Understanding how cells arrange themselves in space -- and why that arrangement matters -- is, in a sense, one of the oldest questions in medicine dressed up in new molecular clothes. Pathologists have known for well over a century that where a cell sits inside a tissue often says as much about its fate as what genes it carries. What has changed, only recently, is our ability to measure that "where" with the same molecular depth we once reserved for dissociated cell suspensions. Cells, after all, do not function as isolated units; they assemble into elaborate, three-dimensional architectures that coordinate everything from organ development to immune surveillance (Wu et al., 2026). The spatial arrangement of genetically and phenotypically distinct subpopulations within their native microenvironment shapes developmental lineages, physiological homeostasis, and -- when things go wrong -- pathological trajectories (Cen et al., 2026). Bulk transcriptomic and proteomic profiling, for all the precision-medicine progress it has enabled, was never built to see this. By averaging signals across dissociated tissue, it necessarily discards the very heterogeneity that gives a tumor its resilience or a failing heart its patchwork of injury and repair (Arora, 2025; Sahu et al., 2026).
Single-cell RNA sequencing (scRNA-seq) was, for a time, the field's best answer to this problem. It gave researchers an unprecedented view of cellular diversity, lineage dynamics, and rare cell populations that bulk methods had simply averaged away (Chen et al., 2023; Sahu et al., 2026). And yet -- and this is easy to forget amid the enthusiasm -- scRNA-seq still requires tissue dissociation. Enzymatic digestion into a single-cell suspension is, almost by definition, an act of spatial erasure: physical cell-to-cell contacts, local signaling gradients, and the broader tissue neighborhood are gone the moment the sample hits the dissociation buffer (Palaganas et al., 2026; Wu et al., 2022). What is lost in that process is not trivial. Localized biochemical gradients, signaling niches, and multicellular neighborhoods are increasingly understood to be causal drivers of cancer progression, neurodegeneration, and cardiovascular disease, not incidental byproducts of it (Arora, 2025; Gong et al., 2024; Kiessling & Kuppe, 2024). Bridging the gap between molecular profiling and tissue histology has therefore become something of a unifying ambition across the life sciences -- less a niche subfield than a shared destination (Shanmugam & Ravikumar, 2026; Wang & Fan, 2021).
Spatial omics technologies were developed, at their core, to preserve that architecture while still measuring molecular features in situ (Lee et al., 2025). The field's modern trajectory can be traced fairly directly to Ståhl et al. (2016), whose array-based in situ RNA capture method laid the technical groundwork for commercial platforms like 10x Genomics Visium (Kiessling & Kuppe, 2024). What followed was rapid branching. Spatial mono-omics technologies split broadly into next-generation sequencing (NGS)-based and imaging-based approaches (Wu et al., 2026). NGS-based platforms -- Visium, Slide-seq, and Stereo-seq (Spatio-Temporal Enhanced REsolution Omics-sequencing) among them -- rely on spatially barcoded capture probes to pull transcripts directly from tissue sections (Lee et al., 2025). Imaging-based techniques take a different route entirely: MERFISH (multiplexed error-robust fluorescence in situ hybridization) and seqFISH+ use iterative cycles of fluorescent labeling to resolve individual transcript molecules at subcellular resolution (Cen et al., 2026).
Useful as spatial transcriptomics has proven, it only tells part of the story. Transcript abundance, it turns out, correlates rather imperfectly with functional protein output, epigenetic state, or downstream metabolic activity -- a gap that single-layer profiling simply cannot close (Cen et al., 2026; Lee et al., 2025). This limitation is what has pushed the field, somewhat inevitably, toward spatial multi-omics: the simultaneous measurement of several biomolecular layers within the same tissue section, which has the added benefit of avoiding the batch effects that plague separately processed samples (Liu et al., 2024; Sahu et al., 2026). Deterministic barcoding in tissue (DBiT-seq), for instance, uses perpendicular microfluidic channels to co-map whole transcriptomes and proteomes on a single slide (Liu et al., 2020), while spatial ATAC-RNA-seq and spatial CUT&Tag-RNA-seq extend this logic to capture transcriptomic and epigenetic landscapes together (Liu et al., 2024). Mass spectrometry-based approaches such as MALDI-MSI (Matrix-Assisted Laser Desorption/Ionization Mass Spectrometry Imaging) round out the picture further still, visualizing small-molecule metabolites and drugs directly in spatial coordinates (Palaganas et al., 2026).
None of this experimental progress, however, has been matched -- at least not yet -- by an equivalent leap in our ability to make sense of the data it produces. That, arguably, is the field's real bottleneck now (Liu et al., 2024). Multi-modal spatial datasets are high-dimensional, non-linear, riddled with missing features, and burdened by platform-specific noise (Cen et al., 2026; Sahu et al., 2026). Integrating them requires computational strategies capable of horizontal (cross-sample), vertical (same-cell), or diagonal (disjoint cells and features) alignment (Kiessling & Kuppe, 2024). Late-integration methods, which analyze each modality independently before merging results, tend to miss the inter-omic regulatory relationships that matter most (Sahu et al., 2026) -- which is part of why intermediate and early integration frameworks, including graph neural networks, variational autoencoders (totalVI, SWITCH among them), and joint latent-space models, have drawn so much recent attention; they are, at least in principle, better suited to learning shared low-dimensional representations without abandoning spatial topography altogether (Cen et al., 2026; Palaganas et al., 2026).
And then there is the clinic, which imposes its own, rather unforgiving, constraints. Moving spatial multi-omics from the research bench into standard practice runs into pre-analytical, analytical, and operational hurdles that are easy to underestimate (Du & Yang, 2025). Current assays remain expensive, low-throughput, and often incompatible with routine sample handling (Wu et al., 2026). Many ultra-high-resolution imaging platforms demand fresh-frozen tissue, which puts them at odds with the formalin-fixed paraffin-embedded (FFPE) archives that most pathology departments actually rely on -- though newer methods such as patho-DBiT and spatial CITE-seq are beginning to chip away at this incompatibility (Wu et al., 2026). Compounding this, the absence of standardized operating procedures, validated inter-laboratory reproducibility, and clinically actionable interpretive thresholds continues to limit how spatial biomarkers are used in real-world patient stratification and trial design (Du & Yang, 2025; Mohr et al., 2024). Closing these gaps is, we would argue, the difference between spatial multi-omics remaining an elegant research curiosity and becoming a genuine companion diagnostic.
To address these scientific, computational, and clinical roadblocks, this work is organized around three guiding questions: (1) how do spatial gradients of epigenetic remodeling and transcriptional activity around pathological lesions -- amyloid-beta plaques in Alzheimer's disease, or hypoxic zones in solid tumors -- dictate localized cellular vulnerability, metabolic zonation, and functional decline; (2) can geometric deep learning architectures and multimodal foundation models reliably perform diagonal data integration and cross-modal imputation across spatial transcriptomics, proteomics, and metabolomics datasets without introducing artifact bias or oversmoothing biologically meaningful microdomains; and (3) what standardized pre-analytical protocols must be established so that spatial multi-omics biomarkers extracted from FFPE clinical cohorts achieve greater than 90% technical reproducibility across multiple centers.
Guided by these questions, this study pursues four objectives: to synthesize the published evidence on spatial co-localization of chromatin accessibility, mRNA transcripts, and protein abundance across healthy and pathological tissue, as reported using spatial CITE-seq and spatial ATAC-RNA-seq frameworks; to propose and specify, as a reproducible computational protocol, an interpretable graph-attention-network architecture for fusing spatial transcriptomic and metabolomic data from consecutive tissue sections, without generating de novo empirical validation, which falls outside the scope of this literature-grounded synthesis; to evaluate, from the published record, the predictive and prognostic utility reported for spatially resolved multi-omic biomarkers — including immune-cell proximity scores and stromal-exclusion indices — in patient stratification and therapy-response prediction; and to outline the design requirements for a standardized, open-access, cloud-native bioinformatic pipeline that could automate preprocessing, spatial domain identification, and cell-cell communication modeling, in service of the broader goal of democratizing clinical access to spatial multi-omics.

