2.1 Heterogeneity as the Central Problem
It is worth pausing on how the field arrived here. Early in the pandemic, trials and treatment protocols concentrated, reasonably, on acute respiratory failure and hyperinflammatory cytokine release. Only gradually did it become clear that a considerable proportion of infected people—up to 15% by some estimates—were left with long-lasting and sometimes disabling physical and cognitive sequelae (Zeraatkar et al., 2025). The symptom spectrum is wide: post-exertional malaise, chronic fatigue, dyspnea, autonomic dysregulation, gastrointestinal disturbance, and cognitive impairment all feature prominently (Pfaff et al., 2023; Zeraatkar et al., 2025).
The difficulty, as several authors have argued, lies partly in how the condition is classified. Diagnostic frameworks that rely on one code, such as ICD-10-CM U09.9, make it easy to count cases but hard to see differences between them (Pfaff et al., 2023). Work in the N3C database illustrates this tension quite clearly: U09.9 supports high-level cohort tracking, but it gathers people with very different organ involvement, symptom trajectories, and presumably different underlying drivers under one label (Pfaff et al., 2023). This, in turn, may be one reason why precision-medicine efforts and trial designs have struggled to gain traction (Halamka et al., 2020; Pfaff et al., 2023).
Systems biology offers an alternative lens. Network medicine views disease as an emergent property of perturbed biological networks rather than as an isolated defect in one cell type or organ (Barabási et al., 2011; Halamka et al., 2020). Within this view, perturbations tend to concentrate in local neighborhoods of the interactome—disease modules (Barabási et al., 2011; Conte et al., 2020). For a post-viral syndrome, the implication is fairly intuitive, even if it is not yet proven: alterations in host–pathogen interactions could spread dysregulation across several tissue-specific sub-networks, which would help explain why a localized viral entry event can end in a systemic, multi-organ illness (Coste et al., 2025; Halamka et al., 2020; Zhou et al., 2020b). Figure 1 sketches the overall framework we use to organize the rest of this review (Figure 1).
2.2 The Host–Pathogen Interactome and Its Tissue Footprint
2.2.1 Mapping the virus–host disease module
Understanding the molecular architecture of long COVID depends on combining datasets generated at different resolutions—bulk, single-cell, and single-nucleus. Multi-omic profiles that include genomic variants, RNA sequencing, liquid chromatography–mass spectrometry (LC-MS) proteomics, and serum metabolomics have been used to chart how viral proteins engage host machinery (Pavel et al., 2021; Tomazou et al., 2021; Zhou et al., 2020b). A key starting point was the affinity-purification mass spectrometry map produced by Gordon et al. (2020), which identified 332 high-confidence human proteins bound by SARS-CoV-2 proteins. When these proteins were projected onto the human interactome, most of them fell into a single large connected component, forming what Morselli Gysi et al. (2021) treated as the core virus–host disease module; Verstraete et al. (2020) reached a broadly similar picture using a multilayer representation. Messina et al. (2020) and Stolfi et al. (2020) likewise used interactome-based models to trace how viral–host contacts extend into wider host pathways.
2.2.2 Pulmonary, gastrointestinal, and metabolic signals
Mapping across tissues has revealed perturbations in several anatomical compartments (Pavel et al., 2021; Stolfi et al., 2020; Tomazou et al., 2021). In the airways and the gut, single-cell data show that the main entry receptors, ACE2 and TMPRSS2, are co-expressed in particular cell populations—airway club cells, type II pneumocytes, and absorptive enterocytes, including those in inflamed ileal tissue (Halamka et al., 2020; Zhou et al., 2020b). In Crohn’s disease models, raised enterocyte co-expression points toward shared sub-networks between viral biology and chronic intestinal inflammation (Zhou et al., 2020b). Metabolic data add another dimension: integrated transcriptomic and metabolomic analyses reported reduced L-arginine and L-citrulline and disturbed arachidonate signaling, patterns that resemble those seen in chronic airway disease such as asthma (Zhou et al., 2020b).
Beyond direct interactors, multi-scale integration using resources such as the Unified Knowledge Space (UKS) suggests that host-response genes one or two steps removed from the virus may be the ones that drive longer-term tissue change (Pavel et al., 2021). These intermediate genes—pro-angiogenic factors such as VEGF and IL-6, and connective-tissue remodeling pathways—could plausibly link acute immune activation to pulmonary vascular endothelialitis, microthrombosis, and fibrosis (Pavel et al., 2021).
2.2.3 The brain: an indirect route to injury
The central nervous system provides perhaps the clearest example of how network analysis can reframe a clinical question. Single-nucleus transcriptomic data across human brain regions show very low baseline neuronal expression of ACE2 and TMPRSS2, which argues against frequent direct neuroinvasion (Zhou et al., 2021). Network proximity analyses, however, indicate that SARS-CoV-2 host factors overlap substantially with modules for neuroinflammation and brain microvascular injury (Zhou et al., 2021). Alternative entry and docking factors—basigin (BSG), furin (FURIN), and neuropilins 1 and 2—together with innate antiviral genes (LY6E, IFITM2, IFITM3, IFNAR1), are expressed at higher levels in brain microvascular endothelial cells and microglia (Zhou et al., 2021). Changes in Alzheimer’s disease biomarkers in cerebrospinal fluid and blood from COVID-19 patients, including TGFB1, SPP1, CXCL10, and TNFRSF1B, further suggest that post-viral cognitive dysfunction may arise from endothelial injury and neuroinflammatory cascades rather than from direct destruction of brain tissue (Zhou et al., 2021). The way these organ-specific signals may fan out from a common interactome module is summarized in Figure 2 (Figure 2); the underlying data layers are compiled in Table 1 (Table 1).
2.3 Machine Learning and Computable Phenotyping
2.3.1 Community detection in large EHR networks
Turning network biology into clinical tools requires methods that can cope with very large, messy real-world datasets (Halamka et al., 2020; Jamshidi et al., 2020; Pfaff et al., 2023). Unsupervised learning is attractive here precisely because it does not assume the categories in advance (Ilbeigipour et al., 2022; Pfaff et al., 2023). In the N3C enclave, which holds records for more than 16 million patients, Pfaff et al. (2023) built networks of diagnoses co-occurring within 0–60 days of a long COVID index date and applied Louvain community detection. The result was a set of age-stratified clusters. A cardiopulmonary pattern

Figure 1. A network medicine framework for moving long COVID from a single diagnostic code to a taxonomy based on disease mechanisms The single ICD-10-CM code U09.9 is useful for identifying patients but hides differences in which organs are involved and what drives the disease. Genomic, transcriptomic, proteomic, metabolomic, and clinical data are mapped onto the human interactome. Network algorithms and machine learning then turn that map into disease modules and patient clusters. The outputs are clinical endotypes, readouts of the underlying mechanisms, and network-guided treatment hypotheses. The dashed arrow shows how prospective cohorts and randomized trials feed back to refine the classification.

Figure 2. Proposed spread of SARS-CoV-2 host–pathogen disturbances through the human interactome into organ-specific sub-networks Viral proteins bind 332 human host proteins through primary receptors (ACE2, TMPRSS2) and alternative docking factors (BSG, FURIN, NRP1). These host targets form one connected disease module. Intermediate genes such as IL6, VEGF, and ICAM1 appear to carry the disturbance into the lungs and blood vessels, heart, gut, and brain vasculature. The organ-level changes shown may together account for the multi-organ endotypes of long COVID. Most of the links shown come from studies of acute infection and remain hypotheses for long COVID.
featured chest pain, dyspnea, palpitations, tachycardia, and exercise intolerance. A neurological and cognitive pattern was dominated by chronic fatigue, brain fog, headache, sleep disturbance, and myalgic encephalomyelitis/chronic fatigue syndrome-like presentations. A gastrointestinal pattern—abdominal pain, nausea, diarrhea, and metabolic disturbance—appeared particularly among people younger than 21 years. An upper-respiratory pattern included persistent anosmia, dysgeusia, rhinitis, and pharyngitis (Halamka et al., 2020; Pfaff et al., 2023). Finally, an age-related comorbid pattern, concentrated in adults aged 65 and older, involved worsening of existing cardiovascular, metabolic, and neurodegenerative conditions such as heart failure and type 2 diabetes (Pfaff et al., 2023).
2.3.2 Clustering clinical trajectories
Graph-based methods are not the only option. K-means, X-means, and SOM neural networks have been used to group patients by clinical variables, laboratory biomarkers, and acute illness trajectories (Benito-León et al., 2021; Ilbeigipour et al., 2022). SOMs, for instance, project inputs such as age, length of hospital stay, intensive care admission, intubation history, and inflammatory marker levels onto a two-dimensional grid (Ilbeigipour et al., 2022). These analyses show that older age and higher inflammatory markers track closely with more severe courses—yet they also show that symptom severity and organ involvement vary a good deal within the same demographic strata (Ilbeigipour et al., 2022). Benito-León et al. (2021) similarly found severity subgroups that were not explained by age or sex alone, which is a small but useful reminder that demographics are an imperfect proxy for biology.
2.3.3 Language models and population networks
Natural language processing (NLP) has opened another door. Deep-learning-augmented curation platforms such as nferX can extract concepts from millions of free-text physician notes (Halamka et al., 2020). When this narrative information is triangulated with structured laboratory data and single-cell expression profiles across roughly 25 human tissues, subtle, uncoded features—early loss of smell or taste, for example—appeared to be better early predictors of outcome than fever or cough (Halamka et al., 2020). At the population level, Coste et al. (2025) modeled long COVID as sitting within a multidimensional web of acute severity, chronic comorbidity, social determinants, and occupational exposure. The combined phenotyping pipeline is illustrated in Figure 3 (Figure 3), and the individual methods are compared in Table 2 (Table 2).
2.4 Network-Based Drug Repurposing
2.4.1 The proximity hypothesis
Because conventional drug development can take a decade or more, network medicine offers a faster, if less certain, route through repurposing (Fiscon et al., 2021; Morselli Gysi et al., 2021; Zhou et al., 2020a). By estimating how approved drugs perturb host networks, algorithms can screen thousands of compounds relatively quickly (Li et al., 2021; Morselli Gysi et al., 2021; Santos et al., 2022). Most of these approaches rest on the network proximity hypothesis: a drug is more likely to be useful if its targets lie within, or close to, the disease module (Fiscon et al., 2021; Morselli Gysi et al., 2021; Stolfi et al., 2020). Proximity is commonly quantified as the average shortest-path distance between a drug’s target set T and the disease protein set S:
dT,S=1/Tt∈Tmins∈Sdt,s
where d(t, s) is the shortest path length between target t and disease protein s in the reference interactome, usually converted to a z-score against randomized target sets (Fiscon et al., 2021; Zhou et al., 2021). Li et al. (2021) provide a concrete example: starting from 34 disease-related genes and a network of 1,344 genes across 24 enriched pathways, their proximity analysis yielded 78 candidates, narrowed by expert review to 30.
2.4.2 Algorithms beyond simple distance
Several algorithms refine this basic idea. SAveRUNNER scores drug–disease pairs with a network similarity measure adjusted for cluster quality and sigmoidal normalization; applied to virus–host networks, it prioritized off-label candidates including anti-inflammatory agents, central nervous system modulators, histamine-receptor antagonists, and sodium-channel blockers (Fiscon et al., 2021). Random walk with restart (RWR) and related propagation methods simulate a walker that moves along interactome edges from disease seed nodes, returning to the seeds with a fixed probability; the resulting visitation scores help identify indirect host targets and functional sub-networks altered by infection (Fiscon et al., 2021; Messina et al., 2020; Stolfi et al., 2020). Graph neural networks and knowledge-graph embeddings go a step further by attempting to capture global topology (Hsieh et al., 2021; Santos et al., 2022). Hsieh et al. (2021), for example, built a COVID-19 knowledge graph linking viral proteins, host genes, pathways, drugs, and phenotypes, learned drug representations with a deep graph model initialized from a general biomedical knowledge graph, and then checked candidates against gene-set enrichment, in vitro screening data, and EHR-based treatment effects. Integrated platforms such as NeDRex combine module-detection methods (DIAMOnD, BiCoN) with drug-ranking methods (TrustRank, closeness centrality) in an expert-in-the-loop environment (Sadegh et al., 2021).
2.4.3 Consensus and the “network drug” observation
No single algorithm seems to capture everything. Recognizing this, Morselli Gysi et al. (2021) fused rankings from multiple pipelines—AI-based, diffusion-based, and proximity-based—into a consensus. Their rank-aggregation approach reportedly reached a hit rate of up to 62% in cell-based viral inhibition assays, compared with roughly 0.8% for unguided screening (Morselli Gysi et al., 2021). Perhaps the more interesting observation is that the great majority of active “network drugs” (over 95%) did not bind viral proteins at all; they acted on host proteins in the neighborhood of the disease module (Morselli Gysi et al., 2021). If that finding generalizes, it would support a strategy of modulating host sub-networks—dampening neuroinflammation, protecting endothelium—rather than chasing the virus alone (Morselli Gysi et al., 2021; Zhou et al., 2021). The full repurposing pipeline is depicted in Figure 4A (Figure 4), and the algorithms are compared in Table 3 (Table 3).
2.5 Network Pharmacology, Combinations, and Patient-Generated Evidence
2.5.1 Complementary Exposure
Because long COVID involves several pathways at once, single-target monotherapy may simply be insufficient (Santos et al., 2022; Zhou et al., 2020a). Network pharmacology offers a rational way to combine drugs so that effects add up while toxicities do not (Hsieh et al., 2021; Zhou et al., 2020a). The Complementary Exposure pattern formalizes this (Zhou et al., 2020a). Two drugs A and B are considered promising partners if both target sets are proximal to the disease module (z_{T_A,S} < 0 and z_{T_B,S} < 0) and, at the same time, their targets occupy separate neighborhoods, reflected by a positive separation score:
sAB=⟨dAB⟩-⟨dAA⟩+⟨dBB⟩/2>0
Here ⟨d_AB⟩ is the mean shortest distance between the targets of A and B, and ⟨d_AA⟩ and ⟨d_BB⟩ describe the internal spread of each target set (Zhou et al., 2020a). In principle, such a pair perturbs different arms of the same module without piling pressure onto the same cellular targets (Figure 4B). Using this logic, Zhou et al. (2020a) proposed pairings such as sirolimus with dactinomycin, mercaptopurine with melatonin, and toremifene with emodin, while Hsieh et al. (2021) highlighted etoposide with sirolimus and hydroxychloroquine with melatonin as complementary combinations.
2.5.2 Knowledge graphs and crowdsourced hypotheses
Interactive resources try to close the distance between prediction and practice (Koss & Bohnet-Joschko, 2022; Santos et al., 2022; Verstraete et al., 2020). CovMulNet19 integrates viral and human proteins, Gene Ontology terms, related diseases, symptoms, and candidate compounds in a single multilayer network, so that molecular perturbations can be read alongside clinical manifestations (Verstraete et al., 2020). CoREx provides a web environment for exploring drug perturbation scores, Connectivity Map (CMAP) expression-reversal signatures, and drug–target relationships within host sub-networks (Santos et al., 2022). Network pharmacology has also been applied to traditional medicine: Sun et al. (2020) combined frequency and association-rule mining of 173 historical prescriptions with network analysis and found considerable overlap with Lianhua Qingwen capsules and Ma Xing Shi Gan decoction, although they stressed the need for experimental confirmation.
Finally, social media mining may complement formal trials by capturing what patients are already trying. Koss and Bohnet-Joschko (2022) applied named-entity recognition and co-occurrence network analysis to posts in a large long COVID community on Reddit, extracting reported use of off-label medicines and supplements. Mapping such crowdsourced candidates back onto host networks could generate hypotheses that are patient-centered—though, of course, self-report is no substitute for controlled evaluation (Koss & Bohnet-Joschko, 2022; Zeraatkar et al., 2025). These platforms and combination models are summarized in Table 4 (Table 4).