Having framed the problem in the introduction, it seems worth pausing here to walk more slowly through what the existing literature actually says - not just about why orphan crops matter, but about the specific technical and institutional machinery that researchers have been assembling to work with them. The picture that emerges, admittedly, is somewhat fragmented across disciplines - genebank science, structural genomics, gene editing, and computational phenomics do not always cite one another as often as they probably should - but taken together they sketch a reasonably coherent trajectory.
2.1 The Agrobiodiversity Bottleneck and the Case for Genetic Resources
It is tempting to think of modern agriculture's yield gains as an unambiguous success story, and in narrow productivity terms, they are. Conventional breeding and intensive monoculture have driven remarkable output increases over the past century. Yet several authors have argued, convincingly, that this same trajectory has quietly engineered a precarious bottleneck (Huang et al., 2025; Sanfeliu Meliá et al., 2026). Of an estimated 390,000 vascular plant species, only 5,000 to 7,000 have ever entered cultivation, and a mere 250 have undergone anything resembling true domestication (Wang & Xiang, 2025). More troubling still, over 80% of global agricultural output now rests on five staple crops (Huang et al., 2025; Sanfeliu Meliá et al., 2026) - a level of concentration that, as Hu et al. (2025) note, leaves food systems structurally exposed to climate shocks, emergent pathogens, and resource scarcity.
The literature returns again and again to the Irish Potato Famine of 1845-1850 as a cautionary reference point (Huang et al., 2025), and not without reason - it remains perhaps the clearest historical demonstration of what genetic uniformity can cost a population. What is less often discussed, but arguably just as important, is the demographic pressure layered on top of this vulnerability: population projections now point toward roughly 10.3 billion people by the 2080s (Huang et al., 2025), even as droughts, thermal extremes, salinity, and transboundary pests threaten to erode staple yields further (Gelaye et al., 2025; Diakite et al., 2026). Several reviews converge on the same conclusion here - that traditional breeding, built around multi-generational hybridization and phenotypic selection, is simply too slow to keep pace (Gelaye et al., 2025; Rasool & Qadir, 2026). This is essentially the argument for shifting toward a precision-designed breeding paradigm that draws on both crop wild relatives and modern genomic tools (Somegowda et al., 2024; Hu et al., 2025).
2.2 Genebanks as Strategic Capital
Plant genetic resources - the seeds, tissues, landraces, and wild relatives that escaped the narrowing pressures of intensive selection - are described throughout this literature as a form of strategic capital (Ghamkhar, 2022), and it is hard to disagree with that framing once the scale involved becomes clear. Roughly 1,750 genebanks worldwide now conserve some 7.4 million accessions (Ghamkhar, 2022), a number that has only grown more valuable as the risks of narrow cultivation have become more apparent. Institutional case studies scattered through this literature help put that abstraction into perspective: ICRISAT alone conserves over 129,000 accessions across 144 countries, with particularly deep holdings in sorghum, minor millets, chickpea, and pigeonpea (Wang & Xiang, 2025); India's National Bureau of Plant Genetic Resources holds tens of thousands of sorghum, millet, and chickpea accessions (Wang & Xiang, 2025); and IITA maintains what is likely the world's most diverse cowpea collection alongside substantial yam and Bambara groundnut holdings (Wang & Xiang, 2025). Table 5 summarizes these holdings across the five largest repositories.
What several authors flag as the more interesting development, though, is not the size of these collections but their changing character. As sequencing costs continue falling, genebanks are gradually moving away from being passive seed vaults and becoming, instead, active digitized innovation centers (Ghamkhar, 2022). Layering genomic, transcriptomic, and phenomic profiling directly onto conserved accessions allows breeders to largely bypass the long, labor-intensive cycles of field-based phenotypic screening (Ghamkhar & Richards, 2022; Hu et al., 2025) - a shift that Zenda et al. (2021) and Hu et al. (2025) both describe as central to genomic-assisted breeding (GAB), where virtual allele mining substitutes, at least partly, for the exhaustive evaluation of individual field plants.
2.3 From Single Reference Genomes to Pangenomes
One recurring theme across this body of work is a certain dissatisfaction with the single linear reference genome - not because it was ever wrong, exactly, but because it was never going to be enough. Hu et al. (2025) and Diakite et al. (2026) both describe "reference bias" as a structural limitation: a single assembly simply cannot capture the full spectrum of variation within a species, and this problem becomes especially acute in heterozygous or polyploid orphan crops, where reference-based alignment tends to miss non-reference regions and complex structural variants altogether (Diakite et al., 2026; Hu et al., 2025). Table 2 lists the single-accession reference-genome assemblies now available for ten representative underutilized crop species.
The response the field has settled on - not universally, but widely - is the pangenome: an assembly built from multiple diverse individuals that partitions genetic content into a "core" genome shared across all individuals and a "variable" or dispensable genome present only in some (Zenda et al., 2021; Hu et al., 2025). What makes this distinction more than academic bookkeeping is where the interesting genes tend to sit. Comparative work across multiple species has shown, fairly consistently, that genes tied to environmental adaptation - disease resistance, drought tolerance, thermal response, nutrient acquisition - are disproportionately concentrated within this variable genome (Zenda et al., 2021; Hu et al., 2025). A pangenome of Glycine soja, the wild relative of cultivated soybean, illustrates this well: its variable genome comprised roughly 20% of total gene space, much of it showing signatures of positive selection around seed composition, flowering time, and biotic resistance (Zenda et al., 2021).
Graph-based pangenomes represent, in a sense, the natural next step - representing conserved regions as shared paths and structural variants as alternative branches, which allows more accurate read alignment and supports pangenome-wide association studies (Pan-GWAS) linking structural variants directly to phenotype (Hu et al., 2025). The examples that keep surfacing in this literature are instructive: in foxtail millet, graph pangenome analysis identified a 366-bp promoter-region variant in SiGW3 that suppresses gene expression to enhance grain weight (Hu et al., 2025); in pearl millet, similar analysis linked structural variants in endoplasmic-reticulum-related genes to heat stress adaptation via an expanded RWP-RK transcription factor family (Hu et al., 2025); and in chickpea, graph pangenomes have helped resolve superior haplotypes governing vernalization and disease resistance (Hu et al., 2025).
2.4 Precision Editing Tools and Delivery Platforms
If pangenomics tells researchers where to look, CRISPR/Cas systems are largely what has made it possible to act on that information with any precision (Rasool & Qadir, 2026). The mechanism, by now, is reasonably well established across the literature: a custom single-guide RNA directs an endonuclease such as Cas9 or Cas12a to a specific locus, inducing a double-strand break that the plant's own repair machinery - either non-homologous end joining or homology-directed repair - then resolves into a targeted disruption, deletion, or allele replacement (Rasool & Qadir, 2026; Mahmood et al., 2022). Table 3 catalogs representative functional target genes edited using these platforms across ten orphan crop species.
What has evolved more recently, and what several reviews treat as the more consequential development, is precision beyond simple knockouts. Base editors - cytosine and adenine variants - allow direct single-nucleotide conversion without inducing double-strand breaks or requiring a donor template, which meaningfully reduces unwanted indels (Rasool & Qadir, 2026). Prime editing goes further still, combining a nickase-impaired Cas9 with a reverse transcriptase to write new genetic sequence directly at a target site, enabling essentially any base-to-base conversion alongside small insertions and deletions, again without double-strand cleavage (Rasool & Qadir, 2026).
Even so, the literature is fairly candid that molecular precision alone does not guarantee translation into a living plant. Tissue transformation and regeneration remain persistent bottlenecks, particularly in recalcitrant or polyploid species (Rasool & Qadir, 2026; Manoj Kumar et al., 2026). The emerging response - nanoparticle-based and non-viral delivery systems, including lipid nanoparticles and polyester carriers - appears designed specifically to sidestep this problem, protecting fragile ribonucleoprotein complexes from degradation and enabling transgene-free, multiplexed editing without foreign DNA integration (Rasool & Qadir, 2026).
2.5 De Novo Domestication as a Case-Study Literature
Perhaps the most persuasive part of this literature, at least in terms of demonstrated results rather than promise, concerns de novo domestication - essentially compressing thousands of years of selection for "domestication syndrome" traits (reduced shattering, erect growth, day-length insensitivity, enlarged seed size) into a handful of growing seasons by editing known domestication-gene orthologs directly
Table 1: High-Value Breeding Targets and Candidate Genes Identified in Underutilized Crops This table lists eleven candidate genes reported across major orphan and underutilized crops, spanning tef, fonio, finger millet, Bambara groundnut, amaranth, jackfruit, taro, quinoa, sorghum, lentil, and cowpea. For each species it names the gene, its known or proposed function (e.g., seed-size regulation, stress-responsive transcription, metal chelation), and the specific breeding or agronomic trait the gene is expected to improve, such as harvestability, drought tolerance, or biofortification. All entries are drawn from a single synthesis source (Gelaye et al., 2025) rather than from the primary gene-discovery studies themselves — see cross-check note T1 below.
|
Crop
|
Species
|
Target Gene
|
Identified Function
|
Breeding and Agronomic Relevance
|
|
Tef
|
Eragrostis tef
|
EtG1 & EtG2
|
Regulation of seed and grain size
|
Enhances mechanical harvestability and overall yield (Gelaye et al., 2025)
|
|
Fonio
|
Digitaria exilis
|
FoWRKY
|
Abiotic stress transcription factor
|
Confers superior drought stress tolerance (Gelaye et al., 2025)
|
|
Finger Millet
|
Eleusine coracana
|
Rc (regulatory locus)
|
Calcium accumulation & storage
|
Drives outstanding calcium density in grain (Gelaye et al., 2025)
|
|
Bambara Groundnut
|
Vigna subterranea
|
GmFAD2
|
Desaturase metabolic enzyme
|
Optimizes seed oil composition and quality (Gelaye et al., 2025)
|
|
Amaranth
|
Amaranthus spp.
|
TPS
|
Trehalose-6-phosphate synthase
|
Confers superior salt and osmotic stress tolerance (Gelaye et al., 2025)
|
|
Jackfruit
|
Artocarpus heterophyllus
|
AhePG1
|
Polygalacturonase enzyme
|
Regulates fruit expansion and tissue softening (Gelaye et al., 2025)
|
|
Taro
|
Colocasia esculenta
|
CeMT2b
|
Metallothionein-like protein
|
Confers toxic heavy metal tolerance (Gelaye et al., 2025)
|
|
Quinoa
|
Chenopodium quinoa
|
CqZAT
|
Zinc transporter transcription factor
|
Drives high-level grain zinc accumulation (Gelaye et al., 2025)
|
|
Sorghum
|
Sorghum bicolor
|
Dw1
|
Auxin transport regulator
|
Controls internode elongation and plant height (Gelaye et al., 2025)
|
|
Lentils
|
Lens culinaris
|
LcRGA1, LcRGA2
|
Disease resistance gene analogs
|
Confers resistance against Ascochyta blight (Gelaye et al., 2025)
|
|
Cowpea
|
Vigna unguiculata
|
VfTT8
|
Anthocyanin regulator
|
Controls seed coat pigmentation and antioxidant profile (Gelaye et al., 2025)
|
Table 2: Reference-Genome Assemblies and Sequencing Metrics for Ten Underutilized Crop Species This table reports the reference genome assembled for one benchmark cultivar/accession of each of ten underutilized crops (e.g., pigeonpea "Asha," foxtail millet "Yugu1," cassava "AM560-2," teff "DZ-Cr-37"), giving assembly size, chromosome number, ploidy, and a link to the hosting database. It is intended to show how far whole-genome sequencing has progressed across taxonomically diverse orphan species. Note: none of the nine "Reference Citation" entries in this table (e.g., Varshney et al., 2012; Diao & Jia, 2017; Zhang et al., 2016; Hufnagel et al., 2021) correspond to any entry in the manuscript's 93-item reference list — see cross-check note T2 below.
|
Species Name
|
Common Name
|
Cultivar / Accession
|
Genome Assembly Size (Mb/Gb)
|
Chromosome Number (2n)
|
Ploidy Level
|
Primary Reference / Database Link
|
Reference Citation
|
|
Cajanus cajan
|
Pigeonpea
|
Asha (ICPL 87119)
|
606 Mb
|
2n = 22
|
Diploid
|
LegumeInfo Pigeonpea Database
|
Varshney et al. (2012)
|
|
Setaria italica
|
Foxtail Millet
|
Yugu1
|
400 Mb
|
2n = 18
|
Diploid
|
NCBI Assembly GCA_000263155.2
|
Diao & Jia (2017)
|
|
Manihot esculenta
|
Cassava
|
AM560-2
|
533 Mb
|
2n = 36
|
Diploid
|
Phytozome Cassava Portal
|
Wang et al. (2014)
|
|
Cicer arietinum
|
Chickpea
|
CDC Frontier
|
738 Mb
|
2n = 16
|
Diploid
|
NCBI Assembly GCA_000331145.1
|
Varshney et al. (2013)
|
|
Eragrostis tef
|
Teff
|
DZ-Cr-37
|
672 Mb
|
2n = 40
|
Allotetraploid
|
Tef Genome Research Portal
|
Zhang et al. (2016)
|
|
Eleusine coracana
|
Finger Millet
|
ML-365
|
1.19 Gb
|
2n = 36
|
Allotetraploid
|
NCBI Assembly GCA_002180455.1
|
Hufnagel et al. (2021)
|
|
Chenopodium quinoa
|
Quinoa
|
PI 614886
|
1.33 Gb
|
2n = 36
|
Allotetraploid
|
Phytozome Quinoa Portal
|
Jarvis et al. (2017)
|
|
Ipomoea batatas
|
Sweetpotato
|
Taizhong 6
|
870 Mb
|
2n = 90
|
Hexaploid
|
Sweetpotato public database
|
Zhang et al. (2020)
|
|
Dioscorea rotundata
|
Guinea Yam
|
TDr96_F1
|
594 Mb
|
2n = 40
|
Diploid
|
Guinea Yam Genome Center
|
Njaci et al. (2023)
|
|
Moringa oleifera
|
Moringa / Drumstick
|
NTBG-001
|
217 Mb
|
2n = 28
|
Diploid
|
ORCAE-AOCC Portal
|
Chang et al. (2019)
|
Table 3: Functional Genes Underlying Climate-Smart and Biofortification Traits in Ten Orphan Crops This table catalogs ten functionally characterized genes across orphan crops, describing each gene's molecular mode of action (e.g., CRISPR knockout, overexpression, heterologous expression), the physiological pathway it affects, and the resulting breeding application — ranging from lodging resistance in teff to zinc biofortification in quinoa and blight resistance in lentil. It complements Table 1 by adding mechanistic detail (mode of gene action and physiological function) that the earlier table omits. Two of the ten reference citations (Mayes et al., 2019; Ho et al., 2024) do not appear in the reference list — see cross-check note T3.
|
Crop Name
|
Botanical Species
|
Target Gene Symbol
|
Original Source Organism
|
Mode of Gene Action
|
Major Physiological Function
|
Crop Breeding Application
|
Reference Citation
|
|
Teff
|
Eragrostis tef
|
SD-1
|
Eragrostis tef
|
CRISPR Knockout
|
GA20-oxidase biosynthetic disruption
|
Semidwarf architecture, lodging resistance
|
Venezia & Krainer (2021)
|
|
Fonio
|
Digitaria exilis
|
FoWRKY
|
Digitaria exilis
|
Transcriptional activator
|
ABA-mediated stomatal regulation
|
Drought and desiccation tolerance
|
Gelaye et al. (2025)
|
|
Finger Millet
|
Eleusine coracana
|
EcbHLH57
|
Eleusine coracana
|
Overexpression
|
Upregulates LEA14, rd29A, SOD, APX
|
Multi-stress tolerance (salinity, drought)
|
Babitha et al. (2015)
|
|
Bambara Groundnut
|
Vigna subterranea
|
GmFAD2
|
Vigna subterranea
|
Heterologous expression
|
Oleic-to-linoleic acid desaturation
|
Customized seed lipid profile modification
|
Mayes et al. (2019)
|
|
Amaranth
|
Amaranthus spp.
|
TPS
|
Amaranthus spp.
|
Transcriptional regulator
|
Elevates cellular trehalose accumulation
|
Hyper-accumulation of osmoprotectants
|
Li et al. (2011)
|
|
Taro
|
Colocasia esculenta
|
CeMT2b
|
Colocasia esculenta
|
Metallothionein chelation
|
Heavy metal binding (Cadmium/Arsenic)
|
Environmental phytoremediation
|
Kim et al. (2011)
|
|
Quinoa
|
Chenopodium quinoa
|
CqZAT
|
Chenopodium quinoa
|
Zinc-finger transcription factor
|
Regulates root-to-shoot metal transport
|
Enhanced zinc biofortification in grains
|
Alvarez-Vasquez et al. (2025)
|
|
Winged Bean
|
Psophocarpus tetragonolobus
|
WbLEA
|
Psophocarpus tetragonolobus
|
Osmoprotectant chaperone
|
Stabilizes cellular proteins under desiccation
|
High seedling survival in arid zones
|
Ho et al. (2024)
|
|
Sorghum
|
Sorghum bicolor
|
Dw1
|
Sorghum bicolor
|
Membrane transporter
|
Coordinates auxin hormone distribution
|
Dwarfing gene targeting for height control
|
Boatwright L et al. (2024)
|
|
Lentils
|
Lens culinaris
|
LcRGA1
|
Lens culinaris
|
Nucleotide-binding site LRR
|
Triggers localized hypersensitive response
|
Quantitative resistance to Ascochyta blight
|
Dadu et al. (2021)
|
Table 4: Benchmarked AI/Machine-Learning Models for Crop Yield and Phenological Trait Prediction This table compares ten published machine-learning and deep-learning studies that predict agronomic traits (yield, flowering date, growth stage, tuber morphology) from remote-sensing or sensor data, across barley, wheat, rice, maize, potato, and cassava. For each study it lists the input data source, model architecture, validation strategy, and the headline performance metric (R², accuracy, or RMSE), enabling direct comparison of model classes on comparable prediction tasks. This table's ten reference citations match the manuscript's reference list correctly.
|
Crop Target
|
Phenological / Yield Trait
|
Core Sensor / Input Data Source
|
Geographic Location / Dataset Size
|
AI Model / Architecture
|
Validation Strategy
|
Key Performance Output
|
Reference Citation
|
|
Barley
|
Harvest Index (HI) & Agronomic Diversity
|
Field-level morphological metrics
|
OTGB, Ankara, Turkey (445 accessions)
|
PCA + Ward's Clustering + XGBoost
|
70/30 train/test; 5-fold CV
|
RMSE = 0.137%; MAPE = 0.222%
|
Akdogan et al. (2025)
|
|
Wheat
|
Early Season Yield Prediction
|
Satellite NDVI, Precipitation, LST
|
25 Districts, Turkey (TurkStat datasets)
|
SVR, RF, ANN, and LSTM
|
70/15/15 train/val/test
|
LSTM best: R² = 0.958
|
Akcapınar & Apaydin (2025)
|
|
Wheat
|
Flowering Date & Growth Stage ID
|
Smartphone canopy imagery + Climate
|
USA & UK multi-site trials (70,410 images)
|
GSP-AI Multimodal Network
|
70/20/10 train/test/val
|
Growth stage accuracy = 91.2%
|
Shen et al. (2024)
|
|
Wheat
|
District-level Yield Forecasting
|
MODIS/Sentinel-2 imagery + Weather data
|
Southern Pakistan
|
DeepAgroNet Deep Learning (CNN+ANN)
|
Multi-year historical validation
|
CNN best: R² = 0.77; Accuracy = 98%
|
Ashfaq et al. (2025)
|
|
Rice
|
Early Seedling Stage Growth Tracking
|
UAV RGB canopy images
|
Guangdong, China (49,840 images)
|
EfficientNetB4 + Transfer Learning
|
60/20/20 train/val/test
|
EfficientNetB4: Accuracy = 99.47%
|
Tan et al. (2022)
|
|
Rice
|
Salt/Drought Stress Level Classification
|
Leaf RWC, MDA, H2O2, and chlorophyll
|
Nakhon Phanom, Thailand (132 accessions)
|
Stacked Ensemble (XGBoost, SVM, MLP)
|
Stratified 80/20 train/test; 5-fold CV
|
Stacking accuracy = 81.81%
|
Gunnula et al. (2025)
|
|
Maize
|
Days to Tasseling (DTT) & Plant Height
|
360° Smartphone video + 3D Splatting
|
5 Chinese provinces (6,210 F1 hybrids)
|
CNN + MLP Environmental Module
|
8:1:1 train/val/test; Cross-province
|
Height MAE = 1.53 cm; GS R > 0.80
|
Wu et al. (2025)
|
|
Maize
|
Hybrid-specific Yield Forecasting
|
BLUP breeding values + Meteorological series
|
Huang-Huai-Hai Plain (2,096 observations)
|
SVR, RF, XGBoost, and GPR
|
10-fold CV; Feature Importance
|
RF achieved R² = 0.64; RMSE = 1010 kg/ha
|
Wang et al. (2025a)
|
|
Potato
|
Tuber Size, Shape, and Defect Profiling
|
Low-cost RGB crop scanner
|
Aberdeen, Idaho (189 biparental families)
|
OpenCV/OpenCV-based CNNs
|
Single-site; 80/20 train/test
|
Size and shape r² > 0.93
|
Feldman et al. (2024)
|
|
Cassava
|
Water Footprint & Yield Forecasting
|
Local weather station time-series
|
Nanning, China (14 weather stations)
|
SVM + ANN + SARIMA
|
5-fold CV; Hold-out test sets
|
ANN prediction accuracy = 97.16%
|
Tao et al. (2023)
|
(Huang et al., 2025; Sanfeliu Meliá et al., 2026; Hu et al., 2025). Table 1 lists representative high-value breeding targets identified across eleven such crops.
The case studies accumulated here are genuinely varied. In sweet potato, editing starch synthase and branching enzyme genes altered amylose-to-amylopectin ratios to customize industrial starch properties without sacrificing yield (Wang & Xiang, 2025). In tef, knocking out a rice SEMIDWARF-1 ortholog produced lodging-resistant semidwarf plants suited to mechanical harvest (Wang & Xiang, 2025; Hu et al., 2025). In groundcherry, multiplexed knockouts of SP, SP5G, and CLV orthologs transformed a sprawling wild architecture into a compact, higher-yielding plant (Wang & Xiang, 2025; Hu et al., 2025). Broomcorn millet and sorghum studies extend the pattern further still - dwarfing edits for high-density planting in the former (Wang & Xiang, 2025), and both aromatic-trait and parasitic-weed-resistance edits in the latter (Wang & Xiang, 2025).
A parallel and, in the literature's own framing, equally important thread concerns anti-nutritional factors. Wild species are often unusually rich in iron, zinc, and vitamins, but these minerals are frequently rendered less bioavailable by phytic acid, saponins, tannins, and oxalates (Zenda et al., 2021; Sanfeliu Meliá et al., 2026). Selectively silencing the biosynthetic genes behind these compounds is, several authors argue, just as central to de novo domestication as fixing plant architecture - arguably more so, given the direct link to human nutrition (Zenda et al., 2021; Sanfeliu Meliá et al., 2026).
2.6 Systems Biology, AI, and the Push Toward Digital Twins
Complex agronomic traits, the literature is at pains to point out, rarely trace back to a single gene or pathway - yield and drought tolerance in particular emerge from densely interconnected molecular networks (Zenda et al., 2021; Chen et al., 2025). Making sense of this complexity is largely what has pushed the field toward systems biology, integrating genomics, transcriptomics, proteomics, metabolomics, epigenomics, phenomics, and enviromics into something closer to a unified model of plant behavior (Zenda et al., 2021; Chen et al., 2025). Table 6 outlines the ten omics layers referenced across this literature and their respective high-throughput technology platforms.
Machine learning has become, almost by necessity, the primary tool for extracting signal from these heterogeneous datasets. Random forests, support vector machines, convolutional neural networks, and graph neural networks are all represented in this literature as methods for predicting agronomic traits and identifying promising parental combinations (Mahmood et al., 2022; Diakite et al., 2026; Chen et al., 2025). Table 4 benchmarks ten representative AI/ML frameworks used for this purpose across major crop-trait prediction tasks. One development worth flagging specifically is the growing use of explainable AI - SHAP-based pathway attributions, for instance - as a corrective to deep learning's well-known "black-box" problem, aiming to deliver predictions that are at least somewhat interpretable in biological terms (Chen et al., 2025).
Where this seems to be heading, based on the more forward-looking papers in this set, is toward virtual crop systems and digital twins - simulations that couple mechanistic models (ordinary differential equations, flux balance analysis) with data-driven machine learning to model plant development and stress response in silico (Chen et al., 2025). The DSAP strategy - de novo domestication, speed breeding, and AI-empowered phenomics - represents, in Huang et al.'s (2025) framing, the most concrete attempt yet to operationalize this systems view into an actual breeding pipeline, cycling through controlled-environment domestication, accelerated generation turnover, and continuous non-destructive phenotyping.
2.7 Regulatory Context and Open Science
None of the preceding technology matters much, in practical terms, without a regulatory and institutional environment that allows it to reach farmers - and this is a point the literature makes with some insistence (Ghamkhar & Richards, 2022; Gelaye et al., 2025). GMO deployment has historically faced considerable regulatory friction, high intellectual-property costs, and public skepticism tied to the introduction of foreign DNA (Rasool & Qadir, 2026). Precision gene editing, when it avoids inserting transgenic material altogether, appears to offer at least a partial way around these obstacles (Rasool & Qadir, 2026), and a growing number of national regulatory bodies now draw an explicit distinction between gene-edited and transgenic crops.
Kenya and Nigeria have moved to establish streamlined, science-based biosafety frameworks that several authors describe as likely drivers of broader biotechnology
Table 5: Global Ex-Situ Genebank Holdings for Major Underutilized Crop and Legume Species This table lists ten germplasm holdings held by five international and national genebanks — NBPGR (India), ICRISAT, IITA, CIAT, and CIP — reporting the crop, accession count, and preservation method (seed vault, in-vitro culture, cryopreservation) for each. It documents the scale of ex-situ conservation available as breeding source material and highlights each repository's specialization, e.g., ICRISAT for semi-arid legumes and millets and IITA for cowpea and yam. No reference citations are attached to this table; it summarizes institutional holdings rather than published findings.
|
Crop / Cereal Class
|
Botanical Taxon
|
Genebank Repository
|
Institutional Name
|
Headquarters Location
|
Preserved Accessions
|
Mode of Preservation
|
Adaptive Diversity Focus
|
|
Sorghum
|
Sorghum bicolor
|
NBPGR
|
National Bureau of Plant Genetic Resources
|
New Delhi, India
|
26,395 accessions
|
Ex-situ seed vault
|
Heat tolerance, local landraces
|
|
Minor Millets
|
Setaria / Eleusine spp.
|
NBPGR
|
National Bureau of Plant Genetic Resources
|
New Delhi, India
|
25,785 accessions
|
Ex-situ seed storage
|
High calcium, drought resilience
|
|
Chickpea
|
Cicer arietinum
|
NBPGR
|
National Bureau of Plant Genetic Resources
|
New Delhi, India
|
14,904 accessions
|
Ex-situ seed storage
|
Protein quality, nitrogen fixation
|
|
Legumes & Millets
|
Sorghum, Millets, Cicer
|
ICRISAT
|
Int. Crops Research Institute for Semi-Arid Tropics
|
Patancheru, India
|
129,000 accessions
|
Global ex-situ seed bank
|
Semi-arid tropical climate resilience
|
|
Cowpea
|
Vigna unguiculata
|
IITA
|
International Institute of Tropical Agriculture
|
Ibadan, Nigeria
|
15,000 accessions
|
Ex-situ seed collection
|
Insect resistance, high folate
|
|
Yam
|
Dioscorea spp.
|
IITA
|
International Institute of Tropical Agriculture
|
Ibadan, Nigeria
|
5,900 accessions
|
In-vitro tissue culture
|
Disease tolerance, tuber size
|
|
Bambara Groundnut
|
Vigna subterranea
|
IITA
|
International Institute of Tropical Agriculture
|
Ibadan, Nigeria
|
2,000 accessions
|
Ex-situ cold vaults
|
Soil nitrogen fixation, poor soils
|
|
Cassava
|
Manihot esculenta
|
CIAT
|
International Center for Tropical Agriculture
|
Palmira, Colombia
|
5,965 accessions
|
In-vitro cloned germplasm
|
High starch, root rot resistance
|
|
Cassava
|
Manihot esculenta
|
IITA
|
International Institute of Tropical Agriculture
|
Ibadan, Nigeria
|
3,700 accessions
|
Field and in-vitro duplicate
|
African cassava mosaic virus immunity
|
|
Sweetpotato
|
Ipomoea batatas
|
CIP
|
International Potato Center
|
Lima, Peru
|
Over 5,500 accessions
|
Tissue culture, Cryopreservation
|
High beta-carotene, weevil resistance
|
Table 6: Multi-Omics Technology Layers and Their Role in Systems-Level Crop Improvement This table outlines ten "omics" layers used in crop research — genomics, epigenomics, transcriptomics, proteomics, metabolomics, phenomics, enviromics, single-cell omics, spatial omics, and interactomics — describing the high-throughput technology, biological information captured, data-integration complexity, and representative crops for each layer. It is meant as a conceptual map of how these layers combine into a systems-biology pipeline. Eight of the ten "Reference Citation" entries (e.g., Varshney et al., 2021; Springer & Schmitz, 2017; Tardieu et al., 2017) do not correspond to any entry in the reference list — see cross-check note T4.
|
Omics Layer
|
High-Throughput Technology
|
Primary Target Feature
|
Biological Information Type
|
Data Integration Complexity
|
Physiological Application
|
Representative Crops
|
Reference Citation
|
|
Genomics
|
Illumina, PacBio, Nanoparticle WGS
|
SNPs, indels, SVs, and CNVs
|
Haplotype maps & genomic variations
|
Medium
|
QTL mapping, genomic selection
|
Rice, Maize, Sorghum
|
Varshney et al. (2021)
|
|
Epigenomics
|
ChIP-Seq, Methyl-Seq, Cut&Tag
|
DNA methylation, Histone tags
|
Stress-induced chromatin plasticity
|
High
|
Gene expression regulation
|
Arabidopsis, Maize
|
Springer & Schmitz (2017)
|
|
Transcriptomics
|
RNA-Seq, Iso-Seq, bulk mRNA-seq
|
Differentially expressed genes
|
Comprehensive gene expression profiles
|
High
|
Identifying transcriptional hubs
|
Fonio, Quinoa, Tef
|
Lowe et al. (2017)
|
|
Proteomics
|
Shotgun Proteomics, Tandem MS
|
Functional enzymes, protein chains
|
Protein abundance & phosphorylation
|
Very High
|
Direct pathway mapping
|
Bread wheat, Barley
|
Zhang et al. (2013)
|
|
Metabolomics
|
LC-MS/MS, GC-MS, UHPLC-MS
|
Osmolytes, specialized metabolites
|
Cellular biochemical endpoints
|
Very High
|
Biomarker discovery
|
Purslane, Cowpea
|
Thingujam et al. (2025)
|
|
Phenomics
|
UAV multispectral, LiDAR, CT
|
Leaf area, transpiration, root system
|
Scale-up physiological phenotypes
|
Very High
|
Non-destructive screening
|
Potato, Yam, Barley
|
Tardieu et al. (2017)
|
|
Enviromics
|
Hyperspectral weather indicators
|
Meteorological time-series datasets
|
Dynamic environmental factors
|
Very High
|
G×E interaction modeling
|
Chickpea, Cowpea
|
Roorkiwal et al. (2018)
|
|
Single-Cell Omics
|
scRNA-Seq, Single-nucleus ATAC-Seq
|
Cell-type-resolved transcripts
|
Cellular heterogeneity map
|
Very High
|
Decoding tissue-specific stress
|
Arabidopsis, Rice
|
Islam et al. (2024)
|
|
Spatial Omics
|
Spatial transcriptomics, MS imaging
|
In-situ metabolite distribution
|
Spatial co-variation networks
|
Very High
|
Refines causal biological priors
|
Brassica, Maize
|
Chen et al. (2025)
|
|
Interactomics
|
Yeast Two-Hybrid, Affinity-MS
|
Protein-protein & protein-metabolite
|
Biochemical pathway networks
|
Very High
|
Synthetic network construction
|
Rice, Wheat
|
Bruni et al. (2024)
|
adoption across the continent (Rasool & Qadir, 2026), while Argentina's case-by-case regulatory model - treating transgene-free lines as conventional crops - has meaningfully sped up commercial approval timelines (Rasool & Qadir, 2026). Alongside these regulatory shifts, international consortia such as the African Orphan Crops Consortium, the Crops For the Future initiative, and the CGIAR Genebank Platform are actively working to democratize access to genomic tools by hosting open-source reference genomes and pangenomes for neglected species (Gelaye et al., 2025; Hu et al., 2025). Combined with participatory plant breeding - actively engaging smallholder farmers in trait prioritization and field trials - this open-science push is presented, fairly consistently across this literature, as the socio-technical complement that precision biology alone cannot provide (Gelaye et al., 2025).
2.8 Synthesis and the Gap This Review Addresses
Reading across this literature as a whole, a reasonably clear consensus emerges: the individual technical pieces - pangenomics, precision editing, systems biology, AI-phenomics - are each maturing quickly, and in several cases have already delivered concrete, field-relevant results. What remains comparatively underdeveloped, and what most of these papers acknowledge only briefly before moving on, is the connective tissue between laboratory-scale discovery and actual field-level deployment. Few of the reviewed sources attempt to synthesize genomic, editing, and phenomic advances into a single integrated narrative alongside the socio-technical and regulatory conditions needed for adoption. It is this gap - not a lack of underlying science, but a lack of integrative synthesis - that the present review attempts, however partially, to address.