1. Introduction
Osteoporosis, though often described in clinical terms as a reduction in bone mass and microarchitectural deterioration, is perhaps better understood as a quietly progressive condition—one that tends to remain unnoticed until it manifests through fragility fractures. These fractures, affecting millions globally each year, represent not only a biological failure of bone strength but also a substantial burden on healthcare systems and quality of life. It has been estimated that nearly 8.9 million fractures annually are attributable to osteoporosis, a figure that continues to rise with aging populations (Cheung et al., 2016). Despite this, early detection remains inconsistent, and in many cases, delayed.
Clinically, osteoporosis is most commonly diagnosed using bone mineral density (BMD) measurements obtained through dual-energy X-ray absorptiometry (DXA), which remains the gold standard. The widely adopted T-score classification—defining osteoporosis at ≤ −2.5 standard deviations and osteopenia between −1.0 and −2.5—provides a standardized diagnostic threshold (Dimai, 2017). Yet, even as this framework offers clarity, it does not fully capture the complexity of fracture risk. Bone strength is not determined by density alone; factors such as bone geometry, trabecular integrity, and patient-specific clinical variables play equally critical roles. Tools like FRAX have attempted to bridge this gap by incorporating clinical risk factors alongside BMD, offering a more holistic risk estimation (Curry et al., 2018). Still, the reliance on DXA introduces practical challenges—limited accessibility, cost constraints, and underutilization in certain populations.
Interestingly, the issue is not merely one of underuse. Some studies suggest a paradoxical pattern: over-screening among low-risk individuals and under-screening among those at higher risk (Amarnath et al., 2015). This imbalance raises questions about how screening strategies are implemented in real-world settings and whether alternative, more accessible diagnostic pathways might help address these disparities. In this context, routine imaging modalities—particularly plain radiographs—begin to appear as an underexplored opportunity.
Lumbar spine X-rays, for instance, are frequently obtained for a variety of clinical indications unrelated to osteoporosis. Yet, embedded within these images may lie subtle structural patterns indicative of reduced bone density—patterns that are not easily discernible through conventional visual assessment. It is here that machine learning, and more specifically deep learning, begins to offer a compelling possibility. Over the past decade, advances in artificial intelligence have enabled models to identify complex, high-dimensional relationships within medical imaging data—relationships that may escape even experienced clinicians (LeCun et al., 2015; Esteva et al., 2019).
Convolutional neural networks (CNNs), in particular, have demonstrated considerable success in image-based diagnostic tasks, ranging from dermatological classification to radiological interpretation. Their ability to automatically extract hierarchical features from raw pixel data makes them especially suited for applications in medical imaging. In the domain of osteoporosis, early investigations suggest that CNN-based approaches can potentially classify bone health status—normal, osteopenic, or osteoporotic—using standard radiographs. However, the extent to which these models can reliably approximate or complement DXA-based measurements remains an open question.
At the same time, it would be overly optimistic to assume that machine learning offers a straightforward solution. Concerns regarding model generalizability, data heterogeneity, and hidden biases continue to challenge the clinical translation of AI systems. Overfitting, in particular, remains a persistent issue, where models perform well on training data but fail to generalize to new populations (Lever et al., 2016). Moreover, variability in imaging protocols, patient demographics, and annotation practices introduces additional layers of complexity that are not easily resolved.
Against this backdrop, the present study seeks to explore a somewhat pragmatic question: can deep learning models, trained on routinely acquired lumbar spine X-ray images, serve as a viable adjunct—or at least a preliminary screening tool—for osteoporosis detection? More specifically, this work evaluates the feasibility of a CNN-based framework to classify osteopenia and osteoporosis using radiographic data, with DXA-derived BMD serving as the reference standard. By leveraging existing imaging data, such an approach may offer a cost-effective and scalable pathway toward improving screening coverage, particularly in settings where access to DXA is limited.
At the same time, this investigation does not assume that deep learning models can replace established diagnostic methods. Rather, it approaches the problem with a degree of caution—recognizing both the promise and the limitations of current AI methodologies. In doing so, the study aims to contribute not only to the technical development of diagnostic models but also to the broader conversation surrounding the responsible integration of artificial intelligence into clinical practice.
Ultimately, the challenge may not lie in choosing between traditional and emerging technologies, but in understanding how they can be meaningfully combined. And perhaps, in that intersection, there is an opportunity—one that is still unfolding—to rethink how osteoporosis is detected, assessed, and, ideally, prevented.





