1. Introduction
Every June, budget officials in Dhaka run into the same familiar tension, and it rarely gets easier with practice. The national budget has to be large enough to cover the government's wage bill in full, yet not so padded that money sits idle while other sectors go wanting. For fiscal year 2020-21, Bangladesh's operating expenditure alone ran into the hundreds of thousands of crore Taka, a substantial share of which was earmarked for pay, allowances, and pension obligations covering well over a million public employees (Ministry of Finance, 2020). Put that way, the figure can feel almost too large to reason about, abstract, distant, more a headline number than a lived reality. But underneath it sits a fairly ordinary administrative problem, one that has less to do with macroeconomics than with bookkeeping: nobody, it seems, has a clean, data-driven picture of who is actually drawing on which benefit, at what age, and in what pattern.
That gap matters more than it might first appear, and perhaps more than officials themselves would readily admit. For years, ministries have leaned on largely manual, precedent-based estimation when projecting pay-related expenditure (Ministry of Finance, 2024), a method that is not unreasonable in the absence of better tools, but one that inevitably drifts as workforce demographics shift underneath it. It is not that officials lack data. Personnel files exist. Provident-fund records exist. Loan ledgers exist. What seems to be missing is the habit of bringing these records together and examining them as a single analytic object, rather than as separate filing exercises. The consequence shows up, somewhat predictably, in the fiscal reporting itself: certain line items sit under-spent for much of the year before requiring late re-appropriation, while others run short and need supplementary allocation as the fiscal period closes (Finance Division, 2021). Neither outcome is catastrophic on its own. Together, though, they chip away at the precision of budget planning, and, in principle at least, they are avoidable.
It is worth stepping back for a moment to note that this tension, matching resource allocation to underlying demand under conditions of uncertainty, is hardly unique to Bangladesh's civil service. A substantial body of research on organisational strategy during periods of economic constraint suggests that firms which respond adaptively to shifting internal and external conditions tend to weather downturns more effectively than those that stick rigidly to precedent (Latham, 2009; Latham & Braun, 2011). Studies of software firms navigating the 2008-2010 global recession, for example, found that human-resource and operational decisions grounded in closer monitoring of workforce data tended to outperform decisions made by habit alone (Wickramasinghe & Perera, 2012; Ahmed et al., 2014). Broadly similar patterns have been documented among small firms (Lai et al., 2016) and across Central and Eastern European economies (Burger et al., 2017), while research on information-technology expenditure during the Greek financial crisis reaches a comparable conclusion (Arvanitis & Loukis, 2017). None of these studies concern public payroll specifically, but together they make a modest, transferable point: organisations that ground planning in their own internal data tend to allocate resources more precisely than those that do not, and there is little reason to assume government offices are the exception.
That transferable point has, over the past decade or so, acquired real institutional backing. Governments around the world have shown growing interest in data analytics and artificial intelligence as tools for public decision-making, not as a replacement for administrative judgement but as a way of sharpening it (Duan et al., 2019). The European Commission's own review of artificial intelligence in the public sphere frames this shift fairly explicitly, treating evidence-based, data-informed governance as increasingly close to standard practice rather than novelty (Craglia et al., 2018). The OECD has reached a broadly similar conclusion, cataloguing a wide range of national efforts to embed data-driven methods into public administration (OECD, 2019), while practitioner-facing analyses describe an emerging, if unevenly adopted, model of "AI-augmented government" (Eggers et al., 2017). Bangladesh's public finance system has, so far, largely sat outside that conversation.
Exploratory data analysis (EDA), as formalised by Tukey (1977), offers a natural, and comparatively low-cost, entry point into that conversation. Rather than beginning with a hypothesis and testing it, EDA asks the analyst to look: to plot distributions, to eyeball relationships between variables, to let the data suggest what deserves closer attention (Behrens, 1997). It is, as Tukey himself put it, closer to detective work than to formal inference, and that is arguably its chief virtue in a setting like this one, where almost no prior systematic analysis exists to build on. Framed this way, EDA is best understood as a first step within, rather than apart from, the wider data-science toolkit described in standard references on data mining and predictive analytics (Witten et al., 2017; Siegel, 2013) and in more recent treatments of the mathematical foundations of data science (Blum et al., 2020) - the exploratory stage that, done carefully, tells you what is worth modelling before you attempt to model it. We searched, admittedly not exhaustively, for prior published work applying this kind of analysis to Bangladeshi civil-service data and found very little; most existing treatments of the national budget remain macro-level and descriptive (Ministry of Finance, 2020), rather than grounded in employee-level administrative records.
This paper reports a small, deliberately modest step in that direction. Using records for 200 employees at a single government office, we apply EDA techniques, histograms, scatter plots, and a correlation heat map, to examine how age relates to pay grade, GPF nominee type, loan category, and religious affiliation. We should be upfront about the ambition here, or rather the lack of it: this is not a national survey, and we make no claim that the patterns we find generalise beyond the office studied. What we do claim, more cautiously, is that this kind of office-level exploratory work is feasible, inexpensive, and informative enough to be worth doing more broadly, and that the patterns observed, however local, illustrate the sort of demographic-financial linkage that budget planners could usefully track. The remainder of the paper sets out our data and methods, reports what the exploratory plots showed, and closes with a discussion of what would be needed to extend this approach toward something closer to a genuine national forecasting tool.







