4.1 The functional landscape of AI/ML servers for ADMET prediction
Read across a decade, the platform literature describes a clear trajectory: single-property QSAR calculators have given way to multi-task, deep-learning prediction engines (Fu et al., 2024). The platforms catalogued in Table 1 are not variations on one design. They differ in what they represent, what they were trained on, and consequently in what they are good for — a point that matters because they are routinely used in combination, as Figure 3 illustrates, and combining tools with correlated blind spots buys less than it appears to.
The first tier remains physicochemical. SwissADME applies hybrid rule-based and empirical engines to estimate human intestinal absorption, passive blood–brain barrier permeability through the BOILED-Egg model, and P-glycoprotein substrate or inhibitor status (Daina et al., 2017). It also flags potential drug–drug interaction liabilities by predicting inhibition across five CYP isoforms — CYP1A2, CYP2C19, CYP2C9, CYP2D6 and CYP3A4 — which is directly relevant in malaria-endemic settings where antiretroviral and antitubercular co-medication is common (Daina et al., 2017; Wu et al., 2020).
Moving beyond two-dimensional descriptors, pkCSM and Deep-PK encode structures as distance-based graph signatures (Pires et al., 2015). This representation supports estimation of steady-state volume of distribution, total clearance, blood–brain barrier partition coefficients, and specific toxicity endpoints including AMES mutagenicity, hERG blockade and drug-induced liver injury (Pires et al., 2015; Amen et al., 2025). ADMETlab 3.0 pushes coverage furthest with multi-task graph neural networks spanning more than 100 endpoints (Fu et al., 2024; Dong et al., 2018).
For toxicity proper, ProTox-II integrates toxicophore fragment propensities with random forest and support vector machine classifiers trained on acute oral toxicity data, assigning molecules to classes I through VI on the basis of predicted LD50 while simultaneously screening hepatotoxicity, carcinogenicity, mutagenicity and cytotoxicity (Banerjee et al., 2018). eToxPred applies extra trees and gradient boosting to return a normalised toxicity probability between 0 and 1 together with a synthetic accessibility score between 1 and 10 (Pu et al., 2019). That pairing deserves more attention than it usually receives: in a resource-constrained antimalarial programme, a safe compound requiring a fourteen-step synthesis is often less useful than a marginally less attractive one that can be made in four.
Operational integration is exemplified by AutoFilter (Ramadoss & Singh, 2025). Chaining Lipinski, Veber and PAINS filtration with AutoDock Vina docking, SwissADME profiling, eToxPred toxicity scoring and 100-ns GROMACS molecular dynamics, the framework processed 2.4 million ChEMBL bioactives and reportedly reduced compound selection timelines and costs by approximately 50% (Ramadoss & Singh, 2025). The saving is an internal estimate and has not been independently replicated, but the architectural point stands: sequencing cheap filters ahead of expensive simulation is what makes screening at this scale tractable at all (Figure 2).
4.2 Target-specific lead prioritisation across P. falciparum enzymatic pathways
Table 2 assembles the four antimalarial campaigns in a common format. Each applied a broadly similar cascade — library, filtration, docking, ADMET gate, dynamics — yet they differ enough in execution that their outputs are not directly comparable, and we treat them individually for that reason.
4.2.1 P. falciparum L-lactate dehydrogenase (PfLDH)
PfLDH (PDB ID: 1LDG) sustains glycolytic energy production during the asexual blood stages on which clinical malaria depends (Ibrahim et al., 2026). Ibrahim et al. (2026) screened 24,316 natural product compounds from Traditional Chinese Medicine repositories through a hierarchical Glide protocol — high-throughput virtual screening, then standard precision, then extra precision — arriving at three leads: ZINC70450989, ZINC85506851 and ZINC85506187 (Table 2).
The docking figures are striking. XP scores of −14.30, −11.10 and −11.05 kcal/mol compare against artemisinin at −5.23, piperaquine at −4.285 and lumefantrine at −4.049 kcal/mol (Ibrahim et al., 2026). MM-GBSA calculations reinforced the ranking, with ZINC70450989 reaching ΔGbind = −91.63 kcal/mol against artemisinin’s −36.59 kcal/mol. A margin of that size invites caution rather than celebration — end-state free energy methods are known to exaggerate differences between strong and weak binders — but the direction of the result was consistent across three independent methods, which is the more meaningful observation.
More persuasive, to our reading, is the resistance work. In silico mutagenesis at Arg109, Lys102, Asp168, His195, Arg171, Pro246 and Pro250 showed that the leads retained substantial affinity against mutated receptors, with ZINC70450989 scoring −12.10 kcal/mol and ΔGbind = −75.29 kcal/mol in the mutant background (Ibrahim et al., 2026). Since resistance is the reason these campaigns exist, testing against anticipated escape mutations before committing to synthesis is a design choice other programmes would do well to adopt.
Across 100-ns simulations the PfLDH Cα backbone in complex with the leads held equilibrium with mean RMSD below 2.0 Å, whereas artemisinin complexes showed excursions to 3.36 Å (Ibrahim et al., 2026). Deep-PK and ProTox-II profiling then confirmed high intestinal absorption, non-penetrant blood–brain barrier behaviour (log BB < −2.5), and toxicity classes 5 and 6 with LD50 between 3,800 and 7,500 mg/kg, with no mutagenic, cytotoxic or carcinogenic flags (Ibrahim et al., 2026; Banerjee et al., 2018). The low CNS penetration is a deliberate design objective for a non-CNS antimalarial, not an incidental finding (Table 4).
4.2.2 Apicoplast DNA polymerase (apPOL)
Ramadoss and Singh (2025) applied AutoFilter to 2.4 million ChEMBL compounds against apPOL, the enzyme responsible for replicating the apicoplast genome (Table 2). Chemical rule filtration and AutoDock Vina grid docking reduced the set to 3,194 active ligands, which were then passed through SwissADME and eToxPred. Five compounds, L1 to L5, were prioritised on docking scores below −9.8 kcal/mol, favourable drug-likeness, low predicted toxicity and accessible synthetic routes (Ramadoss & Singh, 2025; Pu et al., 2019). Subsequent 100-ns CHARMM36 simulations in GROMACS confirmed complex stability.
The distinguishing feature of this campaign is what came next. Preliminary in vitro growth inhibition assays validated the antimalarial activity of the top-ranked lead, with over 50% parasite growth inhibition reported (Ramadoss & Singh, 2025). Among the studies reviewed here, this is the only one to close the loop experimentally, and it therefore carries evidential weight the purely computational campaigns cannot claim — a point we return to in Section 5.
4.2.3 Fatty acid biosynthesis: ACC and FabI
Salifu et al. (2023) targeted two FAS-II enzymes: biotin acetyl-CoA carboxylase (PDB: 1W96) and enoyl-acyl carrier protein reductase (PDB: 3F4B). Rather than docking a library directly, they generated per-residue energy decomposition (PRED) pharmacophores from short MD trajectories of the soraphen A and triclosan reference complexes, then used these to screen 21.7 million ZINC compounds (Salifu et al., 2023). PyRx AutoDock Vina docking and SwissADME profiling yielded nine hits per target.
ProTox-II screening narrowed the ACC set to three non-toxic leads — ZINC38980461, ZINC05378039 and ZINC15772056 — falling into toxicity class IV with LD50 up to 1,500 mg/kg and no mutagenic or cytotoxic flags (Salifu et al., 2023; Banerjee et al., 2018). For FabI, six leads were retained, of which ZINC94919772 showed a class V profile at LD50 = 4,000 mg/kg. Across 100-ns Amber trajectories all complexes remained within the 1.0–3.0 Å RMSD window, and MM/PBSA calculations gave ΔGbind of −36.36 kcal/mol for ZINC38980461 against ACC and −41.42 kcal/mol for the best FabI complex (Salifu et al., 2023).
Two observations follow. The PRED approach, by deriving pharmacophore features from energetic contributions sampled dynamically rather than from a single crystal snapshot, addresses one of docking’s standing weaknesses — receptor rigidity (Table 3). And the toxicity classes obtained here (IV and V) are demonstrably less favourable than those reported for the PfLDH leads (5 and 6), which is a useful reminder that ADMET screening discriminates between campaigns as well as within them.
Table 3. Methodological classes used in computational toxicology and CADD pipelines, with reported performance and known constraints. Each row describes the mechanistic basis of a modelling approach, its role within an antimalarial screening workflow, the accuracy range reported in the source literature, and the limitations that motivate its combination with the other approaches. Performance figures derive from heterogeneous benchmarks published by different groups and should not be read as head-to-head comparisons under matched conditions. The sequence of rows corresponds to the layered pipeline shown in Figure 4.
|
Modelling approach
|
Mechanistic / algorithmic basis
|
Application in the ADMET–CADD workflow
|
Reported accuracy
|
Major strengths
|
Key limitations
|
Reference
|
|
Structure-based virtual screening and docking
|
Rigid or flexible receptor–ligand pose prediction using empirical scoring functions (Glide XP, AutoDock Vina, GOLD)
|
High-throughput screening of chemical libraries against 3D crystal structures to predict binding poses and affinities
|
85–95% pose prediction reliability
|
Rapid prioritisation of very large virtual libraries; atomistic interaction mapping (H-bonds, salt bridges)
|
Neglects protein backbone dynamics unless induced-fit or ensemble docking is used; scoring functions generate false positives
|
Panda et al. (2025); Zhang et al. (2025)
|
|
Quantitative structure–activity relationship (QSAR)
|
Correlates 1D/2D/3D descriptors (log P, TPSA, fingerprints) with bioactivity or toxicity via statistical ML (MLR, PLS, SVM, RF, XGBoost)
|
Efficacy evaluation, pIC50 prediction, ADMET property estimation and rational lead modification
|
70–85% classification / regression accuracy
|
Fast execution at low computational cost; descriptor coefficients indicate what to modify
|
Bounded by the applicability domain of the training set; vulnerable to activity cliffs
|
Panda et al. (2025); Zhang et al. (2025); Wu et al. (2020)
|
|
Molecular dynamics simulation with MM-GBSA / MM-PBSA
|
Newtonian sampling of atomic trajectories over 100–1,000 ns in explicit solvent, with end-state free energy calculation
|
Assessment of conformational stability (RMSD, RMSF, radius of gyration), contact persistence, cryptic pocket discovery, ΔGbind
|
~92% conformational agreement
|
Captures target flexibility, explicit solvation and thermodynamic binding stability over time
|
High computational cost requiring GPU acceleration; accessible timescales may miss slow transitions
|
Panda et al. (2025); Zhang et al. (2025); Salifu et al. (2023)
|
|
Physiologically based pharmacokinetic (PBPK) modelling
|
Systems of differential equations representing absorption, tissue distribution, hepatic metabolism and renal elimination
|
Prediction of human concentration–time profiles (Cmax, AUC, t1/2, clearance) from preclinical in vitro data
|
~65–90% precision for human PK extrapolation
|
Mechanistic representation of human physiology; supports dose optimisation and DDI risk assessment
|
Requires extensive physiological and tissue-partition parameters that are often difficult to measure
|
Panda et al. (2025); Wu et al. (2020)
|
|
Deep learning and graph neural networks
|
Multi-layer networks (DNN, CNN, GNN, D-MPNN, transformers) processing molecular graphs, SMILES and multi-task bioassay data
|
Multi-endpoint ADMET prediction, automated toxicity profiling (DeepTox, ProTox-II, ADMETlab 3.0), structural alert detection
|
85–95% accuracy across multi-task toxicity benchmarks
|
Learns complex non-linear representations automatically; scales to massive chemical datasets
|
Black-box interpretability deficit; requires large, well-curated training sets to avoid overfitting
|
Zhang et al. (2025); Mayr et al. (2016); Fu et al. (2024)
|
|
Per-residue energy decomposition (PRED) pharmacophores
|
Combines short MD trajectories with MM/PBSA energetic decomposition to isolate key residue hotspots
|
Generation of tailored 3D spatial pharmacophores (HBD, HBA, hydrophobic centres) for high-precision virtual screening
|
High enrichment factor in database screening
|
Pharmacophores reflect dynamic energetic contributions rather than a single static crystal snapshot
|
Requires prior MD trajectory generation, which is costly for complex target systems
|
Salifu et al. (2023)
|
Table 4. Major ADMET endpoints, their clinical significance, predictive tools and threshold values applied during antimalarial lead optimisation. For each endpoint the table gives the biological rationale, the servers commonly used to predict it, the threshold or desirable range applied when triaging candidates, and the development consequence of failing that threshold. Threshold values are taken from the decision criteria stated in the source studies rather than imposed here, and ranges are shown where sources disagreed. All leads listed in Table 2 satisfied these criteria, which reflects selective reporting of successful candidates as much as filter performance.
|
ADMET category
|
Endpoint / parameter
|
Biological and clinical significance
|
Predictive tools
|
Threshold / desirable range
|
Failure risk and attrition consequence
|
Reference
|
|
Absorption
|
Human intestinal absorption (HIA) and Caco-2 permeability
|
Governs oral absorption across the gastrointestinal epithelium into systemic circulation
|
SwissADME, pkCSM, ADMETlab 3.0, QikProp
|
HIA > 30% (high absorption > 80%); Caco-2 log Papp > −6.0 cm/s
|
Poor systemic exposure, high oral dosing requirements, erratic absorption
|
Daina et al. (2017); Ibrahim et al. (2026)
|
|
Distribution
|
Blood–brain barrier permeability (log BB) and CNS permeability (log PS)
|
Regulates entry into the central nervous system; critical for non-CNS drugs to avoid neurotoxicity
|
SwissADME, pkCSM, ADMETlab 3.0, QikProp
|
Non-CNS antimalarials: log BB < 0.3; log PS > −2.0 (CNS non-penetrable)
|
Centrally mediated sedation, neurotoxicity, unwanted CNS adverse effects
|
Ibrahim et al. (2024); Amen et al. (2025)
|
|
Distribution
|
Plasma protein binding (%) and volume of distribution (Vss)
|
Determines the unbound fraction available for target engagement versus tissue sequestration
|
pkCSM, Deep-PK, ADMETlab 3.0
|
Moderate binding; Vss > 0.45 L/kg with adequate free fraction
|
Tissue accumulation, low free plasma concentration, unpredictable displacement
|
Amen et al. (2025)
|
|
Metabolism
|
CYP inhibition (CYP1A2, CYP2C19, CYP2D6, CYP3A4)
|
Hepatic biotransformation; CYP3A4 metabolises ~50% of marketed drugs, CYP2D6 ~25%
|
SwissADME, pkCSM, ADMETlab 3.0, SMARTCyp
|
Non-inhibitor status for major isoforms
|
Severe drug–drug interactions, toxicity from co-administered drug accumulation, rapid clearance
|
Amen et al. (2025); Daina et al. (2017)
|
|
Excretion
|
Total clearance (CLtot) and elimination half-life (t1/2)
|
Defines hepatic and renal elimination rate, governing dosing schedule and steady-state exposure
|
pkCSM, Deep-PK, ADMETlab 3.0, PBPK models
|
Clearance 1–100 mL/min/kg; half-life matched to once-daily or single-dose regimens
|
Accumulation toxicity if clearance is too low; loss of efficacy if clearance is excessive
|
Amen et al. (2025)
|
|
Toxicity
|
hERG K⁺ channel inhibition
|
Blockade prolongs the QT interval and can precipitate fatal torsades de pointes
|
pkCSM, Deep-PK, ProTox-II, ADMETlab 3.0, eToxPred
|
Non-blocker; probability < 0.4 or IC50 > 10 μM
|
Leading cause of post-market withdrawal and late-stage clinical attrition
|
Muster et al. (2008); Amen et al. (2025)
|
|
Toxicity
|
Acute oral toxicity (LD50) and AMES mutagenicity
|
Assesses acute systemic lethal dose and potential for genetic damage or carcinogenicity
|
ProTox-II, pkCSM, eToxPred
|
AMES inactive; LD50 > 2,000 mg/kg (ProTox-II class 5 or 6)
|
Genotoxicity, carcinogenicity, acute organ failure, early termination in preclinical safety studies
|
Banerjee et al. (2018); Salifu et al. (2023); Ibrahim et al. (2026)
|
4.2.4 Amodiaquine 4-aminoquinoline derivatives
Ibrahim et al. (2024) took the optimisation route rather than the screening route, building QSAR models for 22 amodiaquine derivatives assayed against the chloroquine-sensitive 3D7 strain (Table 2). Geometry optimisation at DFT/B3LYP/6-31G* in Spartan 14 and PaDEL descriptor calculation yielded a four-descriptor genetic function algorithm model:
pIC50 = −10.8052(MATS5p) − 13.0551(MATS6i) − 10.0899(SpMax1_Bhm) + 0.0907(RDF45v) + 48.5776
The model’s statistics were sound for a set of this size (Ntrain = 16, R² = 0.9243, F = 33.5615, Q²cv = 0.8603, R²test = 0.6640), though the gap between internal Q² and external R² — 0.86 against 0.66 — is the familiar signature of a model that generalises less well than cross-validation suggests (Ibrahim et al., 2024). Using compound A-01 (pIC50 = 9.491) as template, thirteen analogues were designed by shortening substituent chains and introducing cyclic saturation.
Derivative ac — 4-((7-chloroquinolin-4-yl)amino)-2-(cyclohexyl(4-(pyridin-2-yl)piperazin-1-yl)methyl)phenol — returned a predicted pIC50 of 11.4918, against amodiaquine at 8.668 and chloroquine at 8.111 (Ibrahim et al., 2024). We note that the source reports 9.491 for this compound in its summary table, and have retained the value from the primary results narrative while flagging the inconsistency. Either way, the prediction lies outside the potency range of the training set, and extrapolated activity values of this kind should be treated as hypotheses for synthesis rather than as results.
SwissADME and pkCSM profiling confirmed compliance with Lipinski and Veber criteria (TPSA = 64.52 Ų, nRotB = 6), intestinal absorption above 80%, limited blood–brain barrier permeability (log BB < 0.3), CNS non-penetrability (log PS = −2.0), oral rat LD50 of 2.433–3.173 mol/kg, and no skin sensitisation flag (Ibrahim et al., 2024).
4.3 Methodological synergy across CADD pipelines
Table 3 compares the methodological classes deployed across these campaigns, and the pattern in Figure 4 is consistent: single-method screening has largely been abandoned in favour of layered workflows in which each stage compensates for the preceding stage’s characteristic weakness.
Structure-based virtual screening and docking serves as the primary filter, processing million-compound libraries against crystal structures with reported pose reliability in the 85–95% range (Panda et al., 2025; Zhang et al., 2025). Its known failure mode — neglect of backbone flexibility, and scoring functions prone to false positives — is precisely what the later stages exist to catch.
QSAR and ML regression establishes quantitative descriptor–activity relationships, typically achieving 70–85% accuracy, and uniquely among these methods tells the chemist what to change rather than merely which compound to keep (Ibrahim et al., 2024; Wu et al., 2020).
Molecular dynamics tests whether a docked pose survives contact with thermal motion and solvent over 100-ns production runs, with RMSD below 2.0 Å and RMSF below 1.5 Å the conventional stability thresholds (Ibrahim et al., 2026; Panda et al., 2025). Reported conformational agreement of around 92% is good; the cost, requiring GPU acceleration, is what keeps this stage late in the cascade.
End-state free energy calculations (MM-GBSA, MM-PBSA) refine docking scores by incorporating solvation and continuum electrostatics, yielding ΔGbind values that track measured potency more closely than raw docking scores do (Salifu et al., 2023; Ibrahim et al., 2026).
PBPK and systems pharmacokinetics translate endpoint predictions into mechanistic differential-equation models capable of simulating human plasma concentration–time profiles — Cmax, AUC, t1/2 — and thereby informing dose selection (Panda et al., 2025; Wu et al., 2020). Notably, none of the four antimalarial campaigns reviewed here reached this stage. That absence is discussed in Section 5.3.
4.4 Endpoint thresholds governing lead selection
Table 4 consolidates the criteria applied during antimalarial lead triage. These thresholds are conventions rather than physical constants, and they vary somewhat across the source studies, but the consensus ranges are reasonably stable.
Absorption. Oral candidates are expected to show human intestinal absorption above 30%, preferably above 80%, with Caco-2 permeability log Papp > −6.0 cm/s (Daina et al., 2017; Ibrahim et al., 2026). For a disease treated with oral therapy under field conditions, these are not negotiable margins.
Distribution and CNS exposure. Non-CNS antimalarials should maintain log BB < 0.3 and log PS > −2.0 to limit neurological adverse effects, with steady-state volume of distribution above 0.45 L/kg indicating balanced tissue distribution without excessive lipophilic sequestration (Ibrahim et al., 2024; Amen et al., 2025).
Metabolism and clearance. Candidates should be non-inhibitors of CYP3A4, CYP2D6 and CYP2C19 to avoid clinically relevant drug–drug interactions, with clearance and half-life tuned to support once-daily or single-dose regimens (Amen et al., 2025; Daina et al., 2017).
Cardiotoxicity and acute lethality. hERG blockade must be avoided, with inhibition probability below 0.4 or IC50 above 10 μM (Muster et al., 2008; Amen et al., 2025). Candidates should test inactive in AMES and fall within ProTox-II classes 4, 5 or 6 — LD50 from 300 to above 5,000 mg/kg — before any in vivo commitment (Banerjee et al., 2018; Salifu et al., 2023). Across the campaigns in Table 2, every prioritised lead satisfied these criteria, which is encouraging but also, inevitably, a selection effect: compounds that failed the gate were not reported.