An interpretable agent integrating offline imitation and reinforcement learning for cotton irrigation scheduling

Liu, Kejun , Xu, Aojie , Liu, Guilin

2026-08-01 EUROPEAN JOURNAL OF AGRONOMY 2026   179(卷), null(期), (null页)

查看原文

Efficient irrigation scheduling in arid regions requires adapting to dynamic crop-soil-weather interactions while balancing yield and irrigation amount. In practice, irrigation scheduling often relies on agronomic heuristics or process-based crop models, which do not directly provide sequential control policies but still encode agronomic priors and demonstrations, thereby motivating offline imitation learning (IL). Reinforcement learning (RL) provides a framework for irrigation control, yet many studies depend on expensive online exploration and dense sensing impractical on farms. This study therefore developed an offline imitation-reinforcement learning agent that learns cotton irrigation policies from expert demonstrations generated from farm logs using a GA-driven, AquaCrop-calibrated simulator. The policy uses behavior cloning (BC) pretraining and TD3+BC (TD3: Twin Delayed Deep Deterministic Policy Gradient) offline fine-tuning, enhanced by (i) temporally structured state encoder with attention-based fusion, (ii) marginal-yield Shapley reward-shaping, and (iii) value-aware Soft Qfilter. Simulation results showed an improved water-yield trade-off: the agent reduced irrigation by 3-24% compared to standard TD3 while maintaining comparable or higher yields than BC and TD3. Matched ablations identified the Shapley reward and Soft Q-filter as key drivers of stable policy. Under state-space perturbations, the agent maintained low variability (irrigation CV = 2.146%, yield CV = 0.160%), and in cross-site validation retained comparable yield while enhancing water productivity. Shapley Additive Explanations (SHAP)-based interpretability indicates agronomically consistent logic combining soil-moisture thresholds with short-term weather adjustments. Overall, results suggest offline imitation-reinforcement learning can learn physically consistent and interpretable irrigation policies from limited farm logs, supporting scalable decision-making.