Loading...
Development of a Reinforcement Learning Algorithm for Optimizing Blood Allocation under Supply and Demand Uncertainty
Meghdadi, Fatemeh | 2026
75
Viewed
- Type of Document: M.Sc. Thesis
- Language: Farsi
- Document No: 58860 (01)
- University: Sharif University of Technology
- Department: Industrial Engineering
- Advisor(s): Radman, Maryam
- Abstract:
- In this study, the inventory allocation problem in the blood supply chain is investigated as a critical, perishable system characterized by a high level of uncertainty. The objective of this research is to develop an efficient decision-making framework for allocating red blood cells to hospitals within a weighted single-objective function that simultaneously considers minimizing shortages, minimizing blood wastage, prioritizing same-type blood fulfillment, and managing substitution costs by reducing the use of low-priority compatible substitutions. To this end, the problem was first formulated in a deterministic environment with a specified planning horizon and solved under four scenarios with increasing problem sizes. The results showed that the CPLEX solver obtained integer optimal solutions in the first three scenarios; however, in the largest scenario, it was unable to solve the problem within a one-hour time limit due to memory limitations. In the scenarios where the optimal solution was available, the greedy algorithm achieved an average relative deviation of 4.67% from the optimal solution, while reducing the solution time by more than 99% on average. Subsequently, to address the inherent uncertainties in blood supply at the distribution center and demand at hospitals, the problem was modeled as a Markov Decision Process, and an Advantage Actor–Critic algorithm was employed to learn the allocation policy. Within this framework, the training process was initially conducted using data extracted from prior studies in order to derive the allocation policy. The results indicated that, during the final 500 episodes, the learned policy achieved an average service level of approximately 84.5%, a wastage rate of less than 0.1%, and a substitution rate of about 10.5%. The learned policy was then implemented and evaluated using data from the Isfahan Provincial Blood Transfusion Center. The findings showed that the policy attained a 92% service level across the entire supply chain while simultaneously reducing the overall wastage rate from 15% to 7.6%. These results demonstrate that the reinforcement learning policy is capable of making dynamic allocation decisions solely through interaction with the environment and without complete knowledge of future demand, thereby providing a practical framework for decision support in healthcare supply chains
- Keywords:
- Blood Supply Chain ; Inventory Management ; Perishable Products ; Deep Reinforcement Learning ; Dynamic Decision Making ; Demand Uncertainty ; Dynamic Allocation
-
محتواي کتاب
- view
