Journal article

Solving an Order Batching and Sequencing Problem with Reinforcement Learning

Abstract

The purpose of this research is to determine whether a DRL solution would be a suitable solution for the OBSP problem and to compare it with traditional methods. For this purpose, models trained utilizing the PPO algorithm were tested in a complex and realistic warehouse environment, and an attempt was made to measure whether a strategy was developed to decrease the number of orders being late. A heuristic method was also applied and the results were compared on the same environment and data. The results showed that DRL approach that combines heuristics with the PPO algorithm outperforms the heuristics in minimizing the tardy order percentage in all tested scenarios.

Keywords

Pekiştirmeli ÖğrenmeSipariş Gruplama ve Sıralama ProblemiYakın Politika OptimizasyonuDepo Optimizasyonu

122 views · 108 downloads