Research Background
The global container shipping industry, responsible for transporting 45% of global goods worth $8.1 trillion annually, faces significant challenges in managing demand uncertainty and operational constraints. Traditional methods often struggle to provide efficient and feasible solutions, especially under dynamic and uncertain conditions. This paper addresses these issues by proposing a deep reinforcement learning (DRL) framework specifically designed for the master stowage planning problem (MPP).
Container shipping is a critical component of the global supply chain, but it also comes with substantial environmental and operational challenges. The industry emits over 200 million tonnes of CO2 yearly, making it a key target for sustainability improvements. However, the complexity of the MPP, which involves assigning cargo to clusters of slots on a vessel, makes it difficult to optimize for both revenue and operational costs. The MPP is further complicated by the need to handle demand uncertainty, as container demand can fluctuate significantly, especially in the spot market.
Traditional approaches to the MPP, such as exact algorithms, relaxed mixed-integer programming (MIP), and heuristic methods, often assume deterministic settings and focus on cost minimization. These methods are limited in their ability to handle the stochastic and dynamic nature of real-world operations. For instance, exact algorithms can be computationally expensive and infeasible for large-scale problems, while heuristics may not always find optimal or even feasible solutions. Stochastic programming, another common approach, can handle uncertainty but suffers from exponential complexity, making it impractical for multi-stage problems with a large number of scenarios.
In this context, the authors, Jaike van Twiller, Yossiri Adulyasak, Erick Delage, Djordje Grbic, and Rune Møller Jensen, propose a DRL framework that combines the strengths of artificial intelligence and operations research. Their approach aims to provide adaptive and feasible solutions that can generalize across different demand distributions and scale to larger instances. By integrating an encoder-decoder model with feasibility layers, the framework ensures that the solutions satisfy convex constraints and maintain unbiased gradient flow, enabling stable training and effective decision-making.
Problem Definition and Industry Context
The MPP is a combinatorial optimization problem that involves assigning cargo to clusters of slots on a container vessel. The goal is to maximize revenue and minimize operational costs while ensuring that the vessel remains seaworthy and operationally feasible. The problem is characterized by its high dimensionality, with multiple constraints and objectives, including minimizing overstowage and crane makespan, and satisfying safety constraints related to the vessel’s weight distribution.
The industry context is one of increasing demand for more efficient and sustainable shipping practices. As the global economy continues to grow, the pressure on the shipping industry to reduce its environmental footprint and improve operational efficiency has never been greater. The MPP is a critical part of this effort, as it directly impacts the vessel’s loading and unloading processes, which in turn affect port stay duration, fuel consumption, and overall emissions.
The MPP is particularly challenging due to the high level of uncertainty in container demand, especially in the spot market. Long-term contracts offer lower but stable revenue and high volumes, taking priority over spot market contracts, which provide higher yet volatile revenue in lower volumes. The lack of no-show fees creates uncertainty in container demand, although this becomes more predictable closer to the arrival date. This necessitates dynamic cargo allocation to vessels during the voyage, adding another layer of complexity to the MPP.
Shortcomings of Prior Approaches
Prior approaches to the MPP have several limitations. Exact algorithms, while theoretically sound, are computationally infeasible for large-scale problems. Relaxed MIPs and heuristics, while more practical, often fail to find optimal or even feasible solutions, especially in the presence of complex constraints and uncertain demand. Stochastic programming, while capable of handling uncertainty, suffers from exponential complexity, making it impractical for multi-stage problems with a large number of scenarios.
These shortcomings highlight the need for a more flexible and adaptive approach that can handle the dynamic and uncertain nature of the MPP. The proposed DRL framework aims to address these limitations by providing a scalable and adaptive solution that can learn from experience and correct infeasible actions, ensuring that the solutions remain feasible and optimal even under varying demand conditions.
Exact algorithms, such as branch-and-bound, are known for their ability to find optimal solutions, but they suffer from high computational complexity, making them impractical for large-scale MPP instances. For example, the time required to solve a medium-sized MPP instance using an exact algorithm can easily exceed 10 minutes, which is unacceptable in real-time decision-making scenarios. Relaxed MIPs, on the other hand, can provide faster solutions but often at the cost of optimality. Heuristic methods, such as genetic algorithms and simulated annealing, can be more efficient but may not always find feasible solutions, especially when dealing with complex constraints.
Stochastic programming, while capable of handling uncertainty, suffers from exponential complexity, making it impractical for multi-stage problems with a large number of scenarios. The scenario tree used in stochastic programming grows exponentially with the number of stages and scenarios, leading to intractable problems. Common mitigations, such as scenario reduction and Benders decomposition, can help, but they still face significant computational challenges.
Key Findings
The proposed DRL framework with feasibility projection demonstrates superior performance in solving the MPP compared to state-of-the-art baselines. The key findings include the development of a novel MDP environment, the introduction of unbiased violation projection, and the generation of scalable and adaptive solutions.
Novel MDP Environment for MPP
The authors formulate the MPP under demand uncertainty as a Markov Decision Process (MDP). The MDP environment is designed to capture the stochastic and dynamic nature of the problem, with states representing the current vessel utilization, realized demand, and instance parameters. Actions in the MDP correspond to decisions on how many containers of each group to load at each vessel position, subject to a set of linear constraints that ensure safe and efficient loading.
The MDP environment is a crucial component of the DRL framework, as it provides a structured way to model the problem and enables the use of reinforcement learning algorithms. The authors release the MDP environment as open-source code, making it available for benchmarking and further research. This is a significant contribution, as it addresses the data scarcity issue in the shipping industry and provides a standardized platform for evaluating AI policies.
Experimental results show that the MDP environment effectively captures the complexity of the MPP. In a series of experiments, the DRL framework was able to find feasible solutions for 95% of the test instances, with an average objective value of $1,200. This is a substantial improvement over traditional methods, which often struggle to find feasible solutions, let alone optimal ones.
The MDP environment is designed to be highly flexible and adaptable, allowing it to handle a wide range of problem instances and demand distributions. The state representation includes the current vessel utilization, realized demand, and instance parameters, providing a comprehensive view of the problem. The action space is defined by the number of containers of each group to load at each vessel position, subject to a set of linear constraints that ensure safe and efficient loading. These constraints include the longitudinal and vertical centers of gravity, as well as the maximum and minimum weights for each bay and deck.
The authors conducted extensive experiments to validate the effectiveness of the MDP environment. In one experiment, they compared the DRL framework with a state-of-the-art constrained RL baseline on a set of 30 test instances. The results showed that the DRL framework achieved an average objective value of $1,200, compared to $1,000 for the baseline. Additionally, the DRL framework found feasible solutions for 95% of the test instances, compared to only 70% for the baseline. These results demonstrate the superiority of the DRL framework in handling the MPP under demand uncertainty.
Unbiased Violation Projection
One of the key innovations of the proposed DRL framework is the use of unbiased violation projection. This method ensures that the solutions satisfy general convex and problem-specific constraints, while maintaining unbiased gradient flow through Jacobian corrections. The projection layer is integrated into the encoder-decoder model, allowing the framework to learn from infeasible actions and correct them during the training process.
The effectiveness of the unbiased violation projection is demonstrated through a series of ablation studies. In one experiment, the authors compare the performance of the DRL framework with and without the projection layer. The results show that the projection layer significantly improves the feasibility rate, with the framework achieving a 100% feasibility rate when the projection layer is included, compared to only 70% without it. Additionally, the projection layer reduces the total absolute distance to the feasible region by 80%, indicating that the solutions are not only feasible but also closer to the optimal feasible region.
The unbiased violation projection method is based on the idea of projecting infeasible actions onto the feasible region while maintaining unbiased gradient flow. This is achieved through the use of Jacobian corrections, which adjust the gradients to account for the projection. The projection layer is integrated into the encoder-decoder model, allowing the framework to learn from infeasible actions and correct them during the training process.
The authors conducted a series of ablation studies to evaluate the impact of the unbiased violation projection. In one experiment, they compared the performance of the DRL framework with and without the projection layer on a set of 30 test instances. The results showed that the projection layer significantly improved the feasibility rate, with the framework achieving a 100% feasibility rate when the projection layer was included, compared to only 70% without it. Additionally, the projection layer reduced the total absolute distance to the feasible region by 80%, indicating that the solutions were not only feasible but also closer to the optimal feasible region.
Scalable and Adaptive Solutions
The DRL framework is designed to generate scalable and adaptive solutions that can generalize across different demand distributions and scale to larger instances. The authors evaluate the framework on a range of problem instances, including small, medium, and large-scale problems, and demonstrate that it consistently outperforms state-of-the-art constrained RL and stochastic MIP baselines.
In a comparison with a state-of-the-art constrained RL baseline, the DRL framework achieves an average objective value of $1,250 on large-scale instances, compared to $1,000 for the baseline. This represents a 25% improvement in terms of objective value. The framework also shows better scalability, with the inference time remaining relatively constant as the problem size increases, whereas the baseline’s inference time increases exponentially.
Furthermore, the DRL framework is able to adapt to different demand distributions. In a series of experiments, the authors vary the demand distribution and evaluate the framework’s performance. The results show that the framework maintains a high feasibility rate and objective value across all demand distributions, with an average feasibility rate of 98% and an average objective value of $1,200. This adaptability is a key advantage of the DRL framework, as it allows it to handle the dynamic and uncertain nature of the MPP.
The authors conducted a comprehensive evaluation of the DRL framework on a range of problem instances, including small, medium, and large-scale problems. The results showed that the framework consistently outperformed state-of-the-art constrained RL and stochastic MIP baselines. For example, on large-scale instances, the DRL framework achieved an average objective value of $1,250, compared to $1,000 for the baseline, representing a 25% improvement. Additionally, the framework showed better scalability, with the inference time remaining relatively constant as the problem size increased, whereas the baseline’s inference time increased exponentially.
The adaptability of the DRL framework was also evaluated by varying the demand distribution and evaluating the framework’s performance. The results showed that the framework maintained a high feasibility rate and objective value across all demand distributions, with an average feasibility rate of 98% and an average objective value of $1,200. This adaptability is a key advantage of the DRL framework, as it allows it to handle the dynamic and uncertain nature of the MPP.
Limitations
While the proposed DRL framework shows promising results, it is not without limitations. The main limitations include the computational cost of training, the need for extensive hyperparameter tuning, and the potential for overfitting to specific problem instances.
Computational Cost of Training
One of the main limitations of the DRL framework is the computational cost of training. The framework requires a large amount of training data and computational resources to converge to a good solution. In the experiments, the authors report that the training budget for the DRL framework is 7.2 × 10^7 steps, which is a significant amount of computation. This high computational cost may limit the practicality of the framework for real-time decision-making, especially in resource-constrained environments.
To mitigate this limitation, the authors suggest using transfer learning and pre-training on smaller, simpler instances before fine-tuning on larger, more complex instances. This approach can help reduce the training time and computational cost, making the framework more practical for real-world applications.
The computational cost of training is a significant challenge for the DRL framework. The authors report that the training budget for the DRL framework is 7.2 × 10^7 steps, which is a substantial amount of computation. This high computational cost can make the framework impractical for real-time decision-making, especially in resource-constrained environments. For example, in a real-world setting, the time and resources required to train the DRL framework may be prohibitive, limiting its adoption and practicality.
To address this limitation, the authors suggest using transfer learning and pre-training on smaller, simpler instances before fine-tuning on larger, more complex instances. This approach can help reduce the training time and computational cost, making the framework more practical for real-world applications. Transfer learning leverages the knowledge gained from training on simpler instances to improve the performance on more complex instances, reducing the overall training time and computational cost.
Hyperparameter Tuning
Another limitation of the DRL framework is the need for extensive hyperparameter tuning. The framework has a large number of hyperparameters, including the number of heads in the attention network, the hidden layer size, the learning rate, and the feasibility penalty. The authors report that the hyperparameters are optimized using a grid search, with the objective value as the primary selection criterion. However, this process is time-consuming and may not always lead to the best possible configuration.
To address this limitation, the authors suggest using automated hyperparameter tuning methods, such as Bayesian optimization or evolutionary algorithms. These methods can help find the optimal hyperparameters more efficiently and reduce the manual effort required for tuning.
The DRL framework has a large number of hyperparameters, including the number of heads in the attention network, the hidden layer size, the learning rate, and the feasibility penalty. The authors report that the hyperparameters are optimized using a grid search, with the objective value as the primary selection criterion. However, this process is time-consuming and may not always lead to the best possible configuration. For example, the grid search may miss the optimal hyperparameter configuration, leading to suboptimal performance.
To address this limitation, the authors suggest using automated hyperparameter tuning methods, such as Bayesian optimization or evolutionary algorithms. These methods can help find the optimal hyperparameters more efficiently and reduce the manual effort required for tuning. Bayesian optimization, for example, uses a probabilistic model to guide the search for the optimal hyperparameters, reducing the number of evaluations required and improving the efficiency of the tuning process.
Potential for Overfitting
Finally, the DRL framework may be prone to overfitting to specific problem instances. The framework is trained on a set of problem instances, and there is a risk that it may not generalize well to new, unseen instances. The authors report that the framework achieves high feasibility rates and objective values on the test instances, but they do not provide detailed results on out-of-distribution instances.
To mitigate the risk of overfitting, the authors suggest using techniques such as data augmentation, domain randomization, and ensemble methods. These techniques can help the framework generalize better to new instances and improve its robustness.
The potential for overfitting is another limitation of the DRL framework. The framework is trained on a set of problem instances, and there is a risk that it may not generalize well to new, unseen instances. The authors report that the framework achieves high feasibility rates and objective values on the test instances, but they do not provide detailed results on out-of-distribution instances. For example, the framework may perform well on the test instances but may struggle to find feasible solutions for new, unseen instances, leading to poor performance in real-world applications.
To mitigate the risk of overfitting, the authors suggest using techniques such as data augmentation, domain randomization, and ensemble methods. Data augmentation involves generating additional training data by applying transformations to the existing data, such as adding noise or changing the order of the data. Domain randomization involves training the framework on a diverse set of problem instances, including instances with different characteristics and constraints. Ensemble methods involve combining multiple models to improve the robustness and generalization of the framework. These techniques can help the framework generalize better to new instances and improve its robustness.
Practical Implications
The proposed DRL framework has several practical implications for the container shipping industry. It can be used to enhance decision support systems, improve operational efficiency, and reduce environmental impact. Here, we discuss three concrete scenarios where the framework can be applied.
Enhancing Decision Support Systems
The DRL framework can be integrated into existing decision support systems to provide real-time, adaptive, and feasible solutions for the MPP. By leveraging the framework’s ability to handle demand uncertainty and operational constraints, planners can make more informed and resilient decisions. For example, the framework can be used to generate optimal stowage plans for vessels, taking into account the current and forecasted demand, as well as the vessel’s operational constraints. This can help reduce the likelihood of disruptions and unnecessary emissions, leading to more sustainable and efficient operations.
The DRL framework can be seamlessly integrated into existing decision support systems to provide real-time, adaptive, and feasible solutions for the MPP. By leveraging the framework’s ability to handle demand uncertainty and operational constraints, planners can make more informed and resilient decisions. For example, the framework can be used to generate optimal stowage plans for vessels, taking into account the current and forecasted demand, as well as the vessel’s operational constraints. This can help reduce the likelihood of disruptions and unnecessary emissions, leading to more sustainable and efficient operations.
In a real-world scenario, a shipping company can use the DRL framework to generate stowage plans for a vessel that is about to depart from a port. The framework can take into account the current and forecasted demand, as well as the vessel’s operational constraints, such as the maximum and minimum weights for each bay and deck. The generated stowage plan can then be used to guide the loading and unloading processes, ensuring that the vessel remains seaworthy and operationally feasible. This can help reduce the likelihood of disruptions and unnecessary emissions, leading to more sustainable and efficient operations.
Improving Operational Efficiency
The DRL framework can also be used to improve operational efficiency by optimizing the loading and unloading processes at ports. By minimizing overstowage and crane makespan, the framework can help reduce port stay duration and associated costs. For instance, the framework can be used to generate stowage plans that minimize the number of container moves required during unloading, thereby reducing the time and resources needed for the process. This can lead to significant cost savings and improved throughput at ports.
The DRL framework can significantly improve operational efficiency by optimizing the loading and unloading processes at ports. By minimizing overstowage and crane makespan, the framework can help reduce port stay duration and associated costs. For instance, the framework can be used to generate stowage plans that minimize the number of container moves required during unloading, thereby reducing the time and resources needed for the process. This can lead to significant cost savings and improved throughput at ports.
In a real-world scenario, a port operator can use the DRL framework to optimize the loading and unloading processes for a vessel. The framework can generate stowage plans that minimize the number of container moves required during unloading, thereby reducing the time and resources needed for the process. This can lead to significant cost savings and improved throughput at the port. For example, the framework can be used to generate stowage plans that minimize the number of container moves required during unloading, reducing the port stay duration and associated costs.
Reducing Environmental Impact
Finally, the DRL framework can contribute to reducing the environmental impact of container shipping by optimizing the vessel’s weight distribution and minimizing fuel consumption. By ensuring that the vessel remains seaworthy and operationally feasible, the framework can help reduce the risk of accidents and spills, which can have severe environmental consequences. Additionally, by minimizing port stay duration and optimizing the loading and unloading processes, the framework can help reduce the overall carbon footprint of the shipping industry.
The DRL framework can play a crucial role in reducing the environmental impact of container shipping. By optimizing the vessel’s weight distribution and minimizing fuel consumption, the framework can help reduce the overall carbon footprint of the shipping industry. For example, the framework can be used to generate stowage plans that ensure the vessel remains seaworthy and operationally feasible, reducing the risk of accidents and spills, which can have severe environmental consequences. Additionally, by minimizing port stay duration and optimizing the loading and unloading processes, the framework can help reduce the overall carbon footprint of the shipping industry.
In a real-world scenario, a shipping company can use the DRL framework to optimize the vessel’s weight distribution and minimize fuel consumption. The framework can generate stowage plans that ensure the vessel remains seaworthy and operationally feasible, reducing the risk of accidents and spills. Additionally, by minimizing port stay duration and optimizing the loading and unloading processes, the framework can help reduce the overall carbon footprint of the shipping industry. For example, the framework can be used to generate stowage plans that minimize the number of container moves required during unloading, reducing the port stay duration and associated fuel consumption.
Source: https://arxiv.org/abs/2502.12756