Residual SAC for Retail Replenishment Under a 25 m² Limit
A twice-weekly replenishment system reached a 97.81% fill rate over 52 simulated weeks with no violations of the 25 m² floor-space limit.

Overview
This research project addresses replenishment for a consumer-electronics store with five product categories sharing a hard 25 m² floor-space limit. The system simulates 52 weeks, produces continuous replenishment quantities for each category, and exposes operations through a 3D warehouse dashboard.
Problem
A representative week of demand requires approximately 50.39 m² of incoming goods—more than twice the available space. One weekly delivery therefore cannot provide full coverage regardless of additional training. The controller must balance service, inventory, simulated profit, and shortage risk while respecting the physical capacity at every delivery.
Role
I was a contributing author and research contributor. The paper states that the contributing authors contributed equally; it does not provide a more detailed breakdown of individual implementation responsibilities.
Solution
- Formulated the system as an MDP with a 26-dimensional state and a five-dimensional continuous replenishment action.
- Scheduled deliveries on days 0 and 3 so early-week sales release capacity for a second shipment.
- Combined a demand-coverage baseline with Residual SAC-AutoAlpha; the baseline supplies 90% of the proposal and the actor learns a 10% correction.
- Applied rounding, per-SKU clipping, and uniform area scaling so every executed delivery remains within 25 m².
- Separated seeds for coverage selection, residual-weight selection, training, checkpoint selection, and terminal evaluation.
Technical decisions
The delivery schedule was redesigned before tuning the learning algorithm because the primary bottleneck was physical throughput rather than reward design or training time. Reusing one five-dimensional action for both deliveries keeps the policy compact and interpretable, at the cost of preventing independent control of the two shipments.
The residual policy keeps decisions near a service-feasible region while SAC's twin critics and automatically tuned entropy learn a joint correction across all five products. The replay buffer stores continuous proposals before projection to remain consistent with the policy density, although this creates action aliasing when multiple proposals produce the same executed integer action.
Results
On the held-out 52-week rollout using seed 42, the policy achieved a 97.81% aggregate fill rate and a 96.46% minimum per-SKU fill rate. Mean weekly shortage was 4.75 units, P95 shortage was 11.30 units, and mean end-of-week inventory was 1.60 demand-days. Peak post-delivery footprint was 24.79 m², with zero violations of the 25 m² limit. Simulated accounting profit was VND 72.216 billion.
These figures describe one holdout scenario, not a multi-seed guarantee. Two of 52 weeks exceeded the 30-unit shortage threshold, and the worst week reached 121 shortage units.
Lessons learned
- Physical throughput constraints should be addressed in the operating design before optimizing the learning algorithm.
- A hybrid policy combining an interpretable baseline with a learned residual can stabilize service while retaining joint optimization across SKUs.
- Means and P95 values do not fully capture operational risk; maxima and threshold breaches need separate reporting.
Keep exploring