Skip to main content

Joint scheduling training

Select the DFJSP-T graph policy and rolling joint scheduling in the experiment workbench. Models produced by training can be used for testing and factory replay.

The training and parameter-exploration templates use ctde_ppo version 0.3.0. Existing experiment definitions do not automatically adopt new templates. After upgrading the system, reload a template or reselect the algorithm and review parameters.

What the model decides​

The policy uses jobs, machines, vehicles, roads, and current disturbances to select joint production and transport plans. It replans when events change or a decision interval is reached, while updating traffic routes.

Built-in PIBT executes vehicle movement. Training and model testing do not require an additional MAPF algorithm. Existing transport commitments continue while candidate search runs. Domain router and assigner options apply to rule algorithms that only output production dispatch.

Material stations prefer explicit material_handling_config entries, then topology.depot and topology.product. Otherwise, connected traversable non-machine cells supply defaults. Stations, machines, and AGVs must belong to one four-neighbor connected region. Every AGV must have a distinct initial position.

Parameters​

ParameterDefaultPurpose
candidate_limit24Maximum candidate plans per planning round
scenario_count8Predicted uncertainty scenarios used to compare plans
routing_horizon24Traffic prediction horizon in steps
decision_interval5Replanning check interval during steady operation
search_seconds00 uses the normal incremental budget; positive values further limit per-step search time
inference_budget_ms300Policy inference budget per simulation step; search can span steps
sequence_length32Consecutive decision sequence length used for learning
graph_batch_size16Graphs encoded together; independent of PPO minibatch size
input_cache_mb256Graph-input cache cap in MiB; 0 disables it
num_envs4Concurrent sampling environments
dynamic_samplingfalseDynamically assign episodes, discard incomplete ones at synchronization, and prioritize their restart
evaluation_num_envs2Concurrent validation episodes
deviceautoCUDA device for network training, such as cuda:0

More candidates and scenarios can increase the steps needed to complete a planning round. The algorithm can commit an evaluated plan and continue improving while vehicles follow valid routes. Time budgets depend on machine load; use recorded actions for exact playback.

Training progress and save intervals count simulation steps, while learning samples correspond to scheduling decisions. planning_decisions records decisions. inference_seconds, inference_max_seconds, and inference_over_budget_steps record total inference duration, maximum per-step duration, and over-budget steps.

Sampling, updates, and validation​

Sampling and validation use separate CPU cores; network training uses GPU. After the first sampling batch, the next sampling batch overlaps the previous batch's update. Both complete before synchronization. Each batch uses fixed policy weights; data is at most one update behind when training starts.

With dynamic sampling disabled, each worker finishes one episode per round. Enabled workers take more instances while waiting for others; synchronization discards incomplete episodes and prioritizes restarting their original instance/seed next round. Step-limited episodes count as ended, not necessarily completed factories.

Monitoring shows sampling, training, and validation independently. Training updates report epochs, minibatches, optimizer updates, and elapsed time at computation-block boundaries. Synchronization records show retained/discarded trajectories and waiting time.

Graph batching and input caching reduce update overhead. Larger batches increase intermediate GPU-memory use. The input cache excludes model outputs and is released after each update batch; its cap does not bound all training memory. Evicted graph inputs must be prepared again.

Model selection uses frozen models independently of sampling/training synchronization. Validation is scheduled after the first network update, when trained-episode counts cross the interval, and at completion. For an interval of 100 and cumulative counts 29, 50, 72, 97, 126, models after updates 1 and 5 are evaluated. Pending evaluations finish before the final best model is published.

Models and results​

Model files use format 4; older model formats require retraining. Candidate plans must respect buffer capacity and material flow. The network selects executable plans. Healthy vehicles retain committed deliveries; failed unloaded vehicles may be reassigned. Existing urgent-order, reservation, preemption, and failure mechanisms remain active.

When testing a model, inspect completion rate, makespan, reassignment scale, and decision latency. Tail completion-time statistics include a completed-episode count; interpret them together with completion rate.

Keep training, validation, and final benchmarks separate. Final benchmark data must not select models or tune parameters. See the experiment workbench for checkpoints, continued training, tests, and replay.