Trajectory forecasting · 2026

Revisiting Numerical Forecasting Models for Language-Based Trajectory Prediction

MoRE (Mixture of Reward Experts) refines a language-based trajectory predictor with rewards from five frozen numerical models.

1Yonsei University2DGIST

Abstract

Language-based trajectory predictors represent coordinates as discrete tokens and learn auxiliary tasks such as destination and group reasoning. This formulation enables the model to capture behavioral intent and social context beyond coordinate dynamics alone. However, token-level objectives provide only indirect guidance for continuous coordinate-space dynamics. To address this limitation, we introduce MoRE (Mixture Of Reward Experts), a refinement framework that transfers numerical forecasting priors into a pretrained language-based predictor through reinforcement learning. Five frozen numerical predictors provide complementary coordinate-level knowledge of motion and interactions. Their predictions are converted into expert rewards and combined through an uncertainty-weighted consensus that penalizes disagreement. A ground-truth reward anchors the prediction to the target trajectory. To focus refinement on difficult cases, MoRE refines the policy using the top 1% of training samples ranked by predictive entropy. Expert predictions are computed once and cached before PPO training, so the experts are not run during policy updates or inference. In this way, MoRE combines the contextual modeling of the language-based predictor with coordinate-level feedback from numerical experts. On ETH-UCY, MoRE reduces ADE from 0.22 to 0.20 m and FDE from 0.32 to 0.29 m. Relative to the base policy, ADE decreases by 17.9% on SDD and 12.7% on NBA. On ETH-UCY, MoRE also reduces collision rates and better matches ground-truth pedestrian spacing, without increasing measured inference memory or latency.

0.22 → 0.20 mETH-UCY ADE (best-of-20)
0.32 → 0.29 mETH-UCY FDE
−17.9%ADE on SDD vs. base policy
+0 msInference overhead
Method

MoRE training pipeline

MoRE builds on a pretrained language-based forecasting architecture and refines it using rewards from numerical experts.

MoRE training pipeline: uncertainty-driven sample mining, numerical experts, multi-expert consensus reward modeling, and PPO-based reinforcement learning
A frozen base policy ranks training samples by predictive entropy. Numerical expert predictions for the selected subset are cached once. Decoded policy samples receive expert-consensus and ground-truth rewards for PPO refinement, interleaved with supervised learning. Only the refined forecasting policy is needed at inference.
  1. Uncertainty-driven sample mining

    The frozen base policy scores each training sample by the entropy of its token distribution. Refinement focuses on the top 1% most uncertain samples.

  2. Numerical experts

    Five frozen predictors (Social-STGCNN, DMRGCN, GP-Graph, SingularTrajectory, Expert-Trajectory) predict the selected samples once; the predictions are cached.

  3. Multi-expert consensus reward

    Each expert scores a decoded path by Rk = −MSE(Ŝ, Sk). Uncertainty-weighted consensus penalizes expert disagreement, and a ground-truth reward anchors the prediction.

  4. Reinforcement learning

    The policy is refined with PPO, interleaved with supervised learning. The experts are not needed at inference.

Results

Quantitative results

Best-of-20 ADE / FDE: ETH-UCY in meters, SDD and GCS in pixels. Selected rows from Table 1 of the paper; lower is better.

ModelETHHOTELUNIVZARA1ZARA2AVGSDDGCS
Numerical
GP-Graph0.43/0.630.18/0.300.24/0.420.17/0.310.15/0.290.23/0.399.1/13.87.8/13.7
MART0.35/0.470.14/0.220.25/0.450.17/0.290.13/0.220.21/0.337.4/11.810.6/14.1
SingularTrajectory0.35/0.420.13/0.190.25/0.440.19/0.320.15/0.250.21/0.327.58/12.17.9/13.3
MoFlow0.40/0.570.11/0.170.23/0.390.15/0.260.12/0.220.20/0.327.5/12.09.1/11.6
Language-based
LMTraj-SUP0.41/0.500.12/0.160.22/0.340.20/0.320.17/0.270.22/0.327.8/10.17.1/9.6
W2W0.35/0.410.12/0.150.20/0.320.19/0.290.17/0.260.21/0.297.4/10.1–
MoRE (Ours)0.36/0.420.11/0.140.21/0.320.18/0.280.17/0.260.20/0.296.4/9.56.8/8.5
Results

Qualitative results

Numerical models vs. LMTraj vs. MoRE. Pick a dataset, a scene, a camera and a model. Each clip replays the recorded crowd, pauses at the prediction moment while the sampled futures are drawn, lets the future unfold, and then highlights the sample closest to the ground truth. ETH-UCY ZARA and SDD Hyang are replayed in 3D; ETH, HOTEL and UNIV are shown on the real-world videos.

Dataset · 3D replay or real-world video
Model
Camera
Scene

Refinement

From consistent to consistent and accurate

Ego view

Standing behind the pedestrian

Cite

BibTeX

@article{lee2026more,
  title   = {Revisiting Numerical Forecasting Models for Language-Based Trajectory Prediction},
  author  = {Lee, JunGyu and Bae, Inhwan and Jeon, Hae-Gon},
  journal = {arXiv preprint arXiv:2610.07954},
  year    = {2026}
}