Abstract
Textual context such as news, reports, and logs can provide valuable signals for time series forecasting, especially when future dynamics are driven by external events that are not yet visible in historical values. Existing multimodal forecasting methods often either ask large language models (LLMs) to predict numerical values directly or fuse text and time series implicitly, making contextual influence difficult to interpret and control. We propose SCENARIODIFF, a hierarchical contextual reasoning framework for multimodal time series forecasting under noisy and weakly aligned documents. SCENARIODIFF organizes contextual information into three levels: a Historical Context Agent extracts stepwise evidence from raw documents, a Scenario Agent produces a qualitative scenario description for the forecast horizon, and an Anchor Guidance Agent generates sparse anchor points for event-relevant future regions. These structured signals condition a Multimodal Diffusion Transformer, while Anchor Blended Sampling locally refines generated trajectories without retraining. Experiments on the Time-MMD benchmark show that SCENARIODIFF is especially effective in event-driven domains, demonstrating the value of explicit hierarchical scenario guidance for multimodal time series forecasting. Our full implementation is available at https://github.com/ttb06/ScenarioDiff.
Scenario-level guidance
An additional challenge is that contextual text may describe not only smooth trend continuation, but also event-driven deviations such as policy interventions, outages, or abrupt demand shifts. Capturing such cases requires future-oriented hypotheses that can guide a numerical forecaster without asking an LLM to directly output numbers. It also requires a mechanism for translating coarse textual expectations into localized constraints when abrupt changes are likely.
Hierarchical contextual reasoning
In this work, we propose SCENARIODIFF, a hierarchical contextual reasoning framework for MTSF. The key idea is to organize contextual information into three complementary levels of guidance, each serving a different role in guiding the forecast. First, a frozen Historical Context Agent compresses raw documents into stepwise context summaries aligned with the observed history, grounding the model in document-level evidence. Second, a frozen Scenario Agent synthesizes the historical series and these summaries into a scenario description, which serves as a qualitative prior over the forecast horizon. Third, an Anchor Guidance Agent converts the available context into sparse anchor points, providing time-localized guidance for future regions where abrupt changes are likely. This hierarchy separates past evidence, future hypotheses, and local trajectory constraints, making contextual influence more explicit and modular than implicit fusion.
Multimodal Diffusion Transformer and Anchor Blended Sampling
We instantiate the forecaster as a Multimodal Diffusion Transformer, which models future trajectories through iterative denoising conditioned on the structured signals produced by the agents. The stepwise context summaries and scenario description guide the base diffusion process, allowing the model to combine numerical history with scenario-level textual evidence without asking LLMs to directly output numerical forecasts. To make anchor points actionable, we introduce Anchor Blended Sampling, an inference-time refinement procedure that locally edits anchor-relevant regions using a distance-to-band objective while preserving non-anchor regions through blended diffusion. Thus, anchor points act as soft local constraints rather than hard overrides, encouraging event-consistent trajectories without retraining the base forecaster.
SCENARIODIFF / THE FILM
When the news
changes the forecast.
From an event report to a scenario, sparse anchors and a future trajectory. Follow three stories with consequences for public health, daily travel and household costs.
03 EVENTS / 03 PATTERNS OF CHANGE
Three events. Three forecasts.
Anchors provide approximate guidance. Forecasts move partway toward their intervals and can retain a substantial gap from ground truth.
History and ground truth use published observations. Event reports are sourced. Forecast and anchor paths are illustrative retrospective examples, not measured ScenarioDiff results or a point-in-time evaluation.
Results
Main results on Time-MMD
Table I reports average MSE and MAE over prediction horizons across the five Time-MMD domains. Overall, SCENARIODIFF achieves the strongest horizon-level performance, with the largest number of MSE/MAE wins across horizons and domains. The gains are most evident in event-driven domains, especially Economy and Security, while SCENARIODIFF also shows competitive performance on Traffic. These results suggest that noisy documents become more effective for forecasting when transformed into explicit scenario-level signals, rather than fused as unstructured text. Compared with numerical-only and LLM-prior baselines, SCENARIODIFF benefits from using LLM agents for contextual reasoning instead of direct numerical prediction. The agents extract historical evidence, generate scenario descriptions, and produce anchor points that guide a dedicated probabilistic forecaster. This separation of historical context, scenario-level guidance, and anchor-based refinement leads to stronger performance when text contains actionable event signals.
| Models | Hor. 1st | Ovr. 1st | Economy | Energy | Security | Social Good | Traffic | |||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| MSE | MAE | MSE | MAE | MSE | MAE | MSE | MAE | MSE | MAE | |||
| Informer | 0 | 0 | 0.891 | 0.774 | 0.456 | 0.525 | 127.556 | 6.595 | 0.973 | 0.599 | 0.248 | 0.405 |
| Reformer | 0 | 0 | 1.036 | 0.853 | 0.676 | 0.632 | 122.835 | 6.285 | 1.046 | 0.647 | 0.292 | 0.443 |
| Autoformer | 0 | 0 | 0.340 | 0.465 | 0.478 | 0.527 | 112.782 | 5.314 | 1.618 | 0.810 | 0.240 | 0.301 |
| FEDformer | 2 | 0 | 0.286 | 0.410 | 0.394 | 0.457 | 113.725 | 5.382 | 1.238 | 0.666 | 0.231 | 0.271 |
| PatchTST | 5 | 3 | 0.255 | 0.387 | 0.203 | 0.326 | 91.734 | 5.429 | 1.100 | 0.559 | 0.102 | 0.178 |
| iTransformer | 5 | 2 | 0.280 | 0.398 | 0.227 | 0.344 | 113.248 | 5.432 | 1.235 | 0.568 | 0.208 | 0.238 |
| PAttn | 3 | 0 | 0.237 | 0.376 | 0.267 | 0.388 | 83.117 | 4.956 | 1.196 | 0.569 | 0.104 | 0.181 |
| DLinear | 0 | 0 | 0.579 | 0.636 | 0.391 | 0.448 | 106.504 | 4.665 | 1.524 | 0.935 | 0.284 | 0.415 |
| FiLM | 0 | 0 | 0.460 | 0.556 | 0.375 | 0.469 | 108.179 | 5.158 | 1.608 | 0.949 | 0.236 | 0.326 |
| TSMixer | 1 | 0 | 1.973 | 1.155 | 0.410 | 0.482 | 94.503 | 5.735 | 1.301 | 0.796 | 0.818 | 0.714 |
| TiDE | 0 | 0 | 0.463 | 0.548 | 0.484 | 0.518 | 91.498 | 5.548 | 1.939 | 1.045 | 0.230 | 0.378 |
| Time-LLM | 0 | 0 | 0.346 | 0.469 | 0.464 | 0.491 | 79.945 | 4.785 | 1.924 | 1.071 | 0.195 | 0.330 |
| S2IP-LLM | 4 | 0 | 0.273 | 0.417 | 0.224 | 0.343 | 76.184 | 4.381 | 1.025 | 0.594 | 0.191 | 0.310 |
| TaTS | 7 | 3 | 0.924 | 0.732 | 0.457 | 0.540 | 124.457 | 6.368 | 0.886 | 0.541 | 0.184 | 0.312 |
| MM-TSF | 0 | 0 | 0.816 | 0.737 | 0.397 | 0.481 | 127.460 | 6.573 | 0.959 | 0.580 | 0.229 | 0.382 |
| CSDI | 4 | 3 | 1.396 | 0.943 | 0.545 | 0.531 | 96.372 | 6.092 | 0.908 | 0.491 | 0.124 | 0.256 |
| TMDM | 3 | 1 | 1.038 | 0.765 | 0.320 | 0.398 | 76.270 | 4.177 | 1.305 | 0.673 | 0.103 | 0.187 |
| NsDiff | 1 | 0 | 1.781 | 1.200 | 0.245 | 0.403 | 100.422 | 6.356 | 3.109 | 1.527 | 0.255 | 0.437 |
| TimeDiff | 0 | 0 | 3.372 | 1.718 | 1.015 | 0.799 | 102.157 | 6.541 | 2.234 | 1.188 | 2.488 | 1.521 |
| SCENARIODIFF | 13 | 4 | 0.216 | 0.353 | 0.225 | 0.367 | 74.802 | 4.261 | 1.313 | 0.707 | 0.099 | 0.183 |
Probabilistic forecasting
Table II reports CRPS for representative diffusion-based probabilistic forecasting models. SCENARIODIFF achieves the best event-domain average CRPS, indicating that hierarchical scenario guidance can also improve predictive distributions when textual evidence is informative. The improvement is not uniform across all domains, as strong numerical or diffusion-based baselines remain competitive in some settings. This domain-dependent behavior supports our main motivation: scenario-level guidance is most beneficial when external text provides actionable signals about future dynamics.
| Model | Econ. | Ener. | Sec. | Soc. | Traf. | Event Avg. |
|---|---|---|---|---|---|---|
| CSDI | 0.231 | 0.138 | 0.805 | 0.109 | 0.043 | 0.392 |
| TMDM | 0.190 | 0.099 | 0.511 | 0.156 | 0.030 | 0.267 |
| NSDiff | 0.317 | 0.104 | 0.852 | 0.375 | 0.064 | 0.424 |
| TimeDiff | 0.436 | 0.197 | 0.883 | 0.266 | 0.251 | 0.506 |
| SCENARIODIFF | 0.094 | 0.112 | 0.584 | 0.189 | 0.036 | 0.263 |