SCENARIODIFF: A Scenario-level Guidance Framework for Multimodal Time Series Forecasting

Tuan-Binh Tran1, Dat Nguyen-Cong2, Duc-Trong Le3, Thanh Trung Huynh1, Tung Kieu4

1VinUniversity2FPT Software AI Center3VNU University of Engineering and Technology4Aalborg University

ICDM 2026

Abstract

Textual context such as news, reports, and logs can provide valuable signals for time series forecasting, especially when future dynamics are driven by external events that are not yet visible in historical values. Existing multimodal forecasting methods often either ask large language models (LLMs) to predict numerical values directly or fuse text and time series implicitly, making contextual influence difficult to interpret and control. We propose SCENARIODIFF, a hierarchical contextual reasoning framework for multimodal time series forecasting under noisy and weakly aligned documents. SCENARIODIFF organizes contextual information into three levels: a Historical Context Agent extracts stepwise evidence from raw documents, a Scenario Agent produces a qualitative scenario description for the forecast horizon, and an Anchor Guidance Agent generates sparse anchor points for event-relevant future regions. These structured signals condition a Multimodal Diffusion Transformer, while Anchor Blended Sampling locally refines generated trajectories without retraining. Experiments on the Time-MMD benchmark show that SCENARIODIFF is especially effective in event-driven domains, demonstrating the value of explicit hierarchical scenario guidance for multimodal time series forecasting. Our full implementation is available at https://github.com/ttb06/ScenarioDiff.

Scenario-level guidance

An additional challenge is that contextual text may describe not only smooth trend continuation, but also event-driven deviations such as policy interventions, outages, or abrupt demand shifts. Capturing such cases requires future-oriented hypotheses that can guide a numerical forecaster without asking an LLM to directly output numbers. It also requires a mechanism for translating coarse textual expectations into localized constraints when abrupt changes are likely.

Hierarchical contextual reasoning

In this work, we propose SCENARIODIFF, a hierarchical contextual reasoning framework for MTSF. The key idea is to organize contextual information into three complementary levels of guidance, each serving a different role in guiding the forecast. First, a frozen Historical Context Agent compresses raw documents into stepwise context summaries aligned with the observed history, grounding the model in document-level evidence. Second, a frozen Scenario Agent synthesizes the historical series and these summaries into a scenario description, which serves as a qualitative prior over the forecast horizon. Third, an Anchor Guidance Agent converts the available context into sparse anchor points, providing time-localized guidance for future regions where abrupt changes are likely. This hierarchy separates past evidence, future hypotheses, and local trajectory constraints, making contextual influence more explicit and modular than implicit fusion.

Original paper Figure 2 showing time series and observed documents, three guidance agents, the scenario-guided diffusion backbone, anchor blended sampling, and the final forecast.
Fig. 2: SCENARIODIFF overview. Open full-resolution figure.

Multimodal Diffusion Transformer and Anchor Blended Sampling

We instantiate the forecaster as a Multimodal Diffusion Transformer, which models future trajectories through iterative denoising conditioned on the structured signals produced by the agents. The stepwise context summaries and scenario description guide the base diffusion process, allowing the model to combine numerical history with scenario-level textual evidence without asking LLMs to directly output numerical forecasts. To make anchor points actionable, we introduce Anchor Blended Sampling, an inference-time refinement procedure that locally edits anchor-relevant regions using a distance-to-band objective while preserving non-anchor regions through blended diffusion. Thus, anchor points act as soft local constraints rather than hard overrides, encouraging event-consistent trajectories without retraining the base forecaster.

SCENARIODIFF / THE FILM

When the news
changes the forecast.

From an event report to a scenario, sparse anchors and a future trajectory. Follow three stories with consequences for public health, daily travel and household costs.

02:50 / 1080p / English on-screen narrationHealthcare · Mobility · Energy

03 EVENTS / 03 PATTERNS OF CHANGE

Three events. Three forecasts.

Anchors provide approximate guidance. Forecasts move partway toward their intervals and can retain a substantial gap from ground truth.

HistoryBaselineForecastGround truthAnchor interval

History and ground truth use published observations. Event reports are sourced. Forecast and anchor paths are illustrative retrospective examples, not measured ScenarioDiff results or a point-in-time evaluation.

Event walkthrough

Results

Main results on Time-MMD

Table I reports average MSE and MAE over prediction horizons across the five Time-MMD domains. Overall, SCENARIODIFF achieves the strongest horizon-level performance, with the largest number of MSE/MAE wins across horizons and domains. The gains are most evident in event-driven domains, especially Economy and Security, while SCENARIODIFF also shows competitive performance on Traffic. These results suggest that noisy documents become more effective for forecasting when transformed into explicit scenario-level signals, rather than fused as unstructured text. Compared with numerical-only and LLM-prior baselines, SCENARIODIFF benefits from using LLM agents for contextual reasoning instead of direct numerical prediction. The agents extract historical evidence, generate scenario descriptions, and produce anchor points that guide a dedicated probabilistic forecaster. This separation of historical context, scenario-level guidance, and anchor-based refinement leads to stronger performance when text contains actionable event signals.

Table I: Overall results on Time-MMD. The best, runner-up, and third-best mean results are highlighted in red, blue, and bold, respectively. Hor. 1st counts the number of first-place results over all horizon-domain-metric entries, while Ovr. 1st counts first-place results after averaging over horizons within each domain and metric.
ModelsHor. 1stOvr. 1stEconomyEnergySecuritySocial GoodTraffic
MSEMAEMSEMAEMSEMAEMSEMAEMSEMAE
Informer000.8910.7740.4560.525127.5566.5950.9730.5990.2480.405
Reformer001.0360.8530.6760.632122.8356.2851.0460.6470.2920.443
Autoformer000.3400.4650.4780.527112.7825.3141.6180.8100.2400.301
FEDformer200.2860.4100.3940.457113.7255.3821.2380.6660.2310.271
PatchTST530.2550.3870.2030.32691.7345.4291.1000.5590.1020.178
iTransformer520.2800.3980.2270.344113.2485.4321.2350.5680.2080.238
PAttn300.2370.3760.2670.38883.1174.9561.1960.5690.1040.181
DLinear000.5790.6360.3910.448106.5044.6651.5240.9350.2840.415
FiLM000.4600.5560.3750.469108.1795.1581.6080.9490.2360.326
TSMixer101.9731.1550.4100.48294.5035.7351.3010.7960.8180.714
TiDE000.4630.5480.4840.51891.4985.5481.9391.0450.2300.378
Time-LLM000.3460.4690.4640.49179.9454.7851.9241.0710.1950.330
S2IP-LLM400.2730.4170.2240.34376.1844.3811.0250.5940.1910.310
TaTS730.9240.7320.4570.540124.4576.3680.8860.5410.1840.312
MM-TSF000.8160.7370.3970.481127.4606.5730.9590.5800.2290.382
CSDI431.3960.9430.5450.53196.3726.0920.9080.4910.1240.256
TMDM311.0380.7650.3200.39876.2704.1771.3050.6730.1030.187
NsDiff101.7811.2000.2450.403100.4226.3563.1091.5270.2550.437
TimeDiff003.3721.7181.0150.799102.1576.5412.2341.1882.4881.521
SCENARIODIFF1340.2160.3530.2250.36774.8024.2611.3130.7070.0990.183

Probabilistic forecasting

Table II reports CRPS for representative diffusion-based probabilistic forecasting models. SCENARIODIFF achieves the best event-domain average CRPS, indicating that hierarchical scenario guidance can also improve predictive distributions when textual evidence is informative. The improvement is not uniform across all domains, as strong numerical or diffusion-based baselines remain competitive in some settings. This domain-dependent behavior supports our main motivation: scenario-level guidance is most beneficial when external text provides actionable signals about future dynamics.

Table II: Average CRPS over prediction horizons for representative diffusion-based probabilistic forecasting models. We additionally report the average over event-driven domains, where textual scenarios provide actionable future signals.
ModelEcon.Ener.Sec.Soc.Traf.Event Avg.
CSDI0.2310.1380.8050.1090.0430.392
TMDM0.1900.0990.5110.1560.0300.267
NSDiff0.3170.1040.8520.3750.0640.424
TimeDiff0.4360.1970.8830.2660.2510.506
SCENARIODIFF0.0940.1120.5840.1890.0360.263