AAAI 2026Multivariate forecasting

T3Time: Tri-Modal Time Series Forecasting via Adaptive Multi-Head Alignment and Residual Fusion

  • Abdul Monaf Chowdhury1
  • Rabeya Akter1
  • Safaeid Hossain Arib1
  • 1University of Dhaka
TL;DR

A time series can be read in the time domain, in the frequency domain, or through a language model. Most forecasters pick one or two views and fuse them the same way at every horizon. T3Time encodes all three, lets the forecast horizon decide how much weight the temporal and spectral views get, aligns them with the language view through several cross-attention heads weighted per variable, and keeps a learned residual path back to the raw features.

14/16
Best average MSE or MAE across 8 datasets × 2 metrics
−4.4%
Average MSE against TimeCMA, the strongest prompt-based baseline
−11.3%
MSE on ILI against the best baseline, the largest single gain
−4.1%
MSE with only 5% of the training data, against the best baseline per dataset
−3.6% at 10%
Left: a bimodal model fuses a time-series encoder and an LLM statically into a weak representation. Right: T3Time adds a frequency encoder, horizon gating, and cross-modal alignment with adaptive fusion for a robust representation.

Figure 1. Two modalities versus three. (a) Bimodal forecasters fuse time-series and prompt features statically, with no view of frequency and no notion of horizon. (b) T3Time adds a frequency encoder, gates time against frequency by horizon, and fuses the result with the prompt view through adaptive multi-head alignment.

Background

Three ways to read one series

Multivariate forecasting takes the last \(L\) observations of \(N\) variables and predicts the next \(L_p\). Architectures have split along the representation they trust. Time-domain Transformers such as PatchTST and iTransformer learn from raw values; frequency-domain models such as FEDformer model periodicity; and a newer line prompts a pretrained language model, either by reprogramming it (Time-LLM) or by aligning series and prompt embeddings (TimeCMA).

Gap 1Isolated modalities

Most models commit to one representation. Time-language models such as TimeCMA still leave out the spectral view that captures global periodicity.

Gap 2Narrow alignment

When modalities are combined, a single cross-attention head carries all of the alignment, which limits how many different interactions it can capture.

Gap 3Horizon rigidity

The same fusion is applied whether the model forecasts 96 or 720 steps ahead, although near-term forecasts lean on local detail and long ones on periodic structure.

InteractiveWhat the frequency branch sees
A 96-step window and its real FFT, computed live
A time series and its magnitude spectrum Top: a 96-step synthetic series with a daily cycle, a half-day harmonic, a trend and noise. Bottom: the magnitude of its real Fourier transform over 49 frequency bins.

The strongest bin is k = 4, a period of 24 steps.

Each of the \(\lfloor 96/2 \rfloor + 1 = 49\) bins becomes one token for the frequency encoder. Noise spreads across every bin, but a cycle stays in one, which is why a few spectral tokens can summarise periodic structure that would take attention over the whole window to find in the time domain. The series is z-scored first, as the model's input is.

Method

Encode three views, gate by horizon, align, and keep a residual path

T3Time architecture: frequency, time-series and LLM encoding branches; horizon-aware gating; multi-head cross-modal attention with adaptive head fusion; channel-wise residual connection; Transformer decoder.

Figure 2. The T3Time architecture. Three encoding branches (frequency, time series, and a frozen LLM over text prompts), a horizon-aware gate that fuses time and frequency, multi-head cross-modal alignment with per-head importance weights, and a channel-wise residual connection before the decoder.

Tri-modal encoding

For a normalised input \(X \in \mathbb{R}^{B \times N \times L}\), the frequency branch takes the real FFT along time and keeps the magnitude. Each of the \(L_f\) bins is a token: projected, encoded by a one-layer Transformer, and pooled with learned attention weights over the bins.

\[ F = \big|\mathcal{F}_r(X)\big|,\quad L_f = \lfloor L/2 \rfloor + 1, \qquad \tilde F = \sum_{l=1}^{L_f} \alpha_l\, \mathcal{T}\big(\phi(F W_f^\top)\big)_{:,l}, \]
frequency

The time branch embeds each variable's whole window, \(Z_t = X W_t\), and encodes it with a Transformer into \(\tilde Z_t\). The language branch writes one prompt per variable, runs it through a frozen GPT-2, and keeps the last token's embedding as \(Z_{\text{LLM}}\):

# prompt template, one per variable and window
From [t1] to [t2], the values were value1, ..., valuen every f. The total trend value was T.

Horizon-aware gating

Short forecasts should lean on local, time-domain detail and long ones on global periodicity. A small MLP sees the pooled time encoding and the normalised prediction length, and outputs a per-channel gate that mixes the two views:

\[ g = \sigma\big(W_4\, \phi(W_3\, [\,\mathrm{pool}(\tilde Z_t);\ L_p / c\,])\big), \qquad Z_g = g \odot \tilde F + (1-g) \odot \tilde Z_t. \]
horizon gate

Adaptive multi-head cross-modal alignment

Following TimeCMA, the fused temporal–spectral features query the prompt embeddings through cross-attention, but with \(H\) independent heads instead of one. A gating network then scores the heads for each variable and takes their convex combination, so different variables can rely on different alignments:

\[ \begin{aligned} H^{(h)} &= \mathrm{CrossAttn}_h\big(Z_g,\, Z_{\text{LLM}}\big), \quad \pi =\mathrm{softmax}\big(W_6\, \phi(\mathrm{LN}(W_5\, U))\big), \\ \Lambda &= \textstyle\sum_{h=1}^{H} \pi^{(h)}\, H^{(h)}, \end{aligned} \]
head fusion

where \(U\) concatenates the head outputs. The model uses \(H = 4\); a sweep over 1 to 16 heads finds four best on average, with more heads fragmenting the attention rather than helping.

Channel-wise residual fusion

Alignment should add context, not overwrite what the series already says. A learned coefficient per latent channel, \(\gamma_c \in [0,1]\), decides how much of each channel comes from the aligned features and how much from the gated temporal–spectral features, before a Transformer decoder and a linear head produce the forecast:

\[ \Theta_{c} = \gamma_c\, \Lambda_{c} + (1 - \gamma_c)\, Z_{g,c}, \qquad \hat Y = \mathcal{D}(\Theta^\top)\, W_p^\top + b_p. \]
residual
Results

Long-term forecasting on eight benchmarks

Following TimeCMA's protocol, the input window is 96 steps and each score averages four horizons, \(\{96, 192, 336, 720\}\) (for ILI, a 36-step input and \(\{24, 36, 48, 60\}\)). The eight datasets cover electricity transformers (ETTh1/2, ETTm1/2), electricity consumption (ECL), weather, influenza-like illness (ILI) and exchange rates. Each result averages three seeds.

Dataset T3Time MSE Strongest baseline Change in MSE
ETTm1 0.372 TimeCMA · 0.380 −2.1%
ETTm2 0.279 TimeCMA · 0.275 +1.5%
ETTh1 0.418 TimeCMA · 0.423 −1.2%
ETTh2 0.348 TimeCMA · 0.372 −6.5%
ECL 0.170 TimeCMA · 0.174 −2.3%
Weather 0.244 TimeCMA · 0.250 −2.4%
ILI 1.705 TimeCMA · 1.922 −11.3%
Exchange 0.353 DLinear · 0.354 −0.3%

Negative is better. The bar runs from zero on a shared −12% to +4% scale; it is coloured where T3Time has the lower error.

T3Time has the lowest MSE on seven of eight datasets and the lowest MAE on seven of eight. Averaged over all eight it cuts MSE by 4.4% against TimeCMA, 8.9% against iTransformer and 11.2% against Time-LLM.

The exceptions are ETTm2, where TimeCMA's MSE is 1.5% lower, and ETTm1 MAE, where TimeCMA leads by 0.001. The largest gains are on ILI (−11.3%) and ETTh2 (−6.5%).

All baselines · average MSE and MAE 8 datasets × 9 methods

MSE

Dataset T3Time TimeCMA TimeLLM UniTime TimesNet DLinear iTransformer PatchTST OFA
ETTm1 0.372 0.380 0.410 0.385 0.400 0.403 0.407 0.392 0.396
ETTm2 0.279 0.275 0.296 0.293 0.291 0.350 0.288 0.285 0.294
ETTh1 0.418 0.423 0.448 0.442 0.458 0.456 0.454 0.463 0.457
ETTh2 0.348 0.372 0.381 0.378 0.414 0.559 0.383 0.395 0.389
ECL 0.170 0.174 0.195 0.216 0.192 0.212 0.178 0.207 0.217
Weather 0.244 0.250 0.275 0.253 0.259 0.265 0.258 0.257 0.279
ILI 1.705 1.922 2.432 2.108 2.139 2.616 2.444 2.388 2.623
Exchange 0.353 0.395 0.372 0.364 0.416 0.354 0.360 0.390 0.519

MAE

Dataset T3Time TimeCMA TimeLLM UniTime TimesNet DLinear iTransformer PatchTST OFA
ETTm1 0.393 0.392 0.409 0.399 0.406 0.407 0.410 0.402 0.401
ETTm2 0.322 0.323 0.340 0.334 0.333 0.401 0.332 0.328 0.339
ETTh1 0.430 0.431 0.443 0.448 0.450 0.452 0.447 0.449 0.450
ETTh2 0.390 0.397 0.404 0.403 0.427 0.515 0.407 0.414 0.414
ECL 0.266 0.269 0.288 0.306 0.295 0.300 0.270 0.289 0.308
Weather 0.275 0.276 0.291 0.276 0.287 0.317 0.278 0.280 0.297
ILI 0.835 0.921 1.012 0.929 0.931 1.090 1.203 1.011 1.060
Exchange 0.401 0.429 0.416 0.404 0.443 0.414 0.403 0.429 0.500

Lower is better. Bold marks the best in each row and underline the runner-up. OFA is also known as GPT4TS.

Few-shot

With a tenth or a twentieth of the data

Following Time-LLM's few-shot setup, models train on only 10% or 5% of the training time steps, with a 512-step input window. This is where pretrained language priors are supposed to pay off, and where the gap to non-LLM baselines grows: TimesNet, DLinear and PatchTST trail by more than 15% in MSE.

Dataset T3Time MSE Strongest baseline Change in MSE
ETTm1 0.376 TimeCMA · 0.387 −2.8%
ETTm2 0.266 TimeLLM · 0.277 −4.0%
ETTh1 0.449 TimeCMA · 0.480 −6.5%
ETTh2 0.357 TimeLLM · 0.370 −3.5%
Weather 0.226 TimeCMA · 0.229 −1.3%

Lower MSE than the strongest baseline on all five datasets, by 3.6% on average: 7.1% lower MSE than TimeCMA, 7.4% lower than Time-LLM and 13.4% lower than GPT4TS on average.

On ETTm2, Time-LLM keeps a slightly lower MAE; T3Time has the best MAE on the other four.

All baselines · 10% of training data 5 datasets × 8 methods

MSE

Dataset T3Time TimeCMA TimeLLM GPT4TS TimesNet DLinear PatchTST FEDformer
ETTm1 0.376 0.387 0.404 0.464 0.677 0.411 0.501 0.722
ETTm2 0.266 0.312 0.277 0.293 0.320 0.316 0.296 0.463
ETTh1 0.449 0.480 0.556 0.590 0.869 0.691 0.633 0.639
ETTh2 0.357 0.398 0.370 0.397 0.479 0.605 0.415 0.466
Weather 0.226 0.229 0.234 0.238 0.279 0.241 0.242 0.284

MAE

Dataset T3Time TimeCMA TimeLLM GPT4TS TimesNet DLinear PatchTST FEDformer
ETTm1 0.398 0.410 0.427 0.441 0.537 0.429 0.466 0.605
ETTm2 0.327 0.358 0.323 0.335 0.353 0.368 0.343 0.488
ETTh1 0.454 0.479 0.522 0.525 0.628 0.600 0.542 0.561
ETTh2 0.388 0.433 0.394 0.421 0.465 0.538 0.431 0.475
Weather 0.268 0.272 0.273 0.275 0.301 0.283 0.279 0.324

Averages over the four horizons; lower is better.

Dataset T3Time MSE Strongest baseline Change in MSE
ETTm1 0.384 TimeCMA · 0.396 −3.0%
ETTm2 0.267 TimeLLM · 0.274 −2.6%
ETTh1 0.442 TimeCMA · 0.472 −6.4%
ETTh2 0.357 TimeLLM · 0.382 −6.5%
Weather 0.226 TimeCMA · 0.231 −2.2%

Lower MSE than the strongest baseline on all five datasets, by 4.1% on average: 8.0% lower MSE than TimeCMA, 12.3% lower than Time-LLM and 18.4% lower than GPT4TS on average.

On ETTm2, Time-LLM keeps a slightly lower MAE; T3Time has the best MAE on the other four.

All baselines · 5% of training data 5 datasets × 8 methods

MSE

Dataset T3Time TimeCMA TimeLLM GPT4TS TimesNet DLinear PatchTST FEDformer
ETTm1 0.384 0.396 0.425 0.472 0.717 0.400 0.526 0.730
ETTm2 0.267 0.329 0.274 0.308 0.344 0.399 0.314 0.381
ETTh1 0.442 0.472 0.627 0.681 0.925 0.750 0.694 0.658
ETTh2 0.357 0.395 0.382 0.400 0.439 0.694 0.827 0.463
Weather 0.226 0.231 0.260 0.263 0.298 0.263 0.269 0.309

MAE

Dataset T3Time TimeCMA TimeLLM GPT4TS TimesNet DLinear PatchTST FEDformer
ETTm1 0.405 0.416 0.434 0.450 0.561 0.417 0.476 0.592
ETTm2 0.330 0.367 0.323 0.346 0.372 0.426 0.352 0.404
ETTh1 0.451 0.470 0.543 0.560 0.647 0.611 0.569 0.562
ETTh2 0.403 0.430 0.418 0.433 0.448 0.577 0.615 0.454
Weather 0.269 0.273 0.309 0.301 0.318 0.308 0.303 0.353

Averages over the four horizons; lower is better.

Ablations

Every component helps; the residual path most

Removing one component at a time under the long-term setup, the full model is best, or tied for best, in all 14 dataset–metric pairs of the design study.

What each component is worth Mean MSE increase when it is removed (%), 7 datasets
w/o residual connection 9.7
w/o frequency branch 3.4
w/o multi-head CMA 2.9
w/o horizon gating 2.2

Without the channel-wise residual, MSE rises by 9.7% on average and by 28% on ILI: aligned features alone lose information the raw temporal–spectral path still carries. Dropping the frequency branch costs 3.4%, most on Exchange (+5.9%) and ILI (+4.8%).

Values are relative MSE increases per dataset, averaged over the seven datasets of the design study and computed from its table below.

Design-variant table 7 datasets × 5 variants

MSE

Dataset T3Time w/o frequency w/o multi-head w/o residual w/o gating
ETTm1 0.372 0.381 0.374 0.404 0.373
ETTm2 0.279 0.279 0.283 0.288 0.280
ETTh1 0.418 0.433 0.421 0.433 0.425
ETTh2 0.348 0.364 0.355 0.384 0.363
Weather 0.244 0.250 0.249 0.249 0.249
ILI 1.705 1.786 1.813 2.176 1.724
Exchange 0.353 0.374 0.378 0.396 0.373

MAE

Dataset T3Time w/o frequency w/o multi-head w/o residual w/o gating
ETTm1 0.393 0.394 0.395 0.413 0.396
ETTm2 0.322 0.324 0.325 0.329 0.322
ETTh1 0.430 0.433 0.432 0.443 0.432
ETTh2 0.390 0.398 0.392 0.411 0.398
Weather 0.275 0.280 0.277 0.280 0.277
ILI 0.835 0.884 0.865 0.969 0.837
Exchange 0.401 0.410 0.413 0.427 0.412

Averages over the four horizons; lower is better.

What the three views learn

t-SNE of time-series, frequency, prompt and forecast embeddings, coloured by dataset. Time and frequency embeddings cluster by dataset; prompt embeddings are mixed.

Figure 3. t-SNE of each modality's embeddings, coloured by dataset. Time-series and especially frequency embeddings separate the datasets cleanly, since each has its own periodicities. Prompt embeddings are more mixed, reflecting how similar the textual descriptions are. The forecast embeddings cluster much like the time-series ones.

Scope: all results use standard public benchmarks and a frozen GPT-2 for the prompt branch. Larger-scale pretraining and richer representations for each modality are left for future work.

Cite

Citation

If this work is useful to you, please cite it as:

@inproceedings{chowdhury2026t3time,
  title     = {{T3Time}: Tri-Modal Time Series Forecasting via Adaptive
               Multi-Head Alignment and Residual Fusion},
  author    = {Chowdhury, Abdul Monaf and Akter, Rabeya and Arib, Safaeid Hossain},
  booktitle = {Proceedings of the AAAI Conference on Artificial Intelligence},
  volume    = {40},
  number    = {25},
  pages     = {20597--20605},
  year      = {2026},
  doi       = {10.1609/aaai.v40i25.39196}
}