A time series can be read in the time domain, in the frequency domain, or through a language model. Most forecasters pick one or two views and fuse them the same way at every horizon. T3Time encodes all three, lets the forecast horizon decide how much weight the temporal and spectral views get, aligns them with the language view through several cross-attention heads weighted per variable, and keeps a learned residual path back to the raw features.
Figure 1. Two modalities versus three. (a) Bimodal forecasters fuse time-series and prompt features statically, with no view of frequency and no notion of horizon. (b) T3Time adds a frequency encoder, gates time against frequency by horizon, and fuses the result with the prompt view through adaptive multi-head alignment.
Three ways to read one series
Multivariate forecasting takes the last \(L\) observations of \(N\) variables and predicts the next \(L_p\). Architectures have split along the representation they trust. Time-domain Transformers such as PatchTST and iTransformer learn from raw values; frequency-domain models such as FEDformer model periodicity; and a newer line prompts a pretrained language model, either by reprogramming it (Time-LLM) or by aligning series and prompt embeddings (TimeCMA).
Most models commit to one representation. Time-language models such as TimeCMA still leave out the spectral view that captures global periodicity.
When modalities are combined, a single cross-attention head carries all of the alignment, which limits how many different interactions it can capture.
The same fusion is applied whether the model forecasts 96 or 720 steps ahead, although near-term forecasts lean on local detail and long ones on periodic structure.
Each of the \(\lfloor 96/2 \rfloor + 1 = 49\) bins becomes one token for the frequency encoder. Noise spreads across every bin, but a cycle stays in one, which is why a few spectral tokens can summarise periodic structure that would take attention over the whole window to find in the time domain. The series is z-scored first, as the model's input is.
Encode three views, gate by horizon, align, and keep a residual path
Figure 2. The T3Time architecture. Three encoding branches (frequency, time series, and a frozen LLM over text prompts), a horizon-aware gate that fuses time and frequency, multi-head cross-modal alignment with per-head importance weights, and a channel-wise residual connection before the decoder.
Tri-modal encoding
For a normalised input \(X \in \mathbb{R}^{B \times N \times L}\), the frequency branch takes the real FFT along time and keeps the magnitude. Each of the \(L_f\) bins is a token: projected, encoded by a one-layer Transformer, and pooled with learned attention weights over the bins.
The time branch embeds each variable's whole window, \(Z_t = X W_t\), and encodes it with a Transformer into \(\tilde Z_t\). The language branch writes one prompt per variable, runs it through a frozen GPT-2, and keeps the last token's embedding as \(Z_{\text{LLM}}\):
# prompt template, one per variable and window From [t1] to [t2], the values were value1, ..., valuen every f. The total trend value was T.
Horizon-aware gating
Short forecasts should lean on local, time-domain detail and long ones on global periodicity. A small MLP sees the pooled time encoding and the normalised prediction length, and outputs a per-channel gate that mixes the two views:
Adaptive multi-head cross-modal alignment
Following TimeCMA, the fused temporal–spectral features query the prompt embeddings through cross-attention, but with \(H\) independent heads instead of one. A gating network then scores the heads for each variable and takes their convex combination, so different variables can rely on different alignments:
where \(U\) concatenates the head outputs. The model uses \(H = 4\); a sweep over 1 to 16 heads finds four best on average, with more heads fragmenting the attention rather than helping.
Channel-wise residual fusion
Alignment should add context, not overwrite what the series already says. A learned coefficient per latent channel, \(\gamma_c \in [0,1]\), decides how much of each channel comes from the aligned features and how much from the gated temporal–spectral features, before a Transformer decoder and a linear head produce the forecast:
Long-term forecasting on eight benchmarks
Following TimeCMA's protocol, the input window is 96 steps and each score averages four horizons, \(\{96, 192, 336, 720\}\) (for ILI, a 36-step input and \(\{24, 36, 48, 60\}\)). The eight datasets cover electricity transformers (ETTh1/2, ETTm1/2), electricity consumption (ECL), weather, influenza-like illness (ILI) and exchange rates. Each result averages three seeds.
| Dataset | T3Time MSE | Strongest baseline | Change in MSE |
|---|---|---|---|
| ETTm1 | 0.372 | TimeCMA · 0.380 | −2.1% |
| ETTm2 | 0.279 | TimeCMA · 0.275 | +1.5% |
| ETTh1 | 0.418 | TimeCMA · 0.423 | −1.2% |
| ETTh2 | 0.348 | TimeCMA · 0.372 | −6.5% |
| ECL | 0.170 | TimeCMA · 0.174 | −2.3% |
| Weather | 0.244 | TimeCMA · 0.250 | −2.4% |
| ILI | 1.705 | TimeCMA · 1.922 | −11.3% |
| Exchange | 0.353 | DLinear · 0.354 | −0.3% |
Negative is better. The bar runs from zero on a shared −12% to +4% scale; it is coloured where T3Time has the lower error.
T3Time has the lowest MSE on seven of eight datasets and the lowest MAE on seven of eight. Averaged over all eight it cuts MSE by 4.4% against TimeCMA, 8.9% against iTransformer and 11.2% against Time-LLM.
The exceptions are ETTm2, where TimeCMA's MSE is 1.5% lower, and ETTm1 MAE, where TimeCMA leads by 0.001. The largest gains are on ILI (−11.3%) and ETTh2 (−6.5%).
All baselines · average MSE and MAE 8 datasets × 9 methods
MSE
| Dataset | T3Time | TimeCMA | TimeLLM | UniTime | TimesNet | DLinear | iTransformer | PatchTST | OFA |
|---|---|---|---|---|---|---|---|---|---|
| ETTm1 | 0.372 | 0.380 | 0.410 | 0.385 | 0.400 | 0.403 | 0.407 | 0.392 | 0.396 |
| ETTm2 | 0.279 | 0.275 | 0.296 | 0.293 | 0.291 | 0.350 | 0.288 | 0.285 | 0.294 |
| ETTh1 | 0.418 | 0.423 | 0.448 | 0.442 | 0.458 | 0.456 | 0.454 | 0.463 | 0.457 |
| ETTh2 | 0.348 | 0.372 | 0.381 | 0.378 | 0.414 | 0.559 | 0.383 | 0.395 | 0.389 |
| ECL | 0.170 | 0.174 | 0.195 | 0.216 | 0.192 | 0.212 | 0.178 | 0.207 | 0.217 |
| Weather | 0.244 | 0.250 | 0.275 | 0.253 | 0.259 | 0.265 | 0.258 | 0.257 | 0.279 |
| ILI | 1.705 | 1.922 | 2.432 | 2.108 | 2.139 | 2.616 | 2.444 | 2.388 | 2.623 |
| Exchange | 0.353 | 0.395 | 0.372 | 0.364 | 0.416 | 0.354 | 0.360 | 0.390 | 0.519 |
MAE
| Dataset | T3Time | TimeCMA | TimeLLM | UniTime | TimesNet | DLinear | iTransformer | PatchTST | OFA |
|---|---|---|---|---|---|---|---|---|---|
| ETTm1 | 0.393 | 0.392 | 0.409 | 0.399 | 0.406 | 0.407 | 0.410 | 0.402 | 0.401 |
| ETTm2 | 0.322 | 0.323 | 0.340 | 0.334 | 0.333 | 0.401 | 0.332 | 0.328 | 0.339 |
| ETTh1 | 0.430 | 0.431 | 0.443 | 0.448 | 0.450 | 0.452 | 0.447 | 0.449 | 0.450 |
| ETTh2 | 0.390 | 0.397 | 0.404 | 0.403 | 0.427 | 0.515 | 0.407 | 0.414 | 0.414 |
| ECL | 0.266 | 0.269 | 0.288 | 0.306 | 0.295 | 0.300 | 0.270 | 0.289 | 0.308 |
| Weather | 0.275 | 0.276 | 0.291 | 0.276 | 0.287 | 0.317 | 0.278 | 0.280 | 0.297 |
| ILI | 0.835 | 0.921 | 1.012 | 0.929 | 0.931 | 1.090 | 1.203 | 1.011 | 1.060 |
| Exchange | 0.401 | 0.429 | 0.416 | 0.404 | 0.443 | 0.414 | 0.403 | 0.429 | 0.500 |
Lower is better. Bold marks the best in each row and underline the runner-up. OFA is also known as GPT4TS.
With a tenth or a twentieth of the data
Following Time-LLM's few-shot setup, models train on only 10% or 5% of the training time steps, with a 512-step input window. This is where pretrained language priors are supposed to pay off, and where the gap to non-LLM baselines grows: TimesNet, DLinear and PatchTST trail by more than 15% in MSE.
| Dataset | T3Time MSE | Strongest baseline | Change in MSE |
|---|---|---|---|
| ETTm1 | 0.376 | TimeCMA · 0.387 | −2.8% |
| ETTm2 | 0.266 | TimeLLM · 0.277 | −4.0% |
| ETTh1 | 0.449 | TimeCMA · 0.480 | −6.5% |
| ETTh2 | 0.357 | TimeLLM · 0.370 | −3.5% |
| Weather | 0.226 | TimeCMA · 0.229 | −1.3% |
Lower MSE than the strongest baseline on all five datasets, by 3.6% on average: 7.1% lower MSE than TimeCMA, 7.4% lower than Time-LLM and 13.4% lower than GPT4TS on average.
On ETTm2, Time-LLM keeps a slightly lower MAE; T3Time has the best MAE on the other four.
All baselines · 10% of training data 5 datasets × 8 methods
MSE
| Dataset | T3Time | TimeCMA | TimeLLM | GPT4TS | TimesNet | DLinear | PatchTST | FEDformer |
|---|---|---|---|---|---|---|---|---|
| ETTm1 | 0.376 | 0.387 | 0.404 | 0.464 | 0.677 | 0.411 | 0.501 | 0.722 |
| ETTm2 | 0.266 | 0.312 | 0.277 | 0.293 | 0.320 | 0.316 | 0.296 | 0.463 |
| ETTh1 | 0.449 | 0.480 | 0.556 | 0.590 | 0.869 | 0.691 | 0.633 | 0.639 |
| ETTh2 | 0.357 | 0.398 | 0.370 | 0.397 | 0.479 | 0.605 | 0.415 | 0.466 |
| Weather | 0.226 | 0.229 | 0.234 | 0.238 | 0.279 | 0.241 | 0.242 | 0.284 |
MAE
| Dataset | T3Time | TimeCMA | TimeLLM | GPT4TS | TimesNet | DLinear | PatchTST | FEDformer |
|---|---|---|---|---|---|---|---|---|
| ETTm1 | 0.398 | 0.410 | 0.427 | 0.441 | 0.537 | 0.429 | 0.466 | 0.605 |
| ETTm2 | 0.327 | 0.358 | 0.323 | 0.335 | 0.353 | 0.368 | 0.343 | 0.488 |
| ETTh1 | 0.454 | 0.479 | 0.522 | 0.525 | 0.628 | 0.600 | 0.542 | 0.561 |
| ETTh2 | 0.388 | 0.433 | 0.394 | 0.421 | 0.465 | 0.538 | 0.431 | 0.475 |
| Weather | 0.268 | 0.272 | 0.273 | 0.275 | 0.301 | 0.283 | 0.279 | 0.324 |
Averages over the four horizons; lower is better.
| Dataset | T3Time MSE | Strongest baseline | Change in MSE |
|---|---|---|---|
| ETTm1 | 0.384 | TimeCMA · 0.396 | −3.0% |
| ETTm2 | 0.267 | TimeLLM · 0.274 | −2.6% |
| ETTh1 | 0.442 | TimeCMA · 0.472 | −6.4% |
| ETTh2 | 0.357 | TimeLLM · 0.382 | −6.5% |
| Weather | 0.226 | TimeCMA · 0.231 | −2.2% |
Lower MSE than the strongest baseline on all five datasets, by 4.1% on average: 8.0% lower MSE than TimeCMA, 12.3% lower than Time-LLM and 18.4% lower than GPT4TS on average.
On ETTm2, Time-LLM keeps a slightly lower MAE; T3Time has the best MAE on the other four.
All baselines · 5% of training data 5 datasets × 8 methods
MSE
| Dataset | T3Time | TimeCMA | TimeLLM | GPT4TS | TimesNet | DLinear | PatchTST | FEDformer |
|---|---|---|---|---|---|---|---|---|
| ETTm1 | 0.384 | 0.396 | 0.425 | 0.472 | 0.717 | 0.400 | 0.526 | 0.730 |
| ETTm2 | 0.267 | 0.329 | 0.274 | 0.308 | 0.344 | 0.399 | 0.314 | 0.381 |
| ETTh1 | 0.442 | 0.472 | 0.627 | 0.681 | 0.925 | 0.750 | 0.694 | 0.658 |
| ETTh2 | 0.357 | 0.395 | 0.382 | 0.400 | 0.439 | 0.694 | 0.827 | 0.463 |
| Weather | 0.226 | 0.231 | 0.260 | 0.263 | 0.298 | 0.263 | 0.269 | 0.309 |
MAE
| Dataset | T3Time | TimeCMA | TimeLLM | GPT4TS | TimesNet | DLinear | PatchTST | FEDformer |
|---|---|---|---|---|---|---|---|---|
| ETTm1 | 0.405 | 0.416 | 0.434 | 0.450 | 0.561 | 0.417 | 0.476 | 0.592 |
| ETTm2 | 0.330 | 0.367 | 0.323 | 0.346 | 0.372 | 0.426 | 0.352 | 0.404 |
| ETTh1 | 0.451 | 0.470 | 0.543 | 0.560 | 0.647 | 0.611 | 0.569 | 0.562 |
| ETTh2 | 0.403 | 0.430 | 0.418 | 0.433 | 0.448 | 0.577 | 0.615 | 0.454 |
| Weather | 0.269 | 0.273 | 0.309 | 0.301 | 0.318 | 0.308 | 0.303 | 0.353 |
Averages over the four horizons; lower is better.
Every component helps; the residual path most
Removing one component at a time under the long-term setup, the full model is best, or tied for best, in all 14 dataset–metric pairs of the design study.
Without the channel-wise residual, MSE rises by 9.7% on average and by 28% on ILI: aligned features alone lose information the raw temporal–spectral path still carries. Dropping the frequency branch costs 3.4%, most on Exchange (+5.9%) and ILI (+4.8%).
Values are relative MSE increases per dataset, averaged over the seven datasets of the design study and computed from its table below.
Design-variant table 7 datasets × 5 variants
MSE
| Dataset | T3Time | w/o frequency | w/o multi-head | w/o residual | w/o gating |
|---|---|---|---|---|---|
| ETTm1 | 0.372 | 0.381 | 0.374 | 0.404 | 0.373 |
| ETTm2 | 0.279 | 0.279 | 0.283 | 0.288 | 0.280 |
| ETTh1 | 0.418 | 0.433 | 0.421 | 0.433 | 0.425 |
| ETTh2 | 0.348 | 0.364 | 0.355 | 0.384 | 0.363 |
| Weather | 0.244 | 0.250 | 0.249 | 0.249 | 0.249 |
| ILI | 1.705 | 1.786 | 1.813 | 2.176 | 1.724 |
| Exchange | 0.353 | 0.374 | 0.378 | 0.396 | 0.373 |
MAE
| Dataset | T3Time | w/o frequency | w/o multi-head | w/o residual | w/o gating |
|---|---|---|---|---|---|
| ETTm1 | 0.393 | 0.394 | 0.395 | 0.413 | 0.396 |
| ETTm2 | 0.322 | 0.324 | 0.325 | 0.329 | 0.322 |
| ETTh1 | 0.430 | 0.433 | 0.432 | 0.443 | 0.432 |
| ETTh2 | 0.390 | 0.398 | 0.392 | 0.411 | 0.398 |
| Weather | 0.275 | 0.280 | 0.277 | 0.280 | 0.277 |
| ILI | 0.835 | 0.884 | 0.865 | 0.969 | 0.837 |
| Exchange | 0.401 | 0.410 | 0.413 | 0.427 | 0.412 |
Averages over the four horizons; lower is better.
What the three views learn
Figure 3. t-SNE of each modality's embeddings, coloured by dataset. Time-series and especially frequency embeddings separate the datasets cleanly, since each has its own periodicities. Prompt embeddings are more mixed, reflecting how similar the textual descriptions are. The forecast embeddings cluster much like the time-series ones.
Scope: all results use standard public benchmarks and a frozen GPT-2 for the prompt branch. Larger-scale pretraining and richer representations for each modality are left for future work.
Citation
If this work is useful to you, please cite it as:
@inproceedings{chowdhury2026t3time,
title = {{T3Time}: Tri-Modal Time Series Forecasting via Adaptive
Multi-Head Alignment and Residual Fusion},
author = {Chowdhury, Abdul Monaf and Akter, Rabeya and Arib, Safaeid Hossain},
booktitle = {Proceedings of the AAAI Conference on Artificial Intelligence},
volume = {40},
number = {25},
pages = {20597--20605},
year = {2026},
doi = {10.1609/aaai.v40i25.39196}
}