MC-STL: Mask- and Contrast-Enhanced Spatio-Temporal Learning for Urban Flow Prediction

Xu Zhang, Yongshun Gong, Xinxin Zhang, Xiaoming Wu, Chengqi Zhang, Xiangjun Dong

CIKM 2023, pp. 3298–3307 · Free to read in the ACM Digital Library

Also indexed in: Semantic Scholar · OpenAlex · dblp · Wikidata · ACM Digital Library

TL;DR. MC-STL pre-trains two encoders for urban flow prediction: a ViT encoder learns to reconstruct regions masked at different timestamps, and its attention weights also build a GCN adjacency matrix; a global-local cross-attention encoder learns a temporal-order contrastive task. It reaches RMSE 14.53 on full TaxiBJ.

Quick facts

Method nameMC-STL; MCSTL (name of the official code repository, github.com/CodeZx6/MCSTL, and of this site's page)
AuthorsXu Zhang, Yongshun Gong, Xinxin Zhang, Xiaoming Wu, Chengqi Zhang, Xiangjun Dong
DOI10.1145/3583780.3614958
TaskUrban flow prediction (UFP): next-step forecasting of region-level inflow and outflow maps
Mask-enhanced branchViT encoder (Enc_M) pre-trained to reconstruct regions masked at different timestamps; masking rate 0.75 gave the lowest RMSE among 0.25, 0.5 and 0.75
Graph processing module (GPM)Enc_M attention weights, summed over heads and batch, thresholded and normalised, form the adjacency matrix of a two-layer GCN that extracts important regions
Contrast-enhanced branchGlobal-local spatio-temporal cross-attention encoder (Enc_C) pre-trained with a temporal-order triplet task: shuffled anchor as negative, same-order unit from another period as positive
PredictionMask branch (Enc_M with GPM) and contrast branch (Enc_C with GRU) each end in a ResNet decoder; outputs are summed. Embedded holidays and weather are combined with the flow input.
DatasetsTaxiBJ (Beijing taxis, 30-minute interval, 2×32×32, periods P1-P4 plus full set) and BikeNYC (New York bikes, 1-hour interval, 2×16×8): six evaluation sets
EvaluationRMSE and MAE against nine baselines: HA, RNN, ARIMA, ST-ResNet, AGCRN, Multi-STGCnet, ATFM, ST-SSL (BikeNYC and TaxiBJ P3 only) and ST-GSP
Headline resultLowest RMSE and MAE on all six sets. Full TaxiBJ 14.53 / 8.70 (ST-GSP 14.72 / 8.79); BikeNYC 4.81 / 2.38 (Multi-STGCnet 5.52 / 2.47)
AuthorshipXu Zhang and Yongshun Gong contributed equally; Xiangjun Dong is the corresponding author
ConferenceCIKM '23, the 32nd ACM International Conference on Information and Knowledge Management, Birmingham, UK, 21-25 October 2023; pages 3298-3307; ISBN 979-8-4007-0124-5
CodeOfficial PyTorch code at github.com/CodeZx6/MCSTL: evaluation scripts for TaxiBJ and BikeNYC, test data for BikeNYC and TaxiBJ P1, and a BikeNYC checkpoint
AccessFree to read in the ACM Digital Library since ACM opened its whole library on 1 January 2026; publication rights licensed to ACM, no Creative Commons licence recorded
BibTeX keyzhang2023mcstl

Main results

SettingResultCompared with
TaxiBJ, full dataset (last 28 days as test set)RMSE = 14.53, MAE = 8.70
“Our proposed MC-STL model achieved an RMSE of 14.53 and an MAE of 8.70 on the full TaxiBJ dataset.”
Strongest baseline ST-GSP: RMSE 14.72, MAE 8.79
BikeNYC (last 10 days as test set)RMSE = 4.81, MAE = 2.38
“Table 2: The average RMSE and MAE on datasets. The best results are bold and the second best are underlined.”
Strongest baseline Multi-STGCnet: RMSE 5.52, MAE 2.47
TaxiBJ periods P1, P2, P3, P4 (last 10% as test set)RMSE / MAE = 15.94 / 9.55 (P1), 16.63 / 10.12 (P2), 17.11 / 10.14 (P3), 15.56 / 9.05 (P4)
“MC-STL remained superior in prediction outcomes across all six datasets.”
Strongest baseline ATFM on each period: 16.09 / 9.80, 16.73 / 10.33, 17.77 / 10.53, 15.67 / 9.31
TaxiBJ, masking algorithm used in pre-trainingSpatio-temporal random masking (MC-STL-STmask): RMSE = 14.53, MAE = 8.70
“Table 4: Results of two masking algorithms on the TaxiBJ.”
Conventional image random masking (MC-STL-mask): RMSE 14.91, MAE 8.86
TaxiBJ ablation without external factors (MC-STL-w/o-Ext)RMSE = 17.01, MAE = 9.58
“We found that removing external factors resulted in an increase of 2.48 in RMSE”
Full MC-STL: RMSE 14.53, MAE 8.70
TaxiBJ ablation without any pre-training (MC-STL-w/o-Pre)RMSE = 14.65, MAE = 8.75
“MC-STL-w/o-Pre: Removes all pre-training related to mask and contrast-enhanced spatio-temporal learning.”
Full MC-STL: RMSE 14.53, MAE 8.70. The pre-training gain (0.12 RMSE) is smaller than the effect of removing the contrast branch (15.10) or the mask branch (14.77)

Quoted text is verbatim from the paper.

Key points

Abstract

As a critical mission of intelligent transportation systems, urban flow prediction (UFP) benefits in many city services including trip planning, congestion control, and public safety. Despite the achievements of previous studies, limited efforts have been observed on simultaneous investigation of the heterogeneity in both space and time aspects. That is, regional correlations would be variable at different timestamps. In this paper, we propose a spatio-temporal learning framework with mask and contrast enhancements to capture spatio-temporal variabilities among city regions. We devise a mask-enhanced pre-training task to learn latent correlations across the spatial and temporal dimensions, and then a graph-based method is developed to extract the significance of regions by using the inter-regional attention weights. To further acquire contrastive correlations of regions, we elaborate a pre-trained contrastive learning task with the global-local cross-attention mechanism. Thereafter, two well-trained encoders have strong capability to capture latent spatio-temporal representations for the flow forecasting with time-varying. Extensive experiments conducted on real-world urban flow datasets demonstrate that our method compares favorably with other state-of-the-art models.

Frequently asked questions

Is MCSTL the same as MC-STL?

Yes. The CIKM 2023 paper names the model MC-STL (mask- and contrast-enhanced spatio-temporal learning). Its official code repository is github.com/CodeZx6/MCSTL.

What problem does MC-STL address?

The paper notes that most existing urban flow models do not consider spatial and temporal heterogeneity simultaneously, although correlations between regions vary across timestamps. MC-STL uses two pre-training tasks, mask-enhanced and contrast-enhanced, to learn these time-varying spatio-temporal correlations.

How does MC-STL's masking differ from image-style random masking?

Image random masking hides a region at every time step, which the paper argues captures spatial but not temporal relationships between regions. MC-STL masks different regions at different timestamps and reconstructs them from the remaining regions. On TaxiBJ this gives RMSE 14.53 / MAE 8.70, against 14.91 / 8.86 with image random masking.

How does MC-STL build its graph?

A graph processing module (GPM) sums the multi-head self-attention weights of the mask-pre-trained ViT encoder over heads and over the batch, then thresholds and normalises them. The result is the adjacency matrix of a two-layer GCN that extracts important regions.

Which datasets, metrics and baselines does MC-STL use?

MC-STL is evaluated on TaxiBJ (periods P1-P4 and the full set) and BikeNYC, six evaluation sets in total, with RMSE and MAE. The nine baselines are HA, RNN, ARIMA, ST-ResNet, AGCRN, Multi-STGCnet, ATFM, ST-SSL and ST-GSP; ST-SSL is reported only on BikeNYC and TaxiBJ (P3).

How well does MC-STL perform?

In Table 2, MC-STL has the lowest RMSE and MAE on all six evaluation sets. On full TaxiBJ it reaches RMSE 14.53 / MAE 8.70 against 14.72 / 8.79 for ST-GSP, and on BikeNYC 4.81 / 2.38 against 5.52 / 2.47 for Multi-STGCnet.

How do I cite MC-STL?

Cite it as: Zhang, X., Gong, Y., Zhang, X., Wu, X., Zhang, C., & Dong, X. (2023). Mask- and Contrast-Enhanced Spatio-Temporal Learning for Urban Flow Prediction. In Proceedings of the 32nd ACM International Conference on Information and Knowledge Management (pp. 3298–3307). ACM. https://doi.org/10.1145/3583780.3614958 BibTeX is available at https://codezx6.github.io/papers/mcstl.html#cite and in https://codezx6.github.io/publications.bib.

Is the code for MC-STL available?

Yes, at https://github.com/CodeZx6/MCSTL.

中文摘要

掩码与对比增强的时空学习用于城市流量预测(MC-STL)

现有城市流量预测方法大多未同时考虑空间与时间异质性。MC-STL 预训练两个编码器:ViT 编码器学习重建不同时刻被遮掩的区域,其注意力权重构成两层图卷积网络的邻接矩阵;全局-局部时空交叉注意力编码器学习时序对比任务。二者融合预测下一时刻流量,在 TaxiBJ 与 BikeNYC 六个评测集上均优于九个基线,完整 TaxiBJ 上 RMSE 为 14.53。

关键词:城市流量预测;时空预训练;时空随机遮掩;对比学习;图卷积网络(GCN);时空异质性

Keywords

urban flow prediction spatio-temporal predictive modeling spatio-temporal pre-training mask-enhanced learning contrastive learning traffic prediction graph convolutional network (GCN)

Also referred to as: MCSTL; Mask- and Contrast-Enhanced Spatio-Temporal Learning; mask-enhanced spatio-temporal learning; spatio-temporal (ST) random masking; graph processing module (GPM); global-local spatio-temporal cross-attention.

Research area map and related search terms

Field path (broad to narrow): artificial intelligence › machine learning › data mining › spatio-temporal data mining › time series forecasting › intelligent transportation systems › smart cities › urban computing › traffic prediction › crowd flow prediction › urban flow prediction › self-supervised spatio-temporal pre-training › MC-STL

Task

urban flow prediction crowd flow prediction citywide crowd flow forecasting inflow and outflow prediction grid-based traffic flow forecasting spatio-temporal forecasting

Method

masked spatio-temporal pre-training self-supervised pre-training for traffic prediction spatio-temporal contrastive learning Vision Transformer (ViT) encoder attention-derived graph GCN global-local cross-attention

Datasets

TaxiBJ BikeNYC

Compared with

ST-ResNet AGCRN ATFM ST-SSL ST-GSP Multi-STGCnet

中文

城市流量预测 人群流量预测 交通流量预测 流入流出预测 时空预测 掩码预训练 时空对比学习 自监督预训练

Related areas

traffic forecasting spatio-temporal graph neural networks graph neural networks Transformers for time series masked autoencoders self-supervised learning contrastive learning representation learning mobility modeling human mobility spatio-temporal heterogeneity citywide prediction grid-based prediction deep spatio-temporal residual networks

Applications

traffic management congestion control trip planning public safety ride-hailing dispatch bike-sharing rebalancing

Related benchmarks (not used in this paper)

TaxiNYC BikeDC BikeChicago PEMS-BAY METR-LA PeMS04 PeMS08

中文 (扩展)

交通预测 时空数据挖掘 时间序列预测 智能交通 智慧城市 城市计算 时空图神经网络 图神经网络 掩码自编码器 自监督学习 表征学习 人类移动性 网格流量预测 交通管理 共享单车调度

Terms are grouped by role. "Related areas", "Applications" and "Related benchmarks (not used in this paper)" describe the surrounding field, not results of this paper. See the site-wide research area map.

Related papers

Urban flow prediction

Other topics

Cite this paper

@inproceedings{zhang2023mcstl,
  title        = {Mask- and Contrast-Enhanced Spatio-Temporal Learning for Urban Flow Prediction},
  author       = {Zhang, Xu and Gong, Yongshun and Zhang, Xinxin and Wu, Xiaoming and Zhang, Chengqi and Dong, Xiangjun},
  booktitle    = {Proceedings of the 32nd ACM International Conference on Information and Knowledge Management},
  year         = {2023},
  series       = {CIKM '23},
  pages        = {3298--3307},
  publisher    = {ACM},
  doi          = {10.1145/3583780.3614958},
  url          = {https://doi.org/10.1145/3583780.3614958}
}
Zhang, X., Gong, Y., Zhang, X., Wu, X., Zhang, C., & Dong, X. (2023). Mask- and Contrast-Enhanced Spatio-Temporal Learning for Urban Flow Prediction. In Proceedings of the 32nd ACM International Conference on Information and Knowledge Management (pp. 3298–3307). ACM. https://doi.org/10.1145/3583780.3614958