URMDet-SimFire: Trustworthy multimodal detection with uncertainty and reliability modeling for fire and smoke analytics
Zhonghua Dang, Xu Zhang
Array, vol. 31, article 101124 (2026) · Open access, CC BY 4.0
Also indexed in: Semantic Scholar · OpenAlex · Publisher page
Quick facts
| Method name | URMDet-SimFire |
|---|---|
| Authors | Zhonghua Dang, Xu Zhang |
| DOI | 10.1016/j.array.2026.101124 |
| Task | Multimodal fire and smoke object detection (two classes: fire, smoke), validated in simulation |
| Architecture | YOLOv8-style visual branch with edge-aware attention and scale-adaptive fusion, plus four non-visual encoders: thermal cues, environmental sensors, smoke-diffusion attributes, textual scene priors |
| Trust-modeling modules | Uncertainty calibration from classification evidence and localization variance; reliability-guided fusion reweighting each modality by noise level, missingness and cross-modal consistency |
| Datasets (synthetic) | SimFire-Main: 600 scenes of 256x256 RGB images (seed 42), 1 to 3 objects each. SimFire-Corruption: illumination, missing modality, occlusion, sensor noise at five severities |
| Metrics | Precision, recall, F1, AP50, AP75, mAP_proxy = 0.65·AP50 + 0.35·AP75 (a simulation proxy, not COCO mAP), mean IoU, ECE |
| Baselines | Vision-only YOLOv8-sim, multimodal late fusion with fixed weights, uncertainty-aware fusion without reliability weighting |
| Headline result (simulation) | F1 0.4638 and mAP_proxy 0.3231 vs vision-only 0.4348 and 0.2754; ECE 0.2445 is not the lowest |
| Validation setting | Simulation only: no real fire or smoke data for training, validation or testing; single random seed (42) |
| Code and data | No public repository; simulation code and generated outputs available from the corresponding author on reasonable request |
| Corresponding author | Xu Zhang |
| Venue | Array (Elsevier), volume 31, September 2026, article 101124; online 7 August 2026 |
| Open access | Gold open access under CC BY 4.0 |
| BibTeX key | dang2026urmdet |
Main results
| Setting | Result | Compared with |
|---|---|---|
| SimFire-Main (simulation), main comparison, Table 1 | F1 = 0.4638; mAP_proxy = 0.3231“Compared with the vision-only YOLOv8-sim baseline, URMDet-SimFire improves F1-score from 0.4348 to 0.4638 and mAP_proxy from 0.2754 to 0.3231.” | Vision-only YOLOv8-sim: F1 0.4348, mAP_proxy 0.2754 |
| SimFire-Main (simulation), naive multimodal late fusion vs vision-only, Table 1 | Late fusion F1 = 0.3831; mAP_proxy = 0.2211“The multimodal late fusion baseline performs worse than the vision-only baseline. Its F1-score decreases from 0.4348 to 0.3831, and mAP_proxy decreases from 0.2754 to 0.2211.” | Vision-only YOLOv8-sim: F1 0.4348, mAP_proxy 0.2754 |
| SimFire-Main (simulation), calibration, Table 1 | ECE = 0.2445 (lower is better)“The proposed method does not achieve the lowest ECE: its ECE of 0.2445 is higher than those of the vision-only and late-fusion baselines, although lower than that of uncertainty-aware fusion.” | Vision-only 0.2210 and late fusion 0.2290 are lower; uncertainty-aware fusion 0.2623 is higher |
| Ablation (simulation, Table 2): without scale-adaptive fusion | F1 = 0.3686; mAP_proxy = 0.1911“Removing scale-adaptive fusion causes the largest performance degradation among the lightweight ablation variants. The F1-score drops to 0.3686, and mAP_proxy drops to 0.1911.” | Full model in Table 2: F1 0.4583, mAP_proxy 0.3341 |
| Simulated efficiency profile, Table 4 ('Simulated efficiency comparison of different methods.') | 7.6M parameters; 20.4 GFLOPs; 21.4 ms latency; 46.6 FPS (simulated)“The proposed URMDet-SimFire uses 7.6M parameters and 20.4 GFLOPs, with a latency of 21.4 ms and an FPS of 46.6.” | Vision-only YOLOv8-sim: 11.2M parameters, 28.4 GFLOPs, 26.5 ms, 37.7 FPS (simulated) |
Quoted text is verbatim from the paper.
Key points
- In simulation (SimFire-Main), fixed-weight multimodal late fusion scored below vision-only detection: F1 0.3831 vs 0.4348, mAP_proxy 0.2211 vs 0.2754.
- The full model, uncertainty calibration plus reliability-guided fusion, reached F1 0.4638 and mAP_proxy 0.3231, the top detection scores of the four compared methods. Uncertainty-aware fusion without reliability weighting reached mAP_proxy 0.3022.
- Its expected calibration error (ECE 0.2445) is not the lowest: the vision-only (0.2210) and late-fusion (0.2290) baselines score lower, which the authors describe as a calibration trade-off.
- In the ablation, removing scale-adaptive fusion caused the largest drop (mAP_proxy 0.3341 to 0.1911).
- All results come from a controllable simulator with one random seed (42). No real fire or smoke data is used, and no real neural-network training is conducted.
Abstract
Trustworthy multimodal perception is essential for safety-critical analytics, where detection systems must contend with noisy visual observations, missing sensor signals, smoke-like interference, and distributional shifts across environments. Vision-only detectors degrade under precisely these conditions, while naively combining additional modalities can amplify corrupted signals. We propose URMDet-SimFire, an uncertainty- and reliability-aware multimodal object detection framework that addresses this fragility by explicitly modeling the trustworthiness of each input source at inference time. We couple a YOLOv8-style visual branch with four non-visual modality encoders for thermal cues, environmental sensors, smoke-diffusion attributes, and textual scene priors. Two trust-modeling components are then introduced: an uncertainty calibration module that adjusts detection confidence using classification evidence and localization variance, and a reliability-guided fusion module that dynamically reweights each modality according to its current noise level, missingness, and cross-modal consistency. Because collecting real industrial fire data at scale is dangerous and costly, we validate the framework in a controllable simulator that reproduces typical visual degradations, sensor corruption, modality dropouts, and occlusion patterns. We evaluate it through main comparisons, ablations of each trust-modeling component, robustness under four corruption types, and efficiency profiling. The results show that naive multimodal late fusion can underperform vision-only detection when modalities are unreliable, that uncertainty calibration alone is insufficient, and that the proposed reliability-guided design recovers detection quality. The framework offers a reproducible simulation-based paradigm for trust-aware multimodal media analytics.
Frequently asked questions
Was URMDet-SimFire trained or tested on real fire data?
No. All experiments use two synthetic datasets, SimFire-Main (600 scenes) and SimFire-Corruption, generated by a controllable simulator, and no real fire or smoke dataset is used for training, validation or testing. The authors say the results should be read as algorithm validation evidence, not deployment proof.
Which modalities does URMDet-SimFire use?
Five simulated modalities: synthetic RGB images for a YOLOv8-style visual branch, plus thermal cues, environmental sensor readings, smoke-diffusion attributes and textual scene priors. In the implemented code, the non-visual modalities are normalized scalar or vector scores rather than high-dimensional real sensor streams.
Does adding more modalities always help fire and smoke detection?
Not in this study. In the simulator, fixed-weight multimodal late fusion scored below vision-only detection (F1 0.3831 vs 0.4348). The full URMDet-SimFire, which adds uncertainty calibration and reliability-guided fusion, reached F1 0.4638.
Does URMDet-SimFire improve calibration (ECE)?
Not over every baseline. Its ECE of 0.2445 is higher than the vision-only (0.2210) and late-fusion (0.2290) baselines, though lower than uncertainty-aware fusion (0.2623). The authors describe this as a calibration trade-off that warrants further study.
Which component matters most in the ablation?
Scale-adaptive fusion. In the simulated ablation (Table 2), removing it caused the largest drop among the lightweight variants, with F1 falling from 0.4583 to 0.3686 and mAP_proxy from 0.3341 to 0.1911.
Is the URMDet-SimFire code available?
The paper links no public repository. The simulation code and generated outputs are available from the corresponding author upon reasonable request.
How do I cite URMDet-SimFire?
Cite it as: Dang, Z., & Zhang, X. (2026). URMDet-SimFire: Trustworthy multimodal detection with uncertainty and reliability modeling for fire and smoke analytics. Array, 31, Article 101124. https://doi.org/10.1016/j.array.2026.101124 BibTeX is available at https://codezx6.github.io/papers/urmdet-simfire.html#cite and in https://codezx6.github.io/publications.bib.
中文摘要
URMDet-SimFire:面向火灾与烟雾分析的不确定性与可靠性建模的可信多模态检测
URMDet-SimFire 是面向火灾与烟雾的不确定性与可靠性感知多模态目标检测框架,结合 YOLOv8 风格视觉分支与热线索、环境传感器、烟雾扩散属性、文本场景先验四个非视觉模态。不确定性校准调整检测置信度,可靠性引导融合按噪声、缺失与跨模态一致性重加权各模态。实验仅在仿真器中完成:F1 由纯视觉基线 0.4348 升至 0.4638,ECE 未达最低。
关键词:火灾与烟雾检测;多模态融合;不确定性量化;可靠性学习;可信人工智能;目标检测
Keywords
trustworthy AI multimodal fusion uncertainty quantification reliability learning safety-critical perception object detection fire and smoke analytics fire and smoke detection
Also referred to as: uncertainty- and reliability-aware multimodal object detection; trustworthy multimodal object detection; reliability-guided multimodal fusion; uncertainty-aware prediction calibration; SimFire-Main; SimFire-Corruption.
Research area map and related search terms
Field path (broad to narrow): artificial intelligence › computer vision › object detection › multimodal perception › trustworthy AI › uncertainty-aware multimodal object detection › fire and smoke detection › URMDet-SimFire
Task
fire detection smoke detection fire and smoke detection multimodal object detection trustworthy perception robust detection with missing modalities
Method
uncertainty calibration reliability-weighted fusion sensor fusion expected calibration error (ECE) YOLOv8-style detector simulation-based evaluation
Datasets
SimFire-Main (synthetic) SimFire-Corruption (synthetic)
中文
火灾检测 烟雾检测 火灾烟雾检测 多模态目标检测 不确定性估计 可信感知 多传感器融合
Related areas
multimodal fusion sensor fusion uncertainty quantification model calibration robustness missing modalities safety-critical AI thermal imaging IoT sensors simulation and synthetic data anomaly detection YOLO detectors
Applications
industrial fire monitoring wildfire and smoke alarms smart buildings safety surveillance emergency response
Related benchmarks (not used in this paper)
D-Fire FLAME FASDD Foggia fire video dataset
中文 (扩展)
可信人工智能 多模态感知 多传感器融合 不确定性量化 模型校准 鲁棒性 模态缺失 安全关键AI 热成像 物联网传感器 合成数据 工业火灾监测 消防预警 智慧楼宇
Terms are grouped by role. "Related areas", "Applications" and "Related benchmarks (not used in this paper)" describe the surrounding field, not results of this paper. See the site-wide research area map.
Related papers
Object detection
- Automatic visual recognition for leaf disease based on enhanced attention mechanism (PeerJ Computer Science, vol. 10, article e2365, 2024)
Other topics
- DUET: Unified Dual-Space Emotion Control for Diffusion and Flow-Matching Driven Text-to-Speech (2026)
- PhysioSER: Learning Physiology-Informed Vocal Spectrotemporal Representations for Speech Emotion Recognition (2026)
- DSTCN: Exploiting dynamic spatio-temporal correlations for origin-destination demand prediction (2026)
- S2CMEN: A Mutually Enhancement Network for Superpixel Segmentation and Classification of Hyperspectral Image (2025)
- MR-UFP: Enhancing urban flow prediction via mutual reinforcement with multi-scale regional information (2025)
- BiST-IF: Enhancing origin–destination flow prediction via bi-directional spatio-temporal inference and interconnected feature evolution (2025)
- S3CFSL: Spatial-Spectral–Semantic Cross-Domain Few-Shot Learning for Hyperspectral Image Classification (2024)
- ST-FCL: Spatio-temporal fusion and contrastive learning for urban flow prediction (2023)
- MC-STL: Mask- and Contrast-Enhanced Spatio-Temporal Learning for Urban Flow Prediction (2023)
Cite this paper
@article{dang2026urmdet,
title = {{URMDet-SimFire}: Trustworthy multimodal detection with uncertainty and reliability modeling for fire and smoke analytics},
author = {Dang, Zhonghua and Zhang, Xu},
journal = {Array},
year = {2026},
volume = {31},
pages = {101124},
doi = {10.1016/j.array.2026.101124},
issn = {2590-0056},
url = {https://doi.org/10.1016/j.array.2026.101124}
}Dang, Z., & Zhang, X. (2026). URMDet-SimFire: Trustworthy multimodal detection with uncertainty and reliability modeling for fire and smoke analytics. Array, 31, Article 101124. https://doi.org/10.1016/j.array.2026.101124