URMDet-SimFire: Trustworthy multimodal detection with uncertainty and reliability modeling for fire and smoke analytics

Zhonghua Dang, Xu Zhang

Array, vol. 31, article 101124 (2026) · Open access, CC BY 4.0

Also indexed in: Semantic Scholar · OpenAlex · Publisher page

TL;DR. URMDet-SimFire is an uncertainty- and reliability-aware multimodal object detection framework for fire and smoke that reweights each input modality by its estimated reliability and calibrates each detection's confidence by its uncertainty. In simulation only, it raised F1 from 0.4348 (vision-only baseline) to 0.4638.

Quick facts

Method nameURMDet-SimFire
AuthorsZhonghua Dang, Xu Zhang
DOI10.1016/j.array.2026.101124
TaskMultimodal fire and smoke object detection (two classes: fire, smoke), validated in simulation
ArchitectureYOLOv8-style visual branch with edge-aware attention and scale-adaptive fusion, plus four non-visual encoders: thermal cues, environmental sensors, smoke-diffusion attributes, textual scene priors
Trust-modeling modulesUncertainty calibration from classification evidence and localization variance; reliability-guided fusion reweighting each modality by noise level, missingness and cross-modal consistency
Datasets (synthetic)SimFire-Main: 600 scenes of 256x256 RGB images (seed 42), 1 to 3 objects each. SimFire-Corruption: illumination, missing modality, occlusion, sensor noise at five severities
MetricsPrecision, recall, F1, AP50, AP75, mAP_proxy = 0.65·AP50 + 0.35·AP75 (a simulation proxy, not COCO mAP), mean IoU, ECE
BaselinesVision-only YOLOv8-sim, multimodal late fusion with fixed weights, uncertainty-aware fusion without reliability weighting
Headline result (simulation)F1 0.4638 and mAP_proxy 0.3231 vs vision-only 0.4348 and 0.2754; ECE 0.2445 is not the lowest
Validation settingSimulation only: no real fire or smoke data for training, validation or testing; single random seed (42)
Code and dataNo public repository; simulation code and generated outputs available from the corresponding author on reasonable request
Corresponding authorXu Zhang
VenueArray (Elsevier), volume 31, September 2026, article 101124; online 7 August 2026
Open accessGold open access under CC BY 4.0
BibTeX keydang2026urmdet

Main results

SettingResultCompared with
SimFire-Main (simulation), main comparison, Table 1F1 = 0.4638; mAP_proxy = 0.3231
“Compared with the vision-only YOLOv8-sim baseline, URMDet-SimFire improves F1-score from 0.4348 to 0.4638 and mAP_proxy from 0.2754 to 0.3231.”
Vision-only YOLOv8-sim: F1 0.4348, mAP_proxy 0.2754
SimFire-Main (simulation), naive multimodal late fusion vs vision-only, Table 1Late fusion F1 = 0.3831; mAP_proxy = 0.2211
“The multimodal late fusion baseline performs worse than the vision-only baseline. Its F1-score decreases from 0.4348 to 0.3831, and mAP_proxy decreases from 0.2754 to 0.2211.”
Vision-only YOLOv8-sim: F1 0.4348, mAP_proxy 0.2754
SimFire-Main (simulation), calibration, Table 1ECE = 0.2445 (lower is better)
“The proposed method does not achieve the lowest ECE: its ECE of 0.2445 is higher than those of the vision-only and late-fusion baselines, although lower than that of uncertainty-aware fusion.”
Vision-only 0.2210 and late fusion 0.2290 are lower; uncertainty-aware fusion 0.2623 is higher
Ablation (simulation, Table 2): without scale-adaptive fusionF1 = 0.3686; mAP_proxy = 0.1911
“Removing scale-adaptive fusion causes the largest performance degradation among the lightweight ablation variants. The F1-score drops to 0.3686, and mAP_proxy drops to 0.1911.”
Full model in Table 2: F1 0.4583, mAP_proxy 0.3341
Simulated efficiency profile, Table 4 ('Simulated efficiency comparison of different methods.')7.6M parameters; 20.4 GFLOPs; 21.4 ms latency; 46.6 FPS (simulated)
“The proposed URMDet-SimFire uses 7.6M parameters and 20.4 GFLOPs, with a latency of 21.4 ms and an FPS of 46.6.”
Vision-only YOLOv8-sim: 11.2M parameters, 28.4 GFLOPs, 26.5 ms, 37.7 FPS (simulated)

Quoted text is verbatim from the paper.

Key points

Abstract

Trustworthy multimodal perception is essential for safety-critical analytics, where detection systems must contend with noisy visual observations, missing sensor signals, smoke-like interference, and distributional shifts across environments. Vision-only detectors degrade under precisely these conditions, while naively combining additional modalities can amplify corrupted signals. We propose URMDet-SimFire, an uncertainty- and reliability-aware multimodal object detection framework that addresses this fragility by explicitly modeling the trustworthiness of each input source at inference time. We couple a YOLOv8-style visual branch with four non-visual modality encoders for thermal cues, environmental sensors, smoke-diffusion attributes, and textual scene priors. Two trust-modeling components are then introduced: an uncertainty calibration module that adjusts detection confidence using classification evidence and localization variance, and a reliability-guided fusion module that dynamically reweights each modality according to its current noise level, missingness, and cross-modal consistency. Because collecting real industrial fire data at scale is dangerous and costly, we validate the framework in a controllable simulator that reproduces typical visual degradations, sensor corruption, modality dropouts, and occlusion patterns. We evaluate it through main comparisons, ablations of each trust-modeling component, robustness under four corruption types, and efficiency profiling. The results show that naive multimodal late fusion can underperform vision-only detection when modalities are unreliable, that uncertainty calibration alone is insufficient, and that the proposed reliability-guided design recovers detection quality. The framework offers a reproducible simulation-based paradigm for trust-aware multimodal media analytics.

Frequently asked questions

Was URMDet-SimFire trained or tested on real fire data?

No. All experiments use two synthetic datasets, SimFire-Main (600 scenes) and SimFire-Corruption, generated by a controllable simulator, and no real fire or smoke dataset is used for training, validation or testing. The authors say the results should be read as algorithm validation evidence, not deployment proof.

Which modalities does URMDet-SimFire use?

Five simulated modalities: synthetic RGB images for a YOLOv8-style visual branch, plus thermal cues, environmental sensor readings, smoke-diffusion attributes and textual scene priors. In the implemented code, the non-visual modalities are normalized scalar or vector scores rather than high-dimensional real sensor streams.

Does adding more modalities always help fire and smoke detection?

Not in this study. In the simulator, fixed-weight multimodal late fusion scored below vision-only detection (F1 0.3831 vs 0.4348). The full URMDet-SimFire, which adds uncertainty calibration and reliability-guided fusion, reached F1 0.4638.

Does URMDet-SimFire improve calibration (ECE)?

Not over every baseline. Its ECE of 0.2445 is higher than the vision-only (0.2210) and late-fusion (0.2290) baselines, though lower than uncertainty-aware fusion (0.2623). The authors describe this as a calibration trade-off that warrants further study.

Which component matters most in the ablation?

Scale-adaptive fusion. In the simulated ablation (Table 2), removing it caused the largest drop among the lightweight variants, with F1 falling from 0.4583 to 0.3686 and mAP_proxy from 0.3341 to 0.1911.

Is the URMDet-SimFire code available?

The paper links no public repository. The simulation code and generated outputs are available from the corresponding author upon reasonable request.

How do I cite URMDet-SimFire?

Cite it as: Dang, Z., & Zhang, X. (2026). URMDet-SimFire: Trustworthy multimodal detection with uncertainty and reliability modeling for fire and smoke analytics. Array, 31, Article 101124. https://doi.org/10.1016/j.array.2026.101124 BibTeX is available at https://codezx6.github.io/papers/urmdet-simfire.html#cite and in https://codezx6.github.io/publications.bib.

中文摘要

URMDet-SimFire:面向火灾与烟雾分析的不确定性与可靠性建模的可信多模态检测

URMDet-SimFire 是面向火灾与烟雾的不确定性与可靠性感知多模态目标检测框架,结合 YOLOv8 风格视觉分支与热线索、环境传感器、烟雾扩散属性、文本场景先验四个非视觉模态。不确定性校准调整检测置信度,可靠性引导融合按噪声、缺失与跨模态一致性重加权各模态。实验仅在仿真器中完成:F1 由纯视觉基线 0.4348 升至 0.4638,ECE 未达最低。

关键词:火灾与烟雾检测;多模态融合;不确定性量化;可靠性学习;可信人工智能;目标检测

Keywords

trustworthy AI multimodal fusion uncertainty quantification reliability learning safety-critical perception object detection fire and smoke analytics fire and smoke detection

Also referred to as: uncertainty- and reliability-aware multimodal object detection; trustworthy multimodal object detection; reliability-guided multimodal fusion; uncertainty-aware prediction calibration; SimFire-Main; SimFire-Corruption.

Research area map and related search terms

Field path (broad to narrow): artificial intelligence › computer vision › object detection › multimodal perception › trustworthy AI › uncertainty-aware multimodal object detection › fire and smoke detection › URMDet-SimFire

Task

fire detection smoke detection fire and smoke detection multimodal object detection trustworthy perception robust detection with missing modalities

Method

uncertainty calibration reliability-weighted fusion sensor fusion expected calibration error (ECE) YOLOv8-style detector simulation-based evaluation

Datasets

SimFire-Main (synthetic) SimFire-Corruption (synthetic)

中文

火灾检测 烟雾检测 火灾烟雾检测 多模态目标检测 不确定性估计 可信感知 多传感器融合

Related areas

multimodal fusion sensor fusion uncertainty quantification model calibration robustness missing modalities safety-critical AI thermal imaging IoT sensors simulation and synthetic data anomaly detection YOLO detectors

Applications

industrial fire monitoring wildfire and smoke alarms smart buildings safety surveillance emergency response

Related benchmarks (not used in this paper)

D-Fire FLAME FASDD Foggia fire video dataset

中文 (扩展)

可信人工智能 多模态感知 多传感器融合 不确定性量化 模型校准 鲁棒性 模态缺失 安全关键AI 热成像 物联网传感器 合成数据 工业火灾监测 消防预警 智慧楼宇

Terms are grouped by role. "Related areas", "Applications" and "Related benchmarks (not used in this paper)" describe the surrounding field, not results of this paper. See the site-wide research area map.

Related papers

Object detection

Other topics

Cite this paper

@article{dang2026urmdet,
  title        = {{URMDet-SimFire}: Trustworthy multimodal detection with uncertainty and reliability modeling for fire and smoke analytics},
  author       = {Dang, Zhonghua and Zhang, Xu},
  journal      = {Array},
  year         = {2026},
  volume       = {31},
  pages        = {101124},
  doi          = {10.1016/j.array.2026.101124},
  issn         = {2590-0056},
  url          = {https://doi.org/10.1016/j.array.2026.101124}
}
Dang, Z., & Zhang, X. (2026). URMDet-SimFire: Trustworthy multimodal detection with uncertainty and reliability modeling for fire and smoke analytics. Array, 31, Article 101124. https://doi.org/10.1016/j.array.2026.101124