British Machine Vision Conference / 2026

AMGF3D

Adaptive Multi-scale Gated Fusion for Robust Indoor RGB-D 3D Object Detection

Faisal Ahmed John Quarles Kevin Desai

The University of Texas at San Antonio

P2–P5Adaptive scales
2Fusion stages
6Matched seeds
+0.6%Parameters

Abstract

Fully convolutional sparse-voxel detectors are efficient for indoor 3D object detection, but RGB-D detection pipelines can degrade when geometry, appearance, or camera-depth alignment is imperfect. Projection-based 2D-3D fusion can help, yet simple fusion baselines often assume that projected correspondences are reliable and that one fixed image-feature scale is suitable for all object sizes. These assumptions do not hold for indoor scenes with depth discontinuities, occlusion, and large scale variation. We propose Adaptive Multi-scale Gated Fusion (AMGF3D), a selective RGB-D fusion framework for sparse-voxel detectors. AMGF3D aggregates features from a 2D Feature Pyramid Network (P2–P5) using per-voxel scale weights conditioned on local 3D context, and uses a lightweight feature-conditioned gate to modulate projected image evidence before injecting it into the 3D stream. Fusion is applied at two backbone stages to combine lower-level appearance detail with higher-level semantic cues while retaining single-pass inference. Beyond clean accuracy, we evaluate robustness under controlled, multi-seed perturbations. On SUN RGB-D and ARKitScenes, AMGF3D is significantly more robust than additive fusion under the two pre-specified primary perturbations, RGB patch dropout and calibration rotation (six matched seeds, with Wilcoxon and 95% CI support), with gains that persist at mAP@0.50. Across all three benchmarks we observe no statistically significant underperformance within the tested suite. In a SUN RGB-D cross-method comparison, it is best or tied-best on the perturbations that meaningfully degrade accuracy. AMGF3D remains competitive on clean ScanNet, SUN RGB-D, and ARKitScenes evaluation, improves most clearly on small objects, and adds +0.6% parameters and a modest ~9% runtime cost over additive fusion.

Key Contributions

Prior projection-based fusion ties each 3D location to a single, fixed-scale 2D descriptor and injects it unconditionally. AMGF3D removes both assumptions.

01

Adaptive Multi-scale Gated Fusion

A per-voxel, multi-scale, gated RGB-D fusion module that injects 2D pyramid features into a sparse-voxel backbone at two stages, conditioned on local 3D context.

02

Multi-seed Robustness Study

A statistically paired evaluation across three benchmarks — six matched seeds per dataset, pre-specified primary perturbations, and an exact-Wilcoxon with 95% CI criterion.

03

Component-level Attribution

A per-component ablation under perturbation isolates the source: adaptive multi-scale aggregation and dual-stage placement each carry a significant share, gating a smaller one.

04

Small Objects, Low Overhead

Competitive on clean evaluation at both IoU thresholds, with the clearest gains on small objects, at +0.6% parameters and ~9% runtime over additive fusion.

Key Idea

What changes relative to additive fusion (TR3D+FF), at a single fusion stage.

Additive fusion versus AMGF3D at one fusion stage

Selective, not unconditional

The additive baseline samples a single fixed FPN level and adds it to every active voxel with the same weight.

AMGF3D samples P2–P5 and mixes them with per-voxel weights predicted from the voxel's own 3D feature, then gates the resulting residual on the joint geometry–appearance pair.

A voxel whose projected evidence disagrees with its geometry can therefore suppress that evidence, instead of receiving it unconditionally.

Architecture

A sparse 3D stream processes colored point clouds while a frozen ResNet-FPN extracts image features; fusion happens at two backbone stages.

AMGF3D architecture overview

Projected FPN descriptors are fused into active voxels at backbone layers 1 and 3 using adaptive scale weighting and per-voxel gating, before the TR3D neck and detection head.

1
Project
Voxel centers to the image plane
→
2
Sample
All four levels P2–P5
→
3
Weight
Softmax scale mixture per voxel
→
4
Gate
Sigmoid on geometry & appearance
→
5
Fuse
Gated residual update

The module is applied at an early stage (s = 1) and a later stage (s = 3), retaining single-pass inference.

Robustness to Input Perturbations

The central result. Absolute mAP@0.25 for the additive-fusion baseline TR3D+FF and AMGF3D, averaged over six matched training seeds per dataset.

Perturbation (level) ScanNet
TR3D+FF
ScanNet
AMGF3D
SUN RGB-D
TR3D+FF
SUN RGB-D
AMGF3D
ARKitScenes
TR3D+FF
ARKitScenes
AMGF3D
Clean74.575.668.769.176.177.4
Point subsample (50%)69.670.168.668.575.275.8
RGB Gaussian (σ=0.1)71.070.868.668.575.375.9
RGB dropout (50%)71.171.764.564.975.376.0
Calib. rotation (3°)66.767.265.065.961.764.3
Point-coord. noise (σ=0.10 m)0.71.13.38.14.13.5

Bold marks AMGF3D entries significantly more robust than TR3D+FF under the paired-seed criterion. RGB dropout and calibration rotation are the two pre-specified primary perturbations; point-coordinate noise is a secondary descriptive check.

Robustness summary: paired gains and cross-method AP retained

(a) Paired difference Δ = AMGF3D − TR3D+FF in mAP@0.25 at representative severities, averaged over six matched seeds; filled markers are significant, hollow are not. (b) SUN RGB-D cross-method robustness as AP retained relative to clean.

Clean Detection

mAP@0.25 / @0.50 in the FCAF3D/TR3D Best (Mean) convention. AMGF3D stays competitive on clean evaluation without being the headline claim.

Method Venue Inputs ScanNet
@0.25
ScanNet
@0.50
SUN RGB-D
@0.25
SUN RGB-D
@0.50
ARKit
@0.25
ARKit
@0.50
Sparse-voxel reproductions (matched pipeline)
FCAF3DECCV'22PC71.5 (70.7)57.3 (56.0)64.2 (63.8)48.9 (48.2)72.3 (71.7)60.5 (59.6)
TR3DICIP'23PC72.9 (72.0)59.3 (57.4)67.1 (66.3)50.4 (49.6)75.4 (74.9)61.3 (60.7)
TR3D+FFICIP'23PC+RGB75.3 (74.5)60.1 (58.9)69.4 (68.7)53.4 (52.2)76.7 (76.1)62.5 (61.7)
Quoted from source papers
SPGroup3DAAAI'24PC74.3 (73.5)59.6 (58.3)65.4 (64.8)47.1 (46.4)——
V-DETRICLR'24PC77.865.967.550.4——
UniDet3D†AAAI'25PC77.5 (77.0)66.1 (65.0)————
AMGF3D (Ours)—PC+RGB 76.1 (75.6)60.1 (59.0) 69.7 (69.1)53.6 (52.3) 77.8 (77.4)63.9 (62.7)

The sparse-voxel rows (FCAF3D, TR3D, TR3D+FF, AMGF3D) are our own reproductions under one matched pipeline — same frozen 2D backbone, splits, schedule and evaluation code — and are not quotations of those papers' printed numbers, so small differences are expected. Other rows are quoted from their source papers. Bold marks the best entry per column. †For UniDet3D we quote its single-dataset ScanNet result (trained on ScanNet alone) for a like-for-like comparison; its headline numbers use joint training over six datasets, and its ARKitScenes protocol is not comparable to ours.

Per-Component Ablation

On ARKitScenes, mAP@0.25 over six matched training seeds. Each row removes one component from the full model.

Variant Clean Calib. 3° Drop. 50%
TR3D+FF (additive)76.161.775.3
early-only (−late stage)77.161.975.7
late-only (−early stage)76.862.575.4
single-scale (−multi-scale)77.262.775.9
no-gate (−gate)77.163.475.8
AMGF3D (full)77.464.376.0

Under calibration rotation every appearance-side component contributes: removing multi-scale aggregation costs 1.6 mAP@0.25, the early stage 1.8, and the late stage 2.4 (six per-seed deltas of one sign, 95% CI excludes zero, exact Wilcoxon p = 0.03125). Under clean input and RGB dropout all component deltas fall within seed noise (≤ 0.6), which is why these choices are motivated by robustness and not by clean accuracy.

Qualitative Results

Per-scene predictions on ARKitScenes, SUN RGB-D, and ScanNet.

Qualitative comparison across ARKitScenes, SUN RGB-D and ScanNet

Columns: paired RGB frame, bird's-eye view ground truth, TR3D+FF, and AMGF3D. Box colors denote TP/FP/FN at IoU ≥ 0.25. On ScanNet and ARKitScenes the single posed RGB frame covers only part of the reconstructed scene, and the BEV is cropped for legibility.

Citation

If you find this work useful, please cite:

@inproceedings{ahmed2026amgf3d, title = {AMGF3D: Adaptive Multi-scale Gated Fusion for Robust Indoor RGB-D 3D Object Detection}, author = {Ahmed, Faisal and Quarles, John and Desai, Kevin}, booktitle = {British Machine Vision Conference (BMVC)}, year = {2026} }