IEEE ISMAR 2026

SnapPhysics: A Physics-Aware Scene Graph
from a Single View for Interactive Mixed Reality Scenes

1KAIST, UVR Lab    2University of Stuttgart, VISUS
*Co-corresponding authors

SnapPhysics reconstructs 3D objects and estimates their physical properties
(mass, friction, and center of gravity) from a single image.

SnapPhysics teaser: from a single RGB image to interaction-ready 3D objects with physical properties and a physics-aware scene graph

Abstract

Existing methods estimate physical properties from video at high computational cost or use vision-language models (VLMs) without geometric and relational context. SnapPhysics addresses these limitations by combining 3D reconstruction and spatial alignment with a physics-aware scene graph for VLM-based reasoning. It improves scene-level F-Score by 18.6% on 3D-FRONT and, on real scenes, reduces mass error by up to 20.5% and improves log-scale correlation by up to 19.6% over VLM-only estimation, enabling physically interactive MR without manual parameter tuning.

3D Reconstruction
+18.6%
Scene-level F-Score vs.
best learning-based method
(3D-FRONT Dataset)
Mass Estimation
−20.5%
Mass error (mALDE) vs.
VLM-only estimation
+19.6%
Log-scale correlation r²ls
for mass ranking
Physically coherent interactions in MR

Interactive MR Demos

Estimated mass, friction, and CoG drive physically plausible interactions overlaid on the real scene.

Method

SnapPhysics is a two-stage, training-free pipeline:
(A) single-view 3D reconstruction with depth-anchored alignment, (B) physical property inference and physics-aware scene graph.

SnapPhysics pipeline overview

Pipeline overview. (A1) Shape (SAM3D) + metric depth (DA3). (A2) Depth-anchored alignment. (B1) Contextual hierarchy. (B2) VLM-based property estimation into a physics-aware scene graph.

Stage A · Single-View 3D Reconstruction

Depth-anchored alignment registers SAM3D's arbitrary-scale shapes into DA3's metric depth in three steps.

(1) Scene-level alignment. Closed-form Sim(3) + ICP with depth-ratio scale correction.

(2) Per-object alignment. Differentiable refinement with depth, silhouette, Chamfer, and splat losses.

(3) Confidence-guided verification. Low-confidence poses are re-initialized and re-optimized.

Depth-Anchored Alignment

Depth-anchored alignment comparison

(b) DA3: metric depth, no instances. (c) SAM3D: instances, no metric scale. (d) SnapPhysics: both, the geometric cue for plausible physics inference.

Stage B · Physical Property Inference

A contextual hierarchy is built first, then a VLM estimates mass, friction, and CoG from each object's crop, geometry, and relations in support order.

Contextual Hierarchy Construction

Scene graph comparison with ConceptGraphs

Unlike ConceptGraphs (left), objects are split into macro supporters and micro objects, keeping only support (on), front/behind, and left/right relations (right). Support edges set the estimation order and propagate materials.

Physics-Aware Scene Graph

Nodes carry geometry, mass, CoG, and material, and friction lives on contact edges. Built once per scene, directly consumable by physics engines.

Results

1. Scene Reconstruction

Quantitative Comparison

On 3D-FRONT (1,000 scenes), SnapPhysics achieves the best scene-level accuracy despite being training-free.

MethodCD-S ↓F-Score-S ↑CD-O ↓F-Score-O ↑
Total3DSupervised0.27032.900.17936.38
SSRSupervised0.14039.760.17037.79
DiffCADSupervised0.11743.580.19037.45
Gen3DSRTraining-free0.12340.070.15738.11
MIDISupervised0.08050.190.10353.58
SnapPhysics (Ours)Training-free0.07859.530.16843.95

Best per column in bold. Ours is training-free yet improves F-Score-S by 18.6% over MIDI, trained on 3D-FRONT.

Qualitative Comparison

On 3D-FRONT (top) and real-world snapshotsReal captured dataset: four scene categories with per-object ground-truth mass (bottom), SnapPhysics preserves both layout and shape (blue boxes), while baselines show scale misalignment and missing structures (red).

Qualitative comparison on 3D-FRONT and real-world scenes
2. Mass Estimation

Quantitative Comparison

On real captured scenesReal captured dataset: four scene categories with per-object ground-truth mass with ground-truth mass, metric geometry (+DA3) and scene-graph context (+SG) each help, and their combination works best. Each cell reports the mean over 100 evaluations per image (Table 4 in the paper), with the best per column in bold.

Dining Room · 5 scenes / 10 obj
ConditionGeometryContextmAPE ↓mMnRE ↑mALDE ↓r²ls ↑
VLM Only——59.90.6030.5500.901
SAM3Darbitrary scale—49.50.5280.8090.864
+DA3metric—59.30.6040.5400.908
+SGarbitrary scalescene graph53.30.5350.7540.878
Ours (Two-Stage)metricscene graph52.40.6760.4370.928

BibTeX

@inproceedings{kang2026snapphysics,
  title     = {{SnapPhysics}: A Physics-Aware Scene Graph from a Single View
               for Interactive Mixed Reality Scenes},
  author    = {Kang, Suji and Kim, Seok-Young and Kim, Young Bin and Ha, Taewook
               and Schmalstieg, Dieter and Mori, Shohei and Woo, Woontack},
  booktitle = {IEEE International Symposium on Mixed and Augmented Reality (ISMAR)},
  year      = {2026}
}