Existing methods estimate physical properties from video at high computational cost or use vision-language models (VLMs) without geometric and relational context. SnapPhysics addresses these limitations by combining 3D reconstruction and spatial alignment with a physics-aware scene graph for VLM-based reasoning. It improves scene-level F-Score by 18.6% on 3D-FRONT and, on real scenes, reduces mass error by up to 20.5% and improves log-scale correlation by up to 19.6% over VLM-only estimation, enabling physically interactive MR without manual parameter tuning.
Estimated mass, friction, and CoG drive physically plausible interactions overlaid on the real scene.
SnapPhysics is a two-stage, training-free pipeline:
(A) single-view 3D reconstruction with depth-anchored alignment, (B) physical property inference and physics-aware scene graph.

Pipeline overview. (A1) Shape (SAM3D) + metric depth (DA3). (A2) Depth-anchored alignment. (B1) Contextual hierarchy. (B2) VLM-based property estimation into a physics-aware scene graph.
Depth-anchored alignment registers SAM3D's arbitrary-scale shapes into DA3's metric depth in three steps.
(1) Scene-level alignment. Closed-form Sim(3) + ICP with depth-ratio scale correction.
(2) Per-object alignment. Differentiable refinement with depth, silhouette, Chamfer, and splat losses.
(3) Confidence-guided verification. Low-confidence poses are re-initialized and re-optimized.

(b) DA3: metric depth, no instances. (c) SAM3D: instances, no metric scale. (d) SnapPhysics: both, the geometric cue for plausible physics inference.
A contextual hierarchy is built first, then a VLM estimates mass, friction, and CoG from each object's crop, geometry, and relations in support order.

Unlike ConceptGraphs (left), objects are split into macro supporters and micro objects, keeping only support (on), front/behind, and left/right relations (right). Support edges set the estimation order and propagate materials.
Nodes carry geometry, mass, CoG, and material, and friction lives on contact edges. Built once per scene, directly consumable by physics engines.
On 3D-FRONT (1,000 scenes), SnapPhysics achieves the best scene-level accuracy despite being training-free.
| Method | CD-S ↓ | F-Score-S ↑ | CD-O ↓ | F-Score-O ↑ |
|---|---|---|---|---|
| Total3DSupervised | 0.270 | 32.90 | 0.179 | 36.38 |
| SSRSupervised | 0.140 | 39.76 | 0.170 | 37.79 |
| DiffCADSupervised | 0.117 | 43.58 | 0.190 | 37.45 |
| Gen3DSRTraining-free | 0.123 | 40.07 | 0.157 | 38.11 |
| MIDISupervised | 0.080 | 50.19 | 0.103 | 53.58 |
| SnapPhysics (Ours)Training-free | 0.078 | 59.53 | 0.168 | 43.95 |
Best per column in bold. Ours is training-free yet improves F-Score-S by 18.6% over MIDI, trained on 3D-FRONT.
On 3D-FRONT (top) and real-world snapshots
(bottom), SnapPhysics preserves both layout and shape (blue boxes), while baselines show scale misalignment and missing structures (red).

On real captured scenes
with ground-truth mass, metric geometry (+DA3) and scene-graph context (+SG) each help, and their combination works best. Each cell reports the mean over 100 evaluations per image (Table 4 in the paper), with the best per column in bold. *Desk Items contains only objects under 1 kg.
| Condition | Geometry | Context | mAPE ↓ | mMnRE ↑ | mALDE ↓ | r²ls ↑ |
|---|---|---|---|---|---|---|
| VLM Only | — | — | 59.9 | 0.603 | 0.550 | 0.901 |
| SAM3D | arbitrary scale | — | 49.5 | 0.528 | 0.809 | 0.864 |
| +DA3 | metric | — | 59.3 | 0.604 | 0.540 | 0.908 |
| +SG | arbitrary scale | scene graph | 53.3 | 0.535 | 0.754 | 0.878 |
| Ours (Two-Stage) | metric | scene graph | 52.4 | 0.676 | 0.437 | 0.928 |
| Condition | Geometry | Context | mAPE ↓ | mMnRE ↑ | mALDE ↓ | r²ls ↑ |
|---|---|---|---|---|---|---|
| VLM Only | — | — | 102.6 | 0.633 | 0.576 | 0.809 |
| SAM3D | arbitrary scale | — | 99.4 | 0.629 | 0.570 | 0.807 |
| +DA3 | metric | — | 91.1 | 0.649 | 0.534 | 0.876 |
| +SG | arbitrary scale | scene graph | 92.9 | 0.626 | 0.557 | 0.839 |
| Ours (Two-Stage) | metric | scene graph | 87.4 | 0.659 | 0.517 | 0.883 |
| Condition | Geometry | Context | mAPE ↓ | mMnRE ↑ | mALDE ↓ | r²ls ↑ |
|---|---|---|---|---|---|---|
| VLM Only | — | — | 41.6 | 0.701 | 0.397 | 0.603 |
| SAM3D | arbitrary scale | — | 40.4 | 0.616 | 0.508 | 0.836 |
| +DA3 | metric | — | 44.2 | 0.681 | 0.473 | 0.522 |
| +SG | arbitrary scale | scene graph | 43.1 | 0.686 | 0.463 | 0.447 |
| Ours (Two-Stage) | metric | scene graph | 35.9 | 0.735 | 0.333 | 0.721 |
| Condition | Geometry | Context | mAPE ↓ | mMnRE ↑ | mALDE ↓ | r²ls ↑ |
|---|---|---|---|---|---|---|
| VLM Only | — | — | 262.9 | 0.579 | 0.827 | 0.858 |
| SAM3D | arbitrary scale | — | 367.2 | 0.588 | 0.880 | 0.914 |
| +DA3 | metric | — | 213.5 | 0.570 | 0.792 | 0.918 |
| +SG | arbitrary scale | scene graph | 282.2 | 0.542 | 0.887 | 0.917 |
| Ours (Two-Stage) | metric | scene graph | 168.1 | 0.517 | 0.795 | 0.959 |
@inproceedings{kang2026snapphysics,
title = {{SnapPhysics}: A Physics-Aware Scene Graph from a Single View
for Interactive Mixed Reality Scenes},
author = {Kang, Suji and Kim, Seok-Young and Kim, Young Bin and Ha, Taewook
and Schmalstieg, Dieter and Mori, Shohei and Woo, Woontack},
booktitle = {IEEE International Symposium on Mixed and Augmented Reality (ISMAR)},
year = {2026}
}