sceneforge results

/ihub/homedirs/svs_ald/sudhir/real2sim/work/lab
← all runs

lab

session /ihub/homedirs/svs_ald/sudhir/real2sim/captures/session_20260901_180746
work /ihub/homedirs/svs_ald/sudhir/real2sim/work/lab

Stages

10_pose — Pose

Camera poses from unposed images (VGGT / COLMAP / GLOMAP).

not run

20_scale — Metric scale

Converts the scale-free reconstruction to metres using the D455 depth map. This is what replaces Re3Sim's ArUco marker.

How to read it. `spread` is agreement between independent per-frame estimates: <5% consistent, 5-15% loose, >15% do not trust. `baseline_m` should match the real extent of the capture.
backendvggt
methoddense_depth
scale_m_per_unit1.691
spread0.011
verdictconsistent
frames_used8
frames_total8
baseline_m1.867
raw json
{
  "backend": "vggt",
  "method": "dense_depth",
  "scale_m_per_unit": 1.6912546750989792,
  "spread": 0.010579440281930415,
  "verdict": "consistent",
  "frames_used": 8,
  "frames_total": 8,
  "baseline_m": 1.8672663343004527,
  "per_frame": [
    {
      "scale": 1.6965607468549315,
      "median_scale": 1.6962303397381036,
      "n": 116663,
      "valid_frac": 0.699435238255114,
      "rel_mad": 0.021406807291139664,
      "trusted": true,
      "frame": 0
    },
    {
      "scale": 1.7054566941068936,
      "median_scale": 1.7041362693167823,
      "n": 110058,
      "valid_frac": 0.6598359672893834,
      "rel_mad": 0.0247384723749645,
      "trusted": true,
      "frame": 1
    },
    {
      "scale": 1.6859486033430269,
      "median_scale": 1.6855558896480523,
      "n": 109827,
      "valid_frac": 0.6584510419914147,
      "rel_mad": 0.02465091327927448,
      "trusted": true,
      "frame": 2
    },
    {
      "scale": 1.699843251628138,
      "median_scale": 1.7003353414145805,
      "n": 108512,
      "valid_frac": 0.6505671598839301,
      "rel_mad": 0.02623051791420815,
      "trusted": true,
      "frame": 3
    },
    {
      "scale": 1.675267505514713,
      "median_scale": 1.6725795148575342,
      "n": 81818,
      "valid_frac": 0.49052735077579795,
      "rel_mad": 0.03315821632069944,
      "trusted": true,
      "frame": 4
    },
    {
      "scale": 1.710191967653463,
      "median_scale": 1.7110296022256561,
      "n": 101526,
      "valid_frac": 0.6086836614786926,
      "rel_mad": 0.03245943658571427,
      "trusted": true,
      "frame": 5
    },
    {
      "scale": 1.6576749341617,
      "median_scale": 1.6531040452057408,
      "n": 101544,
      "valid_frac": 0.6087915777356772,
      "rel_mad": 0.025595928956704938,
      "trusted": true,
      "frame": 6
    },
    {
      "scale": 1.6663179035497446,
      "median_scale": 1.6633530717233862,
      "n": 115072,
      "valid_frac": 0.6898966402071992,
      "rel_mad": 0.02433516863681764,
      "trusted": true,
      "frame": 7
    }
  ]
}

30_ground — Ground plane / world frame

Fits the table plane and builds a world frame with +Z up and the surface at z=0, so MuJoCo gravity and object placement mean something.

How to read it. `plane.rms_m` under ~10 mm is a clean fit. Check `camera_height_range_m` against reality -- if the cameras are not plausibly above the table, the fitted plane is not the table.
scale_m_per_unit1.691
plane
offset-1.283
inlier_frac0.141
rms_m0.008
raw json
{
  "T_world_recon": [
    [
      0.0,
      1.01046374187277,
      -1.3562099403870456,
      0.0
    ],
    [
      1.6888148448167948,
      -0.07282165388056047,
      -0.05425682166030653,
      0.0
    ],
    [
      -0.09081187130864807,
      -1.3542534508468986,
      -1.0090060311726792,
      1.2827255784466411
    ],
    [
      0.0,
      0.0,
      0.0,
      1.0
    ]
  ],
  "T_rot_only": [
    [
      0.0,
      0.5974639755621864,
      -0.8018957525173874,
      0.0
    ],
    [
      0.9985573844565772,
      -0.04305776944935786,
      -0.03208081104469449,
      0.0
    ],
    [
      -0.05369497134033547,
      -0.8007389252406011,
      -0.5966020647444052,
      1.2827255784466411
    ],
    [
      0.0,
      0.0,
      0.0,
      1.0
    ]
  ],
  "scale_m_per_unit": 1.6912546750989792,
  "plane": {
    "normal": [
      0.05369497134033547,
      0.8007389252406011,
      0.5966020647444052
    ],
    "offset": -1.2827255784466411,
    "inlier_frac": 0.14065909627920553,
    "rms_m": 0.008398504844291953
  },
  "camera_height_range_m": [
    1.2825296438407872,
    1.3217632569496727
  ],
  "warnings": []
}

40_mask — Robot mask

SAM3 text-prompted segmentation of the robot, which must be kept out of the background splat (MuJoCo renders it from the URDF).

How to read it. `empty_masks` should be 0. `median_coverage` above ~50% means the prompt matched the whole scene rather than the robot.
promptrobotic arm, robot
frames8
median_coverage0.058
empty_masks0
raw json
{
  "prompt": "robotic arm, robot",
  "frames": 8,
  "median_coverage": 0.05776513671875,
  "empty_masks": 0,
  "warnings": []
}
000000.png0.0 MB ↓
000001.png0.0 MB ↓
000002.png0.0 MB ↓
000003.png0.0 MB ↓
000004.png0.0 MB ↓
000005.png0.0 MB ↓
000006.png0.0 MB ↓
000007.png0.0 MB ↓
overlay.png1.2 MB ↓

50_splat — Background splat

Gaussian splat trained on the masked images, initialised from the measured depth cloud rather than sparse SfM points.

How to read it. `eval` is on HELD-OUT views -- the honest number. Training-view PSNR only measures overfitting. `eval.depth.bias_mm` is the early warning for a wrong scale factor.
backendgsplat
ply/ihub/homedirs/svs_ald/sudhir/real2sim/work/lab/50_splat/background.ply
n_gaussians100,000
train_seconds22.600
peak_vram_gb0.310
iters600
cap_max200,000
sh_degree3
init_points100,000
train_views7
maskedTrue
eval
n_views1
lpips0.661
psnr15.395
ssim0.525
eval.depth
bias_mm-318.969
inlier_frac_10mm0.006
mean_abs_mm361.597
median_abs_mm275.031
n182,388.000
p95_abs_mm917.487
rms_mm455.307
valid_frac0.712
raw json
{
  "backend": "gsplat",
  "ply": "/ihub/homedirs/svs_ald/sudhir/real2sim/work/lab/50_splat/background.ply",
  "n_gaussians": 100000,
  "train_seconds": 22.6,
  "peak_vram_gb": 0.31,
  "eval": {
    "n_views": 1,
    "lpips": 0.6606054306030273,
    "psnr": 15.394587775008182,
    "ssim": 0.5245035290718079,
    "depth": {
      "bias_mm": -318.96876709955126,
      "inlier_frac_10mm": 0.006458758251639362,
      "mean_abs_mm": 361.5965368355053,
      "median_abs_mm": 275.03082156181335,
      "n": 182388.0,
      "p95_abs_mm": 917.4873560667038,
      "rms_mm": 455.30697381792334,
      "valid_frac": 0.712453125
    }
  },
  "iters": 600,
  "cap_max": 200000,
  "sh_degree": 3,
  "init_points": 100000,
  "train_views": 7,
  "masked": true
}

60_mesh — Static geometry

TSDF fusion of the depth maps into a mesh with real measured thickness, plus convex parts for collision.

not run

70_objects — Objects

Per-object mesh and physics. Mass is hand-tuned; inertia follows from the mesh at that mass.

not run

80_scene — MuJoCo scene

Assembles the MJCF: static geometry, objects, cameras at the real capture poses, robot.

not run

90_verify — Verification vs real photographs

Renders the sim, composites the splat behind it, and diffs against the actual photograph from that pose. The end-to-end check.

not run