Open-Rig3R β€” Waymo (epoch 50)

Unofficial open reimplementation of Rig3R: a rig-aware transformer that takes multiple camera views and predicts per-view pointmaps plus pose and rig raymaps, so a whole camera rig is reconstructed jointly instead of one camera at a time.

Code: engichang1467/Open-Rig3R

This checkpoint is epoch 50 of 50 from a single run on a Waymo mini subset. It is a research checkpoint from a reimplementation, not a reproduction of the paper's published numbers β€” see Limitations before using it for anything.

Results

Held-out Waymo val split, at the end of training:

Metric Value
Pointmap L2 error (scale-normalized) 0.2159
Pose rotation error 1.440Β°
Rig rotation error 1.408Β°
Pose centre error 0.0156
Rig centre error 0.0161

Read the validation curve carefully

val/total rises from 0.395 (epoch 6) to 0.961 (epoch 50). That is not overfitting, and epoch 6 is not the better checkpoint.

The pointmap term is C * err - alpha * log(C), where C is the model's own predicted confidence. Over training C climbs from 1.2 to 9.98, saturating against the conf_max: 10.0 ceiling. The same geometric error therefore costs roughly 10x more at epoch 50 than at epoch 6. Meanwhile the unweighted error stays flat or improves (pointmap_err 0.279 β†’ 0.216) and both angular errors fall monotonically (14.10Β° β†’ 1.44Β°, 14.16Β° β†’ 1.41Β°).

epoch val/total pointmap pointmap_err conf_mean pose_deg rig_deg
6 0.395 0.253 0.2239 3.26 3.01 3.02
25 0.621 0.532 0.2187 6.86 1.63 1.59
50 0.961 0.880 0.2159 9.98 1.440 1.408

val/total is confounded by confidence saturation and should not be used for model selection on this run. Select on pointmap_err, pose_deg, and rig_deg β€” all of which plateau by roughly epoch 45 and are best at epoch 50.

Usage

This is a plain state_dict, not a transformers PreTrainedModel β€” there is no from_pretrained. Load it into the Rig3R class from the repo, with the same architecture arguments the checkpoint was trained with (the class defaults differ and will not load):

from huggingface_hub import hf_hub_download
from safetensors.torch import load_file
from models.rig3r import Rig3R

model = Rig3R(
    encoder_ckpt=None,        # weights come from the checkpoint below, not from DUSt3R
    img_size=128,
    patch_size=16,
    embed_dim=1024,
    num_decoder_layers=2,     # NOT the class default of 6
    num_heads=8,
    mlp_dim=4096,             # NOT the class default of 2048
    metadata_dropout=0.5,
)

path = hf_hub_download("mca183/rig3r-waymo", "model.safetensors")
model.load_state_dict(load_file(path))
model.eval()

# images: (B, V, 3, 128, 128) β€” V = n_frames * num_cameras, e.g. 2 * 5 = 10 for the full Waymo rig
# metadata is optional; the decoder is trained with per-field dropout so it runs without it
outputs = model(images, metadata=None)

forward returns a dict:

key shape notes
pointmap (B, V, H*W, 3) dense per-pixel 3D points
pointmap_conf (B, V, H*W, 1) per-pixel confidence β€” saturated, see Limitations
pose_raymap (B, V, P, 6) centre + unit direction per patch
rig_raymap (B, V, P, 6) centre + unit direction per patch
camera_center_pose (B, V, 3) one centre per view
camera_center_rig (B, V, 3) one centre per view
features (B, V, P, C) decoder patch features, for downstream heads

P = (128 / 16)^2 = 64 patches per view.

Architecture

Total parameters 340.2 M (373 tensors, fp32)
Encoder 303.2 M β€” DUSt3R ViT-L/16, frozen during training
Rig-aware decoder 37.0 M β€” 2 pre-norm transformer layers, 8 heads, MLP dim 4096
Heads pointmap_head, pose_raymap_head, rig_raymap_head
Embedding dim 1024
Input resolution 128 x 128, patch size 16

Encoder initialized from DUSt3R_ViTLarge_BaseDecoder_512_dpt and kept frozen, so only the 37.0 M decoder and heads were trained.

Training

Dataset Waymo Open Dataset, mini subset, full 5-camera rig (FRONT, FRONT_LEFT, FRONT_RIGHT, SIDE_LEFT, SIDE_RIGHT)
Views per sample 10 (n_frames: 2 x 5 cameras)
Epochs 50
Batch size 8
Optimizer AdamW, lr 1e-4, weight decay 0.01
Scheduler Cosine annealing, eta_min 1e-6
Precision bf16 autocast, no grad scaler
Loss weights w_point 1.0, w_pose 1.0, w_rig 1.0
Confidence regularizer alpha 0.2, beta 1.0, conf_max 10.0
Metadata dropout 0.5 per field (frame index exempt)
Seed 0
Hardware NVIDIA A100 80GB PCIe
Wall clock ~3h 20m

conf_max: 10.0 is a deliberate deviation from the paper: Eq. 3 leaves -alpha*log(C) unbounded below, so the ceiling floors the pointmap term at -alpha*log(10) = -0.46. Setting it to null restores Eq. 3 exactly.

Limitations

  • Reimplementation, not the paper. Unofficial; numbers here are not comparable to published Rig3R results.
  • Mini subset. Trained on a small Waymo subset, not the full dataset. Generalization is untested.
  • 128 x 128 input. Well below the paper's resolution; pointmap detail is correspondingly coarse.
  • Shallow decoder. 2 layers rather than 6, and the encoder is frozen throughout.
  • Waymo domain only. Driving scenes, one rig geometry. No indoor, handheld, or non-automotive evaluation.
  • Single run, no held-out test. Metrics are val-split only, one seed, no ablations.
  • Confidence is saturated. conf_mean sits at the conf_max ceiling of 10.0, so predicted confidence is not calibrated and should not be read as an uncertainty estimate.

Dataset terms

Trained on the Waymo Open Dataset, which carries its own license and terms of use. Using this model does not grant any rights to that data β€” obtain it from Waymo directly and comply with their terms.

License

MIT (see the source repo). Copyright 2025 Michael Chang. The DUSt3R encoder initialization and the Waymo Open Dataset carry their own separate licenses.

Citation

Rig3R (original paper):

@article{rig3r,
  title  = {Rig3R: Rig-Aware Conditioning for Learned 3D Reconstruction},
  year   = {2025}
}

This reimplementation:

@software{open_rig3r,
  author = {Chang, Michael},
  title  = {Open-Rig3R: An open reimplementation of Rig3R},
  url    = {https://github.com/engichang1467/Open-Rig3R},
  year   = {2025}
}
Downloads last month

-

Downloads are not tracked for this model. How to track
Safetensors
Model size
0.3B params
Tensor type
F32
Β·
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Evaluation results

  • Pose rotation error (deg) on Waymo Open Dataset (mini subset, held-out val split)
    self-reported
    1.440
  • Rig rotation error (deg) on Waymo Open Dataset (mini subset, held-out val split)
    self-reported
    1.407
  • Pointmap L2 error (scale-normalized) on Waymo Open Dataset (mini subset, held-out val split)
    self-reported
    0.216