Open-Rig3R β Waymo (epoch 50)
Unofficial open reimplementation of Rig3R: a rig-aware transformer that takes multiple camera views and predicts per-view pointmaps plus pose and rig raymaps, so a whole camera rig is reconstructed jointly instead of one camera at a time.
Code: engichang1467/Open-Rig3R
This checkpoint is epoch 50 of 50 from a single run on a Waymo mini subset. It is a research checkpoint from a reimplementation, not a reproduction of the paper's published numbers β see Limitations before using it for anything.
Results
Held-out Waymo val split, at the end of training:
| Metric | Value |
|---|---|
| Pointmap L2 error (scale-normalized) | 0.2159 |
| Pose rotation error | 1.440Β° |
| Rig rotation error | 1.408Β° |
| Pose centre error | 0.0156 |
| Rig centre error | 0.0161 |
Read the validation curve carefully
val/total rises from 0.395 (epoch 6) to 0.961 (epoch 50). That is not overfitting,
and epoch 6 is not the better checkpoint.
The pointmap term is C * err - alpha * log(C), where C is the model's own predicted
confidence. Over training C climbs from 1.2 to 9.98, saturating against the
conf_max: 10.0 ceiling. The same geometric error therefore costs roughly 10x more at
epoch 50 than at epoch 6. Meanwhile the unweighted error stays flat or improves
(pointmap_err 0.279 β 0.216) and both angular errors fall monotonically
(14.10Β° β 1.44Β°, 14.16Β° β 1.41Β°).
| epoch | val/total | pointmap | pointmap_err | conf_mean | pose_deg | rig_deg |
|---|---|---|---|---|---|---|
| 6 | 0.395 | 0.253 | 0.2239 | 3.26 | 3.01 | 3.02 |
| 25 | 0.621 | 0.532 | 0.2187 | 6.86 | 1.63 | 1.59 |
| 50 | 0.961 | 0.880 | 0.2159 | 9.98 | 1.440 | 1.408 |
val/total is confounded by confidence saturation and should not be used for model
selection on this run. Select on pointmap_err, pose_deg, and rig_deg β all of which
plateau by roughly epoch 45 and are best at epoch 50.
Usage
This is a plain state_dict, not a transformers PreTrainedModel β there is no
from_pretrained. Load it into the Rig3R class from the repo, with the same
architecture arguments the checkpoint was trained with (the class defaults differ and
will not load):
from huggingface_hub import hf_hub_download
from safetensors.torch import load_file
from models.rig3r import Rig3R
model = Rig3R(
encoder_ckpt=None, # weights come from the checkpoint below, not from DUSt3R
img_size=128,
patch_size=16,
embed_dim=1024,
num_decoder_layers=2, # NOT the class default of 6
num_heads=8,
mlp_dim=4096, # NOT the class default of 2048
metadata_dropout=0.5,
)
path = hf_hub_download("mca183/rig3r-waymo", "model.safetensors")
model.load_state_dict(load_file(path))
model.eval()
# images: (B, V, 3, 128, 128) β V = n_frames * num_cameras, e.g. 2 * 5 = 10 for the full Waymo rig
# metadata is optional; the decoder is trained with per-field dropout so it runs without it
outputs = model(images, metadata=None)
forward returns a dict:
| key | shape | notes |
|---|---|---|
pointmap |
(B, V, H*W, 3) |
dense per-pixel 3D points |
pointmap_conf |
(B, V, H*W, 1) |
per-pixel confidence β saturated, see Limitations |
pose_raymap |
(B, V, P, 6) |
centre + unit direction per patch |
rig_raymap |
(B, V, P, 6) |
centre + unit direction per patch |
camera_center_pose |
(B, V, 3) |
one centre per view |
camera_center_rig |
(B, V, 3) |
one centre per view |
features |
(B, V, P, C) |
decoder patch features, for downstream heads |
P = (128 / 16)^2 = 64 patches per view.
Architecture
| Total parameters | 340.2 M (373 tensors, fp32) |
| Encoder | 303.2 M β DUSt3R ViT-L/16, frozen during training |
| Rig-aware decoder | 37.0 M β 2 pre-norm transformer layers, 8 heads, MLP dim 4096 |
| Heads | pointmap_head, pose_raymap_head, rig_raymap_head |
| Embedding dim | 1024 |
| Input resolution | 128 x 128, patch size 16 |
Encoder initialized from DUSt3R_ViTLarge_BaseDecoder_512_dpt and kept frozen, so only
the 37.0 M decoder and heads were trained.
Training
| Dataset | Waymo Open Dataset, mini subset, full 5-camera rig (FRONT, FRONT_LEFT, FRONT_RIGHT, SIDE_LEFT, SIDE_RIGHT) |
| Views per sample | 10 (n_frames: 2 x 5 cameras) |
| Epochs | 50 |
| Batch size | 8 |
| Optimizer | AdamW, lr 1e-4, weight decay 0.01 |
| Scheduler | Cosine annealing, eta_min 1e-6 |
| Precision | bf16 autocast, no grad scaler |
| Loss weights | w_point 1.0, w_pose 1.0, w_rig 1.0 |
| Confidence regularizer | alpha 0.2, beta 1.0, conf_max 10.0 |
| Metadata dropout | 0.5 per field (frame index exempt) |
| Seed | 0 |
| Hardware | NVIDIA A100 80GB PCIe |
| Wall clock | ~3h 20m |
conf_max: 10.0 is a deliberate deviation from the paper: Eq. 3 leaves -alpha*log(C)
unbounded below, so the ceiling floors the pointmap term at -alpha*log(10) = -0.46.
Setting it to null restores Eq. 3 exactly.
Limitations
- Reimplementation, not the paper. Unofficial; numbers here are not comparable to published Rig3R results.
- Mini subset. Trained on a small Waymo subset, not the full dataset. Generalization is untested.
- 128 x 128 input. Well below the paper's resolution; pointmap detail is correspondingly coarse.
- Shallow decoder. 2 layers rather than 6, and the encoder is frozen throughout.
- Waymo domain only. Driving scenes, one rig geometry. No indoor, handheld, or non-automotive evaluation.
- Single run, no held-out test. Metrics are val-split only, one seed, no ablations.
- Confidence is saturated.
conf_meansits at theconf_maxceiling of 10.0, so predicted confidence is not calibrated and should not be read as an uncertainty estimate.
Dataset terms
Trained on the Waymo Open Dataset, which carries its own license and terms of use. Using this model does not grant any rights to that data β obtain it from Waymo directly and comply with their terms.
License
MIT (see the source repo). Copyright 2025 Michael Chang. The DUSt3R encoder initialization and the Waymo Open Dataset carry their own separate licenses.
Citation
Rig3R (original paper):
@article{rig3r,
title = {Rig3R: Rig-Aware Conditioning for Learned 3D Reconstruction},
year = {2025}
}
This reimplementation:
@software{open_rig3r,
author = {Chang, Michael},
title = {Open-Rig3R: An open reimplementation of Rig3R},
url = {https://github.com/engichang1467/Open-Rig3R},
year = {2025}
}
Evaluation results
- Pose rotation error (deg) on Waymo Open Dataset (mini subset, held-out val split)self-reported1.440
- Rig rotation error (deg) on Waymo Open Dataset (mini subset, held-out val split)self-reported1.407
- Pointmap L2 error (scale-normalized) on Waymo Open Dataset (mini subset, held-out val split)self-reported0.216