LFG
LFG checkpoint that jointly predicts dense depth / 3D points, camera pose, per-point confidence, object segmentation (7 classes), and per-pixel motion from monocular driving video — no human labels.
Given 3 observed frames, the model predicts all outputs for those frames plus 3 future frames (6 frames total), matching the setting reported in the paper.
LFG: Self-Supervised 4D Learning from In-the-Wild Driving Videos — trained on 58M+ unlabeled dashcam driving frames with frozen teachers generating pseudo ground truth on the fly. See the project page, the paper, and the inference repo.
Model details
| Architecture | LFG (autoregressive transformer with future-frame prediction) |
| Parameters | 1.22B (fp32) |
| Input / output | 3 observed frames in; 3 observed + 3 future frames out |
| Encoder | DINOv2 |
| Heads | point/depth, camera/pose, confidence, segmentation (7 classes), motion |
Files
lfg_seg_motion_m3n3.pt— minimal inference checkpoint:model_state_dict(inference weights only),config(architecture settings),global_step. Loads withtorch.load(..., weights_only=True).
Usage
Use the official inference repo:
python infer.py <video-or-image-dir> \
--checkpoint lfg_seg_motion_m3n3.pt \
--output-dir outputs/ \
--save-visualizations
Outputs per window: points, local_points, camera_poses, conf, segmentation, motion.
Note on depth scale: point maps are predicted up to an unknown scale and shift. Align predictions to a metric reference (e.g. least-squares scale-and-shift against ground-truth depth) before computing metric errors.