LFG

LFG checkpoint that jointly predicts dense depth / 3D points, camera pose, per-point confidence, object segmentation (7 classes), and per-pixel motion from monocular driving video — no human labels.

Given 3 observed frames, the model predicts all outputs for those frames plus 3 future frames (6 frames total), matching the setting reported in the paper.

LFG: Self-Supervised 4D Learning from In-the-Wild Driving Videos — trained on 58M+ unlabeled dashcam driving frames with frozen teachers generating pseudo ground truth on the fly. See the project page, the paper, and the inference repo.

Model details

Architecture LFG (autoregressive transformer with future-frame prediction)
Parameters 1.22B (fp32)
Input / output 3 observed frames in; 3 observed + 3 future frames out
Encoder DINOv2
Heads point/depth, camera/pose, confidence, segmentation (7 classes), motion

Files

  • lfg_seg_motion_m3n3.pt — minimal inference checkpoint: model_state_dict (inference weights only), config (architecture settings), global_step. Loads with torch.load(..., weights_only=True).

Usage

Use the official inference repo:

python infer.py <video-or-image-dir> \
    --checkpoint lfg_seg_motion_m3n3.pt \
    --output-dir outputs/ \
    --save-visualizations

Outputs per window: points, local_points, camera_poses, conf, segmentation, motion.

Note on depth scale: point maps are predicted up to an unknown scale and shift. Align predictions to a metric reference (e.g. least-squares scale-and-shift against ground-truth depth) before computing metric errors.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Paper for AppliedIntuitionResearch/LFG