Instructions to use anhdao69/SimpleMemVLN-R2R-RxR15deg-FullContext-CandidateLogits with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use anhdao69/SimpleMemVLN-R2R-RxR15deg-FullContext-CandidateLogits with Transformers:
# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("anhdao69/SimpleMemVLN-R2R-RxR15deg-FullContext-CandidateLogits", device_map="auto") - Notebooks
- Google Colab
- Kaggle
SimpleMemVLN R2R + RxR_15deg FullContext candidate-logits
Not yet evaluated in Habitat. Training loss is not navigation success. This is NOT the qwen_text FullContext-B checkpoint.
Snapshot and recipe
epoch-1/ is the mid-schedule snapshot at step 3852.
epoch-2/ is the completed two-epoch model at step 7704. Both are retained.
Joint R2R and English-guide RxR_15deg full episodes. Global batch 8 = 4 H100
GPUs x one episode per rank x gradient accumulation 2. LR 5e-6, warmup 232,
cosine decay to 10% of peak, weight decay 0.01, seed 429. BF16, ZeRO-2,
gradient checkpointing and long-sequence activation offload; vision frozen.
Initialized from Qwen/Qwen3.5-4B revision
851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a.
Training source commit: 8a6a28acabd43d378e610345b2ebbe24547c37f8 on
SimpleMemVLN streaming_logits; exact source and dataset hashes are included.
Navigation contract
Serializer vln_candidate_logits_v1, full-context causal attention with
persistent GDN state. Four selected pretrained LM-head rows, trainable at the
backbone LR (lm_rows_trainable); no random classifier and no autoregressive
action generation. Candidate labels A/B/C/D (IDs 32/33/34/35) map to
MOVE_FORWARD/TURN_LEFT/TURN_RIGHT/STOP and Habitat IDs 1/2/3/0.
Explicit candidate-token feedback; unweighted four-way cross entropy.
Epoch-1 action-weighted loss: 0.3663298; epoch-2: 0.1044625. Not comparable numerically to text CE.
Loading and integrity
Use the SimpleMemVLN candidate-aware wrapper loader
qwen_vl.train.vln_runtime.load_checkpoint(epoch_directory, base_model_path)
from streaming_logits, with the pinned pretrained base snapshot.
These are navigation-wrapper weights, not plain AutoModel weights. Do not load
them using the old text-policy serializer. navigation.json is authoritative.
Model/tokenizer/processor metadata only; optimizer states, RNG, images and
credentials remain local. SHA256SUMS.json records file hashes. Publication
is verified by downloading each uploaded file at its immutable Hub revision.