Geoguesser Environment
Interact with a GeoGuessr‑like environment
Interact with a GeoGuessr‑like environment
Note Task splits. 200 held-out eval tasks across 73 countries and 3,452 training tasks, carved from one pool so contamination is enforced once: no shared Mapillary sequence, and no training task within 1 km of an eval task.
Note Run 1. The headline result: 0.6445 mean-of-4 on the held-out split, ahead of gpt-5.4-mini and claude-haiku-4.5, and +0.1620 over its own base (95% CI +/-0.0137, better on 169 of 200 tasks). LoRA r=16 on Qwen3.5-4B, 10 hours on 4xA100.
Note Run 3. The controlled ablation: reverting scale_rewards to group and beta to 0 on top of run 2's config recovered 44% of run 1's gain, at +0.0717.
Show live I/O tracking dashboard
Note Training curves for all four runs in one project, so they sit on a comparable axis. Plot against train/global_step: TRL logs twice per optimiser step here.
From an idea to a trained 4B, with the dead ends left in
Note The write-up: every decision, and what it cost to learn it
Note 22 GB of Mapillary panoramas, the environment's imagery cache