Open-Source 4B VLM RL-Trained to Play GeoGuesser and Beat Much Larger Models
Key Info
Researchers RL-trained a 4B vision-language model to play GeoGuesser, and the full environment, dataset, training recipe, evals, and code are open-sourced. The model runs on a single A100 and reportedly outperforms several much larger models on evals.
Highlights
- Fully open-source: Environment, dataset, training setup, evals, and code are all released for end-to-end reproduction.
- Efficient: Runs on a single A100 GPU.
- Strong performance: Beats GPT 5.4 mini, Claude Haiku, and Qwen 3.5 122B, and comes close to Sonnet on evals.