Open-Source 4B VLM RL-Trained to Play GeoGuesser and Beat Much Larger Models

Hugging Face ·

Key Info

Researchers RL-trained a 4B vision-language model to play GeoGuesser, and the full environment, dataset, training recipe, evals, and code are open-sourced. The model runs on a single A100 and reportedly outperforms several much larger models on evals.

Highlights

  • Fully open-source: Environment, dataset, training setup, evals, and code are all released for end-to-end reproduction.
  • Efficient: Runs on a single A100 GPU.
  • Strong performance: Beats GPT 5.4 mini, Claude Haiku, and Qwen 3.5 122B, and comes close to Sonnet on evals.
Loading...