Co-author · USC
GameBoyWorlds: open-world GenAI & RL agent framework
Under review at ICLR 2027An embodied-AI benchmark of 500 multimodal tasks across 10 games. Frontier and open-source VLMs (GPT-5, Gemini, Claude, Gemma, Qwen) are evaluated through hierarchical agents with supervisor-executor control, dynamic memory, self-reflection, and subgoal planning.
Up to 22%Task-success gain from a curiosity-driven exploration pipeline on select environments