Open-source AI / The AI briefing

AWS uses SkyRL to increase Qwen3-VL-8B maze navigation performance

AWS researchers trained a Qwen3 vision-language model to solve visual mazes with over 95% accuracy using the open-source SkyRL reinforcement learning framework.

AWS publication cover: AWS uses SkyRL to increase Qwen3-VL-8B maze navigation performance
From AWS ML. Original AWS publication cover.

AWS demonstrated training the Qwen3-VL-8B vision-language model using the open-source SkyRL framework on SageMaker HyperPod. By applying Group Relative Policy Optimization post-training to a supervised fine-tuning checkpoint, researchers increased the model's success rate in navigating visual mazes from 43.75% to over 95% on a fixed evaluation set.

Reinforcement learning for vision-language models

The training process utilized Group Relative Policy Optimization to refine the Qwen3-VL-8B model, which initially started from a supervised fine-tuning checkpoint. This multi-turn reinforcement learning approach allows the agent to learn from rewards accumulated over an entire sequence of steps within a visual maze, rather than receiving feedback only for single actions.

To manage the heavy computational demands of these multi-node runs, the framework ran on Amazon SageMaker HyperPod. This infrastructure utilizes cluster resiliency features to monitor node health and automatically replace faulty hardware, allowing long-running reinforcement learning jobs to resume from checkpoints instead of restarting from scratch after a failure.

Evaluation results and infrastructure limits

Testing conducted on a fixed 64-maze evaluation set showed that the reinforcement learning post-training significantly improved performance over the baseline. The solve rate rose from 43.75% to more than 95%. The system integrates with Ray clusters and uses managed dashboards for monitoring training dynamics and observability.

The reported success is based on a specific, fixed set of 64 mazes. The source does not provide performance data for unseen or procedurally generated environments outside of this evaluation set. While the infrastructure supports large-scale workloads, the training duration and total GPU-hours required for these results were not specified.

Original source

This report summarises the source below. Analysis is labelled separately; product and research claims remain attributed to their source.

Read the original at AWS ML

← Back to all AI news