1024-Layer Networks Boost Self-Supervised RL
A NeurIPS Best Paper shows that scaling RL network depth from 2-5 to 1024 layers improves reward-free goal-reaching performance by 2x to 50x.
ImportanceMaterialEvidenceE2 unreplicatedWrite-upQuick
Scaling reinforcement learning network depth from the conventional 2-5 layers to 1024 layers improved agent performance by 2x to 50x in unsupervised, reward-free goal-conditioned tasks.
Recent RL research has largely relied on shallow architectures, limiting the scalability breakthroughs seen in language and vision models. This study by Kevin Wang et al. demonstrates that depth itself is a critical factor for unlocking scalable self-supervised RL.
Experiments were conducted on simulated locomotion and manipulation tasks using a contrastive self-supervised RL algorithm. The results show that increasing depth not only raised success rates but also qualitatively changed the learned behaviors. The paper was accepted as a Best Paper at NeurIPS 2025, though its findings have not yet been reproduced in real-world physical environments.