SWE-Protégé: Learning to Selectively Collaborate With an Expert Unlocks Small Language Models as Software Engineering Agents. NeurIPS 2026.

SWE-Protégé is a post-training framework (SFT + agentic RL) that teaches a small language model to stay the decision-maker on SWE tasks while selectively asking a strong expert model for guidance. It brings Qwen2.5-Coder-7B-Instruct to 42.4% Pass@1 on SWE-bench Verified, +25.4% over the prior SLM state of the art. Work done at Meta.

September 2026 · Patrick Tser Jern Kon, Archana Pradeep, Ang Chen, Alexander P. Ellis, Warren Hunt, Zijian Wang, John Yang and Samuel Thompson