<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>Reinforcement-Learning on Patrick Kon</title><link>https://www.cs-pk.com/tags/reinforcement-learning/</link><description>Recent content in Reinforcement-Learning on Patrick Kon</description><generator>Hugo -- 0.128.0</generator><language>en</language><lastBuildDate>Fri, 25 Sep 2026 00:00:00 +0000</lastBuildDate><atom:link href="https://www.cs-pk.com/tags/reinforcement-learning/index.xml" rel="self" type="application/rss+xml"/><item><title>SWE-Protégé: Learning to Selectively Collaborate With an Expert Unlocks Small Language Models as Software Engineering Agents. NeurIPS 2026.</title><link>https://www.cs-pk.com/papers/14/</link><pubDate>Fri, 25 Sep 2026 00:00:00 +0000</pubDate><guid>https://www.cs-pk.com/papers/14/</guid><description>SWE-Protégé post-trains small language models to selectively ask a strong expert model for help on software engineering tasks.</description></item></channel></rss>