Replit drops model routers, lets the main model decide task delegation
Coding agents are handing delegation back to the model itself; the cost and quality numbers are vendor-run, worth a look for agent builders.
Replit now lets its main model decide subagent tier and thinking effort step by step, replacing a fixed model router.
In an engineering blog post dated September 29, Replit explains that Replit Agent, its AI coding agent, moves away from the common setup where a small model reads each turn and picks a stronger model — a router that is always less capable than the model it chooses for. The new architecture has the core loop pick each subagent's size, effort level and delegation target at every step, adjusting mid-turn.
The company reports that on the DeepSWE and Terminal-Bench software engineering benchmarks, the architecture is Pareto-efficient against Astra running alone, and beats a sidekick architecture with one long-lived worker by 11 and 16 points; each run is a mean of four repetitions.
All these numbers come from Replit's own runs, with published leaderboard figures as baselines rather than same-condition reproductions, so readers should treat them as a vendor's architecture claim, not independently verified performance.