Recursive Self-Improvement in Self-Organizing Multi-Agent Systems
September 26, 2026 · Working paper v0.1

An organization of agents that packages the capabilities it builds as governed, tested and reversible components. Reuse was 3.4 times cheaper in agent time than rebuilding, and a faulty update never reached an agent.
Abstract
Recursive self-improvement of language-model agents changes one agent or one pipeline at a time, is judged by a held-out score, and says nothing about the artifacts it accumulates. We study improvement at the level of an organization of agents. The tool set of the agents’ harness becomes a governed region of the organization’s shared state: agents package the capabilities they build as components with code, libraries, models and a test; an admission gate admits a component only if its test passes without network, together with the tests of the version it replaces and of the components that call it, and only if no other component provides the same capability; every change is an event with an inverse. We show that these rules keep the tool set sound, unique per capability, closed under its calls and reversible, and that its cost in context is bounded, unlike shared text. On a stream of 64 multimodal tasks solved by 16 agents with a text-only model, in six organizational conditions and two replicates, the agents acquired every capability themselves, so accuracy was 95–98 % everywhere and the conditions differ in cost. Using a component another agent had published took 3.4× less agent time and 11× fewer downloaded bytes than acquiring the capability, at equal tokens; publishing cost 4.4× the tokens of an acquisition. Over the whole stream the governed organization downloaded 0.43× the bytes of stateless agents and significantly less than every other organization, but used 1.39× their tokens, and its agent time did not differ significantly at our scale, where publishing large components was slow. Shared notes and a shared directory did not spread capabilities, and a SkillOpt-trained skill accepted no edit. The registry kept one component per capability, kept a faulty update from ever reaching an agent, and after a change of base image retired every failing component and restored itself exactly when the retirements were reverted.
This is a working paper. The full text is in preparation for publication; the abstract above reports its current results.
Cite this work
Yanming Guo (2026). Recursive Self-Improvement in Self-Organizing Multi-Agent Systems. AIDC Research. https://www.ai-dc.ai/research/organizational-self-improvement
@techreport{guo2026organizational,
title = {Recursive Self-Improvement in Self-Organizing Multi-Agent Systems},
author = {Guo, Yanming},
institution = {AIDC Research},
year = {2026},
type = {Working paper},
url = {https://www.ai-dc.ai/research/organizational-self-improvement}
}

