From Language Models to Agent Systems: A Survey of Training, Architectures, Orchestration, and Open Problems
September 23, 2026 · Survey v0.1

A review of 434 works across nine layers, from pre-training to self-improvement and evaluation, with a taxonomy, a 2017–2026 roadmap and three open problems stated as testable hypotheses.
Abstract
Agents built on large language models now write and repair software, operate web and desktop applications, and coordinate with other agents, and usage data show them moving from research prototypes into deployed software. This survey reviews the path from language models to agent systems across nine layers: pre-training and post-training, including reinforcement learning from human feedback and from verifiable rewards; the reasoning, tools and memory of a single agent; coding agents; multi-agent systems; orchestration and communication protocols; agent and environment harnesses; multi-model and multimodal agents; self-improvement; and evaluation. We review 434 works, resolving their bibliographic records against arXiv or Crossref wherever such a record exists; their placement was assisted by a typed classifier that agreed with the author’s independent placement for 82.0% of works (Cohen’s κ = 0.805). We organize the literature into a taxonomy, a 2017–2026 roadmap and genealogies of methods. The synthesis identifies three open problems. First, agentic capability and automated use concentrate in software engineering, where programs can check many outcomes: computer and mathematical tasks are the largest category of Claude usage, at 34–40% of conversations and 44–46% of first-party API traffic, whereas agents validated in domains without verifiable rewards remain rare; usage data alone cannot separate verifiability from the other advantages of software. Second, multi-agent systems that solve tasks with tools typically use two to six agents, one reasoning-graph study organizes over a thousand text-only agents, systems with 104 to 106 agents are simulations, and frameworks are mostly designed for a single process or device. Third, deployment is constrained by consistency, security, privacy, concurrency, portability, maintainability and interoperability, among other properties that lie outside the model. By analogy with the early Web, we argue that agents have passed a first stage, and we state a research agenda as testable hypotheses on rewards for real work, process alignment and human control, and the orchestration of many independently deployed agents.
This is a working paper. The full text is in preparation for publication; the abstract above reports its current results.
Cite this work
Yanming Guo (2026). From Language Models to Agent Systems: A Survey of Training, Architectures, Orchestration, and Open Problems. AIDC Research. https://www.ai-dc.ai/research/agent-systems-survey
@techreport{guo2026agent,
title = {From Language Models to Agent Systems: A Survey of Training, Architectures, Orchestration, and Open Problems},
author = {Guo, Yanming},
institution = {AIDC Research},
year = {2026},
type = {Survey},
url = {https://www.ai-dc.ai/research/agent-systems-survey}
}

