Semantic Graph Modelling: Agents That Build Executable Knowledge Graphs for Long-Horizon Work
October 8, 2026

An agent that works for a company for a week does not fail because one step is hard. It fails, or becomes expensive, because every step needs knowledge that is scattered: what a column in the ERP means, how two tables join, which policy applies to a change, which orders it already put on hold. Today's agents re-derive that knowledge from raw tables, documents and their own growing history, request after request.
We asked what happens if the agent first builds a model of the business and then works on the model. The model is a semantic graph: typed objects and links mapped to the ERP tables (a knowledge graph), plus actions whose rules encode the company's policies and write back to the ERP, functions for derived quantities, and automations that fire when the data change. The agent builds it itself, through a branch, a validation step and a publish step, and then works through a typed command-line interface.
How we tested it
We built OpsWeek, a generator of simulated work weeks at a fictitious distributor: a legacy ERP with coded tables, a data dictionary, six policy documents, events between requests (EDI orders arrive, goods ship, invoices are paid) and forty requests per week. The requests ask for lookups and aggregates, for writes under policies (enter a phone order with a discount, approve or allocate an order), for reports, and for standing instructions ("from now on, check every new EDI order on arrival"). An oracle grades every request on the agent's own state, so one early mistake is not counted again and again.
Seven conditions share the model (DeepSeek V4.1 Flash), the requests and the events. Long-context prompting gets the whole ERP in the prompt; retrieval gets the top 30 chunks; a coding agent gets SQL and the raw ERP transactions. Three conditions give the same agent an expert-built graph with more and more capabilities (read-only, then actions, then functions and automations). The last one, SGM-self, builds its own graph from the ERP and the policies before the week starts. All runs went through one metering proxy on a shared runtime pool, so every token and every second is counted.
What we found
The agent can build the graph. In its preparation request the agent built an initial graph with, on average, 14.7 object types, 15.3 actions and 4.2 automations in 7.3 minutes, for 0.82 CNY, and kept extending it during the week.
Accuracy was at the ceiling. With SQL over the raw ERP, the coding agent already completed 97.7% of requests; with its own graph it completed 99.5%. With six work weeks per condition, a difference this small cannot be told apart from chance. Long-context prompting reached 77.8% at 1.003 CNY per request, and retrieval 24.1%.
The graph did not save tokens. The agent with its own graph used 1.66× the tokens per request. The pattern fits a simple explanation: the transcript of building the graph stays in the conversation and is re-sent with every later model call, while SQL already returns compact answers. Because almost all input tokens were served from the provider's cache, the extra cost was small: 0.010 CNY per request.
The graph changed how the work was done. Every write went through an action that the engine checks, and automations met 78.6% of the standing-instruction obligations without a model call.
The agent without a graph compiled the policies too. In 5 of its 6 weeks, the coding agent wrote its own helper library during preparation (functions that check an order against the credit and approval rules or pick reorder candidates) and ran it about 52 times a week. So both agents turned the policies into code. The difference is that the graph's code is typed, enforced by the engine and triggered by changes, while the helper scripts run only when the agent remembers them.
Where it breaks
Agents of both kinds acted on rules that nobody had asked for yet, because the policy text described them. The graph agent also kept reading through SQL, so the read half of its graph saw little use. These are questions of governance more than of model quality: who may turn policy prose into executed rules, and which interface an agent should read through.
The benchmark has limits. It is synthetic, its policies are clear, and it never asks for a request that must be refused. We tested one inexpensive model. And our ladder adds capabilities to one model of the domain, so it does not tell us whether a policy-checked transaction wrapper and a trigger runtime without a graph would do the same.
What it means
In our setting, an executable knowledge graph changed how an agent writes and reacts more than how accurately or cheaply it works. Its distinct value is enforcement and triggers. Our next steps are harder benchmarks with refusals and longer horizons, stronger models, a comparison with guard and trigger runtimes that have no graph, and graphs that many agents share.
Read the full paper
Read the full working paper: the formal model, the OpsWeek benchmark, seven conditions and all results.
The full paper is in English.
Cite this work
Yanming Guo, Haixin Wang, Yanjun Lu (2026). Semantic Graph Modelling: Agents That Build Executable Knowledge Graphs for Long-Horizon Work. AIDC Research. https://www.ai-dc.ai/research/semantic-graph-modelling
@techreport{guo2026semantic,
title = {Semantic Graph Modelling: Agents That Build Executable Knowledge Graphs for Long-Horizon Work},
author = {Guo, Yanming and Wang, Haixin and Lu, Yanjun},
institution = {AIDC Research},
year = {2026},
type = {Research paper},
url = {https://www.ai-dc.ai/research/semantic-graph-modelling}
}

