Independent R&D
Autonomous Network Agents
The research question was whether AI agents can build useful situational awareness of a live network without a human prompting each step. The network is simulated from public cell-tower locations, with 3GPP-style performance counters and alarms, and faults whose rate follows live public storm warnings. It started with 50 seed cells; the agents grew it to 249 by the time of the first write-up.
Each agent uses a different model chosen for its role. SENTINEL (gpt-oss-120b, every 5 minutes) watches the network. ORACLE (DeepSeek V3.2, every 15 minutes) looks for patterns and writes advisories. ARCHITECT (Qwen 3.5-35B, every 30 minutes) remediates faults and decides where to grow the network.

Engineering work
The system combines a simulated network with a stochastic fault engine, three role-specific agent playbooks, a priority-aware bulletin board, a wake-up trigger, deterministic executor scripts and a fault-fatigue state machine. It runs inside an NVIDIA NeMoClaw/OpenShell sandbox with Landlock, seccomp and a five-endpoint network allowlist.
Coordination through a bulletin board
The agents never call each other directly. Each run starts a fresh session, reads its inputs and posts structured messages to an append-only log. Messages carry a type, a validity period and an optional priority, so stale advice expires instead of blocking later decisions.
Agents can also change each other's pace. In a test, a backhaul fault was injected into three cells at a site in Kerry. SENTINEL detected it, posted an urgent handoff and tightened ORACLE's schedule from 15 minutes to 2. ORACLE advised avoiding Kerry, keeping other counties open and remediating the link, then tightened ARCHITECT's schedule from 30 minutes to 2. ARCHITECT ran the remediation. Detection to action took four minutes, with no human involvement.
Situational awareness in practice
When a backhaul fault hit a Cork city-centre site, SENTINEL classified degradation across five cells as one site-level fault rather than five cell problems. It tracked the degradation streak across cycles by re-deriving it from the data each time, since no session remembers the previous one.
ORACLE ruled out weather, traffic and local events as causes, then told ARCHITECT to defer expansion in Cork while keeping growth open elsewhere. ARCHITECT followed the advisory and resumed Cork growth once the fault cleared. The agents scoped actions geographically and correlated several sources rather than simply forwarding alarms.
When compliance looks like work
The most persistent failure produced no errors. Agents would describe a correct plan in detail, then make zero tool calls. Two separate causes produced the same symptom: instructions that were not explicit enough at the point of action, and tool approvals that were silently dropped in unattended sessions with no one to approve them.
Debugging meant reading plausible output and asking why nothing had changed. Fixes included explicit execution directives in the playbooks, approval rules configured for unattended runs and a pre-flight script that confirms, from inside the sandbox, that agents can actually execute, reach their endpoints and write state.
Models reason, code does bookkeeping
Early versions let ARCHITECT pass cell and site identifiers taken from alarm text. It shortened, invented or reused stale identifiers, and exact-match scripts rejected them without visible failure. The fix removed the model from identifier handling: the agent decides whether to remediate, and a single executor fetches canonical identifiers from the network API and applies the actions.
The same rule applies to growth. ARCHITECT chooses counties, site names, coordinates and cell counts; scripts derive wave numbers, totals and state updates. Over a five-day logged window the network completed 16 growth waves, with every geographic decision made by the agent.
Knowing when to stop
The Kerry remediation failed: the agents' tools could not fix that type of fault. For the next two hours SENTINEL reported it every five minutes, ORACLE reissued its advisory and ARCHITECT retried the fix. Every individual action was reasonable, which is why the loop looked like normal operation.
The cause was missing memory, not poor reasoning. Each run is an isolated session with identical inputs, so any model would have produced the same output. A fault-fatigue lifecycle now moves repeated failures from active to fatigued, then to stuck with an explicit escalation, and periodically to recheck for one more attempt. The aim is to stop retrying while keeping the fault visible.
Three layers
The project settled into three layers. The execution layer — sandbox, network isolation and runtime policy — contains the agents but does not make them correct. The agent layer reasons and decides. Between them sits a governance layer of behavioural limits, decision audit and supervision. That middle layer is where most of the work in this project ended up, and it is the one most agent deployments are missing.
Limitations and next steps
The network, its faults and its management API are simulated. Agents run prescriptive playbooks, and an internal review estimated the system at roughly 15% genuine autonomy and 85% engineered structure. Growth limits such as cells per site and permitted counties are set in prompts and are not enforced in code.
Most remediation calls in the logged window found nothing left to change, because faults had already cleared or been handled. That points to a need for better targeting, not more activity. There are no token budgets or human approval points yet. The next step is measuring decision quality against the simulator's ground truth rather than counting actions.
Where this work fits
Useful for operations workflows where several agents monitor, analyse and act on the same system. The work separates model judgement from deterministic execution and shows how agents can escalate, back off and stay inside a sandbox.
For a similar project, define what each agent may change, run it against a simulator with known ground truth, and verify that every reported action changed state.