Skip to content
← All projects
Multi-agent systems
Research prototype

Independent R&D

Autonomous Network Agents

The research question was whether AI agents can build useful situational awareness of a live network without a human prompting each step. The network is simulated from public cell-tower locations, with 3GPP-style performance counters and alarms, and faults whose rate follows live public storm warnings. It started with 50 seed cells; the agents grew it to 249 by the time of the first write-up.

Each agent uses a different model chosen for its role. SENTINEL (gpt-oss-120b, every 5 minutes) watches the network. ORACLE (DeepSeek V3.2, every 15 minutes) looks for patterns and writes advisories. ARCHITECT (Qwen 3.5-35B, every 30 minutes) remediates faults and decides where to grow the network.

Agents page of the project dashboard. Panels for SENTINEL, ORACLE and ARCHITECT show each agent's latest message, above a feed of bulletin-board posts including an advisory to avoid Dublin and growth wave 16 in Monaghan, Cavan, Tipperary and Wexford.
The agents' shared bulletin board on 7 April 2026, at growth wave 16.

Engineering work

The system combines a simulated network with a stochastic fault engine, three role-specific agent playbooks, a priority-aware bulletin board, a wake-up trigger, deterministic executor scripts and a fault-fatigue state machine. It runs inside an NVIDIA NeMoClaw/OpenShell sandbox with Landlock, seccomp and a five-endpoint network allowlist.

Coordination through a bulletin board

The agents never call each other directly. Each run starts a fresh session, reads its inputs and posts structured messages to an append-only log. Messages carry a type, a validity period and an optional priority, so stale advice expires instead of blocking later decisions.

Agents can also change each other's pace. In a test, a backhaul fault was injected into three cells at a site in Kerry. SENTINEL detected it, posted an urgent handoff and tightened ORACLE's schedule from 15 minutes to 2. ORACLE advised avoiding Kerry, keeping other counties open and remediating the link, then tightened ARCHITECT's schedule from 30 minutes to 2. ARCHITECT ran the remediation. Detection to action took four minutes, with no human involvement.

Situational awareness in practice

When a backhaul fault hit a Cork city-centre site, SENTINEL classified degradation across five cells as one site-level fault rather than five cell problems. It tracked the degradation streak across cycles by re-deriving it from the data each time, since no session remembers the previous one.

ORACLE ruled out weather, traffic and local events as causes, then told ARCHITECT to defer expansion in Cork while keeping growth open elsewhere. ARCHITECT followed the advisory and resumed Cork growth once the fault cleared. The agents scoped actions geographically and correlated several sources rather than simply forwarding alarms.

When compliance looks like work

The most persistent failure produced no errors. Agents would describe a correct plan in detail, then make zero tool calls. Two separate causes produced the same symptom: instructions that were not explicit enough at the point of action, and tool approvals that were silently dropped in unattended sessions with no one to approve them.

Debugging meant reading plausible output and asking why nothing had changed. Fixes included explicit execution directives in the playbooks, approval rules configured for unattended runs and a pre-flight script that confirms, from inside the sandbox, that agents can actually execute, reach their endpoints and write state.

Models reason, code does bookkeeping

Early versions let ARCHITECT pass cell and site identifiers taken from alarm text. It shortened, invented or reused stale identifiers, and exact-match scripts rejected them without visible failure. The fix removed the model from identifier handling: the agent decides whether to remediate, and a single executor fetches canonical identifiers from the network API and applies the actions.

The same rule applies to growth. ARCHITECT chooses counties, site names, coordinates and cell counts; scripts derive wave numbers, totals and state updates. Over a five-day logged window the network completed 16 growth waves, with every geographic decision made by the agent.

Knowing when to stop

The Kerry remediation failed: the agents' tools could not fix that type of fault. For the next two hours SENTINEL reported it every five minutes, ORACLE reissued its advisory and ARCHITECT retried the fix. Every individual action was reasonable, which is why the loop looked like normal operation.

The cause was missing memory, not poor reasoning. Each run is an isolated session with identical inputs, so any model would have produced the same output. A fault-fatigue lifecycle now moves repeated failures from active to fatigued, then to stuck with an explicit escalation, and periodically to recheck for one more attempt. The aim is to stop retrying while keeping the fault visible.

Three layers

The project settled into three layers. The execution layer — sandbox, network isolation and runtime policy — contains the agents but does not make them correct. The agent layer reasons and decides. Between them sits a governance layer of behavioural limits, decision audit and supervision. That middle layer is where most of the work in this project ended up, and it is the one most agent deployments are missing.

Limitations and next steps

The network, its faults and its management API are simulated. Agents run prescriptive playbooks, and an internal review estimated the system at roughly 15% genuine autonomy and 85% engineered structure. Growth limits such as cells per site and permitted counties are set in prompts and are not enforced in code.

Most remediation calls in the logged window found nothing left to change, because faults had already cleared or been handled. That points to a need for better targeting, not more activity. There are no token budgets or human approval points yet. The next step is measuring decision quality against the simulator's ground truth rather than counting actions.

Where this work fits

Useful for operations workflows where several agents monitor, analyse and act on the same system. The work separates model judgement from deterministic execution and shows how agents can escalate, back off and stay inside a sandbox.

For a similar project, define what each agent may change, run it against a simulator with known ground truth, and verify that every reported action changed state.

Code and documentation

View repository

Read the prompt-injection threat model

Discuss a related workflow or integration

Discuss an AI project

Describe the workflow, integration or technical question you want to explore.

Discuss a project

info@genaisolutions.net