This is the public documentation for Engineer Brain, a private project. It is an autonomous network engineering agent built the way I would want a junior engineer to operate: it reads the documentation before it touches anything, it shows its reasoning, it asks before it changes production, and everything it does is on the record. The source stays private because the repo carries my lab topology and operational config. A code walkthrough is available on request.
The agent runs in production on my k3s cluster, deployed by GitOps (GHCR images, ArgoCD Image Updater, merge to main rolls the deployment with no manual steps). It has been validated end to end against a live fault: I broke a BGP session in the lab fabric, and the agent detected it, retrieved the relevant protocol docs from its corpus, diagnosed the root cause, and applied the fix through the human approval gate.
Retrieval quality is gated in CI. The eval harness computes Recall@5/10/15 and MRR against a golden question bank, with per-difficulty and per-protocol breakdowns, and the build fails if Recall@10 drops below 75%. The knowledge corpus is 655 hierarchical chunks ingested from protocol references, troubleshooting trees, runbooks, and caveat docs, with SHA-256 manifest verification at ingestion time.
The system is knowledge-first: the agent's competence comes from a curated corpus, not from hoping the base model knows BGP.
Ingestion. PDF, Markdown, and YAML sources are parsed into hierarchical parent/child chunks, embedded, and indexed twice: dense vectors in ChromaDB and a BM25 index for lexical search. Corpus integrity is enforced by manifest hashes; a modified document without an updated manifest entry is rejected.
Retrieval. Queries run against both indexes, results are fused with Reciprocal Rank Fusion, then focused child chunks are expanded to their parent chunks for context. A cross-encoder reranker is implemented as an opt-in stage, with the model preloaded to a persistent volume so the cluster never pulls models at runtime. It fails open if unavailable.
Reasoning. The orchestrator runs an OODA loop (observe, orient, decide, act) with explicit confidence scoring. Confidence thresholds are policy, not vibes: the agent must clear 85% to act autonomously, 65% to recommend, and anything under 40% escalates to a human.
Execution. All device interaction goes through a tool layer organized in three trust tiers, with 18 tools registered in the trust config: read tools run autonomously, analyze tools are supervised, write tools block on human approval in every interface. A platform driver abstraction handles vendor differences (FRR, Cisco IOS-XE, NX-OS, Arista EOS, Nokia SR Linux), and the lab fabric is a multivendor ContainerLab topology spanning an FRR spine-leaf fabric and Nokia SR Linux border nodes with inter-AS eBGP between them.
Two ideas make the autonomy trustworthy.
Structured Network Autonomy (SNA) is the trust model. Every tool carries a tier (autonomous, supervised, approval, restricted), and the security layer enforces the tier before execution, not after. Security checks live at the tool layer, not in the prompt: blocked destructive commands (reload, write erase, format), injection pattern detection, credential redaction, and per-device write rate limits all run in code that the LLM cannot talk its way around.
GAIT is the audit trail: a git-backed immutable transcript of every tool call, result, and security block. Write tools capture a pre-change baseline before they execute. The rule that matters most: if GAIT is unavailable, the agent halts. No silent degradation.
The threat model is aligned to the OWASP Agentic Top 10. Retrieved corpus chunks and raw device output (MOTD banners included) are scanned for prompt injection before they reach the LLM context, and behavioral monitoring flags anomalous sessions: unusual tool call volume, repeated blocked-command attempts, write operations outside task scope.
Hybrid retrieval over pure vector search, because networking queries are full of exact tokens (RFC numbers, error strings, command syntax) that embeddings blur and BM25 nails. Parent-chunk expansion, because the chunk that matches is rarely the chunk you want to read. Tiered model selection, so routine steps run on cheap models and the frontier model is reserved for the reasoning that earns it. And fail-loud everywhere: the agent stopping is always an acceptable outcome; the agent guessing is not.