ai-memory + AgenticMem · one engineering & development scheme · v1.0.0 GA · epic #3969 · AlphaOne LLC

ai-memory Engineering & Development Scheme

The open-source ai-memory project and AgenticMem, its commercial support vehicle, are built by one team under one engineering scheme. This page is that scheme. It covers the full spectrum of how ai-memory is engineered: planning, development, testing, code and security review, documentation and decision-making. The release is built, tested and landed across two nodes. f1 runs macOS and f2 runs Linux. Each runs a native enterprise-federation data tier (PostgreSQL 18.6 with Apache AGE and pgvector) for CI and acceptance testing. Agents coordinate over ai-memory's own wake plane, and TypeSafe Jev triages the deputy's inbox. Decisions are made by waves of independent AI agents that assess, attack and vote, and a script, not a model, counts the votes. The goal of all of it is to make the best engineering decisions. Every line below is data moving: hover over or tab to a box to trace its links.

2 nodesf1 macOS · f2 Linux
7 runnersself-hosted CI: 4 on f2, 3 on f1
43required checks on release/v1.0.0
PG 18.6AGE 1.8.0 · pgvector 0.8.6 on both nodes
10Codex workers (GPT-6 Astra @ high): 4 on f1, 6 on f2
~200 msone Jev call triages a whole inbox batch
≥4 of 7votes needed to decide in a 3x7 (≥2 of 3 in a 3x3)
code, branches, merges agent messages (ai-memory wake plane) data-tier traffic TypeSafe Jev model and cloud APIs CI and telemetry forward packet return packet planned, not running
01

Fleet map

Everything that runs, on which machine, and what each machine talks to. f2 carries the Claude panes, the team message store and the Linux CI leg. f1 carries a Codex pool, its own message store and the macOS CI legs.

02

Native data tiers and the CI matrix

Every promotion PR runs the correctness matrix against real databases. The enterprise-fed tier is always up on each node, so CI needs no container service. Each job creates its own database and drops it at cleanup.

03

Enterprise federation test fabric

The acceptance swarm exercises the same shape customers deploy: agents talk to an ai-memory daemon, the daemon stores into PostgreSQL, and daemons federate node to node. All three legs are encrypted in transit.

04

Agent comms plane (ai-memory wake plane)

Agents never type into each other's terminals directly. They post to ai-memory's own wake plane, and a per-pane deliverer hands each message to its agent when that agent is idle. Replies take the same path back.

05

TypeSafe Jev inbox triage

Jev classifies each batch of new messages as action, status or noise in one call. Claude spends tokens only on what needs a decision. Jev only routes messages. It never decides correctness, security or merges, and it never sees secrets.

06

From defect to release, and back

Every issue found in testing gets its own GitHub issue, a builder lane, independent review, a merge by GOD and a retest on the native tiers. New findings start the loop again.

07

How we decide: adversarial voting waves

Every consequential decision goes through waves of independent agents with one end state: making the best engineering decisions, measured against the North Star of data integrity, security, performance and reliability. One wave assesses, the next tries to break every claim, and the last votes. A script counts the votes, so no single model can overrule the panel.

SchemeWaves and agentsDecision ruleUsed forRuns published on release/v1.0.0
1x33 independent assessors≥2 of 3A quick adopt or reject on one narrow questionBAML and PAW assessments (2026-09-19)
3x33 assess → 3 attack → 3 vote → synthesis · 10 agent runs≥2 of 3, counted by scriptFocused audits with a clear scopeOpenViking code-level comparison, US agent-memory market audit (2026-09-27); code-reduction plan #4130; #3949 adversarial audit
3x712 evidence dossiers → 7 assess → merge → 7 attack → 7 vote → synthesis → publish · 36 agent runs≥4 of 7, counted by scriptHigh-stakes, many-sided decisionsCompetitor assessment #4055 (2026-09-27), boids emergence (2026-09-23), 3x7 adjudication, Grok dogfood full-spectrum
11- and 21-agent votesOne wide panel of 11 or 21 independent agentsMajority of the panelDirection-setting questions that need many independent viewpointsRed Queen 11- and 21-agent votes, corpus completeness (21 agents)
Documentation 3x77 doc lenses → 7 critics → 7 voters → drift register · 22 agent runs≥4 of 7 per drift itemKeeping every doc and the GitHub Pages site true to the codeFirst full run in progress (started 2026-09-29), see section 09
Chained programsSeveral 3x7 and 3x3 runs in sequence, each feeding the nextEach stage's tally is binding on the nextEnd-to-end programs where one decision shapes the next#4055: 3x7 → two 3x3s → advantage register → ROADMAP-v1.1.0-ADDENDUM
Retest wavesAn independent red → green retest after every fix round, until one comes back clean100% green, with every gap filed 1:1Fixes for issues found in testingMCP lane R10 → R10b → R10c → R10d
08

Decision programs and retest waves

Bigger questions chain several voting runs, so each run starts from a decision the previous one made. Fixes go through repeated retest waves until an independent tester finds nothing left.

09

Documentation fidelity: the 3x7 drift audit

Documentation is part of the product. Every shipped statement in the repo docs and on the GitHub Pages site must be true against the code. The 3x7 drift audit finds where docs and code disagree and votes on what to fix before GA. CI gates then stop each class of drift from coming back.

10

Codebase by the numbers

Measured from git objects at release/v1.0.0 @ 4fff967ce: every tracked text file, physical lines, grouped by language and by Rust module. "Code lines" excludes blank lines and lines that start with a comment marker, so it is an approximation.

1,052,972lines of Rust in 1,575 files
642,177production Rust in src/ (546 files)
400,429lines in tests/, 38% of all Rust
15,957Rust test functions (#[test] and #[tokio::test])
1,428,850lines across 3,104 text files, all languages
188,013lines of Markdown documentation
LanguageFilesLinesCode linesShare
Rust1,5751,052,972708,72573.7%
Markdown800188,013134,96813.2%
Shell17847,71328,6273.3%
JSON15640,38140,4432.8%
HTML7635,45032,9632.5%
Python10334,02127,1682.4%
YAML3110,5315,4840.7%
TypeScript308,3237,1060.6%
SQL1288,2587,9150.6%
TOML242,3271,0570.2%
JavaScript38616490.1%
Total3,1041,428,850995,105100%
Files over 8,000 linesLines
src/store/postgres.rs43,308
src/storage/mod.rs35,239
src/mcp/mod.rs17,437
src/handlers/tests.rs17,378
src/config.rs15,597
src/daemon_runtime.rs14,999
tests/integration.rs13,901
src/cli/doctor.rs9,030
11

CI across GitHub, f2 and f1

Every push or PR to release/v1.0.0 or the carrier fans out to 19 workflows. GitHub-hosted runners take the many fast jobs. The self-hosted runners on f2 and f1 take the few long ones that need the real PostgreSQL tiers. Nothing merges until all 43 required checks pass.

Required-check groupChecksChecks (exact names in branch protection)
Correctness matrix (native tiers)4Check (linux-fed,enterprise-fed) · Check (macos-fed,sqlite) · Check (macos-fed,enterprise-fed) · Check (ubuntu-latest,sqlite)
Native data-tier certification3Enterprise-federation cert-expiry gate (cert §7 / F7) · Certified pg+AGE cells (live PG 18.6 + AGE 1.8.0 + pgvector 0.8.6) · Postgres ignored tests (sal-postgres --ignored)
Build, lint & feature gates9Lint (fmt + clippy) · Postgres feature gate · SAL-only feature gate · vectorlite feature gate · Dockerfile build (no push) · Cross-compile (aarch64-apple-ios) · Cross-compile (aarch64-linux-android) · Build-script custom-build ledger gate (#2635) · MSRV (Rust 1.98)
Coverage2Coverage classify (docs-only short-circuit) · Per-Module Coverage Thresholds
Docs & claims truthfulness gates11Hardcoded-literal duplication ratchet (pm-v3.1) · Docs vs SSOT drift gate · Cloud-init ASCII gate · Doc symbol/path anchor gate (#2629) · SDK-path vs routes.rs membership gate (#2629) · Named-CI-job existence + enforcement-truthfulness gate (#2629) · Doc surface completeness gate (#2839) · Capacity-claim ceiling gate (#2869) · Benchmark-claim canon gate (#2879) · Stale contract-assertion gate (#3688 / #3967) · Claude plugin manifest gate (#3967)
Security & supply-chain gates5actionlint (workflow-injection guard) · Vendor-monoculture + SECS_PER_* lint-gate (#1174 PR10) · Installer checksum fail-closed gate (#2449) · Git-dependency-source supply-chain gate (#2050/#2512) · URL-sink redaction gate (#3688 / #3967)
Architecture & test-hygiene gates6C8 caller-context allowlist check · L3-boundary perma-ban gate (§25.3 S5 / RQ-10 #1853) · Test-reachable stdin-read gate (#1989) · Migration-ladder-uniqueness gate (guardrail-D) · Test-env $HOME-lock gate (#2146) · Non-Rust conformance-reader proof gate (#2452)
Process & governance gates3Classify changes · Required-context + classify-base soundness gate (#2494/#2496/#2508) · External-PR operator-approval gate (author outside team => @alphaonedev review)
Total required on release/v1.0.043Branch protection: strict (up to date with base), enforce_admins on. Lifted only with operator approval, backed up and diffed after.
12

Code review and security review

Claude reviews every change for correctness and security before GOD can merge it. Before every GA release, the entire codebase gets a full code review and a full security review, not just the changes. The operator introduced these full reviews, and they are now a standing part of the scheme.

13

Enterprise federation testing: DigitalOcean, clusters, swarms and hives

Native f1/f2 testing proves the code. Cloud testing proves it installs and federates from zero on machines we have never touched. Swarms and hives of AI agents then prove that many agents can share ai-memory safely. Every DigitalOcean and native f1/f2 run in the v1.0.0 epic was published live on test.agenticmem.co.

14

Planning: the road to the GA tag

One epic, one scope rule and one sequence, set by GOD and gated by the operator. Nothing enters v1.0.0 unless it is in scope, and nothing ships until every stage is green.

15

We eat our own dog food: AI NHI test everything

AI non-human identities (NHI) test ai-memory the way customers will use it. They are wired into the full enterprise federation configuration with every security layer on, and into a singleton SQLite build over MCP. Every feature gets exercised. Every issue they find becomes its own GitHub issue, and the policy is to fix 100% of the issues found in GA release testing. The engineering fleet runs on ai-memory too.

16

Exhaustive A2A testing: low-cost AI agent clusters

We spin up clusters of AI agents on a low-cost model, GLM-5.3-Flash through OpenRouter. They run natively on f1 and f2, and on DigitalOcean droplets created per campaign. Every agent is wired into ai-memory for shared memory and into rust-a2a for agent-to-agent signals. The same wiring runs in both places, so a result on our own nodes can be repeated in a clean room.