mg_
← back

EILAID

· 4 min read

Engineers decide. Agents build. The project remembers.

EILAID (Engineer-In-Loop AI Development) is a local orchestrator. It takes the coding-agent CLIs you already use (Claude Code, Codex, OpenCode or your own adapter) and runs them through a lane-based loop: shape, build, verify, merge. You stay in the loop at two gates, through a chat, and through an inbox of questions the agents couldn’t answer themselves.

The loop

Work items move across a board. Every lane change starts the right agent in the right place. A shaper writes a brief with testable acceptance criteria. A builder works test-first in its own git worktree. A tester, a reviewer and QA verify the work, and the branch merges when the checks pass.

  1. inbox
  2. shaping
  3. gate 1
  4. ready
  5. building
  6. verifying
  7. gate 2
  8. merging
  9. done
TODO-3
a ticket's life. the black lanes are you.

Items never touch your working copy until they merge. When verification fails, the item goes back to building, up to a configurable number of attempts.

Fifteen roles

There’s no single “agent”. There are 15 roles, each a versioned markdown spec, plus 24 shared skills:

role does
conductor the chat front door: triages, files tickets, reports status
architect picks the stack, writes ADRs, breaks epics into a DAG
shaper turns a request into a brief with acceptance criteria
builder writes the code, test-first, in its own worktree
tester / reviewer / qa verify, in that order, the last two in parallel
merger resolves conflicts, keeping both intents
oracle answers questions with citations and a confidence
coach turns review findings into lessons and better specs

…plus researcher, designer, plan-critic, design-reviewer and steward.

A role version is runtime + model + effort + spec. Change any of these and it becomes role@N+1, so the metrics stay comparable across versions.

The Oracle

The Oracle is a knowledge base over your code, docs, decisions and briefs. Agents ask it before they guess. Anything it can’t answer goes to a human once, and the answer is remembered.

Retrieval is hybrid: BM25 via SQLite FTS5, a trigram index for paths and identifiers, exact grep via ripgrep, and optional embeddings. All of these are merged with Reciprocal Rank Fusion. The whole fusion step is this small:

/**
 * Reciprocal Rank Fusion: score(d) = Σ 1 / (k + rank_d), ranks from 1.
 * Rank-based, so bm25 (negative, unbounded) and cosine (-1..1) mix fairly.
 */
export function rrf(lists: Array<Array<{ id: string }>>, k = 60): Array<{ id: string; score: number }> {
  return weightedRrf(lists.map((items) => ({ items, weight: 1 })), k);
}

export function weightedRrf(lists: Array<{ items: Array<{ id: string }>; weight: number }>, k = 60): Array<{ id: string; score: number }> {
  const scores = new Map<string, number>();
  for (const { items, weight } of lists) {
    const seen = new Set<string>();
    items.forEach((it, rank) => {
      if (seen.has(it.id)) return; // a duplicate within one list counts once, at its best rank
      seen.add(it.id);
      scores.set(it.id, (scores.get(it.id) ?? 0) + weight / (k + rank + 1));
    });
  }
  return [...scores.entries()].map(([id, score]) => ({ id, score })).sort((a, b) => b.score - a.score || (a.id < b.id ? -1 : 1));
}

A deadlock in one line

Agents block on the Oracle while they wait for an answer. If Oracle runs had to queue behind the very agents waiting on them, nothing would ever finish. The fix is a set of purposes that never take a concurrency slot:

/** Purposes that never wait for (or take) a concurrency slot. */
export const SLOT_EXEMPT_PURPOSES: ReadonlySet<string> = new Set(["oracle", "oracle-expand", "chat"]);

Boring tech, on purpose

EILAID is one Node process, one SQLite file and your git repo. It’s TypeScript run directly by Node’s built-in type stripping, so the backend has no build step. The backend’s only dependencies are yaml and picomatch; everything else (node:sqlite, node:http, child_process, crypto) ships with Node. The MCP server that agents use to talk to the daemon is hand-written JSON-RPC.

Every run has a budget: per run, per item and per day, enforced live. Each run also ends with a fenced eilaid-result JSON block, validated against that role’s contract. If the block is missing, the run counts as failed.

Does it work?

Here is one real run, starting from an empty repo with Claude Code:

36passing tests
42agent runs
$4.43API-equivalent
0hand edits

The result was a working todo CLI with storage and colours. It’s still an MVP and the repo is private for now, but the loop is real.

↑↓ move · ↵ run · esc close