Metaminds ResearchAI Dark Factory · Open source

uzi: an AI dark factory.

Specs in. Pull requests out. An open-source factory for software that runs with the lights off — label an issue, approve the plan, and it opens a reviewed pull request, never touching main.

1,200+Runs
completed
20B+Tokens
spent
MITOpen-source
licence
A dark factory corridor: a backlit ‘uzi — Specs in. Pull requests out. / Plan · Build · Ship’ sign, a white humanoid robot labelled ‘Lead’ overseeing the floor, and a conveyor of robotic arms labelled ‘coder’ and ‘tester’ assembling glowing amber code crates past a night skyline

01The idea

Issues in,
pull requests out.

Even with today’s agents, “AI coding” keeps you in the driver’s seat of a single session — you kick it off, watch it work, nudge it when it drifts, and start the next task yourself. You are the conveyor belt.

uzi inverts that. The unit of work is an issue, not a message. It reads the issue, plans the change, builds it, reviews it, and opens a pull request from a new branch — while you stay out of the loop until the two decisions that matter.

P1 The issue is the unit You label an issue on your forge; uzi treats it as an order to fulfil — read, plan, build, review, open a PR.
P2 The forge stays the truth Work shows up where your team already looks — as issues, branches and pull requests, kept in two-way sync.
P3 Isolated workers Each run executes in its own container that sees one repo checkout and one run, so a mistake trashes a branch, not your machine.
P4 Two human touchpoints You approve the plan and you merge the PR. Everything between is the factory floor — lights off.
How uzi ships an issue: an issue enters intake and planning, passes the plan gate for human approval, then a coder implements while reviewer, auditor, tester and fact-checker validate in parallel until it holds; a branch and pull request then go to human review and merge. A red pipeline triggers a Fix CI run, and PR comments trigger a rework.
How uzi ships an issue — from a labelled issue through the plan gate, the coder-and-reviewers loop, a pull request and human review, to merge. Two decisions stay human; everything between is the factory floor.

02The pipeline

Two decisions stay human.
Everything between is the floor.

Plan first A worker claims the issue and produces a plan. The run parks at an approval gate. You approve it, or reject with a reason and it re-plans. Nothing is written until you say go.
Implement & review On approval the lead dispatches a coder, then fans out reviewer, auditor, tester and fact-checker in parallel — looping back to the coder until the work holds up. Watch each agent stream live.
Branch & PR, never main On completion uzi opens a branch and a pull request, links them from the run and the card, and moves the issue to human review. main is never touched, by design.
The uzi board: labeled issues as cards moving across columns synced to forge labels
The board is the front door — a per-repo kanban of your forge’s issues. The working columns are forge labels, so moving a card relabels the issue.

03The crew

A lead that delegates, and specialists that check its work

Work is not done just because an agent says so. The lead orchestrates; a role-based crew implements and validates in parallel, looping until the change holds up.

Lead Orchestrator

Plans the run and delegates instead of doing everything itself, working the approved plan one milestone at a time.

Coder Implementer

Writes the change for the current milestone, committing each as its own reviewed slice.

Reviewer · Auditor Validation

Read the diff for correctness and for security, reuse and simplification, fanning out in parallel.

Tester · Fact-checker Validation

Exercise the change and check its claims against the codebase, looping back to the coder until it holds.

A run paused at the plan-approval gate, with the proposed plan in view
The plan gate — uzi stops and asks before it writes anything.
The run activity feed grouped by agent, each with its current step and milestone
A per-agent activity feed — expand any entry to stream that agent’s transcript live.

04Accountable

Every run is reviewed, measured and costed

Nothing about a run is a black box. It is scored after the fact, tracked while it runs, and accounted for to the token.

Run judgeAn optional retrospective reads the whole run trace and returns a verdict plus concrete recommendations. Advice, not a gate — it never changes code.
MilestonesThe lead ticks off each slice as it lands, so even a long run shows honest progress instead of a spinner.
Fully costedEvery run reports tokens in and out, cache hits, duration and dollar cost — broken down per phase and per agent, on your own Anthropic token.
Rate-limit waitHit a cap mid-flight and uzi pauses with a countdown, then resumes on its own when the window resets — same branch, no re-approval.
A finished run’s judge review: a verdict, a retrospective, and token and cost stats
The run judge — a verdict and a retrospective for every finished run.
A run’s cost and token stats, broken down per phase and per agent
Per-run cost and tokens, broken down per phase and per agent.

05Where it stands

Alpha, and it works —
building its own next version.

1,200+
Runs completed to date
20B+
Tokens spent getting here
3
Forges — GitLab, GitHub, Forgejo
0
Commits to main, by design

And the fun part: uzi builds uzi. Issues get filed, plans approved and PRs opened — so a growing share of the factory is written by itself while a human reviews.

06Scheduled jobs

The factory works the factory

Pointed at its own repo, uzi’s standing automations add up to a self-improvement loop — hunting its own bugs, strengthening its tests, keeping its docs honest, and proposing its next feature. Every job falls back to a plain report when it has nothing worth landing, so a quiet week produces no empty pull requests.

Schedule What it does Cadence
bug-triage sweeps bug-labeled issues daily
planned-sweep sweeps Planned-labeled issues daily
docs-hygiene mechanical documentation fixes weekly
test-improvement lands new tests only, no production code weekly
bug-hunt a deep audit of one subsystem, one focused fix weekly
self-improve scans the codebase and opens a self-improvement PR ~2 days
feature-bingo brainstorms one new feature and proposes it weekly

feature-bingo is the factory designing its own next machine.

Once a week it reads the existing ideas, checks what already exists so it doesn’t repeat itself, and proposes exactly one concrete new feature — the problem it solves, a sketch of how it works, and where it lives — as a pull request. A chunk of uzi’s roadmap arrives as PRs to wake up to.

The schedules page, showing the standing automations including feature bingo

07Watch it anywhere

Follow the floor from
the terminal or your phone.

The uzi TUI run view: the crew, milestones, rate-limit meters and the live agent transcript
uzi tui — the crew, milestones, rate-limit meters and the live transcript, in your terminal.
uzi on a phone: the overview dashboard
The whole web UI is responsive — approve a plan straight from your phone.

08Where it came from

An AI research project, released as open source

uzi began as an AI research initiative at Metaminds and has been built on both work and personal time since. Metaminds takes open source seriously, so releasing it was an easy call. It is MIT licensed — a helm install away on Kubernetes, or a docker compose up on a laptop.

Use at your own risk.

uzi runs autonomous agents that read your code, run commands inside their workers, and open pull requests using your own model tokens. Run it against repositories you own. You stay in control by design: review the plan before you approve it and the diff before you merge it — uzi opens pull requests but never merges them and never touches main.

Try the factory

Star it. Break it. Tell us.

uzi is open source and moving fast. Try it, file issues and feature requests, send PRs, and help shape where the dark factory goes next.