Proof
If you distrust demos, start here. This page is the position, the proof, and the honesty in one place — with the caveats left in. Everything below either links to a running product doc or embeds a recording captured from the live fleet. We claim a position, not market share.
The self-hosted governance and verification layer for autonomous coding: the factory that runs your agents' code, tests it for real, and refuses to overclaim. Trust-verified, not code-faster.
The gap we sit in
The 2026 data says buyers already have agents — what they lack is a way to trust them.
Two axes decide the market: where the code runs (someone else’s cloud vs your own infrastructure) and what the product optimizes for (writing code faster vs proving the code is trustworthy). Almost every vendor clusters in the SaaS + code-faster corner. The self-hosted and verification-first quadrant is nearly empty. That is where Factory sits.
The upper-right quadrant — self-hosted and verification-first — is the position. Dot placement is illustrative of category, not a benchmarked score.
Factory vs the alternatives
Scored on the dimensions a skeptical engineering buyer actually asks about. We only mark Yes where the fleet does the thing today.
| Dimension | Factory | SWE-bench-score vendors | SaaS agent platforms |
|---|---|---|---|
| Self-hosted | Yes — runs on your own k8s; local models via Ollama | No — hosted service | Rarely — cloud-first, some enterprise VPC |
| Independent verification | Yes — TFactory is a separate product that grades AIFactory's output; builder never signs its own homework | No — vendor reports its own benchmark score | No — the system that writes also declares "done" |
| Real test execution | Yes — tests generated and run in a per-task sandbox (Nix env), graded on a 5-signal verdict | Benchmark harness only — not your repo | Varies — often lint/build, not a graded test verdict |
| Governance / HITL | Yes — PFactory review gates with citations; human approval before code is emitted; approve/merge in the cockpit | No — score in, patch out | Some — PR review after the fact, not a gate before work |
| Audit / evidence | Yes — HMAC-anchored audit log, completion-event records, screenshot evidence for browser tests | No | Limited — activity logs, rarely tamper-evident |
| Data-egress control | Yes — self-hosted with local models; code and prompts need never leave your cluster | No — code goes to the vendor | No — cloud by default |
| Refuses to overclaim | Yes — verification-core drift guard on the never-overclaim path; VAL assurance levels; false-failed builds fixed rather than hidden | No — optimized to the leaderboard | No — "done" is asserted, not proven |
The proof, recorded
Captured from the running fleet. Nothing here is a mockup.
One task, end to end
Proven on real clouds
What we do NOT claim yet
The caps are the trust signal. We would rather show you the ceiling than let a demo imply we cleared it.
- Verification assurance is capped at VAL-2. Higher assurance levels are designed but not yet the default — we do not claim VAL-3 evidence.
- Deployment is dry-run first. The deploy lane plans and dry-runs real infrastructure; full unattended production rollout is deliberately gated, not "click once and ship."
- Concurrency is bounded. Per-task Nix environments on a single node cap throughput at roughly a handful of concurrent tasks today — scaling is in progress, not finished.
- It is an early, open project. No revenue, market-share or named-customer claims. The scenarios in these docs are illustrative of how a team would use Factory.