Let agents build. Keep control.
Governed infrastructure for autonomous AI agents. Cortex gives agents the runtimes, deployments, and organizational memory to do real work, while capability policy, human approvals, and an append-only audit trail keep them inside the boundaries an organization sets. Each customer gets one isolated stack rather than a shared tenancy.
Cortex
ShippingShipping. Self-serve signup, published pricing, and a public compliance library. The largest thing the lab operates.
Python · TypeScript · Docker Compose · Terraform · WorkOS · Stripe
Agents can now change real systems. That is the whole problem. An agent useful enough to deploy code, provision infrastructure, and run experiments is an agent capable of doing damage, and the usual answer — keep a human in the loop for everything — gives back the leverage that made it worth running.
Cortex separates the two questions. Human roles and agent capabilities are declared independently. Policy bounds which tools and which infrastructure an agent can reach; approvals gate the actions that warrant a person; and every decision lands in an append-only audit trail. Autonomy stays high inside a boundary somebody drew on purpose.
Three things it does
Build
Coding and operations agents get durable projects rather than disposable chat sessions — persistent runtimes, deployments, tools, and organizational memory that survives between tasks.
- Durable project workspaces
- Runtimes and deployments
- Organizational memory
Prove
A hypothesis becomes an experiment. Cortex provisions the isolated world a question needs, runs the software, benchmark, simulator, or prover, and keeps the result reproducible afterwards.
- Isolated experiment worlds
- Reproducible results
- Software, benchmarks, simulators, provers
Govern
Human roles and agent capabilities are separate concerns. Policies bound tools and infrastructure, approvals gate dangerous actions, and every decision reaches an append-only audit trail.
- Capability policy
- Human approvals
- Append-only audit
What Cortex is made of
- EngramA private knowledge graph of how a company actually works, tended nightly by a gardener agent so it stays current rather than accumulating.
- GauntletThe safety and capability benchmark Cortex measures itself against — jailbreak resistance across text, image and audio, code quality, generative capability, and whole-repository builds.
- OMPA pre-configured coding agent that arrives governed rather than needing to be fenced in afterwards.
- Cortex ControlThe operations surface for R&D and delivery — where policies, approvals, deployments and the audit trail are actually read and acted on.
- SemaThe neurosymbolic language underneath, where model inference and formal verification are equal first-class citizens. Pre-production; see the research index.
Measured, with its method
| Measurand | Value | Method | Uncertainty |
|---|---|---|---|
| Benchmark score, inside Cortex | 98 / 100 | Gauntlet, 96 runs | self-reported, undated |
| Benchmark score, same model standalone | 88 / 100 | Gauntlet, 96 runs | self-reported, undated |
| Jailbreak success rate | −47% | text, image and audio tracks | self-reported, undated |
| Requirements met | 81% → 98% | code-quality track | self-reported, undated |
| Whole-repository builds | +72% | full-repo build track | self-reported, undated |
Built for the questions a compliance team asks
One stack per organization
Isolation is the architecture, not a configuration flag. Each organization gets its own provisioned stack rather than a row in a shared table.
EU data residency
Stacks provision into named EU regions. The default posture is European rather than an afterthought bolted onto a US deployment.
Conformity-ready, not certified
Twelve documents in three bands — management system, information security, AI management — mapped to ISO 9001, ISO/IEC 27001, ISO/IEC 42001 and the EU AI Act. Published in full, and honest about what a certificate would still require.
- ISO 9001:2015
- ISO/IEC 27001:2022 and GDPR Art. 28/33/34
- ISO/IEC 42001:2023 and the EU AI Act
If a stack can govern science end to end, your apps and agents are the easy part.
Go and look
- Cortex CloudProduct site, pricing, and self-serve signup
- The Gauntlet benchmark96 runs across four tracks
- Compliance libraryTwelve documents, three bands, publicly readable
- Field notes on agentic engineering
- Install Cortex
- DWG
- A2O-000
- Sheet
- 03 of 09
- Rev
- released
- Scale
- 1:1