All work

Let agents build. Keep control.

Governed infrastructure for autonomous AI agents. Cortex gives agents the runtimes, deployments, and organizational memory to do real work, while capability policy, human approvals, and an append-only audit trail keep them inside the boundaries an organization sets. Each customer gets one isolated stack rather than a shared tenancy.

Cortex

Shipping

Shipping. Self-serve signup, published pricing, and a public compliance library. The largest thing the lab operates.

Python · TypeScript · Docker Compose · Terraform · WorkOS · Stripe

Agents can now change real systems. That is the whole problem. An agent useful enough to deploy code, provision infrastructure, and run experiments is an agent capable of doing damage, and the usual answer — keep a human in the loop for everything — gives back the leverage that made it worth running.

Cortex separates the two questions. Human roles and agent capabilities are declared independently. Policy bounds which tools and which infrastructure an agent can reach; approvals gate the actions that warrant a person; and every decision lands in an append-only audit trail. Autonomy stays high inside a boundary somebody drew on purpose.

Three things it does

  • Build

    Coding and operations agents get durable projects rather than disposable chat sessions — persistent runtimes, deployments, tools, and organizational memory that survives between tasks.

    • Durable project workspaces
    • Runtimes and deployments
    • Organizational memory
  • Prove

    A hypothesis becomes an experiment. Cortex provisions the isolated world a question needs, runs the software, benchmark, simulator, or prover, and keeps the result reproducible afterwards.

    • Isolated experiment worlds
    • Reproducible results
    • Software, benchmarks, simulators, provers
  • Govern

    Human roles and agent capabilities are separate concerns. Policies bound tools and infrastructure, approvals gate dangerous actions, and every decision reaches an append-only audit trail.

    • Capability policy
    • Human approvals
    • Append-only audit

What Cortex is made of

  • EngramA private knowledge graph of how a company actually works, tended nightly by a gardener agent so it stays current rather than accumulating.MemoryShipping
  • GauntletThe safety and capability benchmark Cortex measures itself against — jailbreak resistance across text, image and audio, code quality, generative capability, and whole-repository builds.BenchmarkShipping
  • OMPA pre-configured coding agent that arrives governed rather than needing to be fenced in afterwards.Coding agentShipping
  • Cortex ControlThe operations surface for R&D and delivery — where policies, approvals, deployments and the audit trail are actually read and acted on.Control planeShipping
  • SemaThe neurosymbolic language underneath, where model inference and formal verification are equal first-class citizens. Pre-production; see the research index.Runtime and compilerResearch prototype

Measured, with its method

Self-reported. These are the lab’s own Gauntlet runs, not a third-party evaluation, and the comparison is the same model inside and outside Cortex.
MeasurandValueMethodUncertainty
Benchmark score, inside Cortex98 / 100Gauntlet, 96 runsself-reported, undated
Benchmark score, same model standalone88 / 100Gauntlet, 96 runsself-reported, undated
Jailbreak success rate−47%text, image and audio tracksself-reported, undated
Requirements met81% → 98%code-quality trackself-reported, undated
Whole-repository builds+72%full-repo build trackself-reported, undated

The Gauntlet

Built for the questions a compliance team asks

  • One stack per organization

    Isolation is the architecture, not a configuration flag. Each organization gets its own provisioned stack rather than a row in a shared table.

  • EU data residency

    Stacks provision into named EU regions. The default posture is European rather than an afterthought bolted onto a US deployment.

  • Conformity-ready, not certified

    Twelve documents in three bands — management system, information security, AI management — mapped to ISO 9001, ISO/IEC 27001, ISO/IEC 42001 and the EU AI Act. Published in full, and honest about what a certificate would still require.

    • ISO 9001:2015
    • ISO/IEC 27001:2022 and GDPR Art. 28/33/34
    • ISO/IEC 42001:2023 and the EU AI Act
If a stack can govern science end to end, your apps and agents are the easy part.

Go and look

DWG
A2O-000
Sheet
03 of 09
Rev
released
Scale
1:1