Effects, not wording
Pass means the control held on the system.
Authorized harness
Authorized assurance for AI applications and agents. Authorized tests. Pass or Fail. No exploit cookbook. Aegis Red scores effects — undeclared tools stayed denied, secrets did not leak, a fake policy did not bind — with hashed workpapers.
You describe intents (seeds) and oracles. The harness talks to a system under test, records the transcript, and scores Pass or Fail. Production stays shadow-only unless you write otherwise.
Pass means the control held on the system.
Seeds are fixtures and oracles. Bypass payloads do not ship.
In-process twin first. Same seeds hit HTTP or a model API.
Hashed bundles: utterance, reply, tools, lineage, go / no-go.
The free wheel is a nudge across every industry pack. Complete coverage fills the remainder. Runtime subset of NIST AI RMF, ISO 42001, OWASP LLM, OSI OSAID 1.0, and the EU AI Act — agent behavior under fixtures, not a full management-system certificate.
Safety (Tech / AI): content safety, injection, instruction integrity, hallucination, privacy, tools, summarization, memory, swarm, lineage, change control, plus OWASP LLM, OSI OSAID 1.0, NIST AI 600-1 GAI, ISO/IEC 42001 Annex A, and EU AI Act auditor bodies.
Security: injection, instruction, privacy, tool contract, tool access, policy robustness, lineage.
Banking journeys (named rows), insurance skeleton, community seed store. Harness, portal, twin, live HTTP demo.
The rest of the industry corpus (~300 banking rows plus control functions), auditor folders (OSS, NIST, ISO, OWASP, AI Act) under each industry, repeat campaigns, and the live-hook expectation for tools and state.
Request it with org, contact, industry, and the problem you are trying to solve. We do not publish that remainder in the free wheel.
No product names. The distinction is the job, not the logo.
| Job | Typical elsewhere | Aegis Red |
|---|---|---|
| What is scored | Wording, attack strings, or a questionnaire | Effects on tools, state, and policy bind |
| What ships | Payload lists or a slide crosswalk | Intent + fixtures + oracles; hashed bundles |
| Where it runs | Chat-only or a mock that is not the agent | Twin, then live adapter returning tools and state |
| Mesh / swarm | Single-turn chat | Privilege union and silent handoff are first-class |
| Policy rewrite claim | Often out of scope or a recipe | Intercept: claimed policy, tools that must stay denied |
| Evidence | A score and a screenshot | Lineage (model / prompt / actor / tool) and go / no-go |
| What we will not claim | — | We do not find every vulnerability. We do not issue the certificate. |
Go / no-go on whether AI directives held. Board line with hashed proof. Complete coverage maps to the internal risk register — not a slide.
Install the signed wheel, run the twin, then point the same seeds at a live agent. Adapter work is time, not a payload library.
Portal in plain language. Add a test, run it, download Pass/Fail CSV. No YAML required for the first two seeds.
Free install after a short form. Prove the method on safety and a vertical nudge before you buy depth.
Domain families and examiner-sample bundles when complete coverage is enabled. Runtime subset — not vendor DD or training-data audits.
You sample workpapers. You do not buy the product. Lineage is on the material act.
Tell us whether you are an organization, a startup, or an individual. We give you pip install aegis-red from PyPI (signed wheel, currently 0.2.3). Site: aeigisred.ai. GitHub source stays private. Authorized tests. Pass or Fail. No exploit cookbook.
We need the organization, who to contact, the industry you are trying to solve, and the use case. That is how the remainder is scoped — not a public download.