This document describes a scoped six-to-eight-week design-partner pilot with Ethosure. The goal is to take one AI agent already running in your environment, instrument it with Ethosure’s enforcement layer, and produce an auditor-ready evidence pack.
Three objectives drive the pilot: (1) validate technical fit – confirm the enforcement layer integrates with your agent infrastructure without disrupting legitimate work; (2) establish a compliance baseline – capture the agent’s current policy compliance posture so improvements are measurable; and (3) produce a usable evidence pack mapped to NIST AI RMF, ISO/IEC 42001, and EU AI Act requirements, demonstrating that controls were active throughout.

Scope
In scope: One agent already deployed; one deployment shape (Git hooks, agent tool-use hooks, or MCP proxy, selected in Week 1); a standard coding-ethos policy set (security-by-design, validation-at-the-gate, radical-visibility); the evidence ledger and SARIF output; one stakeholder readout.
Out of scope: Multiple agents; custom policy authoring; SIEM/GRC integration (phase-two); Agent Proxy shape (early development); remediation of violations found during the pilot.
Workflows Covered
Three workflow patterns are observed and enforced: code-write (file writes and code changes evaluated against security and quality policy), command-execution (shell commands intercepted, scoped, and sandboxed), and policy-advisory (agent queries the MCP server via `policy_check_command` or `policy_check_edit` before acting).
Project Plan: Phases and Week-by-Week Timeline
- Phase 1 – Discovery and Setup (Weeks 1–2): Kick-off meeting; review target agent toolchain and access scope; select deployment shape; install enforcement layer in observation mode (records decisions, does not block).
- Phase 2 – Instrumentation and Baseline (Weeks 3–4): Activate enforcement on a limited policy set; weekly finding review to confirm dispositions and adjust thresholds; instrument policy-advisory workflow; begin NIST AI RMF and ISO/IEC 42001 mapping.
- Phase 3 – Live Enforcement (Weeks 5–6): Full enforcement runs without further calibration changes – the record must reflect stable, auditable policy behavior. Escalate dispositions worked through human-review process; outcomes recorded. Compliance team maps SARIF output to EU AI Act requirements.
- Phase 4 – Evidence Pack Assembly (Week 7): Raw enforcement data compiled into the structured evidence pack; reviewed for completeness; readout presentation prepared.
- Phase 5 – Readout and Decision (Week 8): Formal presentation to governance, risk, compliance, and engineering stakeholders. Decision memo produced: proceed to full deployment, expand scope, or defer – based on evidence.
| Week | Phase | Key Milestone |
| 1–2 | Discovery & Setup | Enforcement layer live in observation mode |
| 3–4 | Instrumentation | Enforcement active; calibration complete |
| 5–6 | Live Enforcement | Full enforcement record accumulated |
| 7 | Evidence Pack | Pack assembled and reviewed |
| 8 | Readout | Decision memo issued |
Roles and Resources
Your organization: Engineering lead (~0.5 day/week during active phases); governance/compliance lead (~0.25 day/week); risk or security lead (kick-off and readout). Total commitment: approximately 3–4 hours per week for the engineering lead.
Ethosure: Technical lead (deployment, configuration, evidence output); governance advisor (compliance mapping, evidence pack assembly).
Success Criteria and KPIs
The pilot succeeds when all of the following are met: (1) enforcement layer ran continuously for at least 10 business days; (2) at least 200 agent actions evaluated and recorded (active agents typically produce 500–2,000 actions per week); (3) evidence ledger complete with all findings traceable to policy rules; (4) compliance mapping complete across NIST AI RMF, ISO/IEC 42001, and EU AI Act categories; and (5) stakeholder readout delivered and decision memo issued.
Stretch KPI: if token spend is tracked, a baseline-vs-pilot comparison is included. Per IBM Cost of a Data Breach 2025, the average financial-services breach costs $5.56 million; the evidence pack documents control posture improvements that reduce that exposure.
Deliverables: The Evidence Pack
The evidence pack contains: an enforcement summary report (actions evaluated, disposition breakdown, top policy rules, trend); a SARIF findings file suitable for GRC platform ingestion; a compliance mapping document (NIST AI RMF functions, ISO/IEC 42001 clauses, EU AI Act technical documentation – August 2026 deadline, Article 99 penalties up to €35 million or 7% of global turnover); a policy scope document (versioned, signed policy bundle); a calibration log (Phase 2 adjustments with rationale); an escalation log (human decisions and outcomes); and a decision memo (one-page summary and agreed next step).
Risks and Mitigations
| Risk | Mitigation |
| Irregular agent activity during pilot window | Two-week observation phase confirms activity levels before enforcement activates |
| Policy calibration requires more iteration | Phase 2 is explicitly a calibration phase; enforcement locks in Phase 3 |
| Engineering resource constraints | ~3–4 hours/week ceiling; schedule aligned at kick-off |
| Findings volume too low | Observation-mode data in Weeks 1–2 signals early; scope can expand to a second agent |
| Agent toolchain changes mid-pilot | Engineering lead freezes toolchain changes during Phase 3 (Weeks 5–6) |
The Next Step
The pilot begins with one conversation: which agent is already running, and what does your team need to demonstrate to governance and compliance stakeholders? Ethosure scopes the deployment shape, policy set, and compliance mapping targets within a week. The enforcement layer can be live in observation mode within two weeks. The evidence pack is eight weeks away.