This document describes a scoped six-to-eight-week design-partner pilot with Ethosure. The goal is to take one AI agent already running in your environment, instrument it with Ethosure’s enforcement layer, and produce an auditor-ready evidence pack.

Three objectives drive the pilot: (1) validate technical fit – confirm the enforcement layer integrates with your agent infrastructure without disrupting legitimate work; (2) establish a compliance baseline – capture the agent’s current policy compliance posture so improvements are measurable; and (3) produce a usable evidence pack mapped to NIST AI RMF, ISO/IEC 42001, and EU AI Act requirements, demonstrating that controls were active throughout.

Scope

In scope: One agent already deployed; one deployment shape (Git hooks, agent tool-use hooks, or MCP proxy, selected in Week 1); a standard coding-ethos policy set (security-by-design, validation-at-the-gate, radical-visibility); the evidence ledger and SARIF output; one stakeholder readout.

Out of scope: Multiple agents; custom policy authoring; SIEM/GRC integration (phase-two); Agent Proxy shape (early development); remediation of violations found during the pilot.

Workflows Covered

Three workflow patterns are observed and enforced: code-write (file writes and code changes evaluated against security and quality policy), command-execution (shell commands intercepted, scoped, and sandboxed), and policy-advisory (agent queries the MCP server via `policy_check_command` or `policy_check_edit` before acting).

Project Plan: Phases and Week-by-Week Timeline

  • Phase 1 – Discovery and Setup (Weeks 1–2): Kick-off meeting; review target agent toolchain and access scope; select deployment shape; install enforcement layer in observation mode (records decisions, does not block).
  • Phase 2 – Instrumentation and Baseline (Weeks 3–4): Activate enforcement on a limited policy set; weekly finding review to confirm dispositions and adjust thresholds; instrument policy-advisory workflow; begin NIST AI RMF and ISO/IEC 42001 mapping.
  • Phase 3 – Live Enforcement (Weeks 5–6): Full enforcement runs without further calibration changes – the record must reflect stable, auditable policy behavior. Escalate dispositions worked through human-review process; outcomes recorded. Compliance team maps SARIF output to EU AI Act requirements.
  • Phase 4 – Evidence Pack Assembly (Week 7): Raw enforcement data compiled into the structured evidence pack; reviewed for completeness; readout presentation prepared.
  • Phase 5 – Readout and Decision (Week 8): Formal presentation to governance, risk, compliance, and engineering stakeholders. Decision memo produced: proceed to full deployment, expand scope, or defer – based on evidence.
See also  Why Trust Is Now a Product Feature in Agentic AI
Week Phase Key Milestone
1–2 Discovery & Setup Enforcement layer live in observation mode
3–4 Instrumentation Enforcement active; calibration complete
5–6 Live Enforcement Full enforcement record accumulated
7 Evidence Pack Pack assembled and reviewed
8 Readout Decision memo issued

Roles and Resources

Your organization: Engineering lead (~0.5 day/week during active phases); governance/compliance lead (~0.25 day/week); risk or security lead (kick-off and readout). Total commitment: approximately 3–4 hours per week for the engineering lead.

Ethosure: Technical lead (deployment, configuration, evidence output); governance advisor (compliance mapping, evidence pack assembly).

Success Criteria and KPIs

The pilot succeeds when all of the following are met: (1) enforcement layer ran continuously for at least 10 business days; (2) at least 200 agent actions evaluated and recorded (active agents typically produce 500–2,000 actions per week); (3) evidence ledger complete with all findings traceable to policy rules; (4) compliance mapping complete across NIST AI RMF, ISO/IEC 42001, and EU AI Act categories; and (5) stakeholder readout delivered and decision memo issued.

Stretch KPI: if token spend is tracked, a baseline-vs-pilot comparison is included. Per IBM Cost of a Data Breach 2025, the average financial-services breach costs $5.56 million; the evidence pack documents control posture improvements that reduce that exposure.

Deliverables: The Evidence Pack

The evidence pack contains: an enforcement summary report (actions evaluated, disposition breakdown, top policy rules, trend); a SARIF findings file suitable for GRC platform ingestion; a compliance mapping document (NIST AI RMF functions, ISO/IEC 42001 clauses, EU AI Act technical documentation – August 2026 deadline, Article 99 penalties up to €35 million or 7% of global turnover); a policy scope document (versioned, signed policy bundle); a calibration log (Phase 2 adjustments with rationale); an escalation log (human decisions and outcomes); and a decision memo (one-page summary and agreed next step).

Risks and Mitigations

Risk Mitigation
Irregular agent activity during pilot window Two-week observation phase confirms activity levels before enforcement activates
Policy calibration requires more iteration Phase 2 is explicitly a calibration phase; enforcement locks in Phase 3
Engineering resource constraints ~3–4 hours/week ceiling; schedule aligned at kick-off
Findings volume too low Observation-mode data in Weeks 1–2 signals early; scope can expand to a second agent
Agent toolchain changes mid-pilot Engineering lead freezes toolchain changes during Phase 3 (Weeks 5–6)
See also  Local-First and Sovereign-Ready: Why No Hosted Dependency Matters

 

The Next Step

The pilot begins with one conversation: which agent is already running, and what does your team need to demonstrate to governance and compliance stakeholders? Ethosure scopes the deployment shape, policy set, and compliance mapping targets within a week. The enforcement layer can be live in observation mode within two weeks. The evidence pack is eight weeks away.

Subscribe to Ethosure's Newsletter to get monthly updates on AI Governance

We don’t spam! Read our privacy policy for more info.