AI agents already write production code in your organisation. What they cannot do is prove their work — to you, to an auditor, or to a regulator. This plugin makes AI-assisted delivery verifiable, EU compliance evidence automatic, and your engineering policy enforced — with every claim backed by evidence, starting with this page.
CRA Article 14 reporting obligations apply from 11 September 2026.
Each one is silent until it is expensive. A production incident traced to unverified AI output, a regulator's information request, an audit finding — by the time any of them surfaces, the evidence you need either exists or it does not.
An agent will report the work as done and the tests as passing — sincerely, and sometimes wrongly. Multiply that by every developer, every day. Without an enforced verification step, "done" means "the agent believes it", and that belief is what ships to production.
How delivery is assured →CRA vulnerability-reporting obligations begin 11 September 2026 for anyone shipping software into the EU, with fines up to €15 million or 2.5% of worldwide turnover. GDPR is already live. The AI Act is phasing in. The evidence these regimes demand is tedious to produce by hand — so it isn't produced, until it is suddenly urgent.
The regulatory clock →Every change ships with a test. Migrations are reversible. Endpoints authorize. Your teams agreed to these — and they are checked, when at all, by whoever reviews the pull request and happens to remember. They do not fail loudly; they erode.
How policy is enforced →Five EU regimes now reach software directly. The dates below are fixed in law; the only variable is whether the evidence exists on the day someone asks for it.
GDPR has no start date left to wait for. gdpr-evidence records what personal data a repository actually holds and the legal basis for each item — and then fails a build when a new personal-data column arrives without one.
Answer as things stand today, not as they should be. Nothing leaves this page: no form, no tracking — the count happens in your browser. And to be honest up front: eight answers are not an audit; they are the reason to run one.
This is the real shape of /compliance-status output, on a sample repository. Only three outcomes exist — present, gap, not applicable — each with a pointer or a reason. "Looks fine" is not an outcome this tool produces, and every gap arrives already routed to the skill that closes it.
The report always carries the two dates that make this urgent rather than theoretical: CRA Article 14 reporting obligations apply from 11 September 2026, and full CRA application arrives on 11 December 2027.
Installed once by your platform team, active in every repository. Each track answers one of the three liabilities above — and every output is evidence a third party can check, never a self-assessment.
No dashboard, no new tool to log into. The agent your developers already use simply refuses to declare victory without proof. Here it has just edited a source file and claimed the work is done — and the gate holds the claim until the test suite has run against the code as it now stands.
A gate that blocks on guesswork gets switched off within a week — and takes the checks that did work with it. So a check may only block where its detection is exact; everything else warns. That restraint is why this survives contact with a real engineering organisation.
The contract that makes this safe to put in front of a regulator: the AI produces evidence and drafts, and accountability stays with named people. This is not a limitation we accepted — it is the design.
No artifact — ADR, SBOM, gap report, audit, regulatory filing — is ever marked Accepted, Compliant or Validated by Claude on its own initiative. Everything ships Proposed or Draft until a named person says otherwise.
Every report states one of exactly three outcomes per item — present, gap, or not applicable — each with a pointer or a rationale. "Looks good" is not an output this plugin produces; where it appears, it is a bug.
Nothing here transmits to an authority, emails, or posts to a reporting platform. It writes the text and computes the clock. Submission is an organizational act performed by a named human, and every draft says so on its first line.
Where a question genuinely does not apply — a migration tool with no rollback step, a framework the gates cannot see — the report says so, and says why. A report that never says "not applicable" is a report that has started inventing findings to look useful.
A repository can lower a gate's severity, never raise one — and a lowering is printed as a lowering. A team that switched a gate off must not be able to produce output that looks like a team that passed it. That property survives config edits by design.
A governance product that asserted its own quality would be its own counterexample. So this page is held to the same contract as everything else: 1452 automated checks across 23 suites run on every pull request, on two operating systems — and a script re-derives every number published here from its source and fails the build when a copy disagrees. It exists because we shipped three stale numbers in a single week.
Whether each skill activates when it should is tested blind: 891 blind judgements so far — three independent judges per prompt, no tools, one pinned model. A 2-of-3 split is recorded as FLAKY and never rounded up: the disagreement is the finding.
One round found that the incident-reporting skill would not have engaged on "attackers got into our build server" — a textbook severe incident under the CRA. All three judges missed it, for the same documented reason. The fix was a few words; the point is that we found it by measurement, before a customer did.
We claimed our own catalogue was near its ceiling — then noticed the claim was derived from nothing, and built the instrument instead: 600 selection judgements across 200 row-instances, four arms. Adding three skills cost zero routing regressions, twice. The claim died; the number replaced it.
Every validation round is recorded, including the discarded ones and the rounds that found a real gap. The clean final state is the least informative part of that record, so it is not the only part we keep.
Honest about scale, too. Ten prompts per skill judged three ways is a smoke test for gross failures, not a benchmark — and the ledger says so itself. It has found real defects in material its authors had read carefully several times, which is the entire argument for running it at all.
This is a Claude Code plugin — it rides the tooling your developers already use. No server to run, no data leaving your repositories, no per-seat dashboard. Your platform team installs it; the skills, the gates and the verification hook are active from the next session.
You tell us how your teams use AI today and which regimes apply to you. No slides from our side: the call ends with the repository chosen for the pilot audit.
The harness audit and the first /compliance-status run on one of your repositories. Out comes a real baseline: present, gap or not applicable — the same shape as the report above.
You decide on the gap report of your own code, not on a demo. If you proceed, installation is the two commands below — active from the next session.
claude plugin marketplace add <your-private-repository-url>
claude plugin install oltrematica-skills@oltrematica
Delivery is from a private repository dedicated to your contract: the URL is provided when the engagement starts. It is the same channel through which the 3.x releases have already shipped to our first customers in production.
Say "audit our harness" in any repository. Ten surfaces, each reported present, gap or not applicable — a baseline you can put in a steering-committee deck the same day.
Run /compliance-status. The regulatory position — CRA, GDPR, AI Act, PLD, EAA — inventoried per repository, with each gap routed to the skill that closes it.
The gates run on every change, in CI if you wish. The verification hook holds every "done" claim to a fresh test run. From here on the evidence accumulates as a side effect of normal work.
Nine ecosystems for test-command detection, and thirteen migration tools for the reversibility gate — including the two that have no rollback at all, which are reported as not applicable rather than as a gap. Proprietary licence; delivery is from a private per-customer repository. Talk to Oltrematica about an engagement.