The scarcest skill in AI security is not offensive security — it is understanding how production agent systems actually fail: tool schemas, memory persistence, retrieval scoping, MCP wiring, orchestration loops. Most AI red teamers came from application security and have never shipped an agent. We ship them weekly, which is why our findings come back as pull requests rather than screenshots.
Why teams choose us.
Built By People Who Build Agents
We run a production agentic practice. The review is done by engineers who have debugged tool-permission bugs at 2am, not by testers reading an architecture diagram.
Fixes, Not Findings
Remediation arrives as pull requests against your repository, with a regression eval harness wired into your CI so the same class of failure cannot come back.
The Evidence Pack
The 2026 SIG added an expanded AI governance section and CAIQ now names ISO 42001 and NIST AI RMF. You get answers to those questions, grounded in your actual system.
Independent Of Your Build Team
We do not review agents withRemote built. Independence from the builder is explicit in 2026 practice and sophisticated buyers check for it.
The full menu.
Threat Model & Blast Radius
- Tool surface and permission mapping
- What the agent can reach on its worst day
- Excessive-agency and privilege paths
- Human-in-the-loop gate placement
Adversarial Testing
- Direct and indirect prompt injection
- Tool abuse and unauthorised action chains
- Memory and conversation-state leakage
- Retrieval and vector-store permission leakage
MCP & Tool-Schema Review
- Server trust boundaries
- Tool-description poisoning surface
- Supply chain of connected servers
- Schema validation and output handling
AI Governance Evidence Pack
- SIG and CAIQ AI-section answers
- Model provenance and subprocessor map
- Data flow and training-data position
- Mapped to ISO/IEC 42001 and NIST AI RMF
Remediation & Regression
- Fixes shipped as pull requests
- Guardrail and output-handling layer
- Eval harness running in your CI
- Optional continuous re-test retainer
Our process.
Scope & Authorisation
We agree the systems in scope and sign Rules of Engagement plus a countersigned authorisation letter from a verified asset owner. Default scope is staging and non-destructive.
Model & Map
Architecture review, tool inventory, permission mapping and blast-radius analysis. Most of the serious findings surface here, before a single payload is sent.
Adversarial Pass
Prompt injection, tool abuse, memory and retrieval leakage, MCP trust boundaries — run against the real system, with every attempt logged.
Fix & Evidence
Remediation as pull requests, a regression eval suite in your CI, and the governance evidence pack your buyer's questionnaire is asking for.
What we build with.
Where this review is worth buying
Default scope is staging and non-destructive, which makes this the most portable security service we offer — the constraint is authorisation, not residency.
Full scope. The buying trigger is procurement: the 2026 SIG update added an expanded AI governance section and CAIQ now names ISO 42001 and NIST AI RMF outright.
No production testing without signed Rules of Engagement and a countersigned authorisation letter from a verified asset owner, plus third-party authorisation where the systems sit on AWS or Azure.
Full scope for private-sector clients. A useful secondary trigger: the April 2026 Cyber Essentials Danzell question set brought AI and LLM tools formally into scope as cloud services, so anyone who deployed an agent since their last renewal now has a certification problem they did not have before.
We hold no CREST or CHECK accreditation and will not imply otherwise. UK public sector is closed.
Deliverable, but do not buy it as regulatory readiness. The Digital Omnibus, in force 27 July 2026, deferred Annex III high-risk obligations to 2 December 2027 and Annex I to 2 August 2028. Only Article 50 transparency and Article 4 AI literacy are live.
We will tell you that rather than sell against a deadline that has moved. If you ship a product with digital elements into the EU, the Cyber Resilience Act reporting clock from 11 September 2026 is the real and nearer obligation.
The sharpest regulatory hook for this service anywhere in the world. The DIFC's June 2026 consultation proposes embedding AI safety into processing systems, certification obligations, and an Autonomous Systems Officer role — nobody else is being asked this yet.
We do not certify ISO 42001 and cannot; only an accredited certification body can. Our output is an evidence pack that answers procurement questions, not a certificate.
Every market position we hold, with the legal reason and the transfer mechanism, is on where we work.
Choose this if...
Honest about who this is for.
This will be a fit.
- You have a production agent with real tool access and real data behind it
- Your enterprise deal is blocked on an AI governance section you cannot answer
- You want remediation as pull requests against your codebase
- You want a regression suite so the finding does not come back next quarter
Honestly — not our zone.
- —You want an independent audit, attestation or certification — this is an engineering review and we certify nothing
- —You want ISO 42001 certified or an EU AI Act conformity assessment. Only an accredited body can do the first; we do neither
- —withRemote built the agent. We do not review our own work — our build clients get the Secure Build tier instead
- —You want production testing without signed Rules of Engagement and an authorisation letter. We will not proceed without them































