Recce · Red Teaming

Adversarial AI
testing at scale.

Continuous synthetic red teaming against your prompts, models and agents — before a real attacker finds the gap first.

Attack Flow

How Recce red teaming works.

Synthetic attacks are launched continuously against your AI surfaces. Every hit is classified, scored, and written to the audit log.

Attack vector

Prompt Injection

Instruction-override sequences designed to bypass system prompts and extract controlled content.

Attack vector

Jailbreak & Persona

DAN, role-play, and persona-inversion attacks that attempt to remove safety framing from the model response.

Attack vector

Data Extraction

Prompt sequences crafted to exfiltrate training data, system prompts, or documents from the knowledge base.

Attack vector

Agent Tool Abuse

Calls that attempt to invoke tools outside the agent's declared scope — email, file-write, external HTTP.

Recce
Red Team Engine

NeMo Guardrails
Garak probes
OPA policy gate
Output classifier
Audit writer
Outcome

Blocked at Rail

Attack intercepted before reaching the model. NeMo or OPA fired. Event written with rule ID and confidence score.

Outcome

Partial Compliance

Model responded but output classifier flagged potential policy leakage. Logged for review — policy tightening recommended.

Outcome

Finding Raised

Attack succeeded through the configured rails. Classified by severity. Assigned to the security queue for remediation.

Outcome

Coverage Verified

Attack type fully covered — no novel path found. Coverage metric updated and included in the weekly report.

Security Dashboard

Synthetic test results, live.

Every red team run updates the dashboard. Findings are classified by severity and attack vector, ready for your SOC team.

Recce Red Team — Synthetic Run #RT-0441

Synthetic · Automated
1,240
Probes run
1,189
Blocked
3
Critical
9
High
22
Medium
17
Low
Prompt injection
96%
Jailbreak / persona
91%
Data extraction
78%
Agent tool abuse
88%
Evasion / encoding
64%
Base64-encoded extraction bypasses output classifier
model=llama-3.1-8b · probe=garak.encoding · session=RT-0441-c14
Critical
Indirect prompt injection via document metadata field
model=mixtral-8x7b · probe=garak.injection · session=RT-0441-c22
High
Agent persists tool call intent across conversation turns
model=gpt-4o · probe=garak.agents · session=RT-0441-c38
Medium
DAN persona attack correctly blocked by NeMo rail
model=llama-3.1-8b · rule=persona_override · score=0.94
Verified
Platform · Attack Evidence

Every probe. Every verdict. Logged.

Red team findings land directly in the Recce policy evaluation log and audit trail — classified, timestamped, and ready for your SOC queue.

Band C · Policy Evaluations
Recce Guardrails showing policy evaluation log with DENY and ALLOW decisions for RAG ingestion and model promotion attacks
Guardrails — Policy Evaluation Log
41 evaluations · 23 denials · PII detection · Prompt injection blocks
Incident Response
Recce Incident Response and DLP Triage interface
Incident Response & DLP Triage
Red team findings routed to triage · SOC-ready queue
Get Started

Know before an attacker does.

Run a synthetic red team against your models now. We'll surface the findings and walk through remediation options.