Detection Triage
Does the model correctly distinguish true positives from noise under time pressure?
Measures AI models on remediating active runtime threats — live exploit attempts, anomalous behavior, workload quarantine decisions, and compensating-control authoring.
Each model response is judged on these dimensions by an LLM-as-Judge calibrated against expert annotations.
Does the model correctly distinguish true positives from noise under time pressure?
Are containment actions proportionate and reversible?
Are temporary controls correctly authored when a permanent fix is not yet possible?
Is the recommended response sequence clear, ordered, and operator-ready?
Does the model combine EDR, network, and audit signals to scope the incident?
An illustrative case to show how the model is presented with a problem and how its response is scored.
Suspicious outbound connection from pod payments-api to 185.x.x.x on port 4444; reverse-shell pattern.
Generate a containment plan. Decide whether to kill the pod, isolate it, or apply a NetworkPolicy.
Quarantine via NetworkPolicy denying egress except to in-cluster services. Snapshot memory before kill. Trigger incident workflow. Flag CI image for SBOM re-scan.
Suggestions that immediately kill the pod (destroying evidence) are penalized.
All Remediation Labs benchmarks evolve through three versions, increasing in difficulty and realism.
Models are given the raw security finding and asked to recommend a remediation. No environment context provided.
Models receive the finding plus relevant system, code, and policy context. Evaluated on contextual reasoning quality.
Models must produce a remediation that can be safely executed end-to-end — code patches, IaC changes, runbook steps — with verification.
Get model leaderboards, scoring methodology, dataset details, and roadmap access — or discuss generating proprietary remediation datasets for your environment.
Or email us directly at info@remediationlabs.com