Security tests for LLM apps, written in pytest.

Your app has failure modes your test suite can't see. It'll follow an instruction hidden in a document it just retrieved. It'll repeat a key you told it to keep. It'll act on an authorisation the caller just typed in.

LLMSecTest attacks your running app for all ten OWASP LLM Top 10 categories and tells you what it got out. MIT-licensed, on PyPI, built in the open.

pip install llmsectest Run your first scan »

Just pytest

No separate scanner to run. Your security checks are pytest tests. Same command, same build gate, same plumbing as the rest of your suite.

How it works »

Mapped & scored

Every probe maps to an OWASP LLM Top 10 category and carries a CVSS v4.0 base score. So you can triage a finding instead of reading a wall of text.

The coverage map »

CI-native

Reports come out as SARIF v2.1.0, which your code-scanning dashboard already reads. HTML, JSON and Markdown too. No glue code.

A sample finding »

Quick-start

Get a scan running in seconds.

One line to install, then point it at a model or at your running app. Read the SARIF. There's a full walkthrough in the docs.

bash
~ $ pip install llmsectest
~ $ llmsectest --target anthropic:claude-3-5-haiku
~ $ llmsectest --target app:http://localhost:8000/chat   # your live app
~ $ llmsectest --target app:http://localhost:8000/chat \
      --report-formats=sarif,html,json,markdown
# => OWASP LLM Top 10 categories probed · findings written to results/<target>.sarif

OWASP LLM Top 10 · honest status

What's covered and what isn't yet

All ten OWASP LLM Top 10 (2025) categories are implemented and tested today. Nothing here got marked done before it was. What's left to do is depth. If a run can't reach a category, because you didn't hand it a repo, a model path or an app marker, it says so as an explicit skip. You won't get a silent gap.

A scan reads the whole response, not just the reply field. An answer can come back in a sibling field: one application returned a planted secret in message.reasoning while message.content held a refusal. Probes that plant their marker in your application are scored against the whole body. The finding says where the value turned up.

The numbering below is the 2025 edition. Older editions numbered these differently. Supply chain used to be LLM05 and is now LLM03, so a number from an older list won't line up with what the reports say.

There's a 2026 edition. This tool doesn't implement it yet. It came out on 3 August 2026. I've read it against what's here: nothing was added, dropped, merged or split, but eight of the ten categories change number and System Prompt Leakage becomes Hidden Context Exposure with a wider remit. Every report and every stored baseline this project has published carries 2025 numbers, so renumbering quietly would change what those records mean. OWASP's own category pages still show 2025 too. So I'm staying on 2025, saying the year everywhere, and I'll move once there's a way to keep the old records readable.

  • LLM01Prompt Injectiondone
  • LLM02Sensitive Information Disclosuredone
  • LLM03Supply Chain (Python, npm and Go manifests, OSV)done
  • LLM04Data & Model Poisoningdone
  • LLM05Improper Output Handlingdone
  • LLM06Excessive Agencydone
  • LLM07System Prompt Leakagedone
  • LLM08Vector & Embedding Weaknessesdone
  • LLM09Misinformationdone
  • LLM10Unbounded Consumptiondone

done implemented & tested, 10/10 · depth improvements continue on the roadmap

What "done" still doesn't tell you. Read this before you trust a clean row. LLM06 (excessive agency) can only report what your app says. If yours describes an action in prose instead of emitting the signature you passed in, a clean row means "we didn't see it", not "your app is fine". We write every limitation we find into the changelog. It comes out again when it's fixed.

So the coverage footer names how many cases each category ran. An app-target scan prints exercised: LLM01 (13 cases), LLM05 (3 cases), …, because the two ways in send different amounts: --target app:<url> runs the packaged suite and delivers 38 attacks at full inputs, where the Python API sends 23. Thirteen cases are a different result from one case. A row that renders both the same way is the thing this tool exists to catch.

Roadmap · shipped, building, planned

Where it's going

Built in the open across the funding period, which runs to the end of November 2026. Each phase ships before we claim the next one. The status here tracks what the code really does today.

  1. 01

    Foundation

    shipped
    • pytest-native framework & plugin
    • One adapter for OpenAI · Anthropic · Hugging Face · Ollama & LM Studio (local), with a fail-fast --preflight health check
    • Reports in SARIF v2.1.0 · HTML · JSON · Markdown
    • CVSS v4.0 scoring with OWASP mapping
    • CLI & documentation site
  2. 02

    OWASP coverage

    shipped
    • All 10 of 10 categories live: LLM01 to LLM10, the complete OWASP LLM Top 10 (2025)
    • Black-box testing of a running application, --target app:<url>, reaching 8 categories once you name what the app holds (--app-prompt/-secret/-action/-canary/-rag-poison); LLM01/05/09/10 need nothing
    • White-box scans: dependency manifests with an OSV known-CVE lookup (--repo, --osv, LLM03), and an offline serialization-opcode scan of model files that never unpickles them (--model-scan, LLM04)
    • A white-box LLM08 read of your persisted vector store (--vector-store), reporting how much of the corpus anybody with read access already holds. It leaves your embeddings uninverted
    • LLM02 and LLM06 attack along four mechanisms each rather than four wordings. A leaked secret counts encoded, Unicode-disguised or split character by character
    • RAG-specific LLM08: retrieval exposure (--app-canary) and indirect prompt injection through a poisoned retrieved document (--app-rag-poison), scored so one probe's marker cannot void the other's row
    • Guardrails under load (--app-stress N): N probes at once, reporting only a guardrail that held singly and stopped holding together — and saying so with the numbers when the requests never actually overlapped
    • Red-team jailbreak set (JailbreakBench / AdvBench, --redteam-set) plus a false-refusal rate on the benign twins (--redteam-benign)
    • Reports that can't flatter your app: a probe that got no answer is inconclusive and never a pass, the status reads INCOMPLETE, and a reply identical to what we sent stops with an error instead of scoring. See it on our whole test cohort » and on one result written up alone: your secret-disclosure test passed, the secret still left »
    • Attacks the model writes itself (--redteam-generate N), added to the hand-written set rather than replacing it, with every rewrite that lost its marker or that a hardened target also answers thrown out and counted
    • A PDF report with no dependency at all (--render-pdf), written directly from SARIF rather than converted from HTML: a 60-finding report is 7 pages and 8 KB, with anything that stops it being a pass on page one
    • Standalone HTML from any SARIF file (--render-sarif), proven against committed output from ruff, Bandit and Semgrep

    Release-by-release detail lives in the changelog.

  3. 03

    Depth & reports

    in progress
    • A classifier refusal oracle (GLiGuard / Llama-Guard), replacing the substring rule
    • The last LLM08 white-box dimension: embedding-store poisoning
    • Deeper supply-chain analysis and stress tests
    • CycloneDX SBOM generation, --sbom <path> (LLM03)
    • LLM08 multi-tenant namespace isolation, read from a persisted store with --vector-store <path>: one store holding several tenants' vectors, where a forgotten retrieval filter returns the union
    • A remediation database: every OWASP category carries its own fix steps, CWE ids, compliance mappings and a representative CVSS v4.0 vector, written into the HTML report beside the failures they answer
    • A plugin API — the tool is a pytest11 plugin, so your own probes and oracles register the way any pytest plugin does
    • PDF reports with no dependency at all, --render-pdf, written straight from SARIF
    • Donating the tool to OWASP: the Board accepted the repository donation in principle on 18 September 2026, and the formal process is under way

    The three items at the top are what is still open. Everything marked shipped below them is in the code today. Four of those were planned for after v1.0. They arrived during the funding period instead.

  4. 04

    Integrations & v1.0

    planned, this funding period
    • CI/CD templates for GitHub Actions · GitLab CI · Jenkins
    • v1.0 on PyPI by 2026-11-30, the last day of the funding period. Where it stands today: pip install llmsectest serves 0.3.0, and 0.4.0 is cut and waiting to be uploaded, from 918 passing tests at 86% branch coverage. The remaining distance to 1.0 is this phase's CI/CD templates and the two open Phase 03 promises above it, not new scope.
  5. 05

    Past v1.0

    planned, beyond this funding period
    • Hardening & independent security audit
    • Broader ecosystem integrations beyond the CI templates in phase 04

    Three things that sat here until October 2026 have shipped early. PDF reports, the remediation database and the plugin API are now in phase 03. So is the OWASP community submission.

shipped in the code today · in progress being built now · planned ahead. The last two phases say which side of the funding period they sit on.

Output · the target format

Findings you can act on

A run gives you SARIF that drops straight into GitHub code scanning, GitLab, or any SARIF viewer. Each finding carries its OWASP category, a CVSS vector and a pointer to the fix. Here's the shape of it.

Illustrative. This is the format we target. We didn't run this scan.

Read a finding » The secret-disclosure probes found one thing across the 52 applications scanned on 31 August 2026; 32 of them gave the credential up anyway.

Read another » An output filter that holds only the hash of the credential has to guess where the credential ends. Twelve spellings of one value, seven of them straight through. Both write-ups »

See real ones » Every report from our own regression cohort, unedited: dozens of LLM applications scanned black-box, including the categories that found nothing.

report.sarif
{
  "ruleId": "LLM01-prompt-injection",
  "level": "error",
  "properties": {
    "owasp": "LLM01:2025 Prompt Injection",
    "cvss":  "CVSS:4.0/AV:N/AC:L/.../VC:H",
    "score": 9.2
  },
  "message": {
    "text": "System prompt recovered via instruction override."
  }
}

Built in the open, for the people shipping LLM features.

App developers, security leads, researchers. The repo is public and the roadmap is honest. Watch it, try the adapter, or open an issue.

github.com/wehnsdaefflae/llmsectest »