Benchmark methodology

Proof needs a method, not just a number.

Pentoma's benchmark program is designed to show what each assessment product can validate, how evidence is judged, and where claims are still limited before broader production data exists.

Public benchmark suite

Read the methodology. Then read the code.

Pentoma's benchmark suite is published as an open repository so security teams, researchers, and auditors can review the cases, the pass/fail rules, and the validation criteria we measure ourselves against.

Launch benchmark

Each product gets a separate target corpus.

Pentoma Web, Pentoma Code, and Pentoma AI Red Teaming should not share one vague benchmark. Each product needs targets that match the work customers are buying.

Pentoma Web

Corpus
OWASP Juice Shop, intentionally vulnerable API targets, PortSwigger-style labs, and permitted public validation cases.
Methodology
Authenticated and unauthenticated runs are scored on confirmed exploitability, replay quality, time-to-finding, false-positive rate, and severity accuracy.
Pass criteria
A finding passes only when request/response evidence, safe proof-of-exploit, reproduction steps, and non-destructive validation are present.

Pentoma Code

Corpus
Seeded vulnerable repositories, pull-request fixtures, SAST benchmark cases, and real framework patterns across TypeScript, Python, and API services.
Methodology
Reviews are scored on reachable path explanation, line-level precision, fix quality, duplicate suppression, and developer review usefulness.
Pass criteria
A finding passes only when the vulnerable line, source-to-sink path, rule or LLM rationale, and remediation guidance are reviewable.

Pentoma AI Red Teaming

Corpus
Prompt-injection labs, RAG poisoning scenarios, tool-use sandboxes, agent permission fixtures, and OWASP LLM Top 10 test cases.
Methodology
Runs are scored on replayability, observed behavior, policy/control mapping, severity rationale, and safe reproduction of model or agent failure.
Pass criteria
A finding passes only when prompts, responses, target state, expected control, and deterministic checks where possible are captured.

Measurement principles

The benchmark should make Pentoma more credible, not louder.

The goal is to earn trust before official launch by showing the methodology, the limits, and the evidence standard behind every claim.

Publish the target class, scope, and pass/fail definition before publishing performance claims.
Measure confirmed findings, not raw alert count.
Track false positives, duplicate suppression, and replay completeness alongside coverage.
Separate lab performance from production claims until enough real engagements are validated.

Cadence

New cases ship as the suite matures.

The suite starts narrow and grows on a deliberate cadence. Public claims about Pentoma performance are tied to cases the public can review, not to internal numbers without a methodology readers can audit.