Our engineers set up and run your first chatbot / LLM security scan. Get in touch

MS-2.3: AI system performance is evaluated regularly

Continuous evaluation against benchmarks + drift detection.

Last reviewed July 2026

The gap MS-2.3 closes

In NIST AI 600-1, AI system performance is evaluated regularly addresses measure. Continuous evaluation against benchmarks + drift detection. Penaxtra records this control at high severity and establishes its state by exercising it against the running system, so the result reflects observed behavior rather than a documented assertion.

How Penaxtra delivers MS-2.3

Penaxtra turns this NIST AI 600-1 obligation into recurring, testable evidence: scheduled scans and posture checks produce findings tied to MS-2.3, and the append-only audit log records what was tested and when. The NIST AI 600-1 MS-2.3 identifier is attached when the finding is created, so it appears in the exported evidence pack already mapped to the control. Where the same weakness maps to another framework, the finding carries those control identifiers as well.

MS-2.3 capabilities

Probe and check coverage aligned to MS-2

3 (AI system performance is evaluated regularly).

Findings tagged with the NIST AI 600-1 MS-2

3 identifier.

Penaxtra severity for this control (high)

Cross-framework identifiers attached to the same finding where controls overlap

PDF and JSON evidence export with the control identifier attached

MS-2.3 compliance mapping

Findings for MS-2.3 carry the NIST AI 600-1 MS-2.3 identifier along with the corresponding control identifiers in the other frameworks Penaxtra maps, so one result is reflected across each mapped framework.

Frequently asked

What is MS-2.3 (AI system performance is evaluated regularly)?

Continuous evaluation against benchmarks + drift detection. It is a NIST AI 600-1 control; Penaxtra assesses it at high severity.

How does Penaxtra test for MS-2.3?

Penaxtra turns this NIST AI 600-1 obligation into recurring, testable evidence: scheduled scans and posture checks produce findings tied to MS-2.3, and the append-only audit log records what was tested and when.

Does a finding for MS-2.3 help with an audit?

Each finding is tagged with the NIST AI 600-1 MS-2.3 identifier and exported in the PDF and JSON evidence pack, so it appears on the auditor control list with the identifier already attached.