Public AI Governance Documentation

Evaluation and Monitoring

A practical AI evaluation, release, incident, and quarterly review cadence.

This document is a public governance summary. It does not publish raw AI instructions, exact internal eval cases, source code, tenant data, incident details, database schema, or provider-console settings.

Last reviewed: July 9, 2026

This approach is designed for a solo-founder SaaS. It favors repeatable, evidence-producing checks over heavyweight compliance process.

Evaluation Goals

HireProxy should be evaluated for:

Current Monitoring Signals

Lightweight Grounding Eval Set

Maintain an internal grounding and confabulation evaluation set. The public documentation describes the evaluation categories, while the exact examples, expected answers, and pass/fail notes stay internal.

The eval set should cover:

For each internal example, record the input, allowed facts, facts that must not be stated, expected uncertainty behavior when relevant, model/version context, result, and reviewer notes.

Run the eval set before:

Release Checklist for AI Changes

For each AI-impacting release, answer:

Quarterly Review Checklist

Every quarter:

HireProxy uses a scheduled operator review reminder for this cadence. The operator receives an email when the review opens, and the operator dashboard shows the current review checklist and proof note until the review is marked complete.

Quality Metrics

Use simple metrics that are realistic for a small company:

Escalation Thresholds

Investigate immediately:

Accepted Residual Risks

The following residual risks are accepted with disclosure and monitoring: