AI evaluation methodology

AI earns trust one version at a time.

Mochi does not treat a provider name or rolling model alias as evidence of safety. Eligibility is tied to an exact model version and a frozen evaluation process.

What the benchmark measures

Versioned cases describe an allowed destination set and the expected safe behavior. Reports measure top-1 accuracy, wrong-destination rate, abstention, calibration, latency, privacy exposure, and real request cost. The runner records prompt and schema hashes so results can be traced to the evaluated configuration.

Wrong-destination rate matters separately from overall accuracy because a confident unsafe move is not equivalent to an abstention. Calibration checks whether confidence reflects observed reliability. Latency and cost are measured on the same workload so a cheaper configuration is not accepted by quietly changing the task.

Rules before AI

Deterministic local rules handle files whenever possible. AI sees only ambiguous cases and returns a suggestion from the candidate destinations supplied by the client. It cannot browse the filesystem, create arbitrary paths, or execute a move.

The decision protocol is deliberately narrower than open-ended rule generation. Closed-set classification chooses among known destinations. The client checks identifiers, confidence, and schema validity before presenting a suggestion.

Frozen gates and holdouts

Thresholds are set before the final holdout run. A Jev-first configuration is eligible only when accuracy remains within the documented non-inferiority margin, cost or p95 latency improves materially, fallback remains bounded, and privacy exposure is no worse. Exact values and the benchmark plan are published in the repository.

Freezing the gate before the final run reduces the temptation to tune a passing rule after seeing holdout results. Reports carry the dataset version, resolved model version, prompt hash, schema hash, and timing information needed to compare later runs.

Privacy is part of the evaluation

The benchmark does not treat accuracy as the only outcome. Cases track what fields a configuration needs, how redaction behaves, and whether additional context produces enough value to justify greater exposure. Complete files, absolute filesystem paths, provider credentials, and the operation journal remain outside the provider boundary.

See the privacy policy for user-facing data details and the repository architecture document for protocol boundaries.

Safe rollout

A new model version begins in review-only mode. Rolling aliases are not trusted for automatic behavior. Even after a version passes, the Mac client validates the response, limits it to approved destinations, and performs the filesystem operation itself. Users retain review and undo controls.

A provider changing a model behind an alias does not inherit an earlier result. The new resolved version returns to review-only status until it is evaluated. Filesystem safety and provider-boundary regressions remain release blockers independently of benchmark accuracy.

Read the full benchmark plan and AI architecture.