Technical Documentation // Architecture

A roadmap gated on data and evidence, not a shipped guarantee.

As escalated, labelled cases accumulate, a smaller open-weight model is periodically retrained on that real data, aiming for a faster, self-contained classification path. Architecture changes are never scheduled by calendar dates; progress is mediated entirely by empirical validation gates.

Gating Rules & Deployment Prerequisites
RIGOROUS GATING PROTOCOL
01 // INDUCTION GATE

Data-volume clearing threshold

Constraint: Accumulate sufficient production-calibrated samples

Training begins only once a defined data-volume threshold is cleared: roughly a thousand labelled examples (an illustrative example threshold, not a fixed public commitment). No training pass runs on sparse or unverified logs.

02 // DEPLOYMENT GATE

Held-out benchmark outperformance

Constraint: Strict superiority against current production pipeline

Any fine-tuned model only ships once it demonstrably outperforms the current production path on held-out test data, never before. Parity is insufficient for deployment.

Empirical Evaluation Approach
GROUND-TRUTH COMPARISON

Every prediction is compared directly against hand-labelled ground truth and reported as an accuracy figure alongside per-outcome precision and recall.

Ground truth pairing

Every single candidate prediction is cross-checked directly against hand-labelled ground truth.

Primary accuracy figure

Aggregated classification score across full representative validation splits.

Per-outcome precision

false-positive prevention computed separately across each granular mark classification code.

Per-outcome recall

false-negative containment ensuring subtle student execution slips are not bypassed.

Per-configuration benchmarking

The same held-out test set is scored separately under each pipeline configuration (rule-only, escalation-only, and the combined hybrid path), so any gain from escalation or from a fine-tuned model is measured against the current production path rather than assumed.

Early internal testing on a small validation set showed encouraging results, presented as illustrative preliminary findings, not a verified benchmark.