A roadmap gated on data and evidence, not a shipped guarantee.
As escalated, labelled cases accumulate, a smaller open-weight model is periodically retrained on that real data, aiming for a faster, self-contained classification path. Architecture changes are never scheduled by calendar dates; progress is mediated entirely by empirical validation gates.
Data-volume clearing threshold
Constraint: Accumulate sufficient production-calibrated samples
Training begins only once a defined data-volume threshold is cleared: roughly a thousand labelled examples (an illustrative example threshold, not a fixed public commitment). No training pass runs on sparse or unverified logs.
Held-out benchmark outperformance
Constraint: Strict superiority against current production pipeline
Any fine-tuned model only ships once it demonstrably outperforms the current production path on held-out test data, never before. Parity is insufficient for deployment.
Every prediction is compared directly against hand-labelled ground truth and reported as an accuracy figure alongside per-outcome precision and recall.
Every single candidate prediction is cross-checked directly against hand-labelled ground truth.
Aggregated classification score across full representative validation splits.
false-positive prevention computed separately across each granular mark classification code.
false-negative containment ensuring subtle student execution slips are not bypassed.
The same held-out test set is scored separately under each pipeline configuration (rule-only, escalation-only, and the combined hybrid path), so any gain from escalation or from a fine-tuned model is measured against the current production path rather than assumed.
Early internal testing on a small validation set showed encouraging results, presented as illustrative preliminary findings, not a verified benchmark.