Mark My Words
Empowering Writers
Rubric markingHow close the model sits to an EW trainer
The model drafts a mark for every descriptor on the Empowering Writers narrative and informative rubrics: Met, Partially met or Not met. Here is how often that draft agrees with 5 Empowering Writers trainers, how often the trainers agree with each other, and how both compare with classroom teachers marking the same rubrics.
Measured against EW trainers
5 Empowering Writers trainers marked the same pieces blind, without seeing any suggested mark.
See the validation resultsSet beside classroom teachers
4 classroom teachers marked 419 further pieces on the same rubrics, so trainer and teacher consistency can be compared.
Trainers and teachersOpen about where it leans
We publish the criteria where the model marks higher or lower than a trainer, so a teacher knows where to look first.
Where it leansHow the model marks a piece of work
Four steps from the page to a class set ready for review.
Transcribing the page
Handwriting becomes text. The words, the sentences and the order of ideas carry through.
Reading the rubric
Narrative writing is read against the narrative rubric, informative writing against the informative rubric.
Marking each descriptor
Each descriptor is marked Met, Partially met or Not met, with the sentences from the piece that support the mark.
Teacher review
You confirm or change each mark against the piece and what this student has been taught.
Validation results
The model agrees with an EW trainer as often as the trainers agree
5 Empowering Writers trainers marked 89 descriptors across 65 pieces of narrative and informative writing, Years 1 to 6, without seeing any suggested mark. The model marked the same descriptors. Separately, 4 classroom teachers marked 591 descriptors across 419 pieces on the same rubrics.
Exact agreement
59.5%
trainer ceiling 58.4%
The model and an EW trainer gave the same mark: Met, Partially met or Not met.
Weighted kappa
0.55
trainer ceiling 0.54
Classroom teachers with each other: 0.40. Kappa removes chance agreement; 0.41–0.60 is moderate.
With the trainer majority
64.1%
within one level 98.4%
On 64 descriptors where at least three trainers marked and a majority agreed.
When the trainers agreed
91.7%
22 of 24 descriptors
Every trainer gave the same mark, and the model gave it too.
The model, the trainers and classroom teachers, side by side
Mint is the model against an Empowering Writers trainer. Navy is two trainers against each other on the same pieces: the trainer ceiling. Amber is two classroom teachers against each other. Weighted kappa removes the agreement chance would produce and counts a Met against Not met split as a bigger miss than Met against Partially met, so it is drawn on the same 0–100 axis to be read alongside the percentages.
The model gives the same mark as an EW trainer 59.5% of the time. Two trainers give the same mark 58.4% of the time, and two classroom teachers 58.1%. Exact agreement is close across all three. The difference is in how far apart the marks are when they differ: weighted kappa is 0.55 for the model and 0.54 between trainers, against 0.40 between classroom teachers.
When markers disagree, how far apart are they?
Share of comparisons where one marker said Met and the other said Not met: opposite ends of the rubric. A disagreement between Met and Partially met is a judgement call. A disagreement between Met and Not met changes what the student is told.
Classroom teachers end up at opposite ends 7.8% of the time, about 2.7 times as often as EW trainers (2.9%). This is what Empowering Writers training buys: trainers still differ on borderline pieces, but almost always by one step. The model, trained against a marking guide calibrated to EW trainers, keeps to the same narrow band (3.9%).
95.0%
When trainers split, the model matched one of them
On the 60 descriptors where the trainers disagreed, the model gave a mark at least one trainer had also given.
96.1%
Within one level of a trainer
The model was at most one step from the trainer's mark. Two trainers manage 97.1%; two classroom teachers 92.2%.
0.55 from 0.38
Up on the previous model
Against the same trainers on the same descriptors, the previous model agreed 53.4% of the time. This model, trained on 5,000 pieces marked by AI against a marking guide calibrated to these trainers, agrees 59.5%.
Does the model mark like an EW trainer?
Each trainer is scored against the other 4 trainers on the same pieces. The model is scored against all 5. If the model marks like a trainer, its bar sits among theirs.
The trainers range from 54.9% to 60.6%. The model, at 59.5%, sits inside that range: it reads as one more member of the panel, not an outlier.
By text type
Weighted kappa (× 100) on the criteria specific to each rubric. Organization and Vocabulary, fluency and mechanics appear on both rubrics and are shown with the criteria below.
On narrative writing the model reaches 0.58 against 0.54 between trainers and 0.44 between classroom teachers. On informative writing it is 0.52, 0.51 and 0.35. On both, the model tracks the trainers and both sit above classroom teachers.
Agreement by criterion
Exact agreement on each criterion. Hover for the number of comparisons and the weighted kappa. Each criterion rests on 4 to 13 pieces marked by the trainers, so a few points either way is within the noise; the pattern across criteria is the stronger signal.
The model is level with or above the trainer ceiling on most narrative criteria and on supporting details. Organization is the clear exception: the model agrees with a trainer less often there than trainers agree with each other, and it tends to mark Organization lower. Main ideas and the informative conclusion also sit a few points below the trainer ceiling, on small samples.
Where the model marks higher or lower than a trainer
Average difference between the model's mark and an EW trainer's, in rubric levels. Right of zero, the model marks the criterion higher. Left of zero, it marks it lower.
Organization, suspense, entertaining beginnings and vocabulary are the criteria to read first: the model marks them a little lower than the trainers do. Extended endings and elaborative detail go the other way. None of these averages is as large as half a level.
How to read these numbers
- A comparison is one pair of markers on one descriptor of one piece. A descriptor marked by five trainers gives five model comparisons and ten trainer pairs.
- The trainers and the classroom teachers marked different pieces on the same rubrics and year range, so the two groups are compared as groups, not piece by piece.
- The trainers marked blind. The classroom teachers reviewed with a suggested mark from the previous model on screen and kept it 72.4% of the time, against 53.9% for trainers marking blind. A shared suggestion pulls markers towards each other, so the classroom teachers' agreement with each other is, if anything, generous to them.
- The model figures come from the same training recipe as the model in use, scored at the end of training. On a larger held-out set the model in use is within one point of it.