Mark My Words
Big Write & VCOP
Criterion Scale MarkingWhat AI can see, and what only teachers can
When assessing Cold Write and Big Write samples, our model estimates where each student sits on the Criterion Scale and drafts marks across nearby descriptors. Here is a clear look at how it reads student writing, how its marks behave against real teacher judgements, and why teacher moderation remains essential.
A starting draft, not a final score
The AI fills in baseline marks so you never start from a blank rubric. Confirm, adjust, or reject any mark with a click.
Calibrated by expert markers
A share of samples is independently double-marked by experienced teachers to keep the model aligned with expert Criterion Scale judgement.
See the validation resultsOpen about its limits
We publish how the model performs against real teacher marks, including the strands it gets wrong most often.
How the model marks a piece of work
Four simple steps from handwritten page to a ready-to-review class set.
Transcribing the page
Handwriting converts to digital text. Word choices, punctuation, and sentence structures carry through cleanly.
Finding the working level
The model estimates the student’s Criterion Scale band based on overall vocabulary and sentence control.
Drafting the marks
Descriptors at and around that working level receive suggested marks (tick, dot, or cross).
Teacher review
You review suggested marks in the rubric view, adjust based on student context, and save.
Validation results
The model marks as consistently as a second teacher
Every AI verdict is compared against a teacher verdict on the same descriptor for the same piece of writing. Where two teachers marked the same descriptor independently, we can also measure the AI against the disagreement that already exists between people.
What these figures cover. The model marks from a transcription of the writing, so handwriting, letter formation and visual layout are not part of any result below. It returns a cross for those descriptors because it cannot see them — teachers change these from the original page where the evidence is there.
Exact agreement
70.8%
human ceiling 71.4%
AI and teacher chose the same verdict of Not met, Partially met or Met.
Within one band
91.9%
human ceiling 93.9%
Never more than one band apart — the threshold at which a teacher would call a mark defensible.
Weighted kappa
0.628
human ceiling 0.654
Agreement after removing what chance alone would produce. 0.61–0.80 is substantial agreement.
Mean absolute error
0.37 bands
bias -0.11
Average distance from the teacher. The negative bias means the model leans harsh.
AI vs teacher, next to teacher vs teacher
The bar on the right is the human ceiling: how often two teachers marking the same descriptor on the same essay reached the same verdict. Weighted kappa is shown on the same 0–100 axis for comparison.
The AI sits inside the human disagreement band on every measure. Teachers agreed with each other 71.4% of the time; the AI agreed with a teacher 70.8% of the time — a gap of 0.6 percentage points. The model is statistically indistinguishable from a second marker.
82.9%
When both teachers agreed, the AI agreed too
Across the 35 descriptors where two teachers reached the same verdict, the AI matched that consensus. Clear-cut cases are handled well.
78.6%
When teachers disagreed, the AI matched one of them
On the 14 genuinely ambiguous descriptors, the AI landed on a verdict a real teacher had also chosen. It rarely invents a third opinion.
0.556
Unweighted Cohen's kappa
Because the three verdicts are unevenly distributed, guessing the most common answer would score 34% by chance alone. Kappa strips that out.
Direction and size of the errors
Every judgement placed by how far the AI sat from the teacher. Negative means the AI marked the student down relative to the teacher.
Errors lean one way. The AI is harsher than the teacher 18.2% of the time and softer only 11.0% of the time, an overall bias of -0.11 bands, where teachers marking against each other showed almost no tilt (+0.02). When the model errs, it errs on the side of marking the student down — the safer direction for a teacher reviewing its suggestions.
Agreement by skill strand
Descriptors grouped by the skill they assess, sorted worst first. Strands below the human ceiling are where a draft mark deserves the closest review.
Which strands the AI marks down
Mean difference between the AI verdict and the teacher verdict, in bands. Left of zero the AI is harsher than the teacher; right of zero it is more generous.
Spelling and voice pull hardest to the harsh side — partly because the model cannot know your school's spelling expectations — while the AI is slightly generous on sentence structure and punctuation. During review, harsh spelling and voice marks are the first candidates for an upgrade, and sentence-structure and punctuation ticks are worth a quick confirm.
Reading the draft marks
Six things worth knowing before you review
The first two explain what a draft mark is and what it cannot cover. The rest come from comparing the model's marks against real teacher marks on the same pieces of writing.
It estimates rather than counts
Most descriptors ask how often something appears — mostly, about half the time, at least twice. The model answers by reading the whole piece and forming an overall impression, the way you might after a first read-through. It does not find and count each example, so when an exact number decides the mark, check it on the page.
A cross on anything visual is a blind spot, not a verdict
The model reads a transcription, so handwriting, letter formation and visual layout never reach it. It returns a cross for those descriptors because it cannot see the evidence — not because the evidence is missing. Mark these from the original page, and note they sit outside every figure below.
Start with the dots
It marks harsher than a teacher 18.2% of the time and softer 11.0%, so more of the marks it gets wrong are ones it has held back than ones it has awarded.
Most gaps are one band, not two
91.9% of its marks land within a single band of a teacher’s, so a disagreement is usually a choice between neighbouring verdicts.
Give spelling and voice the closest read
These pull hardest to the harsh side — partly because the model cannot know your school’s word lists or year-level expectations.
Treat a disagreement as a moderation point
On descriptors where two teachers disagreed, the model landed on a verdict one of them had chosen 78.6% of the time. A clash usually means the descriptor is genuinely debatable.
Strand by strand analysis
To make your review fast and focused, here is what the model sees on the page and where teacher input is needed:
Textual evidence
Features that carry across clearly when handwriting is converted into digital text.
VCOP
- Vocabulary
- Spotting descriptive words, topic vocabulary, figurative language, and ambitious word choices.
- Connectives
- Tracking conjunctions, time transitions, and causal linkers that join clauses and connect ideas.
- Openers
- Looking at how sentences start across the piece, including fronted adverbials and varied sentence starters.
- Punctuation
- Checking sentence boundaries and marks across the punctuation pyramid, from capital letters and full stops up to colons, dashes, and semicolons.
Wider text features
- Spelling
- Recognising phonological attempts, high-frequency sight words, prefixes, suffixes, and level-appropriate vocabulary.
- Grammar and sentence control
- Checking subject-verb agreement, tense consistency, and control across simple, compound, and complex sentences.
- Text structure
- Looking at overall text organisation, logical sequencing of ideas, and paragraph flow.
Visual & contextual limits
Elements that need a look at the original scan or knowledge of classroom context.
Visual & layout constraints
- Handwriting and letter formation
- Letter sizing, neatness, and cursive joins require viewing the original handwritten page.
- Visual paragraph spacing
- Indents and line breaks can compress in digital text, though topic shifts remain identifiable.
Context outside the written piece
- Your school's spelling expectations
- The model recognises spelling patterns and plausible attempts, but it does not know your school’s high-frequency word lists or year-level spelling expectations.
- Editing and up-levelling
- A final piece does not show live classroom revisions, drafts, or verbal self-corrections.
- Emergent early-years descriptors
- Early scale indicators often describe pencil grip, oral storytelling, and classroom behaviours.
Qualitative criteria
Areas where textual evidence is visible, but deciding the mark requires teacher discernment and professional judgement.
- Audience, tone, and authorial voice
- Deciding whether writing effectively entertains, persuades, or builds suspense requires teacher judgement.
- Overall level placement
- The model suggests a working level as a starting point; confirming final placement always rests with teachers.
Teacher-led assessment
Faster to the judgement, not instead of it
Mark My Words handles the time-consuming groundwork: transcribing handwriting and drafting initial Criterion Scale descriptors across the class.
Instead of starting from a blank rubric for every student, you start with a structured draft — giving you more time for moderation, student feedback, and planning next steps.
Final assessment decisions always stay with the teacher.