Mark My Words
SRSD Online
TIDE and TREE markingHow close the model sits to a reference marker
On informative writing the model drafts each element of TIDE. On persuasive writing it drafts each element of TREE. Each element is marked present or absent. Here is how that draft is made, how it compares with six reference markers on 50 held-out pieces, and where a teacher's read changes the mark.
A first pass the teacher confirms
The model fills in each element so the rubric opens with a draft. Confirm, change, or clear any element with a click.
Checked on held-out writing
50 held-out pieces, each element marked by 6 reference markers.
See the validation resultsOpen about where it leans
We publish the elements it credits more often than the markers, and the one element it withholds.
How the model marks a piece of work
Four steps from the page to a class set ready for review.
Transcribing the page
Handwriting becomes text. The words, the sentences and the order of ideas carry through.
Reading the strategy
Informative writing is read against TIDE. Persuasive writing is read against TREE.
Marking each element
Topic, ideas or reasons, explanations, ending and linking words are each marked present or absent.
Teacher review
You confirm the draft against the piece, the plan and what this student has been taught.
Validation results
The model agrees with a marker as often as the markers agree
Fifty held-out pieces, half informative and half persuasive, from Years 1 to 7. Six reference markers and the model each marked every included element present or absent. An element left blank by the panel is an element the task did not ask for, and it is left out of every figure below.
Exact agreement
84.7%
marker ceiling 82.5%
The model and a reference marker made the same present or absent call.
Cohen’s kappa
0.594
marker ceiling 0.575
Agreement after removing what chance would produce. 0.41–0.60 is moderate agreement.
With the majority
93.0%
kappa 0.794
On 284 elements where at least three markers scored and were not evenly split, the model matched the majority. 0.61–0.80 is substantial agreement.
When the panel agreed
97.4%
188 of 193 elements
Every marker who scored the element gave the same verdict, and the model gave it too.
Model vs a marker, next to marker vs marker
The navy bar is the human ceiling: how often two reference markers gave the same present or absent verdict on the same element of the same piece. Kappa is drawn on the same 0–100 axis so the two measures can be read together. Kappa removes the agreement you would get by chance, which matters here because most elements are present.
The model agrees with a reference marker 84.7% of the time. Two markers agree with each other 82.5% of the time. Kappa is 0.594 for the model and 0.575 between markers. Informative writing (TIDE) is 84.0% against a ceiling of 82.0%. Persuasive writing (TREE) is 85.4% against a ceiling of 83.0%.
97.4%
When the markers agreed, the model agreed
Across the 193 elements where every marker who scored it reached the same verdict, the model matched that verdict. Clear cases are handled closely.
100%
When markers split, the model matched one of them
On the 118 elements where the panel disagreed, the model landed on a verdict a reference marker had also chosen.
62.3%
Agreement chance alone would produce
Most elements are present, so always answering present would land near this. Cohen's kappa removes it. The model's 0.594 sits beside 0.575 between the markers.
Which way the disagreements go
Every comparison, split by whether the model and the marker chose the same verdict. “Credited” means the model marked the element present and the marker marked it absent. “Missed” means the reverse.
Set against one marker, the model credits an element the marker withheld 11.0% of the time and misses an element the marker found 4.3% of the time. Set against the majority of the panel, those two misses are almost even: 11 extra credits and 9 misses across 284 elements.
Agreement by element
Each bar pair is one element of TIDE or TREE. Mint is the model against a reference marker. Navy is how often two markers agreed on that same element. Who is omitted: it was scored on a single informative piece, where everyone marked it absent.
On almost every element, the model agrees with a marker at least as often as the markers agree with each other. The persuasive topic sentence is the exception: markers agree with each other more often than the model agrees with them. Important ideas, reasons, Do and What are present on nearly every piece, so a high percentage there mostly means both sides found an element that is usually there.
Which elements the model credits more often
Mean difference between the model and a reference marker. Right of zero, the model marks the element present more often than the marker. Left of zero, it withholds the element more often.
Detailed explanations and persuasive explanations are the elements to read first: the model credits them more often than the markers do. Persuasive topic sentences go the other way, and that is the element the model most often withholds.