AIED-Unplugged

Mathematical Exam Diagnostic

Classify a student's handwritten working as correct or by the error it contains.

Question 01 asks how many more 1-real coins Pedro needs, having 20 and needing 45. The student wrote 25, with 45 − 20 as working.
math-50-01Correct answer

The data for this track is a preview sample of the full AIED-Unplugged collection, which is released at a later date. The prediction format below is the basis for the one the full release will use; the release may refine it and carry data this preview does not, as how the preview was annotated sets out.

The dataset splits this track into 471 train, 101 validation and 101 test items. Train and validation carry their targets. Test carries inputs only and decides the ranking.

The test targets are published once the competition closes, and the dataset pages linked here are updated in place. That completes the preview sample. The full collection arrives separately.

Inputs

Two images per item: the printed exam question, and the working a student produced while solving it.

The working is what a student wrote under exam conditions: partial, sometimes restarted, occasionally correct by accident, and sometimes an answer with nothing shown.

Prediction

One label per item, from the thirteen classes listed on this page, naming what happened in the response. Sound work reaching the right answer is Correct answer.

Three classes describe a response with nothing to diagnose. Correct direct answer and Incorrect direct answer cover an answer given with no working shown. Incomplete calculation covers working that starts correctly and stops before it is finished.

Labels are compared exactly as written, so Interpretation error matches and interpretation error does not.

The preview sample contains eleven of the thirteen. Multiplication error and Division error are declared for the full collection and occur in no item here. Both remain permitted values.

Scoring

Macro F1 decides the ranking: the unweighted mean of the per-class F1 scores. The mean is taken over the labels present in the graded reference rather than over all thirteen. A label absent from the graded reference has no per-class score, so predicting one counts only as a miss on the label that was correct.

Overall accuracy is reported alongside it.

Prediction format

Files are checked against this declaration before they are submitted.

ColumnTypeNotes
equation_idstringIdentifier. Must be unique. As the dataset ships it: math-604-01
diagnosticstringOne of the thirteen diagnostic classes, spelled exactly as listed on this page. The preview sample exercises eleven of them; the two it does not contain are still permitted.
CSV, exactly 101 data rows, at most 10 MiB.

A few rows showing the shape of the file. Its identifiers are placeholders, not the dataset’s.

Permitted values

Every value a prediction may carry in the diagnostic column, spelled as a prediction file must spell it. The comparison is exact.

Correct answer
The working is sound and reaches the right answer.
Multiple answers
More than one answer is given, with no indication of which one stands.
Correct direct answer
The right answer, written down with no working shown. Correct, but nothing in the response says how it was reached.
Incorrect direct answer
A wrong answer written down with no working shown. There is no working to diagnose, which is itself what the label records.
Interpretation error
The student solved a different problem from the one the question asks.
Wrong operation
The operation chosen is not the one the question calls for.
Operation setup error
The right operation, assembled wrongly: terms misplaced, digits misaligned, or a value carried into the wrong position.
Partial answer
Part of what the question asks for is answered and the rest is left.
Incomplete calculation
The working starts correctly and stops before reaching an answer.
Addition error
A mistake in adding, in otherwise sound working.
Subtraction error
A mistake in subtracting, in otherwise sound working.
Multiplication error
A mistake in multiplying, in otherwise sound working.
Not in the preview sample.
Division error
A mistake in dividing, in otherwise sound working.
Not in the preview sample.

Baselines

Produced by the organizing team and submitted like any other entry. They are marked on the board.

  • Majority classbaseline-majority

    Predicts the most frequent diagnostic for every item.

  • Zero-shot vision modelbaseline-vlm-zeroshot

    Shows a vision-language model the equation image and the taxonomy, with no fine-tuning.