AIED-Unplugged

Answer Sheet Detection

Read a whole scanned answer sheet and report, for every question on it, what the student marked.

A photographed answer sheet with 21 questions, one option filled in for each.
sheet-1_15068340_2878_76510_CAD01LP521 questions, mostly D (9 of 21)

The data for this track is a preview sample of the full AIED-Unplugged collection, which is released at a later date. The prediction format below is the basis for the one the full release will use; the release may refine it and carry data this preview does not, as how the preview was annotated sets out.

The dataset splits this track into 117 train, 27 validation and 26 test items. Train and validation carry their targets. Test carries inputs only and decides the ranking.

The test targets are published once the competition closes, and the dataset pages linked here are updated in place. That completes the preview sample. The full collection arrives separately.

Inputs

A scan of a complete cartão-resposta, one per row, as collected from schools. Sheets are photographed rather than fed through a scanner, so they carry skew, shadows, folds, and the occasional thumb. A few were photographed a quarter turn round, with the printed header running up the side.

Sheets carry between 16 and 26 questions. The count for each sheet is published alongside it, including for the test split.

Prediction

For every question on the sheet, the value the student marked. The target records the mark itself, not whether it matches the gabarito. A sheet where the student marked D and the key says B is recorded as D.

One row per sheet, with all the questions in a single column as a JSON array:

[
  { "question_number": 1, "label": "A" },
  { "question_number": 2, "label": "Blank" }
]

Every question appears exactly once, and question_number is a number.

A, B, C, and D each mark a single clear choice. Blank means the question was left unmarked, Multiple means more than one bubble was filled in, and Other covers any mark that does not read as one of the options, such as a heavy erasure or a tick in the margin.

Scoring

Cell accuracy decides the ranking: every question on every sheet counts once, pooled across all sheets. A 26-question sheet therefore contributes more than a 16-question one, and a question omitted from a prediction counts as wrong rather than being skipped.

Sheet exact match and macro F1 over the seven values are reported alongside it. Sheet exact match requires every question on a sheet to match.

Prediction format

Files are checked against this declaration before they are submitted.

ColumnTypeNotes
sheet_idstringIdentifier. Must be unique. As the dataset ships it: sheet-1_26069580_2858_76082_CAD01LP2
answersstringA JSON array with one entry per question on the sheet, each {"question_number": n, "label": "..."}. Every question appears exactly once, numbered from 1. A sheet carries between 16 and 26 questions, so the array's length varies row to row.
CSV, exactly 26 data rows, at most 10 MiB.

A few rows showing the shape of the file. Its identifiers are placeholders, not the dataset’s.

Permitted values

Every value a prediction may carry in the answers column, spelled as a prediction file must spell it. The comparison is exact.

A
A single clear mark on option A.
B
A single clear mark on option B.
C
A single clear mark on option C.
D
A single clear mark on option D.
Blank
No mark on any option.
Multiple
More than one option marked.
Other
A mark that is present but not readable as any of the above, such as an erasure or a stray annotation.

Baselines

Produced by the organizing team and submitted like any other entry. They are marked on the board.

  • Pixel thresholdbaseline-threshold

    Classical computer vision: deskew, locate the bubble grid, threshold each bubble's fill.

  • Zero-shot vision modelbaseline-vlm-zeroshot

    Asks a vision-language model to read the sheet and return the marks as JSON.