Task 1: PolOCRBench: Polish Document Understanding
PolOCRBench focuses on evaluating systems that transform images of Polish documents into structured textual representations. The input consists of page images in PNG or JPEG format, while the expected output depends on the selected subtask: (A) full-page transcription into Markdown preserving reading order, headings, lists and paragraphs; (B) table extraction into HTML preserving table structure and cell contents; or (C) key information extraction into JSON according to a predefined schema for a given document type. The benchmark includes challenging contemporary and historical printed documents, handwritten content, mobile-phone photographs, complex tables, multilingual text and mathematical expressions.
Task 2: Polish Language Document Layout Detection
The aim of the task is to advance the research concerning document layout detection in Polish. Document layout recognition involves identifying such elements as the title, section header, main text, table, footnote, image, list, etc. The input to the system consists of page images in graphic file format, while the output is the set of identified areas representing individual structural elements (rectangle coordinates) and the corresponding element category label.
Task 3: Adversarial Style Modification: Probing the Human-AI Perception Gap
This task explores the differences between human and machine perception of writing style through adversarial text modification. Given texts with a distinct stylistic signature, participants are asked to transform them so as to reduce the accuracy of a provided transformer-based style classifier while preserving the original meaning, naturalness and recognizability of the style to human readers. Submissions will be assessed using automatic semantic-equivalence checks, the baseline classifier and human peer annotation measuring both text quality and style recognition.
Task 4: Reverse PolEval: Diagnostic Benchmark Design for Polish LLMs
In this task, participants take on the role of benchmark designers rather than solving a predefined evaluation set. They create original Polish-language diagnostic items consisting of an input and a unique short gold answer, with a focus on linguistic phenomena such as inflection, agreement, case, gender, reference, negation and lexical ambiguity. The submitted benchmarks should reveal meaningful behavioural differences between models rather than merely contain difficult examples. Organizers will run the items on a fixed panel of open-weight language models and evaluate them according to their discriminatory power, behavioural diversity and ability to separate individual model pairs.