Collaborative Dictionary Digitization


In 2027, SIGTYP is hosting a Shared Task on Collaborative Dictionary Digitization.
TBA: AIMS & GOALS (AIM 1: produce high-quality digitized dictionaries that require minimum post-editing effort to be turned into gold standard ones; AIM 2: Perform evaluation of systems on both tasks in terms of their accuracy and the total cost.) TBA: Motivation & Description

Subtasks

Following the established evaluation pipeline (Setiawan et al, 2026; MUDIDI), we will focus on two major subtasks:

Subtask 1: Faithful transcription (OCR)

This stage focuses on vanilla OCR: transforming an input PDF file into a txt one, preserving the layout. Details: TBA (Details on input & output formats; )

Evaluation: WER, CER, Read Order Edit, Markup F1
Baselines: Mathpix, PaddleOCR

Subtask 2: Dictionary Parsing

This stage focuses on dictionary parsing: transforming an input PDF or transcribed text file into a structured format (e.g, MDF). Details: TBA (Details on input & output formats; )

Evaluation: Entry F1 (Entry Detection Accuracy), Field Assignment F1 (Assigning tags/labels), ReadOrderEdit
Baselines: a transformer-based seq2seq model?

On top of that, all systems will be evaluated in terms of their estimated cost and token usage (for API- and subscription-based approaches).

You may submit a system to any of them or both!

Datasets and the Task Setup

The task is organized in two phases: (1) Data collection; (2) System Evaluation. At the start of the data collection phase, we invite EOI from linguists and communities to register their dictionaries for digitization. They will also be asked to provide a gold standard transcripts for a 3-page sample of their dictionaries (the organizerscan provide them with silver-standard transcripts to ease the process). These annotated samples will be used for evaluation of the systems. During the data collection phase, shared task participants will have access to a sample of 27 dictionaries with gold-standard annotations collected ealier as a part of the MUDIDI dataset. They can use the these dictionaries to fine-tune their systems.
Once the data collection phase finishes, systems will be evaluated on the new dictionary samples collected from data contributors. For each dictionary registered for the shared task, the organsers will then list each system's performance and provide its predictions as well as cost estimates. The data contributors will then vote for the best system and choose the one to use for the full dictionary digitization. In addition to this, the shared task allows participants to use external resources as long as they are openly available and can be, in theory, used by other participants for research purposes. In case participants decide to use external resources and data in their system they should use proper citation and provide the details of the dataset in the system description.

Participants will be invited to describe their system in a paper for the SIGTYP workshop proceedings. The task organizers will write an overview paper that describes the task and summarizes the different approaches taken, and analyzes their results.

Important Links

  ↣  Register as a Data Contributor!  

  ↣  Register as a Shared Task Participant!  

  ↣  Shared Task GIT Page (provides the data, baseline models, and evaluation scripts)  

Important Dates

  Data contributors sign up: ↣  TBA
  Data preparation: ↣  TBA
  Training data release: ↣  TBA
  System evaluation phase: ↣  TBA
  Data contributors voting: ↣  TBA
  System descriptions are due: ↣  TBA
  Camera-ready due: ↣  TBA

Task Organizers

Doreen Osmelak, PhD Student, the University of Melbourne
David Setiawan, Research Assistant, the University of Melbourne
Temuulen Khishigsuren, PhD Student, the University of Melbourne
Charles Kemp, Professor, the University of Melbourne
Andreas Shcherbakov, Research Scientist
Terry Regier, Professor, University of California, Berkeley
Nick Thieberger, Assoc Professor, the University of Melbourne
Ryan Cotterell, Asst Professor, ETH Zurich
Ekaterina Vylomova, Senior Lecturer, the University of Melbourne

Contact

    Please contact TBA if you have any questions