Expert-Level Document Parsing
Participants convert difficult document images into structured Markdown while preserving:
Musical notation appears in the challenge data as exploratory content. It does not receive a separate EvalAI metric.
Expert-Level Document Parsing Challenge
Recover structured content from the document pages that challenge today’s strongest parsers—across text, tables, formulas, reading order, and specialized notation.
Competition: August 10–October 10, 2026 · Winners announced at DocInsights 2026
The Dr.DocBench Challenge evaluates expert-level document parsing on pages where strong systems still struggle. It is based on Dr.DocBench, a difficulty-aware benchmark developed by 2077AI and research collaborators, and is part of the DocInsights 2026 workshop at EMNLP 2026.
Participants convert difficult document images into structured Markdown while preserving:
Musical notation appears in the challenge data as exploratory content. It does not receive a separate EvalAI metric.
Shared across the Dr. DocBench and DocSEM challenges. Award details and allocation will be announced soon.
Public EvalAI results will be updated periodically during the competition. Official final rankings will be determined through the Private Test phase.
Ranking is descending by Overall under the official EvalAI scoring contract.
| Rank | Team | Text ED | Table TEDS | Formula CDM | Reading Order | Overall |
|---|---|---|---|---|---|---|
| Leaderboard opens with public EvalAI submissions. | ||||||
The public Dr.DocBench dataset and annotations are available on Hugging Face for development and format reference.
Official evaluation inputs, sample submissions, and validation tools are provided through the Dr.DocBench EvalAI challenge.
Submissions are evaluated using four component metrics plus one Overall aggregate. Final rankings are determined by the Overall score.
Normalized Levenshtein distance over scored text categories, from 0–1. Lower is better.
Mean tree-edit-distance similarity for valid tables, shown on a 0–100 scale. Higher is better.
Mean CDM F1 for isolated display formulas, shown on a 0–100 scale. Higher is better.
(1 − normalized reading-order edit distance) × 100. Higher is better.
Available component scores are converted to a 0–100 scale and averaged for each scorable page, then averaged across all scorable pages. Missing components are excluded from the page denominator rather than scored as zero; non-scorable pages are excluded.
Download the relevant evaluation package and generate one prediction for every released page identifier. The challenge uses a single-page prediction unit, unlike the default two-page window setting described in the Dr.DocBench paper.
The evaluator applies strict checks for missing, extra, or duplicate predictions and invalid identifiers or submission structure.
A canonical mds/ Markdown tree mirroring the released hierarchy is also accepted.
Upload the validated ZIP through EvalAI. Final rankings will be determined through the Private Test phase.
Open to researchers and teams worldwide. There is no team-size limit.
Maximum three submissions per day. Each team may select up to two submissions for final evaluation.
Top-ranked teams must submit a technical report. All external data, models, and resources must be disclosed.