By Interestana AI Editorial — AI-drafted, human-overseen. How we report
Datalab Releases OmniExtractBench for Document Extraction
Datalab released OmniExtractBench on March 18, 2024, an open benchmark designed to evaluate the accuracy of systems in extracting structured data from PDF documents into a JSON schema. This new benchmark aggregates 620 documents sourced from four pre-existing benchmarks, aiming to provide a unified and auditable standard for the field. The release comes at a time when various extraction vendors are publishing their own leaderboards, which Datalab argues are often difficult to compare and scrutinize due to inherent biases and lack of transparency.
OmniExtractBench addresses four primary flaws identified in current extraction benchmarks: bias, opaque evaluation harnesses, unclear scoring mechanisms, and limited document variety. Bias can occur when a benchmark's design or scoring favors the vendor that created it. Opaque harnesses mean that a low model score might be attributable to a faulty evaluation tool rather than the model's performance. Unclear scoring prevents users from understanding why a specific document received a low score. Finally, narrow document variety means some suites focus only on dense tables, while others exclusively use clean, scanned files, failing to represent real-world document complexity.
The benchmark's scoring system is deterministic, meaning a single scorer evaluates all documents and provides explanations for each decision, enhancing transparency and auditability. The scorer is available as an open-source Python package, `omni-extract-bench` (version 0.1.7, compatible with Python 3.11+ and requiring SciPy), distributed under the Apache 2.0 license. To run vendor models, users will need to provide their own API keys and incur associated costs.
The 620 documents within OmniExtractBench are drawn from diverse sources. The largest contribution, 329 documents, comes from Datalab's internal `ExtractBench`, featuring forms, filings, and decks with dense scalar schemas and small document sizes. Datalab also contributed a synthetic suite of 202 documents designed with dense scalar schemas and small document sizes. `LongExtractBench`, sourced from micro1 (commissioned by Reducto), includes 47 documents with very large tables. `LongArray-Extract` provides 42 documents featuring large tables with repeated scalars, often found in regulatory filing forms. Regulatory filing forms constitute the largest single document type category within the benchmark, with 88 instances. The dataset exhibits a range in document length, with 128 documents being a single page, while 33 documents exceed 100 pages and collectively account for 40% of all pages in the benchmark. The code for OmniExtractBench is hosted on GitHub, and the data is available on Hugging Face under a CC BY 4.0 license.
Original source — read the full reporting at the publisher:
Read on MarkTechPostGet the weekly AI digest
AI news + new model releases, weekly. Drafted by our agents, reviewed by humans.