GROBID alternatives for PDF metadata and reference extraction
GROBID is the default choice for turning scholarly PDFs into structured bibliographic data, but it is not the only one, and it is not always the most convenient. This page groups the main open-source alternatives by what you are actually trying to do. Projects change quickly: check each project's repository for current status and licence before you commit to it.
Quick comparison
| Tool | Best at | Output | Runs as |
|---|---|---|---|
| GROBID | Header metadata, reference parsing, full-text structure of scholarly articles | TEI XML (BibTeX for some services) | Java service with REST API, Docker |
| CERMINE | Metadata and references from journal articles | NLM JATS XML | Java library, command line, web service |
| AnyStyle | Parsing raw reference strings; finding references in text | BibTeX, CSL-JSON, JSON, XML | Ruby gem and command line |
| Nougat | Converting academic PDFs (including formulas) to markup with a vision model | Markdown-like text with LaTeX maths | Python, GPU recommended |
| Docling | General document conversion with layout and table structure | Markdown, HTML, JSON | Python library and command line |
| Marker | Fast PDF to Markdown conversion for many document types | Markdown, JSON | Python, CPU or GPU |
| MinerU | PDF to Markdown/JSON with layout, tables and formulas | Markdown, JSON | Python, GPU recommended |
| Reference managers (Zotero, JabRef, Mendeley) | Getting a clean record for one paper | Library entries, any export format | Desktop apps |
| GROBID Tools extractor | A quick, private look at one PDF without installing anything | JSON, BibTeX, RIS, CSL-JSON, Markdown, DOCX | Your browser (heuristic, not GROBID) |
If you need bibliographic metadata and references
GROBID remains the strongest general-purpose open-source option for this job. Its cascade of specialised models handles headers, author names, affiliations and references separately, and optional consolidation against Crossref adds DOIs. See how to run it with Docker.
CERMINE (Content ExtRactor and MINEr), from the University of Warsaw, covers similar ground: it extracts metadata, structured references and content from born-digital journal articles and outputs JATS XML. It is a reasonable second opinion, but development has been much less active than GROBID's in recent years.
AnyStyle focuses on references. Give it raw reference strings (or a plain-text document) and it returns parsed fields as BibTeX or CSL-JSON. It is easy to retrain on your own examples, which makes it a good fit for unusual citation styles. It does not do PDF layout analysis on its own, so pair it with a text extractor.
Older research tools such as ParsCit and AllenAI's Science Parse are still cited in papers but are no longer actively maintained; prefer the options above for new work.
If you need the full text as Markdown or JSON
Many people searching for "GROBID JSON output" really want readable text with structure, for example to feed a search index or a language model. Newer document-conversion tools target exactly that:
- Docling (open-sourced by IBM Research) converts PDFs and other office formats into Markdown, HTML or a lossless JSON document model, with layout analysis and table structure recognition.
- Marker converts PDFs to Markdown and JSON quickly, with optional model-assisted cleanup. Check its licence terms for commercial use.
- Nougat (from Meta AI) is a vision transformer trained on academic papers. It reads page images directly, so it copes with formulas and some scans, but it is slow without a GPU and can hallucinate on unusual pages.
- MinerU combines layout detection, formula and table recognition into Markdown or JSON output.
These tools are not bibliographic parsers: they will give you the reference section as text, not as structured citations. A common pattern is to use one of them for the body text and GROBID (or AnyStyle) for metadata and references. If you already have GROBID TEI and just want JSON, converting the TEI with an XML library is usually simpler than switching tools.
If you only need a clean record for a handful of papers
You may not need a parser at all. Reference managers identify a paper from its DOI, arXiv ID or ISBN and fetch the authoritative record: Zotero's "Retrieve Metadata for PDF" and "Add Item by Identifier", Mendeley's automatic metadata lookup, and JabRef's PDF import (which can use a GROBID service). Crossref's search and its Simple Text Query tool match free-text references to DOIs. See how to find the DOI of a PDF and from extracted records to a reference list.
Choosing
- Thousands of scholarly PDFs, need structured metadata and references: GROBID.
- Messy reference strings in an unusual style: AnyStyle, trained on a few examples.
- Readable Markdown or JSON of the whole document: Docling or Marker; Nougat or MinerU when formulas matter.
- One paper, no installation, file must stay on your machine: the in-browser extractor, then a DOI lookup.