GROBID alternatives for PDF metadata and reference extraction

Last reviewed on October 3, 2026

GROBID is the default choice for turning scholarly PDFs into structured bibliographic data, but it is not the only one, and it is not always the most convenient. This page groups the main open-source alternatives by what you are actually trying to do. Projects change quickly: check each project's repository for current status and licence before you commit to it.

Quick comparison

ToolBest atOutputRuns as
GROBIDHeader metadata, reference parsing, full-text structure of scholarly articlesTEI XML (BibTeX for some services)Java service with REST API, Docker
CERMINEMetadata and references from journal articlesNLM JATS XMLJava library, command line, web service
AnyStyleParsing raw reference strings; finding references in textBibTeX, CSL-JSON, JSON, XMLRuby gem and command line
NougatConverting academic PDFs (including formulas) to markup with a vision modelMarkdown-like text with LaTeX mathsPython, GPU recommended
DoclingGeneral document conversion with layout and table structureMarkdown, HTML, JSONPython library and command line
MarkerFast PDF to Markdown conversion for many document typesMarkdown, JSONPython, CPU or GPU
MinerUPDF to Markdown/JSON with layout, tables and formulasMarkdown, JSONPython, GPU recommended
Reference managers (Zotero, JabRef, Mendeley)Getting a clean record for one paperLibrary entries, any export formatDesktop apps
GROBID Tools extractorA quick, private look at one PDF without installing anythingJSON, BibTeX, RIS, CSL-JSON, Markdown, DOCXYour browser (heuristic, not GROBID)

If you need bibliographic metadata and references

GROBID remains the strongest general-purpose open-source option for this job. Its cascade of specialised models handles headers, author names, affiliations and references separately, and optional consolidation against Crossref adds DOIs. See how to run it with Docker.

CERMINE (Content ExtRactor and MINEr), from the University of Warsaw, covers similar ground: it extracts metadata, structured references and content from born-digital journal articles and outputs JATS XML. It is a reasonable second opinion, but development has been much less active than GROBID's in recent years.

AnyStyle focuses on references. Give it raw reference strings (or a plain-text document) and it returns parsed fields as BibTeX or CSL-JSON. It is easy to retrain on your own examples, which makes it a good fit for unusual citation styles. It does not do PDF layout analysis on its own, so pair it with a text extractor.

Older research tools such as ParsCit and AllenAI's Science Parse are still cited in papers but are no longer actively maintained; prefer the options above for new work.

If you need the full text as Markdown or JSON

Many people searching for "GROBID JSON output" really want readable text with structure, for example to feed a search index or a language model. Newer document-conversion tools target exactly that:

These tools are not bibliographic parsers: they will give you the reference section as text, not as structured citations. A common pattern is to use one of them for the body text and GROBID (or AnyStyle) for metadata and references. If you already have GROBID TEI and just want JSON, converting the TEI with an XML library is usually simpler than switching tools.

If you only need a clean record for a handful of papers

You may not need a parser at all. Reference managers identify a paper from its DOI, arXiv ID or ISBN and fetch the authoritative record: Zotero's "Retrieve Metadata for PDF" and "Add Item by Identifier", Mendeley's automatic metadata lookup, and JabRef's PDF import (which can use a GROBID service). Crossref's search and its Simple Text Query tool match free-text references to DOIs. See how to find the DOI of a PDF and from extracted records to a reference list.

Choosing