What is GROBID? The open-source scholarly PDF parser, explained
Short answer: GROBID (short for GeneRation Of BIbliographic Data) is an open-source machine-learning library that reads scholarly PDFs, such as journal articles, preprints, theses and conference papers, and turns them into structured data. It extracts the header metadata (title, authors, affiliations, abstract, keywords), parses the reference list into structured citations, and can restructure the full text into sections, paragraphs, figures, tables and notes. Its output is TEI XML.
Note on this website: grobid.org is an independent site and is not run by the GROBID developers. The official sources are the GitHub repository github.com/kermitt2/grobid and the documentation at grobid.readthedocs.io. Always check those for the current release, licence and installation instructions.
What GROBID extracts
- Header metadata: title, authors (split into first, middle and last names), affiliations and addresses, abstract, keywords, publication details, and identifiers such as DOI when they are printed in the document.
- References: the bibliography section is located, split into individual references, and each one is parsed into fields (authors, title, journal or book title, volume, issue, pages, year, DOI, and so on).
- Citation contexts: in-text citation callouts such as "[12]" or "(Smith et al., 2019)" can be linked to the matching reference entry.
- Full-text structure: section headings, paragraphs, figures and tables with captions, formulas, footnotes, and acknowledgement and funding statements.
- PDF coordinates: optionally, the page and bounding-box position of extracted elements, which is useful for highlighting results on top of the original PDF.
- Small parsers: separate services parse a single raw citation string, a list of author names, an affiliation string, or a date.
How it works
GROBID does not rely on one big rule set. It chains a cascade of sequence-labelling models, each specialised for one part of a document: one model segments the whole document into zones (front matter, body, references, annexes), another labels the header fields, another splits and labels references, another labels the name parts inside an author string, and so on. Each model only has to solve a narrow problem, which keeps them fast and comparatively accurate.
The default models are linear-chain CRFs (conditional random fields). Recent versions can also use deep-learning models (for example BiLSTM-CRF and transformer-based models via the companion DeLFT library) for some steps, which can improve accuracy at the cost of more memory and, ideally, a GPU. The project ships both a lighter CRF-only Docker image and a larger "full" image that includes the deep-learning models.
Optionally, GROBID can consolidate its results against an external bibliographic database (Crossref, or a self-hosted biblio-glutton service): it looks the extracted record up, and if it finds a confident match it can correct fields and add the DOI.
The TEI XML output
GROBID's main output format is TEI XML, a mature standard for encoding texts. A processed paper becomes a <TEI> document where the header lives in <teiHeader>, the body text in <text><body>, and each reference is a <biblStruct> inside <listBibl>. Some services can also return BibTeX for header and reference results. If you need JSON, the common route is to convert the TEI yourself (any XML library works) or use an existing converter; the alternatives guide mentions tools that output JSON directly.
Who uses GROBID
GROBID has been developed since 2008 by Patrice Lopez and contributors and has been open source (Apache License 2.0) since 2011. It is widely used in scholarly infrastructure, for example in open-access aggregators, repository platforms, digital libraries and large research corpora built from PDFs, and in end-user tools: the JabRef reference manager can call a GROBID service to pull metadata from PDFs, and Apache Tika has a journal parser that uses a GROBID server. The project README lists many more users.
How good is GROBID?
For born-digital scholarly articles, GROBID is generally considered one of the most accurate open-source tools for header metadata and reference parsing, and it is fast enough to process large collections. The documentation publishes end-to-end benchmark results on public article collections, field by field, so you can check how it performs on material similar to yours. Accuracy drops on scanned documents without a good text layer, unusual layouts, non-article documents (books, slides, forms) and some non-English material. As with any extractor, check critical records by hand.
Is GROBID secure to use with confidential papers?
When you run GROBID yourself, for example with Docker on your own machine or server, the PDFs never leave your infrastructure. The only outbound traffic is optional consolidation (lookups sent to Crossref or your own biblio-glutton) which you can turn off. Public demo instances are a different matter: anything you upload there is processed on someone else's server, so do not use them for confidential or unpublished work.
How to try it
- Run it locally with Docker. This is the recommended path and takes a few minutes. Step-by-step commands are in Running GROBID with Docker and the REST API.
- Public demo. The project README links to an online demo instance for quick tests. It is shared and rate-limited, so it is not meant for batch jobs.
- No installation at all. For a quick look at a single PDF, the in-browser extractor on this site pulls out title, authors, DOI, arXiv ID and references without uploading anything. It uses simple heuristics, not GROBID's models, so expect lower accuracy.
Frequently asked questions
Is GROBID free?
Yes. GROBID is open source under the Apache License 2.0. You can use it commercially and run it on your own servers.
What language is GROBID written in?
GROBID is written in Java and runs as a web service with a REST API. Client libraries exist for Python, Java and Node.js, and any language that can send an HTTP request can use it.
Where is the official GROBID documentation?
At grobid.readthedocs.io. Source code, issues and releases are on GitHub.
Can GROBID handle scanned PDFs?
GROBID works from the text layer of the PDF. Image-only scans need OCR first; see Working with scanned PDFs and OCR.