Running GROBID with Docker, the REST API and the Python client

Last reviewed on October 3, 2026

This guide shows the shortest path from zero to a working GROBID service on your own machine, then how to send it PDFs from the command line and from Python. It is an independent summary. The official documentation is the reference for options and for anything that has changed in newer releases.

1. Start GROBID with Docker

Install Docker, then pick an image. The project publishes two flavours:

Use the release tag listed on the GitHub releases page in place of VERSION below (avoid latest, which may not be what you expect):

docker run --rm --init --ulimit core=0 -p 8070:8070 lfoppiano/grobid:VERSION
# or the full image
docker run --rm --init --ulimit core=0 -p 8070:8070 grobid/grobid:VERSION

When the logs settle, open http://localhost:8070 for the built-in web console, or check the service from a terminal:

curl http://localhost:8070/api/isalive    # returns true when ready
curl http://localhost:8070/api/version

Plan for a few gigabytes of RAM. If the container is killed while processing large batches, give Docker more memory or lower the number of parallel requests.

2. The main REST endpoints

Endpoint (POST)InputWhat you get
/api/processHeaderDocumentPDF file in inputTitle, authors, affiliations, abstract, keywords and identifiers as TEI (BibTeX available on request)
/api/processFulltextDocumentPDF file in inputHeader, body structure, figures, tables and parsed references as TEI
/api/processReferencesPDF file in inputOnly the parsed bibliography
/api/processCitationOne raw reference string in citationsThat reference parsed into fields
/api/processCitationListSeveral raw reference stringsEach one parsed

Examples with curl:

# Full text and references of one paper, as TEI XML
curl --form input=@./paper.pdf http://localhost:8070/api/processFulltextDocument > paper.tei.xml

# Header only, with consolidation and raw reference strings kept
curl --form input=@./paper.pdf --form consolidateHeader=1 \
     http://localhost:8070/api/processHeaderDocument

# Parse a single citation string
curl --data-urlencode "citations=Lopez P. GROBID: Combining automatic bibliographic data recognition and term extraction for scholarship publications. ECDL 2009." \
     http://localhost:8070/api/processCitation

Useful parameters

3. Consolidation with Crossref (and the email setting)

Consolidation sends a query for each extracted header or reference to a bibliographic service and uses a confident match to correct fields and add a DOI. With Crossref, identify yourself so that your requests go to Crossref's "polite" pool: set your contact email in GROBID's configuration file (grobid-home/config/grobid.yaml, under the Crossref section, the mailto value). In Docker you can mount your own edited copy of that file into the container. For heavy use, a self-hosted biblio-glutton instance avoids rate limits entirely.

Consolidation slows processing down and sends titles and author names to the external service, so turn it on only when you need DOIs or corrected metadata.

4. Batch processing with the Python client

The official Python client handles concurrency, retries and output files for whole folders of PDFs:

pip install grobid-client-python

# config.json
{
  "grobid_server": "http://localhost:8070",
  "batch_size": 100,
  "sleep_time": 5,
  "timeout": 120
}

# command line: every PDF in ./pdfs becomes a .tei.xml file in ./tei
grobid_client --config ./config.json --input ./pdfs --output ./tei --n 4 processFulltextDocument

Or from Python code:

from grobid_client.grobid_client import GrobidClient

client = GrobidClient(config_path="./config.json")
client.process("processFulltextDocument", "./pdfs", output="./tei",
               consolidate_citations=True, include_raw_citations=True, n=4)

Keep n (parallel requests) close to the number of CPU cores available to the container. Clients also exist for Java and Node.js, and since the service is plain HTTP, any language works.

5. Reading the TEI output

A minimal Python example that pulls the title and the DOIs of references out of a TEI file:

from lxml import etree
ns = {"tei": "http://www.tei-c.org/ns/1.0"}
doc = etree.parse("paper.tei.xml")
title = doc.findtext(".//tei:titleStmt/tei:title", namespaces=ns)
dois = [i.text for i in doc.iterfind(".//tei:listBibl//tei:idno[@type='DOI']", namespaces=ns)]
print(title, len(dois))

Troubleshooting

Using a public demo instead

The GROBID README links to an online demo for quick, small tests. It is shared with everyone, may be slow or unavailable, and processes your file on a remote server. For anything sensitive or in bulk, run your own container. For a quick look at one PDF without uploading it anywhere, the in-browser extractor here is another option (heuristic, not GROBID).