Running GROBID with Docker, the REST API and the Python client
This guide shows the shortest path from zero to a working GROBID service on your own machine, then how to send it PDFs from the command line and from Python. It is an independent summary. The official documentation is the reference for options and for anything that has changed in newer releases.
1. Start GROBID with Docker
Install Docker, then pick an image. The project publishes two flavours:
- CRF-only image (
lfoppiano/grobid): smaller, starts quickly, runs fine on a laptop CPU. A good default. - Full image (
grobid/grobid): includes the deep-learning models. Larger download and more memory, best with a GPU, somewhat better accuracy on some fields.
Use the release tag listed on the GitHub releases page in place of VERSION below (avoid latest, which may not be what you expect):
docker run --rm --init --ulimit core=0 -p 8070:8070 lfoppiano/grobid:VERSION
# or the full image
docker run --rm --init --ulimit core=0 -p 8070:8070 grobid/grobid:VERSION
When the logs settle, open http://localhost:8070 for the built-in web console, or check the service from a terminal:
curl http://localhost:8070/api/isalive # returns true when ready
curl http://localhost:8070/api/version
Plan for a few gigabytes of RAM. If the container is killed while processing large batches, give Docker more memory or lower the number of parallel requests.
2. The main REST endpoints
| Endpoint (POST) | Input | What you get |
|---|---|---|
/api/processHeaderDocument | PDF file in input | Title, authors, affiliations, abstract, keywords and identifiers as TEI (BibTeX available on request) |
/api/processFulltextDocument | PDF file in input | Header, body structure, figures, tables and parsed references as TEI |
/api/processReferences | PDF file in input | Only the parsed bibliography |
/api/processCitation | One raw reference string in citations | That reference parsed into fields |
/api/processCitationList | Several raw reference strings | Each one parsed |
Examples with curl:
# Full text and references of one paper, as TEI XML
curl --form input=@./paper.pdf http://localhost:8070/api/processFulltextDocument > paper.tei.xml
# Header only, with consolidation and raw reference strings kept
curl --form input=@./paper.pdf --form consolidateHeader=1 \
http://localhost:8070/api/processHeaderDocument
# Parse a single citation string
curl --data-urlencode "citations=Lopez P. GROBID: Combining automatic bibliographic data recognition and term extraction for scholarship publications. ECDL 2009." \
http://localhost:8070/api/processCitation
Useful parameters
consolidateHeaderandconsolidateCitations:0turns consolidation off;1looks the record up and merges the matched metadata;2only adds the DOI (and similar identifiers) from the match. Check the docs for the values supported by your version.includeRawCitations=1: keep the original reference string next to each parsed reference, handy for auditing.includeRawAffiliations=1: same idea for affiliation strings.teiCoordinates: request page coordinates for elements (for examplepersName,figure,ref,biblStruct,formula,s), repeated once per element type.segmentSentences=1: wrap sentences of the body text in<s>elements.
3. Consolidation with Crossref (and the email setting)
Consolidation sends a query for each extracted header or reference to a bibliographic service and uses a confident match to correct fields and add a DOI. With Crossref, identify yourself so that your requests go to Crossref's "polite" pool: set your contact email in GROBID's configuration file (grobid-home/config/grobid.yaml, under the Crossref section, the mailto value). In Docker you can mount your own edited copy of that file into the container. For heavy use, a self-hosted biblio-glutton instance avoids rate limits entirely.
Consolidation slows processing down and sends titles and author names to the external service, so turn it on only when you need DOIs or corrected metadata.
4. Batch processing with the Python client
The official Python client handles concurrency, retries and output files for whole folders of PDFs:
pip install grobid-client-python
# config.json
{
"grobid_server": "http://localhost:8070",
"batch_size": 100,
"sleep_time": 5,
"timeout": 120
}
# command line: every PDF in ./pdfs becomes a .tei.xml file in ./tei
grobid_client --config ./config.json --input ./pdfs --output ./tei --n 4 processFulltextDocument
Or from Python code:
from grobid_client.grobid_client import GrobidClient
client = GrobidClient(config_path="./config.json")
client.process("processFulltextDocument", "./pdfs", output="./tei",
consolidate_citations=True, include_raw_citations=True, n=4)
Keep n (parallel requests) close to the number of CPU cores available to the container. Clients also exist for Java and Node.js, and since the service is plain HTTP, any language works.
5. Reading the TEI output
A minimal Python example that pulls the title and the DOIs of references out of a TEI file:
from lxml import etree
ns = {"tei": "http://www.tei-c.org/ns/1.0"}
doc = etree.parse("paper.tei.xml")
title = doc.findtext(".//tei:titleStmt/tei:title", namespaces=ns)
dois = [i.text for i in doc.iterfind(".//tei:listBibl//tei:idno[@type='DOI']", namespaces=ns)]
print(title, len(dois))
Troubleshooting
- HTTP 503: the server is busy; all worker slots are in use. Lower the client concurrency or wait and retry (the Python client retries for you).
- Empty or garbled output: check whether the PDF has a usable text layer. Scanned PDFs need OCR first.
- References missing: try
processReferenceson its own and enableincludeRawCitationsto see what was detected. - Slow processing: disable consolidation, use the CRF-only image, or add CPU cores.
Using a public demo instead
The GROBID README links to an online demo for quick, small tests. It is shared with everyone, may be slow or unavailable, and processes your file on a remote server. For anything sensitive or in bulk, run your own container. For a quick look at one PDF without uploading it anywhere, the in-browser extractor here is another option (heuristic, not GROBID).