Convert pdf to bibtex

Convert PDF to BIBTEX

Extract validated BibTeX citations from scholarly PDFs with Zotero, DOI metadata, OCR, and batch tools.

Convert PDF files online

We don’t have a dedicated online converter for PDF to BIBTEX yet, but you can convert PDF files online to these and more formats:

How to convert pdf to bibtex file

Researchers receive articles as PDFs but need .bib records to cite them in LaTeX, Overleaf, or a shared reference database. The useful result is bibliographic metadata—authors, title, publication details, and DOI—not a conversion of PDF page layout.

PDF and BibTeX files

PDF (Portable Document Format) is a fixed-layout document format created by Adobe and standardized as ISO 32000. Publishers, universities, businesses, scanners, office applications, and web browsers use PDF because it preserves a document’s visual appearance, fonts, images, and pagination across systems.

A scholarly PDF may contain publisher-supplied metadata, selectable text, a DOI, and embedded citation information. A scanned PDF may contain only page images, so software must first apply OCR and then infer metadata from the recognized text.

BibTeX is a plain-text bibliography database format used by BibTeX and BibLaTeX workflows in LaTeX. It stores entries such as @article, @book, and @inproceedings, with fields including author, title, journal, year, volume, pages, doi, and url.

The conventional extension is .bib; although some software accepts .bibtex, use .bib for widest compatibility. BibTeX files can be edited in a text editor, tracked in Git, merged with other libraries, and cited through stable keys such as smith2024.

Extract a record with Zotero

Zotero is the most practical desktop option for individual academic PDFs. It identifies papers from a DOI, embedded metadata, publisher data, or title matching, then exports the resulting library record as BibTeX.

  1. Drag paper.pdf into the Zotero library.
  2. Right-click the PDF and select Retrieve Metadata for PDF. Zotero creates a parent bibliographic item if it identifies a match.
  3. Compare the created item with the PDF’s first page and the publisher or DOI landing page.
  4. Right-click the parent item and select Export Item…, then choose BibTeX and save the output as, for example, references.bib. To export multiple records, place them in a collection and use File → Export Library… or export the collection.

For automatic citation keys and more configurable BibTeX or BibLaTeX output, install Zotero’s Better BibTeX extension. Check its generated key before using it in a manuscript, especially when author names or publication years are incomplete.

cb2Bib is a useful desktop alternative for extracting citation details from PDF text or copied text. It is most effective when the PDF has selectable text and a clear title, author line, DOI, or journal reference; review its proposed entry before saving it to a .bib database.

JabRef can search for metadata from a DOI, ISBN, title, or imported PDF, then save the library directly as a BibTeX database. Its metadata lookup is preferable to manually transcribing a citation, but the resulting entry still requires verification.

Use DOI and online metadata services

If the PDF shows a DOI, use that DOI instead of uploading the document to an unknown converter. Search the DOI or exact title at search.crossref.org, verify the authors, publication, and year, then copy or download the BibTeX citation. Crossref metadata is supplied by publishers and registration agencies, but records can still be incomplete or corrected after publication.

Many DOI registries support BibTeX content negotiation. This command requests a BibTeX record directly from the DOI resolver:

curl -L -H "Accept: application/x-bibtex" https://doi.org/10.1234/example > reference.bib

Replace the sample DOI with the paper’s DOI. Inspect the generated file because the citation key, title capitalization, page range, and entry type may need adjustment.

Google Scholar can provide a fallback record: find the exact paper, select the quotation-mark Cite control, and choose BibTeX. Scholar is useful for papers without a DOI, but its metadata is aggregated from multiple sources and should not be treated as authoritative. Do not upload unpublished, confidential, or copyrighted PDFs to a web converter unless its privacy terms permit that use.

Process scans and batches

Run OCR before importing an image-only scan. Adobe Acrobat provides All tools → Scan & OCR → Recognize text; the local open-source option is:

ocrmypdf scan.pdf searchable.pdf

Import searchable.pdf into Zotero, cb2Bib, or JabRef after OCR. OCR can misread initials, hyphens, accented names, and DOIs, so search the recovered DOI or title against publisher metadata rather than trusting OCR text alone.

For large scholarly-PDF batches, GROBID extracts structured header and reference metadata to TEI XML. After starting its local service, process a file with:

curl -F input=@paper.pdf http://localhost:8070/api/processHeaderDocument > paper.tei.xml

GROBID does not directly create a finished BibTeX database from that endpoint; a TEI-to-BibTeX script or a metadata pipeline is required. Use it when batch extraction justifies setup and validation effort.

Validate the exported entry

A PDF does not inherently contain a BibTeX entry. Identification tools combine embedded metadata, extracted text, DOI lookup, and title matching, so a plausible-looking record can describe the wrong work or contain incorrect fields.

Check the entry type; author spelling and order; title; journal, book, or conference title; year; volume; issue; pages or article number; DOI; and URL. Use the DOI landing page or publisher page as the primary check. Protect acronyms and required capitalization in BibTeX title fields with braces when the selected bibliography style might downcase them, for example title = {Methods for {PDF} and {LaTeX} Archives}.

BibTeX stores citation metadata, not the PDF itself. Keep the PDF as a Zotero attachment or in a project folder, and retain a DOI or stable publisher URL in the .bib record.

PDF vs BIBTEX: format comparison

How the PDF and BIBTEX formats compare on the properties that matter most for this conversion.

Comparison of the PDF and BIBTEX file formats
Property .PDF Portable Document Format .BIBTEX BibTeX bibliography database
Editable text Limited Yes
Fixed page layout Yes No
Text formatting Extensive Basic
Images Yes No
Interactive elements Limited No
Password protection Basic Not supported
Opens in a web browser Yes, natively No
Typical file size Medium Very small
Open standard Yes Partly open
Best used for Publishing Data exchange
Introduced 1993 1985
Developer Adobe Systems (now Adobe Inc.); standardized by ISO Oren Patashnik (original BibTeX system, Donald E. Knuth's TeX ecosystem)
MIME type application/pdf —

Frequently asked questions

Will converting a PDF to BibTeX preserve the whole document?

No. The result contains citation records rather than the PDF's pages, layout, images, or full text. Only bibliographic details that can be identified in the PDF are carried over.

Why are my BibTeX fields missing or incorrect after converting a PDF?

PDFs do not always contain complete, structured citation metadata, so authors, titles, publication dates, page ranges, and identifiers may need to be inferred from text. Scanned PDFs can also produce errors because optical character recognition may misread characters.

Can a PDF be converted to BibTeX if it has several references?

Yes, if the references can be detected and parsed, the result can contain multiple BibTeX entries. Some references may be skipped or combined incorrectly when the PDF has unusual formatting, broken text encoding, or incomplete citation information.

Can I convert the BibTeX file back into the original PDF?

No. BibTeX stores citation metadata and cannot reconstruct the original pages, formatting, figures, or text. The conversion is therefore not reversible.