PDF to Markdown Converter

An automated, zero-server pipeline that converts PDF documents into pristine, Obsidian-ready Markdown using a flexible suite of OCR and AI engines.

Upload PDF via Issues

How to use

  1. Navigate to the Issues tab (or click the button above).
  2. Click New Issue.
  3. Drag and drop your .pdf file into the issue description box.
  4. (Optional) Add slash commands to the issue description to customize the extraction.
  5. Click Submit new issue.

Within a few minutes, a GitHub Action bot will comment on your issue containing direct download links for the raw .md file and a .zip archive containing the Markdown and any extracted images.

Advanced Configuration (Slash Commands)

You can completely customize how the pipeline processes your PDF by dropping these slash commands anywhere in your issue body.

Engine Selection

The pipeline supports three distinct architectural paths. If no engine is specified, the system defaults to marker.

Modifiers & Overrides

Example Workflows

1. The Fast Default (Clean Digital PDFs)
Just drop the PDF in the issue. No commands needed. Runs standard Marker text extraction in seconds.

2. The "Bad Scan" Fixer

/engine=marker-llm
/force-ocr
/page-range=1-5

Forces Marker to visually read a scanned document, then uses an LLM to clean up the messy OCR formatting. Limited to 5 pages to avoid the 15-minute CPU timeout.

3. The Complex Math & Layout Reader

/engine=vision
/page-range=10-25
/timeout=30

Takes pictures of pages 10 through 25 and lets the default Vision AI perfectly transcribe the LaTeX math and tables, giving it an extended 30-minute window to finish the job.

Maintenance

Processed files are stored in the conversions/ directory and are automatically purged after 30 days to optimize repository storage.