Add PDF→markdown batch converter and research-library workflow
convert.py walks pdfs/ (recursing topic subfolders), mirrors a .md tree into md/ via pymupdf4llm, idempotent on mtime. Detects no-text-layer PDFs (needs-ocr.txt) and falls back to plain per-page text when pymupdf4llm's layout pass returns near-empty despite a real text layer. Pin pymupdf4llm==0.3.4 (lightweight line; 1.27.x bundles an ML/OCR pipeline that fails on plain text PDFs). PDFs gitignored (copyrighted, large) — only generated markdown is committed. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
This commit is contained in:
File diff suppressed because it is too large
Load Diff
12731
md/Dark Mirrors_ Azazel and Satanael in Early Jewish Demonology.md
Normal file
12731
md/Dark Mirrors_ Azazel and Satanael in Early Jewish Demonology.md
Normal file
File diff suppressed because it is too large
Load Diff
File diff suppressed because it is too large
Load Diff
File diff suppressed because it is too large
Load Diff
File diff suppressed because it is too large
Load Diff
30314
md/The Encyclopedia Of Demons And Demonology.md
Normal file
30314
md/The Encyclopedia Of Demons And Demonology.md
Normal file
File diff suppressed because it is too large
Load Diff
Reference in New Issue
Block a user