| type | workflow | |||||||
|---|---|---|---|---|---|---|---|---|
| status | active | |||||||
| trigger | /pdf-to-rag | |||||||
| last_verified | 2026-03-12 | |||||||
| tags |
|
Convert a folder of PDFs into a searchable, AI-queryable knowledge base inside Obsidian. Parses text, tables, scans, and optionally describes images using a vision AI model. Verifies Obsidian Copilot is configured correctly before finishing.
Greet the user and give them a plain-English overview before asking anything:
"This pipeline will turn your PDFs into clean, searchable notes inside your Obsidian vault.
Here's what it does:
- Reads every PDF in a folder you specify
- Converts each one to a clean Markdown note (preserves tables, extracts text from scans)
- Adds a summary and tags to the top of every note automatically
- Saves everything into your Obsidian vault
- Checks that your Obsidian Copilot plugin is set up correctly so you can chat with your documents
Before we start, I have a few quick questions."
Ask these questions one at a time, in plain language. Do not ask all at once.
Question 1 — PDF folder:
"Where is the folder of PDFs you want to process? Paste the full folder path. (Example:
C:\Users\You\Documents\ClientFilesor/Users/you/Documents/ClientFiles)"
Question 2 — Obsidian vault:
"What is the full path to your Obsidian vault? This is the root folder where Obsidian stores all your notes. (Example:
C:\Users\You\Documents\MyVaultor/Users/you/Documents/MyVault)"
Question 3 — Vision mode (explain both options clearly):
"Do your PDFs contain images, diagrams, charts, or photos that matter?
Option A — Vision Mode (Recommended for image-heavy docs) An AI vision model looks at each image in your PDFs and writes a description of what it sees — so those visuals become searchable text in your vault.
Cost: roughly $0.001–$0.002 per image (that's a fraction of a cent). A 100-page document with 10 images costs about 1–2 cents total. Very cheap.
Requires: an OpenRouter API key (openrouter.ai — free to sign up, pay only for what you use).
Option B — Text Only (No API key needed) Docling still extracts all text, tables, and OCR content. Images are left as placeholders in the notes. Free to run, no API key required.
Which would you like? Type A for Vision Mode or B for Text Only."
If they choose A — Vision Mode, ask:
"Do you have your OpenRouter API key saved as an environment variable called
OPENROUTER_API_KEY? (Yes/No)If you're not sure, type No and I'll show you how to set it up — it takes about 2 minutes."
If they say No or are unsure, explain:
"No problem. Here's how to set it up:
Windows:
- Search for 'Environment Variables' in your Start menu
- Click 'Edit the system environment variables'
- Click 'Environment Variables...'
- Under 'User variables', click New
- Variable name:
OPENROUTER_API_KEY- Variable value: paste your key from openrouter.ai
- Click OK and restart your terminal
Mac/Linux: Add this line to your
~/.zshrcor~/.bashrcfile:export OPENROUTER_API_KEY=your_key_hereThen runsource ~/.zshrc(or restart your terminal).Let me know when that's done and we'll continue."
Run:
python --versionIf Python is not found, try python3 --version.
If neither works:
"Python isn't installed on this machine yet. You'll need Python 3.10 or newer. Download it from python.org — choose the version that says 'Recommended' and run the installer. Make sure to check the box that says 'Add Python to PATH' during installation. Come back when that's done."
If Python is found but is below 3.10, warn the user and ask them to upgrade.
Run the appropriate install command based on the chosen mode:
Text Only mode:
pip install doclingVision mode:
pip install docling openai pymupdfLet the user know what's happening:
"Installing the required packages — this may take a minute or two. Docling is a document parsing tool from IBM; the others handle image processing and API calls."
If pip fails with a permissions error, try pip install --user or pip3 install.
Based on the chosen mode, write the appropriate Python script to the user's vault root (or a location they specify). Name it pdf_rag_pipeline.py.
Use the paths collected in Step 2 for SOURCE_FOLDER, VAULT_LIBRARY, and VAULT_ROOT.
#!/usr/bin/env python3
"""
pdf_rag_pipeline.py — PDF to Obsidian RAG Pipeline (Text Only mode)
=====================================================================
Converts every PDF in SOURCE_FOLDER into an Obsidian-ready Markdown file.
Pipeline for each PDF:
1. Docling parses the PDF (OCR + table reconstruction)
2. A summary, tags, and YAML frontmatter are generated from the content
3. The finished note is saved to VAULT_LIBRARY/{filename}.md
Requirements:
pip install docling
Usage:
python pdf_rag_pipeline.py
"""
import datetime
import re
from pathlib import Path
from docling.document_converter import DocumentConverter, PdfFormatOption
from docling.datamodel.base_models import InputFormat
from docling.datamodel.pipeline_options import PdfPipelineOptions
# ---------------------------------------------------------------------------
# CONFIGURATION — set these to your actual paths before running
# ---------------------------------------------------------------------------
SOURCE_FOLDER = Path(r"REPLACE_WITH_YOUR_PDF_FOLDER")
VAULT_LIBRARY = Path(r"REPLACE_WITH_YOUR_VAULT_PATH") / "Library"
VAULT_ROOT = Path(r"REPLACE_WITH_YOUR_VAULT_PATH")
COPILOT_CONFIG = VAULT_ROOT / ".obsidian" / "plugins" / "copilot" / "data.json"
# ---------------------------------------------------------------------------
# DOCLING PARSE
# ---------------------------------------------------------------------------
def parse_pdf(pdf_path: Path) -> str:
options = PdfPipelineOptions()
options.do_ocr = True
options.do_table_structure = True
options.table_structure_options.do_cell_matching = True
converter = DocumentConverter(
format_options={InputFormat.PDF: PdfFormatOption(pipeline_options=options)}
)
result = converter.convert(str(pdf_path))
return result.document.export_to_markdown()
# ---------------------------------------------------------------------------
# FRONTMATTER GENERATION (no API — keyword-based)
# ---------------------------------------------------------------------------
def generate_frontmatter(markdown: str, stem: str) -> str:
today = datetime.date.today().isoformat()
word_count = len(markdown.split())
summary = f"Parsed from {stem}.pdf — {word_count:,} words extracted."
return (
f"---\n"
f"title: \"{stem}\"\n"
f"date: {today}\n"
f"source: {stem}.pdf\n"
f"type: document\n"
f"summary: \"{summary}\"\n"
f"---\n"
)
# ---------------------------------------------------------------------------
# COPILOT CONFIG CHECK
# ---------------------------------------------------------------------------
def verify_copilot_config() -> None:
import json
required = {
"embeddingModelKey": "gemini-embedding-001|google",
"indexVaultToVectorStore": "ON STARTUP",
}
if not COPILOT_CONFIG.exists():
print(" Copilot config not found — install the Obsidian Copilot plugin first.")
return
with COPILOT_CONFIG.open("r", encoding="utf-8") as f:
config = json.load(f)
changed = False
for key, expected in required.items():
current = config.get(key)
if current != expected:
print(f" Fixing: {key}")
print(f" was: {current!r}")
print(f" now: {expected!r}")
config[key] = expected
changed = True
else:
print(f" OK: {key} = {current!r}")
if changed:
backup = COPILOT_CONFIG.with_suffix(".json.bak")
backup.write_text(COPILOT_CONFIG.read_text(encoding="utf-8"), encoding="utf-8")
with COPILOT_CONFIG.open("w", encoding="utf-8") as f:
json.dump(config, f, indent=2, ensure_ascii=False)
print(" Copilot config updated. Backup saved.")
else:
print(" Copilot config is correct — no changes needed.")
# ---------------------------------------------------------------------------
# MAIN
# ---------------------------------------------------------------------------
def run():
if not SOURCE_FOLDER.exists():
print(f"ERROR: PDF folder not found: {SOURCE_FOLDER}")
return
pdfs = sorted(SOURCE_FOLDER.glob("*.pdf"))
if not pdfs:
print(f"No PDFs found in: {SOURCE_FOLDER}")
return
VAULT_LIBRARY.mkdir(parents=True, exist_ok=True)
print(f"PDF to Obsidian RAG Pipeline — Text Only Mode")
print(f"Source: {SOURCE_FOLDER}")
print(f"Library: {VAULT_LIBRARY}")
print(f"Found {len(pdfs)} PDF(s)\n")
succeeded, failed = [], []
for pdf in pdfs:
print(f"Processing: {pdf.name}")
try:
markdown = parse_pdf(pdf)
frontmatter = generate_frontmatter(markdown, pdf.stem)
content = frontmatter + "\n" + markdown
out_path = VAULT_LIBRARY / f"{pdf.stem}.md"
out_path.write_text(content, encoding="utf-8")
print(f" Saved: {out_path.name}")
succeeded.append(pdf.name)
except Exception as e:
print(f" ERROR: {e}")
failed.append(pdf.name)
print(f"\nCopilot Config Check")
verify_copilot_config()
print(f"\nDone. {len(succeeded)} succeeded, {len(failed)} failed.")
if failed:
print("Failed:", ", ".join(failed))
if __name__ == "__main__":
run()#!/usr/bin/env python3
"""
pdf_rag_pipeline.py — PDF to Obsidian RAG Pipeline (Vision Mode)
=================================================================
Converts every PDF in SOURCE_FOLDER into an Obsidian-ready Markdown file.
Images in each PDF are described by an AI vision model so they become
searchable text in your vault.
Pipeline for each PDF:
1. Docling parses the PDF (OCR + table reconstruction)
2. Images are extracted and described by a vision model (Claude Haiku via OpenRouter)
3. YAML frontmatter and a Key Takeaways section are generated
4. The finished note is saved to VAULT_LIBRARY/{filename}.md
Requirements:
pip install docling openai pymupdf
Environment:
OPENROUTER_API_KEY — your OpenRouter API key (openrouter.ai)
Cost estimate: ~$0.001–$0.002 per image (fractions of a cent)
"""
import os
import re
import json
import base64
import datetime
from pathlib import Path
from typing import Optional
import fitz # PyMuPDF
from openai import OpenAI
from docling.document_converter import DocumentConverter, PdfFormatOption
from docling.datamodel.base_models import InputFormat
from docling.datamodel.pipeline_options import PdfPipelineOptions
# ---------------------------------------------------------------------------
# CONFIGURATION — set these to your actual paths before running
# ---------------------------------------------------------------------------
SOURCE_FOLDER = Path(r"REPLACE_WITH_YOUR_PDF_FOLDER")
VAULT_LIBRARY = Path(r"REPLACE_WITH_YOUR_VAULT_PATH") / "Library"
VAULT_ROOT = Path(r"REPLACE_WITH_YOUR_VAULT_PATH")
COPILOT_CONFIG = VAULT_ROOT / ".obsidian" / "plugins" / "copilot" / "data.json"
# Vision model — Claude Haiku is the cheapest option that handles images well.
# Swap the slug at openrouter.ai/models if you want a different model.
VISION_MODEL = "anthropic/claude-haiku-4-5"
# Images smaller than this (in bytes) are skipped — covers icons and decorative lines
IMAGE_MIN_BYTES = 5_000
# How much document text to send when generating frontmatter (in characters)
FRONTMATTER_EXCERPT_LIMIT = 6_000
# ---------------------------------------------------------------------------
# DOCLING PARSE
# ---------------------------------------------------------------------------
def parse_pdf(pdf_path: Path) -> str:
options = PdfPipelineOptions()
options.do_ocr = True
options.do_table_structure = True
options.table_structure_options.do_cell_matching = True
converter = DocumentConverter(
format_options={InputFormat.PDF: PdfFormatOption(pipeline_options=options)}
)
result = converter.convert(str(pdf_path))
return result.document.export_to_markdown()
# ---------------------------------------------------------------------------
# IMAGE TRIGGER SYSTEM
# ---------------------------------------------------------------------------
# Patterns Docling uses to mark images it cannot render as text
_PLACEHOLDER_RE = re.compile(
r"(!\[[^\]]*\]\([^)]*\)"
r"|<!-- [Pp]icture \d+ -->"
r"|<!-- [Ii]mage(?::[^>]*)? -->"
r"|\[IMAGE(?::[^\]]*)?\]"
r"|\[Figure \d+[^\]]*\])",
re.IGNORECASE,
)
def _detect_mime(image_bytes: bytes) -> str:
if image_bytes[:3] == b"\xff\xd8\xff": return "image/jpeg"
if image_bytes[:4] == b"\x89PNG": return "image/png"
if image_bytes[8:12] == b"WEBP": return "image/webp"
return "image/png"
def extract_images(pdf_path: Path) -> list[bytes]:
images, seen = [], set()
doc = fitz.open(str(pdf_path))
for page in doc:
for img_info in page.get_images(full=True):
xref = img_info[0]
if xref in seen:
continue
seen.add(xref)
raw = doc.extract_image(xref)["image"]
if len(raw) >= IMAGE_MIN_BYTES:
images.append(raw)
doc.close()
return images
def describe_image(client: OpenAI, image_bytes: bytes, context: str = "", idx: int = 1) -> str:
mime = _detect_mime(image_bytes)
b64 = base64.standard_b64encode(image_bytes).decode()
url = f"data:{mime};base64,{b64}"
user_content = []
if context:
user_content.append({"type": "text", "text": f"Surrounding context:\n{context}\n\nNow describe the image:"})
else:
user_content.append({"type": "text", "text": "Describe this image:"})
user_content.append({"type": "image_url", "image_url": {"url": url}})
resp = client.chat.completions.create(
model=VISION_MODEL,
max_tokens=350,
messages=[
{"role": "system", "content": (
"You are a precise document analyst preparing content for a text-only knowledge base. "
"Describe the provided image clearly: cover the type of visual (chart, photo, diagram, table, etc.), "
"the key data or labels visible, and the main insight it conveys. "
"Write 2–4 sentences. Do not start with 'This image shows' or 'This is a'."
)},
{"role": "user", "content": user_content},
],
)
return resp.choices[0].message.content.strip()
def inject_image_descriptions(markdown: str, pdf_path: Path, client: OpenAI) -> str:
placeholders = list(_PLACEHOLDER_RE.finditer(markdown))
if not placeholders:
return markdown
print(f" Found {len(placeholders)} image(s) — extracting and describing...")
pdf_images = extract_images(pdf_path)
if not pdf_images:
print(" No extractable images found — skipping vision step.")
return markdown
result = markdown
for i, match in enumerate(placeholders):
img_bytes = pdf_images[min(i, len(pdf_images) - 1)]
start = max(0, match.start() - 400)
end = min(len(result), match.end() + 400)
context = result[start:end]
size_kb = len(img_bytes) // 1024
print(f" Image {i+1}/{len(placeholders)} ({size_kb} KB)...")
try:
desc = describe_image(client, img_bytes, context, i + 1)
except Exception as e:
print(f" Warning: vision call failed for image {i+1}: {e}")
desc = f"[Image {i+1} — description unavailable]"
result = result.replace(match.group(0), f"\n> **[Image {i+1}]** {desc}\n", 1)
return result
# ---------------------------------------------------------------------------
# FRONTMATTER + KEY TAKEAWAYS
# ---------------------------------------------------------------------------
def generate_frontmatter_and_takeaways(client: OpenAI, markdown: str, stem: str) -> tuple[str, str]:
today = datetime.date.today().isoformat()
excerpt = markdown[:FRONTMATTER_EXCERPT_LIMIT]
prompt = f"""You are preparing a document for an Obsidian knowledge base.
Filename (no extension): {stem}
Today's date: {today}
Document content (may be truncated):
---
{excerpt}
---
Return EXACTLY two sections separated by the delimiter |||
SECTION 1 — YAML frontmatter (no code fences):
---
title: "<document title>"
date: {today}
category: "<one of: Research, Finance, Legal, Technical, Marketing, Operations, Personal, Other>"
tags: [<3–5 lowercase tags>]
summary: "<1–2 sentence plain-English summary>"
---
SECTION 2 — Key Takeaways (3–5 bullets, no header):
- <insight>
- <insight>
- <insight>
Return format: SECTION_1_YAML|||SECTION_2_BULLETS
No other text."""
resp = client.chat.completions.create(
model=VISION_MODEL,
max_tokens=700,
messages=[{"role": "user", "content": prompt}],
)
raw = resp.choices[0].message.content.strip()
if "|||" in raw:
yaml_part, bullets = raw.split("|||", 1)
else:
yaml_part = f"---\ntitle: \"{stem}\"\ndate: {today}\ncategory: Document\ntags: []\nsummary: \"\"\n---"
bullets = raw if raw.startswith("-") else "- Review document for key insights."
return yaml_part.strip(), bullets.strip()
def assemble(frontmatter: str, takeaways: str, body: str) -> str:
return f"{frontmatter}\n\n## Key Takeaways\n\n{takeaways}\n\n---\n\n{body}"
# ---------------------------------------------------------------------------
# COPILOT CONFIG CHECK
# ---------------------------------------------------------------------------
def verify_copilot_config() -> None:
required = {
"embeddingModelKey": "gemini-embedding-001|google",
"indexVaultToVectorStore": "ON STARTUP",
}
if not COPILOT_CONFIG.exists():
print(" Copilot config not found — install the Obsidian Copilot plugin first.")
return
with COPILOT_CONFIG.open("r", encoding="utf-8") as f:
config = json.load(f)
changed = False
for key, expected in required.items():
current = config.get(key)
if current != expected:
print(f" Fixing: {key}")
print(f" was: {current!r}")
print(f" now: {expected!r}")
config[key] = expected
changed = True
else:
print(f" OK: {key} = {current!r}")
if changed:
backup = COPILOT_CONFIG.with_suffix(".json.bak")
backup.write_text(COPILOT_CONFIG.read_text(encoding="utf-8"), encoding="utf-8")
with COPILOT_CONFIG.open("w", encoding="utf-8") as f:
json.dump(config, f, indent=2, ensure_ascii=False)
print(" Copilot config updated. Backup saved.")
else:
print(" Copilot config is correct — no changes needed.")
# ---------------------------------------------------------------------------
# OPENROUTER CLIENT
# ---------------------------------------------------------------------------
_client: Optional[OpenAI] = None
def get_client() -> OpenAI:
global _client
if _client is None:
key = os.getenv("OPENROUTER_API_KEY")
if not key:
raise EnvironmentError(
"OPENROUTER_API_KEY is not set.\n"
"Set it as an environment variable before running this script.\n"
"See the setup instructions for how to do this."
)
_client = OpenAI(base_url="https://openrouter.ai/api/v1", api_key=key)
return _client
# ---------------------------------------------------------------------------
# SINGLE FILE PIPELINE
# ---------------------------------------------------------------------------
def process_pdf(pdf_path: Path) -> Optional[Path]:
print(f"\n{'─' * 50}")
print(f" {pdf_path.name}")
print(f"{'─' * 50}")
client = get_client()
# 1. Parse
print(" [1/4] Parsing with Docling...")
try:
markdown = parse_pdf(pdf_path)
print(f" {len(markdown.split()):,} words extracted.")
except Exception as e:
print(f" ERROR during parse: {e}")
return None
# 2. Images
print(" [2/4] Checking for images...")
try:
markdown = inject_image_descriptions(markdown, pdf_path, client)
except Exception as e:
print(f" Warning: image step failed ({e}) — continuing without image descriptions.")
# 3. Frontmatter
print(" [3/4] Generating summary and tags...")
today = datetime.date.today().isoformat()
try:
frontmatter, takeaways = generate_frontmatter_and_takeaways(client, markdown, pdf_path.stem)
except Exception as e:
print(f" Warning: summary generation failed ({e}) — using minimal frontmatter.")
frontmatter = f"---\ntitle: \"{pdf_path.stem}\"\ndate: {today}\ncategory: Document\ntags: []\nsummary: \"\"\n---"
takeaways = "- Review document manually."
# 4. Save
print(" [4/4] Saving to vault...")
VAULT_LIBRARY.mkdir(parents=True, exist_ok=True)
content = assemble(frontmatter, takeaways, markdown)
out_path = VAULT_LIBRARY / f"{pdf_path.stem}.md"
try:
out_path.write_text(content, encoding="utf-8")
size_kb = out_path.stat().st_size // 1024
print(f" Saved: {out_path.name} ({size_kb} KB)")
return out_path
except Exception as e:
print(f" ERROR saving: {e}")
return None
# ---------------------------------------------------------------------------
# MAIN
# ---------------------------------------------------------------------------
def run():
if not SOURCE_FOLDER.exists():
print(f"ERROR: PDF folder not found: {SOURCE_FOLDER}")
print("Update SOURCE_FOLDER at the top of this script.")
return
pdfs = sorted(SOURCE_FOLDER.glob("*.pdf"))
if not pdfs:
print(f"No PDFs found in: {SOURCE_FOLDER}")
return
print(f"PDF to Obsidian RAG Pipeline — Vision Mode")
print(f"Source: {SOURCE_FOLDER}")
print(f"Library: {VAULT_LIBRARY}")
print(f"Model: {VISION_MODEL} (via OpenRouter)")
print(f"Found {len(pdfs)} PDF(s)\n")
succeeded, failed = [], []
for pdf in pdfs:
out = process_pdf(pdf)
if out:
succeeded.append(pdf.name)
else:
failed.append(pdf.name)
print(f"\n{'─' * 50}")
print("Copilot Config Check")
verify_copilot_config()
print(f"\nDone. {len(succeeded)} succeeded, {len(failed)} failed.")
if failed:
print("Failed files:", ", ".join(failed))
if __name__ == "__main__":
run()After writing the script, fill in the user's actual paths for SOURCE_FOLDER, VAULT_LIBRARY, and VAULT_ROOT from Step 2.
Tell the user:
"Script saved as
pdf_rag_pipeline.py. I've filled in your folder paths. You're ready to run it."
Run:
python pdf_rag_pipeline.pyMonitor output with the user and explain what each step means as it runs. If errors occur:
| Error | Plain-language fix |
|---|---|
ModuleNotFoundError: docling |
Run pip install docling again |
OPENROUTER_API_KEY not set |
API key not in environment — revisit Step 2 setup |
FileNotFoundError (PDF folder) |
Double-check the SOURCE_FOLDER path for typos |
| Docling OCR errors | Install Tesseract: pip install tesseract (Windows: download from github.com/UB-Mannheim/tesseract) |
embeddingModelKey error in Copilot |
The Copilot config check at the end of the script will fix this automatically |
Tell the user:
"Now restart Obsidian. When it reopens, Copilot will start indexing your new notes automatically (you set it to index on startup).
Once it's done — usually 30 seconds to a few minutes depending on how many files you processed — open the Copilot chat panel, switch to Vault QA mode, and ask a question about your documents.
You should see the answer plus a source link pointing to the exact note it came from."
If this is running inside an Obsidian vault with a daily note, append to today's activity log:
- [HH:MM] /pdf-to-rag — processed [N] PDFs into vault Library. Mode: [Vision/Text Only]. [N] succeeded, [N] failed.
| Issue | Fix |
|---|---|
| Scanned PDF comes out blank | Make sure do_ocr=True is set (it is in this script) — also try pip install tesseract |
| Tables look broken | Docling handles most tables — very unusual layouts may need manual cleanup |
| Copilot not finding new notes | Trigger a manual re-index: Copilot settings → "Force re-index" or restart Obsidian |
| Embeddings 404 error | Check that embeddingModelKey is set to `gemini-embedding-001 |
| Google API key error in Copilot | Paste your Google AI Studio API key directly into Obsidian → Settings → Copilot → API key field |
| Vision model returns generic descriptions | The image might be a simple line or icon — this is expected and harmless |
| What | Cost |
|---|---|
| Per image described | ~$0.001–$0.002 |
| Per document (frontmatter + takeaways) | ~$0.002–$0.004 |
| 100-page doc with 10 images | ~$0.02–$0.04 |
| 100 documents, avg 5 images each | ~$0.35–$0.70 |
| 1,000 documents, avg 5 images each | ~$3.50–$7.00 |
Text Only mode: $0.00 (no API calls required).