Use real text first
If a PDF already contains selectable text, the app can extract it directly and reserve OCR for image-only pages.
Classic local OCR · v1.0.0
Persian PDF OCR turns scanned or selectable Persian and bilingual pages into editable files—page by page, on your Windows computer, without sending documents to an OCR cloud.
Windows x64 · 85.8 MiB classic Tesseract edition · The Surya AI engine is not bundled in v1.0.0.
A controlled extraction workflow
Choose the pages, DPI, languages and output format yourself, then keep both per-page text and a combined result.
If a PDF already contains selectable text, the app can extract it directly and reserve OCR for image-only pages.
The classic engine starts with fas+eng to preserve Persian text while recognising Latin words and numbers.
Select a range, keep separate outputs for each page and preserve the raw extraction before later cleanup.
Save a combined TXT, Markdown, JSON or DOCX file, with UTF-8 handling designed for Persian text.
Documents stay local
The classic workflow renders and recognises pages on your computer. No PDF page or extracted text is submitted to ChatGPT or an online OCR provider.
System requirements
Higher DPI can improve difficult scans, but it also increases processing time and temporary storage use.
Extract the portable release to a local writable folder before launch.
4 GB can be enough for simple pages; large, high-DPI PDFs benefit from more memory.
Keep free space for the source PDF, temporary page images and each requested export.
Version 1.0.0 focuses on the predictable classic engine with Persian and English language data.
No dedicated GPU is required. Page count and DPI have the biggest effect on run time.
Everyday extraction works locally once the portable package is available.
From scan to document
OCR is a draft, not a guarantee. Preserve the original PDF and review names, numbers, tables and right-to-left order.
Persian PDF OCR · v1.0.0
Download the classic local edition for Persian and bilingual PDF extraction on Windows.