a-sh.app ← All products

Classic local OCR · v1.0.0

Persian PDFs in. Useful text out.

Persian PDF OCR turns scanned or selectable Persian and bilingual pages into editable files—page by page, on your Windows computer, without sending documents to an OCR cloud.

Windows x64 · 85.8 MiB classic Tesseract edition · The Surya AI engine is not bundled in v1.0.0.

  • fas + engPersian-first bilingual OCR
  • TXT · MD · JSON · DOCXUseful export choices
  • Local processingNo document upload

A controlled extraction workflow

Keep every page traceable.

Choose the pages, DPI, languages and output format yourself, then keep both per-page text and a combined result.

PDF

Use real text first

If a PDF already contains selectable text, the app can extract it directly and reserve OCR for image-only pages.

FA

Persian-first defaults

The classic engine starts with fas+eng to preserve Persian text while recognising Latin words and numbers.

1…

Page-level control

Select a range, keep separate outputs for each page and preserve the raw extraction before later cleanup.

Four export formats

Save a combined TXT, Markdown, JSON or DOCX file, with UTF-8 handling designed for Persian text.

Documents stay local

Private PDFs do not belong in a public queue.

The classic workflow renders and recognises pages on your computer. No PDF page or extracted text is submitted to ChatGPT or an online OCR provider.

  • PDF pages and OCR text remain in your selected local folders.
  • Temporary images and raw JSON can contain the same sensitive information as the source.
  • Review filenames and absolute-path metadata before sharing JSON or debug output.
  • Medical, legal and personal documents still require appropriate local access control.

System requirements

Built for ordinary Windows hardware.

Higher DPI can improve difficult scans, but it also increases processing time and temporary storage use.

Operating system

Windows 10 or 11 x64

Extract the portable release to a local writable folder before launch.

Memory

8 GB RAM recommended

4 GB can be enough for simple pages; large, high-DPI PDFs benefit from more memory.

Storage

Room for rendered pages

Keep free space for the source PDF, temporary page images and each requested export.

OCR engine

Tesseract classic

Version 1.0.0 focuses on the predictable classic engine with Persian and English language data.

Performance

CPU-based processing

No dedicated GPU is required. Page count and DPI have the biggest effect on run time.

Connectivity

No cloud required

Everyday extraction works locally once the portable package is available.

From scan to document

A workflow you can inspect.

OCR is a draft, not a guarantee. Preserve the original PDF and review names, numbers, tables and right-to-left order.

  1. Download and extractKeep the application and its language data together in the extracted folder.
  2. Choose the PDF and pagesSelect a page range, output folder, DPI and the Persian/English OCR language combination.
  3. Run local extractionLet the app use selectable text first and OCR the pages that need recognition.
  4. Review before relyingCheck proper names, medication, amounts, dates, tables and multi-column reading order against the source.

Persian PDF OCR · v1.0.0

Keep the document. Free the text.

Download the classic local edition for Persian and bilingual PDF extraction on Windows.

Download ZIP ↓