Skip to content

OCR ​

Turn scanned pages into editable, searchable text — entirely on your device.

How it works ​

  1. Open a scanned PDF (one with no selectable text layer).
  2. Pick the Edit text tool (E), then choose OCR this page.
  3. Tesseract runs locally in WebAssembly. Each recognized line becomes a real editable text box with the page background sampled behind it.
  4. Edit the recognized text like any other line, then export — flat scans become clean vector text.

Languages ​

Four languages ship with the app:

LanguageCodeOne-time download
Englisheng~2.8 MB
Españolspa~1.1 MB
Françaisfra~0.6 MB
Deutschdeu~0.8 MB

Each language pack loads once (on first use) and is cached afterwards — OCR works fully offline from then on.

Privacy ​

OCR is fully on-device. The scan never leaves your machine; the only network activity is the one-time language-data download, which you trigger yourself.

TIP

After OCR, recognized lines behave like native text: you can edit them in place, and export produces real searchable, selectable vector text.

Automate it ​

The CLI and MCP server currently focus on text editing, redaction, and page operations. For scanned documents in scripts, pre-OCR in the app first, then automate against the text layer.

MIT licensed — free for personal, commercial, and everything in between.