← Back
Early Access

Interactive OCR

Convert a scanned Tibetan PDF into an editable Word document — one page at a time. Clean pages advance automatically; difficult pages get a focused review loop instead of a global tweak that breaks the rest of the document.

The problem

Traditional OCR pipelines apply one set of settings to an entire document. A page that needs a different model, rotation, or preprocessing either fails quietly or forces a retune that degrades every other page. For translators working from scans, that means hours of cleanup — or starting over.

How this works

  • 1.Each page runs through BDRC Tibetan OCR with a quality score grounded in the same structural rules and word corpus as our spell checker.
  • 2.Pages that look clean are accepted. Pages that don't stay open for retry: an AI diagnostician suggests page-local setting changes — never job-wide overrides.
  • 3.When that isn't enough, an optional AI vision compare (Claude / Gemini) can propose an alternate reading — always as an explicit, per-page action, never a bulk auto-run.
  • 4.Accepted pages export to a clean Word document with page markers, ready for translation work.

Available today

Scanned-PDF OCR already powers the Upload PDF spellcheck path: extract text, flag structural and corpus issues, download an annotated PDF and editable DOCX. The interactive page-by-page assist loop, which leverages an AI diagnostician along with optional AI vision compare leveraging Claude and Gemini, is still in Beta testing.

Word corpus

Spellcheck and OCR quality scoring draw on a syllable inventory built from publicly available Tibetan lexicographic sources — including the Monlam Tibetan Lexicon (Apache-2.0), Botok word lists, and Christian Steinert's public dictionary collection. We use these as a reference inventory for validation on this site, not as a standalone dictionary we redistribute.

Request access

Leave your email if you'd like early access, or if you're evaluating this for translation, archival, or interview purposes. We use this list to gauge interest and reach out when seats open.

Resources for the Orgyen Khandroling Sangha

Source on GitHub

© 2026 Butter Dots Dot Com