- 86kPaddleOCRrag-document
Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports 100+ languages.
86kApache-2.0 - 75kTesseract OCRrag-document
Tesseract Open Source OCR Engine (main repository)
75kApache-2.0 - 39kTesseract.jsrag-document
Pure Javascript OCR for more than 100 Languages ๐๐๐ฅ
39kApache-2.0 - 38kMarkerrag-document
Convert PDF to markdown + JSON quickly with high accuracy
38kApache-2.0 - 30kEasyOCRrag-document
Ready-to-use OCR with 80+ supported languages and all popular writing scripts including Latin, Chinese, Arabic, Devanagari, Cyrillic and etc.
30kApache-2.0 - SponsorReach 50,000+ buyers
Enterprise buyers looking for private AI solutions see your brand here.
- 21kSuryarag-document
OCR, layout analysis, reading order, table recognition in 90+ languages
21kApache-2.0 - 17kLaTeX OCRrag-document
pix2tex: Using a ViT to convert images of equations into LaTeX code.
17kMIT - 6.2kdocTRrag-document
docTR (Document Text Recognition) - a seamless, high-performing & accessible library for OCR-related tasks powered by Deep Learning. Ongoing development and maintenance by t2k.
6.2kApache-2.0
Stop paying for AI APIs. Everything here runs on your hardware.
Enterprise buyers looking for private AI solutions see your brand here.
No fluff. No spam. Join 12,000+ builders.