Skip to main content

Tesseract.js

Pure Javascript OCR for more than 100 Languages 📖🎉🖥

View on GitHub Official site
Rag Document Apache-2.0 Easy setup 38,551 stars

Overview

Description

Extract text from images in over 100 languages with a pure JavaScript library that runs anywhere — no server or GPU required. Add OCR to your web app or Node service with a single npm install or script tag. The big picture: bring the power of Tesseract to your frontend or backend without any native dependencies.

Technical scorecard

License

Apache-2.0

Commercial use

Yes

OpenAI-compatible API

No

REST API

No

Fine-tuning support

No

Quantization support

No

Docker available

No

GUI / no-code available

No

Telemetry

Unknown

Offline after setup

No

Data & Privacy

Does it send data online?

This listing is not marked offline after setup.

Does it store history?

Not verified in this directory yet. Review the upstream docs for persistence, logs, and workspace storage.

License checks?

Commercial use is marked as allowed or likely allowed by the listed license.

Telemetry?

Unknown

Last verified: Jul 22, 2026. Maintainer verification should be treated as directory guidance, not legal advice.

🛡️

Deploying Tesseract.js in production?

Stop silent telemetry leaks before they trigger an unintended data leak. Get our Airgap Certainty Blueprint for Tesseract.js.

Request Privacy and Telemetry Audit

Setup & Installation

Easy

A developer can usually get this running with standard docs.

# Start with the official project repository
# https://github.com/naptha/tesseract.js

Hardware Requirements

Hardware tagsCPU
LanguagesJavaScript

Works Well With