This converter uses Tesseract.js 7, a browser port of the open-source Tesseract OCR engine. The code, WebAssembly core, and selected language data load only after you start extraction. Tesseract performs recognition in its own worker, so the page remains responsive while your device processes the image.
OCR accuracy depends more on the source than on the file format. Upright printed text, even lighting, sharp focus, strong foreground contrast, and enough pixel detail all help. Small screenshots can improve after upscaling, while skewed pages, decorative fonts, columns, handwriting, glare, and blurred camera photos are more likely to need manual correction.