OCR It: open-source tool to pull text from un-copyable documents for LLM processing
TLDR
OCR It is an open-source command-line tool that extracts text from PDFs, images, and scanned documents that cannot be copied directly, producing clean output suitable for feeding into LLM pipelines. The project targets preprocessing steps in document analysis workflows where native text extraction fails. It is lightweight and installable from GitHub.