Category · 45 repos
OCR & Document AI
Turn PDFs, scans and documents into clean, structured data. Ranked by star velocity over the last 24 hours.
Turn any technical book PDF into a Claude Code skill — ready to study, reference, and use while you work.
Python tool for converting files and office documents to Markdown.
Unlimited OCR Works: Welcome the Era of One-shot Long-horizon Parsing.
Automation foundation model for tiny devices: 2-bit, 8-29 MB, tool calls, structured extraction and embeddings on phones, wearables, smart homes, robots, cars and microcontrollers.
A conversational copilot on your phone: understands the other person in QQ / X / Feishu, suggests replies, and fills them into the input box with one tap. You decide whether to send. Non-intrusive; reads the screen only, without hooks or package modification.
Convert documentation websites, GitHub repositories, and PDFs into Claude AI skills with automatic conflict detection
Screen capture, OCR and translation tool.
Windows companion for League of Legends ARAM Mayhem with read-only LCU recommendations, Hextech Augment stats, and in-game OCR overlays.
A free, ad-free, fully offline document-scanning app (Android/Windows): automatic cropping, perspective correction, filters, PDF export, and offline Chinese/English OCR. Rust + Tauri 2
MakeACopy is an open-source document scanner app for Android that allows you to digitize paper documents with OCR functionality. The app is designed to be privacy-friendly, working completely offline without any cloud connection or tracking.
OCR software, free and offline. Open-source, free offline OCR software. Supports screenshots/batch image import, PDF document recognition, watermark/header/footer exclusion, and scanning/generating QR codes. Includes built-in multilingual language packs.
Archive of historical PixPin versions.
Invoice OCR POC
Fast embeddable Rust CSS layout engine for document rendering without a browser.
Tika-Python is a Python binding to the Apache Tika™ REST services allowing Tika to be called natively in the Python community.
The media player for language learning, with dual subtitles, AI-generated subtitles, real-time translation, and more!
A list of open-source AI projects you can use to generate income easily.
Free, private PDF, image, audio, Word, Excel & PowerPoint tools that run 100% in your browser. Compress, convert, edit, redact, OCR, cut MP3, Excel↔JSON, Mermaid to PNG/SVG. No uploads, no sign-up, no watermark. Free forever · 64 tools · 16 languages.
Beautiful pdf components, built on Takumi and Forme. 100% Free, Zero config, one command setup.
Tianshu - enterprise-grade, all-in-one AI data preprocessing platform | PDF/Office to Markdown | MCP protocol support for AI assistant integration | Vue3+FastAPI full-stack solution | document parsing | multimodal information extraction
An offline desktop application based on the WeChat OCR engine that extracts text from images with high accuracy. It provides an intuitive interface and supports various OCR tasks, including screenshots, image rotation, and convenient batch processing.
A simple, elegant dictionary and translation macOS App. Works out of the box; supports offline OCR, Youdao Dictionary, 🍎 Apple system dictionary, 🍎 Apple system translation, OpenAI, Gemini, DeepL, Google, Bing, Tencent, Baidu, Alibaba, Xiaoniu, Caiyun, and Volcano translation. A concise and elegant Dictionary and Translator macOS App for looking up words and translating text.
docling-pipelines
Free downloads of English magazines including The Economist (with audio), The New Yorker, The Guardian, WIRED, and The Atlantic; supports EPUB, MOBI, and PDF formats; updated weekly
DDL Manager v4: Windows tasks, calendar, desktop floating window, local OCR, and cross-computer schedule sync; includes an offline installer with built-in WebView2.
AI comic and manga translator app/browser extension for automatically translating comics, manga, manhwa, BDs, fumetti, and more in multiple languages and formats (Images, PDF, EPUB, CBR, CBZ etc).
The Privacy First PDF Toolkit
Windows screenshot, annotation, offline OCR, clipboard history, GIF and screen recording.
#1 PDF Application on GitHub that lets you edit PDFs on any device anywhere
The first vision plugin for DeepSeek Harness, and the vision bridge for every text-only coding agent. Paste an image, get structured JSON evidence (OCR, layout, semantics). | The most powerful DeepSeek Harness vision plugin, adding vision capabilities to text-only models such as DeepSeek and GLM; paste an image to get structured JSON evidence (OCR, layout, semantics).
Python bindings to PDFium, reasonably cross-platform.
RMT (RuoMengTu) is a free, open-source macro tool built on AHKv2. Let the code handle the tedious work—you have more meaningful things to do.
Local-first Markdown resume builder with live A4 preview, ATS checks, bilingual editing, Auto Fit, and PDF export.
Document scanning app
Daidai Didi general-purpose CAPTCHA recognition OCR, PyPI version
Convert documents to structured data effortlessly. Unstructured is open-source ETL solution for transforming complex documents into clean, structured formats for language models. Visit our website to learn more about our enterprise grade Platform product for production grade workflows, partitioning,
extract text from any document. no muss. no fuss.
Readest is a modern, feature-rich ebook reader designed for avid readers offering seamless cross-platform access, powerful tools, and an intuitive interface to elevate your reading experience.
Open-source Android PDF and document toolkit with PDF tools, Office viewing, XLSX editing, scanning, and image editing. Built with Kotlin and Jetpack Compose.
Transforms complex documents like PDFs and Office docs into LLM-ready markdown/JSON for your Agentic workflows.
A ready-to-use translation and OCR tool developed with WPF.
KCC (a.k.a. Kindle Comic Converter) is a comic and manga converter for ebook readers.
Get your documents ready for gen AI
An automated Electronic Know-Your-Customer (E-KYC) system designed to verify user identities securely. This project combines Optical Character Recognition (OCR) for document text extraction and Biometric Facial Recognition with Liveness Detection to prevent identity fraud.
PDF tooling for Go and the command line.