Project Name | Stars | Downloads | Repos Using This | Packages Using This | Most Recent Commit | Total Releases | Latest Release | Open Issues | License | Language |
---|---|---|---|---|---|---|---|---|---|---|
Node Tika | 128 | 15 | 5 | 4 years ago | 23 | February 22, 2017 | 10 | mit | Java | |
Apache Tika bridge for Node.js. Text and metadata extraction, language detection and more. | ||||||||||
Php Apache Tika | 104 | 3 | 3 | 8 months ago | 38 | April 14, 2023 | mit | PHP | ||
Apache Tika bindings for PHP: extract text and metadata from documents, images and other formats | ||||||||||
Imagecat | 84 | 6 years ago | Java | |||||||
ImageCat is an Apache OODT RADIX application that uses Apache Solr, Apache Tika and Apache OODT to ingest 10s of millions of files (images,but could be extended to other files) in place, and to extract metadata and OCR information from those files/images using Tika and Tesseract OCR. | ||||||||||
Harvester | 59 | 7 years ago | 3 | gpl-3.0 | JavaScript | |||||
Web crawling and document processing through a usable interface. | ||||||||||
Rtika | 52 | 1 | a year ago | 8 | April 25, 2020 | 3 | apache-2.0 | R | ||
R Interface to Apache Tika | ||||||||||
Doc_processing_toolkit | 52 | 7 years ago | 4 | other | Python | |||||
Python library to extract text from PDF, and default to OCR when text extraction fails. | ||||||||||
Cogstack Pipeline | 39 | a year ago | other | Java | ||||||
Distributed, fault tolerant batch processing for Natural Language Applications and Search, using remote partitioning | ||||||||||
Pdf Discovery Demo | 24 | a year ago | 2 | apache-2.0 | JavaScript | |||||
Demonstration of searching PDF document with Solr, Tika, and Tesseract | ||||||||||
Tika Server | 18 | 3 years ago | 2 | apache-2.0 | Java | |||||
Apache Tika Server with Tesseract 4 Docker Setup | ||||||||||
Tika Service | 12 | a year ago | apache-2.0 | Java | ||||||
Apache Tika running as a web service |