tika topic
tika
The Apache Tika toolkit detects and extracts metadata and text from over a thousand different file types (such as PPT, XLS, and PDF).
memex-explorer
Viewers for statistics and dashboarding of Domain Search Engine data
sparkler
Spark-Crawler: Apache Nutch-like crawler that runs on Apache Spark.
extract
A cross-platform command line tool for parallelised content extraction and analysis.
tikaondotnet
Use the Java Tika text extraction library on the .NET platform
fscrawler
Elasticsearch File System Crawler (FS Crawler)
MLwithTensorFlow2ed
Code for Machine Learning with TensorFlow: 2nd Edition Published by Manning Publications
imagecat
ImageCat is an Apache OODT RADIX application that uses Apache Solr, Apache Tika and Apache OODT to ingest 10s of millions of files (images,but could be extended to other files) in place, and to extrac...
php-apache-tika
Apache Tika bindings for PHP: extract text and metadata from documents, images and other formats