tika
There are 142 repositories under tika topic.
apache/tika
The Apache Tika toolkit detects and extracts metadata and text from over a thousand different file types (such as PPT, XLS, and PDF).
dadoonet/fscrawler
Elasticsearch File System Crawler (FS Crawler)
USCDataScience/sparkler
Spark-Crawler: Apache Nutch-like crawler that runs on Apache Spark.
ICIJ/extract
A cross-platform command line tool for parallelised content extraction and analysis.
KevM/tikaondotnet
Use the Java Tika text extraction library on the .NET platform
shebinleo/pdf2html
pdf2html is a module which helps to convert PDF file to HTML pages using Apache Tika. This module also helps to generate thumbnail image for PDF file using Apache PDFBox.
chrismattmann/MLwithTensorFlow2ed
Code for Machine Learning with TensorFlow: 2nd Edition Published by Manning Publications
nasa-jpl-memex/memex-explorer
Viewers for statistics and dashboarding of Domain Search Engine data
vaites/php-apache-tika
Apache Tika bindings for PHP: extract text and metadata from documents, images and other formats
apache/tika-docker
Convenience Docker images for Apache Tika Server
chrismattmann/tika-similarity
Tika-Similarity uses the Tika-Python package (Python port of Apache Tika) to compute file similarity based on Metadata features.
chrismattmann/imagecat
ImageCat is an Apache OODT RADIX application that uses Apache Solr, Apache Tika and Apache OODT to ingest 10s of millions of files (images,but could be extended to other files) in place, and to extract metadata and OCR information from those files/images using Tika and Tesseract OCR.
nasa-jpl-memex/image_space
Interactive Image similarity and Visual Search and Retrieval application
Sotera/newman
Quickly analyze and explore email with advanced analytics and visualization.
ropensci/rtika
R Interface to Apache Tika
nasa-jpl-memex/GeoParser
Extract and Visualize location from any file
OpenSextant/Xponents
Geographic Place, Date/time, and Pattern entity extraction toolkit along with text extraction from unstructured data and GIS outputters.
CogStack/CogStack-Pipeline
Distributed, fault tolerant batch processing for Natural Language Applications and Search, using remote partitioning
tspannhw/nifi-extracttext-processor
Apache NiFi Custom Processor Extracting Text From Files with Apache Tika
ipfs-search/ipfs-tika
Java web application taking IPFS hashes, extracting (textual) content and metadata through Apache's Tika.
sergio11/document_search_engine_architecture
📄🚀 Unleash a powerful Document Search Engine with Apache NiFi for lightning-fast, comprehensive text indexing and search.
mrcsparker/ruby_tika_app
A ruby wrapper for the Tika jar (tika-app.jar) that extracts text in a lot of formats from PDF, xls, doc, etc files
sergeyt/pandora
Small box of pandora to prototype your app with ready for use backend. This is just my compilation of different solutions occasionally applied in hackathons and challenges
apache/tika-helm
A Helm chart to deploy Apache Tika on Kubernetes.
fedelemantuano/tika-app-python
Python bindings for Apache Tika
USCDataScience/tika-dockers
A suite of Machine Learning / Deep Learning Dockerfiles to allow Apache Tika to extract objects and to produce textual captions for images and video
rse/tika-server
Apache Tika Server as a Background Service in Node.js
public-law/oregon-law-parser
Distill information about amendments to the Oregon Revised Statutes.
shelfio/tika-text-extract
Extract text from a document by Apache Tika
catalyst/moodle-search_elastic
An Elasticsearch engine plugin for Moodle's Global Search
chrismattmann/trec-dd-polar
A dataset downloaded from the deep and scientific web across three major Polar data centers for use in research.
Keerthivasan13/CSCI572-Information_Retrieval_And_Web_Search_Engines
Search Engine projects
lagenorhynque/tika
git diff settings for Microsoft Office files
privet56/qDesktopSearch
qDesktopSearch - a Qt5 Desktop App for indexing & searching the files of the local machine
quarkiverse/quarkus-tika
Quarkus Tika extension
hmmh/typo3-solr-file-indexer
TYPO3 Extension: solr_file_indexer