</>MCP Agents Market
Skill

tika

by apache4kJavaUpdated 2026-08-30

The Apache Tika toolkit detects and extracts metadata and text from over a thousand different file types (such as PPT, XLS, and PDF).

Apache Tika is a content analysis toolkit that detects and extracts metadata and structured text from over a thousand file formats including PDF, Microsoft Office documents, and images. Version 4.x is optimized for AI agent workflows, outputting Markdown by default and running parsers in isolated processes to protect against malicious documents. The repository provides ready-to-use agent skills (file-to-markdown and file-to-markdown-docker) that enable AI agents to convert diverse document types into LLM-ready formats. Developers can drop these skills into any agent's skill directory to enable document parsing capabilities without additional dependencies.

Key Features

Supports extraction from over 1,000 file types including PDF, PPT, XLS, DOCX, images, and archives
Outputs Markdown by default (since 4.0) optimized for LLM and RAG pipelines
Crash-isolated forked processes prevent hostile documents from taking down the main service
Vision-language-model parsers (Claude, Gemini, OpenAI) for documents OCR cannot read
Standalone agent skills (file-to-markdown and file-to-markdown-docker) require no repository dependency
Structured recursive extraction (-J flag) returns metadata and content for embedded documents as JSON
Containerized Docker skill with guaranteed OCR support for maximum compatibility
Apache 2.0 licensed open-source project with active development and reproducible builds

Use Cases

  • 01Enabling AI agents to read and extract content from PDF reports, presentations, and spreadsheets
  • 02Building RAG (Retrieval-Augmented Generation) pipelines that ingest documents in multiple formats
  • 03Processing scanned documents and images using vision-language models when traditional OCR fails
  • 04Extracting metadata and text from email attachments for automated analysis and classification
  • 05Converting legacy document formats into Markdown for knowledge base indexing
  • 06Parsing embedded files within archives or compound documents for comprehensive content extraction

Related Skills

View more

tika — FAQ

What is Apache Tika agent skill?+

Apache Tika agent skill provides document parsing capabilities for AI agents, converting over 1,000 file types (PDF, Office documents, images) into Markdown or structured JSON. The skills are standalone modules in the .skills/ directory that can be copied into any agent's skill directory.

How do I install the Tika agent skill?+

Copy the file-to-markdown or file-to-markdown-docker skill folder from the .skills/ directory in the Tika repository into your agent's skill directory. The skills are self-contained and do not require the full Tika repository to function.

What are the prerequisites for using Tika agent skills?+

The file-to-markdown skill requires Java 17 and the tika-app jar or a running tika-server instance. The file-to-markdown-docker skill requires Docker to be installed and provides guaranteed OCR support through containerization.

Is Apache Tika free to use?+

Yes, Apache Tika is open-source software licensed under the Apache License 2.0, which permits free use for both commercial and non-commercial purposes.

Which AI agent clients work with Tika skills?+

Tika agent skills are designed as portable skill modules that can be integrated into any agent framework that supports skill directories. The specific client compatibility depends on where you deploy the skill.

Does Tika require API keys for document parsing?+

Basic document parsing does not require API keys. However, the vision-language-model parsers for advanced OCR (Claude, Gemini, OpenAI) require API keys from the respective providers to access their vision capabilities.

How do I install tika?+

Open the source repository on GitHub and follow its README. tika is a skill — MCP Agents Market links you directly to the official repo.

Is tika free?+

tika is an open-source project hosted on GitHub. Check the repository for its license and any usage requirements.

Related searches