</>MCP Agents Market
Skill

liteparse

by run-llama12.2kRustUpdated 2026-08-21

A fast, helpful, and open-source document parser

LiteParse is an open-source document parsing agent skill that extracts text, tables, and structured data from PDFs and other document formats. Built in Rust with bindings for Node.js, Python, and WebAssembly, it provides fast local parsing with optional OCR capabilities using Tesseract or custom HTTP servers. The tool outputs clean Markdown, JSON with bounding boxes, or plain text—ideal for feeding documents into LLM agents and RAG pipelines. It supports automatic format conversion for Office documents (DOCX, XLSX, PPTX) and images, and includes complexity detection to route documents before heavy processing.

Key Features

Fast spatial text extraction from PDFs using PDFium with precise bounding box coordinates
Markdown output with reconstructed headings, tables, lists, images, and links for LLM consumption
Flexible OCR system with bundled Tesseract or pluggable HTTP servers (EasyOCR, PaddleOCR)
Automatic conversion of Office documents (DOCX, XLSX, PPTX) and images to PDF before parsing
Complexity detection for routing documents and estimating processing costs before full parse
Screenshot generation with form field rendering and solid-rectangle detection
Multi-language support (Rust, Node.js/TypeScript, Python, WebAssembly) with unified CLI
Optional extraction of vector graphics, structure trees, layout blocks, and XFA packets

Use Cases

  • 01Preparing documents for RAG pipelines by converting PDFs to clean Markdown with preserved structure
  • 02Extracting tabular data from invoices, reports, and spreadsheets with bounding box metadata
  • 03Routing complex documents to cloud services (LlamaParse) and simple ones to local parsing
  • 04Generating page screenshots for vision-enabled LLM agents to analyze visual elements
  • 05Processing scanned documents with OCR while preserving spatial layout information
  • 06Batch converting entire directories of mixed-format documents to structured JSON or text

Related Skills

View more

liteparse — FAQ

What is LiteParse?+

LiteParse is an open-source document parser that extracts text, tables, and metadata from PDFs and other formats (DOCX, XLSX, images) with optional OCR. It runs locally and outputs Markdown, JSON with bounding boxes, or plain text suitable for AI agents and RAG systems.

How do I install LiteParse as an agent skill?+

Use the skills CLI with 'npx skills add run-llama/llamaparse-agent-skills --skill liteparse' or manually copy the SKILL.md file to your agent's skills directory. You can also install it as a standalone CLI via npm, pip, or cargo.

Which AI clients work with LiteParse?+

LiteParse works as an agent skill with any client supporting the skills pattern, and can be called from Node.js, Python, Rust, or browser environments via its library APIs. The CLI tool runs independently on Linux, macOS, and Windows.

Does LiteParse require API keys or external services?+

No, LiteParse runs entirely locally with bundled Tesseract OCR. LibreOffice is optional for Office document conversion, and you can optionally connect custom HTTP OCR servers for better accuracy.

Is LiteParse free to use?+

Yes, LiteParse is open-source under the Apache 2.0 license and completely free. For complex documents, the developers recommend their paid cloud service LlamaParse, but LiteParse itself has no usage limits.

What file formats does LiteParse support?+

LiteParse natively parses PDFs and automatically converts DOCX, XLSX, PPTX, ODT, RTF, and common image formats (JPG, PNG, SVG, TIFF) via LibreOffice or built-in image handling. All formats are converted to PDF internally before extraction.

How do I install liteparse?+

Open the source repository on GitHub and follow its README. liteparse is a skill — MCP Agents Market links you directly to the official repo.

Is liteparse free?+

liteparse is an open-source project hosted on GitHub. Check the repository for its license and any usage requirements.

Related searches