Verified project record
opendatalab/MinerU
MinerU converts PDF, image, DOCX, PPTX, and XLSX files into Markdown and JSON for retrieval and extraction workflows. It runs locally across multiple backends via CLI or WebUI and requires at least 16GB of RAM, with optional GPU acceleration depending on the chosen parsing method.
Project overview
MinerU converts PDF, image, DOCX, PPTX, and XLSX files into Markdown and JSON for retrieval and extraction workflows. It runs locally across multiple backends via CLI or WebUI and requires at least 16GB of RAM, with optional GPU acceleration depending on the chosen parsing method.
- Project type
- RAG · Document Processing · Data Processing
- Use cases
- Documents & Office · Knowledge Q&A
- Deployment
- Refer to project documentation
- License
- License pending
Best for
- Developers, data teams, and AI engineers building retrieval, extraction, and processing pipelines for text and image documents.
- Users requiring local orchestration and multi-service deployment through a CLI, FastAPI, or Gradio WebUI.
Key capabilities
- Supports PDF, image, DOCX, PPTX, and XLSX inputs for document conversion.
- Removes headers, footers, footnotes, page numbers, and outputs text in human-readable order for single-column, multi-column, and complex layouts.
- Automatically recognizes and converts formulas to LaTeX and tables to HTML format.
- Automatically detects scanned PDFs and garbled PDFs to enable OCR functionality supporting detection and recognition of 109 languages.
- Extracts images, image descriptions, tables, table titles, and footnotes from documents.
- Provides built-in CLI, FastAPI, and Gradio WebUI for local orchestration and multi-service deployment.
- Preserves the structure of the original document, including headings, paragraphs, and lists.
Limitations and risks
- Document parsing is a difficult and complex task. In scenarios such as complex layouts, scanned pages, and handwritten content, the parsing results may fall short of expectations.
- Docker deployment is only supported on Linux and Windows environments with WSL2 support; macOS users should not use Docker deployment.
- In non-mainline environments, due to the diversity of hardware and software configurations, as well as third-party dependency compatibility issues, 100% project availability cannot be guaranteed.
- Since the key dependency `ray` does not support Python 3.13 on Windows, only versions 3.10~3.12 are supported.
Getting started
- Setup requires command line interface proficiency. Users can start by upgrading pip, installing uv, installing the application package, and running the command-line interface with an input path and output path.
- Users must navigate multiple backend choices and hardware compatibility constraints. The setup requires selecting a parsing backend (pipeline, hybrid, vlm, http-client) based on available CPU or GPU resources.
Evidence and sources
- GitHub project description: Transforms complex documents like PDFs and Office docs into LLM-ready markdown/JSON for your Agentic workflows.
- README: MinerU is a document parsing tool that converts `PDF`, image, `DOCX`, `PPTX`, and `XLSX` inputs into machine-readable formats such as Markdown and JSON for downstream retrieval, e…
- README: 2026/06/18 3.4 Released
- README: The OCR model for the `pipeline` backend has been upgraded to `PP-OCRv6`, improving OCR accuracy by about `11%` on OmniDocBench v1.6.
- README: VLM model upgraded to `MinerU2.5-Pro-2605-1.2B`
AI 搜索
把需求说清楚,让项目选择更有依据
告诉我们你要解决什么、运行在哪里、哪些条件不能妥协。雷达会从已核验项目中给出主推荐、备选和采用前检查。
目标你最终想完成什么
环境本地、云端或现有技术栈
硬条件部署、界面、语言与 License
从一个真实需求开始点击只会填入搜索框,你可以继续修改
今日榜单
0