MonkeyOCRv2 is a visual-text foundation model for Document AI that provides document-native vision encoders, parsing models, and understanding models, achieving top performance on multilingual document parsing and recognition benchmarks.

Stars

336

7-day growth

No data

Forks

39

Open issues

0

License

NOASSERTION

Last updated

2026-07-28

AI repository intelligence
FR-AI / ANALYSIS

Why it is worth attention

It ranks #1 among open-source models on the MDPBench multilingual document parsing leaderboard with 83.3 overall, and releases the largest open-source document pre-training dataset (MonkeyDoc v2) with 113M images across 17 languages, while offering a family of efficient models (S/B/AS) and vLLM integration with DFlash for 2× faster inference.

Who it is for

  • Document AI researchers
  • OCR system developers
  • Multilingual NLP practitioners
  • Computer vision engineers working on text-rich images

Use cases

  • Multilingual document parsing (PDFs, scanned docs) into structured text
  • Document visual question answering (DocVQA, InfoVQA)
  • Text recognition and detection in scenes and documents
  • Document tampering detection and overlapping text segmentation

Strengths

  • Top open-source model on MDPBench (83.3 overall, 17 languages)
  • Open-sourced largest multilingual document pretraining dataset (MonkeyDoc v2, 113M images)
  • Multiple model sizes (S, B, AS) for different tasks with competitive parameter counts
  • Supports vLLM serving with DFlash acceleration for up to 2× speedup

Considerations

  • MonkeyDoc v2 dataset requires ~10 TB disk space to download
  • DFlash acceleration requires vLLM 0.25.1 and CUDA 12.9+, limiting hardware compatibility
  • Model performance on very low-resource languages may vary; evaluation mainly covers 17 languages

README quick start

Quick Start

Description

MonkeyOCRv2 Vision Encoder — A Document-Native Visual Backbone

Related repositories

Similar projects matched by category, topics, and programming language.

lopopolo
Featured
lopopolo GitHub avatar

harness-engineering

Harness Engineering is a methodology for improving coding agent outputs by carefully crafting the environment around them—providing curated context, tools, and executable constraints that encode an organization’s nonfunctional requirements and cumulative lessons.

AI & Machine LearningAI Agents
2,390
slvDev
Featured
slvDev GitHub avatar

esp32-ai

A 28.9 million parameter language model runs on an $8 ESP32-S3 microcontroller entirely on-device, generating simple stories at about 9.5 tokens per second.

AI & Machine LearningLarge Language Models
1,960
littledivy
Featured
littledivy GitHub avatar

mimic

mimic captures traffic from any iOS or web app and automatically generates a Python client library that lets you call the app's API like a regular library.

AI & Machine Learning
1,482