NaviDC-OCR — document parsing, digital and camera-captured

A 1.2B document-parsing VLM that reads layout, text, tables, formulas and code off flat scans and photographed / crumpled pages, and returns Markdown.

model · paper · code

Parsing mode
256 4096
Examples from the NaviDC-OCR model card
Document page Parsing mode Region type