Jina modifies DeepSeek-OCR; in self-tests, it ranks first in page throughput among 14 options.
2026-09-19 11:49:17
According to CoinMeta, Jina has released a document parsing model called jina-ocr-v1. This model is capable of directly converting PDF, scanned documents, tables, and charts into Markdown. It is trained further based on deepseek-ocr and utilizes approximately 3.4 billion total parameters. During decoding, it activates around 570 million parameters per token using a MOE architecture, incorporating fastmtp for inference-based decoding. In self-tests, jina-ocr-v1 achieved a score of 91.14 on omnidocbench v1.6, which is 0.89 points higher than that of deepseek-ocr-2. On olmocr-bench, it scored 83.4, 7.4 points higher than its base deepseek-ocr model. Throughput is one of the main selling points of this model; with a single A100 and 32 concurrent tasks, it can process 2.57 pages per second, making it the highest among the 14 systems tested in Jina, and about 22% higher than deepseek-ocr's rate of 2.10 pages per second. The model weights have been uploaded to Hugging Face, and it uses a CC BY - NC license version 4.0. For commercial use, contact Jina.
Bullish 0
Bearish 0
Source:Internet
This content is for market information only and does not constitute investment advice.