Context Optical Compression 1
Next-gen document intelligence with context optical compression and multilingual support.
DeepSeek OCR is a two-stage transformer-based document AI system that utilizes context optical compression to deliver state-of-the-art document intelligence. It compresses high-resolution documents into lean vision tokens, then decodes them with a 3B-parameter mixture-of-experts model to achieve near-lossless text, layout, and diagram understanding across 100+ languages. It supports GPU-efficient throughput for complex layouts and is trained on 30 million real PDF pages plus synthetic data, preserving layout structure, tables, chemistry (SMILES strings), and geometry tasks.