Targeted Training Kept DharmaOCR Ahead on Brazilian Portuguese
DharmaOCR says its advantage over Mistral OCR4 and Unlimited-OCR came from Portuguese-specific fine-tuning and Direct Preference Optimization.
DharmaOCR says it outperformed Mistral OCR4 and Unlimited-OCR on Brazilian Portuguese OCR through domain specialization and targeted training rather than newer architecture alone.
The model was developed specifically for Brazilian Portuguese. Its training pipeline used two stages: supervised fine-tuning on Portuguese-language files, followed by Direct Preference Optimization, or DPO.
In the first stage, DharmaOCR was fine-tuned on Portuguese-language documents drawn from different sources, formats, and levels of complexity. The company said this aligned the model with Brazilian Portuguese vocabulary, syntax, and document structures, concentrating its capacity on the target language rather than a broader multilingual setting.
The second stage used comparative preference data between competing outputs. Rather than learning only from correct transcriptions, the model was trained to choose the better extraction among alternatives.
DharmaOCR said DPO was intended to improve stability by reducing failure modes that can produce repetitive or incoherent output. According to the company, suppressing those failures reduced inference time and cost while improving production reliability.
The available material does not provide benchmark methodology, exact scores, document composition, inference settings, or independent validation of the comparison.
- Newer Models, Same Advantagehuggingface.co / Release / Published JUL 16, 2026 / Accessed JUL 22, 2026