Open weights
Cohere Labs releases a 2.4B vision model built to read documents at native resolution
North Micro Vision skips the usual downsizing step, processing pages up to A4 size at 200dpi and scoring 0.921 on DocVQA.
Original cover art, generated for this story. THE VISSION does not republish third-party press imagery.
The short version
- Cohere Labs released North Micro Vision, a 2.4-billion-parameter open-weight vision-language model under Apache 2.0.
- It combines a 400M native-resolution vision encoder with a 2B language model, processing images up to A4 dimensions at 200dpi without resizing.
- The model is aimed at document understanding, chart and table interpretation, OCR and visual grounding, and scores 0.921 on the DocVQA benchmark.
- It was trained through a four-stage curriculum emphasizing OCR and document understanding, with community MLX and Nvidia AutoModel fine-tuning support already available.
North Micro Vision skips the usual downsizing step, processing pages up to A4 size at 200dpi and scoring 0.921 on DocVQA.