huggingface.co web signal

Reka ships Edge 2603, a 7B VLM at 331 tokens per 1024px image

TL;DR

  • Reka Edge 2603 is a 7B multimodal model that encodes a 1024x1024 image in 331 input tokens, versus 1063 for Cosmos-Reason2 8B and 1094 for Gemini 3 Pro.
  • The model card lists scores of 88.40 on VQA-V2, 74.30 on MLVU video understanding and 93.13 on RefCOCO-A object detection.
  • Weights are open under a custom license that permits commercial use only for companies with annual revenue under $1 million USD.

Reka's new Edge 2603 model card says the 7B multimodal model encodes a 1024x1024 image in 331 input tokens, against 1063 for Cosmos-Reason2 8B, 1041 for Qwen 3.5 9B and 1094 for Gemini 3 Pro on the same card. On benchmarks Reka published alongside the release, Edge 2603 scores 88.40 on VQA-V2, 74.30 on MLVU video understanding, and 93.13 on RefCOCO-A object detection, with end-to-end latency of 4.69 seconds versus 16.67 for Gemini 3 Pro.

The pitch is physical AI. Reka's announcement frames the release with the line that "frontier intelligence should be fast, lean, and deployable anywhere," and lists Jetson Thor, Jetson AGX Orin and Apple Silicon Macs with 24GB of memory as the target hardware, with quantized builds extending to Jetson Orin Nano, Samsung S25 and iPhone. Reka says Edge 2603 "processes 5.46 images per second," and that a 4-bit quantization "cuts memory consumption significantly, from 13GB to just 5GB," retaining over 98% of original performance.

The catch is in the license. Edge 2603 ships under a custom Reka license that permits commercial use only for companies with annual revenue below $1 million USD; anything larger needs a separate deal the model card does not price. The Hugging Face page shows 688 downloads in the last month.

It lands into a crowded week of open-weight multimodal drops, including Cloudflare's Clef-omni and OneSearch-VL 8B, two of 77 multimodal stories we've tracked over the past 90 days.