North Micro Vision is now free to deploy under Apache 2.0. Let us know what you get up to. Download the model weights, learn more about the model, its architecture, and more: https://t.co/Ho58BvLLUo
Cohere
Tracked through public AI activity and peer connections inside the directory.
- AI signals
- 8 past 30d
- Sources
- 4 distinct domains
- Discussions
- 0 past 30d
- Latest signal
- 9d ago
Articles & links
The model punches above its weight, outperforming Gemma 4 E2B and Ministral 3 3B across a broad range of visual understanding benchmarks and delivers particularly strong results on document understanding and visual Q&A. See the full results: https://t.co/ir2oMAbC6z https:/…
@nvidia Confidential Computing is now available in Model Vault, and we are opening early access to a limited group of beta customers. Beta places are limited. Learn more here: https://t.co/bCt48aZChb
Who Gets to Define the Rules for AI? Artificial intelligence needs evidenced standards, not a cartel. A perspective from @aidangomez, co-founder & CEO of Cohere: https://t.co/QnDmIfsLPM
- Cohere CEO Aidan Gomez published a September 13 essay calling a rival-lab AI standards proposal 'a cartel by any other name.'
- Gomez opposes the 'narrow antitrust waiver' being requested so a handful of Silicon Valley labs can jointly set safety standards.
- He counter-proposes four pillars: evidence-based risk framework, mandatory transparency, capability-scoped independent testing, and conflict-free assurance mechanisms.
Reclaim your time for what truly moves you. Your freedom, your focus. We're here to support you: https://t.co/fXxZdfXo21 https://t.co/2MlxjJXaCZ
@aidangomez For more information, please visit: https://t.co/FwkkZg1hXz
We were founded on advancements in translation. We couldn’t be prouder to continue that legacy. Find the weights on Hugging Face (available in several quants) under a CC BY-NC 4.0 license. https://t.co/sfwJrlhdMS
Don’t know what a megakernel is? Don’t worry. Find out what it is - and how we built the system - on our blog: https://t.co/HF93hrlynu
- Cohere's single-CUDA-file 'megakernel' runs its 30B/3.3B-active North Mini Code model at 292 tokens/second on one H100, 1.58× vLLM's rate.
- End-to-end speedups of 1.25× to 1.41× held across AIME 2025, GPQA, MMLU-Pro CS, SciCode, and LiveCodeBench v6, with matching accuracy.
- The engine caps batch size at 8, runs decode-only, and ships publicly on GitHub with a Blackwell port and FP8/FP4 support next on the roadmap.
A megakernel fuses the entire LLM decode step into a single kernel launch. We build on this by maximizing GPU utilization through kernel fusion while supporting everything a real server needs. Check out how we got there on GitHub: https://t.co/f0P7dkv89S
- Cohere released cohere-megakernel, a serving engine that runs the entire decode forward pass as one persistent CUDA kernel on a single H100.
- The kernel hits 292 tokens/second at batch size 1 in BF16, or 62% of H100 speed-of-light, and 1.58× vLLM's decode throughput.
- End-to-end speedups over vLLM range 1.25× to 1.41× across five benchmarks; batch sizes 1-8 only, decode-only, and model-locked to North Mini Code.
Full article: https://t.co/6oktuoyuvv
“Canada has one of the strongest AI talent pipelines in the world, and the University of Waterloo has played a defining role in building it" - @jpineau1 https://t.co/X7ikKrO179
Are you Cohere? Show it.
Add the Who’s Who of AI badge to your site or bio. It links back to this profile.
Markdown: [](https://aiweekly.co/whos-who/person/cohere)