Gemma 4 12B is live! 🚀 An encoder-free multimodal model (text/img/audio) for local 16GB laptops. Elite reasoning nearing 26B MoE in half the size, fast, and open (Apache 2.0). This is the main reason I was not posting much!! Glad it is launched!! blog.google/innovation-a...
Introducing Gemma 4 12B: a unified, encoder-free multimodal model blog.google
AI Weekly's analysis
→
- The 35M-parameter vision embedder replaces 27 vision transformer layers, keeping the full model inside 16GB with complete image and audio understanding.
- Audio projects from raw 16 kHz waveforms in 40ms frames directly to the LLM backbone, bypassing any separate ASR encoder used in competing designs.
- Single-pass LoRA fine-tuning updates vision, audio, and text weights simultaneously, eliminating the engineering overhead of co-tuning frozen encoders.
Read full analysis →
View on Bluesky ·
♥ 111
↻ 16
↩ 8
·
2 from the directory shared this ·
75d ago
This is a great use of Gemma! having an open model running at +1000 tokens per second can enable some pretty cool use cases! The voice assistant is a good one, but I'm sure there are many others! huggingface.co/blog/cerebra...
Hugging Face and Cerebras bring Gemma 4 to real-time voice AI huggingface.co
AI Weekly's analysis
→
- Hugging Face and Cerebras have shipped an open cascaded speech-to-speech pipeline chaining Nvidia's Parakeet, Google DeepMind's Gemma 4 VLM on Cerebras, and Alibaba's Qwen3TTS.
- The pitch focuses on P95 tail latency stability, not median speed, arguing that occasional multi-second stalls are what break conversational voice apps.
- Hugging Face says the same pipeline already powers more than 9,000 Reachy Mini robots in the wild, giving the demo a real deployment story.
Read full analysis →
View on Bluesky ·
♥ 25
↻ 3
↩ 4
·
2 from the directory shared this ·
47d ago
great post about AI trends specifically mentioning the Gemma 4 success!! 🤩 www.interconnects.ai/p/some-ideas...
Some ideas for what comes next, May 2026 interconnects.ai
👁️ Vision Options: Want to make Gemma see even better? The default vision bucket is 280 for token efficiency. To capture maximum detail (like sharp OCR and 2.51MP resolution), manually bump max_soft_tokens to 1120! Try our new interactive Space to see how it works: huggingface…
Gemma 4 - Vision Token Budget - a Hugging Face Space by google huggingface.co
AI Weekly's analysis
→
- The Space lets users toggle Gemma 4's per-image budget across five preset sizes: 70, 140, 280, 560, and 1120 tokens.
- Gemma 4 launched April 2, 2026 under Apache 2.0 across E2B, E4B, 12B Unified, 26B A4B MoE, and 31B dense sizes.
- The 31B model reportedly hits 76.9% on MMMU Pro and 85.6% on MATH-Vision, with native JSON bounding box output.
Read full analysis →
↻
Gus reposted
Dare Obasanjo
@carnage4life.bsky.social
Gemini now has 1B monthly users, their 14th app to do so, and hit that milestone faster than any other app in its history. The impressive part is this is just counting users of the Gemini app or website, not the integration into their various Workspace apps.
Gemini becomes Google's fastest-growing product ever as it hits 1B users arstechnica.com
AI Weekly's analysis
→
Read full analysis →
View on Bluesky →
I've been playing all day with Hugging Face Chat + Gemma 4 31B (deployed on Cerebras at +1000 tokens per second) and I'm still surprised how amazing it is!! I play with Gemma models basically everyday and this integration + speed still surprised me! take a look: huggingface.co…
google/gemma-4-31B-it - HuggingChat huggingface.co
🫂 A huge shoutout to the community for submitting fixes and finding new ways to make Gemma even better. We couldn't do this without you! ❤️ 🤗 Ready to test the speedup? Download the latest Gemma 4 updates now on Hugging Face: huggingface.co/collections/...
Gemma 4 - a google Collection huggingface.co
This was a pretty good week for Gemma + comunity integrations!!! here is one example: LiveKit: livekit.com/products/inf... conversational assistant using Gemma 31B as the brains. Super fast and great experience. Try their demo on their site!
Gemma 4 31B on LiveKit Inference | A faster, cheaper default for voice livekit.com