Gemma 4 12B is live! 🚀 An encoder-free multimodal model (text/img/audio) for local 16GB laptops. Elite reasoning nearing 26B MoE in half the size, fast, and open (Apache 2.0). This is the main reason I was not posting much!! Glad it is launched!! blog.google/innovation-a...
Introducing Gemma 4 12B: a unified, encoder-free multimodal model blog.google
AI Weekly's analysis
→
- The 35M-parameter vision embedder replaces 27 vision transformer layers, keeping the full model inside 16GB with complete image and audio understanding.
- Audio projects from raw 16 kHz waveforms in 40ms frames directly to the LLM backbone, bypassing any separate ASR encoder used in competing designs.
- Single-pass LoRA fine-tuning updates vision, audio, and text weights simultaneously, eliminating the engineering overhead of co-tuning frozen encoders.
Read full analysis →
View on Bluesky ·
♥ 111
↻ 16
↩ 8
·
2 from the directory shared this ·
55d ago
This is a great use of Gemma! having an open model running at +1000 tokens per second can enable some pretty cool use cases! The voice assistant is a good one, but I'm sure there are many others! huggingface.co/blog/cerebra...
Hugging Face and Cerebras bring Gemma 4 to real-time voice AI huggingface.co
AI Weekly's analysis
→
- Hugging Face and Cerebras have shipped an open cascaded speech-to-speech pipeline chaining Nvidia's Parakeet, Google DeepMind's Gemma 4 VLM on Cerebras, and Alibaba's Qwen3TTS.
- The pitch focuses on P95 tail latency stability, not median speed, arguing that occasional multi-second stalls are what break conversational voice apps.
- Hugging Face says the same pipeline already powers more than 9,000 Reachy Mini robots in the wild, giving the demo a real deployment story.
Read full analysis →
View on Bluesky ·
♥ 25
↻ 3
↩ 4
·
2 from the directory shared this ·
27d ago
great post about AI trends specifically mentioning the Gemma 4 success!! 🤩 www.interconnects.ai/p/some-ideas...
Some ideas for what comes next, May 2026 interconnects.ai
👁️ Vision Options: Want to make Gemma see even better? The default vision bucket is 280 for token efficiency. To capture maximum detail (like sharp OCR and 2.51MP resolution), manually bump max_soft_tokens to 1120! Try our new interactive Space to see how it works: huggingface…
Gemma 4 - Vision Token Budget - a Hugging Face Space by google huggingface.co
AI Weekly's analysis
→
- The Space lets users toggle Gemma 4's per-image budget across five preset sizes: 70, 140, 280, 560, and 1120 tokens.
- Gemma 4 launched April 2, 2026 under Apache 2.0 across E2B, E4B, 12B Unified, 26B A4B MoE, and 31B dense sizes.
- The 31B model reportedly hits 76.9% on MMMU Pro and 85.6% on MATH-Vision, with native JSON bounding box output.
Read full analysis →
I've been playing all day with Hugging Face Chat + Gemma 4 31B (deployed on Cerebras at +1000 tokens per second) and I'm still surprised how amazing it is!! I play with Gemma models basically everyday and this integration + speed still surprised me! take a look: huggingface.co…
google/gemma-4-31B-it - HuggingChat huggingface.co
🫂 A huge shoutout to the community for submitting fixes and finding new ways to make Gemma even better. We couldn't do this without you! ❤️ 🤗 Ready to test the speedup? Download the latest Gemma 4 updates now on Hugging Face: huggingface.co/collections/...
Gemma 4 - a google Collection huggingface.co
This was a pretty good week for Gemma + comunity integrations!!! here is one example: LiveKit: livekit.com/products/inf... conversational assistant using Gemma 31B as the brains. Super fast and great experience. Try their demo on their site!
Gemma 4 31B on LiveKit Inference | A faster, cheaper default for voice livekit.com