Fast H3 implementation for Metal. Enjoy, modify, and so forth: https://t.co/FuyzEtUW7S Contains code from @liuliu which is welcomed in taking back whatever parts he likes for @drawthingsapp in case there are H3 plans there.
GitHub - antirez/h3.c: MiniMax H3 inference engine for Mac computers github.com
AI Weekly's analysis
→
- h3.c is a native Metal implementation of MiniMax H3 multimodal generation for Apple Silicon, with M3 and M5 Max as the tested platforms.
- The transformer checkpoint from Hugging Face is about 33 GB, and peak physical memory hits approximately 40 GB during end-to-end generation.
- On M5 Max, the fast preset renders 512×512 in about 16.69 seconds; the aggressive four-step path finishes in roughly 3.5 seconds.
Read full analysis →
I uploaded the new (0813) Q2 quants of DeepSeek v4 PRO at the following URL. Quality tests still ongoing. https://t.co/424V57IoyO
DeepSeek-V4-Pro-IQ2XXS-w2Q2K-AProjQ8-SExpQ8-OutQ8-Instruct-imatrix-0813.gguf · antirez/deepseek-v4-gguf at main huggingface.co
DeepSeek v4.1 Flash support is now pushed on DwarfStar "main" branch on GitHub, and this is a YouTube video (in English language) where I test both the SSD streamed and the dual MacBook m5 max 128GB setup during a coding session: https://t.co/SAgr2gbFrh
Proviamo DeepSeek v4.1 Flash a 2 bit con DwarfStar (SUB ITA) youtube.com