AMD Acquires Taalas, Startup Etching AI Weights Into Silicon
TL;DR
- AMD announced Thursday it acquired Toronto startup Taalas, whose chips etch AI model weights directly into silicon as mask-ROM alongside SRAM for KV caches.
- Taalas' HC1 test chip on TSMC's 6nm process reportedly served Llama 3.1 8B at 16,960 tokens per second, 48x Nvidia GPUs by its February claim.
- AMD plans to pair its Instinct-based Helios racks with Taalas silicon so GPUs handle prompts and token generation offloads to Taalas accelerators; close is expected in Q4.
The interesting thing about AMD buying Taalas isn't the acquisition itself, it's what it signals about how the biggest AI silicon buyers now think about inference. According to The Register, AMD announced at market close on Thursday that it has acquired the Toronto startup, founded in 2023, whose chips 'bake model weights directly into silicon in a process that promises to boost inference performance by an order of magnitude or more.' In a sense, as the reporting puts it, these are model-specific integrated circuits, or MSICs. Terms weren't disclosed, and subject to regulatory approval the deal is expected to close in the fourth quarter.
The pitch is the benchmarks. Taalas' first test chip, HC1, was fabbed on TSMC's 6nm process tech, and initial numbers had it serving Meta's Llama 3.1 8B at 16,960 tokens a second, which the company said was 48x faster than Nvidia's GPUs and 8.5x faster than Cerebras' accelerators when it was announced last February. A second-generation HC2 chip is due out this summer at 20 billion parameters per chip, with Taalas claiming that 50 accelerators would be enough to serve a trillion-parameter model. AMD reportedly intends to pair its Instinct-based Helios racks with Taalas silicon in a disaggregated setup where compute-heavy prompt processing runs on GPUs and token generation is offloaded to the Taalas accelerators. Vamsi Boppana, AMD's SVP of AI, framed it as building 'a full-stack AI platform that gives customers the flexibility to deploy the right compute solutions for every AI workload.'
The strategic read is that AMD wants its own version of Nvidia's $20 billion licensing deal with Groq from last December, and it's willing to pay for it outright rather than through partnership fees. Inference is where the compute bill actually lands once a model is in production, and if a stable model can be etched into an ASIC an order of magnitude more efficient per token, hyperscalers and enterprises running fixed workloads have a real reason to take the call.
The honest caveats: the performance numbers are Taalas' own from a February announcement, not independently verified in the reporting, and etched-weight silicon has an obvious drawback in that any model change requires a chip re-spin, softened only by the claim that just two layers of metal need to be changed. What the reporting doesn't give you is the deal price, the actual dollars-per-token math versus a comparable GPU rack, or how HC2's 20-billion-parameter ceiling squares with frontier reasoning models that keep growing. The bet is that a meaningful chunk of production inference will settle onto models stable enough to justify a mask set. If it does, this is a smarter use of capital than another generation of chase-Nvidia GPUs.
Originally reported by theregister.com
Read the original article →Original headline: AMD Acquires Taalas, Toronto Startup That Etches AI Model Weights Directly Into Silicon