Written by: Little Pancake
On August 6, AMD announced the acquisition of Toronto-based chip company Taalas, with the transaction price undisclosed. Established in 2023 with a total funding of $219 million, this startup did something that sounds a bit crazy: burn the weights of AI models directly into the metal layer of the chip, creating a dedicated chip that can only run one model.
The way GPUs work is by storing model parameters in high bandwidth memory (HBM) and repeatedly moving data in and out of computation units during inference. This "transport" process consumes a significant amount of energy and time, known in the industry as the "memory wall." Taalas' approach completely eliminates transportation, solidifying weights within the chip's transistors, integrating storage and computation.
The cost is that this chip can only run one model, while the benefit is a magnitude leap in speed and energy efficiency.
What does 17,000 tokens/second mean?
Taalas' technology validation chip HC1 is based on TSMC's 6nm process, with a chip area of 815 square millimeters and 53 billion transistors, and it is built with the model Meta's Llama 3.1 8B.
HC1's benchmark data: single user about 17,000 tokens/second. In comparison, Nvidia H200 achieves about 230 tokens/second on the same model, while H100 achieves about 150 tokens/second. A 73 times speed difference, with power consumption only one-tenth of the latter.
A more intuitive economic calculation: Taalas reports an inference cost of about 0.75 cents per million tokens (Llama 8B), whereas GPU solutions range between 20 to 49 cents. Each HC1 card consumes 250W, and a standard air-cooled rack with 10 cards requires only 12-15 KW, without the need for liquid cooling, while GPU racks commonly consume 120-600 KW.
Forbes analyst Karl Freund, in his evaluation in February, stated: “It is 10 times faster than the fastest inference platform (Cerebras wafer-scale engine) and two orders of magnitude faster than GPUs.”
However, there are several important caveats. HC1 uses Taalas' self-developed 3-bit data format in conjunction with 6-bit parameters, which leads to significant quantization and deterministic quality loss on complex inference tasks. The 17,000 tokens/second data comes from a vendor benchmark test with 1K input/1K output, which does not represent real performance in production environments. Additionally, Llama 3.1 8B is a model released in mid-2024, and a quantized version with 8B parameters can even run on a Raspberry Pi 5. HC1 demonstrates the potential of the architecture, rather than the competitiveness of a mature product.
What about model iteration?
Burning the model into the chip raises the most intuitive question: What if the model iterates? Does the chip become useless?
Taalas' answer is: it won't be useless. Traditional ASIC design cycles take more than two years. Taalas has developed an automated design process that only requires minor modifications to the chip's top metal mask to compile a new model into a new chip, with a cycle time of about 8 weeks. The company has preset the costs for 3 chip updates over 4 years in its cost model.
This answer alleviates some anxiety but cannot eliminate it entirely.
The flexibility of GPUs lies in being able to run Llama today, Claude tomorrow, and a new model that hasn’t been released yet the day after. Taalas's chips cannot do this. If a customer requires multiple model services at the same time (which is very common in production environments), they will need to equip different chips for each model, increasing inventory management and operational complexity.
HC2 is already in development, targeting models with about 20 billion parameters and using 4-bit floating point to improve precision. Taalas initially planned to support cutting-edge models by the end of 2026.
AMD's acquisition makes this path have synergy with the Instinct GPU product line and Helios rack system, allowing customers to use GPUs for flexible and varied workloads while using Taalas chips for stable high-frequency single-model inferences.
The arms race for inference chips
AMD's acquisition of Taalas, viewed in the context of the industry, is the latest in a wave of inference chip acquisitions over the past 8 months.
In December 2025, Nvidia acquired Groq for $20 billion, securing a low-latency inference architecture based on SRAM.
In May 2026, Cerebras went public with an IPO valued at approximately $56 billion, becoming the largest tech IPO since Snowflake in 2020.
In June, Etched ended more than two years of stealth mode, announcing that their first transformer-specific chip has completed tape-out on TSMC's N4P process, with over $1 billion in customer contracts in hand. Intel reportedly signed a letter of intent to acquire SambaNova, while Qualcomm is in talks with Tenstorrent.
These transactions point to the same conclusion: inference is becoming the main battleground for AI computing expenditure.
Industry forecasts predict that by 2026, inference will account for two-thirds of AI computing expenditures, and by 2030, dedicated inference ASICs could capture 45% of the inference market. Nvidia's GPUs still dominate the training side, but on the inference side, its general architecture faces challenges from dedicated chips on multiple fronts.
Taalas is at the extreme end of this spectrum: it has given up even the generality of "a category of models" and has gone straight to making dedicated chips for "one model."
Groq uses SRAM to replace HBM to bypass the memory wall, Cerebras uses wafer-scale area to force breakthroughs, Etched optimizes solely for the transformer architecture, while Taalas's bet is more aggressive: it bets that AI models will gradually stabilize, just as communication protocols ultimately converge into a few standards. When models no longer iterate every month, creating a dedicated chip for the most popular models becomes economically viable.
Betting on the models “solidifying”
AMD Senior Vice President Vamsi Boppana stated in the announcement that the purpose of the acquisition is to provide customers "with the optimal computing solution for every type of AI workload." This translates into a competitive strategy: GPUs handle flexibility, Taalas chips handle extreme efficiency, and customers combine as needed.
Whether this combination logic holds true depends on one core assumption: whether the update cycle of AI models will slow down.
If large models continue to have architectural updates every few months, then burning weights into silicon is like pouring concrete on quicksand; before the chip pays off, the model has already become obsolete. However, if model capabilities tend to converge, and stable service of the top models over one to two years becomes the norm, then Taalas-style dedicated chips will become as natural a choice as baseband chips in the communications industry.
Looking back over the past 12 months, the pace of model updates does seem to be experiencing subtle changes. The intervals between versions of the Llama series from 3.1 to 3.2 to 4 are lengthening, and the architectural changes from GPT 4 to 4o to 5 are narrowing. More new models are being distilled, quantified, and fine-tuned on existing architectures, and the iteration of base models is slowing down.
If this trend continues, Taalas will have made the right bet.
免责声明:本文章仅代表作者个人观点,不代表本平台的立场和观点。本文章仅供信息分享,不构成对任何人的任何投资建议。用户与作者之间的任何争议,与本平台无关。如网页中刊载的文章或图片涉及侵权,请提供相关的权利证明和身份证明发送邮件到support@aicoin.com,本平台相关工作人员将会进行核查。