Why we should care about HBM in AI models
When you hear ‘HBM’, think of a super‑fast, super‑expensive memory stack that sits on top of a GPU or ASIC. It’s the fuel that powers massive language models – but it also costs a fortune. A 40‑GB HBM2e chip can set you back ₹6‑7 lakh, and that’s just for the memory, not the processor. For Indian startups and even big players like Tata Elxsi or Wipro, that price tag is a blocker.
DeepSeek’s secret sauce
DeepSeek, the Chinese AI lab, announced a set of optimisation techniques that cut the memory bandwidth requirement of their 7B‑parameter model by almost 40 %. They achieved this by:
- Switching to grouped‑query attention – it reduces the number of key‑value pairs that need to be fetched from memory.
- Applying activation recomputation – instead of storing every intermediate activation, the model recomputes them on‑the‑fly, shaving off memory usage.
- Using a mixed‑precision format called FP8‑E4M3 for some layers, which keeps accuracy while halving the data width.
The result? A model that runs comfortably on a 16‑GB GDDR6 board, which is about a quarter of the cost of an HBM‑equipped system.
What this means for Indian AI hardware ecosystem
India has a budding AI chip scene – from Wistron’s AI‑focused ASICs to Ineda Systems’ low‑power RISC‑V cores. The common pain point is memory. Most domestic designs still rely on DDR4/5 because HBM supply chains are dominated by Samsung, SK Hynix and Micron, all of which have limited capacity for new customers.
With DeepSeek’s optimisation, a 16‑GB GDDR6 module (₹55 k on the market) can replace a 40‑GB HBM2e stack. That immediately brings the bill of materials (BOM) down by 60‑70 %. For a startup, that could mean the difference between a ₹2 crore prototype and a ₹80 lac one.
Indian use‑case: Edge AI for language services
Imagine a Tamil‑to‑English translation device that runs locally on a small board, no internet needed. Today, that would need a heavy HBM‑based accelerator, making the product price >₹30 k. With DeepSeek’s memory‑light model, the same functionality can be packed into a Raspberry Pi 5‑class board with a 16‑GB LPDDR5 extension, pushing the retail price to under ₹12 k.
Potential roadblocks
1. Licensing & IP – DeepSeek’s techniques are not open‑source. Indian firms will need to negotiate licences or develop in‑house equivalents. 2. Software stack – The model relies on custom kernels for grouped‑query attention. Existing frameworks like TensorFlow Lite or ONNX Runtime may need patches. 3. Performance trade‑off – Activation recomputation adds extra FLOPs, which could raise power consumption on low‑end CPUs.
TamilTech‑ஓட கருத்து
We think this is a turning point. If Indian chip designers can adopt these tricks, the whole ecosystem could shift from being HBM‑dependent to a more democratic, cost‑effective model. That would accelerate home‑grown AI products – from voice assistants in regional languages to real‑time video analytics for smart cities.
What to watch next
- Will DeepSeek open‑source the optimiser? A public repo would fast‑track adoption.
- Can local fabs like STMicroelectronics India or Sahasra Semiconductor produce 16‑GB GDDR6 chips at scale?
- Watch for partnerships – we expect a tie‑up between DeepSeek and an Indian AI startup within the next 3‑6 months.
Bottom line: less HBM means cheaper AI hardware, and that’s music to the ears of Indian innovators. Stay tuned, because the next wave of AI products could be built right here, in our own labs.




Comments (0)
Be the first to comment!