Chips & SupplyUnited States

Thermal Complexity Grows With AI Chips And Photonics

Credited to Semiconductor Engineering · semiengineering.com

Useful

Experts at the Table: Semiconductor Engineering sat down to discuss the impact of heat on photonics, transient behavior, and system reliability, with Lang Lin, director of product management at Synopsys; Jack Berg, vice president of business development at Silvaco; Satish Radhakrishnan, head of semiconductor and electronics at Vinci; Chris Mueth, director of new markets management at Keysight EDA; and John Ferguson, senior director of product management at Siemens EDA. What follows are excerpts of that conversation, which was held behind closed doors at the recent Design Automation Conference. This is the second of a three-part series. Part one is here.

L-R: Silvaco’s Berg; Synopsys’ Lin; Vinci’s Radhakrishnan; Keysight’s Mueth; Siemens’ Ferguson.

SE: Moving data takes power, which generates heat. So how do we keep data moving while still remaining within the power budget? Is the solution photonics? Even there, heat can cause issues.

Mueth: Photonics has a lot of promise for reducing data center bottlenecks, but it can’t be scaled right now. Building a large mesh switch requires many components, and you have high losses. So you have to revisit the whole architecture for how things are done in photonics and do them differently, probably in a modular style that’s scalable. Another issue is heat sensitivity. You have to stabilize that. But even then, your photonic structure is going to have heaters in it because these have to be controlled. And if you have a large mesh switch, the amount of control you need is significant. So how do you calibrate your structures through that chip to get enough performance when you may have thousands of heaters that you have to control? If we can solve that problem, it will open things up a bit. But that still doesn’t get rid of the losses in the |photonics chips. That has to be solved with a modular, scalable architecture.

SE: This leads to another issue, which is that you’re no longer dealing with just one type of processing element. What happens when you go from a GPU to a CPU? Do they both generate the same amount of heat? And is that, in turn, different when you have a high-intensity workload that’s short-lived? How do you cool everything down to keep temperatures consistent?

Radhakrishnan: Co-packaged optics is trying to solve that problem, to some extent, bringing the silicon photonics into the package, keeping it small for EICs/PICs. They are trying to thermally isolate it. But if you try to simulate the whole thing, the part that will be missing is the transient nature, and that’s what’s really needed. If you try to keep something far off, then the heat is going to take more time to get there, and by the time the next wave comes on an ASIC, the heat will still be there. Even a very small change in temperature can shift the ring oscillator. What’s needed is the ability to simulate the whole thing together in a transient domain to make sure it’s going to function reliably and optically. Being able to simulate one part of a system in a transient domain is key to making sure co-packaged optics is viable and that we can push it forward.

Mueth: That’s a good point. The difference in time constants really messes things up. And if you’re doing co-packaged optics, the EIC driver typically would be mounted near or on top of your PIC, so you have heat pollution, which steers your photonics. You’re introducing new stresses, and those stresses have a very big impact on the optical signals — maybe bigger than thermal. You have to know what those trade-offs are, and make sure you’re taking them into consideration so that you’re combining the best of both. Just by moving something, you can make things worse instead of better.

SE: And how do you know your thermal sensors are functioning properly all the time?

Mueth: You can do that in the calibration process. You can put a sniffing loop inside the photonic PIC, but it makes that incredibly complicated.

Ferguson: We can try to move things to reduce some of the thermal that’s built up. But when you’re moving big elements, particularly like what we have in photonics, you’re introducing new stresses. Those have a very big impact on optical signals — maybe bigger than the thermal. So you have to know what those trade-offs are and make sure you’re taking them into consideration so that you’re combining the best of both.

Lin: For silicon photonics, you can think about it as a type of analog circuit. When we deal with analog versus digital, the thermal behavior is very different. Digital transistors involve switching current. The current will supply the transistor’s on and off, so the Joule heating effect is not that bad. But for silicon photonics and analog devices, the DC current is always there, and there’s a big swing of signals. That creates current flowing through the interconnects, like a waveguide. It generates a lot of heat, and we need this heat for the CPO to work. That heat is needed to operate the EIC. But what if there are so many heaters together? How do you place them far enough away so that everything can operate, but not create hot spots within the heater area? Stress is another problem that has to be considered for silicon photonics to operate well.

SE: Heat can cause warpage. If you put two warped chips together and they’re very thin, you may be able to flatten them, but at the same time you’ve probably lost some ability to dissipate heat due to thinner dies, right?

Berg: Thermal variation on a PIC chip generally results in a mismatch of features — much more so than in EIC circuits. If you’re talking about multi-physics and digital twins and surrogate models affiliated with EICs, you essentially can triple that need in photonics ICs because they’re much more sensitive to warpage and thermal gradients. A laser is a heat source that potentially creates a reliability problem. To that end, photonics needs a thermal understanding and a multi-physics confirmation of not only the electrical, mechanical, and thermal, but optical as well. So now you’ve got four solvers that you potentially have to do. It’s a bit more complicated than the electronics, and that’s the reason why photonics has been the next big thing in the semiconductor world for the past 30 years. It’s much more complicated from a physics perspective than anybody thought it would be.

Mueth: Traditionally we were worried about thermal in semiconductors. We ignored the passives. Now, with photonics, you have thermal effects on the passives.

SE: Within a chip you may have multiple dies, because you cannot get the performance you need out of a single die, especially in AI data centers. How do we coordinate the different dies within a package so you aren’t just running everything flat-out all the time?

SE: This is a complex orchestration problem, right? And it needs to be hierarchical.

Mueth: Yes, and when you’re doing the architecture, you’re going to have to model your thermals. One of the corners would be, ‘What are your high-performance use cases?’ You have to model those thermally, electrically, and mechanically, so when you are farther down the design path you have a higher chance of success. But that modeling has to be done up front.

Ferguson: You can do each of those standalone variables, but then you’re not getting the full story. You have to do them all consecutively, which means that what you’re doing is running each one many, many times, one after another, until you get to a steady state. And then you hope that is the correct steady state. That depends on what you’re building.

Radhakrishnan: If you look at the direction AI chips are going, they’re trying to get from training to inference. They are trying to increase the bandwidth. That’s what is driving all the chips to be stacked on top of each other, and that affects their behavior. They want to understand the inference behavior. But it’s not just about peak performance. How much higher is the temperature going to be in transient? If I have much more headroom, because it’s only going to last for a nanosecond, maybe I can go to extra power. But is that nanosecond going to come every millisecond or microsecond? How often am I going to allow peak power? And how much of the thermal budget can I push and let the data center people create a design rule that says, ‘For inference, I can allow a millisecond pulse every 100 milliseconds.’ This is a huge new paradigm that people can exploit when they go to different use cases. It’s all driven by what the end user wants and how they want to achieve the maximum performance they need.

Lin: Going back to the question about hierarchical models, those are important not only to deal with the capacity to simulate thermal for an AI chip, but also to enable an ecosystem with different vendors providing their model hierarchically. For example, HBM vendors have to provide some kind of model for an HBM chip. Then they send that to a system integration house, which will combine different models, build a whole multi-die stack, and run thermal. First, they need to consider the impact of the vendor’s model on thermal. Second, they need to finalize the whole system. They’re the final product delivery hub. They have to make sure all the models are connected, assembled, and then simulate the whole system. So the hierarchical model is a must, and not only for capacity. The industry is working on this. We need the ecosystem to deliver a model, and we have to agree on standards so that everybody can share their models and the integrators can assemble them.

SE: One of the big concerns in data centers is aging of data paths due to electromigration, which is accelerated by heat. How does that impact the overall design process?

Radhakrishnan: I’m seeing a lot of thermal sensors being embedded into the die for aging. Those have to be on-chip to determine degradation over time. Then, you can throttle parts of the chip. But everything is packaged up, so it has to be done by software getting some sensing information back, and then you see what changes you can make. This is even more of an issue with automotive, where chips have to last 10 to 15 years. That’s where you need to do simulation to determine if you’re putting the sensors in the right location. Can I get the right sensing back so a chip can function over time as things change? You need to be able to capture all that, and this is the direction companies are heading.

SE: This can vary by ambient temperature, right?

Berg: Yes, particularly places like Saudi Arabia. The simple rule of thumb is the Arrhenius equation. Every 10 degrees results in a substantial reduction in reliability. This is really important to get right.

Lin: But I don’t see a lot of tools in the industry for accurately predicting thermally induced aging. There’s a gap. We need tools to pair the correct equation so we can predict aging effects, and how to resolve them. Sensors are a way to monitor the aging effects, but what’s happening on the design optimization side? No one has done that for aging-aware design optimization.

Radhakrishnan: It’s an under-design/over-design kind of problem. That’s why, ideally, it has to be dynamic over time. And that is a huge problem.

Read part one of the discussion: Scaling Thermal Analysis From Transistors To Data Centers Rising power density and chiplet complexity drive thermal considerations earlier in the design process.

Related: Photonics Fundamentals For Electronics Engineers: eBook

Original · Semiconductor Engineering

FrontMethod