F5F Stay Refreshed Hardware Desktop Discussion about Intel Xe graphics performance and features.

Discussion about Intel Xe graphics performance and features.

Discussion about Intel Xe graphics performance and features.

Pages (2): Previous 1 2
T
thetalkkari
Member
152
03-01-2023, 05:37 PM
#11
I understand your concern, but fixed-function hardware consumes significant silicon space. Following Raja Kodury's tweets suggests holding peta-operations in hand—something like Nvidia’s tensor cores would be necessary. This likely points to INT8 performance rather than traditional FP32/64 or even TFP32/16. It was meant as a joke, but the paper writer hinted that porting games across systems could become a future application. I’m skeptical this solution is scalable enough, though we can’t say for sure.
T
thetalkkari
03-01-2023, 05:37 PM #11

I understand your concern, but fixed-function hardware consumes significant silicon space. Following Raja Kodury's tweets suggests holding peta-operations in hand—something like Nvidia’s tensor cores would be necessary. This likely points to INT8 performance rather than traditional FP32/64 or even TFP32/16. It was meant as a joke, but the paper writer hinted that porting games across systems could become a future application. I’m skeptical this solution is scalable enough, though we can’t say for sure.

G
Guizk
Member
61
03-06-2023, 12:04 PM
#12
It's actually the design choice behind this that matters. Fixed-function hardware uses less memory and is cheaper than general-purpose units for identical tasks. This approach typically offers around int8 or int4 speedups, though tensor cores aren't strictly necessary. Int4 calculations were already efficient in Kepler (four times faster than regular fp32). Tensor cores introduced a quicker method for FMA matrix operations with those types of numbers, reducing the need for multiple instructions.
G
Guizk
03-06-2023, 12:04 PM #12

It's actually the design choice behind this that matters. Fixed-function hardware uses less memory and is cheaper than general-purpose units for identical tasks. This approach typically offers around int8 or int4 speedups, though tensor cores aren't strictly necessary. Int4 calculations were already efficient in Kepler (four times faster than regular fp32). Tensor cores introduced a quicker method for FMA matrix operations with those types of numbers, reducing the need for multiple instructions.

T
TdmFan92
Senior Member
602
03-08-2023, 03:21 AM
#13
They definitely don't occupy much room on the silicon compared to using GP-units. However, they still consume space and serve a single purpose. I searched for details about the exact size of an FF-block in a current Intel iGPU, but unfortunately couldn't find any and I don’t have much time left to dig deeper. So my initial thought might have been incorrect. But implementing fixed-function hardware generally reduces performance for other tasks, as it takes up die space—though that space is still significant. I’m aware I’ve never seen precise dimensions and only heard descriptions from people unfamiliar with the details, often calling the block large due to its design rather than its actual performance. If I’ve misunderstood, please let me know!

Currently, in the Turing architecture, there are roughly the same number of INT execution units as FP32 units. This means INT32 and FP32 performance are comparable. For a 2080TI at base clock, this would be around 13.4 TOPS. The x4 capability comes from splitting an INT32 calculation into four INT8 calculations, potentially reaching about 53.6 TOPS for INT8 work. I suspect Intel wouldn’t mismatch the INT to float ratio in GP-units. They might even use units that switch between FP and INT operations without doing both in one cycle. In my view, they could enter the POP domain by having dedicated INT8/4 units—what some call tensor cores—not because they also handle matrix multiplication, but simply to keep them separate from GP-units. Refer to them as INT cores if you prefer.

Please correct me if I’m wrong. Also, I’m not sure if this discussion thread is still the right place for this topic.
T
TdmFan92
03-08-2023, 03:21 AM #13

They definitely don't occupy much room on the silicon compared to using GP-units. However, they still consume space and serve a single purpose. I searched for details about the exact size of an FF-block in a current Intel iGPU, but unfortunately couldn't find any and I don’t have much time left to dig deeper. So my initial thought might have been incorrect. But implementing fixed-function hardware generally reduces performance for other tasks, as it takes up die space—though that space is still significant. I’m aware I’ve never seen precise dimensions and only heard descriptions from people unfamiliar with the details, often calling the block large due to its design rather than its actual performance. If I’ve misunderstood, please let me know!

Currently, in the Turing architecture, there are roughly the same number of INT execution units as FP32 units. This means INT32 and FP32 performance are comparable. For a 2080TI at base clock, this would be around 13.4 TOPS. The x4 capability comes from splitting an INT32 calculation into four INT8 calculations, potentially reaching about 53.6 TOPS for INT8 work. I suspect Intel wouldn’t mismatch the INT to float ratio in GP-units. They might even use units that switch between FP and INT operations without doing both in one cycle. In my view, they could enter the POP domain by having dedicated INT8/4 units—what some call tensor cores—not because they also handle matrix multiplication, but simply to keep them separate from GP-units. Refer to them as INT cores if you prefer.

Please correct me if I’m wrong. Also, I’m not sure if this discussion thread is still the right place for this topic.

R
rosaliE65
Member
211
03-08-2023, 08:17 AM
#14
Yeah, I get what you meant. A misused area ends up as wasted space when it comes do chip design, since that same space could be used for something more useful or even just don't exist at all (making the chip cheaper due to the smaller area). Finding actual die shots that explain which part is left for the media engines is really hard, I couldn't get by any from nvidia or intel, but did manage to find one from AMD (from wikichip It's kinda big for that kind of chip, but represents a small area percentage when it comes to big GPUs and whatnot. Anyway, I guess that we kinda steered away from your main point. Intel is probably still going to include media engines into their discrete GPUs, at least for the consumer market, which they are going to attack. Not so likely into their HPC cards tho for obvious reasons. Yes, indeed that how it works. I wouldn't doubt if they had multi purpose execution units with higher INT4/8 throughput. Keep in mind that intel was pushing really hard for low precision ML (going so far as trying single binary values). So maybe that's where Raja got his peta op number from. It's his twitter afterall, not some official marketing campaign. If they come up with special hardware, tensor-like, it'd be pretty cool, but kinda useless for now IMO since most of the market is dominated by nvidia with CUDA, migrating to intel's stack wouldn't be an easy task. Probably isn't, but it's hard to have such kind of discussion in this forum anyway, since it's mostly gaming/consumer focused, and I do enjoy going a bit more technical (thanks for that!).
R
rosaliE65
03-08-2023, 08:17 AM #14

Yeah, I get what you meant. A misused area ends up as wasted space when it comes do chip design, since that same space could be used for something more useful or even just don't exist at all (making the chip cheaper due to the smaller area). Finding actual die shots that explain which part is left for the media engines is really hard, I couldn't get by any from nvidia or intel, but did manage to find one from AMD (from wikichip It's kinda big for that kind of chip, but represents a small area percentage when it comes to big GPUs and whatnot. Anyway, I guess that we kinda steered away from your main point. Intel is probably still going to include media engines into their discrete GPUs, at least for the consumer market, which they are going to attack. Not so likely into their HPC cards tho for obvious reasons. Yes, indeed that how it works. I wouldn't doubt if they had multi purpose execution units with higher INT4/8 throughput. Keep in mind that intel was pushing really hard for low precision ML (going so far as trying single binary values). So maybe that's where Raja got his peta op number from. It's his twitter afterall, not some official marketing campaign. If they come up with special hardware, tensor-like, it'd be pretty cool, but kinda useless for now IMO since most of the market is dominated by nvidia with CUDA, migrating to intel's stack wouldn't be an easy task. Probably isn't, but it's hard to have such kind of discussion in this forum anyway, since it's mostly gaming/consumer focused, and I do enjoy going a bit more technical (thanks for that!).

A
AgentDiamond
Member
95
03-08-2023, 09:06 AM
#15
I also realized the engine is quite large, but this chip is only 200mm², so that should be acceptable. Thanks for locating and sharing it! Absolutely. My main idea was that Intel was proud to add this feature to their server chips. They wouldn’t want to include anything too powerful (big) on them. So they’d need the XE to operate at around 1/8th of a tile, which is only possible if the design allows such scaling. Alternatively, they might just tweak it slightly—they do have engineers to handle it. You’re suggesting that the units used in GPU designs would actually be capable of delivering about 10x or 20x better int8/4 performance compared to int32? Or maybe they’d add an extra int8 module? Well, if they include dedicated blocks, they could still support various integer operations, not just machine learning. There’s also a translation layer called CU2CL, which NVIDIA promoted. In a 2011 paper it performed quite well (see the slides from that year). So adapting this shouldn’t be too difficult at least. They’d still need to invest time in optimization, but that’s always the case. I really appreciate these discussions. Thanks a lot! If you’re interested, we can keep talking about this via PM. But it seems like we’ve reached a point where it’s mostly finished... If you ever come across something cool, feel free to reach out!
A
AgentDiamond
03-08-2023, 09:06 AM #15

I also realized the engine is quite large, but this chip is only 200mm², so that should be acceptable. Thanks for locating and sharing it! Absolutely. My main idea was that Intel was proud to add this feature to their server chips. They wouldn’t want to include anything too powerful (big) on them. So they’d need the XE to operate at around 1/8th of a tile, which is only possible if the design allows such scaling. Alternatively, they might just tweak it slightly—they do have engineers to handle it. You’re suggesting that the units used in GPU designs would actually be capable of delivering about 10x or 20x better int8/4 performance compared to int32? Or maybe they’d add an extra int8 module? Well, if they include dedicated blocks, they could still support various integer operations, not just machine learning. There’s also a translation layer called CU2CL, which NVIDIA promoted. In a 2011 paper it performed quite well (see the slides from that year). So adapting this shouldn’t be too difficult at least. They’d still need to invest time in optimization, but that’s always the case. I really appreciate these discussions. Thanks a lot! If you’re interested, we can keep talking about this via PM. But it seems like we’ve reached a point where it’s mostly finished... If you ever come across something cool, feel free to reach out!

Pages (2): Previous 1 2