F5F Stay Refreshed Hardware Desktop The strength of Apple Silicon comes from its processing capabilities, not primarily its memory.

The strength of Apple Silicon comes from its processing capabilities, not primarily its memory.

The strength of Apple Silicon comes from its processing capabilities, not primarily its memory.

Pages (3): Previous 1 2 3
L
Lord_Sanguine
Member
100
05-18-2023, 04:22 AM
#21
@YoungBlade That's a data center CPU; probably not something consumers will get easily. Also, stop emphasizing more bandwidth doesn't boost CPU speed. Right now, check gaming benchmarks with various RAM rates. Yes, performance drops as memory speed rises, but a) we already handle hundreds of frames per second where the value per frame is smaller at high FPS than low FPS. So tasks limited by CPU will still improve, particularly those constrained by memory. b) Latency and timing matter. And C), our current RAM capacity is reaching its ceiling while maintaining the modularity we've had for decades. To get faster, we need to merge memory into a single unit. Why am I explaining this? Apple has already done it.
L
Lord_Sanguine
05-18-2023, 04:22 AM #21

@YoungBlade That's a data center CPU; probably not something consumers will get easily. Also, stop emphasizing more bandwidth doesn't boost CPU speed. Right now, check gaming benchmarks with various RAM rates. Yes, performance drops as memory speed rises, but a) we already handle hundreds of frames per second where the value per frame is smaller at high FPS than low FPS. So tasks limited by CPU will still improve, particularly those constrained by memory. b) Latency and timing matter. And C), our current RAM capacity is reaching its ceiling while maintaining the modularity we've had for decades. To get faster, we need to merge memory into a single unit. Why am I explaining this? Apple has already done it.

F
fastfrenchie
Junior Member
15
05-19-2023, 05:58 PM
#22
I never claimed extra bandwidth wouldn't improve CPU speed. I mentioned it's about reaching diminishing returns, a point you've recognized... But nothing. If the gains in performance become smaller with faster memory, then the path forward isn't increasing memory speeds, but identifying another key factor limiting overall performance.
F
fastfrenchie
05-19-2023, 05:58 PM #22

I never claimed extra bandwidth wouldn't improve CPU speed. I mentioned it's about reaching diminishing returns, a point you've recognized... But nothing. If the gains in performance become smaller with faster memory, then the path forward isn't increasing memory speeds, but identifying another key factor limiting overall performance.

G
GamerSwag
Junior Member
9
05-24-2023, 04:31 AM
#23
The CPU and GPU components all perform well, with efficient communication between them thanks to their packaging. Apple has heavily optimized both the operating system and software for these chips, plus developed a rapid translation layer (Rosetta 2) for apps not designed for ARM. Other factors also contribute to speed, such as the design of the memory technology. ^^^ The 7980XE uses quad-channel DDR4, offering better memory bandwidth compared to standard dual-channel setups. This advantage isn't enough to completely offset the limitations of Skylake-X cores or the higher latency in mesh architecture compared to Intel's ringbus design. For certain tasks—especially gaming—I've seen performance gains when memory bandwidth is the main constraint. That's been true. High-speed HBM2 can handle up to 1TB/s, and newer versions push even further. However, this only matters up to a point; real benefits depend on matching the hardware with capable processing units.
G
GamerSwag
05-24-2023, 04:31 AM #23

The CPU and GPU components all perform well, with efficient communication between them thanks to their packaging. Apple has heavily optimized both the operating system and software for these chips, plus developed a rapid translation layer (Rosetta 2) for apps not designed for ARM. Other factors also contribute to speed, such as the design of the memory technology. ^^^ The 7980XE uses quad-channel DDR4, offering better memory bandwidth compared to standard dual-channel setups. This advantage isn't enough to completely offset the limitations of Skylake-X cores or the higher latency in mesh architecture compared to Intel's ringbus design. For certain tasks—especially gaming—I've seen performance gains when memory bandwidth is the main constraint. That's been true. High-speed HBM2 can handle up to 1TB/s, and newer versions push even further. However, this only matters up to a point; real benefits depend on matching the hardware with capable processing units.

S
Spawn377
Member
215
05-24-2023, 08:55 PM
#24
The M3 Max features 32-bit memory controllers and its LPDDR5 operates at 6400MT/s (calculated from 32*16*6,400,000,000 divided by 8). The 14900k supports 4 channels of 32 bits each, delivering around 100GB/s. At 6400MT/s, Apple’s approach would waste significant die space on a regular CPU lacking a GPU. These memory controllers don’t adapt to node scaling. With the GPU integrated, Apple needed to ensure the GPU received sufficient data. The N3 process cost per millimeter is extremely high, making it uneconomical to double the channels for similar performance or die area. This explains why the 14900k lacks an additional 32-bit channel. The expense didn’t justify triple the channels on a platform requiring triple-channel DDR5. For workloads demanding high bandwidth, solutions like HEDT with Sapphire Rapid chips or ThreadRipper are viable. Emerald Rapids could help, though no HEDT variant exists yet. The cost pressure comes from needing to support 8 channels instead of just 4. In terms of CPU performance, removing half the memory controllers would have minimal impact. This dramatically reduces die area. When wafer costs hit $20k, every millimeter counts. Numbers are rough estimates—don’t rely on them. Many assumptions are involved, especially regarding defect rates and manufacturing changes. M3 Max measures under 400mm²; exact dimensions aren’t available. With a 50% reduction in controllers, you’d see a 20x increase in chip count per wafer. Apple charges TSMC based on these figures. If your application is bandwidth-intensive, consider specialized options like Sapphire Rapid or ThreadRipper. ARM processors, being RISC-based, require more instructions in memory to perform tasks similar to CISC, which are now handled by micro-ops. This can affect performance if the stack isn’t cached efficiently.
S
Spawn377
05-24-2023, 08:55 PM #24

The M3 Max features 32-bit memory controllers and its LPDDR5 operates at 6400MT/s (calculated from 32*16*6,400,000,000 divided by 8). The 14900k supports 4 channels of 32 bits each, delivering around 100GB/s. At 6400MT/s, Apple’s approach would waste significant die space on a regular CPU lacking a GPU. These memory controllers don’t adapt to node scaling. With the GPU integrated, Apple needed to ensure the GPU received sufficient data. The N3 process cost per millimeter is extremely high, making it uneconomical to double the channels for similar performance or die area. This explains why the 14900k lacks an additional 32-bit channel. The expense didn’t justify triple the channels on a platform requiring triple-channel DDR5. For workloads demanding high bandwidth, solutions like HEDT with Sapphire Rapid chips or ThreadRipper are viable. Emerald Rapids could help, though no HEDT variant exists yet. The cost pressure comes from needing to support 8 channels instead of just 4. In terms of CPU performance, removing half the memory controllers would have minimal impact. This dramatically reduces die area. When wafer costs hit $20k, every millimeter counts. Numbers are rough estimates—don’t rely on them. Many assumptions are involved, especially regarding defect rates and manufacturing changes. M3 Max measures under 400mm²; exact dimensions aren’t available. With a 50% reduction in controllers, you’d see a 20x increase in chip count per wafer. Apple charges TSMC based on these figures. If your application is bandwidth-intensive, consider specialized options like Sapphire Rapid or ThreadRipper. ARM processors, being RISC-based, require more instructions in memory to perform tasks similar to CISC, which are now handled by micro-ops. This can affect performance if the stack isn’t cached efficiently.

Pages (3): Previous 1 2 3