Hope AMD could integrate ten cores into one CCD chip.
Hope AMD could integrate ten cores into one CCD chip.
Since AMD Ryzen was introduced, we've consistently enjoyed 8 cores per chip. I understand this limitation comes from the physical design and arrangement of the chips, preventing more than eight cores on a single unit. Yet, AMD's engineering has advanced significantly since the Ryzen 1000 line was released nearly a decade ago. Back then, it used 14-nanometer transistors, whereas today's Ryzen 9000 series operates with 4-nanometer technology. This means the same transistor count occupies less than 30% of the chip space. Why not simply redesign the chip to add two more cores? There must be a feasible solution. It would also benefit the overall product lineup, especially when compared to Intel's Hybrid P/E-Core approach, which offers substantially more threads per model. While Intel has expanded its thread count over the years, AMD has remained relatively static, missing opportunities that could have been valuable. Picture an X800X-3D processor with integrated 3D V-Cache and 10 cores or 20 threads on a single chip—no core slots needed. Such a design would be highly desirable. Many customers chose Intel's 14900/K/KS for its thread density while delivering top-tier gaming performance, so AMD should seize this chance. Also envision a 20-core/40-thread X950X that easily surpasses any i9 in multi-threaded tasks. AMD could retain its current 8 and 6 core options, resulting in a more comprehensive lineup. I envision an ideal stack like this: Single CCD—Ryzen 3 X300 (4 cores/8 threads), Ryzen 5 X400 (6 cores/12 threads), Ryzen 5 X500 (8 cores/16 threads), Ryzen 7 X600 (10 cores/20 threads), Twin CCD—Ryzen 7 X700 (12 cores/24 threads), Ryzen 9 X800 (16 cores/32 threads), X900 (20 cores/40 threads). Threadripper would then start at 24 cores, offering a solid alternative to desktop Ryzen 9. A 20-core/40-thread X950X could easily outperform any i9 in parallel workloads. AMD could also introduce X3D variants for each chip, and I removed the X950 naming since it doesn’t add much value. This approach would make Threadripper a compelling choice too, as it already supports up to 96 cores and 192 threads. The advantages for desktop CPUs should be obvious. That’s all my thoughts on this.
I believe the benefits are now justified. Back when the Ryzen 1000-3000 line was available, I would have supported your view, but nowadays certain titles can effectively utilize more powerful multi-core processors. Moreover, if you're similar to me and run several background applications while gaming—such as Discord, Chrome browsing, and maybe recording—I see extra cores as a significant advantage for multitasking. Of course, having more threads on workstations is generally beneficial, which is why 16-core desktop CPUs were developed originally.
Zen 1/1+ wasn't built from separate chips; it was a solid design. They didn't use CCDs in Zen 2, which came with RYZEN 3000 chips featuring chiplets on 7nm. Since Zen 2, they moved to 5nm (4nm is part of TSMC's 5nm line). We've only advanced one generation since Zen 2 for the CCD. While the chip size has shrank, the processing units have grown bigger. Zen 5 introduced two extra ALUs inside the core, boosting transistor density per mm². The increase in transistors per square millimeter hasn't changed much from Zen 2 (74mm²) to Zen 5 (70mm²). Zen 3 was 81mm² per CCD, similar to Zen 2 but more compact. In terms of AMD, controlling yields mainly comes from managing CCD size. Going beyond 80 reduces yields due to more defects and fewer wafers per chip, making costs per CCD higher compared to 70nm. Because they can't fit enough production space for all products anymore, they opted not to expand further. This explains why they invested heavily in 7nm (mobile Zen 3) and 6nm (Zen 2 & Zen 3+). For Zen 6 at 3nm, each wafer is about 20k, excluding assembly complexities for chiplets and X3D Vcache. After Zen/Zen 7 on 2nm with improved power delivery, each wafer will be under 30k.
I hadn't realized the original Zen 1 was based on Monolith; I assumed they were CCD from the start and that's why AMD started using Zen/Ryzen names. While the cores themselves have increased, I'm sure they could accommodate ten cores per CCD now. In fact, the I/O section built on a larger node seems to be the main challenge. To my knowledge, the I/O on the Ryzen 9000 still uses either 5nm or 6nm technology, which is significantly larger and consumes more space. Still, using older nodes for I/O makes sense because it greatly reduces manufacturing costs. You wouldn't need I/O features from the newest node, right? The main reason might be that AMD hasn't reached a point where 10 cores per CCD are practical yet. I think AMD could find a middle ground and fit ten cores on a CCD eventually. After seven years of development, it's hard to imagine it wouldn't be possible. But honestly, that was just a thought I had.
I/O remains on older nodes due to minimal scaling in analog technology following recent node reductions. We anticipate a significant increase in scaling when combining GaAs and backside power, though precise estimates are difficult. Vcache also stays on an older node since SRAM scaling is only about 5%, while TSMC 3nm offers no improvement. To accommodate more cores per CCD without increasing size, some compromises are necessary—such as reducing cache performance or delaying optimizations by another year. This would allow a second layout iteration without pre-made blocks, which explains why cloud and dense designs are still around 16x smaller than the original. The SRAM scaling challenge will likely pose the biggest hurdle for Nvidia in the next generation. How do you believe ADA achieved most of its improvements?
@starsmine I often thought shrinking nodes would only help a little for I/O, and beyond a certain point it wouldn’t really matter. I believed the core logic should stay intact, which makes sense in context. What I meant earlier was that reducing node size mainly for I/O isn't very practical. To my knowledge, expanding CCDs slightly could add more cores without major issues. With current CPU advancements, especially after moving past 4-core processors, the need for huge performance jumps via IPC is less pressing. Most users are content with today's speed, and unless demands suddenly spike, this trend should continue. Another concern is shrinking nodes has become harder—just over ten years ago TSMC engineers warned they couldn’t go smaller than 5nm without a major redesign. Although their estimate was slightly off, it took two years to build a working CPU on a node just one nanometer smaller. The IPC gains there aren’t impressive, and single-core apps run fast enough for most people. Multi-core performance is more noticeable now. For high-power desktops, shrinking to 3nm brings its own challenges. Extremely dense transistors generate massive heat quickly, risking thermal meltdowns similar to a nuclear core. Heat dissipation becomes a serious issue, and problems like electromigration and material wear accelerate. These hurdles could affect future desktop CPUs. Yet, with 96-core Threadrippers and 192-core EPYC chips available, it seems AMD can realistically achieve 10 cores on a single Ryzen CCD. It’s clear there are viable solutions now.
It’s worth exploring information on yield rates. Bigger chips often come with bigger monolithic dies, though they aren’t cost-effective for general use.
The issue lies with CCX, not CCDs. Zen 1 through Zen 2 featured four core CCX units, but we only grasped their impact until Zen 3 introduced eight core CCX, where we hit a roadblock. AMD opted to expand via CCX/CCDs for older Cinebench performance, which suits a more traditional design rather than a unified architecture. I’m counting on AMD abandoning the Infinity Fabric strategy in CPU scaling and shifting toward an Intel-style setup with higher bandwidth and lower latency, which could improve unification. If that doesn’t work, we might need to adopt a L4 cache instead of expanding L3 to better combine separate cores.