Question about PCIe bandwidth (resolved)
Question about PCIe bandwidth (resolved)
You're asking about how PCIe lane usage affects performance. With your setup, the two PCIe 4.0 lanes provide higher bandwidth than the three PCIe 3.0 slots. If you use two PCIe 3.0 cards in the same 16X slots, each will operate at full 3.0 speeds. Adding a PCIe 4.0 card in each slot would let it run at 8X speeds, not 4X as you thought. So yes, both GPUs would benefit from the faster lanes.
For that you will need motherboard that support PCIe bifurcation. Though, I don't think there is a way to convert 1x PCIe 4.0 x16 to 2x PCIe 3.0 x16 at the moment. What we can do now is either 2x PCIe 4.0 x8 or 2x PCIe 3.0 x8.
The processor supports 24 PCIe lanes. 16 lanes are assigned to the first PCIe X16 slot, 4 lanes connect to the M.2 connector, and another 4 lanes go to the chipset. The chipset can generate up to 8 PCIe lanes, though this depends on the manufacturer's setup. It's possible you have a second PCIe X16 slot, but it may only function as PCIe X4 electrically. If your device doesn't recognize PCIe 4.0, it will automatically switch to PCIe 3.0, maintaining compatibility at that speed. You won't gain extra lanes or similar benefits. A diagram below illustrates this configuration.
GPUs are still limited in bandwidth usage. Even with a custom motherboard and controller setup that channels all CPU PCIe lanes to the GPU, you'd only notice a minor change in frames per second.
No one wants to change anything like this. I'm just trying to grasp how PCIe lanes function and whether splitting them among multiple GPUs through "bifurcation" (thanks krakkpott) is feasible. I don't think a GPU requires extra bandwidth. My goal is to create a rendering setup for Blender using eight GPUs, each running at double 4.0 speeds or triple 3.0 speeds. There are many choices for x1 speeds, but testing GPU rendering at 3.0 speeds on a x1 to x16 riser showed some limitations. The main issue is minor, so two lanes should resolve it. Frame rates aren't the concern since each frame takes a long time. Thanks everyone for clarifying—I really appreciate your help in understanding this better.
The connections are made straight to the CPU, which contains just 16 lanes that need to be divided among the units. These lanes aren't sent through a splitter or switch, as that would significantly raise the price for a typical consumer setup.
The original text discusses rendering performance and mentions a blog post about using multiple GPUs or commercial GPU servers. It suggests exploring options for faster rendering.