Discussing PCIe bifurcation and optimizing a 3x GPU setup. Looking for expert tips on Mobo advice.
Discussing PCIe bifurcation and optimizing a 3x GPU setup. Looking for expert tips on Mobo advice.
The lane configuration is permanently set in the board, making manual adjustments impossible no matter how many M.2 slots you employ. The manufacturer would need to deliberately exclude CPU lanes to shift them to a separate graphics slot, which isn't currently needed. It seems Epyc 4004 might allow this, but it doesn’t seem practical for consumer boards. I doubt there’s a solid reason for it. When ASUS refers to bifurcation, they actually mean dividing the lanes within an existing hardwired slot, not reallocating them from the CPU. If you install a Hyper M.2 card with several slots, you must divide the PCI-E lanes to ensure each slot functions properly—typically splitting 16 lanes into four per slot.
I'm not comfortable with the location, but I think you're right about not pooling VRAM across GPUs. Each device handles its own memory independently. Using a single GPU will hit PCIe bandwidth quickly. The 4060 Ti with 16GB offers about 288GB/s internally, while over PCIe it's only 16GB/s. So the 16GB should cover what fits on that card. The 4070 TiS is roughly twice as fast, meaning you'll complete roughly double the tasks in the same time compared to one 4060 Ti 16GB.
Surprise finds you. See how this developer performed with a 6,060 Ti 16GB on a Threadripper. If that’s impressive, explore his YouTube channel for ML/DL insights. For the 4060Ti 16GB, it’s clear it struggles with gaming performance, though AI fans are usually open to it. VRAM and CUDA matter most for handling bigger models efficiently. At least from what I’ve read, speed isn’t the only factor. You’ll need more VRAM to run larger models smoothly. It’s possible to limit yourself to a certain VRAM for faster execution of smaller models—maybe gaining a few seconds versus bigger ones. Of course, you can trade off speed for smoother interaction, like in a game where response feels sluggish but still functional. Updated June 15, 2024 by Dennettic
Speed isn't the only factor. When working with ML/DL and AI creation, you need capacity to handle bigger models and data sets. A faster 4070 ti might not give you a noticeable edge over the 4060ti 16gb, but more total RAM lets you run larger models. With just one 4070ti, you're limited to models that fit in 16gb; two 4060tis open you up to 32gb. The performance gaps seem small for everyday tasks, but for massive datasets used by big companies, every bit matters.
That's a shame, but not surprising. I was only bringing it up because they're vague claim that PCIE gen5 will be "Standard" on all mobos. Wishing for more but expecting less seem to be a safe bet for these manuf.s
check if all your GPU needs sufficient bandwidth or you can manage with just two lanes for the cards used in machine learning. If you really need optimal bandwidth, consider high-end options from the Epyc 7002/7003 series with a motherboard that offers most lanes in PCIe slots. That way, simply install the cards and it works fine. The downside with Epyc is that you’ll also need audio and USB add-in cards, since the platform’s choices are limited by design. You might also explore Threadripper, though prices are quite high.