F5F Stay Refreshed Hardware Desktop Discussing PCIe bifurcation and optimizing a 3x GPU setup. Looking for expert tips on Mobo advice.

Discussing PCIe bifurcation and optimizing a 3x GPU setup. Looking for expert tips on Mobo advice.

Discussing PCIe bifurcation and optimizing a 3x GPU setup. Looking for expert tips on Mobo advice.

Pages (3): Previous 1 2 3 Next
K
Kynedee
Posting Freak
784
03-07-2023, 09:05 PM
#11
The lane configuration is permanently set in the board, making manual adjustments impossible no matter how many M.2 slots you employ. The manufacturer would need to deliberately exclude CPU lanes to shift them to a separate graphics slot, which isn't currently needed. It seems Epyc 4004 might allow this, but it doesn’t seem practical for consumer boards. I doubt there’s a solid reason for it. When ASUS refers to bifurcation, they actually mean dividing the lanes within an existing hardwired slot, not reallocating them from the CPU. If you install a Hyper M.2 card with several slots, you must divide the PCI-E lanes to ensure each slot functions properly—typically splitting 16 lanes into four per slot.
K
Kynedee
03-07-2023, 09:05 PM #11

The lane configuration is permanently set in the board, making manual adjustments impossible no matter how many M.2 slots you employ. The manufacturer would need to deliberately exclude CPU lanes to shift them to a separate graphics slot, which isn't currently needed. It seems Epyc 4004 might allow this, but it doesn’t seem practical for consumer boards. I doubt there’s a solid reason for it. When ASUS refers to bifurcation, they actually mean dividing the lanes within an existing hardwired slot, not reallocating them from the CPU. If you install a Hyper M.2 card with several slots, you must divide the PCI-E lanes to ensure each slot functions properly—typically splitting 16 lanes into four per slot.

D
Dylanhtx
Member
156
03-08-2023, 01:15 AM
#12
I'm not comfortable with the location, but I think you're right about not pooling VRAM across GPUs. Each device handles its own memory independently. Using a single GPU will hit PCIe bandwidth quickly. The 4060 Ti with 16GB offers about 288GB/s internally, while over PCIe it's only 16GB/s. So the 16GB should cover what fits on that card. The 4070 TiS is roughly twice as fast, meaning you'll complete roughly double the tasks in the same time compared to one 4060 Ti 16GB.
D
Dylanhtx
03-08-2023, 01:15 AM #12

I'm not comfortable with the location, but I think you're right about not pooling VRAM across GPUs. Each device handles its own memory independently. Using a single GPU will hit PCIe bandwidth quickly. The 4060 Ti with 16GB offers about 288GB/s internally, while over PCIe it's only 16GB/s. So the 16GB should cover what fits on that card. The 4070 TiS is roughly twice as fast, meaning you'll complete roughly double the tasks in the same time compared to one 4060 Ti 16GB.

D
dt118lw
Member
198
03-08-2023, 01:44 AM
#13
I understand now. Thanks for the feedback. It’s disappointing that no one has shared the physical lanes for x8x8x8 support on any board. There’s potential for more affordable, solid ML/DL choices without major hurdles.
D
dt118lw
03-08-2023, 01:44 AM #13

I understand now. Thanks for the feedback. It’s disappointing that no one has shared the physical lanes for x8x8x8 support on any board. There’s potential for more affordable, solid ML/DL choices without major hurdles.

B
220
03-10-2023, 05:42 PM
#14
Surprise finds you. See how this developer performed with a 6,060 Ti 16GB on a Threadripper. If that’s impressive, explore his YouTube channel for ML/DL insights. For the 4060Ti 16GB, it’s clear it struggles with gaming performance, though AI fans are usually open to it. VRAM and CUDA matter most for handling bigger models efficiently. At least from what I’ve read, speed isn’t the only factor. You’ll need more VRAM to run larger models smoothly. It’s possible to limit yourself to a certain VRAM for faster execution of smaller models—maybe gaining a few seconds versus bigger ones. Of course, you can trade off speed for smoother interaction, like in a game where response feels sluggish but still functional. Updated June 15, 2024 by Dennettic
B
bluehypergiant
03-10-2023, 05:42 PM #14

Surprise finds you. See how this developer performed with a 6,060 Ti 16GB on a Threadripper. If that’s impressive, explore his YouTube channel for ML/DL insights. For the 4060Ti 16GB, it’s clear it struggles with gaming performance, though AI fans are usually open to it. VRAM and CUDA matter most for handling bigger models efficiently. At least from what I’ve read, speed isn’t the only factor. You’ll need more VRAM to run larger models smoothly. It’s possible to limit yourself to a certain VRAM for faster execution of smaller models—maybe gaining a few seconds versus bigger ones. Of course, you can trade off speed for smoother interaction, like in a game where response feels sluggish but still functional. Updated June 15, 2024 by Dennettic

K
KidzBeEz
Member
242
03-10-2023, 08:02 PM
#15
K
KidzBeEz
03-10-2023, 08:02 PM #15

I
IPS10
Senior Member
623
03-10-2023, 11:53 PM
#16
X870E functions similarly to X670E, with the main variation being its support for four CPU lanes in USB4.
I
IPS10
03-10-2023, 11:53 PM #16

X870E functions similarly to X670E, with the main variation being its support for four CPU lanes in USB4.

_
_Hundred
Member
51
03-11-2023, 10:47 PM
#17
People expressing concerns note the lack of an inexpensive HEDT solution with many PCI-E lanes since X99. Currently, you're still required to use threads or xeons for high PCI-E capacity, such as W790.
_
_Hundred
03-11-2023, 10:47 PM #17

People expressing concerns note the lack of an inexpensive HEDT solution with many PCI-E lanes since X99. Currently, you're still required to use threads or xeons for high PCI-E capacity, such as W790.

S
ScrewDumpMC
Junior Member
39
03-12-2023, 02:47 AM
#18
Speed isn't the only factor. When working with ML/DL and AI creation, you need capacity to handle bigger models and data sets. A faster 4070 ti might not give you a noticeable edge over the 4060ti 16gb, but more total RAM lets you run larger models. With just one 4070ti, you're limited to models that fit in 16gb; two 4060tis open you up to 32gb. The performance gaps seem small for everyday tasks, but for massive datasets used by big companies, every bit matters.
S
ScrewDumpMC
03-12-2023, 02:47 AM #18

Speed isn't the only factor. When working with ML/DL and AI creation, you need capacity to handle bigger models and data sets. A faster 4070 ti might not give you a noticeable edge over the 4060ti 16gb, but more total RAM lets you run larger models. With just one 4070ti, you're limited to models that fit in 16gb; two 4060tis open you up to 32gb. The performance gaps seem small for everyday tasks, but for massive datasets used by big companies, every bit matters.

H
horselover328
Member
148
03-12-2023, 04:11 AM
#19
That's a shame, but not surprising. I was only bringing it up because they're vague claim that PCIE gen5 will be "Standard" on all mobos. Wishing for more but expecting less seem to be a safe bet for these manuf.s
H
horselover328
03-12-2023, 04:11 AM #19

That's a shame, but not surprising. I was only bringing it up because they're vague claim that PCIE gen5 will be "Standard" on all mobos. Wishing for more but expecting less seem to be a safe bet for these manuf.s

O
OhSteph
Junior Member
19
03-13-2023, 09:37 PM
#20
check if all your GPU needs sufficient bandwidth or you can manage with just two lanes for the cards used in machine learning. If you really need optimal bandwidth, consider high-end options from the Epyc 7002/7003 series with a motherboard that offers most lanes in PCIe slots. That way, simply install the cards and it works fine. The downside with Epyc is that you’ll also need audio and USB add-in cards, since the platform’s choices are limited by design. You might also explore Threadripper, though prices are quite high.
O
OhSteph
03-13-2023, 09:37 PM #20

check if all your GPU needs sufficient bandwidth or you can manage with just two lanes for the cards used in machine learning. If you really need optimal bandwidth, consider high-end options from the Epyc 7002/7003 series with a motherboard that offers most lanes in PCIe slots. That way, simply install the cards and it works fine. The downside with Epyc is that you’ll also need audio and USB add-in cards, since the platform’s choices are limited by design. You might also explore Threadripper, though prices are quite high.

Pages (3): Previous 1 2 3 Next