Newb Hardware Questions

Hello,
Been a while since I’ve done anything with graphics/rendering (POVray, back in the Win98SE days), appears I am going to have the time to get my head around Blender, so taking a couple of cautious steps forward, beginning with a potential hardware upgrade. There is a LOT of information, am feeling rather inundated ATM, thought perhaps a faster route to wisdom would be to pose a few questions; any experienced response is appreciated.
My current workstation is a Corei7-920, 4G RAM, nVidia GT520 (fanless). I looks like it is plenty of machine for the shallow end of the learning curve, but is already outclassed even by a lot of newer, mid-level silicon. I can kind of rationalize a move up to a socket 2011 system, even before indulging Blender, but would prefer to leave an upgrade path to significantly aid Blender’s performance without having to do yet another CPU-mobo upgrade.
FWIW, my bias is for Intel, nVidia, and Linux, am not likely to stray too far from that combination any time soon.

CPU vs GPU :
So, I’m not completely certain of the roles the CPU & GPU play (perhaps due to my lack of knowledge of GPGPU environments/software). My current understanding is that, while working the interface I can use “cycles” to render my scene, and that I can choose either the CPU or the GPU to render the image, but only one. Is that correct? Is it also the case when rendering the completed project? GPUs seem to reign supreme in the interface, probably no CPU would touch a decent card/SLI combo.

SLI & PCIe Bandwidth :
I’ve noticed very few socket 2011 motherboards currently support full 16x bandwidth across 3-way & 4-way SLI configurations. Once I look beyond two video cards, it seems most boards start sharing the bandwidth on one (or both) of the PCIe channels. Does Blender, rendering a complex scene, tend to use most of the PCIe bandwidth? Would running four cards on shared bandwidth be bottlenecked, compared to running four cards on four dedicated 16x channels, or is bandwidth a non-issue when rendering?

CPU Instruction Sets :
Do the new 4.1/4.2 SSE instruction sets on in the current crop of CPUs make a difference? For instance, given a P4@3.0GHz, and a Corei7@3.0GHz, each running one thread, does the enhanced instruction affect rendering speed?

Leveraging Unused Systems :
Are render farms across non-homogeneous systems used/common? I’ve got a wildly varied collection of systems rotting on a shelf, wouldn’t mind putting them to use on a final render, if/when I ever get around to large-ish projects. I seem to recall my last POVray animation taking a month on my 400MHz Celeron…

A bit of clarification on those things would sure help. TIA.

Do the new 4.1/4.2 SSE instruction sets on in the current crop of CPUs make a difference? For instance, given a P4@3.0GHz, and a Corei7@3.0GHz, each running one thread, does the enhanced instruction affect rendering speed?
Get an optimised build from www.graphicall.org

Ah good, thanks for the replies. I believe I understand the process behind rendering a little better now. I had thought the video cards would be communicating with the system RAM in much the way the CPU does, but that kind of manipulation is strictly done within card’s RAM (which is why more RAM on the card is better). So, in the examples from the spreadsheet, if a GTX580 takes a 60 seconds to render a frame, any other cards on the PCIe bus would have that entire time to grab their textures & mesh data. Worst case scenario, if I have several cards and they all try to access system RAM at the same time, they briefly share the bandwidth for that short transfer. That potentially opens up a whole group of less expensive motherboards, good info to know.

mib2berlin - may trouble you with another video card question? You suggest getting a single GTX580 to do the heavy lifting. My local computer store sells a few different high end nVidia cards, and I can get two GTX560Ti w/2G RAM each for the same price as the GTX580 w/3G RAM.
Two GXT560Ti = 384x2 = 768 CUDA cores @822MHz across 4G RAM
One GTX580 = 512 CUDA cores @772MHz on 3G RAM.
The specs would seem to indicate more performance out of an SLI configuration, but this is unknown territory to me. I see no SLI systems in the Blender Cycles Benchmark spreadsheet - is there a reason to avoid SLI with Blender?

From my rather limited knowledge; yes, two cards are faster than one(Does Cycles support two cards yet?).
The issue is more to do with memory. Nowadays 16 GB of RAM for rendering is pretty standard, so while the 580 may be slower, the whole scene will be loaded into memory for each of the other two cards, so the 2GB RAM will fill up pretty fast. Each card needs enough RAM for the whole scene, they can’t share half each, but the rendering would take half the time.
So as long as you are rendering simple scenes 2GB will be enough, but load a few more hi res textures or an .hdr and you will soon want more available memory.

You could not use SLI for CUDA calculation (Cycles). As organic says:

the whole scene will be loaded into memory for each of the other two cards

Another problem is, your system need about 200-300 MB vram for display. This reduce the available vram for CUDA to 1.7 GB.
If you use high res textures (4096) and high poly meshes often you get problems. The textures has to unpack into the vram.
Cycles scales not 100% ATM., this went better in the next blender versions.
A big + for to cards is smooth working with one card for display and one for setup/preview and switch to both cards for final render.
If you work with one card the system lagging a lot during rendering as cycles use the GFX card nearly 100%.
As i am not the high poly guy i would go for the two card setup. You have to work economical with textures an mesh. For example, if you use a texture twice it is only one time in vram.

Cheers, mib.

True, Cycles does not scale 100%, but you can cheat your way around this to improve multi-gpu performance for animations. For example, you can open your scene in two different instances of Blender, then set one instance to render the first half of your scene with one GPU and the other instance of Blender to render the second half of the scene using the other GPU. A bit messy and not optimal if the complexity of the scene is not the same in all frames, but at least it’s doable and you can get 180%-190% performance.

@Danux: Did you overclock your i7 920? Takes a minute to get another 1ghz on air. Another upgrade path would be the Intel Core i7 Extreme Edition 990X - though financially perhaps not ‘the best’ option :wink:

When buying the 5xx card(s), do not be surprised when the opengl performance is quite bad - remember, the opengl performance of the 4xx and 5xx consumer cards is abysmal in Blender. You may want to get an older 285gtx to drive your screen, and a good 3gb 580gtx to take care of the rendering.

Thanks for all the input. I’m going to hold off jumping in to a new motherboard and CPU, as my current system supports dual-SLI, and it would appear that GPUs are yielding the best performance. This was my most immediate concern. Using a card with copious amounts of RAM also seems like the smart direction, and configuring it with SLI is available, if I feel the need, later on. I’ve also got a couple of empty PCI slots, so perhaps I’ll drive my monitor with a Zotac ZT-40605-10L GeForce GT 430 (or some other PCI card) and keep bumping the GPUs up in the PCIe slots. Lots of options without getting into a whole new system.

@Herbert123 - AFAIK, my CPU employs the “turbo” function, so it will overclock when it is able. I used to overclock a lot, but I just can’t be bothered any more. Bumping up to a new socket1366 CPU costs the same as getting into a whole new socket1155 i7 with twice the performance, and a better socket1366 CPU costs the same as a socket 2011 i7-3930k with over three times the performance (but necessitates a new mobo and RAM). If I have to revert to employing CPU power, I believe I’ll have to abandon this socket, as Intel has. It sounds like GPUs rule the roost anyway, so that hardware is getting pushed on to the back burner. As you and others have suggested, I’ll probably start with a good GTX580 (when the time comes) and explore my options from there. It’s too bad they don’t make video cards with memory slots, so we could bump them up.

Thanks again, I’ve got enough to press on here, will lock the thread down (if I am able?) unless others want it left open.

Three tips concerning SLI:

  1. You don’t need SLI to run a multi-gpu enabled renderers such as Cycles or Octane. They natively support multiple GPUs.

  2. It’s been debated among CUDA computing enthusiasts that CUDA performance will actually decrease when SLI is enabled.

  3. SLI will make your GPUs function as one single device. If you don’t enable SLI, your computer will detect each of your GPUs independently. This way, for example, you can tell CUDA enabled apps wich GPU to use for a given task. You could have for example two Blender instaces rendering two different scenes at the same time if you have two GPUs.