Mac: M3 - *Hardware accelerated RT (Part 1)

Very good points

All this need to expand stuff I think is maybe more for gamers needed who swap out gpus or storage

I bought my MacPros added the ram and done and over the years I bought twice different gpus.

But the Mac mini is fine for many tasks and I feel for many it is the same who have normal needs. The m1 is as good as my 12 core Xeon while old that was a workstation.

Yea it sucks ghat I cannot add more ram to the Mac mini now but I would have also bought the 16 GM version if it would be my own (it is a loaner).

Android and iOS is like so I want to play with widgets or just work and be fine with the simplicity.

GPUs also from Nvidia or amd make great moves forward so Apple has to catch up with their own dev and then the advancements of the more powered gpus.

Actually now that you say it, it actually might be a good thing Apple pushing more towards gaming as it might benefit Blender users too.

2 Likes

Yes, the salvation for CG artists will come from the gaming world as the technology and needs overlap greatly.

For instance, Metal 3 accellerates ray tracing rendering, so the two main questions are — can this be applied in Cycles and will it make the need for actual RT cores less critical?

3 Likes

Cycles could use the Neural Engine to speed up raytracing calculations, which is functionally fairly similar to the Tensor Cores on the newer Geforces.

I don’t think it’ll be faster than having dedicated RT cores (I’m not even sure what it is that makes RT cores so special to begin with), but it could give us a nice speedup. The only real question here is whether they will.

1 Like

I do not see much problems with standard RAM extension.
Just more difficult for developers to address memory correctly and useful.
Kind of an Apple ARM optimization like optimizing for Tiled memory, but
easier ?
Or simply as Apple ARM uses swapping extensively with their memory
limited devices.
I see slotted RAM extensions as primary swap memory, just still much
faster than a fast SSD.

Until today I am still not sure if the M1 Mini 16 GB lags in my 3D Apps
when using more memory than available is really the bottleneck.
(when memory pressure gets brownish)
As swapping seems to do not have much impact in Video, Audio and
Photo Apps.
As I do not have one, I can not say if it would really work better with a
128 GB M1 Ultra.

I think all acceleration is just by rendering a smaller image frame to get faster
and do an AI scaling to bring the final full scale frames at a nearly 100% quality
like Nvidia and meanwhile also AMD does.
Also Denoising will help by rendering a faster but noisier image which will be
faster corrected or optimized by AI algorithms.

There are some other tricks and approaches like tiled rendering which allows
M1 to render larger scenes than it can and such,
but there is still the RT process that could run faster on dedicated cores like
Nvidia and now AMD too use.

Hmmh, again one of Apple’s “worst” keynotes.
(The first “worst keynote ever”, again since 2020 ?)
The first hour really was a terrible lost of live time.
I thought I get amused when M2 started but it was as rumored.
M2 is just a slightly improved M1.
Only 18% or so speed improvement after 2 years !?
That is not the typical Apple iOS ARM pace we were used to.
More Intel’s initial competition less pace from 2007 to 2021.

I hoped for the memory increase of the basic M2, but 24 GB is a bit
underwhelming for me personally. I would have been satisfied with 32 GB.
So I could buy another interims M2 Mini to overcome another 1.5 years
until I finally can buy an ARM Mac that is really capable of doing my work
for another few years.

Nevertheless,
typical YouTubers or maybe even the majority of Apple users seem to be
very exited about all the WWDC announced iOS, iPadOS and macOS
features and the bit of hardware as I have seen so far.
(They could not show a Mac Pro teaser as they had so much other stuff to show …)

To give them some benefit of the doubt here, I’d say that most of the biggest improvements Apple has made to ARM have already come via the A chips, of which we started reaping the benefits of a couple years ago.

Plus, there’s only so much they can do with the same die size. 18% single core improvement won’t melt any brains with its awesomeness, but it’s not exactly terrible either.

I recall many stress tests showed that the m1 ram was at times not enough vs the macpro or an intel pc doing massive video edits

I am not sure if it was 8 or 16 gb ram max minis.

The m1 ram like with with iPhone does not need to be as big as with intel too.

Also to be fair this was a software keynote and now they run the workshops

They did some significant changes to ARKit

It can use 4K videos now - that is huge for photogrammetry !

Also room scan will be interesting.

The other iOS improvements are well “was about time” and some of the macOS are actually quite interesting new tools.

The metal 3 news is also good because apples but iOS apps will run on macs did not turn it well

And I don’t want iOS games or macs

Triple A titles would be better

1 Like

Yes, but Apple would not have reached to where they are with A chips
if they had only Intel Pace. There were lots of diagrams where Intel had
only about 15% speed increase per year/generation
(which makes some years - to get a twice as fast CPU - which is about
the least what you can really notice while working as being faster)
while Apple had about 25% increase and so was able to soon take over.
So Covid and Lockdowns aside, 18% for 2 years or one Apple generation
is pretty poor in my eyes, as even Intel was faster, now after 3 (?) years
of being overrun by AMD.

Apple meanwhile brought their ARM SoCs clearly to desktop level,
which is great.
Apple never wanted to lead in raw CPU and (unfortunately) GPU power
but efficiency - which is great.
But Intel (and AMD) seems to be able to deliver better efficiency relatively
fast if they see a market/competition beside their previous raw power
king of gaming at whatever energy cost route they had before.

And AFAIK Nvidias over all GPU render per power is so far not that much
worse than Apples efficiency. Yes, more power need in general but you
also get much more performance.
And if Nvidia (or even AMD or Intel) start to create ARM CPUs or SoCs
seriously for the same purposes, would Apple still be able to keep its
efficiency advantage for very long ?

Apple was thinking about and eliminated bottlenecks, which is great in my
eyes and which no one in X86 seems to have ever felt a need for a change.
But Apple also expects or needs developers to do some extra work to
follow and also optimize their software regularily.
Which in my Mac Pro 5.1 2013 world was already pretty unlikely to happen
in our 3D environment.

Apple does great things, right and at first, in my eyes.
But if AMD, Intel, Nvidia and maybe Qualcom and such somehow decide
that Apple’s way would be the trend to follow, they would easily outpace
Apple and Apple had to do another switch in a decade to adapt to a
extern hardware provider to be able to deliver relatively competitive performant
hardware.

For now Apple M is great with performance and efficiency for Video, Audio,
Photo and overall. Just not for 3D where many would forget about their
power bill to just get their jobs done with a desktop like environment.

1 Like

I am not sure.
The most logic theory for me is that :
any non*-tiled memory* optimized 3D App
(and tiled memory optimization being a larger effort not likely being done)
may overflow the small M1 cache and lead to lagging.
and it will lag.

I am so far not sure if M2, beside some more GPU cores and hopefully
better A15 like cores in M1 design, will also eliminate that too small cache
problem, resulting from a far pre M1 design of A chips,
will have more cache, eliminate that bottleneck if an App is not tiled memory
optimized - which is highly likely …

If M2 does, it would be great.
i would take the extra 8 GB of memory and the 2 extra GPU cores of the M2
as a bonus and buy another M2 Mini while I wait for the real Mac thing.

Given that I’d consider the M2 to be more about the GPU than the CPU, I’d say it’s a fairly decent upgrade on that front. A 35% increase is fairly stout, and could scale even better with the eventual Pro and Max revs to come.

…though I’m still waiting for some benchmarks before I proclaim it the new hotness. Oh, and some RT cores would be REALLY nice right about now.

2 Likes

AMD zen 4 will really show how efficient Apple is as they are 5 nm TSMC too.

2 Likes

RT cores (just like Tensor cores) are a form of ASIC (Application-Specific Integrated Circuit), They do one thing and they do it super fast.

  • Tensor cores do Tensor math operations (AI/ML, like image up-scaling or de-noising, etc…)
  • RT cores do Ray intersection calculations (did the ray hit an object or not ?) the result is sent back to the general purpose cores (CUDA) for the shading phase.

For a more deep dive into how exactly those ray intersections are calculated check this article:

And this article comparing NVIDIA RT cores to AMD’s ray accelerators (they operate a lil bit differently):

On that note, I wouldn’t hold my breath for some form of RT acceleration using apple neural engine as that one is specifically built for AI/ML stuff and trying to make it do other types of calculations would defeat the purpose of it being good at one thing and wouldn’t be different from any other general purpose core.

Here is a comparison between the M1 Ultra Mac studio neural engine vs the RTX 3080ti tensor cores.

Hope this helps.

1 Like

I am guessing RT but s coming, seeing this in their Metal 3 presentation.

Should make Cycles faster I guess.

The other presentations are not up yet but the overview sounds promising.

5 Likes

Yes, and of course what real App optimizations will do.
Which is of course one of the largest bottlenecks and a huge potential over
Win X86 - if it will really happen …

I think Apple ARM CPU core are pretty competitive, just too few if that is what
you are looking for. I think even GPU pretty ok too but maybe also too few so far
but also highly dependent on developers ambitions …
While X86 may work still great by just putting raw power at every cost over
bad software.

1 Like

would definitely be nice if those metal3 changes can make their way to cycles metal.

I know BMW is a bad benchmark, but as someone who does animation and often is gunning for sub 1-minute frame times, I’d love to see those those ~50s render times get down closer to ~30s on my m1 max.

2 Likes

Honestly baby brain is setting in here and what I thought more being a ram issue is more like what you say is a cache issue

A 4090 might make the BMW quasi real time.

My 3060 laptop 2615 overall score 12 sec in BMW
current 3090 6000 score so if linear speed advantage 4-5 sec
4090 said to be almost double performance 2 sec?

I am not sure but for me it is the most reasonable explanation I heard so far.

And I did not get if that will be addressed with M1.5 or maybe only with real M2
in 2023.

If I had seen benchmarks for M1 Ultra or Max that showed a reasonably scaling
over M1 or Pro and could trust that concept I would have already ordered a
BTO Studio around 5k and would just use it for the coming years.
(Or still wait for its delivery)
But somehow Studio still feels strange to me - and WWDC keynote did not help
much to clear things or, as expected, a roadmap where things are going.
on that