Discussion about this post

User's avatar
Ruslan's avatar

The first C&C post which i barely understand 🤯 That low precision large matrix multiplication acceleration bs completely passed me by. Does it has any nonAI use cases?

Peter W.'s avatar

Some examples for applications when these different approaches for the different approaches of ARM vs. X86 vector and matrix implementations would be really helpful. Apparently, SVEs saw a lot less utilization than ARM had envisioned, which is not the case for SMEs. AFAIK, in the case of SVEs, Qualcomm ended up not enabling SVE in their stock-ARM cores (e.g. SD 8 Gen 3), because they figured nobody would actually use it in mobile devices. They did have the circuitry in their big cores (due to the ARM license), but simply let it lie dormant.

How is that now, both in Qualcomm's Oryon cores and in Apple's designs? For the latter, which applications use the 512 bit SMEs that the M CPUs have?

Sorry for all these questions, I find concrete use cases helpful to "get" those more principle differences.

And thanks for another deep dive!

No posts

Ready for more?