Sitemap

TIs the Mac Mini M4 Cluster the Ultimate Machine for Running Large AI Models?

9 min readJan 8, 2025

--

The idea of running a Mac Mini M4 cluster instead of an expensive GPU cluster is exciting, but does it really work?

Press enter or click to view image in full size
The idea of running an M4 Mac Mini cluster: Alex Ziskind (Creative Common Youtube)

What if I told you that you could actually turn a bunch of tiny 4 Mac Minis into a full-fledged, power-efficient, and (sort of) budget-friendly machine learning powerhouse known as cluster. Some of you might think I’m joking, while others are already looking into how to make this great idea a reality.

Well, Machine learning is everywhere these days. The way i feel about it is, It’s in my phone, my car, and probably even when i am recommending which pizza toppings I should order next. But in real world, when it comes to running complex ML models, the cost of the hardware often feels like a punch to the wallet.

There are GPUs like the Nvidia RTX 4090, that are counted as the heavyweights in this game, but they come with the price tags that could make a grown man weep, they’re out of reach for most of us. So, the idea is to set up a high-performance affordable machine learning cluster.( Trust me, a cluster of relatively cheap machines isn’t cheap either), I mean, you might still need to raid your savings for pizza money, but you get the point. So what do You do.

What I Discovered About M4 Mac Mini Clusters?

As I was thinking and researching about the Mac Mini M4 cluster, I came across Alex Ziskind’s grand experiment of setting up a cluster of Apple M4 Mac Minis to run machine learning models.

Its actually a “Creative Commons Attribution license (reuse allowed) video, so, i am sure, neither Alex Ziskind nor Medium will have an issue while I am discussing and commenting on the findings.

Again, you might be thinking, “Mac Minis? Those cute little desktops?”, aren't the clusters (actually) big bulky machines tied together ?

Well, in case you don’t know M4 Mac Mini, this isn’t your average “I just need something to check my email” device. The M4 Mac Mini is a little beast (on its own right) with a base 16GB of RAM at very cheap price (even cheaper with 17% off now on amazon), and this experiment is about to show you that small doesn’t always mean weak.

Why Not Just Use a Powerful GPU?

Let’s talk about GPUs for a second. If you’ve been in the machine learning world, even for just 10 minutes, you’ve probably heard that GPUs are the kings when you are running parallel processing.

GetForce RTX 4090 AERO GPU
GetForce RTX 4090 AERO GPU: Alex Ziskind (Creative Common Youtube)

You can say, GPUs are the bodybuilders of computing, pumping out data like it’s nothing. Meanwhile, if you take the CPUs, they are more like your average office worker, they try their best, but they’re just not built for heavy lifting (Ok, may be underestimating CPUs a bit, but GPUs are very important in parallel processing). And sure, the Nvidia RTX 4090 might be able to bench press your entire ML workload, but it comes with a scary price tag that’ll make your bank account cry (you can check the GPU prices on amazon).

Press enter or click to view image in full size
Mac Mini M4 Cluster
Mac Mini M4 Cluster: Alex Ziskind (Creative Common Youtube)

So whats the alternate? well, in this case we have the Apple M4 Mac Mini. Imagine you have a laptop and a desktop had a baby that was both powerful and adorable. Apple’s has this unified memory system that actually allows the CPU and GPU to share memory, which means it can handle tasks more efficiently than you might actually expect. So, it would be nice to forget about dropping thousands on GPUs that need their own room in the house, these Mac Minis are very compact, efficient, and if every thing comes off, this might be the setup we need.

Alex Ziskind, his YouTube vide, actually made a solid point that Apple’s M4 chips are surprisingly capable for running ML tasks. Well, to me, it’s not going to win any beauty pageants against those massive GPUs at least for now, but in terms of cost-to-performance ratio, it’s got a solid chance of winning the underdog award.

MLX, The Emergence of Apple Silicon for Machine Learning

To write the machine learning code, you would need a library or framework. In 2023, Apple dropped the MLX framework , basically their secret sauce for machine learning, basically optimized for Apple silicon. This framework is like a finely tuned engine that has the capacity to squeeze every bit of performance out of Apple Silicon chips. Benchmarks have shown that it even performs better than PyTorch (you know, that little ML framework everyone swears by) on Apple hardware.

So,whats the main idea behind the use of MLX here

Press enter or click to view image in full size
Running Modules in clustered setup : Alex Ziskind (Creative Common Youtube)

The MLX framework will help M4 Mac Mini to run machine learning code and since machine learning will perform better with a parallel setup(presumably), Alex Ziskin thinks that running parallel Mac Mini M4 machines will distribute the load and we’ll have a better performance overall.

What’s even more mind-blowing is that Apple’s Neural Engine, which was designed with smaller ML models in mind, is now getting a massive performance boost. It’s like Apple went from making a family sedan to a race car, all without adding extra fuel consumption. If you’re into machine learning (in capacity), looking to avoid those expensive GPU setups, Apple Silicon is probably looking like a legitimate contender.

So, it would be great to shout-out to Alex Ziskind for giving us the inside scoop on this , and for showing us that maybe it’s time we all rethink our obsession with GPU powerhouses to certain extent.

Building a Cluster: Does More Mean Better?

Press enter or click to view image in full size
running M4 Mac Minis in parallel
running M4 Mac Minis in parallel : Alex Ziskind (Creative Common Youtube)

So lets discuss the fun from where it starts. We have Mr. Alex Ziskind who decided to take things to the next level and build a cluster of five M4 Mac Minis, (he’s rich or a techie (like me) who doesn’t actually own all these devices, but somehow his job lets him play with them anyway). He connected these machines via thunderbolt ports. The idea was to distribute the load across multiple machines and potentially make things faster.

Press enter or click to view image in full size
Running a small ML module on M4 Mac Mini

The first test was running a small ML model, Llama 3.21 billion parameters on a single base M4 Mac Mini. It handled it like a champ, you have to say, pumping out around 70 tokens per second on average.

While running the same on M4 Pro Mac Mini (alone) it gave around 95 token per second. Ziskind continues on with the experiment and tries out different combinations. Here are the few results, he got.

Performance Results:

Running Llama 3.21 billion parameters (smaller model):

  • Single Base M4: 70 tokens per second (sustained).
  • Single M4 Pro: 95–100 tokens per second.
  • Two Base M4 Machines: 45 tokens per second (due to network bottleneck thunderbolt hub).
  • Two Base M4 Machines (direct Thunderbolt connection): 95 tokens per second.
  • Five Machines (Clustered): 67–74 tokens per second.

Larger Models:

  • 32-billion Quin model: 8 tokens per second on a base M4, 12 tokens per second on M4 Pro.
  • Neotron 70B (70 billion parameters): 4–5 tokens per second on M4 Pro machines.

I am going to discussed what all those result mean and what are the takeaways, but let me discussed something very promising, the power efficiency.

Power Efficiency: The Hidden Benefit of the Cluster

Talking about the power, you might think that running five Mac Minis would suck up the kind of electricity that could power a small whole village, but surprisingly, it doesn’t. Even at full load, the total power consumption for the entire cluster was only around 200 watts. Now, compare that to a single high-end GPU (very expensive anyway), which could easily eat up 600 watts or even more than that.

This is where to me, the Mac Mini cluster really shines. If you’re running ML models at home, avoiding that second mortgage to pay your electric bill is a huge win. So, while you might not be winning any races in terms of raw speed, you are saving some money, and your power bill will thank you.

How cheap is Mac Mini M4

The most exciting thing about the M4 Mac Mini is its base Model with astonishing specs at a very cheap price (even cheaper with a huge 17 % off now on base model). A Mac mini M4 with 16 GB base RAM and the latest M4 chip with 10 core CPU and 10 core GPU performance is very exciting, although things do get started going wild, when you look to upgrade the RAM or the storage particularly.

Press enter or click to view image in full size
Mac Mini M4 Base Price and Specs: Source Apple

Its very cheap, See if have a discount deal running on Amazon right now, to get further discount. If you are confused with the price and feature mechanism, You can find how you can customize your options to find the best deal, and save a lot of money instead of just going wild about the options.

The Takeaways: Is This Practical?

After all the testing and tinkering, Alex Ziskind came to a few very sensible conclusions:

  1. For smaller ML models, a Mac Mini cluster doesn’t offer much more performance than just using one device.
  2. For larger models, more powerful machines like the M4 Pro are going to outshine clusters of base Mac Minis.
  3. Power efficiency is definitely a bit of win

So, what’s the takeaway here? Well, while a Mac Mini cluster might not be the holy grail of ML computing, at least now for now, it’s a pretty cool alternative for anyone looking to experiment on a budget. It’s cost-effective, it’s efficient, and it’s got a certain “David vs. Goliath” charm about it. But let’s not kid ourselves , if you’re planning to tackle truly large-scale models, you’re probably going to need something more powerful than a bunch of little Minis (fair enough they are little beasts in their own right).

Closing Thoughts: The Future of Mac Mini Clusters in ML

In conclusion, Alex Ziskind’s experiment has made it clear that Mac Mini clusters aren’t quite the future of machine learning (yet), they’re certainly worth considering if you’re on a budget or just enjoy playing around with alternative setups. Even in his concluding remarks, Alex said, so far he is not convinced, its a greater idea to run a cluster of Mac Minis, although the concept is great.

“For me, as someone who has already been researching the practicality of Mac Mini M4 clusters, seeing the results others are experiencing has been very insightful. Based on this practical demonstration, I can confidently conclude that using a single Mac Mini (base or Pro, depending on my specific needs) and taking full advantage of Apple’s unified memory system will likely yield better performance than trying to run a Mac Mini cluster, for now. That said, the concept behind clustering is exciting and makes me optimistic for the future. Perhaps Apple will eventually introduce some kind of a built-in feature that allows seamless integration, where we can simply add devices (like Mac Minis) in a cluster, and their specs would automatically sync, functioning as a true cluster with minimal effort, just plug and play, But, how Apple would handle the marketing strategy and manage the financial impact of this clustering thing, it remains to be seen.

So, what do you think? Have you ever tried setting up a cluster for ML? Or are you still waiting for that dream GPU setup to fall from the sky? Let me know in the comments — I’d love to hear about your own experiments, if you had any.

Note: As an Amazon Associate, I may earn a commission from qualifying purchases made through the links provided, without costing you anything extra. This helps support my content and keep it free for you. Thank you for your support!

--

--

Faizan Saghir
Faizan Saghir

Written by Faizan Saghir

I am an IT Expert, who loves to talk and compare latest tech, based on specs, features, benchmarks, real time testing, dimensions and colors etc.