Qwen 35B

Explore how AI models can now run directly on your mobile device.

Running a private Qwen 35B MoE model on a Galaxy S26 Ultra is no small feat, and the early numbers are worth noting.

4 min readMachine Learning

A Reddit post about running a 35B parameter MoE model on a Samsung S26 Ultra is the kind of thing that makes you pause. Not because the numbers are staggering, though an 8-token-per-second output on a phone is nothing to dismiss. What stands out is the framing: no formal PhD, just interest, compute, and a willingness to test. That is the quiet revolution happening in AI. The gatekeepers are still writing papers, but the people actually pushing hardware to its limits are often the ones tinkering in their spare time. This story is less about the specific model and more about who gets to claim expertise anymore.

The post also carries a familiar frustration: four papers stuck on arXiv because the author lacks an institutional affiliation. That detail will resonate with anyone who has tried to enter a field where credentials still matter more than demonstrated ability. We have written before about the strange experience of talking to an AI clone and how it forces you to reconsider what is real. This is the inverse problem. The author has the technical chops, but the system refuses to validate them. Meanwhile, the Forrester function discussions in our community show that curiosity-driven research often outpaces formal coursework. The lesson is not that academia is worthless. It is that the barrier to meaningful contribution has dropped so low that the old validation loops are starting to look obsolete.

What should you take from this practically? First, if you have been waiting for permission to experiment, stop. The author did not wait for a lab or a grant. They used what they had, tested a real constraint, and shared the results. That is a playbook, not a special case. Second, the memory footprint question is the one to watch. Fitting a 35B model on a phone is not just a technical trick. It is a signal that the next wave of useful AI tools will not live in the cloud. They will live in your pocket. We have touched on this in pieces like verifying AI understanding, where the focus is on practical reliability rather than raw capability. The numbers, if they hold up under replication, point to a future where private, on-device inference becomes the default for sensitive work.

Here is the open question we would leave you with. The Reddit poster is not sharing methods, which is their right, but it also means we cannot verify the claims or build on them. In a field that prides itself on reproducibility, that is a tension worth sitting with. If you are reading this and have the resources to test something similar, do not just run the model. Publish the methodology, even if it is on a personal blog or a forum. That is how trust gets built. The author proved they can do the work. The next step is proving the work can be peer-reviewed by people who care, not just by people with titles. That is the shift we are watching.

From Machine Learning

Started testing a private qwen 35B moe capacity LLM runtime on s26 ultra, early testing shows that active model footprint can fit within the device’s memory limits.( not sharing the methods or architecture used) and results suggest roughly 90 input processing t/s achievable after optimisation and output generation is around 8 tokens/s on this mobile.

Point is i learned ai ml based on my interest and no formal PhD , I have compute and resources to test. Anyone willing to join or collab to test on this

Read the original at Machine Learning