Tried testing qwen 35b moe model on s26 ultra , without compromising on precision [R] ,[D]
Our take
The recent Reddit post detailing the successful testing of a Qwen 35B MoE LLM runtime on a Samsung S26 Ultra is a fascinating, and frankly, somewhat improbable development. It speaks to a burgeoning trend of on-device AI processing that's gaining momentum, particularly as we see AI-driven memory crunch jolts India’s smartphone market. The author’s claim of achieving 90 tokens/s input processing and 8 tokens/s output generation, even after acknowledging the need for optimization, is impressive given the constraints of a mobile device. What’s particularly compelling is the self-taught nature of the individual conducting the tests, highlighting the increasing democratization of AI development and experimentation, a sentiment echoed by figures like Neil Rimer who believe Neil Rimer thinks the AI money is coming back out, demonstrating that access to resources, rather than formal credentials, is becoming a more significant driver of innovation. The reluctance to share methodology is understandable, likely to protect intellectual property, but it also underscores the growing prevalence of independent AI researchers pushing boundaries outside of traditional academic or corporate structures.
The broader significance of this goes beyond simply demonstrating a technical feat. It’s a compelling illustration of the direction mobile computing is headed. The current reliance on cloud-based AI processing introduces latency and privacy concerns, and increasingly, users are demanding more responsive and secure experiences. Successfully running a model of this size, even with optimizations, directly on a smartphone suggests a future where sophisticated AI capabilities are seamlessly integrated into our daily lives, without constant reliance on external servers. This aligns with the broader conversation around cloud native infrastructure and its potential to underpin trustworthy agentic AI, where Cloud Native Infrastructure Emerges as the Foundation for Trustworthy Agentic AI becomes increasingly vital for efficient and secure deployments. The author’s challenges in publishing their findings on ArXiv, due to lacking institutional affiliation, also raises an important point about the current academic publishing landscape and the potential for barriers to entry for independent researchers.
This development is particularly noteworthy in the context of the broader AI landscape. While much of the conversation has centered around massive models and computationally intensive training processes, this showcases the equally important work being done in optimizing models for deployment on resource-constrained devices. It’s a testament to the ingenuity of engineers and researchers who are finding creative ways to squeeze maximum performance out of increasingly powerful, yet still limited, hardware. The ability to run powerful LLMs locally on a smartphone opens up a vast range of possibilities, from improved on-device translation and natural language processing to personalized AI assistants that can operate completely offline. It also suggests a shift in the power dynamics within the AI ecosystem, empowering individual developers and researchers to contribute to the field without requiring access to massive data centers or corporate funding.
Looking ahead, the question becomes: how quickly can this type of on-device AI processing become ubiquitous? The author’s call for collaboration is a clear signal of their desire to accelerate progress, and the open-source community will undoubtedly play a crucial role in furthering these efforts. It will be interesting to see if advancements in mobile chipsets, coupled with further innovation in model optimization techniques, will allow for even larger and more capable LLMs to run seamlessly on smartphones and other mobile devices, ultimately blurring the lines between cloud-based and on-device AI and reshaping how we interact with technology.
Started testing a private qwen 35B moe capacity LLM runtime on s26 ultra, early testing shows that active model footprint can fit within the device’s memory limits.( not sharing the methods or architecture used) and results suggest roughly 90 input processing t/s achievable after optimisation and output generation is around 8 tokens/s on this mobile.
Point is i learned ai ml based on my interest and no formal PhD , I have compute and resources to test. Anyone willing to join or collab to test on this
I tried publishing papers on arxiv and 4 papers are still on hold as im first author and from no institution...
[link] [comments]
Read on the original site
Open the publisher's page for the full experience