SineKAN

Explore how sinusoidal activation functions can reshape neural network design.

Some ideas arrive on schedule.

3 min readMachine Learning
Explore how sinusoidal activation functions can reshape neural network design.
[R] SineKAN: Kolmogorov-Arnold Networks Using Sinusoidal Activation Functions

There's a specific kind of restlessness that every practitioner knows. It's the 2 a.m. loop where a question won't let go, and for one researcher, that question was about sinusoids. Could a Kolmogorov-Arnold Network, or KAN, use a simple sinusoidal activation instead of the usual B-spline basis? The answer, as it turns out, was already out there, but the fact that this idea had to be discovered twice is exactly why we should pay attention. This isn't a story about a revolutionary breakthrough; it's a story about the value of questioning the default. If you're spending time wrestling with the limits of Unlock LLM Training: A Practical Guide to Distributed Algorithms, you already know that the most interesting progress often comes from asking what hasn't been tried, not from polishing what has.

The SineKAN paper, shared by u/jacobgorm, is a compact study in intellectual humility. The author openly admits they couldn't sleep, checked if the idea was novel, found it existed, and posted it anyway for discussion. That's a healthy model for how research should work. The core claim is practical: replacing B-splines with sinusoids in a KAN offers a simpler, potentially more efficient path to the same expressive power. We don't need to overstate the results. The honest read is that this is a useful data point, not a verdict. It suggests that the architecture's strength isn't in the specific function family, but in the KAN's core principle of learning the activation itself. That's a meaningful distinction for anyone building models today. It means that when you're designing a system, you have more freedom than the default frameworks suggest. This connects directly to how we think about Exploring Paragraph Structure: How LLMs Navigate Token Space, where the underlying structure matters more than the surface representation.

Our take is straightforward: don't mistake this for a call to abandon B-splines. Do treat it as a reminder that the search space for activation functions is far larger than the handful of names we keep in rotation. The practical implication for you is in the flexibility. If you're hitting a performance ceiling or a computational wall, the answer might not be a deeper network or a larger dataset. It might be a different basis function. The instinct to share the find, even without a dramatic benchmark victory, is the kind of behavior that moves the field forward in small, honest increments. We would tell a reader who asked about this to look at the code, run a quick experiment on their own data, and see if the simplicity translates to their problem. Don't wait for a headline. The specific thing to watch here is whether the research community treats this as a curiosity or as an invitation to explore other exotic activations. If the latter, the next sleepless night might belong to the person who tests a wavelet or a Chebyshev polynomial, and that's a question worth losing sleep over.

From Machine Learning

I couldn't sleep because I couldn't stop wondering if anyone had tried using sinusoids instead of B-splines as activation in a KAN, and fortunately/unfortunately that was already the case. I could not find it posted here, so I though I would share in the hope of some insightful discussion.

Github repo: https://github.com/ereinha/SineKAN

Read the original at Machine Learning