Your 2.7GB Python problem solved with 3300 lines of C

Introducing NoTorch, a streamlined neural network training and inference library crafted entirely in pure C, designed to alleviate the frustrations of heavy installations like PyTorch.

3 min readMachine Learning

The most honest stack for training a 10-million-parameter neural network right now might be a C compiler and a 2019 Intel MacBook. The author of NOTORCH has done something that sounds like a stunt but reads like a much-needed correction. By writing a complete training and inference library in two files of pure C, roughly 3,300 lines, they have demonstrated that the 2.7-gigabyte overhead of importing PyTorch is not a technical necessity. It is a convenience tax that has become invisible. And when that tax is removed, you get a 222-megabyte training load for two concurrent transformer models on hardware that Apple stopped selling years ago.

This matters for anyone who has ever hesitated before spinning up a training job on a laptop. The audience for this library is not the cluster engineer with a dozen A100s. It is the researcher, the student, the side-project builder who wants to run an experiment without first negotiating with their operating system's package manager over disk space. The practical takeaway is that a large fraction of what modern deep learning requires, autograd, AdamW, RoPE, GQA, SwiGLU, even bit-level quantization, can be expressed in a single translation unit and compiled in under a second. The author has ported nanoGPT from scratch, trained it on a Dracula corpus, and gotten coherent output. That is not a demo. It is a proof that the abstraction layer between the developer and the hardware has been adding weight that most users never asked for.

The obvious objection is that this approach will not scale past roughly 100 million parameters on a CPU, and the library's documentation states this plainly. That honesty is refreshing. NOTORCH includes CUDA support for larger models and a BLAS fast path for the ternary quantization, so the ceiling is real but not fixed. What the library does is restore a sense of proportion. It reminds you that the training loop is bounded by memory bandwidth and operation count, not by the number of dynamically loaded Python modules in your environment. Running two transformer trainings simultaneously on an 8-gigabyte Intel machine from 2019 should not be a surprise. It should be the baseline.

The name is a joke, but the engineering is not. By stripping away the Python runtime and the dependency chain that comes with it, the author has produced a tool that makes the act of training a small model feel direct again. You write your model in C, compile it with `-O2`, and run it. If that workflow feels strange, it is only because the industry spent a decade normalizing a far heavier one. The real innovation here is not the library itself. It is the demonstration that you can reclaim those 2.7 gigabytes, the import time, and the cognitive overhead, and still get your gradient updates and your checkpoint files. For anyone building small to medium-scale models, the question is no longer whether you can. It is whether you want to.

From Machine Learning

I'm tired of `pip install torch` eating 2.7 GB every time I want to train a 10m-param model, so I wrote NOTORCH: a complete neural network training/inference library in pure C. Two files (`notorch.h` + `notorch.c`, ~3300 LOC). No Python. Enough.

cc -O2 notorch.c your_model.c -lm -o train

Read the original at Machine Learning