For anyone who has ever watched a single radiomics job chew through minutes of compute time, the numbers from fastrad are not just impressive, they are a practical roadmap. PyRadiomics has been the trusted workhorse, but at roughly three seconds per scan, it is a bottleneck the moment you scale beyond a few dozen patients. Fastrad, built as a PyTorch-native library by Erika, achieves an end-to-end speed of 0.116 seconds on an RTX 4070 Ti. That is a 25× improvement. The per-class gains are even starker, ranging from 12.9× for GLRLM to 49.3× for first-order features. These are not theoretical benchmarks on synthetic data; they come from real validation against the IBSI Phase 1 digital phantom and against PyRadiomics on a TCIA NSCLC CT scan, with all 105 features agreeing to within 10⁻¹¹ percent. The speed is real, and the correctness is verified.
What this means for researchers and clinicians is a shift in what is feasible. When a single scan takes three seconds, a cohort of a thousand scans demands nearly an hour of extraction time. With fastrad, that same cohort drops to under two minutes. That difference changes workflow design. You no longer have to batch jobs overnight or reserve dedicated CPU nodes. You can iterate on feature selection, test different preprocessing parameters, or run quality checks in real time. The library also runs transparently on CPU or GPU, and on Apple Silicon it is 3.56× faster than PyRadiomics running on 32 threads. For labs that already use PyTorch for model training, the integration is natural, the feature extraction pipeline becomes just another tensor operation in the same ecosystem.
The implementation details matter because they reflect genuine engineering discipline. The GLCM and GLSZM kernels were the hardest to match numerically, and the fact that they were solved to within 10⁻¹³ percent deviation on the IBSI phantom tells you this is not a rough approximation. It is a faithful reimplementation. The peak VRAM of 654 MB on an RTX 4070 Ti is also worth noting, this does not require a data-center GPU to be useful. A consumer card handles it comfortably, which lowers the barrier for smaller labs and independent researchers.
Our take is straightforward: fastrad does not promise to revolutionize radiomics; it simply removes a bottleneck that has been silently limiting throughput. The preprint and the code are both open, so anyone can verify the claims and adapt the library to their own pipelines. The practical outcome is that feature extraction is no longer the slow step. That is a concrete, measurable improvement, and it is available today.