The confusion between log transforms and standardization is understandable, and it's worth clearing up because it directly affects how your models perform. You don't have to choose one path forever, nor are you locked into a single sequence. The real question is what your data looks like and what your algorithm needs to function well.
Here's the plain version: log transforms change the shape of your data. They pull in long tails and make skewed distributions more symmetric, which is useful when you have features like income, population, or any metric that spans several orders of magnitude. Standardization, on the other hand, doesn't touch the shape. It rescales the data to have a mean of zero and a standard deviation of one, which is essential for algorithms that rely on distances or gradients, like k-means, PCA, or linear models with regularization. You are correct that standardization keeps the distribution intact. That distinction matters because it means the two techniques are answering different problems, not competing versions of the same fix.
So when do you use one after the other? Sometimes you do both, but only when the data calls for it. If you have a skewed feature that also has a wide range, applying a log transform first to normalize the shape, then standardizing to bring everything onto a comparable scale, is a legitimate and common pipeline. But if your feature is already roughly symmetric, applying a log transform will distort it into something worse. In that case, just standardize. The mistake would be assuming there is a universal order that always works. Instead, look at each feature individually. Ask whether the distribution is heavily skewed. If yes, consider a transform. Then ask whether your algorithm assumes features are on similar scales. If yes, standardize after that transform, or instead of it if the shape is already fine.
What this means for you practically is that you don't need to feel paralyzed by the choice. You need a simple diagnostic habit. Plot your feature's distribution or check its skewness. If the tail is long and the data bunches up on one side, log transform it. Then check the scale across all features. If they vary wildly, standardize. If you're using a tree-based model, none of this matters much because trees don't care about scale or distribution shape. If you're using a distance-based or gradient-based method, both steps matter, but only in that order. The key is to stop thinking of these as either-or rules and start thinking of them as tools you apply based on what your data is telling you. That's the entire logic. No hidden complexity. Just look, decide, and move on.