Excel compatibility

Automated feature engineering evolves with genetic algorithms and open-source simplicity.

Feature engineering remains the quiet battleground where tabular ML models are won or lost.

4 min readMachine Learning

The most honest thing you can say about feature engineering is that it's where good tabular models go to survive or stall, and most of us have felt that stall personally. You've cleaned the data, tuned the hyperparameters, and watched your LightGBM plateau because the raw columns simply don't express the ratio, the group-by mean, or the interaction that your domain intuition says should matter. py-evoFE, a new open-source library from developer tanopereira, takes that frustration seriously by applying genetic algorithms to the search for feature transformations themselves. Instead of handing you another static toolkit, it evolves compact recipes, things like hierarchical chains of log-ratios and target encodings, and then wraps the whole process in a Scikit-Learn-compatible pipeline that plugs straight into your existing workflow. This is not about replacing your judgment; it's about scaling the part of feature engineering that your brain shouldn't have to brute-force alone.

What makes this worth your attention is not the novelty of genetic programming, that idea has been around for decades, but the specific engineering choices that make it practical. The library uses Polars and PyArrow for vectorized computation, which means the search itself doesn't buckle under the weight of its own ambition. It also tackles a problem that quietly kills most automated feature tools: redundant computation across cross-validation folds. By caching stateful projections like UMAP and K-NN lookups via byte-hashing, py-evoFE avoids redoing expensive work every time a fold changes. And the multi-fidelity screening, where cheap CV runs filter out weak candidates before full evaluation, is a smart concession to reality. Brute-force feature generators explode your memory and your overfitting risk; this approach applies evolutionary pressure with complexity penalties, favoring parsimonious recipes that actually generalize. That's a meaningful distinction, and it's one our readers who have struggled with tools like Unlock Data Insights: A Practical Guide to Polars' Performance will recognize: the speed of the underlying computation matters less than how intelligently you spend it.

For the working data scientist, the practical takeaway is that you no longer have to choose between manual feature crafting and black-box automation. The interactive replay viewer, which generates a zero-dependency HTML dashboard of the evolutionary search, is a small but telling feature, it turns a process that feels like a black box into something you can inspect, question, and learn from. That aligns with the ethos we've explored in [Sharing my ML learning repo, NumPy to Transformers, 5 months, daily commits, all notebooks public. [D]](/post/sharing-my-ml-learning-repo-numpy-to-transformers-5-months-d-cmu91wx2d04i95ngm1211p2j9), where the value of open, reproducible learning paths beats polished abstraction every time. But we'd caution against treating this as a silver bullet. Genetic search still requires you to define the search space, choose the evaluator, and set population sizes, it amplifies your judgment, it doesn't replace it. And while the library is MIT-licensed and early in its lifecycle, the real test will come from messy, real-world datasets where memory constraints and feature leakage are not so neatly handled.

The specific thing we'll be watching is how the Caruana ensembling over island winners performs outside of benchmark competitions. That's where the promise of automated feature discovery either earns its keep or becomes another overfit to Kaggle leaderboards. If you're currently stuck in manual feature engineering because you've been burned by tools that promise automation without control, py-evoFE is worth a weekend experiment, not because it's revolutionary, but because it treats feature engineering as a search problem with a budget, which is exactly the right frame. Try it on a dataset where you already know the winning features, and see if it rediscovers them. If it does, you've found a collaborator. If it doesn't, you've learned something about your own data that no brute-force loop would have taught you.

From Machine Learning

I’m excited to announce the release of py-evoFE (v0.3.0) — an open-source Python library that uses genetic algorithms to automatically discover, combine, and optimize feature transformations for tabular datasets.

Feature engineering is still where most tabular ML competitions and production models are won or lost. While GBDTs like LightGBM and XGBoost excel on raw tabular data, they struggle to discover complex ratios, nested group-by aggregations, nonlinear dimensional projections, and interaction graphs on their own.

Read the original at Machine Learning