![[P] Extreme Imbalance Data from 100K dataset only have 56 failure [P]](https://preview.redd.it/plbydmenmm6h1.png?width=140&height=23&auto=webp&s=5ea145541663231254f92fd815bb1237861fd6cb)
Machine Learning
[P] Extreme Imbalance Data from 100K dataset only have 56 failure [P]
Addressing extreme class imbalance—where only 56 failures exist within a 100,000 dataset—presents a significant challenge for predicting machine failure and Remaining Useful Life (RUL). With timestamped data and a binary failure label, effective modeling requires careful algorithm selection. Given you've already identified operating hours and humidity as non-correlated features, consider exploring deep learning approaches such as anomaly detection models or techniques specifically designed for imbalanced datasets like Synthetic Minority Oversampling Technique (SMOT) integrated with recurrent neural networks.
![I Built Paper Deck: A Better Way to Discover AI/ML Papers [P]](https://preview.redd.it/cg32bshjqd6h1.png?width=140&height=75&auto=webp&s=a45be317897998d05be4fedc9b7b57d2e90c4791)




















![Analysis of the results of the "Transforming autoencoders" architecture mentioned by Hilton, for my dissertation. [r]](https://external-preview.redd.it/WDru8xQhJvLHSHkthgxvp2YY8DX7dtS1H1UsgUssaL4.png?width=640&crop=smart&auto=webp&s=1ea2654afb58b4aca7a5717b8200394ce3d64344)





